Audio Front-End for real-world voice AI

Keep the speech engine you already use.Raise recognition where it actually runs.

In front of your ASR, mpWAV cuts echo and noise while preserving the speech recognition needs. Drop it into your existing product — no cloud, no dedicated NPU.

We compare recognition before and after, on your recordings and your environment.

  • Deployed with ROBOCARE
  • Joint research with ETRI
  • Accredited TTA testing
  • NET certification

Before / after

Hear what changes in the same room.

Echo and noise come down; the speech your recognizer needs stays in.

mpWAV Solution Demo

Headphones or a speaker make the difference easier to hear.

  1. BEFORE

    Raw signal from the microphones

  2. mpWAV

    Echo and noise removal (mpAB)

  3. AFTER

    Signal handed to the speech engine

mpWAV does not replace your speech engine. It changes what the engine receives.

Send us a note and a research engineer replies within 2-3 business days.

Diagram: six irregular waveforms from six noise sources pass through the mpWAV gate and come out as three clean waveformsOn the left are six irregular waveforms — nearby conversation, TV and speaker echo, machine and motor noise, vehicle and road noise, far-field speech, and room reverberation. They pass through the mpWAV core (AEC and ANC, on-device) in the middle and emerge on the right as three smooth waveforms feeding an ASR engine, wake word detection, and command recognition.Nearby conversationTV & speaker echoMachine & motor noiseVehicle & road noiseFar-field speechRoom reverberationASR engineWake wordCommandsmpWAVCOREAEC · ANCON-DEVICE

Why audio preprocessing?

It may not be your speech engine.

Speech recognition that works in a quiet office loses accuracy once the product ships into a real environment.

The reason is that the input speech is already damaged by noise and echo before it ever reaches the ASR.

  1. 01

    Nearby conversation

    People talking nearby and the background noise of a store or office arrive together with the user's voice.

  2. 02

    TV & speaker echo

    Sound from a TV or the product's own speaker loops back into the microphone.

  3. 03

    Machine & motor noise

    Motors, fans, and moving parts interfere with recognition.

  4. 04

    Vehicle & road noise

    Driving speed and road surface keep changing the character of the noise.

  5. 05

    Far-field speech

    The farther the user moves away, the quieter the voice and the louder the noise.

  6. 06

    Room reverberation

    Reflected speech mixes with the original and distorts the input signal.

Far-field voice environments in the smart home and the car — how noise and echo reach the microphone

What you need is not just a better ASR. You need an Audio Front-End in front of it.

Why mpWAV

Don't replace your ASR.Change what your ASR hears.

In front of your speech engine, mpWAV turns sound from a real environment into a signal that is easier to recognize.

Instead of swapping in a new speech engine, it gives the ASR you already run a more stable input.

  1. Microphone
  2. mpWAV Audio Front-End
  3. Existing ASR
  1. 01

    Protecting the speech while reducing the noise.

    Ordinary noise removal can sound clean to a person while damaging the very features speech recognition depends on.

    mpWAV focuses on reducing interference while preserving as much of the target speech as recognition needs.

    Speech Preservation

  2. 02

    You keep the ASR you already use.

    There is no need to replace your whole system with a different speech engine.

    mpWAV sits in front of your existing ASR and improves the signal that engine receives.

    Keep Your Existing ASR

  3. 03

    No cloud and no separate NPU required.

    It does not assume a cloud connection or a dedicated AI accelerator, and it is applied with the compute and memory budget of an embedded product in mind.

    You can evaluate whether it fits without significantly changing your current product architecture.

    On-device

  4. 04

    It is fitted to your actual product environment.

    Microphone count, placement, user distance, room structure, and noise characteristics differ for every product.

    mpWAV works out the Audio Front-End configuration your real usage environment calls for.

    Environment-specific Integration

Verified performance

Skip the description.Look at the measurements first.

The core mpWAV preprocessing has been verified in accredited testing.

The figures below come from the TTA test results for mpAB.

  • 42.088 dB

    Stereo echo and noise suppression

  • 64 MHz

    Clock frequency

  • 0.571 MB

    SRAM required

  • 0.852

    Real-Time Factor

Measured on mpAB in accredited TTA testing

Measured in a TTA test report (2023, Ministry of SMEs and Startups Didimdol programme) with mpAB as the article under test.

See the test conditions and full performance

Proven in real environments

Not in a lab —in the product's own environment.

We validate whether the preprocessing holds up in real product conditions, from care robots to far-field speech.

The ROBOCARE Cami-friends care robot — a rounded head with an eye display and a ring light above a wheeled body

ROBOT · HOME ENVIRONMENT

A care robot inreal household noise

In homes where a TV and everyday household noise run together, mpWAV preprocessing was applied so the robot receives the user's wake word and commands reliably.

Environment
A home where TV audio and household noise overlap
Challenge
Not missing the call of an older adult who speaks little and quietly
Applied
Multi-channel audio I/O module supply · speech preprocessing

ROBOCARE Cami-friends

ETRI Conference 2025 — a white humanoid robot in front of visitors, with ETRI AIR ROBOT and a microphone array on its chest

HUMANOID · LIVE DEMO

mpWAV running on a robot,demonstrated on site

At ETRI Conference 2025 we demonstrated voice input and system behaviour through a real robot interface.

Environment
Household noise and reverberation, plus the robot's own actuator noise
Challenge
Recognising the target speaker while the robot itself is moving
Applied
Microphone-array interface hardware design · preprocessing for extreme noise

ETRI

We compare recognition before and after, on your recordings and your environment.

Applications

From speech recognition in the real worldto acoustic problems beyond it.

The mpWAV Audio Front-End is built for products that have to understand a human voice — robots, vehicles, kiosks, and smart devices.

The same acoustic signal processing extends to anomaly detection, defence and security, meetings, and hearing assistance.

Core applications

  • ROBOT

    Catching the user's voiceover the robot's own noise

    Service, home, care, and guide robots have to handle not only ambient noise but the motor and actuator noise they generate themselves.

    mpWAV keeps the wake word and the spoken command reaching the engine even in those conditions.

    Motor Noise · Far-field Speech

    Read about robot self-noise

  • AUTOMOTIVE & MOBILITY

    Voice commands that hold upwhile the vehicle is moving

    Road noise, in-car audio, and passengers talking all arrive at once, and the target voice still has to get through.

    Road Noise · Echo · Multiple Speakers

  • KIOSK

    Clear voice inputin a noisy public space

    In stores, hospitals, and public offices, a barrier-free voice interface has to take questions and commands over the background noise of the room.

    Background Noise · Distance

  • SMART DEVICE & APPLIANCE

    Products that understand youthrough household noise

    In the homes where TVs, appliances, and smart-home devices actually run, conversation, household noise, and speaker echo happen at the same time.

    mpWAV keeps the voice interface working in those conditions.

    TV Audio · Household Noise · Echo

Extended applications

  • INDUSTRIAL

    Factory anomaly detection

    Detecting abnormal motor and equipment sounds acoustically on a noisy production line.

    Acoustic Monitoring

  • DEFENSE & SECURITY

    Defence and security

    Long-range signal detection and multimodal recognition, applied in national R&D projects for environments with extreme noise.

    Long-range Detection · Multimodal Recognition

  • MEETING & VOICE CHAT

    Meetings and voice chat

    Speech separation and speaker diarization for meeting minutes, online meetings, and voice chat.

    Speech Separation · Speaker Diarization

  • HEARING ASSISTANCE

    Hearing assistance

    ClearSense Audio raises speech clarity in noise on ordinary earbuds and phones.

    Speech Enhancement

    Visit ClearSense Audio

The product changes.The acoustic problem repeats.

Noise · Echo · Distance · Reverberation · Source separation

Tell us the product and where it is used, and we will look at whether it fits.

Audio Front-End technology

We combine the preprocessingthe environment actually calls for.

Echo, ambient noise, far-field speech, and a product's own noise are hard to solve with a single algorithm.

mpWAV combines AEC · Beamforming · Noise Cancellation · Adaptive Processing to match the microphone layout and the environment of the real product.

Microphone input

mpWAV Audio Front-End

  • AEC
  • Beamforming
  • Noise Cancellation
  • Adaptive Processing

Existing ASR

  • mpAEC

    Acoustic Echo Cancellation

    Reducing the echo thatreturns from your speaker

    Sound played by the product's own speaker loops back into the microphone, mixes with the user's voice, and can get in the way of recognition.

    mpAEC reduces that acoustic echo so the engine receives a steadier input signal.

    Speaker Echo

    More about mpAEC

  • mpBeamforming

    Beamforming

    Focusing several microphoneson the voice that matters

    Using multiple microphone inputs, it captures speech from the target direction and raises input quality when the speaker is far away.

    Far-field Speech · Directionality

    More about mpBeamforming

  • mpNC

    Noise Cancellation

    Reducing noise from the roomand from the product itself

    Household noise, road noise, and the product's own motor and fan noise all interfere with recognition, and mpNC reduces them.

    The goal is not simply sound that is pleasant to a listener, but an input signal a recognizer can use.

    Environmental Noise · Self Noise

    More about mpNC

  • mpAB

    Adaptive Audio Front-End

    Preserving the signal you needas the environment changes

    Working from the actual input signal, mpAB reduces echo and noise while preserving as much of the target speech as recognition needs.

    Adaptive Processing · Speech Preservation

    More about mpAB

We do not sell one algorithm.We build the Audio Front-End the product needs.

Microphone count, placement, user distance, and noise characteristics decide which combination of techniques applies and how it is integrated.

Integration

From validating the algorithmto shipping it in a product.

mpWAV is not a technology that stops at the PoC.

Depending on your processor, microphone layout, and stage of development, we can look at anything from a software library to hardware porting, an audio module, or production.

mpWAV circular MEMS microphone array boards, a multi-channel audio board, and linear microphone arrays
Multi-channel microphone array and audio I/O boards
  • SOFTWARE LIBRARY

    Applied in softwareon the processor you already have

    We assess applying the mpWAV preprocessing as a software library within your product's existing processor environment.

    That means checking how it integrates with your current audio and recognition pipeline without a major change to the product architecture.

    Minimum Hardware Change

  • PORTING

    Ported to the computeyour product already runs

    For DSP, FPGA, or AP environments already in the product, we assess whether the preprocessing fits and what resources it needs.

    Hardware-specific Integration

  • AUDIO MODULE

    A module for fast validationand product development

    For an early PoC or during development, a module that already includes multi-channel audio I/O and preprocessing can be the faster path.

    How ROBOCARE is supplied

    Fast Prototyping

  • PRODUCTION

    Taking a validated technologyinto production

    Once performance is confirmed through a PoC and product integration, we agree the licensing that fits your stage of development.

    Where volume production and product expansion call for it, dedicated chips and semiconductor IP cooperation are considered too.

    From PoC to Production

Before deciding how to integrate,find out whether it works in your environment.

We look at your product environment first, then propose how to validate and integrate.

From test to production

You do not have to committo a large project up front.

Start with the one environment that matters most, and compare mpWAV before and after on real data.

We look at the product and the problem first, then work through only the scope the validation needs.

  1. 01

    Environment analysis

    We go over the product, the number and placement of microphones, user distance, the main noise environment, and what is going wrong today.

    What we need to start

    Product type · microphone count · user distance · main noise environment

  2. 02

    Data review

    If you have speech recorded in the real environment, we analyse the current problem and the signal characteristics first.

    If there is no recording yet, we can define the test environment and recording conditions together.

  3. 03VALIDATION

    PoC and before/after comparison

    We compare the same conditions with and without mpWAV to see the effect of preprocessing and whether it fits the product.

    Recognition results, latency, and compute load — we define the KPIs the project actually needs together.

  4. 04

    Product integration

    Once the PoC confirms the effect, we decide how to integrate it with your processor and your audio and recognition pipeline.

  5. 05

    Field validation

    The integrated product does not stay in lab conditions — we check it again where it is actually used.

    Distance, room, the product's own noise, and real usage patterns all get reviewed on site.

  6. 06

    Production and expansion

    A validated configuration goes into the product, and can extend to follow-up models or a product line where that is wanted.

    Licensing, technical support, and how it reaches production are agreed per project.

You can start without any recordings.

Product type, microphone count, user distance, and the main noise environment are enough for us to work out how a validation could run.

Tell us about the product environment and we will start from the validation step that fits.

FAQ

Questions we get before a project starts

It is the stage that removes noise and echo before the signal reaches your speech recognition engine.

It raises the quality of the input signal without changing the recognition engine itself, so it applies to a speech system you have already built. mpWAV delivers echo cancellation, beamforming, and noise suppression as a single on-device module.

No. You keep the engine you have.

mpWAV is a preprocessing layer that sits between the microphone input and the recognition engine. It applies the same way to a commercial API or an in-house engine, with no retraining.

It runs in real time at 64 MHz clock, 0.571 MB of SRAM, and a 543 KB binary.

The real-time factor is 0.852, and it needs no cloud connection and no dedicated NPU. It drops into embedded environments such as robots, kiosks, and vehicles.

It was verified in an accredited third-party test.

A TTA test report from the SMBA Didimdol program recorded 42.088 dB of noise reduction and a real-time factor of 0.852 (65 dB speech, 60 dB noise). mpWAV also holds two CES 2024 Innovation Awards and NET new-technology certification from the Ministry of Trade, Industry and Energy.

You can choose a software library, a multi-channel audio I/O preprocessing module, a DSP or FPGA port, a technology license, or an SoC partnership.

Which one fits depends on your development stage and compute platform.

It starts with a PoC on your own recorded speech data.

The sequence is environment analysis, data review, PoC and performance validation, product integration, on-site tuning, and then scaling to production. We can begin with the environment analysis even if you have no field recordings yet.

Verify your product's voice problem with real data

Tell us your product type, audio I/O configuration, usage environment, and current errors — we will recommend the right validation method and integration structure.

You don't have to commit to a large project up front. Start a PoC with your most important environment and data.