Real-World AI Voice Interface Technology

From real-world voice input to recognition and dialogue

mpWAV develops the full voice interface stack — speech enhancement that reduces ambient noise and device echo, wake-word detection, sound source localization, speaker diarization, on-device speech recognition, and a conversational language model.

From single-microphone smart devices to multi-channel robots, kiosks, vehicles, and meeting systems, we provide the right combination of technologies for your product's microphone configuration and compute platform.

Voice Enhancement · Voice Interaction · Speech Recognition · On-Device Dialogue · HW Integration

Image [1]
Full voice interface stack hero

Real-World Audio Input

↓ Voice Enhancement — mpNC · mpAEC · mpBeamforming · mpAB

↓ Voice Interaction — mpWWD · mpS · mpLocalization · mpDiarization

↓ Recognition & Dialogue — mpASR · mpLLM

↓ Product Action — Robot · Kiosk · Vehicle · Smart Device · Meeting

Foundation: Multi-Channel HW · DSP · FPGA · AP · Edge Platform

One Integrated Voice Stack

What does mpWAV technology do?

mpWAV technology makes real-world voice input clearer, detects the user's call and identifies who is speaking from where, and converts speech into text and intent — connecting it to product functions or conversational services.

From small devices with a single microphone to products combining multiple microphones and speakers, you can pick the technologies you need or integrate them as one voice interface stack.

Capture → Enhancement → Interaction Analysis → Recognition → Dialogue → Product Action

Image [2]
Four-layer architecture: enhancement, interaction, recognition, implementation

01 Capture & Enhancement — mpNC · mpAEC · mpBeamforming · mpAB

02 Interaction Intelligence — mpWWD · mpS · mpLocalization · mpDiarization

03 Recognition & Dialogue — mpASR · mpLLM

04 Product Implementation — Multi-Channel HW · DSP · FPGA · AP

Real-World Voice Challenges

Different products face different voice problems — and need different technology

A product's microphone never receives only the user's voice.

Nearby conversation, music, road and motor noise, echo from the product's own speaker, and overlapping speech can all arrive together. Some products cannot fit more than one microphone, and once speech is recognized, the product still has to understand intent and context.

mpWAV does not apply one technology to every product. We combine technologies based on each product's problem and hardware constraints.

Single-microphone products

Earbuds and small devices that cannot add microphones need single-mic noise control.

Multi-microphone products

Robots, kiosks, and vehicles need multi-microphone noise reduction and spatial analysis.

Products with speakers

Response prompts and announcements that loop back into the microphone must be cancelled.

Multi-party conversation

Meetings and voice chat require separating overlapping voices and telling speakers apart.

Conversational services

Kiosks and robots must understand intent and context beyond recognizing the words.

The required technology stack depends on microphone count, number of speakers, product output audio, and the final service.

Image [3]
Five product voice problems (earbuds · robot · kiosk · meeting · vehicle)

Single Mic | Noise + Echo | Multiple Speakers | Far-Field Speech | Dialogue

Voice Enhancement

We improve the input signal before it reaches recognition

Speech recognition performance depends not only on the ASR model but on the quality of the audio arriving at the microphone.

mpWAV provides voice enhancement matched to your product's input conditions — single or multiple microphones, device echo, and ambient noise.

mpAEC

Cancels echo produced by the device's own speaker

mpAEC is a multi-channel acoustic echo cancellation technology that reduces sound played by the product — TV audio, car audio, robot responses, kiosk prompts — from re-entering the microphone.

It is designed for environments where the user's voice and the product's output are present at the same time, and its output can feed beamforming or your existing ASR.

Official material presents stable processing with near-end speech present, no double-talk detection required, no far-end signal decorrelation required, and fast convergence as mpAEC's characteristics.

Key problems

  • Speaker output re-entering the microphone
  • User and product speaking at the same time
  • Changing echo paths
  • Real-time processing

Typical applications

  • Vehicles
  • Robots
  • Kiosks
  • Smart home
  • Meeting systems
Image [4-1]
mpAEC echo cancellation flow

Speaker Output + User Voice → Multi-Channel Mic → mpAEC → Echo-Reduced Speech

mpBeamforming

Reduces ambient noise and strengthens the target voice with multiple microphones

mpBeamforming uses the spatial information across multiple microphones to reduce the impact of ambient noise and deliver the target voice more clearly.

mpWAV optimizes from the actual input signals rather than fixed microphone position data, reducing the repeated tuning burden when microphone configurations change between products.

Official material describes automatic optimization from input signals, no separate tuning when microphone configurations change, and minimal target-voice distortion as its core characteristics.

Key problems

  • Far-field user speech
  • Ambient noise from multiple directions
  • Microphone array changes between products
  • Target-voice distortion during noise removal

Typical applications

  • Robots
  • Kiosks
  • Vehicles
  • Smart home
  • Meeting systems
Image [4-2]
mpBeamforming multi-microphone beam pattern

Noise from all directions (gray) — target voice beam (blue) reinforced

mpAB

Integrates echo cancellation and beamforming into one real-time pipeline

mpAB combines mpAEC and mpBeamforming to process input where product speaker echo and environmental noise are present together.

Multi-channel microphones, the speaker reference, FPGA·DSP·AP platforms, and your existing ASR can be connected into a single product architecture.

Official material presents mpAB as the integration of mpAEC and mpBeamforming, with real-time FPGA implementation and MCU+DSP / AP porting structures.

Key problems

  • Echo and ambient noise occurring together
  • Multi-channel real-time processing
  • Integration with existing ASR
  • Product hardware integration

Typical applications

  • Robots
  • Kiosks
  • Vehicles
  • Smart devices
  • Meeting devices
Image [4-3]
mpAB integrated processing block

Mic Array + Speaker Reference → [ mpAEC + mpBeamforming ] → Enhanced Speech

mpNC

Single-microphone noise control for products that cannot add mics

mpNC is a voice enhancement technology that reduces the impact of ambient noise from a single microphone input.

It provides the technical foundation for improving voice input in earbuds, mobile accessories, small smart devices, and constrained hardware where a microphone array is not an option.

Key problems

  • Product structures that cannot add microphones
  • Space constraints of small devices
  • Everyday ambient noise
  • Limited compute resources

Typical applications

  • Earbuds
  • Wearables
  • Phone accessories
  • Small appliances
  • ClearSense-linked devices
Image [4-4]
mpNC single-microphone processing (earbuds · small devices)

Single Microphone → Speech + Ambient Noise → mpNC → Noise-Reduced Speech

Voice Interaction Intelligence

Beyond cleaning the audio — detecting calls, direction, and speakers

For a product to respond naturally, it needs more than the fact that someone spoke: was the wake word detected, which direction did the voice come from, and who spoke when among several people?

mpWAV's interaction technologies sit between voice enhancement and speech recognition, providing the information a product needs to interact with its users.

mpWWD

Detects the wake word and starts the voice interface

mpWWD detects a predefined wake word so the product can start listening for voice commands.

It connects to robots, smart home devices, vehicles, and edge devices that must decide the moment of activation on-device rather than streaming audio to a server at all times.

Key roles

  • Wake word detection
  • Voice interface activation
  • Always-on standby
  • Starting the on-device command flow

Typical applications

  • Robots
  • Vehicles
  • Smart home
  • Kiosks
  • Wearables
Image [5-1]
mpWWD wake word detection flow

Ambient Audio → mpWWD → Wake Word Detected → mpASR / Listening Mode

mpS

Prepares multi-party voice input for downstream processing

mpS processes audio from environments where multiple voices exist or overlap — meetings and voice chat — into a form suitable for downstream analysis and recognition.

Combined with mpDiarization and mpASR, it forms the input pipeline for multi-party meeting records and voice chat.

Key problems

  • Overlapping voices from multiple speakers
  • Multi-party conversation
  • Complex voice chat input
  • Signal processing before speaker separation

Typical applications

  • Meetings
  • Remote collaboration
  • Voice chat
  • Consultation records
  • Multi-party dialogue
Image [5-2]
mpS multi-party voice processing

Overlapping speaker waveforms → mpS → organized per-voice streams

mpLocalization

Estimates where the voice is coming from

mpLocalization uses multi-microphone input to estimate the direction or position of the user's voice.

It powers interfaces where a robot turns toward the user, a vehicle distinguishes speech by seat, or a multi-directional smart device responds toward the speaker.

Key roles

  • Sound source direction estimation
  • User position information
  • Seat- and direction-based interaction
  • Robot gaze and rotation control

Typical applications

  • Robots
  • Vehicles
  • Smart home
  • Meeting devices
  • Multi-directional kiosks
Image [5-3]
mpLocalization direction estimation

Central mic array with several users — active speaker direction marked in blue

mpDiarization

Tells who spoke when

mpDiarization separates the speech segments of each speaker in multi-party audio.

It organizes per-speaker utterances in meeting minutes, consultation records, and voice chat, and attaches speaker identity to the text produced by mpASR.

Key roles

  • Per-speaker segment separation
  • Speaker-change detection
  • Speaker tagging for minutes
  • Structuring consultation and dialogue records

Typical applications

  • Meeting minutes
  • Remote meetings
  • Consultation records
  • Voice chat
  • Interview analysis
Image [5-4]
mpDiarization speaker timeline

Speaker A ━━━ · Speaker B ━━━ · Speaker C ━━━ (Time →)

Speech Recognition and Dialogue

Turning speech into text — and understanding intent and context

Input that has passed through enhancement and interaction analysis is converted into text or product commands by mpASR.

Where a conversational interface is needed, mpLLM understands the recognition output and dialogue context, connecting questions, responses, and product functions.

mpASR

Noise-robust, on-device end-to-end ASR

mpASR is mpWAV's speech recognition technology, designed to run on servers, PCs, smart devices, and IoT/edge environments.

It supports domain fine-tuning for your product's terminology, menu names, and command set, and combines with mpWAV's speech enhancement front end.

Key roles

  • Speech-to-text conversion
  • Product command recognition
  • Domain fine-tuning
  • Standalone on-device execution
  • Streaming recognition

Typical applications

  • Kiosks
  • Robots
  • Vehicles
  • Smart devices
  • Meeting records
Image [6-1]
mpASR platform structure

Enhanced Speech → mpASR (Server · PC · Smart Device · IoT/Edge) → Text / Command

mpLLM

An on-device language model that understands dialogue inside the product

mpLLM is a conversational language model that connects user intent, dialogue context, and product functions on top of the text produced by mpASR.

Built on a 1B-class on-device model direction, it implements product-specific dialogue flows — such as kiosk voice ordering — while reducing network dependence.

Key roles

  • Understanding user intent
  • Maintaining dialogue context
  • Asking for missing information
  • Interpreting menu, options, and quantity
  • Connecting to product APIs

Typical applications

  • Conversational kiosks
  • Robot dialogue
  • Smart devices
  • Vehicle interfaces
  • Meeting summaries
Image [6-2]
mpASR·mpLLM kiosk voice ordering structure

User Speech → mpAB → mpASR → mpLLM → Intent · Menu · Option · Quantity → Ordering API

Hardware and Embedded Implementation

Turning algorithms into systems that run inside real products

Applying mpWAV technology to a real product requires connecting microphone input, speaker output, the AEC reference, the processing platform, and the speech recognition system together.

mpWAV supports multi-channel microphone arrays, audio I/O, real-time FPGA implementation, MCU+DSP structures, and AP·DSP porting.

Microphone arrays

We design linear and multi-directional arrays matched to product shape and user direction.

Multi-channel audio I/O

Delivers multiple microphones and the speaker reference signal to the processing platform.

FPGA

Implements multi-channel preprocessing in real time and validates the pre-SoC architecture.

DSP·AP

Ports the algorithms to your product's compute environment.

Lightweight & single-mic builds

For devices that cannot fit an array, we evaluate mpNC and lightweight model structures.

Image [7]
Single/multi-channel HW and platform integration

Single Mic / Multi-Channel Array + Speaker Reference

→ Audio Hardware → FPGA · DSP · AP → mpWAV Stack → ASR · LLM · Product API

Application Technology Map

Each application needs a different stack

Robots

  1. 1

    mpWWD

    Wake word detection

  2. 2

    mpLocalization

    User direction estimation

  3. 3

    mpAEC

    Robot response echo cancellation

  4. 4

    mpBeamforming / mpAB

    Motor, fan, and nearby-conversation reduction

  5. 5

    mpASR

    Command recognition

  6. 6

    mpLLM

    Dialogue and intent understanding

Kiosks

  1. 1

    Mic Array

    Frontal far-field voice capture

  2. 2

    mpAB

    Store noise and prompt echo processing

  3. 3

    mpASR

    Menu and option recognition

  4. 4

    mpLLM

    Order context, follow-up questions, and product API

Smart devices & earbuds

  1. 1

    mpNC

    Single-microphone noise reduction

  2. 2

    mpWWD

    Wake word detection

  3. 3

    mpASR

    On-device command recognition

  4. 4

    mpLLM

    Natural-language function execution

Meetings & voice chat

  1. 1

    mpS

    Multi-party voice processing

  2. 2

    mpDiarization

    Per-speaker separation

  3. 3

    mpAEC / Beamforming

    Echo and far-field noise processing

  4. 4

    mpASR

    Meeting transcription

  5. 5

    mpLLM

    Summaries and Q&A

Vehicles & mobility

  1. 1

    mpWWD

    Wake word detection

  2. 2

    mpLocalization

    Seat and speech-direction analysis

  3. 3

    mpAEC

    Car audio echo cancellation

  4. 4

    mpBeamforming / mpAB

    Road noise and passenger conversation

  5. 5

    mpASR

    Vehicle command recognition

  6. 6

    mpLLM

    Natural-language vehicle interface

Image [8]
Five industry pipeline cards

Robot Stack | Kiosk Stack | Smart Device Stack | Meeting Stack | Mobility Stack

Find the Right Technology

Pick your product's problem

Product problemRecommended technology
Echo from the product's own speakermpAEC
Reducing ambient noise with multiple microphonesmpBeamforming
Echo and ambient noise at the same timempAB
Only one microphone availablempNC
Detecting a wake wordmpWWD
Processing multi-party voice inputmpS
Knowing the speaker's direction or positionmpLocalization
Telling who spoke whenmpDiarization
Converting speech to text or commandsmpASR
Understanding dialogue context and intentmpLLM
Multi-channel input and real-time processingMulti-Channel HW

Real products usually face several of these at once, so two or more technologies are often applied together.

Image [9]
Technology selection flowchart

Only one mic? → mpNC

Speaker echo? → mpAEC

Ambient noise / far-field? → mpBeamforming

Echo + noise together? → mpAB

Wake word · direction · speakers? → mpWWD / mpLocalization / mpDiarization

Recognition & dialogue? → mpASR / mpLLM

Performance Validation

We validate with real product results, not technology names

Voice technology performance varies with microphones, speakers, user distance, noise type, the connected ASR, and the execution platform.

mpWAV evaluates before-and-after results against real product data and field conditions.

Voice enhancement

  • Echo reduction
  • Ambient noise reduction
  • Target voice preservation
  • Before/after audio quality

Interaction analysis

  • Wake word detection
  • Direction & position estimation
  • Speaker separation
  • Utterance segmentation

Speech recognition

  • WER
  • CER
  • Command success rate
  • Domain terminology

Dialogue processing

  • Intent understanding
  • Slot & option extraction
  • Context retention
  • Product API execution

Product performance

  • Processing latency
  • Memory footprint
  • DSP·FPGA·AP resources
  • On-device feasibility

Independently measured in accredited testing

SI-SDR after processing
15.59 dB

0 dB input SNR · average of 100 utterances (target ≥ 10 dB)

SNR after processing
15.57 dB

0 dB input SNR · average of 100 utterances (target ≥ 15 dB)

Noise reduction
53.8 dB

Real-world 5 dB SNR · 65 dB speech, 60 dB noise · average of 100 utterances (target ≥ 25 dB)

Real-Time Factor
0.336

Below 1 means real-time capable (target < 1)

Telecommunications Technology Association (TTA) test report TTA-25-1103. The article under test was our own speech enhancement solution, ClearSense Audio v1.1.0 (algorithm v1.0.0), tested 9–12 June 2025 and issued 17 November 2025. Corpora: LibriSpeech ASR and AIHub Korean Speech, with reverberation simulated using gpuRIR.

Image [10]
Full validation dashboard (criteria only where numbers are not yet published)

Echo Reduction | Noise Reduction | Wake Word | Localization | Diarization | WER·CER | Intent | Latency

From Algorithm to Product

Delivered in the form your development stage needs

Software·SDK

Apply mpWAV technology in front of your existing product and ASR.

PoC · performance validation

Verify before-and-after results with your real voice data.

Mic array · HW module

Connect multi-channel voice input and processing to your product.

DSP·FPGA·AP porting

Port the technology to your product's compute platform.

Licensing · joint development

Integrate the technology around your product data and service goals.

On-device model optimization

Fit mpASR and mpLLM to the target platform's memory and compute budget.

SoC · semiconductor IP

Scale into dedicated silicon for volume production and miniaturization.

Image [11]
Algorithm → porting → product → SoC roadmap

Algorithm → Software PoC → HW Evaluation → DSP/FPGA/AP Porting → Product Integration → License/Co-Dev → SoC

Research and Intellectual Property

Connecting research results to real voice interfaces

mpWAV extends speech and audio signal processing and speech recognition research into patents, software, embedded implementations, and product PoCs.

Official material presents international patents and major journal publications for mpAEC and mpBeamforming, the company's domestic and international IP, and the core team's long-term research record.

Research → Patent → Algorithm → Embedded Implementation → Product

Image [12]
Research → patent → productization (papers · patents · algorithms · boards · products)

Papers → Patents → Algorithm waveforms → FPGA·DSP boards → Robot·Kiosk products

FAQ

Frequently asked questions about mpWAV technology

No.

The portfolio spans single- and multi-microphone voice enhancement, wake-word detection, localization, multi-party voice processing, speaker diarization, speech recognition, and an on-device conversational language model.

mpNC reduces the impact of ambient noise from a single microphone input.

It is the foundation for earbuds and small smart devices where a microphone array is not an option.

mpBeamforming for ambient noise reduction, mpAEC for cancelling the product's own speaker echo, and mpAB when both problems must be handled together.

mpLocalization estimates the direction or position of the user's voice using multiple microphones.

Actual accuracy and supported arrays should be validated on your product's structure.

mpDiarization separates per-speaker speech segments in multi-party audio.

Combined with mpS and mpASR, it forms a pipeline that processes meeting audio and organizes it into per-speaker text.

mpASR converts speech to text and mpLLM interprets menu, options, quantity, and dialogue context.

A real deployment also needs the ordering API, menu data, safety policy, and validation on the target hardware.

It is being developed in the direction of a 1B-class on-device dialogue model.

Actual on-device feasibility depends on model version, precision, memory, and accelerator conditions.

Currently every technology is covered as a dedicated section of this overview page.

mpAEC, mpBeamforming, mpAB, mpASR, and multi-channel HW will expand into standalone pages as material is prepared.

Yes.

mpWAV's speech enhancement applies in front of your existing ASR, improving the input signal without replacing the engine.

No.

Results vary with microphone count and placement, speakers, the space, noise, user distance, execution platform, and the connected models. Validate with your real product data.

Build Your Voice Technology Stack

Design the voice technology your product needs as one stack

Tell us your product type, microphone count, speaker structure, dominant noise, and the voice features you want to build — we will review the right mpWAV combination.

From single-mic noise control to multi-channel echo and noise preprocessing, wake-word, localization and speaker analysis, on-device recognition and a conversational LLM — connected to fit your product environment.

Image [13]
Final CTA full stack + product icons

Audio Input → Enhancement → Interaction → ASR → LLM → Product Action

Earbuds · Robot · Kiosk · Vehicle · Meeting · Smart Device

mpWAV develops and licenses voice interface technology that makes speech clear in noisy, real-world environments, built on 25+ years of speech signal processing research.

(주)엠피웨이브·사업자등록번호 859-81-01905·대표 박형민

Teilhard Hall 405, 35 Baekbeom-ro, Mapo-gu, Seoul, Korea

02-705-8916·info@mpwav.com

© 2026 (주)엠피웨이브. All Rights Reserved.