AI voice interfaces built for your product and site
mpWAV provides the full voice interface stack — single-microphone noise control, multi-microphone echo and noise preprocessing, wake-word detection, sound source localization, speaker diarization, on-device speech recognition, and a conversational LLM.
For product environments as different as robots, kiosks, vehicles, smart devices, and meeting systems, we select the required technologies and connect them into one working architecture.
Different industries need different voice technology
mpWAV analyzes your product's microphone count, speaker output, user position, number of speakers, dominant noise, and service goal to compose the right combination of voice technologies.
Earbuds with one microphone start from mpNC; multi-channel robots from mpAEC, mpBeamforming, and mpLocalization; meetings from mpS and mpDiarization; conversational kiosks from mpASR and mpLLM.
We don't force one generic algorithm onto every product — we design the voice stack around the product's conditions.
Image [A2]
Technology selection by product condition
Single Mic → mpNC | Multi-Channel+Noise → mpBeamforming | Speaker+Mic → mpAEC/mpAB
Wake Word·Direction → mpWWD/mpLocalization | Multiple Speakers → mpS/mpDiarization
Recognition → mpASR | Dialogue·Intent → mpLLM
Application Overview
Pick the product and environment you're solving for
Robots
Multi-channel voice interfaces for products that move and speak
Helps robots detect the user's call, find the direction the voice came from, and recognize commands amid their own speakers and motor noise.
Voice ordering and on-device dialogue in store noise
Captures the user's voice with a multi-channel array, reduces store noise and prompt echo, then understands order intent and options through recognition and a conversational language model.
Robot full stack (wake word · direction · noise sources · speaker · mics + pipeline)
Wake word · user direction · motor/fan noise · robot speaker · multi-channel mics
mpWWD → mpLocalization → mpAB → mpASR → mpLLM
Conversational Kiosk
From recognition in store noise to on-device conversational ordering
Accurate recognition alone is not enough for kiosk voice ordering.
Users phrase menus and options in many ways, missing information requires follow-up questions, and the final result must reach the ordering system.
mpWAV connects a linear microphone array, mpAB preprocessing, mpASR recognition, and mpLLM — a 1B-class on-device dialogue model direction — into a conversational ordering pipeline.
Linear mic array
Captures far-field user speech in front of the kiosk on multiple channels.
mpAB
Processes store conversation, background music, and kiosk prompt echo together.
mpASR
Converts menu names, options, quantities, and order utterances into text.
mpLLM
Interprets intent and context, asking follow-up questions for missing information.
Example pipeline
User Speech → Linear Mic Array → mpAB
→ mpASR → mpLLM
→ Menu · Option · Quantity → Ordering API
Example conversation
User“Recommend an iced drink that's not too sweet.”
From wake word to natural-language commands, amid driving noise and car audio
A car cabin combines road noise, wind, HVAC, music, and passenger conversation at once.
The voice interface must recognize the wake word, identify which seat or direction spoke, cancel car audio echo, and then understand commands and natural-language requests.
mpWWD
Detects the automotive wake word.
mpLocalization
Estimates speech direction and seat position.
mpAEC
Reduces echo from car audio and voice prompts.
mpBeamforming·mpAB
Reduces driving noise and passenger conversation.
mpASR
Recognizes navigation, climate, and media commands.
mpLLM
Connects natural-language requests and context to vehicle functions.
Example pipeline
Wake Word → Seat/Direction Detection
→ Cabin Mic Array + Audio Reference → mpAB
→ mpASR → mpLLM / Vehicle Intent
→ Infotainment · Navigation · Control
Official material describes reviewing noise robustness and global expandability in the mobility field.
Direction estimation · mpAEC·Beamforming processing
Smart Devices and Earbuds
AI voice technology for devices that can't add microphones
Earbuds, wearables, and small smart devices often cannot fit a multi-microphone array due to size, layout, and power constraints.
mpNC reduces the impact of ambient noise from a single microphone input, making voice input improvement and a voice interface feasible on single-mic products.
mpNC
Reduces everyday ambient noise from a single microphone input.
mpWWD
Detects the wake word to activate product functions.
mpASR
Recognizes commands and short utterances on-device.
mpLLM
Connects product functions with natural-language requests.
Example pipeline
Single Microphone → mpNC → mpWWD → mpASR → mpLLM → Device Function
Applicable products
Wireless earbuds
Phone accessories
Wearables
Portable smart devices
Small appliances
Voice-controlled IoT devices
Products that cannot fit a multi-channel array are covered by mpNC-based single-microphone processing.
Single Mic → mpNC → Wake Word → ASR → Device Action
Meeting and Voice Chat
Separating voices and telling speakers apart where many people talk
Meetings and voice chat can combine multiple talkers, speaker echo, far-field voices, and overlapping conversation.
mpWAV uses mpS to prepare multi-party input for downstream processing, mpDiarization to tell who spoke when, then mpASR and mpLLM to extend into transcripts, summaries, and Q&A.
mpS
Separates and organizes multi-party or overlapping input for downstream processing.
mpDiarization
Separates per-speaker segments and speaker turns.
mpAEC
Reduces echo from speakers or remote participants re-entering the microphone.
mpBeamforming
Reduces far-field voices and ambient noise in the meeting room.
Analyzing equipment anomalies amid complex production noise
On a production line, many motors and machines run at once, mixing normal and abnormal sounds.
Building on its real-environment noise processing, mpWAV organizes equipment acoustic signals and implements analysis that distinguishes normal from abnormal patterns.
Acoustic preprocessing
Extracts the target equipment's acoustic signature from complex line noise.
Anomaly model
Detects state changes by comparing normal and abnormal data.
Edge processing
Collects and analyzes acoustic data close to the equipment.
License·co-development
Jointly develops data-driven models fit to your equipment and production environment.
Example pipeline
Factory Acoustic Input → Noise-Robust Processing
→ Acoustic Feature Analysis → Normal / Anomaly Classification
→ Inspection · Monitoring · Alert
Applicable areas
Motor anomaly sounds
Rotating equipment
Line acoustic inspection
Equipment condition monitoring
In-line quality inspection
Official material presents validating motor anomaly detection on a noisy production line without a separate test chamber. Detection accuracy, equipment types, and production status are not specified in public material.
Extending voice technology to everyday conversational clarity
ClearSense Audio is a smart listening solution that helps people hear conversation in ambient noise more clearly, using a smartphone and ordinary earphones.
It extends mpWAV's single-microphone noise control and voice enhancement into B2C products and public hearing welfare.
mpNC·Voice enhancement
Reduces ambient noise in conversation captured by the microphone.
Mobile processing
Supports product structures using a smartphone and ordinary earphones.
User-centered interface
Everyday listening that is easier to access than specialist equipment.
Usage environments
Family conversation · restaurants & cafés · welfare center programs · public hearing welfare · everyday listening
Official material presents ClearSense Audio's welfare center validation and public support program experience.
Process timeline (scenario → analysis → data → stack → PoC → integration → API → field → production)
Scenario → Acoustic analysis → Data → Stack → PoC → Integration → API → Field → Production
FAQ
Frequently asked questions about mpWAV applications
Robots, kiosks, vehicles, smart devices and earbuds, meetings and voice chat, factory acoustic anomaly detection, and hearing assistance.
Defense and special environments are handled as dedicated hardware and co-development areas.
Yes.
mpNC reduces ambient noise from a single microphone input — the technical foundation for earbuds and small smart devices.
mpWWD detects the wake word while mpLocalization uses multi-microphone input to estimate the direction or position of the voice.
mpAB processes store noise and prompt echo, mpASR converts speech to text, and mpLLM interprets menu, options, quantity, and context.
The real ordering API and target hardware still need validation.
mpS and mpDiarization process multi-party input and separate per-speaker segments.
Connected to mpASR, this extends to per-speaker transcripts.
Yes.
mpNC, mpAEC, mpBeamforming, and mpAB apply in front of your existing ASR, improving the input audio.
No.
Depending on microphone count, speakers, user position, number of talkers, and required features, you can select only what you need.
Yes.
Starting from the usage scenario and required voice features, we can review microphone structure, the technology combination, and data collection conditions.
No.
Robots, kiosks, mobility, factories, and hearing assistance have deployment or validation experience in official material. Smart devices, meetings and voice chat, and defense are presented as capabilities and expansion directions.
Design Your Application Stack
Design the voice technology your product needs — starting from the application
Tell us your product type, microphone count, speaker structure, dominant noise, and the features you want to build — we will review the right mpWAV stack.
From single-mic noise control to wake·location·speaker analysis, multi-channel preprocessing, on-device ASR, and a conversational LLM — connected to fit your product environment.
mpWAV develops and licenses voice interface technology that makes speech clear in noisy, real-world environments, built on 25+ years of speech signal processing research.