Voice-agent models diverge most beyond choosing the correct tool

MTVA-Bench isolates the language model in cascaded voice systems and tests multilingual, multi-turn tool use under realistic transcript conditions.

Independent
Pritish Mishra · Ishaan Kumar · Akshat Mandoli · Sudarshan Kamath
Research Digest··2 min read
Mishra et al.

The authors built MTVA-Bench around 49 voice-agent configurations and 490 reviewed scenarios spanning seven languages.

Why this paper

Independent · Part of Agent Harness Optimization, now 61 papers

In one line

Most voice agent failures come from argument values and rule compliance, not from selecting the wrong tool.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.