The authors combined a streaming speech encoder with a decoder-only language model and streaming text-to-speech decoder.
Open voice model combines real-time conversation with native tool calling
NVIDIA’s unified streaming system can listen, speak, handle interruptions, and produce structured function calls during full-duplex conversation.
Big Tech
Jagadeesh Balam · Travis Bartley · Edresson Casanova · Sanjay Chauhan · Chen Chen · Zhehuai Chen · +43 more
NVIDIA
Research Digest··2 min read
Thread:RL for Tool Agents
The authors built NemotronLabs VoiceChat, an open speech-to-speech model that processes incoming audio while simultaneously generating speech, text, and tool calls.
Why this paper
From NVIDIA
In one line
NemotronLabs VoiceChat is a full-duplex speech-to-speech model that natively invokes tools during conversation.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§