NVIDIA 全双工语音模型工具调用新架构
中文摘要
NVIDIA 全双工语音新架构:模型发委托token给LLM执行工具调用。结果经轻量机制回传,再由TTS播报。
English Summary
NVIDIA's new full-duplex speech architecture uses delegation tokens to enable LLM tool calling with streaming transcripts. Results return via lightweight mechanism, then TTS output.
原文节选
NVIDIA 提出一种前后端架构,让全双工语音模型通过发出委托 token,将流式转录转发给文本后端 LLM 执行工具调用,再经轻量 prefill-and-repeat 机制回传结果并由 TTS 播报。