Real-Time Streaming Voice Agent in Interactive Scenarios (SAP Collaboration)
Published:
Overview
In collaboration with SAP, this project focused on building an enterprise-ready, low-latency, full-duplex conversational voice agent for smart customer service and anthropomorphic interaction scenarios.
Key Technical Contributions
- End-to-End Pipeline Architecture: Led the design and deployment of a closed-loop system encompassing streaming ASR (Speech Recognition) → LLM multi-turn intent reasoning & task planning → streaming TTS (Speech Synthesis), enabling fluid, human-like voice conversations.
- Context Routing & Barge-in Handling: Developed dynamic context stitching and intent routing with sliding windows and key entity summarization, robustly managing long-horizon context consistency and seamless user barge-in (interruption) flow control.
- Extreme Latency Optimization: Re-architected audio frame ingestion, chunked streaming transport, and token-level incremental synthesis using
asyncio, WebSockets, and multithreaded Ring Buffers; integrated vLLM Chunked Prefill and INT8 quantization, compressing end-to-end first-frame voice response latency from 3.4s to under 0.9s.
Tech Stack & Tools
- Backend & Networking: FastAPI, WebSockets, asyncio, Python
- LLM Serving & Quantization: vLLM (Chunked Prefill, PagedAttention), INT8 Quantization
- Audio Pipelines: Streaming ASR, Streaming TTS, Ring Buffer
