Real-Time Streaming Voice Agent in Interactive Scenarios (SAP Collaboration)

Published:

Overview

In collaboration with SAP, this project focused on building an enterprise-ready, low-latency, full-duplex conversational voice agent for smart customer service and anthropomorphic interaction scenarios.

Key Technical Contributions

  • End-to-End Pipeline Architecture: Led the design and deployment of a closed-loop system encompassing streaming ASR (Speech Recognition) → LLM multi-turn intent reasoning & task planning → streaming TTS (Speech Synthesis), enabling fluid, human-like voice conversations.
  • Context Routing & Barge-in Handling: Developed dynamic context stitching and intent routing with sliding windows and key entity summarization, robustly managing long-horizon context consistency and seamless user barge-in (interruption) flow control.
  • Extreme Latency Optimization: Re-architected audio frame ingestion, chunked streaming transport, and token-level incremental synthesis using asyncio, WebSockets, and multithreaded Ring Buffers; integrated vLLM Chunked Prefill and INT8 quantization, compressing end-to-end first-frame voice response latency from 3.4s to under 0.9s.

Tech Stack & Tools

  • Backend & Networking: FastAPI, WebSockets, asyncio, Python
  • LLM Serving & Quantization: vLLM (Chunked Prefill, PagedAttention), INT8 Quantization
  • Audio Pipelines: Streaming ASR, Streaming TTS, Ring Buffer