A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
See the code中文 | English | User Guide | Quickstart | Paper
Real conversation should not leave you waiting after a single sentence, nor should it grind to a halt just because the Agent is looking something up, calling a tool, or working on a task.
Conversation should keep flowing, and the Agent should always be present.
That is why we built qwen-audio-agent—a realtime voice runtime that keeps Agents talking, working, and present. Whether chatting with you, thinking through a problem, or working on a task, your Agent remains in the conversation. It listens, responds, and when the task is complete, naturally tells you:
"It's ready."
Conversation doesn't stop for background tasks; when a task completes, the result naturally returns to the current conversation:
| Office | Smart Cockpit |
|---|---|
|
|
Questions that can be answered directly are answered immediately; when tools or sustained processing are needed, the task is delegated to the backend Agent. Throughout, the user always faces the same assistant.
For the full design and module breakdown, see the architecture document.
The voice frontend handles realtime conversation; the backend Agent executes tasks. They integrate independently and can be combined as needed.
| Voice frontend | Deployment | Setup | Features |
|---|---|---|---|
| Qwen Audio 3.0 Realtime | Cloud | Bailian API Key | Duplex voice, tool calling |
| GPT-Live / OpenAI Realtime | Cloud | OpenAI API Key | — |
| Google Gemini Live | Cloud | Google API Key | Live video input |
| Qwen3.5-Omni Realtime | Cloud | Bailian API Key | Live video input |
| Qwen3.8 Omni Flash Realtime | Cloud | Bailian API Key + workspace-specific endpoint | Live video input |
| Doubao Seeduplex 3.0 Realtime | Cloud | Volcengine Speech API Key | — |
| StepAudio 3 Realtime | Cloud | StepFun API Key | — |
| Hugging Face Speech-to-Speech | Local | Start the service and set its URL | Configurable STT / LLM / TTS |
| MiniCPM-o 4.5 | Local or cloud | Compatible service URL | Live video input, no tool calling |
To connect another voice service, implement the Realtime Provider interface without changing the Gateway's core voice-session or backend-task logic.
| Backend Agent | Integration | Setup | Rating |
|---|---|---|---|
| None | N/A | Frontend-only mode, no backend config needed | ★★★★★ |
| Qwen Code | Native ACP | One-click install, user config required | ★★★★★ |
| OpenCode | Native ACP | One-click install + Bailian config | ★★★★★ |
| OpenClaw | Built-in ACP bridge | One-click install + Bailian config | ★★★★★ |
| Qoder | Native ACP | One-click install, user config required | ★★★★★ |
| MiniMax Code | Native ACP | One-click install, user config required | ★★★★☆ |
| Kimi Code | Native ACP | One-click install, user config required | ★★★★★ |
| Hermes | Native ACP | One-click install, user config required | ★★★★☆ |
| CodeBuddy | Native ACP | One-click install, user config required | ★★★★☆ |
| Codex | External ACP adapter | One-click install (base + adapter), user config required | ★★★★☆ |
| Claude Code | External ACP adapter | One-click install (base + adapter), user config required | ★★★★☆ |
| DeepSeek Harness | Native ACP | One-click install, DeepSeek API key required | ★★★★☆ |
| Pi | External ACP adapter | One-click install (base + adapter), user config required | ★★★★☆ |
| Muse Code | Native MSP adapter | Install Muse and its optional SDK on demand; user config required | ★★★☆☆ |
Ratings reflect current integration completeness, compatibility, and verification level: five stars indicate a thoroughly tested recommended integration; four stars indicate active development or not yet fully verified. For detailed configuration and capability boundaries, see the backend Agent documentation and configuration guide.
Requires Node.js 22.22.2+ or 24.15.0+, npm 10+. One-click install (recommended):
npm install -g qwen-audio-agent
For building from source, installing from GitHub, and obtaining a DashScope API Key, see the installation guide.
qwenaudio config
DASHSCOPE_API_KEY=your-key
# Voice frontend model: optional, defaults to Qwen Audio 3.0 Realtime Plus
QWEN_AUDIO_REALTIME_MODEL=qwen-audio-3.0-realtime-plus
# Backend Agent: optional, leave empty or set to none for frontend-only mode
AGENT_PROTOCOL=openclaw
# Backend model: optional; explicit values use standard ACP, empty reuses Agent config
QWEN_AUDIO_AGENT_BACKEND_MODEL=qwen3.7-max
Before starting, create a key from the Bailian API Key page. Eligible new users can review the new-user free quota and check remaining usage on the model usage page. Quota and billing rules are subject to the current official Bailian documentation.
The example above uses the default DashScope voice frontend. See Voice Frontends for other cloud and self-hosted options.
With a visual-capable Realtime frontend, WebUI can explicitly stream bounded camera frames alongside live audio. See Realtime frontend configuration.
qwenaudio webui for the browser UI):qwenaudio # Terminal 1: Gateway
qwenaudio tui # Terminal 2: TUI
For full configuration options, local voice frontend setup, and TUI platform notes, see quick start, voice frontends, and TUI notes.
The desktop app provides a persistent floating voice orb with a built-in Gateway, automatic idle sleep, local voice wake, and customizable appearance. Download the installer for your platform from the releases page, or build from source:
npm run desktop:build:local # macOS
npm run desktop:build:win # Windows
npm run desktop:build:linux # Linux (AppImage + deb, no signing)
For visuals, orb behavior, and build instructions, see the desktop documentation.
The current qwen-audio-agent framework focuses on desktop productivity: users can keep talking with the Agent in realtime while delegating tool use, file work, code changes, and long-running tasks to the backend Agent.
This "foreground conversation + background task" design is not limited to desktop use. It can also expand to more scenarios where the Agent can both chat naturally and get real work done.
| Scenario | Description | Link | Status |
|---|---|---|---|
| Desktop | Voice chat, progress follow-up, tools, and background tasks. | Docs | Available |
| Smart cockpit | Vehicle control, navigation, music, weather, and services. | Example | Available |
| X-Omni | Visual conversation, on-demand capture, optional observation and narration. | Example | Available |
| AI Passport | Qwen Voice Bean on a hardware card, with voice conversation and backend tasks. Currently half-duplex only. | Example | Available |
| Customer Service | Voice customer service for retail and airline scenarios. | Example | Available |
| Embodied intelligence | Voice commands, action execution, inspection, and exception feedback. | TBD | Planned |
| Livestream assistant | Audience interaction, product explanation, coupons, and risk reminders. | TBD | Planned |
You can start discussions directly in GitHub Issues.
For users in China, scan the QR codes below to join the WeChat group. If the group QR code is full or expired, scan either maintainer's personal QR code to be invited.
| WeChat Group | Personal | Personal |
|---|---|---|
![]() | ![]() | ![]() |
530 followers · starred Sep 2026
636 followers · starred Aug 2026
166 followers · starred Aug 2026
243 followers · starred Aug 2026
JavaScript
97.5%
CSS
1.2%
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
See the code中文 | English | User Guide | Quickstart | Paper
Real conversation should not leave you waiting after a single sentence, nor should it grind to a halt just because the Agent is looking something up, calling a tool, or working on a task.
Conversation should keep flowing, and the Agent should always be present.
That is why we built qwen-audio-agent—a realtime voice runtime that keeps Agents talking, working, and present. Whether chatting with you, thinking through a problem, or working on a task, your Agent remains in the conversation. It listens, responds, and when the task is complete, naturally tells you:
"It's ready."
Conversation doesn't stop for background tasks; when a task completes, the result naturally returns to the current conversation:
| Office | Smart Cockpit |
|---|---|
|
|
Questions that can be answered directly are answered immediately; when tools or sustained processing are needed, the task is delegated to the backend Agent. Throughout, the user always faces the same assistant.
For the full design and module breakdown, see the architecture document.
The voice frontend handles realtime conversation; the backend Agent executes tasks. They integrate independently and can be combined as needed.
| Voice frontend | Deployment | Setup | Features |
|---|---|---|---|
| Qwen Audio 3.0 Realtime | Cloud | Bailian API Key | Duplex voice, tool calling |
| GPT-Live / OpenAI Realtime | Cloud | OpenAI API Key | — |
| Google Gemini Live | Cloud | Google API Key | Live video input |
| Qwen3.5-Omni Realtime | Cloud | Bailian API Key | Live video input |
| Qwen3.8 Omni Flash Realtime | Cloud | Bailian API Key + workspace-specific endpoint | Live video input |
| Doubao Seeduplex 3.0 Realtime | Cloud | Volcengine Speech API Key | — |
| StepAudio 3 Realtime | Cloud | StepFun API Key | — |
| Hugging Face Speech-to-Speech | Local | Start the service and set its URL | Configurable STT / LLM / TTS |
| MiniCPM-o 4.5 | Local or cloud | Compatible service URL | Live video input, no tool calling |
To connect another voice service, implement the Realtime Provider interface without changing the Gateway's core voice-session or backend-task logic.
| Backend Agent | Integration | Setup | Rating |
|---|---|---|---|
| None | N/A | Frontend-only mode, no backend config needed | ★★★★★ |
| Qwen Code | Native ACP | One-click install, user config required | ★★★★★ |
| OpenCode | Native ACP | One-click install + Bailian config | ★★★★★ |
| OpenClaw | Built-in ACP bridge | One-click install + Bailian config | ★★★★★ |
| Qoder | Native ACP | One-click install, user config required | ★★★★★ |
| MiniMax Code | Native ACP | One-click install, user config required | ★★★★☆ |
| Kimi Code | Native ACP | One-click install, user config required | ★★★★★ |
| Hermes | Native ACP | One-click install, user config required | ★★★★☆ |
| CodeBuddy | Native ACP | One-click install, user config required | ★★★★☆ |
| Codex | External ACP adapter | One-click install (base + adapter), user config required | ★★★★☆ |
| Claude Code | External ACP adapter | One-click install (base + adapter), user config required | ★★★★☆ |
| DeepSeek Harness | Native ACP | One-click install, DeepSeek API key required | ★★★★☆ |
| Pi | External ACP adapter | One-click install (base + adapter), user config required | ★★★★☆ |
| Muse Code | Native MSP adapter | Install Muse and its optional SDK on demand; user config required | ★★★☆☆ |
Ratings reflect current integration completeness, compatibility, and verification level: five stars indicate a thoroughly tested recommended integration; four stars indicate active development or not yet fully verified. For detailed configuration and capability boundaries, see the backend Agent documentation and configuration guide.
Requires Node.js 22.22.2+ or 24.15.0+, npm 10+. One-click install (recommended):
npm install -g qwen-audio-agent
For building from source, installing from GitHub, and obtaining a DashScope API Key, see the installation guide.
qwenaudio config
DASHSCOPE_API_KEY=your-key
# Voice frontend model: optional, defaults to Qwen Audio 3.0 Realtime Plus
QWEN_AUDIO_REALTIME_MODEL=qwen-audio-3.0-realtime-plus
# Backend Agent: optional, leave empty or set to none for frontend-only mode
AGENT_PROTOCOL=openclaw
# Backend model: optional; explicit values use standard ACP, empty reuses Agent config
QWEN_AUDIO_AGENT_BACKEND_MODEL=qwen3.7-max
Before starting, create a key from the Bailian API Key page. Eligible new users can review the new-user free quota and check remaining usage on the model usage page. Quota and billing rules are subject to the current official Bailian documentation.
The example above uses the default DashScope voice frontend. See Voice Frontends for other cloud and self-hosted options.
With a visual-capable Realtime frontend, WebUI can explicitly stream bounded camera frames alongside live audio. See Realtime frontend configuration.
qwenaudio webui for the browser UI):qwenaudio # Terminal 1: Gateway
qwenaudio tui # Terminal 2: TUI
For full configuration options, local voice frontend setup, and TUI platform notes, see quick start, voice frontends, and TUI notes.
The desktop app provides a persistent floating voice orb with a built-in Gateway, automatic idle sleep, local voice wake, and customizable appearance. Download the installer for your platform from the releases page, or build from source:
npm run desktop:build:local # macOS
npm run desktop:build:win # Windows
npm run desktop:build:linux # Linux (AppImage + deb, no signing)
For visuals, orb behavior, and build instructions, see the desktop documentation.
The current qwen-audio-agent framework focuses on desktop productivity: users can keep talking with the Agent in realtime while delegating tool use, file work, code changes, and long-running tasks to the backend Agent.
This "foreground conversation + background task" design is not limited to desktop use. It can also expand to more scenarios where the Agent can both chat naturally and get real work done.
| Scenario | Description | Link | Status |
|---|---|---|---|
| Desktop | Voice chat, progress follow-up, tools, and background tasks. | Docs | Available |
| Smart cockpit | Vehicle control, navigation, music, weather, and services. | Example | Available |
| X-Omni | Visual conversation, on-demand capture, optional observation and narration. | Example | Available |
| AI Passport | Qwen Voice Bean on a hardware card, with voice conversation and backend tasks. Currently half-duplex only. | Example | Available |
| Customer Service | Voice customer service for retail and airline scenarios. | Example | Available |
| Embodied intelligence | Voice commands, action execution, inspection, and exception feedback. | TBD | Planned |
| Livestream assistant | Audience interaction, product explanation, coupons, and risk reminders. | TBD | Planned |
You can start discussions directly in GitHub Issues.
For users in China, scan the QR codes below to join the WeChat group. If the group QR code is full or expired, scan either maintainer's personal QR code to be invited.
| WeChat Group | Personal | Personal |
|---|---|---|
![]() | ![]() | ![]() |
530 followers · starred Sep 2026
636 followers · starred Aug 2026
166 followers · starred Aug 2026
243 followers · starred Aug 2026
JavaScript
97.5%
CSS
1.2%