This repository provides the DuplexOmni model weights. Both the Thinker and Talker components are included in the released checkpoint.
For source code and deployment details, see the DuplexOmni repository.
This model is still under active optimization and may exhibit the following issues:
For low-latency deployment, we recommend using at least 8 NVIDIA H20 GPUs. This recommendation reflects the current optimization level and may change as inference performance improves.
The model should be warmed up after deployment. Without warm-up, the first response packet can have significantly higher latency.
本仓库提供 DuplexOmni 模型权重,其中同时包含 Thinker 和 Talker 两个组件。
源码及部署相关信息请参考 DuplexOmni 仓库。
该模型目前仍在持续优化中,可能存在以下问题:
为获得较低的推理延迟,建议至少使用 8 张 NVIDIA H20 GPU 进行部署。该建议基于当前的优化水平,后续可能会随着推理性能优化而调整。
模型部署并挂载后需要进行预热;否则,首个响应包的延迟可能会明显偏高。
33 commits
This repository provides the DuplexOmni model weights. Both the Thinker and Talker components are included in the released checkpoint.
For source code and deployment details, see the DuplexOmni repository.
This model is still under active optimization and may exhibit the following issues:
For low-latency deployment, we recommend using at least 8 NVIDIA H20 GPUs. This recommendation reflects the current optimization level and may change as inference performance improves.
The model should be warmed up after deployment. Without warm-up, the first response packet can have significantly higher latency.
本仓库提供 DuplexOmni 模型权重,其中同时包含 Thinker 和 Talker 两个组件。
源码及部署相关信息请参考 DuplexOmni 仓库。
该模型目前仍在持续优化中,可能存在以下问题:
为获得较低的推理延迟,建议至少使用 8 张 NVIDIA H20 GPU 进行部署。该建议基于当前的优化水平,后续可能会随着推理性能优化而调整。
模型部署并挂载后需要进行预热;否则,首个响应包的延迟可能会明显偏高。
33 commits