Step-Audio-AQAA: A Fully End-to-End Expressive Large Audio Language Model
49
6 commits
1 linked in READMEs
updated Jun 12, 2025
📚 Paper: Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
Step-Audio-AQAA is a fully end-to-end Large Audio-Language Model (LALM) designed for Audio Query-Audio Answer (AQAA) tasks. It directly processes audio inputs and generates natural, accurate speech responses without relying on traditional ASR and TTS modules, eliminating cascading errors and simplifying the system architecture.
Step-Audio-AQAA consists of three core modules:
@misc{huang2025stepaudioaqaa,
title={Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model},
author={Ailin Huang and Boyong Wu and Bruce Wang and Chao Yan and Chen Hu and Chengli Feng and Fei Tian and Feiyu Shen and Jingbei Li and Mingrui Chen and et al.},
year={2025},
eprint={2506.08967},
archivePrefix={arXiv},
primaryClass={cs.SD}
}
Step-Audio-AQAA is developed by the StepFun team, with contributions from multiple researchers and engineers. For technical support or collaboration, contact the corresponding authors: Daxin Jiang (djiang@stepfun.com), Shuchang Zhou (scotzhou@stepfun.com), Chen Hu (hatcher@stepfun.com).
This model is released under the Apache 2.0 license. For more details, please refer to the license file.
Step-Audio-AQAA: A Fully End-to-End Expressive Large Audio Language Model
49
6 commits
1 linked in READMEs
updated Jun 12, 2025
📚 Paper: Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
Step-Audio-AQAA is a fully end-to-end Large Audio-Language Model (LALM) designed for Audio Query-Audio Answer (AQAA) tasks. It directly processes audio inputs and generates natural, accurate speech responses without relying on traditional ASR and TTS modules, eliminating cascading errors and simplifying the system architecture.
Step-Audio-AQAA consists of three core modules:
@misc{huang2025stepaudioaqaa,
title={Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model},
author={Ailin Huang and Boyong Wu and Bruce Wang and Chao Yan and Chen Hu and Chengli Feng and Fei Tian and Feiyu Shen and Jingbei Li and Mingrui Chen and et al.},
year={2025},
eprint={2506.08967},
archivePrefix={arXiv},
primaryClass={cs.SD}
}
Step-Audio-AQAA is developed by the StepFun team, with contributions from multiple researchers and engineers. For technical support or collaboration, contact the corresponding authors: Daxin Jiang (djiang@stepfun.com), Shuchang Zhou (scotzhou@stepfun.com), Chen Hu (hatcher@stepfun.com).
This model is released under the Apache 2.0 license. For more details, please refer to the license file.