Here to show my experience about playing with ROCm with runable code, step-by-step tutorial to help you reproduce what I have did. If you have iGPU or dGPU of AMD, you may try Machine Learning with them.
NOTICE : For more easier tracking my update, I use π and π₯ to flag the new hot topics.
These projects may not offical announce to support ROCm GPU. But they work fine base on my verification.
| Name | URL | Category | Hands on |
|---|---|---|---|
| CLM-4-Voice | https://github.com/THUDM/GLM-4-Voice | Conversation AI | |
| EchoMimic | https://github.com/BadToBest/EchoMimic | Digital Human GenAI | Run EchoMimic with ROCm |
| Easy-Wav2Lip | https://github.com/anothermartz/Easy-Wav2Lip | Digital Human GenAI | Easy-Wav2Lip-ROCm |
| GOT-OCR2 | https://github.com/Ucas-HaoranWei/GOT-OCR2.0 | end2end OCR | |
| Moshi | https://github.com/kyutai-labs/moshi | Conversation AI | |
| mini-omni | https://github.com/gpt-omni/mini-omni | Conversation AI | |
| mini-omni2 | https://github.com/gpt-omni/mini-omni2 | Conversation AI | |
| Picovoice/orca | https://github.com/Picovoice/orca | Conversation AI | LLM_Voice_Assistant |
| Retrieval-based-Voice-Conversion-WebUI | https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI.git | Easily train a good VC model with voice data <= 10 mins! | |
| Freeze-Omni π π₯ | https://github.com/VITA-MLLM/Freeze-Omni | A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM | Realtime on Radeon W7900, realtime with good response, feel good than Moshi, mini-omni2 |
| Step-Auido π π₯ | https://github.com/stepfun-ai/Step-Audio | Convseration AI | Too big model, not real time |
| Step-Video-T2V π π₯ | https://github.com/stepfun-ai/Step-Video-T2V | Video GenAI | Run with 1xMI300X |
| UI-TARS | https://github.com/bytedance/UI-TARS | Automated GUI Interaction with Native Agentsfrom ByteDance | |
| Qwen2.5-Omni π π₯ | https://github.com/QwenLM/Qwen2.5-Omni | end-to-end multimodal model in the Qwen serie | |
| CosyVoice | https://github.com/FunAudioLLM/CosyVoice | TTS LLM | tutorial , |
@misc{ Playing with ROCm,
author = {He Ye (Alex)},
title = {Playing with ROCm: share my experience and practice},
howpublished = {\url{https://alexhegit.github.io/}},
year = {2024--}
}
142 commits
Jupyter Notebook
97.1%
Python
1.8%
Here to show my experience about playing with ROCm with runable code, step-by-step tutorial to help you reproduce what I have did. If you have iGPU or dGPU of AMD, you may try Machine Learning with them.
NOTICE : For more easier tracking my update, I use π and π₯ to flag the new hot topics.
These projects may not offical announce to support ROCm GPU. But they work fine base on my verification.
| Name | URL | Category | Hands on |
|---|---|---|---|
| CLM-4-Voice | https://github.com/THUDM/GLM-4-Voice | Conversation AI | |
| EchoMimic | https://github.com/BadToBest/EchoMimic | Digital Human GenAI | Run EchoMimic with ROCm |
| Easy-Wav2Lip | https://github.com/anothermartz/Easy-Wav2Lip | Digital Human GenAI | Easy-Wav2Lip-ROCm |
| GOT-OCR2 | https://github.com/Ucas-HaoranWei/GOT-OCR2.0 | end2end OCR | |
| Moshi | https://github.com/kyutai-labs/moshi | Conversation AI | |
| mini-omni | https://github.com/gpt-omni/mini-omni | Conversation AI | |
| mini-omni2 | https://github.com/gpt-omni/mini-omni2 | Conversation AI | |
| Picovoice/orca | https://github.com/Picovoice/orca | Conversation AI | LLM_Voice_Assistant |
| Retrieval-based-Voice-Conversion-WebUI | https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI.git | Easily train a good VC model with voice data <= 10 mins! | |
| Freeze-Omni π π₯ | https://github.com/VITA-MLLM/Freeze-Omni | A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM | Realtime on Radeon W7900, realtime with good response, feel good than Moshi, mini-omni2 |
| Step-Auido π π₯ | https://github.com/stepfun-ai/Step-Audio | Convseration AI | Too big model, not real time |
| Step-Video-T2V π π₯ | https://github.com/stepfun-ai/Step-Video-T2V | Video GenAI | Run with 1xMI300X |
| UI-TARS | https://github.com/bytedance/UI-TARS | Automated GUI Interaction with Native Agentsfrom ByteDance | |
| Qwen2.5-Omni π π₯ | https://github.com/QwenLM/Qwen2.5-Omni | end-to-end multimodal model in the Qwen serie | |
| CosyVoice | https://github.com/FunAudioLLM/CosyVoice | TTS LLM | tutorial , |
@misc{ Playing with ROCm,
author = {He Ye (Alex)},
title = {Playing with ROCm: share my experience and practice},
howpublished = {\url{https://alexhegit.github.io/}},
year = {2024--}
}
142 commits
Jupyter Notebook
97.1%
Python
1.8%