alexhegit/Playing-with-ROCm

See how to play with ROCm, run it with AMD GPUs!

44

stars

142

commits

Jupyter Notebook

primary language

Mar 25, 2026

updated

README

Playing-with-ROCm

Here to show my experience about playing with ROCm with runable code, step-by-step tutorial to help you reproduce what I have did. If you have iGPU or dGPU of AMD, you may try Machine Learning with them.

NOTICE : For more easier tracking my update, I use πŸ†• and πŸ”₯ to flag the new hot topics.

Topics

Pysical AI & Robotics

Training

Finetuning

Inference

MLOPS with ROCm

Application/Demo


Projects work over ROCm

These projects may not offical announce to support ROCm GPU. But they work fine base on my verification.

NameURLCategoryHands on
CLM-4-Voicehttps://github.com/THUDM/GLM-4-VoiceConversation AI
EchoMimichttps://github.com/BadToBest/EchoMimicDigital Human GenAIRun EchoMimic with ROCm
Easy-Wav2Liphttps://github.com/anothermartz/Easy-Wav2LipDigital Human GenAIEasy-Wav2Lip-ROCm
GOT-OCR2https://github.com/Ucas-HaoranWei/GOT-OCR2.0end2end OCR
Moshihttps://github.com/kyutai-labs/moshiConversation AI
mini-omnihttps://github.com/gpt-omni/mini-omniConversation AI
mini-omni2https://github.com/gpt-omni/mini-omni2Conversation AI
Picovoice/orcahttps://github.com/Picovoice/orcaConversation AILLM_Voice_Assistant
Retrieval-based-Voice-Conversion-WebUIhttps://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI.gitEasily train a good VC model with voice data <= 10 mins!
Freeze-Omni πŸ†• πŸ”₯https://github.com/VITA-MLLM/Freeze-OmniA Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLMRealtime on Radeon W7900, realtime with good response, feel good than Moshi, mini-omni2
Step-Auido πŸ†• πŸ”₯https://github.com/stepfun-ai/Step-AudioConvseration AIToo big model, not real time
Step-Video-T2V πŸ†• πŸ”₯https://github.com/stepfun-ai/Step-Video-T2VVideo GenAIRun with 1xMI300X
UI-TARShttps://github.com/bytedance/UI-TARSAutomated GUI Interaction with Native Agentsfrom ByteDance
Qwen2.5-Omni πŸ†• πŸ”₯https://github.com/QwenLM/Qwen2.5-Omniend-to-end multimodal model in the Qwen serie
CosyVoicehttps://github.com/FunAudioLLM/CosyVoiceTTS LLMtutorial , conda-env

Wish List

NameURLCategoryHands on
hertz-devhttps://github.com/Standard-Intelligence/hertz-devConversation AI
Freeze-Omnihttps://github.com/VITA-MLLM/Freeze-OmniConversation AI
LLaMA-Omnihttps://github.com/ictnlp/LLaMA-OmniConversation AI
ichigo Llama 3.1https://github.com/homebrewltd/ichigoConversation AI
ichigo-demohttps://github.com/homebrewltd/ichigo-demo/tree/docker
Exohttps://github.com/exo-explore/exoheterogeneous distribute inference
Perpleicahttps://github.com/ItzCrazyKns/PerplexicaAI Search Engineissue
MiniPerplxhttps://github.com/zaidmukaddam/miniperplxA minimalistic AI-powered search engine
ollama-helmhttps://github.com/otwld/ollama-helm
OpenHandshttps://github.com/All-Hands-AI/OpenHandsa platform for software development agents powered by AI
HayStackhttps://github.com/deepset-ai/haystackend-to-end LLM framework that allows you to build applications powered by LLMs
Bailinghttps://github.com/ictnlp/BayLing
Bailinghttps://github.com/wwbin2017/bailing
BabelDuckhttps://github.com/Orenoid/BabelDuckBeginner-friendly AI conversation practice application
KubeAIhttps://github.com/substratusai/kubeaideploy and manage AI models on Kubernetes
DSPyhttps://dspy.aithe framework for programming
KServehttps://kserve.github.io/website/latest/
Camel-ai/OWLhttps://github.com/camel-ai/owl
VITAhttps://github.com/VITA-MLLM/VITAVITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
DiffRhythmhttps://github.com/ASLP-lab/DiffRhythmEnd-to-End Full-Length Song Generation with Latent Diffusion
Open-Sorahttps://github.com/hpcaitech/Open-Sora
Real-Time-Voice-Cloninghttps://github.com/CorentinJ/Real-Time-Voice-Cloning
OpenVoicehttps://github.com/myshell-ai/OpenVoice
KrilinAIhttps://github.com/krillinai/KrillinAI
RealtimeVoiceChathttps://github.com/KoljaB/RealtimeVoiceChat
pipecathttps://github.com/pipecat-ai/pipecat

Tracing

Misc

MCP

3rd-stuff


@misc{ Playing with ROCm,
  author = {He Ye (Alex)},
  title = {Playing with ROCm: share my experience and practice},
  howpublished = {\url{https://alexhegit.github.io/}},
  year = {2024--}
}

Contributors

alexhegit

142 commits

alexhegit/Playing-with-ROCm

See how to play with ROCm, run it with AMD GPUs!

44

stars

142

commits

Jupyter Notebook

primary language

Mar 25, 2026

updated

README

Playing-with-ROCm

Here to show my experience about playing with ROCm with runable code, step-by-step tutorial to help you reproduce what I have did. If you have iGPU or dGPU of AMD, you may try Machine Learning with them.

NOTICE : For more easier tracking my update, I use πŸ†• and πŸ”₯ to flag the new hot topics.

Topics

Pysical AI & Robotics

Training

Finetuning

Inference

MLOPS with ROCm

Application/Demo


Projects work over ROCm

These projects may not offical announce to support ROCm GPU. But they work fine base on my verification.

NameURLCategoryHands on
CLM-4-Voicehttps://github.com/THUDM/GLM-4-VoiceConversation AI
EchoMimichttps://github.com/BadToBest/EchoMimicDigital Human GenAIRun EchoMimic with ROCm
Easy-Wav2Liphttps://github.com/anothermartz/Easy-Wav2LipDigital Human GenAIEasy-Wav2Lip-ROCm
GOT-OCR2https://github.com/Ucas-HaoranWei/GOT-OCR2.0end2end OCR
Moshihttps://github.com/kyutai-labs/moshiConversation AI
mini-omnihttps://github.com/gpt-omni/mini-omniConversation AI
mini-omni2https://github.com/gpt-omni/mini-omni2Conversation AI
Picovoice/orcahttps://github.com/Picovoice/orcaConversation AILLM_Voice_Assistant
Retrieval-based-Voice-Conversion-WebUIhttps://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI.gitEasily train a good VC model with voice data <= 10 mins!
Freeze-Omni πŸ†• πŸ”₯https://github.com/VITA-MLLM/Freeze-OmniA Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLMRealtime on Radeon W7900, realtime with good response, feel good than Moshi, mini-omni2
Step-Auido πŸ†• πŸ”₯https://github.com/stepfun-ai/Step-AudioConvseration AIToo big model, not real time
Step-Video-T2V πŸ†• πŸ”₯https://github.com/stepfun-ai/Step-Video-T2VVideo GenAIRun with 1xMI300X
UI-TARShttps://github.com/bytedance/UI-TARSAutomated GUI Interaction with Native Agentsfrom ByteDance
Qwen2.5-Omni πŸ†• πŸ”₯https://github.com/QwenLM/Qwen2.5-Omniend-to-end multimodal model in the Qwen serie
CosyVoicehttps://github.com/FunAudioLLM/CosyVoiceTTS LLMtutorial , conda-env

Wish List

NameURLCategoryHands on
hertz-devhttps://github.com/Standard-Intelligence/hertz-devConversation AI
Freeze-Omnihttps://github.com/VITA-MLLM/Freeze-OmniConversation AI
LLaMA-Omnihttps://github.com/ictnlp/LLaMA-OmniConversation AI
ichigo Llama 3.1https://github.com/homebrewltd/ichigoConversation AI
ichigo-demohttps://github.com/homebrewltd/ichigo-demo/tree/docker
Exohttps://github.com/exo-explore/exoheterogeneous distribute inference
Perpleicahttps://github.com/ItzCrazyKns/PerplexicaAI Search Engineissue
MiniPerplxhttps://github.com/zaidmukaddam/miniperplxA minimalistic AI-powered search engine
ollama-helmhttps://github.com/otwld/ollama-helm
OpenHandshttps://github.com/All-Hands-AI/OpenHandsa platform for software development agents powered by AI
HayStackhttps://github.com/deepset-ai/haystackend-to-end LLM framework that allows you to build applications powered by LLMs
Bailinghttps://github.com/ictnlp/BayLing
Bailinghttps://github.com/wwbin2017/bailing
BabelDuckhttps://github.com/Orenoid/BabelDuckBeginner-friendly AI conversation practice application
KubeAIhttps://github.com/substratusai/kubeaideploy and manage AI models on Kubernetes
DSPyhttps://dspy.aithe framework for programming
KServehttps://kserve.github.io/website/latest/
Camel-ai/OWLhttps://github.com/camel-ai/owl
VITAhttps://github.com/VITA-MLLM/VITAVITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
DiffRhythmhttps://github.com/ASLP-lab/DiffRhythmEnd-to-End Full-Length Song Generation with Latent Diffusion
Open-Sorahttps://github.com/hpcaitech/Open-Sora
Real-Time-Voice-Cloninghttps://github.com/CorentinJ/Real-Time-Voice-Cloning
OpenVoicehttps://github.com/myshell-ai/OpenVoice
KrilinAIhttps://github.com/krillinai/KrillinAI
RealtimeVoiceChathttps://github.com/KoljaB/RealtimeVoiceChat
pipecathttps://github.com/pipecat-ai/pipecat

Tracing

Misc

MCP

3rd-stuff


@misc{ Playing with ROCm,
  author = {He Ye (Alex)},
  title = {Playing with ROCm: share my experience and practice},
  howpublished = {\url{https://alexhegit.github.io/}},
  year = {2024--}
}

Contributors

alexhegit

142 commits

Languages

Jupyter Notebook

97.1%

Python

1.8%