🎵 专属你的私人数字调音师|AI 音乐搜索推荐 Agent | 基于大模型 + 知识图谱 + 双模型声学向量的本地智能音乐推荐系统 | LLM-powered Music Recommendation Agent with Hybrid RAG, Neo4j, and Long-term Memory
Python
24
388 commits
updated Sep 11, 2026
A natural-language music recommendation agent
Try the public demo on ModelScope
The Space is configured for AMD MI308X + ROCm; actual GPU availability depends on ModelScope scheduling capacity.
SoulTuner is an open-source music recommendation agent. Describe a mood, scene, sound, artist, or a song you want to avoid in one ordinary sentence. SoulTuner turns that request into a search plan, looks through the music library, and explains why each result fits.
📖 Full feature and interaction details: Feature_Walkthrough.md
![]() | ![]() |
![]() | ![]() |
The language model plans how to search; deterministic application code validates that plan before any retrieval tool runs. The model does not invent a song list and bypass the catalogue.
The default setup can use the Qwen3.7 Plus API. SoulTuner also includes an optional 35B Planner trained specifically for its retrieval contract. On a held-out 500-request planning evaluation, the trained Planner produced valid structured decisions for 99.4% of requests and selected the correct intent and retrieval route for 95.6%. These figures measure planning behaviour, not subjective music quality.
Both Planner options use the same retrieval, memory, ranking, and frontend code. Switching models therefore does not require rewriting the recommendation system.
cd <your project directory>
Copy-Item .env.example .env
notepad .env
Fill in at least these (the default setup uses DashScope / Qwen):
MAIN_LLM_PROVIDER=dashscope
MODEL_NAME=qwen3.7-plus
DASHSCOPE_API_KEY=your DashScope key
NEO4J_PASSWORD=your Neo4j password
MUSIC_DATA_PATH=../data
Then start it and open http://localhost:3003:
.\soultuner.ps1 up gpu
Without an NVIDIA GPU, use .\soultuner.ps1 up cpu. GPU profiles use MuQ as
the primary semantic encoder and OMAR for acoustic reranking; the constrained
CPU profile uses M2D-CLAP instead.
To use another provider (SiliconFlow, Google, Volcengine, or local SGLang / vLLM / Ollama), change MAIN_LLM_PROVIDER and MODEL_NAME and supply the matching key — or adjust it from System Settings in the UI after startup.
SoulTuner accepts an API model or a self-hosted OpenAI-compatible endpoint. The large model can stay on a GPU server while the rest of the application runs on an ordinary computer.
| Option | Best for | What you need |
|---|---|---|
| Qwen3.7 Plus API | the easiest first run | an API key; no large local GPU |
| SoulTuner V4.2 35B | project-specific planning and private hosting | a high-memory inference server |
| Safe demo | UI and retrieval demonstration | CPU only; no external model call |
See the self-hosting package for the model switch, integrity checks, server startup, and benchmark tools. The main Docker deployment supports CPU and NVIDIA CUDA; AMD ROCm deployment is available as an overlay without changing the application code.
| Command | Purpose |
|---|---|
.\soultuner.ps1 doctor | Check that the services are healthy |
.\soultuner.ps1 down | Stop all containers |
.\soultuner.ps1 logs | Tail service logs |
.\soultuner.ps1 test | Run the unit tests |
.\soultuner.ps1 ingest gpu | Process the pending-ingest queue on the GPU worker |
python scripts/dev/start_backend.py | Backend only, for local debugging |
One recommendation request travels this path:
your sentence
│
▼
┌──────────────────────────────────────────────────┐
│ Agent (LangGraph) │
│ recall memory → LLM plan → route by intent │
│ find songs / chat / acquire / clarify │
└──────────────────────┬───────────────────────────┘
▼
┌──────────────────────────────────────────────────┐
│ Retrieval and catalog expansion │
│ graph + MuQ semantics + OMAR rerank → web fill │
└──────────────────────┬───────────────────────────┘
▼
┌──────────────────────────────────────────────────┐
│ Storage: Neo4j (graph + vectors + behaviour) │
│ SQLite (memory ledger + feedback events)│
└──────────────────────┬───────────────────────────┘
▼
SSE streaming → frontend (Next.js)
│
▼
your feedback ─┘ recorded, and updates your taste profile
| Layer | Technology |
|---|---|
| Frontend | Next.js 16 + React 18 |
| Backend | FastAPI + SSE streaming |
| Agent | LangGraph StateGraph |
| Graph database | Neo4j 5.x (relations + native vector index) |
| Text-to-music | MuQ-MuLan + OMAR-RQ on GPU; M2D-CLAP for the CPU profile |
| LLM | dashscope / qwen3.7-plus by default, provider swappable |
| Long-term memory | Local SQLite ledger + Neo4j hot path |
| Ranking | Multi-source fusion → rerank → diversity |
| Deployment | Docker Compose (CPU / GPU entrypoints) |
📖 How to run the recommendation-quality and alignment evaluations: tests/eval/README.md
agent/ LangGraph workflow and intent routing
retrieval/ hybrid retrieval, fusion & ranking, audio encoders, context pipeline
tools/ graph search / text-to-music / web discovery / song acquisition
services/ memory gateway, feedback events, ranking policy, service clients
schemas/ Pydantic contracts (state, query plan, feedback events)
llms/ provider registry and prompts
api/ FastAPI layer
data/ data pipeline and planner distillation harness
web/ Next.js frontend
tests/ unit tests + outcome-oriented evaluation
The planner can be distilled into a local student model. The public repository ships the reproducible harness, while private training data stays outside Git; see data/sft/README.md.
The local CLI, music MCP and optional Gradio adapter reuse the main recommendation API:
These adapters do not start models or databases. Retired GraphZep and standalone search integrations are no longer distributed; see retired services. API-backed validation is not a benchmark of the fine-tuned 35B model, GPU inference, or music playback.
| Variable | Purpose |
|---|---|
DASHSCOPE_API_KEY | Key for the default model (use your provider's key if you switch) |
NEO4J_PASSWORD | Local Neo4j password |
MUSIC_DATA_PATH | Where audio, caches, the ingest queue and feedback logs live |
MUSIC_WEB_SEARCH_ENABLED | Whether web supplementation is allowed |
ADMIN_API_KEY | Optional. Set it and delete / settings / rebuild require the key |
See .env.example for the advanced options; normal use needs none of them.
It listens on 127.0.0.1 only. For remote access use a VPN or SSH tunnel.
Suggestions and bug reports are welcome through GitHub Issues. A pull request is only a proposed change: repository maintainers review it and decide whether it is merged. See CONTRIBUTING.md for the lightweight workflow, SECURITY.md for private vulnerability reports, and CHANGELOG.md for release changes.
The initial architecture came from imagist13/Muisc-Research and has since been substantially rebuilt and extended.
| Project | Used for |
|---|---|
| OpenMuQ/MuQ | MuQ-MuLan, the primary text-to-music model (CC-BY-NC 4.0) |
| nttcslab/m2d | M2D-CLAP encoder for the constrained CPU profile |
| MTG/omar-rq | OMAR-RQ acoustic reranking on GPU profiles |
⚠️ Disclaimer: For study and architecture research. It does not provide, contain or distribute any copyrighted audio or lyrics; obtain audio through lawful channels yourself.
🎵 专属你的私人数字调音师|AI 音乐搜索推荐 Agent | 基于大模型 + 知识图谱 + 双模型声学向量的本地智能音乐推荐系统 | LLM-powered Music Recommendation Agent with Hybrid RAG, Neo4j, and Long-term Memory
Python
24
388 commits
updated Sep 11, 2026
A natural-language music recommendation agent
Try the public demo on ModelScope
The Space is configured for AMD MI308X + ROCm; actual GPU availability depends on ModelScope scheduling capacity.
SoulTuner is an open-source music recommendation agent. Describe a mood, scene, sound, artist, or a song you want to avoid in one ordinary sentence. SoulTuner turns that request into a search plan, looks through the music library, and explains why each result fits.
📖 Full feature and interaction details: Feature_Walkthrough.md
![]() | ![]() |
![]() | ![]() |
The language model plans how to search; deterministic application code validates that plan before any retrieval tool runs. The model does not invent a song list and bypass the catalogue.
The default setup can use the Qwen3.7 Plus API. SoulTuner also includes an optional 35B Planner trained specifically for its retrieval contract. On a held-out 500-request planning evaluation, the trained Planner produced valid structured decisions for 99.4% of requests and selected the correct intent and retrieval route for 95.6%. These figures measure planning behaviour, not subjective music quality.
Both Planner options use the same retrieval, memory, ranking, and frontend code. Switching models therefore does not require rewriting the recommendation system.
cd <your project directory>
Copy-Item .env.example .env
notepad .env
Fill in at least these (the default setup uses DashScope / Qwen):
MAIN_LLM_PROVIDER=dashscope
MODEL_NAME=qwen3.7-plus
DASHSCOPE_API_KEY=your DashScope key
NEO4J_PASSWORD=your Neo4j password
MUSIC_DATA_PATH=../data
Then start it and open http://localhost:3003:
.\soultuner.ps1 up gpu
Without an NVIDIA GPU, use .\soultuner.ps1 up cpu. GPU profiles use MuQ as
the primary semantic encoder and OMAR for acoustic reranking; the constrained
CPU profile uses M2D-CLAP instead.
To use another provider (SiliconFlow, Google, Volcengine, or local SGLang / vLLM / Ollama), change MAIN_LLM_PROVIDER and MODEL_NAME and supply the matching key — or adjust it from System Settings in the UI after startup.
SoulTuner accepts an API model or a self-hosted OpenAI-compatible endpoint. The large model can stay on a GPU server while the rest of the application runs on an ordinary computer.
| Option | Best for | What you need |
|---|---|---|
| Qwen3.7 Plus API | the easiest first run | an API key; no large local GPU |
| SoulTuner V4.2 35B | project-specific planning and private hosting | a high-memory inference server |
| Safe demo | UI and retrieval demonstration | CPU only; no external model call |
See the self-hosting package for the model switch, integrity checks, server startup, and benchmark tools. The main Docker deployment supports CPU and NVIDIA CUDA; AMD ROCm deployment is available as an overlay without changing the application code.
| Command | Purpose |
|---|---|
.\soultuner.ps1 doctor | Check that the services are healthy |
.\soultuner.ps1 down | Stop all containers |
.\soultuner.ps1 logs | Tail service logs |
.\soultuner.ps1 test | Run the unit tests |
.\soultuner.ps1 ingest gpu | Process the pending-ingest queue on the GPU worker |
python scripts/dev/start_backend.py | Backend only, for local debugging |
One recommendation request travels this path:
your sentence
│
▼
┌──────────────────────────────────────────────────┐
│ Agent (LangGraph) │
│ recall memory → LLM plan → route by intent │
│ find songs / chat / acquire / clarify │
└──────────────────────┬───────────────────────────┘
▼
┌──────────────────────────────────────────────────┐
│ Retrieval and catalog expansion │
│ graph + MuQ semantics + OMAR rerank → web fill │
└──────────────────────┬───────────────────────────┘
▼
┌──────────────────────────────────────────────────┐
│ Storage: Neo4j (graph + vectors + behaviour) │
│ SQLite (memory ledger + feedback events)│
└──────────────────────┬───────────────────────────┘
▼
SSE streaming → frontend (Next.js)
│
▼
your feedback ─┘ recorded, and updates your taste profile
| Layer | Technology |
|---|---|
| Frontend | Next.js 16 + React 18 |
| Backend | FastAPI + SSE streaming |
| Agent | LangGraph StateGraph |
| Graph database | Neo4j 5.x (relations + native vector index) |
| Text-to-music | MuQ-MuLan + OMAR-RQ on GPU; M2D-CLAP for the CPU profile |
| LLM | dashscope / qwen3.7-plus by default, provider swappable |
| Long-term memory | Local SQLite ledger + Neo4j hot path |
| Ranking | Multi-source fusion → rerank → diversity |
| Deployment | Docker Compose (CPU / GPU entrypoints) |
📖 How to run the recommendation-quality and alignment evaluations: tests/eval/README.md
agent/ LangGraph workflow and intent routing
retrieval/ hybrid retrieval, fusion & ranking, audio encoders, context pipeline
tools/ graph search / text-to-music / web discovery / song acquisition
services/ memory gateway, feedback events, ranking policy, service clients
schemas/ Pydantic contracts (state, query plan, feedback events)
llms/ provider registry and prompts
api/ FastAPI layer
data/ data pipeline and planner distillation harness
web/ Next.js frontend
tests/ unit tests + outcome-oriented evaluation
The planner can be distilled into a local student model. The public repository ships the reproducible harness, while private training data stays outside Git; see data/sft/README.md.
The local CLI, music MCP and optional Gradio adapter reuse the main recommendation API:
These adapters do not start models or databases. Retired GraphZep and standalone search integrations are no longer distributed; see retired services. API-backed validation is not a benchmark of the fine-tuned 35B model, GPU inference, or music playback.
| Variable | Purpose |
|---|---|
DASHSCOPE_API_KEY | Key for the default model (use your provider's key if you switch) |
NEO4J_PASSWORD | Local Neo4j password |
MUSIC_DATA_PATH | Where audio, caches, the ingest queue and feedback logs live |
MUSIC_WEB_SEARCH_ENABLED | Whether web supplementation is allowed |
ADMIN_API_KEY | Optional. Set it and delete / settings / rebuild require the key |
See .env.example for the advanced options; normal use needs none of them.
It listens on 127.0.0.1 only. For remote access use a VPN or SSH tunnel.
Suggestions and bug reports are welcome through GitHub Issues. A pull request is only a proposed change: repository maintainers review it and decide whether it is merged. See CONTRIBUTING.md for the lightweight workflow, SECURITY.md for private vulnerability reports, and CHANGELOG.md for release changes.
The initial architecture came from imagist13/Muisc-Research and has since been substantially rebuilt and extended.
| Project | Used for |
|---|---|
| OpenMuQ/MuQ | MuQ-MuLan, the primary text-to-music model (CC-BY-NC 4.0) |
| nttcslab/m2d | M2D-CLAP encoder for the constrained CPU profile |
| MTG/omar-rq | OMAR-RQ acoustic reranking on GPU profiles |
⚠️ Disclaimer: For study and architecture research. It does not provide, contain or distribute any copyrighted audio or lyrics; obtain audio through lawful channels yourself.