Turns years of unsorted mp3/flac/wav dumps into one clean library tagged with MusicBrainz data and free of duplicates, then keeps it that way as new music arrives.
Deterministic code and beets do the mechanical work: scanning, fingerprinting, matching and copying. A Claude Code agent handles the judgment calls, such as ambiguous matches, which duplicate to keep, and tags that look wrong. It proposes decisions in batches, and nothing is recorded until you approve.
dump ──► inventory ──► verify tags ──► dedupe ──► match ──► import ──► library
(hash + (vs AcoustID) (3 tiers) (MusicBrainz)
fingerprint) │
└─► review queue ──► you approve
The tools are built to be hard to misuse on a collection you care about:
--apply.You need Python 3.12, uv, Chromaprint (fpcalc) and ffprobe (from ffmpeg).
uv sync
Point musiclib.toml at your directories:
source_dir = "/path/to/unsorted/dump" # read-only
library_dir = "/path/to/clean/library"
inbox_dir = "/path/to/inbox"
state_dir = "state"
Put secrets in a gitignored musiclib.local.toml:
acoustid_key = "..."
beets runs with its own generated config under state/beets/, so it won't touch a beets setup you already have.
Every command prints JSON. A typical first run:
uv run musiclib inventory # scan the dump (read-only, resumable)
uv run musiclib report # formats, tag coverage, duplicates
uv run musiclib acoustid # fingerprint lookups
uv run musiclib verify # check tags against fingerprints
uv run musiclib dupes # propose duplicate keepers
uv run musiclib match # MusicBrainz matching (dry run)
uv run musiclib import # show the import plan
uv run musiclib import --apply # actually import
Run Claude Code in the repo for the parts that need judgment:
/review works through albums that didn't match strongly, in batches you approve./ingest processes new folders in the inbox.CLAUDE.md has the full command reference.
uv run pytest
The tests include a check that the source dump is byte-for-byte unchanged after an inventory.
The design lives in docs/SPEC.md, a synced copy of the canonical spec doc. It covers the architecture, safety rules, duplicate policy, decision log and milestones.
Turns years of unsorted mp3/flac/wav dumps into one clean library tagged with MusicBrainz data and free of duplicates, then keeps it that way as new music arrives.
Deterministic code and beets do the mechanical work: scanning, fingerprinting, matching and copying. A Claude Code agent handles the judgment calls, such as ambiguous matches, which duplicate to keep, and tags that look wrong. It proposes decisions in batches, and nothing is recorded until you approve.
dump ──► inventory ──► verify tags ──► dedupe ──► match ──► import ──► library
(hash + (vs AcoustID) (3 tiers) (MusicBrainz)
fingerprint) │
└─► review queue ──► you approve
The tools are built to be hard to misuse on a collection you care about:
--apply.You need Python 3.12, uv, Chromaprint (fpcalc) and ffprobe (from ffmpeg).
uv sync
Point musiclib.toml at your directories:
source_dir = "/path/to/unsorted/dump" # read-only
library_dir = "/path/to/clean/library"
inbox_dir = "/path/to/inbox"
state_dir = "state"
Put secrets in a gitignored musiclib.local.toml:
acoustid_key = "..."
beets runs with its own generated config under state/beets/, so it won't touch a beets setup you already have.
Every command prints JSON. A typical first run:
uv run musiclib inventory # scan the dump (read-only, resumable)
uv run musiclib report # formats, tag coverage, duplicates
uv run musiclib acoustid # fingerprint lookups
uv run musiclib verify # check tags against fingerprints
uv run musiclib dupes # propose duplicate keepers
uv run musiclib match # MusicBrainz matching (dry run)
uv run musiclib import # show the import plan
uv run musiclib import --apply # actually import
Run Claude Code in the repo for the parts that need judgment:
/review works through albums that didn't match strongly, in batches you approve./ingest processes new folders in the inbox.CLAUDE.md has the full command reference.
uv run pytest
The tests include a check that the source dump is byte-for-byte unchanged after an inventory.
The design lives in docs/SPEC.md, a synced copy of the canonical spec doc. It covers the architecture, safety rules, duplicate policy, decision log and milestones.