nineninesix/gepard

Space

πŸ† GEPARD TTS

40

41 commits

2 linked in READMEs

updated Sep 5, 2026

See the code

README

πŸ† GEPARD TTS

Inference Space for the Gepard autoregressive speech model (nineninesix/gepard-1.0 | nineninesix/gepard-1.1 β€” a Qwen3.5 backbone with 32 FSQ audio heads, a Q-Former voice-cloning compressor, and DPO | GRPO post-training).

Features

  • Preset speakers β€” bundled voices stored as pre-encoded codec tokens (speakers/*.pt), no codec pass needed at request time.
  • Voice cloning β€” record with the microphone or upload a clip; the reference is encoded with the NeMo nano codec and compressed into 8 speaker prefix tokens.
  • Text classifier-free guidance β€” cfg_scale/cfg_frames sharpen text and voice adherence (defaults: temperature=0.3, cfg_scale=3).
  • All generation knobs are exposed under Generation settings.

Architecture

The Space runs on gradio.Server β€” a FastAPI app with Gradio's queueing engine on top. The custom dark-themed UI in index.html is served at / via @app.get("/"); the synthesis pipeline is exposed at @app.api("/synthesize"), so requests flow through Gradio's queue (concurrency control, ZeroGPU allocation, gradio_client compatibility) while the frontend stays a self-contained HTML/CSS/JS bundle.

Configuration

config.yaml selects the checkpoint, codec, preset speakers and default generation parameters β€” re-point the Space without touching code.

Notes

  • The model repo is private: set the HF_TOKEN Space secret.
  • ZeroGPU: the model is loaded once at startup; each request only runs generation inside the GPU context.
  • create_env.py orchestrates a 4-step install at app startup β€” it must run before any ML import in app.py:
    1. Pin huggingface-hub>=1.2,<2.0 (compatible with both gradio 6.20 and nemo-toolkit).
    2. pip install --no-deps nemo-toolkit[tts]==2.4.0 (avoiding a downgrade of hub by the resolver).
    3. Force-reinstall transformers==5.3.0 (Qwen3.5 backbone that the Gepard checkpoint was trained on).
    4. Cap numpy<2.0 so the codec/NeMo stack stays on numpy 1.x.
  • nemo-toolkit is installed at runtime (not in requirements.txt) to keep gradio 6.20's huggingface-hub>=1.2 constraint resolvable at build time β€” NeMo's transformers<=4.52 would otherwise pull hub<1.0.
gradio

Contributors

Simonlob

18 commits

ylankgz

13 commits

CO

nineninesix/gepard

Space

πŸ† GEPARD TTS

40

41 commits

2 linked in READMEs

updated Sep 5, 2026

See the code

README

πŸ† GEPARD TTS

Inference Space for the Gepard autoregressive speech model (nineninesix/gepard-1.0 | nineninesix/gepard-1.1 β€” a Qwen3.5 backbone with 32 FSQ audio heads, a Q-Former voice-cloning compressor, and DPO | GRPO post-training).

Features

  • Preset speakers β€” bundled voices stored as pre-encoded codec tokens (speakers/*.pt), no codec pass needed at request time.
  • Voice cloning β€” record with the microphone or upload a clip; the reference is encoded with the NeMo nano codec and compressed into 8 speaker prefix tokens.
  • Text classifier-free guidance β€” cfg_scale/cfg_frames sharpen text and voice adherence (defaults: temperature=0.3, cfg_scale=3).
  • All generation knobs are exposed under Generation settings.

Architecture

The Space runs on gradio.Server β€” a FastAPI app with Gradio's queueing engine on top. The custom dark-themed UI in index.html is served at / via @app.get("/"); the synthesis pipeline is exposed at @app.api("/synthesize"), so requests flow through Gradio's queue (concurrency control, ZeroGPU allocation, gradio_client compatibility) while the frontend stays a self-contained HTML/CSS/JS bundle.

Configuration

config.yaml selects the checkpoint, codec, preset speakers and default generation parameters β€” re-point the Space without touching code.

Notes

  • The model repo is private: set the HF_TOKEN Space secret.
  • ZeroGPU: the model is loaded once at startup; each request only runs generation inside the GPU context.
  • create_env.py orchestrates a 4-step install at app startup β€” it must run before any ML import in app.py:
    1. Pin huggingface-hub>=1.2,<2.0 (compatible with both gradio 6.20 and nemo-toolkit).
    2. pip install --no-deps nemo-toolkit[tts]==2.4.0 (avoiding a downgrade of hub by the resolver).
    3. Force-reinstall transformers==5.3.0 (Qwen3.5 backbone that the Gepard checkpoint was trained on).
    4. Cap numpy<2.0 so the codec/NeMo stack stays on numpy 1.x.
  • nemo-toolkit is installed at runtime (not in requirements.txt) to keep gradio 6.20's huggingface-hub>=1.2 constraint resolvable at build time β€” NeMo's transformers<=4.52 would otherwise pull hub<1.0.
gradio

Contributors

Simonlob

18 commits

ylankgz

13 commits

CO