cookertron/personaplex-7b-quantized

Run NVIDIA PersonaPlex-7B on consumer GPUs (16GB VRAM) with 4-bit/8-bit quantization fixes and optimized inference.

Python

20

5 commits

updated Jan 19, 2026

See the code

README

PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models

Weights Paper Demo Discord

PersonaPlex is a real-time, full-duplex speech-to-speech conversational model that enables persona control through text-based role prompts and audio-based voice conditioning. Trained on a combination of synthetic and real conversations, it produces natural, low-latency spoken interactions with a consistent persona. PersonaPlex is based on the Moshi architecture and weights.

PersonaPlex Model Architecture
PersonaPlex Architecture

Usage

Installation

  1. Install Dependencies:

    pip install -e .
    

    (This installs the patches and local moshi package)

  2. Authenticate: Log in to your Hugging Face account and accept the PersonaPlex model license.

    export HF_TOKEN=<YOUR_HUGGINGFACE_TOKEN>
    

Launch Server (Optimized for 16GB VRAM)

This release is pre-configured for 4-bit quantization to run efficiently on consumer GPUs (e.g., RTX 4080/4090).

1. Run the helper script:

./run_server.sh

2. Access the Web UI: Open http://localhost:8998 in your browser.

[!IMPORTANT] You MUST use localhost (not an IP address) to ensure the browser grants microphone access in this HTTP mode.

Configuration

The default setup uses 4-bit quantization for maximum stability.

  • 4-bit (Default): Lowest VRAM (~8GB), best stability.
  • 8-bit: Higher precision, slightly more VRAM (~12GB). Edit run_server.sh to change --quantize 4bit to --quantize 8bit if desired.

Offline Evaluation

For offline evaluation use the offline script that streams in an input wav file and produces an output wav file from the captured output stream. The output file will be the same duration as the input file.

Assistant example:

HF_TOKEN=<TOKEN> \
python -m moshi.offline \
  --voice-prompt "NATF2.pt" \
  --input-wav "assets/test/input_assistant.wav" \
  --seed 42424242 \
  --output-wav "output.wav" \
  --output-text "output.json"

Service example:

HF_TOKEN=<TOKEN> \
python -m moshi.offline \
  --voice-prompt "NATM1.pt" \
  --text-prompt "$(cat assets/test/prompt_service.txt)" \
  --input-wav "assets/test/input_service.wav" \
  --seed 42424242 \
  --output-wav "output.wav" \
  --output-text "output.json"

Voices

PersonaPlex supports a wide range of voices; we pre-package embeddings for voices that sound more natural and conversational (NAT) and others that are more varied (VAR). The fixed set of voices are labeled:

Natural(female): NATF0, NATF1, NATF2, NATF3
Natural(male):   NATM0, NATM1, NATM2, NATM3
Variety(female): VARF0, VARF1, VARF2, VARF3, VARF4
Variety(male):   VARM0, VARM1, VARM2, VARM3, VARM4

Prompting Guide

The model is trained on synthetic conversations for a fixed assistant role and varying customer service roles.

Assistant Role

The assistant role has the prompt:

You are a wise and friendly teacher. Answer questions or provide advice in a clear and engaging way.

Use this prompt for the QA assistant focused "User Interruption" evaluation category in FullDuplexBench.

Customer Service Roles

The customer service roles support a variety of prompts. Here are some examples for prompting style reference:

You work for CitySan Services which is a waste management and your name is Ayelen Lucero. Information: Verify customer name Omar Torres. Current schedule: every other week. Upcoming pickup: April 12th. Compost bin service available for $8/month add-on.
You work for Jerusalem Shakshuka which is a restaurant and your name is Owen Foster. Information: There are two shakshuka options: Classic (poached eggs, $9.50) and Spicy (scrambled eggs with jalapenos, $10.25). Sides include warm pita ($2.50) and Israeli salad ($3). No combo offers. Available for drive-through until 9 PM.
You work for AeroRentals Pro which is a drone rental company and your name is Tomaz Novak. Information: AeroRentals Pro has the following availability: PhoenixDrone X ($65/4 hours, $110/8 hours), and the premium SpectraDrone 9 ($95/4 hours, $160/8 hours). Deposit required: $150 for standard models, $300 for premium.

Casual Conversations

The model is also trained on real conversations from the Fisher English Corpus with LLM-labeled prompts for open-ended conversations. Here are some example prompts for casual conversations:

You enjoy having a good conversation.
You enjoy having a good conversation. Have a casual discussion about eating at home versus dining out.
You enjoy having a good conversation. Have an empathetic discussion about the meaning of family amid uncertainty.
You enjoy having a good conversation. Have a reflective conversation about career changes and feeling of home. You have lived in California for 21 years and consider San Francisco your home. You work as a teacher and have traveled a lot. You dislike meetings.
You enjoy having a good conversation. Have a casual conversation about favorite foods and cooking experiences. You are David Green, a former baker now living in Boston. You enjoy cooking diverse international dishes and appreciate many ethnic restaurants.

Use the prompt You enjoy having a good conversation. for the "Pause Handling", "Backchannel" and "Smooth Turn Taking" evaluation categories of FullDuplexBench.

Generalization

Personaplex finetunes Moshi and benefits from the generalization capabilities of the underlying Helium LLM. Thanks to the broad training corpus of the backbone, we find that the model will respond plausibly to out-of-distribution prompts and lead to unexpected or fun conversations. We encourage experimentation with different prompts to test the model's emergent ability to handle scenarios outside its training distribution. As an inspiration we feature the following astronaut prompt in the WebUI:

You enjoy having a good conversation. Have a technical discussion about fixing a reactor core on a spaceship to Mars. You are an astronaut on a Mars mission. Your name is Alex. You are already dealing with a reactor core meltdown on a Mars mission. Several ship systems are failing, and continued instability will lead to catastrophic failure. You explain what is happening and you urgently ask for help thinking through how to stabilize the reactor.

License

The present code is provided under the MIT license. The weights for the models are released under the NVIDIA Open Model license.

Citation

TBD

cookertron/personaplex-7b-quantized

Run NVIDIA PersonaPlex-7B on consumer GPUs (16GB VRAM) with 4-bit/8-bit quantization fixes and optimized inference.

Python

20

5 commits

updated Jan 19, 2026

See the code

README

PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models

Weights Paper Demo Discord

PersonaPlex is a real-time, full-duplex speech-to-speech conversational model that enables persona control through text-based role prompts and audio-based voice conditioning. Trained on a combination of synthetic and real conversations, it produces natural, low-latency spoken interactions with a consistent persona. PersonaPlex is based on the Moshi architecture and weights.

PersonaPlex Model Architecture
PersonaPlex Architecture

Usage

Installation

  1. Install Dependencies:

    pip install -e .
    

    (This installs the patches and local moshi package)

  2. Authenticate: Log in to your Hugging Face account and accept the PersonaPlex model license.

    export HF_TOKEN=<YOUR_HUGGINGFACE_TOKEN>
    

Launch Server (Optimized for 16GB VRAM)

This release is pre-configured for 4-bit quantization to run efficiently on consumer GPUs (e.g., RTX 4080/4090).

1. Run the helper script:

./run_server.sh

2. Access the Web UI: Open http://localhost:8998 in your browser.

[!IMPORTANT] You MUST use localhost (not an IP address) to ensure the browser grants microphone access in this HTTP mode.

Configuration

The default setup uses 4-bit quantization for maximum stability.

  • 4-bit (Default): Lowest VRAM (~8GB), best stability.
  • 8-bit: Higher precision, slightly more VRAM (~12GB). Edit run_server.sh to change --quantize 4bit to --quantize 8bit if desired.

Offline Evaluation

For offline evaluation use the offline script that streams in an input wav file and produces an output wav file from the captured output stream. The output file will be the same duration as the input file.

Assistant example:

HF_TOKEN=<TOKEN> \
python -m moshi.offline \
  --voice-prompt "NATF2.pt" \
  --input-wav "assets/test/input_assistant.wav" \
  --seed 42424242 \
  --output-wav "output.wav" \
  --output-text "output.json"

Service example:

HF_TOKEN=<TOKEN> \
python -m moshi.offline \
  --voice-prompt "NATM1.pt" \
  --text-prompt "$(cat assets/test/prompt_service.txt)" \
  --input-wav "assets/test/input_service.wav" \
  --seed 42424242 \
  --output-wav "output.wav" \
  --output-text "output.json"

Voices

PersonaPlex supports a wide range of voices; we pre-package embeddings for voices that sound more natural and conversational (NAT) and others that are more varied (VAR). The fixed set of voices are labeled:

Natural(female): NATF0, NATF1, NATF2, NATF3
Natural(male):   NATM0, NATM1, NATM2, NATM3
Variety(female): VARF0, VARF1, VARF2, VARF3, VARF4
Variety(male):   VARM0, VARM1, VARM2, VARM3, VARM4

Prompting Guide

The model is trained on synthetic conversations for a fixed assistant role and varying customer service roles.

Assistant Role

The assistant role has the prompt:

You are a wise and friendly teacher. Answer questions or provide advice in a clear and engaging way.

Use this prompt for the QA assistant focused "User Interruption" evaluation category in FullDuplexBench.

Customer Service Roles

The customer service roles support a variety of prompts. Here are some examples for prompting style reference:

You work for CitySan Services which is a waste management and your name is Ayelen Lucero. Information: Verify customer name Omar Torres. Current schedule: every other week. Upcoming pickup: April 12th. Compost bin service available for $8/month add-on.
You work for Jerusalem Shakshuka which is a restaurant and your name is Owen Foster. Information: There are two shakshuka options: Classic (poached eggs, $9.50) and Spicy (scrambled eggs with jalapenos, $10.25). Sides include warm pita ($2.50) and Israeli salad ($3). No combo offers. Available for drive-through until 9 PM.
You work for AeroRentals Pro which is a drone rental company and your name is Tomaz Novak. Information: AeroRentals Pro has the following availability: PhoenixDrone X ($65/4 hours, $110/8 hours), and the premium SpectraDrone 9 ($95/4 hours, $160/8 hours). Deposit required: $150 for standard models, $300 for premium.

Casual Conversations

The model is also trained on real conversations from the Fisher English Corpus with LLM-labeled prompts for open-ended conversations. Here are some example prompts for casual conversations:

You enjoy having a good conversation.
You enjoy having a good conversation. Have a casual discussion about eating at home versus dining out.
You enjoy having a good conversation. Have an empathetic discussion about the meaning of family amid uncertainty.
You enjoy having a good conversation. Have a reflective conversation about career changes and feeling of home. You have lived in California for 21 years and consider San Francisco your home. You work as a teacher and have traveled a lot. You dislike meetings.
You enjoy having a good conversation. Have a casual conversation about favorite foods and cooking experiences. You are David Green, a former baker now living in Boston. You enjoy cooking diverse international dishes and appreciate many ethnic restaurants.

Use the prompt You enjoy having a good conversation. for the "Pause Handling", "Backchannel" and "Smooth Turn Taking" evaluation categories of FullDuplexBench.

Generalization

Personaplex finetunes Moshi and benefits from the generalization capabilities of the underlying Helium LLM. Thanks to the broad training corpus of the backbone, we find that the model will respond plausibly to out-of-distribution prompts and lead to unexpected or fun conversations. We encourage experimentation with different prompts to test the model's emergent ability to handle scenarios outside its training distribution. As an inspiration we feature the following astronaut prompt in the WebUI:

You enjoy having a good conversation. Have a technical discussion about fixing a reactor core on a spaceship to Mars. You are an astronaut on a Mars mission. Your name is Alex. You are already dealing with a reactor core meltdown on a Mars mission. Several ship systems are failing, and continued instability will lead to catastrophic failure. You explain what is happening and you urgently ask for help thinking through how to stabilize the reactor.

License

The present code is provided under the MIT license. The weights for the models are released under the NVIDIA Open Model license.

Citation

TBD

Languages

Python

85.5%

TypeScript

13.9%