laion/voice-tagging-whisper

Model

1

stars

19

commits

1

repos using this model

1

linked in READMEs

Apr 6, 2026

updated

audio
audio-classification
automatic-speech-recognition
endpoints_compatible
paralinguistics
safetensors
speech
transformers
voice
voice-quality
voice-tagging
whisper
Browse cluster: Speech processing and audio analysis

README

Voice-Tagging Whisper

A fine-tuned OpenAI Whisper Small model that generates structured voice attribute tags from speech audio. Instead of transcribing words, this model describes how the voice sounds — its quality, style, loudness, articulation, intonation, and emotional delivery.

Built on top of BUD-E-Whisper (V1.0) and trained for 20 epochs on voice-annotated speech data.

Quick Start

from transformers import WhisperForConditionalGeneration, WhisperProcessor
import torch

# Load model and processor
model = WhisperForConditionalGeneration.from_pretrained(
    "laion/voice-tagging-whisper", torch_dtype=torch.float16
).to("cuda").eval()
processor = WhisperProcessor.from_pretrained("openai/whisper-small")

# Load and process audio (16kHz mono, padded to 30s)
import librosa
waveform, sr = librosa.load("speech.wav", sr=16000, mono=True)

inputs = processor(waveform, sampling_rate=16000, return_tensors="pt")
mel = inputs.input_features.to("cuda", dtype=torch.float16)

with torch.no_grad():
    generated = model.generate(mel, max_new_tokens=256)

tags = processor.batch_decode(generated, skip_special_tokens=True)[0]
print(tags)

Note: This model does not ship its own processor/tokenizer config. Use openai/whisper-small as the processor, which is architecture-compatible.

Output Format

The model outputs a comma-separated sequence of ~8–10 voice attribute tags describing multiple dimensions of the voice simultaneously. There is no free-form caption — the entire output consists of structured tags.

Tag Structure

Based on analysis of 570 diverse audio samples spanning 57 voice taxonomy dimensions, the tags follow a consistent positional order covering these voice dimensions:

PositionDimensionExample ValuesPositional Reliability
1Content safetySuitable for Work, Not Suitable for Work95% consistent
2Naturalnessnatural speaking, naturalness, natural-genuine, slightly unnatural80% consistent
3Fluencyfluent, halting speech, disfluent77% consistent
4Speaking stylecasual speaking style, dramatic style, narrator style, ranting style, ASMR style48% consistent
5Phonation typemodal voice, strained voice, slack voice, rough voice, breathy voice62% consistent
6Airflow / breathinessneutral airflow, pressed voice, breathy, slightly breathy65% consistent
7Loudnessnormal loudness, quiet, loud, very loud, very quiet, whispered47% consistent
8Intonation / prosodyslightly dynamic, dynamic, monotone, falling intonation, irregular intonation34% consistent
9Articulation precisionprecise articulation, neutral articulation, slightly imprecise articulation48% consistent
10Delivery / emotionnatural speaking, crying, screaming, whispering, narration style deliveryFinal position

Typical output length: 9.0 tags on average (range 1–16, most samples produce 8–11).

Tag Vocabulary

Analysis of 570 samples revealed 194 unique tags across the model's vocabulary. Here are the most frequent:

Core Tags (appearing in >10% of samples)

TagFrequencyCategory
Suitable for Work83%Safety
fluent78%Fluency
neutral airflow68%Airflow
modal voice65%Phonation
normal loudness49%Loudness
casual speaking style47%Style
precise articulation47%Articulation
slightly dynamic35%Intonation
natural speaking30%Naturalness / Delivery
neutral articulation20%Articulation
dynamic15%Intonation
very loud12%Loudness
irregular intonation12%Intonation
dramatic style11%Style
pressed voice11%Airflow

Delivery / Emotion Tags (final position)

The last tag typically describes the overall delivery or emotional quality:

TagCountDescription
normal speaking101Neutral, unremarkable delivery
natural speaking71Natural-sounding speech
crying48Emotional, tearful delivery
narration style delivery46Professional narrator tone
screaming31High-energy screaming
high-energy delivery20Energetic, animated speech
strained delivery19Vocally strained
soft speaking14Gentle, quiet delivery
shouting12Loud projected speech
slow deliberate delivery12Measured, intentional pacing
ranting style delivery9Agitated, rant-like speech
out-of-breath delivery9Breathless speech
whispering6Whispered delivery
pleading tone6Pleading, imploring
sad speaking6Sad emotional delivery
laughing while speaking5Speech mixed with laughter
gasping delivery5Gasping or breathless
angry shouting5Angry, aggressive shouting
giggling delivery4Giggly speech
sing-speaking3Semi-melodic delivery

Rare & Specialized Tags

The model also produces less common but descriptive tags:

  • Whisper/ASMR: whisper-talk style, ASMR style, whispery voice, ASMR whisper-delivery
  • Performance: storytelling style, monologue style, formal style, newsreader style, authoritative style
  • Vocal effort: projected voice, tense voice, slightly tense voice
  • Extreme states: out-of-breath delivery, fatigued delivery, gasping delivery, sighing delivery

Examples

Example 1: Calm Narration

Suitable for Work, natural speaking, fluent, narrator style delivery, modal voice,
neutral airflow, normal loudness, monotone, precise articulation, slow deliberate delivery

A clean, professional narrator voice — fluent delivery with precise articulation, balanced loudness, and monotone intonation typical of audiobook or documentary narration.

Example 2: Emotional Crying

Suitable for Work, natural-genuine, halting speech, casual speaking style, rough voice,
breathy, quiet, falling intonation, slightly imprecise articulation, crying

An emotionally charged voice with halting, hesitant speech — quiet and breathy with falling pitch and rough voice quality, characteristic of genuine crying or deep distress.

Example 3: High-Energy Screaming

Suitable for Work, natural pop, fluent, dramatic style, strained voice,
pressed voice, very loud, dynamic, precise articulation, screaming

A loud, high-energy voice with pressed phonation and dramatic style — the kind of intense screaming you might hear in animated entertainment or dramatic performance.

Example 4: ASMR / Whisper

Suitable for Work, natural-Sounding, fluent, ASMR style, breathy voice,
breathy, whispered, monotone, neutral articulation, whispering

A soft, intimate ASMR-style whisper with breathy voice quality, monotone intonation, and very quiet delivery — characteristic of ASMR content or intimate speech.

Example 5: Ranting / Agitated Speech

Suitable for Work, natural-Suitable for Work, fluent, ranting style, strained voice,
pressed voice, very loud, dynamic, precise articulation, screaming

Agitated, high-intensity speech with ranting style, strained phonation, and very loud delivery — characteristic of passionate arguing or emotional outbursts.

Example 6: Casual Conversation

Suitable for Work, natural speaking, fluent, casual speaking style, modal voice,
neutral airflow, normal loudness, slightly dynamic, precise articulation, normal speaking

A typical conversational voice — relaxed, fluent, with balanced loudness and natural modal phonation. This is the most common pattern in everyday speech.

Coverage Analysis (570 Samples)

Detection rates for specific voice characteristics across 570 diverse samples:

Voice CharacteristicSamples DetectedDetection Rate
Screaming / shouting488.4%
Crying488.4%
Ranting539.3%
Disfluent / halting speech5810.2%
Narration style6711.8%
Whisper71.2%
ASMR30.5%
Pleading61.1%
Not Suitable for Work162.8%

These rates reflect the distribution of input samples, not the model's intrinsic sensitivity. The model reliably distinguishes these categories when presented with appropriate audio.

Model Details

PropertyValue
ArchitectureWhisper Small (encoder-decoder)
Parameters~242M
Base modellaion/BUD-E-Whisper
Training20 epochs, 760 steps
Final loss0.018
Input30s mel spectrogram (80 bins, 16kHz audio)
OutputComma-separated voice attribute tags
Max output tokens448
Encoder12 layers, 768 dim, 12 heads
Decoder12 layers, 768 dim, 12 heads
Vocabulary size~194 unique tags observed
Avg output tags9.0 per sample

Architecture Notes

This is a full encoder-decoder Whisper model. The encoder maps audio to rich voice representations (also usable standalone for embedding extraction — see Voice-Taxonomy-57), while the decoder generates the tag sequence auto-regressively.

The model uses the standard Whisper tokenizer and generation config. Audio is processed as 30-second chunks of 80-bin log-mel spectrograms at 16kHz.

Encoder-Only Usage

The encoder of this model also produces high-quality voice embeddings for downstream tasks. It is used as one of the whisper encoders in the Voice-Taxonomy-57 pipeline for 57-dimension voice classification:

from transformers import WhisperModel, WhisperFeatureExtractor

model = WhisperModel.from_pretrained("laion/voice-tagging-whisper", torch_dtype=torch.float16)
encoder = model.encoder.to("cuda").eval()
fe = WhisperFeatureExtractor.from_pretrained("openai/whisper-small")

inputs = fe(waveform, sampling_rate=16000, return_tensors="pt")
with torch.no_grad():
    hidden_states = encoder(inputs.input_features.cuda().half()).last_hidden_state
    # hidden_states shape: (batch, 1500, 768)

License

Apache 2.0

Contributors

laion/voice-tagging-whisper

Model

1

stars

19

commits

1

repos using this model

1

linked in READMEs

Apr 6, 2026

updated

audio
audio-classification
automatic-speech-recognition
endpoints_compatible
paralinguistics
safetensors
speech
transformers
voice
voice-quality
voice-tagging
whisper
Browse cluster: Speech processing and audio analysis

README

Voice-Tagging Whisper

A fine-tuned OpenAI Whisper Small model that generates structured voice attribute tags from speech audio. Instead of transcribing words, this model describes how the voice sounds — its quality, style, loudness, articulation, intonation, and emotional delivery.

Built on top of BUD-E-Whisper (V1.0) and trained for 20 epochs on voice-annotated speech data.

Quick Start

from transformers import WhisperForConditionalGeneration, WhisperProcessor
import torch

# Load model and processor
model = WhisperForConditionalGeneration.from_pretrained(
    "laion/voice-tagging-whisper", torch_dtype=torch.float16
).to("cuda").eval()
processor = WhisperProcessor.from_pretrained("openai/whisper-small")

# Load and process audio (16kHz mono, padded to 30s)
import librosa
waveform, sr = librosa.load("speech.wav", sr=16000, mono=True)

inputs = processor(waveform, sampling_rate=16000, return_tensors="pt")
mel = inputs.input_features.to("cuda", dtype=torch.float16)

with torch.no_grad():
    generated = model.generate(mel, max_new_tokens=256)

tags = processor.batch_decode(generated, skip_special_tokens=True)[0]
print(tags)

Note: This model does not ship its own processor/tokenizer config. Use openai/whisper-small as the processor, which is architecture-compatible.

Output Format

The model outputs a comma-separated sequence of ~8–10 voice attribute tags describing multiple dimensions of the voice simultaneously. There is no free-form caption — the entire output consists of structured tags.

Tag Structure

Based on analysis of 570 diverse audio samples spanning 57 voice taxonomy dimensions, the tags follow a consistent positional order covering these voice dimensions:

PositionDimensionExample ValuesPositional Reliability
1Content safetySuitable for Work, Not Suitable for Work95% consistent
2Naturalnessnatural speaking, naturalness, natural-genuine, slightly unnatural80% consistent
3Fluencyfluent, halting speech, disfluent77% consistent
4Speaking stylecasual speaking style, dramatic style, narrator style, ranting style, ASMR style48% consistent
5Phonation typemodal voice, strained voice, slack voice, rough voice, breathy voice62% consistent
6Airflow / breathinessneutral airflow, pressed voice, breathy, slightly breathy65% consistent
7Loudnessnormal loudness, quiet, loud, very loud, very quiet, whispered47% consistent
8Intonation / prosodyslightly dynamic, dynamic, monotone, falling intonation, irregular intonation34% consistent
9Articulation precisionprecise articulation, neutral articulation, slightly imprecise articulation48% consistent
10Delivery / emotionnatural speaking, crying, screaming, whispering, narration style deliveryFinal position

Typical output length: 9.0 tags on average (range 1–16, most samples produce 8–11).

Tag Vocabulary

Analysis of 570 samples revealed 194 unique tags across the model's vocabulary. Here are the most frequent:

Core Tags (appearing in >10% of samples)

TagFrequencyCategory
Suitable for Work83%Safety
fluent78%Fluency
neutral airflow68%Airflow
modal voice65%Phonation
normal loudness49%Loudness
casual speaking style47%Style
precise articulation47%Articulation
slightly dynamic35%Intonation
natural speaking30%Naturalness / Delivery
neutral articulation20%Articulation
dynamic15%Intonation
very loud12%Loudness
irregular intonation12%Intonation
dramatic style11%Style
pressed voice11%Airflow

Delivery / Emotion Tags (final position)

The last tag typically describes the overall delivery or emotional quality:

TagCountDescription
normal speaking101Neutral, unremarkable delivery
natural speaking71Natural-sounding speech
crying48Emotional, tearful delivery
narration style delivery46Professional narrator tone
screaming31High-energy screaming
high-energy delivery20Energetic, animated speech
strained delivery19Vocally strained
soft speaking14Gentle, quiet delivery
shouting12Loud projected speech
slow deliberate delivery12Measured, intentional pacing
ranting style delivery9Agitated, rant-like speech
out-of-breath delivery9Breathless speech
whispering6Whispered delivery
pleading tone6Pleading, imploring
sad speaking6Sad emotional delivery
laughing while speaking5Speech mixed with laughter
gasping delivery5Gasping or breathless
angry shouting5Angry, aggressive shouting
giggling delivery4Giggly speech
sing-speaking3Semi-melodic delivery

Rare & Specialized Tags

The model also produces less common but descriptive tags:

  • Whisper/ASMR: whisper-talk style, ASMR style, whispery voice, ASMR whisper-delivery
  • Performance: storytelling style, monologue style, formal style, newsreader style, authoritative style
  • Vocal effort: projected voice, tense voice, slightly tense voice
  • Extreme states: out-of-breath delivery, fatigued delivery, gasping delivery, sighing delivery

Examples

Example 1: Calm Narration

Suitable for Work, natural speaking, fluent, narrator style delivery, modal voice,
neutral airflow, normal loudness, monotone, precise articulation, slow deliberate delivery

A clean, professional narrator voice — fluent delivery with precise articulation, balanced loudness, and monotone intonation typical of audiobook or documentary narration.

Example 2: Emotional Crying

Suitable for Work, natural-genuine, halting speech, casual speaking style, rough voice,
breathy, quiet, falling intonation, slightly imprecise articulation, crying

An emotionally charged voice with halting, hesitant speech — quiet and breathy with falling pitch and rough voice quality, characteristic of genuine crying or deep distress.

Example 3: High-Energy Screaming

Suitable for Work, natural pop, fluent, dramatic style, strained voice,
pressed voice, very loud, dynamic, precise articulation, screaming

A loud, high-energy voice with pressed phonation and dramatic style — the kind of intense screaming you might hear in animated entertainment or dramatic performance.

Example 4: ASMR / Whisper

Suitable for Work, natural-Sounding, fluent, ASMR style, breathy voice,
breathy, whispered, monotone, neutral articulation, whispering

A soft, intimate ASMR-style whisper with breathy voice quality, monotone intonation, and very quiet delivery — characteristic of ASMR content or intimate speech.

Example 5: Ranting / Agitated Speech

Suitable for Work, natural-Suitable for Work, fluent, ranting style, strained voice,
pressed voice, very loud, dynamic, precise articulation, screaming

Agitated, high-intensity speech with ranting style, strained phonation, and very loud delivery — characteristic of passionate arguing or emotional outbursts.

Example 6: Casual Conversation

Suitable for Work, natural speaking, fluent, casual speaking style, modal voice,
neutral airflow, normal loudness, slightly dynamic, precise articulation, normal speaking

A typical conversational voice — relaxed, fluent, with balanced loudness and natural modal phonation. This is the most common pattern in everyday speech.

Coverage Analysis (570 Samples)

Detection rates for specific voice characteristics across 570 diverse samples:

Voice CharacteristicSamples DetectedDetection Rate
Screaming / shouting488.4%
Crying488.4%
Ranting539.3%
Disfluent / halting speech5810.2%
Narration style6711.8%
Whisper71.2%
ASMR30.5%
Pleading61.1%
Not Suitable for Work162.8%

These rates reflect the distribution of input samples, not the model's intrinsic sensitivity. The model reliably distinguishes these categories when presented with appropriate audio.

Model Details

PropertyValue
ArchitectureWhisper Small (encoder-decoder)
Parameters~242M
Base modellaion/BUD-E-Whisper
Training20 epochs, 760 steps
Final loss0.018
Input30s mel spectrogram (80 bins, 16kHz audio)
OutputComma-separated voice attribute tags
Max output tokens448
Encoder12 layers, 768 dim, 12 heads
Decoder12 layers, 768 dim, 12 heads
Vocabulary size~194 unique tags observed
Avg output tags9.0 per sample

Architecture Notes

This is a full encoder-decoder Whisper model. The encoder maps audio to rich voice representations (also usable standalone for embedding extraction — see Voice-Taxonomy-57), while the decoder generates the tag sequence auto-regressively.

The model uses the standard Whisper tokenizer and generation config. Audio is processed as 30-second chunks of 80-bin log-mel spectrograms at 16kHz.

Encoder-Only Usage

The encoder of this model also produces high-quality voice embeddings for downstream tasks. It is used as one of the whisper encoders in the Voice-Taxonomy-57 pipeline for 57-dimension voice classification:

from transformers import WhisperModel, WhisperFeatureExtractor

model = WhisperModel.from_pretrained("laion/voice-tagging-whisper", torch_dtype=torch.float16)
encoder = model.encoder.to("cuda").eval()
fe = WhisperFeatureExtractor.from_pretrained("openai/whisper-small")

inputs = fe(waveform, sampling_rate=16000, return_tensors="pt")
with torch.no_grad():
    hidden_states = encoder(inputs.input_features.cuda().half()).last_hidden_state
    # hidden_states shape: (batch, 1500, 768)

License

Apache 2.0

Contributors