Foundation-1 is a family of text-conditioned audio models designed around music production workflows. Rather than treating audio generation as broad caption-to-music generation, Foundation-1 was trained around structured controls for instrument identity, timbre, FX, musical behavior, pitch, timing, and tonality.
The original Foundation-1 checkpoint focuses on structured loop generation. Newer specialized checkpoints extend the same conditioning system into individual one-shots and pitch-consistent playable keybeds.
This means Foundation-1 can now be used at several levels of a production workflow:
The goal is not simply to generate finished audio. The specialized keybed workflow is designed to let the model generate the source material for an instrument, then hand control back to the producer.
| Model | Primary Use | Notes |
|---|---|---|
| Foundation-1 | Loop generation | Original checkpoint for structured, BPM-aware, bar-aware musical loops |
| Foundation-1.2 Samples | Loop Generation / One-Shot generation | Selected earlier in specialized training to preserve stronger short-form loop-generation quality |
| Foundation-1.2 Keybeds | One-Shot generation, Pitch-consistent sample generation and playable instruments | Trained longer to reinforce timbral consistency across pitch and register |
The Foundation-1.2 Keybeds model is not limited to full keybed generation. It remains highly capable at generating individual samples and one-shots. Its specialization reflects the additional training needed to maintain a consistent sonic identity across many generated pitches.
The checkpoints were separated because the longer keybed-focused training improved cross-pitch consistency, while reducing the broader loop-generation quality preserved by the earlier Samples checkpoint.
The Keybed checkpoint extends Foundation-1 from text-to-sample generation into a practical text-to-instrument workflow.
A user can describe a sound using the same instrument and timbre vocabulary used elsewhere in Foundation-1, generate pitch-specific samples across a keyboard, and automatically assemble the result into a playable sampler instrument.
Examples can range from conventional instruments:
Grand Piano, Warm, Gritty, Wet, Low Reverb
to synthetic or hybrid sounds:
Cello, Harp, Rich, Clean, Choir, Pluck, Formant Vocal, Wet, High Reverb
Because pitch is generated rather than conventionally pitch-shifted from a single root sample, the model can create natural timbral variation across registers while maintaining the identity of the prompted sound.
For the intended workflow, the Keybed model is best used with RC Stable Audio Tools, which handles the multi-note generation process and instrument assembly automatically. Generated keybeds can be exported for use as playable sampler instruments rather than remaining as isolated audio generations.
The user-facing prompt is intentionally simple. A descriptor such as:
Grand Piano, Warm, Gritty
is expanded internally into the structured Keybed conditioning grammar used by the model. RC Stable Audio Tools adds the required Keybed / sequence / note information, keeps the user descriptor stable across the instrument, reuses the same resolved seed across generation chunks, then slices and maps the resulting notes into the exported sampler.
This makes the workflow useful not only as a demo interface, but also as a reference implementation for developers interested in building their own VST, sampler, or instrument-generation front end around the model.
For the full prompt-injection and inference design, see the Keybed Training & Inference Strategy.
Foundation-1.2 Keybed generation workflow in RC Stable Audio Tools.
Most audio models can react to broad prompt terms like “warm pad” or “bright synth.” with inconsistent results. Foundation-1 was designed to go further by treating the sound as a layered system:
This layered conditioning approach is a major reason Foundation-1 is able to deliver both high musicality and high prompt control at the same time.
| Prompt | Audio |
|---|---|
| Bass, FM Bass, Medium Delay, Medium Reverb, Low Distortion, Phaser, Sub Bass, Bass, Upper Mids, Acid, Gritty, Wide, Dubstep, Thick, Silky, Warm, Rich, Overdriven, Crisp, Deep, Clean, Pitch Bend, 303, 8 Bars, 140 BPM, E minor |
|
| Sub Bass, Bass, Gritty, Small, Square, Bass, Dark, Digital, Thick, Clean, Simple, Bassline, Epic, Choppy, Melody, 4 Bars, 150 BPM, G# minor |
|
| Flute, Pizzicato, Punchy, Present, Ambient, Nasal, Melody, Epic, Airy, Slow Speed, 8 Bars, 150 BPM, E minor |
|
| High Saw, Spacey, Lead, Warm, Silky, Smooth, 303, Synth Lead, Medium Reverb, Low Distortion, Upper Mids, Mids, Pitch Bend, Arp, 8 Bars, 140 BPM, F minor |
|
| Trumpet, Warm, Complex Arp Melody, High Reverb, Low Distortion, Smooth, Silky, Texture, 8 Bars, 130 BPM, C minor |
|
| Synth, Pad, Chord Progression, Rising, Digital, Bass, Fat, Near, Wide, Silky, Warm, Focused, 8 Bars, 110 BPM, D major |
|
| Piccolo, Flute, Airy, Music Box, plucked, complex melody, 8 Bars, 140 BPM, C# minor |
|
| Synth Lead, Wavetable Bass, Low Distortion, High Reverb, Sub Bass, Upper Mids, Acid, Gritty, Wide, Thick, Silky, Warm, Rich, Overdriven, Crisp, Clean, 303, Complex, 8 Bars, 140 BPM, F minor |
|
| Fiddle, Bowed Strings, Full, Clean, Spacey, Rich, Intimate, Thick, Rolling, Arp, Fast Speed, Complex, 8 Bars, 128 BPM, B minor |
|
| Chiptune, Chord Progression, Pulse Wave, Medium Reverb, 8 Bars, 128 BPM, D minor |
|
| Kalimba, Mallet, Medium Reverb, Overdriven, Wide, Metallic, Thick, Sparkly, Upper Mids, Bright, Airy, Alternating, Chord Progression, Atmosphere, Spacey, Fast Speed, 8 Bars, 120 BPM, B minor |
|
The examples below show short, note-specific one-shot generations across acoustic, synthetic, bass, FX, and hybrid timbres.
| Prompt | Note | Audio |
|---|---|---|
| FX, Sharp, Subdued, Distant, Sparkly, Round, Hit, Deep, Low Reverb, Sub | D#1 |
|
| Synth Bass, Metallic, Rich, Punchy, Digital, Medium Reverb | C2 |
|
| Reese Bass, Hit, Bright, Wavetable, Noisy, Big, Thick, Airy, Deep | F#2 |
|
| Synth Lead, Shiny, Rich, Choir, Bright, Smooth, Short, Synthetic Vox | F5 |
|
| Tuba, Big, Bright, Clean, Biting, Sustained | G#3 |
|
| Clavinet, Wobble, Breathy, Bright, Deep, Gritty, Impact, Sparkly | C#4 |
|
| Cello, Airy, Wide, Woody, Muffled, Smooth | D#3 |
|
| Reese Bass, Subdued, Round, Growl, Hit, Deep, Woody, Fat, Metallic, Medium Delay, Medium Distortion | D#2 |
|
| Digital Piano, Retro, Smooth, Mono, Warm, Low Reverb | E3 |
|
| Supersaw, Big, Smooth, Warm, Vintage, Crisp, Analog, Wide, Muffled, High Reverb | G2 |
|
| Synth Lead, Bright, Shiny, Soft, Warm, Punchy, Smooth, Medium Delay | C#4 |
|
| Sustained, Trumpet, Airy, Wide, High Reverb | F3 |
|
| 808, Thick, Acid, Hit, FX, Rumble, Wide, Big, Distortion | F1 |
|
| Violin, Vintage, Thick, Chiptune, Spacey, Nasal, Medium Delay | A#5 |
|
| Grand Piano, Near, Snappy, Distant, Bell, Bright, Focused, Delay | D3 |
|
The following examples demonstrate generated keybeds across acoustic, synthetic, and hybrid timbres.
A full video playthrough of the Keybed Showcase is available here, which shows the associated MIDI used to audition each generated instrument.
| Prompt | Audio |
|---|---|
| Soft, Silky, Spectral, Smooth, Spacey, Synth, Pad, Subdued, Wet, High Reverb |
|
| Pad, Focused, Buzzy, Big, Deep, Supersaw, Pulse, Dry |
|
| Pan Flute, Hollow, Dark, Smooth, Sustained, Wide, Silky, Wet, Medium Reverb |
|
| Violin, Woody, Pizzicato, Focused, Breathy, Airy, Wet, High Reverb |
|
| Grand Piano, Warm, Gritty, Wet, Low Reverb |
|
| Grand Piano, Cold, Sparkly, Wet, Low Reverb |
|
| Digital Piano, Noisy, Spiccato, Rich, Pluck, Warm, Swell, Clean, Dry |
|
| Cello, Fat, Warm, Metallic, Sharp, Near, Sustained, Dry |
|
| Church Bell, Glassy, Sparkly, Wet, High Reverb |
|
| Thick, Saw, Neuro, Reese Bass, Dry |
|
| Thick, Sine, Reese Bass, Wet, High Reverb, High Phaser |
|
| Xylophone, Sustained, Analog, Warm, Woody, Dry |
|
| Xylophone, Sustained, Analog, Warm, Woody, Bit Crushed, Wet |
|
| Cello, Harp, Rich, Clean, Choir, Pluck, Formant Vocal, Wet, High Reverb |
|
| Acid, Neuro, Synth Lead, Saw, Square, Wide, Focused, Pluck, Sustained, Wet, High Phaser |
|
| Synth Lead, Present, Sharp, Spacey, Digital, Hollow, Focused, Clean, Wet, Cross Delay, High Reverb |
|
| Dubstep, Neuro, Synth Lead, Pitch Bend, Acid, Wet, Cross Delay |
|
Pure waveform keybeds are shown separately to make the learned oscillator shapes easier to compare directly.
| Waveform | Prompt | Audio |
|---|---|---|
| Triangle | Pure Tone, Triangle, Dry |
|
| Square | Pure Tone, Square, Dry |
|
| Pulse | Pure Tone, Pulse, Dry |
|
| Sine | Pure Tone, Sine, Dry |
|
| Saw | Pure Tone, Saw, Dry |
|
The inference pipeline can also combine three independently generated keybeds into a single layered instrument. The examples below use separate Main and Support prompts for each layer. Further the generated instrument allows independent volume control for each layer.
| Layer Prompts | Audio |
|---|---|
| Main: Bright, Airy, Violin Support 1: Thick, Present, Male, Vocal, Choir Support 2: Digital String, Analog |
|
| Main: Marimba, Dark, Sub, Woody, Airy Support 1: Soft, Grand Piano, Dark Support 2: Music Box, Glassy |
|
Foundation-1 was trained to produce structured musical material rather than full music or generic textures. Musical Notation terms can encourage notation, chord progressions, melodies, arps, phrase direction, rhythmic density, and other musically relevant behaviors.
The model supports a broad instrument hierarchy spanning synths, keys, basses, bowed strings, mallets, winds, guitars, brass, vocals, and plucked strings.
Foundation-1 is not limited to broad instrument naming. It also responds to timbral descriptors such as spectral shape, tone, width, density, texture, brightness, warmth, grit, space, and other sonic traits.
Because instrument identity and timbral character were not collapsed into a single flat label, the model is especially strong at timbral hybridization and layered sonic prompting.
The model supports a dedicated FX layer covering multiple forms of reverb, delay, distortion, phaser, and bitcrushing.
Foundation-1 is built for production-ready loop generation, including BPM-aware and bar-aware structure within supported denominations.
The specialized Foundation-1.2 Samples and Keybeds checkpoints generate short, note-specific sounds across acoustic, synthetic, bass, FX, and hybrid timbres. These can be used individually or as raw material for sampler instruments.
The Keybed checkpoint was trained specifically to preserve a prompted timbral identity across changing pitch and register. RC Stable Audio Tools uses this capability to generate multiple note-specific samples and assemble them into playable instruments.
Foundation-1 was trained with a layered tagging hierarchy designed to improve control, composability, and prompt clarity.
This makes it possible to prompt at different levels of abstraction. A user can stay broad with a family-level prompt like Synth or Keys, or get more specific with terms like Synth Lead, Wavetable Bass, Grand Piano, Violin, or Trumpet, then further shape the output using timbral and FX descriptors.
Foundation-1 was trained across the following major instrument families:
Foundation-1 includes a wide sub-family layer covering a broad range of production-relevant instrument roles, including but not limited to:
One of Foundation-1’s main strengths is that it was not trained to treat timbre as an afterthought. Timbral character is directly represented in the prompt system, giving users control over not only what is being generated, but also how it sounds.
Representative timbre descriptors include:
This tagging design makes prompts much more flexible. Instead of only asking for an instrument, users can shape:
This is especially useful for producers who want to guide the output toward a specific role in a mix rather than just a generic instrument label.
For a list of used tags please see the Tag Reference Sheet.
Foundation-1 includes a dedicated FX descriptor layer spanning multiple common production effects.
Representative FX tags include:
Foundation-1 was trained with structured musical descriptors designed to improve phrase coherence, rhythmic intent, melodic motion, and prompt control.
These notation-style prompt terms help steer:
Examples of supported structural ideas may include terms such as:
This notation layer is one of the main reasons Foundation-1 produces unusually coherent musical material instead of static or loosely related phrases. These can be mixed and matched as desired.
Foundation-1 is designed for structured music production workflows and supports:
For best results, use rich prompts built around the model’s tags. These tags can be mixed and matched as needed. The model was trained on a structured hierarchy designed to encourage musically coherent sample generation.
[Instrument Family / Sub-Family], [Timbre], [Musical Behavior / Notation], [FX], [Key], [Bars], [BPM]
Use FX and timbre tags sparingly at first, then layer more once you understand the model’s behavior.
Each row below uses the exact same prompt, but a different random seed.
The timbre tags remain unchanged, so the overall sound character stays consistent while the melodic and musical content varies between generations.
| Prompt | Output A | Output B | Output C |
|---|---|---|---|
| Bass, FM Bass, Medium Delay, Medium Reverb, Low Distortion, Phaser, Acid, Gritty, Wide, Dubstep, Thick, Silky, Warm, Rich, Overdriven, Crisp, Deep, Clean, Triplets, 8 Bars, 150 BPM, A minor |
|
|
|
| Gritty, Acid, Bassline, 303, Synth Lead, FM, Sub, Upper Mids, High Phaser, High Reverb, Pitch Bend, 8 Bars, 140 BPM, E minor |
|
|
|
| Kalimba, Mallet, Medium Reverb, Overdriven, Wide, Metallic, Thick, Sparkly, Upper Mids, Bright, Airy, Small, Alternating Chord Progression, Atmosphere, Spacey, Fast, 4 Bars, 120 BPM, B minor |
|
|
|
Foundation-1 is best used with RC Stable Audio Tools, which is tuned around the model family, its metadata, and its structured prompting system.
RC Stable Audio Tools (Enhanced Fork)
The interface supports different workflows depending on the checkpoint in use.
For the original Foundation-1 loop checkpoint, the interface provides:
The specialized Samples workflow supports:
The Keybed checkpoint can also generate individual one-shots effectively. The dedicated Samples checkpoint is provided because its earlier training endpoint preserves stronger general sample-generation quality.
The Keybed workflow is designed to be used in conjunction with RC Stable Audio Tools rather than as a sequence of manually generated independent notes.
The interface handles the pitch-aware generation pipeline, creates the source files required for the instrument, and can export completed keybeds for sampler use.
The workflow supports:
Generated keybeds can be assembled and exported as playable DecentSampler or SFZ instruments.
RC Stable Audio Tools (Enhanced Fork)
Stable Audio Tools (Original Repository)
Foundation-1 is now distributed as multiple specialized checkpoints built around the same underlying architecture and conditioning system.
The specialized Foundation-1.2 Samples and Keybeds checkpoints retain the same Foundation-1 / SAO architecture. They therefore do not represent separate model architectures; the specialization comes from the training objective and selected checkpoint.
The release uses 16-bit model weights to reduce the model footprint without changing the intended inference quality.
Foundation_1.safetensors — loop-generation checkpointFoundation-1.2-Samples.safetensors — sample-focused checkpointFoundation-1.2-Keybeds.safetensors — keybed-focused checkpointmodel_config.json — shared model configurationmodels directoryFor full keybeds, the RC interface is strongly recommended because it manages the multi-note inference and sampler-export process automatically.
Foundation-1 is designed to run locally on modern GPUs.
Typical VRAM usage during generation is approximately ~7 GB.
For reliable operation, a GPU with at least 8 GB of VRAM is recommended.
Generation speed will vary depending on GPU model, generation mode, sample length, and system configuration.
On an RTX 3090, a standard individual generation is approximately ~7–8 seconds per sample. Full single-layer keybed builds take approximately 50 seconds.
Foundation-1 was built around a structured sample-generation philosophy, rather than generic or genre-based audio captioning. The dataset consists entirely of hand-crafted and labeled audio, produced through a controlled augmentation pipeline.
At a high level, the training design emphasizes:
This design is central to the model’s musical coherence and high degree of sonic control.
For more details on the original dataset and training methodology, see the Training & Dataset Notes.
The Foundation-1.2 Samples and Keybeds checkpoints use the same underlying Foundation-1 architecture, but were trained with a lower learning rate of 1e-5.
| Checkpoint | Training Endpoint | Specialization |
|---|---|---|
| Foundation-1.2 Samples | Epoch 0 / Step 560 | General sample / one-shot generation |
| Foundation-1.2 Keybeds | Epoch 3 / Step 8120 | Cross-pitch timbral consistency and playable keybed generation |
Additional training configuration:
1e-51e-3The longer Keybed run was selected to reinforce cross-note and cross-register consistency. The objective was not simply to improve isolated note quality, but to teach the model how a single sound identity should behave as pitch changes across an instrument.
The Keybed checkpoint used a dedicated multi-note training strategy rather than treating every pitch as an unrelated example. The generation pipeline mirrors that structure by injecting note-sequence conditioning, reusing a common seed across keybed chunks, slicing generated sequences back into individual notes, and packaging the resulting samples into playable instruments.
For a detailed explanation of both the training method and the inference/export pipeline, see the Keybed Training & Inference Strategy.
Foundation-1 is a specialized model family for producer-facing sample and instrument generation, not a general-purpose full-song generator.
Important notes:
The model is also optimized around specific timing relationships between Bars, BPM, and generation duration.
For example:
If the generation duration is shorter than the musical structure implied by the prompt (for example requesting an 8-bar loop but generating only 5 seconds), the model may produce less coherent musical phrases.
The RC Stable Audio Fork automatically handles this timing alignment, making this workflow much easier.
The Keybed model is designed around practical instrument ranges.
Some prompts simply stop making semantic or perceptual sense at extreme registers. For example, a sub-bass at C7 is no longer functioning as a bass, even if the model preserves some aspects of the original waveform or timbral character.
The same applies to many synthetic and acoustic sounds. As pitch rises into extreme upper registers, complex waveforms often become perceptually simpler and can collapse toward high-pitched pure-tone-like behavior. At the opposite end, very low pitches can become increasingly difficult to render consistently, and small pitch or waveform errors become more obvious.
For this reason, RC Stable Audio Tools clamps generation and export ranges to practical regions rather than forcing every generated sound across the full MIDI note range.
Prompt choice also matters. A timbre that has a natural low, mid, or high-register identity may become ambiguous when pushed several octaves outside that range. Some apparent "timbre drift" is therefore not only a model limitation; it is also a consequence of asking for a sound whose defining characteristics change as pitch moves far beyond where that sound is normally perceived.
In my internal testing, roughly 90% of generated instruments maintain a consistent relationship between pitch and timbral identity across the supported keybed ranges. This is a qualitative estimate rather than a programmatic benchmark, because there is no reliable automated metric for determining whether two notes share the same perceived timbre.
The remaining edge cases are most likely to appear with:
The default export ranges were chosen as a practical compromise between keyboard coverage, pitch accuracy, and timbral consistency.
This model is licensed under the Stability AI Community License. It is available for non-commercial use or limited commercial use by entities with annual revenues below USD $1M. For revenues exceeding USD $1M, please refer to the repository license file for full terms.
The original Foundation-1 video covering the model's design philosophy, structured prompting system, timbral control, and loop-generation workflow.
🎥 Watch the Foundation-1 Overview
A companion video covering the new Samples and Keybeds checkpoints, the updated training strategy and general journey to getting this made can be found here.
🎥 Watch the Foundation-1.2 Update
A hands-on walkthrough showing the Keybed model in use, cross-pitch timbral behavior, and the new three-layer instrument exporter.
🎥 Watch the Guided Keybed Demo
Foundation-1 is intended as a producer-facing model family for structured sample and instrument generation, designed to augment music production.
Its goal is to let users explore sound in new ways while retaining precise control over:
The original loop model focuses on structured musical material. The specialized Foundation-1.2 Samples and Keybeds checkpoints extend that same conditioning system toward sound design and playable instrument creation.
That combination of musical structure, instrument identity, timbral control, loop fidelity, and cross-pitch instrument generation is what defines the Foundation-1 family.
48 commits
Foundation-1 is a family of text-conditioned audio models designed around music production workflows. Rather than treating audio generation as broad caption-to-music generation, Foundation-1 was trained around structured controls for instrument identity, timbre, FX, musical behavior, pitch, timing, and tonality.
The original Foundation-1 checkpoint focuses on structured loop generation. Newer specialized checkpoints extend the same conditioning system into individual one-shots and pitch-consistent playable keybeds.
This means Foundation-1 can now be used at several levels of a production workflow:
The goal is not simply to generate finished audio. The specialized keybed workflow is designed to let the model generate the source material for an instrument, then hand control back to the producer.
| Model | Primary Use | Notes |
|---|---|---|
| Foundation-1 | Loop generation | Original checkpoint for structured, BPM-aware, bar-aware musical loops |
| Foundation-1.2 Samples | Loop Generation / One-Shot generation | Selected earlier in specialized training to preserve stronger short-form loop-generation quality |
| Foundation-1.2 Keybeds | One-Shot generation, Pitch-consistent sample generation and playable instruments | Trained longer to reinforce timbral consistency across pitch and register |
The Foundation-1.2 Keybeds model is not limited to full keybed generation. It remains highly capable at generating individual samples and one-shots. Its specialization reflects the additional training needed to maintain a consistent sonic identity across many generated pitches.
The checkpoints were separated because the longer keybed-focused training improved cross-pitch consistency, while reducing the broader loop-generation quality preserved by the earlier Samples checkpoint.
The Keybed checkpoint extends Foundation-1 from text-to-sample generation into a practical text-to-instrument workflow.
A user can describe a sound using the same instrument and timbre vocabulary used elsewhere in Foundation-1, generate pitch-specific samples across a keyboard, and automatically assemble the result into a playable sampler instrument.
Examples can range from conventional instruments:
Grand Piano, Warm, Gritty, Wet, Low Reverb
to synthetic or hybrid sounds:
Cello, Harp, Rich, Clean, Choir, Pluck, Formant Vocal, Wet, High Reverb
Because pitch is generated rather than conventionally pitch-shifted from a single root sample, the model can create natural timbral variation across registers while maintaining the identity of the prompted sound.
For the intended workflow, the Keybed model is best used with RC Stable Audio Tools, which handles the multi-note generation process and instrument assembly automatically. Generated keybeds can be exported for use as playable sampler instruments rather than remaining as isolated audio generations.
The user-facing prompt is intentionally simple. A descriptor such as:
Grand Piano, Warm, Gritty
is expanded internally into the structured Keybed conditioning grammar used by the model. RC Stable Audio Tools adds the required Keybed / sequence / note information, keeps the user descriptor stable across the instrument, reuses the same resolved seed across generation chunks, then slices and maps the resulting notes into the exported sampler.
This makes the workflow useful not only as a demo interface, but also as a reference implementation for developers interested in building their own VST, sampler, or instrument-generation front end around the model.
For the full prompt-injection and inference design, see the Keybed Training & Inference Strategy.
Foundation-1.2 Keybed generation workflow in RC Stable Audio Tools.
Most audio models can react to broad prompt terms like “warm pad” or “bright synth.” with inconsistent results. Foundation-1 was designed to go further by treating the sound as a layered system:
This layered conditioning approach is a major reason Foundation-1 is able to deliver both high musicality and high prompt control at the same time.
| Prompt | Audio |
|---|---|
| Bass, FM Bass, Medium Delay, Medium Reverb, Low Distortion, Phaser, Sub Bass, Bass, Upper Mids, Acid, Gritty, Wide, Dubstep, Thick, Silky, Warm, Rich, Overdriven, Crisp, Deep, Clean, Pitch Bend, 303, 8 Bars, 140 BPM, E minor |
|
| Sub Bass, Bass, Gritty, Small, Square, Bass, Dark, Digital, Thick, Clean, Simple, Bassline, Epic, Choppy, Melody, 4 Bars, 150 BPM, G# minor |
|
| Flute, Pizzicato, Punchy, Present, Ambient, Nasal, Melody, Epic, Airy, Slow Speed, 8 Bars, 150 BPM, E minor |
|
| High Saw, Spacey, Lead, Warm, Silky, Smooth, 303, Synth Lead, Medium Reverb, Low Distortion, Upper Mids, Mids, Pitch Bend, Arp, 8 Bars, 140 BPM, F minor |
|
| Trumpet, Warm, Complex Arp Melody, High Reverb, Low Distortion, Smooth, Silky, Texture, 8 Bars, 130 BPM, C minor |
|
| Synth, Pad, Chord Progression, Rising, Digital, Bass, Fat, Near, Wide, Silky, Warm, Focused, 8 Bars, 110 BPM, D major |
|
| Piccolo, Flute, Airy, Music Box, plucked, complex melody, 8 Bars, 140 BPM, C# minor |
|
| Synth Lead, Wavetable Bass, Low Distortion, High Reverb, Sub Bass, Upper Mids, Acid, Gritty, Wide, Thick, Silky, Warm, Rich, Overdriven, Crisp, Clean, 303, Complex, 8 Bars, 140 BPM, F minor |
|
| Fiddle, Bowed Strings, Full, Clean, Spacey, Rich, Intimate, Thick, Rolling, Arp, Fast Speed, Complex, 8 Bars, 128 BPM, B minor |
|
| Chiptune, Chord Progression, Pulse Wave, Medium Reverb, 8 Bars, 128 BPM, D minor |
|
| Kalimba, Mallet, Medium Reverb, Overdriven, Wide, Metallic, Thick, Sparkly, Upper Mids, Bright, Airy, Alternating, Chord Progression, Atmosphere, Spacey, Fast Speed, 8 Bars, 120 BPM, B minor |
|
The examples below show short, note-specific one-shot generations across acoustic, synthetic, bass, FX, and hybrid timbres.
| Prompt | Note | Audio |
|---|---|---|
| FX, Sharp, Subdued, Distant, Sparkly, Round, Hit, Deep, Low Reverb, Sub | D#1 |
|
| Synth Bass, Metallic, Rich, Punchy, Digital, Medium Reverb | C2 |
|
| Reese Bass, Hit, Bright, Wavetable, Noisy, Big, Thick, Airy, Deep | F#2 |
|
| Synth Lead, Shiny, Rich, Choir, Bright, Smooth, Short, Synthetic Vox | F5 |
|
| Tuba, Big, Bright, Clean, Biting, Sustained | G#3 |
|
| Clavinet, Wobble, Breathy, Bright, Deep, Gritty, Impact, Sparkly | C#4 |
|
| Cello, Airy, Wide, Woody, Muffled, Smooth | D#3 |
|
| Reese Bass, Subdued, Round, Growl, Hit, Deep, Woody, Fat, Metallic, Medium Delay, Medium Distortion | D#2 |
|
| Digital Piano, Retro, Smooth, Mono, Warm, Low Reverb | E3 |
|
| Supersaw, Big, Smooth, Warm, Vintage, Crisp, Analog, Wide, Muffled, High Reverb | G2 |
|
| Synth Lead, Bright, Shiny, Soft, Warm, Punchy, Smooth, Medium Delay | C#4 |
|
| Sustained, Trumpet, Airy, Wide, High Reverb | F3 |
|
| 808, Thick, Acid, Hit, FX, Rumble, Wide, Big, Distortion | F1 |
|
| Violin, Vintage, Thick, Chiptune, Spacey, Nasal, Medium Delay | A#5 |
|
| Grand Piano, Near, Snappy, Distant, Bell, Bright, Focused, Delay | D3 |
|
The following examples demonstrate generated keybeds across acoustic, synthetic, and hybrid timbres.
A full video playthrough of the Keybed Showcase is available here, which shows the associated MIDI used to audition each generated instrument.
| Prompt | Audio |
|---|---|
| Soft, Silky, Spectral, Smooth, Spacey, Synth, Pad, Subdued, Wet, High Reverb |
|
| Pad, Focused, Buzzy, Big, Deep, Supersaw, Pulse, Dry |
|
| Pan Flute, Hollow, Dark, Smooth, Sustained, Wide, Silky, Wet, Medium Reverb |
|
| Violin, Woody, Pizzicato, Focused, Breathy, Airy, Wet, High Reverb |
|
| Grand Piano, Warm, Gritty, Wet, Low Reverb |
|
| Grand Piano, Cold, Sparkly, Wet, Low Reverb |
|
| Digital Piano, Noisy, Spiccato, Rich, Pluck, Warm, Swell, Clean, Dry |
|
| Cello, Fat, Warm, Metallic, Sharp, Near, Sustained, Dry |
|
| Church Bell, Glassy, Sparkly, Wet, High Reverb |
|
| Thick, Saw, Neuro, Reese Bass, Dry |
|
| Thick, Sine, Reese Bass, Wet, High Reverb, High Phaser |
|
| Xylophone, Sustained, Analog, Warm, Woody, Dry |
|
| Xylophone, Sustained, Analog, Warm, Woody, Bit Crushed, Wet |
|
| Cello, Harp, Rich, Clean, Choir, Pluck, Formant Vocal, Wet, High Reverb |
|
| Acid, Neuro, Synth Lead, Saw, Square, Wide, Focused, Pluck, Sustained, Wet, High Phaser |
|
| Synth Lead, Present, Sharp, Spacey, Digital, Hollow, Focused, Clean, Wet, Cross Delay, High Reverb |
|
| Dubstep, Neuro, Synth Lead, Pitch Bend, Acid, Wet, Cross Delay |
|
Pure waveform keybeds are shown separately to make the learned oscillator shapes easier to compare directly.
| Waveform | Prompt | Audio |
|---|---|---|
| Triangle | Pure Tone, Triangle, Dry |
|
| Square | Pure Tone, Square, Dry |
|
| Pulse | Pure Tone, Pulse, Dry |
|
| Sine | Pure Tone, Sine, Dry |
|
| Saw | Pure Tone, Saw, Dry |
|
The inference pipeline can also combine three independently generated keybeds into a single layered instrument. The examples below use separate Main and Support prompts for each layer. Further the generated instrument allows independent volume control for each layer.
| Layer Prompts | Audio |
|---|---|
| Main: Bright, Airy, Violin Support 1: Thick, Present, Male, Vocal, Choir Support 2: Digital String, Analog |
|
| Main: Marimba, Dark, Sub, Woody, Airy Support 1: Soft, Grand Piano, Dark Support 2: Music Box, Glassy |
|
Foundation-1 was trained to produce structured musical material rather than full music or generic textures. Musical Notation terms can encourage notation, chord progressions, melodies, arps, phrase direction, rhythmic density, and other musically relevant behaviors.
The model supports a broad instrument hierarchy spanning synths, keys, basses, bowed strings, mallets, winds, guitars, brass, vocals, and plucked strings.
Foundation-1 is not limited to broad instrument naming. It also responds to timbral descriptors such as spectral shape, tone, width, density, texture, brightness, warmth, grit, space, and other sonic traits.
Because instrument identity and timbral character were not collapsed into a single flat label, the model is especially strong at timbral hybridization and layered sonic prompting.
The model supports a dedicated FX layer covering multiple forms of reverb, delay, distortion, phaser, and bitcrushing.
Foundation-1 is built for production-ready loop generation, including BPM-aware and bar-aware structure within supported denominations.
The specialized Foundation-1.2 Samples and Keybeds checkpoints generate short, note-specific sounds across acoustic, synthetic, bass, FX, and hybrid timbres. These can be used individually or as raw material for sampler instruments.
The Keybed checkpoint was trained specifically to preserve a prompted timbral identity across changing pitch and register. RC Stable Audio Tools uses this capability to generate multiple note-specific samples and assemble them into playable instruments.
Foundation-1 was trained with a layered tagging hierarchy designed to improve control, composability, and prompt clarity.
This makes it possible to prompt at different levels of abstraction. A user can stay broad with a family-level prompt like Synth or Keys, or get more specific with terms like Synth Lead, Wavetable Bass, Grand Piano, Violin, or Trumpet, then further shape the output using timbral and FX descriptors.
Foundation-1 was trained across the following major instrument families:
Foundation-1 includes a wide sub-family layer covering a broad range of production-relevant instrument roles, including but not limited to:
One of Foundation-1’s main strengths is that it was not trained to treat timbre as an afterthought. Timbral character is directly represented in the prompt system, giving users control over not only what is being generated, but also how it sounds.
Representative timbre descriptors include:
This tagging design makes prompts much more flexible. Instead of only asking for an instrument, users can shape:
This is especially useful for producers who want to guide the output toward a specific role in a mix rather than just a generic instrument label.
For a list of used tags please see the Tag Reference Sheet.
Foundation-1 includes a dedicated FX descriptor layer spanning multiple common production effects.
Representative FX tags include:
Foundation-1 was trained with structured musical descriptors designed to improve phrase coherence, rhythmic intent, melodic motion, and prompt control.
These notation-style prompt terms help steer:
Examples of supported structural ideas may include terms such as:
This notation layer is one of the main reasons Foundation-1 produces unusually coherent musical material instead of static or loosely related phrases. These can be mixed and matched as desired.
Foundation-1 is designed for structured music production workflows and supports:
For best results, use rich prompts built around the model’s tags. These tags can be mixed and matched as needed. The model was trained on a structured hierarchy designed to encourage musically coherent sample generation.
[Instrument Family / Sub-Family], [Timbre], [Musical Behavior / Notation], [FX], [Key], [Bars], [BPM]
Use FX and timbre tags sparingly at first, then layer more once you understand the model’s behavior.
Each row below uses the exact same prompt, but a different random seed.
The timbre tags remain unchanged, so the overall sound character stays consistent while the melodic and musical content varies between generations.
| Prompt | Output A | Output B | Output C |
|---|---|---|---|
| Bass, FM Bass, Medium Delay, Medium Reverb, Low Distortion, Phaser, Acid, Gritty, Wide, Dubstep, Thick, Silky, Warm, Rich, Overdriven, Crisp, Deep, Clean, Triplets, 8 Bars, 150 BPM, A minor |
|
|
|
| Gritty, Acid, Bassline, 303, Synth Lead, FM, Sub, Upper Mids, High Phaser, High Reverb, Pitch Bend, 8 Bars, 140 BPM, E minor |
|
|
|
| Kalimba, Mallet, Medium Reverb, Overdriven, Wide, Metallic, Thick, Sparkly, Upper Mids, Bright, Airy, Small, Alternating Chord Progression, Atmosphere, Spacey, Fast, 4 Bars, 120 BPM, B minor |
|
|
|
Foundation-1 is best used with RC Stable Audio Tools, which is tuned around the model family, its metadata, and its structured prompting system.
RC Stable Audio Tools (Enhanced Fork)
The interface supports different workflows depending on the checkpoint in use.
For the original Foundation-1 loop checkpoint, the interface provides:
The specialized Samples workflow supports:
The Keybed checkpoint can also generate individual one-shots effectively. The dedicated Samples checkpoint is provided because its earlier training endpoint preserves stronger general sample-generation quality.
The Keybed workflow is designed to be used in conjunction with RC Stable Audio Tools rather than as a sequence of manually generated independent notes.
The interface handles the pitch-aware generation pipeline, creates the source files required for the instrument, and can export completed keybeds for sampler use.
The workflow supports:
Generated keybeds can be assembled and exported as playable DecentSampler or SFZ instruments.
RC Stable Audio Tools (Enhanced Fork)
Stable Audio Tools (Original Repository)
Foundation-1 is now distributed as multiple specialized checkpoints built around the same underlying architecture and conditioning system.
The specialized Foundation-1.2 Samples and Keybeds checkpoints retain the same Foundation-1 / SAO architecture. They therefore do not represent separate model architectures; the specialization comes from the training objective and selected checkpoint.
The release uses 16-bit model weights to reduce the model footprint without changing the intended inference quality.
Foundation_1.safetensors — loop-generation checkpointFoundation-1.2-Samples.safetensors — sample-focused checkpointFoundation-1.2-Keybeds.safetensors — keybed-focused checkpointmodel_config.json — shared model configurationmodels directoryFor full keybeds, the RC interface is strongly recommended because it manages the multi-note inference and sampler-export process automatically.
Foundation-1 is designed to run locally on modern GPUs.
Typical VRAM usage during generation is approximately ~7 GB.
For reliable operation, a GPU with at least 8 GB of VRAM is recommended.
Generation speed will vary depending on GPU model, generation mode, sample length, and system configuration.
On an RTX 3090, a standard individual generation is approximately ~7–8 seconds per sample. Full single-layer keybed builds take approximately 50 seconds.
Foundation-1 was built around a structured sample-generation philosophy, rather than generic or genre-based audio captioning. The dataset consists entirely of hand-crafted and labeled audio, produced through a controlled augmentation pipeline.
At a high level, the training design emphasizes:
This design is central to the model’s musical coherence and high degree of sonic control.
For more details on the original dataset and training methodology, see the Training & Dataset Notes.
The Foundation-1.2 Samples and Keybeds checkpoints use the same underlying Foundation-1 architecture, but were trained with a lower learning rate of 1e-5.
| Checkpoint | Training Endpoint | Specialization |
|---|---|---|
| Foundation-1.2 Samples | Epoch 0 / Step 560 | General sample / one-shot generation |
| Foundation-1.2 Keybeds | Epoch 3 / Step 8120 | Cross-pitch timbral consistency and playable keybed generation |
Additional training configuration:
1e-51e-3The longer Keybed run was selected to reinforce cross-note and cross-register consistency. The objective was not simply to improve isolated note quality, but to teach the model how a single sound identity should behave as pitch changes across an instrument.
The Keybed checkpoint used a dedicated multi-note training strategy rather than treating every pitch as an unrelated example. The generation pipeline mirrors that structure by injecting note-sequence conditioning, reusing a common seed across keybed chunks, slicing generated sequences back into individual notes, and packaging the resulting samples into playable instruments.
For a detailed explanation of both the training method and the inference/export pipeline, see the Keybed Training & Inference Strategy.
Foundation-1 is a specialized model family for producer-facing sample and instrument generation, not a general-purpose full-song generator.
Important notes:
The model is also optimized around specific timing relationships between Bars, BPM, and generation duration.
For example:
If the generation duration is shorter than the musical structure implied by the prompt (for example requesting an 8-bar loop but generating only 5 seconds), the model may produce less coherent musical phrases.
The RC Stable Audio Fork automatically handles this timing alignment, making this workflow much easier.
The Keybed model is designed around practical instrument ranges.
Some prompts simply stop making semantic or perceptual sense at extreme registers. For example, a sub-bass at C7 is no longer functioning as a bass, even if the model preserves some aspects of the original waveform or timbral character.
The same applies to many synthetic and acoustic sounds. As pitch rises into extreme upper registers, complex waveforms often become perceptually simpler and can collapse toward high-pitched pure-tone-like behavior. At the opposite end, very low pitches can become increasingly difficult to render consistently, and small pitch or waveform errors become more obvious.
For this reason, RC Stable Audio Tools clamps generation and export ranges to practical regions rather than forcing every generated sound across the full MIDI note range.
Prompt choice also matters. A timbre that has a natural low, mid, or high-register identity may become ambiguous when pushed several octaves outside that range. Some apparent "timbre drift" is therefore not only a model limitation; it is also a consequence of asking for a sound whose defining characteristics change as pitch moves far beyond where that sound is normally perceived.
In my internal testing, roughly 90% of generated instruments maintain a consistent relationship between pitch and timbral identity across the supported keybed ranges. This is a qualitative estimate rather than a programmatic benchmark, because there is no reliable automated metric for determining whether two notes share the same perceived timbre.
The remaining edge cases are most likely to appear with:
The default export ranges were chosen as a practical compromise between keyboard coverage, pitch accuracy, and timbral consistency.
This model is licensed under the Stability AI Community License. It is available for non-commercial use or limited commercial use by entities with annual revenues below USD $1M. For revenues exceeding USD $1M, please refer to the repository license file for full terms.
The original Foundation-1 video covering the model's design philosophy, structured prompting system, timbral control, and loop-generation workflow.
🎥 Watch the Foundation-1 Overview
A companion video covering the new Samples and Keybeds checkpoints, the updated training strategy and general journey to getting this made can be found here.
🎥 Watch the Foundation-1.2 Update
A hands-on walkthrough showing the Keybed model in use, cross-pitch timbral behavior, and the new three-layer instrument exporter.
🎥 Watch the Guided Keybed Demo
Foundation-1 is intended as a producer-facing model family for structured sample and instrument generation, designed to augment music production.
Its goal is to let users explore sound in new ways while retaining precise control over:
The original loop model focuses on structured musical material. The specialized Foundation-1.2 Samples and Keybeds checkpoints extend that same conditioning system toward sound design and playable instrument creation.
That combination of musical structure, instrument identity, timbral control, loop fidelity, and cross-pitch instrument generation is what defines the Foundation-1 family.
48 commits