RoyalCities/Foundation-1

Model

363

stars

48

commits

3

linked in READMEs

Sep 9, 2026

updated

audio
Audio-to-Audio
fine-tuning
music-generation
Music Production
sample-generation
stable-audio
text-to-synth

README

Foundation-1 Banner

Foundation-1

Structured text-to-sample and text-to-instrument generation for modern music production


Overview

Foundation-1 is a family of text-conditioned audio models designed around music production workflows. Rather than treating audio generation as broad caption-to-music generation, Foundation-1 was trained around structured controls for instrument identity, timbre, FX, musical behavior, pitch, timing, and tonality.

The original Foundation-1 checkpoint focuses on structured loop generation. Newer specialized checkpoints extend the same conditioning system into individual one-shots and pitch-consistent playable keybeds.

This means Foundation-1 can now be used at several levels of a production workflow:

  • generate a tempo-synced musical loop
  • generate an individual note or sound
  • generate multiple pitch-consistent notes from the same text prompt
  • assemble those notes into a playable sampler instrument
  • combine multiple generated keybeds into layered instruments

The goal is not simply to generate finished audio. The specialized keybed workflow is designed to let the model generate the source material for an instrument, then hand control back to the producer.


Foundation-1 Model Family

ModelPrimary UseNotes
Foundation-1Loop generationOriginal checkpoint for structured, BPM-aware, bar-aware musical loops
Foundation-1.2 SamplesLoop Generation / One-Shot generationSelected earlier in specialized training to preserve stronger short-form loop-generation quality
Foundation-1.2 KeybedsOne-Shot generation, Pitch-consistent sample generation and playable instrumentsTrained longer to reinforce timbral consistency across pitch and register

The Foundation-1.2 Keybeds model is not limited to full keybed generation. It remains highly capable at generating individual samples and one-shots. Its specialization reflects the additional training needed to maintain a consistent sonic identity across many generated pitches.

The checkpoints were separated because the longer keybed-focused training improved cross-pitch consistency, while reducing the broader loop-generation quality preserved by the earlier Samples checkpoint.


Text-to-Synth / Keybed Generation

The Keybed checkpoint extends Foundation-1 from text-to-sample generation into a practical text-to-instrument workflow.

A user can describe a sound using the same instrument and timbre vocabulary used elsewhere in Foundation-1, generate pitch-specific samples across a keyboard, and automatically assemble the result into a playable sampler instrument.

Examples can range from conventional instruments:

Grand Piano, Warm, Gritty, Wet, Low Reverb

to synthetic or hybrid sounds:

Cello, Harp, Rich, Clean, Choir, Pluck, Formant Vocal, Wet, High Reverb

Because pitch is generated rather than conventionally pitch-shifted from a single root sample, the model can create natural timbral variation across registers while maintaining the identity of the prompted sound.

For the intended workflow, the Keybed model is best used with RC Stable Audio Tools, which handles the multi-note generation process and instrument assembly automatically. Generated keybeds can be exported for use as playable sampler instruments rather than remaining as isolated audio generations.

The user-facing prompt is intentionally simple. A descriptor such as:

Grand Piano, Warm, Gritty

is expanded internally into the structured Keybed conditioning grammar used by the model. RC Stable Audio Tools adds the required Keybed / sequence / note information, keeps the user descriptor stable across the instrument, reuses the same resolved seed across generation chunks, then slices and maps the resulting notes into the exported sampler.

This makes the workflow useful not only as a demo interface, but also as a reference implementation for developers interested in building their own VST, sampler, or instrument-generation front end around the model.

For the full prompt-injection and inference design, see the Keybed Training & Inference Strategy.

Foundation-1.2 Exported Keybed

Foundation-1.2 Keybed generation workflow in RC Stable Audio Tools.


What Foundation-1 Does

  • Generates musically coherent loops for production workflows
  • Generates individual note-specific one-shots
  • Generates pitch-consistent multi-note keybeds
  • Builds playable sampler instruments through the RC Stable Audio Tools workflow
  • Supports layered keybeds built from multiple independently prompted sounds
  • Understands BPM and bar count for structured loop generation
  • Locks to major and minor keys across western music theory
  • Supports enharmonic equivalents when prompting scales and keys
  • Separates instrument identity from timbral character
  • Supports timbral mixing by combining instrument and sonic descriptors
  • Responds to FX tags such as reverb, delay, distortion, and modulation
  • Uses notation-style prompt structure to encourage coherent phrasing, melodic shape, rhythmic behavior, and harmonic motion
  • Produces perfect loops within supported BPM / bar denominations
  • Understands Wet vs Dry production context — adding terms like Dry encourages minimal FX processing, while Wet or FX tags produce more processed, spatial, or effected sounds.

Why It Feels Different

Most audio models can react to broad prompt terms like “warm pad” or “bright synth.” with inconsistent results. Foundation-1 was designed to go further by treating the sound as a layered system:

  1. Instrument Family – what broad source category the sound belongs to
  2. Sub-Family – the more specific instrument role or identity
  3. Timbre Tags – the tonal, spectral, or textural character
  4. FX Tags – the processing layer applied to the sound
  5. Notation / Structure Tags – the musical behavior of the generated phrase

This layered conditioning approach is a major reason Foundation-1 is able to deliver both high musicality and high prompt control at the same time.


Audio Showcase

Loop Showcase

PromptAudio
Bass, FM Bass, Medium Delay, Medium Reverb, Low Distortion, Phaser, Sub Bass, Bass, Upper Mids, Acid, Gritty, Wide, Dubstep, Thick, Silky, Warm, Rich, Overdriven, Crisp, Deep, Clean, Pitch Bend, 303, 8 Bars, 140 BPM, E minor
Sub Bass, Bass, Gritty, Small, Square, Bass, Dark, Digital, Thick, Clean, Simple, Bassline, Epic, Choppy, Melody, 4 Bars, 150 BPM, G# minor
Flute, Pizzicato, Punchy, Present, Ambient, Nasal, Melody, Epic, Airy, Slow Speed, 8 Bars, 150 BPM, E minor
High Saw, Spacey, Lead, Warm, Silky, Smooth, 303, Synth Lead, Medium Reverb, Low Distortion, Upper Mids, Mids, Pitch Bend, Arp, 8 Bars, 140 BPM, F minor
Trumpet, Warm, Complex Arp Melody, High Reverb, Low Distortion, Smooth, Silky, Texture, 8 Bars, 130 BPM, C minor
Synth, Pad, Chord Progression, Rising, Digital, Bass, Fat, Near, Wide, Silky, Warm, Focused, 8 Bars, 110 BPM, D major
Piccolo, Flute, Airy, Music Box, plucked, complex melody, 8 Bars, 140 BPM, C# minor
Synth Lead, Wavetable Bass, Low Distortion, High Reverb, Sub Bass, Upper Mids, Acid, Gritty, Wide, Thick, Silky, Warm, Rich, Overdriven, Crisp, Clean, 303, Complex, 8 Bars, 140 BPM, F minor
Fiddle, Bowed Strings, Full, Clean, Spacey, Rich, Intimate, Thick, Rolling, Arp, Fast Speed, Complex, 8 Bars, 128 BPM, B minor
Chiptune, Chord Progression, Pulse Wave, Medium Reverb, 8 Bars, 128 BPM, D minor
Kalimba, Mallet, Medium Reverb, Overdriven, Wide, Metallic, Thick, Sparkly, Upper Mids, Bright, Airy, Alternating, Chord Progression, Atmosphere, Spacey, Fast Speed, 8 Bars, 120 BPM, B minor

One-Shot Showcase

The examples below show short, note-specific one-shot generations across acoustic, synthetic, bass, FX, and hybrid timbres.

PromptNoteAudio
FX, Sharp, Subdued, Distant, Sparkly, Round, Hit, Deep, Low Reverb, SubD#1
Synth Bass, Metallic, Rich, Punchy, Digital, Medium ReverbC2
Reese Bass, Hit, Bright, Wavetable, Noisy, Big, Thick, Airy, DeepF#2
Synth Lead, Shiny, Rich, Choir, Bright, Smooth, Short, Synthetic VoxF5
Tuba, Big, Bright, Clean, Biting, SustainedG#3
Clavinet, Wobble, Breathy, Bright, Deep, Gritty, Impact, SparklyC#4
Cello, Airy, Wide, Woody, Muffled, SmoothD#3
Reese Bass, Subdued, Round, Growl, Hit, Deep, Woody, Fat, Metallic, Medium Delay, Medium DistortionD#2
Digital Piano, Retro, Smooth, Mono, Warm, Low ReverbE3
Supersaw, Big, Smooth, Warm, Vintage, Crisp, Analog, Wide, Muffled, High ReverbG2
Synth Lead, Bright, Shiny, Soft, Warm, Punchy, Smooth, Medium DelayC#4
Sustained, Trumpet, Airy, Wide, High ReverbF3
808, Thick, Acid, Hit, FX, Rumble, Wide, Big, DistortionF1
Violin, Vintage, Thick, Chiptune, Spacey, Nasal, Medium DelayA#5
Grand Piano, Near, Snappy, Distant, Bell, Bright, Focused, DelayD3

Keybed Showcase

The following examples demonstrate generated keybeds across acoustic, synthetic, and hybrid timbres.

A full video playthrough of the Keybed Showcase is available here, which shows the associated MIDI used to audition each generated instrument.

PromptAudio
Soft, Silky, Spectral, Smooth, Spacey, Synth, Pad, Subdued, Wet, High Reverb
Pad, Focused, Buzzy, Big, Deep, Supersaw, Pulse, Dry
Pan Flute, Hollow, Dark, Smooth, Sustained, Wide, Silky, Wet, Medium Reverb
Violin, Woody, Pizzicato, Focused, Breathy, Airy, Wet, High Reverb
Grand Piano, Warm, Gritty, Wet, Low Reverb
Grand Piano, Cold, Sparkly, Wet, Low Reverb
Digital Piano, Noisy, Spiccato, Rich, Pluck, Warm, Swell, Clean, Dry
Cello, Fat, Warm, Metallic, Sharp, Near, Sustained, Dry
Church Bell, Glassy, Sparkly, Wet, High Reverb
Thick, Saw, Neuro, Reese Bass, Dry
Thick, Sine, Reese Bass, Wet, High Reverb, High Phaser
Xylophone, Sustained, Analog, Warm, Woody, Dry
Xylophone, Sustained, Analog, Warm, Woody, Bit Crushed, Wet
Cello, Harp, Rich, Clean, Choir, Pluck, Formant Vocal, Wet, High Reverb
Acid, Neuro, Synth Lead, Saw, Square, Wide, Focused, Pluck, Sustained, Wet, High Phaser
Synth Lead, Present, Sharp, Spacey, Digital, Hollow, Focused, Clean, Wet, Cross Delay, High Reverb
Dubstep, Neuro, Synth Lead, Pitch Bend, Acid, Wet, Cross Delay

Generated Waveforms

Pure waveform keybeds are shown separately to make the learned oscillator shapes easier to compare directly.

WaveformPromptAudio
TrianglePure Tone, Triangle, Dry
SquarePure Tone, Square, Dry
PulsePure Tone, Pulse, Dry
SinePure Tone, Sine, Dry
SawPure Tone, Saw, Dry

Multi-Layered Keybeds

The inference pipeline can also combine three independently generated keybeds into a single layered instrument. The examples below use separate Main and Support prompts for each layer. Further the generated instrument allows independent volume control for each layer.

Layer PromptsAudio
Main: Bright, Airy, Violin
Support 1: Thick, Present, Male, Vocal, Choir
Support 2: Digital String, Analog
Main: Marimba, Dark, Sub, Woody, Airy
Support 1: Soft, Grand Piano, Dark
Support 2: Music Box, Glassy

Core Capabilities

1. Musical Structure

Foundation-1 was trained to produce structured musical material rather than full music or generic textures. Musical Notation terms can encourage notation, chord progressions, melodies, arps, phrase direction, rhythmic density, and other musically relevant behaviors.

2. Instrument Identity

The model supports a broad instrument hierarchy spanning synths, keys, basses, bowed strings, mallets, winds, guitars, brass, vocals, and plucked strings.

3. Timbral Control

Foundation-1 is not limited to broad instrument naming. It also responds to timbral descriptors such as spectral shape, tone, width, density, texture, brightness, warmth, grit, space, and other sonic traits.

4. Timbral Mixing

Because instrument identity and timbral character were not collapsed into a single flat label, the model is especially strong at timbral hybridization and layered sonic prompting.

5. FX Prompting

The model supports a dedicated FX layer covering multiple forms of reverb, delay, distortion, phaser, and bitcrushing.

6. Loop Fidelity

Foundation-1 is built for production-ready loop generation, including BPM-aware and bar-aware structure within supported denominations.

7. One-Shot Generation

The specialized Foundation-1.2 Samples and Keybeds checkpoints generate short, note-specific sounds across acoustic, synthetic, bass, FX, and hybrid timbres. These can be used individually or as raw material for sampler instruments.

8. Pitch-Consistent Keybeds

The Keybed checkpoint was trained specifically to preserve a prompted timbral identity across changing pitch and register. RC Stable Audio Tools uses this capability to generate multiple note-specific samples and assemble them into playable instruments.


Conditioning Architecture

Foundation-1 was trained with a layered tagging hierarchy designed to improve control, composability, and prompt clarity.

Hierarchy Overview

  • Major Family → broad instrument class
  • Sub-Family → more specific instrument role
  • Timbre Tags → tonal / spectral / textural descriptors
  • FX Tags → processing layer
  • Notation Tags → musical behavior and phrasing

This makes it possible to prompt at different levels of abstraction. A user can stay broad with a family-level prompt like Synth or Keys, or get more specific with terms like Synth Lead, Wavetable Bass, Grand Piano, Violin, or Trumpet, then further shape the output using timbral and FX descriptors.


Instrument Coverage

Major Families

Foundation-1 was trained across the following major instrument families:

  • Synth
  • Keys
  • Bass
  • Bowed Strings
  • Mallet
  • Wind
  • Guitar
  • Brass
  • Vocal
  • Plucked Strings

Sub-Family Coverage

Foundation-1 includes a wide sub-family layer covering a broad range of production-relevant instrument roles, including but not limited to:

  • Synth Lead
  • Synth Bass
  • Digital Piano
  • Pluck
  • Grand Piano
  • Bell
  • Pad
  • Atmosphere
  • Digital Strings
  • FM Synth
  • Violin
  • Digital Organ
  • Supersaw
  • Wavetable Bass
  • Rhodes Piano
  • Cello
  • Texture
  • Flute
  • Reese Bass
  • Wavetable Synth
  • Electric Bass
  • Marimba
  • Trumpet
  • Pan Flute
  • Choir
  • Harp
  • Church Organ
  • Acoustic Guitar
  • Hammond Organ
  • Celesta
  • Vibraphone
  • Glockenspiel
  • Ocarina
  • Clarinet
  • French Horn
  • Tuba
  • Oboe
Sub-Family Chart

Timbre System

One of Foundation-1’s main strengths is that it was not trained to treat timbre as an afterthought. Timbral character is directly represented in the prompt system, giving users control over not only what is being generated, but also how it sounds.

Representative timbre descriptors include:

  • Warm
  • Bright
  • Wide
  • Airy
  • Thick
  • Rich
  • Tight
  • Full
  • Gritty
  • Clean
  • Retro
  • Saw
  • Crisp
  • Focused
  • Metallic
  • Chiptune
  • Dark
  • 303
  • Shiny
  • Analog
  • Present
  • Sparkly
  • Ambient
  • Soft
  • Smooth
  • Cold
  • Buzzy
  • Deep
  • Formant Vocal
  • Round
  • Punchy
  • Nasal
  • Vintage
  • Growl
  • Breathy
  • Glassy
  • Noisy
  • Synthetic Vox
  • Supersaw
  • Bitcrushed
  • Dreamy
Timbre Chart

Why This Matters

This tagging design makes prompts much more flexible. Instead of only asking for an instrument, users can shape:

  • tonal balance
  • brightness / darkness
  • width / intimacy
  • clean vs driven character
  • synthetic vs organic feel
  • transient sharpness
  • texture and density
  • spatial character

This is especially useful for producers who want to guide the output toward a specific role in a mix rather than just a generic instrument label.

For a list of used tags please see the Tag Reference Sheet.


FX Layer

Foundation-1 includes a dedicated FX descriptor layer spanning multiple common production effects.

Representative FX tags include:

  • Low Reverb
  • Medium Reverb
  • High Reverb
  • Plate Reverb
  • Low Delay
  • Medium Delay
  • High Delay
  • Ping Pong Delay
  • Stereo Delay
  • Cross Delay
  • Mono Delay
  • Low Distortion
  • Medium Distortion
  • High Distortion
  • Phaser
  • Low Phaser
  • Medium Phaser
  • High Phaser
  • Bitcrush
  • High Bitcrush
FX Chart

Musical Notation and Structure

Foundation-1 was trained with structured musical descriptors designed to improve phrase coherence, rhythmic intent, melodic motion, and prompt control.

These notation-style prompt terms help steer:

  • chord progressions
  • melodies
  • top-line layers
  • arpeggios
  • phrase direction
  • rhythmic density
  • harmonic feel
  • subdivision style
  • simple vs complex motion
  • sustained vs plucked behavior
  • melodic contour and pacing

Examples of supported structural ideas may include terms such as:

  • chord progression
  • melody
  • top melody
  • arp
  • triplets
  • simple
  • complex
  • rising
  • falling
  • strummed
  • sustained
  • catchy
  • epic
  • slow
  • fast

This notation layer is one of the main reasons Foundation-1 produces unusually coherent musical material instead of static or loosely related phrases. These can be mixed and matched as desired.


Tonal and Timing Support

Foundation-1 is designed for structured music production workflows and supports:

Keys and Modes

  • Major keys
  • Minor keys
  • Enharmonic equivalents
  • Western 12-tone chromatic prompting

Loop Structure

  • Supported bar lengths: 4 Bars, 8 Bars
  • Supported BPM denominations: 100 BPM, 110 BPM, 120 BPM, 128 BPM, 130 BPM, 140 BPM, 150 BPM

Prompt Structure

For best results, use rich prompts built around the model’s tags. These tags can be mixed and matched as needed. The model was trained on a structured hierarchy designed to encourage musically coherent sample generation.

Layered Prompt Structure

[Instrument Family / Sub-Family], [Timbre], [Musical Behavior / Notation], [FX], [Key], [Bars], [BPM]

Prompting Notes

  • Start with a clear instrument identity
  • Add 1–3 timbre descriptors for stronger steering
  • Include a notation or musical structure term for better phrase coherence
  • Always include Bars and BPM, which define the musical loop length
  • Ensure the generation duration matches the requested musical structure
  • The RC Stable Audio Fork automatically handles this timing alignment

Use FX and timbre tags sparingly at first, then layer more once you understand the model’s behavior.


One Prompt → Multiple Outputs

Each row below uses the exact same prompt, but a different random seed.
The timbre tags remain unchanged, so the overall sound character stays consistent while the melodic and musical content varies between generations.

PromptOutput AOutput BOutput C
Bass, FM Bass, Medium Delay, Medium Reverb, Low Distortion, Phaser, Acid, Gritty, Wide, Dubstep, Thick, Silky, Warm, Rich, Overdriven, Crisp, Deep, Clean, Triplets, 8 Bars, 150 BPM, A minor
Gritty, Acid, Bassline, 303, Synth Lead, FM, Sub, Upper Mids, High Phaser, High Reverb, Pitch Bend, 8 Bars, 140 BPM, E minor
Kalimba, Mallet, Medium Reverb, Overdriven, Wide, Metallic, Thick, Sparkly, Upper Mids, Bright, Airy, Small, Alternating Chord Progression, Atmosphere, Spacey, Fast, 4 Bars, 120 BPM, B minor

Foundation-1 is best used with RC Stable Audio Tools, which is tuned around the model family, its metadata, and its structured prompting system.

RC Stable Audio Tools (Enhanced Fork)

The interface supports different workflows depending on the checkpoint in use.

Loop Generation

For the original Foundation-1 loop checkpoint, the interface provides:

  • structured prompt building aligned with the training tags
  • random prompt generation
  • automatic BPM / bar timing alignment
  • automatic MIDI extraction from generated audio
  • generation settings tuned for Foundation-1

Samples / One-Shot Generation

The specialized Samples workflow supports:

  • note-specific generation
  • instrument and timbre prompting
  • short sample-focused inference
  • rapid auditioning of generated sounds

The Keybed checkpoint can also generate individual one-shots effectively. The dedicated Samples checkpoint is provided because its earlier training endpoint preserves stronger general sample-generation quality.

Keybed / Text-to-Synth Generation

The Keybed workflow is designed to be used in conjunction with RC Stable Audio Tools rather than as a sequence of manually generated independent notes.

The interface handles the pitch-aware generation pipeline, creates the source files required for the instrument, and can export completed keybeds for sampler use.

The workflow supports:

  • automatic generation across keyboard registers
  • pitch-specific text-conditioned samples
  • playable keybed construction
  • DecentSampler export
  • SFZ export
  • per-instrument ADSR controls
  • built-in sampler effects and tone shaping
  • layered keybeds using independently prompted Main / Support layers
  • independent layer volume control
  • generated instrument previews

Foundation-1.2 Keybed exporter

Generated keybeds can be assembled and exported as playable DecentSampler or SFZ instruments.

RC Stable Audio Tools (Enhanced Fork)

Stable Audio Tools (Original Repository)

Model Files

Foundation-1 is now distributed as multiple specialized checkpoints built around the same underlying architecture and conditioning system.

The specialized Foundation-1.2 Samples and Keybeds checkpoints retain the same Foundation-1 / SAO architecture. They therefore do not represent separate model architectures; the specialization comes from the training objective and selected checkpoint.

The release uses 16-bit model weights to reduce the model footprint without changing the intended inference quality.

  • Foundation_1.safetensors — loop-generation checkpoint
  • Foundation-1.2-Samples.safetensors — sample-focused checkpoint
  • Foundation-1.2-Keybeds.safetensors — keybed-focused checkpoint
  • model_config.json — shared model configuration

Basic Setup for RC Stable Audio Tools

  1. Create a subfolder inside your models directory
  2. Place the desired checkpoint and its compatible config in that folder
  3. Launch the interface
  4. Select the checkpoint from the model selector
  5. Choose the matching Loop, One-Shot, or Keybed workflow
  6. Prompt with layered instrument and timbre descriptors

For full keybeds, the RC interface is strongly recommended because it manages the multi-note inference and sampler-export process automatically.

Hardware Requirements

Foundation-1 is designed to run locally on modern GPUs.

Typical VRAM usage during generation is approximately ~7 GB.
For reliable operation, a GPU with at least 8 GB of VRAM is recommended.

Generation Performance

Generation speed will vary depending on GPU model, generation mode, sample length, and system configuration.

On an RTX 3090, a standard individual generation is approximately ~7–8 seconds per sample. Full single-layer keybed builds take approximately 50 seconds.


Dataset and Training Philosophy

Foundation-1 was built around a structured sample-generation philosophy, rather than generic or genre-based audio captioning. The dataset consists entirely of hand-crafted and labeled audio, produced through a controlled augmentation pipeline.

At a high level, the training design emphasizes:

  • structured musical loops
  • instrument hierarchy
  • explicit timbre representation
  • dedicated FX descriptors
  • notation-aware prompt terms
  • strong production relevance
  • broad reuse for compositional workflows

This design is central to the model’s musical coherence and high degree of sonic control.

For more details on the original dataset and training methodology, see the Training & Dataset Notes.

Specialized Samples and Keybeds Training

The Foundation-1.2 Samples and Keybeds checkpoints use the same underlying Foundation-1 architecture, but were trained with a lower learning rate of 1e-5.

CheckpointTraining EndpointSpecialization
Foundation-1.2 SamplesEpoch 0 / Step 560General sample / one-shot generation
Foundation-1.2 KeybedsEpoch 3 / Step 8120Cross-pitch timbral consistency and playable keybed generation

Additional training configuration:

  • Optimizer: AdamW
  • Learning Rate: 1e-5
  • Weight Decay: 1e-3
  • Scheduler: InverseLR
  • EMA: Enabled
  • Sample Rate: 44,100 Hz
  • Channels: Stereo
  • Training Window: 882,000 samples (~20 seconds)

The longer Keybed run was selected to reinforce cross-note and cross-register consistency. The objective was not simply to improve isolated note quality, but to teach the model how a single sound identity should behave as pitch changes across an instrument.

The Keybed checkpoint used a dedicated multi-note training strategy rather than treating every pitch as an unrelated example. The generation pipeline mirrors that structure by injecting note-sequence conditioning, reusing a common seed across keybed chunks, slicing generated sequences back into individual notes, and packaging the resulting samples into playable instruments.

For a detailed explanation of both the training method and the inference/export pipeline, see the Keybed Training & Inference Strategy.


Limitations

Foundation-1 is a specialized model family for producer-facing sample and instrument generation, not a general-purpose full-song generator.

Important notes:

  • It performs best when prompted using vocabulary aligned with the training design
  • It is optimized for sample-generation workflows, not open-ended genre captioning
  • Only two genre tags were included (Dubstep Growls and Chiptune waveforms), primarily to reinforce waveform behaviors
  • Prompt quality matters — structured layered prompts outperform vague natural language
  • Some timbre tags exert stronger influence than others
  • Certain tag combinations may require iteration to achieve the exact musical role or timbral blend desired
  • Percussion and drum sounds are outside the scope of this release

The model is also optimized around specific timing relationships between Bars, BPM, and generation duration.

For example:

  • an 8-bar loop at 100 BPM ≈ 19 seconds

If the generation duration is shorter than the musical structure implied by the prompt (for example requesting an 8-bar loop but generating only 5 seconds), the model may produce less coherent musical phrases.

The RC Stable Audio Fork automatically handles this timing alignment, making this workflow much easier.

Keybed-Specific Limitations

The Keybed model is designed around practical instrument ranges.

Some prompts simply stop making semantic or perceptual sense at extreme registers. For example, a sub-bass at C7 is no longer functioning as a bass, even if the model preserves some aspects of the original waveform or timbral character.

The same applies to many synthetic and acoustic sounds. As pitch rises into extreme upper registers, complex waveforms often become perceptually simpler and can collapse toward high-pitched pure-tone-like behavior. At the opposite end, very low pitches can become increasingly difficult to render consistently, and small pitch or waveform errors become more obvious.

For this reason, RC Stable Audio Tools clamps generation and export ranges to practical regions rather than forcing every generated sound across the full MIDI note range.

Prompt choice also matters. A timbre that has a natural low, mid, or high-register identity may become ambiguous when pushed several octaves outside that range. Some apparent "timbre drift" is therefore not only a model limitation; it is also a consequence of asking for a sound whose defining characteristics change as pitch moves far beyond where that sound is normally perceived.

In my internal testing, roughly 90% of generated instruments maintain a consistent relationship between pitch and timbral identity across the supported keybed ranges. This is a qualitative estimate rather than a programmatic benchmark, because there is no reliable automated metric for determining whether two notes share the same perceived timbre.

The remaining edge cases are most likely to appear with:

  • extreme low or high registers
  • unconventional bass prompts outside bass ranges
  • strongly formant-dependent sounds
  • heavily processed or hybrid timbres
  • prompts whose intended instrument identity becomes ambiguous at the requested pitch

The default export ranges were chosen as a practical compromise between keyboard coverage, pitch accuracy, and timbral consistency.


License

This model is licensed under the Stability AI Community License. It is available for non-commercial use or limited commercial use by entities with annual revenues below USD $1M. For revenues exceeding USD $1M, please refer to the repository license file for full terms.


Companion Videos

Foundation-1 Overview

The original Foundation-1 video covering the model's design philosophy, structured prompting system, timbral control, and loop-generation workflow.

🎥 Watch the Foundation-1 Overview

Foundation-1.2 Samples & Keybeds Update

A companion video covering the new Samples and Keybeds checkpoints, the updated training strategy and general journey to getting this made can be found here.

🎥 Watch the Foundation-1.2 Update

Keybed Guided Demo

A hands-on walkthrough showing the Keybed model in use, cross-pitch timbral behavior, and the new three-layer instrument exporter.

🎥 Watch the Guided Keybed Demo



Final Notes

Foundation-1 is intended as a producer-facing model family for structured sample and instrument generation, designed to augment music production.

Its goal is to let users explore sound in new ways while retaining precise control over:

  • what the sound is
  • how it behaves musically
  • how it changes across pitch
  • how it sits tonally
  • how it feels sonically
  • how it fits into a production workflow

The original loop model focuses on structured musical material. The specialized Foundation-1.2 Samples and Keybeds checkpoints extend that same conditioning system toward sound design and playable instrument creation.

That combination of musical structure, instrument identity, timbral control, loop fidelity, and cross-pitch instrument generation is what defines the Foundation-1 family.

Contributors

RoyalCities

48 commits

RoyalCities/Foundation-1

Model

363

stars

48

commits

3

linked in READMEs

Sep 9, 2026

updated

audio
Audio-to-Audio
fine-tuning
music-generation
Music Production
sample-generation
stable-audio
text-to-synth

README

Foundation-1 Banner

Foundation-1

Structured text-to-sample and text-to-instrument generation for modern music production


Overview

Foundation-1 is a family of text-conditioned audio models designed around music production workflows. Rather than treating audio generation as broad caption-to-music generation, Foundation-1 was trained around structured controls for instrument identity, timbre, FX, musical behavior, pitch, timing, and tonality.

The original Foundation-1 checkpoint focuses on structured loop generation. Newer specialized checkpoints extend the same conditioning system into individual one-shots and pitch-consistent playable keybeds.

This means Foundation-1 can now be used at several levels of a production workflow:

  • generate a tempo-synced musical loop
  • generate an individual note or sound
  • generate multiple pitch-consistent notes from the same text prompt
  • assemble those notes into a playable sampler instrument
  • combine multiple generated keybeds into layered instruments

The goal is not simply to generate finished audio. The specialized keybed workflow is designed to let the model generate the source material for an instrument, then hand control back to the producer.


Foundation-1 Model Family

ModelPrimary UseNotes
Foundation-1Loop generationOriginal checkpoint for structured, BPM-aware, bar-aware musical loops
Foundation-1.2 SamplesLoop Generation / One-Shot generationSelected earlier in specialized training to preserve stronger short-form loop-generation quality
Foundation-1.2 KeybedsOne-Shot generation, Pitch-consistent sample generation and playable instrumentsTrained longer to reinforce timbral consistency across pitch and register

The Foundation-1.2 Keybeds model is not limited to full keybed generation. It remains highly capable at generating individual samples and one-shots. Its specialization reflects the additional training needed to maintain a consistent sonic identity across many generated pitches.

The checkpoints were separated because the longer keybed-focused training improved cross-pitch consistency, while reducing the broader loop-generation quality preserved by the earlier Samples checkpoint.


Text-to-Synth / Keybed Generation

The Keybed checkpoint extends Foundation-1 from text-to-sample generation into a practical text-to-instrument workflow.

A user can describe a sound using the same instrument and timbre vocabulary used elsewhere in Foundation-1, generate pitch-specific samples across a keyboard, and automatically assemble the result into a playable sampler instrument.

Examples can range from conventional instruments:

Grand Piano, Warm, Gritty, Wet, Low Reverb

to synthetic or hybrid sounds:

Cello, Harp, Rich, Clean, Choir, Pluck, Formant Vocal, Wet, High Reverb

Because pitch is generated rather than conventionally pitch-shifted from a single root sample, the model can create natural timbral variation across registers while maintaining the identity of the prompted sound.

For the intended workflow, the Keybed model is best used with RC Stable Audio Tools, which handles the multi-note generation process and instrument assembly automatically. Generated keybeds can be exported for use as playable sampler instruments rather than remaining as isolated audio generations.

The user-facing prompt is intentionally simple. A descriptor such as:

Grand Piano, Warm, Gritty

is expanded internally into the structured Keybed conditioning grammar used by the model. RC Stable Audio Tools adds the required Keybed / sequence / note information, keeps the user descriptor stable across the instrument, reuses the same resolved seed across generation chunks, then slices and maps the resulting notes into the exported sampler.

This makes the workflow useful not only as a demo interface, but also as a reference implementation for developers interested in building their own VST, sampler, or instrument-generation front end around the model.

For the full prompt-injection and inference design, see the Keybed Training & Inference Strategy.

Foundation-1.2 Exported Keybed

Foundation-1.2 Keybed generation workflow in RC Stable Audio Tools.


What Foundation-1 Does

  • Generates musically coherent loops for production workflows
  • Generates individual note-specific one-shots
  • Generates pitch-consistent multi-note keybeds
  • Builds playable sampler instruments through the RC Stable Audio Tools workflow
  • Supports layered keybeds built from multiple independently prompted sounds
  • Understands BPM and bar count for structured loop generation
  • Locks to major and minor keys across western music theory
  • Supports enharmonic equivalents when prompting scales and keys
  • Separates instrument identity from timbral character
  • Supports timbral mixing by combining instrument and sonic descriptors
  • Responds to FX tags such as reverb, delay, distortion, and modulation
  • Uses notation-style prompt structure to encourage coherent phrasing, melodic shape, rhythmic behavior, and harmonic motion
  • Produces perfect loops within supported BPM / bar denominations
  • Understands Wet vs Dry production context — adding terms like Dry encourages minimal FX processing, while Wet or FX tags produce more processed, spatial, or effected sounds.

Why It Feels Different

Most audio models can react to broad prompt terms like “warm pad” or “bright synth.” with inconsistent results. Foundation-1 was designed to go further by treating the sound as a layered system:

  1. Instrument Family – what broad source category the sound belongs to
  2. Sub-Family – the more specific instrument role or identity
  3. Timbre Tags – the tonal, spectral, or textural character
  4. FX Tags – the processing layer applied to the sound
  5. Notation / Structure Tags – the musical behavior of the generated phrase

This layered conditioning approach is a major reason Foundation-1 is able to deliver both high musicality and high prompt control at the same time.


Audio Showcase

Loop Showcase

PromptAudio
Bass, FM Bass, Medium Delay, Medium Reverb, Low Distortion, Phaser, Sub Bass, Bass, Upper Mids, Acid, Gritty, Wide, Dubstep, Thick, Silky, Warm, Rich, Overdriven, Crisp, Deep, Clean, Pitch Bend, 303, 8 Bars, 140 BPM, E minor
Sub Bass, Bass, Gritty, Small, Square, Bass, Dark, Digital, Thick, Clean, Simple, Bassline, Epic, Choppy, Melody, 4 Bars, 150 BPM, G# minor
Flute, Pizzicato, Punchy, Present, Ambient, Nasal, Melody, Epic, Airy, Slow Speed, 8 Bars, 150 BPM, E minor
High Saw, Spacey, Lead, Warm, Silky, Smooth, 303, Synth Lead, Medium Reverb, Low Distortion, Upper Mids, Mids, Pitch Bend, Arp, 8 Bars, 140 BPM, F minor
Trumpet, Warm, Complex Arp Melody, High Reverb, Low Distortion, Smooth, Silky, Texture, 8 Bars, 130 BPM, C minor
Synth, Pad, Chord Progression, Rising, Digital, Bass, Fat, Near, Wide, Silky, Warm, Focused, 8 Bars, 110 BPM, D major
Piccolo, Flute, Airy, Music Box, plucked, complex melody, 8 Bars, 140 BPM, C# minor
Synth Lead, Wavetable Bass, Low Distortion, High Reverb, Sub Bass, Upper Mids, Acid, Gritty, Wide, Thick, Silky, Warm, Rich, Overdriven, Crisp, Clean, 303, Complex, 8 Bars, 140 BPM, F minor
Fiddle, Bowed Strings, Full, Clean, Spacey, Rich, Intimate, Thick, Rolling, Arp, Fast Speed, Complex, 8 Bars, 128 BPM, B minor
Chiptune, Chord Progression, Pulse Wave, Medium Reverb, 8 Bars, 128 BPM, D minor
Kalimba, Mallet, Medium Reverb, Overdriven, Wide, Metallic, Thick, Sparkly, Upper Mids, Bright, Airy, Alternating, Chord Progression, Atmosphere, Spacey, Fast Speed, 8 Bars, 120 BPM, B minor

One-Shot Showcase

The examples below show short, note-specific one-shot generations across acoustic, synthetic, bass, FX, and hybrid timbres.

PromptNoteAudio
FX, Sharp, Subdued, Distant, Sparkly, Round, Hit, Deep, Low Reverb, SubD#1
Synth Bass, Metallic, Rich, Punchy, Digital, Medium ReverbC2
Reese Bass, Hit, Bright, Wavetable, Noisy, Big, Thick, Airy, DeepF#2
Synth Lead, Shiny, Rich, Choir, Bright, Smooth, Short, Synthetic VoxF5
Tuba, Big, Bright, Clean, Biting, SustainedG#3
Clavinet, Wobble, Breathy, Bright, Deep, Gritty, Impact, SparklyC#4
Cello, Airy, Wide, Woody, Muffled, SmoothD#3
Reese Bass, Subdued, Round, Growl, Hit, Deep, Woody, Fat, Metallic, Medium Delay, Medium DistortionD#2
Digital Piano, Retro, Smooth, Mono, Warm, Low ReverbE3
Supersaw, Big, Smooth, Warm, Vintage, Crisp, Analog, Wide, Muffled, High ReverbG2
Synth Lead, Bright, Shiny, Soft, Warm, Punchy, Smooth, Medium DelayC#4
Sustained, Trumpet, Airy, Wide, High ReverbF3
808, Thick, Acid, Hit, FX, Rumble, Wide, Big, DistortionF1
Violin, Vintage, Thick, Chiptune, Spacey, Nasal, Medium DelayA#5
Grand Piano, Near, Snappy, Distant, Bell, Bright, Focused, DelayD3

Keybed Showcase

The following examples demonstrate generated keybeds across acoustic, synthetic, and hybrid timbres.

A full video playthrough of the Keybed Showcase is available here, which shows the associated MIDI used to audition each generated instrument.

PromptAudio
Soft, Silky, Spectral, Smooth, Spacey, Synth, Pad, Subdued, Wet, High Reverb
Pad, Focused, Buzzy, Big, Deep, Supersaw, Pulse, Dry
Pan Flute, Hollow, Dark, Smooth, Sustained, Wide, Silky, Wet, Medium Reverb
Violin, Woody, Pizzicato, Focused, Breathy, Airy, Wet, High Reverb
Grand Piano, Warm, Gritty, Wet, Low Reverb
Grand Piano, Cold, Sparkly, Wet, Low Reverb
Digital Piano, Noisy, Spiccato, Rich, Pluck, Warm, Swell, Clean, Dry
Cello, Fat, Warm, Metallic, Sharp, Near, Sustained, Dry
Church Bell, Glassy, Sparkly, Wet, High Reverb
Thick, Saw, Neuro, Reese Bass, Dry
Thick, Sine, Reese Bass, Wet, High Reverb, High Phaser
Xylophone, Sustained, Analog, Warm, Woody, Dry
Xylophone, Sustained, Analog, Warm, Woody, Bit Crushed, Wet
Cello, Harp, Rich, Clean, Choir, Pluck, Formant Vocal, Wet, High Reverb
Acid, Neuro, Synth Lead, Saw, Square, Wide, Focused, Pluck, Sustained, Wet, High Phaser
Synth Lead, Present, Sharp, Spacey, Digital, Hollow, Focused, Clean, Wet, Cross Delay, High Reverb
Dubstep, Neuro, Synth Lead, Pitch Bend, Acid, Wet, Cross Delay

Generated Waveforms

Pure waveform keybeds are shown separately to make the learned oscillator shapes easier to compare directly.

WaveformPromptAudio
TrianglePure Tone, Triangle, Dry
SquarePure Tone, Square, Dry
PulsePure Tone, Pulse, Dry
SinePure Tone, Sine, Dry
SawPure Tone, Saw, Dry

Multi-Layered Keybeds

The inference pipeline can also combine three independently generated keybeds into a single layered instrument. The examples below use separate Main and Support prompts for each layer. Further the generated instrument allows independent volume control for each layer.

Layer PromptsAudio
Main: Bright, Airy, Violin
Support 1: Thick, Present, Male, Vocal, Choir
Support 2: Digital String, Analog
Main: Marimba, Dark, Sub, Woody, Airy
Support 1: Soft, Grand Piano, Dark
Support 2: Music Box, Glassy

Core Capabilities

1. Musical Structure

Foundation-1 was trained to produce structured musical material rather than full music or generic textures. Musical Notation terms can encourage notation, chord progressions, melodies, arps, phrase direction, rhythmic density, and other musically relevant behaviors.

2. Instrument Identity

The model supports a broad instrument hierarchy spanning synths, keys, basses, bowed strings, mallets, winds, guitars, brass, vocals, and plucked strings.

3. Timbral Control

Foundation-1 is not limited to broad instrument naming. It also responds to timbral descriptors such as spectral shape, tone, width, density, texture, brightness, warmth, grit, space, and other sonic traits.

4. Timbral Mixing

Because instrument identity and timbral character were not collapsed into a single flat label, the model is especially strong at timbral hybridization and layered sonic prompting.

5. FX Prompting

The model supports a dedicated FX layer covering multiple forms of reverb, delay, distortion, phaser, and bitcrushing.

6. Loop Fidelity

Foundation-1 is built for production-ready loop generation, including BPM-aware and bar-aware structure within supported denominations.

7. One-Shot Generation

The specialized Foundation-1.2 Samples and Keybeds checkpoints generate short, note-specific sounds across acoustic, synthetic, bass, FX, and hybrid timbres. These can be used individually or as raw material for sampler instruments.

8. Pitch-Consistent Keybeds

The Keybed checkpoint was trained specifically to preserve a prompted timbral identity across changing pitch and register. RC Stable Audio Tools uses this capability to generate multiple note-specific samples and assemble them into playable instruments.


Conditioning Architecture

Foundation-1 was trained with a layered tagging hierarchy designed to improve control, composability, and prompt clarity.

Hierarchy Overview

  • Major Family → broad instrument class
  • Sub-Family → more specific instrument role
  • Timbre Tags → tonal / spectral / textural descriptors
  • FX Tags → processing layer
  • Notation Tags → musical behavior and phrasing

This makes it possible to prompt at different levels of abstraction. A user can stay broad with a family-level prompt like Synth or Keys, or get more specific with terms like Synth Lead, Wavetable Bass, Grand Piano, Violin, or Trumpet, then further shape the output using timbral and FX descriptors.


Instrument Coverage

Major Families

Foundation-1 was trained across the following major instrument families:

  • Synth
  • Keys
  • Bass
  • Bowed Strings
  • Mallet
  • Wind
  • Guitar
  • Brass
  • Vocal
  • Plucked Strings

Sub-Family Coverage

Foundation-1 includes a wide sub-family layer covering a broad range of production-relevant instrument roles, including but not limited to:

  • Synth Lead
  • Synth Bass
  • Digital Piano
  • Pluck
  • Grand Piano
  • Bell
  • Pad
  • Atmosphere
  • Digital Strings
  • FM Synth
  • Violin
  • Digital Organ
  • Supersaw
  • Wavetable Bass
  • Rhodes Piano
  • Cello
  • Texture
  • Flute
  • Reese Bass
  • Wavetable Synth
  • Electric Bass
  • Marimba
  • Trumpet
  • Pan Flute
  • Choir
  • Harp
  • Church Organ
  • Acoustic Guitar
  • Hammond Organ
  • Celesta
  • Vibraphone
  • Glockenspiel
  • Ocarina
  • Clarinet
  • French Horn
  • Tuba
  • Oboe
Sub-Family Chart

Timbre System

One of Foundation-1’s main strengths is that it was not trained to treat timbre as an afterthought. Timbral character is directly represented in the prompt system, giving users control over not only what is being generated, but also how it sounds.

Representative timbre descriptors include:

  • Warm
  • Bright
  • Wide
  • Airy
  • Thick
  • Rich
  • Tight
  • Full
  • Gritty
  • Clean
  • Retro
  • Saw
  • Crisp
  • Focused
  • Metallic
  • Chiptune
  • Dark
  • 303
  • Shiny
  • Analog
  • Present
  • Sparkly
  • Ambient
  • Soft
  • Smooth
  • Cold
  • Buzzy
  • Deep
  • Formant Vocal
  • Round
  • Punchy
  • Nasal
  • Vintage
  • Growl
  • Breathy
  • Glassy
  • Noisy
  • Synthetic Vox
  • Supersaw
  • Bitcrushed
  • Dreamy
Timbre Chart

Why This Matters

This tagging design makes prompts much more flexible. Instead of only asking for an instrument, users can shape:

  • tonal balance
  • brightness / darkness
  • width / intimacy
  • clean vs driven character
  • synthetic vs organic feel
  • transient sharpness
  • texture and density
  • spatial character

This is especially useful for producers who want to guide the output toward a specific role in a mix rather than just a generic instrument label.

For a list of used tags please see the Tag Reference Sheet.


FX Layer

Foundation-1 includes a dedicated FX descriptor layer spanning multiple common production effects.

Representative FX tags include:

  • Low Reverb
  • Medium Reverb
  • High Reverb
  • Plate Reverb
  • Low Delay
  • Medium Delay
  • High Delay
  • Ping Pong Delay
  • Stereo Delay
  • Cross Delay
  • Mono Delay
  • Low Distortion
  • Medium Distortion
  • High Distortion
  • Phaser
  • Low Phaser
  • Medium Phaser
  • High Phaser
  • Bitcrush
  • High Bitcrush
FX Chart

Musical Notation and Structure

Foundation-1 was trained with structured musical descriptors designed to improve phrase coherence, rhythmic intent, melodic motion, and prompt control.

These notation-style prompt terms help steer:

  • chord progressions
  • melodies
  • top-line layers
  • arpeggios
  • phrase direction
  • rhythmic density
  • harmonic feel
  • subdivision style
  • simple vs complex motion
  • sustained vs plucked behavior
  • melodic contour and pacing

Examples of supported structural ideas may include terms such as:

  • chord progression
  • melody
  • top melody
  • arp
  • triplets
  • simple
  • complex
  • rising
  • falling
  • strummed
  • sustained
  • catchy
  • epic
  • slow
  • fast

This notation layer is one of the main reasons Foundation-1 produces unusually coherent musical material instead of static or loosely related phrases. These can be mixed and matched as desired.


Tonal and Timing Support

Foundation-1 is designed for structured music production workflows and supports:

Keys and Modes

  • Major keys
  • Minor keys
  • Enharmonic equivalents
  • Western 12-tone chromatic prompting

Loop Structure

  • Supported bar lengths: 4 Bars, 8 Bars
  • Supported BPM denominations: 100 BPM, 110 BPM, 120 BPM, 128 BPM, 130 BPM, 140 BPM, 150 BPM

Prompt Structure

For best results, use rich prompts built around the model’s tags. These tags can be mixed and matched as needed. The model was trained on a structured hierarchy designed to encourage musically coherent sample generation.

Layered Prompt Structure

[Instrument Family / Sub-Family], [Timbre], [Musical Behavior / Notation], [FX], [Key], [Bars], [BPM]

Prompting Notes

  • Start with a clear instrument identity
  • Add 1–3 timbre descriptors for stronger steering
  • Include a notation or musical structure term for better phrase coherence
  • Always include Bars and BPM, which define the musical loop length
  • Ensure the generation duration matches the requested musical structure
  • The RC Stable Audio Fork automatically handles this timing alignment

Use FX and timbre tags sparingly at first, then layer more once you understand the model’s behavior.


One Prompt → Multiple Outputs

Each row below uses the exact same prompt, but a different random seed.
The timbre tags remain unchanged, so the overall sound character stays consistent while the melodic and musical content varies between generations.

PromptOutput AOutput BOutput C
Bass, FM Bass, Medium Delay, Medium Reverb, Low Distortion, Phaser, Acid, Gritty, Wide, Dubstep, Thick, Silky, Warm, Rich, Overdriven, Crisp, Deep, Clean, Triplets, 8 Bars, 150 BPM, A minor
Gritty, Acid, Bassline, 303, Synth Lead, FM, Sub, Upper Mids, High Phaser, High Reverb, Pitch Bend, 8 Bars, 140 BPM, E minor
Kalimba, Mallet, Medium Reverb, Overdriven, Wide, Metallic, Thick, Sparkly, Upper Mids, Bright, Airy, Small, Alternating Chord Progression, Atmosphere, Spacey, Fast, 4 Bars, 120 BPM, B minor

Foundation-1 is best used with RC Stable Audio Tools, which is tuned around the model family, its metadata, and its structured prompting system.

RC Stable Audio Tools (Enhanced Fork)

The interface supports different workflows depending on the checkpoint in use.

Loop Generation

For the original Foundation-1 loop checkpoint, the interface provides:

  • structured prompt building aligned with the training tags
  • random prompt generation
  • automatic BPM / bar timing alignment
  • automatic MIDI extraction from generated audio
  • generation settings tuned for Foundation-1

Samples / One-Shot Generation

The specialized Samples workflow supports:

  • note-specific generation
  • instrument and timbre prompting
  • short sample-focused inference
  • rapid auditioning of generated sounds

The Keybed checkpoint can also generate individual one-shots effectively. The dedicated Samples checkpoint is provided because its earlier training endpoint preserves stronger general sample-generation quality.

Keybed / Text-to-Synth Generation

The Keybed workflow is designed to be used in conjunction with RC Stable Audio Tools rather than as a sequence of manually generated independent notes.

The interface handles the pitch-aware generation pipeline, creates the source files required for the instrument, and can export completed keybeds for sampler use.

The workflow supports:

  • automatic generation across keyboard registers
  • pitch-specific text-conditioned samples
  • playable keybed construction
  • DecentSampler export
  • SFZ export
  • per-instrument ADSR controls
  • built-in sampler effects and tone shaping
  • layered keybeds using independently prompted Main / Support layers
  • independent layer volume control
  • generated instrument previews

Foundation-1.2 Keybed exporter

Generated keybeds can be assembled and exported as playable DecentSampler or SFZ instruments.

RC Stable Audio Tools (Enhanced Fork)

Stable Audio Tools (Original Repository)

Model Files

Foundation-1 is now distributed as multiple specialized checkpoints built around the same underlying architecture and conditioning system.

The specialized Foundation-1.2 Samples and Keybeds checkpoints retain the same Foundation-1 / SAO architecture. They therefore do not represent separate model architectures; the specialization comes from the training objective and selected checkpoint.

The release uses 16-bit model weights to reduce the model footprint without changing the intended inference quality.

  • Foundation_1.safetensors — loop-generation checkpoint
  • Foundation-1.2-Samples.safetensors — sample-focused checkpoint
  • Foundation-1.2-Keybeds.safetensors — keybed-focused checkpoint
  • model_config.json — shared model configuration

Basic Setup for RC Stable Audio Tools

  1. Create a subfolder inside your models directory
  2. Place the desired checkpoint and its compatible config in that folder
  3. Launch the interface
  4. Select the checkpoint from the model selector
  5. Choose the matching Loop, One-Shot, or Keybed workflow
  6. Prompt with layered instrument and timbre descriptors

For full keybeds, the RC interface is strongly recommended because it manages the multi-note inference and sampler-export process automatically.

Hardware Requirements

Foundation-1 is designed to run locally on modern GPUs.

Typical VRAM usage during generation is approximately ~7 GB.
For reliable operation, a GPU with at least 8 GB of VRAM is recommended.

Generation Performance

Generation speed will vary depending on GPU model, generation mode, sample length, and system configuration.

On an RTX 3090, a standard individual generation is approximately ~7–8 seconds per sample. Full single-layer keybed builds take approximately 50 seconds.


Dataset and Training Philosophy

Foundation-1 was built around a structured sample-generation philosophy, rather than generic or genre-based audio captioning. The dataset consists entirely of hand-crafted and labeled audio, produced through a controlled augmentation pipeline.

At a high level, the training design emphasizes:

  • structured musical loops
  • instrument hierarchy
  • explicit timbre representation
  • dedicated FX descriptors
  • notation-aware prompt terms
  • strong production relevance
  • broad reuse for compositional workflows

This design is central to the model’s musical coherence and high degree of sonic control.

For more details on the original dataset and training methodology, see the Training & Dataset Notes.

Specialized Samples and Keybeds Training

The Foundation-1.2 Samples and Keybeds checkpoints use the same underlying Foundation-1 architecture, but were trained with a lower learning rate of 1e-5.

CheckpointTraining EndpointSpecialization
Foundation-1.2 SamplesEpoch 0 / Step 560General sample / one-shot generation
Foundation-1.2 KeybedsEpoch 3 / Step 8120Cross-pitch timbral consistency and playable keybed generation

Additional training configuration:

  • Optimizer: AdamW
  • Learning Rate: 1e-5
  • Weight Decay: 1e-3
  • Scheduler: InverseLR
  • EMA: Enabled
  • Sample Rate: 44,100 Hz
  • Channels: Stereo
  • Training Window: 882,000 samples (~20 seconds)

The longer Keybed run was selected to reinforce cross-note and cross-register consistency. The objective was not simply to improve isolated note quality, but to teach the model how a single sound identity should behave as pitch changes across an instrument.

The Keybed checkpoint used a dedicated multi-note training strategy rather than treating every pitch as an unrelated example. The generation pipeline mirrors that structure by injecting note-sequence conditioning, reusing a common seed across keybed chunks, slicing generated sequences back into individual notes, and packaging the resulting samples into playable instruments.

For a detailed explanation of both the training method and the inference/export pipeline, see the Keybed Training & Inference Strategy.


Limitations

Foundation-1 is a specialized model family for producer-facing sample and instrument generation, not a general-purpose full-song generator.

Important notes:

  • It performs best when prompted using vocabulary aligned with the training design
  • It is optimized for sample-generation workflows, not open-ended genre captioning
  • Only two genre tags were included (Dubstep Growls and Chiptune waveforms), primarily to reinforce waveform behaviors
  • Prompt quality matters — structured layered prompts outperform vague natural language
  • Some timbre tags exert stronger influence than others
  • Certain tag combinations may require iteration to achieve the exact musical role or timbral blend desired
  • Percussion and drum sounds are outside the scope of this release

The model is also optimized around specific timing relationships between Bars, BPM, and generation duration.

For example:

  • an 8-bar loop at 100 BPM ≈ 19 seconds

If the generation duration is shorter than the musical structure implied by the prompt (for example requesting an 8-bar loop but generating only 5 seconds), the model may produce less coherent musical phrases.

The RC Stable Audio Fork automatically handles this timing alignment, making this workflow much easier.

Keybed-Specific Limitations

The Keybed model is designed around practical instrument ranges.

Some prompts simply stop making semantic or perceptual sense at extreme registers. For example, a sub-bass at C7 is no longer functioning as a bass, even if the model preserves some aspects of the original waveform or timbral character.

The same applies to many synthetic and acoustic sounds. As pitch rises into extreme upper registers, complex waveforms often become perceptually simpler and can collapse toward high-pitched pure-tone-like behavior. At the opposite end, very low pitches can become increasingly difficult to render consistently, and small pitch or waveform errors become more obvious.

For this reason, RC Stable Audio Tools clamps generation and export ranges to practical regions rather than forcing every generated sound across the full MIDI note range.

Prompt choice also matters. A timbre that has a natural low, mid, or high-register identity may become ambiguous when pushed several octaves outside that range. Some apparent "timbre drift" is therefore not only a model limitation; it is also a consequence of asking for a sound whose defining characteristics change as pitch moves far beyond where that sound is normally perceived.

In my internal testing, roughly 90% of generated instruments maintain a consistent relationship between pitch and timbral identity across the supported keybed ranges. This is a qualitative estimate rather than a programmatic benchmark, because there is no reliable automated metric for determining whether two notes share the same perceived timbre.

The remaining edge cases are most likely to appear with:

  • extreme low or high registers
  • unconventional bass prompts outside bass ranges
  • strongly formant-dependent sounds
  • heavily processed or hybrid timbres
  • prompts whose intended instrument identity becomes ambiguous at the requested pitch

The default export ranges were chosen as a practical compromise between keyboard coverage, pitch accuracy, and timbral consistency.


License

This model is licensed under the Stability AI Community License. It is available for non-commercial use or limited commercial use by entities with annual revenues below USD $1M. For revenues exceeding USD $1M, please refer to the repository license file for full terms.


Companion Videos

Foundation-1 Overview

The original Foundation-1 video covering the model's design philosophy, structured prompting system, timbral control, and loop-generation workflow.

🎥 Watch the Foundation-1 Overview

Foundation-1.2 Samples & Keybeds Update

A companion video covering the new Samples and Keybeds checkpoints, the updated training strategy and general journey to getting this made can be found here.

🎥 Watch the Foundation-1.2 Update

Keybed Guided Demo

A hands-on walkthrough showing the Keybed model in use, cross-pitch timbral behavior, and the new three-layer instrument exporter.

🎥 Watch the Guided Keybed Demo



Final Notes

Foundation-1 is intended as a producer-facing model family for structured sample and instrument generation, designed to augment music production.

Its goal is to let users explore sound in new ways while retaining precise control over:

  • what the sound is
  • how it behaves musically
  • how it changes across pitch
  • how it sits tonally
  • how it feels sonically
  • how it fits into a production workflow

The original loop model focuses on structured musical material. The specialized Foundation-1.2 Samples and Keybeds checkpoints extend that same conditioning system toward sound design and playable instrument creation.

That combination of musical structure, instrument identity, timbral control, loop fidelity, and cross-pitch instrument generation is what defines the Foundation-1 family.

Contributors

RoyalCities

48 commits