luisquintanilla/model-packages-prototype

Prototype: AI models as small NuGet packages with on-demand fetch, cache, and verify

C#

1

35 commits

updated Sep 3, 2026

See the code

README

Model Packages Prototype

Open in GitHub Codespaces

What if AI models shipped like NuGet packages — small metadata packages that fetch, cache, and verify large model binaries on demand?

The Problem

AI models are large. A typical embedding model is 80–300 MB; LLMs run into the gigabytes. Shipping these inside NuGet packages creates real problems:

  • Bloated restores — every dotnet restore re-downloads hundreds of MB
  • Storage costs — NuGet feeds aren't designed to host large binaries at scale
  • Version control pain — accidentally committing a model binary to git is a common mistake
  • No flexibility — once a model is baked into a .nupkg, you can't redirect consumers to a corporate mirror or a local cache

But .NET developers expect things to just work. dotnet add package SomeModel should give you a working model — no manual downloads, no hunting for URLs, no SHA256 verification by hand.

The Idea

This prototype explores a different approach: model packages contain only code and metadata (~few KB). The heavy model binary is fetched on first use, cached locally, and verified against a SHA256 hash. Think of it like how NuGet itself works — you don't ship source code in a package, you ship compiled artifacts that restore from feeds.

The key insight: just as nuget.config lets you redirect package sources (nuget.org → corporate feed → local folder), a model-sources.json lets you redirect model sources (HuggingFace → corporate mirror → air-gapped local path) — without changing any application code.

Quick Start

Click the badge above, wait for the container to build, then:

# Run the ONNX track — downloads model from HuggingFace on first run (~86 MB)
dotnet run --project samples/SampleConsumer.Onnx

# Run the .mlnet track — uses a pre-built pipeline
dotnet run --project samples/SampleConsumer.MLNet

# Try other task types
dotnet run --project samples/SampleConsumer.Classification
dotnet run --project samples/SampleConsumer.NER
dotnet run --project samples/SampleConsumer.QA
dotnet run --project samples/SampleConsumer.Reranking
dotnet run --project samples/SampleConsumer.TextGeneration

# Audio samples (require a test.wav file in the consumer directory)
dotnet run --project samples/SampleConsumer.WhisperTiny
dotnet run --project samples/SampleConsumer.SileroVad
dotnet run --project samples/SampleConsumer.AstAudioSet
dotnet run --project samples/SampleConsumer.ClapEmbedding
dotnet run --project samples/SampleConsumer.SpeechT5Tts

# Image samples (require a test image; models downloaded on first run)
dotnet run --project samples/SampleConsumer.ImageClassification
dotnet run --project samples/SampleConsumer.ObjectDetection
dotnet run --project samples/SampleConsumer.ImageSegmentation
dotnet run --project samples/SampleConsumer.DepthEstimation
dotnet run --project samples/SampleConsumer.ImageEmbedding
dotnet run --project samples/SampleConsumer.ZeroShotClassification
dotnet run --project samples/SampleConsumer.ImageCaptioning
dotnet run --project samples/SampleConsumer.VisualQA
dotnet run --project samples/SampleConsumer.SegmentAnything
dotnet run --project samples/SampleConsumer.TextToImage

Locally

Prerequisites: .NET 10 SDK

git clone https://github.com/luisquintanilla/model-packages-prototype.git
cd model-packages-prototype
dotnet build
dotnet run --project samples/SampleConsumer.Onnx

Both samples generate text embeddings using all-MiniLM-L6-v2, compute cosine similarities, and display results — proving the full pipeline works end-to-end.

Supported Task Types

The prototype includes sample model packages and consumers for every major ML/AI inference task:

TaskModel PackageConsumerModel
Embedding (ONNX)SampleModelPackage.OnnxSampleConsumer.Onnxall-MiniLM-L6-v2
Embedding (.mlnet)SampleModelPackage.MLNetSampleConsumer.MLNetall-MiniLM-L6-v2
Embedding (BGE)SampleModelPackage.BgeEmbeddingSampleConsumer.BgeEmbeddingBGE-small-en-v1.5
Embedding (E5)SampleModelPackage.E5EmbeddingSampleConsumer.E5EmbeddingE5-small-v2
Embedding (GTE)SampleModelPackage.GteEmbeddingSampleConsumer.GteEmbeddingGTE-small
ClassificationSampleModelPackage.ClassificationSampleConsumer.ClassificationDistilBERT SST-2
Named Entity RecognitionSampleModelPackage.NERSampleConsumer.NERBERT-base NER
Question AnsweringSampleModelPackage.QASampleConsumer.QAMiniLM-Squad2
RerankingSampleModelPackage.RerankingSampleConsumer.RerankingMS MARCO MiniLM
Text Generation (local)SampleModelPackage.TextGenerationSampleConsumer.TextGenerationPhi-3-mini
Text Generation (MEAI)—SampleConsumer.TextGenerationMeaiAny IChatClient provider
Audio Embedding (CLAP)SampleModelPackage.ClapEmbeddingSampleConsumer.ClapEmbeddingCLAP HTSAT-unfused
Audio ClassificationSampleModelPackage.AstAudioSetSampleConsumer.AstAudioSetAST AudioSet
Voice Activity DetectionSampleModelPackage.SileroVadSampleConsumer.SileroVadSilero VAD v4
Speech-to-Text (Tiny)SampleModelPackage.WhisperTinySampleConsumer.WhisperTinyWhisper Tiny
Speech-to-Text (Base)SampleModelPackage.WhisperBaseSampleConsumer.WhisperBaseWhisper Base
Text-to-SpeechSampleModelPackage.SpeechT5TtsSampleConsumer.SpeechT5TtsSpeechT5 TTS
Image ClassificationSampleModelPackage.ImageClassificationSampleConsumer.ImageClassificationViT-Base-Patch16-224
Object DetectionSampleModelPackage.ObjectDetectionSampleConsumer.ObjectDetectionYOLOv8s
Image SegmentationSampleModelPackage.ImageSegmentationSampleConsumer.ImageSegmentationSegFormer-B0 ADE-512
Depth EstimationSampleModelPackage.DepthEstimationSampleConsumer.DepthEstimationDPT-Hybrid-Midas
Image EmbeddingSampleModelPackage.ImageEmbeddingSampleConsumer.ImageEmbeddingCLIP ViT-Base-Patch32
Zero-Shot ClassificationSampleModelPackage.ZeroShotClassificationSampleConsumer.ZeroShotClassificationCLIP ViT-Base-Patch32
Image CaptioningSampleModelPackage.ImageCaptioningSampleConsumer.ImageCaptioningGIT-Base-COCO
Visual QASampleModelPackage.VisualQASampleConsumer.VisualQAGIT-Base-TextVQA
Segment AnythingSampleModelPackage.SegmentAnythingSampleConsumer.SegmentAnythingSAM2-Hiera-Tiny
Text-to-ImageSampleModelPackage.TextToImageSampleConsumer.TextToImageStable Diffusion v1.4

Each model package embeds a manifest and small assets (vocabs, label maps) while large model binaries are fetched on demand through the Core SDK.

Project Map

model-packages-prototype/
│
├── src/
│   ├── ModelPackages/                       ← Core SDK: fetch, cache, verify (format-agnostic)
│   └── ModelPackages.Tool/                  ← CLI tool: prefetch, verify, info, clear-cache
│
├── samples/
│   │  ── Embeddings ──────────────────────────────────────────────────
│   ├── SampleModelPackage.Onnx/             ← MiniLM embedding (raw ONNX from HuggingFace)
│   ├── SampleConsumer.Onnx/                 ← Consumer: cosine similarity demo
│   ├── SampleModelPackage.MLNet/            ← MiniLM embedding (pre-built .mlnet pipeline)
│   ├── SampleConsumer.MLNet/                ← Consumer: same API, different packaging
│   ├── SampleModelPackage.BgeEmbedding/     ← BGE-small-en-v1.5 (query prefix baked in)
│   ├── SampleConsumer.BgeEmbedding/         ← Consumer: asymmetric retrieval demo
│   ├── SampleModelPackage.E5Embedding/      ← E5-small-v2 (dual query/passage prefix)
│   ├── SampleConsumer.E5Embedding/          ← Consumer: dual-prefix retrieval demo
│   ├── SampleModelPackage.GteEmbedding/     ← GTE-small (no prefix needed)
│   ├── SampleConsumer.GteEmbedding/         ← Consumer: semantic search demo
│   │  ── Classification ──────────────────────────────────────────────
│   ├── SampleModelPackage.Classification/   ← DistilBERT sentiment analysis
│   ├── SampleConsumer.Classification/       ← Consumer: classify text sentiment
│   │  ── Named Entity Recognition ────────────────────────────────────
│   ├── SampleModelPackage.NER/              ← BERT-base NER (person, org, location)
│   ├── SampleConsumer.NER/                  ← Consumer: extract named entities
│   │  ── Question Answering ──────────────────────────────────────────
│   ├── SampleModelPackage.QA/               ← MiniLM-Squad2 extractive QA
│   ├── SampleConsumer.QA/                   ← Consumer: answer questions from context
│   │  ── Reranking ───────────────────────────────────────────────────
│   ├── SampleModelPackage.Reranking/        ← MS MARCO MiniLM cross-encoder
│   ├── SampleConsumer.Reranking/            ← Consumer: rerank search results
│   │  ── Text Generation ─────────────────────────────────────────────
│   ├── SampleModelPackage.TextGeneration/   ← Phi-3-mini local ONNX GenAI
│   ├── SampleConsumer.TextGeneration/       ← Consumer: local text generation
│   ├── SampleConsumer.TextGenerationMeai/   ← Consumer: provider-agnostic IChatClient
│   │  ── Audio ────────────────────────────────────────────────────
│   ├── SampleModelPackage.ClapEmbedding/    ← CLAP audio embedding (ONNX)
│   ├── SampleConsumer.ClapEmbedding/        ← Consumer: cosine similarity demo
│   ├── SampleModelPackage.AstAudioSet/      ← AST AudioSet classification (527 labels)
│   ├── SampleConsumer.AstAudioSet/          ← Consumer: classify audio events
│   ├── SampleModelPackage.SileroVad/        ← Silero VAD v4 voice activity detection
│   ├── SampleConsumer.SileroVad/            ← Consumer: detect speech segments
│   ├── SampleModelPackage.WhisperTiny/      ← Whisper Tiny speech-to-text (ONNX)
│   ├── SampleConsumer.WhisperTiny/          ← Consumer: transcribe audio
│   ├── SampleModelPackage.WhisperBase/      ← Whisper Base speech-to-text (ONNX)
│   ├── SampleConsumer.WhisperBase/          ← Consumer: transcribe audio
│   ├── SampleModelPackage.SpeechT5Tts/     ← SpeechT5 text-to-speech (5 ONNX files)
│   ├── SampleConsumer.SpeechT5Tts/         ← Consumer: synthesize speech
│   │  ── Image ────────────────────────────────────────────────────
│   ├── SampleModelPackage.ImageClassification/ ← ViT image classification (ImageNet)
│   ├── SampleConsumer.ImageClassification/     ← Consumer: classify images
│   ├── SampleModelPackage.ObjectDetection/     ← YOLOv8s object detection
│   ├── SampleConsumer.ObjectDetection/         ← Consumer: detect objects in images
│   ├── SampleModelPackage.ImageSegmentation/   ← SegFormer semantic segmentation
│   ├── SampleConsumer.ImageSegmentation/       ← Consumer: segment image pixels
│   ├── SampleModelPackage.DepthEstimation/     ← DPT monocular depth estimation
│   ├── SampleConsumer.DepthEstimation/         ← Consumer: estimate depth maps
│   ├── SampleModelPackage.ImageEmbedding/      ← CLIP image embedding (MEAI)
│   ├── SampleConsumer.ImageEmbedding/          ← Consumer: image similarity
│   ├── SampleModelPackage.ZeroShotClassification/ ← CLIP zero-shot classification
│   ├── SampleConsumer.ZeroShotClassification/     ← Consumer: classify with text labels
│   ├── SampleModelPackage.ImageCaptioning/     ← GIT-Base image captioning (MEAI)
│   ├── SampleConsumer.ImageCaptioning/         ← Consumer: generate captions
│   ├── SampleModelPackage.VisualQA/            ← GIT-Base visual question answering
│   ├── SampleConsumer.VisualQA/                ← Consumer: answer questions about images
│   ├── SampleModelPackage.SegmentAnything/     ← SAM2 segment anything
│   ├── SampleConsumer.SegmentAnything/         ← Consumer: point/box prompted segmentation
│   ├── SampleModelPackage.TextToImage/         ← Stable Diffusion text-to-image
│   └── SampleConsumer.TextToImage/             ← Consumer: generate images from text
│
├── tools/
│   └── PrepareMLNetModel/                   ← Helper to build .mlnet from raw ONNX + vocab
│
└── docs/                                    ← Architecture, design decisions, guides

How It Works

The 4-Layer Architecture

┌─────────────────────────────────────────────────────────┐
│  Layer 4: Consumer App                                  │
│  dotnet add package SampleModelPackage.Onnx             │
│  var gen = await MiniLMModel.CreateEmbeddingGenerator() │
│  var emb = await gen.GenerateAsync(texts)               │
├─────────────────────────────────────────────────────────┤
│  Layer 3: Model Package (authored by model publisher)   │
│  Embeds: model-manifest.json + vocab.txt                │
│  Exposes: MiniLMModel.CreateEmbeddingGeneratorAsync()   │
├─────────────────────────────────────────────────────────┤
│  Layer 2: Inference Library (NuGet packages)            │
│  MLNet.TextInference.Onnx — embeddings, classification,  │
│    NER, QA, reranking via ML.NET + ONNX Runtime           │
│  MLNet.AudioInference.Onnx — audio classification,        │
│    embedding, VAD, TTS, speech-to-text                     │
│  MLNet.ImageInference.Onnx — image classification,        │
│    detection, segmentation, depth, captioning, VQA         │
│  MLNet.ImageGeneration.OnnxGenAI — text-to-image           │
│  MLNet.TextGeneration.OnnxGenAI — local text generation   │
│  IEmbeddingGenerator, IChatClient (MEAI abstractions)     │
├─────────────────────────────────────────────────────────┤
│  Layer 1: Core SDK (ModelPackages)                      │
│  Resolve source → Check cache → Download → SHA256 verify│
│  Named sources, atomic writes, lock files               │
└─────────────────────────────────────────────────────────┘
  1. Core SDK — Knows nothing about ONNX or ML.NET. It fetches files, caches them, and verifies integrity. Works with any model format.
  2. Inference Library — Knows nothing about downloading. It builds ML.NET pipelines and wraps them as IEmbeddingGenerator<string, Embedding<float>> (Microsoft.Extensions.AI).
  3. Model Package — The glue. A model author creates this as a NuGet package. It embeds a manifest (what to download, from where, expected SHA256) and small assets (tokenizer vocabulary). It references both the Core SDK and the inference library.
  4. Consumer — Just installs the model package. One method call to get a working embedding generator. No knowledge of model sources, caching, or verification.

Two Packaging Tracks

Track 1: Raw ONNXTrack 2: Pre-built .mlnet
What's downloadedRaw .onnx file from HuggingFacePre-built .mlnet pipeline zip
Pipeline builtOn consumer's machine (Fit)By model author (ahead of time)
First-run costDownload + Fit (~5s)Download only
FlexibilityConsumer can customize pipelineFixed pipeline
File size86 MB (ONNX only)79 MB (ONNX + vocab + config in zip)

The consumer code is nearly identical for both tracks — that's the proof the abstraction works.

Named Sources (the nuget.config analogy)

Just as nuget.config redirects package feeds, model-sources.json redirects model sources:

{
  "sources": {
    "company-mirror": {
      "type": "mirror",
      "endpoint": "https://models.internal.company.com"
    }
  },
  "defaultSource": "company-mirror"
}

Set MODELPACKAGES_SOURCE=company-mirror or drop a model-sources.json next to your .csproj — no code changes needed.

CLI Tool

# Pre-download a model (e.g., in CI/CD)
dotnet run --project src/ModelPackages.Tool -- prefetch --manifest path/to/model-manifest.json

# Verify cached model integrity
dotnet run --project src/ModelPackages.Tool -- verify --manifest path/to/model-manifest.json

# Show resolved source and cache path
dotnet run --project src/ModelPackages.Tool -- info --manifest path/to/model-manifest.json

Documentation

  • Architecture — Deep dive into the 4-layer design, manifest format, cache layout, source resolution
  • Design Decisions — The "why" behind every major choice
  • Authoring Guide — How to create your own model package
  • CLI Reference — All commands, options, exit codes, CI/CD patterns

Using the Packages

The Core SDK and CLI tool are published to GitHub Packages. Preview builds are published on every push to main.

1. Add the GitHub Packages feed

GitHub Packages requires authentication even for public repos. Create a personal access token with read:packages scope, then:

dotnet nuget add source https://nuget.pkg.github.com/luisquintanilla/index.json \
  --name model-packages \
  --username YOUR_GITHUB_USERNAME \
  --password YOUR_PAT

2. Install the Core SDK (for model package authors)

dotnet add package ModelPackages --prerelease

3. Install the CLI tool

dotnet tool install -g ModelPackages.Tool --prerelease \
  --add-source https://nuget.pkg.github.com/luisquintanilla/index.json

Alternative: Download from GitHub Releases

Release .nupkg files are also attached to GitHub Releases. Download and use with a local NuGet source:

dotnet nuget add source /path/to/downloaded/packages --name local
dotnet add package ModelPackages

Technology

  • .NET 10 (preview)
  • ML.NET — Pipeline construction, ONNX Runtime integration
  • Microsoft.Extensions.AI — IEmbeddingGenerator<string, Embedding<float>> and IChatClient abstractions
  • OnnxRuntime — Model inference (embeddings, classification, NER, QA, reranking)
  • ONNX Runtime GenAI — Local text generation (Phi-3-mini)
  • MLNet.AudioInference.Onnx — Audio inference: classification, embedding, VAD, TTS, speech-to-text
  • MLNet.Audio.Core — Audio primitives: AudioData, AudioIO, MelSpectrogramExtractor
  • MLNet.Audio.Tokenizers — Audio tokenizers: SpeechT5 char tokenizer, Whisper tokenizer
  • MLNet.ImageInference.Onnx — Image inference: classification, detection, segmentation, depth, embedding, captioning, VQA, zero-shot, SAM2
  • MLNet.ImageGeneration.OnnxGenAI — Image generation: Stable Diffusion text-to-image
  • MLNet.Image.Core — Image primitives: BoundingBox, DepthMap, SegmentationMask, preprocessors
  • System.Text.Json — Source-generated serialization for manifests

Status

This is a prototype / proof of concept exploring the design space. Not production-ready. See the Roadmap for planned improvements.

License

MIT

luisquintanilla/model-packages-prototype

Prototype: AI models as small NuGet packages with on-demand fetch, cache, and verify

C#

1

35 commits

updated Sep 3, 2026

See the code

README

Model Packages Prototype

Open in GitHub Codespaces

What if AI models shipped like NuGet packages — small metadata packages that fetch, cache, and verify large model binaries on demand?

The Problem

AI models are large. A typical embedding model is 80–300 MB; LLMs run into the gigabytes. Shipping these inside NuGet packages creates real problems:

  • Bloated restores — every dotnet restore re-downloads hundreds of MB
  • Storage costs — NuGet feeds aren't designed to host large binaries at scale
  • Version control pain — accidentally committing a model binary to git is a common mistake
  • No flexibility — once a model is baked into a .nupkg, you can't redirect consumers to a corporate mirror or a local cache

But .NET developers expect things to just work. dotnet add package SomeModel should give you a working model — no manual downloads, no hunting for URLs, no SHA256 verification by hand.

The Idea

This prototype explores a different approach: model packages contain only code and metadata (~few KB). The heavy model binary is fetched on first use, cached locally, and verified against a SHA256 hash. Think of it like how NuGet itself works — you don't ship source code in a package, you ship compiled artifacts that restore from feeds.

The key insight: just as nuget.config lets you redirect package sources (nuget.org → corporate feed → local folder), a model-sources.json lets you redirect model sources (HuggingFace → corporate mirror → air-gapped local path) — without changing any application code.

Quick Start

Click the badge above, wait for the container to build, then:

# Run the ONNX track — downloads model from HuggingFace on first run (~86 MB)
dotnet run --project samples/SampleConsumer.Onnx

# Run the .mlnet track — uses a pre-built pipeline
dotnet run --project samples/SampleConsumer.MLNet

# Try other task types
dotnet run --project samples/SampleConsumer.Classification
dotnet run --project samples/SampleConsumer.NER
dotnet run --project samples/SampleConsumer.QA
dotnet run --project samples/SampleConsumer.Reranking
dotnet run --project samples/SampleConsumer.TextGeneration

# Audio samples (require a test.wav file in the consumer directory)
dotnet run --project samples/SampleConsumer.WhisperTiny
dotnet run --project samples/SampleConsumer.SileroVad
dotnet run --project samples/SampleConsumer.AstAudioSet
dotnet run --project samples/SampleConsumer.ClapEmbedding
dotnet run --project samples/SampleConsumer.SpeechT5Tts

# Image samples (require a test image; models downloaded on first run)
dotnet run --project samples/SampleConsumer.ImageClassification
dotnet run --project samples/SampleConsumer.ObjectDetection
dotnet run --project samples/SampleConsumer.ImageSegmentation
dotnet run --project samples/SampleConsumer.DepthEstimation
dotnet run --project samples/SampleConsumer.ImageEmbedding
dotnet run --project samples/SampleConsumer.ZeroShotClassification
dotnet run --project samples/SampleConsumer.ImageCaptioning
dotnet run --project samples/SampleConsumer.VisualQA
dotnet run --project samples/SampleConsumer.SegmentAnything
dotnet run --project samples/SampleConsumer.TextToImage

Locally

Prerequisites: .NET 10 SDK

git clone https://github.com/luisquintanilla/model-packages-prototype.git
cd model-packages-prototype
dotnet build
dotnet run --project samples/SampleConsumer.Onnx

Both samples generate text embeddings using all-MiniLM-L6-v2, compute cosine similarities, and display results — proving the full pipeline works end-to-end.

Supported Task Types

The prototype includes sample model packages and consumers for every major ML/AI inference task:

TaskModel PackageConsumerModel
Embedding (ONNX)SampleModelPackage.OnnxSampleConsumer.Onnxall-MiniLM-L6-v2
Embedding (.mlnet)SampleModelPackage.MLNetSampleConsumer.MLNetall-MiniLM-L6-v2
Embedding (BGE)SampleModelPackage.BgeEmbeddingSampleConsumer.BgeEmbeddingBGE-small-en-v1.5
Embedding (E5)SampleModelPackage.E5EmbeddingSampleConsumer.E5EmbeddingE5-small-v2
Embedding (GTE)SampleModelPackage.GteEmbeddingSampleConsumer.GteEmbeddingGTE-small
ClassificationSampleModelPackage.ClassificationSampleConsumer.ClassificationDistilBERT SST-2
Named Entity RecognitionSampleModelPackage.NERSampleConsumer.NERBERT-base NER
Question AnsweringSampleModelPackage.QASampleConsumer.QAMiniLM-Squad2
RerankingSampleModelPackage.RerankingSampleConsumer.RerankingMS MARCO MiniLM
Text Generation (local)SampleModelPackage.TextGenerationSampleConsumer.TextGenerationPhi-3-mini
Text Generation (MEAI)—SampleConsumer.TextGenerationMeaiAny IChatClient provider
Audio Embedding (CLAP)SampleModelPackage.ClapEmbeddingSampleConsumer.ClapEmbeddingCLAP HTSAT-unfused
Audio ClassificationSampleModelPackage.AstAudioSetSampleConsumer.AstAudioSetAST AudioSet
Voice Activity DetectionSampleModelPackage.SileroVadSampleConsumer.SileroVadSilero VAD v4
Speech-to-Text (Tiny)SampleModelPackage.WhisperTinySampleConsumer.WhisperTinyWhisper Tiny
Speech-to-Text (Base)SampleModelPackage.WhisperBaseSampleConsumer.WhisperBaseWhisper Base
Text-to-SpeechSampleModelPackage.SpeechT5TtsSampleConsumer.SpeechT5TtsSpeechT5 TTS
Image ClassificationSampleModelPackage.ImageClassificationSampleConsumer.ImageClassificationViT-Base-Patch16-224
Object DetectionSampleModelPackage.ObjectDetectionSampleConsumer.ObjectDetectionYOLOv8s
Image SegmentationSampleModelPackage.ImageSegmentationSampleConsumer.ImageSegmentationSegFormer-B0 ADE-512
Depth EstimationSampleModelPackage.DepthEstimationSampleConsumer.DepthEstimationDPT-Hybrid-Midas
Image EmbeddingSampleModelPackage.ImageEmbeddingSampleConsumer.ImageEmbeddingCLIP ViT-Base-Patch32
Zero-Shot ClassificationSampleModelPackage.ZeroShotClassificationSampleConsumer.ZeroShotClassificationCLIP ViT-Base-Patch32
Image CaptioningSampleModelPackage.ImageCaptioningSampleConsumer.ImageCaptioningGIT-Base-COCO
Visual QASampleModelPackage.VisualQASampleConsumer.VisualQAGIT-Base-TextVQA
Segment AnythingSampleModelPackage.SegmentAnythingSampleConsumer.SegmentAnythingSAM2-Hiera-Tiny
Text-to-ImageSampleModelPackage.TextToImageSampleConsumer.TextToImageStable Diffusion v1.4

Each model package embeds a manifest and small assets (vocabs, label maps) while large model binaries are fetched on demand through the Core SDK.

Project Map

model-packages-prototype/
│
├── src/
│   ├── ModelPackages/                       ← Core SDK: fetch, cache, verify (format-agnostic)
│   └── ModelPackages.Tool/                  ← CLI tool: prefetch, verify, info, clear-cache
│
├── samples/
│   │  ── Embeddings ──────────────────────────────────────────────────
│   ├── SampleModelPackage.Onnx/             ← MiniLM embedding (raw ONNX from HuggingFace)
│   ├── SampleConsumer.Onnx/                 ← Consumer: cosine similarity demo
│   ├── SampleModelPackage.MLNet/            ← MiniLM embedding (pre-built .mlnet pipeline)
│   ├── SampleConsumer.MLNet/                ← Consumer: same API, different packaging
│   ├── SampleModelPackage.BgeEmbedding/     ← BGE-small-en-v1.5 (query prefix baked in)
│   ├── SampleConsumer.BgeEmbedding/         ← Consumer: asymmetric retrieval demo
│   ├── SampleModelPackage.E5Embedding/      ← E5-small-v2 (dual query/passage prefix)
│   ├── SampleConsumer.E5Embedding/          ← Consumer: dual-prefix retrieval demo
│   ├── SampleModelPackage.GteEmbedding/     ← GTE-small (no prefix needed)
│   ├── SampleConsumer.GteEmbedding/         ← Consumer: semantic search demo
│   │  ── Classification ──────────────────────────────────────────────
│   ├── SampleModelPackage.Classification/   ← DistilBERT sentiment analysis
│   ├── SampleConsumer.Classification/       ← Consumer: classify text sentiment
│   │  ── Named Entity Recognition ────────────────────────────────────
│   ├── SampleModelPackage.NER/              ← BERT-base NER (person, org, location)
│   ├── SampleConsumer.NER/                  ← Consumer: extract named entities
│   │  ── Question Answering ──────────────────────────────────────────
│   ├── SampleModelPackage.QA/               ← MiniLM-Squad2 extractive QA
│   ├── SampleConsumer.QA/                   ← Consumer: answer questions from context
│   │  ── Reranking ───────────────────────────────────────────────────
│   ├── SampleModelPackage.Reranking/        ← MS MARCO MiniLM cross-encoder
│   ├── SampleConsumer.Reranking/            ← Consumer: rerank search results
│   │  ── Text Generation ─────────────────────────────────────────────
│   ├── SampleModelPackage.TextGeneration/   ← Phi-3-mini local ONNX GenAI
│   ├── SampleConsumer.TextGeneration/       ← Consumer: local text generation
│   ├── SampleConsumer.TextGenerationMeai/   ← Consumer: provider-agnostic IChatClient
│   │  ── Audio ────────────────────────────────────────────────────
│   ├── SampleModelPackage.ClapEmbedding/    ← CLAP audio embedding (ONNX)
│   ├── SampleConsumer.ClapEmbedding/        ← Consumer: cosine similarity demo
│   ├── SampleModelPackage.AstAudioSet/      ← AST AudioSet classification (527 labels)
│   ├── SampleConsumer.AstAudioSet/          ← Consumer: classify audio events
│   ├── SampleModelPackage.SileroVad/        ← Silero VAD v4 voice activity detection
│   ├── SampleConsumer.SileroVad/            ← Consumer: detect speech segments
│   ├── SampleModelPackage.WhisperTiny/      ← Whisper Tiny speech-to-text (ONNX)
│   ├── SampleConsumer.WhisperTiny/          ← Consumer: transcribe audio
│   ├── SampleModelPackage.WhisperBase/      ← Whisper Base speech-to-text (ONNX)
│   ├── SampleConsumer.WhisperBase/          ← Consumer: transcribe audio
│   ├── SampleModelPackage.SpeechT5Tts/     ← SpeechT5 text-to-speech (5 ONNX files)
│   ├── SampleConsumer.SpeechT5Tts/         ← Consumer: synthesize speech
│   │  ── Image ────────────────────────────────────────────────────
│   ├── SampleModelPackage.ImageClassification/ ← ViT image classification (ImageNet)
│   ├── SampleConsumer.ImageClassification/     ← Consumer: classify images
│   ├── SampleModelPackage.ObjectDetection/     ← YOLOv8s object detection
│   ├── SampleConsumer.ObjectDetection/         ← Consumer: detect objects in images
│   ├── SampleModelPackage.ImageSegmentation/   ← SegFormer semantic segmentation
│   ├── SampleConsumer.ImageSegmentation/       ← Consumer: segment image pixels
│   ├── SampleModelPackage.DepthEstimation/     ← DPT monocular depth estimation
│   ├── SampleConsumer.DepthEstimation/         ← Consumer: estimate depth maps
│   ├── SampleModelPackage.ImageEmbedding/      ← CLIP image embedding (MEAI)
│   ├── SampleConsumer.ImageEmbedding/          ← Consumer: image similarity
│   ├── SampleModelPackage.ZeroShotClassification/ ← CLIP zero-shot classification
│   ├── SampleConsumer.ZeroShotClassification/     ← Consumer: classify with text labels
│   ├── SampleModelPackage.ImageCaptioning/     ← GIT-Base image captioning (MEAI)
│   ├── SampleConsumer.ImageCaptioning/         ← Consumer: generate captions
│   ├── SampleModelPackage.VisualQA/            ← GIT-Base visual question answering
│   ├── SampleConsumer.VisualQA/                ← Consumer: answer questions about images
│   ├── SampleModelPackage.SegmentAnything/     ← SAM2 segment anything
│   ├── SampleConsumer.SegmentAnything/         ← Consumer: point/box prompted segmentation
│   ├── SampleModelPackage.TextToImage/         ← Stable Diffusion text-to-image
│   └── SampleConsumer.TextToImage/             ← Consumer: generate images from text
│
├── tools/
│   └── PrepareMLNetModel/                   ← Helper to build .mlnet from raw ONNX + vocab
│
└── docs/                                    ← Architecture, design decisions, guides

How It Works

The 4-Layer Architecture

┌─────────────────────────────────────────────────────────┐
│  Layer 4: Consumer App                                  │
│  dotnet add package SampleModelPackage.Onnx             │
│  var gen = await MiniLMModel.CreateEmbeddingGenerator() │
│  var emb = await gen.GenerateAsync(texts)               │
├─────────────────────────────────────────────────────────┤
│  Layer 3: Model Package (authored by model publisher)   │
│  Embeds: model-manifest.json + vocab.txt                │
│  Exposes: MiniLMModel.CreateEmbeddingGeneratorAsync()   │
├─────────────────────────────────────────────────────────┤
│  Layer 2: Inference Library (NuGet packages)            │
│  MLNet.TextInference.Onnx — embeddings, classification,  │
│    NER, QA, reranking via ML.NET + ONNX Runtime           │
│  MLNet.AudioInference.Onnx — audio classification,        │
│    embedding, VAD, TTS, speech-to-text                     │
│  MLNet.ImageInference.Onnx — image classification,        │
│    detection, segmentation, depth, captioning, VQA         │
│  MLNet.ImageGeneration.OnnxGenAI — text-to-image           │
│  MLNet.TextGeneration.OnnxGenAI — local text generation   │
│  IEmbeddingGenerator, IChatClient (MEAI abstractions)     │
├─────────────────────────────────────────────────────────┤
│  Layer 1: Core SDK (ModelPackages)                      │
│  Resolve source → Check cache → Download → SHA256 verify│
│  Named sources, atomic writes, lock files               │
└─────────────────────────────────────────────────────────┘
  1. Core SDK — Knows nothing about ONNX or ML.NET. It fetches files, caches them, and verifies integrity. Works with any model format.
  2. Inference Library — Knows nothing about downloading. It builds ML.NET pipelines and wraps them as IEmbeddingGenerator<string, Embedding<float>> (Microsoft.Extensions.AI).
  3. Model Package — The glue. A model author creates this as a NuGet package. It embeds a manifest (what to download, from where, expected SHA256) and small assets (tokenizer vocabulary). It references both the Core SDK and the inference library.
  4. Consumer — Just installs the model package. One method call to get a working embedding generator. No knowledge of model sources, caching, or verification.

Two Packaging Tracks

Track 1: Raw ONNXTrack 2: Pre-built .mlnet
What's downloadedRaw .onnx file from HuggingFacePre-built .mlnet pipeline zip
Pipeline builtOn consumer's machine (Fit)By model author (ahead of time)
First-run costDownload + Fit (~5s)Download only
FlexibilityConsumer can customize pipelineFixed pipeline
File size86 MB (ONNX only)79 MB (ONNX + vocab + config in zip)

The consumer code is nearly identical for both tracks — that's the proof the abstraction works.

Named Sources (the nuget.config analogy)

Just as nuget.config redirects package feeds, model-sources.json redirects model sources:

{
  "sources": {
    "company-mirror": {
      "type": "mirror",
      "endpoint": "https://models.internal.company.com"
    }
  },
  "defaultSource": "company-mirror"
}

Set MODELPACKAGES_SOURCE=company-mirror or drop a model-sources.json next to your .csproj — no code changes needed.

CLI Tool

# Pre-download a model (e.g., in CI/CD)
dotnet run --project src/ModelPackages.Tool -- prefetch --manifest path/to/model-manifest.json

# Verify cached model integrity
dotnet run --project src/ModelPackages.Tool -- verify --manifest path/to/model-manifest.json

# Show resolved source and cache path
dotnet run --project src/ModelPackages.Tool -- info --manifest path/to/model-manifest.json

Documentation

  • Architecture — Deep dive into the 4-layer design, manifest format, cache layout, source resolution
  • Design Decisions — The "why" behind every major choice
  • Authoring Guide — How to create your own model package
  • CLI Reference — All commands, options, exit codes, CI/CD patterns

Using the Packages

The Core SDK and CLI tool are published to GitHub Packages. Preview builds are published on every push to main.

1. Add the GitHub Packages feed

GitHub Packages requires authentication even for public repos. Create a personal access token with read:packages scope, then:

dotnet nuget add source https://nuget.pkg.github.com/luisquintanilla/index.json \
  --name model-packages \
  --username YOUR_GITHUB_USERNAME \
  --password YOUR_PAT

2. Install the Core SDK (for model package authors)

dotnet add package ModelPackages --prerelease

3. Install the CLI tool

dotnet tool install -g ModelPackages.Tool --prerelease \
  --add-source https://nuget.pkg.github.com/luisquintanilla/index.json

Alternative: Download from GitHub Releases

Release .nupkg files are also attached to GitHub Releases. Download and use with a local NuGet source:

dotnet nuget add source /path/to/downloaded/packages --name local
dotnet add package ModelPackages

Technology

  • .NET 10 (preview)
  • ML.NET — Pipeline construction, ONNX Runtime integration
  • Microsoft.Extensions.AI — IEmbeddingGenerator<string, Embedding<float>> and IChatClient abstractions
  • OnnxRuntime — Model inference (embeddings, classification, NER, QA, reranking)
  • ONNX Runtime GenAI — Local text generation (Phi-3-mini)
  • MLNet.AudioInference.Onnx — Audio inference: classification, embedding, VAD, TTS, speech-to-text
  • MLNet.Audio.Core — Audio primitives: AudioData, AudioIO, MelSpectrogramExtractor
  • MLNet.Audio.Tokenizers — Audio tokenizers: SpeechT5 char tokenizer, Whisper tokenizer
  • MLNet.ImageInference.Onnx — Image inference: classification, detection, segmentation, depth, embedding, captioning, VQA, zero-shot, SAM2
  • MLNet.ImageGeneration.OnnxGenAI — Image generation: Stable Diffusion text-to-image
  • MLNet.Image.Core — Image primitives: BoundingBox, DepthMap, SegmentationMask, preprocessors
  • System.Text.Json — Source-generated serialization for manifests

Status

This is a prototype / proof of concept exploring the design space. Not production-ready. See the Roadmap for planned improvements.

License

MIT

Languages

C#

100.0%