On-device RAG chat for Android using Gemma 3 270M + .NET MAUI + ONNX Runtime
C#
4
3 commits
updated Jun 22, 2026
A fully on-device RAG (Retrieval-Augmented Generation) chat app built with .NET MAUI and Gemma 3. Pick a PDF, ask questions, get streamed answers — no internet required after the first model download, no API keys, no cloud.
┌──────────────────────────────────────────────────────────────┐
│ MAUI Android App │
│ │
│ SetupPage ──navigate──► TabBar │
│ (download models, ├─ SummaryPage │
│ pick + index PDF) └─ ChatPage │
│ │
│ AppSession ← shared state (embedder, generator, store, │
│ document chunks, chat messages) — no DI │
│ container, passed to each page's constructor │
└───────────────────────────────┬───────────────────────────────┘
│ uses
┌───────────────────────────────▼───────────────────────────────┐
│ RagCore │
│ │
│ ModelDownloader ← downloads models from HF, resumable │
│ PdfTextExtractor ← Syncfusion PDF → page strings │
│ TextChunker ← pages → overlapping chunks │
│ MiniLmEmbedder ← chunk/question → 384-dim vector │
│ SqliteChunkRepository ← cache index by PDF hash │
│ InMemoryVectorStore ← cosine similarity search │
│ GemmaAnswerGenerator ← GenerateAsync (chat Q&A) and │
│ SummarizeAsync (document summary) — │
│ separate prompt templates, shared │
│ token-streaming/safety-net plumbing │
└─────────────────────────────────────────────────────────────────┘
PDF
└─ PdfTextExtractor ──► pages
└─ TextChunker ──────► chunks
└─ MiniLmEmbedder ► vectors ──► SqliteChunkRepository (cache)
│
Question │ load on next run
└─ MiniLmEmbedder ──► vector │
└─ InMemoryVectorStore.Search ◄────────────────┘
└─ top 3 chunks + question
└─ GemmaAnswerGenerator.GenerateAsync ──► streamed tokens
Document chunks (page order, not similarity-filtered)
└─ capped at ~3000 words (empirically the largest safe single-call
budget on-device — larger prompts caused a hard, unlogged crash)
└─ GemmaAnswerGenerator.SummarizeAsync ──► streamed tokens
Generated once per indexed PDF and cached on AppSession.Summary — revisiting the Summary tab redisplays the cached text instead of regenerating.
Maui-gemma-3/
├── patches/ ← builder patches (see below)
│ ├── gemma_builder_rope.patch
│ ├── base_builder_rope_theta.patch
│ └── apply.sh
└── src/
├── RagCore/ ← shared class library
│ ├── Chunking/
│ │ ├── TextChunk.cs ← data model: page number + text
│ │ └── TextChunker.cs ← splits pages into overlapping chunks
│ ├── Embedding/
│ │ └── MiniLmEmbedder.cs ← ONNX inference → 384-dim vector
│ ├── Generation/
│ │ └── GemmaAnswerGenerator.cs ← GenerateAsync (chat) + SummarizeAsync
│ │ (summary), separate prompts, shared
│ │ streaming/"Answer:"-strip/fallback logic
│ ├── Ingestion/
│ │ └── PdfTextExtractor.cs ← Syncfusion PDF → page strings
│ ├── Persistence/
│ │ ├── IChunkRepository.cs
│ │ └── SqliteChunkRepository.cs ← SQLite cache keyed by PDF SHA256
│ ├── Retrieval/
│ │ ├── InMemoryVectorStore.cs ← cosine similarity search in RAM
│ │ ├── ScoredChunk.cs
│ │ └── VectorMath.cs
│ └── Services/
│ └── ModelDownloader.cs ← resumable download from HuggingFace
└── MauiApp/ ← Android app
├── MauiProgram.cs
├── App.xaml(.cs) ← creates the one AppSession, passes to AppShell
├── AppShell.xaml(.cs) ← non-tab "setup" route + TabBar(Summary, Ask AI)
├── AppSession.cs ← shared state (embedder/generator/store/chunks/
│ messages/summary) — no DI container, passed
│ directly to each page's constructor
├── SetupPage.xaml(.cs) ← model download/load + PDF picker + indexing
├── SummaryPage.xaml(.cs) ← triggers SummarizeAsync once, caches result
├── ChatPage.xaml(.cs) ← Q&A chat UI, CollectionView of ChatMessage
├── ChatMessage.cs ← INotifyPropertyChanged streaming bubble model
├── MarkdownText.cs ← renders **bold** as FormattedString spans,
│ strips ~~strikethrough~~ markers
└── Platforms/Android/
| Model | Purpose | Size | Source |
|---|---|---|---|
| Gemma 3 270M-it (int4 ONNX) | Text generation | ~864 MB | ihassantariq/gemma-3-270m-it-onnx-int4 |
| all-MiniLM-L6-v2 (ONNX) | Sentence embeddings | ~90 MB | sentence-transformers/all-MiniLM-L6-v2 |
Both are downloaded automatically on first launch. Downloads are resumable — if interrupted, the app picks up from where it left off.
Google's Gemma 3 comes in two architectures:
Gemma3ForCausalLM (text-only). Model type "gemma3_text" in genai_config.Gemma3ForConditionalGeneration (multimodal, text + vision). Model type "gemma3".Our pinned onnxruntime-genai 0.8.3 runtime loads "gemma3" models with the full vision pipeline, adding a ~645MB vision component even when you never send images. The 270M produces a "gemma3_text" bundle — no vision overhead, ~864MB total vs ~6.2GB.
<PackageReference Include="Microsoft.ML.OnnxRuntimeGenAI" Version="0.8.3" />
<PackageReference Include="Microsoft.ML.OnnxRuntime" Version="1.22.0" />
These two must be pinned together. The GenAI execution provider selection breaks with mismatched versions.
The google/gemma-3-270m-it HuggingFace repo ships PyTorch weights (.safetensors), not ONNX. We use Microsoft's onnxruntime-genai Python model builder to convert.
/path/to/python3.10 -m venv /tmp/genai-builder-venv310
source /tmp/genai-builder-venv310/bin/activate
pip install onnxruntime-genai==0.11.4 onnxruntime transformers torch accelerate onnx-ir huggingface_hub
We pin onnxruntime-genai==0.11.4 for the builder because:
"type": "gemma3_text" for Gemma3ForCausalLM architectureshf auth login --token YOUR_HF_WRITE_TOKEN --force
transformers 5.12+ reorganized Gemma 3's RoPE (Rotary Position Embedding) config from direct attributes into a nested dict. The builder at v0.11.4 predates this change. The patches add hasattr guards so the builder handles both old and new transformers versions.
bash patches/apply.sh
What each patch does:
patches/gemma_builder_rope.patch — fixes gemma.py line ~127:
# Before (breaks with transformers 5.12+):
self.rope_local_theta = config.rope_local_base_freq
# After:
if hasattr(config, "rope_local_base_freq"):
self.rope_local_theta = config.rope_local_base_freq
elif hasattr(config, "rope_parameters") and isinstance(config.rope_parameters, dict):
self.rope_local_theta = config.rope_parameters.get("sliding_attention", {}).get("rope_theta", 10000.0)
else:
self.rope_local_theta = 10000.0
patches/base_builder_rope_theta.patch — fixes base.py rope_theta fallback:
# Before:
rope_theta = config.rope_theta if hasattr(config, "rope_theta") else \
config.rope_embedding_base if hasattr(config, "rope_embedding_base") else 10000
# After (also checks rope_parameters dict):
rope_theta = (
config.rope_theta if hasattr(config, "rope_theta")
else config.rope_embedding_base if hasattr(config, "rope_embedding_base")
else config.rope_parameters.get("full_attention", {}).get("rope_theta", 10000)
if hasattr(config, "rope_parameters") and isinstance(config.rope_parameters, dict)
else 10000
)
Relevant builder source (v0.11.4):
python -m onnxruntime_genai.models.builder \
-m google/gemma-3-270m-it \
-o ~/.cache/maui-gemma-3/models/gemma-3-270m-it \
-p int4 -e cpu
| Flag | Value | Meaning |
|---|---|---|
-m | google/gemma-3-270m-it | HuggingFace model ID to download and convert |
-o | output path | Where to write ONNX files |
-p | int4 | 4-bit integer quantization — smallest size, fastest inference |
-e | cpu | CPU execution provider (no GPU required) |
Output files:
gemma-3-270m-it/
├── genai_config.json ← runtime config for onnxruntime-genai
├── chat_template.jinja ← Gemma prompt format template
├── tokenizer.json ← SentencePiece vocabulary
├── tokenizer_config.json ← tokenizer settings
├── model.onnx ← ONNX graph structure (~300KB)
└── model.onnx.data ← quantized weights (~864MB, LFS)
The builder emits "temperature": null. The 0.8.3 runtime requires a number:
sed -i '' 's/"temperature": null/"temperature": 1.0/' \
~/.cache/maui-gemma-3/models/gemma-3-270m-it/genai_config.json
The app's ModelDownloader.PatchGenAiConfigTemperature() applies this automatically at runtime.
hf repos create gemma-3-270m-it-onnx-int4 --type model
hf upload <your-username>/gemma-3-270m-it-onnx-int4 \
~/.cache/maui-gemma-3/models/gemma-3-270m-it .
Then update GemmaBaseUrl in src/RagCore/Services/ModelDownloader.cs:
private const string GemmaBaseUrl =
"https://huggingface.co/<your-username>/gemma-3-270m-it-onnx-int4/resolve/main/";
dotnet build -t:Run -f net10.0-android src/MauiApp/MauiApp.csproj
On first launch the app downloads both models (~950MB total). Downloads are resumable.
C#
96.0%
Shell
4.0%
On-device RAG chat for Android using Gemma 3 270M + .NET MAUI + ONNX Runtime
C#
4
3 commits
updated Jun 22, 2026
A fully on-device RAG (Retrieval-Augmented Generation) chat app built with .NET MAUI and Gemma 3. Pick a PDF, ask questions, get streamed answers — no internet required after the first model download, no API keys, no cloud.
┌──────────────────────────────────────────────────────────────┐
│ MAUI Android App │
│ │
│ SetupPage ──navigate──► TabBar │
│ (download models, ├─ SummaryPage │
│ pick + index PDF) └─ ChatPage │
│ │
│ AppSession ← shared state (embedder, generator, store, │
│ document chunks, chat messages) — no DI │
│ container, passed to each page's constructor │
└───────────────────────────────┬───────────────────────────────┘
│ uses
┌───────────────────────────────▼───────────────────────────────┐
│ RagCore │
│ │
│ ModelDownloader ← downloads models from HF, resumable │
│ PdfTextExtractor ← Syncfusion PDF → page strings │
│ TextChunker ← pages → overlapping chunks │
│ MiniLmEmbedder ← chunk/question → 384-dim vector │
│ SqliteChunkRepository ← cache index by PDF hash │
│ InMemoryVectorStore ← cosine similarity search │
│ GemmaAnswerGenerator ← GenerateAsync (chat Q&A) and │
│ SummarizeAsync (document summary) — │
│ separate prompt templates, shared │
│ token-streaming/safety-net plumbing │
└─────────────────────────────────────────────────────────────────┘
PDF
└─ PdfTextExtractor ──► pages
└─ TextChunker ──────► chunks
└─ MiniLmEmbedder ► vectors ──► SqliteChunkRepository (cache)
│
Question │ load on next run
└─ MiniLmEmbedder ──► vector │
└─ InMemoryVectorStore.Search ◄────────────────┘
└─ top 3 chunks + question
└─ GemmaAnswerGenerator.GenerateAsync ──► streamed tokens
Document chunks (page order, not similarity-filtered)
└─ capped at ~3000 words (empirically the largest safe single-call
budget on-device — larger prompts caused a hard, unlogged crash)
└─ GemmaAnswerGenerator.SummarizeAsync ──► streamed tokens
Generated once per indexed PDF and cached on AppSession.Summary — revisiting the Summary tab redisplays the cached text instead of regenerating.
Maui-gemma-3/
├── patches/ ← builder patches (see below)
│ ├── gemma_builder_rope.patch
│ ├── base_builder_rope_theta.patch
│ └── apply.sh
└── src/
├── RagCore/ ← shared class library
│ ├── Chunking/
│ │ ├── TextChunk.cs ← data model: page number + text
│ │ └── TextChunker.cs ← splits pages into overlapping chunks
│ ├── Embedding/
│ │ └── MiniLmEmbedder.cs ← ONNX inference → 384-dim vector
│ ├── Generation/
│ │ └── GemmaAnswerGenerator.cs ← GenerateAsync (chat) + SummarizeAsync
│ │ (summary), separate prompts, shared
│ │ streaming/"Answer:"-strip/fallback logic
│ ├── Ingestion/
│ │ └── PdfTextExtractor.cs ← Syncfusion PDF → page strings
│ ├── Persistence/
│ │ ├── IChunkRepository.cs
│ │ └── SqliteChunkRepository.cs ← SQLite cache keyed by PDF SHA256
│ ├── Retrieval/
│ │ ├── InMemoryVectorStore.cs ← cosine similarity search in RAM
│ │ ├── ScoredChunk.cs
│ │ └── VectorMath.cs
│ └── Services/
│ └── ModelDownloader.cs ← resumable download from HuggingFace
└── MauiApp/ ← Android app
├── MauiProgram.cs
├── App.xaml(.cs) ← creates the one AppSession, passes to AppShell
├── AppShell.xaml(.cs) ← non-tab "setup" route + TabBar(Summary, Ask AI)
├── AppSession.cs ← shared state (embedder/generator/store/chunks/
│ messages/summary) — no DI container, passed
│ directly to each page's constructor
├── SetupPage.xaml(.cs) ← model download/load + PDF picker + indexing
├── SummaryPage.xaml(.cs) ← triggers SummarizeAsync once, caches result
├── ChatPage.xaml(.cs) ← Q&A chat UI, CollectionView of ChatMessage
├── ChatMessage.cs ← INotifyPropertyChanged streaming bubble model
├── MarkdownText.cs ← renders **bold** as FormattedString spans,
│ strips ~~strikethrough~~ markers
└── Platforms/Android/
| Model | Purpose | Size | Source |
|---|---|---|---|
| Gemma 3 270M-it (int4 ONNX) | Text generation | ~864 MB | ihassantariq/gemma-3-270m-it-onnx-int4 |
| all-MiniLM-L6-v2 (ONNX) | Sentence embeddings | ~90 MB | sentence-transformers/all-MiniLM-L6-v2 |
Both are downloaded automatically on first launch. Downloads are resumable — if interrupted, the app picks up from where it left off.
Google's Gemma 3 comes in two architectures:
Gemma3ForCausalLM (text-only). Model type "gemma3_text" in genai_config.Gemma3ForConditionalGeneration (multimodal, text + vision). Model type "gemma3".Our pinned onnxruntime-genai 0.8.3 runtime loads "gemma3" models with the full vision pipeline, adding a ~645MB vision component even when you never send images. The 270M produces a "gemma3_text" bundle — no vision overhead, ~864MB total vs ~6.2GB.
<PackageReference Include="Microsoft.ML.OnnxRuntimeGenAI" Version="0.8.3" />
<PackageReference Include="Microsoft.ML.OnnxRuntime" Version="1.22.0" />
These two must be pinned together. The GenAI execution provider selection breaks with mismatched versions.
The google/gemma-3-270m-it HuggingFace repo ships PyTorch weights (.safetensors), not ONNX. We use Microsoft's onnxruntime-genai Python model builder to convert.
/path/to/python3.10 -m venv /tmp/genai-builder-venv310
source /tmp/genai-builder-venv310/bin/activate
pip install onnxruntime-genai==0.11.4 onnxruntime transformers torch accelerate onnx-ir huggingface_hub
We pin onnxruntime-genai==0.11.4 for the builder because:
"type": "gemma3_text" for Gemma3ForCausalLM architectureshf auth login --token YOUR_HF_WRITE_TOKEN --force
transformers 5.12+ reorganized Gemma 3's RoPE (Rotary Position Embedding) config from direct attributes into a nested dict. The builder at v0.11.4 predates this change. The patches add hasattr guards so the builder handles both old and new transformers versions.
bash patches/apply.sh
What each patch does:
patches/gemma_builder_rope.patch — fixes gemma.py line ~127:
# Before (breaks with transformers 5.12+):
self.rope_local_theta = config.rope_local_base_freq
# After:
if hasattr(config, "rope_local_base_freq"):
self.rope_local_theta = config.rope_local_base_freq
elif hasattr(config, "rope_parameters") and isinstance(config.rope_parameters, dict):
self.rope_local_theta = config.rope_parameters.get("sliding_attention", {}).get("rope_theta", 10000.0)
else:
self.rope_local_theta = 10000.0
patches/base_builder_rope_theta.patch — fixes base.py rope_theta fallback:
# Before:
rope_theta = config.rope_theta if hasattr(config, "rope_theta") else \
config.rope_embedding_base if hasattr(config, "rope_embedding_base") else 10000
# After (also checks rope_parameters dict):
rope_theta = (
config.rope_theta if hasattr(config, "rope_theta")
else config.rope_embedding_base if hasattr(config, "rope_embedding_base")
else config.rope_parameters.get("full_attention", {}).get("rope_theta", 10000)
if hasattr(config, "rope_parameters") and isinstance(config.rope_parameters, dict)
else 10000
)
Relevant builder source (v0.11.4):
python -m onnxruntime_genai.models.builder \
-m google/gemma-3-270m-it \
-o ~/.cache/maui-gemma-3/models/gemma-3-270m-it \
-p int4 -e cpu
| Flag | Value | Meaning |
|---|---|---|
-m | google/gemma-3-270m-it | HuggingFace model ID to download and convert |
-o | output path | Where to write ONNX files |
-p | int4 | 4-bit integer quantization — smallest size, fastest inference |
-e | cpu | CPU execution provider (no GPU required) |
Output files:
gemma-3-270m-it/
├── genai_config.json ← runtime config for onnxruntime-genai
├── chat_template.jinja ← Gemma prompt format template
├── tokenizer.json ← SentencePiece vocabulary
├── tokenizer_config.json ← tokenizer settings
├── model.onnx ← ONNX graph structure (~300KB)
└── model.onnx.data ← quantized weights (~864MB, LFS)
The builder emits "temperature": null. The 0.8.3 runtime requires a number:
sed -i '' 's/"temperature": null/"temperature": 1.0/' \
~/.cache/maui-gemma-3/models/gemma-3-270m-it/genai_config.json
The app's ModelDownloader.PatchGenAiConfigTemperature() applies this automatically at runtime.
hf repos create gemma-3-270m-it-onnx-int4 --type model
hf upload <your-username>/gemma-3-270m-it-onnx-int4 \
~/.cache/maui-gemma-3/models/gemma-3-270m-it .
Then update GemmaBaseUrl in src/RagCore/Services/ModelDownloader.cs:
private const string GemmaBaseUrl =
"https://huggingface.co/<your-username>/gemma-3-270m-it-onnx-int4/resolve/main/";
dotnet build -t:Run -f net10.0-android src/MauiApp/MauiApp.csproj
On first launch the app downloads both models (~950MB total). Downloads are resumable.
C#
96.0%
Shell
4.0%