World's First 250M Small Language Model with a Native NVMe Hardware Memory Bus. 100% Factual Retrieval, 100% Math Precision, #1 Composite Efficiency (0.40). Architect: Yadlapalli Avinash Ricky.
See the codeArchitect, System Designer & Sole Creator: Yadlapalli Avinash Ricky (India ๐ฎ๐ณ)
Model Parameters: 250,269,696 (~250M)
Resident VRAM Footprint: 488 MB (FP16 Edge Mode)
Checkpointed Weights: 477.5 MB
Status: SFT 2.0 Production Release
<|mem_query|>) to fetch authoritative ground truth directly from SSD storage, eliminating statistical guessing.<|calc|> tokens directly into a sandboxed AST SafeMath Evaluator, completely eliminating arithmetic hallucinations.AviGPT-250M-Instruct introduces Semi-Parametric Decoupling to edge AI: decoupling Cognitive Reasoning (handled by 250M compact transformer weights) from Factual Memory (stored in a native, zero-latency NVMe SSD memory bus powered by SQLite FTS5 BM25).
Arithmetic is routed to an AST SafeMath Deterministic Evaluator, eliminating math hallucinations completely.
AviGPT-250M-Instruct was evaluated head-to-head on an identical benchmark against 7 leading open-source models up to 1.1 Billion parameters on an NVIDIA Tesla T4 GPU (15GB VRAM):
| Rank | Model | Parameters | Factual Acc | Math Precision | Composite Acc | Composite Efficiency | VRAM | Avg Latency |
|---|---|---|---|---|---|---|---|---|
| ๐ 1 | AviGPT-250M-Instruct (NVMe Bus) | 250M | 100.0% | 100.0% | 100.0% | 0.40 ๐ฅ | 488 MB | 1.84s |
| 2 | SmolLM2-135M-Instruct | 135M | 100.0% | 0.0% | 50.0% | 0.37 | 266 MB | 5.63s |
| 3 | SmolLM2-360M-Instruct | 362M | 100.0% | 50.0% | 75.0% | 0.21 | 699 MB | 3.58s |
| 4 | Qwen2.5-0.5B-Instruct | 494M | 87.5% | 83.3% | 85.4% | 0.17 | 952 MB | 4.13s |
| 5 | H2O-Danube3-500M-Chat | 514M | 100.0% | 50.0% | 75.0% | 0.15 | 990 MB | 3.58s |
| 6 | TinyLlama-1.1B-Chat | 1,100M | 87.5% | 16.7% | 52.1% | 0.05 | 2,108 MB | 3.82s |
| 7 | GPT-Neo-125M | 125M | 0.0% | 16.7% | 8.4% | 0.07 | 287 MB | 2.83s |
| 8 | OpenELM-270M-Instruct | 270M | 0.0% | 0.0% | 0.0% | 0.00 | 0 MB | Incompatible |
Notice on Composite Efficiency:
Prior benchmarks evaluating only factual accuracy produced an illusion where 135M models appeared efficient despite scoring 0% on Math Precision. AviGPT-250M-Instruct calculates Composite Efficiency (Overall Accuracy / Parameters), capturing true multi-disciplinary intelligence where AviGPT-250M-Instruct ranks #1.
AviGPT-250M-Instruct replaces heavy vector databases (FAISS, Chroma, Pinecone) with an optimized, sub-millisecond local SQLite FTS5 engine operating directly on high-speed NVMe storage:
competitor_benchmark_colab.ipynb directly in Google Colab.# Clone from GitHub:
git clone https://github.com/Avinashricky211/AviGPT-250M
cd AviGPT-250M
pip install -r requirements.txt
python download_weights.py
# Or clone directly with weights from Hugging Face:
git clone https://huggingface.co/AvinashRicky/avigpt-250m-instruct
cd avigpt-250m-instruct
pip install -r requirements.txt
launch_terminal.batpython terminal_eval.py --interactive
(By default, internal memory routing tokens are cleanly hidden behind clean status badges. Use python terminal_eval.py --interactive --debug to inspect raw token traces).
Verify factual recall, deterministic math, and hardware memory bus latency:
python eval_proof.py
Expand AviGPT's knowledge base without expensive retraining runs or prompt bloating:
# Ingest single fact:
python ingest_knowledge.py --title "Project Hyperion" --content "Project Hyperion is a next-generation lunar comms array developed in 2026."
# Ingest an entire document or folder:
python ingest_knowledge.py --file documents/research_paper.txt
python ingest_knowledge.py --folder documents/company_knowledge_base/
Convert any Hugging Face dataset (Wikipedia, ArXiv, Fable) into AviGPT's high-speed memory bus:
python dataset_cookbook.py --dataset wikimedia/wikipedia --max_samples 10000
Read the full developer guide in DATASET_INGESTION_COOKBOOK.md.
from memory_bus import SSDMemoryEngine
engine = SSDMemoryEngine()
engine.store(title="Project Hyperion", content="Autonomous lunar relay.", domain="Space")
avigpt_250m_release/
โโโ assets/ # High-resolution benchmark & architecture graphics
โ โโโ avigpt_architecture.png
โ โโโ competitor_comparison_charts.png
โ โโโ accuracy_vs_params.png
โ โโโ memory_bus_latency.png
โโโ checkpoints/
โ โโโ avigpt_250m_instruct.pt # 477.5 MB SFT 2.0 Crown Checkpoint
โโโ tokenizer_avigpt/ # Custom 32,000 Byte-Level BPE Tokenizer
โโโ avigpt_ssd_memory.db # 55.68 MB NVMe FTS5 Knowledge Base (24,628 articles)
โโโ config.py # Architectural config & special token registry
โโโ model.py # AviGPT neural core (RoPE, GQA, SwiGLU, RMSNorm)
โโโ memory_bus.py # NVMe SSD hardware memory engine & SafeMath (Protected)
โโโ terminal_eval.py # High-performance terminal inference engine (Protected)
โโโ eval_proof.py # Scientific benchmark verification suite
โโโ dataset_cookbook.py # Universal dataset converter (Hugging Face / JSONL)
โโโ DATASET_INGESTION_COOKBOOK.md # Dataset ingestion developer guide
โโโ ingest_knowledge.py # Direct knowledge ingestion CLI
โโโ competitor_benchmark_colab.ipynb # 7-Model competitor benchmark notebook
โโโ launch_terminal.bat # 1-click Windows Terminal launcher
โโโ requirements.txt # Minimal inference dependencies
โโโ README.md # Hugging Face Model Card & Documentation
| Special Token | Function | Routed Component |
|---|---|---|
<think> ... </think> | Cognitive reasoning & query deconstruction | 250M Neural Weights |
| `< | mem_query | > ... < |
| `< | mem_payload | > ... < |
| `< | calc | > ... < |
| `< | synthesize | >` |
AviGPT-250M-Instruct is an original architecture created, engineered, and trained exclusively by Yadlapalli Avinash Ricky.
@misc{ricky2026avigpt250minstruct,
author = {Yadlapalli Avinash Ricky},
title = {AviGPT-250M-Instruct: Semi-Parametric Edge Intelligence with Native NVMe Hardware Memory Bus},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/AvinashRicky/avigpt-250m-instruct}}
}
Python
59.2%
Jupyter Notebook
39.9%
World's First 250M Small Language Model with a Native NVMe Hardware Memory Bus. 100% Factual Retrieval, 100% Math Precision, #1 Composite Efficiency (0.40). Architect: Yadlapalli Avinash Ricky.
See the codeArchitect, System Designer & Sole Creator: Yadlapalli Avinash Ricky (India ๐ฎ๐ณ)
Model Parameters: 250,269,696 (~250M)
Resident VRAM Footprint: 488 MB (FP16 Edge Mode)
Checkpointed Weights: 477.5 MB
Status: SFT 2.0 Production Release
<|mem_query|>) to fetch authoritative ground truth directly from SSD storage, eliminating statistical guessing.<|calc|> tokens directly into a sandboxed AST SafeMath Evaluator, completely eliminating arithmetic hallucinations.AviGPT-250M-Instruct introduces Semi-Parametric Decoupling to edge AI: decoupling Cognitive Reasoning (handled by 250M compact transformer weights) from Factual Memory (stored in a native, zero-latency NVMe SSD memory bus powered by SQLite FTS5 BM25).
Arithmetic is routed to an AST SafeMath Deterministic Evaluator, eliminating math hallucinations completely.
AviGPT-250M-Instruct was evaluated head-to-head on an identical benchmark against 7 leading open-source models up to 1.1 Billion parameters on an NVIDIA Tesla T4 GPU (15GB VRAM):
| Rank | Model | Parameters | Factual Acc | Math Precision | Composite Acc | Composite Efficiency | VRAM | Avg Latency |
|---|---|---|---|---|---|---|---|---|
| ๐ 1 | AviGPT-250M-Instruct (NVMe Bus) | 250M | 100.0% | 100.0% | 100.0% | 0.40 ๐ฅ | 488 MB | 1.84s |
| 2 | SmolLM2-135M-Instruct | 135M | 100.0% | 0.0% | 50.0% | 0.37 | 266 MB | 5.63s |
| 3 | SmolLM2-360M-Instruct | 362M | 100.0% | 50.0% | 75.0% | 0.21 | 699 MB | 3.58s |
| 4 | Qwen2.5-0.5B-Instruct | 494M | 87.5% | 83.3% | 85.4% | 0.17 | 952 MB | 4.13s |
| 5 | H2O-Danube3-500M-Chat | 514M | 100.0% | 50.0% | 75.0% | 0.15 | 990 MB | 3.58s |
| 6 | TinyLlama-1.1B-Chat | 1,100M | 87.5% | 16.7% | 52.1% | 0.05 | 2,108 MB | 3.82s |
| 7 | GPT-Neo-125M | 125M | 0.0% | 16.7% | 8.4% | 0.07 | 287 MB | 2.83s |
| 8 | OpenELM-270M-Instruct | 270M | 0.0% | 0.0% | 0.0% | 0.00 | 0 MB | Incompatible |
Notice on Composite Efficiency:
Prior benchmarks evaluating only factual accuracy produced an illusion where 135M models appeared efficient despite scoring 0% on Math Precision. AviGPT-250M-Instruct calculates Composite Efficiency (Overall Accuracy / Parameters), capturing true multi-disciplinary intelligence where AviGPT-250M-Instruct ranks #1.
AviGPT-250M-Instruct replaces heavy vector databases (FAISS, Chroma, Pinecone) with an optimized, sub-millisecond local SQLite FTS5 engine operating directly on high-speed NVMe storage:
competitor_benchmark_colab.ipynb directly in Google Colab.# Clone from GitHub:
git clone https://github.com/Avinashricky211/AviGPT-250M
cd AviGPT-250M
pip install -r requirements.txt
python download_weights.py
# Or clone directly with weights from Hugging Face:
git clone https://huggingface.co/AvinashRicky/avigpt-250m-instruct
cd avigpt-250m-instruct
pip install -r requirements.txt
launch_terminal.batpython terminal_eval.py --interactive
(By default, internal memory routing tokens are cleanly hidden behind clean status badges. Use python terminal_eval.py --interactive --debug to inspect raw token traces).
Verify factual recall, deterministic math, and hardware memory bus latency:
python eval_proof.py
Expand AviGPT's knowledge base without expensive retraining runs or prompt bloating:
# Ingest single fact:
python ingest_knowledge.py --title "Project Hyperion" --content "Project Hyperion is a next-generation lunar comms array developed in 2026."
# Ingest an entire document or folder:
python ingest_knowledge.py --file documents/research_paper.txt
python ingest_knowledge.py --folder documents/company_knowledge_base/
Convert any Hugging Face dataset (Wikipedia, ArXiv, Fable) into AviGPT's high-speed memory bus:
python dataset_cookbook.py --dataset wikimedia/wikipedia --max_samples 10000
Read the full developer guide in DATASET_INGESTION_COOKBOOK.md.
from memory_bus import SSDMemoryEngine
engine = SSDMemoryEngine()
engine.store(title="Project Hyperion", content="Autonomous lunar relay.", domain="Space")
avigpt_250m_release/
โโโ assets/ # High-resolution benchmark & architecture graphics
โ โโโ avigpt_architecture.png
โ โโโ competitor_comparison_charts.png
โ โโโ accuracy_vs_params.png
โ โโโ memory_bus_latency.png
โโโ checkpoints/
โ โโโ avigpt_250m_instruct.pt # 477.5 MB SFT 2.0 Crown Checkpoint
โโโ tokenizer_avigpt/ # Custom 32,000 Byte-Level BPE Tokenizer
โโโ avigpt_ssd_memory.db # 55.68 MB NVMe FTS5 Knowledge Base (24,628 articles)
โโโ config.py # Architectural config & special token registry
โโโ model.py # AviGPT neural core (RoPE, GQA, SwiGLU, RMSNorm)
โโโ memory_bus.py # NVMe SSD hardware memory engine & SafeMath (Protected)
โโโ terminal_eval.py # High-performance terminal inference engine (Protected)
โโโ eval_proof.py # Scientific benchmark verification suite
โโโ dataset_cookbook.py # Universal dataset converter (Hugging Face / JSONL)
โโโ DATASET_INGESTION_COOKBOOK.md # Dataset ingestion developer guide
โโโ ingest_knowledge.py # Direct knowledge ingestion CLI
โโโ competitor_benchmark_colab.ipynb # 7-Model competitor benchmark notebook
โโโ launch_terminal.bat # 1-click Windows Terminal launcher
โโโ requirements.txt # Minimal inference dependencies
โโโ README.md # Hugging Face Model Card & Documentation
| Special Token | Function | Routed Component |
|---|---|---|
<think> ... </think> | Cognitive reasoning & query deconstruction | 250M Neural Weights |
| `< | mem_query | > ... < |
| `< | mem_payload | > ... < |
| `< | calc | > ... < |
| `< | synthesize | >` |
AviGPT-250M-Instruct is an original architecture created, engineered, and trained exclusively by Yadlapalli Avinash Ricky.
@misc{ricky2026avigpt250minstruct,
author = {Yadlapalli Avinash Ricky},
title = {AviGPT-250M-Instruct: Semi-Parametric Edge Intelligence with Native NVMe Hardware Memory Bus},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/AvinashRicky/avigpt-250m-instruct}}
}
Python
59.2%
Jupyter Notebook
39.9%