Avinashricky211/AviGPT-250M

World's First 250M Small Language Model with a Native NVMe Hardware Memory Bus. 100% Factual Retrieval, 100% Math Precision, #1 Composite Efficiency (0.40). Architect: Yadlapalli Avinash Ricky.

Python

2

6 commits

updated Sep 30, 2026

See the code

See what people are saying

README

๐Ÿ‘‘ AviGPT-250M-Instruct: Semi-Parametric Edge Intelligence

World's First 250M Small Language Model with a Native NVMe Hardware Memory Bus

Architect, System Designer & Sole Creator: Yadlapalli Avinash Ricky (India ๐Ÿ‡ฎ๐Ÿ‡ณ)
Model Parameters: 250,269,696 (~250M)
Resident VRAM Footprint: 488 MB (FP16 Edge Mode)
Checkpointed Weights: 477.5 MB
Status: SFT 2.0 Production Release


Open In Colab License: Apache 2.0 Parameters: 250M VRAM: 488MB Retrieval: 0.002ms Composite Efficiency: 0.40


๐ŸŒŸ Key Architectural Breakthroughs

  • ๐Ÿ‡ฎ๐Ÿ‡ณ Pioneered in India: Independently architected and engineered from the ground up by Yadlapalli Avinash Ricky as a next-generation breakthrough in edge-tier Small Language Models (SLMs).
  • โšก World's First 250M SLM with Native NVMe Bus: Decouples parametric weights from non-parametric factual storage using an ultra-low latency (0.002 ms) SQLite FTS5 engine operating directly on high-speed NVMe flash storage.
  • ๐Ÿ›ก๏ธ Zero Parametric Hallucination on Indexed Knowledge: Factual queries trigger hardware routing tokens (<|mem_query|>) to fetch authoritative ground truth directly from SSD storage, eliminating statistical guessing.
  • ๐Ÿงฎ 100% Deterministic Arithmetic Accuracy: Emits <|calc|> tokens directly into a sandboxed AST SafeMath Evaluator, completely eliminating arithmetic hallucinations.
  • ๐Ÿ† Global #1 Leaderboard Composite Efficiency (0.40): Outperforms models up to 4.4x its parameter size (including TinyLlama-1.1B) across joint factual recall (100.0%) and mathematical precision (100.0%).

๐Ÿ›๏ธ Architectural Overview

AviGPT-250M-Instruct introduces Semi-Parametric Decoupling to edge AI: decoupling Cognitive Reasoning (handled by 250M compact transformer weights) from Factual Memory (stored in a native, zero-latency NVMe SSD memory bus powered by SQLite FTS5 BM25).

Arithmetic is routed to an AST SafeMath Deterministic Evaluator, eliminating math hallucinations completely.

AviGPT Architecture Blueprint

๐Ÿ† Head-to-Head Competitor Benchmark

AviGPT-250M-Instruct was evaluated head-to-head on an identical benchmark against 7 leading open-source models up to 1.1 Billion parameters on an NVIDIA Tesla T4 GPU (15GB VRAM):

Benchmark Comparison Charts

Official Leaderboard (Verified on Google Colab T4)

RankModelParametersFactual AccMath PrecisionComposite AccComposite EfficiencyVRAMAvg Latency
๐Ÿ‘‘ 1AviGPT-250M-Instruct (NVMe Bus)250M100.0%100.0%100.0%0.40 ๐Ÿฅ‡488 MB1.84s
2SmolLM2-135M-Instruct135M100.0%0.0%50.0%0.37266 MB5.63s
3SmolLM2-360M-Instruct362M100.0%50.0%75.0%0.21699 MB3.58s
4Qwen2.5-0.5B-Instruct494M87.5%83.3%85.4%0.17952 MB4.13s
5H2O-Danube3-500M-Chat514M100.0%50.0%75.0%0.15990 MB3.58s
6TinyLlama-1.1B-Chat1,100M87.5%16.7%52.1%0.052,108 MB3.82s
7GPT-Neo-125M125M0.0%16.7%8.4%0.07287 MB2.83s
8OpenELM-270M-Instruct270M0.0%0.0%0.0%0.000 MBIncompatible
Pareto Efficiency Curve

Notice on Composite Efficiency:
Prior benchmarks evaluating only factual accuracy produced an illusion where 135M models appeared efficient despite scoring 0% on Math Precision. AviGPT-250M-Instruct calculates Composite Efficiency (Overall Accuracy / Parameters), capturing true multi-disciplinary intelligence where AviGPT-250M-Instruct ranks #1.


โšก Hardware-Speed Flash Retrieval (NVMe Bus)

AviGPT-250M-Instruct replaces heavy vector databases (FAISS, Chroma, Pinecone) with an optimized, sub-millisecond local SQLite FTS5 engine operating directly on high-speed NVMe storage:

Flash Retrieval Speed Comparison
  • NVMe Hardware Memory Bus: 0.0020 ms (~496,000 queries/second)
  • Local Vector DBs (Chroma / FAISS): 45.0 ms (22,500x slower)
  • Cloud Vector DBs (Pinecone / Milvus): 100.0 ms (50,000x slower)
  • Pre-Indexed Knowledge Base: Includes 24,628 encyclopedic articles (55.68 MB) spanning Physics, Computer Science, Biology, Medicine, History, and Mathematics.

๐Ÿš€ Quick Start Guides

  1. Open the included competitor_benchmark_colab.ipynb directly in Google Colab.
  2. Select Runtime > Change runtime type > T4 GPU.
  3. Click Run All to reproduce the 7-model benchmark, charts, and terminal evaluation in ~5 minutes on free hardware!

Path B: Local Terminal Cognitive Engine (Windows & Linux)

1. Clone & Install Dependencies

# Clone from GitHub:
git clone https://github.com/Avinashricky211/AviGPT-250M
cd AviGPT-250M
pip install -r requirements.txt
python download_weights.py

# Or clone directly with weights from Hugging Face:
git clone https://huggingface.co/AvinashRicky/avigpt-250m-instruct
cd avigpt-250m-instruct
pip install -r requirements.txt

2. Launch Terminal Engine

  • Windows (1-Click): Double-click launch_terminal.bat
  • Linux / Mac / Windows CLI:
python terminal_eval.py --interactive

(By default, internal memory routing tokens are cleanly hidden behind clean status badges. Use python terminal_eval.py --interactive --debug to inspect raw token traces).

3. Run Scientific Verification Suite

Verify factual recall, deterministic math, and hardware memory bus latency:

python eval_proof.py

๐Ÿ“š Dynamic Knowledge Ingestion (Zero Retraining!)

Expand AviGPT's knowledge base without expensive retraining runs or prompt bloating:

1. Ingest via CLI

# Ingest single fact:
python ingest_knowledge.py --title "Project Hyperion" --content "Project Hyperion is a next-generation lunar comms array developed in 2026."

# Ingest an entire document or folder:
python ingest_knowledge.py --file documents/research_paper.txt
python ingest_knowledge.py --folder documents/company_knowledge_base/

2. Universal Dataset Conversion (Cookbook)

Convert any Hugging Face dataset (Wikipedia, ArXiv, Fable) into AviGPT's high-speed memory bus:

python dataset_cookbook.py --dataset wikimedia/wikipedia --max_samples 10000

Read the full developer guide in DATASET_INGESTION_COOKBOOK.md.

3. Python Ingestion API (2 Lines)

from memory_bus import SSDMemoryEngine

engine = SSDMemoryEngine()
engine.store(title="Project Hyperion", content="Autonomous lunar relay.", domain="Space")

๐Ÿ“ Repository Structure

avigpt_250m_release/
โ”œโ”€โ”€ assets/                            # High-resolution benchmark & architecture graphics
โ”‚   โ”œโ”€โ”€ avigpt_architecture.png
โ”‚   โ”œโ”€โ”€ competitor_comparison_charts.png
โ”‚   โ”œโ”€โ”€ accuracy_vs_params.png
โ”‚   โ””โ”€โ”€ memory_bus_latency.png
โ”œโ”€โ”€ checkpoints/
โ”‚   โ””โ”€โ”€ avigpt_250m_instruct.pt        # 477.5 MB SFT 2.0 Crown Checkpoint
โ”œโ”€โ”€ tokenizer_avigpt/                  # Custom 32,000 Byte-Level BPE Tokenizer
โ”œโ”€โ”€ avigpt_ssd_memory.db               # 55.68 MB NVMe FTS5 Knowledge Base (24,628 articles)
โ”œโ”€โ”€ config.py                          # Architectural config & special token registry
โ”œโ”€โ”€ model.py                           # AviGPT neural core (RoPE, GQA, SwiGLU, RMSNorm)
โ”œโ”€โ”€ memory_bus.py                      # NVMe SSD hardware memory engine & SafeMath (Protected)
โ”œโ”€โ”€ terminal_eval.py                   # High-performance terminal inference engine (Protected)
โ”œโ”€โ”€ eval_proof.py                      # Scientific benchmark verification suite
โ”œโ”€โ”€ dataset_cookbook.py                # Universal dataset converter (Hugging Face / JSONL)
โ”œโ”€โ”€ DATASET_INGESTION_COOKBOOK.md      # Dataset ingestion developer guide
โ”œโ”€โ”€ ingest_knowledge.py                # Direct knowledge ingestion CLI
โ”œโ”€โ”€ competitor_benchmark_colab.ipynb   # 7-Model competitor benchmark notebook
โ”œโ”€โ”€ launch_terminal.bat                # 1-click Windows Terminal launcher
โ”œโ”€โ”€ requirements.txt                   # Minimal inference dependencies
โ””โ”€โ”€ README.md                          # Hugging Face Model Card & Documentation

๐Ÿ”ฌ Special Token Routing Protocol

Special TokenFunctionRouted Component
<think> ... </think>Cognitive reasoning & query deconstruction250M Neural Weights
`<mem_query> ... <
`<mem_payload> ... <
`<calc> ... <
`<synthesize>`

๐Ÿ“œ Authorship & Citation

AviGPT-250M-Instruct is an original architecture created, engineered, and trained exclusively by Yadlapalli Avinash Ricky.

@misc{ricky2026avigpt250minstruct,
  author = {Yadlapalli Avinash Ricky},
  title = {AviGPT-250M-Instruct: Semi-Parametric Edge Intelligence with Native NVMe Hardware Memory Bus},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/AvinashRicky/avigpt-250m-instruct}}
}
bm25
edge-ai
gqa
hardware-memory-bus
huggingface
india-ai
llm
nvme
pytorch
rag
rope
safemath
semi-parametric
slm
small-language-models
sqlite
swiglu
transformers
zero-hallucination

Avinashricky211/AviGPT-250M

World's First 250M Small Language Model with a Native NVMe Hardware Memory Bus. 100% Factual Retrieval, 100% Math Precision, #1 Composite Efficiency (0.40). Architect: Yadlapalli Avinash Ricky.

Python

2

6 commits

updated Sep 30, 2026

See the code

See what people are saying

README

๐Ÿ‘‘ AviGPT-250M-Instruct: Semi-Parametric Edge Intelligence

World's First 250M Small Language Model with a Native NVMe Hardware Memory Bus

Architect, System Designer & Sole Creator: Yadlapalli Avinash Ricky (India ๐Ÿ‡ฎ๐Ÿ‡ณ)
Model Parameters: 250,269,696 (~250M)
Resident VRAM Footprint: 488 MB (FP16 Edge Mode)
Checkpointed Weights: 477.5 MB
Status: SFT 2.0 Production Release


Open In Colab License: Apache 2.0 Parameters: 250M VRAM: 488MB Retrieval: 0.002ms Composite Efficiency: 0.40


๐ŸŒŸ Key Architectural Breakthroughs

  • ๐Ÿ‡ฎ๐Ÿ‡ณ Pioneered in India: Independently architected and engineered from the ground up by Yadlapalli Avinash Ricky as a next-generation breakthrough in edge-tier Small Language Models (SLMs).
  • โšก World's First 250M SLM with Native NVMe Bus: Decouples parametric weights from non-parametric factual storage using an ultra-low latency (0.002 ms) SQLite FTS5 engine operating directly on high-speed NVMe flash storage.
  • ๐Ÿ›ก๏ธ Zero Parametric Hallucination on Indexed Knowledge: Factual queries trigger hardware routing tokens (<|mem_query|>) to fetch authoritative ground truth directly from SSD storage, eliminating statistical guessing.
  • ๐Ÿงฎ 100% Deterministic Arithmetic Accuracy: Emits <|calc|> tokens directly into a sandboxed AST SafeMath Evaluator, completely eliminating arithmetic hallucinations.
  • ๐Ÿ† Global #1 Leaderboard Composite Efficiency (0.40): Outperforms models up to 4.4x its parameter size (including TinyLlama-1.1B) across joint factual recall (100.0%) and mathematical precision (100.0%).

๐Ÿ›๏ธ Architectural Overview

AviGPT-250M-Instruct introduces Semi-Parametric Decoupling to edge AI: decoupling Cognitive Reasoning (handled by 250M compact transformer weights) from Factual Memory (stored in a native, zero-latency NVMe SSD memory bus powered by SQLite FTS5 BM25).

Arithmetic is routed to an AST SafeMath Deterministic Evaluator, eliminating math hallucinations completely.

AviGPT Architecture Blueprint

๐Ÿ† Head-to-Head Competitor Benchmark

AviGPT-250M-Instruct was evaluated head-to-head on an identical benchmark against 7 leading open-source models up to 1.1 Billion parameters on an NVIDIA Tesla T4 GPU (15GB VRAM):

Benchmark Comparison Charts

Official Leaderboard (Verified on Google Colab T4)

RankModelParametersFactual AccMath PrecisionComposite AccComposite EfficiencyVRAMAvg Latency
๐Ÿ‘‘ 1AviGPT-250M-Instruct (NVMe Bus)250M100.0%100.0%100.0%0.40 ๐Ÿฅ‡488 MB1.84s
2SmolLM2-135M-Instruct135M100.0%0.0%50.0%0.37266 MB5.63s
3SmolLM2-360M-Instruct362M100.0%50.0%75.0%0.21699 MB3.58s
4Qwen2.5-0.5B-Instruct494M87.5%83.3%85.4%0.17952 MB4.13s
5H2O-Danube3-500M-Chat514M100.0%50.0%75.0%0.15990 MB3.58s
6TinyLlama-1.1B-Chat1,100M87.5%16.7%52.1%0.052,108 MB3.82s
7GPT-Neo-125M125M0.0%16.7%8.4%0.07287 MB2.83s
8OpenELM-270M-Instruct270M0.0%0.0%0.0%0.000 MBIncompatible
Pareto Efficiency Curve

Notice on Composite Efficiency:
Prior benchmarks evaluating only factual accuracy produced an illusion where 135M models appeared efficient despite scoring 0% on Math Precision. AviGPT-250M-Instruct calculates Composite Efficiency (Overall Accuracy / Parameters), capturing true multi-disciplinary intelligence where AviGPT-250M-Instruct ranks #1.


โšก Hardware-Speed Flash Retrieval (NVMe Bus)

AviGPT-250M-Instruct replaces heavy vector databases (FAISS, Chroma, Pinecone) with an optimized, sub-millisecond local SQLite FTS5 engine operating directly on high-speed NVMe storage:

Flash Retrieval Speed Comparison
  • NVMe Hardware Memory Bus: 0.0020 ms (~496,000 queries/second)
  • Local Vector DBs (Chroma / FAISS): 45.0 ms (22,500x slower)
  • Cloud Vector DBs (Pinecone / Milvus): 100.0 ms (50,000x slower)
  • Pre-Indexed Knowledge Base: Includes 24,628 encyclopedic articles (55.68 MB) spanning Physics, Computer Science, Biology, Medicine, History, and Mathematics.

๐Ÿš€ Quick Start Guides

  1. Open the included competitor_benchmark_colab.ipynb directly in Google Colab.
  2. Select Runtime > Change runtime type > T4 GPU.
  3. Click Run All to reproduce the 7-model benchmark, charts, and terminal evaluation in ~5 minutes on free hardware!

Path B: Local Terminal Cognitive Engine (Windows & Linux)

1. Clone & Install Dependencies

# Clone from GitHub:
git clone https://github.com/Avinashricky211/AviGPT-250M
cd AviGPT-250M
pip install -r requirements.txt
python download_weights.py

# Or clone directly with weights from Hugging Face:
git clone https://huggingface.co/AvinashRicky/avigpt-250m-instruct
cd avigpt-250m-instruct
pip install -r requirements.txt

2. Launch Terminal Engine

  • Windows (1-Click): Double-click launch_terminal.bat
  • Linux / Mac / Windows CLI:
python terminal_eval.py --interactive

(By default, internal memory routing tokens are cleanly hidden behind clean status badges. Use python terminal_eval.py --interactive --debug to inspect raw token traces).

3. Run Scientific Verification Suite

Verify factual recall, deterministic math, and hardware memory bus latency:

python eval_proof.py

๐Ÿ“š Dynamic Knowledge Ingestion (Zero Retraining!)

Expand AviGPT's knowledge base without expensive retraining runs or prompt bloating:

1. Ingest via CLI

# Ingest single fact:
python ingest_knowledge.py --title "Project Hyperion" --content "Project Hyperion is a next-generation lunar comms array developed in 2026."

# Ingest an entire document or folder:
python ingest_knowledge.py --file documents/research_paper.txt
python ingest_knowledge.py --folder documents/company_knowledge_base/

2. Universal Dataset Conversion (Cookbook)

Convert any Hugging Face dataset (Wikipedia, ArXiv, Fable) into AviGPT's high-speed memory bus:

python dataset_cookbook.py --dataset wikimedia/wikipedia --max_samples 10000

Read the full developer guide in DATASET_INGESTION_COOKBOOK.md.

3. Python Ingestion API (2 Lines)

from memory_bus import SSDMemoryEngine

engine = SSDMemoryEngine()
engine.store(title="Project Hyperion", content="Autonomous lunar relay.", domain="Space")

๐Ÿ“ Repository Structure

avigpt_250m_release/
โ”œโ”€โ”€ assets/                            # High-resolution benchmark & architecture graphics
โ”‚   โ”œโ”€โ”€ avigpt_architecture.png
โ”‚   โ”œโ”€โ”€ competitor_comparison_charts.png
โ”‚   โ”œโ”€โ”€ accuracy_vs_params.png
โ”‚   โ””โ”€โ”€ memory_bus_latency.png
โ”œโ”€โ”€ checkpoints/
โ”‚   โ””โ”€โ”€ avigpt_250m_instruct.pt        # 477.5 MB SFT 2.0 Crown Checkpoint
โ”œโ”€โ”€ tokenizer_avigpt/                  # Custom 32,000 Byte-Level BPE Tokenizer
โ”œโ”€โ”€ avigpt_ssd_memory.db               # 55.68 MB NVMe FTS5 Knowledge Base (24,628 articles)
โ”œโ”€โ”€ config.py                          # Architectural config & special token registry
โ”œโ”€โ”€ model.py                           # AviGPT neural core (RoPE, GQA, SwiGLU, RMSNorm)
โ”œโ”€โ”€ memory_bus.py                      # NVMe SSD hardware memory engine & SafeMath (Protected)
โ”œโ”€โ”€ terminal_eval.py                   # High-performance terminal inference engine (Protected)
โ”œโ”€โ”€ eval_proof.py                      # Scientific benchmark verification suite
โ”œโ”€โ”€ dataset_cookbook.py                # Universal dataset converter (Hugging Face / JSONL)
โ”œโ”€โ”€ DATASET_INGESTION_COOKBOOK.md      # Dataset ingestion developer guide
โ”œโ”€โ”€ ingest_knowledge.py                # Direct knowledge ingestion CLI
โ”œโ”€โ”€ competitor_benchmark_colab.ipynb   # 7-Model competitor benchmark notebook
โ”œโ”€โ”€ launch_terminal.bat                # 1-click Windows Terminal launcher
โ”œโ”€โ”€ requirements.txt                   # Minimal inference dependencies
โ””โ”€โ”€ README.md                          # Hugging Face Model Card & Documentation

๐Ÿ”ฌ Special Token Routing Protocol

Special TokenFunctionRouted Component
<think> ... </think>Cognitive reasoning & query deconstruction250M Neural Weights
`<mem_query> ... <
`<mem_payload> ... <
`<calc> ... <
`<synthesize>`

๐Ÿ“œ Authorship & Citation

AviGPT-250M-Instruct is an original architecture created, engineered, and trained exclusively by Yadlapalli Avinash Ricky.

@misc{ricky2026avigpt250minstruct,
  author = {Yadlapalli Avinash Ricky},
  title = {AviGPT-250M-Instruct: Semi-Parametric Edge Intelligence with Native NVMe Hardware Memory Bus},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/AvinashRicky/avigpt-250m-instruct}}
}
bm25
edge-ai
gqa
hardware-memory-bus
huggingface
india-ai
llm
nvme
pytorch
rag
rope
safemath
semi-parametric
slm
small-language-models
sqlite
swiglu
transformers
zero-hallucination

Languages

Python

59.2%

Jupyter Notebook

39.9%