knowthankyew/event-driven-ftaas

Python

0

71 commits

updated Oct 4, 2026

See the code

README

FTaaS: Event-Driven "Fine-Tuning as a Service" (event-driven-ftaas / ml)

An enterprise-grade, asynchronous, event-driven machine learning platform demonstrating a polyglot microservice pattern:

  • Control Plane & Ingestion Gateway: .NET 10 Minimal API / AMQP Producer
  • Compute Worker: Python 3.12, PyTorch, Hugging Face PEFT / LoRA, AMQP Consumer
  • Telemetry & Experiment Tracking: OpenTelemetry ActivitySource + MLflow Tracking Server & Model Registry
  • Message Broker: RabbitMQ (AMQP) with DLX and retry handling
  • Base Models: Multi-model architecture supporting ultra-compact models (HuggingFaceTB/SmolLM2-135M) for zero-cost instant local training, high-capacity reasoning models (google/gemma-2-2b-it) with hardware preflight gates and float16 MPS optimization, and 1.58-bit ternary foundation models (microsoft/BitNet-b1.58-2B-4T) executing natively on CPU via AVX2 integer addition/subtraction.
  • Privacy & Observability Standard: Complies with PRIVACY_TELEMETRY_SCHEMA.md.

πŸ“Ί Interactive Video Demonstration

Watch the complete end-to-end studio workflow in actionβ€”from side-by-side compliance policy comparison to live PII guardrails and event pipeline tracking:

FTaaS Studio Interactive Walkthrough

πŸ“Ή High-Definition Recording: demo.mp4 (Playwright automated run, 1366x860)

To regenerate this demonstration at any time, run: ./scripts/record-demo.sh


Quickstart for Non-Technical Users & Business Leaders

[!TIP] Zero-Code Operation: You do not need to write Python, manage PyTorch, or configure machine learning pipelines to evaluate FTaaS.

Prerequisite: .NET 10 SDK installed on your machine.

Run one command to launch the full interactive web application:

./scripts/start-studio.sh

Then open your browser to http://localhost:5100.

πŸ“– Looking for a guided walkthrough? Read the FTaaS Enterprise Studio Guide for step-by-step instructions, visual explanations, and FAQs.

What You Can Do in the Studio:

  1. ⚑ Side-by-Side Comparison Arena: Select a business persona (e.g. Fintech Support & Compliance, Enterprise SaaS Ops, or Financial Earnings) and test realistic inquiries. Observe how a Generic Foundation Model contrasts with your Company Custom AI (incorporating team SLAs, policy limits, and required regulatory disclaimers).
  2. πŸ›‘οΈ Team Policy & Disclaimer Inspector: Pattern-matching verification checking whether designated training guidelines and statutory tags appear in completions (demonstration aid; not legal advice or statutory regulatory certification).
  3. πŸ› οΈ No-Code Adapter Studio: Define your department's question-and-answer pairs in an intuitive table editor, or drag-and-drop a .csv document. Select your target base model (SmolLM2-135M for lightweight deployment or Gemma 2 2B IT for deeper reasoning capacity) with built-in hardware disclaimers. Click "Train & Deploy Team Adapter" to submit the job in milliseconds without waiting on background compute.
  4. πŸ”„ Live Event Pipeline Visualizer: An animated diagram demonstrating how incoming user requests decouple from background training compute.
  5. πŸ“š Adapter Library: View trained department models with base model indicators, inspect their lightweight footprint (~1.8 MB for SmolLM2, ~12.2 MB for Gemma 2B), and load them into the arena with one click.

[!NOTE] Preview vs. Live Compute: When launched standalone via start-studio.sh, the studio operates in Interactive Preview Mode with pre-formatted demonstration outputs so you can evaluate the interface without spinning up Docker containers or downloading multi-gigabyte models. To run live on-device GPU inference, start Docker and the Python services via ./scripts/dev-up.sh.

[!CAUTION] Data Privacy & Guardrail Scope: Automated heuristic checks detect delimited SSNs and Luhn-valid credit card numbers across all columns (including unmapped metadata). However, this is a best-effort defense, not an exhaustive DLP certification tool. Always sanitize datasets prior to training.


Privacy, Observability & Job Lifecycle Spans

FTaaS follows the portfolio governance standard documented in PRIVACY_TELEMETRY_SCHEMA.md:

  • Consumer Privacy Default: Runs in memory_only mode with zero outbound telemetry egress.
  • Job Lifecycle Spans: Traces the entire training pipeline: job.accepted β†’ job.published β†’ job.consumed β†’ job.training.started β†’ job.training.finished β†’ job.registered.
  • Payload Redaction: Attributes strictly log job IDs, statuses, durations, hardware devices (cuda | mps | cpu), and adapter byte sizes. No raw prompts, completions, or dataset contents are ever permitted in spans.
  • Inspection & Purge API:
    • GET /api/v1/telemetry/privacy-audit: Real-time inspection of active telemetry mode and buffer state.
    • GET /api/v1/telemetry/spans: Inspect buffered in-memory spans.
    • POST /api/v1/telemetry/burn: Immediately purges the in-memory telemetry buffer.
  • Enterprise OTLP Overlay: Setting OTEL_EXPORTER_OTLP_ENDPOINT exports spans to an enterprise collector while maintaining the strict payload scrubbing invariant.

Architecture at a Glance

[Non-Tech Web Studio / Client]
        β”‚
        β–Ό (POST /api/v1/jobs multipart or datasetPath)
 [Ingestion Gateway (.NET 10)] ──► Validates JSONL, Computes SHA-256 Hash, Stores on Disk
        β”‚
        β–Ό (Job Request Event with datasetPath & datasetHash)
 [Async Queue / Message Broker (RabbitMQ)]
        β”‚
        β–Ό
 [Worker / Orchestrator (Python)] ──► Spin up Training Compute (PyTorch / Hugging Face LoRA)
        β”‚                                        β”‚
        β–Ό                                        β–Ό
 [Model Registry (MLflow)]             [Metrics & Loss Telemetry]
        β”‚
        β–Ό
 [Dynamic Inference Engine] ◄──────── Base Model + On-the-fly LoRA Adapter Mounting
        β”‚
        β–Ό
 [Side-by-Side Comparison] ────────── Base Completion vs Fine-Tuned Completion

Why This Architecture? (Engineering Trade-Offs)

Architectural DecisionTrade-Off RationaleEnterprise Reality
Polyglot (.NET 10 + Python).NET offers high-concurrency, strongly-typed contracts, and low-latency API handling; Python owns the cutting-edge ML ecosystem (PyTorch, PEFT).Solves the common anti-pattern of writing web APIs in Python or trying to run ML training in C#. Leverages each ecosystem's primary strength.
Out-of-Band Dataset StorageStoring .jsonl files on disk/object storage and passing only datasetPath + SHA-256 hash over RabbitMQ prevents broker bloat and memory pressure.Real datasets are megabytes to gigabytes. Message brokers degrade rapidly when payloads exceed hundreds of kilobytes.
LoRA (PEFT) vs Full Fine-TuningFreezing 99%+ of base model weights and only training low-rank adapter matrices reduces VRAM requirements by >80% and completes in 2–4 minutes on consumer hardware.In production, training 100 domain adapters on a single base model requires ~50MB per adapter instead of storing 100 full 7GB weights.
Asynchronous DecouplingThe client receives an immediate 202 Accepted with a JobId. Heavy compute executes out-of-process.Synchronous model training over HTTP guarantees timeouts, socket exhaustion, and cascading failures under load.
Side-by-Side Inference GatewayMounting adapters dynamically onto a shared base model allows rapid A/B testing and direct before/after domain comparison.Avoids spinning up dedicated GPU containers for every custom-trained model variant.

Developer Quickstart Workflow (Command Line & APIs)

For engineers wanting to run each microservice manually:

# 1. Bring up Infrastructure (RabbitMQ + MLflow) and seed sample datasets
./scripts/dev-up.sh

# 2. Run the Ingestion API & Web Studio (.NET 10)
cd src/FtaaSService.Api && dotnet run

# 3. In another terminal, run the Training Worker (Python)
cd src/FtaaSService.Worker && source .venv/bin/activate && python consumer.py

# 4. In another terminal, run the Inference Engine
cd src/FtaaSService.Inference && source .venv/bin/activate && python app.py

# 5. Submit a fine-tuning job via curl
curl -X POST http://localhost:5100/api/v1/jobs \
  -F "file=@data/datasets/sample-support-compliance.jsonl" \
  -F "jobName=compliance-v1"

Automated Test Suites & Quality Engineering

The platform includes formal, automated unit and regression test suites across both the .NET control plane and Python compute layers:

1. .NET 10 API Test Suite (xUnit + Coverlet)

Validates compliance boundaries, data sanitization, and state machine idempotency:

  • PII Compliance Gateway: Delimited SSA SSNs, contextual SSNs, and Luhn-valid payment cards.
  • False-Positive Immunity: Verifies that non-contextual 9-digit integers (order IDs, invoice numbers) pass unhindered.
  • Full-Row Unmapped Column Scanning: Scans all metadata columns/properties; strips unmapped fields upon normalization.
  • Zero-Disk In-Memory Guarantee: Verifies zero bytes are written to disk upon compliance rejection.
  • Model Catalog & Validation: Enforces allowlisted base models via GET /api/v1/studio/models, ensuring invalid or unverified model identifiers are rejected at ingestion with 400 Bad Request.
  • Ingestion Ceilings: Enforces 25 MB file size and 50,000 record limits.
  • State Machine Idempotency: Rejects out-of-order sequence updates and prevents terminal state regressions in SQLite.
# Run all .NET unit tests with code coverage collection:
dotnet test tests/FtaaSService.Api.Tests --collect:"XPlat Code Coverage"

2. Python Worker & Inference Tests (unittest / pytest)

Validates model serving resilience, registry lookups, and file integrity:

  • Model Registry & Dynamic LoRA Targeting: Verifies target module selection (["q_proj", "v_proj"] vs ["q_proj", "v_proj", "k_proj", "o_proj"]) and template rendering across supported model architectures.
  • Inference Model Mismatch Protection: Returns actionable 409 Conflict responses when adapters are mounted against mismatched base models.
  • Hardware Preflight & Licensing: Intercepts Hugging Face gated license/auth failures (401/403) and surfaces actionable user remediation instructions.
  • Adapter Weight Integrity: Rejects truncated, zero-byte, or incomplete writes (<100 KB weights, <10 bytes config).
  • Bounded LRU Cache Eviction: Verifies dynamic eviction of least-recently-used LoRA adapters under memory pressure.
# Run all Python unit tests:
PYTHONPATH=src/FtaaSService.Worker:src/FtaaSService.Inference src/FtaaSService.Worker/.venv/bin/python -m pytest tests/

Automated End-to-End Integration Verification

To test the entire live pipeline (Infra $\rightarrow$ .NET 10 Ingestion $\rightarrow$ RabbitMQ $\rightarrow$ PyTorch/PEFT Training on MPS $\rightarrow$ MLflow $\rightarrow$ Dynamic Inference Comparison) in one command:

./scripts/verify-e2e.sh

Roadmap & Implementation Status

  • Architecture Specification & Design Contracts (docs/ARCHITECTURE.md)
  • Phase 1: Local Infrastructure Foundation (docker-compose.yml, scripts/dev-up.sh, RabbitMQ + MLflow healthchecks)
  • Phase 2: Ingestion & Control Plane (.NET 10 Minimal API, JSONL validation, SQLite state machine, AMQP producer/consumer)
  • Phase 3: Python Compute Worker & LoRA Pipeline (AMQP consumer, PyTorch MPS/CPU detection, PEFT/LoRA, MLflow telemetry)
  • Phase 4: Dynamic Model Serving & Side-by-Side Comparison (LoRA dynamic adapter mounting, comparison API)
  • Portfolio Phase 3a (The FTaaS Bridge Exporter): Edge ONNX Exporter (src/FtaaSService.Worker/exporter.py, scripts/export_edge_adapter.py) and Web API export endpoints (GET/POST /api/v1/jobs/{id}/export/edge) compiling LoRA adapters into web-optimized ONNX format with integrity checksums, documented in docs/SMOL_PROTOTYPE_RESULTS.md.
  • Portfolio Phase 3b (In-Browser Execution): Ingesting and executing exported ONNX packages directly inside client browser engines via WebGPU/WASM (onnxruntime-web), single-input ONNX export signature, and dynamic INT8 quantization.
  • Portfolio Phase 3c (Multi-Model Scaling & Gemma 2B Prototype): Dynamic multi-model registry (model_registry.py), google/gemma-2-2b-it support, hardware preflight warnings, float16 MPS optimization with gradient accumulation, and live benchmark evaluation documented in docs/GEMMA_PROTOTYPE_RESULTS.md.
  • Portfolio Phase 3d (Ternary Weight Scaling & BitNet b1.58 Prototype): Microsoft microsoft/BitNet-b1.58-2B-4T 1.58-bit ternary foundation model, native CPU execution via AVX2 integer addition/subtraction (zero GPU required), and empirical scaling benchmarks documented in docs/BITNET_PROTOTYPE_RESULTS.md.
  • BitNet b1.58 Ternary Integration & Engineering Roadmap: Multi-phase engineering roadmap covering native C++ kernel build, ternary LoRA training, multi-model catalog integration, upstream toolchain contributions, and dual-backend inference serving documented in docs/ROADMAP.md.
  • Comparative Performance Analysis: Empirical cross-architecture differentials across all three models documented in docs/MODEL_PERFORMANCE_COMPARISON.md.

πŸ“š Repository Documentation Index

All technical design documents, prototype evaluations, and implementation roadmaps are organized under docs/:

DocumentPurpose
docs/ARCHITECTURE.mdEnd-to-end system architecture, polyglot microservice boundaries, AMQP messaging contracts, and cloud mapping.
docs/ROADMAP.md7-phase BitNet b1.58 ternary integration roadmap, execution status, and upstream contribution tracking.
docs/STUDIO_GUIDE.mdNon-technical visual walkthrough and operational runbook for the FTaaS Enterprise Studio web application.
docs/MODEL_PERFORMANCE_COMPARISON.mdSymmetrical cross-architecture benchmark differentials across SmolLM2-135M, Gemma 2 2B IT, and BitNet b1.58.
docs/SMOL_PROTOTYPE_RESULTS.mdUltra-compact baseline (135M) empirical LoRA training metrics, ONNX export profile, and in-browser WASM execution.
docs/GEMMA_PROTOTYPE_RESULTS.mdHigh-capacity reasoning tier (2.6B) empirical benchmarks, Apple Silicon MPS tuning, and licensing guards.
docs/BITNET_PROTOTYPE_RESULTS.md1.58-bit ternary CPU execution (2.4B) AVX2 SIMD scaling benchmarks, memory profiling, and PyTorch LoRA metrics.
docs/PRIVACY_TELEMETRY_SCHEMA.mdPortfolio governance standard for strict attribute allowlists, burn-to-purge invariants, and OpenTelemetry overlay.
NOTICE.mdRoot statutory disclaimers, consumer privacy invariants, and third-party foundation model attribution.

Production Cloud Parity (AWS & GCP)

Local StackAWS Production ArchitectureGCP Production Architecture
.NET 10 API GatewayAmazon API Gateway + ECS FargateGoogle Cloud Run (.NET 10)
RabbitMQ Broker + DLQAmazon SQS (FIFO) + SQS DLQCloud Pub/Sub + Dead Letter
Storage (Datasets & DB)Amazon S3 + Aurora PostgreSQLGoogle Cloud Storage + Cloud SQL
Compute Worker (Python)SageMaker Training Jobs (Spot Instances)Vertex AI Custom Jobs (Preemptible)
Experiment TelemetryManaged MLflow / SageMaker ExperimentsVertex AI Experiments
Model RegistryMLflow Model Registry / SageMaker RegistryVertex AI Model Registry
Dynamic Inference ServerSageMaker Multi-Model Endpoints / TritonVertex AI Endpoints (vLLM / Triton)

License

This project is licensed under the MIT License.

knowthankyew/event-driven-ftaas

Python

0

71 commits

updated Oct 4, 2026

See the code

README

FTaaS: Event-Driven "Fine-Tuning as a Service" (event-driven-ftaas / ml)

An enterprise-grade, asynchronous, event-driven machine learning platform demonstrating a polyglot microservice pattern:

  • Control Plane & Ingestion Gateway: .NET 10 Minimal API / AMQP Producer
  • Compute Worker: Python 3.12, PyTorch, Hugging Face PEFT / LoRA, AMQP Consumer
  • Telemetry & Experiment Tracking: OpenTelemetry ActivitySource + MLflow Tracking Server & Model Registry
  • Message Broker: RabbitMQ (AMQP) with DLX and retry handling
  • Base Models: Multi-model architecture supporting ultra-compact models (HuggingFaceTB/SmolLM2-135M) for zero-cost instant local training, high-capacity reasoning models (google/gemma-2-2b-it) with hardware preflight gates and float16 MPS optimization, and 1.58-bit ternary foundation models (microsoft/BitNet-b1.58-2B-4T) executing natively on CPU via AVX2 integer addition/subtraction.
  • Privacy & Observability Standard: Complies with PRIVACY_TELEMETRY_SCHEMA.md.

πŸ“Ί Interactive Video Demonstration

Watch the complete end-to-end studio workflow in actionβ€”from side-by-side compliance policy comparison to live PII guardrails and event pipeline tracking:

FTaaS Studio Interactive Walkthrough

πŸ“Ή High-Definition Recording: demo.mp4 (Playwright automated run, 1366x860)

To regenerate this demonstration at any time, run: ./scripts/record-demo.sh


Quickstart for Non-Technical Users & Business Leaders

[!TIP] Zero-Code Operation: You do not need to write Python, manage PyTorch, or configure machine learning pipelines to evaluate FTaaS.

Prerequisite: .NET 10 SDK installed on your machine.

Run one command to launch the full interactive web application:

./scripts/start-studio.sh

Then open your browser to http://localhost:5100.

πŸ“– Looking for a guided walkthrough? Read the FTaaS Enterprise Studio Guide for step-by-step instructions, visual explanations, and FAQs.

What You Can Do in the Studio:

  1. ⚑ Side-by-Side Comparison Arena: Select a business persona (e.g. Fintech Support & Compliance, Enterprise SaaS Ops, or Financial Earnings) and test realistic inquiries. Observe how a Generic Foundation Model contrasts with your Company Custom AI (incorporating team SLAs, policy limits, and required regulatory disclaimers).
  2. πŸ›‘οΈ Team Policy & Disclaimer Inspector: Pattern-matching verification checking whether designated training guidelines and statutory tags appear in completions (demonstration aid; not legal advice or statutory regulatory certification).
  3. πŸ› οΈ No-Code Adapter Studio: Define your department's question-and-answer pairs in an intuitive table editor, or drag-and-drop a .csv document. Select your target base model (SmolLM2-135M for lightweight deployment or Gemma 2 2B IT for deeper reasoning capacity) with built-in hardware disclaimers. Click "Train & Deploy Team Adapter" to submit the job in milliseconds without waiting on background compute.
  4. πŸ”„ Live Event Pipeline Visualizer: An animated diagram demonstrating how incoming user requests decouple from background training compute.
  5. πŸ“š Adapter Library: View trained department models with base model indicators, inspect their lightweight footprint (~1.8 MB for SmolLM2, ~12.2 MB for Gemma 2B), and load them into the arena with one click.

[!NOTE] Preview vs. Live Compute: When launched standalone via start-studio.sh, the studio operates in Interactive Preview Mode with pre-formatted demonstration outputs so you can evaluate the interface without spinning up Docker containers or downloading multi-gigabyte models. To run live on-device GPU inference, start Docker and the Python services via ./scripts/dev-up.sh.

[!CAUTION] Data Privacy & Guardrail Scope: Automated heuristic checks detect delimited SSNs and Luhn-valid credit card numbers across all columns (including unmapped metadata). However, this is a best-effort defense, not an exhaustive DLP certification tool. Always sanitize datasets prior to training.


Privacy, Observability & Job Lifecycle Spans

FTaaS follows the portfolio governance standard documented in PRIVACY_TELEMETRY_SCHEMA.md:

  • Consumer Privacy Default: Runs in memory_only mode with zero outbound telemetry egress.
  • Job Lifecycle Spans: Traces the entire training pipeline: job.accepted β†’ job.published β†’ job.consumed β†’ job.training.started β†’ job.training.finished β†’ job.registered.
  • Payload Redaction: Attributes strictly log job IDs, statuses, durations, hardware devices (cuda | mps | cpu), and adapter byte sizes. No raw prompts, completions, or dataset contents are ever permitted in spans.
  • Inspection & Purge API:
    • GET /api/v1/telemetry/privacy-audit: Real-time inspection of active telemetry mode and buffer state.
    • GET /api/v1/telemetry/spans: Inspect buffered in-memory spans.
    • POST /api/v1/telemetry/burn: Immediately purges the in-memory telemetry buffer.
  • Enterprise OTLP Overlay: Setting OTEL_EXPORTER_OTLP_ENDPOINT exports spans to an enterprise collector while maintaining the strict payload scrubbing invariant.

Architecture at a Glance

[Non-Tech Web Studio / Client]
        β”‚
        β–Ό (POST /api/v1/jobs multipart or datasetPath)
 [Ingestion Gateway (.NET 10)] ──► Validates JSONL, Computes SHA-256 Hash, Stores on Disk
        β”‚
        β–Ό (Job Request Event with datasetPath & datasetHash)
 [Async Queue / Message Broker (RabbitMQ)]
        β”‚
        β–Ό
 [Worker / Orchestrator (Python)] ──► Spin up Training Compute (PyTorch / Hugging Face LoRA)
        β”‚                                        β”‚
        β–Ό                                        β–Ό
 [Model Registry (MLflow)]             [Metrics & Loss Telemetry]
        β”‚
        β–Ό
 [Dynamic Inference Engine] ◄──────── Base Model + On-the-fly LoRA Adapter Mounting
        β”‚
        β–Ό
 [Side-by-Side Comparison] ────────── Base Completion vs Fine-Tuned Completion

Why This Architecture? (Engineering Trade-Offs)

Architectural DecisionTrade-Off RationaleEnterprise Reality
Polyglot (.NET 10 + Python).NET offers high-concurrency, strongly-typed contracts, and low-latency API handling; Python owns the cutting-edge ML ecosystem (PyTorch, PEFT).Solves the common anti-pattern of writing web APIs in Python or trying to run ML training in C#. Leverages each ecosystem's primary strength.
Out-of-Band Dataset StorageStoring .jsonl files on disk/object storage and passing only datasetPath + SHA-256 hash over RabbitMQ prevents broker bloat and memory pressure.Real datasets are megabytes to gigabytes. Message brokers degrade rapidly when payloads exceed hundreds of kilobytes.
LoRA (PEFT) vs Full Fine-TuningFreezing 99%+ of base model weights and only training low-rank adapter matrices reduces VRAM requirements by >80% and completes in 2–4 minutes on consumer hardware.In production, training 100 domain adapters on a single base model requires ~50MB per adapter instead of storing 100 full 7GB weights.
Asynchronous DecouplingThe client receives an immediate 202 Accepted with a JobId. Heavy compute executes out-of-process.Synchronous model training over HTTP guarantees timeouts, socket exhaustion, and cascading failures under load.
Side-by-Side Inference GatewayMounting adapters dynamically onto a shared base model allows rapid A/B testing and direct before/after domain comparison.Avoids spinning up dedicated GPU containers for every custom-trained model variant.

Developer Quickstart Workflow (Command Line & APIs)

For engineers wanting to run each microservice manually:

# 1. Bring up Infrastructure (RabbitMQ + MLflow) and seed sample datasets
./scripts/dev-up.sh

# 2. Run the Ingestion API & Web Studio (.NET 10)
cd src/FtaaSService.Api && dotnet run

# 3. In another terminal, run the Training Worker (Python)
cd src/FtaaSService.Worker && source .venv/bin/activate && python consumer.py

# 4. In another terminal, run the Inference Engine
cd src/FtaaSService.Inference && source .venv/bin/activate && python app.py

# 5. Submit a fine-tuning job via curl
curl -X POST http://localhost:5100/api/v1/jobs \
  -F "file=@data/datasets/sample-support-compliance.jsonl" \
  -F "jobName=compliance-v1"

Automated Test Suites & Quality Engineering

The platform includes formal, automated unit and regression test suites across both the .NET control plane and Python compute layers:

1. .NET 10 API Test Suite (xUnit + Coverlet)

Validates compliance boundaries, data sanitization, and state machine idempotency:

  • PII Compliance Gateway: Delimited SSA SSNs, contextual SSNs, and Luhn-valid payment cards.
  • False-Positive Immunity: Verifies that non-contextual 9-digit integers (order IDs, invoice numbers) pass unhindered.
  • Full-Row Unmapped Column Scanning: Scans all metadata columns/properties; strips unmapped fields upon normalization.
  • Zero-Disk In-Memory Guarantee: Verifies zero bytes are written to disk upon compliance rejection.
  • Model Catalog & Validation: Enforces allowlisted base models via GET /api/v1/studio/models, ensuring invalid or unverified model identifiers are rejected at ingestion with 400 Bad Request.
  • Ingestion Ceilings: Enforces 25 MB file size and 50,000 record limits.
  • State Machine Idempotency: Rejects out-of-order sequence updates and prevents terminal state regressions in SQLite.
# Run all .NET unit tests with code coverage collection:
dotnet test tests/FtaaSService.Api.Tests --collect:"XPlat Code Coverage"

2. Python Worker & Inference Tests (unittest / pytest)

Validates model serving resilience, registry lookups, and file integrity:

  • Model Registry & Dynamic LoRA Targeting: Verifies target module selection (["q_proj", "v_proj"] vs ["q_proj", "v_proj", "k_proj", "o_proj"]) and template rendering across supported model architectures.
  • Inference Model Mismatch Protection: Returns actionable 409 Conflict responses when adapters are mounted against mismatched base models.
  • Hardware Preflight & Licensing: Intercepts Hugging Face gated license/auth failures (401/403) and surfaces actionable user remediation instructions.
  • Adapter Weight Integrity: Rejects truncated, zero-byte, or incomplete writes (<100 KB weights, <10 bytes config).
  • Bounded LRU Cache Eviction: Verifies dynamic eviction of least-recently-used LoRA adapters under memory pressure.
# Run all Python unit tests:
PYTHONPATH=src/FtaaSService.Worker:src/FtaaSService.Inference src/FtaaSService.Worker/.venv/bin/python -m pytest tests/

Automated End-to-End Integration Verification

To test the entire live pipeline (Infra $\rightarrow$ .NET 10 Ingestion $\rightarrow$ RabbitMQ $\rightarrow$ PyTorch/PEFT Training on MPS $\rightarrow$ MLflow $\rightarrow$ Dynamic Inference Comparison) in one command:

./scripts/verify-e2e.sh

Roadmap & Implementation Status

  • Architecture Specification & Design Contracts (docs/ARCHITECTURE.md)
  • Phase 1: Local Infrastructure Foundation (docker-compose.yml, scripts/dev-up.sh, RabbitMQ + MLflow healthchecks)
  • Phase 2: Ingestion & Control Plane (.NET 10 Minimal API, JSONL validation, SQLite state machine, AMQP producer/consumer)
  • Phase 3: Python Compute Worker & LoRA Pipeline (AMQP consumer, PyTorch MPS/CPU detection, PEFT/LoRA, MLflow telemetry)
  • Phase 4: Dynamic Model Serving & Side-by-Side Comparison (LoRA dynamic adapter mounting, comparison API)
  • Portfolio Phase 3a (The FTaaS Bridge Exporter): Edge ONNX Exporter (src/FtaaSService.Worker/exporter.py, scripts/export_edge_adapter.py) and Web API export endpoints (GET/POST /api/v1/jobs/{id}/export/edge) compiling LoRA adapters into web-optimized ONNX format with integrity checksums, documented in docs/SMOL_PROTOTYPE_RESULTS.md.
  • Portfolio Phase 3b (In-Browser Execution): Ingesting and executing exported ONNX packages directly inside client browser engines via WebGPU/WASM (onnxruntime-web), single-input ONNX export signature, and dynamic INT8 quantization.
  • Portfolio Phase 3c (Multi-Model Scaling & Gemma 2B Prototype): Dynamic multi-model registry (model_registry.py), google/gemma-2-2b-it support, hardware preflight warnings, float16 MPS optimization with gradient accumulation, and live benchmark evaluation documented in docs/GEMMA_PROTOTYPE_RESULTS.md.
  • Portfolio Phase 3d (Ternary Weight Scaling & BitNet b1.58 Prototype): Microsoft microsoft/BitNet-b1.58-2B-4T 1.58-bit ternary foundation model, native CPU execution via AVX2 integer addition/subtraction (zero GPU required), and empirical scaling benchmarks documented in docs/BITNET_PROTOTYPE_RESULTS.md.
  • BitNet b1.58 Ternary Integration & Engineering Roadmap: Multi-phase engineering roadmap covering native C++ kernel build, ternary LoRA training, multi-model catalog integration, upstream toolchain contributions, and dual-backend inference serving documented in docs/ROADMAP.md.
  • Comparative Performance Analysis: Empirical cross-architecture differentials across all three models documented in docs/MODEL_PERFORMANCE_COMPARISON.md.

πŸ“š Repository Documentation Index

All technical design documents, prototype evaluations, and implementation roadmaps are organized under docs/:

DocumentPurpose
docs/ARCHITECTURE.mdEnd-to-end system architecture, polyglot microservice boundaries, AMQP messaging contracts, and cloud mapping.
docs/ROADMAP.md7-phase BitNet b1.58 ternary integration roadmap, execution status, and upstream contribution tracking.
docs/STUDIO_GUIDE.mdNon-technical visual walkthrough and operational runbook for the FTaaS Enterprise Studio web application.
docs/MODEL_PERFORMANCE_COMPARISON.mdSymmetrical cross-architecture benchmark differentials across SmolLM2-135M, Gemma 2 2B IT, and BitNet b1.58.
docs/SMOL_PROTOTYPE_RESULTS.mdUltra-compact baseline (135M) empirical LoRA training metrics, ONNX export profile, and in-browser WASM execution.
docs/GEMMA_PROTOTYPE_RESULTS.mdHigh-capacity reasoning tier (2.6B) empirical benchmarks, Apple Silicon MPS tuning, and licensing guards.
docs/BITNET_PROTOTYPE_RESULTS.md1.58-bit ternary CPU execution (2.4B) AVX2 SIMD scaling benchmarks, memory profiling, and PyTorch LoRA metrics.
docs/PRIVACY_TELEMETRY_SCHEMA.mdPortfolio governance standard for strict attribute allowlists, burn-to-purge invariants, and OpenTelemetry overlay.
NOTICE.mdRoot statutory disclaimers, consumer privacy invariants, and third-party foundation model attribution.

Production Cloud Parity (AWS & GCP)

Local StackAWS Production ArchitectureGCP Production Architecture
.NET 10 API GatewayAmazon API Gateway + ECS FargateGoogle Cloud Run (.NET 10)
RabbitMQ Broker + DLQAmazon SQS (FIFO) + SQS DLQCloud Pub/Sub + Dead Letter
Storage (Datasets & DB)Amazon S3 + Aurora PostgreSQLGoogle Cloud Storage + Cloud SQL
Compute Worker (Python)SageMaker Training Jobs (Spot Instances)Vertex AI Custom Jobs (Preemptible)
Experiment TelemetryManaged MLflow / SageMaker ExperimentsVertex AI Experiments
Model RegistryMLflow Model Registry / SageMaker RegistryVertex AI Model Registry
Dynamic Inference ServerSageMaker Multi-Model Endpoints / TritonVertex AI Endpoints (vLLM / Triton)

License

This project is licensed under the MIT License.

Languages

Python

41.5%

C#

29.8%

JavaScript

13.8%

CSS

5.5%

HTML

5.5%

Shell

3.3%