event-driven-ftaas / ml)An enterprise-grade, asynchronous, event-driven machine learning platform demonstrating a polyglot microservice pattern:
HuggingFaceTB/SmolLM2-135M) for zero-cost instant local training, high-capacity reasoning models (google/gemma-2-2b-it) with hardware preflight gates and float16 MPS optimization, and 1.58-bit ternary foundation models (microsoft/BitNet-b1.58-2B-4T) executing natively on CPU via AVX2 integer addition/subtraction.Watch the complete end-to-end studio workflow in actionβfrom side-by-side compliance policy comparison to live PII guardrails and event pipeline tracking:

πΉ High-Definition Recording:
demo.mp4(Playwright automated run, 1366x860)To regenerate this demonstration at any time, run:
./scripts/record-demo.sh
[!TIP] Zero-Code Operation: You do not need to write Python, manage PyTorch, or configure machine learning pipelines to evaluate FTaaS.
Prerequisite: .NET 10 SDK installed on your machine.
Run one command to launch the full interactive web application:
./scripts/start-studio.sh
Then open your browser to http://localhost:5100.
π Looking for a guided walkthrough? Read the FTaaS Enterprise Studio Guide for step-by-step instructions, visual explanations, and FAQs.
.csv document. Select your target base model (SmolLM2-135M for lightweight deployment or Gemma 2 2B IT for deeper reasoning capacity) with built-in hardware disclaimers. Click "Train & Deploy Team Adapter" to submit the job in milliseconds without waiting on background compute.[!NOTE] Preview vs. Live Compute: When launched standalone via
start-studio.sh, the studio operates in Interactive Preview Mode with pre-formatted demonstration outputs so you can evaluate the interface without spinning up Docker containers or downloading multi-gigabyte models. To run live on-device GPU inference, start Docker and the Python services via./scripts/dev-up.sh.
[!CAUTION] Data Privacy & Guardrail Scope: Automated heuristic checks detect delimited SSNs and Luhn-valid credit card numbers across all columns (including unmapped metadata). However, this is a best-effort defense, not an exhaustive DLP certification tool. Always sanitize datasets prior to training.
FTaaS follows the portfolio governance standard documented in PRIVACY_TELEMETRY_SCHEMA.md:
memory_only mode with zero outbound telemetry egress.job.accepted β job.published β job.consumed β job.training.started β job.training.finished β job.registered.cuda | mps | cpu), and adapter byte sizes. No raw prompts, completions, or dataset contents are ever permitted in spans.GET /api/v1/telemetry/privacy-audit: Real-time inspection of active telemetry mode and buffer state.GET /api/v1/telemetry/spans: Inspect buffered in-memory spans.POST /api/v1/telemetry/burn: Immediately purges the in-memory telemetry buffer.OTEL_EXPORTER_OTLP_ENDPOINT exports spans to an enterprise collector while maintaining the strict payload scrubbing invariant.[Non-Tech Web Studio / Client]
β
βΌ (POST /api/v1/jobs multipart or datasetPath)
[Ingestion Gateway (.NET 10)] βββΊ Validates JSONL, Computes SHA-256 Hash, Stores on Disk
β
βΌ (Job Request Event with datasetPath & datasetHash)
[Async Queue / Message Broker (RabbitMQ)]
β
βΌ
[Worker / Orchestrator (Python)] βββΊ Spin up Training Compute (PyTorch / Hugging Face LoRA)
β β
βΌ βΌ
[Model Registry (MLflow)] [Metrics & Loss Telemetry]
β
βΌ
[Dynamic Inference Engine] βββββββββ Base Model + On-the-fly LoRA Adapter Mounting
β
βΌ
[Side-by-Side Comparison] ββββββββββ Base Completion vs Fine-Tuned Completion
| Architectural Decision | Trade-Off Rationale | Enterprise Reality |
|---|---|---|
| Polyglot (.NET 10 + Python) | .NET offers high-concurrency, strongly-typed contracts, and low-latency API handling; Python owns the cutting-edge ML ecosystem (PyTorch, PEFT). | Solves the common anti-pattern of writing web APIs in Python or trying to run ML training in C#. Leverages each ecosystem's primary strength. |
| Out-of-Band Dataset Storage | Storing .jsonl files on disk/object storage and passing only datasetPath + SHA-256 hash over RabbitMQ prevents broker bloat and memory pressure. | Real datasets are megabytes to gigabytes. Message brokers degrade rapidly when payloads exceed hundreds of kilobytes. |
| LoRA (PEFT) vs Full Fine-Tuning | Freezing 99%+ of base model weights and only training low-rank adapter matrices reduces VRAM requirements by >80% and completes in 2β4 minutes on consumer hardware. | In production, training 100 domain adapters on a single base model requires ~50MB per adapter instead of storing 100 full 7GB weights. |
| Asynchronous Decoupling | The client receives an immediate 202 Accepted with a JobId. Heavy compute executes out-of-process. | Synchronous model training over HTTP guarantees timeouts, socket exhaustion, and cascading failures under load. |
| Side-by-Side Inference Gateway | Mounting adapters dynamically onto a shared base model allows rapid A/B testing and direct before/after domain comparison. | Avoids spinning up dedicated GPU containers for every custom-trained model variant. |
For engineers wanting to run each microservice manually:
# 1. Bring up Infrastructure (RabbitMQ + MLflow) and seed sample datasets
./scripts/dev-up.sh
# 2. Run the Ingestion API & Web Studio (.NET 10)
cd src/FtaaSService.Api && dotnet run
# 3. In another terminal, run the Training Worker (Python)
cd src/FtaaSService.Worker && source .venv/bin/activate && python consumer.py
# 4. In another terminal, run the Inference Engine
cd src/FtaaSService.Inference && source .venv/bin/activate && python app.py
# 5. Submit a fine-tuning job via curl
curl -X POST http://localhost:5100/api/v1/jobs \
-F "file=@data/datasets/sample-support-compliance.jsonl" \
-F "jobName=compliance-v1"
The platform includes formal, automated unit and regression test suites across both the .NET control plane and Python compute layers:
Validates compliance boundaries, data sanitization, and state machine idempotency:
GET /api/v1/studio/models, ensuring invalid or unverified model identifiers are rejected at ingestion with 400 Bad Request.# Run all .NET unit tests with code coverage collection:
dotnet test tests/FtaaSService.Api.Tests --collect:"XPlat Code Coverage"
Validates model serving resilience, registry lookups, and file integrity:
["q_proj", "v_proj"] vs ["q_proj", "v_proj", "k_proj", "o_proj"]) and template rendering across supported model architectures.# Run all Python unit tests:
PYTHONPATH=src/FtaaSService.Worker:src/FtaaSService.Inference src/FtaaSService.Worker/.venv/bin/python -m pytest tests/
To test the entire live pipeline (Infra $\rightarrow$ .NET 10 Ingestion $\rightarrow$ RabbitMQ $\rightarrow$ PyTorch/PEFT Training on MPS $\rightarrow$ MLflow $\rightarrow$ Dynamic Inference Comparison) in one command:
./scripts/verify-e2e.sh
docker-compose.yml, scripts/dev-up.sh, RabbitMQ + MLflow healthchecks)src/FtaaSService.Worker/exporter.py, scripts/export_edge_adapter.py) and Web API export endpoints (GET/POST /api/v1/jobs/{id}/export/edge) compiling LoRA adapters into web-optimized ONNX format with integrity checksums, documented in docs/SMOL_PROTOTYPE_RESULTS.md.onnxruntime-web), single-input ONNX export signature, and dynamic INT8 quantization.model_registry.py), google/gemma-2-2b-it support, hardware preflight warnings, float16 MPS optimization with gradient accumulation, and live benchmark evaluation documented in docs/GEMMA_PROTOTYPE_RESULTS.md.microsoft/BitNet-b1.58-2B-4T 1.58-bit ternary foundation model, native CPU execution via AVX2 integer addition/subtraction (zero GPU required), and empirical scaling benchmarks documented in docs/BITNET_PROTOTYPE_RESULTS.md.All technical design documents, prototype evaluations, and implementation roadmaps are organized under docs/:
| Document | Purpose |
|---|---|
docs/ARCHITECTURE.md | End-to-end system architecture, polyglot microservice boundaries, AMQP messaging contracts, and cloud mapping. |
docs/ROADMAP.md | 7-phase BitNet b1.58 ternary integration roadmap, execution status, and upstream contribution tracking. |
docs/STUDIO_GUIDE.md | Non-technical visual walkthrough and operational runbook for the FTaaS Enterprise Studio web application. |
docs/MODEL_PERFORMANCE_COMPARISON.md | Symmetrical cross-architecture benchmark differentials across SmolLM2-135M, Gemma 2 2B IT, and BitNet b1.58. |
docs/SMOL_PROTOTYPE_RESULTS.md | Ultra-compact baseline (135M) empirical LoRA training metrics, ONNX export profile, and in-browser WASM execution. |
docs/GEMMA_PROTOTYPE_RESULTS.md | High-capacity reasoning tier (2.6B) empirical benchmarks, Apple Silicon MPS tuning, and licensing guards. |
docs/BITNET_PROTOTYPE_RESULTS.md | 1.58-bit ternary CPU execution (2.4B) AVX2 SIMD scaling benchmarks, memory profiling, and PyTorch LoRA metrics. |
docs/PRIVACY_TELEMETRY_SCHEMA.md | Portfolio governance standard for strict attribute allowlists, burn-to-purge invariants, and OpenTelemetry overlay. |
NOTICE.md | Root statutory disclaimers, consumer privacy invariants, and third-party foundation model attribution. |
| Local Stack | AWS Production Architecture | GCP Production Architecture |
|---|---|---|
| .NET 10 API Gateway | Amazon API Gateway + ECS Fargate | Google Cloud Run (.NET 10) |
| RabbitMQ Broker + DLQ | Amazon SQS (FIFO) + SQS DLQ | Cloud Pub/Sub + Dead Letter |
| Storage (Datasets & DB) | Amazon S3 + Aurora PostgreSQL | Google Cloud Storage + Cloud SQL |
| Compute Worker (Python) | SageMaker Training Jobs (Spot Instances) | Vertex AI Custom Jobs (Preemptible) |
| Experiment Telemetry | Managed MLflow / SageMaker Experiments | Vertex AI Experiments |
| Model Registry | MLflow Model Registry / SageMaker Registry | Vertex AI Model Registry |
| Dynamic Inference Server | SageMaker Multi-Model Endpoints / Triton | Vertex AI Endpoints (vLLM / Triton) |
This project is licensed under the MIT License.
Python
41.5%
C#
29.8%
JavaScript
13.8%
CSS
5.5%
HTML
5.5%
Shell
3.3%
event-driven-ftaas / ml)An enterprise-grade, asynchronous, event-driven machine learning platform demonstrating a polyglot microservice pattern:
HuggingFaceTB/SmolLM2-135M) for zero-cost instant local training, high-capacity reasoning models (google/gemma-2-2b-it) with hardware preflight gates and float16 MPS optimization, and 1.58-bit ternary foundation models (microsoft/BitNet-b1.58-2B-4T) executing natively on CPU via AVX2 integer addition/subtraction.Watch the complete end-to-end studio workflow in actionβfrom side-by-side compliance policy comparison to live PII guardrails and event pipeline tracking:

πΉ High-Definition Recording:
demo.mp4(Playwright automated run, 1366x860)To regenerate this demonstration at any time, run:
./scripts/record-demo.sh
[!TIP] Zero-Code Operation: You do not need to write Python, manage PyTorch, or configure machine learning pipelines to evaluate FTaaS.
Prerequisite: .NET 10 SDK installed on your machine.
Run one command to launch the full interactive web application:
./scripts/start-studio.sh
Then open your browser to http://localhost:5100.
π Looking for a guided walkthrough? Read the FTaaS Enterprise Studio Guide for step-by-step instructions, visual explanations, and FAQs.
.csv document. Select your target base model (SmolLM2-135M for lightweight deployment or Gemma 2 2B IT for deeper reasoning capacity) with built-in hardware disclaimers. Click "Train & Deploy Team Adapter" to submit the job in milliseconds without waiting on background compute.[!NOTE] Preview vs. Live Compute: When launched standalone via
start-studio.sh, the studio operates in Interactive Preview Mode with pre-formatted demonstration outputs so you can evaluate the interface without spinning up Docker containers or downloading multi-gigabyte models. To run live on-device GPU inference, start Docker and the Python services via./scripts/dev-up.sh.
[!CAUTION] Data Privacy & Guardrail Scope: Automated heuristic checks detect delimited SSNs and Luhn-valid credit card numbers across all columns (including unmapped metadata). However, this is a best-effort defense, not an exhaustive DLP certification tool. Always sanitize datasets prior to training.
FTaaS follows the portfolio governance standard documented in PRIVACY_TELEMETRY_SCHEMA.md:
memory_only mode with zero outbound telemetry egress.job.accepted β job.published β job.consumed β job.training.started β job.training.finished β job.registered.cuda | mps | cpu), and adapter byte sizes. No raw prompts, completions, or dataset contents are ever permitted in spans.GET /api/v1/telemetry/privacy-audit: Real-time inspection of active telemetry mode and buffer state.GET /api/v1/telemetry/spans: Inspect buffered in-memory spans.POST /api/v1/telemetry/burn: Immediately purges the in-memory telemetry buffer.OTEL_EXPORTER_OTLP_ENDPOINT exports spans to an enterprise collector while maintaining the strict payload scrubbing invariant.[Non-Tech Web Studio / Client]
β
βΌ (POST /api/v1/jobs multipart or datasetPath)
[Ingestion Gateway (.NET 10)] βββΊ Validates JSONL, Computes SHA-256 Hash, Stores on Disk
β
βΌ (Job Request Event with datasetPath & datasetHash)
[Async Queue / Message Broker (RabbitMQ)]
β
βΌ
[Worker / Orchestrator (Python)] βββΊ Spin up Training Compute (PyTorch / Hugging Face LoRA)
β β
βΌ βΌ
[Model Registry (MLflow)] [Metrics & Loss Telemetry]
β
βΌ
[Dynamic Inference Engine] βββββββββ Base Model + On-the-fly LoRA Adapter Mounting
β
βΌ
[Side-by-Side Comparison] ββββββββββ Base Completion vs Fine-Tuned Completion
| Architectural Decision | Trade-Off Rationale | Enterprise Reality |
|---|---|---|
| Polyglot (.NET 10 + Python) | .NET offers high-concurrency, strongly-typed contracts, and low-latency API handling; Python owns the cutting-edge ML ecosystem (PyTorch, PEFT). | Solves the common anti-pattern of writing web APIs in Python or trying to run ML training in C#. Leverages each ecosystem's primary strength. |
| Out-of-Band Dataset Storage | Storing .jsonl files on disk/object storage and passing only datasetPath + SHA-256 hash over RabbitMQ prevents broker bloat and memory pressure. | Real datasets are megabytes to gigabytes. Message brokers degrade rapidly when payloads exceed hundreds of kilobytes. |
| LoRA (PEFT) vs Full Fine-Tuning | Freezing 99%+ of base model weights and only training low-rank adapter matrices reduces VRAM requirements by >80% and completes in 2β4 minutes on consumer hardware. | In production, training 100 domain adapters on a single base model requires ~50MB per adapter instead of storing 100 full 7GB weights. |
| Asynchronous Decoupling | The client receives an immediate 202 Accepted with a JobId. Heavy compute executes out-of-process. | Synchronous model training over HTTP guarantees timeouts, socket exhaustion, and cascading failures under load. |
| Side-by-Side Inference Gateway | Mounting adapters dynamically onto a shared base model allows rapid A/B testing and direct before/after domain comparison. | Avoids spinning up dedicated GPU containers for every custom-trained model variant. |
For engineers wanting to run each microservice manually:
# 1. Bring up Infrastructure (RabbitMQ + MLflow) and seed sample datasets
./scripts/dev-up.sh
# 2. Run the Ingestion API & Web Studio (.NET 10)
cd src/FtaaSService.Api && dotnet run
# 3. In another terminal, run the Training Worker (Python)
cd src/FtaaSService.Worker && source .venv/bin/activate && python consumer.py
# 4. In another terminal, run the Inference Engine
cd src/FtaaSService.Inference && source .venv/bin/activate && python app.py
# 5. Submit a fine-tuning job via curl
curl -X POST http://localhost:5100/api/v1/jobs \
-F "file=@data/datasets/sample-support-compliance.jsonl" \
-F "jobName=compliance-v1"
The platform includes formal, automated unit and regression test suites across both the .NET control plane and Python compute layers:
Validates compliance boundaries, data sanitization, and state machine idempotency:
GET /api/v1/studio/models, ensuring invalid or unverified model identifiers are rejected at ingestion with 400 Bad Request.# Run all .NET unit tests with code coverage collection:
dotnet test tests/FtaaSService.Api.Tests --collect:"XPlat Code Coverage"
Validates model serving resilience, registry lookups, and file integrity:
["q_proj", "v_proj"] vs ["q_proj", "v_proj", "k_proj", "o_proj"]) and template rendering across supported model architectures.# Run all Python unit tests:
PYTHONPATH=src/FtaaSService.Worker:src/FtaaSService.Inference src/FtaaSService.Worker/.venv/bin/python -m pytest tests/
To test the entire live pipeline (Infra $\rightarrow$ .NET 10 Ingestion $\rightarrow$ RabbitMQ $\rightarrow$ PyTorch/PEFT Training on MPS $\rightarrow$ MLflow $\rightarrow$ Dynamic Inference Comparison) in one command:
./scripts/verify-e2e.sh
docker-compose.yml, scripts/dev-up.sh, RabbitMQ + MLflow healthchecks)src/FtaaSService.Worker/exporter.py, scripts/export_edge_adapter.py) and Web API export endpoints (GET/POST /api/v1/jobs/{id}/export/edge) compiling LoRA adapters into web-optimized ONNX format with integrity checksums, documented in docs/SMOL_PROTOTYPE_RESULTS.md.onnxruntime-web), single-input ONNX export signature, and dynamic INT8 quantization.model_registry.py), google/gemma-2-2b-it support, hardware preflight warnings, float16 MPS optimization with gradient accumulation, and live benchmark evaluation documented in docs/GEMMA_PROTOTYPE_RESULTS.md.microsoft/BitNet-b1.58-2B-4T 1.58-bit ternary foundation model, native CPU execution via AVX2 integer addition/subtraction (zero GPU required), and empirical scaling benchmarks documented in docs/BITNET_PROTOTYPE_RESULTS.md.All technical design documents, prototype evaluations, and implementation roadmaps are organized under docs/:
| Document | Purpose |
|---|---|
docs/ARCHITECTURE.md | End-to-end system architecture, polyglot microservice boundaries, AMQP messaging contracts, and cloud mapping. |
docs/ROADMAP.md | 7-phase BitNet b1.58 ternary integration roadmap, execution status, and upstream contribution tracking. |
docs/STUDIO_GUIDE.md | Non-technical visual walkthrough and operational runbook for the FTaaS Enterprise Studio web application. |
docs/MODEL_PERFORMANCE_COMPARISON.md | Symmetrical cross-architecture benchmark differentials across SmolLM2-135M, Gemma 2 2B IT, and BitNet b1.58. |
docs/SMOL_PROTOTYPE_RESULTS.md | Ultra-compact baseline (135M) empirical LoRA training metrics, ONNX export profile, and in-browser WASM execution. |
docs/GEMMA_PROTOTYPE_RESULTS.md | High-capacity reasoning tier (2.6B) empirical benchmarks, Apple Silicon MPS tuning, and licensing guards. |
docs/BITNET_PROTOTYPE_RESULTS.md | 1.58-bit ternary CPU execution (2.4B) AVX2 SIMD scaling benchmarks, memory profiling, and PyTorch LoRA metrics. |
docs/PRIVACY_TELEMETRY_SCHEMA.md | Portfolio governance standard for strict attribute allowlists, burn-to-purge invariants, and OpenTelemetry overlay. |
NOTICE.md | Root statutory disclaimers, consumer privacy invariants, and third-party foundation model attribution. |
| Local Stack | AWS Production Architecture | GCP Production Architecture |
|---|---|---|
| .NET 10 API Gateway | Amazon API Gateway + ECS Fargate | Google Cloud Run (.NET 10) |
| RabbitMQ Broker + DLQ | Amazon SQS (FIFO) + SQS DLQ | Cloud Pub/Sub + Dead Letter |
| Storage (Datasets & DB) | Amazon S3 + Aurora PostgreSQL | Google Cloud Storage + Cloud SQL |
| Compute Worker (Python) | SageMaker Training Jobs (Spot Instances) | Vertex AI Custom Jobs (Preemptible) |
| Experiment Telemetry | Managed MLflow / SageMaker Experiments | Vertex AI Experiments |
| Model Registry | MLflow Model Registry / SageMaker Registry | Vertex AI Model Registry |
| Dynamic Inference Server | SageMaker Multi-Model Endpoints / Triton | Vertex AI Endpoints (vLLM / Triton) |
This project is licensed under the MIT License.
Python
41.5%
C#
29.8%
JavaScript
13.8%
CSS
5.5%
HTML
5.5%
Shell
3.3%