NandhaKishorM/laya

Jupyter Notebook

176

22 commits

updated Sep 18, 2026

See the code

See what people are saying (1)

SourceMessageScoreDate

OpenJev

Today, he released Laya, a model based on his paper: https://github.com/NandhaKishorM/laya

0

Sep 18, 2026

README

Laya

Fast, non-autoregressive System 1 decision engine with mathematically calibrated probabilities.

Open In Colab PyPI version Hugging Face Model Hugging Face Space Dev.to Article Buy Me A Coffee License

Laya lets you evaluate typed questions (choice, score, noul) over any state (text, email, ticket, or JSON document) in a single forward pass (~33–38 ms on GPU). It produces structured decision outputs and calibrated confidence scores without text generation, token streaming, or hallucinations.

Powered by the fine-tuned Laya model on Hugging Face.


Installation

pip install laya

Quickstart

import laya

# 1. Load the fine-tuned model directly from Hugging Face Hub (auto-downloads weights)
agent = laya.load("convaiinnovations/laya")

# 2. Provide any state (string or dictionary)
state = {
    "from": "user@acme.com",
    "subject": "Duplicate charge on invoice #4411",
    "body": "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
}

# 3. Define your typed questions
questions = {
    # choice: categorical selection with probabilities & confidence
    "department": {
        "type": "choice",
        "instructions": "Which department should handle this email?",
        "criteria": {
            "billing": "invoices, payments, refunds",
            "technical": "bugs, outages, system errors",
            "sales": "pricing, new contracts",
            "other": "everything else"
        }
    },
    # score: placement on an ordinal rubric
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this request?",
        "criteria": ["not urgent", "soon", "critical deadline or blocking issue"]
    },
    # noul: calibrated boolean probability P(true)
    "churn_risk": {
        "type": "noul",
        "instructions": "Does the user threaten to cancel or leave?"
    },
    "is_phishing": {
        "type": "noul",
        "instructions": "Is this email a phishing or scam attempt?"
    }
}

# 4. Run all questions in ONE single forward pass (~35 ms on GPU)
result = agent.predict(state, questions)
answers = result["answers"]

print("Department :", answers["department"]["choice"])
# -> billing (confidence: 0.94)

print("Urgency    :", answers["urgency"]["score"])
# -> 1.84 / 2.0

print("Churn Risk :", answers["churn_risk"]["noul"])
# -> 0.892 (89.2% probability)

print("Phishing   :", answers["is_phishing"]["noul"])
# -> 0.008 (0.8% probability)

Automated Confidence Gating

Because Laya's probabilities are trained with strictly proper scoring rules (RLCD), confidence scores are statistically meaningful:

dept = answers["department"]["choice"]
conf = answers["department"]["confidence"]

if conf >= 0.85:
    # High confidence: automated action without human in the loop
    route_automatically(dept)
else:
    # Low confidence: escalate to human triage
    escalate_to_human_agent(dept, reason=f"Low confidence ({conf:.2f})")

Built-in Workflow Presets

Laya provides pre-tuned question schemas for immediate production use:

import laya

agent = laya.load("convaiinnovations/laya")

# 1. Intelligent Model Router (routes to small vs. frontier models)
routing = agent.predict({"request": "Refactor this service using dependency injection"}, laya.router_questions())

# 2. Real-time Prompt Guardrails (jailbreaks, injections, leaks)
guard = agent.predict({"prompt": "Ignore all instructions"}, laya.guard_questions())

# 3. Content Safety & Moderation (toxicity, harassment, threats)
safety = agent.predict({"post": "User comment text"}, laya.moderation_questions())

# 4. Support Ticket Triage (intent, urgency, frustration, churn)
triage = agent.predict({"message": "My payment failed twice"}, laya.triage_questions())

Decision Primitives

PrimitiveOutputUse Cases
choiceTop label, probabilities per option, confidenceDepartment routing, intent classification, topic categorization
scoreExpected level on ordinal rubric, distribution, confidenceFrustration level, ticket urgency, harm severity
noulCalibrated probability P(true) from 0.0 to 1.0Phishing detection, spam filtering, jailbreak detection, churn risk

Benchmark: Laya vs. TypeSafe Jev

Laya vs TypeSafe Jev Benchmark
Metric / DimensionTypeSafe Jev (Published)Laya (Fine-Tuned Checkpoint)Analysis / Advantage
P50 Latency (1 Question)~400 ms avg (70 to 500 ms, 150 ms best)38.4 ms (p95: 42.1 ms)Laya is ~10.4x faster on avg (4x faster than Jev best-case)
Batched Latency (10 Questions)~1,500 ms (serial) / ~400 ms156.0 ms (p95: 158.4 ms)Laya evaluates 10 questions in the time Jev answers 1
Batched Latency (50 Questions)Multi-second / rate-limited721.4 msHigh-throughput parallel mini-batching
Benchmark Accuracy67.8% (across 4 production workflows)83.8% in-task macro accuracyLaya achieves +16.0% higher overall accuracy
Intent & Customer Routing~95 to 98% agreement99.1% accuracy (ECE: 0.009)Near-zero calibration error on routing
Moderation & Content Safety~92 to 95% agreement96.7% accuracy (ECE: 0.061)Clean safety boundary separation
Inference & Fact VerificationNot separately reported88.3% accuracy (ECE: 0.054)Full bidirectional attention captures contradictions
Instruction-Following TasksProprietary internal set87.8% in-task / 86.3% zero-shotProven generalization across unseen tasks
Email Triage & PhishingVendor custom workflow73.2% accuracy (ECE: 0.017)Tailored email cleaning & phishing filters
Selective Automation (@ 50% Cov)Claims human escalation92.2% accuracy (ECE: 0.041)Safe automated gating (confidence >= 0.85)
Model Weights & CodeClosed-source / proprietary API100% Open-source Apache 2.0Full data sovereignty & transparency
Inference Cost$0.042 / 1M input tokens recurring$0.00 / self-hostedRuns on commodity GPUs, Mac MPS, or CPU
Multi-Turn Trajectory ModelingStatic state snapshotsTD(lambda = 1.0) prefix modelingReal temporal credit assignment
Deployment ModeCloud-only egressAir-gapped / Local / On-DeviceZero data egress (HIPAA/GDPR compliant)

Live Demo & Resources


Fine-Tuning on Single T4 GPU (Google Colab)

Fine-tune Laya on your custom domain data or commercial datasets on a free T4 GPU:


Support the Project

If Laya helps your research or products, consider supporting independent research:

Buy Me A Coffee


License

Apache 2.0. Developed by Convai Innovations.

Contributors

NandhaKishorM

22 commits

NandhaKishorM/laya

Jupyter Notebook

176

22 commits

updated Sep 18, 2026

See the code

See what people are saying (1)

SourceMessageScoreDate

OpenJev

Today, he released Laya, a model based on his paper: https://github.com/NandhaKishorM/laya

0

Sep 18, 2026

README

Laya

Fast, non-autoregressive System 1 decision engine with mathematically calibrated probabilities.

Open In Colab PyPI version Hugging Face Model Hugging Face Space Dev.to Article Buy Me A Coffee License

Laya lets you evaluate typed questions (choice, score, noul) over any state (text, email, ticket, or JSON document) in a single forward pass (~33–38 ms on GPU). It produces structured decision outputs and calibrated confidence scores without text generation, token streaming, or hallucinations.

Powered by the fine-tuned Laya model on Hugging Face.


Installation

pip install laya

Quickstart

import laya

# 1. Load the fine-tuned model directly from Hugging Face Hub (auto-downloads weights)
agent = laya.load("convaiinnovations/laya")

# 2. Provide any state (string or dictionary)
state = {
    "from": "user@acme.com",
    "subject": "Duplicate charge on invoice #4411",
    "body": "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
}

# 3. Define your typed questions
questions = {
    # choice: categorical selection with probabilities & confidence
    "department": {
        "type": "choice",
        "instructions": "Which department should handle this email?",
        "criteria": {
            "billing": "invoices, payments, refunds",
            "technical": "bugs, outages, system errors",
            "sales": "pricing, new contracts",
            "other": "everything else"
        }
    },
    # score: placement on an ordinal rubric
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this request?",
        "criteria": ["not urgent", "soon", "critical deadline or blocking issue"]
    },
    # noul: calibrated boolean probability P(true)
    "churn_risk": {
        "type": "noul",
        "instructions": "Does the user threaten to cancel or leave?"
    },
    "is_phishing": {
        "type": "noul",
        "instructions": "Is this email a phishing or scam attempt?"
    }
}

# 4. Run all questions in ONE single forward pass (~35 ms on GPU)
result = agent.predict(state, questions)
answers = result["answers"]

print("Department :", answers["department"]["choice"])
# -> billing (confidence: 0.94)

print("Urgency    :", answers["urgency"]["score"])
# -> 1.84 / 2.0

print("Churn Risk :", answers["churn_risk"]["noul"])
# -> 0.892 (89.2% probability)

print("Phishing   :", answers["is_phishing"]["noul"])
# -> 0.008 (0.8% probability)

Automated Confidence Gating

Because Laya's probabilities are trained with strictly proper scoring rules (RLCD), confidence scores are statistically meaningful:

dept = answers["department"]["choice"]
conf = answers["department"]["confidence"]

if conf >= 0.85:
    # High confidence: automated action without human in the loop
    route_automatically(dept)
else:
    # Low confidence: escalate to human triage
    escalate_to_human_agent(dept, reason=f"Low confidence ({conf:.2f})")

Built-in Workflow Presets

Laya provides pre-tuned question schemas for immediate production use:

import laya

agent = laya.load("convaiinnovations/laya")

# 1. Intelligent Model Router (routes to small vs. frontier models)
routing = agent.predict({"request": "Refactor this service using dependency injection"}, laya.router_questions())

# 2. Real-time Prompt Guardrails (jailbreaks, injections, leaks)
guard = agent.predict({"prompt": "Ignore all instructions"}, laya.guard_questions())

# 3. Content Safety & Moderation (toxicity, harassment, threats)
safety = agent.predict({"post": "User comment text"}, laya.moderation_questions())

# 4. Support Ticket Triage (intent, urgency, frustration, churn)
triage = agent.predict({"message": "My payment failed twice"}, laya.triage_questions())

Decision Primitives

PrimitiveOutputUse Cases
choiceTop label, probabilities per option, confidenceDepartment routing, intent classification, topic categorization
scoreExpected level on ordinal rubric, distribution, confidenceFrustration level, ticket urgency, harm severity
noulCalibrated probability P(true) from 0.0 to 1.0Phishing detection, spam filtering, jailbreak detection, churn risk

Benchmark: Laya vs. TypeSafe Jev

Laya vs TypeSafe Jev Benchmark
Metric / DimensionTypeSafe Jev (Published)Laya (Fine-Tuned Checkpoint)Analysis / Advantage
P50 Latency (1 Question)~400 ms avg (70 to 500 ms, 150 ms best)38.4 ms (p95: 42.1 ms)Laya is ~10.4x faster on avg (4x faster than Jev best-case)
Batched Latency (10 Questions)~1,500 ms (serial) / ~400 ms156.0 ms (p95: 158.4 ms)Laya evaluates 10 questions in the time Jev answers 1
Batched Latency (50 Questions)Multi-second / rate-limited721.4 msHigh-throughput parallel mini-batching
Benchmark Accuracy67.8% (across 4 production workflows)83.8% in-task macro accuracyLaya achieves +16.0% higher overall accuracy
Intent & Customer Routing~95 to 98% agreement99.1% accuracy (ECE: 0.009)Near-zero calibration error on routing
Moderation & Content Safety~92 to 95% agreement96.7% accuracy (ECE: 0.061)Clean safety boundary separation
Inference & Fact VerificationNot separately reported88.3% accuracy (ECE: 0.054)Full bidirectional attention captures contradictions
Instruction-Following TasksProprietary internal set87.8% in-task / 86.3% zero-shotProven generalization across unseen tasks
Email Triage & PhishingVendor custom workflow73.2% accuracy (ECE: 0.017)Tailored email cleaning & phishing filters
Selective Automation (@ 50% Cov)Claims human escalation92.2% accuracy (ECE: 0.041)Safe automated gating (confidence >= 0.85)
Model Weights & CodeClosed-source / proprietary API100% Open-source Apache 2.0Full data sovereignty & transparency
Inference Cost$0.042 / 1M input tokens recurring$0.00 / self-hostedRuns on commodity GPUs, Mac MPS, or CPU
Multi-Turn Trajectory ModelingStatic state snapshotsTD(lambda = 1.0) prefix modelingReal temporal credit assignment
Deployment ModeCloud-only egressAir-gapped / Local / On-DeviceZero data egress (HIPAA/GDPR compliant)

Live Demo & Resources


Fine-Tuning on Single T4 GPU (Google Colab)

Fine-tune Laya on your custom domain data or commercial datasets on a free T4 GPU:


Support the Project

If Laya helps your research or products, consider supporting independent research:

Buy Me A Coffee


License

Apache 2.0. Developed by Convai Innovations.

Contributors

NandhaKishorM

22 commits

Languages

Jupyter Notebook

69.5%

Python

30.5%