SamratDuttaOfficial/WaterSheep

A decision model that returns calibrated answers to yes/no, choice, score and multi-label questions.

Python

1

10 commits

updated Oct 2, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built an open-source alternative to Jev, and you can call it straight from JavaScript (r/webdev)

WaterSheep is an open-source alternative to TypeSafe's Jev. You ask it a typed question about some text (yes/no, pick one, rate it, or tag everything that applies) and get back an answer with a probability for every option. Two ways to use it in a web app: * **As a drop-in Jev backend.**…

0

Oct 2, 2026

README

WaterSheep

WaterSheep

Calibrated decisions for any text.

License Hugging Face Demo

Website · Demo · Model

WaterSheep answers yes/no, single-choice, rating and multi-label questions about any text, with a probability for every option.

Usage

pip install transformers torch
from transformers import pipeline

ws = pipeline(model="samratduttaofficial/WaterSheep", trust_remote_code=True)
ws("I was charged twice.", question="Which team should handle this?", options=["billing", "shipping", "support"])
TypeOptionsAnswer
noulnone (yes/no)probability of yes
choiceany labelsthe best option
scorea digit scale, e.g. 1 to 5the expected level
multiany labels, with type="multi"every option above the threshold

Every answer includes a probability for each option.

Using Jev?

WaterSheep is an open-source alternative to Jev. Run it as a local server:

pip install git+https://github.com/SamratDuttaOfficial/WaterSheep
watersheep --model samratduttaofficial/WaterSheep --serve

It answers Jev's POST /v1/systemone requests on your machine, and TypeSafe's Python SDK works against it without code changes:

export TYPESAFE_BASE_URL=http://127.0.0.1:8766

Any API key value works locally. Multi-label questions ("type": "multi") work too, as plain JSON. WaterSheep is independent and not affiliated with TypeSafe AI.

Download

hf download samratduttaofficial/WaterSheep --local-dir WaterSheep

Or with Git (requires Git LFS):

git clone https://huggingface.co/samratduttaofficial/WaterSheep

Then load it from the folder, offline:

ws = pipeline(model="WaterSheep", trust_remote_code=True)

API

Deploy it as a Hugging Face Inference Endpoint, then:

curl https://YOUR-ENDPOINT -H "Authorization: Bearer $HF_TOKEN" -H "Content-Type: application/json" -d '{"inputs": "I was charged twice.", "parameters": {"question": "Which team should handle this?", "options": ["billing", "shipping", "support"]}}'

JavaScript

No install; runs in the browser:

<script type="module">
  import { decide } from "https://samratduttaofficial.github.io/WaterSheep/watersheep.js";
  console.log(await decide("I was charged twice.", "Which team should handle this?", ["billing", "shipping", "support"]));
</script>

With a downloaded copy on your web server, call load({ base: "WaterSheep/" }) first.

Other languages: run onnx/model_quantized.onnx with ONNX Runtime; watersheep.js shows the input format.

Python package

pip install git+https://github.com/SamratDuttaOfficial/WaterSheep
from watersheep import WaterSheep

ws = WaterSheep.load("samratduttaofficial/WaterSheep")
ws.decide("I was charged twice.", "Which team should handle this?", ["billing", "shipping", "support"])

decide returns the answer, its confidence and a probability for every option. ask answers several questions about one text:

ws.ask({
    "state": {"customer": "Priya (premium plan)",
              "message": "Charged twice for order #4411 and the package is 12 days late."},
    "questions": {
        "escalate": {"type": "noul", "instructions": "Should a human agent take over now?"},
        "team": {"type": "choice", "instructions": "Which team should handle this?",
                 "criteria": {"billing": "payments, refunds", "shipping": "delivery problems"}},
        "frustration": {"type": "score", "instructions": "How frustrated is the customer?",
                        "criteria": ["calm", "annoyed", "frustrated", "furious"]},
        "issues": {"type": "multi", "instructions": "Which issues are reported?",
                   "criteria": ["double charge", "late delivery", "damaged item"]},
    },
})
TypeQuestionAnswer
noulyes/noprobability of yes
choicesingle choicethe option, with a probability for each
scorerating scalethe expected level, with a probability for each
multimulti-labelevery option above the threshold, with probabilities

Command line:

watersheep --model samratduttaofficial/WaterSheep --question "Which team should handle this?" --options billing,shipping,support --state "I was charged twice."

--serve runs a local HTTP API on port 8766.

Evaluation

EvaluationAccuracyECE
In-distribution test split77.8%0.026
Held-out datasets, not seen in training61.2%0.043

ECE is the expected calibration error (lower is better).

Calibration by question type

Accuracy against confidence for each question type, before (raw) and after calibration.

Benchmarks

BenchmarkSuiteQuestionsAccuracyECEIn training data
goemotionssentiment200022.4%0.023other split
hatechecksafety200075.1%0.139no
legal_abercrombielegal9521.1%0.316no
legal_contract_nli_confidentiality_of_agreementlegal8269.5%0.177no
legal_corporate_lobbyinglegal49068.4%0.216no
legal_cuad_audit_rightslegal121686.3%0.041no
legal_definition_classificationlegal133756.9%0.279no
legal_function_of_decision_sectionlegal36724.3%0.245no
legal_hearsaylegal9456.4%0.307no
legal_overrulinglegal200062.5%0.151no
legal_personal_jurisdictionlegal5050.0%0.160no
legal_privacy_policy_qalegal200058.9%0.274no
legal_proalegal9551.6%0.379no
legal_ucc_v_common_lawlegal9462.8%0.171no
prompt_injectionsafety11691.4%0.079other split
xstestsafety45073.6%0.140no

Training

  • Base model: answerdotai/ModernBERT-base, fine-tuned with a decision head.
  • Data: openly licensed public datasets (listed in NOTICE) and synthetic decisions from Qwen3.5-4B.
  • Calibration: a temperature per question type, fitted on a validation split.

Training curves

Training loss and learning rate (left); validation accuracy by question type (right).

Synthetic data verification

Share of synthetic examples kept after verification, by question type (left) and by family (right).

Limitations

  • English only.
  • Long inputs are truncated.
  • Rating-scale answers are less accurate than the other types.
  • Probabilities are calibrated on data like the training data; validate them on your own.
  • Not for high-stakes decisions (medical, legal, financial, hiring) on its own.

Train a new model

git clone https://github.com/SamratDuttaOfficial/WaterSheep
cd WaterSheep
./scripts/run.sh

Use scripts\run.bat on Windows.

Results

The figures, tables and data are in results/. To remake them after training and scripts/run-benchmarks.sh (.bat on Windows), run these from the project root with the Python in .venv:

ScriptNeedsWrites to results/
tools/results/data_stats.pya trained modeldata/data_stats.json
tools/results/make_figures.pydata_stats.py, benchmarksfigures/, data/synth_outcomes_by_type.json
tools/results/gen_tables.pydata_stats.pytables/sources.tex, tables/families.tex
tools/results/bench_table.pybenchmarkstables/bench.tex, tables/speed.tex, data/bench_summary.json

Each uses the newest model unless --model is given.

License

Apache 2.0 (LICENSE). Attributions: NOTICE.

Citation

@misc{watersheep,
  author = {Samrat Dutta},
  title  = {WaterSheep: calibrated decisions for any text},
  year   = {2026},
  url    = {https://huggingface.co/samratduttaofficial/WaterSheep}
}

Read the preprint Buy me a coffee

browser
calibration
decision-model
huggingface
javascript
jev
jev-alternative
local-first
machine-learning
multi-label-classification
nlp
onnx
onnxruntime-web
open-source
python
system-one
text-classification
transformers
webassembly
zero-shot-classification

SamratDuttaOfficial/WaterSheep

A decision model that returns calibrated answers to yes/no, choice, score and multi-label questions.

Python

1

10 commits

updated Oct 2, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built an open-source alternative to Jev, and you can call it straight from JavaScript (r/webdev)

WaterSheep is an open-source alternative to TypeSafe's Jev. You ask it a typed question about some text (yes/no, pick one, rate it, or tag everything that applies) and get back an answer with a probability for every option. Two ways to use it in a web app: * **As a drop-in Jev backend.**…

0

Oct 2, 2026

README

WaterSheep

WaterSheep

Calibrated decisions for any text.

License Hugging Face Demo

Website · Demo · Model

WaterSheep answers yes/no, single-choice, rating and multi-label questions about any text, with a probability for every option.

Usage

pip install transformers torch
from transformers import pipeline

ws = pipeline(model="samratduttaofficial/WaterSheep", trust_remote_code=True)
ws("I was charged twice.", question="Which team should handle this?", options=["billing", "shipping", "support"])
TypeOptionsAnswer
noulnone (yes/no)probability of yes
choiceany labelsthe best option
scorea digit scale, e.g. 1 to 5the expected level
multiany labels, with type="multi"every option above the threshold

Every answer includes a probability for each option.

Using Jev?

WaterSheep is an open-source alternative to Jev. Run it as a local server:

pip install git+https://github.com/SamratDuttaOfficial/WaterSheep
watersheep --model samratduttaofficial/WaterSheep --serve

It answers Jev's POST /v1/systemone requests on your machine, and TypeSafe's Python SDK works against it without code changes:

export TYPESAFE_BASE_URL=http://127.0.0.1:8766

Any API key value works locally. Multi-label questions ("type": "multi") work too, as plain JSON. WaterSheep is independent and not affiliated with TypeSafe AI.

Download

hf download samratduttaofficial/WaterSheep --local-dir WaterSheep

Or with Git (requires Git LFS):

git clone https://huggingface.co/samratduttaofficial/WaterSheep

Then load it from the folder, offline:

ws = pipeline(model="WaterSheep", trust_remote_code=True)

API

Deploy it as a Hugging Face Inference Endpoint, then:

curl https://YOUR-ENDPOINT -H "Authorization: Bearer $HF_TOKEN" -H "Content-Type: application/json" -d '{"inputs": "I was charged twice.", "parameters": {"question": "Which team should handle this?", "options": ["billing", "shipping", "support"]}}'

JavaScript

No install; runs in the browser:

<script type="module">
  import { decide } from "https://samratduttaofficial.github.io/WaterSheep/watersheep.js";
  console.log(await decide("I was charged twice.", "Which team should handle this?", ["billing", "shipping", "support"]));
</script>

With a downloaded copy on your web server, call load({ base: "WaterSheep/" }) first.

Other languages: run onnx/model_quantized.onnx with ONNX Runtime; watersheep.js shows the input format.

Python package

pip install git+https://github.com/SamratDuttaOfficial/WaterSheep
from watersheep import WaterSheep

ws = WaterSheep.load("samratduttaofficial/WaterSheep")
ws.decide("I was charged twice.", "Which team should handle this?", ["billing", "shipping", "support"])

decide returns the answer, its confidence and a probability for every option. ask answers several questions about one text:

ws.ask({
    "state": {"customer": "Priya (premium plan)",
              "message": "Charged twice for order #4411 and the package is 12 days late."},
    "questions": {
        "escalate": {"type": "noul", "instructions": "Should a human agent take over now?"},
        "team": {"type": "choice", "instructions": "Which team should handle this?",
                 "criteria": {"billing": "payments, refunds", "shipping": "delivery problems"}},
        "frustration": {"type": "score", "instructions": "How frustrated is the customer?",
                        "criteria": ["calm", "annoyed", "frustrated", "furious"]},
        "issues": {"type": "multi", "instructions": "Which issues are reported?",
                   "criteria": ["double charge", "late delivery", "damaged item"]},
    },
})
TypeQuestionAnswer
noulyes/noprobability of yes
choicesingle choicethe option, with a probability for each
scorerating scalethe expected level, with a probability for each
multimulti-labelevery option above the threshold, with probabilities

Command line:

watersheep --model samratduttaofficial/WaterSheep --question "Which team should handle this?" --options billing,shipping,support --state "I was charged twice."

--serve runs a local HTTP API on port 8766.

Evaluation

EvaluationAccuracyECE
In-distribution test split77.8%0.026
Held-out datasets, not seen in training61.2%0.043

ECE is the expected calibration error (lower is better).

Calibration by question type

Accuracy against confidence for each question type, before (raw) and after calibration.

Benchmarks

BenchmarkSuiteQuestionsAccuracyECEIn training data
goemotionssentiment200022.4%0.023other split
hatechecksafety200075.1%0.139no
legal_abercrombielegal9521.1%0.316no
legal_contract_nli_confidentiality_of_agreementlegal8269.5%0.177no
legal_corporate_lobbyinglegal49068.4%0.216no
legal_cuad_audit_rightslegal121686.3%0.041no
legal_definition_classificationlegal133756.9%0.279no
legal_function_of_decision_sectionlegal36724.3%0.245no
legal_hearsaylegal9456.4%0.307no
legal_overrulinglegal200062.5%0.151no
legal_personal_jurisdictionlegal5050.0%0.160no
legal_privacy_policy_qalegal200058.9%0.274no
legal_proalegal9551.6%0.379no
legal_ucc_v_common_lawlegal9462.8%0.171no
prompt_injectionsafety11691.4%0.079other split
xstestsafety45073.6%0.140no

Training

  • Base model: answerdotai/ModernBERT-base, fine-tuned with a decision head.
  • Data: openly licensed public datasets (listed in NOTICE) and synthetic decisions from Qwen3.5-4B.
  • Calibration: a temperature per question type, fitted on a validation split.

Training curves

Training loss and learning rate (left); validation accuracy by question type (right).

Synthetic data verification

Share of synthetic examples kept after verification, by question type (left) and by family (right).

Limitations

  • English only.
  • Long inputs are truncated.
  • Rating-scale answers are less accurate than the other types.
  • Probabilities are calibrated on data like the training data; validate them on your own.
  • Not for high-stakes decisions (medical, legal, financial, hiring) on its own.

Train a new model

git clone https://github.com/SamratDuttaOfficial/WaterSheep
cd WaterSheep
./scripts/run.sh

Use scripts\run.bat on Windows.

Results

The figures, tables and data are in results/. To remake them after training and scripts/run-benchmarks.sh (.bat on Windows), run these from the project root with the Python in .venv:

ScriptNeedsWrites to results/
tools/results/data_stats.pya trained modeldata/data_stats.json
tools/results/make_figures.pydata_stats.py, benchmarksfigures/, data/synth_outcomes_by_type.json
tools/results/gen_tables.pydata_stats.pytables/sources.tex, tables/families.tex
tools/results/bench_table.pybenchmarkstables/bench.tex, tables/speed.tex, data/bench_summary.json

Each uses the newest model unless --model is given.

License

Apache 2.0 (LICENSE). Attributions: NOTICE.

Citation

@misc{watersheep,
  author = {Samrat Dutta},
  title  = {WaterSheep: calibrated decisions for any text},
  year   = {2026},
  url    = {https://huggingface.co/samratduttaofficial/WaterSheep}
}

Read the preprint Buy me a coffee

browser
calibration
decision-model
huggingface
javascript
jev
jev-alternative
local-first
machine-learning
multi-label-classification
nlp
onnx
onnxruntime-web
open-source
python
system-one
text-classification
transformers
webassembly
zero-shot-classification

Languages

Python

94.6%

TeX

4.0%