samratduttaofficial/WaterSheep

Model

WaterSheep

1

4 commits

1 linked in READMEs

updated Oct 1, 2026

See the code

README

WaterSheep

WaterSheep

Website · Demo · Code

WaterSheep answers yes/no, single-choice, rating and multi-label questions about any text, with a probability for every option. Version 0.1.0 (watersheep-20260928-125452).

Usage

pip install transformers torch
from transformers import pipeline

ws = pipeline(model="samratduttaofficial/WaterSheep", trust_remote_code=True)
ws("I was charged twice.", question="Which team should handle this?", options=["billing", "shipping", "support"])
TypeOptionsAnswer
noulnone (yes/no)probability of yes
choiceany labelsthe best option
scorea digit scale, e.g. 1 to 5the expected level
multiany labels, with type="multi"every option above the threshold

Every answer includes a probability for each option.

Download

hf download samratduttaofficial/WaterSheep --local-dir WaterSheep

Or with Git (requires Git LFS):

git clone https://huggingface.co/samratduttaofficial/WaterSheep

Then load it from the folder, offline:

ws = pipeline(model="WaterSheep", trust_remote_code=True)

API

Deploy it as an Inference Endpoint, then:

curl https://YOUR-ENDPOINT -H "Authorization: Bearer $HF_TOKEN" -H "Content-Type: application/json" -d '{"inputs": "I was charged twice.", "parameters": {"question": "Which team should handle this?", "options": ["billing", "shipping", "support"]}}'

JavaScript

No install; runs in the browser:

<script type="module">
  import { decide } from "https://samratduttaofficial.github.io/WaterSheep/watersheep.js";
  console.log(await decide("I was charged twice.", "Which team should handle this?", ["billing", "shipping", "support"]));
</script>

With a downloaded copy on your web server, call load({ base: "WaterSheep/" }) first.

Other languages: run onnx/model_quantized.onnx with ONNX Runtime; watersheep.js shows the input format.

Evaluation

EvaluationAccuracyECE
In-distribution test split77.8%0.026
Held-out datasets, not seen in training61.2%0.043

ECE is the expected calibration error (lower is better).

Calibration by question type

Accuracy against confidence for each question type, before (raw) and after calibration.

Benchmarks

BenchmarkSuiteQuestionsAccuracyECEIn training data
goemotionssentiment200022.4%0.023other split
hatechecksafety200075.1%0.139no
legal_abercrombielegal9521.1%0.316no
legal_contract_nli_confidentiality_of_agreementlegal8269.5%0.177no
legal_corporate_lobbyinglegal49068.4%0.216no
legal_cuad_audit_rightslegal121686.3%0.041no
legal_definition_classificationlegal133756.9%0.279no
legal_function_of_decision_sectionlegal36724.3%0.245no
legal_hearsaylegal9456.4%0.307no
legal_overrulinglegal200062.5%0.151no
legal_personal_jurisdictionlegal5050.0%0.160no
legal_privacy_policy_qalegal200058.9%0.274no
legal_proalegal9551.6%0.379no
legal_ucc_v_common_lawlegal9462.8%0.171no
prompt_injectionsafety11691.4%0.079other split
xstestsafety45073.6%0.140no

Training

Training curves

Training loss and learning rate (left); validation accuracy by question type (right).

Synthetic data verification

Share of synthetic examples kept after verification, by question type (left) and by family (right).

Limitations

  • English only.
  • Long inputs are truncated.
  • Rating-scale answers are less accurate than the other types.
  • Probabilities are calibrated on data like the training data; validate them on your own.
  • Not for high-stakes decisions (medical, legal, financial, hiring) on its own.

License

Apache 2.0 (LICENSE). Trained on openly licensed data; credits in NOTICE.

Citation

@misc{watersheep,
  author = {Samrat Dutta},
  title  = {WaterSheep: calibrated decisions for any text},
  year   = {2026},
  url    = {https://huggingface.co/samratduttaofficial/WaterSheep}
}

Read the preprint Buy me a coffee

calibration
custom_code
decision-model
endpoints_compatible
feature-extraction
multi-label
onnx
safetensors
transformers
watersheep
zero-shot-classification

samratduttaofficial/WaterSheep

Model

WaterSheep

1

4 commits

1 linked in READMEs

updated Oct 1, 2026

See the code

README

WaterSheep

WaterSheep

Website · Demo · Code

WaterSheep answers yes/no, single-choice, rating and multi-label questions about any text, with a probability for every option. Version 0.1.0 (watersheep-20260928-125452).

Usage

pip install transformers torch
from transformers import pipeline

ws = pipeline(model="samratduttaofficial/WaterSheep", trust_remote_code=True)
ws("I was charged twice.", question="Which team should handle this?", options=["billing", "shipping", "support"])
TypeOptionsAnswer
noulnone (yes/no)probability of yes
choiceany labelsthe best option
scorea digit scale, e.g. 1 to 5the expected level
multiany labels, with type="multi"every option above the threshold

Every answer includes a probability for each option.

Download

hf download samratduttaofficial/WaterSheep --local-dir WaterSheep

Or with Git (requires Git LFS):

git clone https://huggingface.co/samratduttaofficial/WaterSheep

Then load it from the folder, offline:

ws = pipeline(model="WaterSheep", trust_remote_code=True)

API

Deploy it as an Inference Endpoint, then:

curl https://YOUR-ENDPOINT -H "Authorization: Bearer $HF_TOKEN" -H "Content-Type: application/json" -d '{"inputs": "I was charged twice.", "parameters": {"question": "Which team should handle this?", "options": ["billing", "shipping", "support"]}}'

JavaScript

No install; runs in the browser:

<script type="module">
  import { decide } from "https://samratduttaofficial.github.io/WaterSheep/watersheep.js";
  console.log(await decide("I was charged twice.", "Which team should handle this?", ["billing", "shipping", "support"]));
</script>

With a downloaded copy on your web server, call load({ base: "WaterSheep/" }) first.

Other languages: run onnx/model_quantized.onnx with ONNX Runtime; watersheep.js shows the input format.

Evaluation

EvaluationAccuracyECE
In-distribution test split77.8%0.026
Held-out datasets, not seen in training61.2%0.043

ECE is the expected calibration error (lower is better).

Calibration by question type

Accuracy against confidence for each question type, before (raw) and after calibration.

Benchmarks

BenchmarkSuiteQuestionsAccuracyECEIn training data
goemotionssentiment200022.4%0.023other split
hatechecksafety200075.1%0.139no
legal_abercrombielegal9521.1%0.316no
legal_contract_nli_confidentiality_of_agreementlegal8269.5%0.177no
legal_corporate_lobbyinglegal49068.4%0.216no
legal_cuad_audit_rightslegal121686.3%0.041no
legal_definition_classificationlegal133756.9%0.279no
legal_function_of_decision_sectionlegal36724.3%0.245no
legal_hearsaylegal9456.4%0.307no
legal_overrulinglegal200062.5%0.151no
legal_personal_jurisdictionlegal5050.0%0.160no
legal_privacy_policy_qalegal200058.9%0.274no
legal_proalegal9551.6%0.379no
legal_ucc_v_common_lawlegal9462.8%0.171no
prompt_injectionsafety11691.4%0.079other split
xstestsafety45073.6%0.140no

Training

Training curves

Training loss and learning rate (left); validation accuracy by question type (right).

Synthetic data verification

Share of synthetic examples kept after verification, by question type (left) and by family (right).

Limitations

  • English only.
  • Long inputs are truncated.
  • Rating-scale answers are less accurate than the other types.
  • Probabilities are calibrated on data like the training data; validate them on your own.
  • Not for high-stakes decisions (medical, legal, financial, hiring) on its own.

License

Apache 2.0 (LICENSE). Trained on openly licensed data; credits in NOTICE.

Citation

@misc{watersheep,
  author = {Samrat Dutta},
  title  = {WaterSheep: calibrated decisions for any text},
  year   = {2026},
  url    = {https://huggingface.co/samratduttaofficial/WaterSheep}
}

Read the preprint Buy me a coffee

calibration
custom_code
decision-model
endpoints_compatible
feature-extraction
multi-label
onnx
safetensors
transformers
watersheep
zero-shot-classification