WaterSheep

WaterSheep is a small decision model. Give it a text and a question with its possible answers, and it returns the answer with a probability for every option. Four question types are supported: yes/no, single choice, rating scale and select-all-that-apply. It is a 154M-parameter encoder that answers in one forward pass, so it also runs on a CPU.

Version 0.1.0 (watersheep-20260928-125452). Code: https://github.com/SamratDuttaOfficial/WaterSheep

Usage

pip install git+https://github.com/SamratDuttaOfficial/WaterSheep
from watersheep import WaterSheep

ws = WaterSheep.load("samratduttaofficial/WaterSheep")
ws.decide("I was charged twice.", "Which team should handle this?",
          ["billing", "shipping", "support"])

decide returns the answer, its confidence and a probability for every option. Several questions about one text go in a single request:

ws.ask({
    "state": {"customer": "Priya (premium plan)",
              "message": "Charged twice for order #4411 and the package is 12 days late."},
    "questions": {
        "escalate": {"type": "noul", "instructions": "Should a human agent take over now?"},
        "team": {"type": "choice", "instructions": "Which team should handle this?",
                 "criteria": {"billing": "payments, refunds", "shipping": "delivery problems"}},
        "frustration": {"type": "score", "instructions": "How frustrated is the customer?",
                        "criteria": ["calm", "annoyed", "frustrated", "furious"]},
        "issues": {"type": "multi", "instructions": "Which issues are reported?",
                   "criteria": ["double charge", "late delivery", "damaged item"]},
    },
})
Type Question Answer
noul yes/no the probability of yes
choice one of 2-255 options the option, with a probability for each
score a rating scale of 2-10 levels the expected level, with a probability for each
multi select all that apply every option above the threshold, with probabilities

Inputs are cut to 512 tokens. Up to 10 options are scored in one pass; longer lists are scored in rounds.

Evaluation

Evaluation Accuracy ECE
In-distribution test split 77.8% 0.026
Held-out datasets, not seen in training 61.2% 0.043

ECE is the expected calibration error (lower is better).

Benchmarks

Benchmark Suite Questions Accuracy ECE In training data
goemotions sentiment 2000 22.4% 0.023 other split
hatecheck safety 2000 75.1% 0.139 no
legal_abercrombie legal 95 21.1% 0.316 no
legal_contract_nli_confidentiality_of_agreement legal 82 69.5% 0.177 no
legal_corporate_lobbying legal 490 68.4% 0.216 no
legal_cuad_audit_rights legal 1216 86.3% 0.041 no
legal_definition_classification legal 1337 56.9% 0.279 no
legal_function_of_decision_section legal 367 24.3% 0.245 no
legal_hearsay legal 94 56.4% 0.307 no
legal_overruling legal 2000 62.5% 0.151 no
legal_personal_jurisdiction legal 50 50.0% 0.160 no
legal_privacy_policy_qa legal 2000 58.9% 0.274 no
legal_proa legal 95 51.6% 0.379 no
legal_ucc_v_common_law legal 94 62.8% 0.171 no
prompt_injection safety 116 91.4% 0.079 other split
xstest safety 450 73.6% 0.140 no

Training

  • Base model: answerdotai/ModernBERT-base, fine-tuned with a decision head.
  • Data: public datasets with open licenses, mapped to the four question types (listed in NOTICE), plus synthetic decisions written and verified by Qwen3.5-4B.
  • Calibration: a temperature per question type, fitted on a validation split.

Limitations

  • English only.
  • Inputs longer than 512 tokens are truncated.
  • Rating-scale answers are less accurate than the other types.
  • The probabilities are calibrated on data like the training data; check them on your own.
  • Do not use it on its own for medical, legal, financial, hiring or other high-stakes decisions.

License

Apache 2.0 (see LICENSE). NOTICE credits the base model, the teacher model and the public datasets, which keep their own licenses.

Citation

@misc{watersheep,
  author = {Samrat Dutta},
  title  = {WaterSheep: a small decision model with calibrated answers},
  year   = {2026},
  url    = {https://huggingface.co/samratduttaofficial/WaterSheep}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for samratduttaofficial/WaterSheep

Quantized
(77)
this model

Datasets used to train samratduttaofficial/WaterSheep

Space using samratduttaofficial/WaterSheep 1