Post
54
AI evaluation results can look more certain than they actually are.
A run finishes. A dashboard shows PASS. A score gets copied into another system. A few steps later, it may no longer be clear which metric produced it, what was skipped, what population was actually tested, or whether that PASS came from the evaluator at all.
That is the problem behind a new individual IETF Internet-Draft I published today: Claim-Preserving Exchange of AI Evaluation Evidence.
The basic idea is simple:
moving evidence should not make the claim stronger than the evidence itself.
Technical review and counterexamples are very welcome:
https://datatracker.ietf.org/doc/html/draft-abak-ai-evaluation-claim-preservation
A run finishes. A dashboard shows PASS. A score gets copied into another system. A few steps later, it may no longer be clear which metric produced it, what was skipped, what population was actually tested, or whether that PASS came from the evaluator at all.
That is the problem behind a new individual IETF Internet-Draft I published today: Claim-Preserving Exchange of AI Evaluation Evidence.
The basic idea is simple:
moving evidence should not make the claim stronger than the evidence itself.
Technical review and counterexamples are very welcome:
https://datatracker.ietf.org/doc/html/draft-abak-ai-evaluation-claim-preservation