Spaces:
Sleeping
Sleeping
What LLM layer should we measure next?
#1
by genia-dev - opened
This demo measures whether a layer built on top of an LLM improves or destroys individual answers.
It reports paired GAIN, LOSS, or INCONCLUSIVE outcomes instead of relying only on aggregate benchmark scores.
Which system should we support next: a RAG pipeline, an agent workflow, a model router, or another verification layer?