What LLM layer should we measure next?

#1
by genia-dev - opened

This demo measures whether a layer built on top of an LLM improves or destroys individual answers.

It reports paired GAIN, LOSS, or INCONCLUSIVE outcomes instead of relying only on aggregate benchmark scores.

Which system should we support next: a RAG pipeline, an agent workflow, a model router, or another verification layer?

Source:
https://github.com/sxc3030-eng/scaffold-harness

Sign up or log in to comment