Independent benchmark: MN-Violet-Lotus-12B β€” 51.0% across 13 benchmarks

#7
by mergekit-AI - opened

Hi FallenMerick,

We recently evaluated FallenMerick/MN-Violet-Lotus-12B on the MergeKit Cloud benchmark suite.

The model completed all 13 benchmarks, covering 2,560 sampled questions, with a 90% confidence interval.

Overall score: 51.0%

Some of the strongest results were:

  • Global-MMLU-Lite β€” 70.0%
  • PubMedQA β€” 66.1%
  • Shopping MMLU β€” 65.3%
  • MATH-500 β€” 61.1%
  • MedQA β€” 57.0%
  • IFEval β€” 56.8%

Other results:

  • INCLUDE-lite-44 β€” 50.0%
  • CodeXGLUE β€” 50.0%
  • FinQA β€” 48.6%
  • MMLU-Pro β€” 48.5%
  • LegalBench (CUAD) β€” 46.3%
  • GPQA Diamond β€” 31.9%
  • BFCL V3 β€” 11.5%

We thought the capability distribution was particularly interesting, so we wanted to share the results with you and the community.

The evaluation was run through MergeKit Cloud's 13-benchmark evaluation system.

We're happy to provide the full evaluation details/results if useful.

Screenshot 2026-09-26 at 12.14.30β€―AM
Screenshot 2026-09-26 at 12.14.44β€―AM

Sign up or log in to comment