Papers
arxiv:2609.31927

Recipe-Matching, Not Equivalence

Published on Sep 25
Authors:
,
,
,
,

Abstract

MathNet-Retrieve asks a retriever to find, for a math problem, a document stating the same problem. An LLM under one fixed prompt writes each gold document and its near-miss distractors; LLM judges filter them. We call this procedure the "recipe", training on pairs built the same way "recipe-matching", and ask how much score it buys beyond the ability the benchmark claims to test. Two models from one base, matched in rows and settings, differ only in the training file: pairs written under the benchmark's published prompt by another vendor's LLM and judge, or computer-algebra-verified pairs with no LLM anywhere. The first leads by 45 R@1 points on the easy tier. By a non-LLM paraphrase control, half to two thirds of that gap comes from the pairs being LLM-written at all: LLM rewrites under two unrelated prompts, with the verified model's negatives, recover 30 and 22 of the 45 points; back-translations with the same negatives recover almost none. The remaining 15 to 25 points appear only under the benchmark's own prompt and vanish on real duplicates no generator wrote, the same problem in two languages. The hard tier rewards the recipe's pair structure, a deep rewrite against a minimal-edit near-miss: LLM rewrites alone score zero on it, attaching negatives unlocks it, and every negative that does so costs cross-language points; the sets scoring highest on it separate near-misses no LLM wrote worse than LLM rewrites with verified negatives. MELD also moves when a model trains on pairs built its way, without losing retention; on SABER-Math the registered attack fails, and the one gain, from its LLM-written summaries, is small but holds at a matched budget. Only on MathNet-Retrieve could we pin an inversion, benchmark score up and real retention down, to one edit of a training file. We release the generator-free duplicate evaluations, the near-miss test and three trained models.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.31927
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 10

Browse 10 models citing this paper

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.31927 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.