Spaces:
Running
Download raw/method.md from DineshAI/3WPDFjZ1UT: direct link, hf CLI and curl.
- Browser
- Download file 8.34 kB
-
https://huggingface.co/spaces/DineshAI/3WPDFjZ1UT/resolve/main/raw/method.md
- Command line
-
hf download hf://spaces/DineshAI/3WPDFjZ1UT/raw/method.md
-
curl -L -o method.md https://huggingface.co/spaces/DineshAI/3WPDFjZ1UT/resolve/main/raw/method.md
Method
One fixed run command, inherited unchanged by every node of the experiment tree:
uv run python repro/src/run_all.py
All variants live in committed code. Two environment variables do affect what runs, and
an earlier version of this sentence wrongly denied it: CARE_OFFICIAL_DIR selects the
authors' checkout, and CARE_ENTRY selects the shard entrypoint inside the job
bootstrap. Neither changes a claim's result; both are recorded here rather than left as
a false absolute (see Limitations item 19). Prior benchmark shards ran on Hugging Face
cpu-upgrade; the fixed-entrypoint release run used local 8-CPU compute.
repro/src/threads.py is imported before numpy/scipy/torch and pins
every BLAS/OpenMP pool to the container's real cgroup quota, without which these
jobs run 20–40× slower than they should.
Claims 1–3 — Tables 1 and 2
CARE's aggregation is deterministic linear algebra on a fixed n × p judge-score
matrix. Producing that matrix is the expensive step and the paper reports doing it
on an A100. The authors released the matrices for ASSET, CivilComments and
PKU-BETTER and for nothing else. Two of those three columns are reproduced
end-to-end at full scale with the authors' own code (scripts/fully_gaussian_main.py
and scripts/gaussian_mixture_main.py, official repo pinned at 72f5b29), over
five seeds {2024, 2025, 2026, 2027, 2028}, with all nine Table 2 methods
(MV, AVG, WS, UWS, Dawid–Skene, GLAD, MACE, CARE-SVD, CARE-Tensor).
Independently, the arithmetic content of the three claims — 26.8 %, 17.37 %,
12.75 %, 13.4 %, and "best on 5 of 6" — is decided exactly against the published
tables, including which definition of "average relative improvement" the paper
actually used. This is done twice: once in claim_c123_benchmarks.py in floating
point, and once in independent_check.py in exact Fraction arithmetic against separate
hand transcriptions of Tables 1 and 2. Claim 3's official generated 0.814/0.705 pair is
evaluated literally; changing only 0.705 to the actual strongest baseline 0.718 is a
repair control that recovers the paper's nearby 13.4% prose.
Negative control. Each judge column of the ASSET matrix is independently row-permuted. This preserves every judge's marginal distribution but destroys the shared latent structure CARE exploits, so CARE's advantage must disappear. A control that still passed would show the advantage is not coming from confounder-aware aggregation.
Blocked. UltraFeedback, Summarize, FeedbackQA, Review-5K, Yelp, Chatbot-Arena, PKU-SAFER and SHP have no released judge-score matrices. Regenerating one costs 11–20 LLM judges (0.6 B–14 B) over 5,000 examples; Appendix E.2 puts that at up to 3 hours per dataset on an A100. GPU spend is not authorised for this campaign, so those columns are recorded BLOCKED with that exact missing capability rather than substituted by a synthetic proxy.
Claim 4 — Proposition 4.1
Finite experiments cannot settle a universally quantified statement, so the route taken is an independently reconstructed derivation plus assumption-satisfying counterexamples.
- Theorem D.3, symbolically.
K_JHis built as the firsthcolumns of a Householder reflector with a symbolic parameter vector, soK_JHᵀK_JH = I_hholds as a rational identity rather than at one numeric point.sympythen provesL k_i = λ_i k_iandrank(L − λ_i I) = p − 1for(p,h)in{(3,2),(4,2),(4,3),(5,3),(6,4)}. - Theorem D.4's constant, derived not assumed. Writing
M = [K W]ᵀE, rowiand columniofMeach have norm ≤‖E‖₂. Cauchy–Schwarz on the exact first-order eigenvector perturbation givesratio² ≤ (1+s)² + (1−s²) ≤ 4fors ∈ [0,1], i.e. a first-order constant of 2, strictly tighter than the paper's 4. The supremum is then measured by adversarial optimisation overEacross six spectra including near-degenerate gaps, and separately checked at finite‖E‖on 400 random models. - Counterexamples to the main-text restatement. Proposition 4.1 assumes only
orthogonal columns. With
K_JH = [√2 e₁, e₂],K'_JH = [√2 f₁, f₂],f₁ = (e₁+e₂)/√2,f₂ = (e₁−e₂)/√2andK_HH = diag(2,1), both satisfy every stated hypothesis and give the sameL = I₂, yet their columns are not related by sign and permutation. Separately, rescalingK_JH → c K_JHleaves‖K_HH^{-1}‖₂fixed, multipliesδ_ibyc²and the true error by1/c, so the main-text bound is violated by a factor growing linearly inc— the appendix proof needs the‖K_JH‖₂factor that the main text drops.
Claim 5 — Theorem 4.2
- Sign counterexample. At exact recovery,
-uis an equally valid eigenvector. D.5's raw distance is 2 against a zero right-hand side; the sign-aligned control used correctly in D.4 is zero. - Missing-zero-gap family. For
L*=2u₁u₁ᵀ+a u_hu_hᵀandE=(a/r)(u_hvᵀ+vu_hᵀ)withvin the nullspace, the last eigenvector's error is independent ofa. D.5's gap2-amakes the normalized violation diverge like1/a; the required full-spectrum gapakeeps it bounded. A separate 2x2 implementation independently reproduces every row. - Assumption control. Set
n=(r/a)^2; the minimum signal divided by then^-1/2noise is fixed atr/sqrt(p), so a sufficiently large fixedrsatisfies any finite Chandrasekaran signal constant while the violation still diverges. - Estimator-independent lower bound. Two CARE-factorizable Gaussian precision
models rotate only the weak direction. With
n=(c/a)^2their exact n-sample KL stays below 0.033, so Le Cam forces error with probability above 0.436 on one model while D.5 permits only 0.0996 failure and its positive-eigenvalue-gap rate tends to zero. Using the correct gap to zero restores the nonvanishing rate scale.
The prior symbolic composition, constant search, rate sweeps, and negative controls remain in the run and archive. They are not load-bearing for the literal falsification.
Claim 6 — Theorem 4.3
- Derivation. The paper's own eq. (8), (10)–(12) and (11) are composed in
sympy. Composing (10) with (8) reproduces the stated mean bound exactly. Composing (11) with (8) givesC_π C σ_max³ sqrt(p log(p/ε)/n), while the theorem statesC₂ sqrt(p log(p/ε)/n): a factor ofσ_max³is missing. - Measurement of the bound's own quantity. The mixture of Assumption D.8 is
simulated with
μ,π,δ,pandεall frozen and onlyσ_maxvarying, withnset at the theorem's own thresholdn = n₀ σ_max⁶. Along that boundary the stated bound decays likeσ_max^{-3}while the proof chain predicts aσ-free error, because the relative perturbation‖E‖_op/δis constant there. The recovery uses the algorithm the assumption names — multi-view moments plus Anandkumar et al.'s robust tensor power method with deflation — not a nearby substitute. - Robustness to the unknown constants. The violation factor grows like
σ_max³for any fixedC₁, C₂, so no choice of universal constants rescues the stated weight bound.
Negative controls. Over-sampling far past the boundary must drive the error
down; freezing n while raising σ_max must drive it up. Either failing would
mean the measurement is saturated rather than informative.
Independent checker
repro/src/independent_check.py re-derives the load-bearing numbers by different
routes: exact Fraction arithmetic on second transcriptions of Tables 1 and 2;
60-digit mpmath for the Proposition 4.1 counterexample; central finite
differences against the analytic first-order perturbation formula; direct 2x2
diagonalization of the D.5 missing-gap family; direct trace/log-determinant Gaussian KL;
and a Theil–Sen slope for the Theorem 4.3 sweep instead of least squares.
Exit contract
run_all.py exits 1 if any claim contract or the independent checker fails, and
prints the complete verdict JSON to stdout between ===CARE_VERDICT_BEGIN=== and
===CARE_VERDICT_END===; it also writes the local artifact used to render this release.