Spaces:
Running
Running
|
Download raw/method.md from DineshAI/3WPDFjZ1UT: direct link, hf CLI and curl.
- Browser
- Download file 8.34 kB
-
https://huggingface.co/spaces/DineshAI/3WPDFjZ1UT/resolve/main/raw/method.md
- Command line
-
hf download hf://spaces/DineshAI/3WPDFjZ1UT/raw/method.md
-
curl -L -o method.md https://huggingface.co/spaces/DineshAI/3WPDFjZ1UT/resolve/main/raw/method.md
8.34 kB
| # Method | |
| One fixed run command, inherited unchanged by every node of the experiment tree: | |
| ```bash | |
| uv run python repro/src/run_all.py | |
| ``` | |
| All variants live in committed code. Two environment variables do affect what runs, and | |
| an earlier version of this sentence wrongly denied it: `CARE_OFFICIAL_DIR` selects the | |
| authors' checkout, and `CARE_ENTRY` selects the shard entrypoint inside the job | |
| bootstrap. Neither changes a claim's result; both are recorded here rather than left as | |
| a false absolute (see Limitations item 19). Prior benchmark shards ran on Hugging Face | |
| `cpu-upgrade`; the fixed-entrypoint release run used local 8-CPU compute. | |
| `repro/src/threads.py` is imported before numpy/scipy/torch and pins | |
| every BLAS/OpenMP pool to the container's real cgroup quota, without which these | |
| jobs run 20–40× slower than they should. | |
| ## Claims 1–3 — Tables 1 and 2 | |
| CARE's aggregation is deterministic linear algebra on a fixed `n × p` judge-score | |
| matrix. Producing that matrix is the expensive step and the paper reports doing it | |
| on an A100. The authors released the matrices for **ASSET**, **CivilComments** and | |
| **PKU-BETTER** and for nothing else. Two of those three columns are reproduced | |
| end-to-end at full scale with the authors' own code (`scripts/fully_gaussian_main.py` | |
| and `scripts/gaussian_mixture_main.py`, official repo pinned at `72f5b29`), over | |
| five seeds `{2024, 2025, 2026, 2027, 2028}`, with all nine Table 2 methods | |
| (MV, AVG, WS, UWS, Dawid–Skene, GLAD, MACE, CARE-SVD, CARE-Tensor). | |
| Independently, the *arithmetic* content of the three claims — 26.8 %, 17.37 %, | |
| 12.75 %, 13.4 %, and "best on 5 of 6" — is decided exactly against the published | |
| tables, including which definition of "average relative improvement" the paper | |
| actually used. This is done twice: once in `claim_c123_benchmarks.py` in floating | |
| point, and once in `independent_check.py` in exact `Fraction` arithmetic against separate | |
| hand transcriptions of Tables 1 and 2. Claim 3's official generated `0.814/0.705` pair is | |
| evaluated literally; changing only 0.705 to the actual strongest baseline 0.718 is a | |
| repair control that recovers the paper's nearby 13.4% prose. | |
| **Negative control.** Each judge column of the ASSET matrix is independently | |
| row-permuted. This preserves every judge's marginal distribution but destroys the | |
| shared latent structure CARE exploits, so CARE's advantage must disappear. A | |
| control that still passed would show the advantage is not coming from | |
| confounder-aware aggregation. | |
| **Blocked.** UltraFeedback, Summarize, FeedbackQA, Review-5K, Yelp, Chatbot-Arena, | |
| PKU-SAFER and SHP have no released judge-score matrices. Regenerating one costs | |
| 11–20 LLM judges (0.6 B–14 B) over 5,000 examples; Appendix E.2 puts that at up to | |
| 3 hours per dataset on an A100. GPU spend is not authorised for this campaign, so | |
| those columns are recorded BLOCKED with that exact missing capability rather than | |
| substituted by a synthetic proxy. | |
| ## Claim 4 — Proposition 4.1 | |
| Finite experiments cannot settle a universally quantified statement, so the route | |
| taken is an independently reconstructed derivation plus assumption-satisfying | |
| counterexamples. | |
| 1. **Theorem D.3, symbolically.** `K_JH` is built as the first `h` columns of a | |
| Householder reflector with a symbolic parameter vector, so `K_JHᵀK_JH = I_h` | |
| holds as a rational identity rather than at one numeric point. `sympy` then | |
| proves `L k_i = λ_i k_i` and `rank(L − λ_i I) = p − 1` for `(p,h)` in | |
| `{(3,2),(4,2),(4,3),(5,3),(6,4)}`. | |
| 2. **Theorem D.4's constant, derived not assumed.** Writing `M = [K W]ᵀE`, row `i` | |
| and column `i` of `M` each have norm ≤ `‖E‖₂`. Cauchy–Schwarz on the exact | |
| first-order eigenvector perturbation gives | |
| `ratio² ≤ (1+s)² + (1−s²) ≤ 4` for `s ∈ [0,1]`, i.e. a first-order constant of | |
| **2**, strictly tighter than the paper's 4. The supremum is then *measured* by | |
| adversarial optimisation over `E` across six spectra including near-degenerate | |
| gaps, and separately checked at finite `‖E‖` on 400 random models. | |
| 3. **Counterexamples to the main-text restatement.** Proposition 4.1 assumes only | |
| *orthogonal* columns. With `K_JH = [√2 e₁, e₂]`, `K'_JH = [√2 f₁, f₂]`, | |
| `f₁ = (e₁+e₂)/√2`, `f₂ = (e₁−e₂)/√2` and `K_HH = diag(2,1)`, both satisfy every | |
| stated hypothesis and give the same `L = I₂`, yet their columns are not related | |
| by sign and permutation. Separately, rescaling `K_JH → c K_JH` leaves | |
| `‖K_HH^{-1}‖₂` fixed, multiplies `δ_i` by `c²` and the true error by `1/c`, so | |
| the main-text bound is violated by a factor growing linearly in `c` — the | |
| appendix proof needs the `‖K_JH‖₂` factor that the main text drops. | |
| ## Claim 5 — Theorem 4.2 | |
| 1. **Sign counterexample.** At exact recovery, `-u` is an equally valid eigenvector. | |
| D.5's raw distance is 2 against a zero right-hand side; the sign-aligned control used | |
| correctly in D.4 is zero. | |
| 2. **Missing-zero-gap family.** For `L*=2u₁u₁ᵀ+a u_hu_hᵀ` and | |
| `E=(a/r)(u_hvᵀ+vu_hᵀ)` with `v` in the nullspace, the last eigenvector's error is | |
| independent of `a`. D.5's gap `2-a` makes the normalized violation diverge like | |
| `1/a`; the required full-spectrum gap `a` keeps it bounded. A separate 2x2 | |
| implementation independently reproduces every row. | |
| 3. **Assumption control.** Set `n=(r/a)^2`; the minimum signal divided by the | |
| `n^-1/2` noise is fixed at `r/sqrt(p)`, so a sufficiently large fixed `r` satisfies | |
| any finite Chandrasekaran signal constant while the violation still diverges. | |
| 4. **Estimator-independent lower bound.** Two CARE-factorizable Gaussian precision | |
| models rotate only the weak direction. With `n=(c/a)^2` their exact n-sample KL stays | |
| below 0.033, so Le Cam forces error with probability above 0.436 on one model while | |
| D.5 permits only 0.0996 failure and its positive-eigenvalue-gap rate tends to zero. | |
| Using the correct gap to zero restores the nonvanishing rate scale. | |
| The prior symbolic composition, constant search, rate sweeps, and negative controls | |
| remain in the run and archive. They are not load-bearing for the literal falsification. | |
| ## Claim 6 — Theorem 4.3 | |
| 1. **Derivation.** The paper's own eq. (8), (10)–(12) and (11) are composed in | |
| `sympy`. Composing (10) with (8) reproduces the stated mean bound exactly. | |
| Composing (11) with (8) gives `C_π C σ_max³ sqrt(p log(p/ε)/n)`, while the | |
| theorem states `C₂ sqrt(p log(p/ε)/n)`: a factor of `σ_max³` is missing. | |
| 2. **Measurement of the bound's own quantity.** The mixture of Assumption D.8 is | |
| simulated with `μ`, `π`, `δ`, `p` and `ε` all frozen and only `σ_max` varying, | |
| with `n` set at the theorem's own threshold `n = n₀ σ_max⁶`. Along that boundary | |
| the *stated* bound decays like `σ_max^{-3}` while the proof chain predicts a | |
| `σ`-free error, because the relative perturbation `‖E‖_op/δ` is constant there. | |
| The recovery uses the algorithm the assumption names — multi-view moments plus | |
| Anandkumar et al.'s robust tensor power method with deflation — not a nearby | |
| substitute. | |
| 3. **Robustness to the unknown constants.** The violation factor grows like | |
| `σ_max³` for *any* fixed `C₁, C₂`, so no choice of universal constants rescues | |
| the stated weight bound. | |
| **Negative controls.** Over-sampling far past the boundary must drive the error | |
| down; freezing `n` while raising `σ_max` must drive it up. Either failing would | |
| mean the measurement is saturated rather than informative. | |
| ## Independent checker | |
| `repro/src/independent_check.py` re-derives the load-bearing numbers by different | |
| routes: exact `Fraction` arithmetic on second transcriptions of Tables 1 and 2; | |
| 60-digit `mpmath` for the Proposition 4.1 counterexample; central finite | |
| differences against the analytic first-order perturbation formula; direct 2x2 | |
| diagonalization of the D.5 missing-gap family; direct trace/log-determinant Gaussian KL; | |
| and a Theil–Sen slope for the Theorem 4.3 sweep instead of least squares. | |
| ## Exit contract | |
| `run_all.py` exits `1` if any claim contract or the independent checker fails, and | |
| prints the complete verdict JSON to stdout between `===CARE_VERDICT_BEGIN===` and | |
| `===CARE_VERDICT_END===`; it also writes the local artifact used to render this release. | |