Firemedic15's picture
download
raw
64.1 kB
<!DOCTYPE html>
<!--
Vendored + adapted into posterly from ARIS (Auto-claude-code-research-in-sleep),
skill paper-poster-html. Origin: https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep
Upstream tokenization: MIT (c) 2026 wanshuiyin -- ../LICENSES/aris-MIT.txt
This is the tokenized form of posterly's own neutral template; the class names and
content are posterly's own (c) 2026 Ruishuo Chen. As an adapted derivative this file
ships as part of posterly under AGPL-3.0. See ../NOTICE.md.
-->
<!--
============================================================
TEMPLATE: landscape_4col (ARIS fork)
CANVAS: 60in × 36in landscape (ICML / NeurIPS / generic landscape)
LAYOUT: header → optional banner → 4 body columns → optional takeaways → footer
USE WHEN: standard ML conference poster with 3-5 content cards per column,
mix of figures + equations + small tables.
HOW TO USE:
1. Copy this file to your working directory as `poster.html`.
2. Edit the DESIGN TOKENS in `:root` to your lab/venue colors.
3. Replace TODO content placeholders (search for "TODO").
4. Run `python tools/run_gates.py poster.html --tokens design_tokens.json` (the default driver: preflight -> style -> measure -> polish; design_tokens.json is the design-direction pack written at lock time -- SKILL.md Step 2.5) to align columns and check style/structure.
5. Run `python tools/render_preview.py poster.html` to produce the PDF.
NOTE: as shipped, this scaffold passes `preflight` (structure) but is
EXPECTED to FAIL `measure`/`polish` -- figures are commented out and copy
is TODO stubs, so columns only fill the top of the canvas. Those two gates
judge a FILLED poster; balance them once you've added real content. See
templates/README.md ("Scaffolds, not finished posters").
MEASURE ROLES: every layout-critical element carries `data-measure-role`.
Generic measurement scripts depend on these — do not remove.
CANVAS RETARGETING: to print at another size, change the canvas in exactly
TWO places that must stay in sync -- the `@page { size: ... }` rule and the
`.poster { width/height }` (also the `@media print .poster { width/height }`).
Example: ICLR 2026 main conference uses 185cm 90cm landscape (official print
service spec) -- set @page to `185cm 90cm`, width to `1850 * var(--u)` and
height to `900 * var(--u)` (and the print-override .poster to `185cm`/`90cm`).
Adapted from posterly (MIT, © 2026 Ruishuo Chen) — see LICENSES/ & NOTICE.md;
ARIS modifications: flat de-gradient, --fs token scale, zero-inline-style
utilities, data-source/data-color-exempt contracts.
============================================================
-->
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<meta name="generator" content="posterly">
<title>Reproducing Squirrel Benchmark — ICML 2026 Agent Reproduction</title>
<!-- MathJax v3 for inline equations. CDN by default when the file is
opened by hand; the posterly check tools (measure/polish/pack/…)
intercept this request and serve the skill's bundled copy
(assets/mathjax/tex-svg.js, MathJax 3.2.2), so gates render math
deterministically even offline. To make a hand-opened poster
offline too, copy that bundle next to the poster and point the
<script> `src` at it. -->
<script>
window.MathJax = {
tex: {
inlineMath: [['$', '$'], ['\\(', '\\)']],
displayMath: [['$$', '$$'], ['\\[', '\\]']],
packages: {'[+]': ['ams']}
},
svg: { fontCache: 'global' }
};
</script>
<script id="MathJax-script" async src="https://cdn.jsdelivr.net/npm/mathjax@3/es5/tex-svg.js"></script>
<style>
/* =========================================================
CANVAS — 60" × 36" landscape
========================================================= */
@page { size: 60in 36in; margin: 0; }
:root {
/* ===== DESIGN TOKENS ===== */
/* Primary accent (header underline, card-highlight bar, section nums, .keyword) */
--accent: #2D5F8B;
--accent-deep: #1F4566;
--accent-light: #E8F1F8;
--accent-soft: #D7E5F0;
--accent-ink: #FFFFFF; /* ink ON the accent (chips, callout, thead) */
/* Emphasis register for "ours / best" (table .ours row, "★" callouts).
Default is the warm gold; a redesign may choose any register that
keeps 4.5:1 with --emph-ink (see templates/THEMES.md Mechanism 1). */
--emph: #C9A24A;
--emph-soft: #FFF7E0;
--emph-ink: #14314A; /* ink ON the emphasis fill. NOT var(--accent-deep):
#1F4566 on the default gold measures 4.16:1 (< AA);
#14314A measures 5.58:1 (>= 4.5:1). */
/* Contrast-fix variants of --emph for two specific low-contrast hosts
(see style rule 1/3 + polish CONTRAST): a darker gold for text on the
light accent-tint card, a lighter gold tint for text on the solid
accent fill. Both still read as "the gold register", just re-tuned
per host background. */
--emph-on-tint: #8A6A1F; /* key-mark on --bg-emphasis (4.4:1) */
--emph-on-accent: #FFF3D0; /* callout > strong.qmark on --accent (6.1:1) */
/* Text */
--text-primary: #1A1A1A;
--text-secondary: #555555;
--text-muted: #888888;
/* Backgrounds */
--bg-page: #F6F2F0;
--bg-card: #FFFFFF;
--bg-card-tint: #FAFAFB;
--bg-emphasis: var(--accent-light);
/* Borders */
--border-soft: #D8D8D8;
--border-strong: var(--accent);
/* Screen-only dark mat behind the poster (the off-canvas viewport
background). Print resets html/body to white, so this never reaches
paper; kept as a token so no color literal lives outside this block. */
--bg-viewport: #2B2B2B;
/* Base unit. Print: 1mm; screen preview: 1.6px (~3.78px = 1mm at 96dpi). */
--u: 1.6px;
/* Font-size scale (9 archetypes). Every `font-size` in this template
references one of these. `calc(var(--fs-N) * k)` is permitted ONLY for
a COMPONENTS.md-defined variant (e.g. `.eqn--large`). */
--fs-1: calc(9 * var(--u)); /* micro label */
--fs-2: calc(10 * var(--u)); /* small caption */
--fs-3: calc(11 * var(--u)); /* caption / table */
--fs-4: calc(12 * var(--u)); /* body text */
--fs-5: calc(13 * var(--u)); /* equation / emphasis */
--fs-6: calc(15 * var(--u)); /* subtitle */
--fs-7: calc(16 * var(--u)); /* section title */
--fs-8: calc(22 * var(--u)); /* banner number */
--fs-9: calc(32 * var(--u)); /* main title */
/* Fonts. Override at :root if your venue mandates a specific family. */
--font-serif: "Charter", "Source Serif Pro", "Georgia", serif;
--font-sans: "Inter", "Helvetica Neue", sans-serif;
/* Shadows & watermark ink — tokenized so rules 1/3 stay literal-free
(the faint radial page tint is the ONE allowed literal outside this
block; rule 5 validates it WHEN ENABLED -- the style gate ships with rules 4-5 off). */
--shadow-screen: 0 0 60px rgba(0, 0, 0, 0.5);
--shadow-card: 0 calc(2 * var(--u)) calc(6 * var(--u)) rgba(45, 95, 139, 0.05);
--ornament-ink: rgba(45, 95, 139, 0.06);
/* Identity mark ink — the corner-signature ⊕ adopts the LOCAL accent so it
reads native (the woven-signature instead inherits its host's ink). The
glyph GEOMETRY, not the hue, is the posterly signature. Re-theme with the
palette like any other token. */
--ps-mark-ink: var(--accent-deep);
/* Corner-radius scale. Every border-radius calc() multiplies by this;
1 = the shipped soft look, 0 = square/flat (see templates/THEMES.md). */
--rs: 1;
/* Figure mount: the ground behind paper figures (transparent PNGs sit on
it) and the keyline around them (.figure img / .ff-fig img). Re-theme
with the Axis 6 frame decision so figures sit mounted in the design
instead of pasted on it; --fig-frame: transparent = frameless mount.
QR backgrounds stay literal white -- that's scannability, not styling. */
--fig-bg: white;
--fig-frame: var(--border-soft);
/* ===== END DESIGN TOKENS ===== */
}
/* =========================================================
RESET + BASE
========================================================= */
* { box-sizing: border-box; margin: 0; padding: 0; }
/* =========================================================
BASE DEFENSES -- wrap & ink safety. KEEP this block in ANY
skeleton, including a fully custom one (copy it over and
EXTEND the selector lists with your own prose/display
classes -- a hand-rolled skeleton that drops it strands
single-word widows and ragged titles; polish warns as
TEXT-WRAP).
- pretty: fills each line, protects the last-line orphan
(single-word widow). Safe on all prose.
- balance: evens CENTERED display text only -- never pair
it with left-aligned multi-sentence prose (SKILL.md wrap
rules).
- Any inline class you add that paints a background (a
.mark highlight, a keyword chip) must declare its own
`color` (Gate G -- never inherit ink across a ground
change) plus `box-decoration-break: clone;
-webkit-box-decoration-break: clone;` so a wrapped
highlight keeps its padding on both fragments (then it
never needs &nbsp;-gluing to hold one line).
========================================================= */
p, li, dd, figcaption,
.body-text, .caption, .callout, .section-title { text-wrap: pretty; }
.title { text-wrap: balance; }
html, body {
background: var(--bg-viewport);
font-family: var(--font-serif);
color: var(--text-primary);
-webkit-font-smoothing: antialiased;
}
/* =========================================================
UTILITY CLASSES — replace inline `style=` in the body markup.
The template (and any finished poster) must carry ZERO `style=`
attributes. Use these classes instead. Two documented exceptions
exist ONLY as examples and may appear in comments, never live:
- a logo/seal SVG flagged data-color-exempt="logo"
- a paper figure `<img ... data-source="paper" style="width: NN%">`
(AR width tweak; prefer the .w-NN classes below where possible)
========================================================= */
/* Poster container — exact print dimensions; data-measure-role="poster" so poster_check.py
can verify the canvas size. */
.poster {
width: calc(1524 * var(--u));
height: calc(914 * var(--u));
background: var(--bg-page);
/* Page tint kept (radial only, all color stops alpha <= 0.06 per the
de-gradient policy: linear-gradient is banned, this low-alpha radial
wash is the single allowed exception). */
background-image:
radial-gradient(ellipse at top left, rgba(45, 95, 139, 0.06), transparent 40%),
radial-gradient(ellipse at bottom right, rgba(201, 162, 74, 0.05), transparent 50%);
margin: 20px auto;
padding: calc(10 * var(--u)) calc(14 * var(--u));
display: grid;
grid-template-columns: minmax(0, 1fr); /* column-axis twin of the minmax(0,1fr) body-row defense: without it the implicit `auto` column grows to a wide child's max-content and a full-width row overflows the canvas (measure's canvas-overflow gate). Keep on any custom skeleton. */
grid-template-rows: auto auto minmax(0, 1fr) auto auto; /* header | banner | body | takeaways | footer. minmax(0,1fr) (not bare 1fr): an over-tall body compresses instead of pushing takeaways/footer off-canvas */
gap: calc(6 * var(--u));
box-shadow: var(--shadow-screen);
position: relative;
overflow: hidden;
}
/* Decorative top bar — flat solid accent (de-gradient). */
.poster::before {
content: "";
position: absolute; top: 0; left: 0; right: 0;
height: calc(8 * var(--u));
background: var(--accent);
}
/* =========================================================
HEADER
========================================================= */
.header {
display: grid;
grid-template-columns: 1fr minmax(50%, auto) 1fr; /* equal side tracks (1fr) -> the title track is centred on the poster, not just between the side blocks; the centre track is floored at 50% so a one-line title still fills it (else a short title measures narrow and trips a false HEADER/TITLE-SQUEEZED). Best-effort: a side block wide enough to clamp its 1fr track can still pull the title off-centre. */
align-items: center;
gap: calc(16 * var(--u));
padding: calc(2 * var(--u)) calc(4 * var(--u)) calc(5 * var(--u));
border-bottom: calc(2 * var(--u)) solid var(--accent);
}
.venue-badge {
justify-self: start; /* anchor to the far-left edge so the centre track stays centred */
display: flex; flex-direction: column;
align-items: center; justify-content: center;
min-width: calc(95 * var(--u));
text-align: center;
border-right: calc(1 * var(--u)) solid var(--border-soft);
padding-right: calc(12 * var(--u));
}
.venue-badge .vb-venue {
font-family: var(--font-sans);
font-weight: 800;
font-size: var(--fs-9);
color: var(--accent-deep);
line-height: 1;
letter-spacing: -0.5px;
}
.venue-badge .vb-year {
font-family: var(--font-sans);
font-size: var(--fs-5);
color: var(--text-secondary);
margin-top: calc(3 * var(--u));
letter-spacing: 1.2px;
}
.venue-badge .vb-tag {
font-family: var(--font-sans);
font-size: var(--fs-2);
color: var(--accent);
font-weight: 700;
margin-top: calc(2 * var(--u));
letter-spacing: 1.2px;
}
.title-block { text-align: center; min-width: 0; }
.title {
font-family: var(--font-sans);
font-weight: 800;
font-size: var(--fs-9);
line-height: 1.05;
color: var(--accent-deep);
letter-spacing: -0.5px;
}
.title .accent { color: var(--emph); }
.subtitle {
font-family: var(--font-sans);
font-weight: 500;
font-size: var(--fs-6);
color: var(--text-secondary);
margin-top: calc(2 * var(--u));
font-style: italic;
}
.authors-line {
font-family: var(--font-sans);
font-size: var(--fs-4);
color: var(--accent);
font-weight: 600;
margin-top: calc(3 * var(--u));
}
.authors-line .author { margin: 0 calc(4 * var(--u)); }
.authors-line sup { font-size: 0.7em; color: var(--accent); }
.authors-line .aff {
color: var(--text-secondary);
font-weight: 400;
display: block;
margin-top: calc(2 * var(--u));
font-size: var(--fs-4);
}
.right-block {
justify-self: end; /* keep logo+QR in the far-right corner */
display: flex; align-items: center;
gap: calc(10 * var(--u));
}
.qr-block { display: flex; flex-direction: column; align-items: center; gap: calc(2 * var(--u)); }
.qr-block img {
width: calc(85 * var(--u));
height: calc(85 * var(--u));
border: calc(2 * var(--u)) solid var(--accent);
border-radius: calc(4 * var(--u) * var(--rs));
background: white;
padding: calc(2 * var(--u));
}
.qr-label {
font-family: var(--font-sans);
font-size: var(--fs-3);
color: var(--accent);
font-weight: 600;
}
/* Optional logo slot — drop your lab logo here. Keep its rendered height
close to the QR (~85u) so the header doesn't grow disproportionately.
A raster logo needs no exemption. An inline-SVG logo with brand colors
must carry data-color-exempt="logo" so style_check skips its fills, e.g.:
<svg data-color-exempt="logo" ...> ... </svg> */
.logo-slot img { height: calc(85 * var(--u)); width: auto; max-width: calc(360 * var(--u)); object-fit: contain; display: block; }
/* Logo size classes on .logo-slot -- pick from the file's aspect ratio (SKILL.md Gate E).
tall/square sit at the QR-matched slot height; wide is intentionally shorter (~68% of the
QR) and width-capped so a long wordmark doesn't out-mass the title. Do NOT combine these
with logo-stack (that row is width-normalized, below). */
.logo-slot.logo-tall img,
.logo-slot.logo-square img { height: calc(85 * var(--u)); }
.logo-slot.logo-wide img { height: calc(58 * var(--u)); max-width: calc(300 * var(--u)); }
/* Logo chip: a solid backing that keeps a transparent / edge-white logo legible on a
colored/dark header (white chip) or a light header (.logo-chip-dark); the padding + radius
fold a stray white box into a deliberate rounded tile (SKILL.md Gate E, Background). */
.logo-chip {
display: inline-flex; align-items: center; justify-content: center;
background: var(--bg-card); border-radius: calc(3 * var(--u) * var(--rs));
padding: calc(3 * var(--u)) calc(5 * var(--u));
}
.logo-chip.logo-chip-dark { background: var(--text-primary); }
/* logo-row: institution logos in the header right block (REAL logos the
user provided — never fabricate a seal). Each img MUST carry
data-color-exempt="logo"; height pairs with the QR. */
.logo-row { display: flex; align-items: center; gap: calc(5 * var(--u)); }
.logo-row img { height: calc(68 * var(--u)); width: auto; display: block; }
/* --boxed variant: each logo in a labeled tile (logo + institution name) —
more presence at poster distance; white tile + soft border, no new hues. */
.logo-row .lr-item {
display: flex; flex-direction: column; align-items: center;
gap: calc(2 * var(--u));
background: var(--bg-card);
border: 1px solid var(--border-soft);
border-radius: calc(3 * var(--u) * var(--rs));
padding: calc(4 * var(--u)) calc(6 * var(--u));
}
.logo-row .lr-item img { height: calc(58 * var(--u)); }
.logo-row .lr-label {
font-family: var(--font-sans); font-weight: 600; font-size: var(--fs-1);
color: var(--text-secondary); text-align: center; line-height: 1.15;
}
/* logo-stack variant: WIDE WORDMARKS (AR >= ~2) normalized to EQUAL WIDTH
and stacked vertically, left-aligned. Use when two wide wordmarks of
different aspect ratio read unbalanced height-matched in a row (equal
width lets the less-wide mark grow taller and aligns a clean block).
NOT for a square seal or tall mark (equal width blows it up — keep those
height-matched). Do NOT also apply the logo-wide/tall/square classes. */
.logo-row.logo-stack { flex-direction: column; align-items: flex-start; gap: calc(8 * var(--u)); }
.logo-row.logo-stack img { width: calc(170 * var(--u)); height: auto; }
/* venue badge may carry the official venue logo above its text line */
.venue-badge img { height: calc(62 * var(--u)); width: auto; display: block; margin: 0 auto calc(2 * var(--u)); }
.venue-badge .vb-title { font-family: var(--font-sans); font-weight: 800; font-size: var(--fs-5); color: var(--accent-deep); letter-spacing: 0.5px; }
/* =========================================================
OPTIONAL FRAMEWORK BANNER (delete this whole section if you don't need it)
========================================================= */
.framework-banner {
display: flex;
align-items: center;
gap: calc(16 * var(--u));
/* Flat solid emphasis fill (de-gradient). */
background: var(--bg-emphasis);
border: calc(1 * var(--u)) solid var(--border-soft);
border-left: calc(6 * var(--u)) solid var(--accent);
border-radius: calc(6 * var(--u) * var(--rs));
padding: calc(6 * var(--u)) calc(14 * var(--u));
}
.framework-banner img { height: calc(120 * var(--u)); width: auto; display: block; }
.framework-banner .banner-stats {
flex: 1;
display: grid;
grid-template-columns: 1fr 1fr;
gap: calc(6 * var(--u));
}
.framework-banner .bs-item {
background: white;
border: calc(1 * var(--u)) solid var(--border-soft);
border-left: calc(3 * var(--u)) solid var(--accent);
border-radius: calc(3 * var(--u) * var(--rs));
padding: calc(2 * var(--u)) calc(8 * var(--u));
text-align: center;
}
.framework-banner .bs-num {
font-family: var(--font-sans);
font-weight: 800;
font-size: var(--fs-8);
color: var(--accent);
line-height: 1;
}
.framework-banner .bs-label {
font-family: var(--font-sans);
font-size: var(--fs-3);
color: var(--text-secondary);
margin-top: calc(2 * var(--u));
line-height: 1.2;
}
.framework-banner .fb-text {
flex: 1.6;
font-family: var(--font-serif);
font-size: var(--fs-6);
line-height: 1.5;
text-wrap: pretty; /* prose: fill each line, protect last-line orphan. NOT balance — on multi-sentence prose balance shortens+hyphenates line 1 (crammed-left, gap-right). */
text-align: center;
}
.framework-banner .fb-text strong { color: var(--accent-deep); }
.framework-banner .fb-label {
display: inline-block;
background: var(--accent);
color: var(--accent-ink);
font-family: var(--font-sans);
font-size: var(--fs-5);
font-weight: 700;
padding: calc(2 * var(--u)) calc(8 * var(--u));
border-radius: calc(4 * var(--u) * var(--rs));
text-transform: uppercase;
letter-spacing: 1px;
vertical-align: middle;
line-height: 1;
position: relative;
top: calc(-1 * var(--u));
}
/* Method-overview figure for the banner (banner-figure component,
catalogued in COMPONENTS.md). Usually CAPTIONLESS -- the banner text
block beside it is the figure's explanation. width:min-content collapses
the slot to the IMAGE (a captionless figure's slot IS the image); if a
short caption is used it wraps at the image box and can NEVER set the
flex-item width; margin-inline centres a block image (text-align does
not); overflow-wrap guards a long unbreakable token. Do NOT hand-roll a
.fb-fig wrapper or a bare <img class="w-100"> here. */
.framework-banner .banner-figure {
flex: 0 0 auto;
width: min-content;
margin: 0;
text-align: center;
}
.framework-banner .banner-figure img {
height: calc(120 * var(--u));
width: auto;
display: block;
margin-inline: auto;
}
.framework-banner .banner-figure figcaption {
width: 100%;
margin-top: calc(2 * var(--u));
font-family: var(--font-sans);
font-size: var(--fs-2);
line-height: 1.2;
color: var(--text-secondary);
text-align: center;
text-wrap: pretty;
overflow-wrap: anywhere;
}
.framework-banner .banner-figure figcaption strong { color: var(--accent-deep); }
/* =========================================================
BODY: 4 columns, variable-width hint via fr values
========================================================= */
.body-grid {
display: grid;
grid-template-columns: 1fr 1.05fr 1.05fr 1fr;
gap: calc(10 * var(--u));
/* No overflow:hidden here — it would clip card shadows (pitfall #6).
.poster already clips at the page boundary; min-height:0 is the
grid-blowout guard. */
min-height: 0;
}
.column {
display: flex; flex-direction: column;
gap: calc(6 * var(--u));
min-height: 0;
height: 100%;
padding-bottom: calc(4 * var(--u)); /* shadow breathing room above takeaways */
}
/* =========================================================
CARDS
========================================================= */
.card {
background: var(--bg-card);
border-radius: calc(5 * var(--u) * var(--rs));
padding: calc(4 * var(--u)) calc(9 * var(--u));
border: calc(1 * var(--u)) solid var(--border-soft);
box-shadow: var(--shadow-card);
position: relative;
}
.card.tinted { background: var(--bg-card-tint); }
.card.card--compact { padding: calc(3 * var(--u)) calc(6 * var(--u)); } /* predefined variant: tighter padding (fix (f)) */
.card.highlight {
border-left: calc(6 * var(--u)) solid var(--accent);
/* Flat emphasis tint (de-gradient). */
background: var(--bg-emphasis);
}
.section-title {
font-family: var(--font-sans);
font-weight: 700;
font-size: var(--fs-7);
color: var(--accent-deep);
margin-bottom: calc(3 * var(--u));
display: flex; align-items: center;
gap: calc(5 * var(--u));
}
/* Title text + any ★ marker share ONE .st-text span, so the heading wraps as
natural text (hanging indent) instead of flex-wrapping atomic items: the
number badge is never stranded on its own line and the ★ never widows. */
.section-title .st-text { flex: 1; min-width: 0; line-height: 1.18; }
/* Graceful fallback if a title is NOT wrapped in .st-text: float the badge so
bare inline text still wraps beside it rather than dropping below it. */
.section-title:not(:has(.st-text)) { display: block; line-height: 1.18; }
.section-title:not(:has(.st-text)) .num { float: left; margin-right: calc(5 * var(--u)); }
.section-title .num {
display: inline-flex; align-items: center; justify-content: center;
width: calc(22 * var(--u)); height: calc(22 * var(--u));
background: var(--accent); color: var(--accent-ink);
border-radius: 50%;
font-size: var(--fs-5); font-weight: 700;
flex-shrink: 0;
}
/* Inline "★ KEY" / "★ Headline" marker beside a section title. */
.section-title .key-mark { color: var(--emph); font-size: var(--fs-3); }
/* Contrast fix: the shipped gold token is only 2.1:1 on the highlight-card's
light accent tint (floor 3.0:1) -- darken specifically for that host. */
.card.highlight .section-title .key-mark { color: var(--emph-on-tint); }
.body-text, .card p, .card li {
font-family: var(--font-serif);
font-size: var(--fs-4);
line-height: 1.3;
color: var(--text-primary);
}
.card ul, .card ol { padding-left: calc(18 * var(--u)); }
.card li { margin-bottom: calc(2 * var(--u)); }
.keyword { color: var(--accent); font-weight: 700; }
.keyword-emph { color: var(--emph); font-weight: 700; }
.highlight-text {
background: var(--bg-emphasis);
padding: 0 calc(3 * var(--u));
border-radius: calc(2 * var(--u) * var(--rs));
}
/* Equation block */
.eqn {
background: var(--bg-emphasis);
border-left: calc(3 * var(--u)) solid var(--accent);
padding: calc(4 * var(--u)) calc(10 * var(--u));
margin: calc(4 * var(--u)) 0;
font-size: var(--fs-5);
overflow-x: hidden;
}
/* Predefined variant for a larger / more legible display equation.
`calc(var(--fs-N) * k)` is allowed here because this is a
COMPONENTS.md-registered variant (DESIGN_FINAL §A / §10 f). */
.eqn--large { font-size: calc(var(--fs-5) * 1.25); }
.eqn .label {
display: block;
font-family: var(--font-sans);
font-size: var(--fs-2);
color: var(--accent);
font-weight: 600;
margin-bottom: calc(2 * var(--u));
text-transform: uppercase;
letter-spacing: 1px;
}
/* Callout: solid accent for primary, solid emphasis fill for theorems / "★ key" strips */
.callout {
background: var(--accent);
color: var(--accent-ink);
padding: calc(5 * var(--u)) calc(10 * var(--u));
border-radius: calc(4 * var(--u) * var(--rs));
font-size: var(--fs-4);
margin: calc(4 * var(--u)) 0;
}
.callout strong { color: var(--emph); }
/* Contrast fix: the shipped gold token is only 2.8:1 on the solid-accent
callout fill (floor 3.0:1) -- a lighter gold tint clears it (6.1:1). */
.callout strong.qmark { color: var(--emph-on-accent); }
/* Flat solid emphasis fill (de-gradient). */
.callout.emph {
background: var(--emph);
color: var(--emph-ink); /* route through the token — accent-deep here would dodge the register swap AND fail AA on the default gold */
}
.callout.emph strong { color: var(--emph-ink); }
/* Figure container */
.figure { margin: calc(4 * var(--u)) 0; text-align: center; }
.figure img:not([class*="w-"]) { width: 100%; }
.figure--wide img { width: 100%; } /* predefined variant: force full card width over any .w-NN (fix (f)) */
.figure img {
border-radius: calc(4 * var(--u) * var(--rs));
border: calc(1 * var(--u)) solid var(--fig-frame);
background: var(--fig-bg);
}
.figure .caption {
font-family: var(--font-sans);
font-size: var(--fs-3);
color: var(--text-secondary);
margin-top: calc(3 * var(--u));
line-height: 1.3;
text-align: left;
}
.figure .caption strong { color: var(--accent-deep); }
/* Float a figure BESIDE text in a TEXT-RICH card: text wraps to its side
and then below it. The clearfix grows the card to contain the float;
mark the <img> data-fig-layout="beside-text" so the AR gates honour the
intentionally small width. Use ONLY when there is enough text to fill the
figure's height -- a text-sparse card should center .figure instead. A
short-text float leaves an L-shaped void below the text, which polish's
FIG/BESIDE-TEXT-VOID flags. */
.fig-wrap::after { content: ""; display: table; clear: both; }
.ff-fig {
float: right;
width: 48%; max-width: 58%; min-width: 38%;
margin: calc(1 * var(--u)) 0 calc(3 * var(--u)) calc(11 * var(--u));
text-align: center;
}
.ff-fig.left {
float: left;
margin: calc(1 * var(--u)) calc(11 * var(--u)) calc(3 * var(--u)) 0;
}
.ff-fig img {
display: block;
width: 100%;
border-radius: calc(4 * var(--u) * var(--rs));
border: calc(1 * var(--u)) solid var(--fig-frame);
background: var(--fig-bg);
}
.ff-fig .caption {
font-family: var(--font-sans);
font-size: var(--fs-3);
color: var(--text-secondary);
margin-top: calc(3 * var(--u));
line-height: 1.3;
text-align: center;
}
/* Result table with .ours row highlighted with the emphasis tint */
.result-table {
width: 100%;
border-collapse: collapse;
font-family: var(--font-sans);
font-size: var(--fs-3);
margin-top: calc(3 * var(--u));
}
.result-table th, .result-table td {
padding: calc(2 * var(--u)) calc(4 * var(--u));
text-align: center;
border-bottom: calc(1 * var(--u)) solid var(--border-soft);
}
.result-table thead th {
background: var(--accent); color: var(--accent-ink);
font-weight: 600; font-size: var(--fs-2);
}
.result-table tbody tr.group-row td {
background: var(--bg-emphasis); font-weight: 700;
text-align: left;
color: var(--accent-deep);
padding-left: calc(8 * var(--u));
border-bottom: calc(2 * var(--u)) solid var(--accent);
}
.result-table tbody tr.ours td { background: var(--emph-soft); font-weight: 700; }
.result-table tbody tr.ours td:first-child { color: var(--accent-deep); }
/* Greyed reference row (was inline color:#888). */
.result-table tbody tr.reference td { color: var(--text-muted); }
.result-table .method { text-align: left; padding-left: calc(8 * var(--u)); }
.result-table .best { color: var(--accent); font-weight: 700; }
/* 3-up stat box */
.keybox {
display: grid;
grid-template-columns: repeat(3, 1fr);
gap: calc(4 * var(--u));
margin: calc(4 * var(--u)) 0 0;
}
.keybox .kb-item {
background: var(--bg-emphasis);
border-top: calc(2 * var(--u)) solid var(--accent);
padding: calc(3 * var(--u));
text-align: center;
/* The grid stretches every tile to the tallest one; center the content
vertically so a 1-line tile's number aligns with a 2-line neighbour's
instead of top-ragged (SKILL Layout pitfall — stat-tile alignment). */
display: flex; flex-direction: column; justify-content: center;
border-radius: 0 0 calc(3 * var(--u) * var(--rs)) calc(3 * var(--u) * var(--rs));
}
.kb-item .kb-num {
font-family: var(--font-sans);
font-weight: 800;
font-size: var(--fs-6);
color: var(--accent);
line-height: 1;
}
.kb-item .kb-label {
font-family: var(--font-sans);
font-size: var(--fs-1);
color: var(--text-secondary);
margin-top: calc(2 * var(--u));
line-height: 1.1;
/* Reserve two label lines (2lh tracks this line-height) so 1-line
and 2-line labels occupy the same height and the big numbers
align across tiles. Keep labels <= 2 lines (COMPONENTS §keybox). */
min-height: 2lh;
}
/* Optional takeaways strip (delete the section if not needed) */
.takeaways-strip {
display: grid;
grid-template-columns: auto repeat(4, 1fr);
align-items: center;
gap: calc(10 * var(--u));
/* Flat solid emphasis fill (de-gradient). */
background: var(--bg-emphasis);
border: calc(1 * var(--u)) solid var(--border-soft);
border-radius: calc(5 * var(--u) * var(--rs));
padding: calc(8 * var(--u)) calc(12 * var(--u));
}
.takeaways-strip .ts-title {
font-family: var(--font-sans);
font-size: var(--fs-6);
font-weight: 800;
color: var(--accent-deep);
display: flex; align-items: center; gap: calc(6 * var(--u));
}
.takeaways-strip .ts-title .num {
display: inline-flex; align-items: center; justify-content: center;
width: calc(22 * var(--u)); height: calc(22 * var(--u));
background: var(--accent); color: var(--accent-ink);
border-radius: 50%;
font-size: var(--fs-5); font-weight: 700;
}
.takeaways-strip .ts-item {
border-left: calc(3 * var(--u)) solid var(--accent);
padding-left: calc(8 * var(--u));
line-height: 1.3;
text-align: center;
text-wrap: balance;
}
.takeaways-strip .ts-key {
font-family: var(--font-sans);
font-size: var(--fs-3);
font-weight: 700;
color: var(--accent);
text-transform: uppercase;
letter-spacing: 1px;
}
.takeaways-strip .ts-text {
font-family: var(--font-serif);
font-size: var(--fs-4);
margin-left: calc(4 * var(--u));
}
/* Footer */
.footer {
grid-column: 1 / -1;
display: flex; justify-content: space-between; align-items: baseline;
/* When the two blocks can't sit side by side they stack (wrap) instead of
overflowing; row-gap keeps the stacked blocks legible. */
flex-wrap: wrap; gap: calc(2 * var(--u)) calc(10 * var(--u));
padding-top: calc(8 * var(--u));
border-top: calc(1 * var(--u)) solid var(--border-soft);
font-family: var(--font-sans);
font-size: var(--fs-5);
color: var(--text-muted);
}
.footer .method-name { color: var(--accent-deep); }
/* Long repo URL / email breaks mid-token rather than overflowing the edge. */
.footer .repo { color: var(--accent); font-weight: 600; overflow-wrap: anywhere; }
/* Optional watermark */
.ornament {
position: absolute;
right: calc(20 * var(--u));
bottom: calc(20 * var(--u));
font-family: var(--font-sans);
font-size: var(--fs-9);
color: var(--ornament-ink);
font-weight: 900;
letter-spacing: 4px;
pointer-events: none;
user-select: none;
}
/* Identity mark — corner-signature (Axis 8 identity accessory, ALWAYS ON).
Sits in the poster's outer-padding "safe zone" (bottom-right gutter),
clear of the content box; glyph-only ⊕. This SUPERSEDES the legacy
.ornament text watermark above -- leave that disabled (enabling both
just duplicates the corner mark). See COMPONENTS.md. */
.ps-sprite { position: absolute; width: 0; height: 0; overflow: hidden; }
.corner-sig {
position: absolute;
right: calc(4 * var(--u));
bottom: calc(3 * var(--u));
width: calc(7 * var(--u));
height: calc(7 * var(--u));
color: var(--ps-mark-ink);
opacity: 0.5;
pointer-events: none;
user-select: none;
}
.corner-sig svg { display: block; width: 100%; height: 100%; }
/* Dark / colored ground: swap to a light ink so the mark doesn't vanish.
Add `on-dark` to .corner-sig on an Axis-2 dark-base poster. */
.corner-sig.on-dark { color: var(--accent-light); }
/* Identity mark — woven-signature: ONE ⊕ riding EXISTING content (a best /
target / ★ marker, an inline bullet, a terminal sentence period, the
wordmark "o"). Sized to the HOST TEXT so it reads as an inline glyph, not
a block, and inherits the host ink so it blends. Placed per-poster (never
on a data character); aria-hidden; data-color-exempt="logo". See
COMPONENTS.md. */
[data-ps-mark="woven"] {
display: inline-block;
width: 0.82em;
height: 0.82em;
vertical-align: -0.14em;
line-height: 0;
}
[data-ps-mark="woven"] svg { display: block; width: 100%; height: 100%; }
/* =========================================================
DENSITY COMPONENTS (catalogued — see COMPONENTS.md).
Guardrail: no component-local color semantics. Distinction is
carried by labels, order and typography — never by new hues.
========================================================= */
/* equation-stack: compact multi-row formula stack (denser than one big .eqn) */
.equation-stack { margin: calc(2 * var(--u)) 0; }
.equation-stack .eqn { margin: calc(2 * var(--u)) 0; padding: calc(2 * var(--u)) calc(8 * var(--u)); }
/* eqn-anatomy: term-by-term anatomy grid (2x2 default; --row variant = 1x4) */
.eqn-anatomy {
display: grid; grid-template-columns: 1fr 1fr;
gap: calc(2.5 * var(--u)); margin: calc(3 * var(--u)) 0;
}
.eqn-anatomy.eqn-anatomy--row { grid-template-columns: repeat(4, 1fr); }
.eqn-anatomy .ea-item {
background: var(--bg-card-tint);
border: 1px solid var(--border-soft);
border-left: calc(2 * var(--u)) solid var(--accent);
border-radius: calc(2 * var(--u) * var(--rs));
padding: calc(2 * var(--u)) calc(4 * var(--u));
font-size: var(--fs-2); line-height: 1.3;
}
.eqn-anatomy .ea-tag {
display: inline-block; font-family: var(--font-sans); font-weight: 700;
font-size: var(--fs-1); color: var(--accent-deep);
background: var(--accent-light); border-radius: calc(1.5 * var(--u) * var(--rs));
padding: calc(0.5 * var(--u)) calc(2.5 * var(--u));
margin-right: calc(1.5 * var(--u));
}
/* flow-strip: labeled pipeline, ALL steps the same accent; ONLY the
final step may carry the emphasis top bar. Not an "algorithm" unless the
paper has one — see COMPONENTS.md. */
.flow-strip { display: flex; align-items: stretch; gap: calc(1.5 * var(--u)); margin: calc(3 * var(--u)) 0; }
.flow-strip .step {
flex: 1; background: var(--accent-light);
border: 1px solid var(--accent-soft); border-radius: calc(2 * var(--u) * var(--rs));
padding: calc(2 * var(--u)) calc(2.5 * var(--u));
text-align: center; font-size: var(--fs-2); line-height: 1.25;
}
.flow-strip .step .step-name {
display: block; font-family: var(--font-sans); font-weight: 700;
font-size: var(--fs-1); color: var(--accent-deep);
letter-spacing: 0.5px; text-transform: uppercase;
margin-bottom: calc(1 * var(--u));
}
.flow-strip .step--final { border-top: calc(1.5 * var(--u)) solid var(--emph); background: var(--emph-soft); }
.flow-strip .arrow {
align-self: center; color: var(--accent);
font-family: var(--font-sans); font-weight: 700; font-size: var(--fs-4);
flex: 0 0 auto;
}
/* figure--duo: two paper figures sharing one caption; each img MUST be
a data-source="paper" asset and carry .w-45/.w-50 (42-48% each). */
.figure--duo { display: flex; gap: calc(3 * var(--u)); align-items: flex-start; justify-content: center; }
.figure--duo img { max-width: 48%; min-width: 0; } /* hard cap: duo contract is 42-48% each; an oversized w-NN cannot overflow the card */
/* result-table derived column (emph-soft = DERIVED arithmetic; label it) */
.result-table th.derived, .result-table td.derived {
background: var(--emph-soft); font-family: var(--font-sans); font-weight: 700;
}
/* keybox 4-up variant */
.keybox.keybox--4 { grid-template-columns: repeat(4, 1fr); }
/* algo: compact numbered procedure — ONLY when the paper itself states
an explicit algorithm/procedure; never invent steps. */
ol.algo { padding-left: calc(16 * var(--u)); }
ol.algo li { margin-bottom: calc(1.5 * var(--u)); font-size: var(--fs-3); line-height: 1.3; }
/* claim-pills: provenance mini-table (numeric-heavy posters only) */
.claim-pills { width: 100%; border-collapse: collapse; font-family: var(--font-sans); font-size: var(--fs-2); }
.claim-pills td {
border-bottom: 1px solid var(--border-soft);
padding: calc(1.5 * var(--u)) calc(3 * var(--u));
}
.claim-pills .cp-id {
font-weight: 700; color: var(--accent-deep); background: var(--accent-light);
border-radius: calc(1.5 * var(--u) * var(--rs)); padding: 0 calc(2.5 * var(--u) * var(--rs)); white-space: nowrap;
}
.claim-pills .cp-fact { color: var(--accent-deep); font-weight: 700; }
.claim-pills .cp-derived { color: var(--emph); font-weight: 700; }
/* Grid-blowout guards: (1) columns may shrink below min-content so a
wide equation can never widen its track and squeeze siblings;
(2) over-wide display math scales down to fit its box instead of
blowing out the card (SVG keeps aspect via height:auto). */
.body-grid > .column { min-width: 0; }
.card { min-width: 0; }
.eqn mjx-container > svg { max-width: 100%; height: auto; }
/* ── Utility classes — defined LAST so equal-specificity ties
resolve in the utility's favor (source order). ── */
.fs-1 { font-size: var(--fs-1); }
.fs-2 { font-size: var(--fs-2); }
.fs-3 { font-size: var(--fs-3); }
.fs-4 { font-size: var(--fs-4); }
.fs-5 { font-size: var(--fs-5); }
.fs-6 { font-size: var(--fs-6); }
.fs-7 { font-size: var(--fs-7); }
.fs-8 { font-size: var(--fs-8); }
.fs-9 { font-size: var(--fs-9); }
.mt-1 { margin-top: calc(1 * var(--u)); }
.mt-2 { margin-top: calc(2 * var(--u)); }
.mt-3 { margin-top: calc(3 * var(--u)); }
.mt-4 { margin-top: calc(4 * var(--u)); }
.mt-5 { margin-top: calc(5 * var(--u)); }
.mt-6 { margin-top: calc(6 * var(--u)); }
/* Fixed-height balance spacers (column-alignment fine-tuning, measure gate). */
.balance-spacer-23 { height: 23px; }
.balance-spacer-28 { height: 28px; }
.mb-1 { margin-bottom: calc(1 * var(--u)); }
.mb-2 { margin-bottom: calc(2 * var(--u)); }
.mb-3 { margin-bottom: calc(3 * var(--u)); }
.mb-4 { margin-bottom: calc(4 * var(--u)); }
.w-45 { width: 45%; }
.w-50 { width: 50%; }
.w-55 { width: 55%; }
.w-60 { width: 60%; }
.w-65 { width: 65%; }
.w-70 { width: 70%; }
.w-75 { width: 75%; }
.w-80 { width: 80%; }
.w-85 { width: 85%; }
.w-90 { width: 90%; }
.w-95 { width: 95%; }
.w-100 { width: 100%; }
.text-secondary { color: var(--text-secondary); }
.text-muted { color: var(--text-muted); }
.nowrap { white-space: nowrap; }
.text-center { text-align: center; }
/* =========================================================
PRINT OVERRIDE — KEEP LAST so it wins source-order ties.
========================================================= */
@media print {
html, body { background: white; }
.poster { margin: 0; box-shadow: none; width: 60in; height: 36in; page-break-after: avoid; }
:root { --u: 1mm; }
}
</style>
</head>
<body>
<div class="poster" data-measure-role="poster"
data-posterly-contract="identity-v1" data-ps-identity="on">
<!-- Identity sprite: the posterly ⊕ registration glyph, defined ONCE.
Zero-size + absolute so it never enters the grid or shifts layout.
data-color-exempt="logo" -- it IS the posterly logo mark, which keeps
style_check rule 11 (no decorative SVG) and the rule-4 hue gate honest.
Symbol id/viewBox/paths are byte-identical across all templates. -->
<svg class="ps-sprite" data-color-exempt="logo" aria-hidden="true"><defs>
<symbol id="psReg" viewBox="0 0 100 100">
<circle cx="50" cy="50" r="25" fill="none" stroke="currentColor" stroke-width="8"/>
<line x1="50" y1="7" x2="50" y2="93" stroke="currentColor" stroke-width="8"/>
<line x1="7" y1="50" x2="93" y2="50" stroke="currentColor" stroke-width="8"/>
</symbol>
</defs></svg>
<!-- =========================== HEADER =========================== -->
<header class="header" data-measure-role="header">
<!-- LEFT: venue badge. Replace text or swap for an <img> if your venue allows logos. -->
<div class="venue-badge">
<div class="vb-venue">ICML</div>
<div class="vb-year">2026</div>
<div class="vb-tag">REPRO</div>
</div>
<!-- CENTER: title block -->
<div class="title-block">
<h1 class="title">Beyond Text-to-SQL: <span class="accent">Reproducing Squirrel Benchmark</span></h1>
<div class="subtitle">Independently verifying whether LLMs can really debug enterprise ETL SQL (arXiv:2601.18119)</div>
<div class="authors-line">
<span class="author">Agent Reproduction<sup>&#9993;</sup></span>
<span class="aff">ICML-2026-agent-repro Challenge &middot; Hugging Face &times; AlphaXiv Open Reproductions</span>
</div>
</div>
<!-- RIGHT: optional logo + QR. Delete .logo-slot if your venue forbids on-poster logos. -->
<div class="right-block">
<div class="qr-block">
<img data-color-exempt="logo" src="images/qr.png" alt="QR code linking to the published reproduction logbook">
<div class="qr-label">Published logbook</div>
</div>
</div>
</header>
<!-- ===================== OPTIONAL FRAMEWORK BANNER =====================
Delete this entire <section> if your poster doesn't need a banner.
(To do so, also remove `auto` from the grid-template-rows above.) -->
<section class="framework-banner" data-measure-role="banner">
<!-- OPTIONAL method-overview figure on the left. The banner text block
beside it IS its explanation, so it usually needs NO caption. Use the
captionless banner-figure component (NOT a bare <img class="w-100"> --
the `.framework-banner img { width:auto }` rule wins over .w-100):
<figure class="banner-figure">
<img src="assets/paper_figures/framework.png" data-source="paper"
data-asset-id="framework" alt="Method overview">
</figure>
A figcaption here just duplicates the banner text and a long one
stretches the figure slot (polish: BANNER/IMAGE-SLOT). If you truly
need panel labels, bake them into the image or keep them to ONE short
<figcaption> line (the component bounds it to the image width). -->
<div class="fb-text">
<span class="fb-label">Reproduction verdict</span>
&nbsp;All <strong>4 headline claims</strong> of the paper check out verbatim against its own tables and abstract; the real 985-task benchmark and Claude-4-Sonnet are unavailable, so we additionally ran a small <strong>toy</strong> substitute experiment on open models.
</div>
<div class="banner-stats">
<div class="bs-item"><div class="bs-num">4/4</div><div class="bs-label">claims confirmed<br>vs. primary source</div></div>
<div class="bs-item"><div class="bs-num">20</div><div class="bs-label">toy tasks &times; 4 models<br>HF Inference Providers</div></div>
<div class="bs-item"><div class="bs-num">&lt;$1</div><div class="bs-label">compute cost<br>(API calls only)</div></div>
<div class="bs-item"><div class="bs-num">toy</div><div class="bs-label">scale vs. real<br>985-task benchmark</div></div>
</div>
</section>
<!-- ============================ BODY ============================ -->
<div class="body-grid" data-measure-role="body">
<!-- ============ COLUMN 1 ============ -->
<div class="column" data-measure-role="column">
<div class="card highlight" data-measure-role="card" data-logbook-target="executive-summary">
<div class="section-title"><span class="num">1</span><span class="st-text">Why this reproduction</span></div>
<p class="body-text">
Squirrel Benchmark claims enterprise SQL <span class="keyword">debugging</span> (not generation) is where LLMs fail hardest: real ETL scripts run 140+ lines with deep nested joins, yet even the best model tested solves only about a third of tasks.
</p>
<ul class="mt-3 fs-4">
<li>The real 985-task benchmark is <strong>not yet public</strong> (paper: "scheduled for release upon acceptance").</li>
<li>Claude-4-Sonnet, the best model in Table 2, is <strong>unavailable</strong> in this HF-based challenge.</li>
</ul>
<div class="callout mt-4">
<strong class="qmark">Q:</strong> can we verify the paper's numbers from its own primary source, and independently corroborate its qualitative findings with a small toy substitute experiment, since neither the real benchmark nor Claude-4-Sonnet is available to us here?
</div>
</div>
<div class="card" data-measure-role="card" data-logbook-target="claim-1-sqlbench-scale-469-syntax-516-semantic-debugging-tasks-from-1-000-seed-scripts-across-26-business-scenarios">
<div class="section-title"><span class="num">2</span><span class="st-text">Claim 1 &middot; Benchmark scale</span></div>
<p class="body-text">
Built from <span class="keyword">1,000+ seed SQL scripts</span> across 26 business scenarios via an automated reverse-engineering pipeline &mdash; verbatim-confirmed against Section 3.1 / abstract, no discrepancy.
</p>
<div class="keybox mt-3">
<div class="kb-item"><div class="kb-num">469</div><div class="kb-label">Squirrel-Syntax<br>tasks</div></div>
<div class="kb-item"><div class="kb-num">516</div><div class="kb-label">Squirrel-Semantic<br>tasks</div></div>
<div class="kb-item"><div class="kb-num">26</div><div class="kb-label">business<br>scenarios</div></div>
</div>
<p class="body-text mt-3">
Seed scripts alone (before task-level bug injection) already average <strong>120+ lines</strong> with AST <strong>depth &gt; 8</strong> and <strong>width &gt; 12</strong> &mdash; the paper's Section 3.1 figures, cross-checked against its own abstract and intro with no discrepancy found anywhere in the three passages.
</p>
<div class="callout mt-3">
We found no GitHub repo, HF dataset, or checkpoint linked from the paper, arXiv, or OpenReview &mdash; confirmed by directly querying both the arXiv and OpenReview APIs.
</div>
<table class="result-table mt-3">
<thead><tr><th class="method">Corpus stage</th><th># Lines</th><th>AST depth</th><th>AST width</th></tr></thead>
<tbody>
<tr><td class="method">Seed corpus (Sec. 3.1)</td><td>120+</td><td>&gt;8</td><td>&gt;12</td></tr>
<tr class="ours"><td class="method">Final tasks (Table 1)</td><td class="best">140-164</td><td class="best">8.75-8.93</td><td class="best">11.1-11.7</td></tr>
</tbody>
</table>
<div class="balance-spacer-23" aria-hidden="true"></div>
</div>
</div>
<!-- ============ COLUMN 2 ============ -->
<div class="column" data-measure-role="column">
<div class="card" data-measure-role="card" data-logbook-target="claim-4-sqlbench-script-complexity-420-tokens-and-17-34-21-62-functions-per-script-on-average">
<div class="section-title"><span class="num">3</span><span class="st-text">Claim 4 &middot; Script complexity</span></div>
<p class="body-text">Seed scripts must clear a composite complexity threshold before being retained in the corpus, filtering out trivially short queries:</p>
<div class="eqn">
<span class="label">Seed complexity filter</span>
$$\mathcal{C}(q) = \alpha\big(D_{\text{AST}}(q) + W_{\text{AST}}(q)\big) + \beta L(q)$$
</div>
<p class="body-text">Resulting scripts average <strong>420-497 tokens</strong> and <strong>17.34-21.62 functions</strong> each &mdash; verbatim-confirmed against Table 1, an order of magnitude above Spider 1.0's 18.5 tokens per query.</p>
<table class="result-table mt-3">
<thead><tr><th class="method">Corpus</th><th>Tok./SQL</th><th>Func./SQL</th><th>AST depth</th></tr></thead>
<tbody>
<tr><td class="method">Squirrel-Syntax</td><td>496.90</td><td>21.62</td><td>8.93</td></tr>
<tr><td class="method">Squirrel-Semantic</td><td>425.93</td><td>17.34</td><td>8.75</td></tr>
<tr class="ours"><td class="method">Our 10-script toy corpus</td><td class="best">1403</td><td class="best">31</td><td class="best">10</td></tr>
</tbody>
</table>
</div>
<div class="card highlight" data-measure-role="card" data-logbook-target="claim-2-claude-4-sonnet-best-model-36-46-graph-match-on-squirrel-syntax-32-17-on-squirrel-semantic">
<div class="section-title"><span class="num">4</span><span class="st-text">Claim 2 &middot; Best model&nbsp;<span class="key-mark">&#9733;<span data-ps-mark="woven" data-color-exempt="logo" aria-hidden="true"><svg viewBox="0 0 100 100"><use href="#psReg"/></svg></span> KEY</span></span></div>
<p class="body-text">Graph Match credits structurally-equivalent SQL even when strings differ (Eq. 6) &mdash; it is why EM alone undercounts correct fixes:</p>
<div class="eqn eqn--large">
<span class="label">Graph Match score</span>
$$\text{GM} = \tfrac{1}{N}\textstyle\sum_i \mathbf{1}\big[\text{Graph}(\hat q_i) \cong \text{Graph}(q_i)\big]$$
</div>
<div class="callout emph">
<strong>Verbatim-confirmed:</strong> Claude-4-Sonnet is the best of ~30 evaluated models, yet clears only <strong>36.46%</strong> (Syntax) / <strong>32.17%</strong> (Semantic) GM &mdash; it generated the benchmark via reverse engineering and still can't reliably debug it forward, end to end.
</div>
<p class="body-text mt-3 fs-3">
We saw this same EM-vs-GM gap on our own toy run: DeepSeek-V3 scored only 30% EM on Syntax because it dropped a leading SQL comment from an otherwise-correct fix, then jumped straight to 100% GM once our proxy learned to ignore comments too, the same way Calcite's optimized plan already does.
</p>
</div>
</div>
<!-- ============ COLUMN 3 ============ -->
<div class="column" data-measure-role="column">
<div class="card highlight" data-measure-role="card" data-logbook-target="claim-3-deepseek-v3-and-qwen-2-5-coder-32b-scores-most-llms-below-20-success-rate">
<div class="section-title"><span class="num">5</span><span class="st-text">Claim 3 &middot; Model rankings&nbsp;<span class="key-mark">&#9733; Headline</span></span></div>
<p class="body-text fs-3 text-secondary mb-1">
Table 2, Graph Match (GM) across ~30 evaluated LLMs on both splits
</p>
<table class="result-table">
<thead>
<tr>
<th class="method">Model</th>
<th>Syntax GM &#8593;</th>
<th>Semantic GM &#8593;</th>
</tr>
</thead>
<tbody>
<tr class="ours"><td class="method">Claude-4-Sonnet</td><td class="best">36.46</td><td class="best">32.17</td></tr>
<tr><td class="method">DeepSeek-V3</td><td>30.28</td><td>21.32</td></tr>
<tr><td class="method">Qwen-2.5-Coder-32B</td><td>20.26</td><td>23.45</td></tr>
<tr><td class="method">GPT-4o</td><td>4.69</td><td>4.84</td></tr>
<tr><td class="method">Doubao-Seed-1.6</td><td>30.92</td><td>20.93</td></tr>
<tr><td class="method">GPT-5</td><td>18.55</td><td>16.47</td></tr>
</tbody>
</table>
<div class="keybox">
<div class="kb-item"><div class="kb-num">36.46%</div><div class="kb-label">best Syntax<br>(Claude-4-Sonnet)</div></div>
<div class="kb-item"><div class="kb-num">32.17%</div><div class="kb-label">best Semantic<br>(Claude-4-Sonnet)</div></div>
<div class="kb-item"><div class="kb-num">&lt;20%</div><div class="kb-label">most models'<br>GM score</div></div>
</div>
<p class="body-text mt-3 fs-3">
Checked across all 24 main rows of Table 2: <strong>10/24</strong> models score below 20% GM on Syntax, and <strong>15/24</strong> on Semantic (62.5%) &mdash; a clear majority in both cases, matching the paper's own framing of the result.
</p>
</div>
<div class="card" data-measure-role="card">
<div class="section-title"><span class="num">6</span><span class="st-text">Our toy substitute experiment</span></div>
<p class="body-text">985 real tasks and Claude-4-Sonnet unavailable, so we built a 20-task synthetic benchmark (10 domains) and ran 4 open models through Hugging Face's Inference Providers API.</p>
<ul class="mt-3">
<li><strong>Syntax (1 bug pattern)</strong> &mdash; 100% GM-proxy for 3/4 models: too easy/homogeneous vs. the real 469-task, taxonomy-diverse split.</li>
<li><strong>Semantic (3 bug patterns)</strong> &mdash; drops to 50-60% GM-proxy for every model, same ranking as Table 2 (32B &ge; 7B).</li>
</ul>
<div class="callout mt-3">
Directionally reproduces "Semantic &gt; Syntax difficulty" and the model-scale effect &mdash; absolute scores aren&#39;t comparable given the gap between 20 and 985 tasks.
</div>
<table class="result-table mt-3">
<thead><tr><th class="method">Model (toy run)</th><th>Syntax GM</th><th>Semantic GM</th></tr></thead>
<tbody>
<tr><td class="method">Qwen-2.5-Coder-32B</td><td>100%</td><td>60%</td></tr>
<tr><td class="method">DeepSeek-V3-0324</td><td>100%</td><td>50%</td></tr>
<tr><td class="method">Qwen3-235B-Instruct</td><td>100%</td><td>60%</td></tr>
</tbody>
</table>
</div>
</div>
<!-- ============ COLUMN 4 ============ -->
<div class="column" data-measure-role="column">
<div class="card highlight" data-measure-role="card">
<div class="section-title"><span class="num">7</span><span class="st-text">Claim-by-claim verdict</span></div>
<p class="body-text fs-2 text-secondary mb-1">
Every figure checked against arXiv:2601.18119's own Tables 1-2, abstract, and Sections 3-5
</p>
<table class="result-table">
<thead>
<tr>
<th class="method">Claim</th>
<th>Primary source</th>
<th>Toy experiment</th>
</tr>
</thead>
<tbody>
<tr><td class="method">1 &middot; Scale</td><td class="best">verbatim</td><td>order-of-magnitude</td></tr>
<tr><td class="method">2 &middot; Best model</td><td class="best">verbatim</td><td>EM-vs-GM gap replayed</td></tr>
<tr><td class="method">3 &middot; Rankings</td><td class="best">verbatim</td><td>ranking direction holds</td></tr>
<tr class="ours"><td class="method">4 &middot; Complexity</td><td class="best">verbatim</td><td class="best">order-of-magnitude</td></tr>
</tbody>
</table>
<p class="body-text mt-2 fs-3">
Real benchmark/model unavailable &rArr; exact Table 1/2 percentages not independently re-derived, by design. Every other figure on this poster traces to a specific table, section, or equation in the paper &mdash; see the logbook claim pages for the full citation trail.
</p>
</div>
<div class="card" data-measure-role="card" data-logbook-target="conclusion">
<div class="section-title"><span class="num">8</span><span class="st-text">Reproduction bundle</span></div>
<p class="body-text">Every script, generated task, model response, and eval log is published as a Hugging Face Bucket artifact, linked from the logbook's Conclusion page for anyone to re-run or extend.</p>
<ul class="mt-3 fs-4">
<li><code>scripts/generate_benchmark.py</code> &mdash; synthesizes the 10-domain toy corpus and injects taxonomy-guided bugs.</li>
<li><code>scripts/eval_models.py</code> &mdash; runs the paper's own evaluation prompts (Appendix J.3) against any chat-completions model hosted on the Hugging Face Inference Providers API.</li>
<li><code>scripts/metrics.py</code> &mdash; EM / GM-proxy / MB-proxy, with the Calcite-vs-sqlglot substitution documented inline.</li>
</ul>
<div class="callout mt-3">
Bucket: <code>Firemedic15/squirrel-sqlbench-repro</code> &mdash; scripts, generated data, eval outputs, and this poster's own render pipeline, all runnable with just an HF token.
</div>
<p class="body-text mt-2 fs-3 text-secondary">
Both HF Jobs referenced here (evaluation re-run and poster gate/render) are linked from the logbook's Conclusion page alongside their exact commands and hardware flavor, so the whole pipeline is re-runnable end to end by anyone without reading this poster first.
</p>
<div class="balance-spacer-28" aria-hidden="true"></div>
</div>
</div>
</div>
<!-- ============================ OPTIONAL TAKEAWAYS STRIP ============================ -->
<!-- Strip title + slot labels are microcopy, not canon (SKILL.md Step 3):
reword both to the poster's voice; 3-4 slots, any labels. The old fixed
default ("Takeaways" + Idea/Method/Result/Practical) shipped on every
poster is a fingerprint -- treat it as one example set, not the answer. -->
<section class="takeaways-strip" data-measure-role="footer-strip">
<div class="ts-title"><span class="num">9</span> Bottom line</div>
<div class="ts-item"><span class="ts-key">Primary source.</span><span class="ts-text">All 4 claims match Tables 1-2 verbatim, no discrepancy.</span></div>
<div class="ts-item"><span class="ts-key">Real data unavailable.</span><span class="ts-text">985-task benchmark + Claude-4-Sonnet not accessible here.</span></div>
<div class="ts-item"><span class="ts-key">Toy substitute.</span><span class="ts-text">20 tasks &times; 4 open models via HF Inference Providers.</span></div>
<div class="ts-item"><span class="ts-key">Qualitative match.</span><span class="ts-text">Semantic&gt;Syntax difficulty + model-scale effect both hold.</span></div>
</section>
<!-- ============================ FOOTER ============================ -->
<!-- Footer labels ("Code:" / "Contact:" / "Acknowledgements:") are microcopy --
reword or reorder to the poster's voice rather than shipping this exact
skeleton every time (SKILL.md Step 3). -->
<div class="footer" data-measure-role="footer">
<div>
<strong class="method-name">SQUIRREL BENCHMARK REPRODUCTION</strong> &middot; ICML 2026 Agent Reproduction Challenge &middot;
Built with Hugging Face Trackio + posterly.
</div>
<div>
Bundle: <span class="repo">huggingface.co/buckets/Firemedic15/squirrel-sqlbench-repro</span> &nbsp;&middot;&nbsp;
Paper: <span class="repo">arxiv.org/abs/2601.18119</span>
</div>
</div>
<!-- LEGACY decorative watermark (Axis 8): a faint corner TEXT watermark,
SUPERSEDED by the always-on corner-signature below under identity-v1.
Keep it disabled: with data-ps-identity="on" the corner-signature is
the required identity mark (dropping it fails preflight), and in an
anonymous data-ps-identity="off" poster this lab watermark would itself
leak an identifying mark. Kept commented for reference only:
<div class="ornament">LAB &middot; INSTITUTION</div>
-->
<!-- Identity accessory (Axis 8) — corner-signature: posterly's ⊕ in the
bottom-right safe zone. Always on while data-ps-identity="on". -->
<span class="corner-sig" data-ps-mark="corner" data-color-exempt="logo"
aria-hidden="true"><svg viewBox="0 0 100 100"><use href="#psReg"/></svg></span>
</div>
</body>
</html>

Xet Storage Details

Size:
64.1 kB
·
Xet hash:
6eadaf51cfd5e29e83b737e2363a6ca10b9c814b3d79e0ae38bea3d1fec26156

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.