Baladithya Balamurugan Claude Fable 5 commited on
Commit
9a2ce20
·
1 Parent(s): 2a16b30

Wave 21: Stage-0 dataset pipeline — swesmith engine, rollout harness, gates, contract

Browse files

Implements research/deepread/13-synthesis-architecture.md (ADR-016): the
"point at a repo -> dataset" pipeline the critical review architected, closing
every verified P0 from research/deepread/12-verified-findings.md.

New modules (+96 tests, full suite 511 passed / 66 skipped):
- datagen/repo_gate.py: SPDX-ish license detection -> 3 tiers (fail-closed) +
benchmark DECONTAMINATION vs the SWE-bench eval-repo list (closes V3 — zero
decontamination existed anywhere).
- datagen/swesmith_adapter.py + [swesmith] extra: SWE-smith as the synthesis
engine (closes V4 buy-vs-build). Handles the patch-semantics INVERSION
(SWE-smith's patch INTRODUCES the bug; golden_diff = reverse_unified_diff)
+ strategy provenance sidecar + difficulty priors.
- datagen/trajectory.py: canonical trajectory IR (closes D-11). ToolCall
canonical_form = v1 divergence-gate action algebra (replaces the whitespace
stub, D-3); to_policy_row = THE policy-visible serializer, sentinel-tested
to never leak golden_diff/deleted_symbols (D-8).
- datagen/rollout_harness.py: collect_trajectory agent loop over
FeatureDeletionEnv (closes V2 — the SFT corpus finally has a producer; its
env-grounded episodes are also the tree's seeds, fixing D-1) + typed
admission routing (sft/dpo-candidate/quarantine).
- pipeline/{s3_contract,dedup,build_corpus}.py: ONE reconciled dataset layout
(supersedes F1/F2's divergent contracts, V8) with restricted tasks_full
prefix + golden_diff sha256; stable-hash MinHash dedup incl.
cross-generation signatures (D-12); local write-once stage-driver with
holdout-first split + budget ceiling (D-9/D-21).

Fidelity corrections (verified findings, applied to code+docs):
- V1: SDPO "mathematically the same" claim corrected (opsd.py, mapping doc) —
Cursor cites SDPO as background; ours is a third blog-inspired design.
- V5: fabricated numbers struck/tagged (69.3%, "24 generators", "85% compute").
- V7: Streaming DiLoCo citation fixed (Douillard 2501.18512; Kale 2502.12996).
- V11: teacher_replay cost docstring relabeled (synthetic-trace basis).
- V13: kl_in_reward verl wording (default/recommended, not only).
- ADR-016 records the decision + gates; harvested repo_gate from the one
surviving worktree builder (3 of 4 stalled on API capacity; rebuilt directly).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. composer_replication/datagen/repo_gate.py +361 -0
  2. composer_replication/datagen/rollout_harness.py +214 -0
  3. composer_replication/datagen/swesmith_adapter.py +269 -0
  4. composer_replication/datagen/tests/test_repo_gate.py +419 -0
  5. composer_replication/datagen/tests/test_rollout_harness.py +103 -0
  6. composer_replication/datagen/tests/test_swesmith_adapter.py +165 -0
  7. composer_replication/datagen/tests/test_trajectory.py +127 -0
  8. composer_replication/datagen/trajectory.py +203 -0
  9. composer_replication/diloco/__init__.py +6 -3
  10. composer_replication/opsd.py +11 -5
  11. composer_replication/pipeline/__init__.py +38 -0
  12. composer_replication/pipeline/build_corpus.py +137 -0
  13. composer_replication/pipeline/dedup.py +138 -0
  14. composer_replication/pipeline/s3_contract.py +287 -0
  15. composer_replication/pipeline/tests/__init__.py +0 -0
  16. composer_replication/pipeline/tests/test_pipeline.py +223 -0
  17. composer_replication/teacher_replay.py +6 -2
  18. composer_replication/trainer/kl_in_reward.py +3 -1
  19. docs/COMPOSER_RECIPE_MAPPING.md +3 -3
  20. docs/adrs/ADR-016-stage0-dataset-pipeline.md +119 -0
  21. pyproject.toml +8 -0
  22. research/01-composer-2.5.md +3 -3
  23. research/06-feature-deletion-datagen.md +1 -1
  24. research/09-composer-blog-delta-2026.md +1 -1
  25. research/notes/230406767-raft-reward-ranked-finetuning-for-generative-foundation-model-alignmen.md +224 -0
  26. research/notes/230510601-tree-of-thoughts-deliberate-problem-solving-with-large-language-models-2.md +2735 -0
  27. research/notes/230510601-tree-of-thoughts-deliberate-problem-solving-with-large-language-models.md +213 -0
  28. research/notes/231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la-2.md +4095 -0
  29. research/notes/231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la-3.md +4095 -0
  30. research/notes/231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la.md +202 -0
  31. research/notes/231006770-swe-bench-can-language-models-resolve-real-world-github-issues.md +203 -0
  32. research/notes/231108105-diloco-distributed-low-communication-training-of-language-models.md +208 -0
  33. research/notes/231108516-llms-cannot-find-reasoning-errors-but-can-correct-them-given-the-error.md +204 -0
  34. research/notes/231209152-evaluating-augmented-reality-communication-how-can-we-teach-procedural.md +203 -0
  35. research/notes/240201817-llms-cant-plan-but-can-help-planning-in-llm-modulo-frameworks.md +203 -0
  36. research/notes/240203300-deepseekmath-pushing-the-limits-of-mathematical-reasoning-in-open-lang.md +214 -0
  37. research/notes/240411018-many-shot-in-context-learning.md +228 -0
  38. research/notes/240612543-phase-controlled-heat-modulation-with-aharonov-bohm-interferometers.md +197 -0
  39. research/notes/240806195-mutual-reasoning-makes-smaller-llms-stronger-problem-solvers-2.md +2384 -0
  40. research/notes/240806195-mutual-reasoning-makes-smaller-llms-stronger-problem-solvers.md +191 -0
  41. research/notes/241020285-swe-search-enhancing-software-agents-with-monte-carlo-tree-search-and-2.md +2144 -0
  42. research/notes/241221139-training-software-engineering-agents-and-verifiers-with-swe-gym.md +200 -0
  43. research/notes/250104519-rstar-math-small-llms-can-master-math-reasoning-with-self-evolved-deep.md +196 -0
  44. research/notes/250104519-sysname-small-llms-can-master-math-reasoning-with-self-evolved-deep-th.md +3557 -0
  45. research/notes/250109136-agentic-retrieval-augmented-generation-a-survey-on-agentic-rag.md +203 -0
  46. research/notes/250109891-evolving-deeper-llm-thinking.md +196 -0
  47. research/notes/250112599-kimi-k15-scaling-reinforcement-learning-with-llms.md +386 -0
  48. research/notes/250118512-streaming-diloco-with-overlapping-communication-towards-a-distributed.md +206 -0
  49. research/notes/250118639-a-comprehensive-survey-of-the-lean-4-theorem-prover-architecture-appli.md +179 -0
  50. research/notes/250202047-amasquad-a-benchmark-for-amharic-extractive-question-answering.md +183 -0
composer_replication/datagen/repo_gate.py ADDED
@@ -0,0 +1,361 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """repo_gate.py — Stage-0 ingest gate: license tiers + benchmark decontamination.
2
+
3
+ Architecture step 1 of the dataset pipeline (research/deepread/
4
+ 13-synthesis-architecture.md Part B). Closes two verified findings:
5
+
6
+ * V3 / D-5 — ZERO benchmark decontamination existed anywhere in code or
7
+ designs, while the pipeline trains on SWE-bench-family substrates and is
8
+ scored on SWE-bench Verified. ``is_eval_contaminated`` is the hard wall:
9
+ a repo on the eval list is NEVER admitted, regardless of license.
10
+ * V9 / D-13 — the only license filter was a lowercase substring match on a
11
+ task field (``substrates.py`` ``is_redistributable``), with no SPDX
12
+ detection at the repo-ingest path and no trainable-vs-redistributable
13
+ split. ``detect_license`` + ``license_tier`` replace the boolean with a
14
+ three-tier verdict.
15
+
16
+ Why tiers, not a boolean (D-13): weak-copyleft repos (MPL/LGPL) are fine to
17
+ *train on* but we must not *redistribute* derivative diffs from them — a
18
+ boolean "redistributable?" gate either over-excludes them or leaks them into
19
+ published corpora. The tier travels with the verdict so downstream corpus
20
+ steps (step 6) can route TRAINABLE_ONLY rows away from any published split.
21
+
22
+ Why title-anchored matching for the GNU family: GPL-3.0 §13 mentions the
23
+ "GNU Affero General Public License" by name and AGPL-3.0 §13 mentions the
24
+ "GNU General Public License" — naive full-body substring matching
25
+ misclassifies one as the other. We therefore classify the GNU/MPL/Apache
26
+ family from the document HEADER (first ~400 normalized chars, where the
27
+ license title lives) and only use full-body phrases for the short permissive
28
+ licenses whose titles are not distinctive (MIT/ISC/BSD/Unlicense).
29
+
30
+ Stdlib-only on purpose: the gate must run before anything heavy is installed.
31
+ """
32
+ from __future__ import annotations
33
+
34
+ import json
35
+ import re
36
+ from dataclasses import dataclass, field
37
+ from enum import Enum
38
+ from pathlib import Path
39
+
40
+ # ---------------------------------------------------------------------------
41
+ # License detection (V9 / D-13)
42
+ # ---------------------------------------------------------------------------
43
+
44
+ #: License files checked in order; first match wins (case-insensitive on name).
45
+ _LICENSE_FILENAMES: tuple[str, ...] = ("LICENSE", "LICENSE.txt", "LICENSE.md", "COPYING")
46
+
47
+ #: Trove classifier / PEP 639 expression fragments → SPDX id. Secondary signal
48
+ #: only — the classifier cannot distinguish BSD-2 from BSD-3, so it maps to
49
+ #: BSD-3-Clause (the common case) and the LICENSE file is preferred when present.
50
+ _CLASSIFIER_MAP: tuple[tuple[str, str], ...] = (
51
+ ("gnu affero general public license", "AGPL-3.0"),
52
+ ("gnu lesser general public license v3", "LGPL-3.0"),
53
+ ("gnu lesser general public license v2.1", "LGPL-2.1"),
54
+ ("gnu lesser general public license", "LGPL-3.0"),
55
+ ("gnu general public license v3", "GPL-3.0"),
56
+ ("gnu general public license v2", "GPL-2.0"),
57
+ ("mozilla public license 2.0", "MPL-2.0"),
58
+ ("apache software license", "Apache-2.0"),
59
+ ("mit license", "MIT"),
60
+ ("bsd license", "BSD-3-Clause"),
61
+ ("isc license", "ISC"),
62
+ ("the unlicense", "Unlicense"),
63
+ )
64
+
65
+ #: Bare SPDX ids accepted from PEP 639 ``license = "<expr>"`` in pyproject.
66
+ _SPDX_IDS: frozenset[str] = frozenset(
67
+ {
68
+ "MIT", "Apache-2.0", "BSD-2-Clause", "BSD-3-Clause", "ISC",
69
+ "GPL-2.0", "GPL-3.0", "AGPL-3.0", "LGPL-2.1", "LGPL-3.0",
70
+ "MPL-2.0", "Unlicense",
71
+ }
72
+ )
73
+ _SPDX_LOOKUP: dict[str, str] = {s.lower(): s for s in _SPDX_IDS}
74
+ # Common -only/-or-later suffixed forms normalize to the base id we tier on.
75
+ for _base in ("GPL-2.0", "GPL-3.0", "AGPL-3.0", "LGPL-2.1", "LGPL-3.0"):
76
+ _SPDX_LOOKUP[f"{_base.lower()}-only"] = _base
77
+ _SPDX_LOOKUP[f"{_base.lower()}-or-later"] = _base
78
+
79
+
80
+ @dataclass(frozen=True)
81
+ class LicenseInfo:
82
+ """Outcome of license detection: SPDX-ish id + which signal decided it."""
83
+
84
+ spdx_id: str # one of _SPDX_IDS or "unknown"
85
+ signal: str # "license_file" | "classifier" | "none"
86
+ source: str = "" # filename that supplied the winning signal
87
+
88
+
89
+ def _normalize_text(text: str) -> str:
90
+ return re.sub(r"\s+", " ", text).strip().lower()
91
+
92
+
93
+ #: Title strings for the families that cross-cite each other. NOTE: "gnu
94
+ #: affero general public license" does NOT contain "gnu general public
95
+ #: license" as a substring ("affero" splits it), so the titles are disjoint.
96
+ _HEADER_TITLES: tuple[tuple[str, str], ...] = (
97
+ ("gnu affero general public license", "agpl"),
98
+ ("gnu lesser general public license", "lgpl"),
99
+ ("gnu general public license", "gpl"),
100
+ ("mozilla public license", "mpl"),
101
+ ("apache license", "apache"),
102
+ )
103
+
104
+
105
+ def _classify_header(header: str) -> str | None:
106
+ """Title-anchored families (GNU/MPL/Apache). The EARLIEST-occurring title
107
+ wins, because a license document's own title always precedes any
108
+ cross-citation — GPL-3 §13 names the AGPL and AGPL-3 §13 names the GPL,
109
+ so mere presence-matching misclassifies one as the other (the V9 trap)."""
110
+ hits = [(idx, family) for title, family in _HEADER_TITLES if (idx := header.find(title)) >= 0]
111
+ if not hits:
112
+ return None
113
+ family = min(hits)[1]
114
+ if family == "agpl":
115
+ return "AGPL-3.0"
116
+ if family == "lgpl":
117
+ return "LGPL-2.1" if "version 2.1" in header else "LGPL-3.0"
118
+ if family == "gpl":
119
+ return "GPL-2.0" if "version 2" in header and "version 3" not in header else "GPL-3.0"
120
+ if family == "mpl":
121
+ return "MPL-2.0" if "2.0" in header else None
122
+ return "Apache-2.0" if "version 2.0" in header else None
123
+
124
+
125
+ def _classify_body(body: str) -> str | None:
126
+ """Distinctive-phrase matching for the short permissive licenses. Order
127
+ matters: ISC's grant ("permission to use, copy, modify") is checked via
128
+ its unique "and/or distribute … with or without fee" wording so it can't
129
+ be shadowed by MIT's "permission is hereby granted" phrase."""
130
+ if "free and unencumbered software released into the public domain" in body:
131
+ return "Unlicense"
132
+ # Apache boilerplate notice files ("Licensed under the Apache License,
133
+ # Version 2.0") carry the title mid-body, not in a header — the tricky
134
+ # Apache-vs-MIT case: both say "permission"/"license", only Apache names
135
+ # itself with a version.
136
+ if "apache license" in body and "version 2.0" in body:
137
+ return "Apache-2.0"
138
+ if "permission is hereby granted, free of charge, to any person obtaining a copy" in body:
139
+ return "MIT"
140
+ if "with or without fee" in body and "permission to use, copy, modify" in body:
141
+ return "ISC"
142
+ if "redistribution and use in source and binary forms" in body:
143
+ # The third clause ("Neither the name of …") is what separates 3- from 2-.
144
+ return "BSD-3-Clause" if "neither the name of" in body else "BSD-2-Clause"
145
+ return None
146
+
147
+
148
+ def _classify_license_text(text: str) -> str:
149
+ norm = _normalize_text(text)
150
+ return _classify_header(norm[:400]) or _classify_body(norm) or "unknown"
151
+
152
+
153
+ def _classifier_signal(repo_root: Path) -> tuple[str, str] | None:
154
+ """Secondary signal: trove classifiers / PEP 639 license expression in
155
+ pyproject.toml or setup.py. Regex-scan, not a TOML parse — the gate must
156
+ not depend on packaging libs and classifiers are line-shaped in practice."""
157
+ for name in ("pyproject.toml", "setup.py"):
158
+ path = repo_root / name
159
+ if not path.is_file():
160
+ continue
161
+ try:
162
+ text = path.read_text(encoding="utf-8", errors="replace")
163
+ except OSError:
164
+ continue
165
+ # PEP 639: license = "Apache-2.0" (pyproject only, but harmless on setup.py).
166
+ m = re.search(r'license\s*=\s*["\']([A-Za-z0-9.+-]+)["\']', text)
167
+ if m and m.group(1).lower() in _SPDX_LOOKUP:
168
+ return _SPDX_LOOKUP[m.group(1).lower()], name
169
+ low = _normalize_text(text)
170
+ for fragment, spdx in _CLASSIFIER_MAP:
171
+ if f"license :: osi approved :: {fragment}" in low or (
172
+ "license ::" in low and fragment in low
173
+ ):
174
+ return spdx, name
175
+ return None
176
+
177
+
178
+ def detect_license(repo_root: Path) -> LicenseInfo:
179
+ """Detect the repo license. LICENSE-file text is the primary signal;
180
+ packaging classifiers are secondary (used only when the file is absent or
181
+ unclassifiable). The winning signal is recorded so corpus manifests can
182
+ show provenance for the tier decision (V9 closure must be auditable)."""
183
+ for name in _LICENSE_FILENAMES:
184
+ path = repo_root / name
185
+ if not path.is_file():
186
+ # Case-insensitive fallback (e.g. "License.md", "COPYING.txt" not
187
+ # matched here on purpose — only exact-name case variants).
188
+ matches = [p for p in repo_root.glob("*") if p.is_file() and p.name.lower() == name.lower()]
189
+ path = matches[0] if matches else path
190
+ if path.is_file():
191
+ try:
192
+ text = path.read_text(encoding="utf-8", errors="replace")
193
+ except OSError:
194
+ continue
195
+ spdx = _classify_license_text(text)
196
+ if spdx != "unknown":
197
+ return LicenseInfo(spdx_id=spdx, signal="license_file", source=path.name)
198
+ # File exists but unclassifiable → let the classifier signal try
199
+ # before giving up; remember we saw a file for the "none" case.
200
+ fallback = _classifier_signal(repo_root)
201
+ if fallback is not None:
202
+ return LicenseInfo(spdx_id=fallback[0], signal="classifier", source=fallback[1])
203
+ return LicenseInfo(spdx_id="unknown", signal="license_file", source=path.name)
204
+ fallback = _classifier_signal(repo_root)
205
+ if fallback is not None:
206
+ return LicenseInfo(spdx_id=fallback[0], signal="classifier", source=fallback[1])
207
+ return LicenseInfo(spdx_id="unknown", signal="none")
208
+
209
+
210
+ # ---------------------------------------------------------------------------
211
+ # License tiers (D-13: tiers, not a boolean)
212
+ # ---------------------------------------------------------------------------
213
+
214
+
215
+ class Tier(Enum):
216
+ """Three-way license verdict. TRAINABLE_ONLY exists because weak copyleft
217
+ (MPL/LGPL) permits training but redistribution of derivative diffs would
218
+ trigger copyleft obligations — collapsing this to a boolean either loses
219
+ training data or leaks copyleft material into published corpora (D-13)."""
220
+
221
+ REDISTRIBUTABLE = "redistributable"
222
+ TRAINABLE_ONLY = "trainable_only"
223
+ EXCLUDED = "excluded"
224
+
225
+
226
+ _TIER_BY_SPDX: dict[str, Tier] = {
227
+ "MIT": Tier.REDISTRIBUTABLE,
228
+ "Apache-2.0": Tier.REDISTRIBUTABLE,
229
+ "BSD-2-Clause": Tier.REDISTRIBUTABLE,
230
+ "BSD-3-Clause": Tier.REDISTRIBUTABLE,
231
+ "ISC": Tier.REDISTRIBUTABLE,
232
+ "Unlicense": Tier.REDISTRIBUTABLE,
233
+ "MPL-2.0": Tier.TRAINABLE_ONLY,
234
+ "LGPL-2.1": Tier.TRAINABLE_ONLY,
235
+ "LGPL-3.0": Tier.TRAINABLE_ONLY,
236
+ # GPL/AGPL and unknown are EXCLUDED: strong copyleft would bind the model
237
+ # outputs' redistribution story, and "unknown" defaults closed (V9).
238
+ "GPL-2.0": Tier.EXCLUDED,
239
+ "GPL-3.0": Tier.EXCLUDED,
240
+ "AGPL-3.0": Tier.EXCLUDED,
241
+ }
242
+
243
+
244
+ def license_tier(info: LicenseInfo) -> Tier:
245
+ """Map detected license → tier. Anything unrecognized is EXCLUDED — the
246
+ gate fails closed, never open (V9: the old substring filter failed open)."""
247
+ return _TIER_BY_SPDX.get(info.spdx_id, Tier.EXCLUDED)
248
+
249
+
250
+ # ---------------------------------------------------------------------------
251
+ # Benchmark decontamination (V3 / D-5)
252
+ # ---------------------------------------------------------------------------
253
+
254
+ #: The canonical 12 SWE-bench test repos (SWE-bench / -Lite / -Verified /
255
+ #: -Multimodal all draw eval instances from these). Training on ANY of them
256
+ #: contaminates every SWE-bench-family score we report (V3). Lowercase
257
+ #: "org/repo" form. Extend via a JSON file (list of "org/repo" strings)
258
+ #: passed to is_eval_contaminated(extra_list=...) — e.g. SWE-Gym eval splits.
259
+ DECONTAMINATION_LIST: frozenset[str] = frozenset(
260
+ {
261
+ "astropy/astropy",
262
+ "django/django",
263
+ "matplotlib/matplotlib",
264
+ "mwaskom/seaborn",
265
+ "pallets/flask",
266
+ "psf/requests",
267
+ "pydata/xarray",
268
+ "pylint-dev/pylint",
269
+ "pytest-dev/pytest",
270
+ "scikit-learn/scikit-learn",
271
+ "sphinx-doc/sphinx",
272
+ "sympy/sympy",
273
+ }
274
+ )
275
+
276
+
277
+ def load_decontamination_list(path: Path) -> frozenset[str]:
278
+ """Load an extension list from a JSON file: ``["org/repo", ...]``. This is
279
+ THE documented mechanism for adding eval repos (new SWE-bench releases,
280
+ SWE-Gym eval splits) without editing code."""
281
+ entries = json.loads(path.read_text(encoding="utf-8"))
282
+ if not isinstance(entries, list):
283
+ raise ValueError(f"{path}: decontamination JSON must be a list of 'org/repo' strings")
284
+ return frozenset(normalize_repo(str(e)) for e in entries)
285
+
286
+
287
+ def normalize_repo(repo: str) -> str:
288
+ """Reduce any repo spelling — full https/ssh GitHub URL, trailing ``.git``,
289
+ mixed case — to lowercase ``org/repo``. Decontamination must hit no matter
290
+ how the driver spells the repo (V3: a miss here is silent contamination)."""
291
+ r = repo.strip().lower()
292
+ r = re.sub(r"^(https?://|git@)", "", r)
293
+ r = re.sub(r"^[^/]*github\.com[:/]", "", r)
294
+ r = r.rstrip("/")
295
+ r = r.removesuffix(".git")
296
+ parts = [p for p in r.split("/") if p]
297
+ return "/".join(parts[:2]) if len(parts) >= 2 else r
298
+
299
+
300
+ def is_eval_contaminated(repo: str, extra_list: frozenset[str] | None = None) -> bool:
301
+ """True if ``repo`` is in the SWE-bench-family eval set (or the caller's
302
+ extension list). Case-insensitive; accepts URLs and bare org/repo."""
303
+ key = normalize_repo(repo)
304
+ return key in DECONTAMINATION_LIST or (extra_list is not None and key in extra_list)
305
+
306
+
307
+ # ---------------------------------------------------------------------------
308
+ # The gate verdict — single entry point for the pipeline driver
309
+ # ---------------------------------------------------------------------------
310
+
311
+
312
+ @dataclass
313
+ class GateVerdict:
314
+ """Everything the driver needs to admit/reject a repo, with reasons kept
315
+ for the run manifest (step 6's lineage record)."""
316
+
317
+ repo: str
318
+ license_info: LicenseInfo
319
+ tier: Tier
320
+ contaminated: bool
321
+ admitted: bool
322
+ reasons: list[str] = field(default_factory=list)
323
+
324
+
325
+ def gate_repo(repo: str, repo_root: Path | None, extra_decontamination: frozenset[str] | None = None) -> GateVerdict:
326
+ """Architecture step 1: the one call the pipeline driver makes per repo.
327
+
328
+ Hard rules (in priority order):
329
+ 1. Contaminated (V3) → NEVER admitted, even if the license is permissive.
330
+ 2. Tier EXCLUDED (GPL/AGPL/unknown) → not admitted (V9: fail closed).
331
+ 3. Tier TRAINABLE_ONLY → admitted, with the do-not-redistribute
332
+ constraint recorded as a reason so step 6 can route the rows.
333
+ """
334
+ contaminated = is_eval_contaminated(repo, extra_decontamination)
335
+ info = detect_license(repo_root) if repo_root is not None else LicenseInfo("unknown", "none")
336
+ tier = license_tier(info)
337
+
338
+ reasons: list[str] = []
339
+ if contaminated:
340
+ reasons.append(
341
+ f"benchmark decontamination: {normalize_repo(repo)} is a SWE-bench-family eval repo (V3/D-5)"
342
+ )
343
+ if repo_root is None:
344
+ reasons.append("no repo_root provided: license undetectable, failing closed (V9)")
345
+ if tier is Tier.EXCLUDED and not contaminated:
346
+ reasons.append(f"license tier EXCLUDED: spdx={info.spdx_id} (signal={info.signal})")
347
+ if tier is Tier.TRAINABLE_ONLY:
348
+ reasons.append(
349
+ f"license tier TRAINABLE_ONLY: spdx={info.spdx_id} — usable for training, "
350
+ "derivative diffs must NOT be redistributed (D-13)"
351
+ )
352
+
353
+ admitted = (not contaminated) and tier is not Tier.EXCLUDED
354
+ return GateVerdict(
355
+ repo=repo,
356
+ license_info=info,
357
+ tier=tier,
358
+ contaminated=contaminated,
359
+ admitted=admitted,
360
+ reasons=reasons,
361
+ )
composer_replication/datagen/rollout_harness.py ADDED
@@ -0,0 +1,214 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """rollout_harness.py — the agent loop over FeatureDeletionEnv (finding V2).
2
+
3
+ THE critical missing component the design critic identified: nothing in the
4
+ repo ran an agent episode against `FeatureDeletionEnv` to completion, so the
5
+ SFT corpus had NO producer and the tree-of-work had no env-grounded seeds.
6
+ `collect_trajectory` is that producer: prompt → policy.act → env.step → … →
7
+ submit → `_grade()`, emitting a `CanonicalTrajectory` whose steps are real
8
+ executed environment transitions (the seeds the tree needs, fixing the
9
+ seed-trace/oracle disjointness of finding D-1 as a free byproduct).
10
+
11
+ The policy is pluggable (`RolloutPolicy` protocol): a scripted fake for tests,
12
+ a frontier API model for expert-trajectory collection (SWE-Gym/SWE-smith both
13
+ validated this recipe — 491 and 5,016 expert trajectories respectively), or a
14
+ local model later.
15
+ """
16
+ from __future__ import annotations
17
+
18
+ from dataclasses import dataclass
19
+ from typing import Protocol, runtime_checkable
20
+
21
+ from composer_replication.datagen.env import FeatureDeletionEnv, StepResult
22
+ from composer_replication.datagen.schema import FeatureDeletionTask
23
+ from composer_replication.datagen.trajectory import (
24
+ CanonicalTrajectory,
25
+ ToolCall,
26
+ TrajectoryStep,
27
+ )
28
+
29
+
30
+ @runtime_checkable
31
+ class RolloutPolicy(Protocol):
32
+ """Anything that maps (observation, history) → the next action.
33
+
34
+ Returning a `ToolCall` continues the episode (translated to an env action
35
+ dict); returning a plain `str` is the final message — the harness submits.
36
+ """
37
+
38
+ def act(self, observation: str, history: list[TrajectoryStep]) -> ToolCall | str: ...
39
+
40
+
41
+ @dataclass
42
+ class ScriptedPolicy:
43
+ """Test fake: replays a fixed action list, then submits."""
44
+
45
+ actions: list[ToolCall | str]
46
+ _i: int = 0
47
+
48
+ def act(self, observation: str, history: list[TrajectoryStep]) -> ToolCall | str:
49
+ if self._i >= len(self.actions):
50
+ return "done" # str → submit
51
+ a = self.actions[self._i]
52
+ self._i += 1
53
+ return a
54
+
55
+
56
+ class OpenRouterPolicy:
57
+ """Frontier-API policy for expert-trajectory collection (thin stub).
58
+
59
+ Mirrors `teacher_replay._call_teacher`'s payload shape (one chat call,
60
+ temperature 0.2). Lazy-deps on httpx so the module imports without it.
61
+ Deliberately minimal: real expert collection should evaluate adopting
62
+ mini-swe-agent/SWE-agent as the scaffold (deepread 11 finding 2) — this
63
+ class exists so the harness has a live-API path without a new framework.
64
+ """
65
+
66
+ def __init__(self, model_slug: str, api_key: str | None = None,
67
+ max_tokens: int = 512) -> None:
68
+ try:
69
+ import httpx # noqa: F401, PLC0415 — lazy heavy dep
70
+ except ImportError as e:
71
+ raise ImportError(
72
+ "OpenRouterPolicy requires httpx (`pip install httpx` or the "
73
+ "[serverless] extra). For tests use ScriptedPolicy. Got: " + repr(e)
74
+ ) from e
75
+ from composer_replication.teacher_replay import _load_api_key
76
+ self.model_slug = model_slug
77
+ self.api_key = api_key or _load_api_key()
78
+ self.max_tokens = max_tokens
79
+
80
+ def act(self, observation: str, history: list[TrajectoryStep]) -> ToolCall | str:
81
+ import httpx # noqa: PLC0415
82
+
83
+ from composer_replication.teacher_replay import OPENROUTER_URL
84
+ messages = [{"role": "user", "content": observation}]
85
+ r = httpx.post(
86
+ OPENROUTER_URL,
87
+ json={"model": self.model_slug, "messages": messages,
88
+ "max_tokens": self.max_tokens, "temperature": 0.2},
89
+ headers={"Authorization": f"Bearer {self.api_key}"},
90
+ timeout=120.0,
91
+ )
92
+ r.raise_for_status()
93
+ return str(r.json()["choices"][0]["message"]["content"])
94
+
95
+
96
+ def _to_env_action(call: ToolCall) -> dict:
97
+ """ToolCall → FeatureDeletionEnv action dict.
98
+
99
+ CONVENTION (documented here, the single translation point): the env's
100
+ `step()` consumes ``{"type": <tool name>, **args}``; ``type=="submit"``
101
+ triggers grading (env.py:67). A ToolCall named "submit" therefore ends the
102
+ episode through the same path as a plain-text final message.
103
+ """
104
+ return {"type": call.name, **call.args}
105
+
106
+
107
+ def collect_trajectory(
108
+ env: FeatureDeletionEnv,
109
+ task: FeatureDeletionTask,
110
+ policy: RolloutPolicy,
111
+ *,
112
+ max_turns: int = 40,
113
+ budget_usd: float | None = None,
114
+ provenance: dict | None = None,
115
+ ) -> CanonicalTrajectory:
116
+ """Run one episode and return the graded CanonicalTrajectory.
117
+
118
+ The episode ends when the policy emits a plain string (final message →
119
+ submit), a ToolCall named "submit", or `max_turns` is hit (the env grades
120
+ on its own turn limit too — we mirror it here so the harness's history
121
+ stays aligned with the env's accounting).
122
+ """
123
+ obs = env.reset(task)
124
+ steps: list[TrajectoryStep] = []
125
+ final: StepResult | None = None
126
+
127
+ for _ in range(max_turns):
128
+ action = policy.act(obs, steps)
129
+ if isinstance(action, str) or action.name == "submit":
130
+ final = env.step({"type": "submit"})
131
+ steps.append(TrajectoryStep(
132
+ observation=obs, action=action, result=final.observation,
133
+ tool_error=False,
134
+ ))
135
+ break
136
+ res = env.step(_to_env_action(action))
137
+ tool_error = "error" in (res.observation or "").lower()[:200]
138
+ steps.append(TrajectoryStep(
139
+ observation=obs, action=action, result=res.observation,
140
+ tool_error=tool_error,
141
+ ))
142
+ if res.done: # env hit its own turn limit and graded
143
+ final = res
144
+ break
145
+ obs = res.observation
146
+
147
+ if final is None:
148
+ # max_turns exhausted without submit — grade what exists.
149
+ final = env.step({"type": "submit"})
150
+
151
+ info = final.info or {}
152
+ return CanonicalTrajectory(
153
+ task_id=task.task_id,
154
+ steps=steps,
155
+ grade=float(final.reward) if final.reward is not None else None,
156
+ guard_ok=bool(info.get("guard_ok", True)),
157
+ hacked=bool(info.get("hacked", False)),
158
+ provenance={"source": "rollout_harness",
159
+ "policy": type(policy).__name__,
160
+ **(provenance or {})},
161
+ )
162
+
163
+
164
+ # ---------------------------------------------------------------------
165
+ # Admission — type the signal and route it (final report §4)
166
+ # ---------------------------------------------------------------------
167
+
168
+
169
+ @dataclass(frozen=True)
170
+ class AdmissionVerdict:
171
+ """Where a trajectory may go. Routing per the typed-train-on-all verdict:
172
+ clean full passes → SFT; clean near-misses → DPO-candidate (contrastive
173
+ rejected vs a winner, never raw negative gradient); everything else →
174
+ rejected (quarantine-side, full provenance kept for audit)."""
175
+
176
+ sft_admitted: bool
177
+ dpo_candidate: bool
178
+ rejected: bool
179
+ reasons: tuple[str, ...]
180
+
181
+
182
+ def admit(traj: CanonicalTrajectory) -> AdmissionVerdict:
183
+ reasons: list[str] = []
184
+ clean = traj.guard_ok and not traj.hacked
185
+ if not traj.guard_ok:
186
+ reasons.append("pass_to_pass guard broken")
187
+ if traj.hacked:
188
+ reasons.append("hack monitor flagged")
189
+ grade = traj.grade if traj.grade is not None else 0.0
190
+ if traj.grade is None:
191
+ reasons.append("ungraded (no execution oracle)")
192
+
193
+ sft = clean and traj.grade is not None and grade == 1.0
194
+ dpo = clean and traj.grade is not None and 0.0 < grade < 1.0
195
+ if sft:
196
+ reasons.append("clean full pass")
197
+ elif dpo:
198
+ reasons.append(f"clean near-miss (grade={grade:.2f})")
199
+ elif clean and grade == 0.0 and traj.grade is not None:
200
+ reasons.append("clean zero — no partial signal")
201
+ return AdmissionVerdict(
202
+ sft_admitted=sft, dpo_candidate=dpo,
203
+ rejected=not (sft or dpo), reasons=tuple(reasons),
204
+ )
205
+
206
+
207
+ __all__ = [
208
+ "RolloutPolicy",
209
+ "ScriptedPolicy",
210
+ "OpenRouterPolicy",
211
+ "collect_trajectory",
212
+ "AdmissionVerdict",
213
+ "admit",
214
+ ]
composer_replication/datagen/swesmith_adapter.py ADDED
@@ -0,0 +1,269 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """swesmith_adapter.py — adapt SWE-smith instances into Feature-Deletion tasks.
2
+
3
+ THE BUY-VS-BUILD VERDICT (deepread finding V4 / D-6): `pip install swesmith`
4
+ (MIT) already ships what ADR-010's "Option B greenfield generator" would have
5
+ hand-built — env construction from arbitrary GitHub repos (ONE Docker image per
6
+ repo, ~500x more storage-efficient than per-task images), five bug-synthesis
7
+ strategies, issue-text generation, and validation-by-test-execution, at a
8
+ verified $1,360 + ~20 human-hours for 50k tasks. Its **PR Mirror strategy is
9
+ exactly this repo's gold-patch-reversion mechanic** and SWE-smith's own ablation
10
+ (Table 5, arXiv:2504.21798) shows PR-Mirror trajectories train the BEST models
11
+ of its five strategies — independent validation of ADR-010's core approach.
12
+ So SWE-smith is the synthesis ENGINE for "point at a repo"; this module is the
13
+ schema bridge into the existing `FeatureDeletionTask` world.
14
+
15
+ THE SEMANTIC INVERSION (load-bearing — easy to get backwards):
16
+ * SWE-bench-shaped instances: `patch` is the GOLD FIX. broken = HEAD with the
17
+ fix reverted (`git apply -R patch`). `SweBenchAdapter` stores `patch` as
18
+ `golden_diff` directly.
19
+ * SWE-smith instances: `patch` INTRODUCES THE BUG. broken = HEAD with the
20
+ patch APPLIED. The fix — what the agent must produce, the validator's gate-4
21
+ restoration diff — is the REVERSE of the bug patch.
22
+ This adapter therefore stores `golden_diff = reverse_unified_diff(bug_patch)`.
23
+ When mechanical reversal fails (exotic diff features), it falls back to the
24
+ original patch tagged with a provenance marker so downstream gate-4 validation
25
+ knows to use `git apply -R` instead of `git apply`.
26
+
27
+ The adapter itself needs nothing beyond core deps. Live synthesis (building new
28
+ repo profiles / generating new bugs) needs the `swesmith` toolkit + Docker on
29
+ Linux — see the `[swesmith]` extra in pyproject.
30
+ """
31
+ from __future__ import annotations
32
+
33
+ import json
34
+ import re
35
+ from dataclasses import dataclass
36
+
37
+ from composer_replication.datagen.schema import FeatureDeletionTask
38
+ from composer_replication.datagen.substrates import _as_tuple
39
+
40
+ #: Marker prefixed to golden_diff when reverse_unified_diff could not invert the
41
+ #: bug patch mechanically. Consumers (gate-4 validation) must then apply the
42
+ #: remainder with `git apply -R` (it is the FORWARD bug patch, not the fix).
43
+ UNREVERSED_MARKER = "### UNREVERSED-BUG-PATCH (apply with -R) ###\n"
44
+
45
+ #: instance_id substring patterns → synthesis strategy (SWE-smith §2.1 / §B).
46
+ #: Patterns follow the toolkit's naming: e.g.
47
+ #: pandas-dev__pandas.abc123.lm_modify__xyz
48
+ #: ...func_pm_ctrl_invert_if__..., ...combine_file__..., ...pr_1234
49
+ _STRATEGY_PATTERNS: tuple[tuple[str, str], ...] = (
50
+ ("lm_modify", "lm_modify"),
51
+ ("lm_rewrite", "lm_rewrite"),
52
+ ("func_pm", "procedural"), # procedural AST modifications (13 transform types)
53
+ ("func_basic", "procedural"),
54
+ ("combine_file", "combine"),
55
+ ("combine_module", "combine"),
56
+ ("combine", "combine"),
57
+ ("pr_", "pr_mirror"),
58
+ )
59
+
60
+
61
+ def parse_strategy(instance_id: str) -> str:
62
+ """Map a SWE-smith instance_id to its bug-synthesis strategy.
63
+
64
+ Returns one of {lm_modify, lm_rewrite, procedural, combine, pr_mirror,
65
+ unknown}. The strategy matters because SWE-smith's Table 5 ablation found
66
+ trajectory quality differs sharply by strategy (PR Mirror best, LM Modify
67
+ steep drop-off) — we carry it as provenance so corpus builds can weight or
68
+ filter by strategy.
69
+ """
70
+ iid = (instance_id or "").lower()
71
+ for pattern, strategy in _STRATEGY_PATTERNS:
72
+ if pattern in iid:
73
+ return strategy
74
+ return "unknown"
75
+
76
+
77
+ #: Heuristic cold-start difficulty priors per strategy, motivated by SWE-smith
78
+ #: Table 1 medians (PR Mirror: 3 median F2P but 14 lines edited; Combine: 15
79
+ #: F2P / 11 lines = multi-site; procedural: 7 F2P / 5 lines, mechanical).
80
+ #: These only seed DifficultyCurriculum's p-hat before real rollouts exist.
81
+ _DIFFICULTY_PRIOR: dict[str, float] = {
82
+ "pr_mirror": 0.4,
83
+ "combine": 0.4,
84
+ "lm_rewrite": 0.45,
85
+ "lm_modify": 0.55,
86
+ "procedural": 0.6,
87
+ "unknown": 0.5,
88
+ }
89
+
90
+
91
+ _HUNK_RE = re.compile(
92
+ r"^@@ -(?P<old_start>\d+)(?:,(?P<old_count>\d+))? "
93
+ r"\+(?P<new_start>\d+)(?:,(?P<new_count>\d+))? @@(?P<tail>.*)$"
94
+ )
95
+
96
+
97
+ def reverse_unified_diff(patch: str) -> str | None:
98
+ """Mechanically invert a unified diff (swap additions and deletions).
99
+
100
+ Handles the standard unified-diff features SWE-smith patches use:
101
+ ``diff --git`` headers, ``---``/``+++`` file lines, ``@@`` hunk headers
102
+ (old/new ranges swapped), ``+``/``-`` body lines (swapped), context lines,
103
+ and ``\`` markers (kept in place).
104
+
105
+ HONEST LIMITATIONS (returns None — caller falls back to UNREVERSED_MARKER):
106
+ * file mode changes (``old mode``/``new mode``), renames/copies
107
+ (``rename from``...), binary patches (``GIT binary patch``), and
108
+ ``index`` lines with mode suffixes are NOT inverted — reversing them
109
+ correctly requires git plumbing, not text surgery.
110
+ * Within a hunk, a reversed diff's line ORDER for paired -/+ runs is the
111
+ naive swap; `git apply` accepts it, but it is not byte-identical to
112
+ what `git diff` would emit for the reverse change.
113
+ """
114
+ if not patch or "@@" not in patch:
115
+ return None
116
+ unsupported = ("old mode ", "new mode ", "rename from ", "rename to ",
117
+ "copy from ", "copy to ", "GIT binary patch")
118
+ if any(marker in patch for marker in unsupported):
119
+ return None
120
+
121
+ out: list[str] = []
122
+ for line in patch.splitlines():
123
+ if line.startswith("diff --git "):
124
+ # `diff --git a/<old> b/<new>` → swap the two paths.
125
+ m = re.match(r"^diff --git a/(?P<a>.+) b/(?P<b>.+)$", line)
126
+ if m:
127
+ out.append(f"diff --git a/{m.group('b')} b/{m.group('a')}")
128
+ else:
129
+ out.append(line)
130
+ elif line.startswith("--- "):
131
+ out.append("+++ " + line[4:].replace("a/", "b/", 1)
132
+ if line[4:].startswith("a/") else "+++ " + line[4:])
133
+ elif line.startswith("+++ "):
134
+ out.append("--- " + line[4:].replace("b/", "a/", 1)
135
+ if line[4:].startswith("b/") else "--- " + line[4:])
136
+ elif line.startswith("@@"):
137
+ m = _HUNK_RE.match(line)
138
+ if not m:
139
+ return None
140
+ old_start, old_count = m.group("old_start"), m.group("old_count")
141
+ new_start, new_count = m.group("new_start"), m.group("new_count")
142
+ oc = f",{old_count}" if old_count is not None else ""
143
+ nc = f",{new_count}" if new_count is not None else ""
144
+ out.append(f"@@ -{new_start}{nc} +{old_start}{oc} @@{m.group('tail')}")
145
+ elif line.startswith("+"):
146
+ out.append("-" + line[1:])
147
+ elif line.startswith("-"):
148
+ out.append("+" + line[1:])
149
+ else:
150
+ # context lines, `index ...`, `\ No newline...` pass through.
151
+ out.append(line)
152
+ return "\n".join(out) + ("\n" if patch.endswith("\n") else "")
153
+
154
+
155
+ @dataclass(frozen=True)
156
+ class SwesmithMeta:
157
+ """Sidecar provenance for a SWE-smith-derived task.
158
+
159
+ Kept OUT of the frozen `FeatureDeletionTask` schema deliberately — the
160
+ schema is shared with SweBenchAdapter and the trainer; strategy provenance
161
+ is a corpus-construction concern, carried alongside (e.g. into the run
162
+ manifest), never into the policy-visible task row.
163
+ """
164
+
165
+ strategy: str # lm_modify | lm_rewrite | procedural | combine | pr_mirror | unknown
166
+ diff_reversed: bool # True if golden_diff is the mechanical reverse of the bug patch
167
+ source: str = "swesmith"
168
+
169
+
170
+ @dataclass
171
+ class SwesmithAdapter:
172
+ """Convert a SWE-smith instance dict into a FeatureDeletionTask.
173
+
174
+ Mirrors `SweBenchAdapter`'s shape; differs in the patch semantics (see the
175
+ module docstring INVERSION note) and the per-REPO image convention.
176
+ """
177
+
178
+ default_test_command: str = "python -m pytest -q"
179
+
180
+ def image_for(self, instance: dict) -> str:
181
+ # SWE-smith publishes ONE image per repo (not per task). Rows carry
182
+ # `image_name`; some exports use `docker_image`. Fall back to the
183
+ # toolkit's naming convention derived from the repo slug.
184
+ for key in ("image_name", "docker_image"):
185
+ if instance.get(key):
186
+ return str(instance[key])
187
+ repo = str(instance.get("repo", "unknown")).replace("/", "__").lower()
188
+ return f"swesmith.x86_64.{repo}:latest"
189
+
190
+ def to_task(self, instance: dict) -> FeatureDeletionTask:
191
+ task, _meta = self.to_task_with_meta(instance)
192
+ return task
193
+
194
+ def to_task_with_meta(self, instance: dict) -> tuple[FeatureDeletionTask, SwesmithMeta]:
195
+ iid = str(instance.get("instance_id") or instance.get("task_id") or "unknown")
196
+ strategy = parse_strategy(iid)
197
+
198
+ bug_patch = str(instance.get("patch", ""))
199
+ fix = reverse_unified_diff(bug_patch)
200
+ if fix is not None:
201
+ golden_diff = fix
202
+ diff_reversed = True
203
+ else:
204
+ golden_diff = UNREVERSED_MARKER + bug_patch
205
+ diff_reversed = False
206
+
207
+ ftp = _as_tuple(instance.get("FAIL_TO_PASS"))
208
+ ptp = _as_tuple(instance.get("PASS_TO_PASS"))
209
+
210
+ task = FeatureDeletionTask(
211
+ task_id=iid,
212
+ repo=str(instance.get("repo", "unknown")),
213
+ base_commit=str(instance.get("base_commit", "")),
214
+ broken_image=self.image_for(instance),
215
+ test_command=str(instance.get("test_command") or self.default_test_command),
216
+ fail_to_pass=ftp,
217
+ pass_to_pass=ptp,
218
+ golden_diff=golden_diff,
219
+ granularity="feature",
220
+ # SWE-smith rows don't carry per-instance licenses; repo-level
221
+ # licensing is the repo_gate's job (deepread finding V9/D-13).
222
+ upstream_license=str(instance.get("license_name", "unknown")),
223
+ difficulty_prior=_DIFFICULTY_PRIOR.get(strategy, 0.5),
224
+ )
225
+ return task, SwesmithMeta(strategy=strategy, diff_reversed=diff_reversed)
226
+
227
+
228
+ def load_swesmith_instances(
229
+ path_or_hf_id: str,
230
+ *,
231
+ limit: int | None = None,
232
+ ) -> list[dict]:
233
+ """Load SWE-smith instances from a local JSONL file or the HF dataset.
234
+
235
+ Local ``.jsonl`` paths need no extra deps (used by tests/fixtures). HF ids
236
+ (e.g. ``SWE-bench/SWE-smith``) lazy-import `datasets` from the `[datagen]`
237
+ extra.
238
+ """
239
+ if path_or_hf_id.endswith(".jsonl"):
240
+ rows: list[dict] = []
241
+ with open(path_or_hf_id, encoding="utf-8") as f:
242
+ for line in f:
243
+ line = line.strip()
244
+ if not line:
245
+ continue
246
+ rows.append(json.loads(line))
247
+ if limit is not None and len(rows) >= limit:
248
+ break
249
+ return rows
250
+ try:
251
+ from datasets import load_dataset # noqa: PLC0415 — lazy heavy dep
252
+ except ImportError as e:
253
+ raise RuntimeError(
254
+ "Loading SWE-smith from the HF Hub requires `datasets`; install "
255
+ "with `pip install -e .[datagen]`. Got: " + repr(e)
256
+ ) from e
257
+ split = load_dataset(path_or_hf_id, split="train")
258
+ rows = [dict(r) for i, r in enumerate(split) if limit is None or i < limit]
259
+ return rows[: limit if limit is not None else len(rows)]
260
+
261
+
262
+ __all__ = [
263
+ "SwesmithAdapter",
264
+ "SwesmithMeta",
265
+ "UNREVERSED_MARKER",
266
+ "load_swesmith_instances",
267
+ "parse_strategy",
268
+ "reverse_unified_diff",
269
+ ]
composer_replication/datagen/tests/test_repo_gate.py ADDED
@@ -0,0 +1,419 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Tests for the Stage-0 ingest gate (repo_gate.py) — architecture step 1.
2
+
3
+ Coverage targets the two findings the module closes:
4
+ * V9/D-13 — SPDX detection from real license fixture texts (incl. the
5
+ tricky GNU-family cross-citation and Apache-vs-MIT phrasing) and the
6
+ three-tier mapping that replaces the old boolean substring filter.
7
+ * V3/D-5 — decontamination hits in exact, URL, and mixed-case forms, and
8
+ the hard never-admit rule in the composed verdict.
9
+
10
+ CPU-only, stdlib + tmp_path fixtures — no network, no Docker.
11
+ """
12
+ from __future__ import annotations
13
+
14
+ import json
15
+ from pathlib import Path
16
+
17
+ import pytest
18
+
19
+ from composer_replication.datagen.repo_gate import (
20
+ DECONTAMINATION_LIST,
21
+ GateVerdict,
22
+ LicenseInfo,
23
+ Tier,
24
+ detect_license,
25
+ gate_repo,
26
+ is_eval_contaminated,
27
+ license_tier,
28
+ load_decontamination_list,
29
+ normalize_repo,
30
+ )
31
+
32
+ # ---------------------------------------------------------------------
33
+ # License fixture texts — distinctive excerpts of the real license texts.
34
+ # ---------------------------------------------------------------------
35
+
36
+ MIT_TEXT = """\
37
+ MIT License
38
+
39
+ Copyright (c) 2026 Example Org
40
+
41
+ Permission is hereby granted, free of charge, to any person obtaining a copy
42
+ of this software and associated documentation files (the "Software"), to deal
43
+ in the Software without restriction...
44
+ """
45
+
46
+ # The tricky Apache case: the words "permission" and "license" appear in both
47
+ # MIT and Apache; only Apache names itself with a version.
48
+ APACHE_TEXT = """\
49
+ Apache License
50
+ Version 2.0, January 2004
51
+ http://www.apache.org/licenses/
52
+
53
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
54
+
55
+ 1. Definitions.
56
+ "License" shall mean the terms and conditions for use, reproduction,
57
+ and distribution as defined by Sections 1 through 9 of this document.
58
+ """
59
+
60
+ BSD3_TEXT = """\
61
+ BSD 3-Clause License
62
+
63
+ Redistribution and use in source and binary forms, with or without
64
+ modification, are permitted provided that the following conditions are met:
65
+
66
+ 1. Redistributions of source code must retain the above copyright notice...
67
+ 3. Neither the name of the copyright holder nor the names of its
68
+ contributors may be used to endorse or promote products derived from
69
+ this software without specific prior written permission.
70
+ """
71
+
72
+ BSD2_TEXT = """\
73
+ BSD 2-Clause License
74
+
75
+ Redistribution and use in source and binary forms, with or without
76
+ modification, are permitted provided that the following conditions are met:
77
+
78
+ 1. Redistributions of source code must retain the above copyright notice.
79
+ 2. Redistributions in binary form must reproduce the above copyright notice.
80
+ """
81
+
82
+ ISC_TEXT = """\
83
+ ISC License
84
+
85
+ Copyright (c) 2026, Example Org
86
+
87
+ Permission to use, copy, modify, and/or distribute this software for any
88
+ purpose with or without fee is hereby granted, provided that the above
89
+ copyright notice and this permission notice appear in all copies.
90
+ """
91
+
92
+ # GPL-3.0 §13 cross-cites the AGPL by full name — the classic trap for
93
+ # full-body substring matchers. Header anchoring must win.
94
+ GPL3_TEXT = """\
95
+ GNU GENERAL PUBLIC LICENSE
96
+ Version 3, 29 June 2007
97
+
98
+ Copyright (C) 2007 Free Software Foundation, Inc.
99
+
100
+ 13. Use with the GNU Affero General Public License.
101
+ Notwithstanding any other provision of this License, you have
102
+ permission to link or combine any covered work with a work licensed
103
+ under version 3 of the GNU Affero General Public License...
104
+ """
105
+
106
+ GPL2_TEXT = """\
107
+ GNU GENERAL PUBLIC LICENSE
108
+ Version 2, June 1991
109
+
110
+ Copyright (C) 1989, 1991 Free Software Foundation, Inc.
111
+ Everyone is permitted to copy and distribute verbatim copies
112
+ of this license document, but changing it is not allowed.
113
+ """
114
+
115
+ # AGPL-3.0 §13 reciprocally cites the plain GPL by name.
116
+ AGPL3_TEXT = """\
117
+ GNU AFFERO GENERAL PUBLIC LICENSE
118
+ Version 3, 19 November 2007
119
+
120
+ 13. Remote Network Interaction; Use with the GNU General Public License.
121
+ Notwithstanding any other provision of this License...
122
+ """
123
+
124
+ LGPL21_TEXT = """\
125
+ GNU LESSER GENERAL PUBLIC LICENSE
126
+ Version 2.1, February 1999
127
+
128
+ Copyright (C) 1991, 1999 Free Software Foundation, Inc.
129
+ """
130
+
131
+ MPL2_TEXT = """\
132
+ Mozilla Public License Version 2.0
133
+ ==================================
134
+
135
+ 1. Definitions
136
+ --------------
137
+ 1.1. "Contributor"
138
+ means each individual or legal entity that creates, contributes to
139
+ the creation of, or owns Covered Software.
140
+ """
141
+
142
+ UNLICENSE_TEXT = """\
143
+ This is free and unencumbered software released into the public domain.
144
+
145
+ Anyone is free to copy, modify, publish, use, compile, sell, or
146
+ distribute this software, either in source code form or as a compiled
147
+ binary, for any purpose, commercial or non-commercial, and by any means.
148
+ """
149
+
150
+
151
+ def _repo_with_license(tmp_path: Path, text: str, filename: str = "LICENSE") -> Path:
152
+ (tmp_path / filename).write_text(text, encoding="utf-8")
153
+ return tmp_path
154
+
155
+
156
+ # ---------------------------------------------------------------------
157
+ # detect_license — SPDX classification from file text
158
+ # ---------------------------------------------------------------------
159
+
160
+
161
+ @pytest.mark.parametrize(
162
+ ("text", "expected"),
163
+ [
164
+ (MIT_TEXT, "MIT"),
165
+ (APACHE_TEXT, "Apache-2.0"),
166
+ (BSD3_TEXT, "BSD-3-Clause"),
167
+ (BSD2_TEXT, "BSD-2-Clause"),
168
+ (ISC_TEXT, "ISC"),
169
+ (GPL3_TEXT, "GPL-3.0"),
170
+ (GPL2_TEXT, "GPL-2.0"),
171
+ (AGPL3_TEXT, "AGPL-3.0"),
172
+ (LGPL21_TEXT, "LGPL-2.1"),
173
+ (MPL2_TEXT, "MPL-2.0"),
174
+ (UNLICENSE_TEXT, "Unlicense"),
175
+ ],
176
+ ids=["mit", "apache2", "bsd3", "bsd2", "isc", "gpl3", "gpl2", "agpl3", "lgpl21", "mpl2", "unlicense"],
177
+ )
178
+ def test_detect_license_spdx_ids(tmp_path: Path, text: str, expected: str):
179
+ info = detect_license(_repo_with_license(tmp_path, text))
180
+ assert info.spdx_id == expected
181
+ assert info.signal == "license_file"
182
+ assert info.source == "LICENSE"
183
+
184
+
185
+ def test_gpl3_not_misread_as_agpl(tmp_path: Path):
186
+ """GPL-3.0 §13 names the AGPL in its body; header anchoring must keep
187
+ this classified as GPL-3.0 (the V9 substring filter would have tripped)."""
188
+ info = detect_license(_repo_with_license(tmp_path, GPL3_TEXT))
189
+ assert info.spdx_id == "GPL-3.0"
190
+
191
+
192
+ def test_apache_notice_without_header_still_apache(tmp_path: Path):
193
+ """The short 'Licensed under the Apache License, Version 2.0' boilerplate
194
+ has no canonical header — the body fallback must catch it, and must not
195
+ fall through to MIT despite shared 'permission' vocabulary."""
196
+ notice = (
197
+ "Copyright 2026 Example Org\n\n"
198
+ "Licensed under the Apache License, Version 2.0 (the \"License\");\n"
199
+ "you may not use this file except in compliance with the License.\n"
200
+ )
201
+ info = detect_license(_repo_with_license(tmp_path, notice))
202
+ assert info.spdx_id == "Apache-2.0"
203
+
204
+
205
+ def test_detect_license_alternate_filenames(tmp_path: Path):
206
+ info = detect_license(_repo_with_license(tmp_path, GPL2_TEXT, filename="COPYING"))
207
+ assert info.spdx_id == "GPL-2.0"
208
+ assert info.source == "COPYING"
209
+ info2 = detect_license(_repo_with_license(tmp_path, MIT_TEXT, filename="LICENSE.md"))
210
+ # LICENSE.md is also present in tmp_path now alongside COPYING; first
211
+ # filename in priority order (LICENSE/LICENSE.txt/LICENSE.md) wins over COPYING.
212
+ assert info2.spdx_id == "MIT"
213
+ assert info2.source == "LICENSE.md"
214
+
215
+
216
+ def test_detect_license_unknown_text(tmp_path: Path):
217
+ info = detect_license(_repo_with_license(tmp_path, "All rights reserved. Ask legal."))
218
+ assert info.spdx_id == "unknown"
219
+
220
+
221
+ def test_detect_license_no_files(tmp_path: Path):
222
+ info = detect_license(tmp_path)
223
+ assert info == LicenseInfo(spdx_id="unknown", signal="none")
224
+
225
+
226
+ def test_classifier_secondary_signal(tmp_path: Path):
227
+ """No LICENSE file, but pyproject carries a trove classifier — the
228
+ classifier signal must win and be recorded as such."""
229
+ (tmp_path / "pyproject.toml").write_text(
230
+ '[project]\nname = "x"\nclassifiers = [\n'
231
+ ' "License :: OSI Approved :: MIT License",\n]\n',
232
+ encoding="utf-8",
233
+ )
234
+ info = detect_license(tmp_path)
235
+ assert info.spdx_id == "MIT"
236
+ assert info.signal == "classifier"
237
+ assert info.source == "pyproject.toml"
238
+
239
+
240
+ def test_classifier_pep639_expression(tmp_path: Path):
241
+ (tmp_path / "pyproject.toml").write_text(
242
+ '[project]\nname = "x"\nlicense = "Apache-2.0"\n', encoding="utf-8"
243
+ )
244
+ info = detect_license(tmp_path)
245
+ assert info.spdx_id == "Apache-2.0"
246
+ assert info.signal == "classifier"
247
+
248
+
249
+ def test_license_file_beats_classifier(tmp_path: Path):
250
+ """When both signals exist and the file is classifiable, the file wins —
251
+ the classifier is secondary by design (it can't tell BSD-2 from BSD-3)."""
252
+ _repo_with_license(tmp_path, GPL3_TEXT)
253
+ (tmp_path / "pyproject.toml").write_text(
254
+ 'classifiers = ["License :: OSI Approved :: MIT License"]\n', encoding="utf-8"
255
+ )
256
+ info = detect_license(tmp_path)
257
+ assert info.spdx_id == "GPL-3.0"
258
+ assert info.signal == "license_file"
259
+
260
+
261
+ def test_unclassifiable_file_falls_back_to_classifier(tmp_path: Path):
262
+ _repo_with_license(tmp_path, "Custom corporate license, see legal dept.")
263
+ (tmp_path / "pyproject.toml").write_text(
264
+ 'classifiers = ["License :: OSI Approved :: ISC License"]\n', encoding="utf-8"
265
+ )
266
+ info = detect_license(tmp_path)
267
+ assert info.spdx_id == "ISC"
268
+ assert info.signal == "classifier"
269
+
270
+
271
+ # ---------------------------------------------------------------------
272
+ # license_tier — tiers, not a boolean (D-13)
273
+ # ---------------------------------------------------------------------
274
+
275
+
276
+ @pytest.mark.parametrize(
277
+ ("spdx", "tier"),
278
+ [
279
+ ("MIT", Tier.REDISTRIBUTABLE),
280
+ ("Apache-2.0", Tier.REDISTRIBUTABLE),
281
+ ("BSD-2-Clause", Tier.REDISTRIBUTABLE),
282
+ ("BSD-3-Clause", Tier.REDISTRIBUTABLE),
283
+ ("ISC", Tier.REDISTRIBUTABLE),
284
+ ("Unlicense", Tier.REDISTRIBUTABLE),
285
+ ("MPL-2.0", Tier.TRAINABLE_ONLY),
286
+ ("LGPL-2.1", Tier.TRAINABLE_ONLY),
287
+ ("LGPL-3.0", Tier.TRAINABLE_ONLY),
288
+ ("GPL-2.0", Tier.EXCLUDED),
289
+ ("GPL-3.0", Tier.EXCLUDED),
290
+ ("AGPL-3.0", Tier.EXCLUDED),
291
+ ("unknown", Tier.EXCLUDED),
292
+ ("WTFPL", Tier.EXCLUDED), # unrecognized id → fail closed
293
+ ],
294
+ )
295
+ def test_license_tier_mapping(spdx: str, tier: Tier):
296
+ assert license_tier(LicenseInfo(spdx_id=spdx, signal="license_file")) is tier
297
+
298
+
299
+ # ---------------------------------------------------------------------
300
+ # Decontamination (V3 / D-5)
301
+ # ---------------------------------------------------------------------
302
+
303
+
304
+ def test_decontamination_list_has_the_canonical_12():
305
+ assert len(DECONTAMINATION_LIST) == 12
306
+ assert "django/django" in DECONTAMINATION_LIST
307
+ assert "sympy/sympy" in DECONTAMINATION_LIST
308
+
309
+
310
+ @pytest.mark.parametrize(
311
+ "repo",
312
+ [
313
+ "django/django", # exact
314
+ "Django/Django", # case
315
+ "https://github.com/django/django", # https URL
316
+ "https://github.com/django/django.git", # URL + .git
317
+ "git@github.com:django/django.git", # ssh URL
318
+ "https://github.com/django/django/", # trailing slash
319
+ ],
320
+ )
321
+ def test_is_eval_contaminated_hits(repo: str):
322
+ assert is_eval_contaminated(repo) is True
323
+
324
+
325
+ @pytest.mark.parametrize(
326
+ "repo",
327
+ [
328
+ "pandas-dev/pandas",
329
+ "https://github.com/torvalds/linux",
330
+ "someuser/django", # fork-org differs: NOT the eval repo
331
+ ],
332
+ )
333
+ def test_is_eval_contaminated_misses(repo: str):
334
+ assert is_eval_contaminated(repo) is False
335
+
336
+
337
+ def test_normalize_repo_forms():
338
+ assert normalize_repo("git@github.com:PSF/Requests.git") == "psf/requests"
339
+ assert normalize_repo("https://github.com/pydata/xarray/tree/main") == "pydata/xarray"
340
+
341
+
342
+ def test_extension_list_from_json(tmp_path: Path):
343
+ """The documented extension mechanism: extra eval repos load from JSON
344
+ and hit through the same normalized matching."""
345
+ extra_path = tmp_path / "extra.json"
346
+ extra_path.write_text(json.dumps(["SWE-Gym/Extra-Repo"]), encoding="utf-8")
347
+ extra = load_decontamination_list(extra_path)
348
+ assert is_eval_contaminated("https://github.com/swe-gym/extra-repo", extra_list=extra)
349
+ assert not is_eval_contaminated("swe-gym/other-repo", extra_list=extra)
350
+
351
+
352
+ def test_extension_list_rejects_non_list(tmp_path: Path):
353
+ bad = tmp_path / "bad.json"
354
+ bad.write_text('{"repo": "a/b"}', encoding="utf-8")
355
+ with pytest.raises(ValueError):
356
+ load_decontamination_list(bad)
357
+
358
+
359
+ # ---------------------------------------------------------------------
360
+ # gate_repo — verdict composition
361
+ # ---------------------------------------------------------------------
362
+
363
+
364
+ def test_gate_admits_permissive_clean_repo(tmp_path: Path):
365
+ v = gate_repo("example/clean", _repo_with_license(tmp_path, MIT_TEXT))
366
+ assert isinstance(v, GateVerdict)
367
+ assert v.admitted is True
368
+ assert v.tier is Tier.REDISTRIBUTABLE
369
+ assert v.contaminated is False
370
+ assert v.reasons == []
371
+
372
+
373
+ def test_gate_contaminated_never_admitted_even_if_permissive(tmp_path: Path):
374
+ """V3 hard rule: an eval repo with an MIT license is STILL rejected —
375
+ decontamination outranks license."""
376
+ v = gate_repo("https://github.com/pallets/flask", _repo_with_license(tmp_path, MIT_TEXT))
377
+ assert v.contaminated is True
378
+ assert v.admitted is False
379
+ assert any("decontamination" in r for r in v.reasons)
380
+ # license detection still ran and is recorded for the manifest
381
+ assert v.license_info.spdx_id == "MIT"
382
+
383
+
384
+ def test_gate_excluded_tier_never_admitted(tmp_path: Path):
385
+ v = gate_repo("example/agpl-repo", _repo_with_license(tmp_path, AGPL3_TEXT))
386
+ assert v.tier is Tier.EXCLUDED
387
+ assert v.admitted is False
388
+ assert any("EXCLUDED" in r for r in v.reasons)
389
+
390
+
391
+ def test_gate_trainable_only_admitted_with_reason(tmp_path: Path):
392
+ """D-13: weak copyleft is admitted for training, but the verdict must
393
+ carry the do-not-redistribute constraint for step 6 to route on."""
394
+ v = gate_repo("example/mpl-repo", _repo_with_license(tmp_path, MPL2_TEXT))
395
+ assert v.tier is Tier.TRAINABLE_ONLY
396
+ assert v.admitted is True
397
+ assert any("TRAINABLE_ONLY" in r for r in v.reasons)
398
+ assert any("redistributed" in r for r in v.reasons)
399
+
400
+
401
+ def test_gate_no_repo_root_fails_closed():
402
+ """No repo_root → license undetectable → unknown → EXCLUDED → rejected
403
+ (V9: the gate must default closed, never open)."""
404
+ v = gate_repo("example/unfetched", None)
405
+ assert v.license_info.spdx_id == "unknown"
406
+ assert v.tier is Tier.EXCLUDED
407
+ assert v.admitted is False
408
+ assert any("failing closed" in r for r in v.reasons)
409
+
410
+
411
+ def test_gate_extra_decontamination_list(tmp_path: Path):
412
+ extra = frozenset({"my-eval/secret-benchmark"})
413
+ v = gate_repo(
414
+ "https://github.com/My-Eval/Secret-Benchmark.git",
415
+ _repo_with_license(tmp_path, MIT_TEXT),
416
+ extra_decontamination=extra,
417
+ )
418
+ assert v.contaminated is True
419
+ assert v.admitted is False
composer_replication/datagen/tests/test_rollout_harness.py ADDED
@@ -0,0 +1,103 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Tests for the rollout harness (deepread finding V2 — the SFT-corpus producer)."""
2
+ from __future__ import annotations
3
+
4
+ from composer_replication.datagen.env import FeatureDeletionEnv
5
+ from composer_replication.datagen.rollout_harness import (
6
+ ScriptedPolicy,
7
+ admit,
8
+ collect_trajectory,
9
+ )
10
+ from composer_replication.datagen.sandbox import FakeSandbox
11
+ from composer_replication.datagen.schema import FeatureDeletionTask
12
+ from composer_replication.datagen.trajectory import CanonicalTrajectory, ToolCall
13
+
14
+
15
+ def _task() -> FeatureDeletionTask:
16
+ return FeatureDeletionTask(
17
+ task_id="t1", repo="org/repo", base_commit="abc",
18
+ broken_image="img:1", test_command="pytest -q",
19
+ fail_to_pass=("t/a.py::t1", "t/a.py::t2"),
20
+ pass_to_pass=("t/a.py::keep",),
21
+ )
22
+
23
+
24
+ def _env(outcomes: dict[str, bool]) -> FeatureDeletionEnv:
25
+ return FeatureDeletionEnv(FakeSandbox(test_outcomes=outcomes))
26
+
27
+
28
+ def test_collect_trajectory_full_pass():
29
+ """Policy 'fixes' the repo via the FakeSandbox set_outcome pseudo-action,
30
+ then submits — grade 1.0, steps record real env transitions."""
31
+ env = _env({"t/a.py::keep": True})
32
+ policy = ScriptedPolicy(actions=[
33
+ ToolCall("set_outcome", {"outcomes": {"t/a.py::t1": True, "t/a.py::t2": True}}),
34
+ "final answer: implemented the feature",
35
+ ])
36
+ traj = collect_trajectory(env, _task(), policy)
37
+ assert isinstance(traj, CanonicalTrajectory)
38
+ assert traj.grade == 1.0
39
+ assert traj.guard_ok is True and traj.hacked is False
40
+ assert len(traj.steps) == 2
41
+ assert isinstance(traj.steps[0].action, ToolCall)
42
+ assert traj.steps[0].result == "ok" # env.step observation recorded
43
+ assert traj.provenance["source"] == "rollout_harness"
44
+
45
+
46
+ def test_collect_trajectory_guard_broken_zeroes_reward():
47
+ env = _env({"t/a.py::keep": False}) # functional guard broken
48
+ policy = ScriptedPolicy(actions=[
49
+ ToolCall("set_outcome", {"outcomes": {"t/a.py::t1": True, "t/a.py::t2": True,
50
+ "t/a.py::keep": False}}),
51
+ "done",
52
+ ])
53
+ traj = collect_trajectory(env, _task(), policy)
54
+ assert traj.grade == 0.0
55
+ assert traj.guard_ok is False
56
+
57
+
58
+ def test_collect_trajectory_near_miss():
59
+ env = _env({"t/a.py::keep": True})
60
+ policy = ScriptedPolicy(actions=[
61
+ ToolCall("set_outcome", {"outcomes": {"t/a.py::t1": True}}), # 1 of 2
62
+ "done",
63
+ ])
64
+ traj = collect_trajectory(env, _task(), policy)
65
+ assert traj.grade == 0.5
66
+ assert traj.guard_ok is True
67
+
68
+
69
+ def test_collect_trajectory_max_turns_grades_anyway():
70
+ env = _env({"t/a.py::keep": True})
71
+ looping = ScriptedPolicy(actions=[ToolCall("bash", {"command": "ls"})] * 50)
72
+ traj = collect_trajectory(env, _task(), looping, max_turns=3)
73
+ assert traj.grade is not None # graded despite never submitting
74
+
75
+
76
+ # ---------------------------------------------------------------------
77
+ # Admission routing (typed train-on-all, final report §4)
78
+ # ---------------------------------------------------------------------
79
+
80
+
81
+ def _t(grade, guard_ok=True, hacked=False) -> CanonicalTrajectory:
82
+ return CanonicalTrajectory(task_id="x", grade=grade, guard_ok=guard_ok, hacked=hacked)
83
+
84
+
85
+ def test_admit_routes_clean_pass_to_sft():
86
+ v = admit(_t(1.0))
87
+ assert v.sft_admitted and not v.dpo_candidate and not v.rejected
88
+
89
+
90
+ def test_admit_routes_near_miss_to_dpo():
91
+ v = admit(_t(0.5))
92
+ assert v.dpo_candidate and not v.sft_admitted and not v.rejected
93
+
94
+
95
+ def test_admit_rejects_hacked_even_at_full_grade():
96
+ v = admit(_t(1.0, hacked=True))
97
+ assert v.rejected and "hack monitor flagged" in v.reasons
98
+
99
+
100
+ def test_admit_rejects_guard_broken_and_ungraded():
101
+ assert admit(_t(1.0, guard_ok=False)).rejected
102
+ assert admit(_t(None)).rejected
103
+ assert admit(_t(0.0)).rejected
composer_replication/datagen/tests/test_swesmith_adapter.py ADDED
@@ -0,0 +1,165 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Tests for the SWE-smith adapter (deepread finding V4 — buy-vs-build).
2
+
3
+ The load-bearing coverage: the PATCH-SEMANTICS INVERSION (SWE-smith's patch
4
+ introduces the bug; golden_diff must be its reverse) and the mechanical
5
+ reverse_unified_diff round-trip.
6
+ """
7
+ from __future__ import annotations
8
+
9
+ import json
10
+
11
+ import pytest
12
+
13
+ from composer_replication.datagen.schema import FeatureDeletionTask
14
+ from composer_replication.datagen.substrates import SweBenchAdapter
15
+ from composer_replication.datagen.swesmith_adapter import (
16
+ UNREVERSED_MARKER,
17
+ SwesmithAdapter,
18
+ load_swesmith_instances,
19
+ parse_strategy,
20
+ reverse_unified_diff,
21
+ )
22
+
23
+ BUG_PATCH = """\
24
+ diff --git a/pkg/mod.py b/pkg/mod.py
25
+ index 1111111..2222222 100644
26
+ --- a/pkg/mod.py
27
+ +++ b/pkg/mod.py
28
+ @@ -1,4 +1,3 @@
29
+ def add(a, b):
30
+ - return a + b
31
+ + return a - b
32
+ # trailing context
33
+ """
34
+
35
+
36
+ def _instance(**over) -> dict:
37
+ base = {
38
+ "instance_id": "getmoto__moto.abc1234.lm_modify__1a2b",
39
+ "repo": "getmoto/moto",
40
+ "base_commit": "abc1234",
41
+ "patch": BUG_PATCH,
42
+ "FAIL_TO_PASS": json.dumps(["tests/test_mod.py::test_add"]),
43
+ "PASS_TO_PASS": json.dumps(["tests/test_mod.py::test_other"]),
44
+ "image_name": "swesmith.x86_64.getmoto__moto:latest",
45
+ }
46
+ base.update(over)
47
+ return base
48
+
49
+
50
+ # ---------------------------------------------------------------------
51
+ # Strategy parsing
52
+ # ---------------------------------------------------------------------
53
+
54
+
55
+ @pytest.mark.parametrize("iid,expected", [
56
+ ("r__x.abc.lm_modify__1", "lm_modify"),
57
+ ("r__x.abc.lm_rewrite__1", "lm_rewrite"),
58
+ ("r__x.abc.func_pm_ctrl_invert_if__1", "procedural"),
59
+ ("r__x.abc.func_basic__1", "procedural"),
60
+ ("r__x.abc.combine_file__1", "combine"),
61
+ ("r__x.abc.combine_module__2", "combine"),
62
+ ("r__x.abc.pr_1234", "pr_mirror"),
63
+ ("r__x.abc.mystery__1", "unknown"),
64
+ ])
65
+ def test_parse_strategy(iid, expected):
66
+ assert parse_strategy(iid) == expected
67
+
68
+
69
+ # ---------------------------------------------------------------------
70
+ # reverse_unified_diff
71
+ # ---------------------------------------------------------------------
72
+
73
+
74
+ def test_reverse_swaps_adds_and_removes():
75
+ rev = reverse_unified_diff(BUG_PATCH)
76
+ assert rev is not None
77
+ # The bug ADDED "return a - b"; the reverse must REMOVE it.
78
+ assert "- return a - b" in rev
79
+ assert "+ return a + b" in rev
80
+ # Hunk header ranges swapped: -1,4 +1,3 → -1,3 +1,4
81
+ assert "@@ -1,3 +1,4 @@" in rev
82
+ # Context lines untouched.
83
+ assert " def add(a, b):" in rev
84
+ assert " # trailing context" in rev
85
+
86
+
87
+ def test_reverse_round_trip_is_identity_on_body():
88
+ rev = reverse_unified_diff(BUG_PATCH)
89
+ rev2 = reverse_unified_diff(rev)
90
+ # Round trip restores hunks and +/- bodies (headers may normalize).
91
+ orig_body = [ln for ln in BUG_PATCH.splitlines() if ln[:1] in "+-@" and not ln.startswith(("+++", "---"))]
92
+ rt_body = [ln for ln in rev2.splitlines() if ln[:1] in "+-@" and not ln.startswith(("+++", "---"))]
93
+ assert orig_body == rt_body
94
+
95
+
96
+ def test_reverse_refuses_renames_and_binary():
97
+ assert reverse_unified_diff("diff --git a/x b/y\nrename from x\nrename to y\n") is None
98
+ assert reverse_unified_diff("diff --git a/x b/x\nGIT binary patch\nliteral 5\n") is None
99
+ assert reverse_unified_diff("") is None
100
+ assert reverse_unified_diff("no hunks here") is None
101
+
102
+
103
+ # ---------------------------------------------------------------------
104
+ # Adapter
105
+ # ---------------------------------------------------------------------
106
+
107
+
108
+ def test_to_task_golden_diff_is_the_fix_not_the_bug():
109
+ """THE semantic inversion: golden_diff must restore the feature."""
110
+ task, meta = SwesmithAdapter().to_task_with_meta(_instance())
111
+ assert isinstance(task, FeatureDeletionTask)
112
+ assert meta.diff_reversed is True
113
+ assert meta.strategy == "lm_modify"
114
+ # The FIX restores `a + b` (adds it back) and removes the bug.
115
+ assert "+ return a + b" in task.golden_diff
116
+ assert "- return a - b" in task.golden_diff
117
+ assert UNREVERSED_MARKER not in task.golden_diff
118
+
119
+
120
+ def test_to_task_unreversible_patch_gets_marker():
121
+ inst = _instance(patch="diff --git a/x b/y\nrename from x\nrename to y\n@@ -1 +1 @@\n-a\n+b\n")
122
+ task, meta = SwesmithAdapter().to_task_with_meta(inst)
123
+ assert meta.diff_reversed is False
124
+ assert task.golden_diff.startswith(UNREVERSED_MARKER)
125
+
126
+
127
+ def test_image_resolution_prefers_instance_field_then_convention():
128
+ a = SwesmithAdapter()
129
+ assert a.image_for(_instance()) == "swesmith.x86_64.getmoto__moto:latest"
130
+ assert a.image_for(_instance(image_name=None, docker_image="custom:tag")) == "custom:tag"
131
+ inst = _instance(image_name=None)
132
+ inst.pop("docker_image", None)
133
+ assert a.image_for(inst) == "swesmith.x86_64.getmoto__moto:latest"
134
+
135
+
136
+ def test_f2p_p2p_tuple_handling_matches_swebench_semantics():
137
+ task = SwesmithAdapter().to_task(_instance(
138
+ FAIL_TO_PASS=["t/a.py::t1", "t/a.py::t2"], # real list, not JSON string
139
+ PASS_TO_PASS=json.dumps([]),
140
+ ))
141
+ assert task.fail_to_pass == ("t/a.py::t1", "t/a.py::t2")
142
+ assert task.pass_to_pass == ()
143
+
144
+
145
+ def test_difficulty_priors_by_strategy():
146
+ pr = SwesmithAdapter().to_task(_instance(instance_id="r__x.abc.pr_99"))
147
+ proc = SwesmithAdapter().to_task(_instance(instance_id="r__x.abc.func_pm_remove_loop__1"))
148
+ assert pr.difficulty_prior < proc.difficulty_prior # PR Mirror harder prior
149
+
150
+
151
+ def test_redistributable_filter_interplay():
152
+ """repo_gate owns repo-level licensing, but the per-instance filter from
153
+ SweBenchAdapter still composes when a license field IS present."""
154
+ task = SwesmithAdapter().to_task(_instance(license_name="GPL-3.0"))
155
+ assert SweBenchAdapter.is_redistributable(task) is False
156
+ task2 = SwesmithAdapter().to_task(_instance(license_name="MIT"))
157
+ assert SweBenchAdapter.is_redistributable(task2) is True
158
+
159
+
160
+ def test_load_local_jsonl(tmp_path):
161
+ p = tmp_path / "fixtures.jsonl"
162
+ p.write_text("\n".join(json.dumps(_instance(instance_id=f"r__x.abc.pr_{i}")) for i in range(5)))
163
+ rows = load_swesmith_instances(str(p), limit=3)
164
+ assert len(rows) == 3
165
+ assert rows[0]["instance_id"] == "r__x.abc.pr_0"
composer_replication/datagen/tests/test_trajectory.py ADDED
@@ -0,0 +1,127 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Tests for the canonical trajectory IR (deepread findings V2/D-11/D-8).
2
+
3
+ The load-bearing test is the SENTINEL leak guard: `to_policy_row` must never
4
+ emit golden_diff/deleted_symbols, even though the task dataclass carries them.
5
+ """
6
+ from __future__ import annotations
7
+
8
+ import json
9
+
10
+ from composer_replication.datagen.schema import FeatureDeletionTask
11
+ from composer_replication.datagen.trajectory import (
12
+ CanonicalTrajectory,
13
+ ToolCall,
14
+ TrajectoryStep,
15
+ from_trace_states,
16
+ to_policy_row,
17
+ to_sft_messages,
18
+ )
19
+ from composer_replication.teacher_replay import TraceState
20
+
21
+
22
+ def _task(**over) -> FeatureDeletionTask:
23
+ base = dict(
24
+ task_id="t1", repo="org/repo", base_commit="abc",
25
+ broken_image="img:1", test_command="pytest -q",
26
+ fail_to_pass=("t/a.py::t1",), pass_to_pass=("t/a.py::t2",),
27
+ golden_diff="SENTINEL_NEVER_LEAK",
28
+ deleted_symbols=("secret_fn",),
29
+ )
30
+ base.update(over)
31
+ return FeatureDeletionTask(**base)
32
+
33
+
34
+ def _traj() -> CanonicalTrajectory:
35
+ return CanonicalTrajectory(
36
+ task_id="t1",
37
+ steps=[
38
+ TrajectoryStep(observation="repo is broken", action=ToolCall("bash", {"command": "pytest"}),
39
+ result="2 failed", tool_error=False),
40
+ TrajectoryStep(observation="2 failed", action="here is my final patch",
41
+ result="graded", tool_error=False),
42
+ ],
43
+ grade=1.0, guard_ok=True, hacked=False,
44
+ provenance={"source": "test"},
45
+ )
46
+
47
+
48
+ # ---------------------------------------------------------------------
49
+ # ToolCall canonical form — the v1 divergence algebra
50
+ # ---------------------------------------------------------------------
51
+
52
+
53
+ def test_canonical_form_is_order_insensitive_on_args():
54
+ a = ToolCall("edit", {"path": "x.py", "content": "y"})
55
+ b = ToolCall("edit", {"content": "y", "path": "x.py"})
56
+ assert a.canonical_form() == b.canonical_form()
57
+
58
+
59
+ def test_canonical_form_distinguishes_name_and_args():
60
+ assert ToolCall("bash", {"command": "ls"}).canonical_form() != \
61
+ ToolCall("bash", {"command": "ls -la"}).canonical_form()
62
+ assert ToolCall("read", {"f": "x"}).canonical_form() != \
63
+ ToolCall("write", {"f": "x"}).canonical_form()
64
+
65
+
66
+ # ---------------------------------------------------------------------
67
+ # THE leak guard (finding D-8)
68
+ # ---------------------------------------------------------------------
69
+
70
+
71
+ def test_policy_row_never_contains_golden_diff_or_deleted_symbols():
72
+ row = to_policy_row(_traj(), _task())
73
+ blob = json.dumps(row)
74
+ assert "SENTINEL_NEVER_LEAK" not in blob
75
+ assert "secret_fn" not in blob
76
+ assert "golden_diff" not in blob
77
+ assert "deleted_symbols" not in blob
78
+ # And the row still carries what the policy MAY see.
79
+ assert row["repo"] == "org/repo"
80
+ assert row["fail_to_pass"] == ["t/a.py::t1"]
81
+ assert row["grade"] == 1.0
82
+
83
+
84
+ # ---------------------------------------------------------------------
85
+ # IR ↔ SFT messages
86
+ # ---------------------------------------------------------------------
87
+
88
+
89
+ def test_to_sft_messages_alternates_roles():
90
+ msgs = to_sft_messages(_traj())
91
+ assert msgs[0] == {"role": "user", "content": "repo is broken"}
92
+ assert msgs[1]["role"] == "assistant"
93
+ assert "[TOOL_USE] name=bash" in msgs[1]["content"]
94
+ assert msgs[2] == {"role": "user", "content": "2 failed"}
95
+
96
+
97
+ # ---------------------------------------------------------------------
98
+ # Claude Code → IR adapter
99
+ # ---------------------------------------------------------------------
100
+
101
+
102
+ def test_from_trace_states_parses_single_tool_use_and_error_flag():
103
+ states: list[TraceState] = [
104
+ {
105
+ "state_id": "sess1::0000",
106
+ "messages": [
107
+ {"role": "system", "content": "sys"},
108
+ {"role": "user", "content": "[TOOL_RESULT (ERROR)] (id=x)\nboom",
109
+ "tool_error": True},
110
+ ],
111
+ "student_action": '[TOOL_USE] name=Bash input={"command":"ls"}',
112
+ },
113
+ {
114
+ "state_id": "sess1::0001",
115
+ "messages": [{"role": "user", "content": "plain prompt"}],
116
+ "student_action": "I think the fix is...\n\n[TOOL_USE] name=Edit input={\"p\":1}\n\n[TOOL_USE] name=Bash input={\"c\":2}",
117
+ },
118
+ ]
119
+ traj = from_trace_states(states)
120
+ assert traj.task_id == "sess1"
121
+ assert traj.grade is None # ungraded — Claude Code traces have no oracle
122
+ s0, s1 = traj.steps
123
+ assert isinstance(s0.action, ToolCall) and s0.action.name == "Bash"
124
+ assert s0.tool_error is True
125
+ # Multi-tool turn stays as the raw string (honest, not guessed).
126
+ assert isinstance(s1.action, str)
127
+ assert s1.tool_error is False
composer_replication/datagen/trajectory.py ADDED
@@ -0,0 +1,203 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """trajectory.py — the canonical trajectory IR (deepread findings V2/D-11/D-8).
2
+
3
+ THE GAP THIS CLOSES: the repo had 3 (heading to 5) incompatible trajectory
4
+ shapes — Claude Code TraceState text serialization, Bedrock `.jsonl.out` rows,
5
+ the planned tree/rollout/OpenHands shapes — with no shared schema, and the
6
+ divergence gate rested on a whitespace-collapse string normalizer
7
+ (`teacher_replay._normalize_action`, self-admitted skeleton). This module is the
8
+ single intermediate representation every producer adapts INTO and every corpus
9
+ writer reads FROM.
10
+
11
+ SECURITY INVARIANT (finding D-8): `FeatureDeletionTask.golden_diff` uses
12
+ `repr=False`, but `dataclasses.asdict()` and naive JSON serialization still
13
+ include it. `to_policy_row()` is the ONE serializer allowed to produce
14
+ policy-visible rows, and its output is unit-tested to never contain
15
+ `golden_diff` or `deleted_symbols`.
16
+ """
17
+ from __future__ import annotations
18
+
19
+ import json
20
+ import re
21
+ from dataclasses import dataclass, field
22
+ from typing import Any
23
+
24
+ from composer_replication.datagen.schema import FeatureDeletionTask
25
+ from composer_replication.teacher_replay import TraceState
26
+
27
+ #: Bump when the IR shape changes; carried on every trajectory + corpus row.
28
+ SCHEMA_VERSION = "1"
29
+
30
+
31
+ @dataclass(frozen=True)
32
+ class ToolCall:
33
+ """One structured tool invocation — the unit of the action algebra.
34
+
35
+ `canonical_form()` is the v1 divergence-gate algebra (finding D-3): two
36
+ actions are "the same" iff their canonical forms match (tool name + sorted,
37
+ JSON-normalized args). This replaces the whitespace-collapse stub that made
38
+ the divergence gate fire on noise. v2 will need arg-level normalization
39
+ (path equivalence, whitespace-insensitive code args); keep that evolution
40
+ HERE so every consumer inherits it.
41
+ """
42
+
43
+ name: str
44
+ args: dict = field(default_factory=dict)
45
+
46
+ def canonical_form(self) -> str:
47
+ try:
48
+ args_json = json.dumps(self.args, sort_keys=True, separators=(",", ":"))
49
+ except (TypeError, ValueError):
50
+ args_json = str(self.args)
51
+ return f"{self.name}:{args_json}"
52
+
53
+
54
+ @dataclass
55
+ class TrajectoryStep:
56
+ """One agent turn: what it saw, what it did, what came back."""
57
+
58
+ observation: str
59
+ action: ToolCall | str # str = plain text / final message
60
+ result: str | None = None # tool output observed AFTER the action
61
+ tool_error: bool = False
62
+
63
+
64
+ @dataclass
65
+ class CanonicalTrajectory:
66
+ """The IR: an episode (or trace) as a typed step list + outcome + provenance."""
67
+
68
+ task_id: str
69
+ steps: list[TrajectoryStep] = field(default_factory=list)
70
+ grade: float | None = None # _grade() pass-fraction; None = ungraded trace
71
+ guard_ok: bool = True
72
+ hacked: bool = False
73
+ provenance: dict = field(default_factory=dict) # source, policy id, cost_usd, run_id
74
+ schema_version: str = SCHEMA_VERSION
75
+
76
+
77
+ # ---------------------------------------------------------------------
78
+ # Producers → IR
79
+ # ---------------------------------------------------------------------
80
+
81
+ # ClaudeCodeIngester serializes tool calls as "[TOOL_USE] name=<n> input=<json>".
82
+ _TOOL_USE_RE = re.compile(r"^\[TOOL_USE\] name=(?P<name>\S+) input=(?P<input>\{.*\})$")
83
+
84
+
85
+ def _parse_action(student_action: str) -> ToolCall | str:
86
+ """Parse a TraceState student_action back into a ToolCall where possible.
87
+
88
+ Claude Code assistant turns serialize as newline-joined blocks; if exactly
89
+ one [TOOL_USE] block is present we recover the structured call (the common
90
+ case ADR-002 chose one-node-per-turn for). Multi-tool turns and pure-text /
91
+ thinking turns stay as the raw string — honest about what we can't
92
+ structure rather than guessing.
93
+ """
94
+ blocks = [b for b in student_action.split("\n\n") if b.strip()]
95
+ tool_blocks = [b for b in blocks if b.startswith("[TOOL_USE]")]
96
+ if len(tool_blocks) == 1:
97
+ m = _TOOL_USE_RE.match(tool_blocks[0].strip())
98
+ if m:
99
+ try:
100
+ return ToolCall(name=m.group("name"), args=json.loads(m.group("input")))
101
+ except (json.JSONDecodeError, ValueError):
102
+ pass
103
+ return student_action
104
+
105
+
106
+ def from_trace_states(
107
+ states: list[TraceState],
108
+ *,
109
+ task_id: str = "",
110
+ provenance: dict | None = None,
111
+ ) -> CanonicalTrajectory:
112
+ """Adapt a Claude Code trace (TraceState list) into the IR.
113
+
114
+ HONEST CAPABILITY NOTE (finding D-1): these traces carry no executable
115
+ environment — no broken_image, no FAIL_TO_PASS — so the resulting
116
+ trajectory is UNGRADED (`grade=None`) and is admissible only for flat
117
+ Channel-3 / SFT-style uses, never as a tree seed. Env-grounded rollouts
118
+ (rollout_harness.collect_trajectory) are the graded producers.
119
+ """
120
+ steps: list[TrajectoryStep] = []
121
+ for s in states:
122
+ # The observation for step t is the last user message before the turn.
123
+ obs = ""
124
+ tool_error = False
125
+ for msg in reversed(s["messages"]):
126
+ if msg.get("role") == "user":
127
+ obs = str(msg.get("content", ""))
128
+ tool_error = bool(msg.get("tool_error", False))
129
+ break
130
+ steps.append(TrajectoryStep(
131
+ observation=obs,
132
+ action=_parse_action(s["student_action"]),
133
+ result=None,
134
+ tool_error=tool_error,
135
+ ))
136
+ prov = {"source": "claude_code", **(provenance or {})}
137
+ return CanonicalTrajectory(task_id=task_id or (states[0]["state_id"].split("::")[0] if states else ""),
138
+ steps=steps, grade=None, provenance=prov)
139
+
140
+
141
+ # ---------------------------------------------------------------------
142
+ # IR → consumers
143
+ # ---------------------------------------------------------------------
144
+
145
+
146
+ def _action_text(action: ToolCall | str) -> str:
147
+ if isinstance(action, ToolCall):
148
+ return f"[TOOL_USE] name={action.name} input=" + json.dumps(
149
+ action.args, separators=(",", ":")
150
+ )
151
+ return action
152
+
153
+
154
+ def to_sft_messages(traj: CanonicalTrajectory) -> list[dict]:
155
+ """IR → OpenAI-style messages for SFT (obs→user, action→assistant)."""
156
+ messages: list[dict] = []
157
+ for step in traj.steps:
158
+ if step.observation:
159
+ messages.append({"role": "user", "content": step.observation})
160
+ messages.append({"role": "assistant", "content": _action_text(step.action)})
161
+ if step.result:
162
+ messages.append({"role": "user", "content": step.result})
163
+ return messages
164
+
165
+
166
+ #: Task fields the POLICY may see. Everything else (golden_diff,
167
+ #: deleted_symbols) is construction-side and must never reach a corpus row.
168
+ _POLICY_VISIBLE_TASK_FIELDS = (
169
+ "task_id", "repo", "base_commit", "test_command",
170
+ "fail_to_pass", "pass_to_pass", "granularity", "difficulty_prior",
171
+ )
172
+
173
+
174
+ def to_policy_row(traj: CanonicalTrajectory, task: FeatureDeletionTask) -> dict:
175
+ """THE policy-visible corpus serializer (finding D-8 — the leak guard).
176
+
177
+ Builds the row field-by-field from an allowlist; never `asdict(task)`,
178
+ which would include `golden_diff` despite its `repr=False`. Unit-tested
179
+ with a sentinel to prove the absence.
180
+ """
181
+ row: dict[str, Any] = {
182
+ "schema_version": traj.schema_version,
183
+ "messages": to_sft_messages(traj),
184
+ "grade": traj.grade,
185
+ "guard_ok": traj.guard_ok,
186
+ "hacked": traj.hacked,
187
+ "provenance": dict(traj.provenance),
188
+ }
189
+ for f in _POLICY_VISIBLE_TASK_FIELDS:
190
+ v = getattr(task, f)
191
+ row[f] = list(v) if isinstance(v, tuple) else v
192
+ return row
193
+
194
+
195
+ __all__ = [
196
+ "SCHEMA_VERSION",
197
+ "ToolCall",
198
+ "TrajectoryStep",
199
+ "CanonicalTrajectory",
200
+ "from_trace_states",
201
+ "to_sft_messages",
202
+ "to_policy_row",
203
+ ]
composer_replication/diloco/__init__.py CHANGED
@@ -4,9 +4,12 @@ Wraps `torchft.local_sgd.DiLoCo` with the framework's conventions:
4
  - Sign convention is documented LOUDLY here once and tested via Spike 008.
5
  - The wrapper exposes the same constructor shape as torchft's DiLoCo so a
6
  future swap-in of the upstream class is a one-line change.
7
- - Vanilla DiLoCo (Douillard et al. 2023) = `fragment_sync_delay=0`, single
8
- fragment. Streaming DiLoCo (Liu et al. 2025) = non-zero delay, multiple
9
- fragments. Spike 008 uses vanilla; Streaming is configured by the same API.
 
 
 
10
 
11
  Reference: `docs/adrs/ADR-003-diloco-impl.md`.
12
 
 
4
  - Sign convention is documented LOUDLY here once and tested via Spike 008.
5
  - The wrapper exposes the same constructor shape as torchft's DiLoCo so a
6
  future swap-in of the upstream class is a one-line change.
7
+ - Vanilla DiLoCo (Douillard et al. 2023, arXiv:2311.08105) =
8
+ `fragment_sync_delay=0`, single fragment. Streaming DiLoCo (Douillard et
9
+ al., arXiv:2501.18512 "Streaming DiLoCo with overlapping communication";
10
+ the separate Eager-Updates work is Kale et al., arXiv:2502.12996 — citation
11
+ corrected per deepread finding V7) = non-zero delay, multiple fragments.
12
+ Spike 008 uses vanilla; Streaming is configured by the same API.
13
 
14
  Reference: `docs/adrs/ADR-003-diloco-impl.md`.
15
 
composer_replication/opsd.py CHANGED
@@ -10,17 +10,23 @@ Mathematical reference:
10
  - OPSD paper: Zhao et al., "Self-Distilled Reasoner: On-Policy Self-Distillation
11
  for LLMs", arXiv:2601.18734.
12
  - SDPO paper: Hübotter et al., "Reinforcement Learning via Self-Distillation",
13
- arXiv:2601.20802 (formalizes the same loss as Composer 2.5's "Targeted RL with
14
- Textual Feedback").
 
 
 
 
 
15
 
16
  The loss computes JSD/KL divergence between a teacher distribution (model
17
  conditioned on privileged information / a hint) and a student distribution
18
  (model on the original context). Both come from the SAME model — the teacher
19
  is just "the model with hint inserted into context."
20
 
21
- Composer 2.5 uses this with the privileged information being a "hint" inserted
22
- at the error-turn site. We use the same loss; the data collator constructs
23
- ctx_teacher = ctx_student + hint_at_error_turn for us.
 
24
  """
25
 
26
  from __future__ import annotations
 
10
  - OPSD paper: Zhao et al., "Self-Distilled Reasoner: On-Policy Self-Distillation
11
  for LLMs", arXiv:2601.18734.
12
  - SDPO paper: Hübotter et al., "Reinforcement Learning via Self-Distillation",
13
+ arXiv:2601.20802. PROVENANCE (corrected per deepread finding V1): Cursor's
14
+ blog cites SDPO/OPSD only as *background* ("For more background on this
15
+ approach see…"), NOT as its mechanism. Published SDPO distills over the FULL
16
+ rollout with feedback in the prefix and an EMA-regularized teacher; this
17
+ repo's channel is a turn-localized hint-splice with a live (stop-grad,
18
+ non-EMA) teacher — a third, blog-inspired design, neither verbatim SDPO nor
19
+ confirmed-Composer. The kernel below matches OPSD's generalized JSD math.
20
 
21
  The loss computes JSD/KL divergence between a teacher distribution (model
22
  conditioned on privileged information / a hint) and a student distribution
23
  (model on the original context). Both come from the SAME model — the teacher
24
  is just "the model with hint inserted into context."
25
 
26
+ Composer 2.5's blog describes inserting a "hint" at the error-turn site and
27
+ distilling the student toward the hint-conditioned distribution "for that turn
28
+ only". The data collator constructs ctx_teacher = ctx_student +
29
+ hint_at_error_turn for us.
30
  """
31
 
32
  from __future__ import annotations
composer_replication/pipeline/__init__.py ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """composer_replication.pipeline — the Stage-0 dataset-pipeline contract + driver.
2
+
3
+ THE single reconciled dataset contract (supersedes the two divergent layouts in
4
+ research/design-F1 and design-F2 — deepread finding V8/D-7), the pragmatic
5
+ near-duplicate detector, and the local stage-driver that turns
6
+ (tasks, env, policy) into a carded, deduped, holdout-split corpus.
7
+ """
8
+ from composer_replication.pipeline.build_corpus import build_corpus
9
+ from composer_replication.pipeline.dedup import (
10
+ dedup,
11
+ find_near_duplicates,
12
+ jaccard_estimate,
13
+ minhash_signature,
14
+ )
15
+ from composer_replication.pipeline.s3_contract import (
16
+ RunLayout,
17
+ RunManifest,
18
+ write_dataset_card,
19
+ write_dpo_rows,
20
+ write_sft_rows,
21
+ write_tasks,
22
+ write_tasks_full,
23
+ )
24
+
25
+ __all__ = [
26
+ "RunLayout",
27
+ "RunManifest",
28
+ "build_corpus",
29
+ "dedup",
30
+ "find_near_duplicates",
31
+ "jaccard_estimate",
32
+ "minhash_signature",
33
+ "write_dataset_card",
34
+ "write_dpo_rows",
35
+ "write_sft_rows",
36
+ "write_tasks",
37
+ "write_tasks_full",
38
+ ]
composer_replication/pipeline/build_corpus.py ADDED
@@ -0,0 +1,137 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """build_corpus.py — the local Stage-0 stage-driver (architecture step 6-7).
2
+
3
+ One function wires the whole local pipeline: holdout-split the task pool
4
+ (holdout tasks are NEVER rolled out — they are the eval anchor), roll out a
5
+ policy over each train task, admit + type + route trajectories
6
+ (sft / dpo-candidate / quarantine), dedup the SFT rows (within-run AND against
7
+ a prior generation's signatures), and write everything through the single
8
+ s3_contract layout with a manifest + dataset card.
9
+
10
+ Deliberately LOCAL-first (finding D-9): the five-service AWS orchestration is
11
+ Stage 4; this driver must produce one real corpus end-to-end on a laptop with a
12
+ FakeSandbox before anything is distributed. Write-once per layout
13
+ (finding D-21): refuses to run if the manifest already exists.
14
+ """
15
+ from __future__ import annotations
16
+
17
+ from typing import Callable, Sequence
18
+
19
+ from composer_replication.datagen.env import FeatureDeletionEnv
20
+ from composer_replication.datagen.rollout_harness import (
21
+ RolloutPolicy,
22
+ admit,
23
+ collect_trajectory,
24
+ )
25
+ from composer_replication.datagen.schema import FeatureDeletionTask
26
+ from composer_replication.datagen.trajectory import to_policy_row
27
+ from composer_replication.pipeline import s3_contract
28
+ from composer_replication.pipeline.dedup import dedup
29
+ from composer_replication.pipeline.s3_contract import RunLayout, RunManifest
30
+ from composer_replication.safety.holdout import HeldoutSplit
31
+
32
+
33
+ def build_corpus(
34
+ source_tasks: Sequence[FeatureDeletionTask],
35
+ env_factory: Callable[[], FeatureDeletionEnv],
36
+ policy_factory: Callable[[], RolloutPolicy],
37
+ layout: RunLayout,
38
+ manifest: RunManifest,
39
+ *,
40
+ holdout_frac: float = 0.2,
41
+ holdout_seed: int = 0,
42
+ max_tasks: int | None = None,
43
+ cost_per_rollout_usd: float = 0.0,
44
+ prior_signatures: Sequence[Sequence[int]] | None = None,
45
+ dedup_threshold: float = 0.85,
46
+ ) -> RunManifest:
47
+ """Run the Stage-0 pipeline over `source_tasks`; returns the final manifest.
48
+
49
+ Args:
50
+ source_tasks: gate_repo-admitted FeatureDeletionTasks (the caller runs
51
+ `datagen.repo_gate.gate_repo` BEFORE this — the driver assumes the
52
+ license/decontamination gates already passed).
53
+ env_factory: fresh `FeatureDeletionEnv` per rollout (a sandbox is
54
+ stateful; sharing one across episodes leaks trajectory state).
55
+ policy_factory: fresh policy per rollout (ScriptedPolicy is stateful).
56
+ manifest: a `RunManifest` with run_id/created_at/budget set by the
57
+ caller (created_at is caller-passed for reproducibility).
58
+ cost_per_rollout_usd: accounting hook — API policies should report
59
+ real cost; the driver enforces `manifest.budget_usd` with it.
60
+ prior_signatures: previous generation's MinHash signatures
61
+ (cross-generation dedup, finding D-12).
62
+
63
+ Raises:
64
+ FileExistsError: if the layout already has a manifest (write-once).
65
+ """
66
+ if s3_contract.manifest_exists(layout):
67
+ raise FileExistsError(
68
+ f"Run layout already has a manifest at {layout.manifest_path} — "
69
+ "runs are write-once per (root, run_id); mint a new run_id "
70
+ "(finding D-21 idempotency)."
71
+ )
72
+
73
+ # 1. Holdout split FIRST — held-out tasks are never rolled out, so no
74
+ # training signal can derive from them (the HeldoutSplit discipline).
75
+ split = HeldoutSplit.split(source_tasks, holdout_frac=holdout_frac,
76
+ seed=holdout_seed, check_content=True)
77
+ by_id = {t.task_id: t for t in source_tasks}
78
+ holdout_tasks = [by_id[i] for i in sorted(split.holdout_ids)]
79
+ train_tasks = [by_id[i] for i in sorted(split.train_ids)]
80
+ if max_tasks is not None:
81
+ train_tasks = train_tasks[:max_tasks]
82
+
83
+ # 2. Rollouts + admission routing.
84
+ sft_rows: list[dict] = []
85
+ dpo_rows: list[dict] = []
86
+ quarantine_rows: list[dict] = []
87
+ traj_rows: list[dict] = []
88
+ partial = False
89
+ for task in train_tasks:
90
+ if manifest.over_budget:
91
+ partial = True
92
+ break
93
+ traj = collect_trajectory(env_factory(), task, policy_factory(),
94
+ provenance={"run_id": manifest.run_id})
95
+ manifest.spend(cost_per_rollout_usd)
96
+ verdict = admit(traj)
97
+ row = to_policy_row(traj, task)
98
+ traj_rows.append({**row, "admission": list(verdict.reasons)})
99
+ if verdict.sft_admitted:
100
+ sft_rows.append(row)
101
+ elif verdict.dpo_candidate:
102
+ dpo_rows.append(row)
103
+ else:
104
+ quarantine_rows.append({**row, "reasons": list(verdict.reasons)})
105
+
106
+ # 3. Dedup the SFT corpus (within-run + cross-generation).
107
+ def _key(r: dict) -> str:
108
+ return " ".join(m.get("content", "") for m in r.get("messages", []))
109
+
110
+ sft_rows, dedup_stats = dedup(sft_rows, _key, dedup_threshold,
111
+ prior_signatures=prior_signatures)
112
+
113
+ # 4. Write everything through the contract.
114
+ s3_contract.write_tasks(layout, train_tasks)
115
+ s3_contract.write_tasks_full(layout, train_tasks)
116
+ s3_contract.write_holdout(layout, holdout_tasks)
117
+ s3_contract.write_trajectories(layout, traj_rows)
118
+ s3_contract.write_sft_rows(layout, sft_rows)
119
+ s3_contract.write_dpo_rows(layout, dpo_rows)
120
+ s3_contract.write_quarantine(layout, quarantine_rows)
121
+
122
+ manifest.counts = {
123
+ "tasks_train": len(train_tasks),
124
+ "tasks_holdout": len(holdout_tasks),
125
+ "rollouts": len(traj_rows),
126
+ "sft_rows": len(sft_rows),
127
+ "dpo_rows": len(dpo_rows),
128
+ "quarantined": len(quarantine_rows),
129
+ **{f"dedup_{k}": v for k, v in dedup_stats.items()},
130
+ }
131
+ manifest.status = "partial" if partial else "building"
132
+ manifest.write(layout)
133
+ s3_contract.write_dataset_card(layout, manifest, dedup_stats=dedup_stats)
134
+ return manifest
135
+
136
+
137
+ __all__ = ["build_corpus"]
composer_replication/pipeline/dedup.py ADDED
@@ -0,0 +1,138 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """dedup.py — MinHash near-duplicate detection (finding D-12).
2
+
3
+ Cross-generation dedup is a flywheel-collapse mitigation: a self-training loop
4
+ that re-ingests its own outputs accumulates near-identical rows, and per-batch
5
+ `document_deduplicator` (the only dedup the old designs had) never sees across
6
+ runs. This module computes MinHash signatures over word 5-shingles so a run can
7
+ (a) dedup within itself and (b) accept the PRIOR run's signature file and dedup
8
+ against it (lineage threaded by `RunManifest.parent_run_id`).
9
+
10
+ Pragmatic v1: builtin-hash permutation MinHash with N=64 seeds, no banding/LSH
11
+ (O(n^2) pair scan — fine for Stage-0 corpus sizes, thousands of rows).
12
+ `datasketch` (MinHashLSH) is the documented upgrade path when row counts make
13
+ the pair scan bite.
14
+
15
+ NOTE on hash stability: Python's builtin `hash()` over str is salted per
16
+ process (PYTHONHASHSEED), which would make signatures non-portable across
17
+ runs — exactly what cross-generation dedup needs. We therefore use md5-based
18
+ hashing (stable everywhere) despite the small speed cost.
19
+ """
20
+ from __future__ import annotations
21
+
22
+ import hashlib
23
+ import json
24
+ import re
25
+ from typing import IO, Callable, Iterable, Sequence
26
+
27
+ N_PERMUTATIONS = 64
28
+ _SHINGLE_W = 5
29
+ _WORD_RE = re.compile(r"\w+")
30
+ _MAX64 = (1 << 64) - 1
31
+
32
+
33
+ def _shingles(text: str, w: int = _SHINGLE_W) -> set[str]:
34
+ words = _WORD_RE.findall(text.lower())
35
+ if len(words) <= w:
36
+ return {" ".join(words)} if words else set()
37
+ return {" ".join(words[i:i + w]) for i in range(len(words) - w + 1)}
38
+
39
+
40
+ def _stable_hash(s: str, seed: int) -> int:
41
+ h = hashlib.md5(f"{seed}:{s}".encode()).digest()
42
+ return int.from_bytes(h[:8], "big")
43
+
44
+
45
+ def minhash_signature(text: str, n_perm: int = N_PERMUTATIONS) -> tuple[int, ...]:
46
+ """MinHash signature: per-seed minimum over the shingle set."""
47
+ sh = _shingles(text)
48
+ if not sh:
49
+ return tuple([_MAX64] * n_perm)
50
+ return tuple(min(_stable_hash(s, seed) for s in sh) for seed in range(n_perm))
51
+
52
+
53
+ def jaccard_estimate(sig_a: Sequence[int], sig_b: Sequence[int]) -> float:
54
+ """Estimated Jaccard similarity = fraction of agreeing signature slots."""
55
+ if len(sig_a) != len(sig_b) or not sig_a:
56
+ raise ValueError("signatures must be equal-length and non-empty")
57
+ return sum(1 for a, b in zip(sig_a, sig_b) if a == b) / len(sig_a)
58
+
59
+
60
+ def find_near_duplicates(
61
+ rows: Sequence[dict],
62
+ key_fn: Callable[[dict], str],
63
+ threshold: float = 0.85,
64
+ *,
65
+ prior_signatures: Sequence[Sequence[int]] | None = None,
66
+ ) -> list[tuple[int, int]]:
67
+ """All (i, j) index pairs whose estimated Jaccard >= threshold.
68
+
69
+ `prior_signatures` (from a previous run) participate as virtual rows with
70
+ negative indices -(k+1), so a pair (i, -1) means "row i duplicates prior
71
+ signature 0" — the cross-generation case.
72
+ """
73
+ sigs = [minhash_signature(key_fn(r)) for r in rows]
74
+ pairs: list[tuple[int, int]] = []
75
+ for i in range(len(sigs)):
76
+ for j in range(i + 1, len(sigs)):
77
+ if jaccard_estimate(sigs[i], sigs[j]) >= threshold:
78
+ pairs.append((i, j))
79
+ for k, prior in enumerate(prior_signatures or []):
80
+ if jaccard_estimate(sigs[i], prior) >= threshold:
81
+ pairs.append((i, -(k + 1)))
82
+ return pairs
83
+
84
+
85
+ def dedup(
86
+ rows: Sequence[dict],
87
+ key_fn: Callable[[dict], str],
88
+ threshold: float = 0.85,
89
+ *,
90
+ prior_signatures: Sequence[Sequence[int]] | None = None,
91
+ ) -> tuple[list[dict], dict]:
92
+ """Keep-first dedup. Returns (kept_rows, stats).
93
+
94
+ A row duplicating a PRIOR-run signature is dropped outright (the prior run
95
+ already owns it); within-run duplicates keep the earliest occurrence.
96
+ """
97
+ pairs = find_near_duplicates(rows, key_fn, threshold,
98
+ prior_signatures=prior_signatures)
99
+ drop: set[int] = set()
100
+ for i, j in pairs:
101
+ if j < 0:
102
+ drop.add(i) # duplicates a prior-run row
103
+ else:
104
+ drop.add(j) # keep-first within this run
105
+ kept = [r for i, r in enumerate(rows) if i not in drop]
106
+ return kept, {
107
+ "rows_in": len(rows),
108
+ "rows_kept": len(kept),
109
+ "dropped_within_run": len({j for _, j in pairs if j >= 0} & drop),
110
+ "dropped_cross_generation": len({i for i, j in pairs if j < 0} & drop),
111
+ "threshold": threshold,
112
+ }
113
+
114
+
115
+ def signatures_to_jsonl(rows: Sequence[dict], key_fn: Callable[[dict], str],
116
+ fh: IO[str]) -> int:
117
+ """Persist this run's signatures so the NEXT generation can dedup against
118
+ them (pass the loaded list as `prior_signatures`)."""
119
+ n = 0
120
+ for r in rows:
121
+ fh.write(json.dumps(list(minhash_signature(key_fn(r)))) + "\n")
122
+ n += 1
123
+ return n
124
+
125
+
126
+ def load_signatures(fh: IO[str]) -> list[tuple[int, ...]]:
127
+ return [tuple(json.loads(line)) for line in fh if line.strip()]
128
+
129
+
130
+ __all__ = [
131
+ "N_PERMUTATIONS",
132
+ "dedup",
133
+ "find_near_duplicates",
134
+ "jaccard_estimate",
135
+ "load_signatures",
136
+ "minhash_signature",
137
+ "signatures_to_jsonl",
138
+ ]
composer_replication/pipeline/s3_contract.py ADDED
@@ -0,0 +1,287 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """s3_contract.py — THE single dataset layout + manifest (finding V8/D-7/D-8).
2
+
3
+ Supersedes BOTH prior contracts: design-F1's `runs/<id>/{sft_corpus,dpo_pairs,
4
+ rl_task_pool,divergence_pairs,wm_tuples,holdout,diloco_rendezvous}` and
5
+ design-F2's `{traces,tasks,replay,task_grades,corpus}/v1/run_id=<id>` — the two
6
+ were never reconciled and coexisted in the grounding doc. One layout, one
7
+ manifest, two explicit serializers with a unit-tested leak guard.
8
+
9
+ Deliberate exclusions from the run layout:
10
+ * `diloco_rendezvous/` — training-comms state, not dataset; lives in its own
11
+ prefix/bucket (finding D-19).
12
+ * `wm_tuples/` — emitted only when the P4 world-model ablation is scheduled
13
+ (finding D-14); not part of Stage 0.
14
+
15
+ Layout (root = any local path or fsspec URI):
16
+ <root>/runs/<run_id>/
17
+ tasks/manifest.jsonl policy-safe task rows (golden_diff -> sha256)
18
+ tasks_full/manifest.jsonl construction-side full rows (RESTRICTED prefix)
19
+ traj/*.jsonl CanonicalTrajectory records (audit trail)
20
+ corpus_sft/rows.jsonl admitted SFT rows (to_policy_row output)
21
+ corpus_dpo/rows.jsonl DPO-candidate rows
22
+ holdout/tasks.jsonl held-out task ids+rows (never rolled out)
23
+ quarantine/*.jsonl rejected trajectories w/ reasons (audit)
24
+ manifest.json RunManifest
25
+ DATASET_CARD.md human-readable card
26
+ """
27
+ from __future__ import annotations
28
+
29
+ import dataclasses
30
+ import hashlib
31
+ import json
32
+ from dataclasses import dataclass, field
33
+ from typing import IO, Iterable
34
+
35
+ from composer_replication.datagen.schema import FeatureDeletionTask
36
+
37
+ SCHEMA_VERSION = "1"
38
+
39
+
40
+ def _is_local(root: str) -> bool:
41
+ return "://" not in root or root.startswith("file://")
42
+
43
+
44
+ def _open(path: str, mode: str = "w") -> IO[str]:
45
+ """Open a path for text IO; plain `open` locally, fsspec for s3:// etc.
46
+
47
+ fsspec is lazy so the module (and all local-corpus runs) need no extra dep.
48
+ """
49
+ if _is_local(path):
50
+ import os
51
+ local = path.removeprefix("file://")
52
+ os.makedirs(os.path.dirname(local), exist_ok=True)
53
+ return open(local, mode, encoding="utf-8")
54
+ try:
55
+ import fsspec # noqa: PLC0415 — lazy heavy dep
56
+ except ImportError as e:
57
+ raise RuntimeError(
58
+ "Non-local corpus roots require fsspec; install with "
59
+ "`pip install -e .[serverless]`. Got: " + repr(e)
60
+ ) from e
61
+ return fsspec.open(path, mode, encoding="utf-8").open()
62
+
63
+
64
+ def _exists(path: str) -> bool:
65
+ if _is_local(path):
66
+ import os
67
+ return os.path.exists(path.removeprefix("file://"))
68
+ import fsspec # noqa: PLC0415
69
+ fs, _, paths = fsspec.get_fs_token_paths(path)
70
+ return bool(fs.exists(paths[0]))
71
+
72
+
73
+ @dataclass(frozen=True)
74
+ class RunLayout:
75
+ """Pure-path logic for one run's prefixes — testable without any IO."""
76
+
77
+ root: str
78
+ run_id: str
79
+
80
+ def _p(self, *parts: str) -> str:
81
+ base = self.root.rstrip("/")
82
+ return f"{base}/runs/{self.run_id}/" + "/".join(parts)
83
+
84
+ @property
85
+ def tasks_path(self) -> str:
86
+ return self._p("tasks", "manifest.jsonl")
87
+
88
+ @property
89
+ def tasks_full_path(self) -> str:
90
+ # RESTRICTED prefix: carries golden_diff/deleted_symbols. On S3 this
91
+ # prefix gets a deny-by-default policy; locally it is still separated
92
+ # so a naive `corpus_*` glob can never sweep it up.
93
+ return self._p("tasks_full", "manifest.jsonl")
94
+
95
+ @property
96
+ def traj_path(self) -> str:
97
+ return self._p("traj", "trajectories.jsonl")
98
+
99
+ @property
100
+ def sft_path(self) -> str:
101
+ return self._p("corpus_sft", "rows.jsonl")
102
+
103
+ @property
104
+ def dpo_path(self) -> str:
105
+ return self._p("corpus_dpo", "rows.jsonl")
106
+
107
+ @property
108
+ def holdout_path(self) -> str:
109
+ return self._p("holdout", "tasks.jsonl")
110
+
111
+ @property
112
+ def quarantine_path(self) -> str:
113
+ return self._p("quarantine", "rejected.jsonl")
114
+
115
+ @property
116
+ def manifest_path(self) -> str:
117
+ return self._p("manifest.json")
118
+
119
+ @property
120
+ def card_path(self) -> str:
121
+ return self._p("DATASET_CARD.md")
122
+
123
+
124
+ @dataclass
125
+ class RunManifest:
126
+ """Run-level metadata: counts, cost, lineage, budget, acceptance status.
127
+
128
+ `created_at` is CALLER-passed (never datetime.now() in here) so manifests
129
+ are reproducible in tests. `parent_run_id` threads flywheel lineage so
130
+ cross-generation dedup (finding D-12) can find prior signatures.
131
+ """
132
+
133
+ run_id: str
134
+ created_at: str
135
+ source: str = ""
136
+ counts: dict = field(default_factory=dict)
137
+ cost_usd: float = 0.0
138
+ parent_run_id: str | None = None
139
+ schema_version: str = SCHEMA_VERSION
140
+ status: str = "building" # building | accepted | rejected | partial
141
+ budget_usd: float | None = None
142
+
143
+ def spend(self, usd: float) -> None:
144
+ self.cost_usd += usd
145
+
146
+ @property
147
+ def over_budget(self) -> bool:
148
+ return self.budget_usd is not None and self.cost_usd >= self.budget_usd
149
+
150
+ def write(self, layout: RunLayout) -> None:
151
+ with _open(layout.manifest_path) as f:
152
+ json.dump(dataclasses.asdict(self), f, indent=2)
153
+
154
+ @classmethod
155
+ def read(cls, layout: RunLayout) -> RunManifest:
156
+ with _open(layout.manifest_path, "r") as f:
157
+ return cls(**json.load(f))
158
+
159
+
160
+ # ---------------------------------------------------------------------
161
+ # Writers — the leak guard lives here (finding D-8)
162
+ # ---------------------------------------------------------------------
163
+
164
+
165
+ def _task_row_policy_safe(task: FeatureDeletionTask) -> dict:
166
+ """Task row with the construction-side secrets REPLACED, not just hidden.
167
+
168
+ `asdict()` includes `golden_diff` despite `repr=False` — that is exactly
169
+ the leak D-8 flagged. We keep provenance via a sha256 (verifiable, not
170
+ recoverable) and drop `deleted_symbols` entirely (they name the answer).
171
+ """
172
+ row = dataclasses.asdict(task)
173
+ gold = row.pop("golden_diff", "")
174
+ row.pop("deleted_symbols", None)
175
+ row["golden_diff_sha256"] = hashlib.sha256(gold.encode()).hexdigest() if gold else ""
176
+ return row
177
+
178
+
179
+ def write_tasks(layout: RunLayout, tasks: Iterable[FeatureDeletionTask]) -> int:
180
+ """Write the POLICY-SAFE task manifest (the default everything reads)."""
181
+ n = 0
182
+ with _open(layout.tasks_path) as f:
183
+ for t in tasks:
184
+ f.write(json.dumps(_task_row_policy_safe(t)) + "\n")
185
+ n += 1
186
+ return n
187
+
188
+
189
+ def write_tasks_full(layout: RunLayout, tasks: Iterable[FeatureDeletionTask]) -> int:
190
+ """Write FULL task rows (incl. golden_diff) to the RESTRICTED prefix.
191
+
192
+ Only the validator/monitor side reads this; never corpus consumers.
193
+ """
194
+ n = 0
195
+ with _open(layout.tasks_full_path) as f:
196
+ for t in tasks:
197
+ f.write(json.dumps(dataclasses.asdict(t)) + "\n")
198
+ n += 1
199
+ return n
200
+
201
+
202
+ def _write_jsonl(path: str, rows: Iterable[dict]) -> int:
203
+ n = 0
204
+ with _open(path) as f:
205
+ for r in rows:
206
+ f.write(json.dumps(r) + "\n")
207
+ n += 1
208
+ return n
209
+
210
+
211
+ def write_sft_rows(layout: RunLayout, rows: Iterable[dict]) -> int:
212
+ return _write_jsonl(layout.sft_path, rows)
213
+
214
+
215
+ def write_dpo_rows(layout: RunLayout, rows: Iterable[dict]) -> int:
216
+ return _write_jsonl(layout.dpo_path, rows)
217
+
218
+
219
+ def write_quarantine(layout: RunLayout, rows: Iterable[dict]) -> int:
220
+ return _write_jsonl(layout.quarantine_path, rows)
221
+
222
+
223
+ def write_holdout(layout: RunLayout, tasks: Iterable[FeatureDeletionTask]) -> int:
224
+ return _write_jsonl(layout.holdout_path, (_task_row_policy_safe(t) for t in tasks))
225
+
226
+
227
+ def write_trajectories(layout: RunLayout, rows: Iterable[dict]) -> int:
228
+ return _write_jsonl(layout.traj_path, rows)
229
+
230
+
231
+ def write_dataset_card(layout: RunLayout, manifest: RunManifest,
232
+ *, license_tiers: dict[str, int] | None = None,
233
+ dedup_stats: dict | None = None,
234
+ decontamination_note: str = "") -> None:
235
+ """A small human-readable dataset card (finding D-18)."""
236
+ lines = [
237
+ f"# Dataset card — run `{manifest.run_id}`",
238
+ "",
239
+ f"- **created:** {manifest.created_at}",
240
+ f"- **source:** {manifest.source}",
241
+ f"- **status:** {manifest.status}",
242
+ f"- **schema_version:** {manifest.schema_version}",
243
+ f"- **cost (USD):** {manifest.cost_usd:.2f}"
244
+ + (f" / budget {manifest.budget_usd:.2f}" if manifest.budget_usd else ""),
245
+ f"- **lineage:** parent_run_id={manifest.parent_run_id or 'none'}",
246
+ "",
247
+ "## Counts",
248
+ "",
249
+ ]
250
+ for k, v in sorted(manifest.counts.items()):
251
+ lines.append(f"- {k}: {v}")
252
+ if license_tiers:
253
+ lines += ["", "## License tiers seen", ""]
254
+ lines += [f"- {k}: {v}" for k, v in sorted(license_tiers.items())]
255
+ lines += ["", "## Decontamination", "",
256
+ decontamination_note or
257
+ "All source repos checked against the SWE-bench-family eval list "
258
+ "(datagen.repo_gate.DECONTAMINATION_LIST) at ingest."]
259
+ if dedup_stats:
260
+ lines += ["", "## Dedup", ""]
261
+ lines += [f"- {k}: {v}" for k, v in sorted(dedup_stats.items())]
262
+ lines += ["", "Policy-safe rows only: `golden_diff` is sha256-hashed and "
263
+ "`deleted_symbols` dropped in `tasks/`, `corpus_*/`, `holdout/` "
264
+ "(full rows live in the restricted `tasks_full/`).", ""]
265
+ with _open(layout.card_path) as f:
266
+ f.write("\n".join(lines))
267
+
268
+
269
+ def manifest_exists(layout: RunLayout) -> bool:
270
+ """Write-once guard for the driver (finding D-21 idempotency)."""
271
+ return _exists(layout.manifest_path)
272
+
273
+
274
+ __all__ = [
275
+ "SCHEMA_VERSION",
276
+ "RunLayout",
277
+ "RunManifest",
278
+ "manifest_exists",
279
+ "write_dataset_card",
280
+ "write_dpo_rows",
281
+ "write_holdout",
282
+ "write_quarantine",
283
+ "write_sft_rows",
284
+ "write_tasks",
285
+ "write_tasks_full",
286
+ "write_trajectories",
287
+ ]
composer_replication/pipeline/tests/__init__.py ADDED
File without changes
composer_replication/pipeline/tests/test_pipeline.py ADDED
@@ -0,0 +1,223 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Tests for the Stage-0 pipeline: contract, dedup, driver.
2
+
3
+ Load-bearing coverage: the sentinel leak guard on write_tasks (finding D-8),
4
+ holdout exclusion + budget stop + idempotency in build_corpus (D-21), and the
5
+ cross-generation dedup path (D-12).
6
+ """
7
+ from __future__ import annotations
8
+
9
+ import io
10
+ import json
11
+ from pathlib import Path
12
+
13
+ import pytest
14
+
15
+ from composer_replication.datagen.env import FeatureDeletionEnv
16
+ from composer_replication.datagen.rollout_harness import ScriptedPolicy
17
+ from composer_replication.datagen.sandbox import FakeSandbox
18
+ from composer_replication.datagen.schema import FeatureDeletionTask
19
+ from composer_replication.datagen.trajectory import ToolCall
20
+ from composer_replication.pipeline.build_corpus import build_corpus
21
+ from composer_replication.pipeline.dedup import (
22
+ dedup,
23
+ find_near_duplicates,
24
+ jaccard_estimate,
25
+ load_signatures,
26
+ minhash_signature,
27
+ signatures_to_jsonl,
28
+ )
29
+ from composer_replication.pipeline.s3_contract import (
30
+ RunLayout,
31
+ RunManifest,
32
+ write_dataset_card,
33
+ write_tasks,
34
+ write_tasks_full,
35
+ )
36
+
37
+
38
+ def _task(i: int, **over) -> FeatureDeletionTask:
39
+ base = dict(
40
+ task_id=f"task-{i:03d}", repo="org/repo", base_commit="abc",
41
+ broken_image="img:1", test_command="pytest -q",
42
+ fail_to_pass=(f"t/a.py::t{i}",), pass_to_pass=("t/a.py::keep",),
43
+ golden_diff="SENTINEL_NEVER_LEAK", deleted_symbols=("secret_fn",),
44
+ )
45
+ base.update(over)
46
+ return FeatureDeletionTask(**base)
47
+
48
+
49
+ # ---------------------------------------------------------------------
50
+ # RunLayout / RunManifest
51
+ # ---------------------------------------------------------------------
52
+
53
+
54
+ def test_layout_paths_are_pure_and_namespaced():
55
+ lay = RunLayout(root="/data/corpora", run_id="run42")
56
+ assert lay.sft_path == "/data/corpora/runs/run42/corpus_sft/rows.jsonl"
57
+ assert lay.manifest_path == "/data/corpora/runs/run42/manifest.json"
58
+ s3 = RunLayout(root="s3://bucket/prefix/", run_id="r")
59
+ assert s3.tasks_path == "s3://bucket/prefix/runs/r/tasks/manifest.jsonl"
60
+
61
+
62
+ def test_manifest_round_trip_and_budget(tmp_path):
63
+ lay = RunLayout(root=str(tmp_path), run_id="r1")
64
+ m = RunManifest(run_id="r1", created_at="2026-06-09T00:00:00Z",
65
+ source="test", budget_usd=1.0)
66
+ m.spend(0.4)
67
+ assert not m.over_budget
68
+ m.spend(0.6)
69
+ assert m.over_budget
70
+ m.write(lay)
71
+ m2 = RunManifest.read(lay)
72
+ assert m2.cost_usd == pytest.approx(1.0)
73
+ assert m2.budget_usd == 1.0
74
+
75
+
76
+ # ---------------------------------------------------------------------
77
+ # THE leak guard (finding D-8)
78
+ # ---------------------------------------------------------------------
79
+
80
+
81
+ def test_write_tasks_never_leaks_golden_diff(tmp_path):
82
+ lay = RunLayout(root=str(tmp_path), run_id="r1")
83
+ write_tasks(lay, [_task(1)])
84
+ blob = Path(lay.tasks_path).read_text()
85
+ assert "SENTINEL_NEVER_LEAK" not in blob
86
+ assert "secret_fn" not in blob
87
+ row = json.loads(blob.splitlines()[0])
88
+ assert row["golden_diff_sha256"] # provenance preserved as a hash
89
+ # The restricted full writer DOES carry it (construction side only).
90
+ write_tasks_full(lay, [_task(1)])
91
+ assert "SENTINEL_NEVER_LEAK" in Path(lay.tasks_full_path).read_text()
92
+
93
+
94
+ # ---------------------------------------------------------------------
95
+ # MinHash dedup
96
+ # ---------------------------------------------------------------------
97
+
98
+ _TEXT_A = "the quick brown fox jumps over the lazy dog and then runs far away home tonight"
99
+ _TEXT_A2 = "the quick brown fox jumps over the lazy dog and then runs far away home today"
100
+ _TEXT_B = "import numpy as np def main(): return np.zeros(10) print(main()) totally different content here"
101
+
102
+
103
+ def test_jaccard_estimate_near_duplicates_high_disjoint_low():
104
+ sa, sa2, sb = (minhash_signature(t) for t in (_TEXT_A, _TEXT_A2, _TEXT_B))
105
+ assert jaccard_estimate(sa, sa2) > 0.5
106
+ assert jaccard_estimate(sa, sb) < 0.2
107
+ assert jaccard_estimate(sa, sa) == 1.0
108
+
109
+
110
+ def test_dedup_keeps_first_and_drops_near_dup():
111
+ rows = [{"text": _TEXT_A}, {"text": _TEXT_A2}, {"text": _TEXT_B}]
112
+ kept, stats = dedup(rows, lambda r: r["text"], threshold=0.5)
113
+ assert [r["text"] for r in kept] == [_TEXT_A, _TEXT_B]
114
+ assert stats["dropped_within_run"] == 1
115
+
116
+
117
+ def test_cross_generation_dedup_via_signature_file():
118
+ prior_rows = [{"text": _TEXT_A}]
119
+ buf = io.StringIO()
120
+ signatures_to_jsonl(prior_rows, lambda r: r["text"], buf)
121
+ buf.seek(0)
122
+ prior_sigs = load_signatures(buf)
123
+
124
+ rows = [{"text": _TEXT_A2}, {"text": _TEXT_B}]
125
+ kept, stats = dedup(rows, lambda r: r["text"], threshold=0.5,
126
+ prior_signatures=prior_sigs)
127
+ assert [r["text"] for r in kept] == [_TEXT_B]
128
+ assert stats["dropped_cross_generation"] == 1
129
+
130
+
131
+ def test_find_near_duplicates_pairs():
132
+ rows = [{"t": _TEXT_A}, {"t": _TEXT_A2}]
133
+ assert find_near_duplicates(rows, lambda r: r["t"], 0.5) == [(0, 1)]
134
+
135
+
136
+ # ---------------------------------------------------------------------
137
+ # build_corpus end-to-end (FakeSandbox + ScriptedPolicy)
138
+ # ---------------------------------------------------------------------
139
+
140
+
141
+ def _passing_policy():
142
+ # Flips both this task's F2P tests green generically: FakeSandbox's
143
+ # set_outcome takes explicit test names, so the fixture tasks share names
144
+ # via the same fail_to_pass tuple pattern; we set a superset.
145
+ outcomes = {f"t/a.py::t{i}": True for i in range(20)}
146
+ outcomes["t/a.py::keep"] = True
147
+ return ScriptedPolicy(actions=[ToolCall("set_outcome", {"outcomes": outcomes}), "done"])
148
+
149
+
150
+ def _failing_policy():
151
+ return ScriptedPolicy(actions=["gave up immediately"])
152
+
153
+
154
+ def _env():
155
+ return FeatureDeletionEnv(FakeSandbox(test_outcomes={"t/a.py::keep": True}))
156
+
157
+
158
+ def test_build_corpus_end_to_end(tmp_path):
159
+ tasks = [_task(i) for i in range(6)]
160
+ lay = RunLayout(root=str(tmp_path), run_id="e2e")
161
+ manifest = RunManifest(run_id="e2e", created_at="2026-06-09T00:00:00Z", source="fixture")
162
+
163
+ out = build_corpus(tasks, _env, _passing_policy, lay, manifest,
164
+ holdout_frac=0.34, holdout_seed=7)
165
+
166
+ # Holdout exclusion: holdout tasks were never rolled out.
167
+ assert out.counts["tasks_holdout"] >= 1
168
+ assert out.counts["rollouts"] == out.counts["tasks_train"]
169
+ # Full passes routed to SFT (post-dedup near-identical rows collapse —
170
+ # the fixture tasks produce near-identical messages, which is itself a
171
+ # realistic dedup scenario).
172
+ assert out.counts["sft_rows"] >= 1
173
+ assert out.counts["quarantined"] == 0
174
+ # Files exist and the SFT corpus never leaks the sentinel.
175
+ sft_blob = Path(lay.sft_path).read_text()
176
+ assert "SENTINEL_NEVER_LEAK" not in sft_blob
177
+ assert Path(lay.card_path).exists()
178
+ assert Path(lay.holdout_path).exists()
179
+
180
+
181
+ def test_build_corpus_quarantines_failures(tmp_path):
182
+ tasks = [_task(i) for i in range(3)]
183
+ lay = RunLayout(root=str(tmp_path), run_id="fail")
184
+ manifest = RunManifest(run_id="fail", created_at="2026-06-09T00:00:00Z", source="fixture")
185
+ out = build_corpus(tasks, _env, _failing_policy, lay, manifest,
186
+ holdout_frac=0.34, holdout_seed=7)
187
+ assert out.counts["sft_rows"] == 0
188
+ assert out.counts["quarantined"] == out.counts["rollouts"] > 0
189
+
190
+
191
+ def test_build_corpus_budget_stop_marks_partial(tmp_path):
192
+ tasks = [_task(i) for i in range(6)]
193
+ lay = RunLayout(root=str(tmp_path), run_id="budget")
194
+ manifest = RunManifest(run_id="budget", created_at="2026-06-09T00:00:00Z",
195
+ source="fixture", budget_usd=0.25)
196
+ out = build_corpus(tasks, _env, _passing_policy, lay, manifest,
197
+ holdout_frac=0.2, holdout_seed=7,
198
+ cost_per_rollout_usd=0.1)
199
+ assert out.status == "partial"
200
+ assert out.counts["rollouts"] < out.counts["tasks_train"]
201
+
202
+
203
+ def test_build_corpus_is_write_once(tmp_path):
204
+ tasks = [_task(i) for i in range(3)]
205
+ lay = RunLayout(root=str(tmp_path), run_id="once")
206
+ m1 = RunManifest(run_id="once", created_at="2026-06-09T00:00:00Z", source="fixture")
207
+ build_corpus(tasks, _env, _passing_policy, lay, m1, holdout_frac=0.34)
208
+ m2 = RunManifest(run_id="once", created_at="2026-06-09T00:00:01Z", source="fixture")
209
+ with pytest.raises(FileExistsError, match="write-once"):
210
+ build_corpus(tasks, _env, _passing_policy, lay, m2, holdout_frac=0.34)
211
+
212
+
213
+ def test_dataset_card_contents(tmp_path):
214
+ lay = RunLayout(root=str(tmp_path), run_id="card")
215
+ m = RunManifest(run_id="card", created_at="2026-06-09T00:00:00Z",
216
+ source="fixture", counts={"sft_rows": 3})
217
+ write_dataset_card(lay, m, license_tiers={"REDISTRIBUTABLE": 3},
218
+ dedup_stats={"rows_kept": 3})
219
+ card = Path(lay.card_path).read_text()
220
+ assert "run `card`" in card
221
+ assert "sft_rows: 3" in card
222
+ assert "REDISTRIBUTABLE: 3" in card
223
+ assert "Decontamination" in card
composer_replication/teacher_replay.py CHANGED
@@ -4,8 +4,12 @@ This is channel 3 of the integrated trainer: at each step of a frozen agentic
4
  trace, query N pre-trained external teachers (frontier models from different
5
  labs) and convert teacher disagreement into preference pairs for DPO loss.
6
 
7
- Generalized from spike-001's `replay.py`. Verified economic floor (✅ spike 001):
8
- $0.98 mean per-trace cost ungated, $0.30/trace projected with VOI gating.
 
 
 
 
9
 
10
  Usage:
11
  from teacher_replay import replay_trace, extract_dpo_pairs
 
4
  trace, query N pre-trained external teachers (frontier models from different
5
  labs) and convert teacher disagreement into preference pairs for DPO loss.
6
 
7
+ Generalized from spike-001's `replay.py`. Cost calibration (✅ spike 001,
8
+ relabeled per deepread finding V11): $0.98 mean per-TRACE cost ungated was
9
+ measured on a ~50-state SYNTHETIC trace at N=3 teachers; real Claude Code
10
+ sessions run 125–2,830 tool-use messages (ADR-002), so a full real session is
11
+ ~2 orders of magnitude more (~$70–80 flat, before VOI gating). $0.30/trace
12
+ projected with VOI gating, same synthetic basis.
13
 
14
  Usage:
15
  from teacher_replay import replay_trace, extract_dpo_pairs
composer_replication/trainer/kl_in_reward.py CHANGED
@@ -9,7 +9,9 @@ literature says this is not cosmetic:
9
 
10
  * arXiv:2512.21852 ("A Comedy of Estimators") — k1-in-reward improves OOD
11
  generalization; k3-in-reward can collapse.
12
- * verl adopted k1-in-reward as its *only* reverse-KL option.
 
 
13
  * TRL issue #4967 tracks the same divergence.
14
 
15
  OOD generalization is exactly the "take any model to the next level" axis, so
 
9
 
10
  * arXiv:2512.21852 ("A Comedy of Estimators") — k1-in-reward improves OOD
11
  generalization; k3-in-reward can collapse.
12
+ * verl ships k1-in-reward as its default/recommended reverse-KL option
13
+ (it also supports a k3-family "low_var_kl" — wording corrected per
14
+ deepread finding V13).
15
  * TRL issue #4967 tracks the same divergence.
16
 
17
  OOD generalization is exactly the "take any model to the next level" axis, so
docs/COMPOSER_RECIPE_MAPPING.md CHANGED
@@ -22,7 +22,7 @@ The Cursor blog discusses **only three** training innovations explicitly. Everyt
22
 
23
  **Cited prior art** (Cursor's footnote 1):
24
  - **OPSD: Self-Distilled Reasoner — On-Policy Self-Distillation for LLMs** (Zhao et al., 2026, [arXiv:2601.18734](https://arxiv.org/abs/2601.18734), [GitHub: siyan-zhao/OPSD](https://github.com/siyan-zhao/OPSD)). The original on-policy-self-distillation framework: single LLM, teacher conditioned on privileged information (e.g. ground-truth answer), student sees only the question, loss = per-token KL on student's own rollouts.
25
- - **SDPO: Reinforcement Learning via Self-Distillation** (Hübotter et al., 2026, [arXiv:2601.20802](https://arxiv.org/abs/2601.20802), ICLR 2026 Scaling Post-training Workshop). Generalizes OPSD to RL with rich feedback: *"SDPO treats the current model conditioned on feedback as a self-teacher and distills its feedback-informed next-token predictions back into the policy."* This is **mathematically the same** as Composer's targeted-textual-feedback method. **There is published code.** Comparison table from the SDPO paper:
26
 
27
  | Method | Sampling | Signal | Feedback |
28
  |---|---|---|---|
@@ -72,7 +72,7 @@ This is **infrastructure, not algorithm**. It only matters at MoE-1T scale; for
72
  | Composer 2.5 stage | Blog mechanism | Our replication target | v0.0 | v0.1 | v0.2 |
73
  |---|---|---|---|---|---|
74
  | **(a)** Continued pretraining on code | Standard pretraining, code-weighted | Skip — start from already-code-tuned `Qwen3-Coder-7B` or `Qwen3-Coder-30B-A3B` | ✗ | ✗ | ✗ |
75
- | **(b)** Synthetic data at scale | Feature Deletion + 24 other (unnamed) generators | Build 1 generator (Feature Deletion) as OpenEnv-compatible env. Use SWE-bench-lite and SWE-Gym as drop-in alternatives. | ✗ (use SWE-bench-lite only) | ✓ (build Feature Deletion) | scale generator suite |
76
  | **(c)** Realistic-environment RL (RLVR) | Async sandboxes, same tool harness as production | TRL `GRPOTrainer` + verifiers + OpenEnv; SWE-bench-lite env in v0.0; build sandboxed code execution env in v0.1 | ✓ baseline | ✓ + DAPO patches | + decentralized rollouts |
77
  | **(d)** Targeted RL w/ textual feedback (Composer's secret sauce) | Same-model self-distill: insert hint into context → teacher; original → student; on-policy KL at the turn | **Lift the OPSD/SDPO loss directly from `siyan-zhao/OPSD`** (published code, MIT). Generate hints via templates (v0.1) or LLM (v0.2). | ✗ (deferred) | ✓ (this is the Composer-recipe channel) | + learned hint generator |
78
  | **(e)** Trace-replay multi-teacher distill (NOVEL — our addition) | N/A (not in Composer) | N=3 teachers (Opus 4.7, GPT-5, DeepSeek V4 Pro) replay each step; disagreement → DPO pairs | ✓ (this is the v0.0 novelty bet) | ✓ + VOI gating | + tiered teachers |
@@ -148,7 +148,7 @@ Primary sources for each Composer-2.5 component, post-audit:
148
  - **Cursor blog** — [Introducing Composer 2.5](https://cursor.com/blog/composer-2-5) (2026)
149
  - **Cursor blog** — [Composer 2 technical report](https://cursor.com/blog/composer-2-technical-report) (predecessor; named the "Anyrun" environment per subagent — verify if needed)
150
  - **OPSD paper** — Zhao et al., *Self-Distilled Reasoner: On-Policy Self-Distillation for LLMs*, [arXiv:2601.18734](https://arxiv.org/abs/2601.18734), code at [siyan-zhao/OPSD](https://github.com/siyan-zhao/OPSD). MIT.
151
- - **SDPO paper** — Hübotter et al., *Reinforcement Learning via Self-Distillation*, [arXiv:2601.20802](https://arxiv.org/abs/2601.20802), ICLR 2026 Scaling Post-training Workshop. The direct formalization of Composer's hint-distill.
152
  - **Self-Distillation continual-learning** — [arXiv:2601.19897](https://arxiv.org/abs/2601.19897). Cited by Cursor; less directly relevant.
153
  - **Moonshot Kimi K2.5** — base model, [HF model card](https://huggingface.co/moonshotai/Kimi-K2-Thinking).
154
 
 
22
 
23
  **Cited prior art** (Cursor's footnote 1):
24
  - **OPSD: Self-Distilled Reasoner — On-Policy Self-Distillation for LLMs** (Zhao et al., 2026, [arXiv:2601.18734](https://arxiv.org/abs/2601.18734), [GitHub: siyan-zhao/OPSD](https://github.com/siyan-zhao/OPSD)). The original on-policy-self-distillation framework: single LLM, teacher conditioned on privileged information (e.g. ground-truth answer), student sees only the question, loss = per-token KL on student's own rollouts.
25
+ - **SDPO: Reinforcement Learning via Self-Distillation** (Hübotter et al., 2026, [arXiv:2601.20802](https://arxiv.org/abs/2601.20802), ICLR 2026 Scaling Post-training Workshop). Generalizes OPSD to RL with rich feedback: *"SDPO treats the current model conditioned on feedback as a self-teacher and distills its feedback-informed next-token predictions back into the policy."* Cursor's blog cites this paper only as **background** ("For more background on this approach see…") — NOT as its mechanism; SDPO's published loss is full-rollout with feedback-in-prefix and an EMA-regularized teacher, while Composer's blog describes a turn-localized hint splice. Closely related, **not verified identical** (deepread finding V1). **There is published code.** Comparison table from the SDPO paper:
26
 
27
  | Method | Sampling | Signal | Feedback |
28
  |---|---|---|---|
 
72
  | Composer 2.5 stage | Blog mechanism | Our replication target | v0.0 | v0.1 | v0.2 |
73
  |---|---|---|---|---|---|
74
  | **(a)** Continued pretraining on code | Standard pretraining, code-weighted | Skip — start from already-code-tuned `Qwen3-Coder-7B` or `Qwen3-Coder-30B-A3B` | ✗ | ✗ | ✗ |
75
+ | **(b)** Synthetic data at scale | Feature Deletion + an unspecified number of other generators (the blog says only "a range of approaches" — the old "24" was a back-formation from the 25x task multiplier; deepread finding V5) | Build 1 generator (Feature Deletion) as OpenEnv-compatible env. Use SWE-bench-lite and SWE-Gym as drop-in alternatives. | ✗ (use SWE-bench-lite only) | ✓ (build Feature Deletion) | scale generator suite |
76
  | **(c)** Realistic-environment RL (RLVR) | Async sandboxes, same tool harness as production | TRL `GRPOTrainer` + verifiers + OpenEnv; SWE-bench-lite env in v0.0; build sandboxed code execution env in v0.1 | ✓ baseline | ✓ + DAPO patches | + decentralized rollouts |
77
  | **(d)** Targeted RL w/ textual feedback (Composer's secret sauce) | Same-model self-distill: insert hint into context → teacher; original → student; on-policy KL at the turn | **Lift the OPSD/SDPO loss directly from `siyan-zhao/OPSD`** (published code, MIT). Generate hints via templates (v0.1) or LLM (v0.2). | ✗ (deferred) | ✓ (this is the Composer-recipe channel) | + learned hint generator |
78
  | **(e)** Trace-replay multi-teacher distill (NOVEL — our addition) | N/A (not in Composer) | N=3 teachers (Opus 4.7, GPT-5, DeepSeek V4 Pro) replay each step; disagreement → DPO pairs | ✓ (this is the v0.0 novelty bet) | ✓ + VOI gating | + tiered teachers |
 
148
  - **Cursor blog** — [Introducing Composer 2.5](https://cursor.com/blog/composer-2-5) (2026)
149
  - **Cursor blog** — [Composer 2 technical report](https://cursor.com/blog/composer-2-technical-report) (predecessor; named the "Anyrun" environment per subagent — verify if needed)
150
  - **OPSD paper** — Zhao et al., *Self-Distilled Reasoner: On-Policy Self-Distillation for LLMs*, [arXiv:2601.18734](https://arxiv.org/abs/2601.18734), code at [siyan-zhao/OPSD](https://github.com/siyan-zhao/OPSD). MIT.
151
+ - **SDPO paper** — Hübotter et al., *Reinforcement Learning via Self-Distillation*, [arXiv:2601.20802](https://arxiv.org/abs/2601.20802), ICLR 2026 Scaling Post-training Workshop. The closest published formalization; cited by Cursor only as background (deepread finding V1).
152
  - **Self-Distillation continual-learning** — [arXiv:2601.19897](https://arxiv.org/abs/2601.19897). Cited by Cursor; less directly relevant.
153
  - **Moonshot Kimi K2.5** — base model, [HF model card](https://huggingface.co/moonshotai/Kimi-K2-Thinking).
154
 
docs/adrs/ADR-016-stage0-dataset-pipeline.md ADDED
@@ -0,0 +1,119 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ status: accepted
3
+ date: 2026-06-09
4
+ deciders: [Codeseys, ARIA]
5
+ ---
6
+
7
+ # ADR-016: Stage-0 dataset-generation pipeline — SWE-smith engine + rollout harness + ingest gates + single contract
8
+
9
+ ## Context and Problem Statement
10
+
11
+ The user asked to "architect and build a pipeline that builds out a dataset like
12
+ the Composer 2.5 blog mentions, with our vision enhancements — point to an
13
+ open-source repo and use that to build the dataset, or use traces or other
14
+ datasets and enhance them."
15
+
16
+ Before building, a full-source critical review re-read every foundational paper
17
+ and blog (8 source clusters, `research/deepread/01-08`), ground-mapped the repo
18
+ (`00`), ran adversarial fidelity + design critics, and independently VERIFIED
19
+ every finding (`12` — 0 refuted). The verified verdict: the envisioned pipeline
20
+ had four structural breaks (seed-trace/oracle disjointness; no rollout harness —
21
+ the SFT corpus had NO producer; an uncomputable divergence gate; no
22
+ `Sandbox.fork()`), several missing controls (zero benchmark decontamination,
23
+ no secrets gate, a `golden_diff` serialization leak, two unreconciled S3
24
+ contracts, no cross-generation dedup), and a buy-vs-build inversion (the
25
+ planned image-builder duplicates `pip install swesmith`, whose PR-Mirror
26
+ strategy IS this repo's gold-patch-reversion mechanic and is validated best-of-
27
+ five by SWE-smith's own ablation, Table 5 of arXiv:2504.21798).
28
+
29
+ ## Decision
30
+
31
+ Build **Stage 0 local-first** (architecture: `research/deepread/13-synthesis-architecture.md`):
32
+
33
+ 1. **SWE-smith is the synthesis engine** for "point at a repo" (`[swesmith]`
34
+ extra; `datagen/swesmith_adapter.py` bridges its instances into
35
+ `FeatureDeletionTask`, handling the patch-semantics INVERSION — SWE-smith's
36
+ patch introduces the bug, so `golden_diff` = `reverse_unified_diff(patch)`).
37
+ `SweBenchAdapter` remains the bridge for SWE-bench-shaped substrates.
38
+ 2. **Ingest gates before anything else** (`datagen/repo_gate.py`): SPDX-ish
39
+ license detection → three tiers (REDISTRIBUTABLE / TRAINABLE_ONLY /
40
+ EXCLUDED, fail-closed) + **benchmark decontamination** against the
41
+ SWE-bench-family eval-repo list (hard fail).
42
+ 3. **The rollout harness is the corpus producer** (`datagen/rollout_harness.py`):
43
+ `collect_trajectory(env, task, policy)` runs a pluggable policy
44
+ (ScriptedPolicy for tests; OpenRouterPolicy stub; mini-swe-agent/SWE-agent
45
+ adoption is the documented upgrade) through `FeatureDeletionEnv` to
46
+ `_grade()`. Its env-grounded trajectories are ALSO the tree-of-work's seed
47
+ nodes — fixing the seed/oracle disjointness as a byproduct. `admit()` routes
48
+ typed signal: clean full pass → SFT; clean near-miss → DPO candidate;
49
+ guard-broken/hacked → quarantine (never raw negative gradient).
50
+ 4. **One canonical trajectory IR** (`datagen/trajectory.py`): `ToolCall` (whose
51
+ `canonical_form()` is the v1 divergence-gate action algebra, replacing the
52
+ whitespace stub), `CanonicalTrajectory`, adapters from Claude Code traces
53
+ (explicitly UNGRADED — demoted to flat/SFT uses), and `to_policy_row()` —
54
+ the ONE policy-visible serializer, unit-tested to never emit
55
+ `golden_diff`/`deleted_symbols` (sentinel test).
56
+ 5. **One reconciled dataset contract** (`pipeline/s3_contract.py`, supersedes
57
+ design-F1's and design-F2's divergent layouts): `runs/<id>/{tasks,
58
+ tasks_full(RESTRICTED), traj, corpus_sft, corpus_dpo, holdout, quarantine}`
59
+ + `RunManifest` (counts, cost, budget, `parent_run_id` lineage, status) +
60
+ dataset card. Policy-safe task rows carry `golden_diff_sha256`, never the
61
+ diff. DiLoCo rendezvous and `wm_tuples/` are deliberately OUT (separate
62
+ concern; ablation-gated respectively).
63
+ 6. **Cross-generation dedup** (`pipeline/dedup.py`): stable-hash MinHash over
64
+ word 5-shingles; a run can dedup against the prior generation's signature
65
+ file (flywheel-collapse mitigation). datasketch/LSH is the upgrade path.
66
+ 7. **The local stage-driver** (`pipeline/build_corpus.py`): holdout-split FIRST
67
+ (held-out tasks never rolled out), rollouts under a budget ceiling
68
+ (partial-marking), typed routing, dedup, write-once-per-run idempotency.
69
+
70
+ ## Fidelity corrections shipped with this ADR (deepread findings, all verified)
71
+
72
+ - **V1:** "SDPO is mathematically the same as Composer's mechanism" corrected
73
+ in `opsd.py` + `COMPOSER_RECIPE_MAPPING.md` — Cursor cites SDPO/OPSD as
74
+ *background*; our channel is a third, blog-inspired design (turn-localized
75
+ hint splice, live stop-grad teacher, no EMA).
76
+ - **V5:** fabricated numbers struck/tagged: 69.3%/Terminal-Bench parity (no
77
+ primary source), "24 other generators" (back-formed), "85% post-training
78
+ compute" (community speculation) — `research/01`, mapping doc, `research/06`,
79
+ `research/09`.
80
+ - **V7:** Streaming DiLoCo citation fixed (`diloco/__init__.py`): 2501.18512 =
81
+ Douillard et al.; Eager Updates = Kale et al. 2502.12996.
82
+ - **V11:** `teacher_replay.py` cost docstring relabeled ($0.98 = 50-state
83
+ synthetic trace; real sessions ~2 OOM more).
84
+ - **V13:** `kl_in_reward.py` "verl's only reverse-KL option" → default/
85
+ recommended (verl also ships a k3-family option).
86
+
87
+ ## What is deliberately NOT in Stage 0
88
+
89
+ - AWS orchestration (Glue/EMR/Batch/Bedrock-batch/Step Functions) — Stage 4,
90
+ only after local runs are routine (finding D-9).
91
+ - Tree depth>1 — gated on a `Sandbox.fork()` spike + a measured divergence-gate
92
+ firing rate (findings D-3/D-4). Depth-1 multi-candidate rollouts need no fork.
93
+ - World-model `wm_tuples/` emission — gated on the P4 ablation being scheduled
94
+ (finding D-14; CWM evidence is mid-training, not RL-time aux head — V6).
95
+ - Secrets/PII scrub at trace ingest (finding V9) — REQUIRED before any raw
96
+ Claude Code session is uploaded to shared storage; tracked as the next
97
+ pipeline item. Local-only runs are unaffected.
98
+
99
+ ## Acceptance gate
100
+
101
+ - [x] `repo_gate`: 53 tests (license tiers, decontamination, gate verdicts).
102
+ - [x] `swesmith_adapter`: 18 tests (patch INVERSION semantics, reverse round-trip,
103
+ strategy provenance, image conventions).
104
+ - [x] `trajectory` + `rollout_harness`: 13 tests (IR round-trips, SENTINEL leak
105
+ guard, env-grounded episode to grade 1.0 / guard-broken / near-miss,
106
+ admission routing).
107
+ - [x] `pipeline`: 12 tests (layout, manifest+budget, leak guard at the writer,
108
+ MinHash within-run + cross-generation, build_corpus e2e with holdout
109
+ exclusion + budget stop + write-once).
110
+ - [x] Full suite green: 511 passed / 66 skipped.
111
+ - [ ] Live swesmith synthesis on a real pointed-at repo (needs Docker+Linux) —
112
+ the documented `[~]` gate, same shape as ADR-010's Docker e2e.
113
+
114
+ ## More Information
115
+
116
+ - `research/deepread/13-synthesis-architecture.md` — the architecture this implements.
117
+ - `research/deepread/12-verified-findings.md` — the verified finding ledger (V1–V15).
118
+ - `research/deepread/02-swe-task-synthesis.md` — the SWE-smith/R2E-Gym/SWE-Gym deep-read.
119
+ - ADR-010 (the substrate-inversion base this extends), ADR-002 (trace source).
pyproject.toml CHANGED
@@ -83,6 +83,14 @@ aws = [
83
  "boto3>=1.34",
84
  "sagemaker>=2.200,<3",
85
  ]
 
 
 
 
 
 
 
 
86
  # Replaysim dataset normalization (per ADR-004)
87
  #
88
  # NOTE: data-juicer is intentionally NOT pinned as an extra. The package
 
83
  "boto3>=1.34",
84
  "sagemaker>=2.200,<3",
85
  ]
86
+ # SWE-smith task-synthesis engine (deepread finding V4 buy-vs-build verdict):
87
+ # the swesmith toolkit builds env images from arbitrary GitHub repos and
88
+ # synthesizes bugs (PR Mirror = this repo's gold-patch-reversion mechanic).
89
+ # LIVE synthesis needs Docker on Linux (the toolkit does not support macOS/
90
+ # Windows officially); the SwesmithAdapter itself needs nothing beyond core.
91
+ swesmith = [
92
+ "swesmith>=0.1",
93
+ ]
94
  # Replaysim dataset normalization (per ADR-004)
95
  #
96
  # NOTE: data-juicer is intentionally NOT pinned as an extra. The package
research/01-composer-2.5.md CHANGED
@@ -11,7 +11,7 @@
11
  > The targeted-textual-feedback method is correctly described, but this file does **not** cite the three self-distillation papers Cursor cites in footnote 1 (OPSD `arXiv:2601.18734`, SDPO `arXiv:2601.20802`, Self-Distillation Continual Learning `arXiv:2601.19897`). The mapping document does.
12
 
13
  ## Overview
14
- Cursor's Composer 2.5 is an advanced agentic coding model that powers the Cursor IDE. Released in mid-May 2026, it represents a massive leap in agentic capabilities, particularly for long-running, multi-file software engineering tasks. While the base weights are Moonshot AI's open-source **Kimi K2.5** model, roughly 85% of the total compute budget for Composer 2.5 was spent on Cursor's proprietary post-training and Reinforcement Learning (RL) pipeline.
15
 
16
  The resulting model is highly optimized for the exact constraints and tools of the Cursor environment (file edits, terminal usage, LSP interaction). Composer 2.5 is praised for having fewer "false-start" tool calls, avoiding prompt-baiting, and demonstrating a much calmer, more effective collaboration loop than its predecessors.
17
 
@@ -60,8 +60,8 @@ During post-training, Cursor employs **Sharded Muon** and **Dual Mesh HSDP (Hybr
60
  ## Performance Characteristics
61
  Cursor claims Composer 2.5 achieves a Pareto-optimal tradeoff between intelligence and inference cost compared to frontier models (Opus 4.5/4.6, GPT-5.4/5.5).
62
 
63
- * **Intelligence Improvements**: On Cursor's internal *CursorBench* (which tests sweeping, multi-file edits with ambiguous prompts), Composer 2.5 scored 69.3% (or ~61-63% depending on the specific benchmark version cited), a massive jump from Composer 1.5's ~44% and Composer 2's ~52%.
64
- * **Frontier Parity**: On public agentic benchmarks like *Terminal-Bench 2.0*, it hit 69.3%. On *SWE-bench Multilingual*, it achieved parity with or slightly surpassed OpenAI's GPT-5.5.
65
  * **Cost Efficiency**:
66
  * Standard Tier: $0.50 per 1M input / $2.50 per 1M output tokens.
67
  * Fast Tier: $3.00 per 1M input / $15.00 per 1M output tokens.
 
11
  > The targeted-textual-feedback method is correctly described, but this file does **not** cite the three self-distillation papers Cursor cites in footnote 1 (OPSD `arXiv:2601.18734`, SDPO `arXiv:2601.20802`, Self-Distillation Continual Learning `arXiv:2601.19897`). The mapping document does.
12
 
13
  ## Overview
14
+ Cursor's Composer 2.5 is an advanced agentic coding model that powers the Cursor IDE. Released in mid-May 2026, it represents a massive leap in agentic capabilities, particularly for long-running, multi-file software engineering tasks. While the base weights are Moonshot AI's open-source **Kimi K2.5** model, a large share of the compute budget went to Cursor's proprietary post-training/RL pipeline (the widely-circulated "85%" figure is community speculation, in NO primary source — deepread finding V5).
15
 
16
  The resulting model is highly optimized for the exact constraints and tools of the Cursor environment (file edits, terminal usage, LSP interaction). Composer 2.5 is praised for having fewer "false-start" tool calls, avoiding prompt-baiting, and demonstrating a much calmer, more effective collaboration loop than its predecessors.
17
 
 
60
  ## Performance Characteristics
61
  Cursor claims Composer 2.5 achieves a Pareto-optimal tradeoff between intelligence and inference cost compared to frontier models (Opus 4.5/4.6, GPT-5.4/5.5).
62
 
63
+ * **Intelligence Improvements**: On Cursor's internal *CursorBench* (which tests sweeping, multi-file edits with ambiguous prompts), Composer 2.5's score is NOT in any primary source (the circulating 69.3% figure appears in neither the 2.5 blog nor the Composer 2 techreport — deepread finding V5; the techreport's Table 1 gives Composer 2 = 61.3 CursorBench). Treat all 2.5 benchmark numbers as unverified.
64
+ * **Frontier Parity**: Claims of Terminal-Bench 2.0 / SWE-bench Multilingual parity circulate in secondary commentary only; neither primary source contains benchmark numbers for 2.5 (deepread finding V5).
65
  * **Cost Efficiency**:
66
  * Standard Tier: $0.50 per 1M input / $2.50 per 1M output tokens.
67
  * Fast Tier: $3.00 per 1M input / $15.00 per 1M output tokens.
research/06-feature-deletion-datagen.md CHANGED
@@ -327,7 +327,7 @@ Feature-Deletion is **embarrassingly parallel and CPU-bound** — no GPU in the
327
 
328
  1. **Deletion-target selection heuristic** — blog silent (`research/09` §1 "NO CHANGE"). We propose coverage-selectivity (§5 Path B); Cursor's actual heuristic is unknown.
329
  2. **Deleter model vs. program** — blog implies an agent deletes ("asked to delete code… such that the codebase remains functional"); we default to *programmatic* deletion (cheaper, deterministic, no second model). An LLM-deleter is a v0.2 escalation.
330
- 3. **The other ~24 generators** — Feature Deletion is "one synthetic approach… a range of approaches"; the rest are unnamed. Out of scope here; this brief delivers the one named generator.
331
  4. **"Agentic monitoring tools" internals** — unspecified; our §3c monitor is a best-effort programmatic stand-in.
332
  5. **Composer2.pdf (arXiv:2603.24477)** — flagged by `research/09` action-item #1 as the likely home of data-mix % and generator inventory; **not yet extracted**. Recommend a follow-up pull before scaling the generator suite.
333
 
 
327
 
328
  1. **Deletion-target selection heuristic** — blog silent (`research/09` §1 "NO CHANGE"). We propose coverage-selectivity (§5 Path B); Cursor's actual heuristic is unknown.
329
  2. **Deleter model vs. program** — blog implies an agent deletes ("asked to delete code… such that the codebase remains functional"); we default to *programmatic* deletion (cheaper, deterministic, no second model). An LLM-deleter is a v0.2 escalation.
330
+ 3. **The other generators (count UNKNOWN)** — Feature Deletion is "one synthetic approach… a range of approaches"; the rest are unnamed and uncounted (the old "~24" was a back-formation from the 25x task multiplier — deepread finding V5). Out of scope here; this brief delivers the one named generator.
331
  4. **"Agentic monitoring tools" internals** — unspecified; our §3c monitor is a best-effort programmatic stand-in.
332
  5. **Composer2.pdf (arXiv:2603.24477)** — flagged by `research/09` action-item #1 as the likely home of data-mix % and generator inventory; **not yet extracted**. Recommend a follow-up pull before scaling the generator suite.
333
 
research/09-composer-blog-delta-2026.md CHANGED
@@ -20,7 +20,7 @@ The **2.5 blog body is byte-for-byte unchanged** from what the mapping doc captu
20
 
21
  **DELTAS (not in / under-stated in COMPOSER_RECIPE_MAPPING.md):**
22
 
23
- - **[DELTA — new emphasis]** The phrase *"we both **select for** and **create** harder tasks **dynamically throughout the run**"* is a **dynamic curriculum / online task-selection** signal. The mapping doc captured "Feature Deletion + 24 unnamed generators" but did **not** flag that task difficulty is filtered *online* (the model "begins to get most training problems correct," so hard tasks are up-weighted live). This is a data-*mix*/curriculum detail with direct replication impact: our generator suite needs a difficulty filter / pass-rate gate, not just a static task bank.
24
  - **[DELTA — new authoritative source for CPT data mix]** The Composer 2 technical-report blog states the CPT data mix explicitly: *"continued pretraining on a data mix that **emphasizes code** to deepen the base model's coding knowledge"* and *"We find that **reducing pretraining loss improves downstream RL performance**, with better base knowledge reliably translating into a better agent."* The mapping doc marked "continued pretraining on heavily code-weighted data" as `[BLOG-VERIFIED]` from the 2.5 Muon section — but the **causal claim (CPT loss ↓ ⇒ RL performance ↑)** is new and is the stated *justification* for doing CPT at all. Relevant to our "skip CPT, start from Qwen3-Coder" decision: Cursor's own evidence says base-knowledge quality gates RL ceiling, which strengthens the case for starting from an already-code-tuned base.
25
  - **[DELTA — new artifact]** There is now a **full Composer 2 arXiv technical report: [arXiv:2603.24477](https://arxiv.org/abs/2603.24477)** and a downloadable PDF at **`https://cursor.com/resources/Composer2.pdf`** (authored by Sasha Rush et al.). The report explicitly *"covers... ablations on the training recipe, our approach to agent behavior shaping, and the design of our evaluation suite."* The mapping doc cited only the blog stub and never the arXiv ID/PDF. **This PDF is the most likely place to resolve the data-mix weighting %, the RL algorithm name, and the hint-generation mechanism — none of which are in either blog.** → Recommend a dedicated follow-up extraction of Composer2.pdf.
26
  - **[CONFIRM — "Anyrun"]** Mapping doc flagged "Anyrun" as possibly not Cursor-sourced. **Confirmed real:** the Composer 2 report blog says *"**Anyrun**, our internal compute platform for running hundreds of thousands of sandboxed coding environments."* It is a Composer-**2** artifact (carried into 2.5), correctly attributed. Resolves the mapping doc's open flag.
 
20
 
21
  **DELTAS (not in / under-stated in COMPOSER_RECIPE_MAPPING.md):**
22
 
23
+ - **[DELTA — new emphasis]** The phrase *"we both **select for** and **create** harder tasks **dynamically throughout the run**"* is a **dynamic curriculum / online task-selection** signal. The mapping doc captured "Feature Deletion + other unnamed generators" (its old "24" count was a back-formation — deepread finding V5) but did **not** flag that task difficulty is filtered *online* (the model "begins to get most training problems correct," so hard tasks are up-weighted live). This is a data-*mix*/curriculum detail with direct replication impact: our generator suite needs a difficulty filter / pass-rate gate, not just a static task bank.
24
  - **[DELTA — new authoritative source for CPT data mix]** The Composer 2 technical-report blog states the CPT data mix explicitly: *"continued pretraining on a data mix that **emphasizes code** to deepen the base model's coding knowledge"* and *"We find that **reducing pretraining loss improves downstream RL performance**, with better base knowledge reliably translating into a better agent."* The mapping doc marked "continued pretraining on heavily code-weighted data" as `[BLOG-VERIFIED]` from the 2.5 Muon section — but the **causal claim (CPT loss ↓ ⇒ RL performance ↑)** is new and is the stated *justification* for doing CPT at all. Relevant to our "skip CPT, start from Qwen3-Coder" decision: Cursor's own evidence says base-knowledge quality gates RL ceiling, which strengthens the case for starting from an already-code-tuned base.
25
  - **[DELTA — new artifact]** There is now a **full Composer 2 arXiv technical report: [arXiv:2603.24477](https://arxiv.org/abs/2603.24477)** and a downloadable PDF at **`https://cursor.com/resources/Composer2.pdf`** (authored by Sasha Rush et al.). The report explicitly *"covers... ablations on the training recipe, our approach to agent behavior shaping, and the design of our evaluation suite."* The mapping doc cited only the blog stub and never the arXiv ID/PDF. **This PDF is the most likely place to resolve the data-mix weighting %, the RL algorithm name, and the hint-generation mechanism — none of which are in either blog.** → Recommend a dedicated follow-up extraction of Composer2.pdf.
26
  - **[CONFIRM — "Anyrun"]** Mapping doc flagged "Anyrun" as possibly not Cursor-sourced. **Confirmed real:** the Composer 2 report blog says *"**Anyrun**, our internal compute platform for running hundreds of thousands of sandboxed coding environments."* It is a Composer-**2** artifact (carried into 2.5), correctly attributed. Resolves the mapping doc's open flag.
research/notes/230406767-raft-reward-ranked-finetuning-for-generative-foundation-model-alignmen.md ADDED
@@ -0,0 +1,224 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2304.06767] RAFT: Reward rAnked FineTuning for Generative Foundation Model
3
+ Alignment'
4
+ id: 230406767-raft-reward-ranked-finetuning-for-generative-foundation-model-alignmen
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:31:18.566124Z'
8
+ source: https://arxiv.org/abs/2304.06767
9
+ source_domain: arxiv.org
10
+ fetched_at: '2026-06-10T00:31:18.565918Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2304.06767] RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
20
+ Computer Science > Machine Learning
21
+ arXiv:2304.06767
22
+ (cs)
23
+ [Submitted on 13 Apr 2023 (
24
+ v1
25
+ ), last revised 1 Dec 2023 (this version, v4)]
26
+ Title:
27
+ RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
28
+ Authors:
29
+ Hanze Dong
30
+ ,
31
+ Wei Xiong
32
+ ,
33
+ Deepanshu Goyal
34
+ ,
35
+ Yihan Zhang
36
+ ,
37
+ Winnie Chow
38
+ ,
39
+ Rui Pan
40
+ ,
41
+ Shizhe Diao
42
+ ,
43
+ Jipeng Zhang
44
+ ,
45
+ Kashun Shum
46
+ ,
47
+ Tong Zhang
48
+ View a PDF of the paper titled RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment, by Hanze Dong and 9 other authors
49
+ View PDF
50
+ HTML (experimental)
51
+ Abstract:
52
+ Generative foundation models are susceptible to implicit biases that can arise from extensive unsupervised training data. Such biases can produce suboptimal samples, skewed outcomes, and unfairness, with potentially serious consequences. Consequently, aligning these models with human ethics and preferences is an essential step toward ensuring their responsible and effective deployment in real-world applications. Prior research has primarily employed Reinforcement Learning from Human Feedback (RLHF) to address this problem, where generative models are fine-tuned with RL algorithms guided by a human-feedback-informed reward model. However, the inefficiencies and instabilities associated with RL algorithms frequently present substantial obstacles to the successful alignment, necessitating the development of a more robust and streamlined approach. To this end, we introduce a new framework, Reward rAnked FineTuning (RAFT), designed to align generative models effectively. Utilizing a reward model and a sufficient number of samples, our approach selects the high-quality samples, discarding those that exhibit undesired behavior, and subsequently enhancing the model by fine-tuning on these filtered samples. Our studies show that RAFT can effectively improve the model performance in both reward learning and other automated metrics in both large language models and diffusion models.
53
+ Comments:
54
+ 29 pages, 12 figures, Published in Transactions on Machine Learning Research (TMLR)
55
+ Subjects:
56
+ Machine Learning (cs.LG)
57
+ ; Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
58
+ Cite as:
59
+ arXiv:2304.06767
60
+ [cs.LG]
61
+ (or
62
+ arXiv:2304.06767v4
63
+ [cs.LG]
64
+ for this version)
65
+ https://doi.org/10.48550/arXiv.2304.06767
66
+ Focus to learn more
67
+ arXiv-issued DOI via DataCite
68
+ Submission history
69
+ From: Hanze Dong [
70
+ view email
71
+ ]
72
+ [v1]
73
+ Thu, 13 Apr 2023 18:22:40 UTC (62,967 KB)
74
+ [v2]
75
+ Thu, 25 May 2023 06:27:31 UTC (42,022 KB)
76
+ [v3]
77
+ Wed, 30 Aug 2023 01:25:29 UTC (33,955 KB)
78
+ [v4]
79
+ Fri, 1 Dec 2023 14:28:06 UTC (34,049 KB)
80
+ Full-text links:
81
+ Access Paper:
82
+ View a PDF of the paper titled RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment, by Hanze Dong and 9 other authors
83
+ View PDF
84
+ HTML (experimental)
85
+ TeX Source
86
+ view license
87
+ Current browse context:
88
+ cs.LG
89
+ < prev
90
+ |
91
+ next >
92
+ new
93
+ |
94
+ recent
95
+ |
96
+ 2023-04
97
+ Change to browse by:
98
+ cs
99
+ cs.AI
100
+ cs.CL
101
+ cs.CV
102
+ stat
103
+ stat.ML
104
+ References & Citations
105
+ NASA ADS
106
+ Google Scholar
107
+ Semantic Scholar
108
+ export BibTeX citation
109
+ Loading...
110
+ BibTeX formatted citation
111
+ ×
112
+ loading...
113
+ Data provided by:
114
+ Bookmark
115
+ Bibliographic Tools
116
+ Bibliographic and Citation Tools
117
+ Bibliographic Explorer Toggle
118
+ Bibliographic Explorer
119
+ (
120
+ What is the Explorer?
121
+ )
122
+ Connected Papers Toggle
123
+ Connected Papers
124
+ (
125
+ What is Connected Papers?
126
+ )
127
+ Litmaps Toggle
128
+ Litmaps
129
+ (
130
+ What is Litmaps?
131
+ )
132
+ scite.ai Toggle
133
+ scite Smart Citations
134
+ (
135
+ What are Smart Citations?
136
+ )
137
+ Code, Data, Media
138
+ Code, Data and Media Associated with this Article
139
+ alphaXiv Toggle
140
+ alphaXiv
141
+ (
142
+ What is alphaXiv?
143
+ )
144
+ Links to Code Toggle
145
+ CatalyzeX Code Finder for Papers
146
+ (
147
+ What is CatalyzeX?
148
+ )
149
+ DagsHub Toggle
150
+ DagsHub
151
+ (
152
+ What is DagsHub?
153
+ )
154
+ GotitPub Toggle
155
+ Gotit.pub
156
+ (
157
+ What is GotitPub?
158
+ )
159
+ Huggingface Toggle
160
+ Hugging Face
161
+ (
162
+ What is Huggingface?
163
+ )
164
+ Links to Code Toggle
165
+ Papers with Code
166
+ (
167
+ What is Papers with Code?
168
+ )
169
+ ScienceCast Toggle
170
+ ScienceCast
171
+ (
172
+ What is ScienceCast?
173
+ )
174
+ Demos
175
+ Demos
176
+ Replicate Toggle
177
+ Replicate
178
+ (
179
+ What is Replicate?
180
+ )
181
+ Spaces Toggle
182
+ Hugging Face Spaces
183
+ (
184
+ What is Spaces?
185
+ )
186
+ Spaces Toggle
187
+ TXYZ.AI
188
+ (
189
+ What is TXYZ.AI?
190
+ )
191
+ Related Papers
192
+ Recommenders and Search Tools
193
+ Link to Influence Flower
194
+ Influence Flower
195
+ (
196
+ What are Influence Flowers?
197
+ )
198
+ Core recommender toggle
199
+ CORE Recommender
200
+ (
201
+ What is CORE?
202
+ )
203
+ IArxiv recommender toggle
204
+ IArxiv Recommender
205
+ (
206
+ What is IArxiv?
207
+ )
208
+ Author
209
+ Venue
210
+ Institution
211
+ Topic
212
+ About arXivLabs
213
+ arXivLabs: experimental projects with community collaborators
214
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
215
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
216
+ Have an idea for a project that will add value for arXiv's community?
217
+ Learn more about arXivLabs
218
+ .
219
+ Which authors of this paper are endorsers?
220
+ |
221
+ Disable MathJax
222
+ (
223
+ What is MathJax?
224
+ )
research/notes/230510601-tree-of-thoughts-deliberate-problem-solving-with-large-language-models-2.md ADDED
@@ -0,0 +1,2735 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2305.10601] Tree of Thoughts: Deliberate Problem Solving with Large Language
3
+ Models'
4
+ id: 230510601-tree-of-thoughts-deliberate-problem-solving-with-large-language-models-2
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:41:12.876142Z'
8
+ source: https://ar5iv.labs.arxiv.org/html/2305.10601
9
+ source_domain: ar5iv.labs.arxiv.org
10
+ fetched_at: '2026-06-10T00:41:12.875985Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2305.10601] Tree of Thoughts: Deliberate Problem Solving with Large Language Models
20
+ Tree of Thoughts: Deliberate Problem Solving
21
+ with Large Language Models
22
+ Shunyu Yao
23
+ Princeton University
24
+ Dian Yu
25
+ Google DeepMind
26
+ Jeffrey Zhao
27
+ Google DeepMind
28
+ Izhak Shafran
29
+ Google DeepMind
30
+ Thomas L. Griffiths
31
+ Princeton University
32
+ Yuan Cao
33
+ Google DeepMind
34
+ Karthik Narasimhan
35
+ Princeton University
36
+ Abstract
37
+ Language models are increasingly being deployed for general problem solving across a wide range of tasks, but are still confined to token-level, left-to-right decision-making processes during inference. This means they can fall short in tasks that require exploration, strategic lookahead, or where initial decisions play a pivotal role.
38
+ To surmount these challenges, we introduce a new framework for language model inference, “Tree of Thoughts” (ToT), which generalizes over the popular “Chain of Thought” approach to prompting language models, and enables exploration over coherent units of text (“thoughts”) that serve as intermediate steps toward problem solving.
39
+ ToT allows LMs to perform deliberate decision making by considering multiple different reasoning paths and self-evaluating choices to decide the next course of action, as well as looking ahead or backtracking when necessary to make global choices.
40
+ Our experiments show that ToT significantly enhances language models’ problem-solving abilities on three novel tasks requiring non-trivial planning or search: Game of 24, Creative Writing, and Mini Crosswords.
41
+ For instance, in Game of 24, while GPT-4 with chain-of-thought prompting only solved 4% of tasks, our method achieved a success rate of 74%. Code repo with all prompts:
42
+ https://github.com/princeton-nlp/tree-of-thought-llm
43
+ .
44
+ 1
45
+ Introduction
46
+ Originally designed to generate text, scaled-up versions of language models (LMs) such as GPT
47
+ [
48
+ 25
49
+ ,
50
+ 26
51
+ ,
52
+ 1
53
+ ,
54
+ 23
55
+ ]
56
+ and PaLM
57
+ [
58
+ 5
59
+ ]
60
+ have been shown to be increasingly capable of performing an ever wider range of tasks requiring mathematical, symbolic, commonsense, and knowledge reasoning. It is perhaps surprising that underlying all this progress is still the original autoregressive mechanism for generating text, which makes token-level decisions one by one and in a left-to-right fashion.
61
+ Is such a simple mechanism sufficient for a LM to be built toward a general problem solver?
62
+ If not, what problems would challenge the current paradigm, and what should be alternative mechanisms?
63
+ The literature on human cognition provides some clues to answer these questions.
64
+ Research on “dual process” models suggests that people have two modes in which they engage with decisions – a fast, automatic, unconscious mode (“System 1”) and a slow, deliberate, conscious mode (“System 2”)
65
+ [
66
+ 30
67
+ ,
68
+ 31
69
+ ,
70
+ 16
71
+ ,
72
+ 15
73
+ ]
74
+ .
75
+ These two modes have previously been connected to a variety of mathematical models used in machine learning. For example, research on reinforcement learning in humans and other animals has explored the circumstances under which they engage in associative “model free” learning or more deliberative “model based” planning
76
+ [
77
+ 7
78
+ ]
79
+ .
80
+ The simple associative token-level choices of LMs are also reminiscent of “System 1”, and thus might benefit from augmentation by a more deliberate “System 2” planning process that (1) maintains and explores diverse alternatives for current choices instead of just picking one, and (2) evaluates its current status and actively looks ahead or backtracks to make more global decisions.
81
+ To design such a planning process, we return to the origins of artificial intelligence (and cognitive science), drawing inspiration from the planning processes explored by Newell, Shaw, and Simon starting in the 1950s
82
+ [
83
+ 21
84
+ ,
85
+ 22
86
+ ]
87
+ . Newell and colleagues characterized
88
+ problem solving
89
+ [
90
+ 21
91
+ ]
92
+ as search through a combinatorial problem space, represented as a tree. We thus propose the Tree of Thoughts (ToT) framework for general problem solving with language models. As Figure
93
+ 1
94
+ illustrates, while existing methods (detailed below) sample continuous language sequences for problem solving, ToT actively maintains a tree of thoughts, where each
95
+ thought
96
+ is a coherent language sequence that serves as an intermediate step toward problem solving (Table
97
+ 1
98
+ ). Such a high-level semantic unit allows the LM to self-evaluate the progress different intermediate thoughts make towards solving the problem through a deliberate reasoning process that is also instantiated in language (Figures
99
+ 2
100
+ ,
101
+ 4
102
+ ,
103
+ 6
104
+ ). This implementation of search heuristics via LM self-evaluation and deliberation is novel, as previous search heuristics are either programmed or learned. Finally, we combine this language-based capability to generate and evaluate diverse thoughts with search algorithms, such as breadth-first search (BFS) or depth-first search (DFS), which allow systematic exploration of the tree of thoughts with lookahead and backtracking.
105
+ Empirically, we propose three new problems that challenge existing LM inference methods even with the state-of-the-art language model, GPT-4
106
+ [
107
+ 23
108
+ ]
109
+ : Game of 24, Creative Writing, and Crosswords (Table
110
+ 1
111
+ ).
112
+ These tasks require deductive, mathematical, commonsense, lexical reasoning abilities, and a way to incorporate systematic planning or search.
113
+ We show ToT obtains superior results on all three tasks by being general and flexible enough to support different levels of thoughts, different ways to generate and evaluate thoughts, and different search algorithms that adapt to the nature of different problems. We also analyze how such choices affect model performances via systematic ablations and discuss future directions to better train and use LMs.
114
+ Figure 1:
115
+ Schematic illustrating various approaches to problem solving with LLMs. Each rectangle box represents a
116
+ thought
117
+ , which is a coherent language sequence that serves as an intermediate step toward problem solving. See concrete examples of how thoughts are generated, evaluated, and searched in Figures
118
+ 2
119
+ ,
120
+ 4
121
+ ,
122
+ 6
123
+ .
124
+ 2
125
+ Background
126
+ We first formalize some existing methods that use large language models for problem-solving, which our approach is inspired by and later compared with.
127
+ We use
128
+ p
129
+ θ
130
+ subscript
131
+ 𝑝
132
+ 𝜃
133
+ p_{\theta}
134
+ to denote a pre-trained LM with parameters
135
+ θ
136
+ 𝜃
137
+ \theta
138
+ , and
139
+ lowercase letters
140
+ x
141
+ ,
142
+ y
143
+ ,
144
+ z
145
+ ,
146
+ s
147
+ ,
148
+ ⋯
149
+ 𝑥
150
+ 𝑦
151
+ 𝑧
152
+ 𝑠
153
+ ⋯
154
+ x,y,z,s,\cdots
155
+ to denote a language sequence
156
+ , i.e.
157
+ x
158
+ =
159
+ (
160
+ x
161
+ ​
162
+ [
163
+ 1
164
+ ]
165
+ ,
166
+ ⋯
167
+ ,
168
+ x
169
+ ​
170
+ [
171
+ n
172
+ ]
173
+ )
174
+ 𝑥
175
+ 𝑥
176
+ delimited-[]
177
+ 1
178
+ ⋯
179
+ 𝑥
180
+ delimited-[]
181
+ 𝑛
182
+ x=(x[1],\cdots,x[n])
183
+ where each
184
+ x
185
+ ​
186
+ [
187
+ i
188
+ ]
189
+ 𝑥
190
+ delimited-[]
191
+ 𝑖
192
+ x[i]
193
+ is a token, so that
194
+ p
195
+ θ
196
+ ​
197
+ (
198
+ x
199
+ )
200
+ =
201
+ ∏
202
+ i
203
+ =
204
+ 1
205
+ n
206
+ p
207
+ θ
208
+ ​
209
+ (
210
+ x
211
+ ​
212
+ [
213
+ i
214
+ ]
215
+ |
216
+ x
217
+ ​
218
+ [
219
+ 1
220
+ ​
221
+ …
222
+ ​
223
+ i
224
+ ]
225
+ )
226
+ subscript
227
+ 𝑝
228
+ 𝜃
229
+ 𝑥
230
+ superscript
231
+ subscript
232
+ product
233
+ 𝑖
234
+ 1
235
+ 𝑛
236
+ subscript
237
+ 𝑝
238
+ 𝜃
239
+ conditional
240
+ 𝑥
241
+ delimited-[]
242
+ 𝑖
243
+ 𝑥
244
+ delimited-[]
245
+ 1
246
+ …
247
+ 𝑖
248
+ p_{\theta}(x)=\prod_{i=1}^{n}p_{\theta}(x[i]|x[1...i])
249
+ . We use uppercase letters
250
+ S
251
+ ,
252
+ ⋯
253
+ 𝑆
254
+ ⋯
255
+ S,\cdots
256
+ to denote a collection of language sequences.
257
+ Input-output (IO) prompting
258
+ is the most common way to turn a problem input
259
+ x
260
+ 𝑥
261
+ x
262
+ into output
263
+ y
264
+ 𝑦
265
+ y
266
+ with LM:
267
+ y
268
+ ∼
269
+ p
270
+ θ
271
+ ​
272
+ (
273
+ y
274
+ |
275
+ prompt
276
+ I
277
+ ​
278
+ O
279
+ ​
280
+ (
281
+ x
282
+ )
283
+ )
284
+ similar-to
285
+ 𝑦
286
+ subscript
287
+ 𝑝
288
+ 𝜃
289
+ conditional
290
+ 𝑦
291
+ subscript
292
+ prompt
293
+ 𝐼
294
+ 𝑂
295
+ 𝑥
296
+ y\sim p_{\theta}(y|\texttt{prompt}_{{IO}}(x))
297
+ , where
298
+ prompt
299
+ I
300
+ ​
301
+ O
302
+ ​
303
+ (
304
+ x
305
+ )
306
+ subscript
307
+ prompt
308
+ 𝐼
309
+ 𝑂
310
+ 𝑥
311
+ \texttt{prompt}_{IO}(x)
312
+ wraps input
313
+ x
314
+ 𝑥
315
+ x
316
+ with task instructions and/or few-shot input-output examples. For simplicity, let us denote
317
+ p
318
+ θ
319
+ prompt
320
+ ​
321
+ (
322
+ output
323
+ ∣
324
+ input
325
+ )
326
+ =
327
+ p
328
+ θ
329
+ ​
330
+ (
331
+ output
332
+ ∣
333
+ prompt
334
+ ​
335
+ (
336
+ input
337
+ )
338
+ )
339
+ superscript
340
+ subscript
341
+ 𝑝
342
+ 𝜃
343
+ prompt
344
+ conditional
345
+ output
346
+ input
347
+ subscript
348
+ 𝑝
349
+ 𝜃
350
+ conditional
351
+ output
352
+ prompt
353
+ input
354
+ p_{\theta}^{{\rm prompt}}(\texttt{output}\mid\texttt{input})=p_{\theta}(\texttt{output}\mid\texttt{prompt}(\texttt{input}))
355
+ , so that IO prompting can be formulated as
356
+ y
357
+ ∼
358
+ p
359
+ θ
360
+ I
361
+ ​
362
+ O
363
+ ​
364
+ (
365
+ y
366
+ |
367
+ x
368
+ )
369
+ similar-to
370
+ 𝑦
371
+ superscript
372
+ subscript
373
+ 𝑝
374
+ 𝜃
375
+ 𝐼
376
+ 𝑂
377
+ conditional
378
+ 𝑦
379
+ 𝑥
380
+ y\sim p_{\theta}^{IO}(y|x)
381
+ .
382
+ Chain-of-thought (CoT) prompting
383
+ [
384
+ 38
385
+ ]
386
+ was proposed to address cases where the mapping of input
387
+ x
388
+ 𝑥
389
+ x
390
+ to output
391
+ y
392
+ 𝑦
393
+ y
394
+ is non-trivial (e.g. when
395
+ x
396
+ 𝑥
397
+ x
398
+ is a math question and
399
+ y
400
+ 𝑦
401
+ y
402
+ is the final numerical answer). The key idea is to introduce a chain of
403
+ thoughts
404
+ z
405
+ 1
406
+ ,
407
+ ⋯
408
+ ,
409
+ z
410
+ n
411
+ subscript
412
+ 𝑧
413
+ 1
414
+ ⋯
415
+ subscript
416
+ 𝑧
417
+ 𝑛
418
+ z_{1},\cdots,z_{n}
419
+ to bridge
420
+ x
421
+ 𝑥
422
+ x
423
+ and
424
+ y
425
+ 𝑦
426
+ y
427
+ , where each
428
+ z
429
+ i
430
+ subscript
431
+ 𝑧
432
+ 𝑖
433
+ z_{i}
434
+ is a coherent language sequence that serves as a meaningful intermediate step toward problem solving (e.g.
435
+ z
436
+ i
437
+ subscript
438
+ 𝑧
439
+ 𝑖
440
+ z_{i}
441
+ could be an intermediate equation for math QA). To solve problems with CoT, each thought
442
+ z
443
+ i
444
+ ∼
445
+ p
446
+ θ
447
+ C
448
+ ​
449
+ o
450
+ ​
451
+ T
452
+ ​
453
+ (
454
+ z
455
+ i
456
+ ∣
457
+ x
458
+ ,
459
+ z
460
+ 1
461
+ ​
462
+ ⋯
463
+ ​
464
+ i
465
+ −
466
+ 1
467
+ )
468
+ similar-to
469
+ subscript
470
+ 𝑧
471
+ 𝑖
472
+ superscript
473
+ subscript
474
+ 𝑝
475
+ 𝜃
476
+ 𝐶
477
+ 𝑜
478
+ 𝑇
479
+ conditional
480
+ subscript
481
+ 𝑧
482
+ 𝑖
483
+ 𝑥
484
+ subscript
485
+ 𝑧
486
+ 1
487
+ ⋯
488
+ 𝑖
489
+ 1
490
+ z_{i}\sim p_{\theta}^{CoT}(z_{i}\mid x,z_{1\cdots i-1})
491
+ is sampled sequentially, then the output
492
+ y
493
+ ∼
494
+ p
495
+ θ
496
+ C
497
+ ​
498
+ o
499
+ ​
500
+ T
501
+ ​
502
+ (
503
+ y
504
+ |
505
+ x
506
+ ,
507
+ z
508
+ 1
509
+ ​
510
+ ⋯
511
+ ​
512
+ n
513
+ )
514
+ similar-to
515
+ 𝑦
516
+ superscript
517
+ subscript
518
+ 𝑝
519
+ 𝜃
520
+ 𝐶
521
+ 𝑜
522
+ 𝑇
523
+ conditional
524
+ 𝑦
525
+ 𝑥
526
+ subscript
527
+ 𝑧
528
+ 1
529
+ ⋯
530
+ 𝑛
531
+ y\sim p_{\theta}^{CoT}(y|x,z_{1\cdots n})
532
+ . In practice,
533
+ [
534
+ z
535
+ 1
536
+ ​
537
+ ⋯
538
+ ​
539
+ n
540
+ ,
541
+ y
542
+ ]
543
+ ��
544
+ p
545
+ θ
546
+ C
547
+ ​
548
+ o
549
+ ​
550
+ T
551
+ ​
552
+ (
553
+ z
554
+ 1
555
+ ​
556
+ ⋯
557
+ ​
558
+ n
559
+ ,
560
+ y
561
+ |
562
+ x
563
+ )
564
+ similar-to
565
+ subscript
566
+ 𝑧
567
+ 1
568
+ ⋯
569
+ 𝑛
570
+ 𝑦
571
+ superscript
572
+ subscript
573
+ 𝑝
574
+ 𝜃
575
+ 𝐶
576
+ 𝑜
577
+ 𝑇
578
+ subscript
579
+ 𝑧
580
+ 1
581
+ ⋯
582
+ 𝑛
583
+ conditional
584
+ 𝑦
585
+ 𝑥
586
+ [z_{1\cdots n},y]\sim p_{\theta}^{CoT}(z_{1\cdots n},y|x)
587
+ is sampled as a continuous language sequence, and the
588
+ decomposition
589
+ of thoughts (e.g. is each
590
+ z
591
+ i
592
+ subscript
593
+ 𝑧
594
+ 𝑖
595
+ z_{i}
596
+ a phrase, a sentence, or a paragraph) is left ambiguous.
597
+ Self-consistency with CoT (CoT-SC)
598
+ [
599
+ 36
600
+ ]
601
+ is an ensemble approach that samples
602
+ k
603
+ 𝑘
604
+ k
605
+ i.i.d. chains of thought:
606
+ [
607
+ z
608
+ 1
609
+ ​
610
+ ⋯
611
+ ​
612
+ n
613
+ (
614
+ i
615
+ )
616
+ ,
617
+ y
618
+ (
619
+ i
620
+ )
621
+ ]
622
+ ∼
623
+ p
624
+ θ
625
+ C
626
+ ​
627
+ o
628
+ ​
629
+ T
630
+ ​
631
+ (
632
+ z
633
+ 1
634
+ ​
635
+ ⋯
636
+ ​
637
+ n
638
+ ,
639
+ y
640
+ |
641
+ x
642
+ )
643
+ ​
644
+ (
645
+ i
646
+ =
647
+ 1
648
+ ​
649
+ ⋯
650
+ ​
651
+ k
652
+ )
653
+ similar-to
654
+ subscript
655
+ superscript
656
+ 𝑧
657
+ 𝑖
658
+ 1
659
+ ⋯
660
+ 𝑛
661
+ superscript
662
+ 𝑦
663
+ 𝑖
664
+ superscript
665
+ subscript
666
+ 𝑝
667
+ 𝜃
668
+ 𝐶
669
+ 𝑜
670
+ 𝑇
671
+ subscript
672
+ 𝑧
673
+ 1
674
+ ⋯
675
+ 𝑛
676
+ conditional
677
+ 𝑦
678
+ 𝑥
679
+ 𝑖
680
+ 1
681
+ ⋯
682
+ 𝑘
683
+ [z^{(i)}_{1\cdots n},y^{(i)}]\sim p_{\theta}^{CoT}(z_{1\cdots n},y|x)\ (i=1\cdots k)
684
+ , then returns the most frequent output:
685
+ arg
686
+ ⁡
687
+ max
688
+ y
689
+ ⁡
690
+ #
691
+ ​
692
+ {
693
+ i
694
+ ∣
695
+ y
696
+ (
697
+ i
698
+ )
699
+ =
700
+ y
701
+ }
702
+ subscript
703
+ 𝑦
704
+ #
705
+ conditional-set
706
+ 𝑖
707
+ superscript
708
+ 𝑦
709
+ 𝑖
710
+ 𝑦
711
+ \arg\max_{y}\#\{i\mid y^{(i)}=y\}
712
+ . CoT-SC improves upon CoT, because there are generally different thought processes for the same problem (e.g. different ways to prove the same theorem), and the output decision can be more faithful by exploring a richer set of thoughts. However, within each chain there is no local exploration of different thought steps, and the “most frequent” heuristic only applies when the output space is limited (e.g. multi-choice QA).
713
+ 3
714
+ Tree of Thoughts: Deliberate Problem Solving with LM
715
+ A genuine problem-solving process involves the repeated use of available information to initiate exploration, which discloses, in turn, more information until a way to attain the solution is finally discovered.——
716
+ Newell et al. [
717
+ 21
718
+ ]
719
+ Research on human problem-solving suggests that people search through a combinatorial problem-space – a tree where the nodes represent partial solutions, and the branches correspond to operators that modify them
720
+ [
721
+ 21
722
+ ,
723
+ 22
724
+ ]
725
+ . Which branch to take is determined by heuristics that help to navigate the problem-space and guide the problem-solver towards a solution. This perspective highlights two key shortcomings of existing approaches that use LMs to solve general problems: 1) Locally, they do not explore
726
+ different
727
+ continuations within a thought process – the branches of the tree. 2) Globally, they do not incorporate any type of planning, lookahead, or backtracking to help evaluate these different options – the kind of heuristic-guided search that seems characteristic of human problem-solving.
728
+ To address these shortcomings, we introduce
729
+ Tree of Thoughts (ToT)
730
+ , a paradigm that allows LMs to explore multiple reasoning paths over thoughts (Figure
731
+ 1
732
+ (c)). ToT frames any problem as a search over a tree, where each node is a
733
+ state
734
+ s
735
+ =
736
+ [
737
+ x
738
+ ,
739
+ z
740
+ 1
741
+ ​
742
+ ⋯
743
+ ​
744
+ i
745
+ ]
746
+ 𝑠
747
+ 𝑥
748
+ subscript
749
+ 𝑧
750
+ 1
751
+ ⋯
752
+ 𝑖
753
+ s=[x,z_{1\cdots i}]
754
+ representing a partial solution with the input and the sequence of thoughts so far. A specific instantiation of ToT involves answering four questions: 1. How to
755
+ decompose
756
+ the intermediate process into thought steps; 2. How to
757
+ generate
758
+ potential thoughts from each state; 3. How to heuristically
759
+ evaluate
760
+ states; 4. What
761
+ search
762
+ algorithm to use.
763
+ 1. Thought decomposition.
764
+ While CoT samples thoughts coherently without explicit decomposition, ToT leverages problem properties to design and decompose intermediate thought steps. As Table
765
+ 1
766
+ shows, depending on different problems, a thought could be a couple of words (Crosswords), a line of equation (Game of 24), or a whole paragraph of writing plan (Creative Writing). In general, a thought should be “small” enough so that LMs can generate promising and diverse samples (e.g. generating a whole book is usually too “big” to be coherent), yet “big” enough so that LMs can evaluate its prospect toward problem solving (e.g. generating one token is usually too “small” to evaluate).
767
+ 2. Thought generator
768
+ G
769
+ ​
770
+ (
771
+ p
772
+ θ
773
+ ,
774
+ s
775
+ ,
776
+ k
777
+ )
778
+ 𝐺
779
+ subscript
780
+ 𝑝
781
+ 𝜃
782
+ 𝑠
783
+ 𝑘
784
+ G(p_{\theta},s,k)
785
+ .
786
+ Given a tree state
787
+ s
788
+ =
789
+ [
790
+ x
791
+ ,
792
+ z
793
+ 1
794
+ ​
795
+ ⋯
796
+ ​
797
+ i
798
+ ]
799
+ 𝑠
800
+ 𝑥
801
+ subscript
802
+ 𝑧
803
+ 1
804
+ ⋯
805
+ 𝑖
806
+ s=[x,z_{1\cdots i}]
807
+ , we consider two strategies to generate
808
+ k
809
+ 𝑘
810
+ k
811
+ candidates for the next thought step:
812
+ (a)
813
+ Sample
814
+ i.i.d. thoughts from a CoT prompt (Creative Writing, Figure
815
+ 4
816
+ ):
817
+ z
818
+ (
819
+ j
820
+ )
821
+ ∼
822
+ p
823
+ θ
824
+ C
825
+ ​
826
+ o
827
+ ​
828
+ T
829
+ ​
830
+ (
831
+ z
832
+ i
833
+ +
834
+ 1
835
+ |
836
+ s
837
+ )
838
+ =
839
+ p
840
+ θ
841
+ C
842
+ ​
843
+ o
844
+ ​
845
+ T
846
+ ​
847
+ (
848
+ z
849
+ i
850
+ +
851
+ 1
852
+ |
853
+ x
854
+ ,
855
+ z
856
+ 1
857
+ ​
858
+ ⋯
859
+ ​
860
+ i
861
+ )
862
+ ​
863
+ (
864
+ j
865
+ =
866
+ 1
867
+ ​
868
+ ⋯
869
+ ​
870
+ k
871
+ )
872
+ similar-to
873
+ superscript
874
+ 𝑧
875
+ 𝑗
876
+ superscript
877
+ subscript
878
+ 𝑝
879
+ 𝜃
880
+ 𝐶
881
+ 𝑜
882
+ 𝑇
883
+ conditional
884
+ subscript
885
+ 𝑧
886
+ 𝑖
887
+ 1
888
+ 𝑠
889
+ superscript
890
+ subscript
891
+ 𝑝
892
+ 𝜃
893
+ 𝐶
894
+ 𝑜
895
+ 𝑇
896
+ conditional
897
+ subscript
898
+ 𝑧
899
+ 𝑖
900
+ 1
901
+ 𝑥
902
+ subscript
903
+ 𝑧
904
+ 1
905
+ ⋯
906
+ 𝑖
907
+ 𝑗
908
+ 1
909
+ ⋯
910
+ 𝑘
911
+ z^{(j)}\sim p_{\theta}^{CoT}(z_{i+1}|s)=p_{\theta}^{CoT}(z_{i+1}|x,z_{1\cdots i})\ (j=1\cdots k)
912
+ . This works better when the thought space is rich (e.g. each thought is a paragraph), and i.i.d. samples lead to diversity;
913
+ (b)
914
+ Propose
915
+ thoughts sequentially using a “propose prompt” (Game of 24, Figure
916
+ 2
917
+ ; Crosswords, Figure
918
+ 6
919
+ ):
920
+ [
921
+ z
922
+ (
923
+ 1
924
+ )
925
+ ,
926
+ ⋯
927
+ ,
928
+ z
929
+ (
930
+ k
931
+ )
932
+ ]
933
+ ∼
934
+ p
935
+ θ
936
+ p
937
+ ​
938
+ r
939
+ ​
940
+ o
941
+ ​
942
+ p
943
+ ​
944
+ o
945
+ ​
946
+ s
947
+ ​
948
+ e
949
+ ​
950
+ (
951
+ z
952
+ i
953
+ +
954
+ 1
955
+ (
956
+ 1
957
+ ​
958
+ ⋯
959
+ ​
960
+ k
961
+ )
962
+ ∣
963
+ s
964
+ )
965
+ similar-to
966
+ superscript
967
+ 𝑧
968
+ 1
969
+ ⋯
970
+ superscript
971
+ 𝑧
972
+ 𝑘
973
+ superscript
974
+ subscript
975
+ 𝑝
976
+ 𝜃
977
+ 𝑝
978
+ 𝑟
979
+ 𝑜
980
+ 𝑝
981
+ 𝑜
982
+ 𝑠
983
+ 𝑒
984
+ conditional
985
+ superscript
986
+ subscript
987
+ 𝑧
988
+ 𝑖
989
+ 1
990
+ 1
991
+ ⋯
992
+ 𝑘
993
+ 𝑠
994
+ [z^{(1)},\cdots,z^{(k)}]\sim p_{\theta}^{propose}(z_{i+1}^{(1\cdots k)}\mid s)
995
+ . This works better when the thought space is more constrained (e.g. each thought is just a word or a line), so proposing different thoughts in the same context avoids duplication.
996
+ 3. State evaluator
997
+ V
998
+ ​
999
+ (
1000
+ p
1001
+ θ
1002
+ ,
1003
+ S
1004
+ )
1005
+ 𝑉
1006
+ subscript
1007
+ 𝑝
1008
+ 𝜃
1009
+ 𝑆
1010
+ V(p_{\theta},S)
1011
+ .
1012
+ Given a frontier of different states, the state evaluator evaluates the progress they make towards solving the problem, serving as a
1013
+ heuristic
1014
+ for the search algorithm to determine which states to keep exploring and in which order. While heuristics are a standard approach to solving search problems, they are typically either programmed (e.g. DeepBlue
1015
+ [
1016
+ 3
1017
+ ]
1018
+ ) or learned (e.g. AlphaGo
1019
+ [
1020
+ 29
1021
+ ]
1022
+ ). We propose a third alternative, by using the LM to deliberately reason about states. When applicable, such a deliberate heuristic can be more flexible than programmed rules, and more sample-efficient than learned models.
1023
+ Similar to the thought generator, we consider two strategies to evaluate states either independently or together:
1024
+ (a)
1025
+ Value
1026
+ each state independently:
1027
+ V
1028
+ ​
1029
+ (
1030
+ p
1031
+ θ
1032
+ ,
1033
+ S
1034
+ )
1035
+ ​
1036
+ (
1037
+ s
1038
+ )
1039
+ ∼
1040
+ p
1041
+ θ
1042
+ v
1043
+ ​
1044
+ a
1045
+ ​
1046
+ l
1047
+ ​
1048
+ u
1049
+ ​
1050
+ e
1051
+ ​
1052
+ (
1053
+ v
1054
+ |
1055
+ s
1056
+ )
1057
+ ​
1058
+ ∀
1059
+ s
1060
+ ∈
1061
+ S
1062
+ similar-to
1063
+ 𝑉
1064
+ subscript
1065
+ 𝑝
1066
+ 𝜃
1067
+ 𝑆
1068
+ 𝑠
1069
+ superscript
1070
+ subscript
1071
+ 𝑝
1072
+ 𝜃
1073
+ 𝑣
1074
+ 𝑎
1075
+ 𝑙
1076
+ 𝑢
1077
+ 𝑒
1078
+ conditional
1079
+ 𝑣
1080
+ 𝑠
1081
+ for-all
1082
+ 𝑠
1083
+ 𝑆
1084
+ V(p_{\theta},S)(s)\sim p_{\theta}^{value}(v|s)\ \forall s\in S
1085
+ , where a value prompt reasons about the state
1086
+ s
1087
+ 𝑠
1088
+ s
1089
+ to generate a scalar value
1090
+ v
1091
+ 𝑣
1092
+ v
1093
+ (e.g. 1-10) or a classification (e.g. sure/likely/impossible) that could be heuristically turned into a value. The basis of such evaluative reasoning can vary across problems and thought steps. In this work, we explore evaluation via few
1094
+ lookahead
1095
+ simulations (e.g. quickly confirm that 5, 5, 14 can reach 24 via 5 + 5 + 14, or “hot_l” can mean “inn” via filling “e” in “_”) plus commonsense (e.g. 1 2 3 are too small to reach 24, or no word can start with “tzxc”). While the former might promote “good” states, the latter could help eliminate “bad” states. Such valuations do not need to be perfect, and only need to be approximately helpful for decision making.
1096
+ (b)
1097
+ Vote
1098
+ across states:
1099
+ V
1100
+ ​
1101
+ (
1102
+ p
1103
+ θ
1104
+ ,
1105
+ S
1106
+ )
1107
+ ​
1108
+ (
1109
+ s
1110
+ )
1111
+ =
1112
+ 𝟙
1113
+ ​
1114
+ [
1115
+ s
1116
+ =
1117
+ s
1118
+ ∗
1119
+ ]
1120
+ 𝑉
1121
+ subscript
1122
+ 𝑝
1123
+ 𝜃
1124
+ 𝑆
1125
+ 𝑠
1126
+ 1
1127
+ delimited-[]
1128
+ 𝑠
1129
+ superscript
1130
+ 𝑠
1131
+ V(p_{\theta},S)(s)=\mathds{1}[s=s^{*}]
1132
+ , where a “good” state
1133
+ s
1134
+ ∗
1135
+ ∼
1136
+ p
1137
+ θ
1138
+ v
1139
+ ​
1140
+ o
1141
+ ​
1142
+ t
1143
+ ​
1144
+ e
1145
+ ​
1146
+ (
1147
+ s
1148
+ ∗
1149
+ |
1150
+ S
1151
+ )
1152
+ similar-to
1153
+ superscript
1154
+ 𝑠
1155
+ superscript
1156
+ subscript
1157
+ 𝑝
1158
+ 𝜃
1159
+ 𝑣
1160
+ 𝑜
1161
+ 𝑡
1162
+ 𝑒
1163
+ conditional
1164
+ superscript
1165
+ 𝑠
1166
+ 𝑆
1167
+ s^{*}\sim p_{\theta}^{vote}(s^{*}|S)
1168
+ is voted out based on deliberately comparing different states in
1169
+ S
1170
+ 𝑆
1171
+ S
1172
+ in a vote prompt.
1173
+ When problem success is harder to directly value (e.g. passage coherency), it is natural to to instead compare different partial solutions and vote for the most promising one. This is similar in spirit to a “step-wise” self-consistency strategy, i.e. cast “which state to explore” as a multi-choice QA, and use LM samples to vote for it.
1174
+ For both strategies, we could prompt the LM multiple times to aggregate the value or vote results to trade time/resource/cost for more faithful/robust heuristics.
1175
+ Algorithm 1
1176
+ ToT-BFS(
1177
+ x
1178
+ ,
1179
+ p
1180
+ θ
1181
+ ,
1182
+ G
1183
+ ,
1184
+ k
1185
+ ,
1186
+ V
1187
+ ,
1188
+ T
1189
+ ,
1190
+ b
1191
+ 𝑥
1192
+ subscript
1193
+ 𝑝
1194
+ 𝜃
1195
+ 𝐺
1196
+ 𝑘
1197
+ 𝑉
1198
+ 𝑇
1199
+ 𝑏
1200
+ x,p_{\theta},G,k,V,T,b
1201
+ )
1202
+ Input
1203
+ x
1204
+ 𝑥
1205
+ x
1206
+ , LM
1207
+ p
1208
+ θ
1209
+ subscript
1210
+ 𝑝
1211
+ 𝜃
1212
+ p_{\theta}
1213
+ , thought generator
1214
+ G
1215
+ ​
1216
+ (
1217
+ )
1218
+ 𝐺
1219
+ G()
1220
+ & size limit
1221
+ k
1222
+ 𝑘
1223
+ k
1224
+ , states evaluator
1225
+ V
1226
+ ​
1227
+ (
1228
+ )
1229
+ 𝑉
1230
+ V()
1231
+ , step limit
1232
+ T
1233
+ 𝑇
1234
+ T
1235
+ , breadth limit
1236
+ b
1237
+ 𝑏
1238
+ b
1239
+ .
1240
+ S
1241
+ 0
1242
+ ←
1243
+ {
1244
+ x
1245
+ }
1246
+ ←
1247
+ subscript
1248
+ 𝑆
1249
+ 0
1250
+ 𝑥
1251
+ S_{0}\leftarrow\{x\}
1252
+ for
1253
+ t
1254
+ =
1255
+ 1
1256
+ ,
1257
+ ⋯
1258
+ ,
1259
+ T
1260
+ 𝑡
1261
+ 1
1262
+ ⋯
1263
+ 𝑇
1264
+ t=1,\cdots,T
1265
+ do
1266
+ S
1267
+ t
1268
+ ′
1269
+ ←
1270
+ {
1271
+ [
1272
+ s
1273
+ ,
1274
+ z
1275
+ ]
1276
+ ∣
1277
+ s
1278
+ ∈
1279
+ S
1280
+ t
1281
+ −
1282
+ 1
1283
+ ,
1284
+ z
1285
+ t
1286
+ ∈
1287
+ G
1288
+ ​
1289
+ (
1290
+ p
1291
+ θ
1292
+ ,
1293
+ s
1294
+ ,
1295
+ k
1296
+ )
1297
+ }
1298
+ ←
1299
+ subscript
1300
+ superscript
1301
+ 𝑆
1302
+ ′
1303
+ 𝑡
1304
+ conditional-set
1305
+ 𝑠
1306
+ 𝑧
1307
+ formulae-sequence
1308
+ 𝑠
1309
+ subscript
1310
+ 𝑆
1311
+ 𝑡
1312
+ 1
1313
+ subscript
1314
+ 𝑧
1315
+ 𝑡
1316
+ G
1317
+ subscript
1318
+ 𝑝
1319
+ 𝜃
1320
+ 𝑠
1321
+ 𝑘
1322
+ S^{\prime}_{t}\leftarrow\{[s,z]\mid s\in S_{t-1},z_{t}\in{\color[rgb]{0,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,0}\pgfsys@color@gray@stroke{0}\pgfsys@color@gray@fill{0}\mathrm{G}}(p_{\theta},s,k)\}
1323
+ V
1324
+ t
1325
+ ←
1326
+ V
1327
+ ​
1328
+ (
1329
+ p
1330
+ θ
1331
+ ,
1332
+ S
1333
+ t
1334
+ ′
1335
+ )
1336
+ ←
1337
+ subscript
1338
+ 𝑉
1339
+ 𝑡
1340
+ 𝑉
1341
+ subscript
1342
+ 𝑝
1343
+ 𝜃
1344
+ subscript
1345
+ superscript
1346
+ 𝑆
1347
+ ′
1348
+ 𝑡
1349
+ V_{t}\leftarrow V(p_{\theta},S^{\prime}_{t})
1350
+ S
1351
+ t
1352
+ ←
1353
+ arg
1354
+ ⁡
1355
+ max
1356
+ S
1357
+ ⊂
1358
+ S
1359
+ t
1360
+ ′
1361
+ ,
1362
+ |
1363
+ S
1364
+ |
1365
+ =
1366
+ b
1367
+ ​
1368
+ ∑
1369
+ s
1370
+ ∈
1371
+ S
1372
+ V
1373
+ t
1374
+ ​
1375
+ (
1376
+ s
1377
+ )
1378
+ ←
1379
+ subscript
1380
+ 𝑆
1381
+ 𝑡
1382
+ subscript
1383
+ formulae-sequence
1384
+ 𝑆
1385
+ subscript
1386
+ superscript
1387
+ 𝑆
1388
+ ′
1389
+ 𝑡
1390
+ 𝑆
1391
+ 𝑏
1392
+ subscript
1393
+ 𝑠
1394
+ 𝑆
1395
+ subscript
1396
+ 𝑉
1397
+ 𝑡
1398
+ 𝑠
1399
+ S_{t}\leftarrow\arg\max_{S\subset S^{\prime}_{t},|S|=b}\sum_{s\in S}V_{t}(s)
1400
+ end
1401
+ for
1402
+ return
1403
+ G
1404
+ ​
1405
+ (
1406
+ p
1407
+ θ
1408
+ ,
1409
+ arg
1410
+ ⁡
1411
+ max
1412
+ s
1413
+ ∈
1414
+ S
1415
+ T
1416
+ ⁡
1417
+ V
1418
+ T
1419
+ ​
1420
+ (
1421
+ s
1422
+ )
1423
+ ,
1424
+ 1
1425
+ )
1426
+ 𝐺
1427
+ subscript
1428
+ 𝑝
1429
+ 𝜃
1430
+ subscript
1431
+ 𝑠
1432
+ subscript
1433
+ 𝑆
1434
+ 𝑇
1435
+ subscript
1436
+ 𝑉
1437
+ 𝑇
1438
+ 𝑠
1439
+ 1
1440
+ G(p_{\theta},\arg\max_{s\in S_{T}}V_{T}(s),1)
1441
+ Algorithm 2
1442
+ ToT-DFS(
1443
+ s
1444
+ ,
1445
+ t
1446
+ ,
1447
+ p
1448
+ θ
1449
+ ,
1450
+ G
1451
+ ,
1452
+ k
1453
+ ,
1454
+ V
1455
+ ,
1456
+ T
1457
+ ,
1458
+ v
1459
+ t
1460
+ ​
1461
+ h
1462
+ 𝑠
1463
+ 𝑡
1464
+ subscript
1465
+ 𝑝
1466
+ 𝜃
1467
+ 𝐺
1468
+ 𝑘
1469
+ 𝑉
1470
+ 𝑇
1471
+ subscript
1472
+ 𝑣
1473
+ 𝑡
1474
+ ℎ
1475
+ s,t,p_{\theta},G,k,V,T,v_{\small th}
1476
+ )
1477
+ Current state
1478
+ s
1479
+ 𝑠
1480
+ s
1481
+ , step
1482
+ t
1483
+ 𝑡
1484
+ t
1485
+ , LM
1486
+ p
1487
+ θ
1488
+ subscript
1489
+ 𝑝
1490
+ 𝜃
1491
+ p_{\theta}
1492
+ , thought generator
1493
+ G
1494
+ ​
1495
+ (
1496
+ )
1497
+ 𝐺
1498
+ G()
1499
+ and size limit
1500
+ k
1501
+ 𝑘
1502
+ k
1503
+ , states evaluator
1504
+ V
1505
+ ​
1506
+ (
1507
+ )
1508
+ 𝑉
1509
+ V()
1510
+ , step limit
1511
+ T
1512
+ 𝑇
1513
+ T
1514
+ , threshold
1515
+ v
1516
+ t
1517
+ ​
1518
+ h
1519
+ subscript
1520
+ 𝑣
1521
+ 𝑡
1522
+ ℎ
1523
+ v_{\small th}
1524
+ if
1525
+ t
1526
+ >
1527
+ T
1528
+ 𝑡
1529
+ 𝑇
1530
+ t>T
1531
+ then
1532
+ record output
1533
+ G
1534
+ ​
1535
+ (
1536
+ p
1537
+ θ
1538
+ ,
1539
+ s
1540
+ ,
1541
+ 1
1542
+ )
1543
+ 𝐺
1544
+ subscript
1545
+ 𝑝
1546
+ 𝜃
1547
+ 𝑠
1548
+ 1
1549
+ G(p_{\theta},s,1)
1550
+ end
1551
+ if
1552
+ for
1553
+ s
1554
+ ′
1555
+ ∈
1556
+ G
1557
+ ​
1558
+ (
1559
+ p
1560
+ θ
1561
+ ,
1562
+ s
1563
+ ,
1564
+ k
1565
+ )
1566
+ superscript
1567
+ 𝑠
1568
+ ′
1569
+ 𝐺
1570
+ subscript
1571
+ 𝑝
1572
+ 𝜃
1573
+ 𝑠
1574
+ 𝑘
1575
+ s^{\prime}\in G(p_{\theta},s,k)
1576
+ do
1577
+ ▷
1578
+ ▷
1579
+ \triangleright
1580
+ sorted candidates
1581
+ if
1582
+ V
1583
+ ​
1584
+ (
1585
+ p
1586
+ θ
1587
+ ,
1588
+ {
1589
+ s
1590
+ ′
1591
+ }
1592
+ )
1593
+ ​
1594
+ (
1595
+ s
1596
+ )
1597
+ >
1598
+ v
1599
+ t
1600
+ ​
1601
+ h
1602
+ ​
1603
+ r
1604
+ ​
1605
+ e
1606
+ ​
1607
+ s
1608
+ 𝑉
1609
+ subscript
1610
+ 𝑝
1611
+ 𝜃
1612
+ superscript
1613
+ 𝑠
1614
+ ′
1615
+ 𝑠
1616
+ subscript
1617
+ 𝑣
1618
+ 𝑡
1619
+ ℎ
1620
+ 𝑟
1621
+ 𝑒
1622
+ 𝑠
1623
+ V(p_{\theta},\{s^{\prime}\})(s)>v_{\small thres}
1624
+ then
1625
+ ▷
1626
+ ▷
1627
+ \triangleright
1628
+ pruning
1629
+ DFS
1630
+ (
1631
+ s
1632
+ ′
1633
+ ,
1634
+ t
1635
+ +
1636
+ 1
1637
+ )
1638
+ superscript
1639
+ 𝑠
1640
+ ′
1641
+ 𝑡
1642
+ 1
1643
+ (s^{\prime},t+1)
1644
+ end
1645
+ if
1646
+ end
1647
+ for
1648
+ 4. Search algorithm.
1649
+ Finally, within the ToT framework, one can plug and play different search algorithms depending on the tree structure. We explore two relatively simple search algorithms and leave more advanced ones (e.g. A*
1650
+ [
1651
+ 11
1652
+ ]
1653
+ , MCTS
1654
+ [
1655
+ 2
1656
+ ]
1657
+ ) for future work:
1658
+ (a)
1659
+ Breadth-first search (BFS)
1660
+ (Algorithm
1661
+ 1
1662
+ ) maintains a set of the
1663
+ b
1664
+ 𝑏
1665
+ b
1666
+ most promising states per step. This is used for Game of 24 and Creative Writing where the tree depth is limit (
1667
+ T
1668
+ ≤
1669
+ 3
1670
+ 𝑇
1671
+ 3
1672
+ T\leq 3
1673
+ ), and initial thought steps can be evaluated and pruned to a small set (
1674
+ b
1675
+ ≤
1676
+ 5
1677
+ 𝑏
1678
+ 5
1679
+ b\leq 5
1680
+ ).
1681
+ (b)
1682
+ Depth-first search (DFS)
1683
+ (Algorithm
1684
+ 2
1685
+ ) explores the most promising state first, until the final output is reached (
1686
+ t
1687
+ >
1688
+ T
1689
+ 𝑡
1690
+ 𝑇
1691
+ t>T
1692
+ ), or the state evaluator deems it impossible to solve the problem from the current
1693
+ s
1694
+ 𝑠
1695
+ s
1696
+ (
1697
+ V
1698
+ ​
1699
+ (
1700
+ p
1701
+ θ
1702
+ ,
1703
+ {
1704
+ s
1705
+ }
1706
+ )
1707
+ ​
1708
+ (
1709
+ s
1710
+ )
1711
+ ≤
1712
+ v
1713
+ t
1714
+ ​
1715
+ h
1716
+ 𝑉
1717
+ subscript
1718
+ 𝑝
1719
+ 𝜃
1720
+ 𝑠
1721
+ 𝑠
1722
+ subscript
1723
+ 𝑣
1724
+ 𝑡
1725
+ ℎ
1726
+ V(p_{\theta},\{s\})(s)\leq v_{th}
1727
+ for a value threshold
1728
+ v
1729
+ t
1730
+ ​
1731
+ h
1732
+ subscript
1733
+ 𝑣
1734
+ 𝑡
1735
+ ℎ
1736
+ v_{th}
1737
+ ). In the latter case, the subtree from
1738
+ s
1739
+ 𝑠
1740
+ s
1741
+ is
1742
+ pruned
1743
+ to trade exploration for exploitation. In both cases, DFS
1744
+ backtracks
1745
+ to the parent state of
1746
+ s
1747
+ 𝑠
1748
+ s
1749
+ to continue exploration.
1750
+ Conceptually, ToT has several benefits as a method for general problem-solving with LMs: (1)
1751
+ Generality.
1752
+ IO, CoT, CoT-SC, and self-refinement can be seen as special cases of ToT (i.e. trees of limited depth and breadth; Figure
1753
+ 1
1754
+ ). (2)
1755
+ Modularity.
1756
+ The base LM, as well as the thought decomposition, generation, evaluation, and search procedures can all be varied independently. (3)
1757
+ Adaptability
1758
+ . Different problem properties, LM capabilities, and resource constraints can be accommodated. (4)
1759
+ Convenience.
1760
+ No extra training is needed, just a pre-trained LM is sufficient. The next section will show how these conceptual benefits translate to strong empirical performance in different problems.
1761
+ 4
1762
+ Experiments
1763
+ Game of 24
1764
+ Creative Writing
1765
+ 5x5 Crosswords
1766
+ Input
1767
+ 4 numbers
1768
+ (4 9 10 13)
1769
+ 4 random sentences
1770
+ 10 clues
1771
+ (h1. presented;..)
1772
+ Output
1773
+ An equation to reach 24
1774
+ (13-9)*(10-4)=24
1775
+ A passage of 4 paragraphs ending in the 4 sentences
1776
+ 5x5 letters:
1777
+ SHOWN; WIRRA; AVAIL; …
1778
+ Thoughts
1779
+ 3 intermediate equations
1780
+ (13-9=4 (left 4,4,10); 10-4=6 (left 4,6); 4*6=24)
1781
+ A short writing plan
1782
+ (1. Introduce a book that connects…)
1783
+ Words to fill in for clues:
1784
+ (h1. shown; v5. naled; …)
1785
+ #ToT steps
1786
+ 3
1787
+ 1
1788
+ 5-10 (variable)
1789
+ Table 1:
1790
+ Task overview. Input, output, thought examples are in blue.
1791
+ We propose three tasks that are hard even when sampling from the state-of-the-art language model, GPT-4
1792
+ [
1793
+ 23
1794
+ ]
1795
+ , using standard IO prompting or chain-of-thought (CoT) prompting. We show how deliberate search in trees of thoughts (ToT) produces better results, and more importantly, interesting and promising new ways to use language models to solve problems requiring search or planning.
1796
+ Unless otherwise stated, we perform experiments using a Chat Completion mode GPT-4
1797
+ 1
1798
+ 1
1799
+ 1
1800
+ Experiments were done between May 5-16, 2023.
1801
+ with a sampling temperature of 0.7.
1802
+ 4.1
1803
+ Game of 24
1804
+ Game of 24 is a mathematical reasoning challenge, where the goal is to use 4 numbers and basic arithmetic operations (+-*/) to obtain 24.
1805
+ For example, given input “4 9 10 13”, a solution output could be “(10 - 4) * (13 - 9) = 24”.
1806
+ Figure 2:
1807
+ ToT in a game of 24. The LM is prompted for (a) thought generation and (b) valuation.
1808
+ Method
1809
+ Success
1810
+ IO prompt
1811
+ 7.3%
1812
+ CoT prompt
1813
+ 4.0%
1814
+ CoT-SC
1815
+ (k=100)
1816
+ 9.0%
1817
+ ToT (ours)
1818
+ (b=1)
1819
+ 45%
1820
+ ToT (ours)
1821
+ (b=5)
1822
+ 74%
1823
+ IO + Refine
1824
+ (k=10)
1825
+ 27%
1826
+ IO
1827
+ (best of 100)
1828
+ 33%
1829
+ CoT
1830
+ (best of 100)
1831
+ 49%
1832
+ Table 2:
1833
+ Game of 24 Results.
1834
+ Figure 3:
1835
+ Game of 24 (a) scale analysis & (b) error analysis.
1836
+ Task Setup.
1837
+ We scrape data from
1838
+ 4nums.com
1839
+ , which has 1,362 games that are sorted from easy to hard by human solving time, and use a subset of relatively hard games indexed 901-1,000 for testing. For each task, we consider the output as success if it is a valid equation that equals 24 and uses the input numbers each exactly once. We report the success rate across 100 games as the metric.
1840
+ Baselines.
1841
+ We use a standard input-output (IO) prompt with 5 in-context examples. For chain-of-thought (CoT) prompting, we augment each input-output pair with 3 intermediate equations, each operating on two remaining numbers. For example, given input “4 9 10 13”, the thoughts could be “13 - 9 = 4 (left: 4 4 10); 10 - 4 = 6 (left: 4 6); 4 * 6 = 24 (left: 24)”. For each game, we sample IO and CoT prompting for 100 times for average performance.
1842
+ We also consider a CoT self-consistency baseline, which takes the majority output from 100 CoT samples, and an iterative-refine approach on top of an IO sample for at most
1843
+ 10
1844
+ 10
1845
+ 10
1846
+ iterations. At each iteration, the LM is conditioned on all previous history to “reflect on your mistakes and generate a refined answer” if the output is incorrect. Note that it uses groundtruth feedback signals about equation correctness.
1847
+ ToT Setup.
1848
+ To frame Game of 24 into ToT, it is natural to decompose the thoughts into 3 steps, each an intermediate equation. As shown in Figure
1849
+ 2
1850
+ (a), at each tree node, we exact the remaining numbers and prompt the LM to propose some possible next steps.
1851
+ The same “propose prompt” is used for all 3 thought steps, though it only has one example with 4 input numbers.
1852
+ We perform a breadth-first search (BFS) in ToT, where at each step we keep the best
1853
+ b
1854
+ =
1855
+ 5
1856
+ 𝑏
1857
+ 5
1858
+ b=5
1859
+ candidates.
1860
+ To perform deliberate BFS in ToT, as shown in Figure
1861
+ 2
1862
+ (b), we prompt LM to evaluate each thought candidate as “sure/maybe/impossible” with regard to reaching 24. The aim is to promote correct partial solutions that can be verdicted within few lookahead trials, and eliminate impossible partial solutions based on “too big/small” commonsense, and keep the rest “maybe”. We sample values
1863
+ 3
1864
+ 3
1865
+ 3
1866
+ times for each thought.
1867
+ Results.
1868
+ As shown in Table
1869
+ 3
1870
+ , IO, CoT, and CoT-SC prompting methods perform badly on the task, achieving only 7.3%, 4.0%, and 9.0% success rates. In contrast, ToT with a breadth of
1871
+ b
1872
+ =
1873
+ 1
1874
+ 𝑏
1875
+ 1
1876
+ b=1
1877
+ already achieves a success rate of
1878
+ 45
1879
+ %
1880
+ percent
1881
+ 45
1882
+ 45\%
1883
+ , while
1884
+ b
1885
+ =
1886
+ 5
1887
+ 𝑏
1888
+ 5
1889
+ b=5
1890
+ achieves
1891
+ 74
1892
+ %
1893
+ percent
1894
+ 74
1895
+ 74\%
1896
+ .
1897
+ We also consider an oracle setup for IO/CoT, by calculating the success rate using best of
1898
+ k
1899
+ 𝑘
1900
+ k
1901
+ samples
1902
+ (
1903
+ 1
1904
+ ≤
1905
+ k
1906
+ ≤
1907
+ 100
1908
+ )
1909
+ 1
1910
+ 𝑘
1911
+ 100
1912
+ (1\leq k\leq 100)
1913
+ . To compare IO/CoT (best of k) with ToT, we consider calculating the tree nodes visited per task in ToT across
1914
+ b
1915
+ =
1916
+ 1
1917
+ ​
1918
+ ⋯
1919
+ ​
1920
+ 5
1921
+ 𝑏
1922
+ 1
1923
+ ⋯
1924
+ 5
1925
+ b=1\cdots 5
1926
+ , and map the 5 success rates in Figure
1927
+ 3
1928
+ (a), treating IO/CoT (best of
1929
+ k
1930
+ 𝑘
1931
+ k
1932
+ ) as visiting
1933
+ k
1934
+ 𝑘
1935
+ k
1936
+ nodes in a bandit. Not surprisingly, CoT scales better than IO, and best of 100 CoT samples achieve a success rate of
1937
+ 49
1938
+ %
1939
+ percent
1940
+ 49
1941
+ 49\%
1942
+ , but still much worse than exploring more nodes in ToT (
1943
+ b
1944
+ >
1945
+ 1
1946
+ 𝑏
1947
+ 1
1948
+ b>1
1949
+ ).
1950
+ Error analysis.
1951
+ Figure
1952
+ 3
1953
+ (b) breaks down at which step CoT and ToT samples fail the task, i.e. the thought (in CoT) or all
1954
+ b
1955
+ 𝑏
1956
+ b
1957
+ thoughts (in ToT) are invalid or impossible to reach 24. Notably, around 60% of CoT samples already failed the task after generating the first step, or equivalently, the first three words (e.g. “
1958
+ 4
1959
+ +
1960
+ 9
1961
+ 4
1962
+ 9
1963
+ 4+9
1964
+ ”). This highlights the issues with direct left-to-right decoding.
1965
+ 4.2
1966
+ Creative writing
1967
+ Next, we invent a creative writing task where the input is 4 random sentences and the output should be a coherent passage with 4 paragraphs that end in the 4 input sentences respectively.
1968
+ Such a task is open-ended and exploratory, and challenges creative thinking as well as high-level planning.
1969
+ Task setup.
1970
+ We sample random sentences from
1971
+ randomwordgenerator.com
1972
+ to form 100 inputs, and there is no groundtruth passage for each input constraint. As we find that GPT-4 can follow the input constraints most of the time, we focus on evaluating passage coherency in two ways: using a GPT-4 zero-shot prompt to provide a 1-10 scalar score, or using human judgments to compare pairs of outputs from different methods. For the former, we sample 5 scores and average them for each task output, and we find these 5 scores usually consistent, with a standard deviation of around
1973
+ 0.56
1974
+ 0.56
1975
+ 0.56
1976
+ on average across outputs. For the latter, we employ a subset of the authors in a blind study to compare the coherency of CoT vs. ToT generated passage pairs, where the order of passages is random flipped over 100 inputs.
1977
+ Baselines.
1978
+ Given the creative nature of the task, both IO and CoT prompts are zero-shot. While the former prompts the LM to directly generate a coherent passage given input constraints, the latter prompts the LM to first make a brief plan then write the passage, i.e. the plan serves as the intermediate thought step. We generate 10 IO and CoT samples per task.
1979
+ We also consider an iterative-refine (
1980
+ k
1981
+ ≤
1982
+ 5
1983
+ 𝑘
1984
+ 5
1985
+ k\leq 5
1986
+ ) method on top of a random IO sample for each task, where the LM is conditioned on input constraints and the last generated passage to decide if the passage is already “perfectly coherent”, and if not generate a refined one.
1987
+ ToT setup.
1988
+ We build a ToT with depth 2 (and only 1 intermediate thought step) — the LM first generates
1989
+ k
1990
+ =
1991
+ 5
1992
+ 𝑘
1993
+ 5
1994
+ k=5
1995
+ plans and votes for the best one (Figure
1996
+ 4
1997
+ ), then similarly generate
1998
+ k
1999
+ =
2000
+ 5
2001
+ 𝑘
2002
+ 5
2003
+ k=5
2004
+ passages based on the best plan then vote for the best one. Here the breadth limit
2005
+ b
2006
+ =
2007
+ 1
2008
+ 𝑏
2009
+ 1
2010
+ b=1
2011
+ , as only one choice is kept per step. A simple zero-shot vote prompt (“analyze choices below, then conclude which is most promising for the instruction”) is used to sample 5 votes at both steps.
2012
+ Results.
2013
+ Figure
2014
+ 5
2015
+ (a) shows average GPT-4 scores across 100 tasks, where ToT (7.56) is deemed to generate more coherent passages than IO (6.19) and CoT (6.93) on average. While such an automatic metric might be noisy, Figure
2016
+ 5
2017
+ (b) confirms the finding by showing that humans prefer ToT over CoT in 41 out of 100 passage pairs, while only prefer CoT over ToT in 21 (other 38 pairs are found “similarly coherent”). Lastly, iterative-refine is more effective on this natural language task, where it improves IO coherency score from 6.19 to 7.67, and ToT coherency score from 7.56 to 7.91.
2018
+ We believe it could be thought of as a third approach to thought generation in the ToT framework, where new thoughts can arise from refining old thoughts instead of i.i.d. or sequentially generated.
2019
+ Figure 4:
2020
+ A step of deliberate search in a randomly picked Creative Writing task. Given the input, the LM samples 5 different plans, then votes 5 times to decide which plan is best. The majority choice is used to consequently write the output passage with the same sample-vote procedure.
2021
+ Figure 5:
2022
+ Creative Writing results.
2023
+ Method
2024
+ Success Rate (%)
2025
+ Letter
2026
+ Word
2027
+ Game
2028
+ IO
2029
+ 38.7
2030
+ 14
2031
+ 0
2032
+ CoT
2033
+ 40.6
2034
+ 15.6
2035
+ 1
2036
+ ToT (ours)
2037
+ 78
2038
+ 60
2039
+ 20
2040
+ +best state
2041
+ 82.4
2042
+ 67.5
2043
+ 35
2044
+ -prune
2045
+ 65.4
2046
+ 41.5
2047
+ 5
2048
+ -backtrack
2049
+ 54.6
2050
+ 20
2051
+ 5
2052
+ Table 3:
2053
+ Mini Crosswords results.
2054
+ 4.3
2055
+ Mini crosswords
2056
+ Figure 6:
2057
+ In Mini Crosswords, (a) how thoughts are proposed and aggregated in a priority queue for depth-first search (DFS), and (b) how a state is evaluated based on the possibility of filling in each remaining word clue, and pruned if any remaining clue is deemed not possible to fill by the LM. Then DFS backtracks to the parent state and explore the next promising thought for clue.
2058
+ In Game of 24 and Creative Writing, ToT is relatively shallow — at most 3 thought steps are needed to reach the final output. Here we explore
2059
+ 5
2060
+ ×
2061
+ 5
2062
+ 5
2063
+ 5
2064
+ 5\times 5
2065
+ mini crosswords as a harder search problem involving natural language. Again, the goal is not just to solve the task, as more general crosswords can be readily solved with specialized NLP pipelines
2066
+ [
2067
+ 34
2068
+ ]
2069
+ that leverages large-scale retrieval instead of LM. Rather, we aim to explore the limit of LM as a general problem solver that explores its own thoughts and guides its own exploration with deliberate reasoning as heuristics.
2070
+ Task setup.
2071
+ We scrape data from
2072
+ GooBix
2073
+ , which contains 156 games of
2074
+ 5
2075
+ ×
2076
+ 5
2077
+ 5
2078
+ 5
2079
+ 5\times 5
2080
+ mini crosswords. As we observe adjacent games contain similar clues, we use 20 games with indices
2081
+ 1
2082
+ ,
2083
+ 6
2084
+ ,
2085
+ ⋯
2086
+ ,
2087
+ 91
2088
+ ,
2089
+ 96
2090
+ 1
2091
+ 6
2092
+ ⋯
2093
+ 91
2094
+ 96
2095
+ 1,6,\cdots,91,96
2096
+ for testing, and games
2097
+ 136
2098
+ ,
2099
+ 141
2100
+ ,
2101
+ 146
2102
+ ,
2103
+ 151
2104
+ ,
2105
+ 156
2106
+ 136
2107
+ 141
2108
+ 146
2109
+ 151
2110
+ 156
2111
+ 136,141,146,151,156
2112
+ for prompting.
2113
+ For each task, the input describes the 5 horizontal clues and 5 vertical clues, and the output should be a board of
2114
+ 5
2115
+ ×
2116
+ 5
2117
+ =
2118
+ 25
2119
+ 5
2120
+ 5
2121
+ 25
2122
+ 5\times 5=25
2123
+ letters to solve the crosswords. For evaluation, we consider three levels of success: the portion of correct letters (25 per game), words (10 per game), and games.
2124
+ Baselines.
2125
+ We provide 5 example input-output pairs in the IO prompt, and in the CoT prompt additionally include intermediate words in the order h1..5 then v1..5. We run each prompt for 10 samples and average the results.
2126
+ ToT setup.
2127
+ We leverage a depth-first search (Algorithm
2128
+ 2
2129
+ ) that keeps exploring the most promising subsequent word clue until the state is no longer promising, then backtrack to the parent state to explore alternative thoughts.
2130
+ To make search tractable, subsequent thoughts are constrained not to change any filled words or letters, so that the ToT has at most 10 intermediate steps.
2131
+ For thought generation, at each state we translate all existing thoughts (e.g. “h2.motor; h1.tasks” for the state in Figure
2132
+ 6
2133
+ (a)) into letter constraints for remaining clues (e.g. “v1.To heap: tm___;…”) and prompt a proposal prompt
2134
+ 5
2135
+ 5
2136
+ 5
2137
+ times to come up with candidates for where and what to fill in the next word. Importantly, we also prompt the LM to give a confidence level for different thoughts, and aggregate these across proposals to obtain a sorted list of next thoughts to explore (Figure
2138
+ 6
2139
+ (a)).
2140
+ For state evaluations, we similarly translate each state into letter constraints for remaining clues, then evaluate for each clue if it is possible to fill given the constraints. If any remaining clue is deemed “impossible” to fill in (e.g. “v1. To heap: tm_s_”), then the exploration of the state’s subtree is pruned and DFS backtracks to its parent to explore the next promising thought. We limit DFS search steps to 100, and simply render the deepest explored state (the first explored one if multiple) into the final output.
2141
+ Results.
2142
+ As shown in Table
2143
+ 5
2144
+ , IO and CoT prompting methods perform poorly with a word-level success rate less than
2145
+ 16
2146
+ %
2147
+ percent
2148
+ 16
2149
+ 16\%
2150
+ , while ToT significantly improves all metrics, achieving a word-level success rate of
2151
+ 60
2152
+ %
2153
+ percent
2154
+ 60
2155
+ 60\%
2156
+ and solving 4 out of 20 games. Such an improvement is not surprising, given IO and CoT lack mechanisms to try different clues, make changes to decisions, or backtrack.
2157
+ Oracle and ablation studies.
2158
+ When outputting from the oracle best DFS state (instead of the heuristically determined best state) per task, ToT performance is even higher and actually solves 7/20 games (Table
2159
+ 5
2160
+ , “+best state”), indicating our simple output heuristics can be readily improved. Interestingly, sometimes when the crosswords game is actually solved, the state evaluator might still deem some words as “impossible” and prune — possibly because
2161
+ 5
2162
+ ×
2163
+ 5
2164
+ 5
2165
+ 5
2166
+ 5\times 5
2167
+ crosswords by design have some rare or obselete words that GPT-4 cannot recognize
2168
+ 2
2169
+ 2
2170
+ 2
2171
+ For example, “agend” is an obsolete form of “agendum”, but GPT-4 deems it a typo for “agenda”. External retrieval
2172
+ or web interaction
2173
+ could augment LM for problem solving under knowledge uncertainty.
2174
+ .
2175
+ Given the state evaluation as a pruning heuristic is imperfect, we also explore ablating the pruning, and find the performance generally worse (Table
2176
+ 5
2177
+ , “-prune”). However, it could actually find the correct solution for 4/20 games (though only outputting 1 via heuristic), 3 of which are games ToT+pruning cannot solve within 100 steps. Thus, better heuristics for DFS pruning are critical for problem solving in this case.
2178
+ Lastly, we confirm the importance of backtracking by running an ablation that keeps filling the most promising clue for at most 20 steps, allowing overwrites. This is similar to a “greedy” BFS search with breadth limit of
2179
+ b
2180
+ =
2181
+ 1
2182
+ 𝑏
2183
+ 1
2184
+ b=1
2185
+ , and performs poorly with a word level success of only
2186
+ 20
2187
+ %
2188
+ percent
2189
+ 20
2190
+ 20\%
2191
+ (Table
2192
+ 5
2193
+ , “-backtrack”).
2194
+ 5
2195
+ Related Work
2196
+ Planning and decision making.
2197
+ Smart planning and decision making are critical to achieving predefined goals. As they are trained on vast amount of world knowledge and human examples,
2198
+ LMs are known to have already absorbed rich commonsense that makes it possible to propose reasonable plans conditioned on problem setting and environmental states
2199
+ [
2200
+ 12
2201
+ ,
2202
+ 42
2203
+ ,
2204
+ 37
2205
+ ,
2206
+ 13
2207
+ ,
2208
+ 35
2209
+ ,
2210
+ 41
2211
+ ,
2212
+ 40
2213
+ ]
2214
+ . Our proposed ToT approach extends existing planning formulations by considering multiple potentially feasible plans simultaneously at each problem-solving step, and proceeding with the most promising ones. The integration between thought sampling and value feedback organically integrates planning and decision-making mechanisms, enabling effective search inside a solution tree. On the other hand, traditional decision-making procedures usually require training dedicated reward and policy models as in reinforcement learning (for example CHAI
2215
+ [
2216
+ 33
2217
+ ]
2218
+ ), whereas we use the LM itself to provide the value estimates for decision making.
2219
+ RAP
2220
+ [
2221
+ 9
2222
+ ]
2223
+ is a concurrent work that treats language model reasoning as planning with its internal world model, and proposes a MCTS-based method similar to ToT. However, its tasks are simpler than ours, and its framework lacks the modularity to incorporate different tree search algorithms.
2224
+ Self-reflection.
2225
+ Using LLMs to assess the viability of their own predictions is becoming an increasingly important procedure in problem solving.
2226
+ [
2227
+ 28
2228
+ ,
2229
+ 20
2230
+ ,
2231
+ 24
2232
+ ]
2233
+ introduced the “self-reflection” mechanism, in which LMs provide feedback to their generation candidates.
2234
+ [
2235
+ 4
2236
+ ]
2237
+ improves LMs code generation accuracy by injecting feedback messages generated by the LM itself based on its code execution results. Similarly,
2238
+ [
2239
+ 17
2240
+ ]
2241
+ also introduces “critic” or review steps over the actions and states, deciding the next action to take in solving computer operation tasks. Another recent work very relevant to ours is “self-eval guided decoding”
2242
+ [
2243
+ 39
2244
+ ]
2245
+ . Similar to our method, self-eval decoding also follows a tree-search procedure with leaves sampled from stochastic beam search decoding, which are then evaluated by LLM itself with carefully prepared self-eval prompts. Their approach however, uses the PAL formulation
2246
+ [
2247
+ 8
2248
+ ]
2249
+ which represents thoughts as codes, which makes it difficult to tackle challenging tasks like creative writing which we consider in this paper. Our Tree-of-Thought formulation is thus more versatile and handles challenging tasks on which GPT-4 only achieves very low accuracy with standard prompts.
2250
+ Program-guided LLM generation.
2251
+ Our proposal is also related to recent advancements that organize LM’s behavior with systematic procedures
2252
+ [
2253
+ 14
2254
+ ,
2255
+ 44
2256
+ ,
2257
+ 6
2258
+ ,
2259
+ 43
2260
+ ]
2261
+ or symbolic program guidance. For example,
2262
+ Schlag et al. [
2263
+ 27
2264
+ ]
2265
+ embeds LMs in an algorithmic search procedure to help solve problems like question answering step-by-step, in which the search trees are expanded by relevant paragraphs that might provide answers. This approach however differs from ours in that trees are expanded by sampling external paragraphs instead of the LM’s own thoughts, and there is no reflection or voting steps. Another approach, LLM+P
2266
+ [
2267
+ 18
2268
+ ]
2269
+ , goes one step further and delegates the actual planning process to a classical planner.
2270
+ Classical search methods.
2271
+ Last but not least, our approach can be treated as a modern rendition of classical search methods for problem solving. For example it can be considered as a heuristic search algorithm like A*
2272
+ [
2273
+ 10
2274
+ ]
2275
+ , in which the heuristic at each search node is provided by the LM’s self-assessment. From this perspective, our method is also related to NeuroLogic A*esque decoding
2276
+ [
2277
+ 19
2278
+ ]
2279
+ , which is inspired by A* search but introduces look-ahead heuristics that are efficient for LMs to improve the beam-search or top-k sampling decoding. This method however is constrained to sentence generation tasks, whereas our framework are designed for complex, multi-step problem solving guarded by value feedback.
2280
+ 6
2281
+ Discussion
2282
+ Limitations and future directions.
2283
+ Deliberate search such as ToT might not be necessary for many existing tasks that GPT-4 already excels at (see Appendix
2284
+ B.1
2285
+ ), and as an initial step this work only explores three relatively simple tasks that challenges GPT-4 (see Appendix
2286
+ B.2
2287
+ for some GPT-3.5 experiment results) and calls of better search and planning abilities incorporated with LMs. However, as we begin to deploy LMs for more real-world decision making applications (e.g. coding, data analysis, robotics, etc.), more complex tasks could emerge and present new opportunities to study these research questions. Also, search methods like ToT requires more resources (e.g. GPT-4 API cost) than sampling methods in order to improve task performances, but the modular flexibility of ToT allows users to customize such performance-cost tradeoffs, and ongoing open-source efforts
2288
+ [
2289
+ 32
2290
+ ]
2291
+ should readily reduce such costs in the near future. More details about cost and efficiency are in Appendix
2292
+ B.3
2293
+ . Lastly, this work focuses on using an off-the-shelf LM, and fine-tuning LMs using a ToT-style high-level counterfactual decision making (e.g. deliberating over potential choices for the next paragraph, instead of predicting the next token) might present opportunities to enhance the problem-solving capabilities of LMs.
2294
+ Conclusion.
2295
+ The associative “System 1” of LMs can be beneficially augmented by a “System 2” based on searching a tree of possible paths to the solution to a problem. The Tree of Thoughts framework provides a way to translate classical insights about problem-solving into actionable methods for contemporary LMs. At the same time, LMs address a weakness of these classical methods, providing a way to solve complex problems that are not easily formalized, such as creative writing. We see this intersection of LMs with classical approaches to AI as an exciting direction.
2296
+ Broader Impact
2297
+ ToT is a framework that empowers LMs to more autonomously and intelligently make decisions and solve problems. While current tasks are limited to reasoning and search problems, future applications involving interaction with external environments or humans could bring potential danger, e.g. facilitating harmful uses of LMs. On the other hand, ToT also improves the interpretability of model decisions and the opportunity for human alignment, as the resulting representations are readable, high-level language reasoning instead of implicit, low-level token values.
2298
+ Acknowledgements
2299
+ SY and KN acknowledge support from an Oracle Collaborative Research award and the National Science Foundation under Grant No. 2239363. Any opinions, findings, conclusions, or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation. SY is also supported by the Harold W. Dodds Fellowship from Princeton.
2300
+ References
2301
+ Brown et al. [2020]
2302
+ T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal,
2303
+ A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al.
2304
+ Language models are few-shot learners.
2305
+ Advances in neural information processing systems
2306
+ ,
2307
+ 33:1877–1901, 2020.
2308
+ Browne et al. [2012]
2309
+ C. Browne, E. J. Powley, D. Whitehouse, S. M. M. Lucas, P. I. Cowling,
2310
+ P. Rohlfshagen, S. Tavener, D. P. Liebana, S. Samothrakis, and S. Colton.
2311
+ A survey of monte carlo tree search methods.
2312
+ IEEE Transactions on Computational Intelligence and AI in
2313
+ Games
2314
+ , 4:1–43, 2012.
2315
+ Campbell et al. [2002]
2316
+ M. Campbell, A. J. Hoane Jr, and F.-h. Hsu.
2317
+ Deep blue.
2318
+ Artificial intelligence
2319
+ , 134(1-2):57–83,
2320
+ 2002.
2321
+ Chen et al. [2023]
2322
+ X. Chen, M. Lin, N. Schärli, and D. Zhou.
2323
+ Teaching large language models to self-debug, 2023.
2324
+ Chowdhery et al. [2022]
2325
+ A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham,
2326
+ H. W. Chung, C. Sutton, S. Gehrmann, et al.
2327
+ Palm: Scaling language modeling with pathways.
2328
+ arXiv preprint arXiv:2204.02311
2329
+ , 2022.
2330
+ Creswell and Shanahan [2022]
2331
+ A. Creswell and M. Shanahan.
2332
+ Faithful reasoning using large language models.
2333
+ arXiv preprint arXiv:2208.14271
2334
+ , 2022.
2335
+ Daw et al. [2005]
2336
+ N. D. Daw, Y. Niv, and P. Dayan.
2337
+ Uncertainty-based competition between prefrontal and dorsolateral
2338
+ striatal systems for behavioral control.
2339
+ Nature neuroscience
2340
+ , 8(12):1704–1711,
2341
+ 2005.
2342
+ Gao et al. [2023]
2343
+ L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig.
2344
+ Pal: Program-aided language models, 2023.
2345
+ Hao et al. [2023]
2346
+ S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu.
2347
+ Reasoning with language model is planning with world model.
2348
+ arXiv preprint arXiv:2305.14992
2349
+ , 2023.
2350
+ Hart et al. [1968a]
2351
+ P. E. Hart, N. J. Nilsson, and B. Raphael.
2352
+ A formal basis for the heuristic determination of minimum cost paths.
2353
+ IEEE Transactions on Systems Science and Cybernetics
2354
+ ,
2355
+ 4(2):100–107, 1968a.
2356
+ doi:
2357
+ 10.1109/TSSC.1968.300136
2358
+ .
2359
+ Hart et al. [1968b]
2360
+ P. E. Hart, N. J. Nilsson, and B. Raphael.
2361
+ A formal basis for the heuristic determination of minimum cost paths.
2362
+ IEEE transactions on Systems Science and Cybernetics
2363
+ ,
2364
+ 4(2):100–107, 1968b.
2365
+ Huang et al. [2022a]
2366
+ W. Huang, P. Abbeel, D. Pathak, and I. Mordatch.
2367
+ Language models as zero-shot planners: Extracting actionable
2368
+ knowledge for embodied agents, 2022a.
2369
+ Huang et al. [2022b]
2370
+ W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson,
2371
+ I. Mordatch, Y. Chebotar, et al.
2372
+ Inner monologue: Embodied reasoning through planning with language
2373
+ models.
2374
+ arXiv preprint arXiv:2207.05608
2375
+ , 2022b.
2376
+ Jung et al. [2022]
2377
+ J. Jung, L. Qin, S. Welleck, F. Brahman, C. Bhagavatula, R. L. Bras, and
2378
+ Y. Choi.
2379
+ Maieutic prompting: Logically consistent reasoning with recursive
2380
+ explanations.
2381
+ arXiv preprint arXiv:2205.11822
2382
+ , 2022.
2383
+ Kahneman [2011]
2384
+ D. Kahneman.
2385
+ Thinking, fast and slow
2386
+ .
2387
+ Macmillan, 2011.
2388
+ Kahneman et al. [2002]
2389
+ D. Kahneman, S. Frederick, et al.
2390
+ Representativeness revisited: Attribute substitution in intuitive
2391
+ judgment.
2392
+ Heuristics and biases: The psychology of intuitive judgment
2393
+ ,
2394
+ 49(49-81):74, 2002.
2395
+ Kim et al. [2023]
2396
+ G. Kim, P. Baldi, and S. McAleer.
2397
+ Language models can solve computer tasks, 2023.
2398
+ Liu et al. [2023]
2399
+ B. Liu, Y. Jiang, X. Zhang, Q. Liu, S. Zhang, J. Biswas, and P. Stone.
2400
+ Llm+p: Empowering large language models with optimal planning
2401
+ proficiency, 2023.
2402
+ Lu et al. [2021]
2403
+ X. Lu, S. Welleck, P. West, L. Jiang, J. Kasai, D. Khashabi, R. L. Bras,
2404
+ L. Qin, Y. Yu, R. Zellers, N. A. Smith, and Y. Choi.
2405
+ Neurologic a*esque decoding: Constrained text generation with
2406
+ lookahead heuristics.
2407
+ In
2408
+ North American Chapter of the Association for Computational
2409
+ Linguistics
2410
+ , 2021.
2411
+ Madaan et al. [2023]
2412
+ A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon,
2413
+ N. Dziri, S. Prabhumoye, Y. Yang, S. Welleck, B. P. Majumder, S. Gupta,
2414
+ A. Yazdanbakhsh, and P. Clark.
2415
+ Self-refine: Iterative refinement with self-feedback, 2023.
2416
+ Newell et al. [1959]
2417
+ A. Newell, J. C. Shaw, and H. A. Simon.
2418
+ Report on a general problem solving program.
2419
+ In
2420
+ IFIP congress
2421
+ , volume 256, page 64. Pittsburgh, PA, 1959.
2422
+ Newell et al. [1972]
2423
+ A. Newell, H. A. Simon, et al.
2424
+ Human problem solving
2425
+ .
2426
+ Prentice-Hall, 1972.
2427
+ OpenAI [2023]
2428
+ OpenAI.
2429
+ Gpt-4 technical report.
2430
+ ArXiv
2431
+ , abs/2303.08774, 2023.
2432
+ Paul et al. [2023]
2433
+ D. Paul, M. Ismayilzada, M. Peyrard, B. Borges, A. Bosselut, R. West, and
2434
+ B. Faltings.
2435
+ Refiner: Reasoning feedback on intermediate representations, 2023.
2436
+ Radford et al. [2018]
2437
+ A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al.
2438
+ Improving language understanding by generative pre-training.
2439
+ OpenAI blog
2440
+ , 2018.
2441
+ Radford et al. [2019]
2442
+ A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al.
2443
+ Language models are unsupervised multitask learners.
2444
+ OpenAI blog
2445
+ , 1(8):9, 2019.
2446
+ Schlag et al. [2023]
2447
+ I. Schlag, S. Sukhbaatar, A. Celikyilmaz, W. tau Yih, J. Weston,
2448
+ J. Schmidhuber, and X. Li.
2449
+ Large language model programs, 2023.
2450
+ Shinn et al. [2023]
2451
+ N. Shinn, B. Labash, and A. Gopinath.
2452
+ Reflexion: an autonomous agent with dynamic memory and
2453
+ self-reflection, 2023.
2454
+ Silver et al. [2017]
2455
+ D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez,
2456
+ T. Hubert, L. Baker, M. Lai, A. Bolton, et al.
2457
+ Mastering the game of go without human knowledge.
2458
+ nature
2459
+ , 550(7676):354–359, 2017.
2460
+ Sloman [1996]
2461
+ S. A. Sloman.
2462
+ The empirical case for two systems of reasoning.
2463
+ Psychological bulletin
2464
+ , 119(1):3, 1996.
2465
+ Stanovich [1999]
2466
+ K. E. Stanovich.
2467
+ Who is rational? Studies of individual differences in
2468
+ reasoning
2469
+ .
2470
+ Psychology Press, 1999.
2471
+ Touvron et al. [2023]
2472
+ H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix,
2473
+ B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al.
2474
+ Llama: Open and efficient foundation language models.
2475
+ arXiv preprint arXiv:2302.13971
2476
+ , 2023.
2477
+ Verma et al. [2022]
2478
+ S. Verma, J. Fu, S. Yang, and S. Levine.
2479
+ Chai: A chatbot ai for task-oriented dialogue with offline
2480
+ reinforcement learning.
2481
+ In
2482
+ Proceedings of the 2022 Conference of the North American
2483
+ Chapter of the Association for Computational Linguistics: Human Language
2484
+ Technologies
2485
+ , pages 4471–4491, 2022.
2486
+ Wallace et al. [2022]
2487
+ E. Wallace, N. Tomlin, A. Xu, K. Yang, E. Pathak, M. Ginsberg, and D. Klein.
2488
+ Automated crossword solving.
2489
+ arXiv preprint arXiv:2205.09665
2490
+ , 2022.
2491
+ Wang et al. [2023a]
2492
+ L. Wang, W. Xu, Y. Lan, Z. Hu, Y. Lan, R. K.-W. Lee, and E.-P. Lim.
2493
+ Plan-and-solve prompting: Improving zero-shot chain-of-thought
2494
+ reasoning by large language models, 2023a.
2495
+ Wang et al. [2022]
2496
+ X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, and D. Zhou.
2497
+ Self-consistency improves chain of thought reasoning in language
2498
+ models.
2499
+ arXiv preprint arXiv:2203.11171
2500
+ , 2022.
2501
+ Wang et al. [2023b]
2502
+ Z. Wang, S. Cai, A. Liu, X. Ma, and Y. Liang.
2503
+ Describe, explain, plan and select: Interactive planning with large
2504
+ language models enables open-world multi-task agents, 2023b.
2505
+ Wei et al. [2022]
2506
+ J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou.
2507
+ Chain of thought prompting elicits reasoning in large language
2508
+ models.
2509
+ arXiv preprint arXiv:2201.11903
2510
+ , 2022.
2511
+ Xie et al. [2023]
2512
+ Y. Xie, K. Kawaguchi, Y. Zhao, X. Zhao, M.-Y. Kan, J. He, and Q. Xie.
2513
+ Decomposition enhances reasoning via self-evaluation guided decoding,
2514
+ 2023.
2515
+ Yang et al. [2023]
2516
+ S. Yang, O. Nachum, Y. Du, J. Wei, P. Abbeel, and D. Schuurmans.
2517
+ Foundation models for decision making: Problems, methods, and
2518
+ opportunities, 2023.
2519
+ Yao et al. [2022]
2520
+ S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao.
2521
+ ReAct: Synergizing reasoning and acting in language models.
2522
+ arXiv preprint arXiv:2210.03629
2523
+ , 2022.
2524
+ Zhang et al. [2023]
2525
+ S. Zhang, Z. Chen, Y. Shen, M. Ding, J. B. Tenenbaum, and C. Gan.
2526
+ Planning with large language models for code generation.
2527
+ In
2528
+ The Eleventh International Conference on Learning
2529
+ Representations
2530
+ , 2023.
2531
+ URL
2532
+ https://openreview.net/forum?id=Lr8cOOtYbfL
2533
+ .
2534
+ Zhou et al. [2022]
2535
+ D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans,
2536
+ C. Cui, O. Bousquet, Q. Le, et al.
2537
+ Least-to-most prompting enables complex reasoning in large language
2538
+ models.
2539
+ arXiv preprint arXiv:2205.10625
2540
+ , 2022.
2541
+ Zhu et al. [2022]
2542
+ X. Zhu, J. Wang, L. Zhang, Y. Zhang, R. Gan, J. Zhang, and Y. Yang.
2543
+ Solving math word problem via cooperative reasoning induced language
2544
+ models.
2545
+ arXiv preprint arXiv:2210.16257
2546
+ , 2022.
2547
+ Appendix A
2548
+ Code, Prompts, Trajectories
2549
+ All code is available at
2550
+ https://github.com/princeton-nlp/tree-of-thought-llm
2551
+ .
2552
+ All prompts are available at
2553
+ https://github.com/princeton-nlp/tree-of-thought-llm/tree/master/src/tot/prompts
2554
+ .
2555
+ Trajectories are available at
2556
+ https://github.com/princeton-nlp/tree-of-thought-llm/tree/master/logs
2557
+ .
2558
+ Appendix B
2559
+ Additional Experiment Results
2560
+ Given the motivation of exploring and extending the capability frontier of language models, our experiments in the main paper have focused on a setup with the state-of-the-art language model (GPT-4), and three hard tasks invented to challenge it. Here, we report additional experiments with weaker LLM or easier tasks, and discuss cost and efficiency.
2561
+ GSM8K
2562
+ StrategyQA
2563
+ IO
2564
+ 51
2565
+ 73
2566
+ CoT
2567
+ 86
2568
+ 82
2569
+ ToT
2570
+ 90
2571
+ 83
2572
+ Table 4:
2573
+ New tasks with
2574
+ zero-shot ToT and GPT-4.
2575
+ GPT-4
2576
+ GPT-3.5
2577
+ IO
2578
+ 7.3%
2579
+ 6%
2580
+ CoT
2581
+ 4.0%
2582
+ 3%
2583
+ ToT
2584
+ 74%
2585
+ 19%
2586
+ Table 5:
2587
+ Game of 24 with
2588
+ GPT-4 vs GPT-3.5.
2589
+ GPT-4
2590
+ GPT-3.5
2591
+ IO
2592
+ 6.19
2593
+ 4.47
2594
+ CoT
2595
+ 6.93
2596
+ 5.16
2597
+ ToT
2598
+ 7.56
2599
+ 6.62
2600
+ Table 6:
2601
+ Creative Writing with
2602
+ GPT-4 vs. GPT-3.5.
2603
+ B.1
2604
+ Extension to new tasks (GSM8k, StrategyQA) with zero-shot ToT
2605
+ While more common NLP tasks might be too easy for GPT-4 and do not require ToT (which is why we considered harder new tasks), we believe applying ToT to new tasks could be straightforward. For example, we implemented a simple and generic zero-shot ToT-BFS similar to creative writing (sample 5 problem solving strategies then vote for the best one; then sample 5 solutions based on the best strategy then vote for the best one) for GSM8K and StrategyQA with few extra lines of code:
2606
+ # define the answer format of new tasks
2607
+ gsm8k_format = ‘"the answer is n" where n is a number’
2608
+ strategyqa_format = ‘either "the answer is yes" or "the answer is no"’
2609
+
2610
+ # define zero-shot io prompting
2611
+ standard_prompt = ‘Answer the following question with {format}: {input}’
2612
+
2613
+ # define thought format for zero-shot cot and zero-shot tot
2614
+ cot_prompt = ‘‘‘Answer the following question: {input}
2615
+
2616
+ Make a strategy then write. Your output should be of the following format:
2617
+
2618
+ Strategy:
2619
+ Your strategy about how to answer the question.
2620
+
2621
+ Answer:
2622
+ Your answer to the question. It should end with {format}.
2623
+ ’’’
2624
+
2625
+ # define zero-shot voting used for zero-shot tot
2626
+ vote_prompt = ‘‘‘Given an instruction and several choices,
2627
+ decide which choice is most promising.
2628
+ Analyze each choice in detail, then conclude in the last line
2629
+ "The best choice is {s}", where s the integer id of the choice.
2630
+ ’’’
2631
+ We evaluated on a subset of 100 random GSM8K test and StrategyQA dev questions. As shown in Table
2632
+ B
2633
+ and as expected, ToT improves over CoT on both tasks (but only slightly, given GPT-4 + CoT is already very good on such tasks, and StrategyQA’s bottleneck is external knowledge, not reasoning). Considering computational costs, it is more suitable to try smaller LLMs + ToT for traditional NLP tasks, or GPT-4 + ToT for hard tasks that challenge GPT-4 + CoT’s reasoning.
2634
+ B.2
2635
+ Extension to new LMs (GPT-3.5)
2636
+ To understand how ToT works with other LLMs, we also ran GPT-3.5-turbo for Creative Writing (Table
2637
+ B
2638
+ ) and Game of 24 (Table
2639
+ B
2640
+ ).
2641
+ On both tasks, “ToT
2642
+ >
2643
+ >
2644
+ CoT
2645
+ >
2646
+ >
2647
+ IO” remains true for GPT-3.5.
2648
+ On Creative Writing, we find GPT-3.5+ToT outperform GPT-4+IO, and similar to GPT-4+CoT, which suggests ToT could also work well on weaker language models.
2649
+ On Game of 24 (we changed 1-shot proposal prompt to 3-shot to make it work), GPT-3.5+ToT’s 19% is far worse than GPT-4+ToT’s 74%. To further understand the importance of generation vs. evaluation, we ran GPT-4 generation + GPT-3.5 evaluation (64%) and GPT-3.5 generation + GPT-4 evaluation (31%). This suggests the game’s bottleneck is thought generation, and different generation/evaluation language models might attain decent results while reducing costs.
2650
+ B.3
2651
+ Cost and efficiency
2652
+ Running ToT requires significantly more computations than IO or CoT prompting. For example, in Game of 24 (Table
2653
+ 7
2654
+ below), solving a problem with ToT requires 5.5k completion tokens, close to 100 CoT trials (6.7k tokens). But the performance of ToT is better than best of 100 independent CoT trials.
2655
+ Game of 24
2656
+ Generate/Prompt tokens
2657
+ Cost per case
2658
+ Success
2659
+ IO (best of 100)
2660
+ 1.8k / 1.0k
2661
+ $0.13
2662
+ 33%
2663
+ CoT (best of 100)
2664
+ 6.7k / 2.2k
2665
+ $0.47
2666
+ 49%
2667
+ ToT
2668
+ 5.5k / 1.4k
2669
+ $0.74
2670
+ 74%
2671
+ Table 7:
2672
+ Cost analysis on Game of 24.
2673
+ On Creative Writing (Table
2674
+ 8
2675
+ below), we found ToT takes around 5x completion tokens and money cost, which is intuitive as
2676
+ b
2677
+ =
2678
+ 5
2679
+ 𝑏
2680
+ 5
2681
+ b=5
2682
+ and most tokens are generated passages.
2683
+ Creative Writing
2684
+ Generate/Prompt tokens
2685
+ Cost per case
2686
+ IO
2687
+ 0.9k / 0.4k
2688
+ $0.06
2689
+ CoT
2690
+ 0.9k / 0.4k
2691
+ $0.07
2692
+ ToT
2693
+ 4k / 2.9k
2694
+ $0.32
2695
+ Table 8:
2696
+ Cost analysis on Game of 24.
2697
+ So completing Game of 24 and Creative Writing’s main ToT experiments cost around
2698
+ 0.74
2699
+ ×
2700
+ 100
2701
+ +
2702
+ 0.32
2703
+ ×
2704
+ 100
2705
+ =
2706
+ 106
2707
+ 0.74
2708
+ 100
2709
+ 0.32
2710
+ 100
2711
+ 106
2712
+ 0.74\times 100+0.32\times 100=106
2713
+ dollars. Crosswords’ DFS experiments should be also within
2714
+ 100
2715
+ 100
2716
+ 100
2717
+ dollars. In general, cost and efficiency of ToT highly depend on the prompts and search algorithms used, and could require 5-100 times more generated tokens than CoT. Some actionable insights:
2718
+ •
2719
+ We recommend using ToT on tasks requiring deliberate reasoning, on which CoT struggles.
2720
+ •
2721
+ Flexibility of ToT allows some performance-cost tradeoff, e.g., change beam size or vote number in BFS, few-shot vs. zero-shot prompting, GPT-3.5 vs. GPT-4, etc. One could configure the setup based on some resource constraints or performance goal.
2722
+ •
2723
+ There is much space for improving efficiency, e.g., BFS could early stop when solution is found, or trim down beam size to when some thoughts are ”impossible”.
2724
+ •
2725
+ We believe that more computation is indeed required in order for the model to achieve stronger intelligence, and this should not become a blocking issue as in the long run, (open-source) LMs will become much cheaper and more efficient. It is also a great direction how to better train/finetune LMs for thought generation and/or evaluation.
2726
+ ◄
2727
+ Feeling
2728
+ lucky?
2729
+ Conversion
2730
+ report
2731
+ Report
2732
+ an issue
2733
+ View original
2734
+ on arXiv
2735
+ ►
research/notes/230510601-tree-of-thoughts-deliberate-problem-solving-with-large-language-models.md ADDED
@@ -0,0 +1,213 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2305.10601] Tree of Thoughts: Deliberate Problem Solving with Large Language
3
+ Models'
4
+ id: 230510601-tree-of-thoughts-deliberate-problem-solving-with-large-language-models
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:39:55.627654Z'
8
+ source: https://arxiv.org/abs/2305.10601
9
+ source_domain: arxiv.org
10
+ fetched_at: '2026-06-10T00:39:55.627506Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2305.10601] Tree of Thoughts: Deliberate Problem Solving with Large Language Models
20
+ Computer Science > Computation and Language
21
+ arXiv:2305.10601
22
+ (cs)
23
+ [Submitted on 17 May 2023 (
24
+ v1
25
+ ), last revised 3 Dec 2023 (this version, v2)]
26
+ Title:
27
+ Tree of Thoughts: Deliberate Problem Solving with Large Language Models
28
+ Authors:
29
+ Shunyu Yao
30
+ ,
31
+ Dian Yu
32
+ ,
33
+ Jeffrey Zhao
34
+ ,
35
+ Izhak Shafran
36
+ ,
37
+ Thomas L. Griffiths
38
+ ,
39
+ Yuan Cao
40
+ ,
41
+ Karthik Narasimhan
42
+ View a PDF of the paper titled Tree of Thoughts: Deliberate Problem Solving with Large Language Models, by Shunyu Yao and 6 other authors
43
+ View PDF
44
+ HTML (experimental)
45
+ Abstract:
46
+ Language models are increasingly being deployed for general problem solving across a wide range of tasks, but are still confined to token-level, left-to-right decision-making processes during inference. This means they can fall short in tasks that require exploration, strategic lookahead, or where initial decisions play a pivotal role. To surmount these challenges, we introduce a new framework for language model inference, Tree of Thoughts (ToT), which generalizes over the popular Chain of Thought approach to prompting language models, and enables exploration over coherent units of text (thoughts) that serve as intermediate steps toward problem solving. ToT allows LMs to perform deliberate decision making by considering multiple different reasoning paths and self-evaluating choices to decide the next course of action, as well as looking ahead or backtracking when necessary to make global choices. Our experiments show that ToT significantly enhances language models' problem-solving abilities on three novel tasks requiring non-trivial planning or search: Game of 24, Creative Writing, and Mini Crosswords. For instance, in Game of 24, while GPT-4 with chain-of-thought prompting only solved 4% of tasks, our method achieved a success rate of 74%. Code repo with all prompts:
47
+ this https URL
48
+ .
49
+ Comments:
50
+ NeurIPS 2023 camera ready version. Code repo with all prompts:
51
+ this https URL
52
+ Subjects:
53
+ Computation and Language (cs.CL)
54
+ ; Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
55
+ Cite as:
56
+ arXiv:2305.10601
57
+ [cs.CL]
58
+ (or
59
+ arXiv:2305.10601v2
60
+ [cs.CL]
61
+ for this version)
62
+ https://doi.org/10.48550/arXiv.2305.10601
63
+ Focus to learn more
64
+ arXiv-issued DOI via DataCite
65
+ Submission history
66
+ From: Shunyu Yao [
67
+ view email
68
+ ]
69
+ [v1]
70
+ Wed, 17 May 2023 23:16:17 UTC (609 KB)
71
+ [v2]
72
+ Sun, 3 Dec 2023 22:50:35 UTC (623 KB)
73
+ Full-text links:
74
+ Access Paper:
75
+ View a PDF of the paper titled Tree of Thoughts: Deliberate Problem Solving with Large Language Models, by Shunyu Yao and 6 other authors
76
+ View PDF
77
+ HTML (experimental)
78
+ TeX Source
79
+ view license
80
+ Current browse context:
81
+ cs.CL
82
+ < prev
83
+ |
84
+ next >
85
+ new
86
+ |
87
+ recent
88
+ |
89
+ 2023-05
90
+ Change to browse by:
91
+ cs
92
+ cs.AI
93
+ cs.LG
94
+ References & Citations
95
+ NASA ADS
96
+ Google Scholar
97
+ Semantic Scholar
98
+ 5 blog links
99
+ (
100
+ what is this?
101
+ )
102
+ export BibTeX citation
103
+ Loading...
104
+ BibTeX formatted citation
105
+ ×
106
+ loading...
107
+ Data provided by:
108
+ Bookmark
109
+ Bibliographic Tools
110
+ Bibliographic and Citation Tools
111
+ Bibliographic Explorer Toggle
112
+ Bibliographic Explorer
113
+ (
114
+ What is the Explorer?
115
+ )
116
+ Connected Papers Toggle
117
+ Connected Papers
118
+ (
119
+ What is Connected Papers?
120
+ )
121
+ Litmaps Toggle
122
+ Litmaps
123
+ (
124
+ What is Litmaps?
125
+ )
126
+ scite.ai Toggle
127
+ scite Smart Citations
128
+ (
129
+ What are Smart Citations?
130
+ )
131
+ Code, Data, Media
132
+ Code, Data and Media Associated with this Article
133
+ alphaXiv Toggle
134
+ alphaXiv
135
+ (
136
+ What is alphaXiv?
137
+ )
138
+ Links to Code Toggle
139
+ CatalyzeX Code Finder for Papers
140
+ (
141
+ What is CatalyzeX?
142
+ )
143
+ DagsHub Toggle
144
+ DagsHub
145
+ (
146
+ What is DagsHub?
147
+ )
148
+ GotitPub Toggle
149
+ Gotit.pub
150
+ (
151
+ What is GotitPub?
152
+ )
153
+ Huggingface Toggle
154
+ Hugging Face
155
+ (
156
+ What is Huggingface?
157
+ )
158
+ Links to Code Toggle
159
+ Papers with Code
160
+ (
161
+ What is Papers with Code?
162
+ )
163
+ ScienceCast Toggle
164
+ ScienceCast
165
+ (
166
+ What is ScienceCast?
167
+ )
168
+ Demos
169
+ Demos
170
+ Replicate Toggle
171
+ Replicate
172
+ (
173
+ What is Replicate?
174
+ )
175
+ Spaces Toggle
176
+ Hugging Face Spaces
177
+ (
178
+ What is Spaces?
179
+ )
180
+ Spaces Toggle
181
+ TXYZ.AI
182
+ (
183
+ What is TXYZ.AI?
184
+ )
185
+ Related Papers
186
+ Recommenders and Search Tools
187
+ Link to Influence Flower
188
+ Influence Flower
189
+ (
190
+ What are Influence Flowers?
191
+ )
192
+ Core recommender toggle
193
+ CORE Recommender
194
+ (
195
+ What is CORE?
196
+ )
197
+ Author
198
+ Venue
199
+ Institution
200
+ Topic
201
+ About arXivLabs
202
+ arXivLabs: experimental projects with community collaborators
203
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
204
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
205
+ Have an idea for a project that will add value for arXiv's community?
206
+ Learn more about arXivLabs
207
+ .
208
+ Which authors of this paper are endorsers?
209
+ |
210
+ Disable MathJax
211
+ (
212
+ What is MathJax?
213
+ )
research/notes/231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la-2.md ADDED
@@ -0,0 +1,4095 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2310.04406] Language Agent Tree Search Unifies Reasoning Acting and Planning
3
+ in Language Models'
4
+ id: 231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la-2
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:40:44.608943Z'
8
+ source: https://ar5iv.labs.arxiv.org/html/2310.04406
9
+ source_domain: ar5iv.labs.arxiv.org
10
+ fetched_at: '2026-06-10T00:40:44.608803Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2310.04406] Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
20
+ Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
21
+ Andy Zhou
22
+ University of Illinois at Urbana-Champaign
23
+ AI@UIUC
24
+ Kai Yan
25
+ University of Illinois at Urbana-Champaign
26
+ Michal Shlapentokh-Rothman
27
+ University of Illinois at Urbana-Champaign
28
+ Haohan Wang
29
+ University of Illinois at Urbana-Champaign
30
+ Yu-Xiong Wang
31
+ University of Illinois at Urbana-Champaign
32
+ Abstract
33
+ While large language models (LLMs) have demonstrated impressive performance on a range of decision-making tasks, they rely on simple acting processes and fall short of broad deployment as autonomous agents. We introduce LATS (Language Agent Tree Search), a general framework that synergizes the capabilities of LLMs in planning, acting, and reasoning. Drawing inspiration from Monte Carlo tree search commonly used in model-based reinforcement learning, LATS employs LLMs as agents, value functions, and optimizers, repurposing their latent strengths for enhanced decision-making. What is crucial in this method is the use of an environment for external feedback, which offers a more deliberate and adaptive problem-solving mechanism that moves beyond the limitations of existing techniques. Our experimental evaluation across diverse domains, such as programming, HotPotQA, and WebShop, illustrates the applicability of LATS for decision-making while maintaining competitive reasoning performance. In particular, LATS achieves 94.4% for programming on HumanEval with GPT-4 and an average score of 75.9 for web browsing on WebShop with GPT-3.5, demonstrating the effectiveness and generality of our method.
34
+ 1
35
+ Introduction
36
+ General autonomous agents capable of reasoning and decision-making in a variety of environments
37
+ (Wooldridge & Jennings,
38
+ 1995
39
+ )
40
+ have been of longstanding interest in the field of artificial intelligence. While this has traditionally been studied in reinforcement learning, the recent rise of large language models (LLMs)
41
+ (Brown et al.,
42
+ 2020
43
+ ; Chowdhery et al.,
44
+ 2022
45
+ ; Touvron et al.,
46
+ 2023
47
+ ; OpenAI,
48
+ 2023
49
+ )
50
+ with strong reasoning and general adaptability offers an alternative paradigm. Not only have LLMs excelled on standard NLP tasks such as text summarization
51
+ (Nallapati et al.,
52
+ 2016
53
+ )
54
+ or natural language inference
55
+ (Bowman et al.,
56
+ 2015
57
+ )
58
+ , but they have been adapted to an increasingly diverse set of tasks that often require advanced common-sense reasoning or quantitative skills
59
+ (Cobbe et al.,
60
+ 2021
61
+ ; Saparov & He,
62
+ 2022
63
+ )
64
+ . LLMs are also capable of performing in complex environments that involve knowledge and reasoning, such as web navigation
65
+ (Yao et al.,
66
+ 2022
67
+ ; Deng et al.,
68
+ 2023
69
+ )
70
+ , tool-use
71
+ (Schick et al.,
72
+ 2023
73
+ )
74
+ , or open-ended games
75
+ (Fan et al.,
76
+ 2022
77
+ )
78
+ .
79
+ Figure 1:
80
+ An overview of LATS. LATS uses an external environment and self-reflection to improve reasoning and decision-making.
81
+ Reasoning and acting abilities have also been improved by prompting techniques that augment LLMs with feedback or observations from an external environment
82
+ (Yao et al.,
83
+ 2023b
84
+ ; Gao et al.,
85
+ 2022
86
+ ; Shinn et al.,
87
+ 2023
88
+ )
89
+ . This eliminates the need to rely entirely on the base abilities of the Language Model (LM), enhancing it through external tools or semantic feedback. Despite this strength, these methods are reflexive and fall short of humans’ deliberate and thoughtful decision-making characteristics to solve problems
90
+ (Sloman,
91
+ 1996
92
+ ; Evans,
93
+ 2010
94
+ )
95
+ . In particular, such methods fail to consider multiple reasoning paths or to plan ahead. Recent search-guided LLM works
96
+ (Xie et al.,
97
+ 2023
98
+ ; Yao et al.,
99
+ 2023a
100
+ ; Hao et al.,
101
+ 2023
102
+ )
103
+ address this issue by searching over multiple reasoning chains. While these methods enable planning, these methods operate in isolation and do not incorporate external feedback that can improve reasoning.
104
+ To help address these issues, we propose LATS (Language Agent Tree Search), a general framework for decision-making and reasoning with language models. LATS unifies LM planning, acting, and reasoning strategies by expanding ReAct
105
+ (Yao et al.,
106
+ 2023b
107
+ )
108
+ into a search over a combinatorial space of possible reasoning and acting steps. We adapt Monte Carlo tree search (MCTS) from model-based reinforcement learning
109
+ (Silver et al.,
110
+ 2017
111
+ ; Anthony et al.,
112
+ 2017
113
+ ; Jiang et al.,
114
+ 2018
115
+ )
116
+ to language agents, repurposing a pretrained LLM as an agent, value function, and optimizer. Utilizing the strong natural language understanding and in-context learning ability of modern LMs, we use text as an interface between each component of the framework, allowing LATS to adapt planning to environmental conditions without additional training. To the best of our knowledge,
117
+ LATS is the first framework that combines reasoning, acting, and planning to enhance LLMs
118
+ . Notably, LATS doubles the performance of GPT-3.5 on HotPotQA
119
+ (Yang et al.,
120
+ 2018
121
+ )
122
+ over ReAct
123
+ (Yao et al.,
124
+ 2023b
125
+ )
126
+ and raises the average score by
127
+ 22.1
128
+ 22.1
129
+ 22.1
130
+ on WebShop
131
+ (Yao et al.,
132
+ 2022
133
+ )
134
+ . When used with GPT-4, LATS achieves a
135
+ 94.4
136
+ 94.4
137
+ 94.4
138
+ Pass@1 rate for programming on HumanEval
139
+ (Chen et al.,
140
+ 2021
141
+ )
142
+ , setting the state of the art. To summarize, our
143
+ contributions
144
+ are the following:
145
+ •
146
+ We introduce an LM-based Monte Carlo tree search variant to deliberately construct the best trajectory from sampled actions, enabling more flexible and adaptive problem-solving compared to reflexive prompting methods. This is guided by heuristics from the LM.
147
+ •
148
+ By integrating external feedback and self-reflection, LATS enhances model sensibility and enables agents to learn from experience, surpassing reasoning-based search methods.
149
+ •
150
+ Through experiments across diverse domains like programming, interactive QA, and web navigation, we demonstrate the versatility of LATS in harnessing LLMs for autonomous reasoning and decision-making.
151
+ 2
152
+ Related Work
153
+ Approach
154
+ Reasoning
155
+ Acting
156
+ Planning
157
+ Self
158
+ External
159
+ Reflection
160
+ Memory
161
+ CoT
162
+ (Wei et al.,
163
+ 2022
164
+ )
165
+ ✓
166
+ ×
167
+ \times
168
+ ×
169
+ \times
170
+ ×
171
+ \times
172
+ ×
173
+ \times
174
+ ReAct
175
+ (Yao et al.,
176
+ 2023b
177
+ )
178
+ ✓
179
+ ✓
180
+ ×
181
+ \times
182
+ ×
183
+ \times
184
+ ×
185
+ \times
186
+ ToT
187
+ (Yao et al.,
188
+ 2023a
189
+ )
190
+ ✓
191
+ ×
192
+ \times
193
+ ✓
194
+ ✓
195
+ ✓
196
+ RAP
197
+ (Hao et al.,
198
+ 2023
199
+ )
200
+ ✓
201
+ ×
202
+ \times
203
+ ✓
204
+ ×
205
+ \times
206
+ ✓
207
+ Self-Refine
208
+ (Madaan et al.,
209
+ 2023
210
+ )
211
+ ✓
212
+ ×
213
+ \times
214
+ ×
215
+ \times
216
+ ✓
217
+ ×
218
+ \times
219
+ Beam Search
220
+ (Xie et al.,
221
+ 2023
222
+ )
223
+ ✓
224
+ ×
225
+ \times
226
+ ×
227
+ \times
228
+ ✓
229
+ ×
230
+ \times
231
+ Reflexion
232
+ (Shinn et al.,
233
+ 2023
234
+ )
235
+ ✓
236
+ ✓
237
+ ×
238
+ \times
239
+ ✓
240
+ ✓
241
+ LATS (Ours)
242
+ ✓
243
+ ✓
244
+ ✓
245
+ ✓
246
+ ✓
247
+ Table 1:
248
+ A summary of related work on reasoning, acting, and planning. LATS is the first work incorporating designs from all three domains, allowing use in all corresponding tasks. We refer to planning as the use of a search algorithm, self-reflection as the use of LM-generated feedback, and external memory as storaging past text context for future updates of solution.
249
+ a) Tree-of-Thoughts
250
+ b) Reasoning via Planning
251
+ c) Language Agent Tree Search
252
+ Figure 2:
253
+ An overview of the differences between LATS and recently proposed LM search algorithms ToT
254
+ (Yao et al.,
255
+ 2023a
256
+ )
257
+ and RAP
258
+ (Hao et al.,
259
+ 2023
260
+ )
261
+ . LATS leverages environmental feedback and self-reflection to further adapt search and improve performance.
262
+ LLMs for reasoning.
263
+ For LLMs, reasoning typically involves decomposing complex inputs into sequential intermediate steps towards a final answer
264
+ (Cobbe et al.,
265
+ 2021
266
+ )
267
+ , demonstrated with Chain-of-Thought (CoT) prompting
268
+ (Wei et al.,
269
+ 2022
270
+ )
271
+ and its variants
272
+ (Wei et al.,
273
+ 2022
274
+ ; Kojima et al.,
275
+ 2022
276
+ ; Wang et al.,
277
+ 2022
278
+ )
279
+ . However, these methods, which create chains autoregressively in a single step, often suffer from error propagation as the number of steps increases
280
+ (Guo et al.,
281
+ 2018
282
+ ; Chen et al.,
283
+ 2022b
284
+ )
285
+ due to compound errors. Various advancements aim to mitigate this issue; some approaches, such as Self-Consistency
286
+ (Wang et al.,
287
+ 2022
288
+ )
289
+ , employ majority voting over sampled chains, while others focus on multi-step decomposition, such as least-to-most prompting
290
+ (Zhou et al.,
291
+ 2022
292
+ )
293
+ , or use of external tools such as a scratchpad
294
+ (Nye et al.,
295
+ 2021
296
+ )
297
+ or compiler
298
+ (Gao et al.,
299
+ 2022
300
+ )
301
+ . Recently, CoT has been improved with search algorithms
302
+ (Yao et al.,
303
+ 2023a
304
+ ; Hao et al.,
305
+ 2023
306
+ ; Besta et al.,
307
+ 2023
308
+ )
309
+ that can sample trajectories more effectively. Tree-of-thought (ToT) prompting
310
+ (Yao et al.,
311
+ 2023a
312
+ )
313
+ uses DFS or BFS-based search guided by an LM-generated heuristic while Reasoning via Planning (RAP)
314
+ (Hao et al.,
315
+ 2023
316
+ )
317
+ uses MCTS with rollouts simulated by the LM. However, they rely solely on LM internal knowledge and cannot adapt to useful external feedback.
318
+ LLMs for acting.
319
+ The strong reasoning and common-sense abilities of LLMs have also been adapted for decision-making or acting tasks as a policy model in interactive environments. In the realm of robotics LLMs have been employed as high-level controllers of control policies
320
+ (Ahn et al.,
321
+ 2022
322
+ ; Huang et al.,
323
+ 2022
324
+ ; Driess et al.,
325
+ 2023
326
+ )
327
+ . Similar work
328
+ (Baker et al.,
329
+ 2022
330
+ ; Wang et al.,
331
+ 2023
332
+ ; Zhu et al.,
333
+ 2023
334
+ )
335
+ has also adapted LLM agents to complex multimodal games such as Minecraft
336
+ (Guss et al.,
337
+ 2019
338
+ ; Fan et al.,
339
+ 2022
340
+ )
341
+ . LLMs are particularly useful in text-based environments
342
+ (Liu et al.,
343
+ 2018
344
+ ; Shridhar et al.,
345
+ 2020
346
+ ; Liu et al.,
347
+ 2023
348
+ )
349
+ , where acting-based prompting techniques such as ReAct
350
+ (Yao et al.,
351
+ 2023b
352
+ )
353
+ have seen success. Similar to CoT, ReAct is limited by its simplicity and cannot effectively adapt to environment conditions. Many extensions have been proposed to address this, including Self-refine
354
+ (Madaan et al.,
355
+ 2023
356
+ )
357
+ and Reflexion
358
+ (Shinn et al.,
359
+ 2023
360
+ ; Yao et al.,
361
+ 2023c
362
+ )
363
+ , which uses self-reflection to enhance reasoning and decision-making, and AdaPlanner
364
+ (Sun et al.,
365
+ 2023
366
+ )
367
+ , which incorporates both positive and negative environmental feedback. However these methods focus on refining an individual plan or trajectory and do not consider alternative choices at each step. In addition, recent work
368
+ (Huang et al.,
369
+ 2023
370
+ )
371
+ has suggested LLMs cannot self-correct their internal reasoning, making it critical to use external feedback. Alternatively to pure decision-making environments, the reasoning and practical abilities of LLMs have been enhanced by access to external tools, such as APIs, search engines, calculators, or other models
372
+ (Schick et al.,
373
+ 2023
374
+ ; Shen et al.,
375
+ 2023
376
+ ; Surís et al.,
377
+ 2023
378
+ )
379
+ . Contrary to reasoning-based approaches, these methods have not been improved with planning, limiting their effectiveness. We summarize them in Tab.
380
+ 1
381
+ .
382
+ Tree-based search.
383
+ Tree-based search, where multiple branches of outcomes are explored during search, is widely used in many planning algorithms
384
+ (Świechowski et al.,
385
+ 2023
386
+ ; LaValle et al.,
387
+ 2001
388
+ )
389
+ and Reinforcement Learning (RL)
390
+ (Hafner et al.,
391
+ 2019
392
+ ; Du et al.,
393
+ 2023
394
+ ; Wu et al.,
395
+ 2023
396
+ )
397
+ algorithms for its good exploration-exploitation trade-off. Though tree-based search requires an environment model that can expand from arbitrary state
398
+ (Vodopivec et al.,
399
+ 2017
400
+ )
401
+ , which often requires extra training in RL
402
+ (Hafner et al.,
403
+ 2023
404
+ )
405
+ , such problem does not exist for LM tasks as we can conveniently backup to any state by setting the input to be the context and corresponding previous output by the LM. Thus, we work on the tree-based framework and use MCTS
406
+ (Świechowski et al.,
407
+ 2023
408
+ )
409
+ to fully release the potential of LMs, while avoiding the cost of training a value function over language descriptions by leveraging the in-context learning
410
+ (Brown et al.,
411
+ 2020
412
+ )
413
+ abilities of LLMs.
414
+ 3
415
+ Preliminaries
416
+ 3.1
417
+ Problem Setting and Prompting
418
+ Before describing LATS, we first define our problem and outline a few established methods that leverage large language models for reasoning or decision-making. In LM reasoning or decision making, we are given an input
419
+ x
420
+ 𝑥
421
+ x
422
+ in natural language and a pretrained language model
423
+ p
424
+ θ
425
+ ​
426
+ (
427
+ x
428
+ )
429
+ subscript
430
+ 𝑝
431
+ 𝜃
432
+ 𝑥
433
+ p_{\theta}(x)
434
+ parameterized by
435
+ θ
436
+ 𝜃
437
+ \theta
438
+ ; our goal is to generate a final output
439
+ y
440
+ ∼
441
+ p
442
+ θ
443
+ ​
444
+ (
445
+ x
446
+ )
447
+ similar-to
448
+ 𝑦
449
+ subscript
450
+ 𝑝
451
+ 𝜃
452
+ 𝑥
453
+ y\sim p_{\theta}(x)
454
+ corresponding to the answer (reasoning) or completes the task (decision-making). Both
455
+ x
456
+ 𝑥
457
+ x
458
+ and
459
+ y
460
+ 𝑦
461
+ y
462
+ are language
463
+ sequences
464
+ , which are comprised of a list of
465
+ tokens
466
+ (the basic elements of natural language, often words), denoted as
467
+ x
468
+ =
469
+ (
470
+ x
471
+ ​
472
+ [
473
+ 1
474
+ ]
475
+ ,
476
+ …
477
+ ,
478
+ x
479
+ ​
480
+ [
481
+ n
482
+ ]
483
+ )
484
+ 𝑥
485
+ 𝑥
486
+ delimited-[]
487
+ 1
488
+ …
489
+ 𝑥
490
+ delimited-[]
491
+ 𝑛
492
+ x=(x[1],\dots,x[n])
493
+ and
494
+ y
495
+ =
496
+ (
497
+ y
498
+ ​
499
+ [
500
+ 1
501
+ ]
502
+ ,
503
+ …
504
+ ,
505
+ y
506
+ ​
507
+ [
508
+ n
509
+ ]
510
+ )
511
+ 𝑦
512
+ 𝑦
513
+ delimited-[]
514
+ 1
515
+ …
516
+ 𝑦
517
+ delimited-[]
518
+ 𝑛
519
+ y=(y[1],\dots,y[n])
520
+ . The LM decodes text autoregressively, i.e., without other inputs, the probability for an LM to generate a sequence
521
+ x
522
+ 𝑥
523
+ x
524
+ is given by
525
+ p
526
+ θ
527
+ ​
528
+ (
529
+ x
530
+ )
531
+ =
532
+ ∏
533
+ i
534
+ =
535
+ 1
536
+ n
537
+ p
538
+ θ
539
+ ​
540
+ (
541
+ x
542
+ ​
543
+ [
544
+ i
545
+ ]
546
+ |
547
+ x
548
+ ​
549
+ [
550
+ 1
551
+ ​
552
+ …
553
+ ​
554
+ i
555
+ −
556
+ 1
557
+ ]
558
+ )
559
+ subscript
560
+ 𝑝
561
+ 𝜃
562
+ 𝑥
563
+ superscript
564
+ subscript
565
+ product
566
+ 𝑖
567
+ 1
568
+ 𝑛
569
+ subscript
570
+ 𝑝
571
+ 𝜃
572
+ conditional
573
+ 𝑥
574
+ delimited-[]
575
+ 𝑖
576
+ 𝑥
577
+ delimited-[]
578
+ 1
579
+ …
580
+ 𝑖
581
+ 1
582
+ p_{\theta}(x)=\prod_{i=1}^{n}p_{\theta}(x[i]|x[1\dots i-1])
583
+ . Usually, to improve the LM,
584
+ prompts
585
+ are provided along with the input
586
+ x
587
+ 𝑥
588
+ x
589
+ , which are specific instructions or few-shot input-output examples. We denote the generic process where an input
590
+ x
591
+ 𝑥
592
+ x
593
+ is transformed into an output
594
+ y
595
+ 𝑦
596
+ y
597
+ by LM:
598
+ y
599
+ ∼
600
+ p
601
+ θ
602
+ ​
603
+ (
604
+ y
605
+ |
606
+ prompt
607
+ I
608
+ ​
609
+ O
610
+ ​
611
+ (
612
+ x
613
+ )
614
+ )
615
+ similar-to
616
+ 𝑦
617
+ subscript
618
+ 𝑝
619
+ 𝜃
620
+ conditional
621
+ 𝑦
622
+ subscript
623
+ prompt
624
+ 𝐼
625
+ 𝑂
626
+ 𝑥
627
+ y\sim p_{\theta}(y|\texttt{prompt}_{IO}(x))
628
+ , where
629
+ prompt
630
+ I
631
+ ​
632
+ O
633
+ ​
634
+ (
635
+ x
636
+ )
637
+ subscript
638
+ prompt
639
+ 𝐼
640
+ 𝑂
641
+ 𝑥
642
+ \texttt{prompt}_{IO}(x)
643
+ denotes the input
644
+ x
645
+ 𝑥
646
+ x
647
+ .
648
+ Chain-of-thought (CoT) Prompting
649
+ (Wei et al.,
650
+ 2022
651
+ )
652
+ was introduced to cater to scenarios where direct mapping from
653
+ x
654
+ 𝑥
655
+ x
656
+ to
657
+ y
658
+ 𝑦
659
+ y
660
+ is intricate, such as when
661
+ x
662
+ 𝑥
663
+ x
664
+ is from a mathematical query or challenging question. This method hinges on creating
665
+ thoughts
666
+ z
667
+ 1
668
+ ,
669
+ …
670
+ ,
671
+ z
672
+ n
673
+ subscript
674
+ 𝑧
675
+ 1
676
+ …
677
+ subscript
678
+ 𝑧
679
+ 𝑛
680
+ z_{1},\dots,z_{n}
681
+ that act as stepping stones between
682
+ x
683
+ 𝑥
684
+ x
685
+ and
686
+ y
687
+ 𝑦
688
+ y
689
+ ; each thought
690
+ z
691
+ i
692
+ subscript
693
+ 𝑧
694
+ 𝑖
695
+ z_{i}
696
+ is a language sequence. To employ CoT prompting, thoughts are extracted sequentially as
697
+ z
698
+ i
699
+ ∼
700
+ p
701
+ θ
702
+ C
703
+ ​
704
+ o
705
+ ​
706
+ T
707
+ ​
708
+ (
709
+ z
710
+ i
711
+ |
712
+ x
713
+ ,
714
+ z
715
+ 1
716
+ ​
717
+ ⋯
718
+ ​
719
+ i
720
+ −
721
+ 1
722
+ )
723
+ similar-to
724
+ subscript
725
+ 𝑧
726
+ 𝑖
727
+ superscript
728
+ subscript
729
+ 𝑝
730
+ 𝜃
731
+ 𝐶
732
+ 𝑜
733
+ 𝑇
734
+ conditional
735
+ subscript
736
+ 𝑧
737
+ 𝑖
738
+ 𝑥
739
+ subscript
740
+ 𝑧
741
+ 1
742
+ ⋯
743
+ 𝑖
744
+ 1
745
+ z_{i}\sim p_{\theta}^{CoT}(z_{i}|x,z_{1\cdots i-1})
746
+ , with the final output being
747
+ y
748
+ ∼
749
+ p
750
+ θ
751
+ C
752
+ ​
753
+ o
754
+ ​
755
+ T
756
+ ​
757
+ (
758
+ y
759
+ |
760
+ x
761
+ ,
762
+ z
763
+ 1
764
+ ​
765
+ ⋯
766
+ ​
767
+ n
768
+ )
769
+ similar-to
770
+ 𝑦
771
+ superscript
772
+ subscript
773
+ 𝑝
774
+ 𝜃
775
+ 𝐶
776
+ 𝑜
777
+ 𝑇
778
+ conditional
779
+ 𝑦
780
+ 𝑥
781
+ subscript
782
+ 𝑧
783
+ 1
784
+ ⋯
785
+ 𝑛
786
+ y\sim p_{\theta}^{CoT}(y|x,z_{1\cdots n})
787
+ .
788
+ Tree-of-thought (ToT) Prompting
789
+ (Yao et al.,
790
+ 2023a
791
+ )
792
+ extends CoT prompting by exploring multiple reasoning paths over thoughts. It frames problems as a search over a tree where each node
793
+ s
794
+ =
795
+ [
796
+ x
797
+ ,
798
+ z
799
+ 1
800
+ ⋅
801
+ i
802
+ ]
803
+ 𝑠
804
+ 𝑥
805
+ subscript
806
+ 𝑧
807
+ ⋅
808
+ 1
809
+ 𝑖
810
+ s=[x,z_{1\cdot i}]
811
+ represents a partial solution state comprising the original input
812
+ x
813
+ 𝑥
814
+ x
815
+ and thought sequence
816
+ z
817
+ 1
818
+ ​
819
+ ⋯
820
+ ​
821
+ i
822
+ subscript
823
+ 𝑧
824
+ 1
825
+ ⋯
826
+ 𝑖
827
+ z_{1\cdots i}
828
+ . Thoughts
829
+ z
830
+ i
831
+ subscript
832
+ 𝑧
833
+ 𝑖
834
+ z_{i}
835
+ are generated by proposal or sampling with CoT
836
+ z
837
+ i
838
+ ∼
839
+ p
840
+ θ
841
+ C
842
+ ​
843
+ o
844
+ ​
845
+ T
846
+ ​
847
+ (
848
+ z
849
+ i
850
+ |
851
+ x
852
+ ,
853
+ z
854
+ 1
855
+ ​
856
+ ⋯
857
+ ​
858
+ i
859
+ −
860
+ 1
861
+ )
862
+ similar-to
863
+ subscript
864
+ 𝑧
865
+ 𝑖
866
+ superscript
867
+ subscript
868
+ 𝑝
869
+ 𝜃
870
+ 𝐶
871
+ 𝑜
872
+ 𝑇
873
+ conditional
874
+ subscript
875
+ 𝑧
876
+ 𝑖
877
+ 𝑥
878
+ subscript
879
+ 𝑧
880
+ 1
881
+ ⋯
882
+ 𝑖
883
+ 1
884
+ z_{i}\sim p_{\theta}^{CoT}(z_{i}|x,z_{1\cdots i-1})
885
+ . Deliberate search algorithms like breadth-first or depth-first search are used to systematically explore the tree, guided by heuristics based on language model evaluations
886
+ V
887
+ ​
888
+ (
889
+ s
890
+ )
891
+ 𝑉
892
+ 𝑠
893
+ V(s)
894
+ of each state.
895
+ Reasoning via Planning
896
+ (RAP)
897
+ (Hao et al.,
898
+ 2023
899
+ )
900
+ is similar to ToT, except that MCTS is used over DFS or BFS. Heuristics are designed from an LM, such as the likelihood or confidence of an action, and the LM is used as a world model to predict subsequent states during the simulation step.
901
+ ReAct
902
+ (Yao et al.,
903
+ 2023b
904
+ )
905
+ extends language models to tasks where the mapping from
906
+ x
907
+ 𝑥
908
+ x
909
+ to
910
+ y
911
+ 𝑦
912
+ y
913
+ is enhanced by or requires interactions with an external environment, such as a game or API. This technique constructs an action space
914
+ A
915
+ ^
916
+ =
917
+ A
918
+ ∪
919
+ Z
920
+ ^
921
+ 𝐴
922
+ 𝐴
923
+ 𝑍
924
+ \hat{A}=A\cup Z
925
+ that adds permissible actions
926
+ a
927
+ 𝑎
928
+ a
929
+ to the reasoning traces
930
+ z
931
+ 𝑧
932
+ z
933
+ from CoT. Observations
934
+ o
935
+ 𝑜
936
+ o
937
+ from the environment are used to improve both reasoning and acting. To solve problems with ReAct, after each observation, actions are generated from
938
+ p
939
+ θ
940
+ subscript
941
+ 𝑝
942
+ 𝜃
943
+ p_{\theta}
944
+ sequentially as
945
+ a
946
+ i
947
+ ∼
948
+ p
949
+ θ
950
+ R
951
+ ​
952
+ e
953
+ ​
954
+ A
955
+ ​
956
+ c
957
+ ​
958
+ t
959
+ ​
960
+ (
961
+ a
962
+ i
963
+ |
964
+ x
965
+ ,
966
+ o
967
+ 1
968
+ ​
969
+ ⋯
970
+ ​
971
+ i
972
+ −
973
+ 1
974
+ ,
975
+ a
976
+ 1
977
+ ​
978
+ ⋯
979
+ ​
980
+ i
981
+ −
982
+ 1
983
+ )
984
+ similar-to
985
+ subscript
986
+ 𝑎
987
+ 𝑖
988
+ superscript
989
+ subscript
990
+ 𝑝
991
+ 𝜃
992
+ 𝑅
993
+ 𝑒
994
+ 𝐴
995
+ 𝑐
996
+ 𝑡
997
+ conditional
998
+ subscript
999
+ 𝑎
1000
+ 𝑖
1001
+ 𝑥
1002
+ subscript
1003
+ 𝑜
1004
+ 1
1005
+ ⋯
1006
+ 𝑖
1007
+ 1
1008
+ subscript
1009
+ 𝑎
1010
+ 1
1011
+ ⋯
1012
+ 𝑖
1013
+ 1
1014
+ a_{i}\sim p_{\theta}^{ReAct}(a_{i}|x,o_{1\cdots i-1},a_{1\cdots i-1})
1015
+ , with the final output being
1016
+ y
1017
+ ∼
1018
+ p
1019
+ θ
1020
+ R
1021
+ ​
1022
+ e
1023
+ ​
1024
+ A
1025
+ ​
1026
+ c
1027
+ ​
1028
+ t
1029
+ ​
1030
+ (
1031
+ y
1032
+ |
1033
+ x
1034
+ ,
1035
+ o
1036
+ 1
1037
+ ​
1038
+ ⋯
1039
+ ​
1040
+ n
1041
+ ,
1042
+ a
1043
+ 1
1044
+ ​
1045
+ ⋯
1046
+ ​
1047
+ n
1048
+ )
1049
+ similar-to
1050
+ 𝑦
1051
+ superscript
1052
+ subscript
1053
+ 𝑝
1054
+ 𝜃
1055
+ 𝑅
1056
+ 𝑒
1057
+ 𝐴
1058
+ 𝑐
1059
+ 𝑡
1060
+ conditional
1061
+ 𝑦
1062
+ 𝑥
1063
+ subscript
1064
+ 𝑜
1065
+ 1
1066
+ ⋯
1067
+ 𝑛
1068
+ subscript
1069
+ 𝑎
1070
+ 1
1071
+ ⋯
1072
+ 𝑛
1073
+ y\sim p_{\theta}^{ReAct}(y~{}|~{}x,o_{1\cdots n},a_{1\cdots n})
1074
+ .
1075
+ While the previously described prompting techniques improve LM performance on reasoning tasks, they falter on difficult tasks that involve multifaceted decision-making due to several shortcomings: 1)
1076
+ Flexibility
1077
+ : Base prompting methods (CoT or ReAct) autoregressively sample from the LM, neglecting potential alternative continuations from specific states. 2)
1078
+ Sensibility
1079
+ : Reasoning-based methods (CoT, RAP, or ToT) rely solely on the internal representations of the LM and cannot consider external observations. This dependency risks fact hallucination and error propagation while setting a performance ceiling. 3)
1080
+ Adaptability
1081
+ : Current planning frameworks (RAP or ToT) use simple search algorithms such as BFS or cannot leverage environmental feedback to improve planning. Additionally, the agent is static and cannot reuse previous experience or learn from trial and error. While RAP also adopts MCTS, it is constrained to tasks where the LM can become a world model and accurately predict states. These shortcomings limit the ability of LMs to be deployed as general problem-solving agents and form the motivation for LATS.
1082
+ 3.2
1083
+ Monte-Carlo Tree Search (MCTS)
1084
+ Monte-Carlo Tree Search (MCTS) is a heuristic search algorithm that is proved successful on many decision-making environments such as Atari
1085
+ (Ye et al.,
1086
+ 2021
1087
+ )
1088
+ and Go
1089
+ (Silver et al.,
1090
+ 2016
1091
+ )
1092
+ . MCTS builds a decision tree where every node in the tree is a state and edge is an action. MCTS runs for
1093
+ k
1094
+ 𝑘
1095
+ k
1096
+ episodes; for each episode, it starts from the root (i.e., initial state) and iteratively conducts two steps to expand the tree: 1)
1097
+ Expansion
1098
+ , where multiple children states
1099
+ s
1100
+ 𝑠
1101
+ s
1102
+ are explored from the current parent state
1103
+ p
1104
+ 𝑝
1105
+ p
1106
+ by sampling
1107
+ n
1108
+ 𝑛
1109
+ n
1110
+ actions, and 2)
1111
+ Selection
1112
+ , where the children with the highest UCT
1113
+ (Upper Confidence bounds applied to Trees)
1114
+ (Kocsis & Szepesvári,
1115
+ 2006
1116
+ )
1117
+ value is selected by the next iteration. The UCT of a child state
1118
+ s
1119
+ 𝑠
1120
+ s
1121
+ is calculated as follows:
1122
+ U
1123
+ ​
1124
+ C
1125
+ ​
1126
+ T
1127
+ ​
1128
+ (
1129
+ s
1130
+ )
1131
+ =
1132
+ V
1133
+ ​
1134
+ (
1135
+ s
1136
+ )
1137
+ +
1138
+ w
1139
+ ​
1140
+ ln
1141
+ ⁡
1142
+ N
1143
+ ​
1144
+ (
1145
+ p
1146
+ )
1147
+ N
1148
+ ​
1149
+ (
1150
+ s
1151
+ )
1152
+ ,
1153
+ 𝑈
1154
+ 𝐶
1155
+ 𝑇
1156
+ 𝑠
1157
+ 𝑉
1158
+ 𝑠
1159
+ 𝑤
1160
+ 𝑁
1161
+ 𝑝
1162
+ 𝑁
1163
+ 𝑠
1164
+ UCT(s)=V(s)+w\sqrt{\frac{\ln N(p)}{N(s)}},
1165
+ (1)
1166
+ where
1167
+ N
1168
+ ​
1169
+ (
1170
+ s
1171
+ )
1172
+ 𝑁
1173
+ 𝑠
1174
+ N(s)
1175
+ is the number of visits to a node
1176
+ s
1177
+ 𝑠
1178
+ s
1179
+ ,
1180
+ V
1181
+ ​
1182
+ (
1183
+ s
1184
+ )
1185
+ 𝑉
1186
+ 𝑠
1187
+ V(s)
1188
+ is the value function (expected return) from the subtree of
1189
+ s
1190
+ 𝑠
1191
+ s
1192
+ ,
1193
+ w
1194
+ 𝑤
1195
+ w
1196
+ is the exploration weight, and
1197
+ p
1198
+ 𝑝
1199
+ p
1200
+ is the parent node of
1201
+ s
1202
+ 𝑠
1203
+ s
1204
+ . The child node with the highest UCT value is selected for expansion in the next iteration. When the end of an episode is reached, a
1205
+ backpropagation
1206
+ is carried out: the return
1207
+ r
1208
+ 𝑟
1209
+ r
1210
+ is used for updating every
1211
+ V
1212
+ ​
1213
+ (
1214
+ s
1215
+ )
1216
+ 𝑉
1217
+ 𝑠
1218
+ V(s)
1219
+ along the path
1220
+ with the formula
1221
+ V
1222
+ ​
1223
+ (
1224
+ s
1225
+ )
1226
+ =
1227
+ V
1228
+ old
1229
+ ​
1230
+ (
1231
+ s
1232
+ )
1233
+ ​
1234
+ (
1235
+ N
1236
+ ​
1237
+ (
1238
+ s
1239
+ )
1240
+ −
1241
+ 1
1242
+ )
1243
+ +
1244
+ r
1245
+ N
1246
+ ​
1247
+ (
1248
+ s
1249
+ )
1250
+ 𝑉
1251
+ 𝑠
1252
+ subscript
1253
+ 𝑉
1254
+ old
1255
+ 𝑠
1256
+ 𝑁
1257
+ 𝑠
1258
+ 1
1259
+ 𝑟
1260
+ 𝑁
1261
+ 𝑠
1262
+ V(s)=\frac{V_{\text{old}}(s)(N(s)-1)+r}{N(s)}
1263
+ , where
1264
+ V
1265
+ old
1266
+ ​
1267
+ (
1268
+ s
1269
+ )
1270
+ subscript
1271
+ 𝑉
1272
+ old
1273
+ 𝑠
1274
+ V_{\text{old}}(s)
1275
+ is the old value function. Normally, the major shortcoming of MCTS is that it requires an environment model to undo previous steps and form a searching tree, which is often a strong assumption. However, such a limitation does not exist for LMs, as we can conveniently reset to any step by simply copy-pasting historical text input. Such a special property is the key motivation of our work.
1276
+ 4
1277
+ Unifying Planning, Reasoning, and Acting
1278
+ 4.1
1279
+ LM Agent
1280
+ LATS supports sequential reasoning or decision-making tasks on the basis of ReAct. At time step
1281
+ t
1282
+ 𝑡
1283
+ t
1284
+ , an agent receives an observation
1285
+ o
1286
+ t
1287
+ ∈
1288
+ O
1289
+ subscript
1290
+ 𝑜
1291
+ 𝑡
1292
+ 𝑂
1293
+ o_{t}\in O
1294
+ from the environment and takes an action
1295
+ a
1296
+ t
1297
+ ∈
1298
+ A
1299
+ subscript
1300
+ 𝑎
1301
+ 𝑡
1302
+ 𝐴
1303
+ a_{t}\in A
1304
+ following some policy
1305
+ π
1306
+ ​
1307
+ (
1308
+ a
1309
+ t
1310
+ |
1311
+ x
1312
+ ,
1313
+ o
1314
+ 1
1315
+ ​
1316
+ ⋯
1317
+ ​
1318
+ i
1319
+ −
1320
+ 1
1321
+ ,
1322
+ a
1323
+ 1
1324
+ ​
1325
+ ⋯
1326
+ ​
1327
+ i
1328
+ −
1329
+ 1
1330
+ )
1331
+ 𝜋
1332
+ conditional
1333
+ subscript
1334
+ 𝑎
1335
+ 𝑡
1336
+ 𝑥
1337
+ subscript
1338
+ 𝑜
1339
+ 1
1340
+ ⋯
1341
+ 𝑖
1342
+ 1
1343
+ subscript
1344
+ 𝑎
1345
+ 1
1346
+ ⋯
1347
+ 𝑖
1348
+ 1
1349
+ \pi(a_{t}|x,o_{1\cdots i-1},a_{1\cdots i-1})
1350
+ , where
1351
+ x
1352
+ 𝑥
1353
+ x
1354
+ consists of the task instruction and a number of few-shot examples. We initialize the agent with
1355
+ p
1356
+ θ
1357
+ subscript
1358
+ 𝑝
1359
+ 𝜃
1360
+ p_{\theta}
1361
+ to leverage the useful language representations of an LM as a base decision-maker. We follow the ReAct instantiation in which the action space
1362
+ A
1363
+ ^
1364
+ =
1365
+ A
1366
+ ∪
1367
+ Z
1368
+ ^
1369
+ 𝐴
1370
+ 𝐴
1371
+ 𝑍
1372
+ \hat{A}=A\cup Z
1373
+ consists of both the space of permissible actions
1374
+ A
1375
+ 𝐴
1376
+ A
1377
+ and language space of reasoning traces
1378
+ Z
1379
+ 𝑍
1380
+ Z
1381
+ . Actions directly affect the environment and result in observation, while thoughts are used to formalize decisions by organizing information, planning future actions, or injecting internal knowledge. The exact instantiation of the action space depends on the particular environment; for decision-making tasks actions might consist of commands on a website while for reasoning tasks the action space might be limited to a few external tools or APIs.
1382
+ Instead of greedily decoding one trajectory or solution, we sample
1383
+ n
1384
+ 𝑛
1385
+ n
1386
+ actions from
1387
+ p
1388
+ θ
1389
+ subscript
1390
+ 𝑝
1391
+ 𝜃
1392
+ p_{\theta}
1393
+ using the current state. This is based on the intuition that for complex decision-making tasks, there is likely to be a range of potential trajectories or reasoning paths that are correct
1394
+ (Evans,
1395
+ 2010
1396
+ )
1397
+ . Sampling a diverse set of candidates at each step mitigates the stochastic nature of LM text generation and enables greater exploration in both the decision-making and reasoning space. We wrap
1398
+ p
1399
+ θ
1400
+ subscript
1401
+ 𝑝
1402
+ 𝜃
1403
+ p_{\theta}
1404
+ within our proposed search algorithm to deliberately construct the best trajectory from sampled actions.
1405
+ 4.2
1406
+ LATS
1407
+ Figure 3:
1408
+ An overview of the six operations of LATS. A node is
1409
+ selected
1410
+ ,
1411
+ expanded
1412
+ ,
1413
+ evaluated
1414
+ , then
1415
+ simulated
1416
+ until a terminal node is reached, then the resulting value is
1417
+ backpropagated
1418
+ . If the trajectory fails, a
1419
+ reflection
1420
+ is generated and used as additional context for future trials. These operations are performed in succession until the budget is reached or task is successful.
1421
+ The main component of LATS is a search algorithm that controls the overall problem-solving process with deliberate planning. To find the most promising trajectory and systemically balance exploration with exploitation, we adopt a variant of Monte Carlo Tree Search (MCTS) that frames decision-making as a tree search, in which each node
1422
+ s
1423
+ =
1424
+ [
1425
+ x
1426
+ ,
1427
+ a
1428
+ 1
1429
+ ​
1430
+ ⋯
1431
+ ​
1432
+ i
1433
+ ,
1434
+ o
1435
+ 1
1436
+ ​
1437
+ ⋯
1438
+ ​
1439
+ i
1440
+ ]
1441
+ 𝑠
1442
+ 𝑥
1443
+ subscript
1444
+ 𝑎
1445
+ 1
1446
+ ⋯
1447
+ 𝑖
1448
+ subscript
1449
+ 𝑜
1450
+ 1
1451
+ ⋯
1452
+ 𝑖
1453
+ s=[x,a_{1\cdots i},o_{1\cdots i}]
1454
+ represents a state comprising the original input
1455
+ x
1456
+ 𝑥
1457
+ x
1458
+ , action sequence
1459
+ a
1460
+ 1
1461
+ ⋅
1462
+ i
1463
+ subscript
1464
+ 𝑎
1465
+ ⋅
1466
+ 1
1467
+ 𝑖
1468
+ a_{1\cdot i}
1469
+ , and observation sequence
1470
+ o
1471
+ 1
1472
+ ⋅
1473
+ i
1474
+ subscript
1475
+ 𝑜
1476
+ ⋅
1477
+ 1
1478
+ 𝑖
1479
+ o_{1\cdot i}
1480
+ .
1481
+ To adapt MCTS for language agents, LATS repurposes
1482
+ p
1483
+ θ
1484
+ subscript
1485
+ 𝑝
1486
+ 𝜃
1487
+ p_{\theta}
1488
+ as an agent, state evaluator, and feedback generator, leveraging the useful language priors of modern LMs to facilitate planning. While standard MCTS and RAP
1489
+ Hao et al. (
1490
+ 2023
1491
+ )
1492
+ rely on internal dynamics models to facilitate simulation, LATS is model-free and uses environment interaction. LATS consists of a series of operations,
1493
+ selection, expansion, evaluation, simulation, backpropagation, and reflection
1494
+ , performed in succession until the task is successfully completed or a computational limit is reached. The full psuedocode of LATS can be found in Sec.
1495
+ A
1496
+ in the Appendix.
1497
+ Selection.
1498
+ In the first operation, the algorithm identifies a segment of the current tree most suitable for subsequent expansion. Starting from the root node, denoted as the initial state
1499
+ s
1500
+ 0
1501
+ subscript
1502
+ 𝑠
1503
+ 0
1504
+ s_{0}
1505
+ , a child node is selected at each tree level until a leaf node is reached. To balance exploration and exploitation, we use the UCT algorithm as shown in Eq.
1506
+ 1
1507
+ .
1508
+ Expansion.
1509
+ After selecting a node, the second operation expands the tree by sampling
1510
+ n
1511
+ 𝑛
1512
+ n
1513
+ actions from
1514
+ p
1515
+ θ
1516
+ subscript
1517
+ 𝑝
1518
+ 𝜃
1519
+ p_{\theta}
1520
+ , as described in the prior section. The environment receives each action and returns corresponding feedback as an observation. This results in
1521
+ n
1522
+ 𝑛
1523
+ n
1524
+ new child nodes added to the tree. This tree is stored in an external long-term memory structure.
1525
+ Evaluation.
1526
+ The third operation assigns a scalar value to each new child node to be used for selection and backpropagation. This value effectively quantifies the agent’s progress in task completion, serving as a heuristic to steer the search algorithm towards the most promising regions of the tree. Following
1527
+ Yao et al. (
1528
+ 2023a
1529
+ )
1530
+ we repurpose
1531
+ p
1532
+ θ
1533
+ subscript
1534
+ 𝑝
1535
+ 𝜃
1536
+ p_{\theta}
1537
+ into a value function by prompting it to reason about a given state. To obtain a scalar value, we instruct
1538
+ p
1539
+ θ
1540
+ subscript
1541
+ 𝑝
1542
+ 𝜃
1543
+ p_{\theta}
1544
+ to end its reasoning trace with a score indicating the correctness of the trajectory. This method offers enhanced flexibility over programmed heuristics
1545
+ (Campbell et al.,
1546
+ 2002
1547
+ )
1548
+ and greater efficiency than learned heuristics
1549
+ (Silver et al.,
1550
+ 2017
1551
+ )
1552
+ .
1553
+ Simulation.
1554
+ The fourth operation expands the currently selected node until a terminal state is reached. At each depth level we sample and evaluate nodes with the same operations, but prioritize nodes of highest value. Reaching a terminal state provides objective feedback on the correctness of a trajectory. If the task is completed successfully, then LATS terminates the search. If the solution is partially successful or unsuccessful, then we perform two additional operations as described below.
1555
+ Backpropagation.
1556
+ This operation updates the values of the tree based on the outcome of a trajectory. For each node
1557
+ s
1558
+ 0
1559
+ ,
1560
+ s
1561
+ 1
1562
+ ,
1563
+ …
1564
+ ,
1565
+ s
1566
+ n
1567
+ subscript
1568
+ 𝑠
1569
+ 0
1570
+ subscript
1571
+ 𝑠
1572
+ 1
1573
+ …
1574
+ subscript
1575
+ 𝑠
1576
+ 𝑛
1577
+ s_{0},s_{1},\dots,s_{n}
1578
+ in the trajectory from root (initial state
1579
+ s
1580
+ 0
1581
+ subscript
1582
+ 𝑠
1583
+ 0
1584
+ s_{0}
1585
+ ) of the searching tree to leaf (terminal state
1586
+ s
1587
+ n
1588
+ subscript
1589
+ 𝑠
1590
+ 𝑛
1591
+ s_{n}
1592
+ ), its value is updated to reflect the outcome of the simulation by
1593
+ N
1594
+ ​
1595
+ (
1596
+ s
1597
+ i
1598
+ )
1599
+ =
1600
+ N
1601
+ old
1602
+ ​
1603
+ (
1604
+ s
1605
+ i
1606
+ )
1607
+ +
1608
+ 1
1609
+ 𝑁
1610
+ subscript
1611
+ 𝑠
1612
+ 𝑖
1613
+ subscript
1614
+ 𝑁
1615
+ old
1616
+ subscript
1617
+ 𝑠
1618
+ 𝑖
1619
+ 1
1620
+ N(s_{i})=N_{\text{old}}(s_{i})+1
1621
+ and
1622
+ V
1623
+ ​
1624
+ (
1625
+ s
1626
+ i
1627
+ )
1628
+ =
1629
+ r
1630
+ +
1631
+ N
1632
+ old
1633
+ ​
1634
+ (
1635
+ s
1636
+ i
1637
+ )
1638
+ ​
1639
+ V
1640
+ old
1641
+ ​
1642
+ (
1643
+ s
1644
+ i
1645
+ )
1646
+ N
1647
+ ​
1648
+ (
1649
+ s
1650
+ i
1651
+ )
1652
+ 𝑉
1653
+ subscript
1654
+ 𝑠
1655
+ 𝑖
1656
+ 𝑟
1657
+ subscript
1658
+ 𝑁
1659
+ old
1660
+ subscript
1661
+ 𝑠
1662
+ 𝑖
1663
+ subscript
1664
+ 𝑉
1665
+ old
1666
+ subscript
1667
+ 𝑠
1668
+ 𝑖
1669
+ 𝑁
1670
+ subscript
1671
+ 𝑠
1672
+ 𝑖
1673
+ V(s_{i})=\frac{r+N_{\text{old}}(s_{i})V_{\text{old}}(s_{i})}{N(s_{i})}
1674
+ , where
1675
+ r
1676
+ 𝑟
1677
+ r
1678
+ is the return and
1679
+ N
1680
+ old
1681
+ ,
1682
+ V
1683
+ old
1684
+ subscript
1685
+ 𝑁
1686
+ old
1687
+ subscript
1688
+ 𝑉
1689
+ old
1690
+ N_{\text{old}},V_{\text{old}}
1691
+ are the old number of visits and value function. These updated values are used in the UCT formula (Eq.
1692
+ 1
1693
+ ) to guide the selection of the next node for exploration.
1694
+ Reflection.
1695
+ In addition to the environmental feedback, we also leverage
1696
+ self-reflection
1697
+ to further refine the decision-making process
1698
+ (Shinn et al.,
1699
+ 2023
1700
+ ; Madaan et al.,
1701
+ 2023
1702
+ )
1703
+ . Upon encountering an unsuccessful terminal node,
1704
+ p
1705
+ θ
1706
+ subscript
1707
+ 𝑝
1708
+ 𝜃
1709
+ p_{\theta}
1710
+ is prompted with the trajectory and final reward to provide a verbal self-reflection that summarizes the errors in the reasoning or acting process and proposes superior alternatives. We store both failed trajectories and corresponding reflections in the memory. In subsequent iterations, these are integrated as additional context to the agent and value function, refining both through in-context learning. This imparts a semantic gradient signal more useful than a scalar value, enabling the agent to learn from trial and error without the cost of expensive optimization processes such as reinforcement learning.
1711
+ Conceptually, LATS has the following advantages as a general framework for reasoning and decision-making with LM agents.
1712
+ (1)
1713
+ Generality
1714
+ : LATS supports both reasoning and decision-making tasks by defining a shared space of thoughts and actions. (2)
1715
+ Deliberate
1716
+ : The use of MCTS and LM value function ensures a principled search that selects options with high value while exploring promising alternatives. (3)
1717
+ Adaptability
1718
+ : LATS is designed around the use of external feedback through observations and self-reflection, enabling greater adaptation during problem-solving. (4)
1719
+ Flexibility
1720
+ : LATS can accommodate different scenarios, environments, and resource stipulations by modifying state design and tree dimensions. (5)
1721
+ Modularity
1722
+ : The base LM agent, reflection generator, and value function can be independently altered and adapted to individual LM properties.
1723
+ 5
1724
+ Experiments
1725
+ To demonstrate the general applicability of LATS, we evaluate our method on a variety of decision-making domains that requires both reasoning and acting ability: programming
1726
+ (Chen et al.,
1727
+ 2021
1728
+ ; Austin et al.,
1729
+ 2021
1730
+ )
1731
+ , HotPotQA
1732
+ (Yang et al.,
1733
+ 2018
1734
+ )
1735
+ , and WebShop
1736
+ (Yao et al.,
1737
+ 2022
1738
+ )
1739
+ .
1740
+ 5.1
1741
+ HotPotQA
1742
+ For a task that can be approached with both reasoning-based and acting-based strategies, we consider HotPotQA
1743
+ (Yang et al.,
1744
+ 2018
1745
+ )
1746
+ , a multi-hop question-answering benchmark that requires retrieval over two or more Wikipedia passages. For the action space, in addition to LM thoughts we follow the setup from
1747
+ Yao et al. (
1748
+ 2023b
1749
+ )
1750
+ , which provides the agent with API calls to search and lookup information. The output of these API calls and self-generated reflections form the observation space. We use a subset of 100 questions and three few-shot examples for each method. For ToT, we use DFS as the base search algorithm and scoring with the LM as the heuristic. For all methods that involve sampling, including LATS, we sample
1751
+ k
1752
+ =
1753
+ 50
1754
+ 𝑘
1755
+ 50
1756
+ k=50
1757
+ trajectories. More details and prompts can be found in Sec.
1758
+ D
1759
+ and Sec.
1760
+ E
1761
+ in the Appendix.
1762
+ We evaluate internal reasoning strategies by removing actions and observations from the context, corresponding to CoT
1763
+ (Wei et al.,
1764
+ 2022
1765
+ )
1766
+ and its variants, CoT-SC
1767
+ (Wang et al.,
1768
+ 2022
1769
+ )
1770
+ , ToT
1771
+ (Yao et al.,
1772
+ 2023a
1773
+ )
1774
+ , and RAP
1775
+ (Hao et al.,
1776
+ 2023
1777
+ )
1778
+ . These methods rely solely on the agent’s existing knowledge to answer the question. We also consider acting-based methods ReAct, Reflexion, and LATS, which augment the agent with the interactive API environment and primarily evaluate its information retrieval abilities. While LATS is designed for scenarios where external feedback can enhance reasoning, we also implement a reasoning-only version with CoT as the base prompt. We also combine internal and external reasoning in LATS by first prompting with a CoT-based prompt, then switching to a ReAct-based prompt upon failure. This is closer to how humans might approach this task, by using tools to lookup additional information only when the answer is not already known.
1779
+ Prompt Method
1780
+ HotpotQA (EM)
1781
+ I/O
1782
+ 0.32
1783
+ CoT
1784
+ (Wei et al.,
1785
+ 2022
1786
+ )
1787
+ 0.34
1788
+ CoT - SC
1789
+ (Wang et al.,
1790
+ 2022
1791
+ )
1792
+ 0.38
1793
+ ToT
1794
+ (Yao et al.,
1795
+ 2023a
1796
+ )
1797
+ 0.55
1798
+ RAP
1799
+ (Hao et al.,
1800
+ 2023
1801
+ )
1802
+ 0.60
1803
+ RAP (n = 10)
1804
+ 0.60
1805
+ LATS (CoT)
1806
+ 0.60
1807
+ Prompt Method
1808
+ HotpotQA (EM)
1809
+ ReAct
1810
+ (Yao et al.,
1811
+ 2023b
1812
+ )
1813
+ 0.32
1814
+ ReAct (best of k)
1815
+ 0.38
1816
+ Reflexion
1817
+ (Shinn et al.,
1818
+ 2023
1819
+ )
1820
+ 0.51
1821
+ LATS
1822
+ 0.61
1823
+ LATS (n = 3)
1824
+ 0.56
1825
+ LATS (n = 10)
1826
+ 0.64
1827
+ LATS (CoT + ReAct)
1828
+ 0.71
1829
+ Table 2:
1830
+ GPT-3.5 reasoning-based prompting (left) and acting-based prompting (right) results on HotpotQA. LATS achieves the highest exact match (EM) for acting and is competitive on reasoning. Unless otherwise specified, we sample
1831
+ n
1832
+ =
1833
+ 5
1834
+ 𝑛
1835
+ 5
1836
+ n=5
1837
+ nodes during expansion and
1838
+ k
1839
+ =
1840
+ 50
1841
+ 𝑘
1842
+ 50
1843
+ k=50
1844
+ trajectories.
1845
+ Results.
1846
+ We observe in Tab.
1847
+ 2
1848
+ that both internal reasoning and external retrieval strategies perform well on HotPotQA. Due to their large-scale training corpus, modern LLMs already encode factual knowledge and can often directly answer the question correctly. While CoT can slightly enhance performance on questions requiring reasoning, larger gains are observed with search methods ToT and RAP, which can sample and explore more outputs. We observe similar results for acting-based methods. LATS surpasses ReAct, even when sampling the same number of trajectories, by expanding more nodes with principled search (see Fig.
1849
+ 5
1850
+ in Appendix
1851
+ D
1852
+ for a qualitative sample). This is demonstrated when modifying
1853
+ n
1854
+ 𝑛
1855
+ n
1856
+ , the number of nodes expanded during each iteration. Increasing
1857
+ n
1858
+ 𝑛
1859
+ n
1860
+ can consistently improve performance, although at greater computational and inference costs. LATS is also competitive to RAP on internal reasoning but performs worse than acting. Combining internal and external reasoning in LATS results in the highest performance, indicating the importance of external feedback in augmenting reasoning even in tasks the base LM can already perform.
1861
+ 5.2
1862
+ Programming
1863
+ Prompt Method
1864
+ Model
1865
+ Pass@1
1866
+ CoT
1867
+ (Wei et al.,
1868
+ 2022
1869
+ )
1870
+ GPT-3.5
1871
+ 46.9
1872
+ ReAct
1873
+ (Yao et al.,
1874
+ 2023b
1875
+ )
1876
+ GPT-3.5
1877
+ 56.9
1878
+ Reflexion
1879
+ (Shinn et al.,
1880
+ 2023
1881
+ )
1882
+ GPT-3.5
1883
+ 68.1
1884
+ ToT
1885
+ (Yao et al.,
1886
+ 2023a
1887
+ )
1888
+ GPT-3.5
1889
+ 54.4
1890
+ RAP
1891
+ (Hao et al.,
1892
+ 2023
1893
+ )
1894
+ GPT-3.5
1895
+ 63.1
1896
+ LATS (Ours)
1897
+ GPT-3.5
1898
+ 83.8
1899
+ I/O
1900
+ GPT-4
1901
+ 80.1
1902
+ Reflexion
1903
+ GPT-4
1904
+ 91.0
1905
+ LATS
1906
+ GPT-4
1907
+ 94.4
1908
+ Prompt Method
1909
+ Pass@1
1910
+ CoT
1911
+ (Wei et al.,
1912
+ 2022
1913
+ )
1914
+ 54.9
1915
+ ReAct
1916
+ (Wei et al.,
1917
+ 2022
1918
+ )
1919
+ 67.0
1920
+ Reflexion
1921
+ (Shinn et al.,
1922
+ 2023
1923
+ )
1924
+ 70.0
1925
+ ToT
1926
+ (Yao et al.,
1927
+ 2023a
1928
+ )
1929
+ 65.8
1930
+ RAP
1931
+ (Hao et al.,
1932
+ 2023
1933
+ )
1934
+ 71.4
1935
+ LATS (Ours)
1936
+ 81.1
1937
+ Table 3:
1938
+ GPT-3.5 and GPT-4 Pass@1 accuracy on HumanEval
1939
+ (Chen et al.,
1940
+ 2021
1941
+ )
1942
+ and MBPP
1943
+ (Austin et al.,
1944
+ 2021
1945
+ )
1946
+ . Prompting with LATS achieves the highest performance. We sample 5 solutions during expansion for
1947
+ 8
1948
+ iterations.
1949
+ To demonstrate the importance of external observations for complex reasoning tasks, we evaluate the baselines and LATS on programming with Humaneval
1950
+ (Chen et al.,
1951
+ 2021
1952
+ )
1953
+ and MBPP
1954
+ (Austin et al.,
1955
+ 2021
1956
+ )
1957
+ . Both datasets measure the correctness of synthesized programs in Python from natural language docstrings. We use individual solutions as the action space and test suite and compiler feedback as the external observation. We follow
1958
+ Chen et al. (
1959
+ 2022a
1960
+ )
1961
+ and use an LLM to generate a synthetic test suite of syntactically valid “assert” statements for each question. For each step, the solution is evaluated on this test suite, and the results including successful and failed tests and compiler output, are added to the context as an observation. We use the same test suite for Reflexion.
1962
+ For this task, the reasoning and acting baselines share an action space, but acting methods are able to incorporate observations as additional context. For LATS, since each action corresponds to a complete solution, we skip the simulation step of LATS and directly use the percentage of passed tests as the backpropagated reward. We use
1963
+ k
1964
+ =
1965
+ 8
1966
+ 𝑘
1967
+ 8
1968
+ k=8
1969
+ iterations, set the number of generated tests at
1970
+ 4
1971
+ 4
1972
+ 4
1973
+ , and sample
1974
+ n
1975
+ =
1976
+ 5
1977
+ 𝑛
1978
+ 5
1979
+ n=5
1980
+ solutions during expansion. After the search is completed, we select the solution with the highest value and evaluate it on the real test suite for the pass@1 accuracy evaluation. More details and prompts can be found in Sec.
1981
+ D
1982
+ and Sec.
1983
+ F
1984
+ in the Appendix.
1985
+ Results.
1986
+ We find in Tab
1987
+ 3
1988
+ that both search and semantic feedback are crucial for better performance. Despite not using observations, ToT and RAP are competitive with Reflexion. LATS has the highest performance on both datasets. Since RAP uses a similar search algorithm as LATS, this reveals the importance of external feedback for difficult reasoning tasks such as programming. With GPT-4, using LATS sets the state of the art for HumanEval, showing LATS can be used with more advanced LLMs for higher performance.
1989
+ 5.3
1990
+ Webshop
1991
+ For a complex decision-making environment with practical applications, we consider WebShop
1992
+ (Yao et al.,
1993
+ 2022
1994
+ )
1995
+ , an online shopping environment composed of a website with 1.18M real-world products and 12k human instructions. Agents must navigate a website through a variety of commands to purchase an item matching a user specification. We use the preconstructed action space of search and click commands and browser feedback and reflections for the observation. The performance is gauged using two metrics: an average score, reflecting the percentage of user-specified attributes met by the selected product, and a success rate, indicating the frequency with which the chosen product fulfills all given conditions. We compare against acting-based prompting methods and RL-based approaches. We evaluate on 50 instructions, expand
1996
+ n
1997
+ =
1998
+ 5
1999
+ 𝑛
2000
+ 5
2001
+ n=5
2002
+ children for LATS, and set
2003
+ k
2004
+ =
2005
+ 30
2006
+ 𝑘
2007
+ 30
2008
+ k=30
2009
+ for LATS, ReAct best of
2010
+ k
2011
+ 𝑘
2012
+ k
2013
+ , and Reflexion. More details and prompts are in Appendix
2014
+ D
2015
+ and
2016
+ G
2017
+ .
2018
+ Results.
2019
+ We find in Tab.
2020
+ 5
2021
+ that GPT-3.5 with ReAct is competitive to imitation learning, and can exceed reinforcement learning techniques with stronger prompting strategies. Sampling
2022
+ k
2023
+ =
2024
+ 30
2025
+ 𝑘
2026
+ 30
2027
+ k=30
2028
+ trajectories with ReAct and Reflexion results in a similar performance, suggesting the semantic feedback is not as helpful in complex environments like WebShop. Indeed like in
2029
+ Shinn et al. (
2030
+ 2023
2031
+ )
2032
+ , we find that generated reflections are often generic and do not provide useful feedback, resulting in a tendency for the agent to become stuck in local minima. However, using LATS indeed results in a noticeable improvement, indicating a more effective exploration for the same number of iterations.
2033
+ 5.4
2034
+ Additional Observations
2035
+ Method
2036
+ Score
2037
+ SR
2038
+ ReAct
2039
+ (Yao et al.,
2040
+ 2023b
2041
+ )
2042
+ 53.8
2043
+ 28.0
2044
+ ReAct (best of k)
2045
+ 59.1
2046
+ 32.0
2047
+ Reflexion
2048
+ (Shinn et al.,
2049
+ 2023
2050
+ )
2051
+ 64.2
2052
+ 35.0
2053
+ LATS
2054
+ 75.9
2055
+ 38.0
2056
+ IL
2057
+ 59.9
2058
+ 29.1
2059
+ IL+RL
2060
+ 62.4
2061
+ 28.7
2062
+ Fine-tuning
2063
+ (Furuta et al.,
2064
+ 2023
2065
+ )
2066
+ 67.5
2067
+ 45.0
2068
+ Expert
2069
+ 82.1
2070
+ 59.6
2071
+ Table 4:
2072
+ Score and success rate (SR) on Webshop. Table is separated into prompting, RL-based training, and human performance. For the same number of iterations, LATS improves both score and success rate, and surpasses RL-based training. IL/IL+RL taken from
2073
+ Yao et al. (
2074
+ 2022
2075
+ )
2076
+ .
2077
+ Prompt Method
2078
+ HotPotQA (EM)
2079
+ ToT (ReAct)
2080
+ 0.39
2081
+ RAP (ReAct)
2082
+ 0.54
2083
+ LATS (No LM Heuristic)
2084
+ 0.37
2085
+ LATS (DFS)
2086
+ 0.42
2087
+ LATS (No Reflection)
2088
+ 0.56
2089
+ LATS
2090
+ 0.61
2091
+ Table 5:
2092
+ Ablation results on LATS and baseline variants in HotPotQA; we use ReAct as the base prompt and sample
2093
+ n
2094
+ =
2095
+ 5
2096
+ 𝑛
2097
+ 5
2098
+ n=5
2099
+ children and
2100
+ k
2101
+ =
2102
+ 50
2103
+ 𝑘
2104
+ 50
2105
+ k=50
2106
+ maximum trajectories. LATS requires every component and operation for optimal performance.
2107
+ We also conduct additional experiments on HotPotQA to demonstrate the effect of each component of LATS. We also design a version of ToT and RAP with ReAct prompt and can handle external observations. We use HotPotQA as our setup incorporates both reasoning (through thoughts) and acting (through API calls); the results are shown in Tab.
2108
+ 5
2109
+ . More ablations for token consumption on HotPotQA are in Tab.
2110
+ 7
2111
+ in Appendix
2112
+ C
2113
+ . Note that baselines generally perform worse than the reasoning-only setting of HotPotQA, which indicates that the acting-based setting is more challenging and adaption of search algorithms to decision-making scenarios is non-trivial.
2114
+ Self-reflection.
2115
+ We use self-reflection to provide additional semantic signals for the agent. We observe a
2116
+ 0.05
2117
+ 0.05
2118
+ 0.05
2119
+ performance drop when removed from LATS, suggesting this is useful. This is a smaller gain Reflexion
2120
+ (Shinn et al.,
2121
+ 2023
2122
+ )
2123
+ observes over ReAct
2124
+ (Yao et al.,
2125
+ 2023b
2126
+ )
2127
+ as shown in Tab.
2128
+ 2
2129
+ , suggesting overlap between the types of questions where there is an improvement with self-reflection and search. This variant outperforms RAP-ReAct, reflecting our improvements to MCTS.
2130
+ Search Algorithm.
2131
+ MCTS is a more principled search algorithm than variants like A* or DFS search and the basis for observed performance gains. We observe the effects of using DFS, and incorporate the LM-based heuristic used in ToT
2132
+ (Yao et al.,
2133
+ 2023a
2134
+ )
2135
+ in which branches with low values are pruned. This removes the selection and backpropagation operations, and we observe a
2136
+ 0.08
2137
+ 0.08
2138
+ 0.08
2139
+ drop in performance when sampling the same number of nodes, but outperforms ToT-ReAct.
2140
+ 6
2141
+ Conclusion
2142
+ In this work, we introduce Language Agent Tree Search (LATS), the first framework to unify planning, acting, and reasoning for enhanced LLM problem solving. By deliberately constructing trajectories with search algorithms, incorporating external feedback, and enabling agents to learn from experience, LATS addresses key limitations of prior prompting techniques. Our evaluations demonstrate the ability of LATS to harness LLM capabilities for a variety of decision-making tasks while keeping its reasoning ability without additional training. The proposed synergies between search, interaction, and reflection offer a versatile approach to autonomous decision-making, highlighting the potential of LLMs as generalist agents. A full discussion of the limitations and broader impacts is in Appendix
2143
+ B
2144
+ .
2145
+ References
2146
+ Ahn et al. (2022)
2147
+ Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan, and Andy Zeng.
2148
+ Do as i can, not as i say: Grounding language in robotic affordances.
2149
+ arXiv:2204.01691
2150
+ , 2022.
2151
+ Anthony et al. (2017)
2152
+ T. Anthony, Z. Tian, and D. Barber.
2153
+ Thinking fast and slow with deep learning and tree search.
2154
+ In
2155
+ NIPS
2156
+ , 2017.
2157
+ Austin et al. (2021)
2158
+ Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al.
2159
+ Program synthesis with large language models.
2160
+ arXiv:2108.07732
2161
+ , 2021.
2162
+ Baker et al. (2022)
2163
+ Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune.
2164
+ Video pretraining (vpt): Learning to act by watching unlabeled online videos.
2165
+ arXiv:2206.11795
2166
+ , 2022.
2167
+ Besta et al. (2023)
2168
+ Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Michal Podstawski, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler.
2169
+ Graph of thoughts: Solving elaborate problems with large language models.
2170
+ arXiv:2308.09687
2171
+ , 2023.
2172
+ Bowman et al. (2015)
2173
+ Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning.
2174
+ A large annotated corpus for learning natural language inference.
2175
+ In
2176
+ EMNLP
2177
+ , 2015.
2178
+ Brown et al. (2020)
2179
+ Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei.
2180
+ Language models are few-shot learners.
2181
+ In
2182
+ NeurIPS
2183
+ , 2020.
2184
+ Campbell et al. (2002)
2185
+ Murray Campbell, A Joseph Hoane Jr, and Feng-hsiung Hsu.
2186
+ Deep blue.
2187
+ Artificial intelligence
2188
+ , 2002.
2189
+ Chen et al. (2022a)
2190
+ Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen.
2191
+ Codet: Code generation with generated tests.
2192
+ arXiv:2207.10397
2193
+ , 2022a.
2194
+ Chen et al. (2021)
2195
+ Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al.
2196
+ Evaluating large language models trained on code.
2197
+ arXiv:2107.03374
2198
+ , 2021.
2199
+ Chen et al. (2022b)
2200
+ Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen.
2201
+ Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks.
2202
+ arXiv preprint arXiv:2211.12588
2203
+ , 2022b.
2204
+ Chowdhery et al. (2022)
2205
+ Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al.
2206
+ Palm: Scaling language modeling with pathways.
2207
+ arXiv:2204.02311
2208
+ , 2022.
2209
+ Cobbe et al. (2021)
2210
+ Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al.
2211
+ Training verifiers to solve math word problems.
2212
+ arXiv:2110.14168
2213
+ , 2021.
2214
+ Deng et al. (2023)
2215
+ Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su.
2216
+ Mind2web: Towards a generalist agent for the web.
2217
+ arXiv:2306.06070
2218
+ , 2023.
2219
+ Driess et al. (2023)
2220
+ Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence.
2221
+ Palm-e: An embodied multimodal language model.
2222
+ arXiv:2303.03378
2223
+ , 2023.
2224
+ Du et al. (2023)
2225
+ Yilun Du, Mengjiao Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Joshua B. Tenenbaum, Dale Schuurmans, and Pieter Abbeel.
2226
+ Learning universal policies via text-guided video generation.
2227
+ arXiv:2302.00111
2228
+ , 2023.
2229
+ Evans (2010)
2230
+ Jonathan St BT Evans.
2231
+ Intuition and reasoning: A dual-process perspective.
2232
+ Psychological Inquiry
2233
+ , 2010.
2234
+ Fan et al. (2022)
2235
+ Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar.
2236
+ Minedojo: Building open-ended embodied agents with internet-scale knowledge.
2237
+ In
2238
+ NeurIPS Datasets and Benchmarks Track
2239
+ , 2022.
2240
+ Furuta et al. (2023)
2241
+ Hiroki Furuta, Ofir Nachum, Kuang-Huei Lee, Yutaka Matsuo, Shixiang Shane Gu, and Izzeddin Gur.
2242
+ Multimodal web navigation with instruction-finetuned foundation models.
2243
+ arXiv preprint arXiv:2305.11854
2244
+ , 2023.
2245
+ Gao et al. (2022)
2246
+ Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig.
2247
+ Pal: Program-aided language models.
2248
+ arXiv preprint arXiv:2211.10435
2249
+ , 2022.
2250
+ Guo et al. (2018)
2251
+ Jiaxian Guo, Sidi Lu, Han Cai, Weinan Zhang, Yong Yu, and Jun Wang.
2252
+ Long text generation via adversarial training with leaked information.
2253
+ AAAI
2254
+ , 2018.
2255
+ Guss et al. (2019)
2256
+ William H. Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela Veloso, and Ruslan Salakhutdinov.
2257
+ Minerl: A large-scale dataset of minecraft demonstrations.
2258
+ In
2259
+ IJCAI
2260
+ , 2019.
2261
+ Hafner et al. (2019)
2262
+ Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson.
2263
+ Learning latent dynamics for planning from pixels.
2264
+ In
2265
+ ICML
2266
+ , 2019.
2267
+ Hafner et al. (2023)
2268
+ Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap.
2269
+ Mastering diverse domains through world models.
2270
+ arXiv:2301.04104
2271
+ , 2023.
2272
+ Hao et al. (2023)
2273
+ Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu.
2274
+ Reasoning with language model is planning with world model.
2275
+ arXiv:2305.14992
2276
+ , 2023.
2277
+ Huang et al. (2023)
2278
+ Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou.
2279
+ Large language models cannot self-correct reasoning yet.
2280
+ arXiv:2310.01798
2281
+ , 2023.
2282
+ Huang et al. (2022)
2283
+ Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al.
2284
+ Inner monologue: Embodied reasoning through planning with language models.
2285
+ arXiv:2207.05608
2286
+ , 2022.
2287
+ Jiang et al. (2018)
2288
+ D. Jiang, E. Ekwedike, and H. Liu.
2289
+ Feedback-based tree search for reinforcement learning.
2290
+ In
2291
+ ICML
2292
+ , 2018.
2293
+ Kocsis & Szepesvári (2006)
2294
+ Levente Kocsis and Csaba Szepesvári.
2295
+ Bandit based monte-carlo planning.
2296
+ In
2297
+ ECML
2298
+ , 2006.
2299
+ Kojima et al. (2022)
2300
+ Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa.
2301
+ Large language models are zero-shot reasoners.
2302
+ arXiv:2205.11916
2303
+ , 2022.
2304
+ LaValle et al. (2001)
2305
+ Steven M LaValle, James J Kuffner, BR Donald, et al.
2306
+ Rapidly-exploring random trees: Progress and prospects.
2307
+ Algorithmic and computational robotics: new directions
2308
+ , 2001.
2309
+ Liu et al. (2018)
2310
+ Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang.
2311
+ Reinforcement learning on web interfaces using workflow-guided exploration.
2312
+ In
2313
+ ICLR
2314
+ , 2018.
2315
+ Liu et al. (2023)
2316
+ Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang.
2317
+ Agentbench: Evaluating llms as agents.
2318
+ arXiv:2308.03688
2319
+ , 2023.
2320
+ Madaan et al. (2023)
2321
+ Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark.
2322
+ Self-refine: Iterative refinement with self-feedback.
2323
+ arXiv:2303.17651
2324
+ , 2023.
2325
+ Nallapati et al. (2016)
2326
+ Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Caglar Gulcehre, and Bing Xiang.
2327
+ Abstractive text summarization using sequence-to-sequence rnns and beyond.
2328
+ In
2329
+ SIGNLL
2330
+ , 2016.
2331
+ Nye et al. (2021)
2332
+ Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al.
2333
+ Show your work: Scratchpads for intermediate computation with language models.
2334
+ arXiv:2112.00114
2335
+ , 2021.
2336
+ OpenAI (2023)
2337
+ OpenAI.
2338
+ Gpt-4 technical report.
2339
+ arXiv:2303.08774
2340
+ , 2023.
2341
+ Saparov & He (2022)
2342
+ Abulhair Saparov and He He.
2343
+ Language models are greedy reasoners: A systematic formal analysis of chain-of-thought.
2344
+ arXiv:2210.01240
2345
+ , 2022.
2346
+ Schick et al. (2023)
2347
+ Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom.
2348
+ Toolformer: Language models can teach themselves to use tools.
2349
+ arXiv:2302.04761
2350
+ , 2023.
2351
+ Shen et al. (2023)
2352
+ Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang.
2353
+ Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface.
2354
+ arXiv:2303.17580
2355
+ , 2023.
2356
+ Shinn et al. (2023)
2357
+ Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao.
2358
+ Reflexion: Language agents with verbal reinforcement learning.
2359
+ arXiv:2303.11366
2360
+ , 2023.
2361
+ Shridhar et al. (2020)
2362
+ Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht.
2363
+ Alfworld: Aligning text and embodied environments for interactive learning.
2364
+ arXiv:2010.03768
2365
+ , 2020.
2366
+ Silver et al. (2016)
2367
+ David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al.
2368
+ Mastering the game of go with deep neural networks and tree search.
2369
+ nature
2370
+ , 2016.
2371
+ Silver et al. (2017)
2372
+ David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy P. Lillicrap, Fan Hui, L. Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis.
2373
+ Mastering the game of go without human knowledge.
2374
+ Nature
2375
+ , 2017.
2376
+ Sloman (1996)
2377
+ Steven A. Sloman.
2378
+ The empirical case for two systems of reasoning.
2379
+ Psychological Bulletin
2380
+ , 1996.
2381
+ Sun et al. (2023)
2382
+ Haotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai, and Chao Zhang.
2383
+ Adaplanner: Adaptive planning from feedback with language models.
2384
+ arXiv:2305.16653
2385
+ , 2023.
2386
+ Surís et al. (2023)
2387
+ Dídac Surís, Sachit Menon, and Carl Vondrick.
2388
+ Vipergpt: Visual inference via python execution for reasoning.
2389
+ arXiv preprint arXiv:2303.08128
2390
+ , 2023.
2391
+ Świechowski et al. (2023)
2392
+ Maciej Świechowski, Konrad Godlewski, Bartosz Sawicki, and Jacek Mańdziuk.
2393
+ Monte carlo tree search: A review of recent modifications and applications.
2394
+ Artificial Intelligence Review
2395
+ , 2023.
2396
+ Touvron et al. (2023)
2397
+ Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Cantón Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony S. Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel M. Kloumann, A. V. Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, R. Subramanian, Xia Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zhengxu Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and
2398
+ Thomas Scialom.
2399
+ Llama 2: Open foundation and fine-tuned chat models.
2400
+ arXiv:2307.09288
2401
+ , 2023.
2402
+ Vodopivec et al. (2017)
2403
+ Tom Vodopivec, Spyridon Samothrakis, and Branko Ster.
2404
+ On monte carlo tree search and reinforcement learning.
2405
+ Journal of Artificial Intelligence Research
2406
+ , 2017.
2407
+ Wang et al. (2023)
2408
+ Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar.
2409
+ Voyager: An open-ended embodied agent with large language models.
2410
+ arXiv:2305.16291
2411
+ , 2023.
2412
+ Wang et al. (2022)
2413
+ Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou.
2414
+ Self-consistency improves chain of thought reasoning in language models.
2415
+ arXiv:2203.11171
2416
+ , 2022.
2417
+ Wei et al. (2022)
2418
+ Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou.
2419
+ Chain of thought prompting elicits reasoning in large language models.
2420
+ arXiv:2201.11903
2421
+ , 2022.
2422
+ Wooldridge & Jennings (1995)
2423
+ Michael Wooldridge and Nicholas R Jennings.
2424
+ Intelligent agents: Theory and practice.
2425
+ The knowledge engineering review
2426
+ , 1995.
2427
+ Wu et al. (2023)
2428
+ Philipp Wu, Alejandro Escontrela, Danijar Hafner, Pieter Abbeel, and Ken Goldberg.
2429
+ Daydreamer: World models for physical robot learning.
2430
+ In
2431
+ CoRL
2432
+ . PMLR, 2023.
2433
+ Xie et al. (2023)
2434
+ Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao, Min-Yen Kan, Junxian He, and Qizhe Xie.
2435
+ Decomposition enhances reasoning via self-evaluation guided decoding.
2436
+ arXiv:2305.00633
2437
+ , 2023.
2438
+ Yang et al. (2018)
2439
+ Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning.
2440
+ Hotpotqa: A dataset for diverse, explainable multi-hop question answering.
2441
+ arXiv:1809.09600
2442
+ , 2018.
2443
+ Yao et al. (2022)
2444
+ Shunyu Yao, Howard Chen, John Yang, and Karthik R Narasimhan.
2445
+ Webshop: Towards scalable real-world web interaction with grounded language agents.
2446
+ In
2447
+ NeurIPS
2448
+ , 2022.
2449
+ Yao et al. (2023a)
2450
+ Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan.
2451
+ Tree of thoughts: Deliberate problem solving with large language models.
2452
+ arXiv:2305.10601
2453
+ , 2023a.
2454
+ Yao et al. (2023b)
2455
+ Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao.
2456
+ ReAct: Synergizing reasoning and acting in language models.
2457
+ In
2458
+ ICLR
2459
+ , 2023b.
2460
+ Yao et al. (2023c)
2461
+ Weiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu, Yihao Feng, Le Xue, Rithesh Murthy, Zeyuan Chen, Jianguo Zhang, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese.
2462
+ Retroformer: Retrospective large language agents with policy gradient optimization.
2463
+ arXiv preprint arXiv:2308.02151
2464
+ , 2023c.
2465
+ Ye et al. (2021)
2466
+ Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao.
2467
+ Mastering atari games with limited data.
2468
+ In
2469
+ NeurIPS
2470
+ , 2021.
2471
+ Zhou et al. (2022)
2472
+ Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi.
2473
+ Least-to-most prompting enables complex reasoning in large language models.
2474
+ arXiv:2205.10625
2475
+ , 2022.
2476
+ Zhu et al. (2023)
2477
+ Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, Yu Qiao, Zhaoxiang Zhang, and Jifeng Dai.
2478
+ Ghost in the minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory.
2479
+ arXiv:2305.17144
2480
+ , 2023.
2481
+ 7
2482
+ Appendix
2483
+ The appendix is organized as follows. First in Sec.
2484
+ A
2485
+ , we show the pseudocode of our proposed algorithm, LATS; then in Sec.
2486
+ B
2487
+ , we provide further discussion of our method and its limitations, future direction and broader impact; then in Sec.
2488
+ C
2489
+ we provide additional experimental results; then in Sec.
2490
+ D
2491
+ , we specify the environment details in our experiments; finally, we list our prompts used for the three environments in Sec.
2492
+ E
2493
+ (HotPotQA), Sec.
2494
+ F
2495
+ (Programming) and Sec.
2496
+ G
2497
+ (Webshop) respectively.
2498
+ Appendix A
2499
+ LATS Pseudocode
2500
+ Alg.
2501
+ 1
2502
+ shows the pseudocode of our algorithm LATS. Nodes are stored explicitly in the memory. Unless otherwise specified, in all experiments we use
2503
+ n
2504
+ =
2505
+ 5
2506
+ 𝑛
2507
+ 5
2508
+ n=5
2509
+ and
2510
+ w
2511
+ =
2512
+ 1
2513
+ 𝑤
2514
+ 1
2515
+ w=1
2516
+ .
2517
+ Algorithm 1
2518
+ LATS
2519
+ ⁡
2520
+ (
2521
+ S
2522
+ 0
2523
+ ,
2524
+ p
2525
+ θ
2526
+ ,
2527
+ p
2528
+ V
2529
+ ,
2530
+ p
2531
+ ref
2532
+ ,
2533
+ d
2534
+ ,
2535
+ k
2536
+ ,
2537
+ n
2538
+ ,
2539
+ w
2540
+ )
2541
+ LATS
2542
+ subscript
2543
+ 𝑆
2544
+ 0
2545
+ subscript
2546
+ 𝑝
2547
+ 𝜃
2548
+ subscript
2549
+ 𝑝
2550
+ 𝑉
2551
+ subscript
2552
+ 𝑝
2553
+ ref
2554
+ 𝑑
2555
+ 𝑘
2556
+ 𝑛
2557
+ 𝑤
2558
+ \operatorname{LATS}(S_{0},p_{\theta},{p_{V}},p_{\text{ref}},d,k,n,w)
2559
+ Initial state
2560
+ s
2561
+ 1
2562
+ subscript
2563
+ 𝑠
2564
+ 1
2565
+ s_{1}
2566
+ , action generator
2567
+ p
2568
+ θ
2569
+ subscript
2570
+ 𝑝
2571
+ 𝜃
2572
+ p_{\theta}
2573
+ , value function
2574
+ p
2575
+ V
2576
+ subscript
2577
+ 𝑝
2578
+ 𝑉
2579
+ p_{V}
2580
+ , reflection generator
2581
+ p
2582
+ ref
2583
+ subscript
2584
+ 𝑝
2585
+ ref
2586
+ p_{\text{ref}}
2587
+ , number of generated actions
2588
+ n
2589
+ 𝑛
2590
+ n
2591
+ , depth limit
2592
+ L
2593
+ 𝐿
2594
+ L
2595
+ , number of roll-outs
2596
+ K
2597
+ 𝐾
2598
+ K
2599
+ , context
2600
+ c
2601
+ 𝑐
2602
+ c
2603
+ , and exploration weight
2604
+ w
2605
+ 𝑤
2606
+ w
2607
+ Initialize action space
2608
+ A
2609
+ 𝐴
2610
+ A
2611
+ , observation space
2612
+ O
2613
+ 𝑂
2614
+ O
2615
+ Initialize the state-action value function
2616
+ p
2617
+ V
2618
+ :
2619
+ S
2620
+ ×
2621
+ A
2622
+ ↦
2623
+ ℝ
2624
+ :
2625
+ subscript
2626
+ 𝑝
2627
+ 𝑉
2628
+ maps-to
2629
+ 𝑆
2630
+ 𝐴
2631
+ ℝ
2632
+ {p_{V}}:S\times A\mapsto\mathbb{R}
2633
+ and visit counter
2634
+ N
2635
+ :
2636
+ S
2637
+ ↦
2638
+ ℕ
2639
+ :
2640
+ 𝑁
2641
+ maps-to
2642
+ 𝑆
2643
+ ℕ
2644
+ {N}:S\mapsto\mathbb{N}
2645
+ to zero
2646
+ for
2647
+ k
2648
+ ←
2649
+ 0
2650
+ ,
2651
+ …
2652
+ ,
2653
+ K
2654
+ −
2655
+ 1
2656
+ ←
2657
+ 𝑘
2658
+ 0
2659
+ …
2660
+ 𝐾
2661
+ 1
2662
+ k\leftarrow 0,\dots,K-1
2663
+ do
2664
+ for
2665
+ t
2666
+ ←
2667
+ 0
2668
+ ,
2669
+ …
2670
+ ,
2671
+ L
2672
+ −
2673
+ 1
2674
+ ←
2675
+ 𝑡
2676
+ 0
2677
+ …
2678
+ 𝐿
2679
+ 1
2680
+ t\leftarrow 0,\dots,L-1
2681
+ do
2682
+ if
2683
+ s
2684
+ t
2685
+ subscript
2686
+ 𝑠
2687
+ 𝑡
2688
+ s_{t}
2689
+ not terminal
2690
+ then
2691
+ ▷
2692
+ ▷
2693
+ \triangleright
2694
+ Expansion & Simulation
2695
+ for
2696
+ i
2697
+ ←
2698
+ 1
2699
+ ,
2700
+ …
2701
+ ,
2702
+ n
2703
+ ←
2704
+ 𝑖
2705
+ 1
2706
+ …
2707
+ 𝑛
2708
+ i\leftarrow 1,\dots,n
2709
+ do
2710
+ Sample
2711
+ a
2712
+ t
2713
+ (
2714
+ i
2715
+ )
2716
+ ∼
2717
+ p
2718
+ θ
2719
+ ​
2720
+ (
2721
+ a
2722
+ ∣
2723
+ s
2724
+ t
2725
+ )
2726
+ similar-to
2727
+ superscript
2728
+ subscript
2729
+ 𝑎
2730
+ 𝑡
2731
+ 𝑖
2732
+ subscript
2733
+ 𝑝
2734
+ 𝜃
2735
+ conditional
2736
+ 𝑎
2737
+ subscript
2738
+ 𝑠
2739
+ 𝑡
2740
+ a_{t}^{(i)}\sim p_{\theta}(a\mid s_{t})
2741
+ Get
2742
+ o
2743
+ t
2744
+ (
2745
+ i
2746
+ )
2747
+ superscript
2748
+ subscript
2749
+ 𝑜
2750
+ 𝑡
2751
+ 𝑖
2752
+ o_{t}^{(i)}
2753
+ from environment,
2754
+ s
2755
+ t
2756
+ +
2757
+ 1
2758
+ (
2759
+ i
2760
+ )
2761
+ ←
2762
+ (
2763
+ c
2764
+ t
2765
+ (
2766
+ i
2767
+ )
2768
+ ,
2769
+ o
2770
+ t
2771
+ (
2772
+ i
2773
+ )
2774
+ ,
2775
+ a
2776
+ t
2777
+ (
2778
+ i
2779
+ )
2780
+ )
2781
+ ←
2782
+ superscript
2783
+ subscript
2784
+ 𝑠
2785
+ 𝑡
2786
+ 1
2787
+ 𝑖
2788
+ superscript
2789
+ subscript
2790
+ 𝑐
2791
+ 𝑡
2792
+ 𝑖
2793
+ superscript
2794
+ subscript
2795
+ 𝑜
2796
+ 𝑡
2797
+ 𝑖
2798
+ superscript
2799
+ subscript
2800
+ 𝑎
2801
+ 𝑡
2802
+ 𝑖
2803
+ s_{t+1}^{(i)}\leftarrow(c_{t}^{(i)},o_{t}^{(i)},a_{t}^{(i)})
2804
+ ,
2805
+ c
2806
+ t
2807
+ +
2808
+ 1
2809
+ (
2810
+ i
2811
+ )
2812
+ ←
2813
+ (
2814
+ o
2815
+ t
2816
+ (
2817
+ i
2818
+ )
2819
+ ,
2820
+ a
2821
+ t
2822
+ (
2823
+ i
2824
+ )
2825
+ )
2826
+ ←
2827
+ superscript
2828
+ subscript
2829
+ 𝑐
2830
+ 𝑡
2831
+ 1
2832
+ 𝑖
2833
+ superscript
2834
+ subscript
2835
+ 𝑜
2836
+ 𝑡
2837
+ 𝑖
2838
+ superscript
2839
+ subscript
2840
+ 𝑎
2841
+ 𝑡
2842
+ 𝑖
2843
+ c_{t+1}^{(i)}\leftarrow(o_{t}^{(i)},a_{t}^{(i)})
2844
+ Evaluate
2845
+ V
2846
+ t
2847
+ (
2848
+ i
2849
+ )
2850
+ ∼
2851
+ p
2852
+ V
2853
+ ​
2854
+ (
2855
+ s
2856
+ t
2857
+ (
2858
+ i
2859
+ )
2860
+ )
2861
+ similar-to
2862
+ superscript
2863
+ subscript
2864
+ 𝑉
2865
+ 𝑡
2866
+ 𝑖
2867
+ subscript
2868
+ 𝑝
2869
+ 𝑉
2870
+ superscript
2871
+ subscript
2872
+ 𝑠
2873
+ 𝑡
2874
+ 𝑖
2875
+ {V}_{t}^{(i)}\sim{p_{V}}(s_{t}^{(i)})
2876
+ ▷
2877
+ ▷
2878
+ \triangleright
2879
+ Evaluation
2880
+ V
2881
+ ​
2882
+ (
2883
+ s
2884
+ t
2885
+ )
2886
+ ←
2887
+ V
2888
+ t
2889
+ (
2890
+ i
2891
+ )
2892
+ ←
2893
+ 𝑉
2894
+ subscript
2895
+ 𝑠
2896
+ 𝑡
2897
+ superscript
2898
+ subscript
2899
+ 𝑉
2900
+ 𝑡
2901
+ 𝑖
2902
+ {V}(s_{t})\leftarrow{V}_{t}^{(i)}
2903
+ Add
2904
+ s
2905
+ t
2906
+ (
2907
+ i
2908
+ )
2909
+ superscript
2910
+ subscript
2911
+ 𝑠
2912
+ 𝑡
2913
+ 𝑖
2914
+ s_{t}^{(i)}
2915
+ to children
2916
+ end
2917
+ for
2918
+ end
2919
+ if
2920
+ if
2921
+ s
2922
+ t
2923
+ subscript
2924
+ 𝑠
2925
+ 𝑡
2926
+ s_{t}
2927
+ is terminal
2928
+ then
2929
+ ▷
2930
+ ▷
2931
+ \triangleright
2932
+ Reflection
2933
+ Get
2934
+ r
2935
+ 𝑟
2936
+ r
2937
+ from environment
2938
+ if
2939
+ r
2940
+ 𝑟
2941
+ r
2942
+ not success
2943
+ then
2944
+ reflection
2945
+ ←
2946
+ p
2947
+ ref
2948
+ ​
2949
+ (
2950
+ c
2951
+ t
2952
+ )
2953
+ ←
2954
+ reflection
2955
+ subscript
2956
+ 𝑝
2957
+ ref
2958
+ subscript
2959
+ 𝑐
2960
+ 𝑡
2961
+ \text{reflection}\leftarrow p_{\text{ref}}(c_{t})
2962
+ c
2963
+ ←
2964
+ reflection
2965
+ ←
2966
+ 𝑐
2967
+ reflection
2968
+ c\leftarrow\text{reflection}
2969
+ end
2970
+ if
2971
+ end
2972
+ if
2973
+ a
2974
+ t
2975
+ ←
2976
+ arg
2977
+ ⁡
2978
+ max
2979
+ a
2980
+ ∈
2981
+ e
2982
+ ​
2983
+ (
2984
+ s
2985
+ t
2986
+ )
2987
+ ⁡
2988
+ [
2989
+ V
2990
+ ​
2991
+ (
2992
+ s
2993
+ t
2994
+ )
2995
+ +
2996
+ w
2997
+ ​
2998
+ ln
2999
+ ⁡
3000
+ N
3001
+ ​
3002
+ (
3003
+ s
3004
+ t
3005
+ −
3006
+ 1
3007
+ )
3008
+ N
3009
+ ​
3010
+ (
3011
+ s
3012
+ t
3013
+ )
3014
+ ]
3015
+ ←
3016
+ subscript
3017
+ 𝑎
3018
+ 𝑡
3019
+ subscript
3020
+ 𝑎
3021
+ 𝑒
3022
+ subscript
3023
+ 𝑠
3024
+ 𝑡
3025
+ 𝑉
3026
+ subscript
3027
+ 𝑠
3028
+ 𝑡
3029
+ 𝑤
3030
+ 𝑁
3031
+ subscript
3032
+ 𝑠
3033
+ 𝑡
3034
+ 1
3035
+ 𝑁
3036
+ subscript
3037
+ 𝑠
3038
+ 𝑡
3039
+ a_{t}\leftarrow\arg\max_{a\in e(s_{t})}\left[{V(s_{t})}+w\sqrt{\frac{\ln{N}(s_{t-1})}{{N}(s_{t})}}\right]
3040
+ ▷
3041
+ ▷
3042
+ \triangleright
3043
+ Selection
3044
+ N
3045
+ ​
3046
+ (
3047
+ s
3048
+ t
3049
+ +
3050
+ 1
3051
+ )
3052
+ ←
3053
+ N
3054
+ ​
3055
+ (
3056
+ s
3057
+ t
3058
+ +
3059
+ 1
3060
+ )
3061
+ +
3062
+ 1
3063
+ ←
3064
+ 𝑁
3065
+ subscript
3066
+ 𝑠
3067
+ 𝑡
3068
+ 1
3069
+ 𝑁
3070
+ subscript
3071
+ 𝑠
3072
+ 𝑡
3073
+ 1
3074
+ 1
3075
+ {N}(s_{t+1})\leftarrow{N}(s_{t+1})+1
3076
+ if
3077
+ a
3078
+ t
3079
+ subscript
3080
+ 𝑎
3081
+ 𝑡
3082
+ a_{t}
3083
+ is an output action
3084
+ then
3085
+ break
3086
+ end
3087
+ for
3088
+ T
3089
+ ←
3090
+ ←
3091
+ 𝑇
3092
+ absent
3093
+ T\leftarrow
3094
+ the actual number of steps
3095
+ for
3096
+ t
3097
+ ←
3098
+ T
3099
+ −
3100
+ 1
3101
+ ,
3102
+ …
3103
+ ,
3104
+ 0
3105
+ ←
3106
+ 𝑡
3107
+ 𝑇
3108
+ 1
3109
+ …
3110
+ 0
3111
+ t\leftarrow T-1,\dots,0
3112
+ do
3113
+ ▷
3114
+ ▷
3115
+ \triangleright
3116
+ Backpropagation
3117
+ V
3118
+ ​
3119
+ (
3120
+ s
3121
+ t
3122
+ )
3123
+ ←
3124
+ V
3125
+ ​
3126
+ (
3127
+ s
3128
+ t
3129
+ )
3130
+ ​
3131
+ (
3132
+ N
3133
+ ​
3134
+ (
3135
+ s
3136
+ t
3137
+ )
3138
+ −
3139
+ 1
3140
+ )
3141
+ +
3142
+ r
3143
+ N
3144
+ ​
3145
+ (
3146
+ s
3147
+ t
3148
+ )
3149
+ ←
3150
+ 𝑉
3151
+ subscript
3152
+ 𝑠
3153
+ 𝑡
3154
+ 𝑉
3155
+ subscript
3156
+ 𝑠
3157
+ 𝑡
3158
+ 𝑁
3159
+ subscript
3160
+ 𝑠
3161
+ 𝑡
3162
+ 1
3163
+ 𝑟
3164
+ 𝑁
3165
+ subscript
3166
+ 𝑠
3167
+ 𝑡
3168
+ V(s_{t})\leftarrow\frac{V(s_{t})(N(s_{t})-1)+r}{N(s_{t})}
3169
+ end
3170
+ for
3171
+ end
3172
+ for
3173
+ Appendix B
3174
+ Discussion
3175
+ Limitations.
3176
+ Although LATS can improve reasoning and decision-making, this arrives at a higher computational cost relative to simpler prompting methods like ReAct or Reflexion. The search process takes more time than standard prompting or simpler techniques, and requires greater inference costs. While such an issue is mitigated by the fact that the number of nodes
3177
+ n
3178
+ 𝑛
3179
+ n
3180
+ expanded at every step provides a natural trade-off between performance and efficiency (setting
3181
+ n
3182
+ =
3183
+ 1
3184
+ 𝑛
3185
+ 1
3186
+ n=1
3187
+ makes the method as effecient as ReAct with multiple trials or CoT-SC), in practice we recommend using LATS for difficult tasks like programming or for situations where performance is prioritized over efficiency. We hope that continued advancements in LLMs will reduce costs and increase the practicality of LATS.
3188
+ Additionally, the benchmarks we use in this paper are relatively simple and focused on decision-making, compared to the complexity of real-world interactive environments. In addition, some environments might not easily support rollbacks to previous states. However, the design of LATS is flexible and can be adjusted to various resource constraints. Using planning-based prompting methods like LATS in environments like Minecraft
3189
+ (Fan et al.,
3190
+ 2022
3191
+ )
3192
+ and more reasoning benchmarks would be interesting avenues for future work.
3193
+ Broader impact.
3194
+ LATS is a framework that enhances LLM performance through interactions with an environment. This improvement in autonomous decision-making may facilitate harmful uses of LLMs. Alternatively, LATS enhances interpretability and the potential for greater alignment, as it generates understandable, high-level linguistic reasoning and actions through several rounds of decision-making and reflection, rather than relying on implicit, low-level token values.
3195
+ Appendix C
3196
+ Ablations
3197
+ Prompt Method
3198
+ HotpotQA (EM)
3199
+ LATS (w=0.5)
3200
+ 0.55
3201
+ LATS (w=2.0)
3202
+ 0.61
3203
+ LATS (d=4)
3204
+ 0.58
3205
+ LATS (CoT)
3206
+ 0.60
3207
+ LATS (No LM Heuristic)
3208
+ 0.37
3209
+ LATS
3210
+ 0.61
3211
+ Table 6:
3212
+ Ablation results on LATS and baseline variants in HotPotQA measured by Exact Match (EM). We test different depth
3213
+ d
3214
+ 𝑑
3215
+ d
3216
+ , exploration factor
3217
+ w
3218
+ 𝑤
3219
+ w
3220
+ , and versions of LATS using CoT and without the LM value function. We sample
3221
+ n
3222
+ =
3223
+ 5
3224
+ 𝑛
3225
+ 5
3226
+ n=5
3227
+ and
3228
+ k
3229
+ =
3230
+ 50
3231
+ 𝑘
3232
+ 50
3233
+ k=50
3234
+ trajectories.
3235
+ Figure 4:
3236
+ Performance over successive iterations on HumanEval with GPT-3.5.
3237
+ In this section, we ablate various designs of LATS. Experiments are conducted on HotPotQA with a maximum of
3238
+ k
3239
+ =
3240
+ 50
3241
+ 𝑘
3242
+ 50
3243
+ k=50
3244
+ trajectories and sampling size of
3245
+ n
3246
+ =
3247
+ 5
3248
+ 𝑛
3249
+ 5
3250
+ n=5
3251
+ and HumanEval with a maximum of
3252
+ k
3253
+ =
3254
+ 8
3255
+ 𝑘
3256
+ 8
3257
+ k=8
3258
+ trajectories and sampling size of
3259
+ n
3260
+ =
3261
+ 5
3262
+ 𝑛
3263
+ 5
3264
+ n=5
3265
+ . The result for HotPotQA is shown in Tab.
3266
+ 5
3267
+ and HumanEval in Fig.
3268
+ 4
3269
+ .
3270
+ Exploration weight.
3271
+ We find that there is lower performance on HotPotQA when the exploration weight
3272
+ w
3273
+ 𝑤
3274
+ w
3275
+ in the selection formula is decreased to
3276
+ 0.5
3277
+ 0.5
3278
+ 0.5
3279
+ , suggesting that this reduces the effectiveness of the search. Increasing
3280
+ w
3281
+ 𝑤
3282
+ w
3283
+ to
3284
+ 2.0
3285
+ 2.0
3286
+ 2.0
3287
+ does not lead to a performance improvement, but we tend to observe faster convergence. The optimal setting depends on the particular environment and complexity of the state space.
3288
+ Depth.
3289
+ In our main experiments we use a maximum depth of
3290
+ d
3291
+ =
3292
+ 7
3293
+ 𝑑
3294
+ 7
3295
+ d=7
3296
+ on HotPotQA for all methods, following previous work
3297
+ (Yao et al.,
3298
+ 2023b
3299
+ )
3300
+ . We ablate the effect on LATS after reducing it to
3301
+ d
3302
+ =
3303
+ 4
3304
+ 𝑑
3305
+ 4
3306
+ d=4
3307
+ . This results in only a slight drop in performance. We find that most questions can be answered within four steps, and using a greater number of steps tends to force the agent into local minima and rarely improves success.
3308
+ LM value function.
3309
+ The LM value function scores states based on expected future reward. Without this heuristic, the only signal to guide search would be from environment rewards for completed trajectories, which are scarce and often binary. When we remove the evaluation operation, we observe a dramatic
3310
+ 0.24
3311
+ 0.24
3312
+ 0.24
3313
+ drop in performance.
3314
+ Performance over time.
3315
+ To see the effects of increasing the number of trajectories sampled, we change
3316
+ k
3317
+ 𝑘
3318
+ k
3319
+ to different values. We conduct this experiment on HumanEval, which has a more noticeable difference due to sampling less trajectories. The results are shown in Fig.
3320
+ 4
3321
+ , in which LATS scales better with more iterations than Reflexion.
3322
+ Sample complexity and Token cost.
3323
+ One possible concern of LATS is that the tree-structured search might consume much more tokens than existing methods. To further study the computational cost of LATS compared to prior methods, we examine the sample complexity (i.e. asymptotic token cost) of all methods considered in this paper, and count the average number of nodes expanded by our method and other tree-structured methods (ToT and RAP) upon successful search on HotPotQA. We present the results in Tab.
3324
+ 7
3325
+ ; the result shows that our method has the same sample complexity as other tree-based search methods, and has less average number of nodes expanded upon success, which indicates less token cost. The token cost gap will be even larger when taking failed trajectories into account, since our method has higher success rate and reaches computational budget limit less often.
3326
+ Method
3327
+ Performance (
3328
+ ↑
3329
+ ↑
3330
+ \uparrow
3331
+ )
3332
+ Sample complexity (
3333
+ ↓
3334
+ ↓
3335
+ \downarrow
3336
+ )
3337
+ Avg. #nodes upon success (
3338
+ ↓
3339
+ ↓
3340
+ \downarrow
3341
+ )
3342
+ ReAct (Best
3343
+ k
3344
+ =
3345
+ 250
3346
+ 𝑘
3347
+ 250
3348
+ k=250
3349
+ )
3350
+ 0.42
3351
+ 0.42
3352
+ 0.42
3353
+ O
3354
+ ​
3355
+ (
3356
+ k
3357
+ )
3358
+ 𝑂
3359
+ 𝑘
3360
+ O(k)
3361
+ N/A
3362
+ CoT-SC (
3363
+ n
3364
+ =
3365
+ 1
3366
+ ,
3367
+ k
3368
+ =
3369
+ 250
3370
+ formulae-sequence
3371
+ 𝑛
3372
+ 1
3373
+ 𝑘
3374
+ 250
3375
+ n=1,k=250
3376
+ )
3377
+ 0.40
3378
+ 0.40
3379
+ 0.40
3380
+ O
3381
+ ​
3382
+ (
3383
+ k
3384
+ )
3385
+ 𝑂
3386
+ 𝑘
3387
+ O(k)
3388
+ N/A
3389
+ LATS (
3390
+ n
3391
+ =
3392
+ 1
3393
+ ,
3394
+ k
3395
+ =
3396
+ 50
3397
+ formulae-sequence
3398
+ 𝑛
3399
+ 1
3400
+ 𝑘
3401
+ 50
3402
+ n=1,k=50
3403
+ )
3404
+ 0.48
3405
+ 0.48
3406
+ 0.48
3407
+ O
3408
+ ​
3409
+ (
3410
+ k
3411
+ )
3412
+ 𝑂
3413
+ 𝑘
3414
+ O(k)
3415
+ N/A
3416
+ ToT (ReAct)
3417
+ 0.49
3418
+ 0.49
3419
+ 0.49
3420
+ O
3421
+ ​
3422
+ (
3423
+ k
3424
+ ​
3425
+ n
3426
+ )
3427
+ 𝑂
3428
+ 𝑘
3429
+ 𝑛
3430
+ O(kn)
3431
+ 84.05
3432
+ 84.05
3433
+ 84.05
3434
+ RAP (ReAct)
3435
+ 0.54
3436
+ 0.54
3437
+ 0.54
3438
+ O
3439
+ ​
3440
+ (
3441
+ k
3442
+ ​
3443
+ n
3444
+ )
3445
+ 𝑂
3446
+ 𝑘
3447
+ 𝑛
3448
+ O(kn)
3449
+ 70.60
3450
+ 70.60
3451
+ 70.60
3452
+ LATS (
3453
+ n
3454
+ =
3455
+ 5
3456
+ ,
3457
+ k
3458
+ =
3459
+ 50
3460
+ formulae-sequence
3461
+ 𝑛
3462
+ 5
3463
+ 𝑘
3464
+ 50
3465
+ n=5,k=50
3466
+ )
3467
+ 0.61
3468
+ 0.61
3469
+ 0.61
3470
+ O
3471
+ ​
3472
+ (
3473
+ k
3474
+ ​
3475
+ n
3476
+ )
3477
+ 𝑂
3478
+ 𝑘
3479
+ 𝑛
3480
+ O(kn)
3481
+ 66.65
3482
+ 66.65
3483
+ 66.65
3484
+ Table 7:
3485
+ The performance, sample complexity of different methods and average number of nodes expanded upon success by methods with tree-based search.
3486
+ n
3487
+ 𝑛
3488
+ n
3489
+ is the number of children nodes expanded at every step and
3490
+ k
3491
+ 𝑘
3492
+ k
3493
+ is the number of trajectories. Our method has the same sample complexity as other methods with tree-based search and expands less nodes upon success, which indicates lower token cost.
3494
+ Appendix D
3495
+ Environment Details
3496
+ D.1
3497
+ HotPotQA
3498
+ Figure 5:
3499
+ Example trajectories on HotPotQA for ReAct (left) and LATS (right). LATS can sample more actions and avoid failure from previous mistakes by evaluating states with an LM to guide the search toward promising areas of the tree.
3500
+ HotPotQA
3501
+ (Yang et al.,
3502
+ 2018
3503
+ )
3504
+ is a question-answering dataset that requires reasoning over multiple supporting documents to answer questions. It contains 113k Wikipedia-based question-answer pairs crafted by crowdworkers to be diverse, multi-hop, and explainable. Questions cover a range of types like entities, locations, dates, and comparison of shared properties between two entities. Crowdworkers also provide supporting facts from the documents that justify the answer. We use the HotPotQA benchmark setting with all the Wikipedia paragraphs to test retrieval. We use a randomly selected subset of 100 questions for our experiments and a maximum depth limit of 6. Fig.
3505
+ 5
3506
+ illustrates how ReAct and LATS work on an example task of HotPotQA, and gives a qualitative example on how LATS outperforms ReAct on the task.
3507
+ Action Space.
3508
+ We adopt the Wikipedia web API proposed in
3509
+ Yao et al. (
3510
+ 2023b
3511
+ )
3512
+ , with three types of actions to support interactive information retrieval:
3513
+ (1)
3514
+ search
3515
+ [
3516
+ entity
3517
+ ], which returns the first 5 sentences from the corresponding
3518
+ entity
3519
+ wiki page if it exists, or else suggests top-5 similar entities from the Wikipedia search engine,
3520
+ (2)
3521
+ lookup
3522
+ [
3523
+ string
3524
+ ], which returns the next sentence in the page containing
3525
+ string
3526
+ ,
3527
+ (3)
3528
+ finish
3529
+ [
3530
+ answer
3531
+ ], which finishes the current task with
3532
+ answer
3533
+ .
3534
+ These API calls and free-form thoughts form the action space for this environment.
3535
+ D.2
3536
+ Programming
3537
+ The HumanEval dataset
3538
+ (Chen et al.,
3539
+ 2021
3540
+ )
3541
+ is a collection of 164 handwritten programming problems introduced to evaluate the functional correctness of models for synthesizing programs from natural language descriptions. Each problem includes a function signature, docstring description, reference implementation, and multiple unit tests, with an average of 7.7 tests per problem. The programming tasks assess comprehension of natural language, reasoning, algorithms, and basic mathematics, at a difficulty level comparable to simple software interview questions. Pass rates are evaluated with the pass@k metric, where k samples are generated per problem and a problem is considered solved if any sample passes all tests. We use all 164 problems for our experiments and a maximum depth limit of 8.
3542
+ The Mostly Basic Programming Problems (MBPP)
3543
+ Austin et al. (
3544
+ 2021
3545
+ )
3546
+ benchmark contains 974 short Python functions designed to evaluate program synthesis techniques. The dataset was constructed by crowdsourcing from workers with basic Python knowledge. Each data point consists of a natural language description of a programming task, a reference solution implementation, and three test cases for functional correctness. The natural language prompts are typically short, one-sentence descriptions. Solutions cover common programming constructs including mathematical operations, list processing, string manipulation, and usage of the Python standard library. On average, solutions are 6.8 lines of code. The dataset is also supplemented with an additional set of 426 problems that were manually verified for unambiguous specifications, standard function signatures, and accurate test cases. We use a randomly selected subset of 397 problems for our experiments.
3547
+ D.3
3548
+ WebShop
3549
+ WebShop
3550
+ (Yao et al.,
3551
+ 2022
3552
+ )
3553
+ is an interactive web-based environment designed to evaluate agents on grounded language understanding and decision-making. It simulates an e-commerce shopping task by providing agents with over 1 million real-world products scraped from Amazon, spanning 5 categories and 113 subcategories. These products contain rich linguistic information, with an average text length of 262 words and a vocabulary size of 224k. In addition, there are over 800k unique product options available for customization. The environment renders webpages in two modes: HTML mode provides pixel-level observations with interactive elements, while simple mode converts the raw HTML into a structured text observation more amenable for training agents. The action space consists of query searches and button clicks, which transition between 4 page types: search, results, item and item-detail. Instructions are crowdsourced natural language specifying product attributes and options, with a total of 12k collected. Automatic rewards are computed by comparing the product purchased by the agent against the attributes and options specified in the instruction, using both lexical matching and semantic similarity metrics.
3554
+ Type
3555
+ Argument
3556
+ State
3557
+ →
3558
+ →
3559
+ \rightarrow
3560
+ Next State
3561
+ search
3562
+ [
3563
+ Query
3564
+ ]
3565
+ Search
3566
+ →
3567
+ →
3568
+ \rightarrow
3569
+ Results
3570
+ choose
3571
+ Back to search
3572
+ ∗
3573
+ *
3574
+ →
3575
+ →
3576
+ \rightarrow
3577
+ Search
3578
+ choose
3579
+ Prev/Next page
3580
+ Results
3581
+ →
3582
+ →
3583
+ \rightarrow
3584
+ Results
3585
+ choose
3586
+ [
3587
+ Product title
3588
+ ]
3589
+ Results
3590
+ →
3591
+ →
3592
+ \rightarrow
3593
+ Item
3594
+ choose
3595
+ [
3596
+ Option
3597
+ ]
3598
+ Item
3599
+ →
3600
+ →
3601
+ \rightarrow
3602
+ Item
3603
+ choose
3604
+ Desc/Overview
3605
+ Item
3606
+ →
3607
+ →
3608
+ \rightarrow
3609
+ Item-Detail
3610
+ choose
3611
+ Previous
3612
+ Item-Detail
3613
+ →
3614
+ →
3615
+ \rightarrow
3616
+ Item
3617
+ choose
3618
+ Buy
3619
+ Item
3620
+ →
3621
+ →
3622
+ \rightarrow
3623
+ Episode End
3624
+ Table 8:
3625
+ Action space of webshop.
3626
+ There are two evaluation metrics used in WebShop: (1)
3627
+ Task Score
3628
+ : defined as
3629
+ (
3630
+ 100
3631
+ ×
3632
+ avg. reward
3633
+ )
3634
+ 100
3635
+ avg. reward
3636
+ (100\times\text{avg. reward})
3637
+ , which captures the average reward obtained across episodes; and (2)
3638
+ Success Rate (SR)
3639
+ defined as the portion of instructions where
3640
+ r
3641
+ =
3642
+ 1
3643
+ 𝑟
3644
+ 1
3645
+ r=1
3646
+ . The reward is calculated based on the number of attributes satisfied by the selected item. We use 50 environments for our experiments and a maximum depth limit of 15.
3647
+ Appendix E
3648
+ HotPotQA Prompts
3649
+ E.1
3650
+ Base Acting Prompt
3651
+ Solve a question answering task with interleaving Thought, Action, Observation steps. Thought can reason about the current situation, and Action can be three types:
3652
+ (1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
3653
+ (2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
3654
+ (3) Finish[answer], which returns the answer and finishes the task.
3655
+ After each observation, provide the next Thought and next Action.
3656
+ Here are some examples:
3657
+ Question: Which magazine was started first Arthur’s Magazine or First for Women?
3658
+ Thought 1: I need to search Arthur’s Magazine and First for Women, and find which was started first.
3659
+ Action 1: Search[Arthur’s Magazine]
3660
+ Observation 1: Arthur’s Magazine (18440̆0e20̆0800̆0931846) was an American literary periodical published in Philadelphia in the 19th century. Edited by Timothy Shay Arthur, it featured work by Edgar A. Poe, J.H. Ingraham, Sarah Josepha Hale, Thomas G. Spear, and others.[1][2] In May 1846 it was merged into Godey’s Lady’s Book.[3]
3661
+ Thought 2: Arthur’s Magazine was started in 1844. I need to search First for Women next.
3662
+ Action 2: Search[First for Women]
3663
+ Observation 2: First for Women is a woman’s magazine published by Bauer Media Group in the USA.[1] The magazine was started in 1989. It is based in Englewood Cliffs, New Jersey.[2] In 2011 the circulation of the magazine was 1,310,696 copies.[3]
3664
+ Thought 3: First for Women was started in 1989. 1844 (Arthur’s Magazine) ¡ 1989 (First for Women), so Arthur’s Magazine was started first.
3665
+ Action 3: Finish[Arthur’s Magazine]
3666
+ (examples)
3667
+ You have attempted to answer the following question before and failed. The following reflection(s) give a plan to avoid failing to answer the question in the same way you did previously. Use them to improve your strategy of correctly answering the given question.
3668
+ (trajectories)
3669
+ (input)
3670
+ E.2
3671
+ Base Reasoning Prompt
3672
+ Solve a question answering task by having a Thought, then Finish with your answer. Thought can reason about the current situation. Finish[answer] returns the answer and finishes the task. You will be given context that you should use to help you answer the question. Start your response with either Action or an indexed Thought
3673
+ Here are some examples:
3674
+ Question: What is the elevation range for the area that the eastern sector of the Colorado orogeny extends into?
3675
+ Let’s think step by step.
3676
+ Thought 1: The eastern sector of Colorado orogeny extends into the High Plains.
3677
+ Thought 2: High Plains rise in elevation from around 1,800 to 7,000 ft
3678
+ Thought 3: The answer is 1,800 to 7,000 ft.
3679
+ Action: Finish[1,800 to 7,000 ft]
3680
+ (examples)
3681
+ Previous trial:
3682
+ (trajectories)
3683
+ (input)
3684
+ E.3
3685
+ Value Function Prompt
3686
+ Analyze the trajectories of a solution to a question answering task. The trajectories are labeled by environmental observations about the situation, thoughts that can reason about the current situation and actions that can be three types:
3687
+ (1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
3688
+ (2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
3689
+ (3) Finish[answer], which returns the answer and finishes the task.
3690
+ Given a question and a trajectory, evaluate its correctness and provide your reasoning and analysis in detail. Focus on the latest thought, action, and observation. Incomplete trajectories can be correct if the thoughts and actions so far are correct, even if the answer is not found yet. Do not generate additional thoughts or actions. Then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
3691
+ Question: Which magazine was started first Arthur’s Magazine or First for Women?
3692
+ Thought 1: I need to search Arthur’s Magazine and First for Women, and find which was started first.
3693
+ Action 1: Search[Arthur’s Magazine]
3694
+ Observation 1: Arthur’s Magazine (18440̆0e20̆0800̆0931846) was an American literary periodical published in Philadelphia in the 19th century. Edited by Timothy Shay Arthur, it featured work by Edgar A. Poe, J.H. Ingraham, Sarah Josepha Hale, Thomas G. Spear, and others.[1][2] In May 1846 it was merged into Godey’s Lady’s Book.[3]
3695
+ This trajectory is correct as it is reasonable to search for the first magazine provided in the question. It is also better to have simple searches corresponding to a single entity, making this the best action.
3696
+ Thus the correctness score is 10
3697
+ (other examples)
3698
+ (failed trajectories)
3699
+ (context)
3700
+ E.4
3701
+ Reflection Prompt
3702
+ Analyze the trajectories of a solution to a question answering task. The trajectories are labeled by environmental observations about the situation, thoughts that can reason about the current situation and actions that can be three types:
3703
+ (1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
3704
+ (2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
3705
+ (3) Finish[answer], which returns the answer and finishes the task.
3706
+ Given a question and a trajectory, evaluate its correctness and provide your reasoning and analysis in detail. Focus on the latest thought, action, and observation. Incomplete trajectories can be correct if the thoughts and actions so far are correct, even if the answer is not found yet. Do not generate additional thoughts or actions. Then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
3707
+ Question: Which magazine was started first Arthur’s Magazine or First for Women?
3708
+ Thought 1: I need to search Arthur’s Magazine and First for Women, and find which was started first.
3709
+ Action 1: Search[Arthur’s Magazine]
3710
+ Observation 1: Arthur’s Magazine (18440̆0e20̆0800̆0931846) was an American literary periodical published in Philadelphia in the 19th century. Edited by Timothy Shay Arthur, it featured work by Edgar A. Poe, J.H. Ingraham, Sarah Josepha Hale, Thomas G. Spear, and others.[1][2] In May 1846 it was merged into Godey’s Lady’s Book.[3]
3711
+ This trajectory is correct as it is reasonable to search for the first magazine provided in the question. It is also better to have simple searches corresponding to a single entity, making this the best action.
3712
+ Thus the correctness score is 10
3713
+ (other examples)
3714
+ (failed trajectories)
3715
+ (context)
3716
+ Appendix F
3717
+ Programming Prompts
3718
+ F.1
3719
+ HumanEval function implementation example
3720
+ Sample function signature:
3721
+ ⬇
3722
+ def
3723
+ minSubArraySum
3724
+ (
3725
+ nums
3726
+ ):
3727
+ Given
3728
+ an
3729
+ array
3730
+ of
3731
+ integers
3732
+ nums
3733
+ ,
3734
+ find
3735
+ the
3736
+ minimum
3737
+ sum
3738
+ of
3739
+ any
3740
+ non
3741
+ -
3742
+ empty
3743
+ sub
3744
+ -
3745
+ array
3746
+ of
3747
+ nums
3748
+ .
3749
+ Example
3750
+ minSubArraySum
3751
+ ([2,
3752
+ 3,
3753
+ 4,
3754
+ 1,
3755
+ 2,
3756
+ 4])
3757
+ ==
3758
+ 1
3759
+ minSubArraySum
3760
+ ([-1,
3761
+ -2,
3762
+ -3])
3763
+ ==
3764
+ -6
3765
+ Sample function body implementation:
3766
+ ⬇
3767
+ min_sum
3768
+ =
3769
+ float
3770
+ (’
3771
+ inf
3772
+ ’)
3773
+ for
3774
+ i
3775
+ in
3776
+ range
3777
+ (
3778
+ len
3779
+ (
3780
+ nums
3781
+ )):
3782
+ current_sum
3783
+ =
3784
+ 0
3785
+ for
3786
+ j
3787
+ in
3788
+ range
3789
+ (
3790
+ i
3791
+ ,
3792
+ len
3793
+ (
3794
+ nums
3795
+ )):
3796
+ current_sum
3797
+ +=
3798
+ nums
3799
+ [
3800
+ j
3801
+ ]
3802
+ if
3803
+ current_sum
3804
+ <
3805
+ min_sum
3806
+ :
3807
+ min_sum
3808
+ =
3809
+ current_sum
3810
+ return
3811
+ min_sum
3812
+ F.2
3813
+ Base Acting/Reasoning Prompt
3814
+ You are an AI Python assistant. You will be given your previous implementation of a function, a series of unit tests results, and your self-reflection on your previous implementation. Write your full implementation (restate the function signature).
3815
+ Example 1:
3816
+ [previous impl]:
3817
+ ⬇
3818
+ def
3819
+ add
3820
+ (
3821
+ a
3822
+ :
3823
+ int
3824
+ ,
3825
+ b
3826
+ :
3827
+ int
3828
+ )
3829
+ ->
3830
+ int
3831
+ :
3832
+ ”””
3833
+ Given
3834
+ integers
3835
+ a
3836
+ and
3837
+ b
3838
+ ,
3839
+ return
3840
+ the
3841
+ total
3842
+ value
3843
+ of
3844
+ a
3845
+ and
3846
+ b
3847
+ .
3848
+ ”””
3849
+ return
3850
+ a
3851
+ -
3852
+ b
3853
+ [unit test results from previous impl]:
3854
+ Tested passed:
3855
+ Tests failed:
3856
+ assert add(1, 2) == 3 # output: -1
3857
+ assert add(1, 2) == 4 # output: -1
3858
+ [reflection on previous impl]:
3859
+ The implementation failed the test cases where the input integers are 1 and 2. The issue arises because the code does not add the two integers together, but instead subtracts the second integer from the first. To fix this issue, we should change the operator from ‘-‘ to ‘+‘ in the return statement. This will ensure that the function returns the correct output for the given input.
3860
+ [improved impl]:
3861
+ ⬇
3862
+ def
3863
+ add
3864
+ (
3865
+ a
3866
+ :
3867
+ int
3868
+ ,
3869
+ b
3870
+ :
3871
+ int
3872
+ )
3873
+ ->
3874
+ int
3875
+ :
3876
+ ”””
3877
+ Given
3878
+ integers
3879
+ a
3880
+ and
3881
+ b
3882
+ ,
3883
+ return
3884
+ the
3885
+ total
3886
+ value
3887
+ of
3888
+ a
3889
+ and
3890
+ b
3891
+ .
3892
+ ”””
3893
+ return
3894
+ a
3895
+ +
3896
+ b
3897
+ F.3
3898
+ Reflection Prompt
3899
+ You are a Python programming assistant. You will be given a function implementation and a series of unit test results. Your goal is to write a few sentences to explain why your implementation is wrong as indicated by the tests. You will need this as guidance when you try again later. Only provide the few sentence description in your answer, not the implementation. You will be given a few examples by the user.
3900
+ Example 1:
3901
+ [previous impl]:
3902
+ ⬇
3903
+ def
3904
+ add
3905
+ (
3906
+ a
3907
+ :
3908
+ int
3909
+ ,
3910
+ b
3911
+ :
3912
+ int
3913
+ )
3914
+ ->
3915
+ int
3916
+ :
3917
+ ”””
3918
+ Given
3919
+ integers
3920
+ a
3921
+ and
3922
+ b
3923
+ ,
3924
+ return
3925
+ the
3926
+ total
3927
+ value
3928
+ of
3929
+ a
3930
+ and
3931
+ b
3932
+ .
3933
+ ”””
3934
+ return
3935
+ a
3936
+ -
3937
+ b
3938
+ [unit test results from previous impl]:
3939
+ Tested passed:
3940
+ Tests failed:
3941
+ assert add(1, 2) == 3 # output: -1
3942
+ assert add(1, 2) == 4 # output: -1
3943
+ [reflection on previous impl]:
3944
+ The implementation failed the test cases where the input integers are 1 and 2. The issue arises because the code does not add the two integers together, but instead subtracts the second integer from the first. To fix this issue, we should change the operator from ‘-‘ to ‘+‘ in the return statement. This will ensure that the function returns the correct output for the given input.
3945
+ F.4
3946
+ Test Case Generation Prompt
3947
+ You are an AI coding assistant that can write unique, diverse, and intuitive unit tests for functions given the signature and docstring.
3948
+ (input)
3949
+ Appendix G
3950
+ WebShop Prompts
3951
+ G.1
3952
+ Acting Prompt
3953
+ Webshop
3954
+ Instruction:
3955
+ i would like a 3 ounce bottle of bright citrus deodorant for sensitive skin, and price lower than 50.00 dollars
3956
+ [Search]
3957
+ Action: search[3 ounce bright citrus deodorant sensitive skin]
3958
+ Observation:
3959
+ [Back to Search]
3960
+ Page 1 (Total results: 50)
3961
+ [Next
3962
+ ]
3963
+ [B078GWRC1J]
3964
+ Bright Citrus Deodorant by Earth Mama — Natural and Safe for Sensitive Skin, Pregnancy and Breastfeeding, Contains Organic Calendula 3-Ounce
3965
+ $10.99
3966
+ [B078GTKVXY]
3967
+ Ginger Fresh Deodorant by Earth Mama — Natural and Safe for Sensitive Skin, Pregnancy and Breastfeeding, Contains Organic Calendula 3-Ounce
3968
+ $10.99
3969
+ [B08KBVJ4XN]
3970
+ Barrel and Oak - Aluminum-Free Deodorant, Deodorant for Men, Essential Oil-Based Scent, 24-Hour Odor Protection, Cedar & Patchouli Blend, Gentle on Sensitive Skin (Mountain Sage, 2.7 oz, 2-Pack)
3971
+ $15.95
3972
+ Action: think[B078GWRC1J and B078GTKVXY are bright citrus deodorant less then 50 dollars. I can check B078GWRC1J first.]
3973
+ Observation: OK.
3974
+ Action: click[B078GWRC1J]
3975
+ Observation:
3976
+ [Back to Search]
3977
+ [Prev]
3978
+ scent [assorted scents][bright citrus][calming lavender][ginger fresh][simply non-scents]
3979
+ size [travel set (4-pack)][3 ounce (pack of 1)][3-ounce (2-pack)]
3980
+ Bright Citrus Deodorant by Earth Mama — Natural and Safe for Sensitive Skin, Pregnancy and Breastfeeding, Contains Organic Calendula 3-Ounce
3981
+ Price: $10.99
3982
+ Rating: N.A.
3983
+ [Description]
3984
+ [Features]
3985
+ [Reviews]
3986
+ [Buy Now]
3987
+ Action: think[For 3 ounce bottle of bright citrus deodorant for sensitive skin, the item has options ’bright citrus’ and ’3 ounce (pack of 1)’ and seems good to buy.]
3988
+ Observation: OK.
3989
+ Action: click[bright citrus]
3990
+ Observation: You have clicked bright citrus.
3991
+ Action: click[3 ounce (pack of 1)]
3992
+ Observation: You have clicked 3 ounce (pack of 1).
3993
+ Action: click[Buy Now]
3994
+ G.2
3995
+ Value Function Prompt
3996
+ Given an item to purchase and a trajectory that aims to buy an item that exactly matches the specification, which corresponds to the ideal score of 1.0, analyze the following trajectory, then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
3997
+ Here are some examples
3998
+ Webshop
3999
+ Instruction:
4000
+ i am looking for dairy free and apple variety pack of chips, and price lower than 30.00 dollars
4001
+ [Search]
4002
+ Action: search[dairy free and apple variety pack of chips]
4003
+ Observation:
4004
+ [Back to Search]
4005
+ Page 1 (Total results: 50)
4006
+ [Next]
4007
+ [B07HRFJWP8]
4008
+ Enjoy Life Foods Soft Baked Ovals, Breakfast Bars, Nut Free Bars, Soy Free, Dairy Free, Non GMO, Gluten Free, Vegan, Variety Pack, 4 Boxes (20 Bars Total)
4009
+ $100.0
4010
+ [B01KMHY5PG]
4011
+ Enjoy Life Soft Baked Chewy Bars, Variety Pack, Nut Free Bars, Soy Free, Dairy Free, Gluten Free, 6 Boxes (30 Total Bars)
4012
+ $21.49
4013
+ [B008D2X8C4]
4014
+ Enjoy Life Lentil Chips Variety Pack, Dairy Free Chips, Soy Free, Nut Free, Non GMO, Vegan, Gluten Free, 24 Bags (0.8 oz)
4015
+ $100.0
4016
+ Action: think[The first two items are not what I want. The third item is the dairy free and apple variety pack of chips I want.]
4017
+ Observation: OK.
4018
+ Action: click[B008D2X8C4]
4019
+ Observation:
4020
+ [Back to Search]
4021
+ [Prev]
4022
+ flavor name [dill and sour cream][garlic & parmesan][light sea salt][margherita pizza][thai chili lime][variety pack]
4023
+ size [0.8 ounce (pack of 24)][4 ounce (pack of 12)]
4024
+ Price: $100.0
4025
+ Rating: N.A.
4026
+ [Description]
4027
+ [Features]
4028
+ [Reviews]
4029
+ [Buy Now]
4030
+ Action: think[The item has options ’variety pack’ and ’0.8 ounce (pack of 24)’ and seems good to buy.]
4031
+ Observation: OK.
4032
+ Action: click[variety pack]
4033
+ Observation: You have clicked variety pack.
4034
+ Action: click[0.8 ounce (pack of 24)]
4035
+ Observation: You have clicked 0.8 ounce (pack of 24).
4036
+ Action: click[Buy Now]
4037
+ Reflection: In this attempt, I was unsuccessful. I accidentally bought a product that was $100, which is more than my budget of $30. Either way, the initial search results were not good. Next time, I will do search[”variety pack of chips”] and then check if the results meet the dairy free and the $30 budget constraints. I will continue to refine my searches so that I can find more products.
4038
+ Thus the correctness score is 5
4039
+ (input)
4040
+ G.3
4041
+ Reflection Prompt
4042
+ Given an item to purchase and a trajectory that aims to buy an item that exactly matches the specification, which corresponds to the ideal score of 1.0, analyze the following trajectory, then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
4043
+ Here are some examples
4044
+ Webshop
4045
+ Instruction:
4046
+ i am looking for dairy free and apple variety pack of chips, and price lower than 30.00 dollars
4047
+ [Search]
4048
+ Action: search[dairy free and apple variety pack of chips]
4049
+ Observation:
4050
+ [Back to Search]
4051
+ Page 1 (Total results: 50)
4052
+ [Next]
4053
+ [B07HRFJWP8]
4054
+ Enjoy Life Foods Soft Baked Ovals, Breakfast Bars, Nut Free Bars, Soy Free, Dairy Free, Non GMO, Gluten Free, Vegan, Variety Pack, 4 Boxes (20 Bars Total)
4055
+ $100.0
4056
+ [B01KMHY5PG]
4057
+ Enjoy Life Soft Baked Chewy Bars, Variety Pack, Nut Free Bars, Soy Free, Dairy Free, Gluten Free, 6 Boxes (30 Total Bars)
4058
+ $21.49
4059
+ [B008D2X8C4]
4060
+ Enjoy Life Lentil Chips Variety Pack, Dairy Free Chips, Soy Free, Nut Free, Non GMO, Vegan, Gluten Free, 24 Bags (0.8 oz)
4061
+ $100.0
4062
+ Action: think[The first two items are not what I want. The third item is the dairy free and apple variety pack of chips I want.]
4063
+ Observation: OK.
4064
+ Action: click[B008D2X8C4]
4065
+ Observation:
4066
+ [Back to Search]
4067
+ [Prev]
4068
+ flavor name [dill and sour cream][garlic & parmesan][light sea salt][margherita pizza][thai chili lime][variety pack]
4069
+ size [0.8 ounce (pack of 24)][4 ounce (pack of 12)]
4070
+ Price: $100.0
4071
+ Rating: N.A.
4072
+ [Description]
4073
+ [Features]
4074
+ [Reviews]
4075
+ [Buy Now]
4076
+ Action: think[The item has options ’variety pack’ and ’0.8 ounce (pack of 24)’ and seems good to buy.]
4077
+ Observation: OK.
4078
+ Action: click[variety pack]
4079
+ Observation: You have clicked variety pack.
4080
+ Action: click[0.8 ounce (pack of 24)]
4081
+ Observation: You have clicked 0.8 ounce (pack of 24).
4082
+ Action: click[Buy Now]
4083
+ Reflection: In this attempt, I was unsuccessful. I accidentally bought a product that was $100, which is more than my budget of $30. Either way, the initial search results were not good. Next time, I will do search[”variety pack of chips”] and then check if the results meet the dairy free and the $30 budget constraints. I will continue to refine my searches so that I can find more products.
4084
+ (input)
4085
+ Reflection:
4086
+ ◄
4087
+ Feeling
4088
+ lucky?
4089
+ Conversion
4090
+ report
4091
+ Report
4092
+ an issue
4093
+ View original
4094
+ on arXiv
4095
+ ►
research/notes/231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la-3.md ADDED
@@ -0,0 +1,4095 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2310.04406] Language Agent Tree Search Unifies Reasoning Acting and Planning
3
+ in Language Models'
4
+ id: 231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la-3
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:40:52.405072Z'
8
+ source: https://ar5iv.labs.arxiv.org/html/2310.04406
9
+ source_domain: ar5iv.labs.arxiv.org
10
+ fetched_at: '2026-06-10T00:40:52.404928Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2310.04406] Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
20
+ Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
21
+ Andy Zhou
22
+ University of Illinois at Urbana-Champaign
23
+ AI@UIUC
24
+ Kai Yan
25
+ University of Illinois at Urbana-Champaign
26
+ Michal Shlapentokh-Rothman
27
+ University of Illinois at Urbana-Champaign
28
+ Haohan Wang
29
+ University of Illinois at Urbana-Champaign
30
+ Yu-Xiong Wang
31
+ University of Illinois at Urbana-Champaign
32
+ Abstract
33
+ While large language models (LLMs) have demonstrated impressive performance on a range of decision-making tasks, they rely on simple acting processes and fall short of broad deployment as autonomous agents. We introduce LATS (Language Agent Tree Search), a general framework that synergizes the capabilities of LLMs in planning, acting, and reasoning. Drawing inspiration from Monte Carlo tree search commonly used in model-based reinforcement learning, LATS employs LLMs as agents, value functions, and optimizers, repurposing their latent strengths for enhanced decision-making. What is crucial in this method is the use of an environment for external feedback, which offers a more deliberate and adaptive problem-solving mechanism that moves beyond the limitations of existing techniques. Our experimental evaluation across diverse domains, such as programming, HotPotQA, and WebShop, illustrates the applicability of LATS for decision-making while maintaining competitive reasoning performance. In particular, LATS achieves 94.4% for programming on HumanEval with GPT-4 and an average score of 75.9 for web browsing on WebShop with GPT-3.5, demonstrating the effectiveness and generality of our method.
34
+ 1
35
+ Introduction
36
+ General autonomous agents capable of reasoning and decision-making in a variety of environments
37
+ (Wooldridge & Jennings,
38
+ 1995
39
+ )
40
+ have been of longstanding interest in the field of artificial intelligence. While this has traditionally been studied in reinforcement learning, the recent rise of large language models (LLMs)
41
+ (Brown et al.,
42
+ 2020
43
+ ; Chowdhery et al.,
44
+ 2022
45
+ ; Touvron et al.,
46
+ 2023
47
+ ; OpenAI,
48
+ 2023
49
+ )
50
+ with strong reasoning and general adaptability offers an alternative paradigm. Not only have LLMs excelled on standard NLP tasks such as text summarization
51
+ (Nallapati et al.,
52
+ 2016
53
+ )
54
+ or natural language inference
55
+ (Bowman et al.,
56
+ 2015
57
+ )
58
+ , but they have been adapted to an increasingly diverse set of tasks that often require advanced common-sense reasoning or quantitative skills
59
+ (Cobbe et al.,
60
+ 2021
61
+ ; Saparov & He,
62
+ 2022
63
+ )
64
+ . LLMs are also capable of performing in complex environments that involve knowledge and reasoning, such as web navigation
65
+ (Yao et al.,
66
+ 2022
67
+ ; Deng et al.,
68
+ 2023
69
+ )
70
+ , tool-use
71
+ (Schick et al.,
72
+ 2023
73
+ )
74
+ , or open-ended games
75
+ (Fan et al.,
76
+ 2022
77
+ )
78
+ .
79
+ Figure 1:
80
+ An overview of LATS. LATS uses an external environment and self-reflection to improve reasoning and decision-making.
81
+ Reasoning and acting abilities have also been improved by prompting techniques that augment LLMs with feedback or observations from an external environment
82
+ (Yao et al.,
83
+ 2023b
84
+ ; Gao et al.,
85
+ 2022
86
+ ; Shinn et al.,
87
+ 2023
88
+ )
89
+ . This eliminates the need to rely entirely on the base abilities of the Language Model (LM), enhancing it through external tools or semantic feedback. Despite this strength, these methods are reflexive and fall short of humans’ deliberate and thoughtful decision-making characteristics to solve problems
90
+ (Sloman,
91
+ 1996
92
+ ; Evans,
93
+ 2010
94
+ )
95
+ . In particular, such methods fail to consider multiple reasoning paths or to plan ahead. Recent search-guided LLM works
96
+ (Xie et al.,
97
+ 2023
98
+ ; Yao et al.,
99
+ 2023a
100
+ ; Hao et al.,
101
+ 2023
102
+ )
103
+ address this issue by searching over multiple reasoning chains. While these methods enable planning, these methods operate in isolation and do not incorporate external feedback that can improve reasoning.
104
+ To help address these issues, we propose LATS (Language Agent Tree Search), a general framework for decision-making and reasoning with language models. LATS unifies LM planning, acting, and reasoning strategies by expanding ReAct
105
+ (Yao et al.,
106
+ 2023b
107
+ )
108
+ into a search over a combinatorial space of possible reasoning and acting steps. We adapt Monte Carlo tree search (MCTS) from model-based reinforcement learning
109
+ (Silver et al.,
110
+ 2017
111
+ ; Anthony et al.,
112
+ 2017
113
+ ; Jiang et al.,
114
+ 2018
115
+ )
116
+ to language agents, repurposing a pretrained LLM as an agent, value function, and optimizer. Utilizing the strong natural language understanding and in-context learning ability of modern LMs, we use text as an interface between each component of the framework, allowing LATS to adapt planning to environmental conditions without additional training. To the best of our knowledge,
117
+ LATS is the first framework that combines reasoning, acting, and planning to enhance LLMs
118
+ . Notably, LATS doubles the performance of GPT-3.5 on HotPotQA
119
+ (Yang et al.,
120
+ 2018
121
+ )
122
+ over ReAct
123
+ (Yao et al.,
124
+ 2023b
125
+ )
126
+ and raises the average score by
127
+ 22.1
128
+ 22.1
129
+ 22.1
130
+ on WebShop
131
+ (Yao et al.,
132
+ 2022
133
+ )
134
+ . When used with GPT-4, LATS achieves a
135
+ 94.4
136
+ 94.4
137
+ 94.4
138
+ Pass@1 rate for programming on HumanEval
139
+ (Chen et al.,
140
+ 2021
141
+ )
142
+ , setting the state of the art. To summarize, our
143
+ contributions
144
+ are the following:
145
+ •
146
+ We introduce an LM-based Monte Carlo tree search variant to deliberately construct the best trajectory from sampled actions, enabling more flexible and adaptive problem-solving compared to reflexive prompting methods. This is guided by heuristics from the LM.
147
+ •
148
+ By integrating external feedback and self-reflection, LATS enhances model sensibility and enables agents to learn from experience, surpassing reasoning-based search methods.
149
+ •
150
+ Through experiments across diverse domains like programming, interactive QA, and web navigation, we demonstrate the versatility of LATS in harnessing LLMs for autonomous reasoning and decision-making.
151
+ 2
152
+ Related Work
153
+ Approach
154
+ Reasoning
155
+ Acting
156
+ Planning
157
+ Self
158
+ External
159
+ Reflection
160
+ Memory
161
+ CoT
162
+ (Wei et al.,
163
+ 2022
164
+ )
165
+ ✓
166
+ ×
167
+ \times
168
+ ×
169
+ \times
170
+ ×
171
+ \times
172
+ ×
173
+ \times
174
+ ReAct
175
+ (Yao et al.,
176
+ 2023b
177
+ )
178
+ ✓
179
+ ✓
180
+ ×
181
+ \times
182
+ ×
183
+ \times
184
+ ×
185
+ \times
186
+ ToT
187
+ (Yao et al.,
188
+ 2023a
189
+ )
190
+ ✓
191
+ ×
192
+ \times
193
+ ✓
194
+ ✓
195
+ ✓
196
+ RAP
197
+ (Hao et al.,
198
+ 2023
199
+ )
200
+ ✓
201
+ ×
202
+ \times
203
+ ✓
204
+ ×
205
+ \times
206
+ ✓
207
+ Self-Refine
208
+ (Madaan et al.,
209
+ 2023
210
+ )
211
+ ✓
212
+ ×
213
+ \times
214
+ ×
215
+ \times
216
+ ✓
217
+ ×
218
+ \times
219
+ Beam Search
220
+ (Xie et al.,
221
+ 2023
222
+ )
223
+ ✓
224
+ ×
225
+ \times
226
+ ×
227
+ \times
228
+ ✓
229
+ ×
230
+ \times
231
+ Reflexion
232
+ (Shinn et al.,
233
+ 2023
234
+ )
235
+ ✓
236
+ ✓
237
+ ×
238
+ \times
239
+ ✓
240
+ ✓
241
+ LATS (Ours)
242
+ ✓
243
+ ✓
244
+ ✓
245
+ ✓
246
+ ✓
247
+ Table 1:
248
+ A summary of related work on reasoning, acting, and planning. LATS is the first work incorporating designs from all three domains, allowing use in all corresponding tasks. We refer to planning as the use of a search algorithm, self-reflection as the use of LM-generated feedback, and external memory as storaging past text context for future updates of solution.
249
+ a) Tree-of-Thoughts
250
+ b) Reasoning via Planning
251
+ c) Language Agent Tree Search
252
+ Figure 2:
253
+ An overview of the differences between LATS and recently proposed LM search algorithms ToT
254
+ (Yao et al.,
255
+ 2023a
256
+ )
257
+ and RAP
258
+ (Hao et al.,
259
+ 2023
260
+ )
261
+ . LATS leverages environmental feedback and self-reflection to further adapt search and improve performance.
262
+ LLMs for reasoning.
263
+ For LLMs, reasoning typically involves decomposing complex inputs into sequential intermediate steps towards a final answer
264
+ (Cobbe et al.,
265
+ 2021
266
+ )
267
+ , demonstrated with Chain-of-Thought (CoT) prompting
268
+ (Wei et al.,
269
+ 2022
270
+ )
271
+ and its variants
272
+ (Wei et al.,
273
+ 2022
274
+ ; Kojima et al.,
275
+ 2022
276
+ ; Wang et al.,
277
+ 2022
278
+ )
279
+ . However, these methods, which create chains autoregressively in a single step, often suffer from error propagation as the number of steps increases
280
+ (Guo et al.,
281
+ 2018
282
+ ; Chen et al.,
283
+ 2022b
284
+ )
285
+ due to compound errors. Various advancements aim to mitigate this issue; some approaches, such as Self-Consistency
286
+ (Wang et al.,
287
+ 2022
288
+ )
289
+ , employ majority voting over sampled chains, while others focus on multi-step decomposition, such as least-to-most prompting
290
+ (Zhou et al.,
291
+ 2022
292
+ )
293
+ , or use of external tools such as a scratchpad
294
+ (Nye et al.,
295
+ 2021
296
+ )
297
+ or compiler
298
+ (Gao et al.,
299
+ 2022
300
+ )
301
+ . Recently, CoT has been improved with search algorithms
302
+ (Yao et al.,
303
+ 2023a
304
+ ; Hao et al.,
305
+ 2023
306
+ ; Besta et al.,
307
+ 2023
308
+ )
309
+ that can sample trajectories more effectively. Tree-of-thought (ToT) prompting
310
+ (Yao et al.,
311
+ 2023a
312
+ )
313
+ uses DFS or BFS-based search guided by an LM-generated heuristic while Reasoning via Planning (RAP)
314
+ (Hao et al.,
315
+ 2023
316
+ )
317
+ uses MCTS with rollouts simulated by the LM. However, they rely solely on LM internal knowledge and cannot adapt to useful external feedback.
318
+ LLMs for acting.
319
+ The strong reasoning and common-sense abilities of LLMs have also been adapted for decision-making or acting tasks as a policy model in interactive environments. In the realm of robotics LLMs have been employed as high-level controllers of control policies
320
+ (Ahn et al.,
321
+ 2022
322
+ ; Huang et al.,
323
+ 2022
324
+ ; Driess et al.,
325
+ 2023
326
+ )
327
+ . Similar work
328
+ (Baker et al.,
329
+ 2022
330
+ ; Wang et al.,
331
+ 2023
332
+ ; Zhu et al.,
333
+ 2023
334
+ )
335
+ has also adapted LLM agents to complex multimodal games such as Minecraft
336
+ (Guss et al.,
337
+ 2019
338
+ ; Fan et al.,
339
+ 2022
340
+ )
341
+ . LLMs are particularly useful in text-based environments
342
+ (Liu et al.,
343
+ 2018
344
+ ; Shridhar et al.,
345
+ 2020
346
+ ; Liu et al.,
347
+ 2023
348
+ )
349
+ , where acting-based prompting techniques such as ReAct
350
+ (Yao et al.,
351
+ 2023b
352
+ )
353
+ have seen success. Similar to CoT, ReAct is limited by its simplicity and cannot effectively adapt to environment conditions. Many extensions have been proposed to address this, including Self-refine
354
+ (Madaan et al.,
355
+ 2023
356
+ )
357
+ and Reflexion
358
+ (Shinn et al.,
359
+ 2023
360
+ ; Yao et al.,
361
+ 2023c
362
+ )
363
+ , which uses self-reflection to enhance reasoning and decision-making, and AdaPlanner
364
+ (Sun et al.,
365
+ 2023
366
+ )
367
+ , which incorporates both positive and negative environmental feedback. However these methods focus on refining an individual plan or trajectory and do not consider alternative choices at each step. In addition, recent work
368
+ (Huang et al.,
369
+ 2023
370
+ )
371
+ has suggested LLMs cannot self-correct their internal reasoning, making it critical to use external feedback. Alternatively to pure decision-making environments, the reasoning and practical abilities of LLMs have been enhanced by access to external tools, such as APIs, search engines, calculators, or other models
372
+ (Schick et al.,
373
+ 2023
374
+ ; Shen et al.,
375
+ 2023
376
+ ; Surís et al.,
377
+ 2023
378
+ )
379
+ . Contrary to reasoning-based approaches, these methods have not been improved with planning, limiting their effectiveness. We summarize them in Tab.
380
+ 1
381
+ .
382
+ Tree-based search.
383
+ Tree-based search, where multiple branches of outcomes are explored during search, is widely used in many planning algorithms
384
+ (Świechowski et al.,
385
+ 2023
386
+ ; LaValle et al.,
387
+ 2001
388
+ )
389
+ and Reinforcement Learning (RL)
390
+ (Hafner et al.,
391
+ 2019
392
+ ; Du et al.,
393
+ 2023
394
+ ; Wu et al.,
395
+ 2023
396
+ )
397
+ algorithms for its good exploration-exploitation trade-off. Though tree-based search requires an environment model that can expand from arbitrary state
398
+ (Vodopivec et al.,
399
+ 2017
400
+ )
401
+ , which often requires extra training in RL
402
+ (Hafner et al.,
403
+ 2023
404
+ )
405
+ , such problem does not exist for LM tasks as we can conveniently backup to any state by setting the input to be the context and corresponding previous output by the LM. Thus, we work on the tree-based framework and use MCTS
406
+ (Świechowski et al.,
407
+ 2023
408
+ )
409
+ to fully release the potential of LMs, while avoiding the cost of training a value function over language descriptions by leveraging the in-context learning
410
+ (Brown et al.,
411
+ 2020
412
+ )
413
+ abilities of LLMs.
414
+ 3
415
+ Preliminaries
416
+ 3.1
417
+ Problem Setting and Prompting
418
+ Before describing LATS, we first define our problem and outline a few established methods that leverage large language models for reasoning or decision-making. In LM reasoning or decision making, we are given an input
419
+ x
420
+ 𝑥
421
+ x
422
+ in natural language and a pretrained language model
423
+ p
424
+ θ
425
+ ​
426
+ (
427
+ x
428
+ )
429
+ subscript
430
+ 𝑝
431
+ 𝜃
432
+ 𝑥
433
+ p_{\theta}(x)
434
+ parameterized by
435
+ θ
436
+ 𝜃
437
+ \theta
438
+ ; our goal is to generate a final output
439
+ y
440
+ ∼
441
+ p
442
+ θ
443
+ ​
444
+ (
445
+ x
446
+ )
447
+ similar-to
448
+ 𝑦
449
+ subscript
450
+ 𝑝
451
+ 𝜃
452
+ 𝑥
453
+ y\sim p_{\theta}(x)
454
+ corresponding to the answer (reasoning) or completes the task (decision-making). Both
455
+ x
456
+ 𝑥
457
+ x
458
+ and
459
+ y
460
+ 𝑦
461
+ y
462
+ are language
463
+ sequences
464
+ , which are comprised of a list of
465
+ tokens
466
+ (the basic elements of natural language, often words), denoted as
467
+ x
468
+ =
469
+ (
470
+ x
471
+ ​
472
+ [
473
+ 1
474
+ ]
475
+ ,
476
+ …
477
+ ,
478
+ x
479
+ ​
480
+ [
481
+ n
482
+ ]
483
+ )
484
+ 𝑥
485
+ 𝑥
486
+ delimited-[]
487
+ 1
488
+ …
489
+ 𝑥
490
+ delimited-[]
491
+ 𝑛
492
+ x=(x[1],\dots,x[n])
493
+ and
494
+ y
495
+ =
496
+ (
497
+ y
498
+ ​
499
+ [
500
+ 1
501
+ ]
502
+ ,
503
+ …
504
+ ,
505
+ y
506
+ ​
507
+ [
508
+ n
509
+ ]
510
+ )
511
+ 𝑦
512
+ 𝑦
513
+ delimited-[]
514
+ 1
515
+ …
516
+ 𝑦
517
+ delimited-[]
518
+ 𝑛
519
+ y=(y[1],\dots,y[n])
520
+ . The LM decodes text autoregressively, i.e., without other inputs, the probability for an LM to generate a sequence
521
+ x
522
+ 𝑥
523
+ x
524
+ is given by
525
+ p
526
+ θ
527
+ ​
528
+ (
529
+ x
530
+ )
531
+ =
532
+ ∏
533
+ i
534
+ =
535
+ 1
536
+ n
537
+ p
538
+ θ
539
+ ​
540
+ (
541
+ x
542
+ ​
543
+ [
544
+ i
545
+ ]
546
+ |
547
+ x
548
+ ​
549
+ [
550
+ 1
551
+ ​
552
+ …
553
+ ​
554
+ i
555
+ −
556
+ 1
557
+ ]
558
+ )
559
+ subscript
560
+ 𝑝
561
+ 𝜃
562
+ 𝑥
563
+ superscript
564
+ subscript
565
+ product
566
+ 𝑖
567
+ 1
568
+ 𝑛
569
+ subscript
570
+ 𝑝
571
+ 𝜃
572
+ conditional
573
+ 𝑥
574
+ delimited-[]
575
+ 𝑖
576
+ 𝑥
577
+ delimited-[]
578
+ 1
579
+ …
580
+ 𝑖
581
+ 1
582
+ p_{\theta}(x)=\prod_{i=1}^{n}p_{\theta}(x[i]|x[1\dots i-1])
583
+ . Usually, to improve the LM,
584
+ prompts
585
+ are provided along with the input
586
+ x
587
+ 𝑥
588
+ x
589
+ , which are specific instructions or few-shot input-output examples. We denote the generic process where an input
590
+ x
591
+ 𝑥
592
+ x
593
+ is transformed into an output
594
+ y
595
+ 𝑦
596
+ y
597
+ by LM:
598
+ y
599
+ ∼
600
+ p
601
+ θ
602
+ ​
603
+ (
604
+ y
605
+ |
606
+ prompt
607
+ I
608
+ ​
609
+ O
610
+ ​
611
+ (
612
+ x
613
+ )
614
+ )
615
+ similar-to
616
+ 𝑦
617
+ subscript
618
+ 𝑝
619
+ 𝜃
620
+ conditional
621
+ 𝑦
622
+ subscript
623
+ prompt
624
+ 𝐼
625
+ 𝑂
626
+ 𝑥
627
+ y\sim p_{\theta}(y|\texttt{prompt}_{IO}(x))
628
+ , where
629
+ prompt
630
+ I
631
+ ​
632
+ O
633
+ ​
634
+ (
635
+ x
636
+ )
637
+ subscript
638
+ prompt
639
+ 𝐼
640
+ 𝑂
641
+ 𝑥
642
+ \texttt{prompt}_{IO}(x)
643
+ denotes the input
644
+ x
645
+ 𝑥
646
+ x
647
+ .
648
+ Chain-of-thought (CoT) Prompting
649
+ (Wei et al.,
650
+ 2022
651
+ )
652
+ was introduced to cater to scenarios where direct mapping from
653
+ x
654
+ 𝑥
655
+ x
656
+ to
657
+ y
658
+ 𝑦
659
+ y
660
+ is intricate, such as when
661
+ x
662
+ 𝑥
663
+ x
664
+ is from a mathematical query or challenging question. This method hinges on creating
665
+ thoughts
666
+ z
667
+ 1
668
+ ,
669
+ …
670
+ ,
671
+ z
672
+ n
673
+ subscript
674
+ 𝑧
675
+ 1
676
+ …
677
+ subscript
678
+ 𝑧
679
+ 𝑛
680
+ z_{1},\dots,z_{n}
681
+ that act as stepping stones between
682
+ x
683
+ 𝑥
684
+ x
685
+ and
686
+ y
687
+ 𝑦
688
+ y
689
+ ; each thought
690
+ z
691
+ i
692
+ subscript
693
+ 𝑧
694
+ 𝑖
695
+ z_{i}
696
+ is a language sequence. To employ CoT prompting, thoughts are extracted sequentially as
697
+ z
698
+ i
699
+ ∼
700
+ p
701
+ θ
702
+ C
703
+ ​
704
+ o
705
+ ​
706
+ T
707
+ ​
708
+ (
709
+ z
710
+ i
711
+ |
712
+ x
713
+ ,
714
+ z
715
+ 1
716
+ ​
717
+ ⋯
718
+ ​
719
+ i
720
+ −
721
+ 1
722
+ )
723
+ similar-to
724
+ subscript
725
+ 𝑧
726
+ 𝑖
727
+ superscript
728
+ subscript
729
+ 𝑝
730
+ 𝜃
731
+ 𝐶
732
+ 𝑜
733
+ 𝑇
734
+ conditional
735
+ subscript
736
+ 𝑧
737
+ 𝑖
738
+ 𝑥
739
+ subscript
740
+ 𝑧
741
+ 1
742
+ ⋯
743
+ 𝑖
744
+ 1
745
+ z_{i}\sim p_{\theta}^{CoT}(z_{i}|x,z_{1\cdots i-1})
746
+ , with the final output being
747
+ y
748
+ ∼
749
+ p
750
+ θ
751
+ C
752
+ ​
753
+ o
754
+ ​
755
+ T
756
+ ​
757
+ (
758
+ y
759
+ |
760
+ x
761
+ ,
762
+ z
763
+ 1
764
+ ​
765
+ ⋯
766
+ ​
767
+ n
768
+ )
769
+ similar-to
770
+ 𝑦
771
+ superscript
772
+ subscript
773
+ 𝑝
774
+ 𝜃
775
+ 𝐶
776
+ 𝑜
777
+ 𝑇
778
+ conditional
779
+ 𝑦
780
+ 𝑥
781
+ subscript
782
+ 𝑧
783
+ 1
784
+ ⋯
785
+ 𝑛
786
+ y\sim p_{\theta}^{CoT}(y|x,z_{1\cdots n})
787
+ .
788
+ Tree-of-thought (ToT) Prompting
789
+ (Yao et al.,
790
+ 2023a
791
+ )
792
+ extends CoT prompting by exploring multiple reasoning paths over thoughts. It frames problems as a search over a tree where each node
793
+ s
794
+ =
795
+ [
796
+ x
797
+ ,
798
+ z
799
+ 1
800
+ ⋅
801
+ i
802
+ ]
803
+ 𝑠
804
+ 𝑥
805
+ subscript
806
+ 𝑧
807
+ ⋅
808
+ 1
809
+ 𝑖
810
+ s=[x,z_{1\cdot i}]
811
+ represents a partial solution state comprising the original input
812
+ x
813
+ 𝑥
814
+ x
815
+ and thought sequence
816
+ z
817
+ 1
818
+ ​
819
+ ⋯
820
+ ​
821
+ i
822
+ subscript
823
+ 𝑧
824
+ 1
825
+ ⋯
826
+ 𝑖
827
+ z_{1\cdots i}
828
+ . Thoughts
829
+ z
830
+ i
831
+ subscript
832
+ 𝑧
833
+ 𝑖
834
+ z_{i}
835
+ are generated by proposal or sampling with CoT
836
+ z
837
+ i
838
+ ∼
839
+ p
840
+ θ
841
+ C
842
+ ​
843
+ o
844
+ ​
845
+ T
846
+ ​
847
+ (
848
+ z
849
+ i
850
+ |
851
+ x
852
+ ,
853
+ z
854
+ 1
855
+ ​
856
+ ⋯
857
+ ​
858
+ i
859
+ −
860
+ 1
861
+ )
862
+ similar-to
863
+ subscript
864
+ 𝑧
865
+ 𝑖
866
+ superscript
867
+ subscript
868
+ 𝑝
869
+ 𝜃
870
+ 𝐶
871
+ 𝑜
872
+ 𝑇
873
+ conditional
874
+ subscript
875
+ 𝑧
876
+ 𝑖
877
+ 𝑥
878
+ subscript
879
+ 𝑧
880
+ 1
881
+ ⋯
882
+ 𝑖
883
+ 1
884
+ z_{i}\sim p_{\theta}^{CoT}(z_{i}|x,z_{1\cdots i-1})
885
+ . Deliberate search algorithms like breadth-first or depth-first search are used to systematically explore the tree, guided by heuristics based on language model evaluations
886
+ V
887
+ ​
888
+ (
889
+ s
890
+ )
891
+ 𝑉
892
+ 𝑠
893
+ V(s)
894
+ of each state.
895
+ Reasoning via Planning
896
+ (RAP)
897
+ (Hao et al.,
898
+ 2023
899
+ )
900
+ is similar to ToT, except that MCTS is used over DFS or BFS. Heuristics are designed from an LM, such as the likelihood or confidence of an action, and the LM is used as a world model to predict subsequent states during the simulation step.
901
+ ReAct
902
+ (Yao et al.,
903
+ 2023b
904
+ )
905
+ extends language models to tasks where the mapping from
906
+ x
907
+ 𝑥
908
+ x
909
+ to
910
+ y
911
+ 𝑦
912
+ y
913
+ is enhanced by or requires interactions with an external environment, such as a game or API. This technique constructs an action space
914
+ A
915
+ ^
916
+ =
917
+ A
918
+ ∪
919
+ Z
920
+ ^
921
+ 𝐴
922
+ 𝐴
923
+ 𝑍
924
+ \hat{A}=A\cup Z
925
+ that adds permissible actions
926
+ a
927
+ 𝑎
928
+ a
929
+ to the reasoning traces
930
+ z
931
+ 𝑧
932
+ z
933
+ from CoT. Observations
934
+ o
935
+ 𝑜
936
+ o
937
+ from the environment are used to improve both reasoning and acting. To solve problems with ReAct, after each observation, actions are generated from
938
+ p
939
+ θ
940
+ subscript
941
+ 𝑝
942
+ 𝜃
943
+ p_{\theta}
944
+ sequentially as
945
+ a
946
+ i
947
+ ∼
948
+ p
949
+ θ
950
+ R
951
+ ​
952
+ e
953
+ ​
954
+ A
955
+ ​
956
+ c
957
+ ​
958
+ t
959
+ ​
960
+ (
961
+ a
962
+ i
963
+ |
964
+ x
965
+ ,
966
+ o
967
+ 1
968
+ ​
969
+ ⋯
970
+ ​
971
+ i
972
+ −
973
+ 1
974
+ ,
975
+ a
976
+ 1
977
+ ​
978
+ ⋯
979
+ ​
980
+ i
981
+ −
982
+ 1
983
+ )
984
+ similar-to
985
+ subscript
986
+ 𝑎
987
+ 𝑖
988
+ superscript
989
+ subscript
990
+ 𝑝
991
+ 𝜃
992
+ 𝑅
993
+ 𝑒
994
+ 𝐴
995
+ 𝑐
996
+ 𝑡
997
+ conditional
998
+ subscript
999
+ 𝑎
1000
+ 𝑖
1001
+ 𝑥
1002
+ subscript
1003
+ 𝑜
1004
+ 1
1005
+ ⋯
1006
+ 𝑖
1007
+ 1
1008
+ subscript
1009
+ 𝑎
1010
+ 1
1011
+ ⋯
1012
+ 𝑖
1013
+ 1
1014
+ a_{i}\sim p_{\theta}^{ReAct}(a_{i}|x,o_{1\cdots i-1},a_{1\cdots i-1})
1015
+ , with the final output being
1016
+ y
1017
+ ∼
1018
+ p
1019
+ θ
1020
+ R
1021
+ ​
1022
+ e
1023
+ ​
1024
+ A
1025
+ ​
1026
+ c
1027
+ ​
1028
+ t
1029
+ ​
1030
+ (
1031
+ y
1032
+ |
1033
+ x
1034
+ ,
1035
+ o
1036
+ 1
1037
+ ​
1038
+ ⋯
1039
+ ​
1040
+ n
1041
+ ,
1042
+ a
1043
+ 1
1044
+ ​
1045
+ ⋯
1046
+ ​
1047
+ n
1048
+ )
1049
+ similar-to
1050
+ 𝑦
1051
+ superscript
1052
+ subscript
1053
+ 𝑝
1054
+ 𝜃
1055
+ 𝑅
1056
+ 𝑒
1057
+ 𝐴
1058
+ 𝑐
1059
+ 𝑡
1060
+ conditional
1061
+ 𝑦
1062
+ 𝑥
1063
+ subscript
1064
+ 𝑜
1065
+ 1
1066
+ ⋯
1067
+ 𝑛
1068
+ subscript
1069
+ 𝑎
1070
+ 1
1071
+ ⋯
1072
+ 𝑛
1073
+ y\sim p_{\theta}^{ReAct}(y~{}|~{}x,o_{1\cdots n},a_{1\cdots n})
1074
+ .
1075
+ While the previously described prompting techniques improve LM performance on reasoning tasks, they falter on difficult tasks that involve multifaceted decision-making due to several shortcomings: 1)
1076
+ Flexibility
1077
+ : Base prompting methods (CoT or ReAct) autoregressively sample from the LM, neglecting potential alternative continuations from specific states. 2)
1078
+ Sensibility
1079
+ : Reasoning-based methods (CoT, RAP, or ToT) rely solely on the internal representations of the LM and cannot consider external observations. This dependency risks fact hallucination and error propagation while setting a performance ceiling. 3)
1080
+ Adaptability
1081
+ : Current planning frameworks (RAP or ToT) use simple search algorithms such as BFS or cannot leverage environmental feedback to improve planning. Additionally, the agent is static and cannot reuse previous experience or learn from trial and error. While RAP also adopts MCTS, it is constrained to tasks where the LM can become a world model and accurately predict states. These shortcomings limit the ability of LMs to be deployed as general problem-solving agents and form the motivation for LATS.
1082
+ 3.2
1083
+ Monte-Carlo Tree Search (MCTS)
1084
+ Monte-Carlo Tree Search (MCTS) is a heuristic search algorithm that is proved successful on many decision-making environments such as Atari
1085
+ (Ye et al.,
1086
+ 2021
1087
+ )
1088
+ and Go
1089
+ (Silver et al.,
1090
+ 2016
1091
+ )
1092
+ . MCTS builds a decision tree where every node in the tree is a state and edge is an action. MCTS runs for
1093
+ k
1094
+ 𝑘
1095
+ k
1096
+ episodes; for each episode, it starts from the root (i.e., initial state) and iteratively conducts two steps to expand the tree: 1)
1097
+ Expansion
1098
+ , where multiple children states
1099
+ s
1100
+ 𝑠
1101
+ s
1102
+ are explored from the current parent state
1103
+ p
1104
+ 𝑝
1105
+ p
1106
+ by sampling
1107
+ n
1108
+ 𝑛
1109
+ n
1110
+ actions, and 2)
1111
+ Selection
1112
+ , where the children with the highest UCT
1113
+ (Upper Confidence bounds applied to Trees)
1114
+ (Kocsis & Szepesvári,
1115
+ 2006
1116
+ )
1117
+ value is selected by the next iteration. The UCT of a child state
1118
+ s
1119
+ 𝑠
1120
+ s
1121
+ is calculated as follows:
1122
+ U
1123
+ ​
1124
+ C
1125
+ ​
1126
+ T
1127
+ ​
1128
+ (
1129
+ s
1130
+ )
1131
+ =
1132
+ V
1133
+ ​
1134
+ (
1135
+ s
1136
+ )
1137
+ +
1138
+ w
1139
+ ​
1140
+ ln
1141
+ ⁡
1142
+ N
1143
+ ​
1144
+ (
1145
+ p
1146
+ )
1147
+ N
1148
+ ​
1149
+ (
1150
+ s
1151
+ )
1152
+ ,
1153
+ 𝑈
1154
+ 𝐶
1155
+ 𝑇
1156
+ 𝑠
1157
+ 𝑉
1158
+ 𝑠
1159
+ 𝑤
1160
+ 𝑁
1161
+ 𝑝
1162
+ 𝑁
1163
+ 𝑠
1164
+ UCT(s)=V(s)+w\sqrt{\frac{\ln N(p)}{N(s)}},
1165
+ (1)
1166
+ where
1167
+ N
1168
+ ​
1169
+ (
1170
+ s
1171
+ )
1172
+ 𝑁
1173
+ 𝑠
1174
+ N(s)
1175
+ is the number of visits to a node
1176
+ s
1177
+ 𝑠
1178
+ s
1179
+ ,
1180
+ V
1181
+ ​
1182
+ (
1183
+ s
1184
+ )
1185
+ 𝑉
1186
+ 𝑠
1187
+ V(s)
1188
+ is the value function (expected return) from the subtree of
1189
+ s
1190
+ 𝑠
1191
+ s
1192
+ ,
1193
+ w
1194
+ 𝑤
1195
+ w
1196
+ is the exploration weight, and
1197
+ p
1198
+ 𝑝
1199
+ p
1200
+ is the parent node of
1201
+ s
1202
+ 𝑠
1203
+ s
1204
+ . The child node with the highest UCT value is selected for expansion in the next iteration. When the end of an episode is reached, a
1205
+ backpropagation
1206
+ is carried out: the return
1207
+ r
1208
+ 𝑟
1209
+ r
1210
+ is used for updating every
1211
+ V
1212
+ ​
1213
+ (
1214
+ s
1215
+ )
1216
+ 𝑉
1217
+ 𝑠
1218
+ V(s)
1219
+ along the path
1220
+ with the formula
1221
+ V
1222
+ ​
1223
+ (
1224
+ s
1225
+ )
1226
+ =
1227
+ V
1228
+ old
1229
+ ​
1230
+ (
1231
+ s
1232
+ )
1233
+ ​
1234
+ (
1235
+ N
1236
+ ​
1237
+ (
1238
+ s
1239
+ )
1240
+ −
1241
+ 1
1242
+ )
1243
+ +
1244
+ r
1245
+ N
1246
+ ​
1247
+ (
1248
+ s
1249
+ )
1250
+ 𝑉
1251
+ 𝑠
1252
+ subscript
1253
+ 𝑉
1254
+ old
1255
+ 𝑠
1256
+ 𝑁
1257
+ 𝑠
1258
+ 1
1259
+ 𝑟
1260
+ 𝑁
1261
+ 𝑠
1262
+ V(s)=\frac{V_{\text{old}}(s)(N(s)-1)+r}{N(s)}
1263
+ , where
1264
+ V
1265
+ old
1266
+ ​
1267
+ (
1268
+ s
1269
+ )
1270
+ subscript
1271
+ 𝑉
1272
+ old
1273
+ 𝑠
1274
+ V_{\text{old}}(s)
1275
+ is the old value function. Normally, the major shortcoming of MCTS is that it requires an environment model to undo previous steps and form a searching tree, which is often a strong assumption. However, such a limitation does not exist for LMs, as we can conveniently reset to any step by simply copy-pasting historical text input. Such a special property is the key motivation of our work.
1276
+ 4
1277
+ Unifying Planning, Reasoning, and Acting
1278
+ 4.1
1279
+ LM Agent
1280
+ LATS supports sequential reasoning or decision-making tasks on the basis of ReAct. At time step
1281
+ t
1282
+ 𝑡
1283
+ t
1284
+ , an agent receives an observation
1285
+ o
1286
+ t
1287
+ ∈
1288
+ O
1289
+ subscript
1290
+ 𝑜
1291
+ 𝑡
1292
+ 𝑂
1293
+ o_{t}\in O
1294
+ from the environment and takes an action
1295
+ a
1296
+ t
1297
+ ∈
1298
+ A
1299
+ subscript
1300
+ 𝑎
1301
+ 𝑡
1302
+ 𝐴
1303
+ a_{t}\in A
1304
+ following some policy
1305
+ π
1306
+ ​
1307
+ (
1308
+ a
1309
+ t
1310
+ |
1311
+ x
1312
+ ,
1313
+ o
1314
+ 1
1315
+ ​
1316
+ ⋯
1317
+ ​
1318
+ i
1319
+ −
1320
+ 1
1321
+ ,
1322
+ a
1323
+ 1
1324
+ ​
1325
+ ⋯
1326
+ ​
1327
+ i
1328
+ −
1329
+ 1
1330
+ )
1331
+ 𝜋
1332
+ conditional
1333
+ subscript
1334
+ 𝑎
1335
+ 𝑡
1336
+ 𝑥
1337
+ subscript
1338
+ 𝑜
1339
+ 1
1340
+ ⋯
1341
+ 𝑖
1342
+ 1
1343
+ subscript
1344
+ 𝑎
1345
+ 1
1346
+ ⋯
1347
+ 𝑖
1348
+ 1
1349
+ \pi(a_{t}|x,o_{1\cdots i-1},a_{1\cdots i-1})
1350
+ , where
1351
+ x
1352
+ 𝑥
1353
+ x
1354
+ consists of the task instruction and a number of few-shot examples. We initialize the agent with
1355
+ p
1356
+ θ
1357
+ subscript
1358
+ 𝑝
1359
+ 𝜃
1360
+ p_{\theta}
1361
+ to leverage the useful language representations of an LM as a base decision-maker. We follow the ReAct instantiation in which the action space
1362
+ A
1363
+ ^
1364
+ =
1365
+ A
1366
+ ∪
1367
+ Z
1368
+ ^
1369
+ 𝐴
1370
+ 𝐴
1371
+ 𝑍
1372
+ \hat{A}=A\cup Z
1373
+ consists of both the space of permissible actions
1374
+ A
1375
+ 𝐴
1376
+ A
1377
+ and language space of reasoning traces
1378
+ Z
1379
+ 𝑍
1380
+ Z
1381
+ . Actions directly affect the environment and result in observation, while thoughts are used to formalize decisions by organizing information, planning future actions, or injecting internal knowledge. The exact instantiation of the action space depends on the particular environment; for decision-making tasks actions might consist of commands on a website while for reasoning tasks the action space might be limited to a few external tools or APIs.
1382
+ Instead of greedily decoding one trajectory or solution, we sample
1383
+ n
1384
+ 𝑛
1385
+ n
1386
+ actions from
1387
+ p
1388
+ θ
1389
+ subscript
1390
+ 𝑝
1391
+ 𝜃
1392
+ p_{\theta}
1393
+ using the current state. This is based on the intuition that for complex decision-making tasks, there is likely to be a range of potential trajectories or reasoning paths that are correct
1394
+ (Evans,
1395
+ 2010
1396
+ )
1397
+ . Sampling a diverse set of candidates at each step mitigates the stochastic nature of LM text generation and enables greater exploration in both the decision-making and reasoning space. We wrap
1398
+ p
1399
+ θ
1400
+ subscript
1401
+ 𝑝
1402
+ 𝜃
1403
+ p_{\theta}
1404
+ within our proposed search algorithm to deliberately construct the best trajectory from sampled actions.
1405
+ 4.2
1406
+ LATS
1407
+ Figure 3:
1408
+ An overview of the six operations of LATS. A node is
1409
+ selected
1410
+ ,
1411
+ expanded
1412
+ ,
1413
+ evaluated
1414
+ , then
1415
+ simulated
1416
+ until a terminal node is reached, then the resulting value is
1417
+ backpropagated
1418
+ . If the trajectory fails, a
1419
+ reflection
1420
+ is generated and used as additional context for future trials. These operations are performed in succession until the budget is reached or task is successful.
1421
+ The main component of LATS is a search algorithm that controls the overall problem-solving process with deliberate planning. To find the most promising trajectory and systemically balance exploration with exploitation, we adopt a variant of Monte Carlo Tree Search (MCTS) that frames decision-making as a tree search, in which each node
1422
+ s
1423
+ =
1424
+ [
1425
+ x
1426
+ ,
1427
+ a
1428
+ 1
1429
+ ​
1430
+ ⋯
1431
+ ​
1432
+ i
1433
+ ,
1434
+ o
1435
+ 1
1436
+ ​
1437
+ ⋯
1438
+ ​
1439
+ i
1440
+ ]
1441
+ 𝑠
1442
+ 𝑥
1443
+ subscript
1444
+ 𝑎
1445
+ 1
1446
+ ⋯
1447
+ 𝑖
1448
+ subscript
1449
+ 𝑜
1450
+ 1
1451
+ ⋯
1452
+ 𝑖
1453
+ s=[x,a_{1\cdots i},o_{1\cdots i}]
1454
+ represents a state comprising the original input
1455
+ x
1456
+ 𝑥
1457
+ x
1458
+ , action sequence
1459
+ a
1460
+ 1
1461
+ ⋅
1462
+ i
1463
+ subscript
1464
+ 𝑎
1465
+ ⋅
1466
+ 1
1467
+ 𝑖
1468
+ a_{1\cdot i}
1469
+ , and observation sequence
1470
+ o
1471
+ 1
1472
+ ⋅
1473
+ i
1474
+ subscript
1475
+ 𝑜
1476
+ ⋅
1477
+ 1
1478
+ 𝑖
1479
+ o_{1\cdot i}
1480
+ .
1481
+ To adapt MCTS for language agents, LATS repurposes
1482
+ p
1483
+ θ
1484
+ subscript
1485
+ 𝑝
1486
+ 𝜃
1487
+ p_{\theta}
1488
+ as an agent, state evaluator, and feedback generator, leveraging the useful language priors of modern LMs to facilitate planning. While standard MCTS and RAP
1489
+ Hao et al. (
1490
+ 2023
1491
+ )
1492
+ rely on internal dynamics models to facilitate simulation, LATS is model-free and uses environment interaction. LATS consists of a series of operations,
1493
+ selection, expansion, evaluation, simulation, backpropagation, and reflection
1494
+ , performed in succession until the task is successfully completed or a computational limit is reached. The full psuedocode of LATS can be found in Sec.
1495
+ A
1496
+ in the Appendix.
1497
+ Selection.
1498
+ In the first operation, the algorithm identifies a segment of the current tree most suitable for subsequent expansion. Starting from the root node, denoted as the initial state
1499
+ s
1500
+ 0
1501
+ subscript
1502
+ 𝑠
1503
+ 0
1504
+ s_{0}
1505
+ , a child node is selected at each tree level until a leaf node is reached. To balance exploration and exploitation, we use the UCT algorithm as shown in Eq.
1506
+ 1
1507
+ .
1508
+ Expansion.
1509
+ After selecting a node, the second operation expands the tree by sampling
1510
+ n
1511
+ 𝑛
1512
+ n
1513
+ actions from
1514
+ p
1515
+ θ
1516
+ subscript
1517
+ 𝑝
1518
+ 𝜃
1519
+ p_{\theta}
1520
+ , as described in the prior section. The environment receives each action and returns corresponding feedback as an observation. This results in
1521
+ n
1522
+ 𝑛
1523
+ n
1524
+ new child nodes added to the tree. This tree is stored in an external long-term memory structure.
1525
+ Evaluation.
1526
+ The third operation assigns a scalar value to each new child node to be used for selection and backpropagation. This value effectively quantifies the agent’s progress in task completion, serving as a heuristic to steer the search algorithm towards the most promising regions of the tree. Following
1527
+ Yao et al. (
1528
+ 2023a
1529
+ )
1530
+ we repurpose
1531
+ p
1532
+ θ
1533
+ subscript
1534
+ 𝑝
1535
+ 𝜃
1536
+ p_{\theta}
1537
+ into a value function by prompting it to reason about a given state. To obtain a scalar value, we instruct
1538
+ p
1539
+ θ
1540
+ subscript
1541
+ 𝑝
1542
+ 𝜃
1543
+ p_{\theta}
1544
+ to end its reasoning trace with a score indicating the correctness of the trajectory. This method offers enhanced flexibility over programmed heuristics
1545
+ (Campbell et al.,
1546
+ 2002
1547
+ )
1548
+ and greater efficiency than learned heuristics
1549
+ (Silver et al.,
1550
+ 2017
1551
+ )
1552
+ .
1553
+ Simulation.
1554
+ The fourth operation expands the currently selected node until a terminal state is reached. At each depth level we sample and evaluate nodes with the same operations, but prioritize nodes of highest value. Reaching a terminal state provides objective feedback on the correctness of a trajectory. If the task is completed successfully, then LATS terminates the search. If the solution is partially successful or unsuccessful, then we perform two additional operations as described below.
1555
+ Backpropagation.
1556
+ This operation updates the values of the tree based on the outcome of a trajectory. For each node
1557
+ s
1558
+ 0
1559
+ ,
1560
+ s
1561
+ 1
1562
+ ,
1563
+ …
1564
+ ,
1565
+ s
1566
+ n
1567
+ subscript
1568
+ 𝑠
1569
+ 0
1570
+ subscript
1571
+ 𝑠
1572
+ 1
1573
+ …
1574
+ subscript
1575
+ 𝑠
1576
+ 𝑛
1577
+ s_{0},s_{1},\dots,s_{n}
1578
+ in the trajectory from root (initial state
1579
+ s
1580
+ 0
1581
+ subscript
1582
+ 𝑠
1583
+ 0
1584
+ s_{0}
1585
+ ) of the searching tree to leaf (terminal state
1586
+ s
1587
+ n
1588
+ subscript
1589
+ 𝑠
1590
+ 𝑛
1591
+ s_{n}
1592
+ ), its value is updated to reflect the outcome of the simulation by
1593
+ N
1594
+ ​
1595
+ (
1596
+ s
1597
+ i
1598
+ )
1599
+ =
1600
+ N
1601
+ old
1602
+ ​
1603
+ (
1604
+ s
1605
+ i
1606
+ )
1607
+ +
1608
+ 1
1609
+ 𝑁
1610
+ subscript
1611
+ 𝑠
1612
+ 𝑖
1613
+ subscript
1614
+ 𝑁
1615
+ old
1616
+ subscript
1617
+ 𝑠
1618
+ 𝑖
1619
+ 1
1620
+ N(s_{i})=N_{\text{old}}(s_{i})+1
1621
+ and
1622
+ V
1623
+ ​
1624
+ (
1625
+ s
1626
+ i
1627
+ )
1628
+ =
1629
+ r
1630
+ +
1631
+ N
1632
+ old
1633
+ ​
1634
+ (
1635
+ s
1636
+ i
1637
+ )
1638
+ ​
1639
+ V
1640
+ old
1641
+ ​
1642
+ (
1643
+ s
1644
+ i
1645
+ )
1646
+ N
1647
+ ​
1648
+ (
1649
+ s
1650
+ i
1651
+ )
1652
+ 𝑉
1653
+ subscript
1654
+ 𝑠
1655
+ 𝑖
1656
+ 𝑟
1657
+ subscript
1658
+ 𝑁
1659
+ old
1660
+ subscript
1661
+ 𝑠
1662
+ 𝑖
1663
+ subscript
1664
+ 𝑉
1665
+ old
1666
+ subscript
1667
+ 𝑠
1668
+ 𝑖
1669
+ 𝑁
1670
+ subscript
1671
+ 𝑠
1672
+ 𝑖
1673
+ V(s_{i})=\frac{r+N_{\text{old}}(s_{i})V_{\text{old}}(s_{i})}{N(s_{i})}
1674
+ , where
1675
+ r
1676
+ 𝑟
1677
+ r
1678
+ is the return and
1679
+ N
1680
+ old
1681
+ ,
1682
+ V
1683
+ old
1684
+ subscript
1685
+ 𝑁
1686
+ old
1687
+ subscript
1688
+ 𝑉
1689
+ old
1690
+ N_{\text{old}},V_{\text{old}}
1691
+ are the old number of visits and value function. These updated values are used in the UCT formula (Eq.
1692
+ 1
1693
+ ) to guide the selection of the next node for exploration.
1694
+ Reflection.
1695
+ In addition to the environmental feedback, we also leverage
1696
+ self-reflection
1697
+ to further refine the decision-making process
1698
+ (Shinn et al.,
1699
+ 2023
1700
+ ; Madaan et al.,
1701
+ 2023
1702
+ )
1703
+ . Upon encountering an unsuccessful terminal node,
1704
+ p
1705
+ θ
1706
+ subscript
1707
+ 𝑝
1708
+ 𝜃
1709
+ p_{\theta}
1710
+ is prompted with the trajectory and final reward to provide a verbal self-reflection that summarizes the errors in the reasoning or acting process and proposes superior alternatives. We store both failed trajectories and corresponding reflections in the memory. In subsequent iterations, these are integrated as additional context to the agent and value function, refining both through in-context learning. This imparts a semantic gradient signal more useful than a scalar value, enabling the agent to learn from trial and error without the cost of expensive optimization processes such as reinforcement learning.
1711
+ Conceptually, LATS has the following advantages as a general framework for reasoning and decision-making with LM agents.
1712
+ (1)
1713
+ Generality
1714
+ : LATS supports both reasoning and decision-making tasks by defining a shared space of thoughts and actions. (2)
1715
+ Deliberate
1716
+ : The use of MCTS and LM value function ensures a principled search that selects options with high value while exploring promising alternatives. (3)
1717
+ Adaptability
1718
+ : LATS is designed around the use of external feedback through observations and self-reflection, enabling greater adaptation during problem-solving. (4)
1719
+ Flexibility
1720
+ : LATS can accommodate different scenarios, environments, and resource stipulations by modifying state design and tree dimensions. (5)
1721
+ Modularity
1722
+ : The base LM agent, reflection generator, and value function can be independently altered and adapted to individual LM properties.
1723
+ 5
1724
+ Experiments
1725
+ To demonstrate the general applicability of LATS, we evaluate our method on a variety of decision-making domains that requires both reasoning and acting ability: programming
1726
+ (Chen et al.,
1727
+ 2021
1728
+ ; Austin et al.,
1729
+ 2021
1730
+ )
1731
+ , HotPotQA
1732
+ (Yang et al.,
1733
+ 2018
1734
+ )
1735
+ , and WebShop
1736
+ (Yao et al.,
1737
+ 2022
1738
+ )
1739
+ .
1740
+ 5.1
1741
+ HotPotQA
1742
+ For a task that can be approached with both reasoning-based and acting-based strategies, we consider HotPotQA
1743
+ (Yang et al.,
1744
+ 2018
1745
+ )
1746
+ , a multi-hop question-answering benchmark that requires retrieval over two or more Wikipedia passages. For the action space, in addition to LM thoughts we follow the setup from
1747
+ Yao et al. (
1748
+ 2023b
1749
+ )
1750
+ , which provides the agent with API calls to search and lookup information. The output of these API calls and self-generated reflections form the observation space. We use a subset of 100 questions and three few-shot examples for each method. For ToT, we use DFS as the base search algorithm and scoring with the LM as the heuristic. For all methods that involve sampling, including LATS, we sample
1751
+ k
1752
+ =
1753
+ 50
1754
+ 𝑘
1755
+ 50
1756
+ k=50
1757
+ trajectories. More details and prompts can be found in Sec.
1758
+ D
1759
+ and Sec.
1760
+ E
1761
+ in the Appendix.
1762
+ We evaluate internal reasoning strategies by removing actions and observations from the context, corresponding to CoT
1763
+ (Wei et al.,
1764
+ 2022
1765
+ )
1766
+ and its variants, CoT-SC
1767
+ (Wang et al.,
1768
+ 2022
1769
+ )
1770
+ , ToT
1771
+ (Yao et al.,
1772
+ 2023a
1773
+ )
1774
+ , and RAP
1775
+ (Hao et al.,
1776
+ 2023
1777
+ )
1778
+ . These methods rely solely on the agent’s existing knowledge to answer the question. We also consider acting-based methods ReAct, Reflexion, and LATS, which augment the agent with the interactive API environment and primarily evaluate its information retrieval abilities. While LATS is designed for scenarios where external feedback can enhance reasoning, we also implement a reasoning-only version with CoT as the base prompt. We also combine internal and external reasoning in LATS by first prompting with a CoT-based prompt, then switching to a ReAct-based prompt upon failure. This is closer to how humans might approach this task, by using tools to lookup additional information only when the answer is not already known.
1779
+ Prompt Method
1780
+ HotpotQA (EM)
1781
+ I/O
1782
+ 0.32
1783
+ CoT
1784
+ (Wei et al.,
1785
+ 2022
1786
+ )
1787
+ 0.34
1788
+ CoT - SC
1789
+ (Wang et al.,
1790
+ 2022
1791
+ )
1792
+ 0.38
1793
+ ToT
1794
+ (Yao et al.,
1795
+ 2023a
1796
+ )
1797
+ 0.55
1798
+ RAP
1799
+ (Hao et al.,
1800
+ 2023
1801
+ )
1802
+ 0.60
1803
+ RAP (n = 10)
1804
+ 0.60
1805
+ LATS (CoT)
1806
+ 0.60
1807
+ Prompt Method
1808
+ HotpotQA (EM)
1809
+ ReAct
1810
+ (Yao et al.,
1811
+ 2023b
1812
+ )
1813
+ 0.32
1814
+ ReAct (best of k)
1815
+ 0.38
1816
+ Reflexion
1817
+ (Shinn et al.,
1818
+ 2023
1819
+ )
1820
+ 0.51
1821
+ LATS
1822
+ 0.61
1823
+ LATS (n = 3)
1824
+ 0.56
1825
+ LATS (n = 10)
1826
+ 0.64
1827
+ LATS (CoT + ReAct)
1828
+ 0.71
1829
+ Table 2:
1830
+ GPT-3.5 reasoning-based prompting (left) and acting-based prompting (right) results on HotpotQA. LATS achieves the highest exact match (EM) for acting and is competitive on reasoning. Unless otherwise specified, we sample
1831
+ n
1832
+ =
1833
+ 5
1834
+ 𝑛
1835
+ 5
1836
+ n=5
1837
+ nodes during expansion and
1838
+ k
1839
+ =
1840
+ 50
1841
+ 𝑘
1842
+ 50
1843
+ k=50
1844
+ trajectories.
1845
+ Results.
1846
+ We observe in Tab.
1847
+ 2
1848
+ that both internal reasoning and external retrieval strategies perform well on HotPotQA. Due to their large-scale training corpus, modern LLMs already encode factual knowledge and can often directly answer the question correctly. While CoT can slightly enhance performance on questions requiring reasoning, larger gains are observed with search methods ToT and RAP, which can sample and explore more outputs. We observe similar results for acting-based methods. LATS surpasses ReAct, even when sampling the same number of trajectories, by expanding more nodes with principled search (see Fig.
1849
+ 5
1850
+ in Appendix
1851
+ D
1852
+ for a qualitative sample). This is demonstrated when modifying
1853
+ n
1854
+ 𝑛
1855
+ n
1856
+ , the number of nodes expanded during each iteration. Increasing
1857
+ n
1858
+ 𝑛
1859
+ n
1860
+ can consistently improve performance, although at greater computational and inference costs. LATS is also competitive to RAP on internal reasoning but performs worse than acting. Combining internal and external reasoning in LATS results in the highest performance, indicating the importance of external feedback in augmenting reasoning even in tasks the base LM can already perform.
1861
+ 5.2
1862
+ Programming
1863
+ Prompt Method
1864
+ Model
1865
+ Pass@1
1866
+ CoT
1867
+ (Wei et al.,
1868
+ 2022
1869
+ )
1870
+ GPT-3.5
1871
+ 46.9
1872
+ ReAct
1873
+ (Yao et al.,
1874
+ 2023b
1875
+ )
1876
+ GPT-3.5
1877
+ 56.9
1878
+ Reflexion
1879
+ (Shinn et al.,
1880
+ 2023
1881
+ )
1882
+ GPT-3.5
1883
+ 68.1
1884
+ ToT
1885
+ (Yao et al.,
1886
+ 2023a
1887
+ )
1888
+ GPT-3.5
1889
+ 54.4
1890
+ RAP
1891
+ (Hao et al.,
1892
+ 2023
1893
+ )
1894
+ GPT-3.5
1895
+ 63.1
1896
+ LATS (Ours)
1897
+ GPT-3.5
1898
+ 83.8
1899
+ I/O
1900
+ GPT-4
1901
+ 80.1
1902
+ Reflexion
1903
+ GPT-4
1904
+ 91.0
1905
+ LATS
1906
+ GPT-4
1907
+ 94.4
1908
+ Prompt Method
1909
+ Pass@1
1910
+ CoT
1911
+ (Wei et al.,
1912
+ 2022
1913
+ )
1914
+ 54.9
1915
+ ReAct
1916
+ (Wei et al.,
1917
+ 2022
1918
+ )
1919
+ 67.0
1920
+ Reflexion
1921
+ (Shinn et al.,
1922
+ 2023
1923
+ )
1924
+ 70.0
1925
+ ToT
1926
+ (Yao et al.,
1927
+ 2023a
1928
+ )
1929
+ 65.8
1930
+ RAP
1931
+ (Hao et al.,
1932
+ 2023
1933
+ )
1934
+ 71.4
1935
+ LATS (Ours)
1936
+ 81.1
1937
+ Table 3:
1938
+ GPT-3.5 and GPT-4 Pass@1 accuracy on HumanEval
1939
+ (Chen et al.,
1940
+ 2021
1941
+ )
1942
+ and MBPP
1943
+ (Austin et al.,
1944
+ 2021
1945
+ )
1946
+ . Prompting with LATS achieves the highest performance. We sample 5 solutions during expansion for
1947
+ 8
1948
+ iterations.
1949
+ To demonstrate the importance of external observations for complex reasoning tasks, we evaluate the baselines and LATS on programming with Humaneval
1950
+ (Chen et al.,
1951
+ 2021
1952
+ )
1953
+ and MBPP
1954
+ (Austin et al.,
1955
+ 2021
1956
+ )
1957
+ . Both datasets measure the correctness of synthesized programs in Python from natural language docstrings. We use individual solutions as the action space and test suite and compiler feedback as the external observation. We follow
1958
+ Chen et al. (
1959
+ 2022a
1960
+ )
1961
+ and use an LLM to generate a synthetic test suite of syntactically valid “assert” statements for each question. For each step, the solution is evaluated on this test suite, and the results including successful and failed tests and compiler output, are added to the context as an observation. We use the same test suite for Reflexion.
1962
+ For this task, the reasoning and acting baselines share an action space, but acting methods are able to incorporate observations as additional context. For LATS, since each action corresponds to a complete solution, we skip the simulation step of LATS and directly use the percentage of passed tests as the backpropagated reward. We use
1963
+ k
1964
+ =
1965
+ 8
1966
+ 𝑘
1967
+ 8
1968
+ k=8
1969
+ iterations, set the number of generated tests at
1970
+ 4
1971
+ 4
1972
+ 4
1973
+ , and sample
1974
+ n
1975
+ =
1976
+ 5
1977
+ 𝑛
1978
+ 5
1979
+ n=5
1980
+ solutions during expansion. After the search is completed, we select the solution with the highest value and evaluate it on the real test suite for the pass@1 accuracy evaluation. More details and prompts can be found in Sec.
1981
+ D
1982
+ and Sec.
1983
+ F
1984
+ in the Appendix.
1985
+ Results.
1986
+ We find in Tab
1987
+ 3
1988
+ that both search and semantic feedback are crucial for better performance. Despite not using observations, ToT and RAP are competitive with Reflexion. LATS has the highest performance on both datasets. Since RAP uses a similar search algorithm as LATS, this reveals the importance of external feedback for difficult reasoning tasks such as programming. With GPT-4, using LATS sets the state of the art for HumanEval, showing LATS can be used with more advanced LLMs for higher performance.
1989
+ 5.3
1990
+ Webshop
1991
+ For a complex decision-making environment with practical applications, we consider WebShop
1992
+ (Yao et al.,
1993
+ 2022
1994
+ )
1995
+ , an online shopping environment composed of a website with 1.18M real-world products and 12k human instructions. Agents must navigate a website through a variety of commands to purchase an item matching a user specification. We use the preconstructed action space of search and click commands and browser feedback and reflections for the observation. The performance is gauged using two metrics: an average score, reflecting the percentage of user-specified attributes met by the selected product, and a success rate, indicating the frequency with which the chosen product fulfills all given conditions. We compare against acting-based prompting methods and RL-based approaches. We evaluate on 50 instructions, expand
1996
+ n
1997
+ =
1998
+ 5
1999
+ 𝑛
2000
+ 5
2001
+ n=5
2002
+ children for LATS, and set
2003
+ k
2004
+ =
2005
+ 30
2006
+ 𝑘
2007
+ 30
2008
+ k=30
2009
+ for LATS, ReAct best of
2010
+ k
2011
+ 𝑘
2012
+ k
2013
+ , and Reflexion. More details and prompts are in Appendix
2014
+ D
2015
+ and
2016
+ G
2017
+ .
2018
+ Results.
2019
+ We find in Tab.
2020
+ 5
2021
+ that GPT-3.5 with ReAct is competitive to imitation learning, and can exceed reinforcement learning techniques with stronger prompting strategies. Sampling
2022
+ k
2023
+ =
2024
+ 30
2025
+ 𝑘
2026
+ 30
2027
+ k=30
2028
+ trajectories with ReAct and Reflexion results in a similar performance, suggesting the semantic feedback is not as helpful in complex environments like WebShop. Indeed like in
2029
+ Shinn et al. (
2030
+ 2023
2031
+ )
2032
+ , we find that generated reflections are often generic and do not provide useful feedback, resulting in a tendency for the agent to become stuck in local minima. However, using LATS indeed results in a noticeable improvement, indicating a more effective exploration for the same number of iterations.
2033
+ 5.4
2034
+ Additional Observations
2035
+ Method
2036
+ Score
2037
+ SR
2038
+ ReAct
2039
+ (Yao et al.,
2040
+ 2023b
2041
+ )
2042
+ 53.8
2043
+ 28.0
2044
+ ReAct (best of k)
2045
+ 59.1
2046
+ 32.0
2047
+ Reflexion
2048
+ (Shinn et al.,
2049
+ 2023
2050
+ )
2051
+ 64.2
2052
+ 35.0
2053
+ LATS
2054
+ 75.9
2055
+ 38.0
2056
+ IL
2057
+ 59.9
2058
+ 29.1
2059
+ IL+RL
2060
+ 62.4
2061
+ 28.7
2062
+ Fine-tuning
2063
+ (Furuta et al.,
2064
+ 2023
2065
+ )
2066
+ 67.5
2067
+ 45.0
2068
+ Expert
2069
+ 82.1
2070
+ 59.6
2071
+ Table 4:
2072
+ Score and success rate (SR) on Webshop. Table is separated into prompting, RL-based training, and human performance. For the same number of iterations, LATS improves both score and success rate, and surpasses RL-based training. IL/IL+RL taken from
2073
+ Yao et al. (
2074
+ 2022
2075
+ )
2076
+ .
2077
+ Prompt Method
2078
+ HotPotQA (EM)
2079
+ ToT (ReAct)
2080
+ 0.39
2081
+ RAP (ReAct)
2082
+ 0.54
2083
+ LATS (No LM Heuristic)
2084
+ 0.37
2085
+ LATS (DFS)
2086
+ 0.42
2087
+ LATS (No Reflection)
2088
+ 0.56
2089
+ LATS
2090
+ 0.61
2091
+ Table 5:
2092
+ Ablation results on LATS and baseline variants in HotPotQA; we use ReAct as the base prompt and sample
2093
+ n
2094
+ =
2095
+ 5
2096
+ 𝑛
2097
+ 5
2098
+ n=5
2099
+ children and
2100
+ k
2101
+ =
2102
+ 50
2103
+ 𝑘
2104
+ 50
2105
+ k=50
2106
+ maximum trajectories. LATS requires every component and operation for optimal performance.
2107
+ We also conduct additional experiments on HotPotQA to demonstrate the effect of each component of LATS. We also design a version of ToT and RAP with ReAct prompt and can handle external observations. We use HotPotQA as our setup incorporates both reasoning (through thoughts) and acting (through API calls); the results are shown in Tab.
2108
+ 5
2109
+ . More ablations for token consumption on HotPotQA are in Tab.
2110
+ 7
2111
+ in Appendix
2112
+ C
2113
+ . Note that baselines generally perform worse than the reasoning-only setting of HotPotQA, which indicates that the acting-based setting is more challenging and adaption of search algorithms to decision-making scenarios is non-trivial.
2114
+ Self-reflection.
2115
+ We use self-reflection to provide additional semantic signals for the agent. We observe a
2116
+ 0.05
2117
+ 0.05
2118
+ 0.05
2119
+ performance drop when removed from LATS, suggesting this is useful. This is a smaller gain Reflexion
2120
+ (Shinn et al.,
2121
+ 2023
2122
+ )
2123
+ observes over ReAct
2124
+ (Yao et al.,
2125
+ 2023b
2126
+ )
2127
+ as shown in Tab.
2128
+ 2
2129
+ , suggesting overlap between the types of questions where there is an improvement with self-reflection and search. This variant outperforms RAP-ReAct, reflecting our improvements to MCTS.
2130
+ Search Algorithm.
2131
+ MCTS is a more principled search algorithm than variants like A* or DFS search and the basis for observed performance gains. We observe the effects of using DFS, and incorporate the LM-based heuristic used in ToT
2132
+ (Yao et al.,
2133
+ 2023a
2134
+ )
2135
+ in which branches with low values are pruned. This removes the selection and backpropagation operations, and we observe a
2136
+ 0.08
2137
+ 0.08
2138
+ 0.08
2139
+ drop in performance when sampling the same number of nodes, but outperforms ToT-ReAct.
2140
+ 6
2141
+ Conclusion
2142
+ In this work, we introduce Language Agent Tree Search (LATS), the first framework to unify planning, acting, and reasoning for enhanced LLM problem solving. By deliberately constructing trajectories with search algorithms, incorporating external feedback, and enabling agents to learn from experience, LATS addresses key limitations of prior prompting techniques. Our evaluations demonstrate the ability of LATS to harness LLM capabilities for a variety of decision-making tasks while keeping its reasoning ability without additional training. The proposed synergies between search, interaction, and reflection offer a versatile approach to autonomous decision-making, highlighting the potential of LLMs as generalist agents. A full discussion of the limitations and broader impacts is in Appendix
2143
+ B
2144
+ .
2145
+ References
2146
+ Ahn et al. (2022)
2147
+ Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan, and Andy Zeng.
2148
+ Do as i can, not as i say: Grounding language in robotic affordances.
2149
+ arXiv:2204.01691
2150
+ , 2022.
2151
+ Anthony et al. (2017)
2152
+ T. Anthony, Z. Tian, and D. Barber.
2153
+ Thinking fast and slow with deep learning and tree search.
2154
+ In
2155
+ NIPS
2156
+ , 2017.
2157
+ Austin et al. (2021)
2158
+ Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al.
2159
+ Program synthesis with large language models.
2160
+ arXiv:2108.07732
2161
+ , 2021.
2162
+ Baker et al. (2022)
2163
+ Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune.
2164
+ Video pretraining (vpt): Learning to act by watching unlabeled online videos.
2165
+ arXiv:2206.11795
2166
+ , 2022.
2167
+ Besta et al. (2023)
2168
+ Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Michal Podstawski, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler.
2169
+ Graph of thoughts: Solving elaborate problems with large language models.
2170
+ arXiv:2308.09687
2171
+ , 2023.
2172
+ Bowman et al. (2015)
2173
+ Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning.
2174
+ A large annotated corpus for learning natural language inference.
2175
+ In
2176
+ EMNLP
2177
+ , 2015.
2178
+ Brown et al. (2020)
2179
+ Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei.
2180
+ Language models are few-shot learners.
2181
+ In
2182
+ NeurIPS
2183
+ , 2020.
2184
+ Campbell et al. (2002)
2185
+ Murray Campbell, A Joseph Hoane Jr, and Feng-hsiung Hsu.
2186
+ Deep blue.
2187
+ Artificial intelligence
2188
+ , 2002.
2189
+ Chen et al. (2022a)
2190
+ Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen.
2191
+ Codet: Code generation with generated tests.
2192
+ arXiv:2207.10397
2193
+ , 2022a.
2194
+ Chen et al. (2021)
2195
+ Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al.
2196
+ Evaluating large language models trained on code.
2197
+ arXiv:2107.03374
2198
+ , 2021.
2199
+ Chen et al. (2022b)
2200
+ Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen.
2201
+ Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks.
2202
+ arXiv preprint arXiv:2211.12588
2203
+ , 2022b.
2204
+ Chowdhery et al. (2022)
2205
+ Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al.
2206
+ Palm: Scaling language modeling with pathways.
2207
+ arXiv:2204.02311
2208
+ , 2022.
2209
+ Cobbe et al. (2021)
2210
+ Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al.
2211
+ Training verifiers to solve math word problems.
2212
+ arXiv:2110.14168
2213
+ , 2021.
2214
+ Deng et al. (2023)
2215
+ Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su.
2216
+ Mind2web: Towards a generalist agent for the web.
2217
+ arXiv:2306.06070
2218
+ , 2023.
2219
+ Driess et al. (2023)
2220
+ Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence.
2221
+ Palm-e: An embodied multimodal language model.
2222
+ arXiv:2303.03378
2223
+ , 2023.
2224
+ Du et al. (2023)
2225
+ Yilun Du, Mengjiao Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Joshua B. Tenenbaum, Dale Schuurmans, and Pieter Abbeel.
2226
+ Learning universal policies via text-guided video generation.
2227
+ arXiv:2302.00111
2228
+ , 2023.
2229
+ Evans (2010)
2230
+ Jonathan St BT Evans.
2231
+ Intuition and reasoning: A dual-process perspective.
2232
+ Psychological Inquiry
2233
+ , 2010.
2234
+ Fan et al. (2022)
2235
+ Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar.
2236
+ Minedojo: Building open-ended embodied agents with internet-scale knowledge.
2237
+ In
2238
+ NeurIPS Datasets and Benchmarks Track
2239
+ , 2022.
2240
+ Furuta et al. (2023)
2241
+ Hiroki Furuta, Ofir Nachum, Kuang-Huei Lee, Yutaka Matsuo, Shixiang Shane Gu, and Izzeddin Gur.
2242
+ Multimodal web navigation with instruction-finetuned foundation models.
2243
+ arXiv preprint arXiv:2305.11854
2244
+ , 2023.
2245
+ Gao et al. (2022)
2246
+ Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig.
2247
+ Pal: Program-aided language models.
2248
+ arXiv preprint arXiv:2211.10435
2249
+ , 2022.
2250
+ Guo et al. (2018)
2251
+ Jiaxian Guo, Sidi Lu, Han Cai, Weinan Zhang, Yong Yu, and Jun Wang.
2252
+ Long text generation via adversarial training with leaked information.
2253
+ AAAI
2254
+ , 2018.
2255
+ Guss et al. (2019)
2256
+ William H. Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela Veloso, and Ruslan Salakhutdinov.
2257
+ Minerl: A large-scale dataset of minecraft demonstrations.
2258
+ In
2259
+ IJCAI
2260
+ , 2019.
2261
+ Hafner et al. (2019)
2262
+ Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson.
2263
+ Learning latent dynamics for planning from pixels.
2264
+ In
2265
+ ICML
2266
+ , 2019.
2267
+ Hafner et al. (2023)
2268
+ Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap.
2269
+ Mastering diverse domains through world models.
2270
+ arXiv:2301.04104
2271
+ , 2023.
2272
+ Hao et al. (2023)
2273
+ Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu.
2274
+ Reasoning with language model is planning with world model.
2275
+ arXiv:2305.14992
2276
+ , 2023.
2277
+ Huang et al. (2023)
2278
+ Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou.
2279
+ Large language models cannot self-correct reasoning yet.
2280
+ arXiv:2310.01798
2281
+ , 2023.
2282
+ Huang et al. (2022)
2283
+ Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al.
2284
+ Inner monologue: Embodied reasoning through planning with language models.
2285
+ arXiv:2207.05608
2286
+ , 2022.
2287
+ Jiang et al. (2018)
2288
+ D. Jiang, E. Ekwedike, and H. Liu.
2289
+ Feedback-based tree search for reinforcement learning.
2290
+ In
2291
+ ICML
2292
+ , 2018.
2293
+ Kocsis & Szepesvári (2006)
2294
+ Levente Kocsis and Csaba Szepesvári.
2295
+ Bandit based monte-carlo planning.
2296
+ In
2297
+ ECML
2298
+ , 2006.
2299
+ Kojima et al. (2022)
2300
+ Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa.
2301
+ Large language models are zero-shot reasoners.
2302
+ arXiv:2205.11916
2303
+ , 2022.
2304
+ LaValle et al. (2001)
2305
+ Steven M LaValle, James J Kuffner, BR Donald, et al.
2306
+ Rapidly-exploring random trees: Progress and prospects.
2307
+ Algorithmic and computational robotics: new directions
2308
+ , 2001.
2309
+ Liu et al. (2018)
2310
+ Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang.
2311
+ Reinforcement learning on web interfaces using workflow-guided exploration.
2312
+ In
2313
+ ICLR
2314
+ , 2018.
2315
+ Liu et al. (2023)
2316
+ Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang.
2317
+ Agentbench: Evaluating llms as agents.
2318
+ arXiv:2308.03688
2319
+ , 2023.
2320
+ Madaan et al. (2023)
2321
+ Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark.
2322
+ Self-refine: Iterative refinement with self-feedback.
2323
+ arXiv:2303.17651
2324
+ , 2023.
2325
+ Nallapati et al. (2016)
2326
+ Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Caglar Gulcehre, and Bing Xiang.
2327
+ Abstractive text summarization using sequence-to-sequence rnns and beyond.
2328
+ In
2329
+ SIGNLL
2330
+ , 2016.
2331
+ Nye et al. (2021)
2332
+ Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al.
2333
+ Show your work: Scratchpads for intermediate computation with language models.
2334
+ arXiv:2112.00114
2335
+ , 2021.
2336
+ OpenAI (2023)
2337
+ OpenAI.
2338
+ Gpt-4 technical report.
2339
+ arXiv:2303.08774
2340
+ , 2023.
2341
+ Saparov & He (2022)
2342
+ Abulhair Saparov and He He.
2343
+ Language models are greedy reasoners: A systematic formal analysis of chain-of-thought.
2344
+ arXiv:2210.01240
2345
+ , 2022.
2346
+ Schick et al. (2023)
2347
+ Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom.
2348
+ Toolformer: Language models can teach themselves to use tools.
2349
+ arXiv:2302.04761
2350
+ , 2023.
2351
+ Shen et al. (2023)
2352
+ Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang.
2353
+ Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface.
2354
+ arXiv:2303.17580
2355
+ , 2023.
2356
+ Shinn et al. (2023)
2357
+ Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao.
2358
+ Reflexion: Language agents with verbal reinforcement learning.
2359
+ arXiv:2303.11366
2360
+ , 2023.
2361
+ Shridhar et al. (2020)
2362
+ Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht.
2363
+ Alfworld: Aligning text and embodied environments for interactive learning.
2364
+ arXiv:2010.03768
2365
+ , 2020.
2366
+ Silver et al. (2016)
2367
+ David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al.
2368
+ Mastering the game of go with deep neural networks and tree search.
2369
+ nature
2370
+ , 2016.
2371
+ Silver et al. (2017)
2372
+ David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy P. Lillicrap, Fan Hui, L. Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis.
2373
+ Mastering the game of go without human knowledge.
2374
+ Nature
2375
+ , 2017.
2376
+ Sloman (1996)
2377
+ Steven A. Sloman.
2378
+ The empirical case for two systems of reasoning.
2379
+ Psychological Bulletin
2380
+ , 1996.
2381
+ Sun et al. (2023)
2382
+ Haotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai, and Chao Zhang.
2383
+ Adaplanner: Adaptive planning from feedback with language models.
2384
+ arXiv:2305.16653
2385
+ , 2023.
2386
+ Surís et al. (2023)
2387
+ Dídac Surís, Sachit Menon, and Carl Vondrick.
2388
+ Vipergpt: Visual inference via python execution for reasoning.
2389
+ arXiv preprint arXiv:2303.08128
2390
+ , 2023.
2391
+ Świechowski et al. (2023)
2392
+ Maciej Świechowski, Konrad Godlewski, Bartosz Sawicki, and Jacek Mańdziuk.
2393
+ Monte carlo tree search: A review of recent modifications and applications.
2394
+ Artificial Intelligence Review
2395
+ , 2023.
2396
+ Touvron et al. (2023)
2397
+ Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Cantón Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony S. Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel M. Kloumann, A. V. Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, R. Subramanian, Xia Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zhengxu Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and
2398
+ Thomas Scialom.
2399
+ Llama 2: Open foundation and fine-tuned chat models.
2400
+ arXiv:2307.09288
2401
+ , 2023.
2402
+ Vodopivec et al. (2017)
2403
+ Tom Vodopivec, Spyridon Samothrakis, and Branko Ster.
2404
+ On monte carlo tree search and reinforcement learning.
2405
+ Journal of Artificial Intelligence Research
2406
+ , 2017.
2407
+ Wang et al. (2023)
2408
+ Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar.
2409
+ Voyager: An open-ended embodied agent with large language models.
2410
+ arXiv:2305.16291
2411
+ , 2023.
2412
+ Wang et al. (2022)
2413
+ Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou.
2414
+ Self-consistency improves chain of thought reasoning in language models.
2415
+ arXiv:2203.11171
2416
+ , 2022.
2417
+ Wei et al. (2022)
2418
+ Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou.
2419
+ Chain of thought prompting elicits reasoning in large language models.
2420
+ arXiv:2201.11903
2421
+ , 2022.
2422
+ Wooldridge & Jennings (1995)
2423
+ Michael Wooldridge and Nicholas R Jennings.
2424
+ Intelligent agents: Theory and practice.
2425
+ The knowledge engineering review
2426
+ , 1995.
2427
+ Wu et al. (2023)
2428
+ Philipp Wu, Alejandro Escontrela, Danijar Hafner, Pieter Abbeel, and Ken Goldberg.
2429
+ Daydreamer: World models for physical robot learning.
2430
+ In
2431
+ CoRL
2432
+ . PMLR, 2023.
2433
+ Xie et al. (2023)
2434
+ Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao, Min-Yen Kan, Junxian He, and Qizhe Xie.
2435
+ Decomposition enhances reasoning via self-evaluation guided decoding.
2436
+ arXiv:2305.00633
2437
+ , 2023.
2438
+ Yang et al. (2018)
2439
+ Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning.
2440
+ Hotpotqa: A dataset for diverse, explainable multi-hop question answering.
2441
+ arXiv:1809.09600
2442
+ , 2018.
2443
+ Yao et al. (2022)
2444
+ Shunyu Yao, Howard Chen, John Yang, and Karthik R Narasimhan.
2445
+ Webshop: Towards scalable real-world web interaction with grounded language agents.
2446
+ In
2447
+ NeurIPS
2448
+ , 2022.
2449
+ Yao et al. (2023a)
2450
+ Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan.
2451
+ Tree of thoughts: Deliberate problem solving with large language models.
2452
+ arXiv:2305.10601
2453
+ , 2023a.
2454
+ Yao et al. (2023b)
2455
+ Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao.
2456
+ ReAct: Synergizing reasoning and acting in language models.
2457
+ In
2458
+ ICLR
2459
+ , 2023b.
2460
+ Yao et al. (2023c)
2461
+ Weiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu, Yihao Feng, Le Xue, Rithesh Murthy, Zeyuan Chen, Jianguo Zhang, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese.
2462
+ Retroformer: Retrospective large language agents with policy gradient optimization.
2463
+ arXiv preprint arXiv:2308.02151
2464
+ , 2023c.
2465
+ Ye et al. (2021)
2466
+ Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao.
2467
+ Mastering atari games with limited data.
2468
+ In
2469
+ NeurIPS
2470
+ , 2021.
2471
+ Zhou et al. (2022)
2472
+ Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi.
2473
+ Least-to-most prompting enables complex reasoning in large language models.
2474
+ arXiv:2205.10625
2475
+ , 2022.
2476
+ Zhu et al. (2023)
2477
+ Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, Yu Qiao, Zhaoxiang Zhang, and Jifeng Dai.
2478
+ Ghost in the minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory.
2479
+ arXiv:2305.17144
2480
+ , 2023.
2481
+ 7
2482
+ Appendix
2483
+ The appendix is organized as follows. First in Sec.
2484
+ A
2485
+ , we show the pseudocode of our proposed algorithm, LATS; then in Sec.
2486
+ B
2487
+ , we provide further discussion of our method and its limitations, future direction and broader impact; then in Sec.
2488
+ C
2489
+ we provide additional experimental results; then in Sec.
2490
+ D
2491
+ , we specify the environment details in our experiments; finally, we list our prompts used for the three environments in Sec.
2492
+ E
2493
+ (HotPotQA), Sec.
2494
+ F
2495
+ (Programming) and Sec.
2496
+ G
2497
+ (Webshop) respectively.
2498
+ Appendix A
2499
+ LATS Pseudocode
2500
+ Alg.
2501
+ 1
2502
+ shows the pseudocode of our algorithm LATS. Nodes are stored explicitly in the memory. Unless otherwise specified, in all experiments we use
2503
+ n
2504
+ =
2505
+ 5
2506
+ 𝑛
2507
+ 5
2508
+ n=5
2509
+ and
2510
+ w
2511
+ =
2512
+ 1
2513
+ 𝑤
2514
+ 1
2515
+ w=1
2516
+ .
2517
+ Algorithm 1
2518
+ LATS
2519
+ ⁡
2520
+ (
2521
+ S
2522
+ 0
2523
+ ,
2524
+ p
2525
+ θ
2526
+ ,
2527
+ p
2528
+ V
2529
+ ,
2530
+ p
2531
+ ref
2532
+ ,
2533
+ d
2534
+ ,
2535
+ k
2536
+ ,
2537
+ n
2538
+ ,
2539
+ w
2540
+ )
2541
+ LATS
2542
+ subscript
2543
+ 𝑆
2544
+ 0
2545
+ subscript
2546
+ 𝑝
2547
+ 𝜃
2548
+ subscript
2549
+ 𝑝
2550
+ 𝑉
2551
+ subscript
2552
+ 𝑝
2553
+ ref
2554
+ 𝑑
2555
+ 𝑘
2556
+ 𝑛
2557
+ 𝑤
2558
+ \operatorname{LATS}(S_{0},p_{\theta},{p_{V}},p_{\text{ref}},d,k,n,w)
2559
+ Initial state
2560
+ s
2561
+ 1
2562
+ subscript
2563
+ 𝑠
2564
+ 1
2565
+ s_{1}
2566
+ , action generator
2567
+ p
2568
+ θ
2569
+ subscript
2570
+ 𝑝
2571
+ 𝜃
2572
+ p_{\theta}
2573
+ , value function
2574
+ p
2575
+ V
2576
+ subscript
2577
+ 𝑝
2578
+ 𝑉
2579
+ p_{V}
2580
+ , reflection generator
2581
+ p
2582
+ ref
2583
+ subscript
2584
+ 𝑝
2585
+ ref
2586
+ p_{\text{ref}}
2587
+ , number of generated actions
2588
+ n
2589
+ 𝑛
2590
+ n
2591
+ , depth limit
2592
+ L
2593
+ 𝐿
2594
+ L
2595
+ , number of roll-outs
2596
+ K
2597
+ 𝐾
2598
+ K
2599
+ , context
2600
+ c
2601
+ 𝑐
2602
+ c
2603
+ , and exploration weight
2604
+ w
2605
+ 𝑤
2606
+ w
2607
+ Initialize action space
2608
+ A
2609
+ 𝐴
2610
+ A
2611
+ , observation space
2612
+ O
2613
+ 𝑂
2614
+ O
2615
+ Initialize the state-action value function
2616
+ p
2617
+ V
2618
+ :
2619
+ S
2620
+ ×
2621
+ A
2622
+ ↦
2623
+ ℝ
2624
+ :
2625
+ subscript
2626
+ 𝑝
2627
+ 𝑉
2628
+ maps-to
2629
+ 𝑆
2630
+ 𝐴
2631
+ ℝ
2632
+ {p_{V}}:S\times A\mapsto\mathbb{R}
2633
+ and visit counter
2634
+ N
2635
+ :
2636
+ S
2637
+ ↦
2638
+ ℕ
2639
+ :
2640
+ 𝑁
2641
+ maps-to
2642
+ 𝑆
2643
+ ℕ
2644
+ {N}:S\mapsto\mathbb{N}
2645
+ to zero
2646
+ for
2647
+ k
2648
+ ←
2649
+ 0
2650
+ ,
2651
+ …
2652
+ ,
2653
+ K
2654
+ −
2655
+ 1
2656
+ ←
2657
+ 𝑘
2658
+ 0
2659
+ …
2660
+ 𝐾
2661
+ 1
2662
+ k\leftarrow 0,\dots,K-1
2663
+ do
2664
+ for
2665
+ t
2666
+ ←
2667
+ 0
2668
+ ,
2669
+ …
2670
+ ,
2671
+ L
2672
+ −
2673
+ 1
2674
+ ←
2675
+ 𝑡
2676
+ 0
2677
+ …
2678
+ 𝐿
2679
+ 1
2680
+ t\leftarrow 0,\dots,L-1
2681
+ do
2682
+ if
2683
+ s
2684
+ t
2685
+ subscript
2686
+ 𝑠
2687
+ 𝑡
2688
+ s_{t}
2689
+ not terminal
2690
+ then
2691
+ ▷
2692
+ ▷
2693
+ \triangleright
2694
+ Expansion & Simulation
2695
+ for
2696
+ i
2697
+ ←
2698
+ 1
2699
+ ,
2700
+ …
2701
+ ,
2702
+ n
2703
+ ←
2704
+ 𝑖
2705
+ 1
2706
+ …
2707
+ 𝑛
2708
+ i\leftarrow 1,\dots,n
2709
+ do
2710
+ Sample
2711
+ a
2712
+ t
2713
+ (
2714
+ i
2715
+ )
2716
+ ∼
2717
+ p
2718
+ θ
2719
+ ​
2720
+ (
2721
+ a
2722
+ ∣
2723
+ s
2724
+ t
2725
+ )
2726
+ similar-to
2727
+ superscript
2728
+ subscript
2729
+ 𝑎
2730
+ 𝑡
2731
+ 𝑖
2732
+ subscript
2733
+ 𝑝
2734
+ 𝜃
2735
+ conditional
2736
+ 𝑎
2737
+ subscript
2738
+ 𝑠
2739
+ 𝑡
2740
+ a_{t}^{(i)}\sim p_{\theta}(a\mid s_{t})
2741
+ Get
2742
+ o
2743
+ t
2744
+ (
2745
+ i
2746
+ )
2747
+ superscript
2748
+ subscript
2749
+ 𝑜
2750
+ 𝑡
2751
+ 𝑖
2752
+ o_{t}^{(i)}
2753
+ from environment,
2754
+ s
2755
+ t
2756
+ +
2757
+ 1
2758
+ (
2759
+ i
2760
+ )
2761
+ ←
2762
+ (
2763
+ c
2764
+ t
2765
+ (
2766
+ i
2767
+ )
2768
+ ,
2769
+ o
2770
+ t
2771
+ (
2772
+ i
2773
+ )
2774
+ ,
2775
+ a
2776
+ t
2777
+ (
2778
+ i
2779
+ )
2780
+ )
2781
+ ←
2782
+ superscript
2783
+ subscript
2784
+ 𝑠
2785
+ 𝑡
2786
+ 1
2787
+ 𝑖
2788
+ superscript
2789
+ subscript
2790
+ 𝑐
2791
+ 𝑡
2792
+ 𝑖
2793
+ superscript
2794
+ subscript
2795
+ 𝑜
2796
+ 𝑡
2797
+ 𝑖
2798
+ superscript
2799
+ subscript
2800
+ 𝑎
2801
+ 𝑡
2802
+ 𝑖
2803
+ s_{t+1}^{(i)}\leftarrow(c_{t}^{(i)},o_{t}^{(i)},a_{t}^{(i)})
2804
+ ,
2805
+ c
2806
+ t
2807
+ +
2808
+ 1
2809
+ (
2810
+ i
2811
+ )
2812
+ ←
2813
+ (
2814
+ o
2815
+ t
2816
+ (
2817
+ i
2818
+ )
2819
+ ,
2820
+ a
2821
+ t
2822
+ (
2823
+ i
2824
+ )
2825
+ )
2826
+ ←
2827
+ superscript
2828
+ subscript
2829
+ 𝑐
2830
+ 𝑡
2831
+ 1
2832
+ 𝑖
2833
+ superscript
2834
+ subscript
2835
+ 𝑜
2836
+ 𝑡
2837
+ 𝑖
2838
+ superscript
2839
+ subscript
2840
+ 𝑎
2841
+ 𝑡
2842
+ 𝑖
2843
+ c_{t+1}^{(i)}\leftarrow(o_{t}^{(i)},a_{t}^{(i)})
2844
+ Evaluate
2845
+ V
2846
+ t
2847
+ (
2848
+ i
2849
+ )
2850
+ ∼
2851
+ p
2852
+ V
2853
+ ​
2854
+ (
2855
+ s
2856
+ t
2857
+ (
2858
+ i
2859
+ )
2860
+ )
2861
+ similar-to
2862
+ superscript
2863
+ subscript
2864
+ 𝑉
2865
+ 𝑡
2866
+ 𝑖
2867
+ subscript
2868
+ 𝑝
2869
+ 𝑉
2870
+ superscript
2871
+ subscript
2872
+ 𝑠
2873
+ 𝑡
2874
+ 𝑖
2875
+ {V}_{t}^{(i)}\sim{p_{V}}(s_{t}^{(i)})
2876
+ ▷
2877
+ ▷
2878
+ \triangleright
2879
+ Evaluation
2880
+ V
2881
+ ​
2882
+ (
2883
+ s
2884
+ t
2885
+ )
2886
+ ←
2887
+ V
2888
+ t
2889
+ (
2890
+ i
2891
+ )
2892
+ ←
2893
+ 𝑉
2894
+ subscript
2895
+ 𝑠
2896
+ 𝑡
2897
+ superscript
2898
+ subscript
2899
+ 𝑉
2900
+ 𝑡
2901
+ 𝑖
2902
+ {V}(s_{t})\leftarrow{V}_{t}^{(i)}
2903
+ Add
2904
+ s
2905
+ t
2906
+ (
2907
+ i
2908
+ )
2909
+ superscript
2910
+ subscript
2911
+ 𝑠
2912
+ 𝑡
2913
+ 𝑖
2914
+ s_{t}^{(i)}
2915
+ to children
2916
+ end
2917
+ for
2918
+ end
2919
+ if
2920
+ if
2921
+ s
2922
+ t
2923
+ subscript
2924
+ 𝑠
2925
+ 𝑡
2926
+ s_{t}
2927
+ is terminal
2928
+ then
2929
+ ▷
2930
+ ▷
2931
+ \triangleright
2932
+ Reflection
2933
+ Get
2934
+ r
2935
+ 𝑟
2936
+ r
2937
+ from environment
2938
+ if
2939
+ r
2940
+ 𝑟
2941
+ r
2942
+ not success
2943
+ then
2944
+ reflection
2945
+ ←
2946
+ p
2947
+ ref
2948
+ ​
2949
+ (
2950
+ c
2951
+ t
2952
+ )
2953
+ ←
2954
+ reflection
2955
+ subscript
2956
+ 𝑝
2957
+ ref
2958
+ subscript
2959
+ 𝑐
2960
+ 𝑡
2961
+ \text{reflection}\leftarrow p_{\text{ref}}(c_{t})
2962
+ c
2963
+ ←
2964
+ reflection
2965
+ ←
2966
+ 𝑐
2967
+ reflection
2968
+ c\leftarrow\text{reflection}
2969
+ end
2970
+ if
2971
+ end
2972
+ if
2973
+ a
2974
+ t
2975
+ ←
2976
+ arg
2977
+ ⁡
2978
+ max
2979
+ a
2980
+ ∈
2981
+ e
2982
+ ​
2983
+ (
2984
+ s
2985
+ t
2986
+ )
2987
+ ⁡
2988
+ [
2989
+ V
2990
+ ​
2991
+ (
2992
+ s
2993
+ t
2994
+ )
2995
+ +
2996
+ w
2997
+ ​
2998
+ ln
2999
+ ⁡
3000
+ N
3001
+ ​
3002
+ (
3003
+ s
3004
+ t
3005
+ −
3006
+ 1
3007
+ )
3008
+ N
3009
+ ​
3010
+ (
3011
+ s
3012
+ t
3013
+ )
3014
+ ]
3015
+ ←
3016
+ subscript
3017
+ 𝑎
3018
+ 𝑡
3019
+ subscript
3020
+ 𝑎
3021
+ 𝑒
3022
+ subscript
3023
+ 𝑠
3024
+ 𝑡
3025
+ 𝑉
3026
+ subscript
3027
+ 𝑠
3028
+ 𝑡
3029
+ 𝑤
3030
+ 𝑁
3031
+ subscript
3032
+ 𝑠
3033
+ 𝑡
3034
+ 1
3035
+ 𝑁
3036
+ subscript
3037
+ 𝑠
3038
+ 𝑡
3039
+ a_{t}\leftarrow\arg\max_{a\in e(s_{t})}\left[{V(s_{t})}+w\sqrt{\frac{\ln{N}(s_{t-1})}{{N}(s_{t})}}\right]
3040
+ ▷
3041
+ ▷
3042
+ \triangleright
3043
+ Selection
3044
+ N
3045
+ ​
3046
+ (
3047
+ s
3048
+ t
3049
+ +
3050
+ 1
3051
+ )
3052
+ ←
3053
+ N
3054
+ ​
3055
+ (
3056
+ s
3057
+ t
3058
+ +
3059
+ 1
3060
+ )
3061
+ +
3062
+ 1
3063
+ ←
3064
+ 𝑁
3065
+ subscript
3066
+ 𝑠
3067
+ 𝑡
3068
+ 1
3069
+ 𝑁
3070
+ subscript
3071
+ 𝑠
3072
+ 𝑡
3073
+ 1
3074
+ 1
3075
+ {N}(s_{t+1})\leftarrow{N}(s_{t+1})+1
3076
+ if
3077
+ a
3078
+ t
3079
+ subscript
3080
+ 𝑎
3081
+ 𝑡
3082
+ a_{t}
3083
+ is an output action
3084
+ then
3085
+ break
3086
+ end
3087
+ for
3088
+ T
3089
+ ←
3090
+ ←
3091
+ 𝑇
3092
+ absent
3093
+ T\leftarrow
3094
+ the actual number of steps
3095
+ for
3096
+ t
3097
+ ←
3098
+ T
3099
+ −
3100
+ 1
3101
+ ,
3102
+ …
3103
+ ,
3104
+ 0
3105
+ ←
3106
+ 𝑡
3107
+ 𝑇
3108
+ 1
3109
+ …
3110
+ 0
3111
+ t\leftarrow T-1,\dots,0
3112
+ do
3113
+ ▷
3114
+ ▷
3115
+ \triangleright
3116
+ Backpropagation
3117
+ V
3118
+ ​
3119
+ (
3120
+ s
3121
+ t
3122
+ )
3123
+ ←
3124
+ V
3125
+ ​
3126
+ (
3127
+ s
3128
+ t
3129
+ )
3130
+ ​
3131
+ (
3132
+ N
3133
+ ​
3134
+ (
3135
+ s
3136
+ t
3137
+ )
3138
+ −
3139
+ 1
3140
+ )
3141
+ +
3142
+ r
3143
+ N
3144
+ ​
3145
+ (
3146
+ s
3147
+ t
3148
+ )
3149
+ ←
3150
+ 𝑉
3151
+ subscript
3152
+ 𝑠
3153
+ 𝑡
3154
+ 𝑉
3155
+ subscript
3156
+ 𝑠
3157
+ 𝑡
3158
+ 𝑁
3159
+ subscript
3160
+ 𝑠
3161
+ 𝑡
3162
+ 1
3163
+ 𝑟
3164
+ 𝑁
3165
+ subscript
3166
+ 𝑠
3167
+ 𝑡
3168
+ V(s_{t})\leftarrow\frac{V(s_{t})(N(s_{t})-1)+r}{N(s_{t})}
3169
+ end
3170
+ for
3171
+ end
3172
+ for
3173
+ Appendix B
3174
+ Discussion
3175
+ Limitations.
3176
+ Although LATS can improve reasoning and decision-making, this arrives at a higher computational cost relative to simpler prompting methods like ReAct or Reflexion. The search process takes more time than standard prompting or simpler techniques, and requires greater inference costs. While such an issue is mitigated by the fact that the number of nodes
3177
+ n
3178
+ 𝑛
3179
+ n
3180
+ expanded at every step provides a natural trade-off between performance and efficiency (setting
3181
+ n
3182
+ =
3183
+ 1
3184
+ 𝑛
3185
+ 1
3186
+ n=1
3187
+ makes the method as effecient as ReAct with multiple trials or CoT-SC), in practice we recommend using LATS for difficult tasks like programming or for situations where performance is prioritized over efficiency. We hope that continued advancements in LLMs will reduce costs and increase the practicality of LATS.
3188
+ Additionally, the benchmarks we use in this paper are relatively simple and focused on decision-making, compared to the complexity of real-world interactive environments. In addition, some environments might not easily support rollbacks to previous states. However, the design of LATS is flexible and can be adjusted to various resource constraints. Using planning-based prompting methods like LATS in environments like Minecraft
3189
+ (Fan et al.,
3190
+ 2022
3191
+ )
3192
+ and more reasoning benchmarks would be interesting avenues for future work.
3193
+ Broader impact.
3194
+ LATS is a framework that enhances LLM performance through interactions with an environment. This improvement in autonomous decision-making may facilitate harmful uses of LLMs. Alternatively, LATS enhances interpretability and the potential for greater alignment, as it generates understandable, high-level linguistic reasoning and actions through several rounds of decision-making and reflection, rather than relying on implicit, low-level token values.
3195
+ Appendix C
3196
+ Ablations
3197
+ Prompt Method
3198
+ HotpotQA (EM)
3199
+ LATS (w=0.5)
3200
+ 0.55
3201
+ LATS (w=2.0)
3202
+ 0.61
3203
+ LATS (d=4)
3204
+ 0.58
3205
+ LATS (CoT)
3206
+ 0.60
3207
+ LATS (No LM Heuristic)
3208
+ 0.37
3209
+ LATS
3210
+ 0.61
3211
+ Table 6:
3212
+ Ablation results on LATS and baseline variants in HotPotQA measured by Exact Match (EM). We test different depth
3213
+ d
3214
+ 𝑑
3215
+ d
3216
+ , exploration factor
3217
+ w
3218
+ 𝑤
3219
+ w
3220
+ , and versions of LATS using CoT and without the LM value function. We sample
3221
+ n
3222
+ =
3223
+ 5
3224
+ 𝑛
3225
+ 5
3226
+ n=5
3227
+ and
3228
+ k
3229
+ =
3230
+ 50
3231
+ 𝑘
3232
+ 50
3233
+ k=50
3234
+ trajectories.
3235
+ Figure 4:
3236
+ Performance over successive iterations on HumanEval with GPT-3.5.
3237
+ In this section, we ablate various designs of LATS. Experiments are conducted on HotPotQA with a maximum of
3238
+ k
3239
+ =
3240
+ 50
3241
+ 𝑘
3242
+ 50
3243
+ k=50
3244
+ trajectories and sampling size of
3245
+ n
3246
+ =
3247
+ 5
3248
+ 𝑛
3249
+ 5
3250
+ n=5
3251
+ and HumanEval with a maximum of
3252
+ k
3253
+ =
3254
+ 8
3255
+ 𝑘
3256
+ 8
3257
+ k=8
3258
+ trajectories and sampling size of
3259
+ n
3260
+ =
3261
+ 5
3262
+ 𝑛
3263
+ 5
3264
+ n=5
3265
+ . The result for HotPotQA is shown in Tab.
3266
+ 5
3267
+ and HumanEval in Fig.
3268
+ 4
3269
+ .
3270
+ Exploration weight.
3271
+ We find that there is lower performance on HotPotQA when the exploration weight
3272
+ w
3273
+ 𝑤
3274
+ w
3275
+ in the selection formula is decreased to
3276
+ 0.5
3277
+ 0.5
3278
+ 0.5
3279
+ , suggesting that this reduces the effectiveness of the search. Increasing
3280
+ w
3281
+ 𝑤
3282
+ w
3283
+ to
3284
+ 2.0
3285
+ 2.0
3286
+ 2.0
3287
+ does not lead to a performance improvement, but we tend to observe faster convergence. The optimal setting depends on the particular environment and complexity of the state space.
3288
+ Depth.
3289
+ In our main experiments we use a maximum depth of
3290
+ d
3291
+ =
3292
+ 7
3293
+ 𝑑
3294
+ 7
3295
+ d=7
3296
+ on HotPotQA for all methods, following previous work
3297
+ (Yao et al.,
3298
+ 2023b
3299
+ )
3300
+ . We ablate the effect on LATS after reducing it to
3301
+ d
3302
+ =
3303
+ 4
3304
+ 𝑑
3305
+ 4
3306
+ d=4
3307
+ . This results in only a slight drop in performance. We find that most questions can be answered within four steps, and using a greater number of steps tends to force the agent into local minima and rarely improves success.
3308
+ LM value function.
3309
+ The LM value function scores states based on expected future reward. Without this heuristic, the only signal to guide search would be from environment rewards for completed trajectories, which are scarce and often binary. When we remove the evaluation operation, we observe a dramatic
3310
+ 0.24
3311
+ 0.24
3312
+ 0.24
3313
+ drop in performance.
3314
+ Performance over time.
3315
+ To see the effects of increasing the number of trajectories sampled, we change
3316
+ k
3317
+ 𝑘
3318
+ k
3319
+ to different values. We conduct this experiment on HumanEval, which has a more noticeable difference due to sampling less trajectories. The results are shown in Fig.
3320
+ 4
3321
+ , in which LATS scales better with more iterations than Reflexion.
3322
+ Sample complexity and Token cost.
3323
+ One possible concern of LATS is that the tree-structured search might consume much more tokens than existing methods. To further study the computational cost of LATS compared to prior methods, we examine the sample complexity (i.e. asymptotic token cost) of all methods considered in this paper, and count the average number of nodes expanded by our method and other tree-structured methods (ToT and RAP) upon successful search on HotPotQA. We present the results in Tab.
3324
+ 7
3325
+ ; the result shows that our method has the same sample complexity as other tree-based search methods, and has less average number of nodes expanded upon success, which indicates less token cost. The token cost gap will be even larger when taking failed trajectories into account, since our method has higher success rate and reaches computational budget limit less often.
3326
+ Method
3327
+ Performance (
3328
+ ↑
3329
+ ↑
3330
+ \uparrow
3331
+ )
3332
+ Sample complexity (
3333
+ ↓
3334
+ ↓
3335
+ \downarrow
3336
+ )
3337
+ Avg. #nodes upon success (
3338
+ ↓
3339
+ ↓
3340
+ \downarrow
3341
+ )
3342
+ ReAct (Best
3343
+ k
3344
+ =
3345
+ 250
3346
+ 𝑘
3347
+ 250
3348
+ k=250
3349
+ )
3350
+ 0.42
3351
+ 0.42
3352
+ 0.42
3353
+ O
3354
+ ​
3355
+ (
3356
+ k
3357
+ )
3358
+ 𝑂
3359
+ 𝑘
3360
+ O(k)
3361
+ N/A
3362
+ CoT-SC (
3363
+ n
3364
+ =
3365
+ 1
3366
+ ,
3367
+ k
3368
+ =
3369
+ 250
3370
+ formulae-sequence
3371
+ 𝑛
3372
+ 1
3373
+ 𝑘
3374
+ 250
3375
+ n=1,k=250
3376
+ )
3377
+ 0.40
3378
+ 0.40
3379
+ 0.40
3380
+ O
3381
+ ​
3382
+ (
3383
+ k
3384
+ )
3385
+ 𝑂
3386
+ 𝑘
3387
+ O(k)
3388
+ N/A
3389
+ LATS (
3390
+ n
3391
+ =
3392
+ 1
3393
+ ,
3394
+ k
3395
+ =
3396
+ 50
3397
+ formulae-sequence
3398
+ 𝑛
3399
+ 1
3400
+ 𝑘
3401
+ 50
3402
+ n=1,k=50
3403
+ )
3404
+ 0.48
3405
+ 0.48
3406
+ 0.48
3407
+ O
3408
+ ​
3409
+ (
3410
+ k
3411
+ )
3412
+ 𝑂
3413
+ 𝑘
3414
+ O(k)
3415
+ N/A
3416
+ ToT (ReAct)
3417
+ 0.49
3418
+ 0.49
3419
+ 0.49
3420
+ O
3421
+ ​
3422
+ (
3423
+ k
3424
+ ​
3425
+ n
3426
+ )
3427
+ 𝑂
3428
+ 𝑘
3429
+ 𝑛
3430
+ O(kn)
3431
+ 84.05
3432
+ 84.05
3433
+ 84.05
3434
+ RAP (ReAct)
3435
+ 0.54
3436
+ 0.54
3437
+ 0.54
3438
+ O
3439
+ ​
3440
+ (
3441
+ k
3442
+ ​
3443
+ n
3444
+ )
3445
+ 𝑂
3446
+ 𝑘
3447
+ 𝑛
3448
+ O(kn)
3449
+ 70.60
3450
+ 70.60
3451
+ 70.60
3452
+ LATS (
3453
+ n
3454
+ =
3455
+ 5
3456
+ ,
3457
+ k
3458
+ =
3459
+ 50
3460
+ formulae-sequence
3461
+ 𝑛
3462
+ 5
3463
+ 𝑘
3464
+ 50
3465
+ n=5,k=50
3466
+ )
3467
+ 0.61
3468
+ 0.61
3469
+ 0.61
3470
+ O
3471
+ ​
3472
+ (
3473
+ k
3474
+ ​
3475
+ n
3476
+ )
3477
+ 𝑂
3478
+ 𝑘
3479
+ 𝑛
3480
+ O(kn)
3481
+ 66.65
3482
+ 66.65
3483
+ 66.65
3484
+ Table 7:
3485
+ The performance, sample complexity of different methods and average number of nodes expanded upon success by methods with tree-based search.
3486
+ n
3487
+ 𝑛
3488
+ n
3489
+ is the number of children nodes expanded at every step and
3490
+ k
3491
+ 𝑘
3492
+ k
3493
+ is the number of trajectories. Our method has the same sample complexity as other methods with tree-based search and expands less nodes upon success, which indicates lower token cost.
3494
+ Appendix D
3495
+ Environment Details
3496
+ D.1
3497
+ HotPotQA
3498
+ Figure 5:
3499
+ Example trajectories on HotPotQA for ReAct (left) and LATS (right). LATS can sample more actions and avoid failure from previous mistakes by evaluating states with an LM to guide the search toward promising areas of the tree.
3500
+ HotPotQA
3501
+ (Yang et al.,
3502
+ 2018
3503
+ )
3504
+ is a question-answering dataset that requires reasoning over multiple supporting documents to answer questions. It contains 113k Wikipedia-based question-answer pairs crafted by crowdworkers to be diverse, multi-hop, and explainable. Questions cover a range of types like entities, locations, dates, and comparison of shared properties between two entities. Crowdworkers also provide supporting facts from the documents that justify the answer. We use the HotPotQA benchmark setting with all the Wikipedia paragraphs to test retrieval. We use a randomly selected subset of 100 questions for our experiments and a maximum depth limit of 6. Fig.
3505
+ 5
3506
+ illustrates how ReAct and LATS work on an example task of HotPotQA, and gives a qualitative example on how LATS outperforms ReAct on the task.
3507
+ Action Space.
3508
+ We adopt the Wikipedia web API proposed in
3509
+ Yao et al. (
3510
+ 2023b
3511
+ )
3512
+ , with three types of actions to support interactive information retrieval:
3513
+ (1)
3514
+ search
3515
+ [
3516
+ entity
3517
+ ], which returns the first 5 sentences from the corresponding
3518
+ entity
3519
+ wiki page if it exists, or else suggests top-5 similar entities from the Wikipedia search engine,
3520
+ (2)
3521
+ lookup
3522
+ [
3523
+ string
3524
+ ], which returns the next sentence in the page containing
3525
+ string
3526
+ ,
3527
+ (3)
3528
+ finish
3529
+ [
3530
+ answer
3531
+ ], which finishes the current task with
3532
+ answer
3533
+ .
3534
+ These API calls and free-form thoughts form the action space for this environment.
3535
+ D.2
3536
+ Programming
3537
+ The HumanEval dataset
3538
+ (Chen et al.,
3539
+ 2021
3540
+ )
3541
+ is a collection of 164 handwritten programming problems introduced to evaluate the functional correctness of models for synthesizing programs from natural language descriptions. Each problem includes a function signature, docstring description, reference implementation, and multiple unit tests, with an average of 7.7 tests per problem. The programming tasks assess comprehension of natural language, reasoning, algorithms, and basic mathematics, at a difficulty level comparable to simple software interview questions. Pass rates are evaluated with the pass@k metric, where k samples are generated per problem and a problem is considered solved if any sample passes all tests. We use all 164 problems for our experiments and a maximum depth limit of 8.
3542
+ The Mostly Basic Programming Problems (MBPP)
3543
+ Austin et al. (
3544
+ 2021
3545
+ )
3546
+ benchmark contains 974 short Python functions designed to evaluate program synthesis techniques. The dataset was constructed by crowdsourcing from workers with basic Python knowledge. Each data point consists of a natural language description of a programming task, a reference solution implementation, and three test cases for functional correctness. The natural language prompts are typically short, one-sentence descriptions. Solutions cover common programming constructs including mathematical operations, list processing, string manipulation, and usage of the Python standard library. On average, solutions are 6.8 lines of code. The dataset is also supplemented with an additional set of 426 problems that were manually verified for unambiguous specifications, standard function signatures, and accurate test cases. We use a randomly selected subset of 397 problems for our experiments.
3547
+ D.3
3548
+ WebShop
3549
+ WebShop
3550
+ (Yao et al.,
3551
+ 2022
3552
+ )
3553
+ is an interactive web-based environment designed to evaluate agents on grounded language understanding and decision-making. It simulates an e-commerce shopping task by providing agents with over 1 million real-world products scraped from Amazon, spanning 5 categories and 113 subcategories. These products contain rich linguistic information, with an average text length of 262 words and a vocabulary size of 224k. In addition, there are over 800k unique product options available for customization. The environment renders webpages in two modes: HTML mode provides pixel-level observations with interactive elements, while simple mode converts the raw HTML into a structured text observation more amenable for training agents. The action space consists of query searches and button clicks, which transition between 4 page types: search, results, item and item-detail. Instructions are crowdsourced natural language specifying product attributes and options, with a total of 12k collected. Automatic rewards are computed by comparing the product purchased by the agent against the attributes and options specified in the instruction, using both lexical matching and semantic similarity metrics.
3554
+ Type
3555
+ Argument
3556
+ State
3557
+ →
3558
+ →
3559
+ \rightarrow
3560
+ Next State
3561
+ search
3562
+ [
3563
+ Query
3564
+ ]
3565
+ Search
3566
+ →
3567
+ →
3568
+ \rightarrow
3569
+ Results
3570
+ choose
3571
+ Back to search
3572
+ ∗
3573
+ *
3574
+ →
3575
+ →
3576
+ \rightarrow
3577
+ Search
3578
+ choose
3579
+ Prev/Next page
3580
+ Results
3581
+ →
3582
+ →
3583
+ \rightarrow
3584
+ Results
3585
+ choose
3586
+ [
3587
+ Product title
3588
+ ]
3589
+ Results
3590
+ →
3591
+ →
3592
+ \rightarrow
3593
+ Item
3594
+ choose
3595
+ [
3596
+ Option
3597
+ ]
3598
+ Item
3599
+ →
3600
+ →
3601
+ \rightarrow
3602
+ Item
3603
+ choose
3604
+ Desc/Overview
3605
+ Item
3606
+ →
3607
+ →
3608
+ \rightarrow
3609
+ Item-Detail
3610
+ choose
3611
+ Previous
3612
+ Item-Detail
3613
+ →
3614
+ →
3615
+ \rightarrow
3616
+ Item
3617
+ choose
3618
+ Buy
3619
+ Item
3620
+ →
3621
+ →
3622
+ \rightarrow
3623
+ Episode End
3624
+ Table 8:
3625
+ Action space of webshop.
3626
+ There are two evaluation metrics used in WebShop: (1)
3627
+ Task Score
3628
+ : defined as
3629
+ (
3630
+ 100
3631
+ ×
3632
+ avg. reward
3633
+ )
3634
+ 100
3635
+ avg. reward
3636
+ (100\times\text{avg. reward})
3637
+ , which captures the average reward obtained across episodes; and (2)
3638
+ Success Rate (SR)
3639
+ defined as the portion of instructions where
3640
+ r
3641
+ =
3642
+ 1
3643
+ 𝑟
3644
+ 1
3645
+ r=1
3646
+ . The reward is calculated based on the number of attributes satisfied by the selected item. We use 50 environments for our experiments and a maximum depth limit of 15.
3647
+ Appendix E
3648
+ HotPotQA Prompts
3649
+ E.1
3650
+ Base Acting Prompt
3651
+ Solve a question answering task with interleaving Thought, Action, Observation steps. Thought can reason about the current situation, and Action can be three types:
3652
+ (1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
3653
+ (2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
3654
+ (3) Finish[answer], which returns the answer and finishes the task.
3655
+ After each observation, provide the next Thought and next Action.
3656
+ Here are some examples:
3657
+ Question: Which magazine was started first Arthur’s Magazine or First for Women?
3658
+ Thought 1: I need to search Arthur’s Magazine and First for Women, and find which was started first.
3659
+ Action 1: Search[Arthur’s Magazine]
3660
+ Observation 1: Arthur’s Magazine (18440̆0e20̆0800̆0931846) was an American literary periodical published in Philadelphia in the 19th century. Edited by Timothy Shay Arthur, it featured work by Edgar A. Poe, J.H. Ingraham, Sarah Josepha Hale, Thomas G. Spear, and others.[1][2] In May 1846 it was merged into Godey’s Lady’s Book.[3]
3661
+ Thought 2: Arthur’s Magazine was started in 1844. I need to search First for Women next.
3662
+ Action 2: Search[First for Women]
3663
+ Observation 2: First for Women is a woman’s magazine published by Bauer Media Group in the USA.[1] The magazine was started in 1989. It is based in Englewood Cliffs, New Jersey.[2] In 2011 the circulation of the magazine was 1,310,696 copies.[3]
3664
+ Thought 3: First for Women was started in 1989. 1844 (Arthur’s Magazine) ¡ 1989 (First for Women), so Arthur’s Magazine was started first.
3665
+ Action 3: Finish[Arthur’s Magazine]
3666
+ (examples)
3667
+ You have attempted to answer the following question before and failed. The following reflection(s) give a plan to avoid failing to answer the question in the same way you did previously. Use them to improve your strategy of correctly answering the given question.
3668
+ (trajectories)
3669
+ (input)
3670
+ E.2
3671
+ Base Reasoning Prompt
3672
+ Solve a question answering task by having a Thought, then Finish with your answer. Thought can reason about the current situation. Finish[answer] returns the answer and finishes the task. You will be given context that you should use to help you answer the question. Start your response with either Action or an indexed Thought
3673
+ Here are some examples:
3674
+ Question: What is the elevation range for the area that the eastern sector of the Colorado orogeny extends into?
3675
+ Let’s think step by step.
3676
+ Thought 1: The eastern sector of Colorado orogeny extends into the High Plains.
3677
+ Thought 2: High Plains rise in elevation from around 1,800 to 7,000 ft
3678
+ Thought 3: The answer is 1,800 to 7,000 ft.
3679
+ Action: Finish[1,800 to 7,000 ft]
3680
+ (examples)
3681
+ Previous trial:
3682
+ (trajectories)
3683
+ (input)
3684
+ E.3
3685
+ Value Function Prompt
3686
+ Analyze the trajectories of a solution to a question answering task. The trajectories are labeled by environmental observations about the situation, thoughts that can reason about the current situation and actions that can be three types:
3687
+ (1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
3688
+ (2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
3689
+ (3) Finish[answer], which returns the answer and finishes the task.
3690
+ Given a question and a trajectory, evaluate its correctness and provide your reasoning and analysis in detail. Focus on the latest thought, action, and observation. Incomplete trajectories can be correct if the thoughts and actions so far are correct, even if the answer is not found yet. Do not generate additional thoughts or actions. Then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
3691
+ Question: Which magazine was started first Arthur’s Magazine or First for Women?
3692
+ Thought 1: I need to search Arthur’s Magazine and First for Women, and find which was started first.
3693
+ Action 1: Search[Arthur’s Magazine]
3694
+ Observation 1: Arthur’s Magazine (18440̆0e20̆0800̆0931846) was an American literary periodical published in Philadelphia in the 19th century. Edited by Timothy Shay Arthur, it featured work by Edgar A. Poe, J.H. Ingraham, Sarah Josepha Hale, Thomas G. Spear, and others.[1][2] In May 1846 it was merged into Godey’s Lady’s Book.[3]
3695
+ This trajectory is correct as it is reasonable to search for the first magazine provided in the question. It is also better to have simple searches corresponding to a single entity, making this the best action.
3696
+ Thus the correctness score is 10
3697
+ (other examples)
3698
+ (failed trajectories)
3699
+ (context)
3700
+ E.4
3701
+ Reflection Prompt
3702
+ Analyze the trajectories of a solution to a question answering task. The trajectories are labeled by environmental observations about the situation, thoughts that can reason about the current situation and actions that can be three types:
3703
+ (1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
3704
+ (2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
3705
+ (3) Finish[answer], which returns the answer and finishes the task.
3706
+ Given a question and a trajectory, evaluate its correctness and provide your reasoning and analysis in detail. Focus on the latest thought, action, and observation. Incomplete trajectories can be correct if the thoughts and actions so far are correct, even if the answer is not found yet. Do not generate additional thoughts or actions. Then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
3707
+ Question: Which magazine was started first Arthur’s Magazine or First for Women?
3708
+ Thought 1: I need to search Arthur’s Magazine and First for Women, and find which was started first.
3709
+ Action 1: Search[Arthur’s Magazine]
3710
+ Observation 1: Arthur’s Magazine (18440̆0e20̆0800̆0931846) was an American literary periodical published in Philadelphia in the 19th century. Edited by Timothy Shay Arthur, it featured work by Edgar A. Poe, J.H. Ingraham, Sarah Josepha Hale, Thomas G. Spear, and others.[1][2] In May 1846 it was merged into Godey’s Lady’s Book.[3]
3711
+ This trajectory is correct as it is reasonable to search for the first magazine provided in the question. It is also better to have simple searches corresponding to a single entity, making this the best action.
3712
+ Thus the correctness score is 10
3713
+ (other examples)
3714
+ (failed trajectories)
3715
+ (context)
3716
+ Appendix F
3717
+ Programming Prompts
3718
+ F.1
3719
+ HumanEval function implementation example
3720
+ Sample function signature:
3721
+ ⬇
3722
+ def
3723
+ minSubArraySum
3724
+ (
3725
+ nums
3726
+ ):
3727
+ Given
3728
+ an
3729
+ array
3730
+ of
3731
+ integers
3732
+ nums
3733
+ ,
3734
+ find
3735
+ the
3736
+ minimum
3737
+ sum
3738
+ of
3739
+ any
3740
+ non
3741
+ -
3742
+ empty
3743
+ sub
3744
+ -
3745
+ array
3746
+ of
3747
+ nums
3748
+ .
3749
+ Example
3750
+ minSubArraySum
3751
+ ([2,
3752
+ 3,
3753
+ 4,
3754
+ 1,
3755
+ 2,
3756
+ 4])
3757
+ ==
3758
+ 1
3759
+ minSubArraySum
3760
+ ([-1,
3761
+ -2,
3762
+ -3])
3763
+ ==
3764
+ -6
3765
+ Sample function body implementation:
3766
+ ⬇
3767
+ min_sum
3768
+ =
3769
+ float
3770
+ (’
3771
+ inf
3772
+ ’)
3773
+ for
3774
+ i
3775
+ in
3776
+ range
3777
+ (
3778
+ len
3779
+ (
3780
+ nums
3781
+ )):
3782
+ current_sum
3783
+ =
3784
+ 0
3785
+ for
3786
+ j
3787
+ in
3788
+ range
3789
+ (
3790
+ i
3791
+ ,
3792
+ len
3793
+ (
3794
+ nums
3795
+ )):
3796
+ current_sum
3797
+ +=
3798
+ nums
3799
+ [
3800
+ j
3801
+ ]
3802
+ if
3803
+ current_sum
3804
+ <
3805
+ min_sum
3806
+ :
3807
+ min_sum
3808
+ =
3809
+ current_sum
3810
+ return
3811
+ min_sum
3812
+ F.2
3813
+ Base Acting/Reasoning Prompt
3814
+ You are an AI Python assistant. You will be given your previous implementation of a function, a series of unit tests results, and your self-reflection on your previous implementation. Write your full implementation (restate the function signature).
3815
+ Example 1:
3816
+ [previous impl]:
3817
+ ⬇
3818
+ def
3819
+ add
3820
+ (
3821
+ a
3822
+ :
3823
+ int
3824
+ ,
3825
+ b
3826
+ :
3827
+ int
3828
+ )
3829
+ ->
3830
+ int
3831
+ :
3832
+ ”””
3833
+ Given
3834
+ integers
3835
+ a
3836
+ and
3837
+ b
3838
+ ,
3839
+ return
3840
+ the
3841
+ total
3842
+ value
3843
+ of
3844
+ a
3845
+ and
3846
+ b
3847
+ .
3848
+ ”””
3849
+ return
3850
+ a
3851
+ -
3852
+ b
3853
+ [unit test results from previous impl]:
3854
+ Tested passed:
3855
+ Tests failed:
3856
+ assert add(1, 2) == 3 # output: -1
3857
+ assert add(1, 2) == 4 # output: -1
3858
+ [reflection on previous impl]:
3859
+ The implementation failed the test cases where the input integers are 1 and 2. The issue arises because the code does not add the two integers together, but instead subtracts the second integer from the first. To fix this issue, we should change the operator from ‘-‘ to ‘+‘ in the return statement. This will ensure that the function returns the correct output for the given input.
3860
+ [improved impl]:
3861
+ ⬇
3862
+ def
3863
+ add
3864
+ (
3865
+ a
3866
+ :
3867
+ int
3868
+ ,
3869
+ b
3870
+ :
3871
+ int
3872
+ )
3873
+ ->
3874
+ int
3875
+ :
3876
+ ”””
3877
+ Given
3878
+ integers
3879
+ a
3880
+ and
3881
+ b
3882
+ ,
3883
+ return
3884
+ the
3885
+ total
3886
+ value
3887
+ of
3888
+ a
3889
+ and
3890
+ b
3891
+ .
3892
+ ”””
3893
+ return
3894
+ a
3895
+ +
3896
+ b
3897
+ F.3
3898
+ Reflection Prompt
3899
+ You are a Python programming assistant. You will be given a function implementation and a series of unit test results. Your goal is to write a few sentences to explain why your implementation is wrong as indicated by the tests. You will need this as guidance when you try again later. Only provide the few sentence description in your answer, not the implementation. You will be given a few examples by the user.
3900
+ Example 1:
3901
+ [previous impl]:
3902
+ ⬇
3903
+ def
3904
+ add
3905
+ (
3906
+ a
3907
+ :
3908
+ int
3909
+ ,
3910
+ b
3911
+ :
3912
+ int
3913
+ )
3914
+ ->
3915
+ int
3916
+ :
3917
+ ”””
3918
+ Given
3919
+ integers
3920
+ a
3921
+ and
3922
+ b
3923
+ ,
3924
+ return
3925
+ the
3926
+ total
3927
+ value
3928
+ of
3929
+ a
3930
+ and
3931
+ b
3932
+ .
3933
+ ”””
3934
+ return
3935
+ a
3936
+ -
3937
+ b
3938
+ [unit test results from previous impl]:
3939
+ Tested passed:
3940
+ Tests failed:
3941
+ assert add(1, 2) == 3 # output: -1
3942
+ assert add(1, 2) == 4 # output: -1
3943
+ [reflection on previous impl]:
3944
+ The implementation failed the test cases where the input integers are 1 and 2. The issue arises because the code does not add the two integers together, but instead subtracts the second integer from the first. To fix this issue, we should change the operator from ‘-‘ to ‘+‘ in the return statement. This will ensure that the function returns the correct output for the given input.
3945
+ F.4
3946
+ Test Case Generation Prompt
3947
+ You are an AI coding assistant that can write unique, diverse, and intuitive unit tests for functions given the signature and docstring.
3948
+ (input)
3949
+ Appendix G
3950
+ WebShop Prompts
3951
+ G.1
3952
+ Acting Prompt
3953
+ Webshop
3954
+ Instruction:
3955
+ i would like a 3 ounce bottle of bright citrus deodorant for sensitive skin, and price lower than 50.00 dollars
3956
+ [Search]
3957
+ Action: search[3 ounce bright citrus deodorant sensitive skin]
3958
+ Observation:
3959
+ [Back to Search]
3960
+ Page 1 (Total results: 50)
3961
+ [Next
3962
+ ]
3963
+ [B078GWRC1J]
3964
+ Bright Citrus Deodorant by Earth Mama — Natural and Safe for Sensitive Skin, Pregnancy and Breastfeeding, Contains Organic Calendula 3-Ounce
3965
+ $10.99
3966
+ [B078GTKVXY]
3967
+ Ginger Fresh Deodorant by Earth Mama — Natural and Safe for Sensitive Skin, Pregnancy and Breastfeeding, Contains Organic Calendula 3-Ounce
3968
+ $10.99
3969
+ [B08KBVJ4XN]
3970
+ Barrel and Oak - Aluminum-Free Deodorant, Deodorant for Men, Essential Oil-Based Scent, 24-Hour Odor Protection, Cedar & Patchouli Blend, Gentle on Sensitive Skin (Mountain Sage, 2.7 oz, 2-Pack)
3971
+ $15.95
3972
+ Action: think[B078GWRC1J and B078GTKVXY are bright citrus deodorant less then 50 dollars. I can check B078GWRC1J first.]
3973
+ Observation: OK.
3974
+ Action: click[B078GWRC1J]
3975
+ Observation:
3976
+ [Back to Search]
3977
+ [Prev]
3978
+ scent [assorted scents][bright citrus][calming lavender][ginger fresh][simply non-scents]
3979
+ size [travel set (4-pack)][3 ounce (pack of 1)][3-ounce (2-pack)]
3980
+ Bright Citrus Deodorant by Earth Mama — Natural and Safe for Sensitive Skin, Pregnancy and Breastfeeding, Contains Organic Calendula 3-Ounce
3981
+ Price: $10.99
3982
+ Rating: N.A.
3983
+ [Description]
3984
+ [Features]
3985
+ [Reviews]
3986
+ [Buy Now]
3987
+ Action: think[For 3 ounce bottle of bright citrus deodorant for sensitive skin, the item has options ’bright citrus’ and ’3 ounce (pack of 1)’ and seems good to buy.]
3988
+ Observation: OK.
3989
+ Action: click[bright citrus]
3990
+ Observation: You have clicked bright citrus.
3991
+ Action: click[3 ounce (pack of 1)]
3992
+ Observation: You have clicked 3 ounce (pack of 1).
3993
+ Action: click[Buy Now]
3994
+ G.2
3995
+ Value Function Prompt
3996
+ Given an item to purchase and a trajectory that aims to buy an item that exactly matches the specification, which corresponds to the ideal score of 1.0, analyze the following trajectory, then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
3997
+ Here are some examples
3998
+ Webshop
3999
+ Instruction:
4000
+ i am looking for dairy free and apple variety pack of chips, and price lower than 30.00 dollars
4001
+ [Search]
4002
+ Action: search[dairy free and apple variety pack of chips]
4003
+ Observation:
4004
+ [Back to Search]
4005
+ Page 1 (Total results: 50)
4006
+ [Next]
4007
+ [B07HRFJWP8]
4008
+ Enjoy Life Foods Soft Baked Ovals, Breakfast Bars, Nut Free Bars, Soy Free, Dairy Free, Non GMO, Gluten Free, Vegan, Variety Pack, 4 Boxes (20 Bars Total)
4009
+ $100.0
4010
+ [B01KMHY5PG]
4011
+ Enjoy Life Soft Baked Chewy Bars, Variety Pack, Nut Free Bars, Soy Free, Dairy Free, Gluten Free, 6 Boxes (30 Total Bars)
4012
+ $21.49
4013
+ [B008D2X8C4]
4014
+ Enjoy Life Lentil Chips Variety Pack, Dairy Free Chips, Soy Free, Nut Free, Non GMO, Vegan, Gluten Free, 24 Bags (0.8 oz)
4015
+ $100.0
4016
+ Action: think[The first two items are not what I want. The third item is the dairy free and apple variety pack of chips I want.]
4017
+ Observation: OK.
4018
+ Action: click[B008D2X8C4]
4019
+ Observation:
4020
+ [Back to Search]
4021
+ [Prev]
4022
+ flavor name [dill and sour cream][garlic & parmesan][light sea salt][margherita pizza][thai chili lime][variety pack]
4023
+ size [0.8 ounce (pack of 24)][4 ounce (pack of 12)]
4024
+ Price: $100.0
4025
+ Rating: N.A.
4026
+ [Description]
4027
+ [Features]
4028
+ [Reviews]
4029
+ [Buy Now]
4030
+ Action: think[The item has options ’variety pack’ and ’0.8 ounce (pack of 24)’ and seems good to buy.]
4031
+ Observation: OK.
4032
+ Action: click[variety pack]
4033
+ Observation: You have clicked variety pack.
4034
+ Action: click[0.8 ounce (pack of 24)]
4035
+ Observation: You have clicked 0.8 ounce (pack of 24).
4036
+ Action: click[Buy Now]
4037
+ Reflection: In this attempt, I was unsuccessful. I accidentally bought a product that was $100, which is more than my budget of $30. Either way, the initial search results were not good. Next time, I will do search[”variety pack of chips”] and then check if the results meet the dairy free and the $30 budget constraints. I will continue to refine my searches so that I can find more products.
4038
+ Thus the correctness score is 5
4039
+ (input)
4040
+ G.3
4041
+ Reflection Prompt
4042
+ Given an item to purchase and a trajectory that aims to buy an item that exactly matches the specification, which corresponds to the ideal score of 1.0, analyze the following trajectory, then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
4043
+ Here are some examples
4044
+ Webshop
4045
+ Instruction:
4046
+ i am looking for dairy free and apple variety pack of chips, and price lower than 30.00 dollars
4047
+ [Search]
4048
+ Action: search[dairy free and apple variety pack of chips]
4049
+ Observation:
4050
+ [Back to Search]
4051
+ Page 1 (Total results: 50)
4052
+ [Next]
4053
+ [B07HRFJWP8]
4054
+ Enjoy Life Foods Soft Baked Ovals, Breakfast Bars, Nut Free Bars, Soy Free, Dairy Free, Non GMO, Gluten Free, Vegan, Variety Pack, 4 Boxes (20 Bars Total)
4055
+ $100.0
4056
+ [B01KMHY5PG]
4057
+ Enjoy Life Soft Baked Chewy Bars, Variety Pack, Nut Free Bars, Soy Free, Dairy Free, Gluten Free, 6 Boxes (30 Total Bars)
4058
+ $21.49
4059
+ [B008D2X8C4]
4060
+ Enjoy Life Lentil Chips Variety Pack, Dairy Free Chips, Soy Free, Nut Free, Non GMO, Vegan, Gluten Free, 24 Bags (0.8 oz)
4061
+ $100.0
4062
+ Action: think[The first two items are not what I want. The third item is the dairy free and apple variety pack of chips I want.]
4063
+ Observation: OK.
4064
+ Action: click[B008D2X8C4]
4065
+ Observation:
4066
+ [Back to Search]
4067
+ [Prev]
4068
+ flavor name [dill and sour cream][garlic & parmesan][light sea salt][margherita pizza][thai chili lime][variety pack]
4069
+ size [0.8 ounce (pack of 24)][4 ounce (pack of 12)]
4070
+ Price: $100.0
4071
+ Rating: N.A.
4072
+ [Description]
4073
+ [Features]
4074
+ [Reviews]
4075
+ [Buy Now]
4076
+ Action: think[The item has options ’variety pack’ and ’0.8 ounce (pack of 24)’ and seems good to buy.]
4077
+ Observation: OK.
4078
+ Action: click[variety pack]
4079
+ Observation: You have clicked variety pack.
4080
+ Action: click[0.8 ounce (pack of 24)]
4081
+ Observation: You have clicked 0.8 ounce (pack of 24).
4082
+ Action: click[Buy Now]
4083
+ Reflection: In this attempt, I was unsuccessful. I accidentally bought a product that was $100, which is more than my budget of $30. Either way, the initial search results were not good. Next time, I will do search[”variety pack of chips”] and then check if the results meet the dairy free and the $30 budget constraints. I will continue to refine my searches so that I can find more products.
4084
+ (input)
4085
+ Reflection:
4086
+ ◄
4087
+ Feeling
4088
+ lucky?
4089
+ Conversion
4090
+ report
4091
+ Report
4092
+ an issue
4093
+ View original
4094
+ on arXiv
4095
+ ►
research/notes/231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la.md ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2310.04406] Language Agent Tree Search Unifies Reasoning Acting and Planning
3
+ in Language Models'
4
+ id: 231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:39:54.848871Z'
8
+ source: https://arxiv.org/abs/2310.04406
9
+ source_domain: arxiv.org
10
+ fetched_at: '2026-06-10T00:39:54.848723Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2310.04406] Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
20
+ Computer Science > Artificial Intelligence
21
+ arXiv:2310.04406
22
+ (cs)
23
+ [Submitted on 6 Oct 2023 (
24
+ v1
25
+ ), last revised 6 Jun 2024 (this version, v3)]
26
+ Title:
27
+ Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
28
+ Authors:
29
+ Andy Zhou
30
+ ,
31
+ Kai Yan
32
+ ,
33
+ Michal Shlapentokh-Rothman
34
+ ,
35
+ Haohan Wang
36
+ ,
37
+ Yu-Xiong Wang
38
+ View a PDF of the paper titled Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models, by Andy Zhou and 4 other authors
39
+ View PDF
40
+ HTML (experimental)
41
+ Abstract:
42
+ While language models (LMs) have shown potential across a range of decision-making tasks, their reliance on simple acting processes limits their broad deployment as autonomous agents. In this paper, we introduce Language Agent Tree Search (LATS) -- the first general framework that synergizes the capabilities of LMs in reasoning, acting, and planning. By leveraging the in-context learning ability of LMs, we integrate Monte Carlo Tree Search into LATS to enable LMs as agents, along with LM-powered value functions and self-reflections for proficient exploration and enhanced decision-making. A key feature of our approach is the incorporation of an environment for external feedback, which offers a more deliberate and adaptive problem-solving mechanism that surpasses the constraints of existing techniques. Our experimental evaluation across diverse domains, including programming, interactive question-answering (QA), web navigation, and math, validates the effectiveness and generality of LATS in decision-making while maintaining competitive or improved reasoning performance. Notably, LATS achieves state-of-the-art pass@1 accuracy (92.7%) for programming on HumanEval with GPT-4 and demonstrates gradient-free performance (average score of 75.9) comparable to gradient-based fine-tuning for web navigation on WebShop with GPT-3.5. Code can be found at
43
+ this https URL
44
+ Comments:
45
+ Code at
46
+ this https URL
47
+ Subjects:
48
+ Artificial Intelligence (cs.AI)
49
+ ; Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
50
+ Cite as:
51
+ arXiv:2310.04406
52
+ [cs.AI]
53
+ (or
54
+ arXiv:2310.04406v3
55
+ [cs.AI]
56
+ for this version)
57
+ https://doi.org/10.48550/arXiv.2310.04406
58
+ Focus to learn more
59
+ arXiv-issued DOI via DataCite
60
+ Submission history
61
+ From: Andy Zhou [
62
+ view email
63
+ ]
64
+ [v1]
65
+ Fri, 6 Oct 2023 17:55:11 UTC (371 KB)
66
+ [v2]
67
+ Tue, 5 Dec 2023 05:25:55 UTC (465 KB)
68
+ [v3]
69
+ Thu, 6 Jun 2024 02:51:17 UTC (960 KB)
70
+ Full-text links:
71
+ Access Paper:
72
+ View a PDF of the paper titled Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models, by Andy Zhou and 4 other authors
73
+ View PDF
74
+ HTML (experimental)
75
+ TeX Source
76
+ view license
77
+ Current browse context:
78
+ cs.AI
79
+ < prev
80
+ |
81
+ next >
82
+ new
83
+ |
84
+ recent
85
+ |
86
+ 2023-10
87
+ Change to browse by:
88
+ cs
89
+ cs.CL
90
+ cs.CV
91
+ cs.LG
92
+ References & Citations
93
+ NASA ADS
94
+ Google Scholar
95
+ Semantic Scholar
96
+ export BibTeX citation
97
+ Loading...
98
+ BibTeX formatted citation
99
+ ×
100
+ loading...
101
+ Data provided by:
102
+ Bookmark
103
+ Bibliographic Tools
104
+ Bibliographic and Citation Tools
105
+ Bibliographic Explorer Toggle
106
+ Bibliographic Explorer
107
+ (
108
+ What is the Explorer?
109
+ )
110
+ Connected Papers Toggle
111
+ Connected Papers
112
+ (
113
+ What is Connected Papers?
114
+ )
115
+ Litmaps Toggle
116
+ Litmaps
117
+ (
118
+ What is Litmaps?
119
+ )
120
+ scite.ai Toggle
121
+ scite Smart Citations
122
+ (
123
+ What are Smart Citations?
124
+ )
125
+ Code, Data, Media
126
+ Code, Data and Media Associated with this Article
127
+ alphaXiv Toggle
128
+ alphaXiv
129
+ (
130
+ What is alphaXiv?
131
+ )
132
+ Links to Code Toggle
133
+ CatalyzeX Code Finder for Papers
134
+ (
135
+ What is CatalyzeX?
136
+ )
137
+ DagsHub Toggle
138
+ DagsHub
139
+ (
140
+ What is DagsHub?
141
+ )
142
+ GotitPub Toggle
143
+ Gotit.pub
144
+ (
145
+ What is GotitPub?
146
+ )
147
+ Huggingface Toggle
148
+ Hugging Face
149
+ (
150
+ What is Huggingface?
151
+ )
152
+ ScienceCast Toggle
153
+ ScienceCast
154
+ (
155
+ What is ScienceCast?
156
+ )
157
+ Demos
158
+ Demos
159
+ Replicate Toggle
160
+ Replicate
161
+ (
162
+ What is Replicate?
163
+ )
164
+ Spaces Toggle
165
+ Hugging Face Spaces
166
+ (
167
+ What is Spaces?
168
+ )
169
+ Spaces Toggle
170
+ TXYZ.AI
171
+ (
172
+ What is TXYZ.AI?
173
+ )
174
+ Related Papers
175
+ Recommenders and Search Tools
176
+ Link to Influence Flower
177
+ Influence Flower
178
+ (
179
+ What are Influence Flowers?
180
+ )
181
+ Core recommender toggle
182
+ CORE Recommender
183
+ (
184
+ What is CORE?
185
+ )
186
+ Author
187
+ Venue
188
+ Institution
189
+ Topic
190
+ About arXivLabs
191
+ arXivLabs: experimental projects with community collaborators
192
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
193
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
194
+ Have an idea for a project that will add value for arXiv's community?
195
+ Learn more about arXivLabs
196
+ .
197
+ Which authors of this paper are endorsers?
198
+ |
199
+ Disable MathJax
200
+ (
201
+ What is MathJax?
202
+ )
research/notes/231006770-swe-bench-can-language-models-resolve-real-world-github-issues.md ADDED
@@ -0,0 +1,203 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2310.06770] SWE-bench: Can Language Models Resolve Real-World GitHub Issues?'
3
+ id: 231006770-swe-bench-can-language-models-resolve-real-world-github-issues
4
+ tags:
5
+ - deepread
6
+ created: '2026-06-10T00:23:35.577828Z'
7
+ source: https://arxiv.org/abs/2310.06770
8
+ source_domain: arxiv.org
9
+ fetched_at: '2026-06-10T00:23:35.577638Z'
10
+ fetch_provider: builtin
11
+ status: draft
12
+ type: note
13
+ tier: institutional
14
+ content_type: paper
15
+ deprecated: false
16
+ ---
17
+
18
+ [2310.06770] SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
19
+ Computer Science > Computation and Language
20
+ arXiv:2310.06770
21
+ (cs)
22
+ [Submitted on 10 Oct 2023 (
23
+ v1
24
+ ), last revised 11 Nov 2024 (this version, v3)]
25
+ Title:
26
+ SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
27
+ Authors:
28
+ Carlos E. Jimenez
29
+ ,
30
+ John Yang
31
+ ,
32
+ Alexander Wettig
33
+ ,
34
+ Shunyu Yao
35
+ ,
36
+ Kexin Pei
37
+ ,
38
+ Ofir Press
39
+ ,
40
+ Karthik Narasimhan
41
+ View a PDF of the paper titled SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, by Carlos E. Jimenez and 6 other authors
42
+ View PDF
43
+ Abstract:
44
+ Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models. To this end, we introduce SWE-bench, an evaluation framework consisting of $2,294$ software engineering problems drawn from real GitHub issues and corresponding pull requests across $12$ popular Python repositories. Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks. Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere $1.96$% of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous.
45
+ Comments:
46
+ Data, code, and leaderboard are available at
47
+ this https URL
48
+ ICLR 2024,
49
+ this https URL
50
+ Subjects:
51
+ Computation and Language (cs.CL)
52
+ ; Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
53
+ Cite as:
54
+ arXiv:2310.06770
55
+ [cs.CL]
56
+ (or
57
+ arXiv:2310.06770v3
58
+ [cs.CL]
59
+ for this version)
60
+ https://doi.org/10.48550/arXiv.2310.06770
61
+ Focus to learn more
62
+ arXiv-issued DOI via DataCite
63
+ Submission history
64
+ From: Carlos E. Jimenez [
65
+ view email
66
+ ]
67
+ [v1]
68
+ Tue, 10 Oct 2023 16:47:29 UTC (2,003 KB)
69
+ [v2]
70
+ Fri, 5 Apr 2024 18:16:29 UTC (2,258 KB)
71
+ [v3]
72
+ Mon, 11 Nov 2024 23:05:04 UTC (2,398 KB)
73
+ Full-text links:
74
+ Access Paper:
75
+ View a PDF of the paper titled SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, by Carlos E. Jimenez and 6 other authors
76
+ View PDF
77
+ TeX Source
78
+ view license
79
+ Current browse context:
80
+ cs.CL
81
+ < prev
82
+ |
83
+ next >
84
+ new
85
+ |
86
+ recent
87
+ |
88
+ 2023-10
89
+ Change to browse by:
90
+ cs
91
+ cs.AI
92
+ cs.SE
93
+ References & Citations
94
+ NASA ADS
95
+ Google Scholar
96
+ Semantic Scholar
97
+ export BibTeX citation
98
+ Loading...
99
+ BibTeX formatted citation
100
+ ×
101
+ loading...
102
+ Data provided by:
103
+ Bookmark
104
+ Bibliographic Tools
105
+ Bibliographic and Citation Tools
106
+ Bibliographic Explorer Toggle
107
+ Bibliographic Explorer
108
+ (
109
+ What is the Explorer?
110
+ )
111
+ Connected Papers Toggle
112
+ Connected Papers
113
+ (
114
+ What is Connected Papers?
115
+ )
116
+ Litmaps Toggle
117
+ Litmaps
118
+ (
119
+ What is Litmaps?
120
+ )
121
+ scite.ai Toggle
122
+ scite Smart Citations
123
+ (
124
+ What are Smart Citations?
125
+ )
126
+ Code, Data, Media
127
+ Code, Data and Media Associated with this Article
128
+ alphaXiv Toggle
129
+ alphaXiv
130
+ (
131
+ What is alphaXiv?
132
+ )
133
+ Links to Code Toggle
134
+ CatalyzeX Code Finder for Papers
135
+ (
136
+ What is CatalyzeX?
137
+ )
138
+ DagsHub Toggle
139
+ DagsHub
140
+ (
141
+ What is DagsHub?
142
+ )
143
+ GotitPub Toggle
144
+ Gotit.pub
145
+ (
146
+ What is GotitPub?
147
+ )
148
+ Huggingface Toggle
149
+ Hugging Face
150
+ (
151
+ What is Huggingface?
152
+ )
153
+ ScienceCast Toggle
154
+ ScienceCast
155
+ (
156
+ What is ScienceCast?
157
+ )
158
+ Demos
159
+ Demos
160
+ Replicate Toggle
161
+ Replicate
162
+ (
163
+ What is Replicate?
164
+ )
165
+ Spaces Toggle
166
+ Hugging Face Spaces
167
+ (
168
+ What is Spaces?
169
+ )
170
+ Spaces Toggle
171
+ TXYZ.AI
172
+ (
173
+ What is TXYZ.AI?
174
+ )
175
+ Related Papers
176
+ Recommenders and Search Tools
177
+ Link to Influence Flower
178
+ Influence Flower
179
+ (
180
+ What are Influence Flowers?
181
+ )
182
+ Core recommender toggle
183
+ CORE Recommender
184
+ (
185
+ What is CORE?
186
+ )
187
+ Author
188
+ Venue
189
+ Institution
190
+ Topic
191
+ About arXivLabs
192
+ arXivLabs: experimental projects with community collaborators
193
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
194
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
195
+ Have an idea for a project that will add value for arXiv's community?
196
+ Learn more about arXivLabs
197
+ .
198
+ Which authors of this paper are endorsers?
199
+ |
200
+ Disable MathJax
201
+ (
202
+ What is MathJax?
203
+ )
research/notes/231108105-diloco-distributed-low-communication-training-of-language-models.md ADDED
@@ -0,0 +1,208 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2311.08105] DiLoCo: Distributed Low-Communication Training of Language Models'
3
+ id: 231108105-diloco-distributed-low-communication-training-of-language-models
4
+ tags:
5
+ - deepread
6
+ created: '2026-06-10T00:30:20.411067Z'
7
+ source: https://arxiv.org/abs/2311.08105
8
+ source_domain: arxiv.org
9
+ fetched_at: '2026-06-10T00:30:20.410923Z'
10
+ fetch_provider: builtin
11
+ status: draft
12
+ type: note
13
+ tier: institutional
14
+ content_type: paper
15
+ deprecated: false
16
+ ---
17
+
18
+ [2311.08105] DiLoCo: Distributed Low-Communication Training of Language Models
19
+ Computer Science > Machine Learning
20
+ arXiv:2311.08105
21
+ (cs)
22
+ [Submitted on 14 Nov 2023 (
23
+ v1
24
+ ), last revised 23 Sep 2024 (this version, v3)]
25
+ Title:
26
+ DiLoCo: Distributed Low-Communication Training of Language Models
27
+ Authors:
28
+ Arthur Douillard
29
+ ,
30
+ Qixuan Feng
31
+ ,
32
+ Andrei A. Rusu
33
+ ,
34
+ Rachita Chhaparia
35
+ ,
36
+ Yani Donchev
37
+ ,
38
+ Adhiguna Kuncoro
39
+ ,
40
+ Marc'Aurelio Ranzato
41
+ ,
42
+ Arthur Szlam
43
+ ,
44
+ Jiajun Shen
45
+ View a PDF of the paper titled DiLoCo: Distributed Low-Communication Training of Language Models, by Arthur Douillard and 8 other authors
46
+ View PDF
47
+ HTML (experimental)
48
+ Abstract:
49
+ Large language models (LLM) have become a critical component in many applications of machine learning. However, standard approaches to training LLM require a large number of tightly interconnected accelerators, with devices exchanging gradients and other intermediate states at each optimization step. While it is difficult to build and maintain a single computing cluster hosting many accelerators, it might be easier to find several computing clusters each hosting a smaller number of devices. In this work, we propose a distributed optimization algorithm, Distributed Low-Communication (DiLoCo), that enables training of language models on islands of devices that are poorly connected. The approach is a variant of federated averaging, where the number of inner steps is large, the inner optimizer is AdamW, and the outer optimizer is Nesterov momentum. On the widely used C4 dataset, we show that DiLoCo on 8 workers performs as well as fully synchronous optimization while communicating 500 times less. DiLoCo exhibits great robustness to the data distribution of each worker. It is also robust to resources becoming unavailable over time, and vice versa, it can seamlessly leverage resources that become available during training.
50
+ Subjects:
51
+ Machine Learning (cs.LG)
52
+ ; Computation and Language (cs.CL)
53
+ Cite as:
54
+ arXiv:2311.08105
55
+ [cs.LG]
56
+ (or
57
+ arXiv:2311.08105v3
58
+ [cs.LG]
59
+ for this version)
60
+ https://doi.org/10.48550/arXiv.2311.08105
61
+ Focus to learn more
62
+ arXiv-issued DOI via DataCite
63
+ Submission history
64
+ From: Arthur Douillard [
65
+ view email
66
+ ]
67
+ [v1]
68
+ Tue, 14 Nov 2023 12:05:45 UTC (1,609 KB)
69
+ [v2]
70
+ Sat, 2 Dec 2023 14:10:14 UTC (1,610 KB)
71
+ [v3]
72
+ Mon, 23 Sep 2024 10:41:27 UTC (1,610 KB)
73
+ Full-text links:
74
+ Access Paper:
75
+ View a PDF of the paper titled DiLoCo: Distributed Low-Communication Training of Language Models, by Arthur Douillard and 8 other authors
76
+ View PDF
77
+ HTML (experimental)
78
+ TeX Source
79
+ view license
80
+ Current browse context:
81
+ cs.LG
82
+ < prev
83
+ |
84
+ next >
85
+ new
86
+ |
87
+ recent
88
+ |
89
+ 2023-11
90
+ Change to browse by:
91
+ cs
92
+ cs.CL
93
+ References & Citations
94
+ NASA ADS
95
+ Google Scholar
96
+ Semantic Scholar
97
+ export BibTeX citation
98
+ Loading...
99
+ BibTeX formatted citation
100
+ ×
101
+ loading...
102
+ Data provided by:
103
+ Bookmark
104
+ Bibliographic Tools
105
+ Bibliographic and Citation Tools
106
+ Bibliographic Explorer Toggle
107
+ Bibliographic Explorer
108
+ (
109
+ What is the Explorer?
110
+ )
111
+ Connected Papers Toggle
112
+ Connected Papers
113
+ (
114
+ What is Connected Papers?
115
+ )
116
+ Litmaps Toggle
117
+ Litmaps
118
+ (
119
+ What is Litmaps?
120
+ )
121
+ scite.ai Toggle
122
+ scite Smart Citations
123
+ (
124
+ What are Smart Citations?
125
+ )
126
+ Code, Data, Media
127
+ Code, Data and Media Associated with this Article
128
+ alphaXiv Toggle
129
+ alphaXiv
130
+ (
131
+ What is alphaXiv?
132
+ )
133
+ Links to Code Toggle
134
+ CatalyzeX Code Finder for Papers
135
+ (
136
+ What is CatalyzeX?
137
+ )
138
+ DagsHub Toggle
139
+ DagsHub
140
+ (
141
+ What is DagsHub?
142
+ )
143
+ GotitPub Toggle
144
+ Gotit.pub
145
+ (
146
+ What is GotitPub?
147
+ )
148
+ Huggingface Toggle
149
+ Hugging Face
150
+ (
151
+ What is Huggingface?
152
+ )
153
+ ScienceCast Toggle
154
+ ScienceCast
155
+ (
156
+ What is ScienceCast?
157
+ )
158
+ Demos
159
+ Demos
160
+ Replicate Toggle
161
+ Replicate
162
+ (
163
+ What is Replicate?
164
+ )
165
+ Spaces Toggle
166
+ Hugging Face Spaces
167
+ (
168
+ What is Spaces?
169
+ )
170
+ Spaces Toggle
171
+ TXYZ.AI
172
+ (
173
+ What is TXYZ.AI?
174
+ )
175
+ Related Papers
176
+ Recommenders and Search Tools
177
+ Link to Influence Flower
178
+ Influence Flower
179
+ (
180
+ What are Influence Flowers?
181
+ )
182
+ Core recommender toggle
183
+ CORE Recommender
184
+ (
185
+ What is CORE?
186
+ )
187
+ IArxiv recommender toggle
188
+ IArxiv Recommender
189
+ (
190
+ What is IArxiv?
191
+ )
192
+ Author
193
+ Venue
194
+ Institution
195
+ Topic
196
+ About arXivLabs
197
+ arXivLabs: experimental projects with community collaborators
198
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
199
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
200
+ Have an idea for a project that will add value for arXiv's community?
201
+ Learn more about arXivLabs
202
+ .
203
+ Which authors of this paper are endorsers?
204
+ |
205
+ Disable MathJax
206
+ (
207
+ What is MathJax?
208
+ )
research/notes/231108516-llms-cannot-find-reasoning-errors-but-can-correct-them-given-the-error.md ADDED
@@ -0,0 +1,204 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2311.08516] LLMs cannot find reasoning errors, but can correct them given
3
+ the error location'
4
+ id: 231108516-llms-cannot-find-reasoning-errors-but-can-correct-them-given-the-error
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:40:16.980357Z'
8
+ source: https://arxiv.org/abs/2311.08516
9
+ source_domain: arxiv.org
10
+ fetched_at: '2026-06-10T00:40:16.980220Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2311.08516] LLMs cannot find reasoning errors, but can correct them given the error location
20
+ Computer Science > Artificial Intelligence
21
+ arXiv:2311.08516
22
+ (cs)
23
+ [Submitted on 14 Nov 2023 (
24
+ v1
25
+ ), last revised 4 Jun 2024 (this version, v3)]
26
+ Title:
27
+ LLMs cannot find reasoning errors, but can correct them given the error location
28
+ Authors:
29
+ Gladys Tyen
30
+ ,
31
+ Hassan Mansoor
32
+ ,
33
+ Victor Cărbune
34
+ ,
35
+ Peter Chen
36
+ ,
37
+ Tony Mak
38
+ View a PDF of the paper titled LLMs cannot find reasoning errors, but can correct them given the error location, by Gladys Tyen and 4 other authors
39
+ View PDF
40
+ HTML (experimental)
41
+ Abstract:
42
+ While self-correction has shown promise in improving LLM outputs in terms of style and quality (e.g. Chen et al., 2023b; Madaan et al., 2023), recent attempts to self-correct logical or reasoning errors often cause correct answers to become incorrect, resulting in worse performances overall (Huang et al., 2023). In this paper, we show that poor self-correction performance stems from LLMs' inability to find logical mistakes, rather than their ability to correct a known mistake. Firstly, we benchmark several state-of-the-art LLMs on their mistake-finding ability and demonstrate that they generally struggle with the task, even in highly objective, unambiguous cases. Secondly, we test the correction abilities of LLMs -- separately from mistake finding -- using a backtracking setup that feeds ground truth mistake location information to the model. We show that this boosts downstream task performance across our 5 reasoning tasks, indicating that LLMs' correction abilities are robust. Finally, we show that it is possible to obtain mistake location information without ground truth labels or in-domain training data. We train a small classifier with out-of-domain data, which exhibits stronger mistake-finding performance than prompting a large model. We release our dataset of LLM-generated logical mistakes, BIG-Bench Mistake, to enable further research into locating LLM reasoning mistakes.
43
+ Comments:
44
+ ACL 2024 Findings
45
+ Subjects:
46
+ Artificial Intelligence (cs.AI)
47
+ ; Computation and Language (cs.CL); Machine Learning (cs.LG)
48
+ Cite as:
49
+ arXiv:2311.08516
50
+ [cs.AI]
51
+ (or
52
+ arXiv:2311.08516v3
53
+ [cs.AI]
54
+ for this version)
55
+ https://doi.org/10.48550/arXiv.2311.08516
56
+ Focus to learn more
57
+ arXiv-issued DOI via DataCite
58
+ Submission history
59
+ From: Gladys Tyen [
60
+ view email
61
+ ]
62
+ [v1]
63
+ Tue, 14 Nov 2023 20:12:38 UTC (7,191 KB)
64
+ [v2]
65
+ Tue, 9 Jan 2024 03:32:32 UTC (7,191 KB)
66
+ [v3]
67
+ Tue, 4 Jun 2024 10:25:13 UTC (7,319 KB)
68
+ Full-text links:
69
+ Access Paper:
70
+ View a PDF of the paper titled LLMs cannot find reasoning errors, but can correct them given the error location, by Gladys Tyen and 4 other authors
71
+ View PDF
72
+ HTML (experimental)
73
+ TeX Source
74
+ view license
75
+ Current browse context:
76
+ cs.AI
77
+ < prev
78
+ |
79
+ next >
80
+ new
81
+ |
82
+ recent
83
+ |
84
+ 2023-11
85
+ Change to browse by:
86
+ cs
87
+ cs.CL
88
+ cs.LG
89
+ References & Citations
90
+ NASA ADS
91
+ Google Scholar
92
+ Semantic Scholar
93
+ export BibTeX citation
94
+ Loading...
95
+ BibTeX formatted citation
96
+ ×
97
+ loading...
98
+ Data provided by:
99
+ Bookmark
100
+ Bibliographic Tools
101
+ Bibliographic and Citation Tools
102
+ Bibliographic Explorer Toggle
103
+ Bibliographic Explorer
104
+ (
105
+ What is the Explorer?
106
+ )
107
+ Connected Papers Toggle
108
+ Connected Papers
109
+ (
110
+ What is Connected Papers?
111
+ )
112
+ Litmaps Toggle
113
+ Litmaps
114
+ (
115
+ What is Litmaps?
116
+ )
117
+ scite.ai Toggle
118
+ scite Smart Citations
119
+ (
120
+ What are Smart Citations?
121
+ )
122
+ Code, Data, Media
123
+ Code, Data and Media Associated with this Article
124
+ alphaXiv Toggle
125
+ alphaXiv
126
+ (
127
+ What is alphaXiv?
128
+ )
129
+ Links to Code Toggle
130
+ CatalyzeX Code Finder for Papers
131
+ (
132
+ What is CatalyzeX?
133
+ )
134
+ DagsHub Toggle
135
+ DagsHub
136
+ (
137
+ What is DagsHub?
138
+ )
139
+ GotitPub Toggle
140
+ Gotit.pub
141
+ (
142
+ What is GotitPub?
143
+ )
144
+ Huggingface Toggle
145
+ Hugging Face
146
+ (
147
+ What is Huggingface?
148
+ )
149
+ Links to Code Toggle
150
+ Papers with Code
151
+ (
152
+ What is Papers with Code?
153
+ )
154
+ ScienceCast Toggle
155
+ ScienceCast
156
+ (
157
+ What is ScienceCast?
158
+ )
159
+ Demos
160
+ Demos
161
+ Replicate Toggle
162
+ Replicate
163
+ (
164
+ What is Replicate?
165
+ )
166
+ Spaces Toggle
167
+ Hugging Face Spaces
168
+ (
169
+ What is Spaces?
170
+ )
171
+ Spaces Toggle
172
+ TXYZ.AI
173
+ (
174
+ What is TXYZ.AI?
175
+ )
176
+ Related Papers
177
+ Recommenders and Search Tools
178
+ Link to Influence Flower
179
+ Influence Flower
180
+ (
181
+ What are Influence Flowers?
182
+ )
183
+ Core recommender toggle
184
+ CORE Recommender
185
+ (
186
+ What is CORE?
187
+ )
188
+ Author
189
+ Venue
190
+ Institution
191
+ Topic
192
+ About arXivLabs
193
+ arXivLabs: experimental projects with community collaborators
194
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
195
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
196
+ Have an idea for a project that will add value for arXiv's community?
197
+ Learn more about arXivLabs
198
+ .
199
+ Which authors of this paper are endorsers?
200
+ |
201
+ Disable MathJax
202
+ (
203
+ What is MathJax?
204
+ )
research/notes/231209152-evaluating-augmented-reality-communication-how-can-we-teach-procedural.md ADDED
@@ -0,0 +1,203 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2312.09152] Evaluating Augmented Reality Communication: How Can We Teach
3
+ Procedural Skill in AR?'
4
+ id: 231209152-evaluating-augmented-reality-communication-how-can-we-teach-procedural
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:40:10.692238Z'
8
+ source: https://arxiv.org/abs/2312.09152
9
+ source_domain: arxiv.org
10
+ fetched_at: '2026-06-10T00:40:10.692096Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2312.09152] Evaluating Augmented Reality Communication: How Can We Teach Procedural Skill in AR?
20
+ Computer Science > Human-Computer Interaction
21
+ arXiv:2312.09152
22
+ (cs)
23
+ [Submitted on 14 Dec 2023]
24
+ Title:
25
+ Evaluating Augmented Reality Communication: How Can We Teach Procedural Skill in AR?
26
+ Authors:
27
+ Manuel Rebol
28
+ ,
29
+ Krzysztof Pietroszek
30
+ ,
31
+ Neal Sikka
32
+ ,
33
+ Claudia Ranniger
34
+ ,
35
+ Colton Hood
36
+ ,
37
+ Adam Rutenberg
38
+ ,
39
+ Puja Sasankan
40
+ ,
41
+ Christian Gütl
42
+ View a PDF of the paper titled Evaluating Augmented Reality Communication: How Can We Teach Procedural Skill in AR?, by Manuel Rebol and 7 other authors
43
+ View PDF
44
+ HTML (experimental)
45
+ Abstract:
46
+ Augmented reality (AR) has great potential for use in healthcare applications, especially remote medical training and supervision. In this paper, we analyze the usage of an AR communication system to teach a medical procedure, the placement of a central venous catheter (CVC) under ultrasound guidance. We examine various AR communication and collaboration components, including gestural communication, volumetric information, annotations, augmented objects, and augmented screens. We compare how teaching in AR differs from teaching through videoconferencing-based communication. Our results include a detailed medical training steps analysis in which we compare how verbal and visual communication differs between video and AR training. We identify procedural steps in which medical experts give visual instructions utilizing AR components. We examine the change in AR usage and interaction over time and recognize patterns between users. Moreover, AR design recommendations are given based on post-training interviews.
47
+ Comments:
48
+ this https URL
49
+ Subjects:
50
+ Human-Computer Interaction (cs.HC)
51
+ Cite as:
52
+ arXiv:2312.09152
53
+ [cs.HC]
54
+ (or
55
+ arXiv:2312.09152v1
56
+ [cs.HC]
57
+ for this version)
58
+ https://doi.org/10.48550/arXiv.2312.09152
59
+ Focus to learn more
60
+ arXiv-issued DOI via DataCite
61
+ Journal reference:
62
+ Proceedings of the 29th ACM Symposium on Virtual Reality Software and Technology (VRST 2023)
63
+ Related DOI
64
+ :
65
+ https://doi.org/10.1145/3611659.3615685
66
+ Focus to learn more
67
+ DOI(s) linking to related resources
68
+ Submission history
69
+ From: Manuel Rebol [
70
+ view email
71
+ ]
72
+ [v1]
73
+ Thu, 14 Dec 2023 17:22:22 UTC (2,671 KB)
74
+ Full-text links:
75
+ Access Paper:
76
+ View a PDF of the paper titled Evaluating Augmented Reality Communication: How Can We Teach Procedural Skill in AR?, by Manuel Rebol and 7 other authors
77
+ View PDF
78
+ HTML (experimental)
79
+ TeX Source
80
+ view license
81
+ Current browse context:
82
+ cs.HC
83
+ < prev
84
+ |
85
+ next >
86
+ new
87
+ |
88
+ recent
89
+ |
90
+ 2023-12
91
+ Change to browse by:
92
+ cs
93
+ References & Citations
94
+ NASA ADS
95
+ Google Scholar
96
+ Semantic Scholar
97
+ export BibTeX citation
98
+ Loading...
99
+ BibTeX formatted citation
100
+ ×
101
+ loading...
102
+ Data provided by:
103
+ Bookmark
104
+ Bibliographic Tools
105
+ Bibliographic and Citation Tools
106
+ Bibliographic Explorer Toggle
107
+ Bibliographic Explorer
108
+ (
109
+ What is the Explorer?
110
+ )
111
+ Connected Papers Toggle
112
+ Connected Papers
113
+ (
114
+ What is Connected Papers?
115
+ )
116
+ Litmaps Toggle
117
+ Litmaps
118
+ (
119
+ What is Litmaps?
120
+ )
121
+ scite.ai Toggle
122
+ scite Smart Citations
123
+ (
124
+ What are Smart Citations?
125
+ )
126
+ Code, Data, Media
127
+ Code, Data and Media Associated with this Article
128
+ alphaXiv Toggle
129
+ alphaXiv
130
+ (
131
+ What is alphaXiv?
132
+ )
133
+ Links to Code Toggle
134
+ CatalyzeX Code Finder for Papers
135
+ (
136
+ What is CatalyzeX?
137
+ )
138
+ DagsHub Toggle
139
+ DagsHub
140
+ (
141
+ What is DagsHub?
142
+ )
143
+ GotitPub Toggle
144
+ Gotit.pub
145
+ (
146
+ What is GotitPub?
147
+ )
148
+ Huggingface Toggle
149
+ Hugging Face
150
+ (
151
+ What is Huggingface?
152
+ )
153
+ ScienceCast Toggle
154
+ ScienceCast
155
+ (
156
+ What is ScienceCast?
157
+ )
158
+ Demos
159
+ Demos
160
+ Replicate Toggle
161
+ Replicate
162
+ (
163
+ What is Replicate?
164
+ )
165
+ Spaces Toggle
166
+ Hugging Face Spaces
167
+ (
168
+ What is Spaces?
169
+ )
170
+ Spaces Toggle
171
+ TXYZ.AI
172
+ (
173
+ What is TXYZ.AI?
174
+ )
175
+ Related Papers
176
+ Recommenders and Search Tools
177
+ Link to Influence Flower
178
+ Influence Flower
179
+ (
180
+ What are Influence Flowers?
181
+ )
182
+ Core recommender toggle
183
+ CORE Recommender
184
+ (
185
+ What is CORE?
186
+ )
187
+ Author
188
+ Venue
189
+ Institution
190
+ Topic
191
+ About arXivLabs
192
+ arXivLabs: experimental projects with community collaborators
193
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
194
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
195
+ Have an idea for a project that will add value for arXiv's community?
196
+ Learn more about arXivLabs
197
+ .
198
+ Which authors of this paper are endorsers?
199
+ |
200
+ Disable MathJax
201
+ (
202
+ What is MathJax?
203
+ )
research/notes/240201817-llms-cant-plan-but-can-help-planning-in-llm-modulo-frameworks.md ADDED
@@ -0,0 +1,203 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2402.01817] LLMs Can''t Plan, But Can Help Planning in LLM-Modulo Frameworks'
3
+ id: 240201817-llms-cant-plan-but-can-help-planning-in-llm-modulo-frameworks
4
+ tags:
5
+ - deepread
6
+ created: '2026-06-10T00:40:16.134883Z'
7
+ source: https://arxiv.org/abs/2402.01817
8
+ source_domain: arxiv.org
9
+ fetched_at: '2026-06-10T00:40:16.134739Z'
10
+ fetch_provider: builtin
11
+ status: draft
12
+ type: note
13
+ tier: institutional
14
+ content_type: paper
15
+ deprecated: false
16
+ ---
17
+
18
+ [2402.01817] LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
19
+ Computer Science > Artificial Intelligence
20
+ arXiv:2402.01817
21
+ (cs)
22
+ [Submitted on 2 Feb 2024 (
23
+ v1
24
+ ), last revised 12 Jun 2024 (this version, v3)]
25
+ Title:
26
+ LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
27
+ Authors:
28
+ Subbarao Kambhampati
29
+ ,
30
+ Karthik Valmeekam
31
+ ,
32
+ Lin Guan
33
+ ,
34
+ Mudit Verma
35
+ ,
36
+ Kaya Stechly
37
+ ,
38
+ Siddhant Bhambri
39
+ ,
40
+ Lucas Saldyt
41
+ ,
42
+ Anil Murthy
43
+ View a PDF of the paper titled LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks, by Subbarao Kambhampati and 7 other authors
44
+ View PDF
45
+ HTML (experimental)
46
+ Abstract:
47
+ There is considerable confusion about the role of Large Language Models (LLMs) in planning and reasoning tasks. On one side are over-optimistic claims that LLMs can indeed do these tasks with just the right prompting or self-verification strategies. On the other side are perhaps over-pessimistic claims that all that LLMs are good for in planning/reasoning tasks are as mere translators of the problem specification from one syntactic format to another, and ship the problem off to external symbolic solvers. In this position paper, we take the view that both these extremes are misguided. We argue that auto-regressive LLMs cannot, by themselves, do planning or self-verification (which is after all a form of reasoning), and shed some light on the reasons for misunderstandings in the literature. We will also argue that LLMs should be viewed as universal approximate knowledge sources that have much more meaningful roles to play in planning/reasoning tasks beyond simple front-end/back-end format translators. We present a vision of {\bf LLM-Modulo Frameworks} that combine the strengths of LLMs with external model-based verifiers in a tighter bi-directional interaction regime. We will show how the models driving the external verifiers themselves can be acquired with the help of LLMs. We will also argue that rather than simply pipelining LLMs and symbolic components, this LLM-Modulo Framework provides a better neuro-symbolic approach that offers tighter integration between LLMs and symbolic components, and allows extending the scope of model-based planning/reasoning regimes towards more flexible knowledge, problem and preference specifications.
48
+ Subjects:
49
+ Artificial Intelligence (cs.AI)
50
+ ; Machine Learning (cs.LG)
51
+ Cite as:
52
+ arXiv:2402.01817
53
+ [cs.AI]
54
+ (or
55
+ arXiv:2402.01817v3
56
+ [cs.AI]
57
+ for this version)
58
+ https://doi.org/10.48550/arXiv.2402.01817
59
+ Focus to learn more
60
+ arXiv-issued DOI via DataCite
61
+ Journal reference:
62
+ Proceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024
63
+ Submission history
64
+ From: Subbarao Kambhampati [
65
+ view email
66
+ ]
67
+ [v1]
68
+ Fri, 2 Feb 2024 14:43:18 UTC (4,551 KB)
69
+ [v2]
70
+ Tue, 6 Feb 2024 01:29:37 UTC (4,552 KB)
71
+ [v3]
72
+ Wed, 12 Jun 2024 01:13:11 UTC (6,405 KB)
73
+ Full-text links:
74
+ Access Paper:
75
+ View a PDF of the paper titled LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks, by Subbarao Kambhampati and 7 other authors
76
+ View PDF
77
+ HTML (experimental)
78
+ TeX Source
79
+ view license
80
+ Current browse context:
81
+ cs.AI
82
+ < prev
83
+ |
84
+ next >
85
+ new
86
+ |
87
+ recent
88
+ |
89
+ 2024-02
90
+ Change to browse by:
91
+ cs
92
+ cs.LG
93
+ References & Citations
94
+ NASA ADS
95
+ Google Scholar
96
+ Semantic Scholar
97
+ export BibTeX citation
98
+ Loading...
99
+ BibTeX formatted citation
100
+ ×
101
+ loading...
102
+ Data provided by:
103
+ Bookmark
104
+ Bibliographic Tools
105
+ Bibliographic and Citation Tools
106
+ Bibliographic Explorer Toggle
107
+ Bibliographic Explorer
108
+ (
109
+ What is the Explorer?
110
+ )
111
+ Connected Papers Toggle
112
+ Connected Papers
113
+ (
114
+ What is Connected Papers?
115
+ )
116
+ Litmaps Toggle
117
+ Litmaps
118
+ (
119
+ What is Litmaps?
120
+ )
121
+ scite.ai Toggle
122
+ scite Smart Citations
123
+ (
124
+ What are Smart Citations?
125
+ )
126
+ Code, Data, Media
127
+ Code, Data and Media Associated with this Article
128
+ alphaXiv Toggle
129
+ alphaXiv
130
+ (
131
+ What is alphaXiv?
132
+ )
133
+ Links to Code Toggle
134
+ CatalyzeX Code Finder for Papers
135
+ (
136
+ What is CatalyzeX?
137
+ )
138
+ DagsHub Toggle
139
+ DagsHub
140
+ (
141
+ What is DagsHub?
142
+ )
143
+ GotitPub Toggle
144
+ Gotit.pub
145
+ (
146
+ What is GotitPub?
147
+ )
148
+ Huggingface Toggle
149
+ Hugging Face
150
+ (
151
+ What is Huggingface?
152
+ )
153
+ ScienceCast Toggle
154
+ ScienceCast
155
+ (
156
+ What is ScienceCast?
157
+ )
158
+ Demos
159
+ Demos
160
+ Replicate Toggle
161
+ Replicate
162
+ (
163
+ What is Replicate?
164
+ )
165
+ Spaces Toggle
166
+ Hugging Face Spaces
167
+ (
168
+ What is Spaces?
169
+ )
170
+ Spaces Toggle
171
+ TXYZ.AI
172
+ (
173
+ What is TXYZ.AI?
174
+ )
175
+ Related Papers
176
+ Recommenders and Search Tools
177
+ Link to Influence Flower
178
+ Influence Flower
179
+ (
180
+ What are Influence Flowers?
181
+ )
182
+ Core recommender toggle
183
+ CORE Recommender
184
+ (
185
+ What is CORE?
186
+ )
187
+ Author
188
+ Venue
189
+ Institution
190
+ Topic
191
+ About arXivLabs
192
+ arXivLabs: experimental projects with community collaborators
193
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
194
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
195
+ Have an idea for a project that will add value for arXiv's community?
196
+ Learn more about arXivLabs
197
+ .
198
+ Which authors of this paper are endorsers?
199
+ |
200
+ Disable MathJax
201
+ (
202
+ What is MathJax?
203
+ )
research/notes/240203300-deepseekmath-pushing-the-limits-of-mathematical-reasoning-in-open-lang.md ADDED
@@ -0,0 +1,214 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2402.03300] DeepSeekMath: Pushing the Limits of Mathematical Reasoning in
3
+ Open Language Models'
4
+ id: 240203300-deepseekmath-pushing-the-limits-of-mathematical-reasoning-in-open-lang
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-09T23:28:29.232007Z'
8
+ source: https://arxiv.org/abs/2402.03300
9
+ source_domain: arxiv.org
10
+ fetched_at: '2026-06-09T23:28:29.231807Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2402.03300] DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
20
+ Computer Science > Computation and Language
21
+ arXiv:2402.03300
22
+ (cs)
23
+ [Submitted on 5 Feb 2024 (
24
+ v1
25
+ ), last revised 27 Apr 2024 (this version, v3)]
26
+ Title:
27
+ DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
28
+ Authors:
29
+ Zhihong Shao
30
+ ,
31
+ Peiyi Wang
32
+ ,
33
+ Qihao Zhu
34
+ ,
35
+ Runxin Xu
36
+ ,
37
+ Junxiao Song
38
+ ,
39
+ Xiao Bi
40
+ ,
41
+ Haowei Zhang
42
+ ,
43
+ Mingchuan Zhang
44
+ ,
45
+ Y.K. Li
46
+ ,
47
+ Y. Wu
48
+ ,
49
+ Daya Guo
50
+ View a PDF of the paper titled DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, by Zhihong Shao and 10 other authors
51
+ View PDF
52
+ HTML (experimental)
53
+ Abstract:
54
+ Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pre-training DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens sourced from Common Crawl, together with natural language and code data. DeepSeekMath 7B has achieved an impressive score of 51.7% on the competition-level MATH benchmark without relying on external toolkits and voting techniques, approaching the performance level of Gemini-Ultra and GPT-4. Self-consistency over 64 samples from DeepSeekMath 7B achieves 60.9% on MATH. The mathematical reasoning capability of DeepSeekMath is attributed to two key factors: First, we harness the significant potential of publicly available web data through a meticulously engineered data selection pipeline. Second, we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.
55
+ Subjects:
56
+ Computation and Language (cs.CL)
57
+ ; Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
58
+ Cite as:
59
+ arXiv:2402.03300
60
+ [cs.CL]
61
+ (or
62
+ arXiv:2402.03300v3
63
+ [cs.CL]
64
+ for this version)
65
+ https://doi.org/10.48550/arXiv.2402.03300
66
+ Focus to learn more
67
+ arXiv-issued DOI via DataCite
68
+ Submission history
69
+ From: Zhihong Shao [
70
+ view email
71
+ ]
72
+ [v1]
73
+ Mon, 5 Feb 2024 18:55:32 UTC (3,417 KB)
74
+ [v2]
75
+ Tue, 6 Feb 2024 18:39:38 UTC (3,417 KB)
76
+ [v3]
77
+ Sat, 27 Apr 2024 15:25:53 UTC (3,417 KB)
78
+ Full-text links:
79
+ Access Paper:
80
+ View a PDF of the paper titled DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, by Zhihong Shao and 10 other authors
81
+ View PDF
82
+ HTML (experimental)
83
+ TeX Source
84
+ view license
85
+ Current browse context:
86
+ cs.CL
87
+ < prev
88
+ |
89
+ next >
90
+ new
91
+ |
92
+ recent
93
+ |
94
+ 2024-02
95
+ Change to browse by:
96
+ cs
97
+ cs.AI
98
+ cs.LG
99
+ References & Citations
100
+ NASA ADS
101
+ Google Scholar
102
+ Semantic Scholar
103
+ export BibTeX citation
104
+ Loading...
105
+ BibTeX formatted citation
106
+ ×
107
+ loading...
108
+ Data provided by:
109
+ Bookmark
110
+ Bibliographic Tools
111
+ Bibliographic and Citation Tools
112
+ Bibliographic Explorer Toggle
113
+ Bibliographic Explorer
114
+ (
115
+ What is the Explorer?
116
+ )
117
+ Connected Papers Toggle
118
+ Connected Papers
119
+ (
120
+ What is Connected Papers?
121
+ )
122
+ Litmaps Toggle
123
+ Litmaps
124
+ (
125
+ What is Litmaps?
126
+ )
127
+ scite.ai Toggle
128
+ scite Smart Citations
129
+ (
130
+ What are Smart Citations?
131
+ )
132
+ Code, Data, Media
133
+ Code, Data and Media Associated with this Article
134
+ alphaXiv Toggle
135
+ alphaXiv
136
+ (
137
+ What is alphaXiv?
138
+ )
139
+ Links to Code Toggle
140
+ CatalyzeX Code Finder for Papers
141
+ (
142
+ What is CatalyzeX?
143
+ )
144
+ DagsHub Toggle
145
+ DagsHub
146
+ (
147
+ What is DagsHub?
148
+ )
149
+ GotitPub Toggle
150
+ Gotit.pub
151
+ (
152
+ What is GotitPub?
153
+ )
154
+ Huggingface Toggle
155
+ Hugging Face
156
+ (
157
+ What is Huggingface?
158
+ )
159
+ Links to Code Toggle
160
+ Papers with Code
161
+ (
162
+ What is Papers with Code?
163
+ )
164
+ ScienceCast Toggle
165
+ ScienceCast
166
+ (
167
+ What is ScienceCast?
168
+ )
169
+ Demos
170
+ Demos
171
+ Replicate Toggle
172
+ Replicate
173
+ (
174
+ What is Replicate?
175
+ )
176
+ Spaces Toggle
177
+ Hugging Face Spaces
178
+ (
179
+ What is Spaces?
180
+ )
181
+ Spaces Toggle
182
+ TXYZ.AI
183
+ (
184
+ What is TXYZ.AI?
185
+ )
186
+ Related Papers
187
+ Recommenders and Search Tools
188
+ Link to Influence Flower
189
+ Influence Flower
190
+ (
191
+ What are Influence Flowers?
192
+ )
193
+ Core recommender toggle
194
+ CORE Recommender
195
+ (
196
+ What is CORE?
197
+ )
198
+ Author
199
+ Venue
200
+ Institution
201
+ Topic
202
+ About arXivLabs
203
+ arXivLabs: experimental projects with community collaborators
204
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
205
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
206
+ Have an idea for a project that will add value for arXiv's community?
207
+ Learn more about arXivLabs
208
+ .
209
+ Which authors of this paper are endorsers?
210
+ |
211
+ Disable MathJax
212
+ (
213
+ What is MathJax?
214
+ )
research/notes/240411018-many-shot-in-context-learning.md ADDED
@@ -0,0 +1,228 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2404.11018] Many-Shot In-Context Learning'
3
+ id: 240411018-many-shot-in-context-learning
4
+ tags:
5
+ - deepread
6
+ created: '2026-06-10T00:40:15.011649Z'
7
+ source: https://arxiv.org/abs/2404.11018
8
+ source_domain: arxiv.org
9
+ fetched_at: '2026-06-10T00:40:15.011513Z'
10
+ fetch_provider: builtin
11
+ status: draft
12
+ type: note
13
+ tier: institutional
14
+ content_type: paper
15
+ deprecated: false
16
+ ---
17
+
18
+ [2404.11018] Many-Shot In-Context Learning
19
+ Computer Science > Machine Learning
20
+ arXiv:2404.11018
21
+ (cs)
22
+ [Submitted on 17 Apr 2024 (
23
+ v1
24
+ ), last revised 17 Oct 2024 (this version, v3)]
25
+ Title:
26
+ Many-Shot In-Context Learning
27
+ Authors:
28
+ Rishabh Agarwal
29
+ ,
30
+ Avi Singh
31
+ ,
32
+ Lei M. Zhang
33
+ ,
34
+ Bernd Bohnet
35
+ ,
36
+ Luis Rosias
37
+ ,
38
+ Stephanie Chan
39
+ ,
40
+ Biao Zhang
41
+ ,
42
+ Ankesh Anand
43
+ ,
44
+ Zaheer Abbas
45
+ ,
46
+ Azade Nova
47
+ ,
48
+ John D. Co-Reyes
49
+ ,
50
+ Eric Chu
51
+ ,
52
+ Feryal Behbahani
53
+ ,
54
+ Aleksandra Faust
55
+ ,
56
+ Hugo Larochelle
57
+ View a PDF of the paper titled Many-Shot In-Context Learning, by Rishabh Agarwal and 13 other authors
58
+ View PDF
59
+ HTML (experimental)
60
+ Abstract:
61
+ Large language models (LLMs) excel at few-shot in-context learning (ICL) -- learning from a few examples provided in context at inference, without any weight updates. Newly expanded context windows allow us to investigate ICL with hundreds or thousands of examples -- the many-shot regime. Going from few-shot to many-shot, we observe significant performance gains across a wide variety of generative and discriminative tasks. While promising, many-shot ICL can be bottlenecked by the available amount of human-generated examples. To mitigate this limitation, we explore two new settings: Reinforced and Unsupervised ICL. Reinforced ICL uses model-generated chain-of-thought rationales in place of human examples. Unsupervised ICL removes rationales from the prompt altogether, and prompts the model only with domain-specific questions. We find that both Reinforced and Unsupervised ICL can be quite effective in the many-shot regime, particularly on complex reasoning tasks. Finally, we demonstrate that, unlike few-shot learning, many-shot learning is effective at overriding pretraining biases, can learn high-dimensional functions with numerical inputs, and performs comparably to fine-tuning. We also find that inference cost increases linearly in the many-shot regime, and frontier LLMs benefit from many-shot ICL to varying degrees. Our analysis also reveals the limitations of next-token prediction loss as an indicator of downstream ICL performance.
62
+ Comments:
63
+ NeurIPS (Spotlight)
64
+ Subjects:
65
+ Machine Learning (cs.LG)
66
+ ; Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
67
+ Cite as:
68
+ arXiv:2404.11018
69
+ [cs.LG]
70
+ (or
71
+ arXiv:2404.11018v3
72
+ [cs.LG]
73
+ for this version)
74
+ https://doi.org/10.48550/arXiv.2404.11018
75
+ Focus to learn more
76
+ arXiv-issued DOI via DataCite
77
+ Submission history
78
+ From: Rishabh Agarwal [
79
+ view email
80
+ ]
81
+ [v1]
82
+ Wed, 17 Apr 2024 02:49:26 UTC (327 KB)
83
+ [v2]
84
+ Wed, 22 May 2024 17:06:10 UTC (370 KB)
85
+ [v3]
86
+ Thu, 17 Oct 2024 17:45:09 UTC (414 KB)
87
+ Full-text links:
88
+ Access Paper:
89
+ View a PDF of the paper titled Many-Shot In-Context Learning, by Rishabh Agarwal and 13 other authors
90
+ View PDF
91
+ HTML (experimental)
92
+ TeX Source
93
+ view license
94
+ Current browse context:
95
+ cs.LG
96
+ < prev
97
+ |
98
+ next >
99
+ new
100
+ |
101
+ recent
102
+ |
103
+ 2024-04
104
+ Change to browse by:
105
+ cs
106
+ cs.AI
107
+ cs.CL
108
+ References & Citations
109
+ NASA ADS
110
+ Google Scholar
111
+ Semantic Scholar
112
+ export BibTeX citation
113
+ Loading...
114
+ BibTeX formatted citation
115
+ ×
116
+ loading...
117
+ Data provided by:
118
+ Bookmark
119
+ Bibliographic Tools
120
+ Bibliographic and Citation Tools
121
+ Bibliographic Explorer Toggle
122
+ Bibliographic Explorer
123
+ (
124
+ What is the Explorer?
125
+ )
126
+ Connected Papers Toggle
127
+ Connected Papers
128
+ (
129
+ What is Connected Papers?
130
+ )
131
+ Litmaps Toggle
132
+ Litmaps
133
+ (
134
+ What is Litmaps?
135
+ )
136
+ scite.ai Toggle
137
+ scite Smart Citations
138
+ (
139
+ What are Smart Citations?
140
+ )
141
+ Code, Data, Media
142
+ Code, Data and Media Associated with this Article
143
+ alphaXiv Toggle
144
+ alphaXiv
145
+ (
146
+ What is alphaXiv?
147
+ )
148
+ Links to Code Toggle
149
+ CatalyzeX Code Finder for Papers
150
+ (
151
+ What is CatalyzeX?
152
+ )
153
+ DagsHub Toggle
154
+ DagsHub
155
+ (
156
+ What is DagsHub?
157
+ )
158
+ GotitPub Toggle
159
+ Gotit.pub
160
+ (
161
+ What is GotitPub?
162
+ )
163
+ Huggingface Toggle
164
+ Hugging Face
165
+ (
166
+ What is Huggingface?
167
+ )
168
+ Links to Code Toggle
169
+ Papers with Code
170
+ (
171
+ What is Papers with Code?
172
+ )
173
+ ScienceCast Toggle
174
+ ScienceCast
175
+ (
176
+ What is ScienceCast?
177
+ )
178
+ Demos
179
+ Demos
180
+ Replicate Toggle
181
+ Replicate
182
+ (
183
+ What is Replicate?
184
+ )
185
+ Spaces Toggle
186
+ Hugging Face Spaces
187
+ (
188
+ What is Spaces?
189
+ )
190
+ Spaces Toggle
191
+ TXYZ.AI
192
+ (
193
+ What is TXYZ.AI?
194
+ )
195
+ Related Papers
196
+ Recommenders and Search Tools
197
+ Link to Influence Flower
198
+ Influence Flower
199
+ (
200
+ What are Influence Flowers?
201
+ )
202
+ Core recommender toggle
203
+ CORE Recommender
204
+ (
205
+ What is CORE?
206
+ )
207
+ IArxiv recommender toggle
208
+ IArxiv Recommender
209
+ (
210
+ What is IArxiv?
211
+ )
212
+ Author
213
+ Venue
214
+ Institution
215
+ Topic
216
+ About arXivLabs
217
+ arXivLabs: experimental projects with community collaborators
218
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
219
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
220
+ Have an idea for a project that will add value for arXiv's community?
221
+ Learn more about arXivLabs
222
+ .
223
+ Which authors of this paper are endorsers?
224
+ |
225
+ Disable MathJax
226
+ (
227
+ What is MathJax?
228
+ )
research/notes/240612543-phase-controlled-heat-modulation-with-aharonov-bohm-interferometers.md ADDED
@@ -0,0 +1,197 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2406.12543] Phase-controlled heat modulation with Aharonov-Bohm interferometers'
3
+ id: 240612543-phase-controlled-heat-modulation-with-aharonov-bohm-interferometers
4
+ tags:
5
+ - deepread
6
+ created: '2026-06-10T00:40:09.876451Z'
7
+ source: https://arxiv.org/abs/2406.12543
8
+ source_domain: arxiv.org
9
+ fetched_at: '2026-06-10T00:40:09.876309Z'
10
+ fetch_provider: builtin
11
+ status: draft
12
+ type: note
13
+ tier: institutional
14
+ content_type: paper
15
+ deprecated: false
16
+ ---
17
+
18
+ [2406.12543] Phase-controlled heat modulation with Aharonov-Bohm interferometers
19
+ Condensed Matter > Mesoscale and Nanoscale Physics
20
+ arXiv:2406.12543
21
+ (cond-mat)
22
+ [Submitted on 18 Jun 2024]
23
+ Title:
24
+ Phase-controlled heat modulation with Aharonov-Bohm interferometers
25
+ Authors:
26
+ Sun-Yong Hwang
27
+ ,
28
+ Björn Sothmann
29
+ ,
30
+ Rosa López
31
+ View a PDF of the paper titled Phase-controlled heat modulation with Aharonov-Bohm interferometers, by Sun-Yong Hwang and 2 other authors
32
+ View PDF
33
+ HTML (experimental)
34
+ Abstract:
35
+ A heat modulator is proposed based on a voltage-biased Aharonov-Bohm interferometer. Once an electrical bias is applied, Peltier effects give rise to a flow of heat that can be modulated by a magnetic flux. We determine the corresponding temperature changes using a simple thermal model. Our calculations demonstrate that the modulated temperature difference can be as large as 80 mK at base temperature about 600 mK with relative temperature variations reaching 10\%. Our model also predicts, quite generally, the emergence of spin-polarized heat flows without any ferromagnetic contacts, if Rashba spin-orbit interaction is combined with the applied magnetic flux, which potentially paves the way towards caloritronic information processing.
36
+ Comments:
37
+ 8 pages, 4 figures
38
+ Subjects:
39
+ Mesoscale and Nanoscale Physics (cond-mat.mes-hall)
40
+ Cite as:
41
+ arXiv:2406.12543
42
+ [cond-mat.mes-hall]
43
+ (or
44
+ arXiv:2406.12543v1
45
+ [cond-mat.mes-hall]
46
+ for this version)
47
+ https://doi.org/10.48550/arXiv.2406.12543
48
+ Focus to learn more
49
+ arXiv-issued DOI via DataCite
50
+ Journal reference:
51
+ Phys. Rev. Research 6, 013215 (2024)
52
+ Related DOI
53
+ :
54
+ https://doi.org/10.1103/PhysRevResearch.6.013215
55
+ Focus to learn more
56
+ DOI(s) linking to related resources
57
+ Submission history
58
+ From: Sun-Yong Hwang [
59
+ view email
60
+ ]
61
+ [v1]
62
+ Tue, 18 Jun 2024 12:22:44 UTC (1,894 KB)
63
+ Full-text links:
64
+ Access Paper:
65
+ View a PDF of the paper titled Phase-controlled heat modulation with Aharonov-Bohm interferometers, by Sun-Yong Hwang and 2 other authors
66
+ View PDF
67
+ HTML (experimental)
68
+ TeX Source
69
+ view license
70
+ Current browse context:
71
+ cond-mat.mes-hall
72
+ < prev
73
+ |
74
+ next >
75
+ new
76
+ |
77
+ recent
78
+ |
79
+ 2024-06
80
+ Change to browse by:
81
+ cond-mat
82
+ References & Citations
83
+ NASA ADS
84
+ Google Scholar
85
+ Semantic Scholar
86
+ export BibTeX citation
87
+ Loading...
88
+ BibTeX formatted citation
89
+ ×
90
+ loading...
91
+ Data provided by:
92
+ Bookmark
93
+ Bibliographic Tools
94
+ Bibliographic and Citation Tools
95
+ Bibliographic Explorer Toggle
96
+ Bibliographic Explorer
97
+ (
98
+ What is the Explorer?
99
+ )
100
+ Connected Papers Toggle
101
+ Connected Papers
102
+ (
103
+ What is Connected Papers?
104
+ )
105
+ Litmaps Toggle
106
+ Litmaps
107
+ (
108
+ What is Litmaps?
109
+ )
110
+ scite.ai Toggle
111
+ scite Smart Citations
112
+ (
113
+ What are Smart Citations?
114
+ )
115
+ Code, Data, Media
116
+ Code, Data and Media Associated with this Article
117
+ alphaXiv Toggle
118
+ alphaXiv
119
+ (
120
+ What is alphaXiv?
121
+ )
122
+ Links to Code Toggle
123
+ CatalyzeX Code Finder for Papers
124
+ (
125
+ What is CatalyzeX?
126
+ )
127
+ DagsHub Toggle
128
+ DagsHub
129
+ (
130
+ What is DagsHub?
131
+ )
132
+ GotitPub Toggle
133
+ Gotit.pub
134
+ (
135
+ What is GotitPub?
136
+ )
137
+ Huggingface Toggle
138
+ Hugging Face
139
+ (
140
+ What is Huggingface?
141
+ )
142
+ ScienceCast Toggle
143
+ ScienceCast
144
+ (
145
+ What is ScienceCast?
146
+ )
147
+ Demos
148
+ Demos
149
+ Replicate Toggle
150
+ Replicate
151
+ (
152
+ What is Replicate?
153
+ )
154
+ Spaces Toggle
155
+ Hugging Face Spaces
156
+ (
157
+ What is Spaces?
158
+ )
159
+ Spaces Toggle
160
+ TXYZ.AI
161
+ (
162
+ What is TXYZ.AI?
163
+ )
164
+ Related Papers
165
+ Recommenders and Search Tools
166
+ Link to Influence Flower
167
+ Influence Flower
168
+ (
169
+ What are Influence Flowers?
170
+ )
171
+ Core recommender toggle
172
+ CORE Recommender
173
+ (
174
+ What is CORE?
175
+ )
176
+ IArxiv recommender toggle
177
+ IArxiv Recommender
178
+ (
179
+ What is IArxiv?
180
+ )
181
+ Author
182
+ Venue
183
+ Institution
184
+ Topic
185
+ About arXivLabs
186
+ arXivLabs: experimental projects with community collaborators
187
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
188
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
189
+ Have an idea for a project that will add value for arXiv's community?
190
+ Learn more about arXivLabs
191
+ .
192
+ Which authors of this paper are endorsers?
193
+ |
194
+ Disable MathJax
195
+ (
196
+ What is MathJax?
197
+ )
research/notes/240806195-mutual-reasoning-makes-smaller-llms-stronger-problem-solvers-2.md ADDED
@@ -0,0 +1,2384 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2408.06195] Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers'
3
+ id: 240806195-mutual-reasoning-makes-smaller-llms-stronger-problem-solvers-2
4
+ tags:
5
+ - deepread
6
+ created: '2026-06-10T00:40:45.751799Z'
7
+ source: https://ar5iv.labs.arxiv.org/html/2408.06195
8
+ source_domain: ar5iv.labs.arxiv.org
9
+ fetched_at: '2026-06-10T00:40:45.751621Z'
10
+ fetch_provider: builtin
11
+ status: draft
12
+ type: note
13
+ tier: institutional
14
+ content_type: paper
15
+ deprecated: false
16
+ ---
17
+
18
+ [2408.06195] Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
19
+ Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
20
+ Zhenting Qi
21
+ ∗‡†
22
+ Mingyuan Ma
23
+ ∗‡†
24
+ Jiahang Xu
25
+ ∗‡
26
+ Li Lyna Zhang
27
+ ‡⋄
28
+ Fan Yang
29
+ ‡
30
+ Mao Yang
31
+ ‡
32
+ ‡
33
+ Microsoft Research Asia
34
+ †
35
+ Harvard University
36
+ Abstract
37
+ This paper introduces rStar, a self-play mutual reasoning approach that significantly improves reasoning capabilities of small language models (SLMs) without fine-tuning or superior models. rStar decouples reasoning into a self-play mutual generation-discrimination process. First, a target SLM augments the Monte Carlo Tree Search (MCTS) with
38
+ a rich set of human-like reasoning actions
39
+ to construct higher quality reasoning trajectories. Next, another SLM, with capabilities similar to the target SLM, acts as a discriminator to verify each trajectory generated by the target SLM. The mutually agreed reasoning trajectories are considered
40
+ mutual consistent
41
+ , thus are more likely to be correct. Extensive experiments across five SLMs demonstrate rStar can effectively solve diverse reasoning problems, including GSM8K, GSM-Hard, MATH, SVAMP, and StrategyQA. Remarkably, rStar boosts GSM8K accuracy from 12.51% to 63.91% for LLaMA2-7B, from 36.46% to 81.88% for Mistral-7B, from 74.53% to 91.13% for LLaMA3-8B-Instruct. Code will be available at
42
+ here
43
+ .
44
+ $*$
45
+ $*$
46
+ footnotetext:
47
+ Equal contribution. Zhenting Qi and Mingyuan Ma did the work during an internship at MSRA
48
+ $\diamond$
49
+ $\diamond$
50
+ footnotetext:
51
+ Corresponding author: lzhani@microsoft.com
52
+ 1
53
+ Introduction
54
+ Despite their success, large language models (LLMs) face significant challenges in complex reasoning
55
+ (Valmeekam et al.,
56
+ 2022
57
+ ; Weng et al.,
58
+ 2023
59
+ )
60
+ . For example, state of the art models like Mistral-7B
61
+ (Jiang et al.,
62
+ 2023
63
+ )
64
+ can only achieve 36.5% accuracy on the GSM8K dataset, even with techniques like Chain-of-Throught (CoT)
65
+ (Wei et al.,
66
+ 2022
67
+ )
68
+ . Although fine-tuning is shown to be an effective way to improve reasoning capability, most LLMs rely on fine-tuning data distilled or synthesized by
69
+ superior
70
+ models like GPT-4
71
+ (Wang et al.,
72
+ 2024a
73
+ ; Gou et al.,
74
+ 2023
75
+ )
76
+ . Meanwhile, the community has been actively working on a complimentary and yet more challenging approach: Reasoning improvements
77
+ without
78
+ a superior teacher LLM.
79
+ Figure 1:
80
+ With 32 rounds of inference, rStar makes SLMs highly capable problem-solvers, matching or even surpassing the reasoning performance achieved after domain-specialized SFT.
81
+ A promising paradigm to improve reasoning without superior models is to leverage the knowledge within LLMs themselves
82
+ (Wang et al.,
83
+ 2023
84
+ ; Hao et al.,
85
+ 2023
86
+ ; Madaan et al.,
87
+ 2024
88
+ )
89
+ . For example, RAP
90
+ (Hao et al.,
91
+ 2023
92
+ )
93
+ adopts a self-exploration solution to iteratively improve LLM’s reasoning performance through self-rewarded feedback. Unfortunately, study suggests that this paradigm often suffers from two fundamental issues.
94
+ First, LLMs often struggle to effectively explore the solution space during reasoning. The self-exploration often traps in a solution space with low-quality reasoning steps even after many attempts. For example, our experiments reveal that after 32 rounds of self-exploration with RAP
95
+ (Hao et al.,
96
+ 2023
97
+ )
98
+ , only 24% of the trajectories generated by LLaMA2-7B on GSM8K are correct.
99
+ Second, even the self-exploration can find high quality reasoning steps, it is difficult for SLMs to tell which reasoning steps are of higher quality or determine which final answers are correct, thus it is hard to effectively guide the self-exploration. Our study shows that a naïve reward-based self-exploration guidance can lead to results no better than random guesses (see Appendix
100
+ A.1
101
+ ).
102
+ A more troublesome fact is that the above two issues are more pronounced in the smaller version of LLMs, i.e.,
103
+ SLM
104
+ s, due to their weaker capabilities. For instance, while GPT-4 can improve by self-refining its output
105
+ (Madaan et al.,
106
+ 2024
107
+ ; Wu et al.,
108
+ 2024
109
+ ; Zhou et al.,
110
+ 2024
111
+ )
112
+ , the approaches are less effective in SLMs and may even lead to worse performance
113
+ (Forsman,
114
+ 2024
115
+ )
116
+ . This significantly hinders the adoption of neural language models.
117
+ This paper introduces
118
+ S
119
+ elf-play mu
120
+ T
121
+ u
122
+ A
123
+ l
124
+ R
125
+ easoning
126
+ (rStar), a novel approach that boosts SLMs’ reasoning capability during inference without fine-tuning or superior models. To address the aforementioned challenges, rStar decouples reasoning into a self-play mutual generation-discrimination process as illustrated in Fig.
127
+ 2
128
+ .
129
+ Specifically, rStar is unique in the following approaches. First, although relying on a conventional Monte Carlo Tree Search (MCTS) for SLMs to self-generate reasoning steps, rStar advocates
130
+ a richer set of reasoning actions
131
+ in the self-exploration. The new proposed actions simulate human reasoning behaviors given the current reasoning state, such as decomposing and searching for a specific reasoning step, proposing a new sub-question, or rephrasing the given question. This enables SLMs to generate high-quality candidate reasoning trajectories during self-exploration.
132
+ Second, to effectively guide the exploration among the generated reasoning trajectories, rStar augments the MCTS process with a new discrimination process called
133
+ mutual consistency
134
+ . In particular, rStar employs a second SLM with the similar capability, acting as a discriminator to provide unsupervised feedback on each candidate reasoning trajectory generated by MCTS. To improve the accuracy of the feedback, rStar hints the second SLM with sampled partial reasoning trajectories, asking it to complete the remaining reasoning steps. And rStar deems the mutually agreed reasoning trajectories of higher quality. Mutual consistency mirrors the common human practice in the absence of supervision, where agreement among peers (i.e., two SLMs) on derived answers suggests a higher likelihood of correctness.
135
+ As a result, mutual consistency offers more effective reasoning across diverse tasks than other approaches like self-consistency
136
+ (Wang et al.,
137
+ 2023
138
+ )
139
+ and avoids the risk of overfitting when training a reward model
140
+ (Chen et al.,
141
+ 2024a
142
+ ; Wang et al.,
143
+ 2024b
144
+ )
145
+ .
146
+ Figure 2:
147
+ Our self-play mutual reasoning is a generation-discrimination process: (1) a self-generator augments the target SLM to generate candidate reasoning trajectories using MCTS; (2) the discriminator uses another SLM to provide unsupervised feedback on each trajectory based on partial hints; (3) based on this feedback, the target SLM decides a final reasoning trajectory as the solution.
148
+ Extensive experiments across five SLMs and five diverse reasoning tasks demonstrate the effectiveness of rStar. With just 32 rounds of MCTS inference, rStar significantly enhances SLMs’ reasoning capabilities, matching or even surpassing the accuracy achieved after fine-tuning. For example, rStar boosts GSM8K accuracy from 12.51% to 63.91% for LLaMA2-7B, from 36.46% to 81.88% for Mistral, and from 47.23% to 85.52% for LLaMA3-8B. Furthermore, we conduct comprehensive experiments to verify rStar’s superiority over state-of-the-art baselines, including single-round inference techniques like few-shot CoT, multi-round prompting approaches such as self-consistency, and self-improvement techniques such as RAP, ToT, self-evaluation and self-verification.
149
+ 2
150
+ Related Work
151
+ Prompting Language Models to Reason
152
+ .
153
+ Prompting-based methods, such as Chain-of-Thought
154
+ (Wei et al.,
155
+ 2022
156
+ )
157
+ , focus on designing instructions and pipelines to enhance LLMs’ reasoning performance during inference. Recent advances include planning
158
+ (Hao et al.,
159
+ 2023
160
+ ; Ding et al.,
161
+ 2023
162
+ )
163
+ , problem decomposition
164
+ (Zhou et al.,
165
+ 2022
166
+ ; Khot et al.,
167
+ 2022
168
+ ; Hao et al.,
169
+ 2023
170
+ )
171
+ , abstraction
172
+ (Zheng et al.,
173
+ 2023
174
+ )
175
+ , programming
176
+ (Chen et al.,
177
+ 2022
178
+ ; Zhou et al.,
179
+ 2023
180
+ )
181
+ .
182
+ These methods aim to improve single-round inference performance and are orthogonal to ours.
183
+ LLM Self-improvement
184
+ . Recently, research on the self-improvement of LLMs has rapidly increased.
185
+ Fine-tuning based methods
186
+ (Chen et al.,
187
+ 2024b
188
+ ;
189
+ a
190
+ )
191
+ leverage the capabilities of a well-pretrained LLM to synthesize data and progressively enhance its performance. Advanced prompting techniques, such as self-verification
192
+ (Gero et al.,
193
+ 2023
194
+ ; Zhou et al.,
195
+ 2023
196
+ )
197
+ , and RAP
198
+ (Hao et al.,
199
+ 2023
200
+ )
201
+ , improve performance through iterative self-exploring based on self-diagnosed feedback at inference time.
202
+ However, as illustrated in previous section, the achieved performance often depend on the LLM’s inherent capabilities, and for SLMs, their weaker instruction-following ability and unreliable self-rewarding can mislead self-improvement.
203
+ Sampling Reasoning Paths
204
+ . Recent works
205
+ (Brown et al.,
206
+ 2024
207
+ ; Li et al.,
208
+ 2024
209
+ ; Snell et al.,
210
+ 2024
211
+ )
212
+ on mathematical reasoning have shown that sampling diverse reasoning paths can significantly enhance performance compared to greedy one-time decoding. Self-Consistency
213
+ (Wang et al.,
214
+ 2023
215
+ )
216
+ sample a complete CoT path each time. Tree-search approaches
217
+ (Yao et al.,
218
+ 2024
219
+ ; Hao et al.,
220
+ 2023
221
+ ; Zhang et al.,
222
+ 2024
223
+ )
224
+ , like MCTS, further improve the performance by breaking down tasks and sampling simpler, individual intermediate reasoning steps. However, most approaches have limited action spaces. For example, RAP
225
+ (Hao et al.,
226
+ 2023
227
+ )
228
+ decomposes only subproblems, while AlphaMath
229
+ (Chen et al.,
230
+ 2024a
231
+ )
232
+ searches only for one CoT step, limiting effectiveness in generating better trajectories.
233
+ Answer Verification
234
+ . To select correct reasoning trajectories, majority voting
235
+ (Wang et al.,
236
+ 2023
237
+ )
238
+ is a widely-used approach. To improve accuracy, some works train value or rewards model for verification
239
+ (Wang et al.,
240
+ 2024b
241
+ ; Chen et al.,
242
+ 2024a
243
+ )
244
+ , but these require additional annotations and have risks in overfitting to specific tasks. Self-verification
245
+ (Weng et al.,
246
+ 2023
247
+ )
248
+ leverages LLM capabilities for backward self-verification. Nevertheless, its effectiveness hinges on its inherent ability to reason effectively. Recent studies have shown that LLM struggles to evaluate itself and rectify its initial responses without any external feedbacks
249
+ (Huang et al.,
250
+ 2023
251
+ ; Feng et al.,
252
+ 2023
253
+ )
254
+ .
255
+ 3
256
+ Methodology
257
+ 3.1
258
+ Overview
259
+ Problem Formulation
260
+ . To solve a reasoning problem by SLMs, we formulate it as a multi-step reasoning generation task, which breaks
261
+ the problem into simpler sub-tasks. This is more effective than traditional CoT-based reasoning
262
+ (Wei et al.,
263
+ 2022
264
+ ; Wang et al.,
265
+ 2023
266
+ )
267
+ , as it is much easier for SLMs to correctly generate one step than complete reasoning steps in a single inference. We leverage the Monte-Carlo Tree Search (MCTS) algorithm
268
+ (Kocsis & Szepesvári,
269
+ 2006
270
+ )
271
+ to augment the target SLM for self-generating multi-step reasoning solutions.
272
+ Formally, for a given problem
273
+ x
274
+ 𝑥
275
+ x
276
+ and a target SLM
277
+ M
278
+ 𝑀
279
+ M
280
+ , the MCTS augments
281
+ M
282
+ 𝑀
283
+ M
284
+ to incrementally build a search tree
285
+ 𝒯
286
+ 𝒯
287
+ \mathcal{T}
288
+ . As illustrated in Fig.
289
+ 3
290
+ , the root node represents the question
291
+ x
292
+ 𝑥
293
+ x
294
+ , an edge represents an action
295
+ a
296
+ 𝑎
297
+ a
298
+ , each child node is an intermediate step
299
+ s
300
+ 𝑠
301
+ s
302
+ generated by
303
+ M
304
+ 𝑀
305
+ M
306
+ under the corresponding action. A path from the root node to a leaf node (denoted as
307
+ s
308
+ d
309
+ subscript
310
+ 𝑠
311
+ 𝑑
312
+ s_{d}
313
+ , also called a terminal node) constitutes a candidate solution trajectory
314
+ 𝐭
315
+ =
316
+ x
317
+ ⊕
318
+ s
319
+ 1
320
+ ⊕
321
+ s
322
+ 2
323
+ ⊕
324
+ …
325
+ ⊕
326
+ s
327
+ d
328
+ 𝐭
329
+ direct-sum
330
+ 𝑥
331
+ subscript
332
+ 𝑠
333
+ 1
334
+ subscript
335
+ 𝑠
336
+ 2
337
+ …
338
+ subscript
339
+ 𝑠
340
+ 𝑑
341
+ \mathbf{t}=x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{d}
342
+ . From the search tree
343
+ 𝒯
344
+ 𝒯
345
+ \mathcal{T}
346
+ , we can extract a set of solution trajectories
347
+ 𝕋
348
+ =
349
+ {
350
+ 𝐭
351
+ 1
352
+ ,
353
+ 𝐭
354
+ 2
355
+ ,
356
+ …
357
+ ,
358
+ 𝐭
359
+ n
360
+ }
361
+ ​
362
+ (
363
+ n
364
+ ≥
365
+ 1
366
+ )
367
+ 𝕋
368
+ superscript
369
+ 𝐭
370
+ 1
371
+ superscript
372
+ 𝐭
373
+ 2
374
+ …
375
+ superscript
376
+ 𝐭
377
+ 𝑛
378
+ 𝑛
379
+ 1
380
+ \mathbb{T}=\{\mathbf{t}^{1},\mathbf{t}^{2},...,\mathbf{t}^{n}\}(n\geq 1)
381
+ . Our goal is to find the trajectories that can achieve the correct answer for the given question.
382
+ Challenges in SLM Self-Improvement
383
+ . MCTS allows an SLM to explore and evaluate multiple potential solutions. Ideally, by balancing exploration of new possibilities with the exploitation of high-reward actions, the SLM can gradually refine its reasoning steps to generate a final correct reasoning trajectory. However, due to the limited capabilities in SLMs, traditional MCTS yields minimal improvement. First, the vast solution space makes it challenging for SLMs to generate effective solutions. Existing MCTS-based methods
384
+ (Hao et al.,
385
+ 2023
386
+ ; Kang et al.,
387
+ 2024
388
+ )
389
+ that use single actions limit diversity and struggle to generalize across tasks. Approaches like self-consistency
390
+ (Wang et al.,
391
+ 2023
392
+ )
393
+ use random sampling ensure diversity, SLMs often produce poor-quality solutions, requiring many attempts to find a correct solution, thereby increasing inference costs.
394
+ Second, it’s challenging to accurately reward each action. Without ground truth labels, it’s difficult to verify the correctness for each intermediate step
395
+ s
396
+ i
397
+ subscript
398
+ 𝑠
399
+ 𝑖
400
+ s_{i}
401
+ and the final answer in
402
+ s
403
+ d
404
+ subscript
405
+ 𝑠
406
+ 𝑑
407
+ s_{d}
408
+ . Majority voting in self-consistency requires most traces to be correct, which is often not the case for SLMs. Methods like RAP
409
+ (Hao et al.,
410
+ 2023
411
+ )
412
+ use self-rewarding, but our study shows SLMs perform near-random self-rewarding (Appendix
413
+ A.1
414
+ ). Training a reward model, as in M
415
+ ∗
416
+ (Kang et al.,
417
+ 2024
418
+ )
419
+ , can address this challenge but faces difficulties in collecting training data and generalizing across various tasks.
420
+ Overview
421
+ .
422
+ To address these challenges, this section introduces our methodology, rStar, which decomposes reasoning into solution generation and mutual verification in Fig.
423
+ 2
424
+ . To tackle the first challenge, we introduce a richer set of human-like reasoning actions that allows for thorough space exploration across diverse reasoning tasks. To address the second challenge, we design an SLM-tailored reward function to evaluate intermediate steps, avoiding reliance on their often unreliable self-evaluations. Moreover, we use another SLM as a discriminator to augment the MCTS process, mutually verifying the correctness of each trajectory with the generator SLM.
425
+ 3.2
426
+ Self-generating Reasoning Trajectory with MCTS Rollout
427
+ Figure 3:
428
+ An example to illustrate the process of self-generator. Highlighted nodes from top to bottom constitute a complete reasoning trace.
429
+ Given a question, MCTS augments the target SLM to explore a rich, human-like reasoning action space and generate the next steps based on the current state.
430
+ A Rich Set of Human-like Reasoning Actions
431
+ . At the core of MCTS generation lies the action space, which defines the scope of tree exploration. Most MCTS-based methods use a single action type to build the tree. For instance, in RAP, the action is to propose the next sub-question, whereas in AlphaMath
432
+ (Chen et al.,
433
+ 2024a
434
+ )
435
+ and MindStar
436
+ (Kang et al.,
437
+ 2024
438
+ )
439
+ , the action is to generate the next reasoning step.
440
+ However, relying on a single action type can easily lead to ineffective space exploration.
441
+ To address this, we revisit how humans approach reasoning.
442
+ Different people solve problems using diverse actions: some break into sub-questions, others solve it directly, and some might rephrase the problem to focus on key conditions. Moreover, people adjust their approach based on current states, choosing different actions as needed. Inspired by this human reasoning process, we introduce a richer set of 5 actions to maximize the SLM’s potential for correctly solving complex reasoning problems.
443
+ ⋄
444
+ ⋄
445
+ \diamond
446
+ A1
447
+ : Propose an one-step thought
448
+ . This action prompts the LLM to generate the next one-step thought for a given question, by considering the existing reasoning steps. Unlike the CoT, which generates complete thoughts, this approach simplifies the reasoning process and allows the LLM to perform better decision making
449
+ (Yao et al.,
450
+ 2024
451
+ ; Besta et al.,
452
+ 2024
453
+ )
454
+ .
455
+ ⋄
456
+ ⋄
457
+ \diamond
458
+ A2
459
+ : Propose the remaining thought steps.
460
+ Instead of generating only one step thought per state, this action aligns with standard CoT, enabling “fast thinking” to solve simple question in fewer steps. Given the already generated reasoning steps, it prompts the LLM to directly produce the remaining steps until reaching the final answer.
461
+ ⋄
462
+ ⋄
463
+ \diamond
464
+ A3
465
+ : Propose next sub-question along with its answer.
466
+ This action is inspired by
467
+ least-to-most prompting
468
+ (Zhou et al.,
469
+ 2022
470
+ )
471
+ , which breaks down a complex problem into a series of simpler sub-questions and solves them sequentially. Following RAP’s implementation, we prompt the LLM to ask and then answer the next sub-question.
472
+ ⋄
473
+ ⋄
474
+ \diamond
475
+ A4
476
+ : Answer the sub-question again.
477
+ Considering that a sub-question might not be answered correctly by
478
+ A3
479
+ , we propose this action to re-answer it. To improve accuracy, this action prompts the LLM to use few-shot CoT. Note that the original answer generated by
480
+ A3
481
+ did not use a CoT-like prompt but instead followed the least-to-most problem decomposition prompt
482
+ (Zhou et al.,
483
+ 2022
484
+ )
485
+ .
486
+ ⋄
487
+ ⋄
488
+ \diamond
489
+ A5
490
+ : Rephrase the question/sub-question.
491
+ When analyzing incorrect cases, we found that many of them are due the LLM misunderstanding the question. For example, it might miss a specific condition provided in the question. Therefore, we propose a new action to rephrase the question more simply. Specifically, we prompt the LLM to clearly list all conditions given in the problem statement.
492
+ Table 1:
493
+ Ablation study on the effectiveness of our rich action space: we evaluate LLaMA3-8B on 200 sampled GSM8K questions.
494
+ Action Space
495
+ Accuracy
496
+ A
497
+ 3
498
+ subscript
499
+ 𝐴
500
+ 3
501
+ A_{3}
502
+ (i.e., RAP)
503
+ 70.5
504
+ A
505
+ 3
506
+ subscript
507
+ 𝐴
508
+ 3
509
+ A_{3}
510
+ +
511
+ A
512
+ 5
513
+ subscript
514
+ 𝐴
515
+ 5
516
+ A_{5}
517
+ 72.5
518
+ A
519
+ 3
520
+ subscript
521
+ 𝐴
522
+ 3
523
+ A_{3}
524
+ +
525
+ A
526
+ 4
527
+ subscript
528
+ 𝐴
529
+ 4
530
+ A_{4}
531
+ +
532
+ A
533
+ 5
534
+ subscript
535
+ 𝐴
536
+ 5
537
+ A_{5}
538
+ 73.5
539
+ A
540
+ 2
541
+ subscript
542
+ 𝐴
543
+ 2
544
+ A_{2}
545
+ +
546
+ A
547
+ 3
548
+ subscript
549
+ 𝐴
550
+ 3
551
+ A_{3}
552
+ +
553
+ A
554
+ 4
555
+ subscript
556
+ 𝐴
557
+ 4
558
+ A_{4}
559
+ +
560
+ A
561
+ 5
562
+ subscript
563
+ 𝐴
564
+ 5
565
+ A_{5}
566
+ 74.0
567
+ All (
568
+ A
569
+ 1
570
+ subscript
571
+ 𝐴
572
+ 1
573
+ A_{1}
574
+ +
575
+ A
576
+ 2
577
+ subscript
578
+ 𝐴
579
+ 2
580
+ A_{2}
581
+ +
582
+ A
583
+ 3
584
+ subscript
585
+ 𝐴
586
+ 3
587
+ A_{3}
588
+ +
589
+ A
590
+ 4
591
+ subscript
592
+ 𝐴
593
+ 4
594
+ A_{4}
595
+ +
596
+ A
597
+ 5
598
+ subscript
599
+ 𝐴
600
+ 5
601
+ A_{5}
602
+ )
603
+ 75.0
604
+ The above 5 actions define a highly diverse action space
605
+ {
606
+ A
607
+ 1
608
+ ,
609
+ A
610
+ 2
611
+ ,
612
+ A
613
+ 3
614
+ ,
615
+ A
616
+ 4
617
+ ,
618
+ A
619
+ 5
620
+ }
621
+ subscript
622
+ 𝐴
623
+ 1
624
+ subscript
625
+ 𝐴
626
+ 2
627
+ subscript
628
+ 𝐴
629
+ 3
630
+ subscript
631
+ 𝐴
632
+ 4
633
+ subscript
634
+ 𝐴
635
+ 5
636
+ \{A_{1},A_{2},A_{3},A_{4},A_{5}\}
637
+ .
638
+ At each step
639
+ i
640
+ 𝑖
641
+ i
642
+ , MCTS selects an action
643
+ a
644
+ i
645
+ subscript
646
+ 𝑎
647
+ 𝑖
648
+ a_{i}
649
+ from this space. We then use this action
650
+ a
651
+ i
652
+ subscript
653
+ 𝑎
654
+ 𝑖
655
+ a_{i}
656
+ to prompt the LLM to generate the next reasoning step
657
+ s
658
+ i
659
+ subscript
660
+ 𝑠
661
+ 𝑖
662
+ s_{i}
663
+ , based on the current state, which is the previous generated trajectory
664
+ x
665
+ ⊕
666
+ s
667
+ 1
668
+ ⊕
669
+ s
670
+ 2
671
+ ⊕
672
+ …
673
+ ⊕
674
+ s
675
+ i
676
+ −
677
+ 1
678
+ direct-sum
679
+ 𝑥
680
+ subscript
681
+ 𝑠
682
+ 1
683
+ subscript
684
+ 𝑠
685
+ 2
686
+ …
687
+ subscript
688
+ 𝑠
689
+ 𝑖
690
+ 1
691
+ x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{i-1}
692
+ . Note that certain actions require orders. For example,
693
+ A4
694
+ can only happen after
695
+ A3
696
+ , and
697
+ A5
698
+ can only happen after the root question. As shown in Table
699
+ 1
700
+ , each action plays a crucial role in improving the final reasoning accuracy.
701
+ Reward Function.
702
+ Another critical component in MCTS is the reward function, which evaluates the value of each action and directs the tree expansion.
703
+ We design a simple yet effective reward function for SLMs. First, we exclude self-rewarding techniques for any intermediate nodes due to the limited capabilities of SLMs. Second, to ensure generalization across different reasoning tasks, we avoid introducing external supervision (e.g., tools or trained value models). Our approach draws inspiration from AlphaGo
704
+ (Silver et al.,
705
+ 2017
706
+ )
707
+ , where we score each intermediate node based on its contribution to the final correct answer. Consequently, actions that frequently lead to correct answers receive higher rewards, making them more likely to be selected in future MCTS tree expansions.
708
+ We define
709
+ Q
710
+ ​
711
+ (
712
+ s
713
+ ,
714
+ a
715
+ )
716
+ 𝑄
717
+ 𝑠
718
+ 𝑎
719
+ Q(s,a)
720
+ as the reward value for node
721
+ s
722
+ 𝑠
723
+ s
724
+ generated under action
725
+ a
726
+ 𝑎
727
+ a
728
+ .
729
+ Initially, all unexplored nodes are assigned
730
+ Q
731
+ ​
732
+ (
733
+ s
734
+ i
735
+ ,
736
+ a
737
+ i
738
+ )
739
+ =
740
+ 0
741
+ 𝑄
742
+ subscript
743
+ 𝑠
744
+ 𝑖
745
+ subscript
746
+ 𝑎
747
+ 𝑖
748
+ 0
749
+ Q(s_{i},a_{i})=0
750
+ , leading to random tree expansions. Upon reaching the first terminal node
751
+ n
752
+ d
753
+ subscript
754
+ 𝑛
755
+ 𝑑
756
+ n_{d}
757
+ , we compute a reward score
758
+ Q
759
+ ​
760
+ (
761
+ s
762
+ d
763
+ ,
764
+ a
765
+ d
766
+ )
767
+ 𝑄
768
+ subscript
769
+ 𝑠
770
+ 𝑑
771
+ subscript
772
+ 𝑎
773
+ 𝑑
774
+ Q(s_{d},a_{d})
775
+ based on whether it reaches the correct answer.
776
+ This score is then back-propagated to each intermediate node along the trajectory
777
+ 𝐭
778
+ =
779
+ x
780
+ ⊕
781
+ s
782
+ 1
783
+ ⊕
784
+ s
785
+ 2
786
+ ⊕
787
+ …
788
+ ⊕
789
+ s
790
+ d
791
+ 𝐭
792
+ direct-sum
793
+ 𝑥
794
+ subscript
795
+ 𝑠
796
+ 1
797
+ subscript
798
+ 𝑠
799
+ 2
800
+ …
801
+ subscript
802
+ 𝑠
803
+ 𝑑
804
+ \mathbf{t}=x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{d}
805
+ . Specifically, for each
806
+ s
807
+ i
808
+ subscript
809
+ 𝑠
810
+ 𝑖
811
+ s_{i}
812
+ (for
813
+ i
814
+ =
815
+ 1
816
+ ,
817
+ 2
818
+ ,
819
+ …
820
+ ,
821
+ d
822
+ −
823
+ 1
824
+ 𝑖
825
+ 1
826
+ 2
827
+ …
828
+ 𝑑
829
+ 1
830
+ i=1,2,...,d-1
831
+ ), its
832
+ Q
833
+ 𝑄
834
+ Q
835
+ value is updated as follows:
836
+ Q
837
+ ​
838
+ (
839
+ s
840
+ i
841
+ ,
842
+ a
843
+ i
844
+ )
845
+ =
846
+ Q
847
+ ​
848
+ (
849
+ s
850
+ i
851
+ ,
852
+ a
853
+ i
854
+ )
855
+ +
856
+ Q
857
+ ​
858
+ (
859
+ s
860
+ d
861
+ ,
862
+ a
863
+ d
864
+ )
865
+ 𝑄
866
+ subscript
867
+ 𝑠
868
+ 𝑖
869
+ subscript
870
+ 𝑎
871
+ 𝑖
872
+ 𝑄
873
+ subscript
874
+ 𝑠
875
+ 𝑖
876
+ subscript
877
+ 𝑎
878
+ 𝑖
879
+ 𝑄
880
+ subscript
881
+ 𝑠
882
+ 𝑑
883
+ subscript
884
+ 𝑎
885
+ 𝑑
886
+ Q(s_{i},a_{i})=Q(s_{i},a_{i})+Q(s_{d},a_{d})
887
+ . To compute the
888
+ Q
889
+ ​
890
+ (
891
+ s
892
+ d
893
+ ,
894
+ a
895
+ d
896
+ )
897
+ 𝑄
898
+ subscript
899
+ 𝑠
900
+ 𝑑
901
+ subscript
902
+ 𝑎
903
+ 𝑑
904
+ Q(s_{d},a_{d})
905
+ for the terminal node, we use the likelihood (confidence) of self-consistency majority voting as the reward value.
906
+ Figure 4:
907
+ The prompt example for mutual reasoning consistency.
908
+ Solution Generation with MCTS Rollout
909
+ . We now describe how our MCTS generates candidate reasoning trajectories. Starting from the initial root node
910
+ s
911
+ 0
912
+ subscript
913
+ 𝑠
914
+ 0
915
+ s_{0}
916
+ , we perform multiple searches consisting of
917
+ selection
918
+ ,
919
+ expansion
920
+ ,
921
+ simulations
922
+ and
923
+ back-propagation
924
+ . Specifically, the simulation is performed using the default
925
+ rollout
926
+ policy, and to achieve more accurate reward estimation, we perform multiple rollouts. To balance the exploration and exploitation, we use the well-known Upper Confidence Bounds applied to Trees (UCT)
927
+ (Kocsis & Szepesvári,
928
+ 2006
929
+ )
930
+ to select each node. This selection process is mathematically represented as:
931
+ UCT
932
+ ​
933
+ (
934
+ s
935
+ ,
936
+ a
937
+ )
938
+ =
939
+ Q
940
+ ​
941
+ (
942
+ s
943
+ ,
944
+ a
945
+ )
946
+ N
947
+ ​
948
+ (
949
+ s
950
+ ,
951
+ a
952
+ )
953
+ +
954
+ c
955
+ ​
956
+ ln
957
+ ⁡
958
+ N
959
+ p
960
+ ​
961
+ a
962
+ ​
963
+ r
964
+ ​
965
+ e
966
+ ​
967
+ n
968
+ ​
969
+ t
970
+ ​
971
+ (
972
+ s
973
+ )
974
+ N
975
+ ​
976
+ (
977
+ s
978
+ ,
979
+ a
980
+ )
981
+ .
982
+ UCT
983
+ 𝑠
984
+ 𝑎
985
+ 𝑄
986
+ 𝑠
987
+ 𝑎
988
+ 𝑁
989
+ 𝑠
990
+ 𝑎
991
+ 𝑐
992
+ subscript
993
+ 𝑁
994
+ 𝑝
995
+ 𝑎
996
+ 𝑟
997
+ 𝑒
998
+ 𝑛
999
+ 𝑡
1000
+ 𝑠
1001
+ 𝑁
1002
+ 𝑠
1003
+ 𝑎
1004
+ \text{UCT}(s,a)=\frac{Q(s,a)}{N(s,a)}+c\sqrt{\frac{\ln N_{parent}(s)}{N(s,a)}}.
1005
+ where
1006
+ N
1007
+ ​
1008
+ (
1009
+ s
1010
+ ,
1011
+ a
1012
+ )
1013
+ 𝑁
1014
+ 𝑠
1015
+ 𝑎
1016
+ N(s,a)
1017
+ is the number of times node
1018
+ s
1019
+ 𝑠
1020
+ s
1021
+ has been visited in previous iterations, and
1022
+ N
1023
+ p
1024
+ ​
1025
+ a
1026
+ ​
1027
+ r
1028
+ ​
1029
+ e
1030
+ ​
1031
+ n
1032
+ ​
1033
+ t
1034
+ ​
1035
+ (
1036
+ s
1037
+ )
1038
+ subscript
1039
+ 𝑁
1040
+ 𝑝
1041
+ 𝑎
1042
+ 𝑟
1043
+ 𝑒
1044
+ 𝑛
1045
+ 𝑡
1046
+ 𝑠
1047
+ N_{parent}(s)
1048
+ represents the visiting count of the parent node of
1049
+ s
1050
+ 𝑠
1051
+ s
1052
+ .
1053
+ Q
1054
+ ​
1055
+ (
1056
+ s
1057
+ ,
1058
+ a
1059
+ )
1060
+ 𝑄
1061
+ 𝑠
1062
+ 𝑎
1063
+ Q(s,a)
1064
+ is the estimated reward value and will be updated through back-propagation.
1065
+ c
1066
+ 𝑐
1067
+ c
1068
+ is a constant that balances exploitation and exploration.
1069
+ Once the search reaches a terminal node, either a terminal state or a predetermined maximum tree depth
1070
+ d
1071
+ 𝑑
1072
+ d
1073
+ , we obtain a trajectory from the root to terminal node. We collect all trajectories from the rollout iterations as candidate solutions. The next section explains how we verify each of them.
1074
+ 3.3
1075
+ Reasoning Trajectory Selection with Mutual Consistency
1076
+ In traditional MCTS, typically only one trajectory is selected as the final solution based on a specific metric, such as choosing the path with the highest reward from the rollout iterations. Unfortunately, after trying various existing methods, we found it challenging to define a single metric that reliably selects the trajectory containing the correct answer.
1077
+ Therefore, we collect all trajectories and propose mutual reasoning consistency for answer selection.
1078
+ Mutual Reasoning Consistency by Discriminator SLM
1079
+ 2
1080
+ . As shown in Fig.
1081
+ 2
1082
+ , in addition to the target SLM
1083
+ M
1084
+ 𝑀
1085
+ M
1086
+ , we introduce another SLM
1087
+ M
1088
+ ^
1089
+ ^
1090
+ 𝑀
1091
+ \hat{M}
1092
+ to serve as a discriminator, providing external unsupervised feedback for each candidate trajectory.
1093
+ Specifically, for
1094
+ 𝐭
1095
+ =
1096
+ x
1097
+ ⊕
1098
+ s
1099
+ 1
1100
+ ⊕
1101
+ s
1102
+ 2
1103
+ ⊕
1104
+ …
1105
+ ⊕
1106
+ s
1107
+ d
1108
+ 𝐭
1109
+ direct-sum
1110
+ 𝑥
1111
+ subscript
1112
+ 𝑠
1113
+ 1
1114
+ subscript
1115
+ 𝑠
1116
+ 2
1117
+ …
1118
+ subscript
1119
+ 𝑠
1120
+ 𝑑
1121
+ \mathbf{t}=x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{d}
1122
+ , we mask the reasoning steps starting from a randomly sampled step
1123
+ i
1124
+ 𝑖
1125
+ i
1126
+ (
1127
+ i
1128
+ <
1129
+ d
1130
+ 𝑖
1131
+ 𝑑
1132
+ i<d
1133
+ ). We then provide the earlier reasoning trajectory
1134
+ 𝐭
1135
+ =
1136
+ x
1137
+ ⊕
1138
+ s
1139
+ 1
1140
+ ⊕
1141
+ s
1142
+ 2
1143
+ ⊕
1144
+ …
1145
+ ⊕
1146
+ s
1147
+ i
1148
+ −
1149
+ 1
1150
+ 𝐭
1151
+ direct-sum
1152
+ 𝑥
1153
+ subscript
1154
+ 𝑠
1155
+ 1
1156
+ subscript
1157
+ 𝑠
1158
+ 2
1159
+ …
1160
+ subscript
1161
+ 𝑠
1162
+ 𝑖
1163
+ 1
1164
+ \mathbf{t}=x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{i-1}
1165
+ as a prompt to
1166
+ M
1167
+ ^
1168
+ ^
1169
+ 𝑀
1170
+ \hat{M}
1171
+ to complete the remaining steps for the question. Due to the provision of the earlier
1172
+ i
1173
+ −
1174
+ 1
1175
+ 𝑖
1176
+ 1
1177
+ i-1
1178
+ reasoning steps as a hint, we reduce the difficulty, thereby increasing the likelihood that SLM
1179
+ M
1180
+ ^
1181
+ ^
1182
+ 𝑀
1183
+ \hat{M}
1184
+ can provide the correct answer.
1185
+ As shown in Fig.
1186
+ 4
1187
+ , we compare whether the answer completed by
1188
+ M
1189
+ ^
1190
+ ^
1191
+ 𝑀
1192
+ \hat{M}
1193
+ matches the original trajectory
1194
+ 𝐭
1195
+ 𝐭
1196
+ \mathbf{t}
1197
+ . If they are consistent, we consider
1198
+ t
1199
+ 𝑡
1200
+ t
1201
+ as an validate trajectory for final selection.
1202
+ We provide an intuitive explanation to illustrate the rational behind our approach. Consider students solving a problem without a teacher’s feedback. A student (SLM
1203
+ 1
1204
+ ) unsure of their solution might ask a peer (SLM
1205
+ 2
1206
+ ) to review their reasoning. If the peer, given the same initial steps, arrives at the same answer, the student gains confidence in their solution. This peer verification process reflects the mutual reasoning consistency we aim to achieve.
1207
+ Final Trajectory Selection by SLM
1208
+ 1
1209
+ . After applying mutual reasoning consistency to all candidate trajectories, we return to the target SLM
1210
+ M
1211
+ 𝑀
1212
+ M
1213
+ to select the final trajectory from the validated ones. We compute each trajectory’s final score by multiplying its reward with the terminal node’s confidence score achieved from rollouts. The trajectory with the highest final score is chosen as the solution.
1214
+ Table 2:
1215
+ rStar greatly improves reasoning accuracy across various SLMs and tasks. rStar (generator@maj): uses majority voting for answer verification to show the MCTS generator’s effectiveness.
1216
+ Method
1217
+ LLaMA2-7B
1218
+ Mistral-7B
1219
+ LLaMA3-8B
1220
+ LLaMA3-8B-Instruct
1221
+ Phi3-mini-4k
1222
+ GSM8K
1223
+ Zero-shot CoT
1224
+ 1.44
1225
+ 17.89
1226
+ 22.66
1227
+ 68.38
1228
+ 20.17
1229
+ Few-shot CoT
1230
+ 12.51
1231
+ 36.46
1232
+ 47.23
1233
+ 74.53
1234
+ 83.45
1235
+ SC@maj8
1236
+ 15.31
1237
+ 42.91
1238
+ 54.21
1239
+ 78.39
1240
+ 86.35
1241
+ SC@maj64
1242
+ 20.77
1243
+ 52.84
1244
+ 64.37
1245
+ 83.24
1246
+ 88.02
1247
+ SC@maj128
1248
+ 23.05
1249
+ 57.25
1250
+ 67.55
1251
+ 84.69
1252
+ 88.68
1253
+ ToT
1254
+ 12.96
1255
+ 38.89
1256
+ 36.01
1257
+ 69.07
1258
+ 79.68
1259
+ RAP
1260
+ 24.34
1261
+ 56.25
1262
+ 57.99
1263
+ 80.59
1264
+ 81.88
1265
+ rStar (generator @maj)
1266
+ 27.22
1267
+ 64.59
1268
+ 74.38
1269
+ 88.70
1270
+ 90.44
1271
+ rStar
1272
+ 63.91
1273
+ 81.88
1274
+ 85.52
1275
+ 91.13
1276
+ 90.67
1277
+ GSM-Hard
1278
+ Zero-shot CoT
1279
+ 0.83
1280
+ 5.16
1281
+ 6.44
1282
+ 14.94
1283
+ 33.73
1284
+ Few-shot CoT
1285
+ 3.71
1286
+ 13.57
1287
+ 13.80
1288
+ 25.63
1289
+ 40.63
1290
+ SC@maj8
1291
+ 4.39
1292
+ 17.36
1293
+ 18.20
1294
+ 28.51
1295
+ 42.00
1296
+ SC@maj64
1297
+ 6.52
1298
+ 22.59
1299
+ 23.73
1300
+ 30.33
1301
+ 44.80
1302
+ SC@maj128
1303
+ 6.89
1304
+ 25.01
1305
+ 25.47
1306
+ 31.16
1307
+ 45.56
1308
+ ToT
1309
+ 2.35
1310
+ 11.47
1311
+ 10.61
1312
+ 19.64
1313
+ 32.68
1314
+ RAP
1315
+ 7.28
1316
+ 22.52
1317
+ 18.95
1318
+ 29.64
1319
+ 40.94
1320
+ rStar (generator @maj)
1321
+ 8.64
1322
+ 29.26
1323
+ 26.76
1324
+ 33.35
1325
+ 46.55
1326
+ rStar
1327
+ 18.57
1328
+ 37.91
1329
+ 32.97
1330
+ 37.53
1331
+ 46.55
1332
+ SVAMP
1333
+ Zero-shot CoT
1334
+ 8.90
1335
+ 26.10
1336
+ 40.20
1337
+ 70.90
1338
+ 84.70
1339
+ Few-shot CoT
1340
+ 48.10
1341
+ 72.80
1342
+ 76.90
1343
+ 89.20
1344
+ 92.80
1345
+ SC@maj8
1346
+ 49.90
1347
+ 74.60
1348
+ 79.10
1349
+ 89.20
1350
+ 93.50
1351
+ SC@maj64
1352
+ 54.10
1353
+ 76.70
1354
+ 80.70
1355
+ 90.50
1356
+ 93.30
1357
+ SC@maj128
1358
+ 54.50
1359
+ 76.60
1360
+ 80.80
1361
+ 90.60
1362
+ 93.70
1363
+ ToT
1364
+ 33.40
1365
+ 56.30
1366
+ 62.20
1367
+ 79.80
1368
+ 84.90
1369
+ RAP
1370
+ 41.00
1371
+ 71.80
1372
+ 73.10
1373
+ 85.70
1374
+ 91.50
1375
+ rStar (generator @maj)
1376
+ 60.30
1377
+ 83.10
1378
+ 86.20
1379
+ 91.89
1380
+ 93.80
1381
+ rStar
1382
+ 74.90
1383
+ 86.40
1384
+ 90.00
1385
+ 94.29
1386
+ 94.10
1387
+ StrategyQA
1388
+ Zero-shot CoT
1389
+ 52.67
1390
+ 57.20
1391
+ 41.48
1392
+ 57.21
1393
+ 54.68
1394
+ Few-shot CoT
1395
+ 58.82
1396
+ 65.65
1397
+ 64.05
1398
+ 68.41
1399
+ 63.61
1400
+ SC@maj8
1401
+ 59.10
1402
+ 65.50
1403
+ 63.76
1404
+ 68.26
1405
+ 64.34
1406
+ SC@maj64
1407
+ 58.51
1408
+ 63.61
1409
+ 63.46
1410
+ 67.39
1411
+ 62.74
1412
+ SC@maj128
1413
+ 58.37
1414
+ 62.01
1415
+ 63.31
1416
+ 66.67
1417
+ 59.53
1418
+ ToT
1419
+ 45.27
1420
+ 55.75
1421
+ 57.64
1422
+ 60.41
1423
+ 40.47
1424
+ RAP
1425
+ 59.68
1426
+ 64.48
1427
+ 63.32
1428
+ 68.71
1429
+ 60.26
1430
+ rStar (generator @maj)
1431
+ 61.57
1432
+ 69.43
1433
+ 65.50
1434
+ 71.47
1435
+ 65.50
1436
+ rStar
1437
+ 67.25
1438
+ 70.31
1439
+ 67.69
1440
+ 71.57
1441
+ 67.25
1442
+ 4
1443
+ Experiments
1444
+ 4.1
1445
+ Setup
1446
+ Models and Datasets
1447
+ . rStar is a general approach applicable to various LLMs and reasoning tasks. We evaluate 5 SLMs: Phi3-mini (3.8B)
1448
+ (Abdin et al.,
1449
+ 2024
1450
+ )
1451
+ , LLaMA2-7B, Mistral-7B
1452
+ (Jiang et al.,
1453
+ 2023
1454
+ )
1455
+ , LLaMA3-8B, and LLaMA3-8B-Instruct
1456
+ (Meta,
1457
+ 2024
1458
+ )
1459
+ . We test across 5 reasoning tasks, including 4 mathematical tasks (GSM8K
1460
+ (Cobbe et al.,
1461
+ 2021
1462
+ )
1463
+ , GSM-Hard
1464
+ (Gao et al.,
1465
+ 2022
1466
+ )
1467
+ , MATH
1468
+ (Hendrycks et al.,
1469
+ 2021
1470
+ )
1471
+ , SVAMP
1472
+ (Patel et al.,
1473
+ 2021
1474
+ )
1475
+ ) and one commonsense reasoning task (StrategyQA
1476
+ (Geva et al.,
1477
+ 2021
1478
+ )
1479
+ ).
1480
+ Implementation Details.
1481
+ In the trajectory self-generation stage, we augment each target SLM with our MCTS, performing 32 rollouts. Except for MATH, where we set the depth
1482
+ d
1483
+ 𝑑
1484
+ d
1485
+ to 8, all other tasks have a
1486
+ d
1487
+ 𝑑
1488
+ d
1489
+ =5. Actions
1490
+ A
1491
+ 1
1492
+ subscript
1493
+ 𝐴
1494
+ 1
1495
+ A_{1}
1496
+ and
1497
+ A
1498
+ 3
1499
+ subscript
1500
+ 𝐴
1501
+ 3
1502
+ A_{3}
1503
+ have a maximum of 5 nodes per depth, while the other actions have a default node count of 1. In the trajectory discrimination stage, we use Phi3-mini-4k as the discriminator, which has only 3.8B parameters, for effective inference. Moreover, the discriminator performs inference in a parallelized manner, making the verification process highly efficient.
1504
+ Notably, when Phi3 is the target SLM, it performs self-discrimination. For a given trajectory, we randomly split it between 20% and 80% of its steps, providing the first half of the steps as input to the discriminator SLM, which then completes the remaining steps.
1505
+ Detailed prompts are available in appendix
1506
+ A.3
1507
+ .
1508
+ 4.2
1509
+ Main Results
1510
+ Baselines
1511
+ . We compare rStar against three strong baseline types:
1512
+ (i)
1513
+ single-round CoT prompting
1514
+ , including zero-shot CoT
1515
+ (Kojima et al.,
1516
+ 2022
1517
+ )
1518
+ and few-shot CoT
1519
+ (Wei et al.,
1520
+ 2022
1521
+ )
1522
+ ;
1523
+ (ii)
1524
+ multi-round CoT prompting
1525
+ using the widely adopted self-consistency (SC) method
1526
+ (Wang et al.,
1527
+ 2023
1528
+ )
1529
+ . We sample answers 8, 64, and 128 times, employing majority voting for answer selection; and
1530
+ (iii)
1531
+ multi-round self-improving approaches
1532
+ . For this, we select ToT
1533
+ (Yao et al.,
1534
+ 2024
1535
+ )
1536
+ and RAP
1537
+ (Hao et al.,
1538
+ 2023
1539
+ )
1540
+ as baselines, using BFS and MCTS for tree search, respectively. Note that the action in ToT corresponds to our action
1541
+ A
1542
+ 1
1543
+ subscript
1544
+ 𝐴
1545
+ 1
1546
+ A_{1}
1547
+ , while RAP corresponds to our action
1548
+ A
1549
+ 3
1550
+ subscript
1551
+ 𝐴
1552
+ 3
1553
+ A_{3}
1554
+ . For the answer selection, we follow their original implementations.
1555
+ Results on diverse reasoning benchmarks
1556
+ . We start by evaluating the effectiveness of rStar on general reasoning benchmarks. Table
1557
+ 2
1558
+ compares its accuracy with state-of-the-art baselines on diverse SLMs and reasoning datasets. To demonstrate the effectiveness of our generator, we also provide the accuracy of rStar (gen. @maj), which do not apply our discriminator and use majority voting for answer verification. We highlight three key observations:
1559
+ (1)
1560
+ SLMs empowered with rStar demonstrate highly capable problem-solving abilities. For example, LLaMA2-7B originally had an accuracy of only 12.51% on GSM8K using few-shot CoT. However, with improvements from rStar, its accuracy increased to 63.91%, nearly matching the accuracy achieved with fine-tuning as shown in Fig.
1561
+ 1
1562
+ . Similarly, Mistral with rStar can even outperform fine-tuned MetaMath by +4.18%. This improvement shows that SLMs already have strong reasoning capabilities but need guidance to generate and select the correct solutions.
1563
+ (2)
1564
+ rStar consistently improves the reasoning accuracy of various evaluated SLMs across different tasks to a state-of-the-art level. In contrast, none of the baseline approaches consistently perform well across all four benchmarks. For example, while SC excels in three mathematical tasks, it is less effective on the logical reasoning task of StrategyQA. Specifically, SC with more sampling can even lower the score on StrategyQA. RAP performs better than SC on StrategyQA but falls short compared to SC on most mathematical reasoning tasks.
1565
+ (3)
1566
+ Even without our proposed discriminator for reasoning trajectory verification, our MCTS generator demonstrates greater effectiveness in improving reasoning accuracy for SLMs compared to existing multi-round inference baselines. For example, rStar (generator @maj) achieves up to 2.88%-16.39% higher accuracy than RAP, 10.60%- 38.37% higher accuracy than ToT, and 1.69% - 7.34% higher accuracy than SC on the GSM8K dataset.
1567
+ Table 3:
1568
+ Reasoning performance comparison on the challenging MATH-500 dataset. Due to the extensive LaTeX syntax in the dataset, which is challenging for pre-trained LLMs in instruction following, we evaluate only on LLaMA3-8B-instruct and Phi3-Mini-4k-Instruct.
1569
+ Method
1570
+ LLaMA3-8b-Instruct
1571
+ Phi3-mini-4k
1572
+ Zeroshot CoT
1573
+ 5.80
1574
+ 3.60
1575
+ Fewshot CoT
1576
+ 17.80
1577
+ 32.20
1578
+ SC@maj8
1579
+ 30.00
1580
+ 40.40
1581
+ SC@maj64
1582
+ 33.00
1583
+ 45.20
1584
+ SC@maj128
1585
+ 33.80
1586
+ 45.60
1587
+ ToT
1588
+ 13.60
1589
+ 18.20
1590
+ RAP
1591
+ 18.80
1592
+ 27.80
1593
+ rStar (generator @maj)
1594
+ 38.30
1595
+ 48.40
1596
+ rStar
1597
+ 42.94
1598
+ 48.60
1599
+ Figure 5:
1600
+ Performance comparison on the GSM8K dataset under different number of rollouts. rStar can significantly improve reasoning accuracy with just 2 rollouts.
1601
+ Results on challenging mathematical dataset
1602
+ . We also evaluate the effectiveness of rStar on more challenging mathematical datasets. In particular, we select the GSM-Hard and MATH datasets. Following
1603
+ (Wang et al.,
1604
+ 2024b
1605
+ ; Lightman et al.,
1606
+ 2023
1607
+ )
1608
+ , we use MATH-500, a subset of representative problems from the MATH dataset, to speedup the evaluation. As shown in Table
1609
+ 2
1610
+ and Table
1611
+ 3
1612
+ , rStar
1613
+ is capable of significantly improve the reasoning accuracy of SLMs on these challenging mathematical datasets.
1614
+ Remarkably, when compared to SOTA baselines, we observe a significant improvements of up to +12.9% and +9.14% on GSM-Hard and MATH-500, respectively.
1615
+ 4.3
1616
+ Ablation Study
1617
+ Table 4:
1618
+ Ablation study on the effectiveness of our MCTS generator. Ours+self-eval: we apply self-evaluation to prompt model for self-rewarding each intermediate action in our generator.
1619
+ Generator
1620
+ LLaMA3-8B
1621
+ LLaMA3-8B-Instruct
1622
+ GSM8K
1623
+ StrategyQA
1624
+ GSM8K
1625
+ StrategyQA
1626
+ Answer verification
1627
+ Answer verification
1628
+ Answer verification
1629
+ Answer verification
1630
+ Maj
1631
+ Ours
1632
+ Maj
1633
+ Ours
1634
+ Maj
1635
+ Ours
1636
+ Maj
1637
+ Ours
1638
+ RAP
1639
+ 56.56
1640
+ 57.31
1641
+ 62.30
1642
+ 64.63
1643
+ 81.35
1644
+ 84.69
1645
+ 69.43
1646
+ 70.60
1647
+ SC (@128)
1648
+ 67.55
1649
+ 85.06
1650
+ 63.31
1651
+ 65.65
1652
+ 84.69
1653
+ 89.99
1654
+ 66.67
1655
+ 68.56
1656
+ Ours+Self-eval
1657
+ 70.28
1658
+ 82.18
1659
+ 65.07
1660
+ 66.22
1661
+ 88.07
1662
+ 89.92
1663
+ 69.28
1664
+ 69.43
1665
+ Ours
1666
+ 74.38
1667
+ 85.52
1668
+ 65.50
1669
+ 67.69
1670
+ 88.70
1671
+ 91.13
1672
+ 71.47
1673
+ 71.57
1674
+ Table 5:
1675
+ Ablation study on discriminator effectiveness. We evaluate accuracy on GSM8K.
1676
+ Left
1677
+ : Our discriminator consistently outperforms others in verifying solution trajectories generated by different generators.
1678
+ Right
1679
+ : The ablation study on the choice of discriminator model.
1680
+ Discriminator
1681
+ LLaMA3-8B
1682
+ LLaMA3-8B-Instruct
1683
+ Generator
1684
+ Generator
1685
+ SC
1686
+ Ours
1687
+ SC
1688
+ Ours
1689
+ Maj
1690
+ 67.55
1691
+ 74.38
1692
+ 84.69
1693
+ 88.70
1694
+ Self-verification
1695
+ 74.00
1696
+ 75.52
1697
+ 83.02
1698
+ 86.63
1699
+ Ours
1700
+ 85.06
1701
+ 85.52
1702
+ 89.99
1703
+ 91.13
1704
+ Model
1705
+ Discriminator SLM
1706
+ Accuracy
1707
+ LLaMA3-8B-Instruct
1708
+ Maj
1709
+ 88.70
1710
+ LLaMA3-8B-Instruct
1711
+ 88.78
1712
+ LLaMA3.1-8B-Instruct
1713
+ 89.52
1714
+ Phi3-Mini-Instruct
1715
+ 91.13
1716
+ GPT-4 (2024-05-01)
1717
+ 92.57
1718
+ Effectiveness under different rollouts
1719
+ . rStar uses a rollout policy for MCTS tree expansion. More rollouts generate more candidate solution trajectories but increase inference cost. In Fig.
1720
+ 5
1721
+ , we compare the accuracy of SC, RAP, and our rStar across different rollouts on GSM8K. For SC, we sample solutions based on each number of rollouts and use majority voting to select the answer. We highlight two key observations:
1722
+ (1)
1723
+ Even with just 2 rollouts, rStar significantly improves reasoning accuracy for SLMs, demonstrating its effectiveness;
1724
+ (2)
1725
+ Both rStar and SC benefit from more rollouts, whereas RAP tends to saturate and even decline after 4 rollouts on LLaMA3-8B-Instruct. One reason is that the single-type action space in RAP limits the effective MCTS exploration.
1726
+ The effectiveness of MCTS generator
1727
+ . We compare our MCTS generator with three baselines: (i) the MCTS generator used in RAP; (ii) SC with 128 randomly sampled solutions; and (iii) our generator with Self-evaluation, a popular technique that self-evaluates the reward score for each action. Baseline (iii) specifically evaluates the effectiveness of our reward function.
1728
+ To isolate the impact of answer verification methods, each generator is evaluated under both majority voting and our discriminator for trajectory selection. As shown in Table
1729
+ 4
1730
+ , our generator consistently outperforms the baseline generators across different answer verification methods. More, we demonstrate the effectiveness of our SLM-tailored reward function, as self-evaluation reduces our generator’s accuracy.
1731
+ The effectiveness of discriminator
1732
+ . We setup two experiments for evaluation. First, we compare our discrimination approach with two baselines: the majority voting and self-verification
1733
+ (Weng et al.,
1734
+ 2023
1735
+ )
1736
+ . Specifically, we follow the key idea in
1737
+ Weng et al. (
1738
+ 2023
1739
+ )
1740
+ to prompt the SLM (i.e., generator SLM) to self-verify the correctness of each trajectory. To demonstrate the generalization ability of our discriminator, we used candidate solutions from different generators for evaluation.
1741
+ As shown in Table
1742
+ 5
1743
+ (Left), our discriminator significantly improves reasoning accuracy when performing answer verification on trajectories generated by different generators. Similar to the previous self-evaluation experiment, self-verification on SLMs is ineffective.
1744
+ Second, we study the impact of discriminator model selection. Our current discriminator models are all Phi3-Mini-Instruct. We tested various LLMs, both stronger and weaker, as discriminators for LLaMA3-8B-Instruct. As shown in Table
1745
+ 5
1746
+ (Right), the choice of discriminator model generally does not affect the effectiveness of our mutual reasoning consistency for answer verification. Notably, using the powerful GPT-4 as the discriminator only slightly improves performance (91.13% to 92.57%), demonstrating that mutual reasoning consistency can effectively verify answers using SLMs.
1747
+ 5
1748
+ Conclusion
1749
+ In this work, we present rStar, a generator-discriminator self-play approach that significantly grow the reasoning capabilities for SLMs at the inference time. Our approach reveals that SLMs, such as LLaMA2-7B, already exhibit strong reasoning capabilities prior to domain specialized supervised fine-tuning. rStar achieves state-of-the-art performance across five SLMs and five diverse reasoning tasks, substantially outperforming existing multi-round prompting and self-improvement approaches. Furthermore, we conduct extensive ablation studies and analysis, contributing to the development of more advanced SLM self-improved reasoning.
1750
+ References
1751
+ Abdin et al. (2024)
1752
+ Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah,
1753
+ Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl,
1754
+ et al.
1755
+ Phi-3 technical report: A highly capable language model locally on
1756
+ your phone.
1757
+ arXiv preprint arXiv:2404.14219
1758
+ , 2024.
1759
+ Besta et al. (2024)
1760
+ Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal
1761
+ Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert
1762
+ Niewiadomski, Piotr Nyczyk, et al.
1763
+ Graph of thoughts: Solving elaborate problems with large language
1764
+ models.
1765
+ In
1766
+ Proceedings of the AAAI Conference on Artificial
1767
+ Intelligence
1768
+ , volume 38, pp.  17682–17690, 2024.
1769
+ Brown et al. (2024)
1770
+ Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le,
1771
+ Christopher Ré, and Azalia Mirhoseini.
1772
+ Large language monkeys: Scaling inference compute with repeated
1773
+ sampling.
1774
+ arXiv preprint arXiv:2407.21787
1775
+ , 2024.
1776
+ Chen et al. (2024a)
1777
+ Guoxin Chen, Minpeng Liao, Chengxi Li, and Kai Fan.
1778
+ Alphamath almost zero: process supervision without process,
1779
+ 2024a.
1780
+ Chen et al. (2022)
1781
+ Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen.
1782
+ Program of thoughts prompting: Disentangling computation from
1783
+ reasoning for numerical reasoning tasks.
1784
+ arXiv preprint arXiv:2211.12588
1785
+ , 2022.
1786
+ Chen et al. (2024b)
1787
+ Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu.
1788
+ Self-play fine-tuning converts weak language models to strong
1789
+ language models.
1790
+ arXiv preprint arXiv:2401.01335
1791
+ , 2024b.
1792
+ Cobbe et al. (2021)
1793
+ Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz
1794
+ Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano,
1795
+ et al.
1796
+ Training verifiers to solve math word problems.
1797
+ arXiv preprint arXiv:2110.14168
1798
+ , 2021.
1799
+ Ding et al. (2023)
1800
+ Ruomeng Ding, Chaoyun Zhang, Lu Wang, Yong Xu, Minghua Ma, Wei Zhang, Si Qin,
1801
+ Saravan Rajmohan, Qingwei Lin, and Dongmei Zhang.
1802
+ Everything of thoughts: Defying the law of penrose triangle for
1803
+ thought generation.
1804
+ arXiv preprint arXiv:2311.04254
1805
+ , 2023.
1806
+ Feng et al. (2023)
1807
+ Xidong Feng, Ziyu Wan, Muning Wen, Ying Wen, Weinan Zhang, and Jun Wang.
1808
+ Alphazero-like tree-search can guide large language model decoding
1809
+ and training.
1810
+ arXiv preprint arXiv:2309.17179
1811
+ , 2023.
1812
+ Forsman (2024)
1813
+ Anton Forsman.
1814
+ Analyzing the performance of self-refine on different large language
1815
+ models.
1816
+ 2024.
1817
+ URL
1818
+ https://github.com/anforsm/self-refine/blob/main/report.pdf
1819
+ .
1820
+ Gao et al. (2022)
1821
+ Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie
1822
+ Callan, and Graham Neubig.
1823
+ Pal: Program-aided language models.
1824
+ arXiv preprint arXiv:2211.10435
1825
+ , 2022.
1826
+ Gero et al. (2023)
1827
+ Zelalem Gero, Chandan Singh, Hao Cheng, Tristan Naumann, Michel Galley,
1828
+ Jianfeng Gao, and Hoifung Poon.
1829
+ Self-verification improves few-shot clinical information extraction.
1830
+ arXiv preprint arXiv:2306.00024
1831
+ , 2023.
1832
+ Geva et al. (2021)
1833
+ Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan
1834
+ Berant.
1835
+ Did aristotle use a laptop? a question answering benchmark with
1836
+ implicit reasoning strategies.
1837
+ Transactions of the Association for Computational Linguistics
1838
+ ,
1839
+ 9:346–361, 2021.
1840
+ URL
1841
+ https://huggingface.co/datasets/ChilleD/StrategyQA
1842
+ .
1843
+ Gou et al. (2023)
1844
+ Zhibin Gou, Zhihong Shao, Yeyun Gong, Yujiu Yang, Minlie Huang, Nan Duan,
1845
+ Weizhu Chen, et al.
1846
+ Tora: A tool-integrated reasoning agent for mathematical problem
1847
+ solving.
1848
+ arXiv preprint arXiv:2309.17452
1849
+ , 2023.
1850
+ Hao et al. (2023)
1851
+ Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and
1852
+ Zhiting Hu.
1853
+ Reasoning with language model is planning with world model.
1854
+ arXiv preprint arXiv:2305.14992
1855
+ , 2023.
1856
+ Hendrycks et al. (2021)
1857
+ Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric
1858
+ Tang, Dawn Song, and Jacob Steinhardt.
1859
+ Measuring mathematical problem solving with the math dataset.
1860
+ arXiv preprint arXiv:2103.03874
1861
+ , 2021.
1862
+ Huang et al. (2023)
1863
+ Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu,
1864
+ Xinying Song, and Denny Zhou.
1865
+ Large language models cannot self-correct reasoning yet.
1866
+ arXiv preprint arXiv:2310.01798
1867
+ , 2023.
1868
+ Jiang et al. (2023)
1869
+ Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford,
1870
+ Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel,
1871
+ Guillaume Lample, Lucile Saulnier, et al.
1872
+ Mistral 7b.
1873
+ arXiv preprint arXiv:2310.06825
1874
+ , 2023.
1875
+ Kang et al. (2024)
1876
+ Jikun Kang, Xin Zhe Li, Xi Chen, Amirreza Kazemi, and Boxing Chen.
1877
+ Mindstar: Enhancing math reasoning in pre-trained llms at inference
1878
+ time.
1879
+ arXiv preprint arXiv:2405.16265
1880
+ , 2024.
1881
+ Khot et al. (2022)
1882
+ Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter
1883
+ Clark, and Ashish Sabharwal.
1884
+ Decomposed prompting: A modular approach for solving complex tasks.
1885
+ arXiv preprint arXiv:2210.02406
1886
+ , 2022.
1887
+ Kocsis & Szepesvári (2006)
1888
+ Levente Kocsis and Csaba Szepesvári.
1889
+ Bandit based monte-carlo planning.
1890
+ volume 2006, pp.  282–293, 09 2006.
1891
+ ISBN 978-3-540-45375-8.
1892
+ doi:
1893
+ 10.1007/11871842_29
1894
+ .
1895
+ Kojima et al. (2022)
1896
+ Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke
1897
+ Iwasawa.
1898
+ Large language models are zero-shot reasoners.
1899
+ Advances in neural information processing systems
1900
+ ,
1901
+ 35:22199–22213, 2022.
1902
+ Li et al. (2024)
1903
+ Chen Li, Weiqi Wang, Jingcheng Hu, Yixuan Wei, Nanning Zheng, Han Hu, Zheng
1904
+ Zhang, and Houwen Peng.
1905
+ Common 7b language models already possess strong math capabilities.
1906
+ arXiv preprint arXiv:2403.04706
1907
+ , 2024.
1908
+ Lightman et al. (2023)
1909
+ Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy
1910
+ Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe.
1911
+ Let’s verify step by step.
1912
+ arXiv preprint arXiv:2305.20050
1913
+ , 2023.
1914
+ Madaan et al. (2024)
1915
+ Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah
1916
+ Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al.
1917
+ Self-refine: Iterative refinement with self-feedback.
1918
+ Advances in Neural Information Processing Systems
1919
+ , 36, 2024.
1920
+ Meta (2024)
1921
+ Meta.
1922
+ Introducing meta llama3: The most capable openly available llm to
1923
+ date, 2024.
1924
+ URL
1925
+ https://ai.meta.com/blog/meta-llama-3/
1926
+ .
1927
+ Patel et al. (2021)
1928
+ Arkil Patel, Satwik Bhattamishra, and Navin Goyal.
1929
+ Are nlp models really able to solve simple math word problems?
1930
+ In
1931
+ Proceedings of the 2021 Conference of the North American
1932
+ Chapter of the Association for Computational Linguistics: Human Language
1933
+ Technologies
1934
+ , pp.  2080–2094, 2021.
1935
+ Roy & Roth (2015)
1936
+ Subhro Roy and Dan Roth.
1937
+ Solving General Arithmetic Word Problems.
1938
+ In
1939
+ Proc. of the Conference on Empirical Methods in Natural
1940
+ Language Processing (EMNLP)
1941
+ , 2015.
1942
+ URL
1943
+ http://cogcomp.org/papers/arithmetic.pdf
1944
+ .
1945
+ Silver et al. (2017)
1946
+ David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew
1947
+ Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore
1948
+ Graepel, et al.
1949
+ Mastering chess and shogi by self-play with a general reinforcement
1950
+ learning algorithm.
1951
+ arXiv preprint arXiv:1712.01815
1952
+ , 2017.
1953
+ Snell et al. (2024)
1954
+ Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar.
1955
+ Scaling llm test-time compute optimally can be more effective than
1956
+ scaling model parameters, 2024.
1957
+ URL
1958
+ https://arxiv.org/abs/2408.03314
1959
+ .
1960
+ Valmeekam et al. (2022)
1961
+ Karthik Valmeekam, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati.
1962
+ Large language models still can’t plan (a benchmark for LLMs on
1963
+ planning and reasoning about change).
1964
+ In
1965
+ NeurIPS 2022 Foundation Models for Decision Making
1966
+ Workshop
1967
+ , 2022.
1968
+ URL
1969
+ https://openreview.net/forum?id=wUU-7XTL5XO
1970
+ .
1971
+ Wang et al. (2024a)
1972
+ Ke Wang, Houxing Ren, Aojun Zhou, Zimu Lu, Sichun Luo, Weikang Shi, Renrui
1973
+ Zhang, Linqi Song, Mingjie Zhan, and Hongsheng Li.
1974
+ Mathcoder: Seamless code integration in LLMs for enhanced
1975
+ mathematical reasoning.
1976
+ In
1977
+ The Twelfth International Conference on Learning
1978
+ Representations
1979
+ , 2024a.
1980
+ URL
1981
+ https://openreview.net/forum?id=z8TW0ttBPp
1982
+ .
1983
+ Wang et al. (2024b)
1984
+ Peiyi Wang, Lei Li, Zhihong Shao, R. X. Xu, Damai Dai, Yifei Li, Deli Chen,
1985
+ Y. Wu, and Zhifang Sui.
1986
+ Math-shepherd: Verify and reinforce llms step-by-step without human
1987
+ annotations, 2024b.
1988
+ Wang et al. (2023)
1989
+ Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang,
1990
+ Aakanksha Chowdhery, and Denny Zhou.
1991
+ Self-consistency improves chain of thought reasoning in language
1992
+ models.
1993
+ In
1994
+ The Eleventh International Conference on Learning
1995
+ Representations
1996
+ , 2023.
1997
+ URL
1998
+ https://openreview.net/forum?id=1PL1NIMMrw
1999
+ .
2000
+ Wei et al. (2022)
2001
+ Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V
2002
+ Le, Denny Zhou, et al.
2003
+ Chain-of-thought prompting elicits reasoning in large language
2004
+ models.
2005
+ Advances in Neural Information Processing Systems
2006
+ ,
2007
+ 35:24824–24837, 2022.
2008
+ Weng et al. (2023)
2009
+ Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun,
2010
+ Kang Liu, and Jun Zhao.
2011
+ Large language models are better reasoners with self-verification.
2012
+ 2023.
2013
+ Wu et al. (2024)
2014
+ Zhenyu Wu, Qingkai Zeng, Zhihan Zhang, Zhaoxuan Tan, Chao Shen, and Meng Jiang.
2015
+ Large language models can self-correct with minimal effort.
2016
+ arXiv preprint arXiv:2405.14092
2017
+ , 2024.
2018
+ Yao et al. (2024)
2019
+ Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and
2020
+ Karthik Narasimhan.
2021
+ Tree of thoughts: Deliberate problem solving with large language
2022
+ models.
2023
+ Advances in Neural Information Processing Systems
2024
+ , 36, 2024.
2025
+ Zhang et al. (2024)
2026
+ Di Zhang, Jiatong Li, Xiaoshui Huang, Dongzhan Zhou, Yuqiang Li, and Wanli
2027
+ Ouyang.
2028
+ Accessing gpt-4 level mathematical olympiad solutions via monte carlo
2029
+ tree self-refine with llama-3 8b.
2030
+ arXiv preprint arXiv:2406.07394
2031
+ , 2024.
2032
+ Zheng et al. (2023)
2033
+ Huaixiu Steven Zheng, Swaroop Mishra, Xinyun Chen, Heng-Tze Cheng, Ed H Chi,
2034
+ Quoc V Le, and Denny Zhou.
2035
+ Take a step back: Evoking reasoning via abstraction in large language
2036
+ models.
2037
+ arXiv preprint arXiv:2310.06117
2038
+ , 2023.
2039
+ Zhou et al. (2023)
2040
+ Aojun Zhou, Ke Wang, Zimu Lu, Weikang Shi, Sichun Luo, Zipeng Qin, Shaoqing Lu,
2041
+ Anya Jia, Linqi Song, Mingjie Zhan, et al.
2042
+ Solving challenging math word problems using gpt-4 code interpreter
2043
+ with code-based self-verification.
2044
+ arXiv preprint arXiv:2308.07921
2045
+ , 2023.
2046
+ Zhou et al. (2022)
2047
+ Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi
2048
+ Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, et al.
2049
+ Least-to-most prompting enables complex reasoning in large language
2050
+ models.
2051
+ In
2052
+ The Eleventh International Conference on Learning
2053
+ Representations
2054
+ , 2022.
2055
+ Zhou et al. (2024)
2056
+ Pei Zhou, Jay Pujara, Xiang Ren, Xinyun Chen, Heng-Tze Cheng, Quoc V Le, Ed H
2057
+ Chi, Denny Zhou, Swaroop Mishra, and Huaixiu Steven Zheng.
2058
+ Self-discover: Large language models self-compose reasoning
2059
+ structures.
2060
+ arXiv preprint arXiv:2402.03620
2061
+ , 2024.
2062
+ Appendix A
2063
+ Appendix
2064
+ A.1
2065
+ Experiments to evaluate the self-rewarding in SLMs
2066
+ Table 6:
2067
+ Analysis on the effectiveness of SLMs’ self-rewarding. The original
2068
+ r
2069
+ 1
2070
+ subscript
2071
+ 𝑟
2072
+ 1
2073
+ r_{1}
2074
+ is a self-evaluation of the helpfulness of the new proposed subquestion, while
2075
+ r
2076
+ 2
2077
+ subscript
2078
+ 𝑟
2079
+ 2
2080
+ r_{2}
2081
+ measures the confidence in answering the subquestion through self-consistency majority voting. Results show that replacing the self-evaluated
2082
+ r
2083
+ 1
2084
+ subscript
2085
+ 𝑟
2086
+ 1
2087
+ r_{1}
2088
+ to random values does not significantly impact the final reasoning performance.
2089
+ Method
2090
+ LLaMA2-7B
2091
+ Mistral
2092
+ GSM8K
2093
+ RAP
2094
+ 24.34
2095
+ 56.25
2096
+ RAP + random
2097
+ r
2098
+ 1
2099
+ subscript
2100
+ 𝑟
2101
+ 1
2102
+ r_{1}
2103
+ 22.90
2104
+ 55.50
2105
+ RAP + random
2106
+ r
2107
+ 2
2108
+ subscript
2109
+ 𝑟
2110
+ 2
2111
+ r_{2}
2112
+ 22.67
2113
+ 49.66
2114
+ Multiarith
2115
+ RAP
2116
+ 57.22
2117
+ 91.11
2118
+ RAP + random
2119
+ r
2120
+ 1
2121
+ subscript
2122
+ 𝑟
2123
+ 1
2124
+ r_{1}
2125
+ 52.78
2126
+ 90.56
2127
+ RAP + random
2128
+ r
2129
+ 2
2130
+ subscript
2131
+ 𝑟
2132
+ 2
2133
+ r_{2}
2134
+ 47.22
2135
+ 81.11
2136
+ Ablation study on self-rewarding in RAP
2137
+ . RAP rewards both intermediate and terminal nodes. For each node generated by its action, it combines two scores,
2138
+ r
2139
+ 1
2140
+ subscript
2141
+ 𝑟
2142
+ 1
2143
+ r_{1}
2144
+ and
2145
+ r
2146
+ 2
2147
+ subscript
2148
+ 𝑟
2149
+ 2
2150
+ r_{2}
2151
+ , to determine the final reward score. Formally,
2152
+ r
2153
+ =
2154
+ r
2155
+ 1
2156
+ ×
2157
+ r
2158
+ 2
2159
+ 𝑟
2160
+ subscript
2161
+ 𝑟
2162
+ 1
2163
+ subscript
2164
+ 𝑟
2165
+ 2
2166
+ r=r_{1}\times r_{2}
2167
+ .
2168
+ r
2169
+ 1
2170
+ subscript
2171
+ 𝑟
2172
+ 1
2173
+ r_{1}
2174
+ is a self-evaluation score that evaluates the LLM’s own estimation of the helpfulness of the current node. Specifically, it prompts the LLM with the question "
2175
+ Is the new question useful
2176
+ ?".
2177
+ r
2178
+ 2
2179
+ subscript
2180
+ 𝑟
2181
+ 2
2182
+ r_{2}
2183
+ is the confidence of correctly answering the proposed new question, measured by self-consistency majority voting.
2184
+ To evaluate the effectiveness of self-rewarding in RAP, we replace
2185
+ r
2186
+ 1
2187
+ subscript
2188
+ 𝑟
2189
+ 1
2190
+ r_{1}
2191
+ and
2192
+ r
2193
+ 2
2194
+ subscript
2195
+ 𝑟
2196
+ 2
2197
+ r_{2}
2198
+ with random values sampled from (0,1)and re-run RAP on LLaMA2-7B and Mistral-7B. We select a challenging dataset, GSM8K and an easy mathematical reasoning dataset, Multiarith
2199
+ (Roy & Roth,
2200
+ 2015
2201
+ )
2202
+ , for evaluation.
2203
+ Table
2204
+ 6
2205
+ compares the results with original RAP. We can see that replacing
2206
+ r
2207
+ 1
2208
+ subscript
2209
+ 𝑟
2210
+ 1
2211
+ r_{1}
2212
+ with random values has minimal impact on RAP’s performance across different SLMs and datasets. However, replacing
2213
+ r
2214
+ 2
2215
+ subscript
2216
+ 𝑟
2217
+ 2
2218
+ r_{2}
2219
+ with random values result in a noticeable drop in accuracy on Mistral and Multiarith. This indicates that self-evaluation
2220
+ r
2221
+ 1
2222
+ subscript
2223
+ 𝑟
2224
+ 1
2225
+ r_{1}
2226
+ has minimal effect, suggesting that LLaMA2-7B and Mistral are essentially performing near-random self-evaluations.
2227
+ A.2
2228
+ Discussions
2229
+ Discussions on the importance of generator and discriminator
2230
+ . In our experiments, we found that on certain SLMs, the discriminator yields more significant improvement than the generator. For instance, on LLaMA2-7B, rStar (generator @maj) can improves accuracy by +4.17% on GSM8K, while our discriminator can further boosts accuracy by +36.69%. However, both the generator and discriminator are crucial to final performance. The generator must effectively produce the correct solution first, and the discriminator then provide reasonable feedback to select the correct solution.
2231
+ The importance of the generator and discriminator varies based on the SLM’s solution generation effectiveness. For LLaMA2-7B, the proportion of correct solution trajectories is low, leading to a lower majority voting score. In this case, the discriminator is crucial for identifying the correct solution, significantly boosting accuracy. Conversely, for stronger models like LLaMA3-8B-instruct, our generator produces a higher proportion of correct solutions. Therefore, improving the generator results in greater accuracy improvements.
2232
+ Table 7:
2233
+ Inference costs of rStar on GSM8K. We show the average number of inferences and generated tokens required to solve a question.
2234
+ LLaMA2-7B
2235
+ Mistral
2236
+ Avg. calls
2237
+ 166.81
2238
+ 148.90
2239
+ Avg. generated tokens
2240
+ 367.1k
2241
+ 348.6k
2242
+ Inference cost
2243
+ . rStar grows SLMs reasoning capabilities at inference time. The primary inference cost arises from our MCTS self-generator. Table
2244
+ 7
2245
+ shows the average number of inferences and tokens generated for solving a GSM8K question after 32 rollouts.
2246
+ On LLaMA2-7B and Mistral, this averages 166 and 148 model calls to solve a question, respectively. Currently, completing the 32 rollouts for the entire GSM8K test set takes about 4.5 days on a single A100 GPU per model. These costs can be significantly reduced by distributing tasks across multiple GPUs or batching model calls within each rollout.
2247
+ A.3
2248
+ Prompt examples in MTCS
2249
+ Action 1. Propose an one-step thought.
2250
+ ### Instruction:
2251
+ There are 15 trees in the grove. Grove workers will plant trees in the grove today. After they are done, there will be 21 trees. How many trees did the grove workers plant today?
2252
+ ### Response:
2253
+ Let’s think step by step.
2254
+ Step 1: Identify the initial number of trees. The problem states there are 15 trees in the grove.
2255
+ Step 2: Identify the final number of trees. The problem states there will be 21 trees after the workers are done planting.
2256
+ Step 3: Subtract the initial number of trees from the final number of trees to find out how many trees were planted.
2257
+ Step 4: Therefore, the grove workers planted 21 (final number of trees) - 15 (initial number of trees) = 6 trees today.
2258
+ Step 5: The answer is 6.
2259
+ ### Instruction:
2260
+ If there are 3 cars in the parking lot and 2 more cars arrive, how many cars are in the parking lot?
2261
+ ### Response:
2262
+ Let’s think step by step.
2263
+ Step 1: Start with the number of cars that are already in the parking lot, which is 3 cars.
2264
+ Step 2: Add the number of cars that arrive, which is 2 cars.
2265
+ Step 3: Add the numbers together. there are 3 cars + 2 cars = 5 cars in the parking lot.
2266
+ Step 4: The answer is 5.
2267
+ ### Instruction:
2268
+ Leah had 32 chocolates and her sister had 42. If they ate 35, how many pieces do they have left in total?
2269
+ ### Response:
2270
+ Let’s think step by step.
2271
+ Step 1: Add the number of chocolates Leah and her sister had initially. Leah had 32 chocolates and her sister had 42 chocolates. So, they had 32 + 42 = 74 chocolates in total.
2272
+ Step 2: Subtract the number of chocolates they ate from the total number of chocolates they had. They ate 35 chocolates. So, they have 74 - 35 = 39 chocolates left.
2273
+ Step 3: The answer is 39.
2274
+ ### Instruction:
2275
+ Jason had 20 lollipops. He gave Denny some lollipops. Now Jason has 12 lollipops. How many lollipops did Jason give to Denny?
2276
+ ### Response:
2277
+ Let’s think step by step.
2278
+ Step 1: Identify the initial number of lollipops. The problem states Jason had 20 lollipops.
2279
+ Step 2: Identify the final number of lollipops. The problem states Jason now has 12 lollipops.
2280
+ Step 3: Subtract the final number of lollipops from the initial number to find out how many lollipops Jason gave to Denny.
2281
+ Step 4: Therefore, Jason gave Denny 20 (initial number of lollipops) - 12 (final number of lollipops) = 8 lollipops.
2282
+ Step 5: The answer is 8.
2283
+ ###
2284
+ Instruction:
2285
+ {user question}
2286
+ ###
2287
+ Response:
2288
+ Let’s think step by step.
2289
+ Action 2: Propose the remaining thought steps /A4: Answer the sub-question again.
2290
+ ### Instruction:
2291
+ There are 15 trees in the grove. Grove workers will plant trees in the grove today. After they are done, there will be 21 trees. How many trees did the grove workers plant today?
2292
+ ### Response:
2293
+ Let’s think step by step. There are 15 trees originally. Then there were 21 trees after some more were planted. So there must have been 21 - 15 = 6. The answer is: 6.
2294
+ ### Instruction:
2295
+ If there are 3 cars in the parking lot and 2 more cars arrive, how many cars are in the parking lot?
2296
+ ### Response:
2297
+ Let’s think step by step. There are originally 3 cars. 2 more cars arrive. 3 + 2 = 5. The answer is: 5.
2298
+ ### Instruction:
2299
+ Leah had 32 chocolates and her sister had 42. If they ate 35, how many pieces do they have left in total?
2300
+ ### Response:
2301
+ Let’s think step by step. Originally, Leah had 32 chocolates. Her sister had 42. So in total they had 32 + 42 = 74. After eating 35, they had 74 - 35 = 39. The answer is: 39.
2302
+ ### Instruction:
2303
+ Jason had 20 lollipops. He gave Denny some lollipops. Now Jason has 12 lollipops. How many lollipops did Jason give to Denny?
2304
+ ### Response:
2305
+ Let’s think step by step. Jason started with 20 lollipops. Then he had 12 after giving some to Denny. So he gave Denny 20 - 12 = 8. The answer is: 8.
2306
+ ### Instruction:
2307
+ Shawn has five toys. For Christmas, he got two toys each from his mom and dad. How many toys does he have now?
2308
+ ### Response:
2309
+ Let’s think step by step. Shawn started with 5 toys. If he got 2 toys each from his mom and dad, then that is 4 more toys. 5 + 4 = 9. The answer is: 9.
2310
+ ### Instruction:
2311
+ There were nine computers in the server room. Five more computers were installed each day, from monday to thursday. How many computers are now in the server room?
2312
+ ### Response:
2313
+ Let’s think step by step. There were originally 9 computers. For each of 4 days, 5 more computers were added. So 5 * 4 = 20 computers were added. 9 + 20 is 29. The answer is: 29.
2314
+ ### Instruction:
2315
+ Michael had 58 golf balls. On tuesday, he lost 23 golf balls. On wednesday, he lost 2 more. How many golf balls did he have at the end of wednesday?
2316
+ ### Response:
2317
+ Let’s think step by step. Michael started with 58 golf balls. After losing 23 on tuesday, he had 58 - 23 = 35. After losing 2 more, he had 35 - 2 = 33 golf balls. The answer is: 33.
2318
+ ### Instruction:
2319
+ Olivia has $23. She bought five bagels for $3 each. How much money does she have left?
2320
+ ### Response:
2321
+ Let’s think step by step. Olivia had 23 dollars. 5 bagels for 3 dollars each will be 5 x 3 = 15 dollars. So she has 23 - 15 dollars left. 23 - 15 is 8. The answer is: 8.
2322
+ ###
2323
+ Instruction:
2324
+ {user question}
2325
+ ###
2326
+ Response
2327
+ :
2328
+ Action 3: Propose next sub-question along with its answer.
2329
+ Given a question, please decompose it into sub-questions. For each sub-question, please answer it in a complete sentence, ending with "The answer is <a numeric answer>". When the original question is answerable, please start the subquestion with "Now we can answer the question: <original question>".
2330
+ Question 1: Four years ago, Kody was only half as old as Mohamed. If Mohamed is currently twice as 30 years old, how old is Kody?
2331
+ Question 1.1: How old is Mohamed currently?
2332
+ Answer 1.1: Mohamed is twice as old as 30 years, which means he is 30 * 2 = 60 years old.
2333
+ Question 1.2: What was Kody’s age four years ago, given that it was half of Mohamed’s age at that time?
2334
+ Answer 1.2: Four years ago, Mohamed was 60 - 4 = 56 years old, so Kody was half of that, which is 56 / 2 = 28 years old.
2335
+ Question 1.3: Now we can answer the question: How old is Kody?
2336
+ Answer 1.3: Kody is currently 28 + 4 = 32 years old. The answer is 32.
2337
+ Question 2: On a moonless night, three fireflies danced in the evening breeze. They were joined by four less than a dozen more fireflies before two of the fireflies flew away. How many fireflies remained?
2338
+ Question 2.1: How many fireflies joined?
2339
+ Answer 2.1: The fireflies were joined by four less than a dozen more fireflies, which are 12 - 4 = 8 fireflies. The answer is 8.
2340
+ Question 2.2: Now we can answer the question: How many fireflies remained?
2341
+ Answer 2.2: Three fireflies were dancing originally. They were joined by 8 fireflies before two of them flew away. So there were 3 + 8 - 2 = 9 remaining. The answer is 9.
2342
+ Question 3: Ali has four $10 bills and six $20 bills that he saved after working for Mr. James on his farm. Ali gives her sister half of the total money he has and uses 3/5 of the remaining amount of money to buy dinner. Calculate the amount of money he has after buying the dinner.
2343
+ Question 3.1: How much money does Ali have after giving half of his total money to his sister?
2344
+ Answer 3.1: Ali initially has four $10 bills and six $20 bills, totaling 4 * 10 + 6 * 20 = 160 dollars. Giving half of this to his sister leaves him with 160 / 2 = 80 dollars. The answer is 80.
2345
+ Question 3.2: How much money does Ali spend on dinner?
2346
+ Answer 3.2: Ali uses 3/5 of his remaining money, which is 80 dollars, to buy dinner. Therefore, he spends 80 * 3/5 = 48 dollars on dinner. The answer is 48.
2347
+ Question 3.3: Now we can answer the question: How much money does Ali have after buying the dinner?
2348
+ Answer 3.3: After buying the dinner, Ali has 80 - 48 = 32 dollars left. The answer is 32.
2349
+ Question 4: A car is driving through a tunnel with many turns. After a while, the car must travel through a ring that requires a total of 4 right-hand turns. After the 1st turn, it travels 5 meters. After the 2nd turn, it travels 8 meters. After the 3rd turn, it travels a little further and at the 4th turn, it immediately exits the tunnel. If the car has driven a total of 23 meters around the ring, how far did it have to travel after the 3rd turn?
2350
+ Question 4.1: How far did the car travel except for the 3rd turn?
2351
+ Answer 4.1: It travels 5 meters after the 1st, 8 meters after the 2nd, and 0 meters after the 4th turn. It’s a total of 5 + 8 + 0 = 13 meters. The answer is 13.
2352
+ Question 4.2: Now we can answer the question: How far did the car have to travel after the 3rd turn?
2353
+ Answer 4.2: The car has driven a total of 23 meters around the ring. It travels 13 meters except for the 3rd turn. So it has to travel 23 - 13 = 10 meters after the 3rd turn. The answer is 10.
2354
+ Question 5: {user question}
2355
+ Action 5: Rephrase the question/sub-question.
2356
+ You are an AI assistant to help me rephrase questions by splitting the question context into conditions. In your rephrased question, remember to fully express the information in the original question.
2357
+ Original Question: Olivia has $23. She bought five bagels for $3 each. How much money does she have left?
2358
+ Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Olivia starts with $23. Condition 2: She buys five bagels, each costing $3. Question: How much money does Olivia have remaining after her purchase?
2359
+ Original Question: Michael had 58 golf balls. On Tuesday, he lost 23 golf balls. On Wednesday, he lost 2 more. How many golf balls did he have at the end of Wednesday?
2360
+ Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Michael initially has 58 golf balls. Condition 2: On Tuesday, he loses 23 golf balls. Condition 3: On Wednesday, he loses 2 additional golf balls. Question: What is the total number of golf balls Michael has left at the end of Wednesday?
2361
+ Original Question: Angelo and Melanie want to plan how many hours over the next week they should study together for their test next week. They have 2 chapters of their textbook to study and 4 worksheets to memorize. They figure out that they should dedicate 3 hours to each chapter of their textbook and 1.5 hours for each worksheet. If they plan to study no more than 4 hours each day, how many days should they plan to study total over the next week if they take a 10-minute break every hour, include 3 10-minute snack breaks each day, and 30 minutes for lunch each day?
2362
+ Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Angelo and Melanie need to study 2 textbook chapters and 4 worksheets. Condition 2: They allocate 3 hours per textbook chapter and 1.5 hours per worksheet. Condition 3: Their daily study limit is 4 hours, with a 10-minute break every hour, three 10-minute snack breaks, and a 30-minute lunch break each day. Question: Over the next week, for how many days should they plan to study to cover all their materials?
2363
+ Original Question: Leah had 32 chocolates and her sister had 42. If they ate 35, how many pieces do they have left in total?
2364
+ Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Leah has 32 chocolates. Condition 2: Her sister has 42 chocolates. Condition 3: Together, they consume 35 chocolates. Question: How many chocolates remain between them after they have eaten some?
2365
+ Original Question: There were nine computers in the server room. Five more computers were installed each day, from Monday to Thursday. How many computers are now in the server room?
2366
+ Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Initially, there are nine computers in the server room. Condition 2: Each day, from Monday to Thursday, five additional computers are installed. Question: What is the total number of computers in the server room after these installations?
2367
+ Original Question: Jason had 20 lollipops. He gave Denny some lollipops. Now Jason has 12 lollipops. How many lollipops did Jason give to Denny?
2368
+ Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Jason starts with 20 lollipops. Condition 2: After giving some lollipops to Denny, Jason has 12 lollipops left. Question: How many lollipops did Jason give to Denny?
2369
+ Original Question: Sam bought a dozen boxes, each with 30 highlighter pens inside, for $10 each box. He rearranged five of these boxes into packages of six highlighters each and sold them for $3 per package. He sold the rest of the highlighters separately at the rate of three pens for $2. How much profit did he make in total, in dollars?
2370
+ Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Sam purchases a dozen boxes of highlighters, with each box containing 30 pens, at $10 per box. Condition 2: He repackages five boxes into packages of six highlighters, selling each package for $3. Condition 3: He sells the remaining highlighters at a rate of three for $2. Question: What is Sam’s total profit from these transactions?
2371
+ Original Question: There are 15 trees in the grove. Grove workers will plant trees in the grove today. After they are done, there will be 21 trees. How many trees did the grove workers plant today?
2372
+ Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Initially, there are 15 trees in the grove. Condition 2: Grove workers will add more trees to the grove today. Condition 3: After planting, the total number of trees in the grove will increase to 21. Question: How many trees did the grove workers plant today?
2373
+ Original Question: {user question}
2374
+ Rephrased Question:
2375
+ ◄
2376
+ Feeling
2377
+ lucky?
2378
+ Conversion
2379
+ report
2380
+ Report
2381
+ an issue
2382
+ View original
2383
+ on arXiv
2384
+ ►
research/notes/240806195-mutual-reasoning-makes-smaller-llms-stronger-problem-solvers.md ADDED
@@ -0,0 +1,191 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2408.06195] Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers'
3
+ id: 240806195-mutual-reasoning-makes-smaller-llms-stronger-problem-solvers
4
+ tags:
5
+ - deepread
6
+ created: '2026-06-10T00:39:56.384617Z'
7
+ source: https://arxiv.org/abs/2408.06195
8
+ source_domain: arxiv.org
9
+ fetched_at: '2026-06-10T00:39:56.384488Z'
10
+ fetch_provider: builtin
11
+ status: draft
12
+ type: note
13
+ tier: institutional
14
+ content_type: paper
15
+ deprecated: false
16
+ ---
17
+
18
+ [2408.06195] Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
19
+ Computer Science > Computation and Language
20
+ arXiv:2408.06195
21
+ (cs)
22
+ [Submitted on 12 Aug 2024]
23
+ Title:
24
+ Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
25
+ Authors:
26
+ Zhenting Qi
27
+ ,
28
+ Mingyuan Ma
29
+ ,
30
+ Jiahang Xu
31
+ ,
32
+ Li Lyna Zhang
33
+ ,
34
+ Fan Yang
35
+ ,
36
+ Mao Yang
37
+ View a PDF of the paper titled Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers, by Zhenting Qi and 5 other authors
38
+ View PDF
39
+ HTML (experimental)
40
+ Abstract:
41
+ This paper introduces rStar, a self-play mutual reasoning approach that significantly improves reasoning capabilities of small language models (SLMs) without fine-tuning or superior models. rStar decouples reasoning into a self-play mutual generation-discrimination process. First, a target SLM augments the Monte Carlo Tree Search (MCTS) with a rich set of human-like reasoning actions to construct higher quality reasoning trajectories. Next, another SLM, with capabilities similar to the target SLM, acts as a discriminator to verify each trajectory generated by the target SLM. The mutually agreed reasoning trajectories are considered mutual consistent, thus are more likely to be correct. Extensive experiments across five SLMs demonstrate rStar can effectively solve diverse reasoning problems, including GSM8K, GSM-Hard, MATH, SVAMP, and StrategyQA. Remarkably, rStar boosts GSM8K accuracy from 12.51% to 63.91% for LLaMA2-7B, from 36.46% to 81.88% for Mistral-7B, from 74.53% to 91.13% for LLaMA3-8B-Instruct. Code will be available at
42
+ this https URL
43
+ .
44
+ Subjects:
45
+ Computation and Language (cs.CL)
46
+ Cite as:
47
+ arXiv:2408.06195
48
+ [cs.CL]
49
+ (or
50
+ arXiv:2408.06195v1
51
+ [cs.CL]
52
+ for this version)
53
+ https://doi.org/10.48550/arXiv.2408.06195
54
+ Focus to learn more
55
+ arXiv-issued DOI via DataCite
56
+ Submission history
57
+ From: Li Lyna Zhang [
58
+ view email
59
+ ]
60
+ [v1]
61
+ Mon, 12 Aug 2024 14:42:13 UTC (1,140 KB)
62
+ Full-text links:
63
+ Access Paper:
64
+ View a PDF of the paper titled Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers, by Zhenting Qi and 5 other authors
65
+ View PDF
66
+ HTML (experimental)
67
+ TeX Source
68
+ view license
69
+ Current browse context:
70
+ cs.CL
71
+ < prev
72
+ |
73
+ next >
74
+ new
75
+ |
76
+ recent
77
+ |
78
+ 2024-08
79
+ Change to browse by:
80
+ cs
81
+ References & Citations
82
+ NASA ADS
83
+ Google Scholar
84
+ Semantic Scholar
85
+ export BibTeX citation
86
+ Loading...
87
+ BibTeX formatted citation
88
+ ×
89
+ loading...
90
+ Data provided by:
91
+ Bookmark
92
+ Bibliographic Tools
93
+ Bibliographic and Citation Tools
94
+ Bibliographic Explorer Toggle
95
+ Bibliographic Explorer
96
+ (
97
+ What is the Explorer?
98
+ )
99
+ Connected Papers Toggle
100
+ Connected Papers
101
+ (
102
+ What is Connected Papers?
103
+ )
104
+ Litmaps Toggle
105
+ Litmaps
106
+ (
107
+ What is Litmaps?
108
+ )
109
+ scite.ai Toggle
110
+ scite Smart Citations
111
+ (
112
+ What are Smart Citations?
113
+ )
114
+ Code, Data, Media
115
+ Code, Data and Media Associated with this Article
116
+ alphaXiv Toggle
117
+ alphaXiv
118
+ (
119
+ What is alphaXiv?
120
+ )
121
+ Links to Code Toggle
122
+ CatalyzeX Code Finder for Papers
123
+ (
124
+ What is CatalyzeX?
125
+ )
126
+ DagsHub Toggle
127
+ DagsHub
128
+ (
129
+ What is DagsHub?
130
+ )
131
+ GotitPub Toggle
132
+ Gotit.pub
133
+ (
134
+ What is GotitPub?
135
+ )
136
+ Huggingface Toggle
137
+ Hugging Face
138
+ (
139
+ What is Huggingface?
140
+ )
141
+ ScienceCast Toggle
142
+ ScienceCast
143
+ (
144
+ What is ScienceCast?
145
+ )
146
+ Demos
147
+ Demos
148
+ Replicate Toggle
149
+ Replicate
150
+ (
151
+ What is Replicate?
152
+ )
153
+ Spaces Toggle
154
+ Hugging Face Spaces
155
+ (
156
+ What is Spaces?
157
+ )
158
+ Spaces Toggle
159
+ TXYZ.AI
160
+ (
161
+ What is TXYZ.AI?
162
+ )
163
+ Related Papers
164
+ Recommenders and Search Tools
165
+ Link to Influence Flower
166
+ Influence Flower
167
+ (
168
+ What are Influence Flowers?
169
+ )
170
+ Core recommender toggle
171
+ CORE Recommender
172
+ (
173
+ What is CORE?
174
+ )
175
+ Author
176
+ Venue
177
+ Institution
178
+ Topic
179
+ About arXivLabs
180
+ arXivLabs: experimental projects with community collaborators
181
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
182
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
183
+ Have an idea for a project that will add value for arXiv's community?
184
+ Learn more about arXivLabs
185
+ .
186
+ Which authors of this paper are endorsers?
187
+ |
188
+ Disable MathJax
189
+ (
190
+ What is MathJax?
191
+ )
research/notes/241020285-swe-search-enhancing-software-agents-with-monte-carlo-tree-search-and-2.md ADDED
@@ -0,0 +1,2144 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2410.20285] SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search
3
+ and Iterative Refinement'
4
+ id: 241020285-swe-search-enhancing-software-agents-with-monte-carlo-tree-search-and-2
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:41:20.614811Z'
8
+ source: https://ar5iv.labs.arxiv.org/html/2410.20285
9
+ source_domain: ar5iv.labs.arxiv.org
10
+ fetched_at: '2026-06-10T00:41:20.614654Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2410.20285] SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement
20
+ SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement
21
+ Antonis Antoniades
22
+ 1∗
23
+ , Albert Örwall
24
+ 2
25
+ ,
26
+ Kexun Zhang
27
+ 3
28
+ ,
29
+ Yuxi Xie
30
+ 4
31
+ , Anirudh Goyal
32
+ 5
33
+ , William Wang
34
+ 1
35
+ 1
36
+ University of California, Santa Barbara,
37
+ 2
38
+ Moatless AI,
39
+ 3
40
+ Carnegie Mellon University,
41
+ 4
42
+ National University of Singapore,
43
+ 5
44
+ Mila
45
+ Denotes equal contribution.
46
+ Correspondence to:
47
+ antonis@ucsb.edu
48
+ ,
49
+ albert@moatless.ai
50
+ .
51
+ Code:
52
+ github.com/aorwall/moatless-tree-search
53
+ Abstract
54
+ Software engineers operating in complex and dynamic environments must continuously adapt to evolving requirements, learn iteratively from experience, and reconsider their approaches based on new insights. However, current large language model (LLM)-based software agents often rely on rigid processes and tend to repeat ineffective actions without the capacity to evaluate their performance or adapt their strategies over time. To address these challenges, we propose SWE-Search, a multi-agent framework that integrates Monte Carlo Tree Search (MCTS) with a self-improvement mechanism to enhance software agents’ performance on repository-level software tasks. SWE-Search extends traditional MCTS by incorporating a hybrid value function that leverages LLMs for both numerical value estimation and qualitative evaluation. This enables self-feedback loops where agents iteratively refine their strategies based on both quantitative numerical evaluations and qualitative natural language assessments of pursued trajectories. The framework includes a SWE-Agent for adaptive exploration, a Value Agent for iterative feedback, and a Discriminator Agent that facilitates multi-agent debate for collaborative decision-making. Applied to the SWE-bench benchmark, our approach demonstrates a 23% relative improvement in performance across five models compared to standard open-source agents without MCTS. Our analysis reveals how performance scales with increased search depth and identifies key factors that facilitate effective self-evaluation in software agents. This work highlights the potential of self-evaluation driven search techniques to enhance agent reasoning and planning in complex, dynamic software engineering environments.
55
+ 1
56
+ Introduction
57
+ Software engineering is a complex and iterative process involving exploration, problem-solving, and decision-making under uncertainty. Tasks such as debugging, feature development, and code refactoring require continuous assessment of different approaches, frequent backtracking, and the incorporation of new information. While machine learning has made progress in automating parts of this workflow
58
+ (Li et al.,
59
+ 2022
60
+ ; OpenAI et al.,
61
+ 2024
62
+ ; Ouyang et al.,
63
+ 2022
64
+ ; Yang et al.,
65
+ 2024b
66
+ )
67
+ , replicating the adaptive and strategic behavior of human engineers remains a significant challenge. This is due to the inherently non-linear and iterative nature of software engineering, where engineers dynamically explore various solutions, refine strategies based on feedback, and collaborate to identify the most effective path forward. Current large language model (LLM)-based software agents
68
+ (Xia et al.,
69
+ 2024
70
+ ; Zhang et al.,
71
+ 2024d
72
+ )
73
+ , while powerful, often struggle with complex, long-horizon tasks that require adaptive strategies and flexible reassessment over time. These agents can become trapped in repetitive patterns, limiting their effectiveness in tackling more intricate software engineering problems.
74
+ To address these challenges, we introduce
75
+ SWE-Search
76
+ , a multi-agent system that replicates the adaptability, iterative learning, and collaborative decision-making of human engineers. SWE-Search is designed to address three critical needs in software engineering:
77
+ Flexible Exploration and Adaptation
78
+ : Engineering problems often require exploring multiple approaches and adapting strategies based on evolving information
79
+ (Li et al.,
80
+ 2022
81
+ )
82
+ . SWE-Search’s SWE-Agent operates in a flexible state space, allowing it to fluidly transition between actions such as planning, searching, and editing. This design mirrors the way engineers backtrack and adjust their approach dynamically, ensuring the agent can revise its course when faced with new challenges or information, and points towards the direction of more general, open-ended systems
83
+ (Wang et al.,
84
+ 2023
85
+ ; Ma et al.,
86
+ 2024a
87
+ ; Lu et al.,
88
+ 2024b
89
+ ; Faldor et al.,
90
+ 2024
91
+ ; Hu et al.,
92
+ 2024
93
+ ; Lu et al.,
94
+ 2024a
95
+ )
96
+ .
97
+ Iterative Learning through Feedback
98
+ : Effective engineering relies heavily on continuous testing and refinement. To replicate this, SWE-Search integrates a Monte Carlo Tree Search (MCTS)
99
+ (Silver et al.,
100
+ 2016b
101
+ )
102
+ planning module paired with a Value Agent. The MCTS module balances exploration and exploitation to guide the agent through complex solution spaces. The Value Agent augments this process by providing both utility estimates and qualitative feedback, allowing the agent to iteratively improve its decision-making based on past experiences, similar to how engineers refine their work through feedback and debugging.
103
+ Collaborative Decision-Making
104
+ : Complex problems often benefit from diverse perspectives
105
+ (Khan et al.,
106
+ 2024
107
+ ; Amayuelas et al.,
108
+ 2024
109
+ ; Du et al.,
110
+ 2023
111
+ ; Zhang et al.,
112
+ 2024c
113
+ )
114
+ . In SWE-Search, once a set of potential solutions is generated, the Discriminator Agent facilitates a multi-agent debate. Each agent advocates for different solutions by presenting arguments, which are critically evaluated by a judge agent. This process mirrors real-world engineering collaboration, where teams deliberate to refine and select the most robust solutions.
115
+ The architecture of SWE-Search is designed to automate software engineering tasks through these adaptive, feedback-driven, and collaborative processes. The SWE-Agent serves as the system’s problem solver, operating in a dynamic environment where it can backtrack and adapt its actions as necessary. The MCTS Planning Module efficiently guides exploration and exploitation, ensuring that the agent balances the need for innovation with the need to focus on promising solutions. The Value Agent provides continual feedback, offering both quantitative assessments and qualitative insights, helping the agent refine its strategy iteratively. Finally, the Discriminator Agent ensures that the final decision is rigorously vetted through a multi-agent debate, simulating the collaborative decision-making processes commonly found in engineering teams.
116
+ We evaluate SWE-Search on the SWE-bench benchmark, a comprehensive dataset from real-world open-source repositories. SWE-bench tests agents’ ability to resolve software issues by generating code patches that fix failing tests. SWE-Search demonstrates a
117
+ 23
118
+ %
119
+ percent
120
+ 23
121
+ 23\%
122
+ relative performance improvement across five models compared to standard open-source agents, highlighting the effectiveness of strategic search and iterative self-evaluation. Through detailed analysis, we explore how performance scales with increased search depth and identify key factors that enhance self-assessment in software agents. Our work demonstrates the potential of MCTS and iterative learning to improve agent reasoning and planning in dynamic, complex domains like software engineering, introducing a new paradigm for autonomous software development.
123
+ 2
124
+ Related Work
125
+ Search methods
126
+ Various search approaches have been applied to Large Language Models (LLMs) to facilitate System 2
127
+ (Kahneman,
128
+ 2011
129
+ ; Saha et al.,
130
+ 2024
131
+ ; Pan et al.,
132
+ 2023
133
+ ; Bounsi et al.,
134
+ 2024
135
+ )
136
+ thinking in non-linear reasoning structures. A critical feature of these approaches is their ability to backtrack. Unlike greedy processes
137
+ (Black,
138
+ 2005
139
+ )
140
+ , search algorithms explore multiple branches at each step, potentially escaping paths that lead to dead ends. These methods differ in their strategies for exploring and memorizing possible choices, and in their heuristics for switching between them. Breadth-first search
141
+ (Moore,
142
+ 1959
143
+ )
144
+ maintains all possible search paths, incurring significant memory and computational costs. Depth-first search
145
+ (Cormen et al.,
146
+ 2009
147
+ )
148
+ , in contrast, prioritizes the most promising path in a more greedy manner. When applied to LLMs, these methods demonstrate a trade-off between diversity and quality in text generation
149
+ (Yao et al.,
150
+ 2023
151
+ )
152
+ . The A
153
+ ∗
154
+ algorithm
155
+ (Hart et al.,
156
+ 1968
157
+ )
158
+ combines aspects of breadth-first and greedy search to find optimal solutions using a predetermined evaluation function. In this work, we adopt Monte Carlo Tree Search (MCTS)
159
+ (Silver et al.,
160
+ 2016b
161
+ )
162
+ , an advanced search algorithm that conducts statistical tree search without requiring dedicated evaluation heuristics for each state. MCTS has achieved impressive results in complex strategy games
163
+ (Silver et al.,
164
+ 2016a
165
+ )
166
+ , protein folding
167
+ (Jumper et al.,
168
+ 2021
169
+ )
170
+ , and algorithm discovery
171
+ (Fawzi et al.,
172
+ 2022
173
+ )
174
+ .
175
+ Software Agents
176
+ Software agents are designed to perform autonomous actions within large codebases. Given a repository-level task, these agents typically locate relevant files and code segments before implementing necessary changes. We focus on the SWE-bench task
177
+ (Jimenez et al.,
178
+ 2024
179
+ )
180
+ , which involves resolving real-world GitHub issues. Among the agents with disclosed technical details on SWE-bench,
181
+ Yang et al. (
182
+ 2024b
183
+ )
184
+ introduced the concept of agent-computer interfaces with SWE-agent. OpenDevin
185
+ (Wang et al.,
186
+ 2024b
187
+ )
188
+ presents a collection of community-driven agents, including CodeAct
189
+ (Wang et al.,
190
+ 2024a
191
+ )
192
+ . The Agentless approach demonstrated competitive performance using a simple two-step process of localization and repair. AutoCodeRover
193
+ (Zhang et al.,
194
+ 2024d
195
+ )
196
+ incorporated advanced code tools such as abstract syntax trees and spectrum-based fault localization. The Alibaba Lingma Agent
197
+ (Ma et al.,
198
+ 2024b
199
+ )
200
+ introduced a search-based approach for repository exploration, followed by a structured editing phase. While effective, it constitutes a more hand-designed solution specifically designed to interface with the search functionality of their agent.
201
+ 3
202
+ Methodology
203
+ SWE-Search is a multi-agent system designed to tackle complex software engineering tasks by integrating dynamic planning, value estimation, and deliberative decision-making. The core motivation behind this method is to emulate the sophisticated, iterative workflows of human software engineers, where exploration, planning, and collaboration are crucial to solving intricate problems. By leveraging the strengths of Monte Carlo Tree Search (MCTS) for planning, a Value Agent for utility estimation and feedback, and a Discriminator Agent for final decision-making through debate, SWE-Search provides a comprehensive, adaptive framework capable of navigating and solving real-world software engineering challenges.
204
+ SWE-Search consists of four primary components that work in synergy:
205
+ SWE-Search Framework and Action Agent
206
+ : Building on the moatless-tools framework
207
+ (Örwall,
208
+ 2024
209
+ )
210
+ , SWE-Search operates in a dynamic code environment with a flexible state-space and a git-like commit tree structure. This design facilitates efficient backtracking to previous states, enabling the Action Agent to explore diverse solution trajectories. The adaptable state-space enhances the system’s ability to exploit the MCTS module effectively.
211
+ Search Algorithm
212
+ : The core of SWE-Search’s exploration strategy is based on a Monte Carlo Tree Search (MCTS) which uses a heuristic-based selection process similar to AlphaZero
213
+ (Silver et al.,
214
+ 2016a
215
+ )
216
+ , specifically tailored for software engineering tasks. This modified MCTS algorithm effectively balances exploration and exploitation, helping the agent explore a diverse set of solutions and converge quickly on the most promising strategies.
217
+ Value (Function) Agent
218
+ : To approximate the utility of each observation, we employ an LLM-based value function, which in addition to outputting a value, also generates an explanation in natural language. This explanation can be leveraged to improve subsequent actions from parent nodes, enabling iterative self-improvement of the search process.
219
+ Discriminator Agent
220
+ : In the final stage of SWE-Search, the Discriminator Agent evaluates the solutions generated by the search process. Inspired by multi-agent debate frameworks
221
+ Du et al. (
222
+ 2023
223
+ ); Khan et al. (
224
+ 2024
225
+ ); Amayuelas et al. (
226
+ 2024
227
+ )
228
+ , this agent engages in a structured debate, where multiple agents argue for or against the proposed solutions. The debate process not only surfaces diverse perspectives but also leads to a more rigorously justified final decision.
229
+ This system architecture combines the strengths of dynamic action selection, strategic planning, and collaborative deliberation, creating a comprehensive tool capable of handling the complexity and iterative nature of software engineering tasks.
230
+ 3.1
231
+ Problem Formulation
232
+ Figure 1:
233
+ SWE-Search Overview.
234
+ Tree search.
235
+ Each state is represented as a node from which the agent can expand from, and each corresponding action is presented as an edge.
236
+ Evaluation.
237
+ Uses all relevant context including trajectory information, file context, and executed tests, to provide a quantitative value estimation and qualitative explanation in natural language.
238
+ Expansion.
239
+ Nodes can be expanded using value function feedback from future actions.
240
+ The task of the SWE agent can be formalized as a tuple
241
+ ℳ
242
+ =
243
+ (
244
+ 𝒮
245
+ ,
246
+ 𝒞
247
+ ,
248
+ 𝒜
249
+ ,
250
+ 𝒱
251
+ ,
252
+ 𝒫
253
+ ,
254
+ p
255
+ 0
256
+ ,
257
+ ρ
258
+ )
259
+ ℳ
260
+ 𝒮
261
+ 𝒞
262
+ 𝒜
263
+ 𝒱
264
+ 𝒫
265
+ subscript
266
+ 𝑝
267
+ 0
268
+ 𝜌
269
+ \mathcal{M}=(\mathcal{S},\mathcal{C},\mathcal{A},\mathcal{V},\mathcal{P},p_{0},\rho)
270
+ . Here,
271
+ 𝒮
272
+ 𝒮
273
+ \mathcal{S}
274
+ represents the state space, encompassing all possible states such as the current context of the files the agent is working on and the overall status of the codebase. The context space, denoted as
275
+ 𝒞
276
+ 𝒞
277
+ \mathcal{C}
278
+ , includes metadata about the repository and the initial problem description. The value function
279
+ 𝒱
280
+ 𝒱
281
+ \mathcal{V}
282
+ assigns a utility score to each state-action pair
283
+ O
284
+ ​
285
+ (
286
+ a
287
+ ,
288
+ t
289
+ )
290
+ 𝑂
291
+ 𝑎
292
+ 𝑡
293
+ O(a,t)
294
+ , guiding the agent’s decisions.
295
+ The environment’s dynamics are defined by a context-dependent transition function
296
+ 𝒫
297
+ :
298
+ 𝒮
299
+ ×
300
+ 𝒜
301
+ ×
302
+ 𝒞
303
+ →
304
+ Δ
305
+ ​
306
+ (
307
+ 𝒮
308
+ )
309
+ :
310
+ 𝒫
311
+ →
312
+ 𝒮
313
+ 𝒜
314
+ 𝒞
315
+ Δ
316
+ 𝒮
317
+ \mathcal{P}:\mathcal{S}\times\mathcal{A}\times\mathcal{C}\rightarrow\Delta(\mathcal{S})
318
+ , which models the evolution of the repository’s state after each action. The initial state distribution,
319
+ p
320
+ 0
321
+ :
322
+ 𝒞
323
+ →
324
+ Δ
325
+ ​
326
+ (
327
+ 𝒮
328
+ )
329
+ :
330
+ subscript
331
+ 𝑝
332
+ 0
333
+ →
334
+ 𝒞
335
+ Δ
336
+ ���
337
+ p_{0}:\mathcal{C}\rightarrow\Delta(\mathcal{S})
338
+ , specifies how the initial state depends on the given context, while
339
+ ρ
340
+ ∈
341
+ Δ
342
+ ​
343
+ (
344
+ 𝒞
345
+ )
346
+ 𝜌
347
+ Δ
348
+ 𝒞
349
+ \rho\in\Delta(\mathcal{C})
350
+ defines the distribution over contexts.
351
+ Given an initial context
352
+ c
353
+ ∼
354
+ ρ
355
+ similar-to
356
+ 𝑐
357
+ 𝜌
358
+ c\sim\rho
359
+ and an initial state
360
+ s
361
+ 0
362
+ ∼
363
+ p
364
+ 0
365
+ (
366
+ ⋅
367
+ ∣
368
+ c
369
+ )
370
+ s_{0}\sim p_{0}(\cdot\mid c)
371
+ , the SWE agent executes its policy
372
+ π
373
+ :
374
+ 𝒮
375
+ ×
376
+ 𝒞
377
+ →
378
+ Δ
379
+ ​
380
+ (
381
+ 𝒜
382
+ )
383
+ :
384
+ 𝜋
385
+ →
386
+ 𝒮
387
+ 𝒞
388
+ Δ
389
+ 𝒜
390
+ \pi:\mathcal{S}\times\mathcal{C}\rightarrow\Delta(\mathcal{A})
391
+ , which selects actions based on the current state and context. At each time step
392
+ t
393
+ 𝑡
394
+ t
395
+ , the agent takes an action
396
+ a
397
+ t
398
+ ∼
399
+ π
400
+ ​
401
+ (
402
+ s
403
+ t
404
+ ,
405
+ c
406
+ )
407
+ similar-to
408
+ subscript
409
+ 𝑎
410
+ 𝑡
411
+ 𝜋
412
+ subscript
413
+ 𝑠
414
+ 𝑡
415
+ 𝑐
416
+ a_{t}\sim\pi(s_{t},c)
417
+ and receives a corresponding reward
418
+ ℛ
419
+ ​
420
+ (
421
+ s
422
+ t
423
+ ,
424
+ a
425
+ t
426
+ ,
427
+ c
428
+ )
429
+ ℛ
430
+ subscript
431
+ 𝑠
432
+ 𝑡
433
+ subscript
434
+ 𝑎
435
+ 𝑡
436
+ 𝑐
437
+ \mathcal{R}(s_{t},a_{t},c)
438
+ . The environment then transitions to a new state
439
+ s
440
+ t
441
+ +
442
+ 1
443
+ ∼
444
+ 𝒫
445
+ (
446
+ ⋅
447
+ ∣
448
+ s
449
+ t
450
+ ,
451
+ a
452
+ t
453
+ ,
454
+ c
455
+ )
456
+ s_{t+1}\sim\mathcal{P}(\cdot\mid s_{t},a_{t},c)
457
+ , and the agent continues to observe this updated state. Over time, this process generates a trajectory
458
+ τ
459
+ :=
460
+ {
461
+ s
462
+ t
463
+ ,
464
+ a
465
+ t
466
+ ,
467
+ r
468
+ t
469
+ }
470
+ t
471
+ =
472
+ 0
473
+ T
474
+ assign
475
+ 𝜏
476
+ superscript
477
+ subscript
478
+ subscript
479
+ 𝑠
480
+ 𝑡
481
+ subscript
482
+ 𝑎
483
+ 𝑡
484
+ subscript
485
+ 𝑟
486
+ 𝑡
487
+ 𝑡
488
+ 0
489
+ 𝑇
490
+ \tau:=\{s_{t},a_{t},r_{t}\}_{t=0}^{T}
491
+ as the agent interacts with the environment.
492
+ The agent’s objective is to maximize the cumulative reward over the trajectory, which is captured by the value function
493
+ v
494
+ ​
495
+ (
496
+ s
497
+ t
498
+ ,
499
+ a
500
+ t
501
+ ,
502
+ {
503
+ s
504
+ i
505
+ }
506
+ i
507
+ =
508
+ 0
509
+ t
510
+ −
511
+ 1
512
+ ,
513
+ {
514
+ a
515
+ i
516
+ }
517
+ i
518
+ =
519
+ 0
520
+ t
521
+ −
522
+ 1
523
+ )
524
+ 𝑣
525
+ subscript
526
+ 𝑠
527
+ 𝑡
528
+ subscript
529
+ 𝑎
530
+ 𝑡
531
+ superscript
532
+ subscript
533
+ subscript
534
+ 𝑠
535
+ 𝑖
536
+ 𝑖
537
+ 0
538
+ 𝑡
539
+ 1
540
+ superscript
541
+ subscript
542
+ subscript
543
+ 𝑎
544
+ 𝑖
545
+ 𝑖
546
+ 0
547
+ 𝑡
548
+ 1
549
+ v(s_{t},a_{t},\{s_{i}\}_{i=0}^{t-1},\{a_{i}\}_{i=0}^{t-1})
550
+ . This value function depends not only on the current state and action but also on the history of previous states and actions, which deviates from the assumptions of a Markovian process. Formally, the agent seeks to maximize the expected cumulative reward, defined as:
551
+ max
552
+ π
553
+ ⁡
554
+ V
555
+ T
556
+ ​
557
+ (
558
+ ρ
559
+ )
560
+ =
561
+ max
562
+ π
563
+ ⁡
564
+ 𝔼
565
+ ​
566
+ τ
567
+ ​
568
+ [
569
+ ∑
570
+ t
571
+ =
572
+ 0
573
+ T
574
+ ℛ
575
+ ​
576
+ (
577
+ s
578
+ t
579
+ ,
580
+ a
581
+ t
582
+ ,
583
+ c
584
+ )
585
+ ∣
586
+ c
587
+ ∼
588
+ ρ
589
+ ;
590
+ π
591
+ ]
592
+ subscript
593
+ 𝜋
594
+ superscript
595
+ 𝑉
596
+ 𝑇
597
+ 𝜌
598
+ subscript
599
+ 𝜋
600
+ 𝔼
601
+ 𝜏
602
+ delimited-[]
603
+ similar-to
604
+ conditional
605
+ superscript
606
+ subscript
607
+ 𝑡
608
+ 0
609
+ 𝑇
610
+ ℛ
611
+ subscript
612
+ 𝑠
613
+ 𝑡
614
+ subscript
615
+ 𝑎
616
+ 𝑡
617
+ 𝑐
618
+ 𝑐
619
+ 𝜌
620
+ 𝜋
621
+ \max_{\pi}V^{T}(\rho)=\max_{\pi}\mathbb{E}{\tau}\left[\sum_{t=0}^{T}\mathcal{R}(s_{t},a_{t},c)\mid c\sim\rho;\pi\right]
622
+ .
623
+ This optimization captures the agent’s (in-context) process, as it adjusts its policy
624
+ π
625
+ 𝜋
626
+ \pi
627
+ to achieve the highest expected return across multiple trajectories, considering both current and historical information.
628
+ 3.2
629
+ SWE-Search Framework and Action Agent
630
+ The SWE-Search Action Agent builds on the moatless-tools framework
631
+ (Örwall,
632
+ 2024
633
+ )
634
+ . Its action space,
635
+ 𝒜
636
+ 𝒜
637
+ \mathcal{A}
638
+ , is organized as a two-tier hierarchy, comprising both action types and their corresponding specific actions. Formally, this can be expressed as
639
+ 𝒜
640
+ =
641
+ (
642
+ t
643
+ ,
644
+ a
645
+ )
646
+ ∣
647
+ t
648
+ ∈
649
+ 𝒯
650
+ ,
651
+ a
652
+ ∈
653
+ 𝒜
654
+ t
655
+ formulae-sequence
656
+ 𝒜
657
+ conditional
658
+ 𝑡
659
+ 𝑎
660
+ 𝑡
661
+ 𝒯
662
+ 𝑎
663
+ subscript
664
+ 𝒜
665
+ 𝑡
666
+ \mathcal{A}={(t,a)\mid t\in\mathcal{T},a\in\mathcal{A}_{t}}
667
+ , where
668
+ 𝒯
669
+ 𝒯
670
+ \mathcal{T}
671
+ represents the set of action types (e.g.,
672
+ Search
673
+ ,
674
+ Plan
675
+ ,
676
+ Edit
677
+ ), and
678
+ 𝒜
679
+ t
680
+ subscript
681
+ 𝒜
682
+ 𝑡
683
+ \mathcal{A}_{t}
684
+ is the set of possible actions corresponding to each type
685
+ t
686
+ 𝑡
687
+ t
688
+ . These actions range from tool invocations and code modifications to the generation of structured text. To enhance the agent’s effectiveness in search-driven tasks, we introduced the following modifications:
689
+ One key modification we implemented is the expansion of the
690
+ Plan
691
+ state, allowing it to transition flexibly to any other state, rather than being limited to transitioning only to
692
+ Edit
693
+ . This change is motivated by the need to enable more dynamic and adaptive problem-solving behaviors within the agent. In the context of software engineering, rigid state transitions can be overly restrictive, forcing the agent into predetermined pathways that may not always align with the complexities of real-world scenarios. For instance, during code modification tasks, an agent might recognize mid-process that further planning, additional searches, or different types of analysis are necessary before proceeding with edits. Restricting transitions only to editing would artificially constrain the agent, potentially leading it to suboptimal actions or causing it to become stuck in unproductive loops. By allowing transitions to any state, we empower the agent to adapt to new information as it arises (
694
+ Fig.
695
+ 2
696
+ ), exploring a wider variety of trajectories. This enhanced flexibility reflects the iterative and often non-linear nature of real software engineering workflows, where engineers frequently revisit planning, testing, and research phases before committing to edits.
697
+ Second, the agent is empowered to execute any tests within the codebase at its discretion, as well as to create and implement new tests. The results of these tests are incorporated into both the value function and the agent’s subsequent decision-making process. It is crucial to highlight that the tests required to resolve a given instance (i.e., fail-to-pass tests) are not explicitly revealed to the agent. However, the agent can leverage any pre-existing tests within the repository, simulating the behavior of a real-world software engineer.
698
+ 1
699
+ 1
700
+ 1
701
+ This approach aligns with the practices of other SWE agents, and has been validated by the authors of SWE-bench, who confirmed its legitimacy as long as the fail-to-pass tests remain concealed from the model.
702
+ 3.3
703
+ Value (Function) Agent
704
+ The role of the Value Agent extends beyond simply estimating the expected utility of a given state-action pair
705
+ O
706
+ n
707
+ ​
708
+ (
709
+ s
710
+ n
711
+ ,
712
+ a
713
+ n
714
+ )
715
+ subscript
716
+ 𝑂
717
+ 𝑛
718
+ subscript
719
+ 𝑠
720
+ 𝑛
721
+ subscript
722
+ 𝑎
723
+ 𝑛
724
+ O_{n}(s_{n},a_{n})
725
+ . In addition to calculating the value
726
+ v
727
+ n
728
+ subscript
729
+ 𝑣
730
+ 𝑛
731
+ v_{n}
732
+ , the Value Agent generates a written explanation, denoted as
733
+ ε
734
+ 𝜀
735
+ \varepsilon
736
+ . This explanation serves a dual purpose: it provides transparency into the decision-making process and functions as feedback for the Action Agent, which can leverage this explanation when re-expanding from the parent node of
737
+ O
738
+ n
739
+ subscript
740
+ 𝑂
741
+ 𝑛
742
+ O_{n}
743
+ (see
744
+ Figure
745
+ 1
746
+ ,
747
+ hindsight feedback
748
+ ). This approach enables the system to iteratively refine its decision-making process, mirroring how a human software engineer continuously re-evaluates their approach based on new information to improve their problem-solving strategy.
749
+ The input to the value function consists of all state-action pairs up to and including the current state being evaluated, alongside specific instructions on how to assess the state. This allows the Value Agent to contextualize the decision within the trajectory, accounting for the sequence of actions and states leading up to the present. The final output of the value function can be formalized as:
750
+ (
751
+ v
752
+ t
753
+ ,
754
+ ε
755
+ t
756
+ )
757
+ =
758
+ V
759
+ ​
760
+ (
761
+ s
762
+ t
763
+ ,
764
+ a
765
+ t
766
+ ,
767
+ {
768
+ s
769
+ i
770
+ }
771
+ i
772
+ =
773
+ 0
774
+ ​
775
+ …
776
+ ​
777
+ t
778
+ −
779
+ 1
780
+ ,
781
+ {
782
+ a
783
+ i
784
+ }
785
+ i
786
+ =
787
+ 0
788
+ ​
789
+ …
790
+ ​
791
+ t
792
+ −
793
+ 1
794
+ )
795
+ subscript
796
+ 𝑣
797
+ 𝑡
798
+ subscript
799
+ 𝜀
800
+ 𝑡
801
+ 𝑉
802
+ subscript
803
+ 𝑠
804
+ 𝑡
805
+ subscript
806
+ 𝑎
807
+ 𝑡
808
+ subscript
809
+ subscript
810
+ 𝑠
811
+ 𝑖
812
+ 𝑖
813
+ 0
814
+ …
815
+ 𝑡
816
+ 1
817
+ subscript
818
+ subscript
819
+ 𝑎
820
+ 𝑖
821
+ 𝑖
822
+ 0
823
+ …
824
+ 𝑡
825
+ 1
826
+ (v_{t},\varepsilon_{t})=V(s_{t},a_{t},\{s_{i}\}_{i=0{\ldots}t-1},\{a_{i}\}_{i=0{\ldots}t-1})
827
+ (1)
828
+ Here,
829
+ v
830
+ t
831
+ subscript
832
+ 𝑣
833
+ 𝑡
834
+ v_{t}
835
+ represents the expected utility of the current state-action pair, while
836
+ ε
837
+ t
838
+ subscript
839
+ 𝜀
840
+ 𝑡
841
+ \varepsilon_{t}
842
+ is the accompanying explanation.
843
+ In practice, the Value Agent is tasked with analyzing the entire trajectory leading up to the current state-action pair, providing not only the required utility estimate
844
+ v
845
+ t
846
+ subscript
847
+ 𝑣
848
+ 𝑡
849
+ v_{t}
850
+ , but also a detailed explanation
851
+ ε
852
+ t
853
+ subscript
854
+ 𝜀
855
+ 𝑡
856
+ \varepsilon_{t}
857
+ . This explanation is critical for the agent’s overall performance, as it offers insight into the reasoning behind utility estimates, which in turn informs the Action Agent’s future decisions. We have observed that one of the key factors driving the effectiveness of the Value Agent lies in the clarity and specificity of these explanations. A well-articulated explanation can illuminate the strengths and limitations of different state types (e.g.,
858
+ Search
859
+ ,
860
+ Edit
861
+ ,
862
+ Plan
863
+ ), helping the Action Agent better understand which types of states are more promising or risky to pursue.
864
+ By providing detailed feedback on the potential utility of different actions and contextualizing them within the broader trajectory, the Value Agent enables more informed and strategic decision-making by the Action Agent. This integration of both quantitative and qualitative feedback leads to improved performance and more adaptive behavior throughout the task (
865
+ Fig.
866
+ 4
867
+ a
868
+ ).
869
+ 3.4
870
+ Search Algorithm
871
+ Our search tree is structured with nodes representing states
872
+ 𝒮
873
+ t
874
+ subscript
875
+ 𝒮
876
+ 𝑡
877
+ \mathcal{S}_{t}
878
+ and edges representing actions
879
+ 𝒜
880
+ t
881
+ subscript
882
+ 𝒜
883
+ 𝑡
884
+ \mathcal{A}_{t}
885
+ . The search algorithm employed is a modified Monte Carlo Tree Search (MCTS), specifically adapted for the tasks of the SWE-Agent. Unlike prior approaches for web agents that utilize language models in the selection process
886
+ Koh et al. (
887
+ 2024
888
+ ); Zhang et al. (
889
+ 2024b
890
+ )
891
+ , we deliberately choose not to rely on language models for node selection. Instead, we adopt a more straightforward heuristic-based selection function, similar to the approach used in AlphaZero
892
+ Silver et al. (
893
+ 2016a
894
+ ;
895
+ 2018
896
+ )
897
+ . This decision is driven by the need for interpretability, efficiency, and the focus on tasks where heuristic-based exploration suffices to guide the agent effectively through complex software engineering environments.
898
+ At the core of our algorithm is a modified Upper Confidence Bound for Trees (UCT) selection criterion
899
+ Kocsis & Szepesvári (
900
+ 2006
901
+ )
902
+ , which determines the next node to expand. This criterion balances exploitation of known high-reward actions with exploration of less-visited states. We introduce additional terms to encourage strategic exploration early in the search process, and to penalize over-exploration at later stages when convergence on the optimal solution is desired. The modified UCT function is expressed as:
903
+ U
904
+ ​
905
+ C
906
+ ​
907
+ T
908
+ ​
909
+ (
910
+ s
911
+ ,
912
+ a
913
+ )
914
+ =
915
+ e
916
+ ​
917
+ x
918
+ ​
919
+ p
920
+ ​
921
+ l
922
+ ​
923
+ o
924
+ ​
925
+ i
926
+ ​
927
+ t
928
+ ​
929
+ a
930
+ ​
931
+ t
932
+ ​
933
+ i
934
+ ​
935
+ o
936
+ ​
937
+ n
938
+ +
939
+ e
940
+ ​
941
+ x
942
+ ​
943
+ p
944
+ ​
945
+ l
946
+ ​
947
+ o
948
+ ​
949
+ r
950
+ ​
951
+ a
952
+ ​
953
+ t
954
+ ​
955
+ i
956
+ ​
957
+ o
958
+ ​
959
+ n
960
+ +
961
+ e
962
+ ​
963
+ a
964
+ ​
965
+ r
966
+ ​
967
+ l
968
+ ​
969
+ y
970
+ ​
971
+ _
972
+ ​
973
+ d
974
+ ​
975
+ e
976
+ ​
977
+ p
978
+ ​
979
+ t
980
+ ​
981
+ h
982
+ ​
983
+ _
984
+ ​
985
+ b
986
+ ​
987
+ o
988
+ ​
989
+ n
990
+ ​
991
+ u
992
+ ​
993
+ s
994
+ −
995
+ l
996
+ ​
997
+ a
998
+ ​
999
+ t
1000
+ ​
1001
+ e
1002
+ ​
1003
+ _
1004
+ ​
1005
+ d
1006
+ ​
1007
+ e
1008
+ ​
1009
+ p
1010
+ ​
1011
+ t
1012
+ ​
1013
+ h
1014
+ ​
1015
+ _
1016
+ ​
1017
+ p
1018
+ ​
1019
+ e
1020
+ ​
1021
+ n
1022
+ ​
1023
+ a
1024
+ ​
1025
+ l
1026
+ ​
1027
+ t
1028
+ ​
1029
+ y
1030
+ 𝑈
1031
+ 𝐶
1032
+ 𝑇
1033
+ 𝑠
1034
+ 𝑎
1035
+ 𝑒
1036
+ 𝑥
1037
+ 𝑝
1038
+ 𝑙
1039
+ 𝑜
1040
+ 𝑖
1041
+ 𝑡
1042
+ 𝑎
1043
+ 𝑡
1044
+ 𝑖
1045
+ 𝑜
1046
+ 𝑛
1047
+ 𝑒
1048
+ 𝑥
1049
+ 𝑝
1050
+ 𝑙
1051
+ 𝑜
1052
+ 𝑟
1053
+ 𝑎
1054
+ 𝑡
1055
+ 𝑖
1056
+ 𝑜
1057
+ 𝑛
1058
+ 𝑒
1059
+ 𝑎
1060
+ 𝑟
1061
+ 𝑙
1062
+ 𝑦
1063
+ _
1064
+ 𝑑
1065
+ 𝑒
1066
+ 𝑝
1067
+ 𝑡
1068
+ ℎ
1069
+ _
1070
+ 𝑏
1071
+ 𝑜
1072
+ 𝑛
1073
+ 𝑢
1074
+ 𝑠
1075
+ 𝑙
1076
+ 𝑎
1077
+ 𝑡
1078
+ 𝑒
1079
+ _
1080
+ 𝑑
1081
+ 𝑒
1082
+ 𝑝
1083
+ 𝑡
1084
+ ℎ
1085
+ _
1086
+ 𝑝
1087
+ 𝑒
1088
+ 𝑛
1089
+ 𝑎
1090
+ 𝑙
1091
+ 𝑡
1092
+ 𝑦
1093
+ UCT(s,a)=exploitation+exploration+early\_depth\_bonus-late\_depth\_penalty
1094
+ (2)
1095
+ This can be expressed more formally as:
1096
+ U
1097
+ ​
1098
+ C
1099
+ ​
1100
+ T
1101
+ ​
1102
+ (
1103
+ s
1104
+ ,
1105
+ a
1106
+ )
1107
+ =
1108
+ V
1109
+ ​
1110
+ (
1111
+ s
1112
+ ,
1113
+ a
1114
+ )
1115
+ +
1116
+ C
1117
+ ​
1118
+ ln
1119
+ ⁡
1120
+ N
1121
+ ​
1122
+ (
1123
+ s
1124
+ )
1125
+ N
1126
+ ​
1127
+ (
1128
+ s
1129
+ ,
1130
+ a
1131
+ )
1132
+ +
1133
+ α
1134
+ ​
1135
+ e
1136
+ −
1137
+ β
1138
+ ​
1139
+ (
1140
+ d
1141
+ −
1142
+ 1
1143
+ )
1144
+ −
1145
+ γ
1146
+ ​
1147
+ d
1148
+ 𝑈
1149
+ 𝐶
1150
+ 𝑇
1151
+ 𝑠
1152
+ 𝑎
1153
+ 𝑉
1154
+ 𝑠
1155
+ 𝑎
1156
+ 𝐶
1157
+ 𝑁
1158
+ 𝑠
1159
+ 𝑁
1160
+ 𝑠
1161
+ 𝑎
1162
+ 𝛼
1163
+ superscript
1164
+ 𝑒
1165
+ 𝛽
1166
+ 𝑑
1167
+ 1
1168
+ 𝛾
1169
+ 𝑑
1170
+ UCT(s,a)=V(s,a)+C\sqrt{\frac{\ln N(s)}{N(s,a)}}+\alpha e^{-\beta(d-1)}-\gamma\sqrt{d}
1171
+ (3)
1172
+ V
1173
+ ​
1174
+ (
1175
+ s
1176
+ ,
1177
+ a
1178
+ )
1179
+ 𝑉
1180
+ 𝑠
1181
+ 𝑎
1182
+ V(s,a)
1183
+ is the value estimate of the state-action pair
1184
+ ,
1185
+ N
1186
+ ​
1187
+ (
1188
+ s
1189
+ ,
1190
+ a
1191
+ )
1192
+ 𝑁
1193
+ 𝑠
1194
+ 𝑎
1195
+ N(s,a)
1196
+ is the number of times the state-action pair
1197
+ (
1198
+ s
1199
+ ,
1200
+ a
1201
+ )
1202
+ 𝑠
1203
+ 𝑎
1204
+ (s,a)
1205
+ has been visited,
1206
+ N
1207
+ ​
1208
+ (
1209
+ s
1210
+ )
1211
+ 𝑁
1212
+ 𝑠
1213
+ N(s)
1214
+ is the visit count of state
1215
+ s
1216
+ 𝑠
1217
+ s
1218
+ ,
1219
+ d
1220
+ 𝑑
1221
+ d
1222
+ is the depth of the node in the search tree, and
1223
+ C
1224
+ 𝐶
1225
+ C
1226
+ ,
1227
+ α
1228
+ 𝛼
1229
+ \alpha
1230
+ ,
1231
+ β
1232
+ 𝛽
1233
+ \beta
1234
+ , and
1235
+ γ
1236
+ 𝛾
1237
+ \gamma
1238
+ are constants that control the balance between exploration, exploitation, and depth-dependent rewards and penalties.
1239
+ This formulation is inspired by the way software engineers explore potential solutions to a task. In practice, an engineer’s search process can be broken down into the following key phases, which our algorithm mirrors:
1240
+ Early Exploration
1241
+ : Initially, an engineer explores a wide variety of potential approaches to fully understand the problem and identify promising strategies. This is encouraged in our algorithm by the
1242
+ e
1243
+ ​
1244
+ a
1245
+ ​
1246
+ r
1247
+ ​
1248
+ l
1249
+ ​
1250
+ y
1251
+ ​
1252
+ _
1253
+ ​
1254
+ d
1255
+ ​
1256
+ e
1257
+ ​
1258
+ p
1259
+ ​
1260
+ t
1261
+ ​
1262
+ h
1263
+ ​
1264
+ _
1265
+ ​
1266
+ b
1267
+ ​
1268
+ o
1269
+ ​
1270
+ n
1271
+ ​
1272
+ u
1273
+ ​
1274
+ s
1275
+ 𝑒
1276
+ 𝑎
1277
+ 𝑟
1278
+ 𝑙
1279
+ 𝑦
1280
+ _
1281
+ 𝑑
1282
+ 𝑒
1283
+ 𝑝
1284
+ 𝑡
1285
+ ℎ
1286
+ _
1287
+ 𝑏
1288
+ 𝑜
1289
+ 𝑛
1290
+ 𝑢
1291
+ 𝑠
1292
+ early\_depth\_bonus
1293
+ , represented by the term
1294
+ α
1295
+ ​
1296
+ e
1297
+ −
1298
+ β
1299
+ ​
1300
+ (
1301
+ d
1302
+ −
1303
+ 1
1304
+ )
1305
+ 𝛼
1306
+ superscript
1307
+ 𝑒
1308
+ 𝛽
1309
+ 𝑑
1310
+ 1
1311
+ \alpha e^{-\beta(d-1)}
1312
+ , which rewards exploration at shallow depths, simulating the early phases of wide exploration.
1313
+ Convergence and Exploitation
1314
+ : As the engineer gains more information and narrows down the options, the focus shifts to exploiting the most effective solution paths. This transition is handled by the standard UCT exploitation term
1315
+ V
1316
+ ​
1317
+ (
1318
+ s
1319
+ ,
1320
+ a
1321
+ )
1322
+ 𝑉
1323
+ 𝑠
1324
+ 𝑎
1325
+ V(s,a)
1326
+ and is further reinforced by the
1327
+ l
1328
+ ​
1329
+ a
1330
+ ​
1331
+ t
1332
+ ​
1333
+ e
1334
+ ​
1335
+ _
1336
+ ​
1337
+ d
1338
+ ​
1339
+ e
1340
+ ​
1341
+ p
1342
+ ​
1343
+ t
1344
+ ​
1345
+ h
1346
+ ​
1347
+ _
1348
+ ​
1349
+ p
1350
+ ​
1351
+ e
1352
+ ​
1353
+ n
1354
+ ​
1355
+ a
1356
+ ​
1357
+ l
1358
+ ​
1359
+ t
1360
+ ​
1361
+ y
1362
+ 𝑙
1363
+ 𝑎
1364
+ 𝑡
1365
+ 𝑒
1366
+ _
1367
+ 𝑑
1368
+ 𝑒
1369
+ 𝑝
1370
+ 𝑡
1371
+ ℎ
1372
+ _
1373
+ 𝑝
1374
+ 𝑒
1375
+ 𝑛
1376
+ 𝑎
1377
+ 𝑙
1378
+ 𝑡
1379
+ 𝑦
1380
+ late\_depth\_penalty
1381
+ (
1382
+ −
1383
+ γ
1384
+ ​
1385
+ d
1386
+ 𝛾
1387
+ 𝑑
1388
+ -\gamma\sqrt{d}
1389
+ ), which discourages over-exploration as the agent delves deeper into the search tree.
1390
+ Quick Abandonment of Poor Strategies
1391
+ : Software engineers are also adept at abandoning poor strategies when new information indicates that a particular approach is not viable. We capture this behavior by implementing a simple heuristic rule that abandons nodes associated with consecutive low rewards, ensuring that the agent does not waste resources on unproductive trajectories.
1392
+ At each step, the node with the highest UCT value is selected for expansion, formalized as:
1393
+ s
1394
+ ∗
1395
+ =
1396
+ arg
1397
+ ​
1398
+ max
1399
+ (
1400
+ s
1401
+ ,
1402
+ a
1403
+ )
1404
+ ⁡
1405
+ U
1406
+ ​
1407
+ C
1408
+ ​
1409
+ T
1410
+ ​
1411
+ (
1412
+ s
1413
+ ,
1414
+ a
1415
+ )
1416
+ superscript
1417
+ 𝑠
1418
+ subscript
1419
+ arg
1420
+ max
1421
+ 𝑠
1422
+ 𝑎
1423
+ 𝑈
1424
+ 𝐶
1425
+ 𝑇
1426
+ 𝑠
1427
+ 𝑎
1428
+ s^{*}=\operatorname*{arg\,max}_{(s,a)}UCT(s,a)
1429
+ (4)
1430
+ This approach effectively mimics the decision-making process of a software engineer, who balances exploration of potential strategies with a focus on converging towards the optimal solution, while remaining flexible enough to backtrack when necessary. By incorporating heuristic feedback and depth-based adjustments, the algorithm avoids getting stuck in unproductive paths and enhances the agent’s ability to identify high-reward strategies with minimal computational overhead
1431
+ Appendix
1432
+ 6
1433
+ .
1434
+ 3.4.1
1435
+ Discriminator Agent
1436
+ The final stage of SWE-Search involves the Discriminator Agent, whose role is to evaluate the candidate solutions generated by the search process and select the one most likely to resolve the issue at hand. This module accepts up to five final solutions produced by the search and engages in a multi-agent debate to determine the most promising option. Drawing inspiration from recent work on persuasive multi-agent debates
1437
+ (Khan et al.,
1438
+ 2024
1439
+ ; Amayuelas et al.,
1440
+ 2024
1441
+ )
1442
+ , the Discriminator leverages the collective reasoning of multiple agents to ensure a more robust final selection. Configuration and hyperparameter details can be found in
1443
+ Table
1444
+ 2
1445
+ .
1446
+ In this stage, agents are presented with the original problem statement and candidate solutions. They engage in a structured debate to determine the most effective solution, supporting their choices with logical reasoning and evidence from the search process. This debate encourages a thorough exploration of trade-offs between solutions, potentially uncovering strengths or weaknesses not evident during individual searches. Finally, a judge agent evaluates the arguments and selects the solution deemed most likely to resolve the issue. This process simulates the collaborative decision-making in software engineering teams, where diverse perspectives lead to a more thorough evaluation of candidate solutions, ultimately increasing the likelihood of identifying the most optimal outcome.
1447
+ The discriminator process not only enhances the robustness of the final solution but also adds transparency, as the reasoning behind the choice is clearly articulated and evaluated. This ensures that the selected solution is well-reasoned and thoroughly vetted before implementation.
1448
+ Figure 2:
1449
+ Hindsight feedback error correction.
1450
+ Instance sympy__sympy-15678, SWE-Search with Qwen2.5-72B-Instruct. Initially, the Action Agent performs edits and runs tests, which pass. It prematurely concludes the search. Without actually knowing the proposed solution does not resolve the issue, the Value Agent identifies potentially missed tests and assigns a low reward. Upon re-expansion using the Value Agent’s feedback, new tests fail, prompting the Action Agent to make additional edits, which result in a preferred solution which ultimately resolves the issue.
1451
+ 4
1452
+ Experiments
1453
+ Benchmark
1454
+ For our experiments, we utilize SWE-bench Lite, a curated subset of the official SWE-bench, containing 300 instances. This dataset is specifically designed to be self-contained and focuses primarily on evaluating functional bug fixes, providing a controlled environment to assess the performance of our system.
1455
+ Evaluation Metrics
1456
+ We use two metrics: resolve rate (
1457
+ Pass@1
1458
+ ) and
1459
+ Pass@5
1460
+ . Resolve rate is the percentage of issues successfully resolved, measuring overall effectiveness. Pass@5 is the percentage of issues where a correct solution is found within five attempts. This allows us to assess the efficiency of the search in identifying successful bug fixes within a limited number of iterations.
1461
+ Baselines
1462
+ Software agents leverage diverse tools, architectures, and models, leading to variability in their performance on subsets of the SWE-bench Lite dataset
1463
+ (Zhang et al.,
1464
+ 2024a
1465
+ )
1466
+ . For comparison, we build upon the moatless-tools framework
1467
+ (Örwall,
1468
+ 2024
1469
+ )
1470
+ , a high-performing open-source agent commonly used in research settings
1471
+ (Chowdhury et al.,
1472
+ 2024
1473
+ )
1474
+ . To isolate the impact of our search approach, we adapt moatless-tools as our baseline, referred to as moatless-adapted. This allows us to fairly compare the performance of SWE-Search against moatless-adapted across various models, including two closed-source models (GPT-4o, GPT-4o-mini) and three open-source models (Qwen2.5-72B-Instruct
1475
+ (Yang et al.,
1476
+ 2024a
1477
+ )
1478
+ , Llama-3.1-70B-Instruct
1479
+ (Dubey et al.,
1480
+ 2024
1481
+ )
1482
+ , and DeepSeek-V2.5
1483
+ (DeepSeek-AI et al.,
1484
+ 2024
1485
+ )
1486
+ ). We also reference official moatless-tools GPT-4o results on SWE-bench Lite to ensure a fair and consistent comparison.
1487
+ Implementation Details
1488
+ For consistency, we use identical prompts across all models. In SWE-Search, we limit each node to a maximum of three expansions and cap the total search iterations at 100. Further details on model hyperparameters can be found in
1489
+ Appendix,
1490
+ 2
1491
+ .
1492
+ Table 1:
1493
+ Resolve Rate Comparison, SWE-bench Lite
1494
+ Model
1495
+ Moatless-v1
1496
+ Moatless-adapted
1497
+ SWE-Search
1498
+ %
1499
+ Δ
1500
+ Δ
1501
+ \Delta
1502
+ GPT-4o
1503
+ 24.3
1504
+ 25.7
1505
+ 31.0
1506
+ +17
1507
+ GPT-4o-mini
1508
+ –
1509
+ 13.0
1510
+ 17.0
1511
+ +24
1512
+ Qwen-2.5-72b-Instruct
1513
+ –
1514
+ 18.0
1515
+ 24.7
1516
+ +27
1517
+ Deepseek-V2.5
1518
+ –
1519
+ 16.3
1520
+ 21.0
1521
+ +22
1522
+ Llama-3.1-70b-Instruct
1523
+ –
1524
+ 13.6
1525
+ 17.7
1526
+ +23
1527
+ Mean %
1528
+ Δ
1529
+ Δ
1530
+ \Delta
1531
+ +23
1532
+ 4.1
1533
+ Experimental Results
1534
+ 4.1.1
1535
+ SWE-Search Surpasses all Corresponding Base Agents and Enables Smaller, Open Source Models to Approach GPT-4o
1536
+ On average, SWE-Search outperforms the baseline agent across all five models, achieving a 23% relative improvement
1537
+ (Table
1538
+ 1
1539
+ )
1540
+ . Notably, SWE-Search with Qwen-2.5-72B-Instruct exceeds the performance of GPT-4o using the original Moatless-v1 framework, and closely matches its performance when compared with the Moatless-adapted agent, with only a slight difference (
1541
+ Δ
1542
+ =
1543
+ −
1544
+ 1
1545
+ %
1546
+ Δ
1547
+ percent
1548
+ 1
1549
+ \Delta=-1\%
1550
+ ). Interestingly, all five models demonstrate significant improvement when utilizing the proposed approach, with consistent gains across different models.
1551
+ 4.1.2
1552
+ Search Enables Agents to Make Better Use of More Flexibility
1553
+ To prevent goal divergence, most agents, including moatless-tools, rely on strict transition rules, where state transitions follow predetermined sequences (e.g., Search
1554
+ →
1555
+ →
1556
+ \rightarrow
1557
+ Identify, Plan
1558
+ →
1559
+ →
1560
+ \rightarrow
1561
+ Edit). In moatless-adapted, we introduce a more flexible transition logic that allows a Plan state to transition into any other state type. This added flexibility has both advantages and drawbacks. On the positive side, it enables the agent to autonomously correct its trajectory without external feedback, particularly when the necessary adjustments span only a limited portion of the task. However, this increased flexibility also introduces the risk of the agent becoming trapped in infinite loops. Without a high-level control mechanism to detect and mitigate these situations, the agent may fail to recover from such loops. This trade-off is evident in the modest performance difference between Moatless-v1 and moatless-adapted, with a slight performance improvement of only 1.4% (
1562
+ Table
1563
+ 1
1564
+ ).
1565
+ 4.1.3
1566
+ Impact of Hindsight Feedback on Agent Performance
1567
+ One key advantage of utilizing LLMs as general value functions is their dual ability to provide both quantitative value estimates and qualitative assessments in natural language. These qualitative insights can significantly enhance the agent’s action generation and search process by offering detailed feedback on potential errors or overlooked aspects of the task. In practice, feedback was also crucial in eliciting diversity in the actions taken by the agent, as without it, the agent would often take very similar actions when re-expanding from a parent node.
1568
+ As shown in
1569
+ Figure
1570
+ 2
1571
+ , this mechanism plays a critical role in improving the agent’s performance. During the initial expansion, the agent prematurely concludes that the task is complete. However, the value function correctly identifies gaps in the test coverage, specifically in addressing potential corner cases, and assigns a low reward. This feedback prompts the agent to re-expand the parent state, leading to the introduction of new tests, which subsequently fail. The agent then performs a series of edits (summarized in the figure for brevity), ultimately resolving the task correctly. Empirically, we observe that the instances unresolved by moatless-adapted but successfully solved by SWE-Search are often attributed to this search-and-feedback loop, where iterative feedback drives the agent toward a correct solution.
1572
+ 4.2
1573
+ Importance of Comprehensive State Information for Value Function Performance
1574
+ Model
1575
+ Pass@1
1576
+ Pass@5
1577
+ GPT-4o
1578
+ 31.0
1579
+ 34.0
1580
+ GPT-4o-mini
1581
+ 17.0
1582
+ 22.3
1583
+ Qwen-2.5-72b-Instruct
1584
+ 24.7
1585
+ 25.7
1586
+ Deepseek-V2.5
1587
+ 21.0
1588
+ 23.3
1589
+ Llama-3.1-70b-Instruct
1590
+ 21.0
1591
+ 22.3
1592
+ Figure 3:
1593
+ SWE-bench SWE-Search results
1594
+ The effectiveness of SWE-Search hinges on the value function’s ability to accurately differentiate between desirable and undesirable states, and to provide actionable feedback that drives improvement. However, our experiments revealed that the value function sometimes failed to recognize critical decision points in the search tree. It frequently misinterpreted the purpose of certain actions, leading to the undervaluation of effective strategies by assigning low rewards. As shown in
1595
+ Figure
1596
+ 4
1597
+ a
1598
+ , before the introduction of state-specific value prompts, the agent consistently assigned low rewards even when the Action Agent correctly identified the need for additional context, such as locating relevant files. This issue persisted despite the agent successfully identifying the files later. By implementing state-specific prompts across core state clusters (Searching, Planning, Editing), the value function became significantly more adept at interpreting the intent behind actions and evaluating their outcomes within each state. For further details on experiments distinguishing between effective and ineffective states, refer to
1599
+ Appendix
1600
+ 8
1601
+ .
1602
+ Figure 4:
1603
+ (a) Importance of state-specific value prompts.
1604
+ On the left and right are the respective Value Agents’ outputs with and without state-specific prompts. While the action in both cases is effective in finding the right file, the non-state-specific scenario does not recognize this and assigns a low reward. On the contrary, the state-specific prompt correctly assigns a high reward to this state.
1605
+ (b) Performance scaling with search depth across different language models.
1606
+ The graph shows the number of issues resolved as a function of the number of transitions (search iterations) for all models used.
1607
+ Scaling SWE agents with Inference-time Compute
1608
+ The success of large language models (LLMs) has traditionally been attributed to the expansion of training data and model size, i.e., training-time compute
1609
+ (Wei et al.,
1610
+ 2022
1611
+ ; Chung et al.,
1612
+ 2022
1613
+ )
1614
+ . Recently, researchers have started exploring how different methods scale with inference-time
1615
+ (OpenAI,
1616
+ 2024
1617
+ ; Snell et al.,
1618
+ 2024
1619
+ ; Dubey et al.,
1620
+ 2024
1621
+ )
1622
+ . Here, we study the performance of software engineering agents through increased inference-time compute. As shown in
1623
+ Figure
1624
+ 4
1625
+ b
1626
+ , increasing search iterations leads to a consistent rise in the number of resolved issues. To ensure experimental feasibility across the 300 instances in the SWE-bench Lite dataset, we applied conservative parameters (maximum iterations
1627
+ =
1628
+ 100
1629
+ absent
1630
+ 100
1631
+ =100
1632
+ , maximum expansions per node
1633
+ =
1634
+ 3
1635
+ absent
1636
+ 3
1637
+ =3
1638
+ ). Approaches like SWE-Search enable the allocation of greater resources to specific challenges, such as addressing critical software vulnerabilities
1639
+ (Rigaki et al.,
1640
+ 2024
1641
+ ; Fang et al.,
1642
+ 2024
1643
+ )
1644
+ , offering a scalable solution to complex tasks.
1645
+ Figure 5:
1646
+ (a) Value Function vs. Discriminator Comparison.
1647
+ Comparison of value function vs. discriminator ability to discern the final solution that resolved the issue when there is one. The discriminator performs better across all models except GPT-4o-mini. DeepSeek-V2.5 had the smallest disparity between the two methods, suggesting an ability to act as a well-calibrated value function.
1648
+ (b) Model-Specific Issue Resolution.
1649
+ Venn diagram of resolved issues by model. Each model can solve a handful of unique instances.
1650
+ Convergence of Value Function and Discriminator to Right Solution
1651
+ The search process can yield multiple proposed solutions. Ideally, the mean trajectory value of the the proposed solution that resolves the issue will always be the highest, which would yields the ideal performance of the agent
1652
+ (Table
1653
+ 3
1654
+ )
1655
+ . In practice, the value function successfully converged on the correct solution 73% of the time on average across the five models. The discriminator module performed even better, increasing the proportion of correct solutions selected to 84%. While in typical large action spaces, Monte Carlo Tree Search (MCTS) is run for thousands of iterations
1656
+ (Silver et al.,
1657
+ 2016b
1658
+ )
1659
+ , the value function’s success rate remains impressive given the computational constraints. However, SWE-Search could further benefit from enhanced methods for identifying the correct solutions more consistently, allowing it to fully reach its potential.
1660
+ Different Models can Resolve Vastly Different Issue Subsets
1661
+ When comparing the resolved instances across the five models, we observed significant diversity in the subsets of issues each model successfully solved. As shown in
1662
+ Figure
1663
+ 5
1664
+ , each model managed to resolve at least one unique instance. Notably, a surprising number of issues (33) were solved by other models but not by GPT-4o. This suggests that model diversity could play an important role, at least in the short term, in enhancing the performance of SWE-agents.
1665
+ 5
1666
+ Discussion and Conclusion
1667
+ In this paper, we introduced SWE-Search, a general framework that integrates Monte Carlo Tree Search (MCTS) and qualitative feedback to enhance the performance of software engineering agents. The proposed approach demonstrated improvements over different baseline models, highlighting the potential of search-based methods in software engineering tasks.
1668
+ One of the key advantages of search-based approaches, as demonstrated in our work, is their ability to scale performance with increased inference-time compute. This flexibility allows the system to adapt to problems that require higher computational resources, such as discovering software vulnerabilities or even generating large codebases from scratch. Future research should focus on two main directions: (a) investigating how search agents scale with computational resources, and (b) expanding the application of software agent search to a broader range of complex use cases.
1669
+ Given that search techniques like MCTS closely resemble the problem-solving processes of human software engineers, we expect these methods to become increasingly prevalent in agent-driven systems. As the nature of software engineering tasks evolves, system architectures will need to become more fluid and adaptable, fully leveraging the potential of search-based techniques. This evolution will likely lead to the development of larger, more general agentic systems capable of tackling a wide array of software engineering challenges.
1670
+ References
1671
+ Amayuelas et al. (2024)
1672
+ Alfonso Amayuelas, Xianjun Yang, Antonis Antoniades, Wenyue Hua, Liangming Pan, and William Wang.
1673
+ Multiagent collaboration attack: Investigating adversarial attacks in large language model collaborations via debate, 2024.
1674
+ URL
1675
+ https://arxiv.org/abs/2406.14711
1676
+ .
1677
+ Black (2005)
1678
+ Paul E. Black.
1679
+ greedy algorithm, feb 2005.
1680
+ URL
1681
+ https://www.nist.gov/dads/HTML/greedyalgo.html
1682
+ .
1683
+ Accessed: TODAY.
1684
+ Bounsi et al. (2024)
1685
+ Wilfried Bounsi, Borja Ibarz, Andrew Dudzik, Jessica B. Hamrick, Larisa Markeeva, Alex Vitvitskyi, Razvan Pascanu, and Petar Veličković.
1686
+ Transformers meet neural algorithmic reasoners, 2024.
1687
+ URL
1688
+ https://arxiv.org/abs/2406.09308
1689
+ .
1690
+ Chowdhury et al. (2024)
1691
+ Neil Chowdhury, James Aung, Chan Jun Shern, Oliver Jaffe, Dane Sherburn, Giulio Starace, Evan Mays, Rachel Dias, Marwan Aljubeh, Mia Glaese, Carlos E. Jimenez, John Yang, Kevin Liu, and Aleksander Madry.
1692
+ Introducing SWE-bench verified, August 2024.
1693
+ URL
1694
+ https://openai.com/research/introducing-swe-bench-verified
1695
+ .
1696
+ OpenAI Blog.
1697
+ Chung et al. (2022)
1698
+ Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei.
1699
+ Scaling instruction-finetuned language models, 2022.
1700
+ URL
1701
+ https://arxiv.org/abs/2210.11416
1702
+ .
1703
+ Cormen et al. (2009)
1704
+ Thomas H. Cormen, Charles E. Leiserson, Ronald L. Rivest, and Clifford Stein.
1705
+ Introduction to Algorithms, Third Edition
1706
+ .
1707
+ The MIT Press, 3rd edition, 2009.
1708
+ ISBN 0262033844.
1709
+ DeepSeek-AI et al. (2024)
1710
+ DeepSeek-AI, Qihao Zhu, Daya Guo, Zhihong Shao, Dejian Yang, Peiyi Wang, Runxin Xu, Y. Wu, Yukun Li, Huazuo Gao, Shirong Ma, Wangding Zeng, Xiao Bi, Zihui Gu, Hanwei Xu, Damai Dai, Kai Dong, Liyue Zhang, Yishi Piao, Zhibin Gou, Zhenda Xie, Zhewen Hao, Bingxuan Wang, Junxiao Song, Deli Chen, Xin Xie, Kang Guan, Yuxiang You, Aixin Liu, Qiushi Du, Wenjun Gao, Xuan Lu, Qinyu Chen, Yaohui Wang, Chengqi Deng, Jiashi Li, Chenggang Zhao, Chong Ruan, Fuli Luo, and Wenfeng Liang.
1711
+ Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence, 2024.
1712
+ URL
1713
+ https://arxiv.org/abs/2406.11931
1714
+ .
1715
+ Du et al. (2023)
1716
+ Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch.
1717
+ Improving factuality and reasoning in language models through multiagent debate, 2023.
1718
+ URL
1719
+ https://arxiv.org/abs/2305.14325
1720
+ .
1721
+ Dubey et al. (2024)
1722
+ Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov,
1723
+ Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Olivier
1724
+ Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vladan Petrovic, Weiwei Chu,
1725
+ Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaoqing Ellen Tan, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aaron Grattafiori, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alex Vaughan, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Franco, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl
1726
+ Parker, Carly Burton, Catalina Mejia, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, Danny Wyatt, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Firat Ozgenel, Francesco Caggioni, Francisco Guzmán, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Govind Thattai, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Irina-Elena Veliche, Itai Gat, Jake Weissman, James
1727
+ Geboski, James Kohli, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Karthik Prasad, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kun Huang, Kunal Chawla, Kushal Lakhotia, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Maria Tsimpoukelli, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini
1728
+ Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikolay Pavlovich Laptev, Ning Dong, Ning Zhang, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Rohan Maheswari, Russ Howes, Ruty Rinott, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta,
1729
+ Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Kohler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vítor Albiero, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaofang Wang, Xiaojian Wu, Xiaolan Wang, Xide Xia, Xilun Wu, Xinbo Gao, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yuchen Hao, Yundi Qian, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, and Zhiwei Zhao.
1730
+ The llama 3 herd of models, 2024.
1731
+ URL
1732
+ https://arxiv.org/abs/2407.21783
1733
+ .
1734
+ Faldor et al. (2024)
1735
+ Maxence Faldor, Jenny Zhang, Antoine Cully, and Jeff Clune.
1736
+ Omni-epic: Open-endedness via models of human notions of interestingness with environments programmed in code, 2024.
1737
+ URL
1738
+ https://arxiv.org/abs/2405.15568
1739
+ .
1740
+ Fang et al. (2024)
1741
+ Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang.
1742
+ Teams of llm agents can exploit zero-day vulnerabilities, 2024.
1743
+ URL
1744
+ https://arxiv.org/abs/2406.01637
1745
+ .
1746
+ Fawzi et al. (2022)
1747
+ A. Fawzi, M. Balog, A. Huang, T. Hubert, B. Romera-Paredes, M. Barekatain, A. Novikov, F. J. R. Ruiz, J. Schrittwieser, G. Swirszcz, D. Silver, D. Hassabis, and P. Kohli.
1748
+ Discovering faster matrix multiplication algorithms with reinforcement learning.
1749
+ Nature
1750
+ , 610(7930):47–53, 2022.
1751
+ doi:
1752
+ 10.1038/s41586-022-05172-4
1753
+ .
1754
+ Hart et al. (1968)
1755
+ Peter E. Hart, Nils J. Nilsson, and Bertram Raphael.
1756
+ A formal basis for the heuristic determination of minimum cost paths.
1757
+ IEEE Trans. Syst. Sci. Cybern.
1758
+ , 4(2):100–107, 1968.
1759
+ doi:
1760
+ 10.1109/TSSC.1968.300136
1761
+ .
1762
+ URL
1763
+ https://doi.org/10.1109/TSSC.1968.300136
1764
+ .
1765
+ Hu et al. (2024)
1766
+ Shengran Hu, Cong Lu, and Jeff Clune.
1767
+ Automated design of agentic systems, 2024.
1768
+ URL
1769
+ https://arxiv.org/abs/2408.08435
1770
+ .
1771
+ Jimenez et al. (2024)
1772
+ Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.
1773
+ Swe-bench: Can language models resolve real-world github issues?, 2024.
1774
+ URL
1775
+ https://arxiv.org/abs/2310.06770
1776
+ .
1777
+ Jumper et al. (2021)
1778
+ J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera-Paredes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Petersen, D. Reiman, E. Clancy, M. Zielinski, M. Steinegger, M. Pacholska, T. Berghammer, S. Bodenstein, D. Silver, O. Vinyals, A. W. Senior, K. Kavukcuoglu, P. Kohli, and D. Hassabis.
1779
+ Highly accurate protein structure prediction with AlphaFold.
1780
+ Nature
1781
+ , 596(7873):583–589, 2021.
1782
+ doi:
1783
+ 10.1038/s41586-021-03819-2
1784
+ .
1785
+ Kahneman (2011)
1786
+ Daniel Kahneman.
1787
+ Thinking, fast and slow
1788
+ .
1789
+ Farrar, Straus and Giroux, New York, NY, US, 2011.
1790
+ ISBN 978-0-374-27563-1.
1791
+ Khan et al. (2024)
1792
+ Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, and Ethan Perez.
1793
+ Debating with more persuasive llms leads to more truthful answers, 2024.
1794
+ URL
1795
+ https://arxiv.org/abs/2402.06782
1796
+ .
1797
+ Kocsis & Szepesvári (2006)
1798
+ Levente Kocsis and Csaba Szepesvári.
1799
+ Bandit based monte-carlo planning.
1800
+ In Johannes Fürnkranz, Tobias Scheffer, and Myra Spiliopoulou (eds.),
1801
+ Machine Learning: ECML 2006
1802
+ , pp.  282–293, Berlin, Heidelberg, 2006. Springer Berlin Heidelberg.
1803
+ ISBN 978-3-540-46056-5.
1804
+ Koh et al. (2024)
1805
+ Jing Yu Koh, Stephen McAleer, Daniel Fried, and Ruslan Salakhutdinov.
1806
+ Tree search for language model agents, 2024.
1807
+ URL
1808
+ https://arxiv.org/abs/2407.01476
1809
+ .
1810
+ Li et al. (2022)
1811
+ Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals.
1812
+ Competition-level code generation with alphacode.
1813
+ Science
1814
+ , 378(6624):1092–1097, 2022.
1815
+ doi:
1816
+ 10.1126/science.abq1158
1817
+ .
1818
+ URL
1819
+ https://www.science.org/doi/abs/10.1126/science.abq1158
1820
+ .
1821
+ Lu et al. (2024a)
1822
+ Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha.
1823
+ The ai scientist: Towards fully automated open-ended scientific discovery, 2024a.
1824
+ URL
1825
+ https://arxiv.org/abs/2408.06292
1826
+ .
1827
+ Lu et al. (2024b)
1828
+ Cong Lu, Shengran Hu, and Jeff Clune.
1829
+ Intelligent go-explore: Standing on the shoulders of giant foundation models, 2024b.
1830
+ URL
1831
+ https://arxiv.org/abs/2405.15143
1832
+ .
1833
+ Ma et al. (2024a)
1834
+ Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar.
1835
+ Eureka: Human-level reward design via coding large language models, 2024a.
1836
+ URL
1837
+ https://arxiv.org/abs/2310.12931
1838
+ .
1839
+ Ma et al. (2024b)
1840
+ Yingwei Ma, Qingping Yang, Rongyu Cao, Binhua Li, Fei Huang, and Yongbin Li.
1841
+ How to understand whole software repository?, 2024b.
1842
+ URL
1843
+ https://arxiv.org/abs/2406.01422
1844
+ .
1845
+ Moore (1959)
1846
+ E.F. Moore.
1847
+ The Shortest Path Through a Maze
1848
+ .
1849
+ Bell Telephone System. Technical publications. monograph. Bell Telephone System., 1959.
1850
+ URL
1851
+ https://books.google.com/books?id=IVZBHAAACAAJ
1852
+ .
1853
+ OpenAI (2024)
1854
+ OpenAI.
1855
+ OpenAI o1 System Card, September 2024.
1856
+ URL
1857
+ https://openai.com/research/o1-system-card
1858
+ .
1859
+ Online report.
1860
+ OpenAI et al. (2024)
1861
+ OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo
1862
+ Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan
1863
+ Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez,
1864
+ Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia
1865
+ Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph.
1866
+ Gpt-4 technical report, 2024.
1867
+ URL
1868
+ https://arxiv.org/abs/2303.08774
1869
+ .
1870
+ Örwall (2024)
1871
+ Albert Örwall.
1872
+ Moatless tools, jun 2024.
1873
+ URL
1874
+ https://github.com/aorwall/moatless-tools
1875
+ .
1876
+ Accessed: 2024-07-16.
1877
+ Ouyang et al. (2022)
1878
+ Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe.
1879
+ Training language models to follow instructions with human feedback, 2022.
1880
+ URL
1881
+ https://arxiv.org/abs/2203.02155
1882
+ .
1883
+ Pan et al. (2023)
1884
+ Liangming Pan, Alon Albalak, Xinyi Wang, and William Yang Wang.
1885
+ Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning, 2023.
1886
+ URL
1887
+ https://arxiv.org/abs/2305.12295
1888
+ .
1889
+ Rigaki et al. (2024)
1890
+ Maria Rigaki, Carlos Catania, and Sebastian Garcia.
1891
+ Hackphyr: A local fine-tuned llm agent for network security environments, 2024.
1892
+ URL
1893
+ https://arxiv.org/abs/2409.11276
1894
+ .
1895
+ Saha et al. (2024)
1896
+ Swarnadeep Saha, Archiki Prasad, Justin Chih-Yao Chen, Peter Hase, Elias Stengel-Eskin, and Mohit Bansal.
1897
+ System-1.x: Learning to balance fast and slow planning with language models, 2024.
1898
+ URL
1899
+ https://arxiv.org/abs/2407.14414
1900
+ .
1901
+ Silver et al. (2016a)
1902
+ David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis.
1903
+ Mastering the game of go with deep neural networks and tree search.
1904
+ Nature
1905
+ , 529(7587):484–489, 1 2016a.
1906
+ ISSN 1476-4687.
1907
+ doi:
1908
+ 10.1038/nature16961
1909
+ .
1910
+ URL
1911
+ https://doi.org/10.1038/nature16961
1912
+ .
1913
+ Silver et al. (2016b)
1914
+ David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis.
1915
+ Mastering the game of go with deep neural networks and tree search.
1916
+ Nat.
1917
+ , 529(7587):484–489, 2016b.
1918
+ doi:
1919
+ 10.1038/NATURE16961
1920
+ .
1921
+ URL
1922
+ https://doi.org/10.1038/nature16961
1923
+ .
1924
+ Silver et al. (2018)
1925
+ David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis.
1926
+ A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.
1927
+ Science
1928
+ , 362(6419):1140–1144, 2018.
1929
+ doi:
1930
+ 10.1126/science.aar6404
1931
+ .
1932
+ URL
1933
+ https://www.science.org/doi/abs/10.1126/science.aar6404
1934
+ .
1935
+ Snell et al. (2024)
1936
+ Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar.
1937
+ Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024.
1938
+ URL
1939
+ https://arxiv.org/abs/2408.03314
1940
+ .
1941
+ Wang et al. (2023)
1942
+ Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar.
1943
+ Voyager: An open-ended embodied agent with large language models, 2023.
1944
+ URL
1945
+ https://arxiv.org/abs/2305.16291
1946
+ .
1947
+ Wang et al. (2024a)
1948
+ Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji.
1949
+ Executable code actions elicit better llm agents, 2024a.
1950
+ URL
1951
+ https://arxiv.org/abs/2402.01030
1952
+ .
1953
+ Wang et al. (2024b)
1954
+ Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, Robert Brennan, Hao Peng, Heng Ji, and Graham Neubig.
1955
+ Opendevin: An open platform for ai software developers as generalist agents, 2024b.
1956
+ URL
1957
+ https://arxiv.org/abs/2407.16741
1958
+ .
1959
+ Wei et al. (2022)
1960
+ Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al.
1961
+ Emergent abilities of large language models.
1962
+ arXiv preprint arXiv:2206.07682
1963
+ , 2022.
1964
+ Xia et al. (2024)
1965
+ Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang.
1966
+ Agentless: Demystifying llm-based software engineering agents, 2024.
1967
+ URL
1968
+ https://arxiv.org/abs/2407.01489
1969
+ .
1970
+ Yang et al. (2024a)
1971
+ An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Xuejing Liu, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, Zhifang Guo, and Zhihao Fan.
1972
+ Qwen2 technical report, 2024a.
1973
+ URL
1974
+ https://arxiv.org/abs/2407.10671
1975
+ .
1976
+ Yang et al. (2024b)
1977
+ John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press.
1978
+ Swe-agent: Agent-computer interfaces enable automated software engineering, 2024b.
1979
+ URL
1980
+ https://arxiv.org/abs/2405.15793
1981
+ .
1982
+ Yao et al. (2023)
1983
+ Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan.
1984
+ Tree of thoughts: Deliberate problem solving with large language models.
1985
+ In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.),
1986
+ Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023
1987
+ , 2023.
1988
+ URL
1989
+ http://papers.nips.cc/paper_files/paper/2023/hash/271db9922b8d1f4dd7aaef84ed5ac703-Abstract-Conference.html
1990
+ .
1991
+ Zhang et al. (2024a)
1992
+ Kexun Zhang, Weiran Yao, Zuxin Liu, Yihao Feng, Zhiwei Liu, Rithesh Murthy, Tian Lan, Lei Li, Renze Lou, Jiacheng Xu, Bo Pang, Yingbo Zhou, Shelby Heinecke, Silvio Savarese, Huan Wang, and Caiming Xiong.
1993
+ Diversity empowers intelligence: Integrating expertise of software engineering agents, 2024a.
1994
+ URL
1995
+ https://arxiv.org/abs/2408.07060
1996
+ .
1997
+ Zhang et al. (2024b)
1998
+ Yao Zhang, Zijian Ma, Yunpu Ma, Zhen Han, Yu Wu, and Volker Tresp.
1999
+ Webpilot: A versatile and autonomous multi-agent system for web task execution with strategic exploration, 2024b.
2000
+ URL
2001
+ https://arxiv.org/abs/2408.15978
2002
+ .
2003
+ Zhang et al. (2024c)
2004
+ Yiqun Zhang, Xiaocui Yang, Shi Feng, Daling Wang, Yifei Zhang, and Kaisong Song.
2005
+ Can llms beat humans in debating? a dynamic multi-agent framework for competitive debate, 2024c.
2006
+ URL
2007
+ https://arxiv.org/abs/2408.04472
2008
+ .
2009
+ Zhang et al. (2024d)
2010
+ Yuntong Zhang, Haifeng Ruan, Zhiyu Fan, and Abhik Roychoudhury.
2011
+ Autocoderover: Autonomous program improvement, 2024d.
2012
+ URL
2013
+ https://arxiv.org/abs/2404.05427
2014
+ .
2015
+ Appendix A
2016
+ Reproducibility
2017
+ All models and data used in our work are publicly available. We additionally provide hyperparameter details in
2018
+ Appendix
2019
+ 2
2020
+ . The code will be released as a public repository upon publication.
2021
+ Appendix B
2022
+ Additional Implementation Details
2023
+ Moatless-adapted is an extended version of the moatless-tools library with support for a tree structure, the ability to revert to earlier versions of the codebase, and the capability to run tests.
2024
+ The standard implementation of moatless-tools is based on a finite state machine structure where a state holds information about file context and properties set in the configuration or from previous states. It can then transition to a new state when an action is executed. The request that initiates the action is created by an LLM. This follows a linear structure where one state can transition to another state. In moatless-adapted, this model is extended so that a state can expand by using actions to create more states. The connections between states are then represented in a tree structure with nodes.
2025
+ Each state has a file context associated with it. This file context will be included in the prompt sent to an LLM. To limit the size of the prompt, files are divided into ”spans,” where a span could be, for example, a section of code (e.g., imports), a class, or a function. These are identified by span IDs. Thus, the LLM sees a limited part of the code at a time but can request more context by searching for or adding files and spans. The file context therefore changes over time, and a specific state of file context is linked to a specific state.
2026
+ In the standard implementation of moatless-tools, changes to the codebase are made linearly, and each change is saved directly to the file system. In moatless-adapted, however, there is a need to be able to revert to earlier states and thus return to a previous version of the codebase. To handle this, the code is stored in a git repository where each change is committed, and each state has a reference to a commit as well as the current patch of the diff from the initial commit that existed before starting. This way, one can go back to an earlier state by specifying the state ID, and the commit that was current at that time will be checked out.
2027
+ The test files present in the file context are run each time the Plan state is initiated, and the test results are provided to the state. The tests are then run in Docker images built via the SWE-bench library. To use this approach in a benchmark where a larger number of instances should be able to run simultaneously, a solution is used where these images are run as pods in a Kubernetes cluster. Moatless-tools communicates with the testbed by applying patches and running commands via an API. When a new instance starts, a pod is created which is then reset at each run, applying the current patch and running tests according to the test command specified in the SWE-bench library. It’s important to add here that the agent is not aware of the
2028
+ PASS_TO_PASS
2029
+ or
2030
+ FAIL_TO_PASS
2031
+ tests in the SWE-bench harness, but only knows how to run the tests. This corresponds to a real engineering environment where each project can have its own test commands.
2032
+ Appendix C
2033
+ MCTS Hyperparameters
2034
+ The Monte Carlo Tree Search (MCTS) algorithm used in this study employs several hyperparameters.
2035
+ Table 2:
2036
+ MCTS Hyperparameters
2037
+ Hyperparameter
2038
+ Description
2039
+ Default
2040
+ c_param
2041
+ UCT exploration parameter
2042
+ 1.41
2043
+ max_expansions
2044
+ Max children per node
2045
+ 5
2046
+ max_iterations
2047
+ Max MCTS iterations
2048
+ 100
2049
+ provide_feedback
2050
+ Enable feedback
2051
+ True
2052
+ best_first
2053
+ Use best-first strategy
2054
+ True
2055
+ value_function_temperature
2056
+ Value function temperature
2057
+ 0.2
2058
+ max_depth
2059
+ Max tree depth
2060
+ 20
2061
+ UCT Score Calculation Parameters
2062
+ exploration_weight
2063
+ UCT exploration weight
2064
+ 1.0
2065
+ depth_weight
2066
+ Depth penalty weight
2067
+ 0.8
2068
+ depth_bonus_factor
2069
+ Depth bonus factor
2070
+ 200.0
2071
+ high_value_threshold
2072
+ High-value node threshold
2073
+ 55.0
2074
+ low_value_threshold
2075
+ Low-value node threshold
2076
+ 50.0
2077
+ very_high_value_threshold
2078
+ Very high-value threshold
2079
+ 75.0
2080
+ high_value_leaf_bonus_constant
2081
+ High-value leaf bonus
2082
+ 20.0
2083
+ high_value_bad_children_bonus_constant
2084
+ High-value bad children bonus
2085
+ 20.0
2086
+ high_value_child_penalty_constant
2087
+ High-value child penalty
2088
+ 5.0
2089
+ Action Model Parameters
2090
+ action_model_temperature
2091
+ Action model temperature
2092
+ 0.2
2093
+ Discriminator Parameters
2094
+ number_of_agents
2095
+ Number of Discriminator Agents
2096
+ 5
2097
+ number_of_round
2098
+ Number of debate rounds
2099
+ 3
2100
+ discriminator_temperature
2101
+ Discriminator temperature
2102
+ 1.0
2103
+ These hyperparameters can be adjusted to fine-tune the MCTS algorithm’s performance for specific problem domains or computational constraints. The values listed here are the defaults as defined in the
2104
+ TreeSearchSettings
2105
+ class and the MCTS implementation.
2106
+ Appendix D
2107
+ Ability of MCTS to Escape Unproductive Loops vs. Baseline
2108
+ Figure 6:
2109
+ Avoiding Repetitive Actions, django__django__10914.
2110
+ We found that the base agent can often get stuck performing repetitive actions
2111
+ that do not bring it closer to solving the issue, and which commonly lead to unresolvable dead-ends. In this example, the base agent
2112
+ was stuck implementing wrong tests which continuously returned errors. In contrast, when this happens in
2113
+ SWE-Search, the Value Agent recognizes this, terminating these trajectories quickly,
2114
+ as happens in Node 73 (orange).
2115
+ Appendix E
2116
+ Model Instance Resolution Uniqueness
2117
+ To understand the complementary strengths of different models in resolving software issues, we analyzed how unique their resolved issue subsets where. Figure
2118
+ 7
2119
+ illustrates the resolution patterns for each model across five of the codebases in SWE-bench-lite.
2120
+ Figure 7:
2121
+ Unique Issue Resolution Patterns Across Models and Libraries.
2122
+ Each column represents a different Python reposiroty, and each row within a column represents a specific issue. Colored blocks indicate successful resolution by the corresponding model (see legend). White spaces denote unresolved issues. This visualization highlights the diverse problem-solving capabilities of different models across various software domains, demonstrating that no single model dominates across all issues and libraries.
2123
+ Appendix F
2124
+ Ability of Value Function to Discern Successful Trajectories
2125
+ Before implementing SWE-Search, we conducted a general study across many models to evaluate the models’ ability to differentiate states which led to resolved vs. unresolved issues. Figure
2126
+ 8
2127
+ shows the results of this study. We found that in general, models assigned higher rewards to states which eventually led to resolved issues. Of particular interest was the Deepseek model, which seemed to identify critical errors in trajectories effectively. This was also observed in the final agent (see Fig.
2128
+ 5
2129
+ a).
2130
+ Figure 8:
2131
+ Average State Reward Comparison Across Models.
2132
+ This graph compares the average state rewards assigned by different language models for resolved (green) and unresolved (red) issues. Error bars indicate standard deviation. Most models consistently assign higher rewards to states leading to resolved issues, with the exception of the. The ’Average’ column represents the mean across all models, demonstrating a clear distinction between resolved and unresolved states.
2133
+ Appendix G
2134
+ Value Function Prompts
2135
+ ◄
2136
+ Feeling
2137
+ lucky?
2138
+ Conversion
2139
+ report
2140
+ Report
2141
+ an issue
2142
+ View original
2143
+ on arXiv
2144
+ ►
research/notes/241221139-training-software-engineering-agents-and-verifiers-with-swe-gym.md ADDED
@@ -0,0 +1,200 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2412.21139] Training Software Engineering Agents and Verifiers with SWE-Gym'
3
+ id: 241221139-training-software-engineering-agents-and-verifiers-with-swe-gym
4
+ tags:
5
+ - deepread
6
+ created: '2026-06-10T00:22:57.639430Z'
7
+ source: https://arxiv.org/abs/2412.21139
8
+ source_domain: arxiv.org
9
+ fetched_at: '2026-06-10T00:22:57.639309Z'
10
+ fetch_provider: builtin
11
+ status: draft
12
+ type: note
13
+ tier: institutional
14
+ content_type: paper
15
+ deprecated: false
16
+ ---
17
+
18
+ [2412.21139] Training Software Engineering Agents and Verifiers with SWE-Gym
19
+ Computer Science > Software Engineering
20
+ arXiv:2412.21139
21
+ (cs)
22
+ [Submitted on 30 Dec 2024 (
23
+ v1
24
+ ), last revised 6 Jun 2025 (this version, v2)]
25
+ Title:
26
+ Training Software Engineering Agents and Verifiers with SWE-Gym
27
+ Authors:
28
+ Jiayi Pan
29
+ ,
30
+ Xingyao Wang
31
+ ,
32
+ Graham Neubig
33
+ ,
34
+ Navdeep Jaitly
35
+ ,
36
+ Heng Ji
37
+ ,
38
+ Alane Suhr
39
+ ,
40
+ Yizhe Zhang
41
+ View a PDF of the paper titled Training Software Engineering Agents and Verifiers with SWE-Gym, by Jiayi Pan and 6 other authors
42
+ View PDF
43
+ HTML (experimental)
44
+ Abstract:
45
+ We present SWE-Gym, the first environment for training real-world software engineering (SWE) agents. SWE-Gym contains 2,438 real-world Python task instances, each comprising a codebase with an executable runtime environment, unit tests, and a task specified in natural language. We use SWE-Gym to train language model based SWE agents, achieving up to 19% absolute gains in resolve rate on the popular SWE-Bench Verified and Lite test sets. We also experiment with inference-time scaling through verifiers trained on agent trajectories sampled from SWE-Gym. When combined with our fine-tuned SWE agents, we achieve 32.0% and 26.0% on SWE-Bench Verified and Lite, respectively, reflecting a new state-of-the-art for open-weight SWE agents. To facilitate further research, we publicly release SWE-Gym, models, and agent trajectories.
46
+ Comments:
47
+ Accepted at ICML 2025. Code at
48
+ this https URL
49
+ Subjects:
50
+ Software Engineering (cs.SE)
51
+ ; Computation and Language (cs.CL)
52
+ Cite as:
53
+ arXiv:2412.21139
54
+ [cs.SE]
55
+ (or
56
+ arXiv:2412.21139v2
57
+ [cs.SE]
58
+ for this version)
59
+ https://doi.org/10.48550/arXiv.2412.21139
60
+ Focus to learn more
61
+ arXiv-issued DOI via DataCite
62
+ Submission history
63
+ From: Jiayi Pan [
64
+ view email
65
+ ]
66
+ [v1]
67
+ Mon, 30 Dec 2024 18:15:39 UTC (156 KB)
68
+ [v2]
69
+ Fri, 6 Jun 2025 07:53:20 UTC (295 KB)
70
+ Full-text links:
71
+ Access Paper:
72
+ View a PDF of the paper titled Training Software Engineering Agents and Verifiers with SWE-Gym, by Jiayi Pan and 6 other authors
73
+ View PDF
74
+ HTML (experimental)
75
+ TeX Source
76
+ view license
77
+ Current browse context:
78
+ cs.SE
79
+ < prev
80
+ |
81
+ next >
82
+ new
83
+ |
84
+ recent
85
+ |
86
+ 2024-12
87
+ Change to browse by:
88
+ cs
89
+ cs.CL
90
+ References & Citations
91
+ NASA ADS
92
+ Google Scholar
93
+ Semantic Scholar
94
+ export BibTeX citation
95
+ Loading...
96
+ BibTeX formatted citation
97
+ ×
98
+ loading...
99
+ Data provided by:
100
+ Bookmark
101
+ Bibliographic Tools
102
+ Bibliographic and Citation Tools
103
+ Bibliographic Explorer Toggle
104
+ Bibliographic Explorer
105
+ (
106
+ What is the Explorer?
107
+ )
108
+ Connected Papers Toggle
109
+ Connected Papers
110
+ (
111
+ What is Connected Papers?
112
+ )
113
+ Litmaps Toggle
114
+ Litmaps
115
+ (
116
+ What is Litmaps?
117
+ )
118
+ scite.ai Toggle
119
+ scite Smart Citations
120
+ (
121
+ What are Smart Citations?
122
+ )
123
+ Code, Data, Media
124
+ Code, Data and Media Associated with this Article
125
+ alphaXiv Toggle
126
+ alphaXiv
127
+ (
128
+ What is alphaXiv?
129
+ )
130
+ Links to Code Toggle
131
+ CatalyzeX Code Finder for Papers
132
+ (
133
+ What is CatalyzeX?
134
+ )
135
+ DagsHub Toggle
136
+ DagsHub
137
+ (
138
+ What is DagsHub?
139
+ )
140
+ GotitPub Toggle
141
+ Gotit.pub
142
+ (
143
+ What is GotitPub?
144
+ )
145
+ Huggingface Toggle
146
+ Hugging Face
147
+ (
148
+ What is Huggingface?
149
+ )
150
+ ScienceCast Toggle
151
+ ScienceCast
152
+ (
153
+ What is ScienceCast?
154
+ )
155
+ Demos
156
+ Demos
157
+ Replicate Toggle
158
+ Replicate
159
+ (
160
+ What is Replicate?
161
+ )
162
+ Spaces Toggle
163
+ Hugging Face Spaces
164
+ (
165
+ What is Spaces?
166
+ )
167
+ Spaces Toggle
168
+ TXYZ.AI
169
+ (
170
+ What is TXYZ.AI?
171
+ )
172
+ Related Papers
173
+ Recommenders and Search Tools
174
+ Link to Influence Flower
175
+ Influence Flower
176
+ (
177
+ What are Influence Flowers?
178
+ )
179
+ Core recommender toggle
180
+ CORE Recommender
181
+ (
182
+ What is CORE?
183
+ )
184
+ Author
185
+ Venue
186
+ Institution
187
+ Topic
188
+ About arXivLabs
189
+ arXivLabs: experimental projects with community collaborators
190
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
191
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
192
+ Have an idea for a project that will add value for arXiv's community?
193
+ Learn more about arXivLabs
194
+ .
195
+ Which authors of this paper are endorsers?
196
+ |
197
+ Disable MathJax
198
+ (
199
+ What is MathJax?
200
+ )
research/notes/250104519-rstar-math-small-llms-can-master-math-reasoning-with-self-evolved-deep.md ADDED
@@ -0,0 +1,196 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2501.04519] rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved
3
+ Deep Thinking'
4
+ id: 250104519-rstar-math-small-llms-can-master-math-reasoning-with-self-evolved-deep
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:40:01.597011Z'
8
+ source: https://arxiv.org/abs/2501.04519
9
+ source_domain: arxiv.org
10
+ fetched_at: '2026-06-10T00:40:01.596873Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2501.04519] rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
20
+ Computer Science > Computation and Language
21
+ arXiv:2501.04519
22
+ (cs)
23
+ [Submitted on 8 Jan 2025]
24
+ Title:
25
+ rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
26
+ Authors:
27
+ Xinyu Guan
28
+ ,
29
+ Li Lyna Zhang
30
+ ,
31
+ Yifei Liu
32
+ ,
33
+ Ning Shang
34
+ ,
35
+ Youran Sun
36
+ ,
37
+ Yi Zhu
38
+ ,
39
+ Fan Yang
40
+ ,
41
+ Mao Yang
42
+ View a PDF of the paper titled rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking, by Xinyu Guan and 7 other authors
43
+ View PDF
44
+ HTML (experimental)
45
+ Abstract:
46
+ We present rStar-Math to demonstrate that small language models (SLMs) can rival or even surpass the math reasoning capability of OpenAI o1, without distillation from superior models. rStar-Math achieves this by exercising "deep thinking" through Monte Carlo Tree Search (MCTS), where a math policy SLM performs test-time search guided by an SLM-based process reward model. rStar-Math introduces three innovations to tackle the challenges in training the two SLMs: (1) a novel code-augmented CoT data sythesis method, which performs extensive MCTS rollouts to generate step-by-step verified reasoning trajectories used to train the policy SLM; (2) a novel process reward model training method that avoids naïve step-level score annotation, yielding a more effective process preference model (PPM); (3) a self-evolution recipe in which the policy SLM and PPM are built from scratch and iteratively evolved to improve reasoning capabilities. Through 4 rounds of self-evolution with millions of synthesized solutions for 747k math problems, rStar-Math boosts SLMs' math reasoning to state-of-the-art levels. On the MATH benchmark, it improves Qwen2.5-Math-7B from 58.8% to 90.0% and Phi3-mini-3.8B from 41.4% to 86.4%, surpassing o1-preview by +4.5% and +0.9%. On the USA Math Olympiad (AIME), rStar-Math solves an average of 53.3% (8/15) of problems, ranking among the top 20% the brightest high school math students. Code and data will be available at
47
+ this https URL
48
+ .
49
+ Subjects:
50
+ Computation and Language (cs.CL)
51
+ Cite as:
52
+ arXiv:2501.04519
53
+ [cs.CL]
54
+ (or
55
+ arXiv:2501.04519v1
56
+ [cs.CL]
57
+ for this version)
58
+ https://doi.org/10.48550/arXiv.2501.04519
59
+ Focus to learn more
60
+ arXiv-issued DOI via DataCite
61
+ Submission history
62
+ From: Li Lyna Zhang [
63
+ view email
64
+ ]
65
+ [v1]
66
+ Wed, 8 Jan 2025 14:12:57 UTC (632 KB)
67
+ Full-text links:
68
+ Access Paper:
69
+ View a PDF of the paper titled rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking, by Xinyu Guan and 7 other authors
70
+ View PDF
71
+ HTML (experimental)
72
+ TeX Source
73
+ view license
74
+ Current browse context:
75
+ cs.CL
76
+ < prev
77
+ |
78
+ next >
79
+ new
80
+ |
81
+ recent
82
+ |
83
+ 2025-01
84
+ Change to browse by:
85
+ cs
86
+ References & Citations
87
+ NASA ADS
88
+ Google Scholar
89
+ Semantic Scholar
90
+ export BibTeX citation
91
+ Loading...
92
+ BibTeX formatted citation
93
+ ×
94
+ loading...
95
+ Data provided by:
96
+ Bookmark
97
+ Bibliographic Tools
98
+ Bibliographic and Citation Tools
99
+ Bibliographic Explorer Toggle
100
+ Bibliographic Explorer
101
+ (
102
+ What is the Explorer?
103
+ )
104
+ Connected Papers Toggle
105
+ Connected Papers
106
+ (
107
+ What is Connected Papers?
108
+ )
109
+ Litmaps Toggle
110
+ Litmaps
111
+ (
112
+ What is Litmaps?
113
+ )
114
+ scite.ai Toggle
115
+ scite Smart Citations
116
+ (
117
+ What are Smart Citations?
118
+ )
119
+ Code, Data, Media
120
+ Code, Data and Media Associated with this Article
121
+ alphaXiv Toggle
122
+ alphaXiv
123
+ (
124
+ What is alphaXiv?
125
+ )
126
+ Links to Code Toggle
127
+ CatalyzeX Code Finder for Papers
128
+ (
129
+ What is CatalyzeX?
130
+ )
131
+ DagsHub Toggle
132
+ DagsHub
133
+ (
134
+ What is DagsHub?
135
+ )
136
+ GotitPub Toggle
137
+ Gotit.pub
138
+ (
139
+ What is GotitPub?
140
+ )
141
+ Huggingface Toggle
142
+ Hugging Face
143
+ (
144
+ What is Huggingface?
145
+ )
146
+ ScienceCast Toggle
147
+ ScienceCast
148
+ (
149
+ What is ScienceCast?
150
+ )
151
+ Demos
152
+ Demos
153
+ Replicate Toggle
154
+ Replicate
155
+ (
156
+ What is Replicate?
157
+ )
158
+ Spaces Toggle
159
+ Hugging Face Spaces
160
+ (
161
+ What is Spaces?
162
+ )
163
+ Spaces Toggle
164
+ TXYZ.AI
165
+ (
166
+ What is TXYZ.AI?
167
+ )
168
+ Related Papers
169
+ Recommenders and Search Tools
170
+ Link to Influence Flower
171
+ Influence Flower
172
+ (
173
+ What are Influence Flowers?
174
+ )
175
+ Core recommender toggle
176
+ CORE Recommender
177
+ (
178
+ What is CORE?
179
+ )
180
+ Author
181
+ Venue
182
+ Institution
183
+ Topic
184
+ About arXivLabs
185
+ arXivLabs: experimental projects with community collaborators
186
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
187
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
188
+ Have an idea for a project that will add value for arXiv's community?
189
+ Learn more about arXivLabs
190
+ .
191
+ Which authors of this paper are endorsers?
192
+ |
193
+ Disable MathJax
194
+ (
195
+ What is MathJax?
196
+ )
research/notes/250104519-sysname-small-llms-can-master-math-reasoning-with-self-evolved-deep-th.md ADDED
@@ -0,0 +1,3557 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2501.04519] \sysname: Small LLMs Can Master Math Reasoning with Self-Evolved
3
+ Deep Thinking'
4
+ id: 250104519-sysname-small-llms-can-master-math-reasoning-with-self-evolved-deep-th
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:40:46.873514Z'
8
+ source: https://ar5iv.labs.arxiv.org/html/2501.04519
9
+ source_domain: ar5iv.labs.arxiv.org
10
+ fetched_at: '2026-06-10T00:40:46.873327Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2501.04519] \sysname: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
20
+ \sysname
21
+ : Small LLMs Can Master Math Reasoning
22
+ with Self-Evolved Deep Thinking
23
+ Xinyu Guan
24
+ ∗
25
+ Li Lyna Zhang
26
+ ∗⋄
27
+ Yifei Liu
28
+ Ning Shang   Youran Sun    Yi Zhu    Fan Yang    Mao Yang
29
+ Microsoft Research Asia
30
+ Abstract
31
+ We present
32
+ \sysname
33
+ to demonstrate that small language models (SLMs) can rival or even surpass the math reasoning capability of OpenAI o1, without distillation from superior models.
34
+ \sysname
35
+ achieves this by exercising “deep thinking” through Monte Carlo Tree Search (MCTS), where a math
36
+ policy SLM
37
+ performs test-time search guided by an SLM-based
38
+ process reward model
39
+ .
40
+ \sysname
41
+ introduces three innovations to tackle the challenges in training the two SLMs:
42
+ (1)
43
+ a novel code-augmented CoT data sythesis method, which performs extensive MCTS rollouts to generate
44
+ step-by-step verified reasoning trajectories
45
+ used to train the policy SLM;
46
+ (2)
47
+ a novel process reward model training method that avoids naïve step-level score annotation, yielding a more effective
48
+ process preference model (PPM)
49
+ ;
50
+ (3)
51
+ a
52
+ self-evolution recipe
53
+ in which the policy SLM and PPM are built from scratch and iteratively evolved to improve reasoning capabilities.
54
+ Through 4 rounds of self-evolution with millions of synthesized solutions for 747k math problems,
55
+ \sysname
56
+ boosts SLMs’ math reasoning to state-of-the-art levels. On the MATH benchmark, it improves Qwen2.5-Math-7B from 58.8% to 90.0% and Phi3-mini-3.8B from 41.4% to 86.4%, surpassing o1-preview by +4.5% and +0.9%. On the USA Math Olympiad (AIME),
57
+ \sysname
58
+ solves an average of 53.3% (8/15) of problems, ranking among the top 20% the brightest high school math students. Code and data will be available at
59
+ https://github.com/microsoft/rStar
60
+ .
61
+ Task
62
+ (pass@1 Acc)
63
+ rStar-Math
64
+ (Qwen-7B)
65
+ rStar-Math
66
+ (Qwen-1.5B)
67
+ rStar-Math
68
+ (Phi3-mini)
69
+ OpenAI
70
+ o1-preview
71
+ OpenAI
72
+ o1-mini
73
+ QWQ
74
+ 32B-preview
75
+ GPT-4o
76
+ DeepSeek-V3
77
+ MATH
78
+ 90.0
79
+ 88.6
80
+ 86.4
81
+ 85.5
82
+ 90.0
83
+ 90.6
84
+ 76.6
85
+ 90.2
86
+ AIME 2024
87
+ 53.3
88
+ 46.7
89
+ 43.3
90
+ 44.6
91
+ 56.7
92
+ 50.0
93
+ 9.3
94
+ 39.2
95
+ Olympiad Bench
96
+ 65.6
97
+ 64.6
98
+ 60.3
99
+ -
100
+ 65.3
101
+ 61.2
102
+ 43.3
103
+ 55.4
104
+ College Math
105
+ 60.5
106
+ 59.3
107
+ 59.1
108
+ -
109
+ 57.8
110
+ 55.8
111
+ 48.5
112
+ 58.9
113
+ Omni-Math
114
+ 50.5
115
+ 48.5
116
+ 46.0
117
+ 52.5
118
+ 60.5
119
+ 49.6
120
+ 30.5
121
+ 35.9
122
+ Table 1:
123
+ \sysname
124
+ enables frontier math reasoning in SLMs via deep thinking over 64 trajectories.
125
+ $*$
126
+ $*$
127
+ footnotetext:
128
+ Equal contribution.
129
+ $\diamond$
130
+ $\diamond$
131
+ footnotetext:
132
+ Project leader; correspondence to lzhani@microsoft.com
133
+ $\S$
134
+ $\S$
135
+ footnotetext:
136
+ Xinyu Guan and Youran Sun did this work during the internship at MSRA. Xinyu Guan (2001gxy@gmail.com) is with Peking University, Youran Sun is with Tsinghua University.
137
+ 1
138
+ Introduction
139
+ Recent studies have demonstrated that large language models (LLMs) are capable of tackling mathematical problems
140
+ (Team,
141
+ 2024a
142
+ ; Yang et al.,
143
+ 2024
144
+ ; OpenAI,
145
+ 2024
146
+ ; Liu et al.,
147
+ 2024
148
+ )
149
+ . However, the conventional approach of having LLMs generate complete solutions in a single inference – akin to System 1 thinking
150
+ (Daniel,
151
+ 2011
152
+ )
153
+ – often yields fast but error-prone results
154
+ (Valmeekam et al.,
155
+ 2023
156
+ ; OpenAI,
157
+ 2023
158
+ )
159
+ . In response, test-time compute scaling
160
+ (Snell et al.,
161
+ 2024
162
+ ; Qi et al.,
163
+ 2024
164
+ )
165
+ suggests a paradigm shift toward a System 2-style thinking, which emulates human reasoning through a slower and deeper thought process. In this paradigm, an LLM serves as a policy model to generate multiple math reasoning steps, which are then evaluated by another LLM acting as a reward model
166
+ (OpenAI,
167
+ 2024
168
+ )
169
+ . The steps and solutions deemed more likely to be correct are selected. The process repeats iteratively and ultimately derives the final answer.
170
+ In the test-time compute paradigm, the key is to train a powerful policy model that generates promising solution steps and a reliable reward model that accurately evaluates them, both of which depend on
171
+ high-quality
172
+ training data. Unfortunately, it is well-known that off-the-shelf high-quality math reasoning data is scarce, and synthesizing high-quality math data faces fundamental challenges.
173
+ For the policy model, it is challenging to distinguish erroneous reasoning steps from the correct ones, complicating the elimination of low-quality data. It is worth noting that in math reasoning, a correct final answer does not ensure the correctness of the entire reasoning trace
174
+ (Lanham et al.,
175
+ 2023
176
+ )
177
+ . Incorrect intermediate steps significantly decrease data quality.
178
+ As for the reward model, process reward modeling (PRM) shows a great potential by providing fine-grained feedback on intermediate steps
179
+ (Lightman et al.,
180
+ 2023
181
+ )
182
+ . However, the training data is even scarcer in this regard: accurate step-by-step feedback requires intense human labeling efforts and is impractical to scale, while those automatic annotation attempts show limited gains due to noisy reward scores
183
+ (Luo et al.,
184
+ 2024
185
+ ; Wang et al.,
186
+ 2024c
187
+ ; Chen et al.,
188
+ 2024
189
+ )
190
+ .
191
+ Due to the above challenges, existing distill-based data synthesis approaches to training policy models, e.g., scaling up GPT4-distilled CoT data
192
+ (Tang et al.,
193
+ 2024
194
+ ; Huang et al.,
195
+ 2024
196
+ )
197
+ , have shown diminishing returns and cannot exceed the capability of their teacher model; meanwhile, as of today, training reliable PRMs for math reasoning remains an open question.
198
+ Figure 1:
199
+ The overview of
200
+ \sysname
201
+ .
202
+ In this work, we introduce
203
+ \sysname
204
+ , a self-evolvable System 2-style reasoning approach that achieves the state-of-the-art math reasoning, rivaling and sometimes even surpassing OpenAI o1 on challenging math competition benchmarks with a model size as small as 7 billion. Unlike solutions relying on superior LLMs for data synthesis,
205
+ \sysname
206
+ leverages smaller language models (SLMs) with Monte Carlo Tree Search (MCTS) to establish a self-evolutionary process, iteratively generating higher-quality training data. To achieve self-evolution,
207
+ \sysname
208
+ introduces three key innovations.
209
+ First, a novel code-augmented CoT data synthesis method, which performs
210
+ extensive
211
+ MCTS rollouts to generate
212
+ step-by-step verified reasoning trajectories
213
+ with
214
+ self-annotated MCTS Q-values
215
+ . Specifically, math problem-solving is decomposed into multi-step generation within MCTS. At each step, the SLM serving as the policy model samples candidate nodes, each generating a one-step CoT and the corresponding Python code. To verify the generation quality, only nodes with successful Python code execution are retained, thus mitigating errors in intermediate steps. Moreover, extensive MCTS rollouts automatically assign a Q-value to each intermediate step based on its contribution: steps contributing to more trajectories that lead to the correct answer are given higher Q-values and considered higher quality. This ensures that the reasoning trajectories generated by SLMs consist of correct, high-quality intermediate steps.
216
+ Second, a novel method that trains an SLM acting as a
217
+ process preference model
218
+ , i.e., a PPM to implement the desired PRM, that reliably predicts a reward label for each math reasoning step. The PPM leverages the fact that, although Q-values are still not precise enough to score each reasoning step despite using extensive MCTS rollouts, the Q-values can reliably distinguish positive (correct) steps from negative (irrelevant/incorrect) ones. Thus the training method constructs preference pairs for each step based on Q-values and uses a pairwise ranking loss
219
+ (Ouyang et al.,
220
+ 2022
221
+ )
222
+ to optimize PPM’s score prediction for each reasoning step, achieving reliable labeling. This approach avoids conventional methods that directly use Q-values as reward labels
223
+ (Luo et al.,
224
+ 2024
225
+ ; Chen et al.,
226
+ 2024
227
+ )
228
+ , which are inherently noisy and imprecise in stepwise reward assignment.
229
+ Finally, a four-round self-evolution recipe that progressively builds both a frontier policy model and PPM from scratch. We begin by curating a dataset of 747k math word problems from publicly available sources. In each round, we use the latest policy model and PPM to perform MCTS, generating increasingly high-quality training data using the above two methods to train a stronger policy model and PPM for next round. Each round achieves progressive refinement: (1) a stronger policy SLM, (2) a more reliable PPM, (3) generating better reasoning trajectories via PPM-augmented MCTS, and (4) improving training data coverage to tackle more challenging and even competition-level math problems.
230
+ Extensive experiments across four SLMs (1.5B-7B) and seven math reasoning tasks demonstrate the effectiveness of
231
+ \sysname
232
+ . Remarkably,
233
+ \sysname
234
+ improves all four SLMs, matching or even surpassing OpenAI o1 on challenging math benchmarks. On MATH benchmark, with 8 search trajectories,
235
+ \sysname
236
+ boosts Qwen2.5-Math-7B from 58.8% to 89.4% and Qwen2.5-Math-1.5B from 51.2% to 87.8%. With 64 trajectories, the scores rise to 90% and 88.4%, outperforming o1-preview by 4.5% and 2.6% and matching o1-mini’s 90%. On the Olympiad-level AIME 2024,
237
+ \sysname
238
+ solves on average 53.3% (8/15) of the problems, exceeding o1-preview by 8.7% and all other open-sourced LLMs. We further conduct comprehensive experiments to verify the superiority of step-by-step verified reasoning trajectories over state-of-the-art data synthesis baselines, as well as the PPM’s effectiveness compared to outcome reward models and Q value-based PRMs. Finally, we present key findings from
239
+ \sysname
240
+ deep thinking, including the intrinsic self-reflection capability and PPM’s preference for theorem-applications intermediate steps.
241
+ 2
242
+ Related Works
243
+ Math Data Synthesis
244
+ . Advancements in LLM math reasoning have largely relied on curating high-quality CoT data, with most leading approaches being GPT-distilled, using frontier models like GPT-4 for synthesis
245
+ (Wang et al.,
246
+ 2024b
247
+ ; Gou et al.,
248
+ 2023
249
+ ; Luo et al.,
250
+ 2023
251
+ )
252
+ . Notable works include NuminaMath
253
+ (Jia LI and Polu,
254
+ 2024a
255
+ )
256
+ and
257
+ MetaMath
258
+ (Yu et al.,
259
+ 2023b
260
+ )
261
+ . While effective, this limits reasoning to the capabilities of the teacher LLM.
262
+ Hard problems that the teacher LLM cannot solve are excluded in the training set.
263
+ Even solvable problems may contain error-prone intermediate steps, which are hard to detect. Although rejection sampling methods
264
+ (Yuan et al.,
265
+ 2023
266
+ ; Brown et al.,
267
+ 2024
268
+ )
269
+ can improve data quality,
270
+ they do not guarantee correct intermediate steps. As a result, scaling up CoT data has diminishing returns, with gains nearing saturation—e.g., OpenMathInstruct-2
271
+ (Toshniwal et al.,
272
+ 2024
273
+ )
274
+ only sees a 3.9% boost on MATH despite an 8× increase in dataset size.
275
+ Scaling Test-time Compute
276
+ has introduced new scaling laws, allowing LLMs to improve performance across by generating multiple samples and using reward models for best-solution selection
277
+ (Snell et al.,
278
+ 2024
279
+ ; Wu et al.,
280
+ 2024
281
+ ; Brown et al.,
282
+ 2024
283
+ )
284
+ . Various test-time search methods have been proposed
285
+ (Kang et al.,
286
+ 2024
287
+ ; Wang et al.,
288
+ 2024a
289
+ )
290
+ , including random sampling
291
+ (Wang et al.,
292
+ 2023
293
+ )
294
+ and tree-search methods
295
+ (Yao et al.,
296
+ 2024
297
+ ; Hao et al.,
298
+ 2023
299
+ ; Zhang et al.,
300
+ 2024b
301
+ ; Qi et al.,
302
+ 2024
303
+ )
304
+ like MCTS. However, open-source methods for scaling test-time computation have shown limited gains in math reasoning, often due to policy LLM or reward model limitations.
305
+ \sysname
306
+ addresses this by iteratively evolving the policy LLM and reward model, achieving System 2 mathematical reasoning performance comparable to OpenAI o1
307
+ (OpenAI,
308
+ 2024
309
+ )
310
+ .
311
+ Reward Models
312
+ are crucial for effective System 2 reasoning but are challenging to obtain. Recent works include LLM-as-a-Judge for verification
313
+ (Zheng et al.,
314
+ 2023
315
+ ; Qi et al.,
316
+ 2024
317
+ )
318
+ and specialized reward models like Outcome Reward Model
319
+ (Yang et al.,
320
+ 2024
321
+ ; Yu et al.,
322
+ 2023a
323
+ )
324
+ and Process Reward Model (PRM)
325
+ (Lightman et al.,
326
+ 2024
327
+ )
328
+ . While PRMs offer promising dense, step-level reward signals for
329
+ complex reasoning
330
+ (Luo et al.,
331
+ 2024
332
+ ; Wang et al.,
333
+ 2024c
334
+ )
335
+ , collecting step-level annotations remains an obstacle. While
336
+ Kang et al. (
337
+ 2024
338
+ ); Wang et al. (
339
+ 2024a
340
+ )
341
+ rely on costly human-annotated datasets like PRM800k
342
+ (Lightman et al.,
343
+ 2024
344
+ )
345
+ ,
346
+ recent approaches
347
+ (Wang et al.,
348
+ 2024c
349
+ ; Luo et al.,
350
+ 2024
351
+ )
352
+ explore automated annotation via Monte Carlo Sampling or MCTS. However, they struggle to generate precise reward scores, which limits performance gains.
353
+ \sysname
354
+ introduces a novel process preference reward (PPM) that eliminates the need for accurate step-level reward score annotation.
355
+ 3
356
+ Methodology
357
+ 3.1
358
+ Design Choices
359
+ MCTS for Effective System 2 Reasoning
360
+ .
361
+ We aim to train a math policy SLM and a process reward model (PRM), and integrating both within Monte Carlo Tree Search (MCTS) for System 2 deep thinking. MCTS is chosen for two key reasons. First, it breaks down complex math problems into simpler single-step generation tasks, reducing the difficulty for the policy SLM compared to other System 2 methods like Best-of-N
362
+ (Brown et al.,
363
+ 2024
364
+ )
365
+ or self-consistency
366
+ (Wang et al.,
367
+ 2023
368
+ )
369
+ , which require generating full solutions in one inference.
370
+ Second, the step-by-step generation in MCTS naturally yields step-level training data for both models. Standard MCTS rollout automatically assign Q-value to each step based on its contribution to the final correct answer, obviating the need for human-generated step-level annotations for process reward model training.
371
+ Ideally, advanced LLMs such as GPT-4 could be integrated within MCTS to generate training data. However, this approach faces two key challenges. First, even these powerful models struggle to consistently solve difficult problems, such as Olympiad-level mathematics. Consequently, the resulting training data would primarily consist of simpler solvable problems, limiting its diversity and quality. Second, annotating per-step Q-values demands extensive MCTS rollouts; insufficient tree exploration can lead to spurious Q-value assignments, such as overestimating suboptimal steps. Given that each rollout involves multiple single-step generations and these models are computationally expensive, increasing rollouts significantly raises inference costs.
372
+ Overview
373
+ . To this end, we explore using two 7B SLMs (a policy SLM and a PRM) to generate higher-quality training data, with their smaller size allowing for extensive MCTS rollouts on accessible hardware (e.g., 4
374
+ ×
375
+ \times
376
+ 40GB A100 GPUs). However, self-generating data presents greater challenges for SLMs, due to their weaker capabilities.
377
+ SLMs frequently fail to generate correct solutions, and even when the final answer is correct, the intermediate steps are often flawed or of poor quality. Moreover, SLMs solve fewer challenging problems compared to advanced models like GPT-4.
378
+ This section introduces our methodology, as illustrated in Fig.
379
+ 1
380
+ . To mitigate errors and low-quality intermediate steps, we introduce a code-augmented CoT synthetic method, which performs extensive MCTS rollouts to generate step-by-step verified reasoning trajectories, annotated with Q-values. To further improve SLM performance on challenging problems, we introduce a four-round self-evolution recipe. In each round, both the policy SLM and the reward model are updated to stronger versions, progressively tackling more difficult problems and generating higher-quality training data. Finally, we present a novel process reward model training approach that eliminates the need for precise per-step reward annotations, yielding the more
381
+ effective process preference model (PPM).
382
+ 3.2
383
+ Step-by-Step Verified Reasoning Trajectory
384
+ We start by introducing our method for generating step-by-step verified reasoning trajectories with per-step Q-value annotations. Given a problem
385
+ x
386
+ x
387
+ and a policy model
388
+ M
389
+ M
390
+ , we run the standard MCTS to incrementally construct a search tree for step-by-step solution exploration. As shown in Fig.
391
+ 1
392
+ (a),
393
+ the root node represents question
394
+ x
395
+ x
396
+ , while child nodes correspond to intermediate steps
397
+ s
398
+ s
399
+ generated by
400
+ M
401
+ M
402
+ . A root-to-leaf path ending at terminal node
403
+ s
404
+ d
405
+ s_{d}
406
+ forms a trajectory
407
+ 𝐭
408
+ =
409
+ x
410
+ ⊕
411
+ s
412
+ 1
413
+ ⊕
414
+ s
415
+ 2
416
+ ⊕
417
+ …
418
+ ⊕
419
+ s
420
+ d
421
+ \mathbf{t}=x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{d}
422
+ , with each step
423
+ s
424
+ i
425
+ s_{i}
426
+ assigned a Q-value
427
+ Q
428
+ ​
429
+ (
430
+ s
431
+ i
432
+ )
433
+ Q(s_{i})
434
+ .
435
+ From the search tree
436
+ 𝒯
437
+ \mathcal{T}
438
+ , we extract solution trajectories
439
+ 𝕋
440
+ =
441
+ {
442
+ 𝐭
443
+ 1
444
+ ,
445
+ 𝐭
446
+ 2
447
+ ,
448
+ …
449
+ ,
450
+ 𝐭
451
+ n
452
+ }
453
+ ​
454
+ (
455
+ n
456
+ ≥
457
+ 1
458
+ )
459
+ \mathbb{T}=\{\mathbf{t}^{1},\mathbf{t}^{2},...,\mathbf{t}^{n}\}(n\geq 1)
460
+ . Our goal is to select high-quality trajectories from
461
+ 𝒯
462
+ \mathcal{T}
463
+ to construct the training set. For this purpose, we introduce code-augmented CoT synthesis method to filter out low-quality generations and perform extensive rollouts to improve the reliability of Q-value accuracy.
464
+ Code-augmented CoT Generation
465
+ . Prior MCTS approaches primarily generate natural language (NL) CoTs
466
+ (Qi et al.,
467
+ 2024
468
+ ; Zhang et al.,
469
+ 2024a
470
+ )
471
+ . However, LLMs often suffer from hallucination, producing incorrect or irrelevant steps yet still arrive at the correct answer by chance
472
+ (Lanham et al.,
473
+ 2023
474
+ )
475
+ . These flawed steps are challenging to detect and eliminate. To address this, we propose a novel code execution augmented CoT. As shown in Fig.
476
+ 2
477
+ , the policy model generates a one-step NL CoT alongside its corresponding Python code, where the NL CoT is embedded as a Python comment. Only generations with successfully executed Python code are retained as valid candidates.
478
+ Figure 2:
479
+ An example of Code-augmented CoT.
480
+ Specifically, starting from the initial root node
481
+ x
482
+ x
483
+ , we perform multiple MCTS iterations through
484
+ selection
485
+ ,
486
+ expansion
487
+ ,
488
+ rollout
489
+ , and
490
+ back-propagation
491
+ . At step
492
+ i
493
+ i
494
+ , we collect the latest reasoning trajectory
495
+ x
496
+ ⊕
497
+ s
498
+ 1
499
+ ⊕
500
+ s
501
+ 2
502
+ ⊕
503
+ …
504
+ ⊕
505
+ s
506
+ i
507
+ −
508
+ 1
509
+ x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{i-1}
510
+ as the current state. Based on this state, we prompt (see Appendix
511
+ A.3
512
+ ) the policy model to generate
513
+ n
514
+ n
515
+ candidates
516
+ s
517
+ i
518
+ ,
519
+ 0
520
+ ,
521
+ …
522
+ ,
523
+ s
524
+ i
525
+ ,
526
+ n
527
+ −
528
+ 1
529
+ s_{i,0},...,s_{i,n-1}
530
+ for step
531
+ i
532
+ i
533
+ . Python code execution is then employed to filter valid nodes. As shown in Fig.
534
+ 2
535
+ , each generation
536
+ s
537
+ i
538
+ ,
539
+ j
540
+ s_{i,j}
541
+ is concatenated with the code from all previous steps, forming
542
+ s
543
+ 1
544
+ ⊕
545
+ s
546
+ 2
547
+ ⊕
548
+ …
549
+ ⊕
550
+ s
551
+ i
552
+ −
553
+ 1
554
+ ⊕
555
+ s
556
+ i
557
+ ,
558
+ j
559
+ s_{1}\oplus s_{2}\oplus...\oplus s_{i-1}\oplus s_{i,j}
560
+ . Candidates that execute successfully are retained as valid nodes and scored by the PPM, which assigns a Q-value
561
+ q
562
+ ​
563
+ (
564
+ s
565
+ i
566
+ )
567
+ q(s_{i})
568
+ .
569
+ Then, we use the well-known Upper Confidence bounds for Trees (UCT)
570
+ (Kocsis and Szepesvári,
571
+ 2006
572
+ )
573
+ to select the best node among the
574
+ n
575
+ n
576
+ candidates. This selection process is mathematically represented as:
577
+ UCT
578
+ ​
579
+ (
580
+ s
581
+ )
582
+ =
583
+ Q
584
+ ​
585
+ (
586
+ s
587
+ )
588
+ +
589
+ c
590
+ ​
591
+ ln
592
+ ⁡
593
+ N
594
+ p
595
+ ​
596
+ a
597
+ ​
598
+ r
599
+ ​
600
+ e
601
+ ​
602
+ n
603
+ ​
604
+ t
605
+ ​
606
+ (
607
+ s
608
+ )
609
+ N
610
+ ​
611
+ (
612
+ s
613
+ )
614
+ ;
615
+ where
616
+ Q
617
+ ​
618
+ (
619
+ s
620
+ )
621
+ =
622
+ q
623
+ ​
624
+ (
625
+ s
626
+ )
627
+ N
628
+ ​
629
+ (
630
+ s
631
+ )
632
+ \displaystyle\text{UCT}(s)=Q(s)+c\sqrt{\frac{\ln N_{parent}(s)}{N(s)}};\quad\text{where}\quad Q(s)=\frac{q(s)}{N(s)}
633
+ (1)
634
+ where
635
+ N
636
+ ​
637
+ (
638
+ s
639
+ )
640
+ N(s)
641
+ denotes the number of visits to node
642
+ s
643
+ s
644
+ , and
645
+ N
646
+ parent
647
+ ​
648
+ (
649
+ s
650
+ )
651
+ N_{\text{parent}}(s)
652
+ is the visit count of
653
+ s
654
+ s
655
+ ’s parent node. The predicted reward
656
+ q
657
+ ​
658
+ (
659
+ s
660
+ )
661
+ q(s)
662
+ is provided by the PPM and will be updated through back-propagation.
663
+ c
664
+ c
665
+ is a constant that balances exploitation and exploration.
666
+ Extensive Rollouts for Q-value Annotation
667
+ . Accurate Q-value
668
+ Q
669
+ ​
670
+ (
671
+ s
672
+ )
673
+ Q(s)
674
+ annotation in Eq.
675
+ 1
676
+ is crucial for guiding MCTS node selection towards correct problem-solving paths and identifying high-quality steps within trajectories.
677
+ To improve Q-value reliability, we draw inspiration from Go players, who retrospectively evaluate the reward of each move based on game outcomes. Although initial estimates may be imprecise, repeated gameplay refines these evaluations over time. Similarly, in each rollout, we update the Q-value of each step based on its contribution to achieving the correct final answer. After extensive MCTS rollouts, steps consistently leading to correct answers achieve higher Q-values, occasional successes yield moderate Q-values, and consistently incorrect steps receive low Q-values. Specifically, we introduce two self-annotation methods to obtain these step-level Q-values. Fig.
678
+ 1
679
+ (c) shows the detailed setting in the four rounds of self-evolution.
680
+ Terminal-guided annotation
681
+ . During the first two rounds, when the PPM is unavailable or insufficiently accurate, we use terminal-guided annotation. Formally, let
682
+ q
683
+ ​
684
+ (
685
+ s
686
+ i
687
+ )
688
+ k
689
+ q(s_{i})^{k}
690
+ denote the q value for step
691
+ s
692
+ i
693
+ s_{i}
694
+ after back-propagation in the
695
+ k
696
+ t
697
+ ​
698
+ h
699
+ k^{th}
700
+ rollout. Following AlphaGo
701
+ (Silver et al.,
702
+ 2017
703
+ )
704
+ and rStar
705
+ (Qi et al.,
706
+ 2024
707
+ )
708
+ , we score each intermediate node based on its contribution to the final correct answer:
709
+ q
710
+ ​
711
+ (
712
+ s
713
+ i
714
+ )
715
+ k
716
+ =
717
+ q
718
+ ​
719
+ (
720
+ s
721
+ i
722
+ )
723
+ k
724
+ −
725
+ 1
726
+ +
727
+ q
728
+ ​
729
+ (
730
+ s
731
+ d
732
+ )
733
+ k
734
+ ;
735
+ \displaystyle q(s_{i})^{k}=q(s_{i})^{k-1}+q(s_{d})^{k};
736
+ (2)
737
+ where the initial q value
738
+ q
739
+ ​
740
+ (
741
+ s
742
+ i
743
+ )
744
+ 0
745
+ =
746
+ 0
747
+ q(s_{i})^{0}=0
748
+ in the first rollout. If this step frequently leads to a correct answer, its
749
+ q
750
+ q
751
+ value will increase; otherwise, it decreases. Terminal nodes are scored as
752
+ q
753
+ ​
754
+ (
755
+ s
756
+ d
757
+ )
758
+ =
759
+ 1
760
+ q(s_{d})=1
761
+ for correct answers and
762
+ q
763
+ ​
764
+ (
765
+ s
766
+ d
767
+ )
768
+ =
769
+ −
770
+ 1
771
+ q(s_{d})=-1
772
+ otherwise, as shown in Fig.
773
+ 1
774
+ .
775
+ PRM-augmented annotation
776
+ . Starting from the third round, we use PPM to score each step for more effective generation. Compared to terminal-guided annotation, which requires multiple rollouts for a meaningful
777
+ q
778
+ q
779
+ value, PPM directly predicts a non-zero initial
780
+ q
781
+ q
782
+ value.
783
+ PPM-augmented MCTS also helps the policy model to generate higher-quality steps, guiding solutions towards correct paths. Formally, for step
784
+ s
785
+ i
786
+ s_{i}
787
+ , PPM predicts an initial
788
+ q
789
+ ​
790
+ (
791
+ s
792
+ i
793
+ )
794
+ 0
795
+ q(s_{i})^{0}
796
+ value based on the partial trajectory:
797
+ q
798
+ ​
799
+ (
800
+ s
801
+ i
802
+ )
803
+ 0
804
+ =
805
+ P
806
+ ​
807
+ P
808
+ ​
809
+ M
810
+ ​
811
+ (
812
+ x
813
+ ⊕
814
+ s
815
+ 1
816
+ ⊕
817
+ s
818
+ 2
819
+ ⊕
820
+ …
821
+ ⊕
822
+ s
823
+ i
824
+ −
825
+ 1
826
+ ⊕
827
+ s
828
+ i
829
+ )
830
+ \displaystyle q(s_{i})^{0}=PPM(x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{i-1}\oplus s_{i})
831
+ (3)
832
+ This
833
+ q
834
+ q
835
+ value will be updated based on terminal node’s
836
+ q
837
+ ​
838
+ (
839
+ s
840
+ d
841
+ )
842
+ q(s_{d})
843
+ value through MCTS
844
+ back-propagation
845
+ in Eq.
846
+ 2
847
+ .
848
+ For terminal node
849
+ s
850
+ d
851
+ s_{d}
852
+ , we do not use PRM for scoring during training data generation. Instead, we assign a more accurate score based on ground truth labels as terminal-guided rewarding.
853
+ 3.3
854
+ Process Preference Model
855
+ Process reward models, which provide granular step-level reward signals, is highly desirable for solving challenging math problems. However, obtaining high-quality step-level training data remains an open challenge. Existing methods rely on human annotations
856
+ (Lightman et al.,
857
+ 2023
858
+ )
859
+ or MCTS-generated scores
860
+ (Zhang et al.,
861
+ 2024a
862
+ ; Chen et al.,
863
+ 2024
864
+ )
865
+ to assign a score for each step. These scores then serve as training targets, with methods such as MSE loss
866
+ (Chen et al.,
867
+ 2024
868
+ )
869
+ or pointwise loss
870
+ (Wang et al.,
871
+ 2024c
872
+ ; Luo et al.,
873
+ 2024
874
+ ; Zhang et al.,
875
+ 2024a
876
+ )
877
+ used to minimize the difference between predicted and labeled scores.
878
+ As a result, the precision of these annotated step-level reward scores directly determines the effectiveness of the resulting process reward model.
879
+ Unfortunately, precise per-step scoring remains a unsolved challenge. Although our extensive MCTS rollouts improve the reliability of Q-values, precisely evaluating fine-grained step quality presents a major obstacle. For instance, among a set of correct steps, it is difficult to rank them as best, second-best, or average and then assign precise scores. Similarly, among incorrect steps, differentiating the worst from moderately poor steps poses analogous challenges. Even expert human annotation struggles with consistency, particularly at scale, leading to inherent noise in training labels.
880
+ We introduce a novel training method that trains a process preference model (PPM) by constructing step-level positive-negative preference pairs. As shown in Fig.
881
+ 1
882
+ (b), instead of using Q-values as direct reward labels, we use them to select steps from MCTS tree for preference pair construction. For each step, we select two candidates with the highest Q-values as positive steps and two with the lowest as negative steps. Critically, the selected positive steps must lead to a correct final answer, while negative steps must lead to incorrect answers. For intermediate steps (except the final answer step), the positive and negative pairs share the same preceding steps. For the final answer step, where identical reasoning trajectories rarely yield different final answers, we relax this restriction.
883
+ We select two correct trajectories with the highest average Q-values as positive examples and two incorrect trajectories with the lowest average Q-values as negative examples. Following
884
+ (Ouyang et al.,
885
+ 2022
886
+ )
887
+ , we define our loss function using the standard Bradley-Terry model with a pairwise ranking loss:
888
+ ℒ
889
+ p
890
+ ​
891
+ p
892
+ ​
893
+ m
894
+ ​
895
+ (
896
+ θ
897
+ )
898
+ =
899
+ −
900
+ 1
901
+ 2
902
+ ×
903
+ 2
904
+ ​
905
+ E
906
+ (
907
+ x
908
+ ,
909
+ y
910
+ i
911
+ p
912
+ ​
913
+ o
914
+ ​
915
+ s
916
+ ,
917
+ y
918
+ i
919
+ n
920
+ ​
921
+ e
922
+ ​
923
+ g
924
+ ∈
925
+ 𝔻
926
+ )
927
+ ​
928
+ [
929
+ l
930
+ ​
931
+ o
932
+ ​
933
+ g
934
+ ​
935
+ (
936
+ σ
937
+ ​
938
+ (
939
+ r
940
+ θ
941
+ ​
942
+ (
943
+ x
944
+ ,
945
+ y
946
+ i
947
+ p
948
+ ​
949
+ o
950
+ ​
951
+ s
952
+ )
953
+ −
954
+ r
955
+ θ
956
+ ​
957
+ (
958
+ x
959
+ ,
960
+ y
961
+ i
962
+ n
963
+ ​
964
+ e
965
+ ​
966
+ g
967
+ )
968
+ )
969
+ )
970
+ ]
971
+ \displaystyle\mathcal{L}_{ppm}(\theta)=-\frac{1}{2\times 2}E_{(x,y_{i}^{pos},y_{i}^{neg}\in\mathbb{D})}[log(\sigma(r_{\theta}(x,y_{i}^{pos})-r_{\theta}(x,y_{i}^{neg})))]
972
+ (4)
973
+ when
974
+ i
975
+ is not final answer step
976
+ ,
977
+ y
978
+ i
979
+ p
980
+ ​
981
+ o
982
+ ​
983
+ s
984
+ =
985
+ s
986
+ 1
987
+ ⊕
988
+ …
989
+ ⊕
990
+ s
991
+ i
992
+ −
993
+ 1
994
+ ⊕
995
+ s
996
+ i
997
+ p
998
+ ​
999
+ o
1000
+ ​
1001
+ s
1002
+ ;
1003
+ y
1004
+ i
1005
+ n
1006
+ ​
1007
+ e
1008
+ ​
1009
+ g
1010
+ =
1011
+ s
1012
+ 1
1013
+ ⊕
1014
+ …
1015
+ ⊕
1016
+ s
1017
+ i
1018
+ −
1019
+ 1
1020
+ ⊕
1021
+ s
1022
+ i
1023
+ n
1024
+ ​
1025
+ e
1026
+ ​
1027
+ g
1028
+ \displaystyle\text{when $i$ is not final answer step},y_{i}^{pos}=s_{1}\oplus...\oplus s_{i-1}\oplus s_{i}^{pos};y_{i}^{neg}=s_{1}\oplus...\oplus s_{i-1}\oplus s_{i}^{neg}\vskip-4.30554pt
1029
+ (5)
1030
+ Here,
1031
+ r
1032
+ θ
1033
+ ​
1034
+ (
1035
+ x
1036
+ ,
1037
+ y
1038
+ i
1039
+ )
1040
+ r_{\theta}(x,y_{i})
1041
+ denotes the output of the PPM, where
1042
+ x
1043
+ x
1044
+ is the problem and
1045
+ y
1046
+ y
1047
+ is the trajectory from the first step to the
1048
+ i
1049
+ t
1050
+ ​
1051
+ h
1052
+ i^{th}
1053
+ step.
1054
+ 3.4
1055
+ Self-Evolved Deep Thinking
1056
+ 3.4.1
1057
+ Training with Step-by-Step Verified Reasoning Trajectory
1058
+ Math Problems Collection
1059
+ . We collect a large dataset of 747k math word problems with final answer ground-truth labels, primarily from NuminaMath
1060
+ (Jia LI and Polu,
1061
+ 2024a
1062
+ )
1063
+ and MetaMath
1064
+ (Yu et al.,
1065
+ 2023b
1066
+ )
1067
+ . Notably, only competition-level problems (e.g., Olympiads and AIME/AMC) from NuminaMath are included, as we observe that grade-school-level problems do not significantly improve LLM complex math reasoning. To augment the limited competition-level problems, we follow
1068
+ (Li et al.,
1069
+ 2024
1070
+ )
1071
+ and use GPT-4 to synthesize new problems based on the seed problems in 7.5k MATH train set and 3.6k AMC-AIME training split. However, GPT-4 often generated unsolvable problems or incorrect solutions for challenging seed problems. To filter these, we prompt GPT-4 to generate 10 solutions per problem, retaining only those with at least 3 consistent solutions.
1072
+ Reasoning Trajectories Collection
1073
+ . Instead of using the original solutions in the 747k math dataset, we conduct extensive MCTS rollouts (Sec.
1074
+ 3.2
1075
+ ) to generate higher-quality step-by-step verified reasoning trajectories. In each self-evolution round, we perform 16 rollouts per math problem, which leads to 16 reasoning trajectories. Problems are then categories by difficulty based on the correct ratio of the generated trajectories:
1076
+ easy
1077
+ (all solutions are correct),
1078
+ medium
1079
+ (a mix of correct and incorrect solutions) and
1080
+ hard
1081
+ (all solutions are incorrect). For
1082
+ hard
1083
+ problems with no correct trajectories, an additional MCTS with 16 rollouts is performed. After that, all step-by-step trajectories and their annotated Q-values are collected and filtered to train the policy SLM and process preference model.
1084
+ Supervised Fine-tuning the Policy SLM
1085
+ . Through extensive experiments, we find that selecting high-quality reasoning trajectories is the key for fine-tuning a frontier math LLM. While methods such as GPT-distillation and Best-of-N can include low-quality or erroneous intermediate steps, a more effective approach ensures that every step in the trajectory is of high quality. To achieve this, we use per-step Q-values to select optimal trajectories from MCTS rollouts. Specifically, for each math problem, we select the top-2 trajectories with the highest average Q-values among those leading to correct answers as SFT training data.
1086
+ Training PPM
1087
+ . The PPM is initialized from the fine-tuned policy model, with its next-token prediction head replaced by a scalar-value head consisting of a linear layer and a tanh function to constrain outputs to the range [-1, 1]. We filter out math problems where all solution trajectories are fully correct or incorrect. For problems with mixed outcomes, we select two positive and two negative examples for each step based on Q-values, which are used as preference pairs for training data.
1088
+ 3.4.2
1089
+ Recipe for Self-Evolution
1090
+ Table 2:
1091
+ Percentage of the 747k math problems correctly solved in each round. Only problems have correct solutions are included in the training set. The first round uses DeepSeek-Coder-Instruct as the policy LLM, while later rounds use our fine-tuned 7B policy SLM.
1092
+ #
1093
+ models in MCTS
1094
+ GSM-level
1095
+ MATH-level
1096
+ Olympiad-level
1097
+ All
1098
+ Round 1
1099
+ DeepSeek-Coder-V2-Instruct
1100
+ 96.61%
1101
+ 67.36%
1102
+ 20.99%
1103
+ 60.17%
1104
+ Round 2
1105
+ policy SLM-r1
1106
+ 97.88%
1107
+ 67.40%
1108
+ 56.04%
1109
+ 66.60%
1110
+ Round 3
1111
+ policy SLM-r2, PPM-r2
1112
+ 98.15%
1113
+ 88.69%
1114
+ 62.16%
1115
+ 77.86%
1116
+ Round 4
1117
+ policy SLM-r3, PPM-r3
1118
+ 98.15%
1119
+ 94.53%
1120
+ 80.58%
1121
+ 90.25%
1122
+ Table 3:
1123
+ Pass@1 accuracy of the resulting policy SLM in each round, showing continuous improvement until surpassing the bootstrap model.
1124
+ Round#
1125
+ MATH
1126
+ AIME 2024
1127
+ AMC 2023
1128
+ Olympiad Bench
1129
+ College Math
1130
+ GSM8K
1131
+ GaokaoEn 2023
1132
+ DeepSeek-Coder-V2-Instruct
1133
+ (bootstrap model)
1134
+ 75.3
1135
+ 13.3
1136
+ 57.5
1137
+ 37.6
1138
+ 46.2
1139
+ 94.9
1140
+ 64.7
1141
+ Base (Qwen2.5-Math-7B)
1142
+ 58.8
1143
+ 0.0
1144
+ 22.5
1145
+ 21.8
1146
+ 41.6
1147
+ 91.6
1148
+ 51.7
1149
+ \hdashline
1150
+ policy SLM-r1
1151
+ 69.6
1152
+ 3.3
1153
+ 30.0
1154
+ 34.7
1155
+ 44.5
1156
+ 88.4
1157
+ 57.4
1158
+ policy SLM-r2
1159
+ 73.6
1160
+ 10.0
1161
+ 35.0
1162
+ 39.0
1163
+ 45.7
1164
+ 89.1
1165
+ 59.7
1166
+ policy SLM-r3
1167
+ 75.8
1168
+ 16.7
1169
+ 45.0
1170
+ 44.1
1171
+ 49.6
1172
+ 89.3
1173
+ 62.8
1174
+ policy SLM-r4
1175
+ 78.4
1176
+ 26.7
1177
+ 47.5
1178
+ 47.1
1179
+ 52.5
1180
+ 89.7
1181
+ 65.7
1182
+ Table 4:
1183
+ The quality of PPM consistently improves across rounds. The policy model has been fixed with policy SLM-r1 for a fair comparison.
1184
+ Round#
1185
+ MATH
1186
+ AIME 2024
1187
+ AMC 2023
1188
+ Olympiad Bench
1189
+ College Math
1190
+ GSM8K
1191
+ GaokaoEn 2023
1192
+ PPM-r1
1193
+ 75.2
1194
+ 10.0
1195
+ 57.5
1196
+ 35.7
1197
+ 45.4
1198
+ 90.9
1199
+ 60.3
1200
+ PPM-r2
1201
+ 84.1
1202
+ 26.7
1203
+ 75.0
1204
+ 52.7
1205
+ 54.2
1206
+ 93.3
1207
+ 73.0
1208
+ PPM-r3
1209
+ 85.2
1210
+ 33.3
1211
+ 77.5
1212
+ 59.5
1213
+ 55.6
1214
+ 93.9
1215
+ 76.6
1216
+ PPM-r4
1217
+ 87.0
1218
+ 43.3
1219
+ 77.5
1220
+ 61.5
1221
+ 56.8
1222
+ 94.2
1223
+ 77.8
1224
+ Due to the weaker capabilities of SLMs, we perform four rounds of MCTS deep thinking to progressively generate higher-quality data and expand the training set with more challenging math problems. Each round uses MCTS to generate step-by-step verified reasoning trajectories, which are then used to train the new policy SLM and PPM. The new models are then applied in next round to generate higher-quality training data. Fig.
1225
+ 1
1226
+ (c) and Table
1227
+ 2
1228
+ detail the models used for data generation in each round, along with the identifiers of the trained policy model and PPM. Next, we outline the details and specific improvements targeted in each round.
1229
+ Round 1: Bootstrapping an initial strong policy SLM-r1
1230
+ . To enable SLMs to self-generate reasonably good training data, we perform a bootstrap round to fine-tune an initial strong policy model, denoted as SLM-r1.
1231
+ As shown in Table
1232
+ 2
1233
+ , we run MCTS with DeepSeek-Coder-V2-Instruct (236B) to collect the SFT data. With no available reward model in this round, we use terminal-guided annotation for Q-values and limit MCTS to 8 rollouts for efficiency. For correct solutions, the top-2 trajectories with the highest average Q-values are selected as SFT data. We also train PPM-r1, but the limited rollouts yields unreliable Q-values, affecting the effectiveness of PPM-r1 ( Table
1234
+ 4
1235
+ ).
1236
+ Round 2: Training a reliable PPM-r2
1237
+ . In this round, with the policy model updated to the 7B SLM-r1, we conduct extensive MCTS rollouts for more reliable Q-value annotation and train the first reliable reward model, PPM-r2. Specifically, we perform 16 MCTS rollouts per problem. The resulting step-by-step verified reasoning trajectories show significant improvements in both quality and Q-value precision. As shown in Table
1238
+ 4
1239
+ , PPM-r2 is notably more effective than in the bootstrap round. Moreover, the policy SLM-r2 also continues to improve as expected (Table
1240
+ 3
1241
+ ).
1242
+ Round 3: PPM-augmented MCTS to significantly improve data quality
1243
+ . With the reliable PPM-r2, we perform PPM-augmented MCTS in this round to generate data, leading to significantly higher-quality trajectories that cover more math and Olympiad-level problems in the training set (Table
1244
+ 2
1245
+ ). The generated reasoning trajectories and self-annotated Q-values are then used to train the new policy SLM-r3 and PPM-r3, both of which show significant improvements.
1246
+ Round 4: Solving challenging math problems
1247
+ . After the third round, while grade school and MATH problems achieve high success rates, only 62.16% of Olympiad-level problems are included in the training set. This is
1248
+ NOT
1249
+ solely due to weak reasoning abilities in our SLMs, as many Olympiad problems remain unsolved by GPT-4 or o1. To improve coverage, we adopt a straightforward strategy. For unsolved problems after 16 MCTS rollouts, we perform an additional 64 rollouts, and if needed, increase to 128. We also conduct multiple MCTS tree expansions with different random seeds. This boosts the success rate of Olympiad-level problems to 80.58%.
1250
+ After four rounds of self-evolution, 90.25% of the 747k math problems are successfully covered into the training set, as shown in Table
1251
+ 2
1252
+ . Among the remaining unsolved problems, a significant portion consists of synthetic questions. We manually review a random sample of 20 problems and find that 19 are incorrectly labeled with wrong answers. Based on this, we conclude that the remaining unsolved problems are of low quality and thus terminate the self-evolution at round 4.
1253
+ 4
1254
+ Evaluation
1255
+ 4.1
1256
+ Setup
1257
+ Evaluation Datasets
1258
+ . We evaluate
1259
+ \sysname
1260
+ on diverse mathematical benchmarks. In addition to the widely-used GSM8K
1261
+ (Cobbe et al.,
1262
+ 2021
1263
+ )
1264
+ , we include challenging benchmarks from multiple domains:
1265
+ (i)
1266
+ competition and Olympiad-level benchmarks, such as MATH-500
1267
+ (Lightman et al.,
1268
+ 2023
1269
+ )
1270
+ , AIME 2024
1271
+ (AI-MO,
1272
+ 2024a
1273
+ )
1274
+ , AMC 2023
1275
+ (AI-MO,
1276
+ 2024b
1277
+ )
1278
+ and Olympiad Bench
1279
+ (He et al.,
1280
+ 2024
1281
+ )
1282
+ . Specifically, AIME is the exams designed to challenge the brightest high school math students in American, with the 2024 dataset comprising 30 problems from AIME I and II exams;
1283
+ (ii)
1284
+ college-level math problems from College Math
1285
+ (Tang et al.,
1286
+ 2024
1287
+ )
1288
+ and
1289
+ (iii)
1290
+ out-of-domain math benchmark: GaoKao (Chinese
1291
+ College Entrance Exam) En 2023
1292
+ (Liao et al.,
1293
+ 2024
1294
+ )
1295
+ .
1296
+ Base Models and Setup
1297
+ .
1298
+ \sysname
1299
+ is a general approach applicable to various LLMs. To show its effectiveness and generalizability, we use SLMs of different sizes as the base policy models:
1300
+ Qwen2.5-Math-1.5B
1301
+ (Qwen,
1302
+ 2024b
1303
+ )
1304
+ , Phi3-mini-Instruct (3B)
1305
+ (Microsoft,
1306
+ 2024
1307
+ ; Abdin et al.,
1308
+ 2024
1309
+ )
1310
+ , Qwen2-Math-7B
1311
+ (Qwen,
1312
+ 2024a
1313
+ )
1314
+ and Qwen2.5-Math-7B
1315
+ (Qwen,
1316
+ 2024c
1317
+ )
1318
+ . Among these, Phi3-mini-Instruct is a general-purpose SLM without specialization in math reasoning.
1319
+ Due to limited GPU resources, we performed 4 rounds of self-evolution exclusively on Qwen2.5-Math-7B, yielding 4 evolved policy SLMs (Table
1320
+ 3
1321
+ ) and 4 PPMs (Table
1322
+ 4
1323
+ ). For the other 3 policy LLMs, we fine-tune them using step-by-step verified trajectories generated from Qwen2.5-Math-7B’s 4th round. The final PPM from this round is then used as the reward model for the 3 policy SLMs.
1324
+ Baselines
1325
+ .
1326
+ \sysname
1327
+ is a System 2 method. We compare it against three strong baselines representing both System 1 and System 2 approaches:
1328
+ (i)
1329
+ Frontier LLMs
1330
+ , including GPT-4o, the latest Claude, OpenAI o1-preview and o1-mini.
1331
+ We measure their accuracy on AMC 2023, Olympiad Bench, College Math, Gaokao and GSM8K, with accuracy numbers for other benchmarks are taken from public technical reports
1332
+ (Team,
1333
+ 2024a
1334
+ )
1335
+ .
1336
+ (ii)
1337
+ Open-sourced superior reasoning models
1338
+ , including DeepSeek-Coder-v2-Instruct, Mathstral
1339
+ (Team,
1340
+ 2024b
1341
+ )
1342
+ , NuminaMath-72B
1343
+ (Jia LI and Polu,
1344
+ 2024a
1345
+ )
1346
+ , and LLaMA3.1
1347
+ (Dubey et al.,
1348
+ 2024
1349
+ )
1350
+ , which represent the current mainstream System 1 approaches for improving LLM math reasoning.
1351
+ (iii)
1352
+ Both System 1 and System 2 performance of the base models trained from the original models teams
1353
+ , including Instruct versions (e.g., Qwen2.5-Math-7B-Instruct) and Best-of-N (e.g., Qwen2.5-Math-72B-Instruct+Qwen2.5-Math-RM-72B). Notably, the reward model used for the three Qwen base models is a 72B ORM, significantly larger than our 7B PPM.
1354
+ Evaluation Metric
1355
+ . We report Pass@1 accuracy for all baselines. For System 2 baselines, we use default evaluation settings, such as default thinking time for o1-mini and o1-preview. For Qwen models with Best-of-N, we re-evaluate MATH-500, AIME/AMC accuracy; other benchmarks results are from their technical reports. For a fair comparison,
1356
+ \sysname
1357
+ run MCTS to generate the same number of solutions as Qwen. Specifically, for AIME/AMC, we generate 16 trajectories for AIME/AMC and 8 for other benchmarks, using PPM to select the best solution. We also report performance with increased test-time computation using 64 trajectories, denoted as
1358
+ \sysname
1359
+ 64
1360
+ .
1361
+ Table 5:
1362
+ The results of
1363
+ \sysname
1364
+ and other frontier LLMs on the most challenging math benchmarks.
1365
+ \sysname
1366
+ 64
1367
+ shows the Pass@1 accuracy achieved when sampling 64 trajectories.
1368
+ Competition and College Level
1369
+ OOD
1370
+ Model
1371
+ Method
1372
+ MATH
1373
+ AIME
1374
+ 2024
1375
+ AMC
1376
+ 2023
1377
+ Olympiad
1378
+ Bench
1379
+ College
1380
+ Math
1381
+ GSM8K
1382
+ Gaokao
1383
+ En 2023
1384
+ Frontier LLMs
1385
+ GPT-4o
1386
+ System 1
1387
+ 76.6
1388
+ 9.3
1389
+ 47.5
1390
+ 43.3
1391
+ 48.5
1392
+ 92.9
1393
+ 67.5
1394
+ Claude3.5-Sonnet
1395
+ System 1
1396
+ 78.3
1397
+ 16.0
1398
+ -
1399
+ -
1400
+ -
1401
+ 96.4
1402
+ -
1403
+ GPT-o1-preview
1404
+ -
1405
+ 85.5
1406
+ 44.6
1407
+ 90.0
1408
+ -
1409
+ -
1410
+ -
1411
+ -
1412
+ GPT-o1-mini
1413
+ -
1414
+ 90.0
1415
+ 56.7
1416
+ 95.0
1417
+ 65.3
1418
+ 57.8
1419
+ 94.8
1420
+ 78.4
1421
+ Open-Sourced Reasoning LLMs
1422
+ DeepSeek-Coder-V2-Instruct
1423
+ System 1
1424
+ 75.3
1425
+ 13.3
1426
+ 57.5
1427
+ 37.6
1428
+ 46.2
1429
+ 94.9
1430
+ 64.7
1431
+ Mathstral-7B-v0.1
1432
+ System 1
1433
+ 57.8
1434
+ 0.0
1435
+ 37.5
1436
+ 21.5
1437
+ 33.7
1438
+ 84.9
1439
+ 46.0
1440
+ NuminaMath-72B-CoT
1441
+ System 1
1442
+ 64.0
1443
+ 3.3
1444
+ 70.0
1445
+ 32.6
1446
+ 39.7
1447
+ 90.8
1448
+ 58.4
1449
+ LLaMA3.1-8B-Instruct
1450
+ System 1
1451
+ 51.4
1452
+ 6.7
1453
+ 25.0
1454
+ 15.4
1455
+ 33.8
1456
+ 76.6
1457
+ 38.4
1458
+ LLaMA3.1-70B-Instruct
1459
+ System 1
1460
+ 65.4
1461
+ 23.3
1462
+ 50.0
1463
+ 27.7
1464
+ 42.5
1465
+ 94.1
1466
+ 54.0
1467
+ Qwen2.5-Math-72B-Instruct
1468
+ System 1
1469
+ 85.6
1470
+ 30.0
1471
+ 70.0
1472
+ 49.0
1473
+ 49.5
1474
+ 95.9
1475
+ 71.9
1476
+ Qwen2.5-Math-72B-Instruct+72B ORM
1477
+ System 2
1478
+ 85.8
1479
+ 36.7
1480
+ 72.5
1481
+ 54.5
1482
+ 50.6
1483
+ 96.4
1484
+ 76.9
1485
+ General Base Model: Phi3-mini-Instruct (3.8B)
1486
+ Phi3-mini-Instruct (base model)
1487
+ System 1
1488
+ 41.4
1489
+ 3.33
1490
+ 7.5
1491
+ 12.3
1492
+ 33.1
1493
+ 85.7
1494
+ 37.1
1495
+ \sysname
1496
+ (3.8B SLM+7B PPM)
1497
+ System 2
1498
+ 85.4
1499
+ 40.0
1500
+ 77.5
1501
+ 59.3
1502
+ 58.0
1503
+ 94.5
1504
+ 77.1
1505
+ \sysname
1506
+ 64
1507
+ (3.8B SLM+7B PPM)
1508
+ System 2
1509
+ 86.4
1510
+ 43.3
1511
+ 80.0
1512
+ 60.3
1513
+ 59.1
1514
+ 94.7
1515
+ 77.7
1516
+ Math-Specialized Base Model: Qwen2.5-Math-1.5B
1517
+ Qwen2.5-Math-1.5B (base model)
1518
+ System 1
1519
+ 51.2
1520
+ 0.0
1521
+ 22.5
1522
+ 16.7
1523
+ 38.4
1524
+ 74.6
1525
+ 46.5
1526
+ Qwen2.5-Math-1.5B-Instruct
1527
+ System 1
1528
+ 60.0
1529
+ 10.0
1530
+ 60.0
1531
+ 38.1
1532
+ 47.7
1533
+ 84.8
1534
+ 65.5
1535
+ Qwen2.5-Math-1.5B-Instruct+72B ORM
1536
+ System 2
1537
+ 83.4
1538
+ 20.0
1539
+ 72.5
1540
+ 47.3
1541
+ 50.2
1542
+ 94.1
1543
+ 73.0
1544
+ \sysname
1545
+ (1.5B SLM+7B PPM)
1546
+ System 2
1547
+ 87.8
1548
+ 46.7
1549
+ 80.0
1550
+ 63.5
1551
+ 59.0
1552
+ 94.3
1553
+ 77.7
1554
+ \sysname
1555
+ 64
1556
+ (1.5B SLM+7B PPM)
1557
+ System 2
1558
+ 88.6
1559
+ 46.7
1560
+ 85.0
1561
+ 64.6
1562
+ 59.3
1563
+ 94.8
1564
+ 79.5
1565
+ Math-Specialized Base Model: Qwen2-Math-7B
1566
+ Qwen2-Math-7B (base model)
1567
+ System 1
1568
+ 53.4
1569
+ 3.3
1570
+ 25.0
1571
+ 17.3
1572
+ 39.4
1573
+ 80.4
1574
+ 47.3
1575
+ Qwen2-Math-7B-Instruct
1576
+ System 1
1577
+ 73.2
1578
+ 13.3
1579
+ 62.5
1580
+ 38.2
1581
+ 45.9
1582
+ 89.9
1583
+ 62.1
1584
+ Qwen2-Math-7B-Instruct+72B ORM
1585
+ System 2
1586
+ 83.4
1587
+ 23.3
1588
+ 62.5
1589
+ 47.6
1590
+ 47.9
1591
+ 95.1
1592
+ 71.9
1593
+ \sysname
1594
+ (7B SLM+7B PPM)
1595
+ System 2
1596
+ 88.2
1597
+ 43.3
1598
+ 80.0
1599
+ 63.1
1600
+ 58.4
1601
+ 94.6
1602
+ 78.2
1603
+ \sysname
1604
+ 64
1605
+ (7B SLM+7B PPM)
1606
+ System 2
1607
+ 88.6
1608
+ 46.7
1609
+ 85.0
1610
+ 63.4
1611
+ 59.3
1612
+ 94.8
1613
+ 79.2
1614
+ Math-Specialized Base Model: Qwen2.5-Math-7B
1615
+ Qwen2.5-Math-7B (base model)
1616
+ System 1
1617
+ 58.8
1618
+ 0.0
1619
+ 22.5
1620
+ 21.8
1621
+ 41.6
1622
+ 91.6
1623
+ 51.7
1624
+ Qwen2.5-Math-7B-Instruct
1625
+ System 1
1626
+ 82.6
1627
+ 6.0
1628
+ 62.5
1629
+ 41.6
1630
+ 46.8
1631
+ 95.2
1632
+ 66.8
1633
+ Qwen2.5-Math-7B-Instruct+72B ORM
1634
+ System 2
1635
+ 88.4
1636
+ 26.7
1637
+ 75.0
1638
+ 49.9
1639
+ 49.6
1640
+ 97.9
1641
+ 75.1
1642
+ \sysname
1643
+ (7B SLM+7B PPM)
1644
+ System 2
1645
+ 89.4
1646
+ 50.0
1647
+ 87.5
1648
+ 65.3
1649
+ 59.0
1650
+ 95.0
1651
+ 80.5
1652
+ \sysname
1653
+ 64
1654
+ (7B SLM+7B PPM)
1655
+ System 2
1656
+ 90.0
1657
+ 53.3
1658
+ 87.5
1659
+ 65.6
1660
+ 60.5
1661
+ 95.2
1662
+ 81.3
1663
+ 4.2
1664
+ Main Results
1665
+ Results on diverse challenging math benchmarks
1666
+ . Table
1667
+ 5
1668
+ shows the results of
1669
+ \sysname
1670
+ with comparing to state-of-the-art reasoning models. We highlight three key observations:
1671
+ (1)
1672
+ \sysname
1673
+ significantly improves SLMs math reasoning capabilities, achieving performance comparable to or surpassing OpenAI o1 with substantially smaller model size (1.5B-7B). For example, Qwen2.5-Math-7B, originally at 58.8% accuracy on MATH, improved dramatically to 90.0% with
1674
+ \sysname
1675
+ , outperforming o1-preview and Claude 3.5 Sonnet while matching o1-mini. On the College Math benchmark,
1676
+ \sysname
1677
+ exceeds o1-mini by 2.7%. On AIME 2024,
1678
+ \sysname
1679
+ scored 53.3%, ranking just below o1-mini, with the 7B model solving 8/15 problems in both AIME I and II, placing in the top 20% of the brightest high school math students.
1680
+ Notably, 8 of the unsolved problems were geometry-based, requiring visual understanding, a capability
1681
+ \sysname
1682
+ currently does not support.
1683
+ (2)
1684
+ Despite using smaller policy models (1.5B-7B) and reward models (7B),
1685
+ \sysname
1686
+ significantly outperforms state-of-the-art System 2 baselines. Compared to Qwen Best-of-N baselines, which use the same base models (Qwen2-Math-7B, Qwen2.5-Math-1.5B/7B) but a 10
1687
+ ×
1688
+ \times
1689
+ larger reward model (Qwen2.5-Math-RM-72B),
1690
+ \sysname
1691
+ consistently improves the reasoning accuracy of all base models to state-of-the-art levels. Even against Best-of-N with a 10
1692
+ ×
1693
+ \times
1694
+ larger Qwen2.5-Math-72B-Instruct policy model,
1695
+ \sysname
1696
+ surpasses it on all benchmarks except GSM8K, using the same number of sampled solutions.
1697
+ (3)
1698
+ Beyond well-known benchmarks like MATH, GSM8K, and AIME, which may risk over-optimization,
1699
+ \sysname
1700
+ shows strong generalizability on other challenging math benchmarks, including Olympiad Bench, College Math, and the Chinese College Entrance Math Exam (Gaokao), setting new state-of-the-art scores. As discussed in Sec.
1701
+ 3.4
1702
+ , our training set is primarily sourced from public datasets, with no specific optimizations for these benchmarks.
1703
+ Figure 3:
1704
+ Reasoning performance under scaling up the test-time compute.
1705
+ Scaling up test-time computation
1706
+ .
1707
+ \sysname
1708
+ uses MCTS to augment the policy model, searching solutions guided by the PPM. By increasing test-time computation, it explores more trajectories, potentially improving performance.
1709
+ In Fig.
1710
+ 3
1711
+ , we show the impact of test-time compute scaling by comparing the accuracy of the official Qwen Best-of-N across different numbers of sampled trajectories on four challenging math benchmarks. Sampling only one trajectory corresponds to the policy LLM’s Pass@1 accuracy, indicating a fallback to System 1 reasoning. We highlight two key observations:
1712
+ (1)
1713
+ With only 4 trajectories,
1714
+ \sysname
1715
+ significantly outperforms Best-of-N baselines, exceeding o1-preview and approaching o1-mini, demonstrating its effectiveness.
1716
+ (2)
1717
+ Scaling test-time compute improves reasoning accuracy across all benchmarks, though with varying trends. On Math, AIME, and Olympiad Bench,
1718
+ \sysname
1719
+ shows saturation or slow improvement at 64 trajectories, while on College Math, performance continues to improve steadily.
1720
+ 4.3
1721
+ Ablation Study and Analysis
1722
+ We ablate the effectiveness of our three innovations. For System 2-style inference, Pass@1 accuracy is measured with 16 trajectories for AIME and AMC, and 8 for other benchmarks.
1723
+ Table 6:
1724
+ The continuously improved math reasoning capabilities through
1725
+ \sysname
1726
+ self-evolved deep thinking. Starting from round 2, the 7B base model powered by
1727
+ \sysname
1728
+ surpasses GPT-4o.
1729
+ Round#
1730
+ MATH
1731
+ AIME 2024
1732
+ AMC 2023
1733
+ Olympiad Bench
1734
+ College Math
1735
+ GSM8K
1736
+ GaokaoEn 2023
1737
+ GPT-4o
1738
+ 76.6
1739
+ 9.3
1740
+ 47.5
1741
+ 43.3
1742
+ 48.5
1743
+ 92.9
1744
+ 67.5
1745
+ Base 7B model
1746
+ 58.8
1747
+ 0.0
1748
+ 22.5
1749
+ 21.8
1750
+ 41.6
1751
+ 91.6
1752
+ 51.7
1753
+ \sysname
1754
+ Round 1
1755
+ 75.2
1756
+ 10.0
1757
+ 57.5
1758
+ 35.7
1759
+ 45.4
1760
+ 90.9
1761
+ 60.3
1762
+ \sysname
1763
+ Round 2
1764
+ 86.6
1765
+ 43.3
1766
+ 75.0
1767
+ 59.4
1768
+ 55.6
1769
+ 94.0
1770
+ 76.4
1771
+ \sysname
1772
+ Round 3
1773
+ 87.0
1774
+ 46.7
1775
+ 80.0
1776
+ 61.6
1777
+ 56.5
1778
+ 94.2
1779
+ 77.1
1780
+ \sysname
1781
+ Round 4
1782
+ 89.4
1783
+ 50.0
1784
+ 87.5
1785
+ 65.3
1786
+ 59.0
1787
+ 95.0
1788
+ 80.5
1789
+ The effectiveness of self-evolution
1790
+ . The impressive results in Table
1791
+ 5
1792
+ are achieved after 4 rounds of
1793
+ \sysname
1794
+ self-evolved deep thinking. Table
1795
+ 6
1796
+ shows the math reasoning performance in each round, demonstrating a continuous improvement in accuracy.
1797
+ In round 1, the main improvement comes from applying SFT to the base model. Round 2 brings a significant boost with the application of a stronger PPM in MCTS, which unlocks the full potential of System 2 deep reasoning. Notably, starting from round 2,
1798
+ \sysname
1799
+ outperforms GPT-4o. Rounds 3 and 4 show further improvements, driven by stronger System 2 reasoning through better policy SLMs and PPMs.
1800
+ The effectiveness of step-by-step verified reasoning trajectory
1801
+ .
1802
+ \sysname
1803
+ generates step-by-step verified reasoning trajectories, which eliminate error intermediate steps and further expand training set with more challenging problems. To evaluate its effectiveness, we use the data generated from round 4 as SFT training data and compare it against
1804
+ three strong baselines:
1805
+ (i)
1806
+ GPT-distillation, which includes open-sourced CoT solutions synthesized using GPT-4, such as MetaMath
1807
+ (Yu et al.,
1808
+ 2023b
1809
+ )
1810
+ , NuminaMath-CoT
1811
+ (Jia LI and Polu,
1812
+ 2024b
1813
+ )
1814
+ ;
1815
+ (ii)
1816
+ Random sampling from self-generation,
1817
+ which use the same policy model (i.e., policy SLM-r3) to randomly generate trajectories;
1818
+ (iii)
1819
+ Rejection sampling, where 32 trajectories are randomly sampled from the policy model, with high-quality solutions ranked by our trained ORM (appendix
1820
+ A.1
1821
+ ). For fairness, we select two correct trajectories for each math problem in baseline (ii) and (iii). All SFT experiments use the same training recipe.
1822
+ Table 7:
1823
+ Ablation study on the effectiveness of our step-by-step verified reasoning trajectories as the SFT dataset. We report the SFT accuracy of Qwen2.5-Math-7B fine-tuned with different datasets.
1824
+ Dataset
1825
+ MATH
1826
+ AIME
1827
+ AMC
1828
+ Olympiad Bench
1829
+ College Math
1830
+ GSM8K
1831
+ GaokaoEn 2023
1832
+ GPT-4o
1833
+ -
1834
+ 76.6
1835
+ 9.3
1836
+ 47.5
1837
+ 43.3
1838
+ 48.5
1839
+ 92.9
1840
+ 67.5
1841
+ GPT4-distillation
1842
+ (Open-sourced)
1843
+ MetaMath
1844
+ 55.2
1845
+ 3.33
1846
+ 32.5
1847
+ 19.1
1848
+ 39.2
1849
+ 85.1
1850
+ 43.6
1851
+ NuminaMath-CoT
1852
+ 69.6
1853
+ 10.0
1854
+ 50.0
1855
+ 37.2
1856
+ 43.4
1857
+ 89.8
1858
+ 59.5
1859
+ Self-generation
1860
+ by policy SLM-r3
1861
+ Random sample
1862
+ 72.4
1863
+ 10.0
1864
+ 45.0
1865
+ 41.0
1866
+ 48.0
1867
+ 87.5
1868
+ 57.1
1869
+ Rejection sampling
1870
+ 73.4
1871
+ 13.3
1872
+ 47.5
1873
+ 44.7
1874
+ 50.8
1875
+ 89.3
1876
+ 61.7
1877
+ Step-by-step verified (ours)
1878
+ 78.4
1879
+ 26.7
1880
+ 47.5
1881
+ 47.1
1882
+ 52.5
1883
+ 89.7
1884
+ 65.7
1885
+ Table
1886
+ 7
1887
+ shows the math reasoning accuracy of Qwen2.5-Math-7B fine-tuned on different datasets. We highlight two observations:
1888
+ (i)
1889
+ Fine-tuning with our step-by-step verified trajectories significantly outperforms all other baselines. This is primarily due to our PPM-augmented MCTS for code-augmented CoT synthesis, which provides denser verification during math solution generation. It proves more effective than both random sampling, which lacks verification, and rejection sampling, where ORM provides only sparse verification.
1890
+ (ii)
1891
+ Even randomly sampled code-augmented CoT solutions from our SLM yields comparable or better performance than GPT-4 synthesized NuminaMath and MetaMath datasets.
1892
+ This indicates that our policy SLMs, after rounds of self-evolution, can generate high-quality math solutions. These results demonstrates the huge potential of our method to self-generate higher-quality reasoning data without relying on advanced LLM distillation.
1893
+ The effectiveness of PPM
1894
+ . We train both a strong ORM and Q-value score-based PRM (PQM) for comparison. To ensure a fair evaluation, we use the highest-quality training data: the step-by-step verified trajectories generated in round 4, with selected math problems matching those used for PPM training. Similar to PPM, we use step-level Q-values as to select positive and negative trajectories for each math problem.
1895
+ The ORM is trained using a pairwise ranking loss
1896
+ (Ouyang et al.,
1897
+ 2022
1898
+ )
1899
+ , while the PQM follows
1900
+ (Chen et al.,
1901
+ 2024
1902
+ ; Zhang et al.,
1903
+ 2024a
1904
+ )
1905
+ to use Q-values as reward labels and optimize with MSE loss. Detailed training settings are provided in Appendix
1906
+ A.1
1907
+ .
1908
+ Table 8:
1909
+ Ablation study on the reward model. Process reward models (PQM and PPM) outperform ORM, with PPM pushing the frontier of math reasoning capabilities.
1910
+ RM
1911
+ Inference
1912
+ MATH
1913
+ AIME
1914
+ AMC
1915
+ Olympiad Bench
1916
+ College Math
1917
+ GSM8K
1918
+ GaokaoEn
1919
+ o1-mini
1920
+ -
1921
+ 90.0
1922
+ 56.7
1923
+ 95.0
1924
+ 65.3
1925
+ 55.6
1926
+ 94.8
1927
+ 78.6
1928
+ ORM
1929
+ Best-of-N
1930
+ 82.6
1931
+ 26.7
1932
+ 65.0
1933
+ 55.1
1934
+ 55.5
1935
+ 92.3
1936
+ 72.5
1937
+ PQM
1938
+ MCTS
1939
+ 88.2
1940
+ 46.7
1941
+ 85.0
1942
+ 62.9
1943
+ 57.6
1944
+ 94.6
1945
+ 79.5
1946
+ PPM
1947
+ MCTS
1948
+ 89.4
1949
+ 50.0
1950
+ 87.5
1951
+ 65.3
1952
+ 59.0
1953
+ 95.0
1954
+ 80.5
1955
+ Table
1956
+ 8
1957
+ compares the performance of ORM, PQM, and PPM for System 2 reasoning using our final round policy model. ORM provides reward signals only at the end of problem solving, so we use the Best-of-N method, while PRM and PPM leverage MCTS-driven search. As shown in Table
1958
+ 8
1959
+ , both PQM and PPM outperform ORM by providing denser step-level reward signals, leading to higher accuracy on complex math reasoning tasks. However, PQM struggles on more challenging benchmarks, such as MATH and Olympiad Bench, due to the inherent imprecision of Q-values.
1960
+ In contrast, PPM constructs step-level preference data for training, enabling our 7B policy model to achieve comparable or superior performance to o1-mini across all benchmarks.
1961
+ 5
1962
+ Findings and Discussions
1963
+ Figure 4:
1964
+ An example of intrinsic self-reflection during
1965
+ \sysname
1966
+ deep thinking.
1967
+ The emergence of intrinsic self-reflection capability
1968
+ . A key breakthrough in OpenAI o1 is its intrinsic self-reflection capability. When the model makes an error, it recognizes the mistake and can self-correct with a correct answer
1969
+ (Noam Brown and Lightman,
1970
+ 2024
1971
+ )
1972
+ . Yet it has consistently
1973
+ been found to be largely ineffective in open-sourced LLMs. The community has actively explored various approaches, including self-correction
1974
+ (Huang et al.,
1975
+ 2023
1976
+ ; Kumar et al.,
1977
+ 2024
1978
+ )
1979
+ , self-reflection
1980
+ (Renze and Guven,
1981
+ 2024
1982
+ ; Shinn et al.,
1983
+ 2024
1984
+ )
1985
+ , to explicitly train or prompt LLMs to develop such capability.
1986
+ In our experiments, we unexpectedly observe that our MCTS-driven deep thinking exhibits self-reflection during problem-solving. As shown in Fig.
1987
+ 4
1988
+ , the model initially formalizes an equation using
1989
+ SymPy
1990
+ in the first three steps, which would lead to an incorrect answer (left branch). Interestingly, in the fourth step (right branch), the policy model recognizes the low quality of its earlier steps and refrains from continuing along the initial problem-solving path. Instead, it backtracks and resolves the problem using a new, simpler approach, ultimately arriving at the correct answer. An additional example of self-correction is provided in Appendix
1991
+ A.2
1992
+ . Notably, no self-reflection training data or prompt was included, suggesting that advanced System 2 reasoning can foster intrinsic self-reflection.
1993
+ Figure 5:
1994
+ Pass@1 accuracy of policy models and their accuracy after applying System 2 reasoning with various reward models, shows that reward models primarily determine the final performance.
1995
+ PPM shapes the reasoning boundary in System 2 deep thinking
1996
+ . Both the policy and reward models are crucial for System 2 deep reasoning. Our experiments show that once the policy model attains a reasonably strong capability level,
1997
+ (see Appendix
1998
+ A.1
1999
+ ), the PPM becomes the key determinant of the upper performance limit.
2000
+ Fig.
2001
+ 5
2002
+ summarizes the accuracy of policy models of different sizes, as well as the improvements achieved with reward models. Despite variations in Pass@1 accuracy due to differences in training strategies, datasets, and model scales, the reward model proves to be the dominant factor in System 2 reasoning. For instance, although the SFT accuracy of
2003
+ \sysname
2004
+ -7B is lower than Qwen2.5-Math-72B-Instruct, pairing it with our 7B PPM allows
2005
+ \sysname
2006
+ to outperform the 72B policy model with Qwen 72B ORM. Moreover, despite varying Pass@1 accuracy across our three policy SLM sizes, the final reasoning accuracy converges after applying the PPM.
2007
+ PPM spots theorem-application steps
2008
+ . When solving challenging math problems, identifying and applying relevant theorems or key conclusions often form the cornerstone of successful problem-solving
2009
+ (Xin et al.,
2010
+ 2024
2011
+ )
2012
+ . In our experiments, we find that during
2013
+ \sysname
2014
+ problem-solving, our PPM effectively identifies critical theorem-application intermediate steps within policy model’s deep thinking process. These steps are predicted with high reward scores, guiding the policy model to generate the correct solution. Appendix
2015
+ A.2
2016
+ provides examples where the PPM successfully identifies key theorems such as Fermat’s little theorem
2017
+ (Weisstein,
2018
+ a
2019
+ )
2020
+ , Vieta’s formulas
2021
+ (Weisstein,
2022
+ b
2023
+ )
2024
+ , the AM-GM inequality
2025
+ (
2026
+ amg,
2027
+ )
2028
+ , the Pythagorean theorem
2029
+ (
2030
+ pyt,
2031
+ )
2032
+ , and the Shoelace Theorem
2033
+ (
2034
+ sho,
2035
+ )
2036
+ , etc.
2037
+ Generalization discussions
2038
+ .
2039
+ \sysname
2040
+ offers a general methodology for improving LLM reasoning applicable to various domains. First,
2041
+ \sysname
2042
+ can generalize to more challenging math tasks, such as theorem proving, though its current focus is on word problems due to dataset limitations. Nonetheless,
2043
+ \sysname
2044
+ demonstrates the potential to prove mathematical statements. As shown in Appendix
2045
+ A.2
2046
+ , it successfully proves an Olympiad-level problem involving Fermat’s Little Theorem, providing a step-by-step correct proof through its deep reasoning process. Second,
2047
+ \sysname
2048
+ can generalize to other domains, such as code and commonsense reasoning. Notably, synthesizing step-by-step verified training trajectories for general reasoning requires a mechanism to provide feedback on whether a given trajectory reaches the desired output at the end of MCTS rollout. For instance, in code reasoning, this could involve designing extensive test cases; in general reasoning, feedback could be obtained through human labeling or mutual verification with another LLM
2049
+ (Qi et al.,
2050
+ 2024
2051
+ )
2052
+ .
2053
+ 6
2054
+ Conclusion
2055
+ In this work, we present
2056
+ \sysname
2057
+ , a self-evolved System 2 deep thinking approach that significantly boosts the math reasoning capabilities of small LLMs, achieving state-of-the-art OpenAI o1-level performance. Our approach demonstrates that SLMs can self-generate high-quality training data for frontier-level math reasoning. Extensive experiments across four different-sized SLMs and challenging math benchmarks demonstrate the superiority of
2058
+ \sysname
2059
+ , with achieving leading results while outperforming existing math reasoning LLMs and Best-of-N baselines. We also reveal key findings, including the emergence of self-reflection and the effectiveness of the PPM in identifying critical intermediate steps, such as theorem-application steps. Finally,
2060
+ \sysname
2061
+ can achieve further improvements by collecting more challenging math problems, we leave this as future work.
2062
+ Acknowledgement
2063
+ In the early stages of this work, we faced significant challenges due to limited GPU resources and restricted access to the GPT-4 API. We are deeply grateful to Qiufeng Yin and Chengmin Chi for their assistance in collecting math problems and providing GPT-4 resources for new math problem synthesis. Special thanks go to my colleagues, Lingxiao Ma, Ying Cao, Baotong Lu, Jing Liu, Jiahang Xu, Chengruidong Zhang, Siyuan Wang, Gaokai Zhang, Yujian Li, and Yang Wang, for generously sharing their GPU quotas with us.
2064
+ References
2065
+ [1]
2066
+ Inequality of arithmetic and geometric means.
2067
+ URL
2068
+ https://artofproblemsolving.com/wiki/index.php/AM-GM_Inequality
2069
+ .
2070
+ [2]
2071
+ Pythagorean theorem.
2072
+ URL
2073
+ https://en.wikipedia.org/wiki/Pythagorean_theorem
2074
+ .
2075
+ [3]
2076
+ Shoelace theorem.
2077
+ URL
2078
+ https://artofproblemsolving.com/wiki/index.php/Shoelace_Theorem
2079
+ .
2080
+ Abdin et al. [2024]
2081
+ Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah,
2082
+ Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl,
2083
+ et al.
2084
+ Phi-3 technical report: A highly capable language model locally on
2085
+ your phone.
2086
+ arXiv preprint arXiv:2404.14219
2087
+ , 2024.
2088
+ AI-MO [2024a]
2089
+ AI-MO.
2090
+ Aime 2024, 2024a.
2091
+ URL
2092
+ https://huggingface.co/datasets/AI-MO/aimo-validation-aime
2093
+ .
2094
+ AI-MO [2024b]
2095
+ AI-MO.
2096
+ Amc 2023, 2024b.
2097
+ URL
2098
+ https://huggingface.co/datasets/AI-MO/aimo-validation-amc
2099
+ .
2100
+ Brown et al. [2024]
2101
+ Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le,
2102
+ Christopher Ré, and Azalia Mirhoseini.
2103
+ Large language monkeys: Scaling inference compute with repeated
2104
+ sampling.
2105
+ arXiv preprint arXiv:2407.21787
2106
+ , 2024.
2107
+ Chen et al. [2024]
2108
+ Guoxin Chen, Minpeng Liao, Chengxi Li, and Kai Fan.
2109
+ Alphamath almost zero: process supervision without process, 2024.
2110
+ Cobbe et al. [2021]
2111
+ Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz
2112
+ Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano,
2113
+ et al.
2114
+ Training verifiers to solve math word problems.
2115
+ arXiv preprint arXiv:2110.14168
2116
+ , 2021.
2117
+ Daniel [2011]
2118
+ Kahneman Daniel.
2119
+ Thinking, fast and slow.
2120
+ Macmillan
2121
+ , 2011.
2122
+ Dubey et al. [2024]
2123
+ Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad
2124
+ Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan,
2125
+ et al.
2126
+ The llama 3 herd of models.
2127
+ arXiv preprint arXiv:2407.21783
2128
+ , 2024.
2129
+ Gou et al. [2023]
2130
+ Zhibin Gou, Zhihong Shao, Yeyun Gong, Yujiu Yang, Minlie Huang, Nan Duan,
2131
+ Weizhu Chen, et al.
2132
+ Tora: A tool-integrated reasoning agent for mathematical problem
2133
+ solving.
2134
+ arXiv preprint arXiv:2309.17452
2135
+ , 2023.
2136
+ Hao et al. [2023]
2137
+ Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and
2138
+ Zhiting Hu.
2139
+ Reasoning with language model is planning with world model.
2140
+ arXiv preprint arXiv:2305.14992
2141
+ , 2023.
2142
+ He et al. [2024]
2143
+ Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen,
2144
+ Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, et al.
2145
+ Olympiadbench: A challenging benchmark for promoting agi with
2146
+ olympiad-level bilingual multimodal scientific problems.
2147
+ arXiv preprint arXiv:2402.14008
2148
+ , 2024.
2149
+ Huang et al. [2023]
2150
+ Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu,
2151
+ Xinying Song, and Denny Zhou.
2152
+ Large language models cannot self-correct reasoning yet.
2153
+ arXiv preprint arXiv:2310.01798
2154
+ , 2023.
2155
+ Huang et al. [2024]
2156
+ Zhen Huang, Haoyang Zou, Xuefeng Li, Yixiu Liu, Yuxiang Zheng, Ethan Chern,
2157
+ Shijie Xia, Yiwei Qin, Weizhe Yuan, and Pengfei Liu.
2158
+ O1 replication journey – part 2: Surpassing o1-preview through
2159
+ simple distillation big progress or bitter lesson?
2160
+ Github
2161
+ , 2024.
2162
+ URL
2163
+ https://github.com/GAIR-NLP/O1-Journey
2164
+ .
2165
+ Jia LI and Polu [2024a]
2166
+ Lewis Tunstall Ben Lipkin Roman Soletskyi Shengyi Costa Huang Kashif Rasul
2167
+ Longhui Yu Albert Jiang Ziju Shen Zihan Qin Bin Dong Li Zhou Yann Fleureau
2168
+ Guillaume Lample Jia LI, Edward Beeching and Stanislas Polu.
2169
+ Numinamath.
2170
+ [https://github.com/project-numina/aimo-progress-prize](https://github.com/project-numina/aimo-progress-prize/blob/main/report/numina_dataset.pdf)
2171
+ ,
2172
+ 2024a.
2173
+ Jia LI and Polu [2024b]
2174
+ Lewis Tunstall Ben Lipkin Roman Soletskyi Shengyi Costa Huang Kashif Rasul
2175
+ Longhui Yu Albert Jiang Ziju Shen Zihan Qin Bin Dong Li Zhou Yann Fleureau
2176
+ Guillaume Lample Jia LI, Edward Beeching and Stanislas Polu.
2177
+ Numinamath cot, 2024b.
2178
+ URL
2179
+ https://huggingface.co/datasets/AI-MO/NuminaMath-CoT
2180
+ .
2181
+ Kang et al. [2024]
2182
+ Jikun Kang, Xin Zhe Li, Xi Chen, Amirreza Kazemi, and Boxing Chen.
2183
+ Mindstar: Enhancing math reasoning in pre-trained llms at inference
2184
+ time.
2185
+ arXiv preprint arXiv:2405.16265
2186
+ , 2024.
2187
+ Kocsis and Szepesvári [2006]
2188
+ Levente Kocsis and Csaba Szepesvári.
2189
+ Bandit based monte-carlo planning.
2190
+ volume 2006, pages 282–293, 09 2006.
2191
+ ISBN 978-3-540-45375-8.
2192
+ doi:
2193
+ 10.1007/11871842_29
2194
+ .
2195
+ Kumar et al. [2024]
2196
+ Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi
2197
+ Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, et al.
2198
+ Training language models to self-correct via reinforcement learning.
2199
+ arXiv preprint arXiv:2409.12917
2200
+ , 2024.
2201
+ Lanham et al. [2023]
2202
+ Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison,
2203
+ Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion,
2204
+ et al.
2205
+ Measuring faithfulness in chain-of-thought reasoning.
2206
+ arXiv preprint arXiv:2307.13702
2207
+ , 2023.
2208
+ Li et al. [2024]
2209
+ Chen Li, Weiqi Wang, Jingcheng Hu, Yixuan Wei, Nanning Zheng, Han Hu, Zheng
2210
+ Zhang, and Houwen Peng.
2211
+ Common 7b language models already possess strong math capabilities.
2212
+ arXiv preprint arXiv:2403.04706
2213
+ , 2024.
2214
+ Liao et al. [2024]
2215
+ Minpeng Liao, Wei Luo, Chengxi Li, Jing Wu, and Kai Fan.
2216
+ Mario: Math reasoning with code interpreter output–a reproducible
2217
+ pipeline.
2218
+ arXiv preprint arXiv:2401.08190
2219
+ , 2024.
2220
+ Lightman et al. [2023]
2221
+ Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy
2222
+ Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe.
2223
+ Let’s verify step by step.
2224
+ arXiv preprint arXiv:2305.20050
2225
+ , 2023.
2226
+ Lightman et al. [2024]
2227
+ Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker,
2228
+ Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe.
2229
+ Let’s verify step by step.
2230
+ In
2231
+ The Twelfth International Conference on Learning
2232
+ Representations
2233
+ , 2024.
2234
+ URL
2235
+ https://openreview.net/forum?id=v8L0pN6EOi
2236
+ .
2237
+ Liu et al. [2024]
2238
+ Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang
2239
+ Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al.
2240
+ Deepseek-v3 technical report.
2241
+ arXiv preprint arXiv:2412.19437
2242
+ , 2024.
2243
+ Luo et al. [2023]
2244
+ Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo
2245
+ Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang.
2246
+ Wizardmath: Empowering mathematical reasoning for large language
2247
+ models via reinforced evol-instruct.
2248
+ arXiv preprint arXiv:2308.09583
2249
+ , 2023.
2250
+ Luo et al. [2024]
2251
+ Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Phatale, Harsh Lara, Yunxuan
2252
+ Li, Lei Shu, Yun Zhu, Lei Meng, Jiao Sun, et al.
2253
+ Improve mathematical reasoning in language models by automated
2254
+ process supervision.
2255
+ arXiv preprint arXiv:2406.06592
2256
+ , 2024.
2257
+ Microsoft [2024]
2258
+ Microsoft.
2259
+ Phi-3-mini-4k-instruct, 2024.
2260
+ URL
2261
+ https://huggingface.co/microsoft/Phi-3-mini-4k-instruct
2262
+ .
2263
+ Noam Brown and Lightman [2024]
2264
+ Ilge Akkaya Noam Brown and Hunter Lightman.
2265
+ Openai’s noam brown, ilge akkaya and hunter lightman on o1 and
2266
+ teaching llms to reason better, 2024.
2267
+ URL
2268
+ https://www.youtube.com/watch?v=jPluSXJpdrA
2269
+ .
2270
+ OpenAI [2023]
2271
+ OpenAI.
2272
+ Gpt-4 technical report.
2273
+ 2023.
2274
+ OpenAI [2024]
2275
+ OpenAI.
2276
+ Openai o1 system card.
2277
+ preprint
2278
+ , 2024.
2279
+ Ouyang et al. [2022]
2280
+ Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela
2281
+ Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al.
2282
+ Training language models to follow instructions with human feedback.
2283
+ Advances in Neural Information Processing Systems
2284
+ ,
2285
+ 35:27730–27744, 2022.
2286
+ Qi et al. [2024]
2287
+ Zhenting Qi, Mingyuan Ma, Jiahang Xu, Li Lyna Zhang, Fan Yang, and Mao Yang.
2288
+ Mutual reasoning makes smaller llms stronger problem-solvers.
2289
+ arXiv preprint arXiv:2408.06195
2290
+ , 2024.
2291
+ Qwen [2024a]
2292
+ Qwen.
2293
+ Qwen2-math-7b, 2024a.
2294
+ URL
2295
+ https://huggingface.co/Qwen/Qwen2-Math-7B
2296
+ .
2297
+ Qwen [2024b]
2298
+ Qwen.
2299
+ Qwen2.5-math-1.5b, 2024b.
2300
+ URL
2301
+ https://huggingface.co/Qwen/Qwen2.5-Math-1.5B
2302
+ .
2303
+ Qwen [2024c]
2304
+ Qwen.
2305
+ Qwen2.5-math-7b, 2024c.
2306
+ URL
2307
+ https://huggingface.co/Qwen/Qwen2.5-Math-7B
2308
+ .
2309
+ Renze and Guven [2024]
2310
+ Matthew Renze and Erhan Guven.
2311
+ Self-reflection in llm agents: Effects on problem-solving
2312
+ performance.
2313
+ arXiv preprint arXiv:2405.06682
2314
+ , 2024.
2315
+ Shinn et al. [2024]
2316
+ Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu
2317
+ Yao.
2318
+ Reflexion: Language agents with verbal reinforcement learning.
2319
+ Advances in Neural Information Processing Systems
2320
+ , 36, 2024.
2321
+ Silver et al. [2017]
2322
+ David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew
2323
+ Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore
2324
+ Graepel, et al.
2325
+ Mastering chess and shogi by self-play with a general reinforcement
2326
+ learning algorithm.
2327
+ arXiv preprint arXiv:1712.01815
2328
+ , 2017.
2329
+ Snell et al. [2024]
2330
+ Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar.
2331
+ Scaling llm test-time compute optimally can be more effective than
2332
+ scaling model parameters.
2333
+ arXiv preprint arXiv:2408.03314
2334
+ , 2024.
2335
+ Tang et al. [2024]
2336
+ Zhengyang Tang, Xingxing Zhang, Benyou Wan, and Furu Wei.
2337
+ Mathscale: Scaling instruction tuning for mathematical reasoning.
2338
+ arXiv preprint arXiv:2403.02884
2339
+ , 2024.
2340
+ Team [2024a]
2341
+ Qwen Team.
2342
+ Qwq: Reflect deeply on the boundaries of the unknown, November
2343
+ 2024a.
2344
+ URL
2345
+ https://qwenlm.github.io/blog/qwq-32b-preview/
2346
+ .
2347
+ Team [2024b]
2348
+ The Mistral AI Team.
2349
+ Mathstral-7b-v0.1, 2024b.
2350
+ URL
2351
+ https://huggingface.co/mistralai/Mathstral-7B-v0.1
2352
+ .
2353
+ Toshniwal et al. [2024]
2354
+ Shubham Toshniwal, Wei Du, Ivan Moshkov, Branislav Kisacanin, Alexan
2355
+ Ayrapetyan, and Igor Gitman.
2356
+ Openmathinstruct-2: Accelerating ai for math with massive open-source
2357
+ instruction data.
2358
+ arXiv preprint arXiv:2410.01560
2359
+ , 2024.
2360
+ Valmeekam et al. [2023]
2361
+ Karthik Valmeekam, Sarath Sreedharan, Matthew Marquez, Alberto Olmo, and
2362
+ Subbarao Kambhampati.
2363
+ On the planning abilities of large language models (a critical
2364
+ investigation with a proposed benchmark).
2365
+ arXiv preprint arXiv:2302.06706
2366
+ , 2023.
2367
+ Wang et al. [2024a]
2368
+ Chaojie Wang, Yanchen Deng, Zhiyi Lv, Shuicheng Yan, and An Bo.
2369
+ Q*: Improving multi-step reasoning for llms with deliberative
2370
+ planning, 2024a.
2371
+ Wang et al. [2024b]
2372
+ Ke Wang, Houxing Ren, Aojun Zhou, Zimu Lu, Sichun Luo, Weikang Shi, Renrui
2373
+ Zhang, Linqi Song, Mingjie Zhan, and Hongsheng Li.
2374
+ Mathcoder: Seamless code integration in LLMs for enhanced
2375
+ mathematical reasoning.
2376
+ In
2377
+ The Twelfth International Conference on Learning
2378
+ Representations
2379
+ , 2024b.
2380
+ URL
2381
+ https://openreview.net/forum?id=z8TW0ttBPp
2382
+ .
2383
+ Wang et al. [2024c]
2384
+ Peiyi Wang, Lei Li, Zhihong Shao, R. X. Xu, Damai Dai, Yifei Li, Deli Chen,
2385
+ Y. Wu, and Zhifang Sui.
2386
+ Math-shepherd: Verify and reinforce llms step-by-step without human
2387
+ annotations, 2024c.
2388
+ Wang et al. [2023]
2389
+ Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang,
2390
+ Aakanksha Chowdhery, and Denny Zhou.
2391
+ Self-consistency improves chain of thought reasoning in language
2392
+ models.
2393
+ In
2394
+ The Eleventh International Conference on Learning
2395
+ Representations
2396
+ , 2023.
2397
+ URL
2398
+ https://openreview.net/forum?id=1PL1NIMMrw
2399
+ .
2400
+ Weisstein [a]
2401
+ Eric W. Weisstein.
2402
+ Fermat’s little theorem, a.
2403
+ URL
2404
+ https://mathworld.wolfram.com/FermatsLittleTheorem.html
2405
+ .
2406
+ Weisstein [b]
2407
+ Eric W. Weisstein.
2408
+ Vieta’s formulas, from mathworld—a wolfram web resource,
2409
+ b.
2410
+ URL
2411
+ http://mathworld.wolfram.com/Tree.html
2412
+ .
2413
+ Wu et al. [2024]
2414
+ Yangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck, and Yiming Yang.
2415
+ An empirical analysis of compute-optimal inference for
2416
+ problem-solving with language models.
2417
+ arXiv preprint arXiv:2408.00724
2418
+ , 2024.
2419
+ Xin et al. [2024]
2420
+ Huajian Xin, Daya Guo, Zhihong Shao, Zhizhou Ren, Qihao Zhu, Bo Liu, Chong
2421
+ Ruan, Wenda Li, and Xiaodan Liang.
2422
+ Deepseek-prover: Advancing theorem proving in llms through
2423
+ large-scale synthetic data.
2424
+ arXiv preprint arXiv:2405.14333
2425
+ , 2024.
2426
+ Yang et al. [2024]
2427
+ An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li,
2428
+ Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, et al.
2429
+ Qwen2. 5-math technical report: Toward mathematical expert model via
2430
+ self-improvement.
2431
+ arXiv preprint arXiv:2409.12122
2432
+ , 2024.
2433
+ Yao et al. [2024]
2434
+ Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and
2435
+ Karthik Narasimhan.
2436
+ Tree of thoughts: Deliberate problem solving with large language
2437
+ models.
2438
+ Advances in Neural Information Processing Systems
2439
+ , 36, 2024.
2440
+ Yu et al. [2023a]
2441
+ Fei Yu, Anningzhe Gao, and Benyou Wang.
2442
+ Outcome-supervised verifiers for planning in mathematical reasoning.
2443
+ arXiv preprint arXiv:2311.09724
2444
+ , 2023a.
2445
+ Yu et al. [2023b]
2446
+ Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang,
2447
+ James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu.
2448
+ Metamath: Bootstrap your own mathematical questions for large
2449
+ language models.
2450
+ arXiv preprint arXiv:2309.12284
2451
+ , 2023b.
2452
+ Yuan et al. [2023]
2453
+ Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Keming Lu, Chuanqi Tan,
2454
+ Chang Zhou, and Jingren Zhou.
2455
+ Scaling relationship on learning mathematical reasoning with large
2456
+ language models.
2457
+ arXiv preprint arXiv:2308.01825
2458
+ , 2023.
2459
+ Zhang et al. [2024a]
2460
+ Dan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue, Yuxiao Dong, and Jie Tang.
2461
+ Rest-mcts*: Llm self-training via process reward guided tree search.
2462
+ arXiv preprint arXiv:2406.03816
2463
+ , 2024a.
2464
+ Zhang et al. [2024b]
2465
+ Di Zhang, Jiatong Li, Xiaoshui Huang, Dongzhan Zhou, Yuqiang Li, and Wanli
2466
+ Ouyang.
2467
+ Accessing gpt-4 level mathematical olympiad solutions via monte carlo
2468
+ tree self-refine with llama-3 8b.
2469
+ arXiv preprint arXiv:2406.07394
2470
+ , 2024b.
2471
+ Zheng et al. [2023]
2472
+ Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao
2473
+ Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E.
2474
+ Gonzalez, and Ion Stoica.
2475
+ Judging LLM-as-a-judge with MT-bench and chatbot arena.
2476
+ In
2477
+ Thirty-seventh Conference on Neural Information Processing
2478
+ Systems Datasets and Benchmarks Track
2479
+ , 2023.
2480
+ Appendix A
2481
+ Appendix
2482
+ A.1
2483
+ Additional Experiments and Details
2484
+ Data Generation Details
2485
+ . As detailed in Sec.
2486
+ 3.4
2487
+ , each round starts by self-generating step-by-step verified trajectories for 747k math word problems. The maximum tree depth
2488
+ d
2489
+ d
2490
+ is set to 16, with 16 MCTS rollouts conducted per problem by default. At each step, we allow to explore 8 candidate nodes, and the constant
2491
+ c
2492
+ c
2493
+ in Eq.
2494
+ 1
2495
+ is set to 2 to promote greater exploration. In the bootstrap round, due to the large size of the initial policy model (236B), we used smaller parameters: 8 rollouts and 5 candidate nodes per step. To improve the accuracy of solving challenging problems in round 4, we increase the number of candidate nodes to 16 and conduct 2 MCTS tree expansions per problem using different random seeds. Detailed prompts are available in Appendix
2496
+ A.3
2497
+ .
2498
+ Training Details
2499
+ . In each round, we collect step-by-step verified trajectories to fine-tune the policy LLM and train the PPM. To reduce noise
2500
+ in synthetic math problems (e.g., incorrect ground-truth answers labeled by GPT-4), we remove synthetic problems with trajectories achieving less than 50% accuracy. Based on our extensive experiments, the policy LLM is fine-tuned from the initial base model in each round, rather than training incrementally on the model from the previous round.
2501
+ All policy SLMs are trained for 2 epochs with a sequence length of 4096 tokens and a batch size of 128. We use AdamW optimizer with a linear learning rate scheduler, setting the initial learning rate to 7e-6 for Qwen models, and a cosine scheduler with an initial learning rate of 5e-6 for Phi3-mini-Instruct.
2502
+ The PPM is trained for 1 epoch with a batch size of 512 and an initial learning rate of 7e-6.
2503
+ Training the ORM and PQM
2504
+ . The Outcome Reward Model (ORM) and the Q-value-based Process Reward Model (PQM) share the same model architecture and training parameters with our PPM. To train the ORM, we collect trajectories from math problems containing both correct and incorrect solutions. Specifically, the two trajectories with the highest average Q-values are selected as positive examples, while the two with the lowest are chosen as negative examples. Following Qwen2.5-Math
2505
+ (Yang et al.,
2506
+ 2024
2507
+ )
2508
+ , we adopt the pairwise ranking loss
2509
+ (Ouyang et al.,
2510
+ 2022
2511
+ )
2512
+ to optimize the ORM. To train the PQM, we follow
2513
+ Chen et al. (
2514
+ 2024
2515
+ )
2516
+ to use step-level Q-values as reward labels. Let
2517
+ 𝐱
2518
+ =
2519
+ x
2520
+ ⊕
2521
+ s
2522
+ 1
2523
+ ⊕
2524
+ s
2525
+ 2
2526
+ ⊕
2527
+ …
2528
+ ⊕
2529
+ s
2530
+ d
2531
+ \mathbf{x}=x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{d}
2532
+ be the trajectory, with annotated Q-values
2533
+ 𝐐
2534
+ =
2535
+ (
2536
+ Q
2537
+ ​
2538
+ (
2539
+ s
2540
+ 1
2541
+ )
2542
+ ,
2543
+ Q
2544
+ ​
2545
+ (
2546
+ s
2547
+ 1
2548
+ )
2549
+ ,
2550
+ …
2551
+ ,
2552
+ Q
2553
+ ​
2554
+ (
2555
+ s
2556
+ d
2557
+ )
2558
+ )
2559
+ \mathbf{Q}=(Q(s_{1}),Q(s_{1}),...,Q(s_{d}))
2560
+ and predicted Q-values
2561
+ 𝐐
2562
+ ′
2563
+ =
2564
+ (
2565
+ Q
2566
+ ′
2567
+ ​
2568
+ (
2569
+ s
2570
+ 1
2571
+ )
2572
+ ,
2573
+ Q
2574
+ ′
2575
+ ​
2576
+ (
2577
+ s
2578
+ 1
2579
+ )
2580
+ ,
2581
+ …
2582
+ ,
2583
+ Q
2584
+ ′
2585
+ ​
2586
+ (
2587
+ s
2588
+ d
2589
+ )
2590
+ )
2591
+ \mathbf{Q^{\prime}}=(Q^{\prime}(s_{1}),Q^{\prime}(s_{1}),...,Q^{\prime}(s_{d}))
2592
+ for each step. To stabilize PQM training, we treat each trajectory as a single training sample and predict Q-values for all steps simultaneously, rather than splitting it into individual per-step samples. Specifically, to predict the Q-value
2593
+ Q
2594
+ ′
2595
+ ​
2596
+ (
2597
+ s
2598
+ i
2599
+ )
2600
+ Q^{\prime}(s_{i})
2601
+ for step
2602
+ s
2603
+ i
2604
+ s_{i}
2605
+ , PQM takes the trajectory from the question up to step
2606
+ s
2607
+ i
2608
+ s_{i}
2609
+ (i.e.,
2610
+ x
2611
+ ⊕
2612
+ s
2613
+ 1
2614
+ ⊕
2615
+ s
2616
+ 2
2617
+ ⊕
2618
+ …
2619
+ ⊕
2620
+ s
2621
+ i
2622
+ x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{i}
2623
+ ) as input and outputs a value between -1 and 1. We use a mean squared error (MSE) loss for PQM training:
2624
+ ℒ
2625
+ p
2626
+ ​
2627
+ r
2628
+ ​
2629
+ m
2630
+ ​
2631
+ (
2632
+ 𝐱
2633
+ )
2634
+ =
2635
+ ‖
2636
+ 𝐐
2637
+ −
2638
+ 𝐐
2639
+ ′
2640
+ ‖
2641
+ 𝟐
2642
+ \mathcal{L}_{prm}(\bf{x})=\|\bf{Q}-\bf{Q^{\prime}}\|^{2}
2643
+ (6)
2644
+ Self-evolution Inference Costs.
2645
+ In the initial bootstrap round, we use DeepSeek-Coder-v2-Instruct (236B) as the policy model, using 10 nodes of 8×80GB H100 GPUs with 8 MCTS rollouts. This required approximately two weeks to finish the data generation. For rounds 2–4, using our fine-tuned 7B SLM as the policy model, data generation was performed on 15 nodes of 4×40GB A100 GPUs,
2646
+ with each round completed in three days. In the final round, to include more challenging problems, we increased the number of MCTS rollouts to 64, extending the data generation time to one week.
2647
+ Table 9:
2648
+ Inference costs of
2649
+ \sysname
2650
+ . We show the average number of generated tokens required to generate a trajectory for a given question.
2651
+ MATH
2652
+ AIME 2024
2653
+ AMC 2023
2654
+ Olympiad Bench
2655
+ College Math
2656
+ GSM8K
2657
+ GaokaoEn 2023
2658
+ 5453
2659
+ 15693
2660
+ 14544
2661
+ 7889
2662
+ 4503
2663
+ 3299
2664
+ 6375
2665
+ Inference Setting
2666
+ . In our evaluation, we run multiple MCTS to generate candidate solution trajectories. For each problem, we generate 32 candidate nodes at each step and use the PPM to score each node. Since the PPM effectively provides step-level quality evaluations, we limit MCTS to just 4 rollouts per step to update the Q-values. After completing MCTS, the trajectory with the highest PPM score is selected as the final answer. Table
2667
+ 9
2668
+ presents the average number of tokens generated to produce a trajectory in MCTS.
2669
+ Table 10:
2670
+ Pass@1 (greedy) accuracy of our fine-tuned policy models for Phi3-mini, Qwen2.5-Math-1.5B, Qwen2-Math-7B and Qwen2.5-Math-7B.
2671
+ Model
2672
+ MATH
2673
+ AIME 2024
2674
+ AMC 2023
2675
+ Olympiad Bench
2676
+ College Math
2677
+ GSM8K
2678
+ GaokaoEn 2023
2679
+ General Base Model: Phi3-mini-Instruct (3.8B)
2680
+ Phi3-mini-Instruct
2681
+ 41.4
2682
+ 3.33
2683
+ 7.5
2684
+ 12.3
2685
+ 33.1
2686
+ 85.7
2687
+ 37.1
2688
+ Our policy model
2689
+ 68.0
2690
+ 10.0
2691
+ 37.5
2692
+ 36.6
2693
+ 48.7
2694
+ 87.9
2695
+ 53.2
2696
+ Math-Specialized Base Model: Qwen2.5-Math-1.5B
2697
+ Qwen2.5-Math-1.5B
2698
+ 51.2
2699
+ 0.0
2700
+ 22.5
2701
+ 16.7
2702
+ 38.4
2703
+ 74.6
2704
+ 46.5
2705
+ Qwen2.5-Math-1.5B-Instruct
2706
+ 60.0
2707
+ 10.0
2708
+ 60.0
2709
+ 38.1
2710
+ 47.7
2711
+ 84.8
2712
+ 65.5
2713
+ Our policy model
2714
+ 74.8
2715
+ 13.3
2716
+ 47.5
2717
+ 42.5
2718
+ 50.1
2719
+ 83.1
2720
+ 58.7
2721
+ Math-Specialized Base Model: Qwen2-Math-7B
2722
+ Qwen2-Math-7B
2723
+ 53.4
2724
+ 3.3
2725
+ 25.0
2726
+ 17.3
2727
+ 39.4
2728
+ 80.4
2729
+ 47.3
2730
+ Qwen2-Math-7B-Instruct
2731
+ 73.2
2732
+ 13.3
2733
+ 62.5
2734
+ 38.2
2735
+ 45.9
2736
+ 89.9
2737
+ 62.1
2738
+ Our policy model
2739
+ 73.8
2740
+ 16.7
2741
+ 45.0
2742
+ 43.9
2743
+ 52.0
2744
+ 88.3
2745
+ 65.2
2746
+ Math-Specialized Base Model: Qwen2.5-Math-7B
2747
+ Qwen2.5-Math-7B
2748
+ 58.8
2749
+ 0.0
2750
+ 22.5
2751
+ 21.8
2752
+ 41.6
2753
+ 91.6
2754
+ 51.7
2755
+ Qwen2.5-Math-7B-Instruct
2756
+ 82.6
2757
+ 6.0
2758
+ 62.5
2759
+ 41.6
2760
+ 46.8
2761
+ 95.2
2762
+ 66.8
2763
+ Our policy model
2764
+ 78.4
2765
+ 26.7
2766
+ 47.5
2767
+ 47.1
2768
+ 52.5
2769
+ 89.7
2770
+ 65.7
2771
+ Figure 6:
2772
+ Pass@N accuracy with random sampling from different policy models. Compared to the official Qwen instruct version, our policy model exhibits a stronger ability to sample correct solutions.
2773
+ Figure 7:
2774
+ Pass@N accuracy with PPM-augmented MCTS. Under the same PPM guidance, the four policy models of varying sizes demonstrate convergent capabilities in sampling correct solutions.
2775
+ Pass@N.
2776
+ Table
2777
+ 10
2778
+ compares the math reasoning performance of our policy models with the instruct versions developed by the original model team. Our policy models do not consistently outperform the instruct versions. For example, on the Qwen2.5-Math-7B base model, Qwen2.5-Math-7B-Instruct achieves 4.2% higher accuracy on the MATH benchmark. However, in System 2 deep thinking paradigm, the pass@1 accuracy alone does not fully reflect the reasoning capabilities for the policy model. To provide a more comprehensive evaluation, Fig.
2779
+ 6
2780
+ and Fig.
2781
+ 7
2782
+ present the pass@N accuracy. In this metric, the policy model generates
2783
+ N
2784
+ N
2785
+ solutions under two settings: random sampling (Fig.
2786
+ 6
2787
+ ) and PPM-augmented MCTS deep thinking (Fig.
2788
+ 7
2789
+ ). A problem is considered solved if one solution is correct.
2790
+ As shown in Figure
2791
+ 6
2792
+ , our policy model achieves comparable Pass@64 accuracy with Qwen2.5-Math-7B-Instruct on MATH and significantly outperforms it on other benchmarks.
2793
+ Figure
2794
+ 7
2795
+ further compares the pass@N accuracy of our four policy models (different sizes) after MCTS deep thinking under the same PPM guidance. We can observe that after generating 64 trajectories, the pass@N accuracy of different policy models becomes comparable.
2796
+ A.2
2797
+ Examples of Reasoning Trajectories
2798
+ Self-correction Example
2799
+ Question: The coordinates of a parallelogram are (5, 3), (6, 8), (7, 4) and
2800
+ (
2801
+ x
2802
+ ,
2803
+ y
2804
+ )
2805
+ (x,y)
2806
+ and
2807
+ x
2808
+ >
2809
+ 7
2810
+ x>7
2811
+ . What is the value of
2812
+ x
2813
+ +
2814
+ y
2815
+ x+y
2816
+ Our trajectory
2817
+ :
2818
+ <code>
2819
+ # Import necessary libraries
2820
+ ⬇
2821
+ import
2822
+ numpy
2823
+ as
2824
+ np
2825
+ <end_of_step>
2826
+ # Define the coordinates of the parallelogram
2827
+ ⬇
2828
+ A
2829
+ =
2830
+ np
2831
+ .
2832
+ array
2833
+ ([5,
2834
+ 3])
2835
+ B
2836
+ =
2837
+ np
2838
+ .
2839
+ array
2840
+ ([6,
2841
+ 8])
2842
+ C
2843
+ =
2844
+ np
2845
+ .
2846
+ array
2847
+ ([7,
2848
+ 4])
2849
+ <end_of_step>
2850
+ # Calculate the vector differences to find the missing coordinate D
2851
+ # In a parallelogram, the vector difference between two adjacent sides is equal to the vector difference between the other two adjacent sides
2852
+ # We can use the formula: D = A + (C - B)
2853
+ ⬇
2854
+ D
2855
+ =
2856
+ A
2857
+ +
2858
+ (
2859
+ C
2860
+ -
2861
+ B
2862
+ )
2863
+ <end_of_step>
2864
+ # Since the coordinates provided are (5, 3), (6, 8), (7, 4) and (x, y), we can infer that the vector from B to C should be the same as the vector from A to the missing point (x, y)
2865
+ # This means: (x - 5, y - 3) = (1, -4) because the vector from B to C is (7 - 6, 4 - 8) = (1, -4)
2866
+ # Solve for x and y
2867
+ ⬇
2868
+ x
2869
+ =
2870
+ 5
2871
+ +
2872
+ 1
2873
+ y
2874
+ =
2875
+ 3
2876
+ +
2877
+ (-4)
2878
+ <end_of_step>
2879
+ # Verify that the point (x, y) exists and satisfies the condition x > 7
2880
+ # Since x = 6 is not greater than 7, we need to use the correct formula that maintains the parallelogram property and x > 7
2881
+ # Correcting the formula to find the other possible D [
2882
+ Self-correction!
2883
+ ]
2884
+ ⬇
2885
+ D
2886
+ =
2887
+ C
2888
+ +
2889
+ (
2890
+ B
2891
+ -
2892
+ A
2893
+ )
2894
+ <end_of_step>
2895
+ # Solve for x and y again
2896
+ ⬇
2897
+ x
2898
+ =
2899
+ 7
2900
+ +
2901
+ (6
2902
+ -
2903
+ 5)
2904
+ y
2905
+ =
2906
+ 4
2907
+ +
2908
+ (8
2909
+ -
2910
+ 3)
2911
+ <end_of_step>
2912
+ # Verify that the point (x, y) exists and satisfies the condition x > 7
2913
+ # Since x = 8 is greater than 7, which satisfies the condition given, we can now find the sum of x and y
2914
+ ⬇
2915
+ x_plus_y
2916
+ =
2917
+ x
2918
+ +
2919
+ y
2920
+ <end_of_step>
2921
+ # Now print the final answer
2922
+ ⬇
2923
+ print
2924
+ (
2925
+ x_plus_y
2926
+ )
2927
+ # This should print the value of x + y
2928
+ <end_of_code>
2929
+ <output>17<end_of_output>
2930
+ <answer>The value of x + y is
2931
+ 17
2932
+ <end_of_answer>
2933
+ Math statement proof using Fermat’s little theorem
2934
+ Question: If ( pt+a ) is a perfect square for some value of ( t ), then ( a
2935
+ p
2936
+ −
2937
+ 1
2938
+ 2
2939
+ {}^{\frac{p-1}{2}}
2940
+ - 1 ) is divisible by ( p ) (where ( p ) is a prime number).
2941
+ Our trajectory for proving
2942
+ :
2943
+ <code>
2944
+ # Step 1: Define the variables
2945
+ ⬇
2946
+ from
2947
+ sympy
2948
+ import
2949
+ symbols
2950
+ ,
2951
+ isprime
2952
+ ,
2953
+ sqrt
2954
+ ,
2955
+ mod_inverse
2956
+ # Define the variables
2957
+ ⬇
2958
+ p
2959
+ ,
2960
+ t
2961
+ ,
2962
+ a
2963
+ =
2964
+ symbols
2965
+ (’
2966
+ p
2967
+ t
2968
+ a
2969
+ ’)\\
2970
+ <end_of_step>
2971
+ # Step 2: Assume p*t + a is a perfect square
2972
+ # This means there exists an integer k such that p*t + a = k
2973
+ 2
2974
+ # Rearrange to find a in terms of p, t, and k
2975
+ # a = k
2976
+ 2
2977
+ - p*t
2978
+ <end_of_step>
2979
+ #
2980
+ Step 3: Use Fermat’s Little Theorem
2981
+ # Fermat’s Little Theorem states that if p is a prime and a is an integer not divisible by p,
2982
+ then a
2983
+ (p-1)
2984
+ ≡
2985
+ \equiv
2986
+ 1 (mod p)
2987
+ # We need to show that a
2988
+ ((p-1)/2)
2989
+ - 1 is divisible by p
2990
+ <end_of_step>
2991
+ # Step 4: Substitute a = k
2992
+ 2
2993
+ - p*t into Fermat’s Little Theorem
2994
+ # a
2995
+ ((p-1)/2)
2996
+ ≡
2997
+ \equiv
2998
+ (k
2999
+ 2
3000
+ - p*t)
3001
+ ((p-1)/2)
3002
+ (mod p)
3003
+ # Since p*t is divisible by p, (k
3004
+ 2
3005
+ - p*t)
3006
+ ((p-1)/2)
3007
+ ≡
3008
+ \equiv
3009
+ k
3010
+ (p-1)
3011
+ (mod p)
3012
+ # By Fermat’s Little Theorem, k
3013
+ (p-1)
3014
+ ≡
3015
+ \equiv
3016
+ 1 (mod p)
3017
+ <end_of_step>
3018
+ # Step 5: Therefore, a
3019
+ ((p-1)/2)
3020
+ ≡
3021
+ \equiv
3022
+ 1 (mod p)
3023
+ # This implies a
3024
+ ((p-1)/2)
3025
+ - 1
3026
+ ≡
3027
+ \equiv
3028
+ 0 (mod p)
3029
+ # Hence, a
3030
+ ((p-1)/2)
3031
+ - 1 is divisible by p
3032
+ <end_of_step>
3033
+ # Now print the final answer
3034
+ print("The expression a
3035
+ ((p-1)/2)
3036
+ - 1 is divisible by p.")
3037
+ <end_of_code>
3038
+ <output>
3039
+ The expression
3040
+ a
3041
+ (
3042
+ (
3043
+ p
3044
+ −
3045
+ 1
3046
+ )
3047
+ /
3048
+ 2
3049
+ )
3050
+ a^{((p-1)/2)}
3051
+ - 1 is divisible by p.
3052
+ The statement is true.
3053
+ <end_of_output>
3054
+ Problem-solving with application of the Vieta’s formulas and the AM-GM inequality
3055
+ Question: Let
3056
+ a
3057
+ a
3058
+ be a positive real number such that all the roots of
3059
+ x
3060
+ 3
3061
+ +
3062
+ a
3063
+ ​
3064
+ x
3065
+ 2
3066
+ +
3067
+ a
3068
+ ​
3069
+ x
3070
+ +
3071
+ 1
3072
+ =
3073
+ 0
3074
+ x^{3}+ax^{2}+ax+1=0
3075
+ are real. Find the smallest possible value of
3076
+ a
3077
+ .
3078
+ a.
3079
+ Our trajectory for proving
3080
+ :
3081
+ <code>
3082
+ # Step 1: Define the polynomial equation
3083
+ from sympy import symbols, solve, Eq
3084
+ # Define the variable and the parameter
3085
+ x, a = symbols(’x a’)
3086
+ <end_of_step>
3087
+ [PPM Score: -0.46430]
3088
+ # Step 2: Define the polynomial equation
3089
+ polynomial = x**3 + a*x**2 + a*x + 1
3090
+ <end_of_step>
3091
+ [PPM Score: 0.53937]
3092
+ #
3093
+ Step 3: Use Vieta’s formulas to relate the coefficients to the roots
3094
+ # Since the polynomial is monic (leading coefficient is 1), the sum of the roots (taken one at a time) is -a
3095
+ # The sum of the products of the roots taken two at a time is a
3096
+ # The product of the roots is -1
3097
+ # Since all roots are real, we can denote them as r1, r2, r3
3098
+ # From Vieta’s formulas:
3099
+ # r1 + r2 + r3 = -a
3100
+ # r1*r2 + r2*r3 + r3*r1 = a
3101
+ # r1*r2*r3 = -1
3102
+ <end_of_step>
3103
+ [PPM Score:
3104
+ 0.9989]
3105
+ #
3106
+ Step 4: Use the AM-GM inequality to find the minimum value of a
3107
+ # By AM-GM, (r1*r2 + r2*r3 + r3*r1) / 3 >= (r1*r2*r3)
3108
+ (2/3)
3109
+ # Substituting the known values:
3110
+ # a / 3 >= (-1)
3111
+ (2/3)
3112
+ # a / 3 >= 1
3113
+ # a >= 3
3114
+ <end_of_step>
3115
+ [PPM Score:
3116
+ 0.9999]
3117
+ # Step 5: Verify if a = 3 is achievable
3118
+ # If a = 3, then r1 + r2 + r3 = -3 and r1*r2 + r2*r3 + r3*r1 = 3
3119
+ # We need to check if there exist real roots that satisfy these conditions
3120
+ # Consider the case where r1 = r2 = r3 = -1
3121
+ # Then r1 + r2 + r3 = -3 and r1*r2 + r2*r3 + r3*r1 = 3, and r1*r2*r3 = -1
3122
+ # Thus, a = 3 is achievable with real roots
3123
+ <end_of_step>
3124
+ [PPM Score: 1.0]
3125
+ # Step 6: Print the final answer
3126
+ print(3)
3127
+ <end_of_code>
3128
+ <output>3<end_of_output>
3129
+ <answer>The smallest possible value of a is
3130
+ 3
3131
+ .
3132
+ <end_of_answer>
3133
+ Problem-solving with application the Shoelace Theorem
3134
+ Question: In the diagram,
3135
+ D
3136
+ D
3137
+ and
3138
+ E
3139
+ E
3140
+ are the midpoints of
3141
+ A
3142
+ ​
3143
+ B
3144
+ ¯
3145
+ \overline{AB}
3146
+ and
3147
+ B
3148
+ ​
3149
+ C
3150
+ ¯
3151
+ \overline{BC}
3152
+ respectively. Determine the area of quadrilateral
3153
+ D
3154
+ ​
3155
+ B
3156
+ ​
3157
+ E
3158
+ ​
3159
+ F
3160
+ DBEF
3161
+ . [asy]
3162
+ size(180); defaultpen(linewidth(.7pt)+fontsize(10pt));
3163
+ pair A, B, C, D, E, F;
3164
+ A=(0,6);
3165
+ B=(0,0);
3166
+ C=(8,0);
3167
+ D=(0,3);
3168
+ E=(4,0);
3169
+ F=(8/3,2);
3170
+ draw(E–A–C–D);
3171
+ draw((-1,0)–(10,0), EndArrow);
3172
+ draw((0,-1)–(0,8), EndArrow);
3173
+ label("
3174
+ A
3175
+ ​
3176
+ (
3177
+ 0
3178
+ ,
3179
+ 6
3180
+ )
3181
+ A(0,6)
3182
+ ", A, W);
3183
+ label("
3184
+ B
3185
+ ​
3186
+ (
3187
+ 0
3188
+ ,
3189
+ 0
3190
+ )
3191
+ B(0,0)
3192
+ ", B, SW);
3193
+ label("
3194
+ C
3195
+ ​
3196
+ (
3197
+ 8
3198
+ ,
3199
+ 0
3200
+ )
3201
+ C(8,0)
3202
+ ", C, S);
3203
+ label("
3204
+ D
3205
+ D
3206
+ ", D, W);
3207
+ label("
3208
+ E
3209
+ E
3210
+ ", E, S);
3211
+ label("
3212
+ F
3213
+ F
3214
+ ", F, SW);
3215
+ label("
3216
+ x
3217
+ x
3218
+ ", (10,0), dir(0));
3219
+ label("
3220
+ y
3221
+ y
3222
+ ", (0,8), dir(90));
3223
+ [/asy]
3224
+ Our trajectory for proving
3225
+ :
3226
+ <code>
3227
+ # Step 1: Define the coordinates of points A, B, C, D, and E
3228
+ ⬇
3229
+ A
3230
+ =
3231
+ (0,
3232
+ 6)
3233
+ B
3234
+ =
3235
+ (0,
3236
+ 0)
3237
+ C
3238
+ =
3239
+ (8,
3240
+ 0)
3241
+ D
3242
+ =
3243
+ (0,
3244
+ 3)
3245
+ #
3246
+ Midpoint
3247
+ of
3248
+ AB
3249
+ E
3250
+ =
3251
+ (4,
3252
+ 0)
3253
+ #
3254
+ Midpoint
3255
+ of
3256
+ BC
3257
+ <end_of_step>
3258
+ [PPM Score: -0.76639]
3259
+ # Step 2: Define the coordinates of point F using the section formula (known to be (8/3, 2))
3260
+ ⬇
3261
+ F
3262
+ =
3263
+ (8/3,
3264
+ 2)
3265
+ <end_of_step>
3266
+ [PPM Score: 0.8429]
3267
+ #
3268
+ Step 3: Use the Shoelace Theorem to find the area of quadrilateral DBEF
3269
+ # Shoelace Theorem formula: Area = 0.5 * |x1y2 + x2y3 + x3y4 + x4y1 - (y1x2 + y2x3 + y3x4 + y4x1)|
3270
+ # Coordinates of D, B, E, F are (0, 3), (0, 0), (4, 0), (8/3, 2) respectively
3271
+ ⬇
3272
+ x1
3273
+ ,
3274
+ y1
3275
+ =
3276
+ D
3277
+ x2
3278
+ ,
3279
+ y2
3280
+ =
3281
+ B
3282
+ x3
3283
+ ,
3284
+ y3
3285
+ =
3286
+ E
3287
+ x4
3288
+ ,
3289
+ y4
3290
+ =
3291
+ F
3292
+ area
3293
+ =
3294
+ 0.5
3295
+ *
3296
+ abs
3297
+ (
3298
+ x1
3299
+ *
3300
+ y2
3301
+ +
3302
+ x2
3303
+ *
3304
+ y3
3305
+ +
3306
+ x3
3307
+ *
3308
+ y4
3309
+ +
3310
+ x4
3311
+ *
3312
+ y1
3313
+ -
3314
+ (
3315
+ y1
3316
+ *
3317
+ x2
3318
+ +
3319
+ y2
3320
+ *
3321
+ x3
3322
+ +
3323
+ y3
3324
+ *
3325
+ x4
3326
+ +
3327
+ y4
3328
+ *
3329
+ x1
3330
+ ))
3331
+ <end_of_step>
3332
+ [PPM Score:
3333
+ 0.9972]
3334
+ # Step 4: Print the area of quadrilateral DBEF
3335
+ ⬇
3336
+ print
3337
+ (
3338
+ area
3339
+ )
3340
+ <end_of_code>
3341
+ <output>8.0<end_of_output>
3342
+ <answer>
3343
+ 8
3344
+ <end_of_answer>
3345
+ A.3
3346
+ Prompt Examples
3347
+ MCTS Prompt for DeepSeek-V2-Coder-Instruct (bootstrap round)
3348
+ You are a powerful agent with broad math knowledge and great python programming skills. You need to use python interpreter to do accurate calculation on math equations.
3349
+ !!! Remember:
3350
+ 1. Use code solve the problem step by step. The solution should include three parts: <code>, <output>, and <answer>.
3351
+ 2. All calculations should be done in python code. Provide concise reasoning and thinking in the comments of the code.
3352
+ 3. The most related python packages include ‘math‘, ‘sympy‘, ‘scipy‘, and ‘numpy‘.
3353
+ 4. Please use the following template:
3354
+ Question: the input question
3355
+ <code>Construct the code step by step. Use <end_of_step> to indicate the end of each step. Ensure your code can execute correctly(excluding <end_of_step>) and print the answer. Avoid undefined variables (NameError), unimported packages, or formatting errors (SyntaxError, TypeError). In the last step of the code, print the final answer and add a comment: Now print the final answer.<end_of_code>
3356
+ <output>Execute the code in using the Python interpreter and display the printed results.<end_of_output>
3357
+ <answer>The concise answer without verbose context, put your final answer’s numerical part (without unit, only focus on the numerical part if it’s a choice question) in
3358
+ boxed.<end_of_answer> Now! It’s your turn.
3359
+ Question:
3360
+ {input}
3361
+ The following are 2 demonstration examples:
3362
+ Question: Terrell usually lifts two 20-pound weights 12 times. If he uses two 15-pound weights instead, how many times must Terrell lift them in order to lift the same total weight?
3363
+ <code>
3364
+ # Step 1: Calculate the total weight lifted with two 20-pound weights
3365
+ total_weight_20 = 2 * 20 * 12
3366
+ <end_of_step>
3367
+ # Step 2: Calculate the weight lifted per repetition with two 15-pound weights
3368
+ weight_per_rep_15 = 2 * 15
3369
+ <end_of_step>
3370
+ # Step 3: Calculate the number of repetitions needed to lift the same total weight with two 15-pound weights
3371
+ reps_needed = total_weight_20 / weight_per_rep_15
3372
+ <end_of_step>
3373
+ # Now print the final answer
3374
+ print(reps_needed)
3375
+ <end_of_code>
3376
+ <output>16.0 <end_of_output> <answer>From the result, we can see that Terrell must lift the 15-pound weights
3377
+ boxed16 times to lift the same total weight.
3378
+ <end_of_answer>,
3379
+ Question: Find the value of
3380
+ x
3381
+ x
3382
+ that satisfies
3383
+ 3
3384
+ ​
3385
+ x
3386
+ +
3387
+ 5
3388
+ 6
3389
+ ​
3390
+ x
3391
+ +
3392
+ 5
3393
+ =
3394
+ 5
3395
+ 3
3396
+ \frac{\sqrt{3x+5}}{\sqrt{6x+5}}=\frac{\sqrt{5}}{3}
3397
+ . Express your answer as a common fraction.
3398
+ <code>
3399
+ from sympy import symbols, Eq, solve, sqrt
3400
+ # Define the variable x
3401
+ x = symbols(’x’)
3402
+ <end_of_step>
3403
+ # Define the equation
3404
+ equation = Eq(sqrt(3*x + 5) / sqrt(6*x + 5), sqrt(5) / 3)
3405
+ <end_of_step>
3406
+ # Solve the equation for x
3407
+ solution = solve(equation, x)
3408
+ <end_of_step>
3409
+ # Now print the final answer
3410
+ print(solution)
3411
+ <end_of_code>
3412
+ <output>[20/3]
3413
+ <end_of_output> <answer>From the result, we can see that the value of x is
3414
+ 20
3415
+ 3
3416
+ <end_of_answer>
3417
+ MCTS Prompt for our policy SLM
3418
+ MCTS prompt for next step generation
3419
+ :
3420
+ <|user|>:
3421
+ {input question}
3422
+ <|assistant|>: Let’s think step by step and solve the problem with code.
3423
+ {trajectory of previous steps}
3424
+ SFT Prompt Example
3425
+ :
3426
+ <|user|>:
3427
+ How many numbers can you get by multiplying two or more distinct members of the set
3428
+ 1
3429
+ ,
3430
+ 2
3431
+ ,
3432
+ 3
3433
+ ,
3434
+ 5
3435
+ ,
3436
+ 11
3437
+ \\
3438
+ {1,2,3,5,11\\
3439
+ }
3440
+ together?
3441
+ <|assistant|>: Let’s think step by step and solve the problem with code.
3442
+ <code>
3443
+ # Step 1: Create the set of numbers
3444
+ ⬇
3445
+ numbers
3446
+ =
3447
+ {1,
3448
+ 2,
3449
+ 3,
3450
+ 5,
3451
+ 11}
3452
+ <end_of_step>
3453
+ # Step 2: Define a function to generate all possible products of two or more distinct members of the set
3454
+ ⬇
3455
+ from
3456
+ itertools
3457
+ import
3458
+ combinations
3459
+ def
3460
+ generate_products
3461
+ (
3462
+ numbers
3463
+ ):
3464
+ products
3465
+ =
3466
+ set
3467
+ ()
3468
+ for
3469
+ r
3470
+ in
3471
+ range
3472
+ (2,
3473
+ len
3474
+ (
3475
+ numbers
3476
+ )
3477
+ +
3478
+ 1):
3479
+ #
3480
+ For
3481
+ combinations
3482
+ of
3483
+ 2,
3484
+ 3,
3485
+ 4,
3486
+ and
3487
+ 5
3488
+ elements
3489
+ for
3490
+ combo
3491
+ in
3492
+ combinations
3493
+ (
3494
+ numbers
3495
+ ,
3496
+ r
3497
+ ):
3498
+ product
3499
+ =
3500
+ 1
3501
+ for
3502
+ num
3503
+ in
3504
+ combo
3505
+ :
3506
+ product
3507
+ *=
3508
+ num
3509
+ products
3510
+ .
3511
+ add
3512
+ (
3513
+ product
3514
+ )
3515
+ return
3516
+ products
3517
+ <end_of_step>
3518
+ # Step 3: Calculate the products
3519
+ ⬇
3520
+ unique_products
3521
+ =
3522
+ generate_products
3523
+ (
3524
+ numbers
3525
+ )
3526
+ <end_of_step>
3527
+ # Step 4: Calculate the number of unique products
3528
+ ⬇
3529
+ num_unique_products
3530
+ =
3531
+ len
3532
+ (
3533
+ unique_products
3534
+ )
3535
+ <end_of_step>
3536
+ # Now print the final answer
3537
+ ⬇
3538
+ print
3539
+ (
3540
+ num_unique_products
3541
+ )
3542
+ <end_of_code>
3543
+ <output>15<end_of_output>
3544
+ <answer>From the result, we can see that there are
3545
+ 15
3546
+ unique products.
3547
+ <end_of_answer>
3548
+ ◄
3549
+ Feeling
3550
+ lucky?
3551
+ Conversion
3552
+ report
3553
+ Report
3554
+ an issue
3555
+ View original
3556
+ on arXiv
3557
+ ►
research/notes/250109136-agentic-retrieval-augmented-generation-a-survey-on-agentic-rag.md ADDED
@@ -0,0 +1,203 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2501.09136] Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG'
3
+ id: 250109136-agentic-retrieval-augmented-generation-a-survey-on-agentic-rag
4
+ tags:
5
+ - deepread
6
+ created: '2026-06-10T00:24:47.557837Z'
7
+ source: https://arxiv.org/abs/2501.09136
8
+ source_domain: arxiv.org
9
+ fetched_at: '2026-06-10T00:24:47.557707Z'
10
+ fetch_provider: builtin
11
+ status: draft
12
+ type: note
13
+ tier: institutional
14
+ content_type: paper
15
+ deprecated: false
16
+ ---
17
+
18
+ [2501.09136] Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
19
+ Computer Science > Artificial Intelligence
20
+ arXiv:2501.09136
21
+ (cs)
22
+ [Submitted on 15 Jan 2025 (
23
+ v1
24
+ ), last revised 1 Apr 2026 (this version, v4)]
25
+ Title:
26
+ Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
27
+ Authors:
28
+ Aditi Singh
29
+ ,
30
+ Abul Ehtesham
31
+ ,
32
+ Saket Kumar
33
+ ,
34
+ Tala Talaei Khoei
35
+ ,
36
+ Athanasios V. Vasilakos
37
+ View a PDF of the paper titled Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG, by Aditi Singh and 4 other authors
38
+ View PDF
39
+ HTML (experimental)
40
+ Abstract:
41
+ Large Language Models (LLMs) have advanced artificial intelligence by enabling human-like text generation and natural language understanding. However, their reliance on static training data limits their ability to respond to dynamic, real-time queries, resulting in outdated or inaccurate outputs. Retrieval-Augmented Generation (RAG) has emerged as a solution, enhancing LLMs by integrating real-time data retrieval to provide contextually relevant and up-to-date responses. Despite its promise, traditional RAG systems are constrained by static workflows and lack the adaptability required for multi-step reasoning and complex task management. Agentic Retrieval-Augmented Generation (Agentic RAG) transcends these limitations by embedding autonomous AI agents into the RAG pipeline. These agents leverage agentic design patterns reflection, planning, tool use, and multi-agent collaboration to dynamically manage retrieval strategies, iteratively refine contextual understanding, and adapt workflows through operational structures ranging from sequential steps to adaptive collaboration. This integration enables Agentic RAG systems to deliver flexibility, scalability, and context-awareness across diverse applications. This paper presents an analytical survey of Agentic RAG systems. It traces the evolution of RAG paradigms, introduces a principled taxonomy of Agentic RAG architectures based on agent cardinality, control structure, autonomy, and knowledge representation, and provides a comparative analysis of design trade-offs across existing frameworks. The survey examines applications in healthcare, finance, education, and enterprise document processing, and distills practical lessons for system designers and practitioners. Finally, it identifies key open research challenges related to evaluation, coordination, memory management, efficiency, and governance, outlining directions for future research.
42
+ Subjects:
43
+ Artificial Intelligence (cs.AI)
44
+ ; Computation and Language (cs.CL); Information Retrieval (cs.IR)
45
+ Cite as:
46
+ arXiv:2501.09136
47
+ [cs.AI]
48
+ (or
49
+ arXiv:2501.09136v4
50
+ [cs.AI]
51
+ for this version)
52
+ https://doi.org/10.48550/arXiv.2501.09136
53
+ Focus to learn more
54
+ arXiv-issued DOI via DataCite
55
+ Submission history
56
+ From: Abul Ehtesham [
57
+ view email
58
+ ]
59
+ [v1]
60
+ Wed, 15 Jan 2025 20:40:25 UTC (20,962 KB)
61
+ [v2]
62
+ Mon, 3 Feb 2025 04:01:36 UTC (22,453 KB)
63
+ [v3]
64
+ Tue, 4 Feb 2025 04:48:00 UTC (22,430 KB)
65
+ [v4]
66
+ Wed, 1 Apr 2026 15:51:06 UTC (13,996 KB)
67
+ Full-text links:
68
+ Access Paper:
69
+ View a PDF of the paper titled Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG, by Aditi Singh and 4 other authors
70
+ View PDF
71
+ HTML (experimental)
72
+ TeX Source
73
+ view license
74
+ Current browse context:
75
+ cs.AI
76
+ < prev
77
+ |
78
+ next >
79
+ new
80
+ |
81
+ recent
82
+ |
83
+ 2025-01
84
+ Change to browse by:
85
+ cs
86
+ cs.CL
87
+ cs.IR
88
+ References & Citations
89
+ NASA ADS
90
+ Google Scholar
91
+ Semantic Scholar
92
+ export BibTeX citation
93
+ Loading...
94
+ BibTeX formatted citation
95
+ ×
96
+ loading...
97
+ Data provided by:
98
+ Bookmark
99
+ Bibliographic Tools
100
+ Bibliographic and Citation Tools
101
+ Bibliographic Explorer Toggle
102
+ Bibliographic Explorer
103
+ (
104
+ What is the Explorer?
105
+ )
106
+ Connected Papers Toggle
107
+ Connected Papers
108
+ (
109
+ What is Connected Papers?
110
+ )
111
+ Litmaps Toggle
112
+ Litmaps
113
+ (
114
+ What is Litmaps?
115
+ )
116
+ scite.ai Toggle
117
+ scite Smart Citations
118
+ (
119
+ What are Smart Citations?
120
+ )
121
+ Code, Data, Media
122
+ Code, Data and Media Associated with this Article
123
+ alphaXiv Toggle
124
+ alphaXiv
125
+ (
126
+ What is alphaXiv?
127
+ )
128
+ Links to Code Toggle
129
+ CatalyzeX Code Finder for Papers
130
+ (
131
+ What is CatalyzeX?
132
+ )
133
+ DagsHub Toggle
134
+ DagsHub
135
+ (
136
+ What is DagsHub?
137
+ )
138
+ GotitPub Toggle
139
+ Gotit.pub
140
+ (
141
+ What is GotitPub?
142
+ )
143
+ Huggingface Toggle
144
+ Hugging Face
145
+ (
146
+ What is Huggingface?
147
+ )
148
+ Links to Code Toggle
149
+ Papers with Code
150
+ (
151
+ What is Papers with Code?
152
+ )
153
+ ScienceCast Toggle
154
+ ScienceCast
155
+ (
156
+ What is ScienceCast?
157
+ )
158
+ Demos
159
+ Demos
160
+ Replicate Toggle
161
+ Replicate
162
+ (
163
+ What is Replicate?
164
+ )
165
+ Spaces Toggle
166
+ Hugging Face Spaces
167
+ (
168
+ What is Spaces?
169
+ )
170
+ Spaces Toggle
171
+ TXYZ.AI
172
+ (
173
+ What is TXYZ.AI?
174
+ )
175
+ Related Papers
176
+ Recommenders and Search Tools
177
+ Link to Influence Flower
178
+ Influence Flower
179
+ (
180
+ What are Influence Flowers?
181
+ )
182
+ Core recommender toggle
183
+ CORE Recommender
184
+ (
185
+ What is CORE?
186
+ )
187
+ Author
188
+ Venue
189
+ Institution
190
+ Topic
191
+ About arXivLabs
192
+ arXivLabs: experimental projects with community collaborators
193
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
194
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
195
+ Have an idea for a project that will add value for arXiv's community?
196
+ Learn more about arXivLabs
197
+ .
198
+ Which authors of this paper are endorsers?
199
+ |
200
+ Disable MathJax
201
+ (
202
+ What is MathJax?
203
+ )
research/notes/250109891-evolving-deeper-llm-thinking.md ADDED
@@ -0,0 +1,196 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2501.09891] Evolving Deeper LLM Thinking'
3
+ id: 250109891-evolving-deeper-llm-thinking
4
+ tags:
5
+ - deepread
6
+ created: '2026-06-10T00:24:58.674469Z'
7
+ source: https://arxiv.org/abs/2501.09891
8
+ source_domain: arxiv.org
9
+ fetched_at: '2026-06-10T00:24:58.674344Z'
10
+ fetch_provider: builtin
11
+ status: draft
12
+ type: note
13
+ tier: institutional
14
+ content_type: paper
15
+ deprecated: false
16
+ ---
17
+
18
+ [2501.09891] Evolving Deeper LLM Thinking
19
+ Computer Science > Artificial Intelligence
20
+ arXiv:2501.09891
21
+ (cs)
22
+ [Submitted on 17 Jan 2025]
23
+ Title:
24
+ Evolving Deeper LLM Thinking
25
+ Authors:
26
+ Kuang-Huei Lee
27
+ ,
28
+ Ian Fischer
29
+ ,
30
+ Yueh-Hua Wu
31
+ ,
32
+ Dave Marwood
33
+ ,
34
+ Shumeet Baluja
35
+ ,
36
+ Dale Schuurmans
37
+ ,
38
+ Xinyun Chen
39
+ View a PDF of the paper titled Evolving Deeper LLM Thinking, by Kuang-Huei Lee and 6 other authors
40
+ View PDF
41
+ HTML (experimental)
42
+ Abstract:
43
+ We explore an evolutionary search strategy for scaling inference time compute in Large Language Models. The proposed approach, Mind Evolution, uses a language model to generate, recombine and refine candidate responses. The proposed approach avoids the need to formalize the underlying inference problem whenever a solution evaluator is available. Controlling for inference cost, we find that Mind Evolution significantly outperforms other inference strategies such as Best-of-N and Sequential Revision in natural language planning tasks. In the TravelPlanner and Natural Plan benchmarks, Mind Evolution solves more than 98% of the problem instances using Gemini 1.5 Pro without the use of a formal solver.
44
+ Subjects:
45
+ Artificial Intelligence (cs.AI)
46
+ Cite as:
47
+ arXiv:2501.09891
48
+ [cs.AI]
49
+ (or
50
+ arXiv:2501.09891v1
51
+ [cs.AI]
52
+ for this version)
53
+ https://doi.org/10.48550/arXiv.2501.09891
54
+ Focus to learn more
55
+ arXiv-issued DOI via DataCite
56
+ Submission history
57
+ From: Dale Schuurmans [
58
+ view email
59
+ ]
60
+ [v1]
61
+ Fri, 17 Jan 2025 00:41:44 UTC (3,183 KB)
62
+ Full-text links:
63
+ Access Paper:
64
+ View a PDF of the paper titled Evolving Deeper LLM Thinking, by Kuang-Huei Lee and 6 other authors
65
+ View PDF
66
+ HTML (experimental)
67
+ TeX Source
68
+ view license
69
+ Current browse context:
70
+ cs.AI
71
+ < prev
72
+ |
73
+ next >
74
+ new
75
+ |
76
+ recent
77
+ |
78
+ 2025-01
79
+ Change to browse by:
80
+ cs
81
+ References & Citations
82
+ NASA ADS
83
+ Google Scholar
84
+ Semantic Scholar
85
+ export BibTeX citation
86
+ Loading...
87
+ BibTeX formatted citation
88
+ ×
89
+ loading...
90
+ Data provided by:
91
+ Bookmark
92
+ Bibliographic Tools
93
+ Bibliographic and Citation Tools
94
+ Bibliographic Explorer Toggle
95
+ Bibliographic Explorer
96
+ (
97
+ What is the Explorer?
98
+ )
99
+ Connected Papers Toggle
100
+ Connected Papers
101
+ (
102
+ What is Connected Papers?
103
+ )
104
+ Litmaps Toggle
105
+ Litmaps
106
+ (
107
+ What is Litmaps?
108
+ )
109
+ scite.ai Toggle
110
+ scite Smart Citations
111
+ (
112
+ What are Smart Citations?
113
+ )
114
+ Code, Data, Media
115
+ Code, Data and Media Associated with this Article
116
+ alphaXiv Toggle
117
+ alphaXiv
118
+ (
119
+ What is alphaXiv?
120
+ )
121
+ Links to Code Toggle
122
+ CatalyzeX Code Finder for Papers
123
+ (
124
+ What is CatalyzeX?
125
+ )
126
+ DagsHub Toggle
127
+ DagsHub
128
+ (
129
+ What is DagsHub?
130
+ )
131
+ GotitPub Toggle
132
+ Gotit.pub
133
+ (
134
+ What is GotitPub?
135
+ )
136
+ Huggingface Toggle
137
+ Hugging Face
138
+ (
139
+ What is Huggingface?
140
+ )
141
+ Links to Code Toggle
142
+ Papers with Code
143
+ (
144
+ What is Papers with Code?
145
+ )
146
+ ScienceCast Toggle
147
+ ScienceCast
148
+ (
149
+ What is ScienceCast?
150
+ )
151
+ Demos
152
+ Demos
153
+ Replicate Toggle
154
+ Replicate
155
+ (
156
+ What is Replicate?
157
+ )
158
+ Spaces Toggle
159
+ Hugging Face Spaces
160
+ (
161
+ What is Spaces?
162
+ )
163
+ Spaces Toggle
164
+ TXYZ.AI
165
+ (
166
+ What is TXYZ.AI?
167
+ )
168
+ Related Papers
169
+ Recommenders and Search Tools
170
+ Link to Influence Flower
171
+ Influence Flower
172
+ (
173
+ What are Influence Flowers?
174
+ )
175
+ Core recommender toggle
176
+ CORE Recommender
177
+ (
178
+ What is CORE?
179
+ )
180
+ Author
181
+ Venue
182
+ Institution
183
+ Topic
184
+ About arXivLabs
185
+ arXivLabs: experimental projects with community collaborators
186
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
187
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
188
+ Have an idea for a project that will add value for arXiv's community?
189
+ Learn more about arXivLabs
190
+ .
191
+ Which authors of this paper are endorsers?
192
+ |
193
+ Disable MathJax
194
+ (
195
+ What is MathJax?
196
+ )
research/notes/250112599-kimi-k15-scaling-reinforcement-learning-with-llms.md ADDED
@@ -0,0 +1,386 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2501.12599] Kimi k1.5: Scaling Reinforcement Learning with LLMs'
3
+ id: 250112599-kimi-k15-scaling-reinforcement-learning-with-llms
4
+ tags:
5
+ - deepread
6
+ created: '2026-06-10T00:24:53.655467Z'
7
+ source: https://arxiv.org/abs/2501.12599
8
+ source_domain: arxiv.org
9
+ fetched_at: '2026-06-10T00:24:53.655188Z'
10
+ fetch_provider: builtin
11
+ status: draft
12
+ type: note
13
+ tier: institutional
14
+ content_type: paper
15
+ deprecated: false
16
+ ---
17
+
18
+ [2501.12599] Kimi k1.5: Scaling Reinforcement Learning with LLMs
19
+ Computer Science > Artificial Intelligence
20
+ arXiv:2501.12599
21
+ (cs)
22
+ [Submitted on 22 Jan 2025 (
23
+ v1
24
+ ), last revised 3 Jun 2025 (this version, v4)]
25
+ Title:
26
+ Kimi k1.5: Scaling Reinforcement Learning with LLMs
27
+ Authors:
28
+ Kimi Team
29
+ ,
30
+ Angang Du
31
+ ,
32
+ Bofei Gao
33
+ ,
34
+ Bowei Xing
35
+ ,
36
+ Changjiu Jiang
37
+ ,
38
+ Cheng Chen
39
+ ,
40
+ Cheng Li
41
+ ,
42
+ Chenjun Xiao
43
+ ,
44
+ Chenzhuang Du
45
+ ,
46
+ Chonghua Liao
47
+ ,
48
+ Chuning Tang
49
+ ,
50
+ Congcong Wang
51
+ ,
52
+ Dehao Zhang
53
+ ,
54
+ Enming Yuan
55
+ ,
56
+ Enzhe Lu
57
+ ,
58
+ Fengxiang Tang
59
+ ,
60
+ Flood Sung
61
+ ,
62
+ Guangda Wei
63
+ ,
64
+ Guokun Lai
65
+ ,
66
+ Haiqing Guo
67
+ ,
68
+ Han Zhu
69
+ ,
70
+ Hao Ding
71
+ ,
72
+ Hao Hu
73
+ ,
74
+ Hao Yang
75
+ ,
76
+ Hao Zhang
77
+ ,
78
+ Haotian Yao
79
+ ,
80
+ Haotian Zhao
81
+ ,
82
+ Haoyu Lu
83
+ ,
84
+ Haoze Li
85
+ ,
86
+ Haozhen Yu
87
+ ,
88
+ Hongcheng Gao
89
+ ,
90
+ Huabin Zheng
91
+ ,
92
+ Huan Yuan
93
+ ,
94
+ Jia Chen
95
+ ,
96
+ Jianhang Guo
97
+ ,
98
+ Jianlin Su
99
+ ,
100
+ Jianzhou Wang
101
+ ,
102
+ Jie Zhao
103
+ ,
104
+ Jin Zhang
105
+ ,
106
+ Jingyuan Liu
107
+ ,
108
+ Junjie Yan
109
+ ,
110
+ Junyan Wu
111
+ ,
112
+ Lidong Shi
113
+ ,
114
+ Ling Ye
115
+ ,
116
+ Longhui Yu
117
+ ,
118
+ Mengnan Dong
119
+ ,
120
+ Neo Zhang
121
+ ,
122
+ Ningchen Ma
123
+ ,
124
+ Qiwei Pan
125
+ ,
126
+ Qucheng Gong
127
+ ,
128
+ Shaowei Liu
129
+ ,
130
+ Shengling Ma
131
+ ,
132
+ Shupeng Wei
133
+ ,
134
+ Sihan Cao
135
+ ,
136
+ Siying Huang
137
+ ,
138
+ Tao Jiang
139
+ ,
140
+ Weihao Gao
141
+ ,
142
+ Weimin Xiong
143
+ ,
144
+ Weiran He
145
+ ,
146
+ Weixiao Huang
147
+ ,
148
+ Weixin Xu
149
+ ,
150
+ Wenhao Wu
151
+ ,
152
+ Wenyang He
153
+ ,
154
+ Xianghui Wei
155
+ ,
156
+ Xianqing Jia
157
+ ,
158
+ Xingzhe Wu
159
+ ,
160
+ Xinran Xu
161
+ ,
162
+ Xinxing Zu
163
+ ,
164
+ Xinyu Zhou
165
+ ,
166
+ Xuehai Pan
167
+ ,
168
+ Y. Charles
169
+ ,
170
+ Yang Li
171
+ ,
172
+ Yangyang Hu
173
+ ,
174
+ Yangyang Liu
175
+ ,
176
+ Yanru Chen
177
+ ,
178
+ Yejie Wang
179
+ ,
180
+ Yibo Liu
181
+ ,
182
+ Yidao Qin
183
+ ,
184
+ Yifeng Liu
185
+ ,
186
+ Ying Yang
187
+ ,
188
+ Yiping Bao
189
+ ,
190
+ Yulun Du
191
+ ,
192
+ Yuxin Wu
193
+ ,
194
+ Yuzhi Wang
195
+ ,
196
+ Zaida Zhou
197
+ ,
198
+ Zhaoji Wang
199
+ ,
200
+ Zhaowei Li
201
+ ,
202
+ Zhen Zhu
203
+ ,
204
+ Zheng Zhang
205
+ ,
206
+ Zhexu Wang
207
+ ,
208
+ Zhilin Yang
209
+ ,
210
+ Zhiqi Huang
211
+ ,
212
+ Zihao Huang
213
+ ,
214
+ Ziyao Xu
215
+ ,
216
+ Zonghan Yang
217
+ ,
218
+ Zongyu Lin
219
+ View a PDF of the paper titled Kimi k1.5: Scaling Reinforcement Learning with LLMs, by Kimi Team and 95 other authors
220
+ View PDF
221
+ HTML (experimental)
222
+ Abstract:
223
+ Language model pretraining with next token prediction has proved effective for scaling compute but is limited to the amount of available training data. Scaling reinforcement learning (RL) unlocks a new axis for the continued improvement of artificial intelligence, with the promise that large language models (LLMs) can scale their training data by learning to explore with rewards. However, prior published work has not produced competitive results. In light of this, we report on the training practice of Kimi k1.5, our latest multi-modal LLM trained with RL, including its RL training techniques, multi-modal data recipes, and infrastructure optimization. Long context scaling and improved policy optimization methods are key ingredients of our approach, which establishes a simplistic, effective RL framework without relying on more complex techniques such as Monte Carlo tree search, value functions, and process reward models. Notably, our system achieves state-of-the-art reasoning performance across multiple benchmarks and modalities -- e.g., 77.5 on AIME, 96.2 on MATH 500, 94-th percentile on Codeforces, 74.9 on MathVista -- matching OpenAI's o1. Moreover, we present effective long2short methods that use long-CoT techniques to improve short-CoT models, yielding state-of-the-art short-CoT reasoning results -- e.g., 60.8 on AIME, 94.6 on MATH500, 47.3 on LiveCodeBench -- outperforming existing short-CoT models such as GPT-4o and Claude Sonnet 3.5 by a large margin (up to +550%).
224
+ Comments:
225
+ 25 pages
226
+ Subjects:
227
+ Artificial Intelligence (cs.AI)
228
+ ; Machine Learning (cs.LG)
229
+ Cite as:
230
+ arXiv:2501.12599
231
+ [cs.AI]
232
+ (or
233
+ arXiv:2501.12599v4
234
+ [cs.AI]
235
+ for this version)
236
+ https://doi.org/10.48550/arXiv.2501.12599
237
+ Focus to learn more
238
+ arXiv-issued DOI via DataCite
239
+ Submission history
240
+ From: Flood Sung [
241
+ view email
242
+ ]
243
+ [v1]
244
+ Wed, 22 Jan 2025 02:48:14 UTC (614 KB)
245
+ [v2]
246
+ Wed, 5 Mar 2025 02:16:32 UTC (614 KB)
247
+ [v3]
248
+ Wed, 28 May 2025 03:57:30 UTC (614 KB)
249
+ [v4]
250
+ Tue, 3 Jun 2025 02:14:54 UTC (603 KB)
251
+ Full-text links:
252
+ Access Paper:
253
+ View a PDF of the paper titled Kimi k1.5: Scaling Reinforcement Learning with LLMs, by Kimi Team and 95 other authors
254
+ View PDF
255
+ HTML (experimental)
256
+ TeX Source
257
+ view license
258
+ Current browse context:
259
+ cs.AI
260
+ < prev
261
+ |
262
+ next >
263
+ new
264
+ |
265
+ recent
266
+ |
267
+ 2025-01
268
+ Change to browse by:
269
+ cs
270
+ cs.LG
271
+ References & Citations
272
+ NASA ADS
273
+ Google Scholar
274
+ Semantic Scholar
275
+ export BibTeX citation
276
+ Loading...
277
+ BibTeX formatted citation
278
+ ×
279
+ loading...
280
+ Data provided by:
281
+ Bookmark
282
+ Bibliographic Tools
283
+ Bibliographic and Citation Tools
284
+ Bibliographic Explorer Toggle
285
+ Bibliographic Explorer
286
+ (
287
+ What is the Explorer?
288
+ )
289
+ Connected Papers Toggle
290
+ Connected Papers
291
+ (
292
+ What is Connected Papers?
293
+ )
294
+ Litmaps Toggle
295
+ Litmaps
296
+ (
297
+ What is Litmaps?
298
+ )
299
+ scite.ai Toggle
300
+ scite Smart Citations
301
+ (
302
+ What are Smart Citations?
303
+ )
304
+ Code, Data, Media
305
+ Code, Data and Media Associated with this Article
306
+ alphaXiv Toggle
307
+ alphaXiv
308
+ (
309
+ What is alphaXiv?
310
+ )
311
+ Links to Code Toggle
312
+ CatalyzeX Code Finder for Papers
313
+ (
314
+ What is CatalyzeX?
315
+ )
316
+ DagsHub Toggle
317
+ DagsHub
318
+ (
319
+ What is DagsHub?
320
+ )
321
+ GotitPub Toggle
322
+ Gotit.pub
323
+ (
324
+ What is GotitPub?
325
+ )
326
+ Huggingface Toggle
327
+ Hugging Face
328
+ (
329
+ What is Huggingface?
330
+ )
331
+ Links to Code Toggle
332
+ Papers with Code
333
+ (
334
+ What is Papers with Code?
335
+ )
336
+ ScienceCast Toggle
337
+ ScienceCast
338
+ (
339
+ What is ScienceCast?
340
+ )
341
+ Demos
342
+ Demos
343
+ Replicate Toggle
344
+ Replicate
345
+ (
346
+ What is Replicate?
347
+ )
348
+ Spaces Toggle
349
+ Hugging Face Spaces
350
+ (
351
+ What is Spaces?
352
+ )
353
+ Spaces Toggle
354
+ TXYZ.AI
355
+ (
356
+ What is TXYZ.AI?
357
+ )
358
+ Related Papers
359
+ Recommenders and Search Tools
360
+ Link to Influence Flower
361
+ Influence Flower
362
+ (
363
+ What are Influence Flowers?
364
+ )
365
+ Core recommender toggle
366
+ CORE Recommender
367
+ (
368
+ What is CORE?
369
+ )
370
+ Author
371
+ Venue
372
+ Institution
373
+ Topic
374
+ About arXivLabs
375
+ arXivLabs: experimental projects with community collaborators
376
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
377
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
378
+ Have an idea for a project that will add value for arXiv's community?
379
+ Learn more about arXivLabs
380
+ .
381
+ Which authors of this paper are endorsers?
382
+ |
383
+ Disable MathJax
384
+ (
385
+ What is MathJax?
386
+ )
research/notes/250118512-streaming-diloco-with-overlapping-communication-towards-a-distributed.md ADDED
@@ -0,0 +1,206 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2501.18512] Streaming DiLoCo with overlapping communication: Towards a Distributed
3
+ Free Lunch'
4
+ id: 250118512-streaming-diloco-with-overlapping-communication-towards-a-distributed
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:30:21.211856Z'
8
+ source: https://arxiv.org/abs/2501.18512
9
+ source_domain: arxiv.org
10
+ fetched_at: '2026-06-10T00:30:21.211709Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2501.18512] Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
20
+ Computer Science > Computation and Language
21
+ arXiv:2501.18512
22
+ (cs)
23
+ [Submitted on 30 Jan 2025]
24
+ Title:
25
+ Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
26
+ Authors:
27
+ Arthur Douillard
28
+ ,
29
+ Yanislav Donchev
30
+ ,
31
+ Keith Rush
32
+ ,
33
+ Satyen Kale
34
+ ,
35
+ Zachary Charles
36
+ ,
37
+ Zachary Garrett
38
+ ,
39
+ Gabriel Teston
40
+ ,
41
+ Dave Lacey
42
+ ,
43
+ Ross McIlroy
44
+ ,
45
+ Jiajun Shen
46
+ ,
47
+ Alexandre Ramé
48
+ ,
49
+ Arthur Szlam
50
+ ,
51
+ Marc'Aurelio Ranzato
52
+ ,
53
+ Paul Barham
54
+ View a PDF of the paper titled Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch, by Arthur Douillard and Yanislav Donchev and Keith Rush and Satyen Kale and Zachary Charles and Zachary Garrett and Gabriel Teston and Dave Lacey and Ross McIlroy and Jiajun Shen and Alexandre Ram\'e and Arthur Szlam and Marc'Aurelio Ranzato and Paul Barham
55
+ View PDF
56
+ HTML (experimental)
57
+ Abstract:
58
+ Training of large language models (LLMs) is typically distributed across a large number of accelerators to reduce training time. Since internal states and parameter gradients need to be exchanged at each and every single gradient step, all devices need to be co-located using low-latency high-bandwidth communication links to support the required high volume of exchanged bits. Recently, distributed algorithms like DiLoCo have relaxed such co-location constraint: accelerators can be grouped into ``workers'', where synchronizations between workers only occur infrequently. This in turn means that workers can afford being connected by lower bandwidth communication links without affecting learning quality. However, in these methods, communication across workers still requires the same peak bandwidth as before, as the synchronizations require all parameters to be exchanged across all workers. In this paper, we improve DiLoCo in three ways. First, we synchronize only subsets of parameters in sequence, rather than all at once, which greatly reduces peak bandwidth. Second, we allow workers to continue training while synchronizing, which decreases wall clock time. Third, we quantize the data exchanged by workers, which further reduces bandwidth across workers. By properly combining these modifications, we show experimentally that we can distribute training of billion-scale parameters and reach similar quality as before, but reducing required bandwidth by two orders of magnitude.
59
+ Subjects:
60
+ Computation and Language (cs.CL)
61
+ Cite as:
62
+ arXiv:2501.18512
63
+ [cs.CL]
64
+ (or
65
+ arXiv:2501.18512v1
66
+ [cs.CL]
67
+ for this version)
68
+ https://doi.org/10.48550/arXiv.2501.18512
69
+ Focus to learn more
70
+ arXiv-issued DOI via DataCite
71
+ Submission history
72
+ From: Arthur Douillard [
73
+ view email
74
+ ]
75
+ [v1]
76
+ Thu, 30 Jan 2025 17:23:50 UTC (3,278 KB)
77
+ Full-text links:
78
+ Access Paper:
79
+ View a PDF of the paper titled Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch, by Arthur Douillard and Yanislav Donchev and Keith Rush and Satyen Kale and Zachary Charles and Zachary Garrett and Gabriel Teston and Dave Lacey and Ross McIlroy and Jiajun Shen and Alexandre Ram\'e and Arthur Szlam and Marc'Aurelio Ranzato and Paul Barham
80
+ View PDF
81
+ HTML (experimental)
82
+ TeX Source
83
+ view license
84
+ Current browse context:
85
+ cs.CL
86
+ < prev
87
+ |
88
+ next >
89
+ new
90
+ |
91
+ recent
92
+ |
93
+ 2025-01
94
+ Change to browse by:
95
+ cs
96
+ References & Citations
97
+ NASA ADS
98
+ Google Scholar
99
+ Semantic Scholar
100
+ export BibTeX citation
101
+ Loading...
102
+ BibTeX formatted citation
103
+ ×
104
+ loading...
105
+ Data provided by:
106
+ Bookmark
107
+ Bibliographic Tools
108
+ Bibliographic and Citation Tools
109
+ Bibliographic Explorer Toggle
110
+ Bibliographic Explorer
111
+ (
112
+ What is the Explorer?
113
+ )
114
+ Connected Papers Toggle
115
+ Connected Papers
116
+ (
117
+ What is Connected Papers?
118
+ )
119
+ Litmaps Toggle
120
+ Litmaps
121
+ (
122
+ What is Litmaps?
123
+ )
124
+ scite.ai Toggle
125
+ scite Smart Citations
126
+ (
127
+ What are Smart Citations?
128
+ )
129
+ Code, Data, Media
130
+ Code, Data and Media Associated with this Article
131
+ alphaXiv Toggle
132
+ alphaXiv
133
+ (
134
+ What is alphaXiv?
135
+ )
136
+ Links to Code Toggle
137
+ CatalyzeX Code Finder for Papers
138
+ (
139
+ What is CatalyzeX?
140
+ )
141
+ DagsHub Toggle
142
+ DagsHub
143
+ (
144
+ What is DagsHub?
145
+ )
146
+ GotitPub Toggle
147
+ Gotit.pub
148
+ (
149
+ What is GotitPub?
150
+ )
151
+ Huggingface Toggle
152
+ Hugging Face
153
+ (
154
+ What is Huggingface?
155
+ )
156
+ ScienceCast Toggle
157
+ ScienceCast
158
+ (
159
+ What is ScienceCast?
160
+ )
161
+ Demos
162
+ Demos
163
+ Replicate Toggle
164
+ Replicate
165
+ (
166
+ What is Replicate?
167
+ )
168
+ Spaces Toggle
169
+ Hugging Face Spaces
170
+ (
171
+ What is Spaces?
172
+ )
173
+ Spaces Toggle
174
+ TXYZ.AI
175
+ (
176
+ What is TXYZ.AI?
177
+ )
178
+ Related Papers
179
+ Recommenders and Search Tools
180
+ Link to Influence Flower
181
+ Influence Flower
182
+ (
183
+ What are Influence Flowers?
184
+ )
185
+ Core recommender toggle
186
+ CORE Recommender
187
+ (
188
+ What is CORE?
189
+ )
190
+ Author
191
+ Venue
192
+ Institution
193
+ Topic
194
+ About arXivLabs
195
+ arXivLabs: experimental projects with community collaborators
196
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
197
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
198
+ Have an idea for a project that will add value for arXiv's community?
199
+ Learn more about arXivLabs
200
+ .
201
+ Which authors of this paper are endorsers?
202
+ |
203
+ Disable MathJax
204
+ (
205
+ What is MathJax?
206
+ )
research/notes/250118639-a-comprehensive-survey-of-the-lean-4-theorem-prover-architecture-appli.md ADDED
@@ -0,0 +1,179 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2501.18639] A Comprehensive Survey of the Lean 4 Theorem Prover: Architecture,
3
+ Applications, and Advances'
4
+ id: 250118639-a-comprehensive-survey-of-the-lean-4-theorem-prover-architecture-appli
5
+ tags:
6
+ - deepread
7
+ created: '2026-06-10T00:25:14.929249Z'
8
+ source: https://arxiv.org/abs/2501.18639
9
+ source_domain: arxiv.org
10
+ fetched_at: '2026-06-10T00:25:14.929112Z'
11
+ fetch_provider: builtin
12
+ status: draft
13
+ type: note
14
+ tier: institutional
15
+ content_type: paper
16
+ deprecated: false
17
+ ---
18
+
19
+ [2501.18639] A Comprehensive Survey of the Lean 4 Theorem Prover: Architecture, Applications, and Advances
20
+ Computer Science > Logic in Computer Science
21
+ arXiv:2501.18639
22
+ (cs)
23
+ [Submitted on 28 Jan 2025]
24
+ Title:
25
+ A Comprehensive Survey of the Lean 4 Theorem Prover: Architecture, Applications, and Advances
26
+ Authors:
27
+ Xichen Tang
28
+ View a PDF of the paper titled A Comprehensive Survey of the Lean 4 Theorem Prover: Architecture, Applications, and Advances, by Xichen Tang
29
+ View PDF
30
+ Abstract:
31
+ This comprehensive survey examines Lean 4, a state-of-the-art interactive theorem prover and functional programming language. We analyze its architectural design, type system, metaprogramming capabilities, and practical applications in formal verification and mathematics. Through detailed comparisons with other proof assistants and extensive case studies, we demonstrate Lean 4's unique advantages in proof automation, performance, and usability. The paper also explores recent developments in its ecosystem, including libraries, tools, and educational applications, providing insights into its growing impact on formal methods and mathematical formalization.
32
+ Subjects:
33
+ Logic in Computer Science (cs.LO)
34
+ ; Programming Languages (cs.PL)
35
+ Cite as:
36
+ arXiv:2501.18639
37
+ [cs.LO]
38
+ (or
39
+ arXiv:2501.18639v1
40
+ [cs.LO]
41
+ for this version)
42
+ https://doi.org/10.48550/arXiv.2501.18639
43
+ Focus to learn more
44
+ arXiv-issued DOI via DataCite
45
+ Submission history
46
+ From: Xichen Tang [
47
+ view email
48
+ ]
49
+ [v1]
50
+ Tue, 28 Jan 2025 17:15:54 UTC (2,729 KB)
51
+ Full-text links:
52
+ Access Paper:
53
+ View a PDF of the paper titled A Comprehensive Survey of the Lean 4 Theorem Prover: Architecture, Applications, and Advances, by Xichen Tang
54
+ View PDF
55
+ view license
56
+ Current browse context:
57
+ cs.LO
58
+ < prev
59
+ |
60
+ next >
61
+ new
62
+ |
63
+ recent
64
+ |
65
+ 2025-01
66
+ Change to browse by:
67
+ cs
68
+ cs.PL
69
+ References & Citations
70
+ NASA ADS
71
+ Google Scholar
72
+ Semantic Scholar
73
+ export BibTeX citation
74
+ Loading...
75
+ BibTeX formatted citation
76
+ ×
77
+ loading...
78
+ Data provided by:
79
+ Bookmark
80
+ Bibliographic Tools
81
+ Bibliographic and Citation Tools
82
+ Bibliographic Explorer Toggle
83
+ Bibliographic Explorer
84
+ (
85
+ What is the Explorer?
86
+ )
87
+ Connected Papers Toggle
88
+ Connected Papers
89
+ (
90
+ What is Connected Papers?
91
+ )
92
+ Litmaps Toggle
93
+ Litmaps
94
+ (
95
+ What is Litmaps?
96
+ )
97
+ scite.ai Toggle
98
+ scite Smart Citations
99
+ (
100
+ What are Smart Citations?
101
+ )
102
+ Code, Data, Media
103
+ Code, Data and Media Associated with this Article
104
+ alphaXiv Toggle
105
+ alphaXiv
106
+ (
107
+ What is alphaXiv?
108
+ )
109
+ Links to Code Toggle
110
+ CatalyzeX Code Finder for Papers
111
+ (
112
+ What is CatalyzeX?
113
+ )
114
+ DagsHub Toggle
115
+ DagsHub
116
+ (
117
+ What is DagsHub?
118
+ )
119
+ GotitPub Toggle
120
+ Gotit.pub
121
+ (
122
+ What is GotitPub?
123
+ )
124
+ Huggingface Toggle
125
+ Hugging Face
126
+ (
127
+ What is Huggingface?
128
+ )
129
+ ScienceCast Toggle
130
+ ScienceCast
131
+ (
132
+ What is ScienceCast?
133
+ )
134
+ Demos
135
+ Demos
136
+ Replicate Toggle
137
+ Replicate
138
+ (
139
+ What is Replicate?
140
+ )
141
+ Spaces Toggle
142
+ Hugging Face Spaces
143
+ (
144
+ What is Spaces?
145
+ )
146
+ Spaces Toggle
147
+ TXYZ.AI
148
+ (
149
+ What is TXYZ.AI?
150
+ )
151
+ Related Papers
152
+ Recommenders and Search Tools
153
+ Link to Influence Flower
154
+ Influence Flower
155
+ (
156
+ What are Influence Flowers?
157
+ )
158
+ Core recommender toggle
159
+ CORE Recommender
160
+ (
161
+ What is CORE?
162
+ )
163
+ Author
164
+ Venue
165
+ Institution
166
+ Topic
167
+ About arXivLabs
168
+ arXivLabs: experimental projects with community collaborators
169
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
170
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
171
+ Have an idea for a project that will add value for arXiv's community?
172
+ Learn more about arXivLabs
173
+ .
174
+ Which authors of this paper are endorsers?
175
+ |
176
+ Disable MathJax
177
+ (
178
+ What is MathJax?
179
+ )
research/notes/250202047-amasquad-a-benchmark-for-amharic-extractive-question-answering.md ADDED
@@ -0,0 +1,183 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: '[2502.02047] AmaSQuAD: A Benchmark for Amharic Extractive Question Answering'
3
+ id: 250202047-amasquad-a-benchmark-for-amharic-extractive-question-answering
4
+ tags:
5
+ - deepread
6
+ created: '2026-06-10T00:24:12.341628Z'
7
+ source: https://arxiv.org/abs/2502.02047
8
+ source_domain: arxiv.org
9
+ fetched_at: '2026-06-10T00:24:12.341471Z'
10
+ fetch_provider: builtin
11
+ status: draft
12
+ type: note
13
+ tier: institutional
14
+ content_type: paper
15
+ deprecated: false
16
+ ---
17
+
18
+ [2502.02047] AmaSQuAD: A Benchmark for Amharic Extractive Question Answering
19
+ Computer Science > Computation and Language
20
+ arXiv:2502.02047
21
+ (cs)
22
+ [Submitted on 4 Feb 2025]
23
+ Title:
24
+ AmaSQuAD: A Benchmark for Amharic Extractive Question Answering
25
+ Authors:
26
+ Nebiyou Daniel Hailemariam
27
+ ,
28
+ Blessed Guda
29
+ ,
30
+ Tsegazeab Tefferi
31
+ View a PDF of the paper titled AmaSQuAD: A Benchmark for Amharic Extractive Question Answering, by Nebiyou Daniel Hailemariam and 2 other authors
32
+ View PDF
33
+ HTML (experimental)
34
+ Abstract:
35
+ This research presents a novel framework for translating extractive question-answering datasets into low-resource languages, as demonstrated by the creation of the AmaSQuAD dataset, a translation of SQuAD 2.0 into Amharic. The methodology addresses challenges related to misalignment between translated questions and answers, as well as the presence of multiple answer instances in the translated context. For this purpose, we used cosine similarity utilizing embeddings from a fine-tuned BERT-based model for Amharic and Longest Common Subsequence (LCS). Additionally, we fine-tune the XLM-R model on the AmaSQuAD synthetic dataset for Amharic Question-Answering. The results show an improvement in baseline performance, with the fine-tuned model achieving an increase in the F1 score from 36.55% to 44.41% and 50.01% to 57.5% on the AmaSQuAD development dataset. Moreover, the model demonstrates improvement on the human-curated AmQA dataset, increasing the F1 score from 67.80% to 68.80% and the exact match score from 52.50% to 52.66%.The AmaSQuAD dataset is publicly available Datasets
36
+ Subjects:
37
+ Computation and Language (cs.CL)
38
+ Cite as:
39
+ arXiv:2502.02047
40
+ [cs.CL]
41
+ (or
42
+ arXiv:2502.02047v1
43
+ [cs.CL]
44
+ for this version)
45
+ https://doi.org/10.48550/arXiv.2502.02047
46
+ Focus to learn more
47
+ arXiv-issued DOI via DataCite
48
+ Submission history
49
+ From: Blessed Guda [
50
+ view email
51
+ ]
52
+ [v1]
53
+ Tue, 4 Feb 2025 06:27:39 UTC (778 KB)
54
+ Full-text links:
55
+ Access Paper:
56
+ View a PDF of the paper titled AmaSQuAD: A Benchmark for Amharic Extractive Question Answering, by Nebiyou Daniel Hailemariam and 2 other authors
57
+ View PDF
58
+ HTML (experimental)
59
+ TeX Source
60
+ view license
61
+ Current browse context:
62
+ cs.CL
63
+ < prev
64
+ |
65
+ next >
66
+ new
67
+ |
68
+ recent
69
+ |
70
+ 2025-02
71
+ Change to browse by:
72
+ cs
73
+ References & Citations
74
+ NASA ADS
75
+ Google Scholar
76
+ Semantic Scholar
77
+ export BibTeX citation
78
+ Loading...
79
+ BibTeX formatted citation
80
+ ×
81
+ loading...
82
+ Data provided by:
83
+ Bookmark
84
+ Bibliographic Tools
85
+ Bibliographic and Citation Tools
86
+ Bibliographic Explorer Toggle
87
+ Bibliographic Explorer
88
+ (
89
+ What is the Explorer?
90
+ )
91
+ Connected Papers Toggle
92
+ Connected Papers
93
+ (
94
+ What is Connected Papers?
95
+ )
96
+ Litmaps Toggle
97
+ Litmaps
98
+ (
99
+ What is Litmaps?
100
+ )
101
+ scite.ai Toggle
102
+ scite Smart Citations
103
+ (
104
+ What are Smart Citations?
105
+ )
106
+ Code, Data, Media
107
+ Code, Data and Media Associated with this Article
108
+ alphaXiv Toggle
109
+ alphaXiv
110
+ (
111
+ What is alphaXiv?
112
+ )
113
+ Links to Code Toggle
114
+ CatalyzeX Code Finder for Papers
115
+ (
116
+ What is CatalyzeX?
117
+ )
118
+ DagsHub Toggle
119
+ DagsHub
120
+ (
121
+ What is DagsHub?
122
+ )
123
+ GotitPub Toggle
124
+ Gotit.pub
125
+ (
126
+ What is GotitPub?
127
+ )
128
+ Huggingface Toggle
129
+ Hugging Face
130
+ (
131
+ What is Huggingface?
132
+ )
133
+ ScienceCast Toggle
134
+ ScienceCast
135
+ (
136
+ What is ScienceCast?
137
+ )
138
+ Demos
139
+ Demos
140
+ Replicate Toggle
141
+ Replicate
142
+ (
143
+ What is Replicate?
144
+ )
145
+ Spaces Toggle
146
+ Hugging Face Spaces
147
+ (
148
+ What is Spaces?
149
+ )
150
+ Spaces Toggle
151
+ TXYZ.AI
152
+ (
153
+ What is TXYZ.AI?
154
+ )
155
+ Related Papers
156
+ Recommenders and Search Tools
157
+ Link to Influence Flower
158
+ Influence Flower
159
+ (
160
+ What are Influence Flowers?
161
+ )
162
+ Core recommender toggle
163
+ CORE Recommender
164
+ (
165
+ What is CORE?
166
+ )
167
+ Author
168
+ Venue
169
+ Institution
170
+ Topic
171
+ About arXivLabs
172
+ arXivLabs: experimental projects with community collaborators
173
+ arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
174
+ Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
175
+ Have an idea for a project that will add value for arXiv's community?
176
+ Learn more about arXivLabs
177
+ .
178
+ Which authors of this paper are endorsers?
179
+ |
180
+ Disable MathJax
181
+ (
182
+ What is MathJax?
183
+ )