Instructions to use Codeseys/composer-replication-framework with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Codeseys/composer-replication-framework with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Codeseys/composer-replication-framework", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Wave 21: Stage-0 dataset pipeline — swesmith engine, rollout harness, gates, contract
Browse filesImplements research/deepread/13-synthesis-architecture.md (ADR-016): the
"point at a repo -> dataset" pipeline the critical review architected, closing
every verified P0 from research/deepread/12-verified-findings.md.
New modules (+96 tests, full suite 511 passed / 66 skipped):
- datagen/repo_gate.py: SPDX-ish license detection -> 3 tiers (fail-closed) +
benchmark DECONTAMINATION vs the SWE-bench eval-repo list (closes V3 — zero
decontamination existed anywhere).
- datagen/swesmith_adapter.py + [swesmith] extra: SWE-smith as the synthesis
engine (closes V4 buy-vs-build). Handles the patch-semantics INVERSION
(SWE-smith's patch INTRODUCES the bug; golden_diff = reverse_unified_diff)
+ strategy provenance sidecar + difficulty priors.
- datagen/trajectory.py: canonical trajectory IR (closes D-11). ToolCall
canonical_form = v1 divergence-gate action algebra (replaces the whitespace
stub, D-3); to_policy_row = THE policy-visible serializer, sentinel-tested
to never leak golden_diff/deleted_symbols (D-8).
- datagen/rollout_harness.py: collect_trajectory agent loop over
FeatureDeletionEnv (closes V2 — the SFT corpus finally has a producer; its
env-grounded episodes are also the tree's seeds, fixing D-1) + typed
admission routing (sft/dpo-candidate/quarantine).
- pipeline/{s3_contract,dedup,build_corpus}.py: ONE reconciled dataset layout
(supersedes F1/F2's divergent contracts, V8) with restricted tasks_full
prefix + golden_diff sha256; stable-hash MinHash dedup incl.
cross-generation signatures (D-12); local write-once stage-driver with
holdout-first split + budget ceiling (D-9/D-21).
Fidelity corrections (verified findings, applied to code+docs):
- V1: SDPO "mathematically the same" claim corrected (opsd.py, mapping doc) —
Cursor cites SDPO as background; ours is a third blog-inspired design.
- V5: fabricated numbers struck/tagged (69.3%, "24 generators", "85% compute").
- V7: Streaming DiLoCo citation fixed (Douillard 2501.18512; Kale 2502.12996).
- V11: teacher_replay cost docstring relabeled (synthetic-trace basis).
- V13: kl_in_reward verl wording (default/recommended, not only).
- ADR-016 records the decision + gates; harvested repo_gate from the one
surviving worktree builder (3 of 4 stalled on API capacity; rebuilt directly).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- composer_replication/datagen/repo_gate.py +361 -0
- composer_replication/datagen/rollout_harness.py +214 -0
- composer_replication/datagen/swesmith_adapter.py +269 -0
- composer_replication/datagen/tests/test_repo_gate.py +419 -0
- composer_replication/datagen/tests/test_rollout_harness.py +103 -0
- composer_replication/datagen/tests/test_swesmith_adapter.py +165 -0
- composer_replication/datagen/tests/test_trajectory.py +127 -0
- composer_replication/datagen/trajectory.py +203 -0
- composer_replication/diloco/__init__.py +6 -3
- composer_replication/opsd.py +11 -5
- composer_replication/pipeline/__init__.py +38 -0
- composer_replication/pipeline/build_corpus.py +137 -0
- composer_replication/pipeline/dedup.py +138 -0
- composer_replication/pipeline/s3_contract.py +287 -0
- composer_replication/pipeline/tests/__init__.py +0 -0
- composer_replication/pipeline/tests/test_pipeline.py +223 -0
- composer_replication/teacher_replay.py +6 -2
- composer_replication/trainer/kl_in_reward.py +3 -1
- docs/COMPOSER_RECIPE_MAPPING.md +3 -3
- docs/adrs/ADR-016-stage0-dataset-pipeline.md +119 -0
- pyproject.toml +8 -0
- research/01-composer-2.5.md +3 -3
- research/06-feature-deletion-datagen.md +1 -1
- research/09-composer-blog-delta-2026.md +1 -1
- research/notes/230406767-raft-reward-ranked-finetuning-for-generative-foundation-model-alignmen.md +224 -0
- research/notes/230510601-tree-of-thoughts-deliberate-problem-solving-with-large-language-models-2.md +2735 -0
- research/notes/230510601-tree-of-thoughts-deliberate-problem-solving-with-large-language-models.md +213 -0
- research/notes/231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la-2.md +4095 -0
- research/notes/231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la-3.md +4095 -0
- research/notes/231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la.md +202 -0
- research/notes/231006770-swe-bench-can-language-models-resolve-real-world-github-issues.md +203 -0
- research/notes/231108105-diloco-distributed-low-communication-training-of-language-models.md +208 -0
- research/notes/231108516-llms-cannot-find-reasoning-errors-but-can-correct-them-given-the-error.md +204 -0
- research/notes/231209152-evaluating-augmented-reality-communication-how-can-we-teach-procedural.md +203 -0
- research/notes/240201817-llms-cant-plan-but-can-help-planning-in-llm-modulo-frameworks.md +203 -0
- research/notes/240203300-deepseekmath-pushing-the-limits-of-mathematical-reasoning-in-open-lang.md +214 -0
- research/notes/240411018-many-shot-in-context-learning.md +228 -0
- research/notes/240612543-phase-controlled-heat-modulation-with-aharonov-bohm-interferometers.md +197 -0
- research/notes/240806195-mutual-reasoning-makes-smaller-llms-stronger-problem-solvers-2.md +2384 -0
- research/notes/240806195-mutual-reasoning-makes-smaller-llms-stronger-problem-solvers.md +191 -0
- research/notes/241020285-swe-search-enhancing-software-agents-with-monte-carlo-tree-search-and-2.md +2144 -0
- research/notes/241221139-training-software-engineering-agents-and-verifiers-with-swe-gym.md +200 -0
- research/notes/250104519-rstar-math-small-llms-can-master-math-reasoning-with-self-evolved-deep.md +196 -0
- research/notes/250104519-sysname-small-llms-can-master-math-reasoning-with-self-evolved-deep-th.md +3557 -0
- research/notes/250109136-agentic-retrieval-augmented-generation-a-survey-on-agentic-rag.md +203 -0
- research/notes/250109891-evolving-deeper-llm-thinking.md +196 -0
- research/notes/250112599-kimi-k15-scaling-reinforcement-learning-with-llms.md +386 -0
- research/notes/250118512-streaming-diloco-with-overlapping-communication-towards-a-distributed.md +206 -0
- research/notes/250118639-a-comprehensive-survey-of-the-lean-4-theorem-prover-architecture-appli.md +179 -0
- research/notes/250202047-amasquad-a-benchmark-for-amharic-extractive-question-answering.md +183 -0
|
@@ -0,0 +1,361 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""repo_gate.py — Stage-0 ingest gate: license tiers + benchmark decontamination.
|
| 2 |
+
|
| 3 |
+
Architecture step 1 of the dataset pipeline (research/deepread/
|
| 4 |
+
13-synthesis-architecture.md Part B). Closes two verified findings:
|
| 5 |
+
|
| 6 |
+
* V3 / D-5 — ZERO benchmark decontamination existed anywhere in code or
|
| 7 |
+
designs, while the pipeline trains on SWE-bench-family substrates and is
|
| 8 |
+
scored on SWE-bench Verified. ``is_eval_contaminated`` is the hard wall:
|
| 9 |
+
a repo on the eval list is NEVER admitted, regardless of license.
|
| 10 |
+
* V9 / D-13 — the only license filter was a lowercase substring match on a
|
| 11 |
+
task field (``substrates.py`` ``is_redistributable``), with no SPDX
|
| 12 |
+
detection at the repo-ingest path and no trainable-vs-redistributable
|
| 13 |
+
split. ``detect_license`` + ``license_tier`` replace the boolean with a
|
| 14 |
+
three-tier verdict.
|
| 15 |
+
|
| 16 |
+
Why tiers, not a boolean (D-13): weak-copyleft repos (MPL/LGPL) are fine to
|
| 17 |
+
*train on* but we must not *redistribute* derivative diffs from them — a
|
| 18 |
+
boolean "redistributable?" gate either over-excludes them or leaks them into
|
| 19 |
+
published corpora. The tier travels with the verdict so downstream corpus
|
| 20 |
+
steps (step 6) can route TRAINABLE_ONLY rows away from any published split.
|
| 21 |
+
|
| 22 |
+
Why title-anchored matching for the GNU family: GPL-3.0 §13 mentions the
|
| 23 |
+
"GNU Affero General Public License" by name and AGPL-3.0 §13 mentions the
|
| 24 |
+
"GNU General Public License" — naive full-body substring matching
|
| 25 |
+
misclassifies one as the other. We therefore classify the GNU/MPL/Apache
|
| 26 |
+
family from the document HEADER (first ~400 normalized chars, where the
|
| 27 |
+
license title lives) and only use full-body phrases for the short permissive
|
| 28 |
+
licenses whose titles are not distinctive (MIT/ISC/BSD/Unlicense).
|
| 29 |
+
|
| 30 |
+
Stdlib-only on purpose: the gate must run before anything heavy is installed.
|
| 31 |
+
"""
|
| 32 |
+
from __future__ import annotations
|
| 33 |
+
|
| 34 |
+
import json
|
| 35 |
+
import re
|
| 36 |
+
from dataclasses import dataclass, field
|
| 37 |
+
from enum import Enum
|
| 38 |
+
from pathlib import Path
|
| 39 |
+
|
| 40 |
+
# ---------------------------------------------------------------------------
|
| 41 |
+
# License detection (V9 / D-13)
|
| 42 |
+
# ---------------------------------------------------------------------------
|
| 43 |
+
|
| 44 |
+
#: License files checked in order; first match wins (case-insensitive on name).
|
| 45 |
+
_LICENSE_FILENAMES: tuple[str, ...] = ("LICENSE", "LICENSE.txt", "LICENSE.md", "COPYING")
|
| 46 |
+
|
| 47 |
+
#: Trove classifier / PEP 639 expression fragments → SPDX id. Secondary signal
|
| 48 |
+
#: only — the classifier cannot distinguish BSD-2 from BSD-3, so it maps to
|
| 49 |
+
#: BSD-3-Clause (the common case) and the LICENSE file is preferred when present.
|
| 50 |
+
_CLASSIFIER_MAP: tuple[tuple[str, str], ...] = (
|
| 51 |
+
("gnu affero general public license", "AGPL-3.0"),
|
| 52 |
+
("gnu lesser general public license v3", "LGPL-3.0"),
|
| 53 |
+
("gnu lesser general public license v2.1", "LGPL-2.1"),
|
| 54 |
+
("gnu lesser general public license", "LGPL-3.0"),
|
| 55 |
+
("gnu general public license v3", "GPL-3.0"),
|
| 56 |
+
("gnu general public license v2", "GPL-2.0"),
|
| 57 |
+
("mozilla public license 2.0", "MPL-2.0"),
|
| 58 |
+
("apache software license", "Apache-2.0"),
|
| 59 |
+
("mit license", "MIT"),
|
| 60 |
+
("bsd license", "BSD-3-Clause"),
|
| 61 |
+
("isc license", "ISC"),
|
| 62 |
+
("the unlicense", "Unlicense"),
|
| 63 |
+
)
|
| 64 |
+
|
| 65 |
+
#: Bare SPDX ids accepted from PEP 639 ``license = "<expr>"`` in pyproject.
|
| 66 |
+
_SPDX_IDS: frozenset[str] = frozenset(
|
| 67 |
+
{
|
| 68 |
+
"MIT", "Apache-2.0", "BSD-2-Clause", "BSD-3-Clause", "ISC",
|
| 69 |
+
"GPL-2.0", "GPL-3.0", "AGPL-3.0", "LGPL-2.1", "LGPL-3.0",
|
| 70 |
+
"MPL-2.0", "Unlicense",
|
| 71 |
+
}
|
| 72 |
+
)
|
| 73 |
+
_SPDX_LOOKUP: dict[str, str] = {s.lower(): s for s in _SPDX_IDS}
|
| 74 |
+
# Common -only/-or-later suffixed forms normalize to the base id we tier on.
|
| 75 |
+
for _base in ("GPL-2.0", "GPL-3.0", "AGPL-3.0", "LGPL-2.1", "LGPL-3.0"):
|
| 76 |
+
_SPDX_LOOKUP[f"{_base.lower()}-only"] = _base
|
| 77 |
+
_SPDX_LOOKUP[f"{_base.lower()}-or-later"] = _base
|
| 78 |
+
|
| 79 |
+
|
| 80 |
+
@dataclass(frozen=True)
|
| 81 |
+
class LicenseInfo:
|
| 82 |
+
"""Outcome of license detection: SPDX-ish id + which signal decided it."""
|
| 83 |
+
|
| 84 |
+
spdx_id: str # one of _SPDX_IDS or "unknown"
|
| 85 |
+
signal: str # "license_file" | "classifier" | "none"
|
| 86 |
+
source: str = "" # filename that supplied the winning signal
|
| 87 |
+
|
| 88 |
+
|
| 89 |
+
def _normalize_text(text: str) -> str:
|
| 90 |
+
return re.sub(r"\s+", " ", text).strip().lower()
|
| 91 |
+
|
| 92 |
+
|
| 93 |
+
#: Title strings for the families that cross-cite each other. NOTE: "gnu
|
| 94 |
+
#: affero general public license" does NOT contain "gnu general public
|
| 95 |
+
#: license" as a substring ("affero" splits it), so the titles are disjoint.
|
| 96 |
+
_HEADER_TITLES: tuple[tuple[str, str], ...] = (
|
| 97 |
+
("gnu affero general public license", "agpl"),
|
| 98 |
+
("gnu lesser general public license", "lgpl"),
|
| 99 |
+
("gnu general public license", "gpl"),
|
| 100 |
+
("mozilla public license", "mpl"),
|
| 101 |
+
("apache license", "apache"),
|
| 102 |
+
)
|
| 103 |
+
|
| 104 |
+
|
| 105 |
+
def _classify_header(header: str) -> str | None:
|
| 106 |
+
"""Title-anchored families (GNU/MPL/Apache). The EARLIEST-occurring title
|
| 107 |
+
wins, because a license document's own title always precedes any
|
| 108 |
+
cross-citation — GPL-3 §13 names the AGPL and AGPL-3 §13 names the GPL,
|
| 109 |
+
so mere presence-matching misclassifies one as the other (the V9 trap)."""
|
| 110 |
+
hits = [(idx, family) for title, family in _HEADER_TITLES if (idx := header.find(title)) >= 0]
|
| 111 |
+
if not hits:
|
| 112 |
+
return None
|
| 113 |
+
family = min(hits)[1]
|
| 114 |
+
if family == "agpl":
|
| 115 |
+
return "AGPL-3.0"
|
| 116 |
+
if family == "lgpl":
|
| 117 |
+
return "LGPL-2.1" if "version 2.1" in header else "LGPL-3.0"
|
| 118 |
+
if family == "gpl":
|
| 119 |
+
return "GPL-2.0" if "version 2" in header and "version 3" not in header else "GPL-3.0"
|
| 120 |
+
if family == "mpl":
|
| 121 |
+
return "MPL-2.0" if "2.0" in header else None
|
| 122 |
+
return "Apache-2.0" if "version 2.0" in header else None
|
| 123 |
+
|
| 124 |
+
|
| 125 |
+
def _classify_body(body: str) -> str | None:
|
| 126 |
+
"""Distinctive-phrase matching for the short permissive licenses. Order
|
| 127 |
+
matters: ISC's grant ("permission to use, copy, modify") is checked via
|
| 128 |
+
its unique "and/or distribute … with or without fee" wording so it can't
|
| 129 |
+
be shadowed by MIT's "permission is hereby granted" phrase."""
|
| 130 |
+
if "free and unencumbered software released into the public domain" in body:
|
| 131 |
+
return "Unlicense"
|
| 132 |
+
# Apache boilerplate notice files ("Licensed under the Apache License,
|
| 133 |
+
# Version 2.0") carry the title mid-body, not in a header — the tricky
|
| 134 |
+
# Apache-vs-MIT case: both say "permission"/"license", only Apache names
|
| 135 |
+
# itself with a version.
|
| 136 |
+
if "apache license" in body and "version 2.0" in body:
|
| 137 |
+
return "Apache-2.0"
|
| 138 |
+
if "permission is hereby granted, free of charge, to any person obtaining a copy" in body:
|
| 139 |
+
return "MIT"
|
| 140 |
+
if "with or without fee" in body and "permission to use, copy, modify" in body:
|
| 141 |
+
return "ISC"
|
| 142 |
+
if "redistribution and use in source and binary forms" in body:
|
| 143 |
+
# The third clause ("Neither the name of …") is what separates 3- from 2-.
|
| 144 |
+
return "BSD-3-Clause" if "neither the name of" in body else "BSD-2-Clause"
|
| 145 |
+
return None
|
| 146 |
+
|
| 147 |
+
|
| 148 |
+
def _classify_license_text(text: str) -> str:
|
| 149 |
+
norm = _normalize_text(text)
|
| 150 |
+
return _classify_header(norm[:400]) or _classify_body(norm) or "unknown"
|
| 151 |
+
|
| 152 |
+
|
| 153 |
+
def _classifier_signal(repo_root: Path) -> tuple[str, str] | None:
|
| 154 |
+
"""Secondary signal: trove classifiers / PEP 639 license expression in
|
| 155 |
+
pyproject.toml or setup.py. Regex-scan, not a TOML parse — the gate must
|
| 156 |
+
not depend on packaging libs and classifiers are line-shaped in practice."""
|
| 157 |
+
for name in ("pyproject.toml", "setup.py"):
|
| 158 |
+
path = repo_root / name
|
| 159 |
+
if not path.is_file():
|
| 160 |
+
continue
|
| 161 |
+
try:
|
| 162 |
+
text = path.read_text(encoding="utf-8", errors="replace")
|
| 163 |
+
except OSError:
|
| 164 |
+
continue
|
| 165 |
+
# PEP 639: license = "Apache-2.0" (pyproject only, but harmless on setup.py).
|
| 166 |
+
m = re.search(r'license\s*=\s*["\']([A-Za-z0-9.+-]+)["\']', text)
|
| 167 |
+
if m and m.group(1).lower() in _SPDX_LOOKUP:
|
| 168 |
+
return _SPDX_LOOKUP[m.group(1).lower()], name
|
| 169 |
+
low = _normalize_text(text)
|
| 170 |
+
for fragment, spdx in _CLASSIFIER_MAP:
|
| 171 |
+
if f"license :: osi approved :: {fragment}" in low or (
|
| 172 |
+
"license ::" in low and fragment in low
|
| 173 |
+
):
|
| 174 |
+
return spdx, name
|
| 175 |
+
return None
|
| 176 |
+
|
| 177 |
+
|
| 178 |
+
def detect_license(repo_root: Path) -> LicenseInfo:
|
| 179 |
+
"""Detect the repo license. LICENSE-file text is the primary signal;
|
| 180 |
+
packaging classifiers are secondary (used only when the file is absent or
|
| 181 |
+
unclassifiable). The winning signal is recorded so corpus manifests can
|
| 182 |
+
show provenance for the tier decision (V9 closure must be auditable)."""
|
| 183 |
+
for name in _LICENSE_FILENAMES:
|
| 184 |
+
path = repo_root / name
|
| 185 |
+
if not path.is_file():
|
| 186 |
+
# Case-insensitive fallback (e.g. "License.md", "COPYING.txt" not
|
| 187 |
+
# matched here on purpose — only exact-name case variants).
|
| 188 |
+
matches = [p for p in repo_root.glob("*") if p.is_file() and p.name.lower() == name.lower()]
|
| 189 |
+
path = matches[0] if matches else path
|
| 190 |
+
if path.is_file():
|
| 191 |
+
try:
|
| 192 |
+
text = path.read_text(encoding="utf-8", errors="replace")
|
| 193 |
+
except OSError:
|
| 194 |
+
continue
|
| 195 |
+
spdx = _classify_license_text(text)
|
| 196 |
+
if spdx != "unknown":
|
| 197 |
+
return LicenseInfo(spdx_id=spdx, signal="license_file", source=path.name)
|
| 198 |
+
# File exists but unclassifiable → let the classifier signal try
|
| 199 |
+
# before giving up; remember we saw a file for the "none" case.
|
| 200 |
+
fallback = _classifier_signal(repo_root)
|
| 201 |
+
if fallback is not None:
|
| 202 |
+
return LicenseInfo(spdx_id=fallback[0], signal="classifier", source=fallback[1])
|
| 203 |
+
return LicenseInfo(spdx_id="unknown", signal="license_file", source=path.name)
|
| 204 |
+
fallback = _classifier_signal(repo_root)
|
| 205 |
+
if fallback is not None:
|
| 206 |
+
return LicenseInfo(spdx_id=fallback[0], signal="classifier", source=fallback[1])
|
| 207 |
+
return LicenseInfo(spdx_id="unknown", signal="none")
|
| 208 |
+
|
| 209 |
+
|
| 210 |
+
# ---------------------------------------------------------------------------
|
| 211 |
+
# License tiers (D-13: tiers, not a boolean)
|
| 212 |
+
# ---------------------------------------------------------------------------
|
| 213 |
+
|
| 214 |
+
|
| 215 |
+
class Tier(Enum):
|
| 216 |
+
"""Three-way license verdict. TRAINABLE_ONLY exists because weak copyleft
|
| 217 |
+
(MPL/LGPL) permits training but redistribution of derivative diffs would
|
| 218 |
+
trigger copyleft obligations — collapsing this to a boolean either loses
|
| 219 |
+
training data or leaks copyleft material into published corpora (D-13)."""
|
| 220 |
+
|
| 221 |
+
REDISTRIBUTABLE = "redistributable"
|
| 222 |
+
TRAINABLE_ONLY = "trainable_only"
|
| 223 |
+
EXCLUDED = "excluded"
|
| 224 |
+
|
| 225 |
+
|
| 226 |
+
_TIER_BY_SPDX: dict[str, Tier] = {
|
| 227 |
+
"MIT": Tier.REDISTRIBUTABLE,
|
| 228 |
+
"Apache-2.0": Tier.REDISTRIBUTABLE,
|
| 229 |
+
"BSD-2-Clause": Tier.REDISTRIBUTABLE,
|
| 230 |
+
"BSD-3-Clause": Tier.REDISTRIBUTABLE,
|
| 231 |
+
"ISC": Tier.REDISTRIBUTABLE,
|
| 232 |
+
"Unlicense": Tier.REDISTRIBUTABLE,
|
| 233 |
+
"MPL-2.0": Tier.TRAINABLE_ONLY,
|
| 234 |
+
"LGPL-2.1": Tier.TRAINABLE_ONLY,
|
| 235 |
+
"LGPL-3.0": Tier.TRAINABLE_ONLY,
|
| 236 |
+
# GPL/AGPL and unknown are EXCLUDED: strong copyleft would bind the model
|
| 237 |
+
# outputs' redistribution story, and "unknown" defaults closed (V9).
|
| 238 |
+
"GPL-2.0": Tier.EXCLUDED,
|
| 239 |
+
"GPL-3.0": Tier.EXCLUDED,
|
| 240 |
+
"AGPL-3.0": Tier.EXCLUDED,
|
| 241 |
+
}
|
| 242 |
+
|
| 243 |
+
|
| 244 |
+
def license_tier(info: LicenseInfo) -> Tier:
|
| 245 |
+
"""Map detected license → tier. Anything unrecognized is EXCLUDED — the
|
| 246 |
+
gate fails closed, never open (V9: the old substring filter failed open)."""
|
| 247 |
+
return _TIER_BY_SPDX.get(info.spdx_id, Tier.EXCLUDED)
|
| 248 |
+
|
| 249 |
+
|
| 250 |
+
# ---------------------------------------------------------------------------
|
| 251 |
+
# Benchmark decontamination (V3 / D-5)
|
| 252 |
+
# ---------------------------------------------------------------------------
|
| 253 |
+
|
| 254 |
+
#: The canonical 12 SWE-bench test repos (SWE-bench / -Lite / -Verified /
|
| 255 |
+
#: -Multimodal all draw eval instances from these). Training on ANY of them
|
| 256 |
+
#: contaminates every SWE-bench-family score we report (V3). Lowercase
|
| 257 |
+
#: "org/repo" form. Extend via a JSON file (list of "org/repo" strings)
|
| 258 |
+
#: passed to is_eval_contaminated(extra_list=...) — e.g. SWE-Gym eval splits.
|
| 259 |
+
DECONTAMINATION_LIST: frozenset[str] = frozenset(
|
| 260 |
+
{
|
| 261 |
+
"astropy/astropy",
|
| 262 |
+
"django/django",
|
| 263 |
+
"matplotlib/matplotlib",
|
| 264 |
+
"mwaskom/seaborn",
|
| 265 |
+
"pallets/flask",
|
| 266 |
+
"psf/requests",
|
| 267 |
+
"pydata/xarray",
|
| 268 |
+
"pylint-dev/pylint",
|
| 269 |
+
"pytest-dev/pytest",
|
| 270 |
+
"scikit-learn/scikit-learn",
|
| 271 |
+
"sphinx-doc/sphinx",
|
| 272 |
+
"sympy/sympy",
|
| 273 |
+
}
|
| 274 |
+
)
|
| 275 |
+
|
| 276 |
+
|
| 277 |
+
def load_decontamination_list(path: Path) -> frozenset[str]:
|
| 278 |
+
"""Load an extension list from a JSON file: ``["org/repo", ...]``. This is
|
| 279 |
+
THE documented mechanism for adding eval repos (new SWE-bench releases,
|
| 280 |
+
SWE-Gym eval splits) without editing code."""
|
| 281 |
+
entries = json.loads(path.read_text(encoding="utf-8"))
|
| 282 |
+
if not isinstance(entries, list):
|
| 283 |
+
raise ValueError(f"{path}: decontamination JSON must be a list of 'org/repo' strings")
|
| 284 |
+
return frozenset(normalize_repo(str(e)) for e in entries)
|
| 285 |
+
|
| 286 |
+
|
| 287 |
+
def normalize_repo(repo: str) -> str:
|
| 288 |
+
"""Reduce any repo spelling — full https/ssh GitHub URL, trailing ``.git``,
|
| 289 |
+
mixed case — to lowercase ``org/repo``. Decontamination must hit no matter
|
| 290 |
+
how the driver spells the repo (V3: a miss here is silent contamination)."""
|
| 291 |
+
r = repo.strip().lower()
|
| 292 |
+
r = re.sub(r"^(https?://|git@)", "", r)
|
| 293 |
+
r = re.sub(r"^[^/]*github\.com[:/]", "", r)
|
| 294 |
+
r = r.rstrip("/")
|
| 295 |
+
r = r.removesuffix(".git")
|
| 296 |
+
parts = [p for p in r.split("/") if p]
|
| 297 |
+
return "/".join(parts[:2]) if len(parts) >= 2 else r
|
| 298 |
+
|
| 299 |
+
|
| 300 |
+
def is_eval_contaminated(repo: str, extra_list: frozenset[str] | None = None) -> bool:
|
| 301 |
+
"""True if ``repo`` is in the SWE-bench-family eval set (or the caller's
|
| 302 |
+
extension list). Case-insensitive; accepts URLs and bare org/repo."""
|
| 303 |
+
key = normalize_repo(repo)
|
| 304 |
+
return key in DECONTAMINATION_LIST or (extra_list is not None and key in extra_list)
|
| 305 |
+
|
| 306 |
+
|
| 307 |
+
# ---------------------------------------------------------------------------
|
| 308 |
+
# The gate verdict — single entry point for the pipeline driver
|
| 309 |
+
# ---------------------------------------------------------------------------
|
| 310 |
+
|
| 311 |
+
|
| 312 |
+
@dataclass
|
| 313 |
+
class GateVerdict:
|
| 314 |
+
"""Everything the driver needs to admit/reject a repo, with reasons kept
|
| 315 |
+
for the run manifest (step 6's lineage record)."""
|
| 316 |
+
|
| 317 |
+
repo: str
|
| 318 |
+
license_info: LicenseInfo
|
| 319 |
+
tier: Tier
|
| 320 |
+
contaminated: bool
|
| 321 |
+
admitted: bool
|
| 322 |
+
reasons: list[str] = field(default_factory=list)
|
| 323 |
+
|
| 324 |
+
|
| 325 |
+
def gate_repo(repo: str, repo_root: Path | None, extra_decontamination: frozenset[str] | None = None) -> GateVerdict:
|
| 326 |
+
"""Architecture step 1: the one call the pipeline driver makes per repo.
|
| 327 |
+
|
| 328 |
+
Hard rules (in priority order):
|
| 329 |
+
1. Contaminated (V3) → NEVER admitted, even if the license is permissive.
|
| 330 |
+
2. Tier EXCLUDED (GPL/AGPL/unknown) → not admitted (V9: fail closed).
|
| 331 |
+
3. Tier TRAINABLE_ONLY → admitted, with the do-not-redistribute
|
| 332 |
+
constraint recorded as a reason so step 6 can route the rows.
|
| 333 |
+
"""
|
| 334 |
+
contaminated = is_eval_contaminated(repo, extra_decontamination)
|
| 335 |
+
info = detect_license(repo_root) if repo_root is not None else LicenseInfo("unknown", "none")
|
| 336 |
+
tier = license_tier(info)
|
| 337 |
+
|
| 338 |
+
reasons: list[str] = []
|
| 339 |
+
if contaminated:
|
| 340 |
+
reasons.append(
|
| 341 |
+
f"benchmark decontamination: {normalize_repo(repo)} is a SWE-bench-family eval repo (V3/D-5)"
|
| 342 |
+
)
|
| 343 |
+
if repo_root is None:
|
| 344 |
+
reasons.append("no repo_root provided: license undetectable, failing closed (V9)")
|
| 345 |
+
if tier is Tier.EXCLUDED and not contaminated:
|
| 346 |
+
reasons.append(f"license tier EXCLUDED: spdx={info.spdx_id} (signal={info.signal})")
|
| 347 |
+
if tier is Tier.TRAINABLE_ONLY:
|
| 348 |
+
reasons.append(
|
| 349 |
+
f"license tier TRAINABLE_ONLY: spdx={info.spdx_id} — usable for training, "
|
| 350 |
+
"derivative diffs must NOT be redistributed (D-13)"
|
| 351 |
+
)
|
| 352 |
+
|
| 353 |
+
admitted = (not contaminated) and tier is not Tier.EXCLUDED
|
| 354 |
+
return GateVerdict(
|
| 355 |
+
repo=repo,
|
| 356 |
+
license_info=info,
|
| 357 |
+
tier=tier,
|
| 358 |
+
contaminated=contaminated,
|
| 359 |
+
admitted=admitted,
|
| 360 |
+
reasons=reasons,
|
| 361 |
+
)
|
|
@@ -0,0 +1,214 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""rollout_harness.py — the agent loop over FeatureDeletionEnv (finding V2).
|
| 2 |
+
|
| 3 |
+
THE critical missing component the design critic identified: nothing in the
|
| 4 |
+
repo ran an agent episode against `FeatureDeletionEnv` to completion, so the
|
| 5 |
+
SFT corpus had NO producer and the tree-of-work had no env-grounded seeds.
|
| 6 |
+
`collect_trajectory` is that producer: prompt → policy.act → env.step → … →
|
| 7 |
+
submit → `_grade()`, emitting a `CanonicalTrajectory` whose steps are real
|
| 8 |
+
executed environment transitions (the seeds the tree needs, fixing the
|
| 9 |
+
seed-trace/oracle disjointness of finding D-1 as a free byproduct).
|
| 10 |
+
|
| 11 |
+
The policy is pluggable (`RolloutPolicy` protocol): a scripted fake for tests,
|
| 12 |
+
a frontier API model for expert-trajectory collection (SWE-Gym/SWE-smith both
|
| 13 |
+
validated this recipe — 491 and 5,016 expert trajectories respectively), or a
|
| 14 |
+
local model later.
|
| 15 |
+
"""
|
| 16 |
+
from __future__ import annotations
|
| 17 |
+
|
| 18 |
+
from dataclasses import dataclass
|
| 19 |
+
from typing import Protocol, runtime_checkable
|
| 20 |
+
|
| 21 |
+
from composer_replication.datagen.env import FeatureDeletionEnv, StepResult
|
| 22 |
+
from composer_replication.datagen.schema import FeatureDeletionTask
|
| 23 |
+
from composer_replication.datagen.trajectory import (
|
| 24 |
+
CanonicalTrajectory,
|
| 25 |
+
ToolCall,
|
| 26 |
+
TrajectoryStep,
|
| 27 |
+
)
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
@runtime_checkable
|
| 31 |
+
class RolloutPolicy(Protocol):
|
| 32 |
+
"""Anything that maps (observation, history) → the next action.
|
| 33 |
+
|
| 34 |
+
Returning a `ToolCall` continues the episode (translated to an env action
|
| 35 |
+
dict); returning a plain `str` is the final message — the harness submits.
|
| 36 |
+
"""
|
| 37 |
+
|
| 38 |
+
def act(self, observation: str, history: list[TrajectoryStep]) -> ToolCall | str: ...
|
| 39 |
+
|
| 40 |
+
|
| 41 |
+
@dataclass
|
| 42 |
+
class ScriptedPolicy:
|
| 43 |
+
"""Test fake: replays a fixed action list, then submits."""
|
| 44 |
+
|
| 45 |
+
actions: list[ToolCall | str]
|
| 46 |
+
_i: int = 0
|
| 47 |
+
|
| 48 |
+
def act(self, observation: str, history: list[TrajectoryStep]) -> ToolCall | str:
|
| 49 |
+
if self._i >= len(self.actions):
|
| 50 |
+
return "done" # str → submit
|
| 51 |
+
a = self.actions[self._i]
|
| 52 |
+
self._i += 1
|
| 53 |
+
return a
|
| 54 |
+
|
| 55 |
+
|
| 56 |
+
class OpenRouterPolicy:
|
| 57 |
+
"""Frontier-API policy for expert-trajectory collection (thin stub).
|
| 58 |
+
|
| 59 |
+
Mirrors `teacher_replay._call_teacher`'s payload shape (one chat call,
|
| 60 |
+
temperature 0.2). Lazy-deps on httpx so the module imports without it.
|
| 61 |
+
Deliberately minimal: real expert collection should evaluate adopting
|
| 62 |
+
mini-swe-agent/SWE-agent as the scaffold (deepread 11 finding 2) — this
|
| 63 |
+
class exists so the harness has a live-API path without a new framework.
|
| 64 |
+
"""
|
| 65 |
+
|
| 66 |
+
def __init__(self, model_slug: str, api_key: str | None = None,
|
| 67 |
+
max_tokens: int = 512) -> None:
|
| 68 |
+
try:
|
| 69 |
+
import httpx # noqa: F401, PLC0415 — lazy heavy dep
|
| 70 |
+
except ImportError as e:
|
| 71 |
+
raise ImportError(
|
| 72 |
+
"OpenRouterPolicy requires httpx (`pip install httpx` or the "
|
| 73 |
+
"[serverless] extra). For tests use ScriptedPolicy. Got: " + repr(e)
|
| 74 |
+
) from e
|
| 75 |
+
from composer_replication.teacher_replay import _load_api_key
|
| 76 |
+
self.model_slug = model_slug
|
| 77 |
+
self.api_key = api_key or _load_api_key()
|
| 78 |
+
self.max_tokens = max_tokens
|
| 79 |
+
|
| 80 |
+
def act(self, observation: str, history: list[TrajectoryStep]) -> ToolCall | str:
|
| 81 |
+
import httpx # noqa: PLC0415
|
| 82 |
+
|
| 83 |
+
from composer_replication.teacher_replay import OPENROUTER_URL
|
| 84 |
+
messages = [{"role": "user", "content": observation}]
|
| 85 |
+
r = httpx.post(
|
| 86 |
+
OPENROUTER_URL,
|
| 87 |
+
json={"model": self.model_slug, "messages": messages,
|
| 88 |
+
"max_tokens": self.max_tokens, "temperature": 0.2},
|
| 89 |
+
headers={"Authorization": f"Bearer {self.api_key}"},
|
| 90 |
+
timeout=120.0,
|
| 91 |
+
)
|
| 92 |
+
r.raise_for_status()
|
| 93 |
+
return str(r.json()["choices"][0]["message"]["content"])
|
| 94 |
+
|
| 95 |
+
|
| 96 |
+
def _to_env_action(call: ToolCall) -> dict:
|
| 97 |
+
"""ToolCall → FeatureDeletionEnv action dict.
|
| 98 |
+
|
| 99 |
+
CONVENTION (documented here, the single translation point): the env's
|
| 100 |
+
`step()` consumes ``{"type": <tool name>, **args}``; ``type=="submit"``
|
| 101 |
+
triggers grading (env.py:67). A ToolCall named "submit" therefore ends the
|
| 102 |
+
episode through the same path as a plain-text final message.
|
| 103 |
+
"""
|
| 104 |
+
return {"type": call.name, **call.args}
|
| 105 |
+
|
| 106 |
+
|
| 107 |
+
def collect_trajectory(
|
| 108 |
+
env: FeatureDeletionEnv,
|
| 109 |
+
task: FeatureDeletionTask,
|
| 110 |
+
policy: RolloutPolicy,
|
| 111 |
+
*,
|
| 112 |
+
max_turns: int = 40,
|
| 113 |
+
budget_usd: float | None = None,
|
| 114 |
+
provenance: dict | None = None,
|
| 115 |
+
) -> CanonicalTrajectory:
|
| 116 |
+
"""Run one episode and return the graded CanonicalTrajectory.
|
| 117 |
+
|
| 118 |
+
The episode ends when the policy emits a plain string (final message →
|
| 119 |
+
submit), a ToolCall named "submit", or `max_turns` is hit (the env grades
|
| 120 |
+
on its own turn limit too — we mirror it here so the harness's history
|
| 121 |
+
stays aligned with the env's accounting).
|
| 122 |
+
"""
|
| 123 |
+
obs = env.reset(task)
|
| 124 |
+
steps: list[TrajectoryStep] = []
|
| 125 |
+
final: StepResult | None = None
|
| 126 |
+
|
| 127 |
+
for _ in range(max_turns):
|
| 128 |
+
action = policy.act(obs, steps)
|
| 129 |
+
if isinstance(action, str) or action.name == "submit":
|
| 130 |
+
final = env.step({"type": "submit"})
|
| 131 |
+
steps.append(TrajectoryStep(
|
| 132 |
+
observation=obs, action=action, result=final.observation,
|
| 133 |
+
tool_error=False,
|
| 134 |
+
))
|
| 135 |
+
break
|
| 136 |
+
res = env.step(_to_env_action(action))
|
| 137 |
+
tool_error = "error" in (res.observation or "").lower()[:200]
|
| 138 |
+
steps.append(TrajectoryStep(
|
| 139 |
+
observation=obs, action=action, result=res.observation,
|
| 140 |
+
tool_error=tool_error,
|
| 141 |
+
))
|
| 142 |
+
if res.done: # env hit its own turn limit and graded
|
| 143 |
+
final = res
|
| 144 |
+
break
|
| 145 |
+
obs = res.observation
|
| 146 |
+
|
| 147 |
+
if final is None:
|
| 148 |
+
# max_turns exhausted without submit — grade what exists.
|
| 149 |
+
final = env.step({"type": "submit"})
|
| 150 |
+
|
| 151 |
+
info = final.info or {}
|
| 152 |
+
return CanonicalTrajectory(
|
| 153 |
+
task_id=task.task_id,
|
| 154 |
+
steps=steps,
|
| 155 |
+
grade=float(final.reward) if final.reward is not None else None,
|
| 156 |
+
guard_ok=bool(info.get("guard_ok", True)),
|
| 157 |
+
hacked=bool(info.get("hacked", False)),
|
| 158 |
+
provenance={"source": "rollout_harness",
|
| 159 |
+
"policy": type(policy).__name__,
|
| 160 |
+
**(provenance or {})},
|
| 161 |
+
)
|
| 162 |
+
|
| 163 |
+
|
| 164 |
+
# ---------------------------------------------------------------------
|
| 165 |
+
# Admission — type the signal and route it (final report §4)
|
| 166 |
+
# ---------------------------------------------------------------------
|
| 167 |
+
|
| 168 |
+
|
| 169 |
+
@dataclass(frozen=True)
|
| 170 |
+
class AdmissionVerdict:
|
| 171 |
+
"""Where a trajectory may go. Routing per the typed-train-on-all verdict:
|
| 172 |
+
clean full passes → SFT; clean near-misses → DPO-candidate (contrastive
|
| 173 |
+
rejected vs a winner, never raw negative gradient); everything else →
|
| 174 |
+
rejected (quarantine-side, full provenance kept for audit)."""
|
| 175 |
+
|
| 176 |
+
sft_admitted: bool
|
| 177 |
+
dpo_candidate: bool
|
| 178 |
+
rejected: bool
|
| 179 |
+
reasons: tuple[str, ...]
|
| 180 |
+
|
| 181 |
+
|
| 182 |
+
def admit(traj: CanonicalTrajectory) -> AdmissionVerdict:
|
| 183 |
+
reasons: list[str] = []
|
| 184 |
+
clean = traj.guard_ok and not traj.hacked
|
| 185 |
+
if not traj.guard_ok:
|
| 186 |
+
reasons.append("pass_to_pass guard broken")
|
| 187 |
+
if traj.hacked:
|
| 188 |
+
reasons.append("hack monitor flagged")
|
| 189 |
+
grade = traj.grade if traj.grade is not None else 0.0
|
| 190 |
+
if traj.grade is None:
|
| 191 |
+
reasons.append("ungraded (no execution oracle)")
|
| 192 |
+
|
| 193 |
+
sft = clean and traj.grade is not None and grade == 1.0
|
| 194 |
+
dpo = clean and traj.grade is not None and 0.0 < grade < 1.0
|
| 195 |
+
if sft:
|
| 196 |
+
reasons.append("clean full pass")
|
| 197 |
+
elif dpo:
|
| 198 |
+
reasons.append(f"clean near-miss (grade={grade:.2f})")
|
| 199 |
+
elif clean and grade == 0.0 and traj.grade is not None:
|
| 200 |
+
reasons.append("clean zero — no partial signal")
|
| 201 |
+
return AdmissionVerdict(
|
| 202 |
+
sft_admitted=sft, dpo_candidate=dpo,
|
| 203 |
+
rejected=not (sft or dpo), reasons=tuple(reasons),
|
| 204 |
+
)
|
| 205 |
+
|
| 206 |
+
|
| 207 |
+
__all__ = [
|
| 208 |
+
"RolloutPolicy",
|
| 209 |
+
"ScriptedPolicy",
|
| 210 |
+
"OpenRouterPolicy",
|
| 211 |
+
"collect_trajectory",
|
| 212 |
+
"AdmissionVerdict",
|
| 213 |
+
"admit",
|
| 214 |
+
]
|
|
@@ -0,0 +1,269 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""swesmith_adapter.py — adapt SWE-smith instances into Feature-Deletion tasks.
|
| 2 |
+
|
| 3 |
+
THE BUY-VS-BUILD VERDICT (deepread finding V4 / D-6): `pip install swesmith`
|
| 4 |
+
(MIT) already ships what ADR-010's "Option B greenfield generator" would have
|
| 5 |
+
hand-built — env construction from arbitrary GitHub repos (ONE Docker image per
|
| 6 |
+
repo, ~500x more storage-efficient than per-task images), five bug-synthesis
|
| 7 |
+
strategies, issue-text generation, and validation-by-test-execution, at a
|
| 8 |
+
verified $1,360 + ~20 human-hours for 50k tasks. Its **PR Mirror strategy is
|
| 9 |
+
exactly this repo's gold-patch-reversion mechanic** and SWE-smith's own ablation
|
| 10 |
+
(Table 5, arXiv:2504.21798) shows PR-Mirror trajectories train the BEST models
|
| 11 |
+
of its five strategies — independent validation of ADR-010's core approach.
|
| 12 |
+
So SWE-smith is the synthesis ENGINE for "point at a repo"; this module is the
|
| 13 |
+
schema bridge into the existing `FeatureDeletionTask` world.
|
| 14 |
+
|
| 15 |
+
THE SEMANTIC INVERSION (load-bearing — easy to get backwards):
|
| 16 |
+
* SWE-bench-shaped instances: `patch` is the GOLD FIX. broken = HEAD with the
|
| 17 |
+
fix reverted (`git apply -R patch`). `SweBenchAdapter` stores `patch` as
|
| 18 |
+
`golden_diff` directly.
|
| 19 |
+
* SWE-smith instances: `patch` INTRODUCES THE BUG. broken = HEAD with the
|
| 20 |
+
patch APPLIED. The fix — what the agent must produce, the validator's gate-4
|
| 21 |
+
restoration diff — is the REVERSE of the bug patch.
|
| 22 |
+
This adapter therefore stores `golden_diff = reverse_unified_diff(bug_patch)`.
|
| 23 |
+
When mechanical reversal fails (exotic diff features), it falls back to the
|
| 24 |
+
original patch tagged with a provenance marker so downstream gate-4 validation
|
| 25 |
+
knows to use `git apply -R` instead of `git apply`.
|
| 26 |
+
|
| 27 |
+
The adapter itself needs nothing beyond core deps. Live synthesis (building new
|
| 28 |
+
repo profiles / generating new bugs) needs the `swesmith` toolkit + Docker on
|
| 29 |
+
Linux — see the `[swesmith]` extra in pyproject.
|
| 30 |
+
"""
|
| 31 |
+
from __future__ import annotations
|
| 32 |
+
|
| 33 |
+
import json
|
| 34 |
+
import re
|
| 35 |
+
from dataclasses import dataclass
|
| 36 |
+
|
| 37 |
+
from composer_replication.datagen.schema import FeatureDeletionTask
|
| 38 |
+
from composer_replication.datagen.substrates import _as_tuple
|
| 39 |
+
|
| 40 |
+
#: Marker prefixed to golden_diff when reverse_unified_diff could not invert the
|
| 41 |
+
#: bug patch mechanically. Consumers (gate-4 validation) must then apply the
|
| 42 |
+
#: remainder with `git apply -R` (it is the FORWARD bug patch, not the fix).
|
| 43 |
+
UNREVERSED_MARKER = "### UNREVERSED-BUG-PATCH (apply with -R) ###\n"
|
| 44 |
+
|
| 45 |
+
#: instance_id substring patterns → synthesis strategy (SWE-smith §2.1 / §B).
|
| 46 |
+
#: Patterns follow the toolkit's naming: e.g.
|
| 47 |
+
#: pandas-dev__pandas.abc123.lm_modify__xyz
|
| 48 |
+
#: ...func_pm_ctrl_invert_if__..., ...combine_file__..., ...pr_1234
|
| 49 |
+
_STRATEGY_PATTERNS: tuple[tuple[str, str], ...] = (
|
| 50 |
+
("lm_modify", "lm_modify"),
|
| 51 |
+
("lm_rewrite", "lm_rewrite"),
|
| 52 |
+
("func_pm", "procedural"), # procedural AST modifications (13 transform types)
|
| 53 |
+
("func_basic", "procedural"),
|
| 54 |
+
("combine_file", "combine"),
|
| 55 |
+
("combine_module", "combine"),
|
| 56 |
+
("combine", "combine"),
|
| 57 |
+
("pr_", "pr_mirror"),
|
| 58 |
+
)
|
| 59 |
+
|
| 60 |
+
|
| 61 |
+
def parse_strategy(instance_id: str) -> str:
|
| 62 |
+
"""Map a SWE-smith instance_id to its bug-synthesis strategy.
|
| 63 |
+
|
| 64 |
+
Returns one of {lm_modify, lm_rewrite, procedural, combine, pr_mirror,
|
| 65 |
+
unknown}. The strategy matters because SWE-smith's Table 5 ablation found
|
| 66 |
+
trajectory quality differs sharply by strategy (PR Mirror best, LM Modify
|
| 67 |
+
steep drop-off) — we carry it as provenance so corpus builds can weight or
|
| 68 |
+
filter by strategy.
|
| 69 |
+
"""
|
| 70 |
+
iid = (instance_id or "").lower()
|
| 71 |
+
for pattern, strategy in _STRATEGY_PATTERNS:
|
| 72 |
+
if pattern in iid:
|
| 73 |
+
return strategy
|
| 74 |
+
return "unknown"
|
| 75 |
+
|
| 76 |
+
|
| 77 |
+
#: Heuristic cold-start difficulty priors per strategy, motivated by SWE-smith
|
| 78 |
+
#: Table 1 medians (PR Mirror: 3 median F2P but 14 lines edited; Combine: 15
|
| 79 |
+
#: F2P / 11 lines = multi-site; procedural: 7 F2P / 5 lines, mechanical).
|
| 80 |
+
#: These only seed DifficultyCurriculum's p-hat before real rollouts exist.
|
| 81 |
+
_DIFFICULTY_PRIOR: dict[str, float] = {
|
| 82 |
+
"pr_mirror": 0.4,
|
| 83 |
+
"combine": 0.4,
|
| 84 |
+
"lm_rewrite": 0.45,
|
| 85 |
+
"lm_modify": 0.55,
|
| 86 |
+
"procedural": 0.6,
|
| 87 |
+
"unknown": 0.5,
|
| 88 |
+
}
|
| 89 |
+
|
| 90 |
+
|
| 91 |
+
_HUNK_RE = re.compile(
|
| 92 |
+
r"^@@ -(?P<old_start>\d+)(?:,(?P<old_count>\d+))? "
|
| 93 |
+
r"\+(?P<new_start>\d+)(?:,(?P<new_count>\d+))? @@(?P<tail>.*)$"
|
| 94 |
+
)
|
| 95 |
+
|
| 96 |
+
|
| 97 |
+
def reverse_unified_diff(patch: str) -> str | None:
|
| 98 |
+
"""Mechanically invert a unified diff (swap additions and deletions).
|
| 99 |
+
|
| 100 |
+
Handles the standard unified-diff features SWE-smith patches use:
|
| 101 |
+
``diff --git`` headers, ``---``/``+++`` file lines, ``@@`` hunk headers
|
| 102 |
+
(old/new ranges swapped), ``+``/``-`` body lines (swapped), context lines,
|
| 103 |
+
and ``\`` markers (kept in place).
|
| 104 |
+
|
| 105 |
+
HONEST LIMITATIONS (returns None — caller falls back to UNREVERSED_MARKER):
|
| 106 |
+
* file mode changes (``old mode``/``new mode``), renames/copies
|
| 107 |
+
(``rename from``...), binary patches (``GIT binary patch``), and
|
| 108 |
+
``index`` lines with mode suffixes are NOT inverted — reversing them
|
| 109 |
+
correctly requires git plumbing, not text surgery.
|
| 110 |
+
* Within a hunk, a reversed diff's line ORDER for paired -/+ runs is the
|
| 111 |
+
naive swap; `git apply` accepts it, but it is not byte-identical to
|
| 112 |
+
what `git diff` would emit for the reverse change.
|
| 113 |
+
"""
|
| 114 |
+
if not patch or "@@" not in patch:
|
| 115 |
+
return None
|
| 116 |
+
unsupported = ("old mode ", "new mode ", "rename from ", "rename to ",
|
| 117 |
+
"copy from ", "copy to ", "GIT binary patch")
|
| 118 |
+
if any(marker in patch for marker in unsupported):
|
| 119 |
+
return None
|
| 120 |
+
|
| 121 |
+
out: list[str] = []
|
| 122 |
+
for line in patch.splitlines():
|
| 123 |
+
if line.startswith("diff --git "):
|
| 124 |
+
# `diff --git a/<old> b/<new>` → swap the two paths.
|
| 125 |
+
m = re.match(r"^diff --git a/(?P<a>.+) b/(?P<b>.+)$", line)
|
| 126 |
+
if m:
|
| 127 |
+
out.append(f"diff --git a/{m.group('b')} b/{m.group('a')}")
|
| 128 |
+
else:
|
| 129 |
+
out.append(line)
|
| 130 |
+
elif line.startswith("--- "):
|
| 131 |
+
out.append("+++ " + line[4:].replace("a/", "b/", 1)
|
| 132 |
+
if line[4:].startswith("a/") else "+++ " + line[4:])
|
| 133 |
+
elif line.startswith("+++ "):
|
| 134 |
+
out.append("--- " + line[4:].replace("b/", "a/", 1)
|
| 135 |
+
if line[4:].startswith("b/") else "--- " + line[4:])
|
| 136 |
+
elif line.startswith("@@"):
|
| 137 |
+
m = _HUNK_RE.match(line)
|
| 138 |
+
if not m:
|
| 139 |
+
return None
|
| 140 |
+
old_start, old_count = m.group("old_start"), m.group("old_count")
|
| 141 |
+
new_start, new_count = m.group("new_start"), m.group("new_count")
|
| 142 |
+
oc = f",{old_count}" if old_count is not None else ""
|
| 143 |
+
nc = f",{new_count}" if new_count is not None else ""
|
| 144 |
+
out.append(f"@@ -{new_start}{nc} +{old_start}{oc} @@{m.group('tail')}")
|
| 145 |
+
elif line.startswith("+"):
|
| 146 |
+
out.append("-" + line[1:])
|
| 147 |
+
elif line.startswith("-"):
|
| 148 |
+
out.append("+" + line[1:])
|
| 149 |
+
else:
|
| 150 |
+
# context lines, `index ...`, `\ No newline...` pass through.
|
| 151 |
+
out.append(line)
|
| 152 |
+
return "\n".join(out) + ("\n" if patch.endswith("\n") else "")
|
| 153 |
+
|
| 154 |
+
|
| 155 |
+
@dataclass(frozen=True)
|
| 156 |
+
class SwesmithMeta:
|
| 157 |
+
"""Sidecar provenance for a SWE-smith-derived task.
|
| 158 |
+
|
| 159 |
+
Kept OUT of the frozen `FeatureDeletionTask` schema deliberately — the
|
| 160 |
+
schema is shared with SweBenchAdapter and the trainer; strategy provenance
|
| 161 |
+
is a corpus-construction concern, carried alongside (e.g. into the run
|
| 162 |
+
manifest), never into the policy-visible task row.
|
| 163 |
+
"""
|
| 164 |
+
|
| 165 |
+
strategy: str # lm_modify | lm_rewrite | procedural | combine | pr_mirror | unknown
|
| 166 |
+
diff_reversed: bool # True if golden_diff is the mechanical reverse of the bug patch
|
| 167 |
+
source: str = "swesmith"
|
| 168 |
+
|
| 169 |
+
|
| 170 |
+
@dataclass
|
| 171 |
+
class SwesmithAdapter:
|
| 172 |
+
"""Convert a SWE-smith instance dict into a FeatureDeletionTask.
|
| 173 |
+
|
| 174 |
+
Mirrors `SweBenchAdapter`'s shape; differs in the patch semantics (see the
|
| 175 |
+
module docstring INVERSION note) and the per-REPO image convention.
|
| 176 |
+
"""
|
| 177 |
+
|
| 178 |
+
default_test_command: str = "python -m pytest -q"
|
| 179 |
+
|
| 180 |
+
def image_for(self, instance: dict) -> str:
|
| 181 |
+
# SWE-smith publishes ONE image per repo (not per task). Rows carry
|
| 182 |
+
# `image_name`; some exports use `docker_image`. Fall back to the
|
| 183 |
+
# toolkit's naming convention derived from the repo slug.
|
| 184 |
+
for key in ("image_name", "docker_image"):
|
| 185 |
+
if instance.get(key):
|
| 186 |
+
return str(instance[key])
|
| 187 |
+
repo = str(instance.get("repo", "unknown")).replace("/", "__").lower()
|
| 188 |
+
return f"swesmith.x86_64.{repo}:latest"
|
| 189 |
+
|
| 190 |
+
def to_task(self, instance: dict) -> FeatureDeletionTask:
|
| 191 |
+
task, _meta = self.to_task_with_meta(instance)
|
| 192 |
+
return task
|
| 193 |
+
|
| 194 |
+
def to_task_with_meta(self, instance: dict) -> tuple[FeatureDeletionTask, SwesmithMeta]:
|
| 195 |
+
iid = str(instance.get("instance_id") or instance.get("task_id") or "unknown")
|
| 196 |
+
strategy = parse_strategy(iid)
|
| 197 |
+
|
| 198 |
+
bug_patch = str(instance.get("patch", ""))
|
| 199 |
+
fix = reverse_unified_diff(bug_patch)
|
| 200 |
+
if fix is not None:
|
| 201 |
+
golden_diff = fix
|
| 202 |
+
diff_reversed = True
|
| 203 |
+
else:
|
| 204 |
+
golden_diff = UNREVERSED_MARKER + bug_patch
|
| 205 |
+
diff_reversed = False
|
| 206 |
+
|
| 207 |
+
ftp = _as_tuple(instance.get("FAIL_TO_PASS"))
|
| 208 |
+
ptp = _as_tuple(instance.get("PASS_TO_PASS"))
|
| 209 |
+
|
| 210 |
+
task = FeatureDeletionTask(
|
| 211 |
+
task_id=iid,
|
| 212 |
+
repo=str(instance.get("repo", "unknown")),
|
| 213 |
+
base_commit=str(instance.get("base_commit", "")),
|
| 214 |
+
broken_image=self.image_for(instance),
|
| 215 |
+
test_command=str(instance.get("test_command") or self.default_test_command),
|
| 216 |
+
fail_to_pass=ftp,
|
| 217 |
+
pass_to_pass=ptp,
|
| 218 |
+
golden_diff=golden_diff,
|
| 219 |
+
granularity="feature",
|
| 220 |
+
# SWE-smith rows don't carry per-instance licenses; repo-level
|
| 221 |
+
# licensing is the repo_gate's job (deepread finding V9/D-13).
|
| 222 |
+
upstream_license=str(instance.get("license_name", "unknown")),
|
| 223 |
+
difficulty_prior=_DIFFICULTY_PRIOR.get(strategy, 0.5),
|
| 224 |
+
)
|
| 225 |
+
return task, SwesmithMeta(strategy=strategy, diff_reversed=diff_reversed)
|
| 226 |
+
|
| 227 |
+
|
| 228 |
+
def load_swesmith_instances(
|
| 229 |
+
path_or_hf_id: str,
|
| 230 |
+
*,
|
| 231 |
+
limit: int | None = None,
|
| 232 |
+
) -> list[dict]:
|
| 233 |
+
"""Load SWE-smith instances from a local JSONL file or the HF dataset.
|
| 234 |
+
|
| 235 |
+
Local ``.jsonl`` paths need no extra deps (used by tests/fixtures). HF ids
|
| 236 |
+
(e.g. ``SWE-bench/SWE-smith``) lazy-import `datasets` from the `[datagen]`
|
| 237 |
+
extra.
|
| 238 |
+
"""
|
| 239 |
+
if path_or_hf_id.endswith(".jsonl"):
|
| 240 |
+
rows: list[dict] = []
|
| 241 |
+
with open(path_or_hf_id, encoding="utf-8") as f:
|
| 242 |
+
for line in f:
|
| 243 |
+
line = line.strip()
|
| 244 |
+
if not line:
|
| 245 |
+
continue
|
| 246 |
+
rows.append(json.loads(line))
|
| 247 |
+
if limit is not None and len(rows) >= limit:
|
| 248 |
+
break
|
| 249 |
+
return rows
|
| 250 |
+
try:
|
| 251 |
+
from datasets import load_dataset # noqa: PLC0415 — lazy heavy dep
|
| 252 |
+
except ImportError as e:
|
| 253 |
+
raise RuntimeError(
|
| 254 |
+
"Loading SWE-smith from the HF Hub requires `datasets`; install "
|
| 255 |
+
"with `pip install -e .[datagen]`. Got: " + repr(e)
|
| 256 |
+
) from e
|
| 257 |
+
split = load_dataset(path_or_hf_id, split="train")
|
| 258 |
+
rows = [dict(r) for i, r in enumerate(split) if limit is None or i < limit]
|
| 259 |
+
return rows[: limit if limit is not None else len(rows)]
|
| 260 |
+
|
| 261 |
+
|
| 262 |
+
__all__ = [
|
| 263 |
+
"SwesmithAdapter",
|
| 264 |
+
"SwesmithMeta",
|
| 265 |
+
"UNREVERSED_MARKER",
|
| 266 |
+
"load_swesmith_instances",
|
| 267 |
+
"parse_strategy",
|
| 268 |
+
"reverse_unified_diff",
|
| 269 |
+
]
|
|
@@ -0,0 +1,419 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Tests for the Stage-0 ingest gate (repo_gate.py) — architecture step 1.
|
| 2 |
+
|
| 3 |
+
Coverage targets the two findings the module closes:
|
| 4 |
+
* V9/D-13 — SPDX detection from real license fixture texts (incl. the
|
| 5 |
+
tricky GNU-family cross-citation and Apache-vs-MIT phrasing) and the
|
| 6 |
+
three-tier mapping that replaces the old boolean substring filter.
|
| 7 |
+
* V3/D-5 — decontamination hits in exact, URL, and mixed-case forms, and
|
| 8 |
+
the hard never-admit rule in the composed verdict.
|
| 9 |
+
|
| 10 |
+
CPU-only, stdlib + tmp_path fixtures — no network, no Docker.
|
| 11 |
+
"""
|
| 12 |
+
from __future__ import annotations
|
| 13 |
+
|
| 14 |
+
import json
|
| 15 |
+
from pathlib import Path
|
| 16 |
+
|
| 17 |
+
import pytest
|
| 18 |
+
|
| 19 |
+
from composer_replication.datagen.repo_gate import (
|
| 20 |
+
DECONTAMINATION_LIST,
|
| 21 |
+
GateVerdict,
|
| 22 |
+
LicenseInfo,
|
| 23 |
+
Tier,
|
| 24 |
+
detect_license,
|
| 25 |
+
gate_repo,
|
| 26 |
+
is_eval_contaminated,
|
| 27 |
+
license_tier,
|
| 28 |
+
load_decontamination_list,
|
| 29 |
+
normalize_repo,
|
| 30 |
+
)
|
| 31 |
+
|
| 32 |
+
# ---------------------------------------------------------------------
|
| 33 |
+
# License fixture texts — distinctive excerpts of the real license texts.
|
| 34 |
+
# ---------------------------------------------------------------------
|
| 35 |
+
|
| 36 |
+
MIT_TEXT = """\
|
| 37 |
+
MIT License
|
| 38 |
+
|
| 39 |
+
Copyright (c) 2026 Example Org
|
| 40 |
+
|
| 41 |
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
| 42 |
+
of this software and associated documentation files (the "Software"), to deal
|
| 43 |
+
in the Software without restriction...
|
| 44 |
+
"""
|
| 45 |
+
|
| 46 |
+
# The tricky Apache case: the words "permission" and "license" appear in both
|
| 47 |
+
# MIT and Apache; only Apache names itself with a version.
|
| 48 |
+
APACHE_TEXT = """\
|
| 49 |
+
Apache License
|
| 50 |
+
Version 2.0, January 2004
|
| 51 |
+
http://www.apache.org/licenses/
|
| 52 |
+
|
| 53 |
+
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
| 54 |
+
|
| 55 |
+
1. Definitions.
|
| 56 |
+
"License" shall mean the terms and conditions for use, reproduction,
|
| 57 |
+
and distribution as defined by Sections 1 through 9 of this document.
|
| 58 |
+
"""
|
| 59 |
+
|
| 60 |
+
BSD3_TEXT = """\
|
| 61 |
+
BSD 3-Clause License
|
| 62 |
+
|
| 63 |
+
Redistribution and use in source and binary forms, with or without
|
| 64 |
+
modification, are permitted provided that the following conditions are met:
|
| 65 |
+
|
| 66 |
+
1. Redistributions of source code must retain the above copyright notice...
|
| 67 |
+
3. Neither the name of the copyright holder nor the names of its
|
| 68 |
+
contributors may be used to endorse or promote products derived from
|
| 69 |
+
this software without specific prior written permission.
|
| 70 |
+
"""
|
| 71 |
+
|
| 72 |
+
BSD2_TEXT = """\
|
| 73 |
+
BSD 2-Clause License
|
| 74 |
+
|
| 75 |
+
Redistribution and use in source and binary forms, with or without
|
| 76 |
+
modification, are permitted provided that the following conditions are met:
|
| 77 |
+
|
| 78 |
+
1. Redistributions of source code must retain the above copyright notice.
|
| 79 |
+
2. Redistributions in binary form must reproduce the above copyright notice.
|
| 80 |
+
"""
|
| 81 |
+
|
| 82 |
+
ISC_TEXT = """\
|
| 83 |
+
ISC License
|
| 84 |
+
|
| 85 |
+
Copyright (c) 2026, Example Org
|
| 86 |
+
|
| 87 |
+
Permission to use, copy, modify, and/or distribute this software for any
|
| 88 |
+
purpose with or without fee is hereby granted, provided that the above
|
| 89 |
+
copyright notice and this permission notice appear in all copies.
|
| 90 |
+
"""
|
| 91 |
+
|
| 92 |
+
# GPL-3.0 §13 cross-cites the AGPL by full name — the classic trap for
|
| 93 |
+
# full-body substring matchers. Header anchoring must win.
|
| 94 |
+
GPL3_TEXT = """\
|
| 95 |
+
GNU GENERAL PUBLIC LICENSE
|
| 96 |
+
Version 3, 29 June 2007
|
| 97 |
+
|
| 98 |
+
Copyright (C) 2007 Free Software Foundation, Inc.
|
| 99 |
+
|
| 100 |
+
13. Use with the GNU Affero General Public License.
|
| 101 |
+
Notwithstanding any other provision of this License, you have
|
| 102 |
+
permission to link or combine any covered work with a work licensed
|
| 103 |
+
under version 3 of the GNU Affero General Public License...
|
| 104 |
+
"""
|
| 105 |
+
|
| 106 |
+
GPL2_TEXT = """\
|
| 107 |
+
GNU GENERAL PUBLIC LICENSE
|
| 108 |
+
Version 2, June 1991
|
| 109 |
+
|
| 110 |
+
Copyright (C) 1989, 1991 Free Software Foundation, Inc.
|
| 111 |
+
Everyone is permitted to copy and distribute verbatim copies
|
| 112 |
+
of this license document, but changing it is not allowed.
|
| 113 |
+
"""
|
| 114 |
+
|
| 115 |
+
# AGPL-3.0 §13 reciprocally cites the plain GPL by name.
|
| 116 |
+
AGPL3_TEXT = """\
|
| 117 |
+
GNU AFFERO GENERAL PUBLIC LICENSE
|
| 118 |
+
Version 3, 19 November 2007
|
| 119 |
+
|
| 120 |
+
13. Remote Network Interaction; Use with the GNU General Public License.
|
| 121 |
+
Notwithstanding any other provision of this License...
|
| 122 |
+
"""
|
| 123 |
+
|
| 124 |
+
LGPL21_TEXT = """\
|
| 125 |
+
GNU LESSER GENERAL PUBLIC LICENSE
|
| 126 |
+
Version 2.1, February 1999
|
| 127 |
+
|
| 128 |
+
Copyright (C) 1991, 1999 Free Software Foundation, Inc.
|
| 129 |
+
"""
|
| 130 |
+
|
| 131 |
+
MPL2_TEXT = """\
|
| 132 |
+
Mozilla Public License Version 2.0
|
| 133 |
+
==================================
|
| 134 |
+
|
| 135 |
+
1. Definitions
|
| 136 |
+
--------------
|
| 137 |
+
1.1. "Contributor"
|
| 138 |
+
means each individual or legal entity that creates, contributes to
|
| 139 |
+
the creation of, or owns Covered Software.
|
| 140 |
+
"""
|
| 141 |
+
|
| 142 |
+
UNLICENSE_TEXT = """\
|
| 143 |
+
This is free and unencumbered software released into the public domain.
|
| 144 |
+
|
| 145 |
+
Anyone is free to copy, modify, publish, use, compile, sell, or
|
| 146 |
+
distribute this software, either in source code form or as a compiled
|
| 147 |
+
binary, for any purpose, commercial or non-commercial, and by any means.
|
| 148 |
+
"""
|
| 149 |
+
|
| 150 |
+
|
| 151 |
+
def _repo_with_license(tmp_path: Path, text: str, filename: str = "LICENSE") -> Path:
|
| 152 |
+
(tmp_path / filename).write_text(text, encoding="utf-8")
|
| 153 |
+
return tmp_path
|
| 154 |
+
|
| 155 |
+
|
| 156 |
+
# ---------------------------------------------------------------------
|
| 157 |
+
# detect_license — SPDX classification from file text
|
| 158 |
+
# ---------------------------------------------------------------------
|
| 159 |
+
|
| 160 |
+
|
| 161 |
+
@pytest.mark.parametrize(
|
| 162 |
+
("text", "expected"),
|
| 163 |
+
[
|
| 164 |
+
(MIT_TEXT, "MIT"),
|
| 165 |
+
(APACHE_TEXT, "Apache-2.0"),
|
| 166 |
+
(BSD3_TEXT, "BSD-3-Clause"),
|
| 167 |
+
(BSD2_TEXT, "BSD-2-Clause"),
|
| 168 |
+
(ISC_TEXT, "ISC"),
|
| 169 |
+
(GPL3_TEXT, "GPL-3.0"),
|
| 170 |
+
(GPL2_TEXT, "GPL-2.0"),
|
| 171 |
+
(AGPL3_TEXT, "AGPL-3.0"),
|
| 172 |
+
(LGPL21_TEXT, "LGPL-2.1"),
|
| 173 |
+
(MPL2_TEXT, "MPL-2.0"),
|
| 174 |
+
(UNLICENSE_TEXT, "Unlicense"),
|
| 175 |
+
],
|
| 176 |
+
ids=["mit", "apache2", "bsd3", "bsd2", "isc", "gpl3", "gpl2", "agpl3", "lgpl21", "mpl2", "unlicense"],
|
| 177 |
+
)
|
| 178 |
+
def test_detect_license_spdx_ids(tmp_path: Path, text: str, expected: str):
|
| 179 |
+
info = detect_license(_repo_with_license(tmp_path, text))
|
| 180 |
+
assert info.spdx_id == expected
|
| 181 |
+
assert info.signal == "license_file"
|
| 182 |
+
assert info.source == "LICENSE"
|
| 183 |
+
|
| 184 |
+
|
| 185 |
+
def test_gpl3_not_misread_as_agpl(tmp_path: Path):
|
| 186 |
+
"""GPL-3.0 §13 names the AGPL in its body; header anchoring must keep
|
| 187 |
+
this classified as GPL-3.0 (the V9 substring filter would have tripped)."""
|
| 188 |
+
info = detect_license(_repo_with_license(tmp_path, GPL3_TEXT))
|
| 189 |
+
assert info.spdx_id == "GPL-3.0"
|
| 190 |
+
|
| 191 |
+
|
| 192 |
+
def test_apache_notice_without_header_still_apache(tmp_path: Path):
|
| 193 |
+
"""The short 'Licensed under the Apache License, Version 2.0' boilerplate
|
| 194 |
+
has no canonical header — the body fallback must catch it, and must not
|
| 195 |
+
fall through to MIT despite shared 'permission' vocabulary."""
|
| 196 |
+
notice = (
|
| 197 |
+
"Copyright 2026 Example Org\n\n"
|
| 198 |
+
"Licensed under the Apache License, Version 2.0 (the \"License\");\n"
|
| 199 |
+
"you may not use this file except in compliance with the License.\n"
|
| 200 |
+
)
|
| 201 |
+
info = detect_license(_repo_with_license(tmp_path, notice))
|
| 202 |
+
assert info.spdx_id == "Apache-2.0"
|
| 203 |
+
|
| 204 |
+
|
| 205 |
+
def test_detect_license_alternate_filenames(tmp_path: Path):
|
| 206 |
+
info = detect_license(_repo_with_license(tmp_path, GPL2_TEXT, filename="COPYING"))
|
| 207 |
+
assert info.spdx_id == "GPL-2.0"
|
| 208 |
+
assert info.source == "COPYING"
|
| 209 |
+
info2 = detect_license(_repo_with_license(tmp_path, MIT_TEXT, filename="LICENSE.md"))
|
| 210 |
+
# LICENSE.md is also present in tmp_path now alongside COPYING; first
|
| 211 |
+
# filename in priority order (LICENSE/LICENSE.txt/LICENSE.md) wins over COPYING.
|
| 212 |
+
assert info2.spdx_id == "MIT"
|
| 213 |
+
assert info2.source == "LICENSE.md"
|
| 214 |
+
|
| 215 |
+
|
| 216 |
+
def test_detect_license_unknown_text(tmp_path: Path):
|
| 217 |
+
info = detect_license(_repo_with_license(tmp_path, "All rights reserved. Ask legal."))
|
| 218 |
+
assert info.spdx_id == "unknown"
|
| 219 |
+
|
| 220 |
+
|
| 221 |
+
def test_detect_license_no_files(tmp_path: Path):
|
| 222 |
+
info = detect_license(tmp_path)
|
| 223 |
+
assert info == LicenseInfo(spdx_id="unknown", signal="none")
|
| 224 |
+
|
| 225 |
+
|
| 226 |
+
def test_classifier_secondary_signal(tmp_path: Path):
|
| 227 |
+
"""No LICENSE file, but pyproject carries a trove classifier — the
|
| 228 |
+
classifier signal must win and be recorded as such."""
|
| 229 |
+
(tmp_path / "pyproject.toml").write_text(
|
| 230 |
+
'[project]\nname = "x"\nclassifiers = [\n'
|
| 231 |
+
' "License :: OSI Approved :: MIT License",\n]\n',
|
| 232 |
+
encoding="utf-8",
|
| 233 |
+
)
|
| 234 |
+
info = detect_license(tmp_path)
|
| 235 |
+
assert info.spdx_id == "MIT"
|
| 236 |
+
assert info.signal == "classifier"
|
| 237 |
+
assert info.source == "pyproject.toml"
|
| 238 |
+
|
| 239 |
+
|
| 240 |
+
def test_classifier_pep639_expression(tmp_path: Path):
|
| 241 |
+
(tmp_path / "pyproject.toml").write_text(
|
| 242 |
+
'[project]\nname = "x"\nlicense = "Apache-2.0"\n', encoding="utf-8"
|
| 243 |
+
)
|
| 244 |
+
info = detect_license(tmp_path)
|
| 245 |
+
assert info.spdx_id == "Apache-2.0"
|
| 246 |
+
assert info.signal == "classifier"
|
| 247 |
+
|
| 248 |
+
|
| 249 |
+
def test_license_file_beats_classifier(tmp_path: Path):
|
| 250 |
+
"""When both signals exist and the file is classifiable, the file wins —
|
| 251 |
+
the classifier is secondary by design (it can't tell BSD-2 from BSD-3)."""
|
| 252 |
+
_repo_with_license(tmp_path, GPL3_TEXT)
|
| 253 |
+
(tmp_path / "pyproject.toml").write_text(
|
| 254 |
+
'classifiers = ["License :: OSI Approved :: MIT License"]\n', encoding="utf-8"
|
| 255 |
+
)
|
| 256 |
+
info = detect_license(tmp_path)
|
| 257 |
+
assert info.spdx_id == "GPL-3.0"
|
| 258 |
+
assert info.signal == "license_file"
|
| 259 |
+
|
| 260 |
+
|
| 261 |
+
def test_unclassifiable_file_falls_back_to_classifier(tmp_path: Path):
|
| 262 |
+
_repo_with_license(tmp_path, "Custom corporate license, see legal dept.")
|
| 263 |
+
(tmp_path / "pyproject.toml").write_text(
|
| 264 |
+
'classifiers = ["License :: OSI Approved :: ISC License"]\n', encoding="utf-8"
|
| 265 |
+
)
|
| 266 |
+
info = detect_license(tmp_path)
|
| 267 |
+
assert info.spdx_id == "ISC"
|
| 268 |
+
assert info.signal == "classifier"
|
| 269 |
+
|
| 270 |
+
|
| 271 |
+
# ---------------------------------------------------------------------
|
| 272 |
+
# license_tier — tiers, not a boolean (D-13)
|
| 273 |
+
# ---------------------------------------------------------------------
|
| 274 |
+
|
| 275 |
+
|
| 276 |
+
@pytest.mark.parametrize(
|
| 277 |
+
("spdx", "tier"),
|
| 278 |
+
[
|
| 279 |
+
("MIT", Tier.REDISTRIBUTABLE),
|
| 280 |
+
("Apache-2.0", Tier.REDISTRIBUTABLE),
|
| 281 |
+
("BSD-2-Clause", Tier.REDISTRIBUTABLE),
|
| 282 |
+
("BSD-3-Clause", Tier.REDISTRIBUTABLE),
|
| 283 |
+
("ISC", Tier.REDISTRIBUTABLE),
|
| 284 |
+
("Unlicense", Tier.REDISTRIBUTABLE),
|
| 285 |
+
("MPL-2.0", Tier.TRAINABLE_ONLY),
|
| 286 |
+
("LGPL-2.1", Tier.TRAINABLE_ONLY),
|
| 287 |
+
("LGPL-3.0", Tier.TRAINABLE_ONLY),
|
| 288 |
+
("GPL-2.0", Tier.EXCLUDED),
|
| 289 |
+
("GPL-3.0", Tier.EXCLUDED),
|
| 290 |
+
("AGPL-3.0", Tier.EXCLUDED),
|
| 291 |
+
("unknown", Tier.EXCLUDED),
|
| 292 |
+
("WTFPL", Tier.EXCLUDED), # unrecognized id → fail closed
|
| 293 |
+
],
|
| 294 |
+
)
|
| 295 |
+
def test_license_tier_mapping(spdx: str, tier: Tier):
|
| 296 |
+
assert license_tier(LicenseInfo(spdx_id=spdx, signal="license_file")) is tier
|
| 297 |
+
|
| 298 |
+
|
| 299 |
+
# ---------------------------------------------------------------------
|
| 300 |
+
# Decontamination (V3 / D-5)
|
| 301 |
+
# ---------------------------------------------------------------------
|
| 302 |
+
|
| 303 |
+
|
| 304 |
+
def test_decontamination_list_has_the_canonical_12():
|
| 305 |
+
assert len(DECONTAMINATION_LIST) == 12
|
| 306 |
+
assert "django/django" in DECONTAMINATION_LIST
|
| 307 |
+
assert "sympy/sympy" in DECONTAMINATION_LIST
|
| 308 |
+
|
| 309 |
+
|
| 310 |
+
@pytest.mark.parametrize(
|
| 311 |
+
"repo",
|
| 312 |
+
[
|
| 313 |
+
"django/django", # exact
|
| 314 |
+
"Django/Django", # case
|
| 315 |
+
"https://github.com/django/django", # https URL
|
| 316 |
+
"https://github.com/django/django.git", # URL + .git
|
| 317 |
+
"git@github.com:django/django.git", # ssh URL
|
| 318 |
+
"https://github.com/django/django/", # trailing slash
|
| 319 |
+
],
|
| 320 |
+
)
|
| 321 |
+
def test_is_eval_contaminated_hits(repo: str):
|
| 322 |
+
assert is_eval_contaminated(repo) is True
|
| 323 |
+
|
| 324 |
+
|
| 325 |
+
@pytest.mark.parametrize(
|
| 326 |
+
"repo",
|
| 327 |
+
[
|
| 328 |
+
"pandas-dev/pandas",
|
| 329 |
+
"https://github.com/torvalds/linux",
|
| 330 |
+
"someuser/django", # fork-org differs: NOT the eval repo
|
| 331 |
+
],
|
| 332 |
+
)
|
| 333 |
+
def test_is_eval_contaminated_misses(repo: str):
|
| 334 |
+
assert is_eval_contaminated(repo) is False
|
| 335 |
+
|
| 336 |
+
|
| 337 |
+
def test_normalize_repo_forms():
|
| 338 |
+
assert normalize_repo("git@github.com:PSF/Requests.git") == "psf/requests"
|
| 339 |
+
assert normalize_repo("https://github.com/pydata/xarray/tree/main") == "pydata/xarray"
|
| 340 |
+
|
| 341 |
+
|
| 342 |
+
def test_extension_list_from_json(tmp_path: Path):
|
| 343 |
+
"""The documented extension mechanism: extra eval repos load from JSON
|
| 344 |
+
and hit through the same normalized matching."""
|
| 345 |
+
extra_path = tmp_path / "extra.json"
|
| 346 |
+
extra_path.write_text(json.dumps(["SWE-Gym/Extra-Repo"]), encoding="utf-8")
|
| 347 |
+
extra = load_decontamination_list(extra_path)
|
| 348 |
+
assert is_eval_contaminated("https://github.com/swe-gym/extra-repo", extra_list=extra)
|
| 349 |
+
assert not is_eval_contaminated("swe-gym/other-repo", extra_list=extra)
|
| 350 |
+
|
| 351 |
+
|
| 352 |
+
def test_extension_list_rejects_non_list(tmp_path: Path):
|
| 353 |
+
bad = tmp_path / "bad.json"
|
| 354 |
+
bad.write_text('{"repo": "a/b"}', encoding="utf-8")
|
| 355 |
+
with pytest.raises(ValueError):
|
| 356 |
+
load_decontamination_list(bad)
|
| 357 |
+
|
| 358 |
+
|
| 359 |
+
# ---------------------------------------------------------------------
|
| 360 |
+
# gate_repo — verdict composition
|
| 361 |
+
# ---------------------------------------------------------------------
|
| 362 |
+
|
| 363 |
+
|
| 364 |
+
def test_gate_admits_permissive_clean_repo(tmp_path: Path):
|
| 365 |
+
v = gate_repo("example/clean", _repo_with_license(tmp_path, MIT_TEXT))
|
| 366 |
+
assert isinstance(v, GateVerdict)
|
| 367 |
+
assert v.admitted is True
|
| 368 |
+
assert v.tier is Tier.REDISTRIBUTABLE
|
| 369 |
+
assert v.contaminated is False
|
| 370 |
+
assert v.reasons == []
|
| 371 |
+
|
| 372 |
+
|
| 373 |
+
def test_gate_contaminated_never_admitted_even_if_permissive(tmp_path: Path):
|
| 374 |
+
"""V3 hard rule: an eval repo with an MIT license is STILL rejected —
|
| 375 |
+
decontamination outranks license."""
|
| 376 |
+
v = gate_repo("https://github.com/pallets/flask", _repo_with_license(tmp_path, MIT_TEXT))
|
| 377 |
+
assert v.contaminated is True
|
| 378 |
+
assert v.admitted is False
|
| 379 |
+
assert any("decontamination" in r for r in v.reasons)
|
| 380 |
+
# license detection still ran and is recorded for the manifest
|
| 381 |
+
assert v.license_info.spdx_id == "MIT"
|
| 382 |
+
|
| 383 |
+
|
| 384 |
+
def test_gate_excluded_tier_never_admitted(tmp_path: Path):
|
| 385 |
+
v = gate_repo("example/agpl-repo", _repo_with_license(tmp_path, AGPL3_TEXT))
|
| 386 |
+
assert v.tier is Tier.EXCLUDED
|
| 387 |
+
assert v.admitted is False
|
| 388 |
+
assert any("EXCLUDED" in r for r in v.reasons)
|
| 389 |
+
|
| 390 |
+
|
| 391 |
+
def test_gate_trainable_only_admitted_with_reason(tmp_path: Path):
|
| 392 |
+
"""D-13: weak copyleft is admitted for training, but the verdict must
|
| 393 |
+
carry the do-not-redistribute constraint for step 6 to route on."""
|
| 394 |
+
v = gate_repo("example/mpl-repo", _repo_with_license(tmp_path, MPL2_TEXT))
|
| 395 |
+
assert v.tier is Tier.TRAINABLE_ONLY
|
| 396 |
+
assert v.admitted is True
|
| 397 |
+
assert any("TRAINABLE_ONLY" in r for r in v.reasons)
|
| 398 |
+
assert any("redistributed" in r for r in v.reasons)
|
| 399 |
+
|
| 400 |
+
|
| 401 |
+
def test_gate_no_repo_root_fails_closed():
|
| 402 |
+
"""No repo_root → license undetectable → unknown → EXCLUDED → rejected
|
| 403 |
+
(V9: the gate must default closed, never open)."""
|
| 404 |
+
v = gate_repo("example/unfetched", None)
|
| 405 |
+
assert v.license_info.spdx_id == "unknown"
|
| 406 |
+
assert v.tier is Tier.EXCLUDED
|
| 407 |
+
assert v.admitted is False
|
| 408 |
+
assert any("failing closed" in r for r in v.reasons)
|
| 409 |
+
|
| 410 |
+
|
| 411 |
+
def test_gate_extra_decontamination_list(tmp_path: Path):
|
| 412 |
+
extra = frozenset({"my-eval/secret-benchmark"})
|
| 413 |
+
v = gate_repo(
|
| 414 |
+
"https://github.com/My-Eval/Secret-Benchmark.git",
|
| 415 |
+
_repo_with_license(tmp_path, MIT_TEXT),
|
| 416 |
+
extra_decontamination=extra,
|
| 417 |
+
)
|
| 418 |
+
assert v.contaminated is True
|
| 419 |
+
assert v.admitted is False
|
|
@@ -0,0 +1,103 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Tests for the rollout harness (deepread finding V2 — the SFT-corpus producer)."""
|
| 2 |
+
from __future__ import annotations
|
| 3 |
+
|
| 4 |
+
from composer_replication.datagen.env import FeatureDeletionEnv
|
| 5 |
+
from composer_replication.datagen.rollout_harness import (
|
| 6 |
+
ScriptedPolicy,
|
| 7 |
+
admit,
|
| 8 |
+
collect_trajectory,
|
| 9 |
+
)
|
| 10 |
+
from composer_replication.datagen.sandbox import FakeSandbox
|
| 11 |
+
from composer_replication.datagen.schema import FeatureDeletionTask
|
| 12 |
+
from composer_replication.datagen.trajectory import CanonicalTrajectory, ToolCall
|
| 13 |
+
|
| 14 |
+
|
| 15 |
+
def _task() -> FeatureDeletionTask:
|
| 16 |
+
return FeatureDeletionTask(
|
| 17 |
+
task_id="t1", repo="org/repo", base_commit="abc",
|
| 18 |
+
broken_image="img:1", test_command="pytest -q",
|
| 19 |
+
fail_to_pass=("t/a.py::t1", "t/a.py::t2"),
|
| 20 |
+
pass_to_pass=("t/a.py::keep",),
|
| 21 |
+
)
|
| 22 |
+
|
| 23 |
+
|
| 24 |
+
def _env(outcomes: dict[str, bool]) -> FeatureDeletionEnv:
|
| 25 |
+
return FeatureDeletionEnv(FakeSandbox(test_outcomes=outcomes))
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+
def test_collect_trajectory_full_pass():
|
| 29 |
+
"""Policy 'fixes' the repo via the FakeSandbox set_outcome pseudo-action,
|
| 30 |
+
then submits — grade 1.0, steps record real env transitions."""
|
| 31 |
+
env = _env({"t/a.py::keep": True})
|
| 32 |
+
policy = ScriptedPolicy(actions=[
|
| 33 |
+
ToolCall("set_outcome", {"outcomes": {"t/a.py::t1": True, "t/a.py::t2": True}}),
|
| 34 |
+
"final answer: implemented the feature",
|
| 35 |
+
])
|
| 36 |
+
traj = collect_trajectory(env, _task(), policy)
|
| 37 |
+
assert isinstance(traj, CanonicalTrajectory)
|
| 38 |
+
assert traj.grade == 1.0
|
| 39 |
+
assert traj.guard_ok is True and traj.hacked is False
|
| 40 |
+
assert len(traj.steps) == 2
|
| 41 |
+
assert isinstance(traj.steps[0].action, ToolCall)
|
| 42 |
+
assert traj.steps[0].result == "ok" # env.step observation recorded
|
| 43 |
+
assert traj.provenance["source"] == "rollout_harness"
|
| 44 |
+
|
| 45 |
+
|
| 46 |
+
def test_collect_trajectory_guard_broken_zeroes_reward():
|
| 47 |
+
env = _env({"t/a.py::keep": False}) # functional guard broken
|
| 48 |
+
policy = ScriptedPolicy(actions=[
|
| 49 |
+
ToolCall("set_outcome", {"outcomes": {"t/a.py::t1": True, "t/a.py::t2": True,
|
| 50 |
+
"t/a.py::keep": False}}),
|
| 51 |
+
"done",
|
| 52 |
+
])
|
| 53 |
+
traj = collect_trajectory(env, _task(), policy)
|
| 54 |
+
assert traj.grade == 0.0
|
| 55 |
+
assert traj.guard_ok is False
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
def test_collect_trajectory_near_miss():
|
| 59 |
+
env = _env({"t/a.py::keep": True})
|
| 60 |
+
policy = ScriptedPolicy(actions=[
|
| 61 |
+
ToolCall("set_outcome", {"outcomes": {"t/a.py::t1": True}}), # 1 of 2
|
| 62 |
+
"done",
|
| 63 |
+
])
|
| 64 |
+
traj = collect_trajectory(env, _task(), policy)
|
| 65 |
+
assert traj.grade == 0.5
|
| 66 |
+
assert traj.guard_ok is True
|
| 67 |
+
|
| 68 |
+
|
| 69 |
+
def test_collect_trajectory_max_turns_grades_anyway():
|
| 70 |
+
env = _env({"t/a.py::keep": True})
|
| 71 |
+
looping = ScriptedPolicy(actions=[ToolCall("bash", {"command": "ls"})] * 50)
|
| 72 |
+
traj = collect_trajectory(env, _task(), looping, max_turns=3)
|
| 73 |
+
assert traj.grade is not None # graded despite never submitting
|
| 74 |
+
|
| 75 |
+
|
| 76 |
+
# ---------------------------------------------------------------------
|
| 77 |
+
# Admission routing (typed train-on-all, final report §4)
|
| 78 |
+
# ---------------------------------------------------------------------
|
| 79 |
+
|
| 80 |
+
|
| 81 |
+
def _t(grade, guard_ok=True, hacked=False) -> CanonicalTrajectory:
|
| 82 |
+
return CanonicalTrajectory(task_id="x", grade=grade, guard_ok=guard_ok, hacked=hacked)
|
| 83 |
+
|
| 84 |
+
|
| 85 |
+
def test_admit_routes_clean_pass_to_sft():
|
| 86 |
+
v = admit(_t(1.0))
|
| 87 |
+
assert v.sft_admitted and not v.dpo_candidate and not v.rejected
|
| 88 |
+
|
| 89 |
+
|
| 90 |
+
def test_admit_routes_near_miss_to_dpo():
|
| 91 |
+
v = admit(_t(0.5))
|
| 92 |
+
assert v.dpo_candidate and not v.sft_admitted and not v.rejected
|
| 93 |
+
|
| 94 |
+
|
| 95 |
+
def test_admit_rejects_hacked_even_at_full_grade():
|
| 96 |
+
v = admit(_t(1.0, hacked=True))
|
| 97 |
+
assert v.rejected and "hack monitor flagged" in v.reasons
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
def test_admit_rejects_guard_broken_and_ungraded():
|
| 101 |
+
assert admit(_t(1.0, guard_ok=False)).rejected
|
| 102 |
+
assert admit(_t(None)).rejected
|
| 103 |
+
assert admit(_t(0.0)).rejected
|
|
@@ -0,0 +1,165 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Tests for the SWE-smith adapter (deepread finding V4 — buy-vs-build).
|
| 2 |
+
|
| 3 |
+
The load-bearing coverage: the PATCH-SEMANTICS INVERSION (SWE-smith's patch
|
| 4 |
+
introduces the bug; golden_diff must be its reverse) and the mechanical
|
| 5 |
+
reverse_unified_diff round-trip.
|
| 6 |
+
"""
|
| 7 |
+
from __future__ import annotations
|
| 8 |
+
|
| 9 |
+
import json
|
| 10 |
+
|
| 11 |
+
import pytest
|
| 12 |
+
|
| 13 |
+
from composer_replication.datagen.schema import FeatureDeletionTask
|
| 14 |
+
from composer_replication.datagen.substrates import SweBenchAdapter
|
| 15 |
+
from composer_replication.datagen.swesmith_adapter import (
|
| 16 |
+
UNREVERSED_MARKER,
|
| 17 |
+
SwesmithAdapter,
|
| 18 |
+
load_swesmith_instances,
|
| 19 |
+
parse_strategy,
|
| 20 |
+
reverse_unified_diff,
|
| 21 |
+
)
|
| 22 |
+
|
| 23 |
+
BUG_PATCH = """\
|
| 24 |
+
diff --git a/pkg/mod.py b/pkg/mod.py
|
| 25 |
+
index 1111111..2222222 100644
|
| 26 |
+
--- a/pkg/mod.py
|
| 27 |
+
+++ b/pkg/mod.py
|
| 28 |
+
@@ -1,4 +1,3 @@
|
| 29 |
+
def add(a, b):
|
| 30 |
+
- return a + b
|
| 31 |
+
+ return a - b
|
| 32 |
+
# trailing context
|
| 33 |
+
"""
|
| 34 |
+
|
| 35 |
+
|
| 36 |
+
def _instance(**over) -> dict:
|
| 37 |
+
base = {
|
| 38 |
+
"instance_id": "getmoto__moto.abc1234.lm_modify__1a2b",
|
| 39 |
+
"repo": "getmoto/moto",
|
| 40 |
+
"base_commit": "abc1234",
|
| 41 |
+
"patch": BUG_PATCH,
|
| 42 |
+
"FAIL_TO_PASS": json.dumps(["tests/test_mod.py::test_add"]),
|
| 43 |
+
"PASS_TO_PASS": json.dumps(["tests/test_mod.py::test_other"]),
|
| 44 |
+
"image_name": "swesmith.x86_64.getmoto__moto:latest",
|
| 45 |
+
}
|
| 46 |
+
base.update(over)
|
| 47 |
+
return base
|
| 48 |
+
|
| 49 |
+
|
| 50 |
+
# ---------------------------------------------------------------------
|
| 51 |
+
# Strategy parsing
|
| 52 |
+
# ---------------------------------------------------------------------
|
| 53 |
+
|
| 54 |
+
|
| 55 |
+
@pytest.mark.parametrize("iid,expected", [
|
| 56 |
+
("r__x.abc.lm_modify__1", "lm_modify"),
|
| 57 |
+
("r__x.abc.lm_rewrite__1", "lm_rewrite"),
|
| 58 |
+
("r__x.abc.func_pm_ctrl_invert_if__1", "procedural"),
|
| 59 |
+
("r__x.abc.func_basic__1", "procedural"),
|
| 60 |
+
("r__x.abc.combine_file__1", "combine"),
|
| 61 |
+
("r__x.abc.combine_module__2", "combine"),
|
| 62 |
+
("r__x.abc.pr_1234", "pr_mirror"),
|
| 63 |
+
("r__x.abc.mystery__1", "unknown"),
|
| 64 |
+
])
|
| 65 |
+
def test_parse_strategy(iid, expected):
|
| 66 |
+
assert parse_strategy(iid) == expected
|
| 67 |
+
|
| 68 |
+
|
| 69 |
+
# ---------------------------------------------------------------------
|
| 70 |
+
# reverse_unified_diff
|
| 71 |
+
# ---------------------------------------------------------------------
|
| 72 |
+
|
| 73 |
+
|
| 74 |
+
def test_reverse_swaps_adds_and_removes():
|
| 75 |
+
rev = reverse_unified_diff(BUG_PATCH)
|
| 76 |
+
assert rev is not None
|
| 77 |
+
# The bug ADDED "return a - b"; the reverse must REMOVE it.
|
| 78 |
+
assert "- return a - b" in rev
|
| 79 |
+
assert "+ return a + b" in rev
|
| 80 |
+
# Hunk header ranges swapped: -1,4 +1,3 → -1,3 +1,4
|
| 81 |
+
assert "@@ -1,3 +1,4 @@" in rev
|
| 82 |
+
# Context lines untouched.
|
| 83 |
+
assert " def add(a, b):" in rev
|
| 84 |
+
assert " # trailing context" in rev
|
| 85 |
+
|
| 86 |
+
|
| 87 |
+
def test_reverse_round_trip_is_identity_on_body():
|
| 88 |
+
rev = reverse_unified_diff(BUG_PATCH)
|
| 89 |
+
rev2 = reverse_unified_diff(rev)
|
| 90 |
+
# Round trip restores hunks and +/- bodies (headers may normalize).
|
| 91 |
+
orig_body = [ln for ln in BUG_PATCH.splitlines() if ln[:1] in "+-@" and not ln.startswith(("+++", "---"))]
|
| 92 |
+
rt_body = [ln for ln in rev2.splitlines() if ln[:1] in "+-@" and not ln.startswith(("+++", "---"))]
|
| 93 |
+
assert orig_body == rt_body
|
| 94 |
+
|
| 95 |
+
|
| 96 |
+
def test_reverse_refuses_renames_and_binary():
|
| 97 |
+
assert reverse_unified_diff("diff --git a/x b/y\nrename from x\nrename to y\n") is None
|
| 98 |
+
assert reverse_unified_diff("diff --git a/x b/x\nGIT binary patch\nliteral 5\n") is None
|
| 99 |
+
assert reverse_unified_diff("") is None
|
| 100 |
+
assert reverse_unified_diff("no hunks here") is None
|
| 101 |
+
|
| 102 |
+
|
| 103 |
+
# ---------------------------------------------------------------------
|
| 104 |
+
# Adapter
|
| 105 |
+
# ---------------------------------------------------------------------
|
| 106 |
+
|
| 107 |
+
|
| 108 |
+
def test_to_task_golden_diff_is_the_fix_not_the_bug():
|
| 109 |
+
"""THE semantic inversion: golden_diff must restore the feature."""
|
| 110 |
+
task, meta = SwesmithAdapter().to_task_with_meta(_instance())
|
| 111 |
+
assert isinstance(task, FeatureDeletionTask)
|
| 112 |
+
assert meta.diff_reversed is True
|
| 113 |
+
assert meta.strategy == "lm_modify"
|
| 114 |
+
# The FIX restores `a + b` (adds it back) and removes the bug.
|
| 115 |
+
assert "+ return a + b" in task.golden_diff
|
| 116 |
+
assert "- return a - b" in task.golden_diff
|
| 117 |
+
assert UNREVERSED_MARKER not in task.golden_diff
|
| 118 |
+
|
| 119 |
+
|
| 120 |
+
def test_to_task_unreversible_patch_gets_marker():
|
| 121 |
+
inst = _instance(patch="diff --git a/x b/y\nrename from x\nrename to y\n@@ -1 +1 @@\n-a\n+b\n")
|
| 122 |
+
task, meta = SwesmithAdapter().to_task_with_meta(inst)
|
| 123 |
+
assert meta.diff_reversed is False
|
| 124 |
+
assert task.golden_diff.startswith(UNREVERSED_MARKER)
|
| 125 |
+
|
| 126 |
+
|
| 127 |
+
def test_image_resolution_prefers_instance_field_then_convention():
|
| 128 |
+
a = SwesmithAdapter()
|
| 129 |
+
assert a.image_for(_instance()) == "swesmith.x86_64.getmoto__moto:latest"
|
| 130 |
+
assert a.image_for(_instance(image_name=None, docker_image="custom:tag")) == "custom:tag"
|
| 131 |
+
inst = _instance(image_name=None)
|
| 132 |
+
inst.pop("docker_image", None)
|
| 133 |
+
assert a.image_for(inst) == "swesmith.x86_64.getmoto__moto:latest"
|
| 134 |
+
|
| 135 |
+
|
| 136 |
+
def test_f2p_p2p_tuple_handling_matches_swebench_semantics():
|
| 137 |
+
task = SwesmithAdapter().to_task(_instance(
|
| 138 |
+
FAIL_TO_PASS=["t/a.py::t1", "t/a.py::t2"], # real list, not JSON string
|
| 139 |
+
PASS_TO_PASS=json.dumps([]),
|
| 140 |
+
))
|
| 141 |
+
assert task.fail_to_pass == ("t/a.py::t1", "t/a.py::t2")
|
| 142 |
+
assert task.pass_to_pass == ()
|
| 143 |
+
|
| 144 |
+
|
| 145 |
+
def test_difficulty_priors_by_strategy():
|
| 146 |
+
pr = SwesmithAdapter().to_task(_instance(instance_id="r__x.abc.pr_99"))
|
| 147 |
+
proc = SwesmithAdapter().to_task(_instance(instance_id="r__x.abc.func_pm_remove_loop__1"))
|
| 148 |
+
assert pr.difficulty_prior < proc.difficulty_prior # PR Mirror harder prior
|
| 149 |
+
|
| 150 |
+
|
| 151 |
+
def test_redistributable_filter_interplay():
|
| 152 |
+
"""repo_gate owns repo-level licensing, but the per-instance filter from
|
| 153 |
+
SweBenchAdapter still composes when a license field IS present."""
|
| 154 |
+
task = SwesmithAdapter().to_task(_instance(license_name="GPL-3.0"))
|
| 155 |
+
assert SweBenchAdapter.is_redistributable(task) is False
|
| 156 |
+
task2 = SwesmithAdapter().to_task(_instance(license_name="MIT"))
|
| 157 |
+
assert SweBenchAdapter.is_redistributable(task2) is True
|
| 158 |
+
|
| 159 |
+
|
| 160 |
+
def test_load_local_jsonl(tmp_path):
|
| 161 |
+
p = tmp_path / "fixtures.jsonl"
|
| 162 |
+
p.write_text("\n".join(json.dumps(_instance(instance_id=f"r__x.abc.pr_{i}")) for i in range(5)))
|
| 163 |
+
rows = load_swesmith_instances(str(p), limit=3)
|
| 164 |
+
assert len(rows) == 3
|
| 165 |
+
assert rows[0]["instance_id"] == "r__x.abc.pr_0"
|
|
@@ -0,0 +1,127 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Tests for the canonical trajectory IR (deepread findings V2/D-11/D-8).
|
| 2 |
+
|
| 3 |
+
The load-bearing test is the SENTINEL leak guard: `to_policy_row` must never
|
| 4 |
+
emit golden_diff/deleted_symbols, even though the task dataclass carries them.
|
| 5 |
+
"""
|
| 6 |
+
from __future__ import annotations
|
| 7 |
+
|
| 8 |
+
import json
|
| 9 |
+
|
| 10 |
+
from composer_replication.datagen.schema import FeatureDeletionTask
|
| 11 |
+
from composer_replication.datagen.trajectory import (
|
| 12 |
+
CanonicalTrajectory,
|
| 13 |
+
ToolCall,
|
| 14 |
+
TrajectoryStep,
|
| 15 |
+
from_trace_states,
|
| 16 |
+
to_policy_row,
|
| 17 |
+
to_sft_messages,
|
| 18 |
+
)
|
| 19 |
+
from composer_replication.teacher_replay import TraceState
|
| 20 |
+
|
| 21 |
+
|
| 22 |
+
def _task(**over) -> FeatureDeletionTask:
|
| 23 |
+
base = dict(
|
| 24 |
+
task_id="t1", repo="org/repo", base_commit="abc",
|
| 25 |
+
broken_image="img:1", test_command="pytest -q",
|
| 26 |
+
fail_to_pass=("t/a.py::t1",), pass_to_pass=("t/a.py::t2",),
|
| 27 |
+
golden_diff="SENTINEL_NEVER_LEAK",
|
| 28 |
+
deleted_symbols=("secret_fn",),
|
| 29 |
+
)
|
| 30 |
+
base.update(over)
|
| 31 |
+
return FeatureDeletionTask(**base)
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
def _traj() -> CanonicalTrajectory:
|
| 35 |
+
return CanonicalTrajectory(
|
| 36 |
+
task_id="t1",
|
| 37 |
+
steps=[
|
| 38 |
+
TrajectoryStep(observation="repo is broken", action=ToolCall("bash", {"command": "pytest"}),
|
| 39 |
+
result="2 failed", tool_error=False),
|
| 40 |
+
TrajectoryStep(observation="2 failed", action="here is my final patch",
|
| 41 |
+
result="graded", tool_error=False),
|
| 42 |
+
],
|
| 43 |
+
grade=1.0, guard_ok=True, hacked=False,
|
| 44 |
+
provenance={"source": "test"},
|
| 45 |
+
)
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
# ---------------------------------------------------------------------
|
| 49 |
+
# ToolCall canonical form — the v1 divergence algebra
|
| 50 |
+
# ---------------------------------------------------------------------
|
| 51 |
+
|
| 52 |
+
|
| 53 |
+
def test_canonical_form_is_order_insensitive_on_args():
|
| 54 |
+
a = ToolCall("edit", {"path": "x.py", "content": "y"})
|
| 55 |
+
b = ToolCall("edit", {"content": "y", "path": "x.py"})
|
| 56 |
+
assert a.canonical_form() == b.canonical_form()
|
| 57 |
+
|
| 58 |
+
|
| 59 |
+
def test_canonical_form_distinguishes_name_and_args():
|
| 60 |
+
assert ToolCall("bash", {"command": "ls"}).canonical_form() != \
|
| 61 |
+
ToolCall("bash", {"command": "ls -la"}).canonical_form()
|
| 62 |
+
assert ToolCall("read", {"f": "x"}).canonical_form() != \
|
| 63 |
+
ToolCall("write", {"f": "x"}).canonical_form()
|
| 64 |
+
|
| 65 |
+
|
| 66 |
+
# ---------------------------------------------------------------------
|
| 67 |
+
# THE leak guard (finding D-8)
|
| 68 |
+
# ---------------------------------------------------------------------
|
| 69 |
+
|
| 70 |
+
|
| 71 |
+
def test_policy_row_never_contains_golden_diff_or_deleted_symbols():
|
| 72 |
+
row = to_policy_row(_traj(), _task())
|
| 73 |
+
blob = json.dumps(row)
|
| 74 |
+
assert "SENTINEL_NEVER_LEAK" not in blob
|
| 75 |
+
assert "secret_fn" not in blob
|
| 76 |
+
assert "golden_diff" not in blob
|
| 77 |
+
assert "deleted_symbols" not in blob
|
| 78 |
+
# And the row still carries what the policy MAY see.
|
| 79 |
+
assert row["repo"] == "org/repo"
|
| 80 |
+
assert row["fail_to_pass"] == ["t/a.py::t1"]
|
| 81 |
+
assert row["grade"] == 1.0
|
| 82 |
+
|
| 83 |
+
|
| 84 |
+
# ---------------------------------------------------------------------
|
| 85 |
+
# IR ↔ SFT messages
|
| 86 |
+
# ---------------------------------------------------------------------
|
| 87 |
+
|
| 88 |
+
|
| 89 |
+
def test_to_sft_messages_alternates_roles():
|
| 90 |
+
msgs = to_sft_messages(_traj())
|
| 91 |
+
assert msgs[0] == {"role": "user", "content": "repo is broken"}
|
| 92 |
+
assert msgs[1]["role"] == "assistant"
|
| 93 |
+
assert "[TOOL_USE] name=bash" in msgs[1]["content"]
|
| 94 |
+
assert msgs[2] == {"role": "user", "content": "2 failed"}
|
| 95 |
+
|
| 96 |
+
|
| 97 |
+
# ---------------------------------------------------------------------
|
| 98 |
+
# Claude Code → IR adapter
|
| 99 |
+
# ---------------------------------------------------------------------
|
| 100 |
+
|
| 101 |
+
|
| 102 |
+
def test_from_trace_states_parses_single_tool_use_and_error_flag():
|
| 103 |
+
states: list[TraceState] = [
|
| 104 |
+
{
|
| 105 |
+
"state_id": "sess1::0000",
|
| 106 |
+
"messages": [
|
| 107 |
+
{"role": "system", "content": "sys"},
|
| 108 |
+
{"role": "user", "content": "[TOOL_RESULT (ERROR)] (id=x)\nboom",
|
| 109 |
+
"tool_error": True},
|
| 110 |
+
],
|
| 111 |
+
"student_action": '[TOOL_USE] name=Bash input={"command":"ls"}',
|
| 112 |
+
},
|
| 113 |
+
{
|
| 114 |
+
"state_id": "sess1::0001",
|
| 115 |
+
"messages": [{"role": "user", "content": "plain prompt"}],
|
| 116 |
+
"student_action": "I think the fix is...\n\n[TOOL_USE] name=Edit input={\"p\":1}\n\n[TOOL_USE] name=Bash input={\"c\":2}",
|
| 117 |
+
},
|
| 118 |
+
]
|
| 119 |
+
traj = from_trace_states(states)
|
| 120 |
+
assert traj.task_id == "sess1"
|
| 121 |
+
assert traj.grade is None # ungraded — Claude Code traces have no oracle
|
| 122 |
+
s0, s1 = traj.steps
|
| 123 |
+
assert isinstance(s0.action, ToolCall) and s0.action.name == "Bash"
|
| 124 |
+
assert s0.tool_error is True
|
| 125 |
+
# Multi-tool turn stays as the raw string (honest, not guessed).
|
| 126 |
+
assert isinstance(s1.action, str)
|
| 127 |
+
assert s1.tool_error is False
|
|
@@ -0,0 +1,203 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""trajectory.py — the canonical trajectory IR (deepread findings V2/D-11/D-8).
|
| 2 |
+
|
| 3 |
+
THE GAP THIS CLOSES: the repo had 3 (heading to 5) incompatible trajectory
|
| 4 |
+
shapes — Claude Code TraceState text serialization, Bedrock `.jsonl.out` rows,
|
| 5 |
+
the planned tree/rollout/OpenHands shapes — with no shared schema, and the
|
| 6 |
+
divergence gate rested on a whitespace-collapse string normalizer
|
| 7 |
+
(`teacher_replay._normalize_action`, self-admitted skeleton). This module is the
|
| 8 |
+
single intermediate representation every producer adapts INTO and every corpus
|
| 9 |
+
writer reads FROM.
|
| 10 |
+
|
| 11 |
+
SECURITY INVARIANT (finding D-8): `FeatureDeletionTask.golden_diff` uses
|
| 12 |
+
`repr=False`, but `dataclasses.asdict()` and naive JSON serialization still
|
| 13 |
+
include it. `to_policy_row()` is the ONE serializer allowed to produce
|
| 14 |
+
policy-visible rows, and its output is unit-tested to never contain
|
| 15 |
+
`golden_diff` or `deleted_symbols`.
|
| 16 |
+
"""
|
| 17 |
+
from __future__ import annotations
|
| 18 |
+
|
| 19 |
+
import json
|
| 20 |
+
import re
|
| 21 |
+
from dataclasses import dataclass, field
|
| 22 |
+
from typing import Any
|
| 23 |
+
|
| 24 |
+
from composer_replication.datagen.schema import FeatureDeletionTask
|
| 25 |
+
from composer_replication.teacher_replay import TraceState
|
| 26 |
+
|
| 27 |
+
#: Bump when the IR shape changes; carried on every trajectory + corpus row.
|
| 28 |
+
SCHEMA_VERSION = "1"
|
| 29 |
+
|
| 30 |
+
|
| 31 |
+
@dataclass(frozen=True)
|
| 32 |
+
class ToolCall:
|
| 33 |
+
"""One structured tool invocation — the unit of the action algebra.
|
| 34 |
+
|
| 35 |
+
`canonical_form()` is the v1 divergence-gate algebra (finding D-3): two
|
| 36 |
+
actions are "the same" iff their canonical forms match (tool name + sorted,
|
| 37 |
+
JSON-normalized args). This replaces the whitespace-collapse stub that made
|
| 38 |
+
the divergence gate fire on noise. v2 will need arg-level normalization
|
| 39 |
+
(path equivalence, whitespace-insensitive code args); keep that evolution
|
| 40 |
+
HERE so every consumer inherits it.
|
| 41 |
+
"""
|
| 42 |
+
|
| 43 |
+
name: str
|
| 44 |
+
args: dict = field(default_factory=dict)
|
| 45 |
+
|
| 46 |
+
def canonical_form(self) -> str:
|
| 47 |
+
try:
|
| 48 |
+
args_json = json.dumps(self.args, sort_keys=True, separators=(",", ":"))
|
| 49 |
+
except (TypeError, ValueError):
|
| 50 |
+
args_json = str(self.args)
|
| 51 |
+
return f"{self.name}:{args_json}"
|
| 52 |
+
|
| 53 |
+
|
| 54 |
+
@dataclass
|
| 55 |
+
class TrajectoryStep:
|
| 56 |
+
"""One agent turn: what it saw, what it did, what came back."""
|
| 57 |
+
|
| 58 |
+
observation: str
|
| 59 |
+
action: ToolCall | str # str = plain text / final message
|
| 60 |
+
result: str | None = None # tool output observed AFTER the action
|
| 61 |
+
tool_error: bool = False
|
| 62 |
+
|
| 63 |
+
|
| 64 |
+
@dataclass
|
| 65 |
+
class CanonicalTrajectory:
|
| 66 |
+
"""The IR: an episode (or trace) as a typed step list + outcome + provenance."""
|
| 67 |
+
|
| 68 |
+
task_id: str
|
| 69 |
+
steps: list[TrajectoryStep] = field(default_factory=list)
|
| 70 |
+
grade: float | None = None # _grade() pass-fraction; None = ungraded trace
|
| 71 |
+
guard_ok: bool = True
|
| 72 |
+
hacked: bool = False
|
| 73 |
+
provenance: dict = field(default_factory=dict) # source, policy id, cost_usd, run_id
|
| 74 |
+
schema_version: str = SCHEMA_VERSION
|
| 75 |
+
|
| 76 |
+
|
| 77 |
+
# ---------------------------------------------------------------------
|
| 78 |
+
# Producers → IR
|
| 79 |
+
# ---------------------------------------------------------------------
|
| 80 |
+
|
| 81 |
+
# ClaudeCodeIngester serializes tool calls as "[TOOL_USE] name=<n> input=<json>".
|
| 82 |
+
_TOOL_USE_RE = re.compile(r"^\[TOOL_USE\] name=(?P<name>\S+) input=(?P<input>\{.*\})$")
|
| 83 |
+
|
| 84 |
+
|
| 85 |
+
def _parse_action(student_action: str) -> ToolCall | str:
|
| 86 |
+
"""Parse a TraceState student_action back into a ToolCall where possible.
|
| 87 |
+
|
| 88 |
+
Claude Code assistant turns serialize as newline-joined blocks; if exactly
|
| 89 |
+
one [TOOL_USE] block is present we recover the structured call (the common
|
| 90 |
+
case ADR-002 chose one-node-per-turn for). Multi-tool turns and pure-text /
|
| 91 |
+
thinking turns stay as the raw string — honest about what we can't
|
| 92 |
+
structure rather than guessing.
|
| 93 |
+
"""
|
| 94 |
+
blocks = [b for b in student_action.split("\n\n") if b.strip()]
|
| 95 |
+
tool_blocks = [b for b in blocks if b.startswith("[TOOL_USE]")]
|
| 96 |
+
if len(tool_blocks) == 1:
|
| 97 |
+
m = _TOOL_USE_RE.match(tool_blocks[0].strip())
|
| 98 |
+
if m:
|
| 99 |
+
try:
|
| 100 |
+
return ToolCall(name=m.group("name"), args=json.loads(m.group("input")))
|
| 101 |
+
except (json.JSONDecodeError, ValueError):
|
| 102 |
+
pass
|
| 103 |
+
return student_action
|
| 104 |
+
|
| 105 |
+
|
| 106 |
+
def from_trace_states(
|
| 107 |
+
states: list[TraceState],
|
| 108 |
+
*,
|
| 109 |
+
task_id: str = "",
|
| 110 |
+
provenance: dict | None = None,
|
| 111 |
+
) -> CanonicalTrajectory:
|
| 112 |
+
"""Adapt a Claude Code trace (TraceState list) into the IR.
|
| 113 |
+
|
| 114 |
+
HONEST CAPABILITY NOTE (finding D-1): these traces carry no executable
|
| 115 |
+
environment — no broken_image, no FAIL_TO_PASS — so the resulting
|
| 116 |
+
trajectory is UNGRADED (`grade=None`) and is admissible only for flat
|
| 117 |
+
Channel-3 / SFT-style uses, never as a tree seed. Env-grounded rollouts
|
| 118 |
+
(rollout_harness.collect_trajectory) are the graded producers.
|
| 119 |
+
"""
|
| 120 |
+
steps: list[TrajectoryStep] = []
|
| 121 |
+
for s in states:
|
| 122 |
+
# The observation for step t is the last user message before the turn.
|
| 123 |
+
obs = ""
|
| 124 |
+
tool_error = False
|
| 125 |
+
for msg in reversed(s["messages"]):
|
| 126 |
+
if msg.get("role") == "user":
|
| 127 |
+
obs = str(msg.get("content", ""))
|
| 128 |
+
tool_error = bool(msg.get("tool_error", False))
|
| 129 |
+
break
|
| 130 |
+
steps.append(TrajectoryStep(
|
| 131 |
+
observation=obs,
|
| 132 |
+
action=_parse_action(s["student_action"]),
|
| 133 |
+
result=None,
|
| 134 |
+
tool_error=tool_error,
|
| 135 |
+
))
|
| 136 |
+
prov = {"source": "claude_code", **(provenance or {})}
|
| 137 |
+
return CanonicalTrajectory(task_id=task_id or (states[0]["state_id"].split("::")[0] if states else ""),
|
| 138 |
+
steps=steps, grade=None, provenance=prov)
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
# ---------------------------------------------------------------------
|
| 142 |
+
# IR → consumers
|
| 143 |
+
# ---------------------------------------------------------------------
|
| 144 |
+
|
| 145 |
+
|
| 146 |
+
def _action_text(action: ToolCall | str) -> str:
|
| 147 |
+
if isinstance(action, ToolCall):
|
| 148 |
+
return f"[TOOL_USE] name={action.name} input=" + json.dumps(
|
| 149 |
+
action.args, separators=(",", ":")
|
| 150 |
+
)
|
| 151 |
+
return action
|
| 152 |
+
|
| 153 |
+
|
| 154 |
+
def to_sft_messages(traj: CanonicalTrajectory) -> list[dict]:
|
| 155 |
+
"""IR → OpenAI-style messages for SFT (obs→user, action→assistant)."""
|
| 156 |
+
messages: list[dict] = []
|
| 157 |
+
for step in traj.steps:
|
| 158 |
+
if step.observation:
|
| 159 |
+
messages.append({"role": "user", "content": step.observation})
|
| 160 |
+
messages.append({"role": "assistant", "content": _action_text(step.action)})
|
| 161 |
+
if step.result:
|
| 162 |
+
messages.append({"role": "user", "content": step.result})
|
| 163 |
+
return messages
|
| 164 |
+
|
| 165 |
+
|
| 166 |
+
#: Task fields the POLICY may see. Everything else (golden_diff,
|
| 167 |
+
#: deleted_symbols) is construction-side and must never reach a corpus row.
|
| 168 |
+
_POLICY_VISIBLE_TASK_FIELDS = (
|
| 169 |
+
"task_id", "repo", "base_commit", "test_command",
|
| 170 |
+
"fail_to_pass", "pass_to_pass", "granularity", "difficulty_prior",
|
| 171 |
+
)
|
| 172 |
+
|
| 173 |
+
|
| 174 |
+
def to_policy_row(traj: CanonicalTrajectory, task: FeatureDeletionTask) -> dict:
|
| 175 |
+
"""THE policy-visible corpus serializer (finding D-8 — the leak guard).
|
| 176 |
+
|
| 177 |
+
Builds the row field-by-field from an allowlist; never `asdict(task)`,
|
| 178 |
+
which would include `golden_diff` despite its `repr=False`. Unit-tested
|
| 179 |
+
with a sentinel to prove the absence.
|
| 180 |
+
"""
|
| 181 |
+
row: dict[str, Any] = {
|
| 182 |
+
"schema_version": traj.schema_version,
|
| 183 |
+
"messages": to_sft_messages(traj),
|
| 184 |
+
"grade": traj.grade,
|
| 185 |
+
"guard_ok": traj.guard_ok,
|
| 186 |
+
"hacked": traj.hacked,
|
| 187 |
+
"provenance": dict(traj.provenance),
|
| 188 |
+
}
|
| 189 |
+
for f in _POLICY_VISIBLE_TASK_FIELDS:
|
| 190 |
+
v = getattr(task, f)
|
| 191 |
+
row[f] = list(v) if isinstance(v, tuple) else v
|
| 192 |
+
return row
|
| 193 |
+
|
| 194 |
+
|
| 195 |
+
__all__ = [
|
| 196 |
+
"SCHEMA_VERSION",
|
| 197 |
+
"ToolCall",
|
| 198 |
+
"TrajectoryStep",
|
| 199 |
+
"CanonicalTrajectory",
|
| 200 |
+
"from_trace_states",
|
| 201 |
+
"to_sft_messages",
|
| 202 |
+
"to_policy_row",
|
| 203 |
+
]
|
|
@@ -4,9 +4,12 @@ Wraps `torchft.local_sgd.DiLoCo` with the framework's conventions:
|
|
| 4 |
- Sign convention is documented LOUDLY here once and tested via Spike 008.
|
| 5 |
- The wrapper exposes the same constructor shape as torchft's DiLoCo so a
|
| 6 |
future swap-in of the upstream class is a one-line change.
|
| 7 |
-
- Vanilla DiLoCo (Douillard et al. 2023) =
|
| 8 |
-
fragment. Streaming DiLoCo (
|
| 9 |
-
|
|
|
|
|
|
|
|
|
|
| 10 |
|
| 11 |
Reference: `docs/adrs/ADR-003-diloco-impl.md`.
|
| 12 |
|
|
|
|
| 4 |
- Sign convention is documented LOUDLY here once and tested via Spike 008.
|
| 5 |
- The wrapper exposes the same constructor shape as torchft's DiLoCo so a
|
| 6 |
future swap-in of the upstream class is a one-line change.
|
| 7 |
+
- Vanilla DiLoCo (Douillard et al. 2023, arXiv:2311.08105) =
|
| 8 |
+
`fragment_sync_delay=0`, single fragment. Streaming DiLoCo (Douillard et
|
| 9 |
+
al., arXiv:2501.18512 "Streaming DiLoCo with overlapping communication";
|
| 10 |
+
the separate Eager-Updates work is Kale et al., arXiv:2502.12996 — citation
|
| 11 |
+
corrected per deepread finding V7) = non-zero delay, multiple fragments.
|
| 12 |
+
Spike 008 uses vanilla; Streaming is configured by the same API.
|
| 13 |
|
| 14 |
Reference: `docs/adrs/ADR-003-diloco-impl.md`.
|
| 15 |
|
|
@@ -10,17 +10,23 @@ Mathematical reference:
|
|
| 10 |
- OPSD paper: Zhao et al., "Self-Distilled Reasoner: On-Policy Self-Distillation
|
| 11 |
for LLMs", arXiv:2601.18734.
|
| 12 |
- SDPO paper: Hübotter et al., "Reinforcement Learning via Self-Distillation",
|
| 13 |
-
arXiv:2601.20802 (
|
| 14 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 15 |
|
| 16 |
The loss computes JSD/KL divergence between a teacher distribution (model
|
| 17 |
conditioned on privileged information / a hint) and a student distribution
|
| 18 |
(model on the original context). Both come from the SAME model — the teacher
|
| 19 |
is just "the model with hint inserted into context."
|
| 20 |
|
| 21 |
-
Composer 2.5
|
| 22 |
-
|
| 23 |
-
ctx_teacher = ctx_student +
|
|
|
|
| 24 |
"""
|
| 25 |
|
| 26 |
from __future__ import annotations
|
|
|
|
| 10 |
- OPSD paper: Zhao et al., "Self-Distilled Reasoner: On-Policy Self-Distillation
|
| 11 |
for LLMs", arXiv:2601.18734.
|
| 12 |
- SDPO paper: Hübotter et al., "Reinforcement Learning via Self-Distillation",
|
| 13 |
+
arXiv:2601.20802. PROVENANCE (corrected per deepread finding V1): Cursor's
|
| 14 |
+
blog cites SDPO/OPSD only as *background* ("For more background on this
|
| 15 |
+
approach see…"), NOT as its mechanism. Published SDPO distills over the FULL
|
| 16 |
+
rollout with feedback in the prefix and an EMA-regularized teacher; this
|
| 17 |
+
repo's channel is a turn-localized hint-splice with a live (stop-grad,
|
| 18 |
+
non-EMA) teacher — a third, blog-inspired design, neither verbatim SDPO nor
|
| 19 |
+
confirmed-Composer. The kernel below matches OPSD's generalized JSD math.
|
| 20 |
|
| 21 |
The loss computes JSD/KL divergence between a teacher distribution (model
|
| 22 |
conditioned on privileged information / a hint) and a student distribution
|
| 23 |
(model on the original context). Both come from the SAME model — the teacher
|
| 24 |
is just "the model with hint inserted into context."
|
| 25 |
|
| 26 |
+
Composer 2.5's blog describes inserting a "hint" at the error-turn site and
|
| 27 |
+
distilling the student toward the hint-conditioned distribution "for that turn
|
| 28 |
+
only". The data collator constructs ctx_teacher = ctx_student +
|
| 29 |
+
hint_at_error_turn for us.
|
| 30 |
"""
|
| 31 |
|
| 32 |
from __future__ import annotations
|
|
@@ -0,0 +1,38 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""composer_replication.pipeline — the Stage-0 dataset-pipeline contract + driver.
|
| 2 |
+
|
| 3 |
+
THE single reconciled dataset contract (supersedes the two divergent layouts in
|
| 4 |
+
research/design-F1 and design-F2 — deepread finding V8/D-7), the pragmatic
|
| 5 |
+
near-duplicate detector, and the local stage-driver that turns
|
| 6 |
+
(tasks, env, policy) into a carded, deduped, holdout-split corpus.
|
| 7 |
+
"""
|
| 8 |
+
from composer_replication.pipeline.build_corpus import build_corpus
|
| 9 |
+
from composer_replication.pipeline.dedup import (
|
| 10 |
+
dedup,
|
| 11 |
+
find_near_duplicates,
|
| 12 |
+
jaccard_estimate,
|
| 13 |
+
minhash_signature,
|
| 14 |
+
)
|
| 15 |
+
from composer_replication.pipeline.s3_contract import (
|
| 16 |
+
RunLayout,
|
| 17 |
+
RunManifest,
|
| 18 |
+
write_dataset_card,
|
| 19 |
+
write_dpo_rows,
|
| 20 |
+
write_sft_rows,
|
| 21 |
+
write_tasks,
|
| 22 |
+
write_tasks_full,
|
| 23 |
+
)
|
| 24 |
+
|
| 25 |
+
__all__ = [
|
| 26 |
+
"RunLayout",
|
| 27 |
+
"RunManifest",
|
| 28 |
+
"build_corpus",
|
| 29 |
+
"dedup",
|
| 30 |
+
"find_near_duplicates",
|
| 31 |
+
"jaccard_estimate",
|
| 32 |
+
"minhash_signature",
|
| 33 |
+
"write_dataset_card",
|
| 34 |
+
"write_dpo_rows",
|
| 35 |
+
"write_sft_rows",
|
| 36 |
+
"write_tasks",
|
| 37 |
+
"write_tasks_full",
|
| 38 |
+
]
|
|
@@ -0,0 +1,137 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""build_corpus.py — the local Stage-0 stage-driver (architecture step 6-7).
|
| 2 |
+
|
| 3 |
+
One function wires the whole local pipeline: holdout-split the task pool
|
| 4 |
+
(holdout tasks are NEVER rolled out — they are the eval anchor), roll out a
|
| 5 |
+
policy over each train task, admit + type + route trajectories
|
| 6 |
+
(sft / dpo-candidate / quarantine), dedup the SFT rows (within-run AND against
|
| 7 |
+
a prior generation's signatures), and write everything through the single
|
| 8 |
+
s3_contract layout with a manifest + dataset card.
|
| 9 |
+
|
| 10 |
+
Deliberately LOCAL-first (finding D-9): the five-service AWS orchestration is
|
| 11 |
+
Stage 4; this driver must produce one real corpus end-to-end on a laptop with a
|
| 12 |
+
FakeSandbox before anything is distributed. Write-once per layout
|
| 13 |
+
(finding D-21): refuses to run if the manifest already exists.
|
| 14 |
+
"""
|
| 15 |
+
from __future__ import annotations
|
| 16 |
+
|
| 17 |
+
from typing import Callable, Sequence
|
| 18 |
+
|
| 19 |
+
from composer_replication.datagen.env import FeatureDeletionEnv
|
| 20 |
+
from composer_replication.datagen.rollout_harness import (
|
| 21 |
+
RolloutPolicy,
|
| 22 |
+
admit,
|
| 23 |
+
collect_trajectory,
|
| 24 |
+
)
|
| 25 |
+
from composer_replication.datagen.schema import FeatureDeletionTask
|
| 26 |
+
from composer_replication.datagen.trajectory import to_policy_row
|
| 27 |
+
from composer_replication.pipeline import s3_contract
|
| 28 |
+
from composer_replication.pipeline.dedup import dedup
|
| 29 |
+
from composer_replication.pipeline.s3_contract import RunLayout, RunManifest
|
| 30 |
+
from composer_replication.safety.holdout import HeldoutSplit
|
| 31 |
+
|
| 32 |
+
|
| 33 |
+
def build_corpus(
|
| 34 |
+
source_tasks: Sequence[FeatureDeletionTask],
|
| 35 |
+
env_factory: Callable[[], FeatureDeletionEnv],
|
| 36 |
+
policy_factory: Callable[[], RolloutPolicy],
|
| 37 |
+
layout: RunLayout,
|
| 38 |
+
manifest: RunManifest,
|
| 39 |
+
*,
|
| 40 |
+
holdout_frac: float = 0.2,
|
| 41 |
+
holdout_seed: int = 0,
|
| 42 |
+
max_tasks: int | None = None,
|
| 43 |
+
cost_per_rollout_usd: float = 0.0,
|
| 44 |
+
prior_signatures: Sequence[Sequence[int]] | None = None,
|
| 45 |
+
dedup_threshold: float = 0.85,
|
| 46 |
+
) -> RunManifest:
|
| 47 |
+
"""Run the Stage-0 pipeline over `source_tasks`; returns the final manifest.
|
| 48 |
+
|
| 49 |
+
Args:
|
| 50 |
+
source_tasks: gate_repo-admitted FeatureDeletionTasks (the caller runs
|
| 51 |
+
`datagen.repo_gate.gate_repo` BEFORE this — the driver assumes the
|
| 52 |
+
license/decontamination gates already passed).
|
| 53 |
+
env_factory: fresh `FeatureDeletionEnv` per rollout (a sandbox is
|
| 54 |
+
stateful; sharing one across episodes leaks trajectory state).
|
| 55 |
+
policy_factory: fresh policy per rollout (ScriptedPolicy is stateful).
|
| 56 |
+
manifest: a `RunManifest` with run_id/created_at/budget set by the
|
| 57 |
+
caller (created_at is caller-passed for reproducibility).
|
| 58 |
+
cost_per_rollout_usd: accounting hook — API policies should report
|
| 59 |
+
real cost; the driver enforces `manifest.budget_usd` with it.
|
| 60 |
+
prior_signatures: previous generation's MinHash signatures
|
| 61 |
+
(cross-generation dedup, finding D-12).
|
| 62 |
+
|
| 63 |
+
Raises:
|
| 64 |
+
FileExistsError: if the layout already has a manifest (write-once).
|
| 65 |
+
"""
|
| 66 |
+
if s3_contract.manifest_exists(layout):
|
| 67 |
+
raise FileExistsError(
|
| 68 |
+
f"Run layout already has a manifest at {layout.manifest_path} — "
|
| 69 |
+
"runs are write-once per (root, run_id); mint a new run_id "
|
| 70 |
+
"(finding D-21 idempotency)."
|
| 71 |
+
)
|
| 72 |
+
|
| 73 |
+
# 1. Holdout split FIRST — held-out tasks are never rolled out, so no
|
| 74 |
+
# training signal can derive from them (the HeldoutSplit discipline).
|
| 75 |
+
split = HeldoutSplit.split(source_tasks, holdout_frac=holdout_frac,
|
| 76 |
+
seed=holdout_seed, check_content=True)
|
| 77 |
+
by_id = {t.task_id: t for t in source_tasks}
|
| 78 |
+
holdout_tasks = [by_id[i] for i in sorted(split.holdout_ids)]
|
| 79 |
+
train_tasks = [by_id[i] for i in sorted(split.train_ids)]
|
| 80 |
+
if max_tasks is not None:
|
| 81 |
+
train_tasks = train_tasks[:max_tasks]
|
| 82 |
+
|
| 83 |
+
# 2. Rollouts + admission routing.
|
| 84 |
+
sft_rows: list[dict] = []
|
| 85 |
+
dpo_rows: list[dict] = []
|
| 86 |
+
quarantine_rows: list[dict] = []
|
| 87 |
+
traj_rows: list[dict] = []
|
| 88 |
+
partial = False
|
| 89 |
+
for task in train_tasks:
|
| 90 |
+
if manifest.over_budget:
|
| 91 |
+
partial = True
|
| 92 |
+
break
|
| 93 |
+
traj = collect_trajectory(env_factory(), task, policy_factory(),
|
| 94 |
+
provenance={"run_id": manifest.run_id})
|
| 95 |
+
manifest.spend(cost_per_rollout_usd)
|
| 96 |
+
verdict = admit(traj)
|
| 97 |
+
row = to_policy_row(traj, task)
|
| 98 |
+
traj_rows.append({**row, "admission": list(verdict.reasons)})
|
| 99 |
+
if verdict.sft_admitted:
|
| 100 |
+
sft_rows.append(row)
|
| 101 |
+
elif verdict.dpo_candidate:
|
| 102 |
+
dpo_rows.append(row)
|
| 103 |
+
else:
|
| 104 |
+
quarantine_rows.append({**row, "reasons": list(verdict.reasons)})
|
| 105 |
+
|
| 106 |
+
# 3. Dedup the SFT corpus (within-run + cross-generation).
|
| 107 |
+
def _key(r: dict) -> str:
|
| 108 |
+
return " ".join(m.get("content", "") for m in r.get("messages", []))
|
| 109 |
+
|
| 110 |
+
sft_rows, dedup_stats = dedup(sft_rows, _key, dedup_threshold,
|
| 111 |
+
prior_signatures=prior_signatures)
|
| 112 |
+
|
| 113 |
+
# 4. Write everything through the contract.
|
| 114 |
+
s3_contract.write_tasks(layout, train_tasks)
|
| 115 |
+
s3_contract.write_tasks_full(layout, train_tasks)
|
| 116 |
+
s3_contract.write_holdout(layout, holdout_tasks)
|
| 117 |
+
s3_contract.write_trajectories(layout, traj_rows)
|
| 118 |
+
s3_contract.write_sft_rows(layout, sft_rows)
|
| 119 |
+
s3_contract.write_dpo_rows(layout, dpo_rows)
|
| 120 |
+
s3_contract.write_quarantine(layout, quarantine_rows)
|
| 121 |
+
|
| 122 |
+
manifest.counts = {
|
| 123 |
+
"tasks_train": len(train_tasks),
|
| 124 |
+
"tasks_holdout": len(holdout_tasks),
|
| 125 |
+
"rollouts": len(traj_rows),
|
| 126 |
+
"sft_rows": len(sft_rows),
|
| 127 |
+
"dpo_rows": len(dpo_rows),
|
| 128 |
+
"quarantined": len(quarantine_rows),
|
| 129 |
+
**{f"dedup_{k}": v for k, v in dedup_stats.items()},
|
| 130 |
+
}
|
| 131 |
+
manifest.status = "partial" if partial else "building"
|
| 132 |
+
manifest.write(layout)
|
| 133 |
+
s3_contract.write_dataset_card(layout, manifest, dedup_stats=dedup_stats)
|
| 134 |
+
return manifest
|
| 135 |
+
|
| 136 |
+
|
| 137 |
+
__all__ = ["build_corpus"]
|
|
@@ -0,0 +1,138 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""dedup.py — MinHash near-duplicate detection (finding D-12).
|
| 2 |
+
|
| 3 |
+
Cross-generation dedup is a flywheel-collapse mitigation: a self-training loop
|
| 4 |
+
that re-ingests its own outputs accumulates near-identical rows, and per-batch
|
| 5 |
+
`document_deduplicator` (the only dedup the old designs had) never sees across
|
| 6 |
+
runs. This module computes MinHash signatures over word 5-shingles so a run can
|
| 7 |
+
(a) dedup within itself and (b) accept the PRIOR run's signature file and dedup
|
| 8 |
+
against it (lineage threaded by `RunManifest.parent_run_id`).
|
| 9 |
+
|
| 10 |
+
Pragmatic v1: builtin-hash permutation MinHash with N=64 seeds, no banding/LSH
|
| 11 |
+
(O(n^2) pair scan — fine for Stage-0 corpus sizes, thousands of rows).
|
| 12 |
+
`datasketch` (MinHashLSH) is the documented upgrade path when row counts make
|
| 13 |
+
the pair scan bite.
|
| 14 |
+
|
| 15 |
+
NOTE on hash stability: Python's builtin `hash()` over str is salted per
|
| 16 |
+
process (PYTHONHASHSEED), which would make signatures non-portable across
|
| 17 |
+
runs — exactly what cross-generation dedup needs. We therefore use md5-based
|
| 18 |
+
hashing (stable everywhere) despite the small speed cost.
|
| 19 |
+
"""
|
| 20 |
+
from __future__ import annotations
|
| 21 |
+
|
| 22 |
+
import hashlib
|
| 23 |
+
import json
|
| 24 |
+
import re
|
| 25 |
+
from typing import IO, Callable, Iterable, Sequence
|
| 26 |
+
|
| 27 |
+
N_PERMUTATIONS = 64
|
| 28 |
+
_SHINGLE_W = 5
|
| 29 |
+
_WORD_RE = re.compile(r"\w+")
|
| 30 |
+
_MAX64 = (1 << 64) - 1
|
| 31 |
+
|
| 32 |
+
|
| 33 |
+
def _shingles(text: str, w: int = _SHINGLE_W) -> set[str]:
|
| 34 |
+
words = _WORD_RE.findall(text.lower())
|
| 35 |
+
if len(words) <= w:
|
| 36 |
+
return {" ".join(words)} if words else set()
|
| 37 |
+
return {" ".join(words[i:i + w]) for i in range(len(words) - w + 1)}
|
| 38 |
+
|
| 39 |
+
|
| 40 |
+
def _stable_hash(s: str, seed: int) -> int:
|
| 41 |
+
h = hashlib.md5(f"{seed}:{s}".encode()).digest()
|
| 42 |
+
return int.from_bytes(h[:8], "big")
|
| 43 |
+
|
| 44 |
+
|
| 45 |
+
def minhash_signature(text: str, n_perm: int = N_PERMUTATIONS) -> tuple[int, ...]:
|
| 46 |
+
"""MinHash signature: per-seed minimum over the shingle set."""
|
| 47 |
+
sh = _shingles(text)
|
| 48 |
+
if not sh:
|
| 49 |
+
return tuple([_MAX64] * n_perm)
|
| 50 |
+
return tuple(min(_stable_hash(s, seed) for s in sh) for seed in range(n_perm))
|
| 51 |
+
|
| 52 |
+
|
| 53 |
+
def jaccard_estimate(sig_a: Sequence[int], sig_b: Sequence[int]) -> float:
|
| 54 |
+
"""Estimated Jaccard similarity = fraction of agreeing signature slots."""
|
| 55 |
+
if len(sig_a) != len(sig_b) or not sig_a:
|
| 56 |
+
raise ValueError("signatures must be equal-length and non-empty")
|
| 57 |
+
return sum(1 for a, b in zip(sig_a, sig_b) if a == b) / len(sig_a)
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
def find_near_duplicates(
|
| 61 |
+
rows: Sequence[dict],
|
| 62 |
+
key_fn: Callable[[dict], str],
|
| 63 |
+
threshold: float = 0.85,
|
| 64 |
+
*,
|
| 65 |
+
prior_signatures: Sequence[Sequence[int]] | None = None,
|
| 66 |
+
) -> list[tuple[int, int]]:
|
| 67 |
+
"""All (i, j) index pairs whose estimated Jaccard >= threshold.
|
| 68 |
+
|
| 69 |
+
`prior_signatures` (from a previous run) participate as virtual rows with
|
| 70 |
+
negative indices -(k+1), so a pair (i, -1) means "row i duplicates prior
|
| 71 |
+
signature 0" — the cross-generation case.
|
| 72 |
+
"""
|
| 73 |
+
sigs = [minhash_signature(key_fn(r)) for r in rows]
|
| 74 |
+
pairs: list[tuple[int, int]] = []
|
| 75 |
+
for i in range(len(sigs)):
|
| 76 |
+
for j in range(i + 1, len(sigs)):
|
| 77 |
+
if jaccard_estimate(sigs[i], sigs[j]) >= threshold:
|
| 78 |
+
pairs.append((i, j))
|
| 79 |
+
for k, prior in enumerate(prior_signatures or []):
|
| 80 |
+
if jaccard_estimate(sigs[i], prior) >= threshold:
|
| 81 |
+
pairs.append((i, -(k + 1)))
|
| 82 |
+
return pairs
|
| 83 |
+
|
| 84 |
+
|
| 85 |
+
def dedup(
|
| 86 |
+
rows: Sequence[dict],
|
| 87 |
+
key_fn: Callable[[dict], str],
|
| 88 |
+
threshold: float = 0.85,
|
| 89 |
+
*,
|
| 90 |
+
prior_signatures: Sequence[Sequence[int]] | None = None,
|
| 91 |
+
) -> tuple[list[dict], dict]:
|
| 92 |
+
"""Keep-first dedup. Returns (kept_rows, stats).
|
| 93 |
+
|
| 94 |
+
A row duplicating a PRIOR-run signature is dropped outright (the prior run
|
| 95 |
+
already owns it); within-run duplicates keep the earliest occurrence.
|
| 96 |
+
"""
|
| 97 |
+
pairs = find_near_duplicates(rows, key_fn, threshold,
|
| 98 |
+
prior_signatures=prior_signatures)
|
| 99 |
+
drop: set[int] = set()
|
| 100 |
+
for i, j in pairs:
|
| 101 |
+
if j < 0:
|
| 102 |
+
drop.add(i) # duplicates a prior-run row
|
| 103 |
+
else:
|
| 104 |
+
drop.add(j) # keep-first within this run
|
| 105 |
+
kept = [r for i, r in enumerate(rows) if i not in drop]
|
| 106 |
+
return kept, {
|
| 107 |
+
"rows_in": len(rows),
|
| 108 |
+
"rows_kept": len(kept),
|
| 109 |
+
"dropped_within_run": len({j for _, j in pairs if j >= 0} & drop),
|
| 110 |
+
"dropped_cross_generation": len({i for i, j in pairs if j < 0} & drop),
|
| 111 |
+
"threshold": threshold,
|
| 112 |
+
}
|
| 113 |
+
|
| 114 |
+
|
| 115 |
+
def signatures_to_jsonl(rows: Sequence[dict], key_fn: Callable[[dict], str],
|
| 116 |
+
fh: IO[str]) -> int:
|
| 117 |
+
"""Persist this run's signatures so the NEXT generation can dedup against
|
| 118 |
+
them (pass the loaded list as `prior_signatures`)."""
|
| 119 |
+
n = 0
|
| 120 |
+
for r in rows:
|
| 121 |
+
fh.write(json.dumps(list(minhash_signature(key_fn(r)))) + "\n")
|
| 122 |
+
n += 1
|
| 123 |
+
return n
|
| 124 |
+
|
| 125 |
+
|
| 126 |
+
def load_signatures(fh: IO[str]) -> list[tuple[int, ...]]:
|
| 127 |
+
return [tuple(json.loads(line)) for line in fh if line.strip()]
|
| 128 |
+
|
| 129 |
+
|
| 130 |
+
__all__ = [
|
| 131 |
+
"N_PERMUTATIONS",
|
| 132 |
+
"dedup",
|
| 133 |
+
"find_near_duplicates",
|
| 134 |
+
"jaccard_estimate",
|
| 135 |
+
"load_signatures",
|
| 136 |
+
"minhash_signature",
|
| 137 |
+
"signatures_to_jsonl",
|
| 138 |
+
]
|
|
@@ -0,0 +1,287 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""s3_contract.py — THE single dataset layout + manifest (finding V8/D-7/D-8).
|
| 2 |
+
|
| 3 |
+
Supersedes BOTH prior contracts: design-F1's `runs/<id>/{sft_corpus,dpo_pairs,
|
| 4 |
+
rl_task_pool,divergence_pairs,wm_tuples,holdout,diloco_rendezvous}` and
|
| 5 |
+
design-F2's `{traces,tasks,replay,task_grades,corpus}/v1/run_id=<id>` — the two
|
| 6 |
+
were never reconciled and coexisted in the grounding doc. One layout, one
|
| 7 |
+
manifest, two explicit serializers with a unit-tested leak guard.
|
| 8 |
+
|
| 9 |
+
Deliberate exclusions from the run layout:
|
| 10 |
+
* `diloco_rendezvous/` — training-comms state, not dataset; lives in its own
|
| 11 |
+
prefix/bucket (finding D-19).
|
| 12 |
+
* `wm_tuples/` — emitted only when the P4 world-model ablation is scheduled
|
| 13 |
+
(finding D-14); not part of Stage 0.
|
| 14 |
+
|
| 15 |
+
Layout (root = any local path or fsspec URI):
|
| 16 |
+
<root>/runs/<run_id>/
|
| 17 |
+
tasks/manifest.jsonl policy-safe task rows (golden_diff -> sha256)
|
| 18 |
+
tasks_full/manifest.jsonl construction-side full rows (RESTRICTED prefix)
|
| 19 |
+
traj/*.jsonl CanonicalTrajectory records (audit trail)
|
| 20 |
+
corpus_sft/rows.jsonl admitted SFT rows (to_policy_row output)
|
| 21 |
+
corpus_dpo/rows.jsonl DPO-candidate rows
|
| 22 |
+
holdout/tasks.jsonl held-out task ids+rows (never rolled out)
|
| 23 |
+
quarantine/*.jsonl rejected trajectories w/ reasons (audit)
|
| 24 |
+
manifest.json RunManifest
|
| 25 |
+
DATASET_CARD.md human-readable card
|
| 26 |
+
"""
|
| 27 |
+
from __future__ import annotations
|
| 28 |
+
|
| 29 |
+
import dataclasses
|
| 30 |
+
import hashlib
|
| 31 |
+
import json
|
| 32 |
+
from dataclasses import dataclass, field
|
| 33 |
+
from typing import IO, Iterable
|
| 34 |
+
|
| 35 |
+
from composer_replication.datagen.schema import FeatureDeletionTask
|
| 36 |
+
|
| 37 |
+
SCHEMA_VERSION = "1"
|
| 38 |
+
|
| 39 |
+
|
| 40 |
+
def _is_local(root: str) -> bool:
|
| 41 |
+
return "://" not in root or root.startswith("file://")
|
| 42 |
+
|
| 43 |
+
|
| 44 |
+
def _open(path: str, mode: str = "w") -> IO[str]:
|
| 45 |
+
"""Open a path for text IO; plain `open` locally, fsspec for s3:// etc.
|
| 46 |
+
|
| 47 |
+
fsspec is lazy so the module (and all local-corpus runs) need no extra dep.
|
| 48 |
+
"""
|
| 49 |
+
if _is_local(path):
|
| 50 |
+
import os
|
| 51 |
+
local = path.removeprefix("file://")
|
| 52 |
+
os.makedirs(os.path.dirname(local), exist_ok=True)
|
| 53 |
+
return open(local, mode, encoding="utf-8")
|
| 54 |
+
try:
|
| 55 |
+
import fsspec # noqa: PLC0415 — lazy heavy dep
|
| 56 |
+
except ImportError as e:
|
| 57 |
+
raise RuntimeError(
|
| 58 |
+
"Non-local corpus roots require fsspec; install with "
|
| 59 |
+
"`pip install -e .[serverless]`. Got: " + repr(e)
|
| 60 |
+
) from e
|
| 61 |
+
return fsspec.open(path, mode, encoding="utf-8").open()
|
| 62 |
+
|
| 63 |
+
|
| 64 |
+
def _exists(path: str) -> bool:
|
| 65 |
+
if _is_local(path):
|
| 66 |
+
import os
|
| 67 |
+
return os.path.exists(path.removeprefix("file://"))
|
| 68 |
+
import fsspec # noqa: PLC0415
|
| 69 |
+
fs, _, paths = fsspec.get_fs_token_paths(path)
|
| 70 |
+
return bool(fs.exists(paths[0]))
|
| 71 |
+
|
| 72 |
+
|
| 73 |
+
@dataclass(frozen=True)
|
| 74 |
+
class RunLayout:
|
| 75 |
+
"""Pure-path logic for one run's prefixes — testable without any IO."""
|
| 76 |
+
|
| 77 |
+
root: str
|
| 78 |
+
run_id: str
|
| 79 |
+
|
| 80 |
+
def _p(self, *parts: str) -> str:
|
| 81 |
+
base = self.root.rstrip("/")
|
| 82 |
+
return f"{base}/runs/{self.run_id}/" + "/".join(parts)
|
| 83 |
+
|
| 84 |
+
@property
|
| 85 |
+
def tasks_path(self) -> str:
|
| 86 |
+
return self._p("tasks", "manifest.jsonl")
|
| 87 |
+
|
| 88 |
+
@property
|
| 89 |
+
def tasks_full_path(self) -> str:
|
| 90 |
+
# RESTRICTED prefix: carries golden_diff/deleted_symbols. On S3 this
|
| 91 |
+
# prefix gets a deny-by-default policy; locally it is still separated
|
| 92 |
+
# so a naive `corpus_*` glob can never sweep it up.
|
| 93 |
+
return self._p("tasks_full", "manifest.jsonl")
|
| 94 |
+
|
| 95 |
+
@property
|
| 96 |
+
def traj_path(self) -> str:
|
| 97 |
+
return self._p("traj", "trajectories.jsonl")
|
| 98 |
+
|
| 99 |
+
@property
|
| 100 |
+
def sft_path(self) -> str:
|
| 101 |
+
return self._p("corpus_sft", "rows.jsonl")
|
| 102 |
+
|
| 103 |
+
@property
|
| 104 |
+
def dpo_path(self) -> str:
|
| 105 |
+
return self._p("corpus_dpo", "rows.jsonl")
|
| 106 |
+
|
| 107 |
+
@property
|
| 108 |
+
def holdout_path(self) -> str:
|
| 109 |
+
return self._p("holdout", "tasks.jsonl")
|
| 110 |
+
|
| 111 |
+
@property
|
| 112 |
+
def quarantine_path(self) -> str:
|
| 113 |
+
return self._p("quarantine", "rejected.jsonl")
|
| 114 |
+
|
| 115 |
+
@property
|
| 116 |
+
def manifest_path(self) -> str:
|
| 117 |
+
return self._p("manifest.json")
|
| 118 |
+
|
| 119 |
+
@property
|
| 120 |
+
def card_path(self) -> str:
|
| 121 |
+
return self._p("DATASET_CARD.md")
|
| 122 |
+
|
| 123 |
+
|
| 124 |
+
@dataclass
|
| 125 |
+
class RunManifest:
|
| 126 |
+
"""Run-level metadata: counts, cost, lineage, budget, acceptance status.
|
| 127 |
+
|
| 128 |
+
`created_at` is CALLER-passed (never datetime.now() in here) so manifests
|
| 129 |
+
are reproducible in tests. `parent_run_id` threads flywheel lineage so
|
| 130 |
+
cross-generation dedup (finding D-12) can find prior signatures.
|
| 131 |
+
"""
|
| 132 |
+
|
| 133 |
+
run_id: str
|
| 134 |
+
created_at: str
|
| 135 |
+
source: str = ""
|
| 136 |
+
counts: dict = field(default_factory=dict)
|
| 137 |
+
cost_usd: float = 0.0
|
| 138 |
+
parent_run_id: str | None = None
|
| 139 |
+
schema_version: str = SCHEMA_VERSION
|
| 140 |
+
status: str = "building" # building | accepted | rejected | partial
|
| 141 |
+
budget_usd: float | None = None
|
| 142 |
+
|
| 143 |
+
def spend(self, usd: float) -> None:
|
| 144 |
+
self.cost_usd += usd
|
| 145 |
+
|
| 146 |
+
@property
|
| 147 |
+
def over_budget(self) -> bool:
|
| 148 |
+
return self.budget_usd is not None and self.cost_usd >= self.budget_usd
|
| 149 |
+
|
| 150 |
+
def write(self, layout: RunLayout) -> None:
|
| 151 |
+
with _open(layout.manifest_path) as f:
|
| 152 |
+
json.dump(dataclasses.asdict(self), f, indent=2)
|
| 153 |
+
|
| 154 |
+
@classmethod
|
| 155 |
+
def read(cls, layout: RunLayout) -> RunManifest:
|
| 156 |
+
with _open(layout.manifest_path, "r") as f:
|
| 157 |
+
return cls(**json.load(f))
|
| 158 |
+
|
| 159 |
+
|
| 160 |
+
# ---------------------------------------------------------------------
|
| 161 |
+
# Writers — the leak guard lives here (finding D-8)
|
| 162 |
+
# ---------------------------------------------------------------------
|
| 163 |
+
|
| 164 |
+
|
| 165 |
+
def _task_row_policy_safe(task: FeatureDeletionTask) -> dict:
|
| 166 |
+
"""Task row with the construction-side secrets REPLACED, not just hidden.
|
| 167 |
+
|
| 168 |
+
`asdict()` includes `golden_diff` despite `repr=False` — that is exactly
|
| 169 |
+
the leak D-8 flagged. We keep provenance via a sha256 (verifiable, not
|
| 170 |
+
recoverable) and drop `deleted_symbols` entirely (they name the answer).
|
| 171 |
+
"""
|
| 172 |
+
row = dataclasses.asdict(task)
|
| 173 |
+
gold = row.pop("golden_diff", "")
|
| 174 |
+
row.pop("deleted_symbols", None)
|
| 175 |
+
row["golden_diff_sha256"] = hashlib.sha256(gold.encode()).hexdigest() if gold else ""
|
| 176 |
+
return row
|
| 177 |
+
|
| 178 |
+
|
| 179 |
+
def write_tasks(layout: RunLayout, tasks: Iterable[FeatureDeletionTask]) -> int:
|
| 180 |
+
"""Write the POLICY-SAFE task manifest (the default everything reads)."""
|
| 181 |
+
n = 0
|
| 182 |
+
with _open(layout.tasks_path) as f:
|
| 183 |
+
for t in tasks:
|
| 184 |
+
f.write(json.dumps(_task_row_policy_safe(t)) + "\n")
|
| 185 |
+
n += 1
|
| 186 |
+
return n
|
| 187 |
+
|
| 188 |
+
|
| 189 |
+
def write_tasks_full(layout: RunLayout, tasks: Iterable[FeatureDeletionTask]) -> int:
|
| 190 |
+
"""Write FULL task rows (incl. golden_diff) to the RESTRICTED prefix.
|
| 191 |
+
|
| 192 |
+
Only the validator/monitor side reads this; never corpus consumers.
|
| 193 |
+
"""
|
| 194 |
+
n = 0
|
| 195 |
+
with _open(layout.tasks_full_path) as f:
|
| 196 |
+
for t in tasks:
|
| 197 |
+
f.write(json.dumps(dataclasses.asdict(t)) + "\n")
|
| 198 |
+
n += 1
|
| 199 |
+
return n
|
| 200 |
+
|
| 201 |
+
|
| 202 |
+
def _write_jsonl(path: str, rows: Iterable[dict]) -> int:
|
| 203 |
+
n = 0
|
| 204 |
+
with _open(path) as f:
|
| 205 |
+
for r in rows:
|
| 206 |
+
f.write(json.dumps(r) + "\n")
|
| 207 |
+
n += 1
|
| 208 |
+
return n
|
| 209 |
+
|
| 210 |
+
|
| 211 |
+
def write_sft_rows(layout: RunLayout, rows: Iterable[dict]) -> int:
|
| 212 |
+
return _write_jsonl(layout.sft_path, rows)
|
| 213 |
+
|
| 214 |
+
|
| 215 |
+
def write_dpo_rows(layout: RunLayout, rows: Iterable[dict]) -> int:
|
| 216 |
+
return _write_jsonl(layout.dpo_path, rows)
|
| 217 |
+
|
| 218 |
+
|
| 219 |
+
def write_quarantine(layout: RunLayout, rows: Iterable[dict]) -> int:
|
| 220 |
+
return _write_jsonl(layout.quarantine_path, rows)
|
| 221 |
+
|
| 222 |
+
|
| 223 |
+
def write_holdout(layout: RunLayout, tasks: Iterable[FeatureDeletionTask]) -> int:
|
| 224 |
+
return _write_jsonl(layout.holdout_path, (_task_row_policy_safe(t) for t in tasks))
|
| 225 |
+
|
| 226 |
+
|
| 227 |
+
def write_trajectories(layout: RunLayout, rows: Iterable[dict]) -> int:
|
| 228 |
+
return _write_jsonl(layout.traj_path, rows)
|
| 229 |
+
|
| 230 |
+
|
| 231 |
+
def write_dataset_card(layout: RunLayout, manifest: RunManifest,
|
| 232 |
+
*, license_tiers: dict[str, int] | None = None,
|
| 233 |
+
dedup_stats: dict | None = None,
|
| 234 |
+
decontamination_note: str = "") -> None:
|
| 235 |
+
"""A small human-readable dataset card (finding D-18)."""
|
| 236 |
+
lines = [
|
| 237 |
+
f"# Dataset card — run `{manifest.run_id}`",
|
| 238 |
+
"",
|
| 239 |
+
f"- **created:** {manifest.created_at}",
|
| 240 |
+
f"- **source:** {manifest.source}",
|
| 241 |
+
f"- **status:** {manifest.status}",
|
| 242 |
+
f"- **schema_version:** {manifest.schema_version}",
|
| 243 |
+
f"- **cost (USD):** {manifest.cost_usd:.2f}"
|
| 244 |
+
+ (f" / budget {manifest.budget_usd:.2f}" if manifest.budget_usd else ""),
|
| 245 |
+
f"- **lineage:** parent_run_id={manifest.parent_run_id or 'none'}",
|
| 246 |
+
"",
|
| 247 |
+
"## Counts",
|
| 248 |
+
"",
|
| 249 |
+
]
|
| 250 |
+
for k, v in sorted(manifest.counts.items()):
|
| 251 |
+
lines.append(f"- {k}: {v}")
|
| 252 |
+
if license_tiers:
|
| 253 |
+
lines += ["", "## License tiers seen", ""]
|
| 254 |
+
lines += [f"- {k}: {v}" for k, v in sorted(license_tiers.items())]
|
| 255 |
+
lines += ["", "## Decontamination", "",
|
| 256 |
+
decontamination_note or
|
| 257 |
+
"All source repos checked against the SWE-bench-family eval list "
|
| 258 |
+
"(datagen.repo_gate.DECONTAMINATION_LIST) at ingest."]
|
| 259 |
+
if dedup_stats:
|
| 260 |
+
lines += ["", "## Dedup", ""]
|
| 261 |
+
lines += [f"- {k}: {v}" for k, v in sorted(dedup_stats.items())]
|
| 262 |
+
lines += ["", "Policy-safe rows only: `golden_diff` is sha256-hashed and "
|
| 263 |
+
"`deleted_symbols` dropped in `tasks/`, `corpus_*/`, `holdout/` "
|
| 264 |
+
"(full rows live in the restricted `tasks_full/`).", ""]
|
| 265 |
+
with _open(layout.card_path) as f:
|
| 266 |
+
f.write("\n".join(lines))
|
| 267 |
+
|
| 268 |
+
|
| 269 |
+
def manifest_exists(layout: RunLayout) -> bool:
|
| 270 |
+
"""Write-once guard for the driver (finding D-21 idempotency)."""
|
| 271 |
+
return _exists(layout.manifest_path)
|
| 272 |
+
|
| 273 |
+
|
| 274 |
+
__all__ = [
|
| 275 |
+
"SCHEMA_VERSION",
|
| 276 |
+
"RunLayout",
|
| 277 |
+
"RunManifest",
|
| 278 |
+
"manifest_exists",
|
| 279 |
+
"write_dataset_card",
|
| 280 |
+
"write_dpo_rows",
|
| 281 |
+
"write_holdout",
|
| 282 |
+
"write_quarantine",
|
| 283 |
+
"write_sft_rows",
|
| 284 |
+
"write_tasks",
|
| 285 |
+
"write_tasks_full",
|
| 286 |
+
"write_trajectories",
|
| 287 |
+
]
|
|
File without changes
|
|
@@ -0,0 +1,223 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Tests for the Stage-0 pipeline: contract, dedup, driver.
|
| 2 |
+
|
| 3 |
+
Load-bearing coverage: the sentinel leak guard on write_tasks (finding D-8),
|
| 4 |
+
holdout exclusion + budget stop + idempotency in build_corpus (D-21), and the
|
| 5 |
+
cross-generation dedup path (D-12).
|
| 6 |
+
"""
|
| 7 |
+
from __future__ import annotations
|
| 8 |
+
|
| 9 |
+
import io
|
| 10 |
+
import json
|
| 11 |
+
from pathlib import Path
|
| 12 |
+
|
| 13 |
+
import pytest
|
| 14 |
+
|
| 15 |
+
from composer_replication.datagen.env import FeatureDeletionEnv
|
| 16 |
+
from composer_replication.datagen.rollout_harness import ScriptedPolicy
|
| 17 |
+
from composer_replication.datagen.sandbox import FakeSandbox
|
| 18 |
+
from composer_replication.datagen.schema import FeatureDeletionTask
|
| 19 |
+
from composer_replication.datagen.trajectory import ToolCall
|
| 20 |
+
from composer_replication.pipeline.build_corpus import build_corpus
|
| 21 |
+
from composer_replication.pipeline.dedup import (
|
| 22 |
+
dedup,
|
| 23 |
+
find_near_duplicates,
|
| 24 |
+
jaccard_estimate,
|
| 25 |
+
load_signatures,
|
| 26 |
+
minhash_signature,
|
| 27 |
+
signatures_to_jsonl,
|
| 28 |
+
)
|
| 29 |
+
from composer_replication.pipeline.s3_contract import (
|
| 30 |
+
RunLayout,
|
| 31 |
+
RunManifest,
|
| 32 |
+
write_dataset_card,
|
| 33 |
+
write_tasks,
|
| 34 |
+
write_tasks_full,
|
| 35 |
+
)
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
def _task(i: int, **over) -> FeatureDeletionTask:
|
| 39 |
+
base = dict(
|
| 40 |
+
task_id=f"task-{i:03d}", repo="org/repo", base_commit="abc",
|
| 41 |
+
broken_image="img:1", test_command="pytest -q",
|
| 42 |
+
fail_to_pass=(f"t/a.py::t{i}",), pass_to_pass=("t/a.py::keep",),
|
| 43 |
+
golden_diff="SENTINEL_NEVER_LEAK", deleted_symbols=("secret_fn",),
|
| 44 |
+
)
|
| 45 |
+
base.update(over)
|
| 46 |
+
return FeatureDeletionTask(**base)
|
| 47 |
+
|
| 48 |
+
|
| 49 |
+
# ---------------------------------------------------------------------
|
| 50 |
+
# RunLayout / RunManifest
|
| 51 |
+
# ---------------------------------------------------------------------
|
| 52 |
+
|
| 53 |
+
|
| 54 |
+
def test_layout_paths_are_pure_and_namespaced():
|
| 55 |
+
lay = RunLayout(root="/data/corpora", run_id="run42")
|
| 56 |
+
assert lay.sft_path == "/data/corpora/runs/run42/corpus_sft/rows.jsonl"
|
| 57 |
+
assert lay.manifest_path == "/data/corpora/runs/run42/manifest.json"
|
| 58 |
+
s3 = RunLayout(root="s3://bucket/prefix/", run_id="r")
|
| 59 |
+
assert s3.tasks_path == "s3://bucket/prefix/runs/r/tasks/manifest.jsonl"
|
| 60 |
+
|
| 61 |
+
|
| 62 |
+
def test_manifest_round_trip_and_budget(tmp_path):
|
| 63 |
+
lay = RunLayout(root=str(tmp_path), run_id="r1")
|
| 64 |
+
m = RunManifest(run_id="r1", created_at="2026-06-09T00:00:00Z",
|
| 65 |
+
source="test", budget_usd=1.0)
|
| 66 |
+
m.spend(0.4)
|
| 67 |
+
assert not m.over_budget
|
| 68 |
+
m.spend(0.6)
|
| 69 |
+
assert m.over_budget
|
| 70 |
+
m.write(lay)
|
| 71 |
+
m2 = RunManifest.read(lay)
|
| 72 |
+
assert m2.cost_usd == pytest.approx(1.0)
|
| 73 |
+
assert m2.budget_usd == 1.0
|
| 74 |
+
|
| 75 |
+
|
| 76 |
+
# ---------------------------------------------------------------------
|
| 77 |
+
# THE leak guard (finding D-8)
|
| 78 |
+
# ---------------------------------------------------------------------
|
| 79 |
+
|
| 80 |
+
|
| 81 |
+
def test_write_tasks_never_leaks_golden_diff(tmp_path):
|
| 82 |
+
lay = RunLayout(root=str(tmp_path), run_id="r1")
|
| 83 |
+
write_tasks(lay, [_task(1)])
|
| 84 |
+
blob = Path(lay.tasks_path).read_text()
|
| 85 |
+
assert "SENTINEL_NEVER_LEAK" not in blob
|
| 86 |
+
assert "secret_fn" not in blob
|
| 87 |
+
row = json.loads(blob.splitlines()[0])
|
| 88 |
+
assert row["golden_diff_sha256"] # provenance preserved as a hash
|
| 89 |
+
# The restricted full writer DOES carry it (construction side only).
|
| 90 |
+
write_tasks_full(lay, [_task(1)])
|
| 91 |
+
assert "SENTINEL_NEVER_LEAK" in Path(lay.tasks_full_path).read_text()
|
| 92 |
+
|
| 93 |
+
|
| 94 |
+
# ---------------------------------------------------------------------
|
| 95 |
+
# MinHash dedup
|
| 96 |
+
# ---------------------------------------------------------------------
|
| 97 |
+
|
| 98 |
+
_TEXT_A = "the quick brown fox jumps over the lazy dog and then runs far away home tonight"
|
| 99 |
+
_TEXT_A2 = "the quick brown fox jumps over the lazy dog and then runs far away home today"
|
| 100 |
+
_TEXT_B = "import numpy as np def main(): return np.zeros(10) print(main()) totally different content here"
|
| 101 |
+
|
| 102 |
+
|
| 103 |
+
def test_jaccard_estimate_near_duplicates_high_disjoint_low():
|
| 104 |
+
sa, sa2, sb = (minhash_signature(t) for t in (_TEXT_A, _TEXT_A2, _TEXT_B))
|
| 105 |
+
assert jaccard_estimate(sa, sa2) > 0.5
|
| 106 |
+
assert jaccard_estimate(sa, sb) < 0.2
|
| 107 |
+
assert jaccard_estimate(sa, sa) == 1.0
|
| 108 |
+
|
| 109 |
+
|
| 110 |
+
def test_dedup_keeps_first_and_drops_near_dup():
|
| 111 |
+
rows = [{"text": _TEXT_A}, {"text": _TEXT_A2}, {"text": _TEXT_B}]
|
| 112 |
+
kept, stats = dedup(rows, lambda r: r["text"], threshold=0.5)
|
| 113 |
+
assert [r["text"] for r in kept] == [_TEXT_A, _TEXT_B]
|
| 114 |
+
assert stats["dropped_within_run"] == 1
|
| 115 |
+
|
| 116 |
+
|
| 117 |
+
def test_cross_generation_dedup_via_signature_file():
|
| 118 |
+
prior_rows = [{"text": _TEXT_A}]
|
| 119 |
+
buf = io.StringIO()
|
| 120 |
+
signatures_to_jsonl(prior_rows, lambda r: r["text"], buf)
|
| 121 |
+
buf.seek(0)
|
| 122 |
+
prior_sigs = load_signatures(buf)
|
| 123 |
+
|
| 124 |
+
rows = [{"text": _TEXT_A2}, {"text": _TEXT_B}]
|
| 125 |
+
kept, stats = dedup(rows, lambda r: r["text"], threshold=0.5,
|
| 126 |
+
prior_signatures=prior_sigs)
|
| 127 |
+
assert [r["text"] for r in kept] == [_TEXT_B]
|
| 128 |
+
assert stats["dropped_cross_generation"] == 1
|
| 129 |
+
|
| 130 |
+
|
| 131 |
+
def test_find_near_duplicates_pairs():
|
| 132 |
+
rows = [{"t": _TEXT_A}, {"t": _TEXT_A2}]
|
| 133 |
+
assert find_near_duplicates(rows, lambda r: r["t"], 0.5) == [(0, 1)]
|
| 134 |
+
|
| 135 |
+
|
| 136 |
+
# ---------------------------------------------------------------------
|
| 137 |
+
# build_corpus end-to-end (FakeSandbox + ScriptedPolicy)
|
| 138 |
+
# ---------------------------------------------------------------------
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
def _passing_policy():
|
| 142 |
+
# Flips both this task's F2P tests green generically: FakeSandbox's
|
| 143 |
+
# set_outcome takes explicit test names, so the fixture tasks share names
|
| 144 |
+
# via the same fail_to_pass tuple pattern; we set a superset.
|
| 145 |
+
outcomes = {f"t/a.py::t{i}": True for i in range(20)}
|
| 146 |
+
outcomes["t/a.py::keep"] = True
|
| 147 |
+
return ScriptedPolicy(actions=[ToolCall("set_outcome", {"outcomes": outcomes}), "done"])
|
| 148 |
+
|
| 149 |
+
|
| 150 |
+
def _failing_policy():
|
| 151 |
+
return ScriptedPolicy(actions=["gave up immediately"])
|
| 152 |
+
|
| 153 |
+
|
| 154 |
+
def _env():
|
| 155 |
+
return FeatureDeletionEnv(FakeSandbox(test_outcomes={"t/a.py::keep": True}))
|
| 156 |
+
|
| 157 |
+
|
| 158 |
+
def test_build_corpus_end_to_end(tmp_path):
|
| 159 |
+
tasks = [_task(i) for i in range(6)]
|
| 160 |
+
lay = RunLayout(root=str(tmp_path), run_id="e2e")
|
| 161 |
+
manifest = RunManifest(run_id="e2e", created_at="2026-06-09T00:00:00Z", source="fixture")
|
| 162 |
+
|
| 163 |
+
out = build_corpus(tasks, _env, _passing_policy, lay, manifest,
|
| 164 |
+
holdout_frac=0.34, holdout_seed=7)
|
| 165 |
+
|
| 166 |
+
# Holdout exclusion: holdout tasks were never rolled out.
|
| 167 |
+
assert out.counts["tasks_holdout"] >= 1
|
| 168 |
+
assert out.counts["rollouts"] == out.counts["tasks_train"]
|
| 169 |
+
# Full passes routed to SFT (post-dedup near-identical rows collapse —
|
| 170 |
+
# the fixture tasks produce near-identical messages, which is itself a
|
| 171 |
+
# realistic dedup scenario).
|
| 172 |
+
assert out.counts["sft_rows"] >= 1
|
| 173 |
+
assert out.counts["quarantined"] == 0
|
| 174 |
+
# Files exist and the SFT corpus never leaks the sentinel.
|
| 175 |
+
sft_blob = Path(lay.sft_path).read_text()
|
| 176 |
+
assert "SENTINEL_NEVER_LEAK" not in sft_blob
|
| 177 |
+
assert Path(lay.card_path).exists()
|
| 178 |
+
assert Path(lay.holdout_path).exists()
|
| 179 |
+
|
| 180 |
+
|
| 181 |
+
def test_build_corpus_quarantines_failures(tmp_path):
|
| 182 |
+
tasks = [_task(i) for i in range(3)]
|
| 183 |
+
lay = RunLayout(root=str(tmp_path), run_id="fail")
|
| 184 |
+
manifest = RunManifest(run_id="fail", created_at="2026-06-09T00:00:00Z", source="fixture")
|
| 185 |
+
out = build_corpus(tasks, _env, _failing_policy, lay, manifest,
|
| 186 |
+
holdout_frac=0.34, holdout_seed=7)
|
| 187 |
+
assert out.counts["sft_rows"] == 0
|
| 188 |
+
assert out.counts["quarantined"] == out.counts["rollouts"] > 0
|
| 189 |
+
|
| 190 |
+
|
| 191 |
+
def test_build_corpus_budget_stop_marks_partial(tmp_path):
|
| 192 |
+
tasks = [_task(i) for i in range(6)]
|
| 193 |
+
lay = RunLayout(root=str(tmp_path), run_id="budget")
|
| 194 |
+
manifest = RunManifest(run_id="budget", created_at="2026-06-09T00:00:00Z",
|
| 195 |
+
source="fixture", budget_usd=0.25)
|
| 196 |
+
out = build_corpus(tasks, _env, _passing_policy, lay, manifest,
|
| 197 |
+
holdout_frac=0.2, holdout_seed=7,
|
| 198 |
+
cost_per_rollout_usd=0.1)
|
| 199 |
+
assert out.status == "partial"
|
| 200 |
+
assert out.counts["rollouts"] < out.counts["tasks_train"]
|
| 201 |
+
|
| 202 |
+
|
| 203 |
+
def test_build_corpus_is_write_once(tmp_path):
|
| 204 |
+
tasks = [_task(i) for i in range(3)]
|
| 205 |
+
lay = RunLayout(root=str(tmp_path), run_id="once")
|
| 206 |
+
m1 = RunManifest(run_id="once", created_at="2026-06-09T00:00:00Z", source="fixture")
|
| 207 |
+
build_corpus(tasks, _env, _passing_policy, lay, m1, holdout_frac=0.34)
|
| 208 |
+
m2 = RunManifest(run_id="once", created_at="2026-06-09T00:00:01Z", source="fixture")
|
| 209 |
+
with pytest.raises(FileExistsError, match="write-once"):
|
| 210 |
+
build_corpus(tasks, _env, _passing_policy, lay, m2, holdout_frac=0.34)
|
| 211 |
+
|
| 212 |
+
|
| 213 |
+
def test_dataset_card_contents(tmp_path):
|
| 214 |
+
lay = RunLayout(root=str(tmp_path), run_id="card")
|
| 215 |
+
m = RunManifest(run_id="card", created_at="2026-06-09T00:00:00Z",
|
| 216 |
+
source="fixture", counts={"sft_rows": 3})
|
| 217 |
+
write_dataset_card(lay, m, license_tiers={"REDISTRIBUTABLE": 3},
|
| 218 |
+
dedup_stats={"rows_kept": 3})
|
| 219 |
+
card = Path(lay.card_path).read_text()
|
| 220 |
+
assert "run `card`" in card
|
| 221 |
+
assert "sft_rows: 3" in card
|
| 222 |
+
assert "REDISTRIBUTABLE: 3" in card
|
| 223 |
+
assert "Decontamination" in card
|
|
@@ -4,8 +4,12 @@ This is channel 3 of the integrated trainer: at each step of a frozen agentic
|
|
| 4 |
trace, query N pre-trained external teachers (frontier models from different
|
| 5 |
labs) and convert teacher disagreement into preference pairs for DPO loss.
|
| 6 |
|
| 7 |
-
Generalized from spike-001's `replay.py`.
|
| 8 |
-
$0.98 mean per-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
|
| 10 |
Usage:
|
| 11 |
from teacher_replay import replay_trace, extract_dpo_pairs
|
|
|
|
| 4 |
trace, query N pre-trained external teachers (frontier models from different
|
| 5 |
labs) and convert teacher disagreement into preference pairs for DPO loss.
|
| 6 |
|
| 7 |
+
Generalized from spike-001's `replay.py`. Cost calibration (✅ spike 001,
|
| 8 |
+
relabeled per deepread finding V11): $0.98 mean per-TRACE cost ungated was
|
| 9 |
+
measured on a ~50-state SYNTHETIC trace at N=3 teachers; real Claude Code
|
| 10 |
+
sessions run 125–2,830 tool-use messages (ADR-002), so a full real session is
|
| 11 |
+
~2 orders of magnitude more (~$70–80 flat, before VOI gating). $0.30/trace
|
| 12 |
+
projected with VOI gating, same synthetic basis.
|
| 13 |
|
| 14 |
Usage:
|
| 15 |
from teacher_replay import replay_trace, extract_dpo_pairs
|
|
@@ -9,7 +9,9 @@ literature says this is not cosmetic:
|
|
| 9 |
|
| 10 |
* arXiv:2512.21852 ("A Comedy of Estimators") — k1-in-reward improves OOD
|
| 11 |
generalization; k3-in-reward can collapse.
|
| 12 |
-
* verl
|
|
|
|
|
|
|
| 13 |
* TRL issue #4967 tracks the same divergence.
|
| 14 |
|
| 15 |
OOD generalization is exactly the "take any model to the next level" axis, so
|
|
|
|
| 9 |
|
| 10 |
* arXiv:2512.21852 ("A Comedy of Estimators") — k1-in-reward improves OOD
|
| 11 |
generalization; k3-in-reward can collapse.
|
| 12 |
+
* verl ships k1-in-reward as its default/recommended reverse-KL option
|
| 13 |
+
(it also supports a k3-family "low_var_kl" — wording corrected per
|
| 14 |
+
deepread finding V13).
|
| 15 |
* TRL issue #4967 tracks the same divergence.
|
| 16 |
|
| 17 |
OOD generalization is exactly the "take any model to the next level" axis, so
|
|
@@ -22,7 +22,7 @@ The Cursor blog discusses **only three** training innovations explicitly. Everyt
|
|
| 22 |
|
| 23 |
**Cited prior art** (Cursor's footnote 1):
|
| 24 |
- **OPSD: Self-Distilled Reasoner — On-Policy Self-Distillation for LLMs** (Zhao et al., 2026, [arXiv:2601.18734](https://arxiv.org/abs/2601.18734), [GitHub: siyan-zhao/OPSD](https://github.com/siyan-zhao/OPSD)). The original on-policy-self-distillation framework: single LLM, teacher conditioned on privileged information (e.g. ground-truth answer), student sees only the question, loss = per-token KL on student's own rollouts.
|
| 25 |
-
- **SDPO: Reinforcement Learning via Self-Distillation** (Hübotter et al., 2026, [arXiv:2601.20802](https://arxiv.org/abs/2601.20802), ICLR 2026 Scaling Post-training Workshop). Generalizes OPSD to RL with rich feedback: *"SDPO treats the current model conditioned on feedback as a self-teacher and distills its feedback-informed next-token predictions back into the policy."*
|
| 26 |
|
| 27 |
| Method | Sampling | Signal | Feedback |
|
| 28 |
|---|---|---|---|
|
|
@@ -72,7 +72,7 @@ This is **infrastructure, not algorithm**. It only matters at MoE-1T scale; for
|
|
| 72 |
| Composer 2.5 stage | Blog mechanism | Our replication target | v0.0 | v0.1 | v0.2 |
|
| 73 |
|---|---|---|---|---|---|
|
| 74 |
| **(a)** Continued pretraining on code | Standard pretraining, code-weighted | Skip — start from already-code-tuned `Qwen3-Coder-7B` or `Qwen3-Coder-30B-A3B` | ✗ | ✗ | ✗ |
|
| 75 |
-
| **(b)** Synthetic data at scale | Feature Deletion +
|
| 76 |
| **(c)** Realistic-environment RL (RLVR) | Async sandboxes, same tool harness as production | TRL `GRPOTrainer` + verifiers + OpenEnv; SWE-bench-lite env in v0.0; build sandboxed code execution env in v0.1 | ✓ baseline | ✓ + DAPO patches | + decentralized rollouts |
|
| 77 |
| **(d)** Targeted RL w/ textual feedback (Composer's secret sauce) | Same-model self-distill: insert hint into context → teacher; original → student; on-policy KL at the turn | **Lift the OPSD/SDPO loss directly from `siyan-zhao/OPSD`** (published code, MIT). Generate hints via templates (v0.1) or LLM (v0.2). | ✗ (deferred) | ✓ (this is the Composer-recipe channel) | + learned hint generator |
|
| 78 |
| **(e)** Trace-replay multi-teacher distill (NOVEL — our addition) | N/A (not in Composer) | N=3 teachers (Opus 4.7, GPT-5, DeepSeek V4 Pro) replay each step; disagreement → DPO pairs | ✓ (this is the v0.0 novelty bet) | ✓ + VOI gating | + tiered teachers |
|
|
@@ -148,7 +148,7 @@ Primary sources for each Composer-2.5 component, post-audit:
|
|
| 148 |
- **Cursor blog** — [Introducing Composer 2.5](https://cursor.com/blog/composer-2-5) (2026)
|
| 149 |
- **Cursor blog** — [Composer 2 technical report](https://cursor.com/blog/composer-2-technical-report) (predecessor; named the "Anyrun" environment per subagent — verify if needed)
|
| 150 |
- **OPSD paper** — Zhao et al., *Self-Distilled Reasoner: On-Policy Self-Distillation for LLMs*, [arXiv:2601.18734](https://arxiv.org/abs/2601.18734), code at [siyan-zhao/OPSD](https://github.com/siyan-zhao/OPSD). MIT.
|
| 151 |
-
- **SDPO paper** — Hübotter et al., *Reinforcement Learning via Self-Distillation*, [arXiv:2601.20802](https://arxiv.org/abs/2601.20802), ICLR 2026 Scaling Post-training Workshop. The
|
| 152 |
- **Self-Distillation continual-learning** — [arXiv:2601.19897](https://arxiv.org/abs/2601.19897). Cited by Cursor; less directly relevant.
|
| 153 |
- **Moonshot Kimi K2.5** — base model, [HF model card](https://huggingface.co/moonshotai/Kimi-K2-Thinking).
|
| 154 |
|
|
|
|
| 22 |
|
| 23 |
**Cited prior art** (Cursor's footnote 1):
|
| 24 |
- **OPSD: Self-Distilled Reasoner — On-Policy Self-Distillation for LLMs** (Zhao et al., 2026, [arXiv:2601.18734](https://arxiv.org/abs/2601.18734), [GitHub: siyan-zhao/OPSD](https://github.com/siyan-zhao/OPSD)). The original on-policy-self-distillation framework: single LLM, teacher conditioned on privileged information (e.g. ground-truth answer), student sees only the question, loss = per-token KL on student's own rollouts.
|
| 25 |
+
- **SDPO: Reinforcement Learning via Self-Distillation** (Hübotter et al., 2026, [arXiv:2601.20802](https://arxiv.org/abs/2601.20802), ICLR 2026 Scaling Post-training Workshop). Generalizes OPSD to RL with rich feedback: *"SDPO treats the current model conditioned on feedback as a self-teacher and distills its feedback-informed next-token predictions back into the policy."* Cursor's blog cites this paper only as **background** ("For more background on this approach see…") — NOT as its mechanism; SDPO's published loss is full-rollout with feedback-in-prefix and an EMA-regularized teacher, while Composer's blog describes a turn-localized hint splice. Closely related, **not verified identical** (deepread finding V1). **There is published code.** Comparison table from the SDPO paper:
|
| 26 |
|
| 27 |
| Method | Sampling | Signal | Feedback |
|
| 28 |
|---|---|---|---|
|
|
|
|
| 72 |
| Composer 2.5 stage | Blog mechanism | Our replication target | v0.0 | v0.1 | v0.2 |
|
| 73 |
|---|---|---|---|---|---|
|
| 74 |
| **(a)** Continued pretraining on code | Standard pretraining, code-weighted | Skip — start from already-code-tuned `Qwen3-Coder-7B` or `Qwen3-Coder-30B-A3B` | ✗ | ✗ | ✗ |
|
| 75 |
+
| **(b)** Synthetic data at scale | Feature Deletion + an unspecified number of other generators (the blog says only "a range of approaches" — the old "24" was a back-formation from the 25x task multiplier; deepread finding V5) | Build 1 generator (Feature Deletion) as OpenEnv-compatible env. Use SWE-bench-lite and SWE-Gym as drop-in alternatives. | ✗ (use SWE-bench-lite only) | ✓ (build Feature Deletion) | scale generator suite |
|
| 76 |
| **(c)** Realistic-environment RL (RLVR) | Async sandboxes, same tool harness as production | TRL `GRPOTrainer` + verifiers + OpenEnv; SWE-bench-lite env in v0.0; build sandboxed code execution env in v0.1 | ✓ baseline | ✓ + DAPO patches | + decentralized rollouts |
|
| 77 |
| **(d)** Targeted RL w/ textual feedback (Composer's secret sauce) | Same-model self-distill: insert hint into context → teacher; original → student; on-policy KL at the turn | **Lift the OPSD/SDPO loss directly from `siyan-zhao/OPSD`** (published code, MIT). Generate hints via templates (v0.1) or LLM (v0.2). | ✗ (deferred) | ✓ (this is the Composer-recipe channel) | + learned hint generator |
|
| 78 |
| **(e)** Trace-replay multi-teacher distill (NOVEL — our addition) | N/A (not in Composer) | N=3 teachers (Opus 4.7, GPT-5, DeepSeek V4 Pro) replay each step; disagreement → DPO pairs | ✓ (this is the v0.0 novelty bet) | ✓ + VOI gating | + tiered teachers |
|
|
|
|
| 148 |
- **Cursor blog** — [Introducing Composer 2.5](https://cursor.com/blog/composer-2-5) (2026)
|
| 149 |
- **Cursor blog** — [Composer 2 technical report](https://cursor.com/blog/composer-2-technical-report) (predecessor; named the "Anyrun" environment per subagent — verify if needed)
|
| 150 |
- **OPSD paper** — Zhao et al., *Self-Distilled Reasoner: On-Policy Self-Distillation for LLMs*, [arXiv:2601.18734](https://arxiv.org/abs/2601.18734), code at [siyan-zhao/OPSD](https://github.com/siyan-zhao/OPSD). MIT.
|
| 151 |
+
- **SDPO paper** — Hübotter et al., *Reinforcement Learning via Self-Distillation*, [arXiv:2601.20802](https://arxiv.org/abs/2601.20802), ICLR 2026 Scaling Post-training Workshop. The closest published formalization; cited by Cursor only as background (deepread finding V1).
|
| 152 |
- **Self-Distillation continual-learning** — [arXiv:2601.19897](https://arxiv.org/abs/2601.19897). Cited by Cursor; less directly relevant.
|
| 153 |
- **Moonshot Kimi K2.5** — base model, [HF model card](https://huggingface.co/moonshotai/Kimi-K2-Thinking).
|
| 154 |
|
|
@@ -0,0 +1,119 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
status: accepted
|
| 3 |
+
date: 2026-06-09
|
| 4 |
+
deciders: [Codeseys, ARIA]
|
| 5 |
+
---
|
| 6 |
+
|
| 7 |
+
# ADR-016: Stage-0 dataset-generation pipeline — SWE-smith engine + rollout harness + ingest gates + single contract
|
| 8 |
+
|
| 9 |
+
## Context and Problem Statement
|
| 10 |
+
|
| 11 |
+
The user asked to "architect and build a pipeline that builds out a dataset like
|
| 12 |
+
the Composer 2.5 blog mentions, with our vision enhancements — point to an
|
| 13 |
+
open-source repo and use that to build the dataset, or use traces or other
|
| 14 |
+
datasets and enhance them."
|
| 15 |
+
|
| 16 |
+
Before building, a full-source critical review re-read every foundational paper
|
| 17 |
+
and blog (8 source clusters, `research/deepread/01-08`), ground-mapped the repo
|
| 18 |
+
(`00`), ran adversarial fidelity + design critics, and independently VERIFIED
|
| 19 |
+
every finding (`12` — 0 refuted). The verified verdict: the envisioned pipeline
|
| 20 |
+
had four structural breaks (seed-trace/oracle disjointness; no rollout harness —
|
| 21 |
+
the SFT corpus had NO producer; an uncomputable divergence gate; no
|
| 22 |
+
`Sandbox.fork()`), several missing controls (zero benchmark decontamination,
|
| 23 |
+
no secrets gate, a `golden_diff` serialization leak, two unreconciled S3
|
| 24 |
+
contracts, no cross-generation dedup), and a buy-vs-build inversion (the
|
| 25 |
+
planned image-builder duplicates `pip install swesmith`, whose PR-Mirror
|
| 26 |
+
strategy IS this repo's gold-patch-reversion mechanic and is validated best-of-
|
| 27 |
+
five by SWE-smith's own ablation, Table 5 of arXiv:2504.21798).
|
| 28 |
+
|
| 29 |
+
## Decision
|
| 30 |
+
|
| 31 |
+
Build **Stage 0 local-first** (architecture: `research/deepread/13-synthesis-architecture.md`):
|
| 32 |
+
|
| 33 |
+
1. **SWE-smith is the synthesis engine** for "point at a repo" (`[swesmith]`
|
| 34 |
+
extra; `datagen/swesmith_adapter.py` bridges its instances into
|
| 35 |
+
`FeatureDeletionTask`, handling the patch-semantics INVERSION — SWE-smith's
|
| 36 |
+
patch introduces the bug, so `golden_diff` = `reverse_unified_diff(patch)`).
|
| 37 |
+
`SweBenchAdapter` remains the bridge for SWE-bench-shaped substrates.
|
| 38 |
+
2. **Ingest gates before anything else** (`datagen/repo_gate.py`): SPDX-ish
|
| 39 |
+
license detection → three tiers (REDISTRIBUTABLE / TRAINABLE_ONLY /
|
| 40 |
+
EXCLUDED, fail-closed) + **benchmark decontamination** against the
|
| 41 |
+
SWE-bench-family eval-repo list (hard fail).
|
| 42 |
+
3. **The rollout harness is the corpus producer** (`datagen/rollout_harness.py`):
|
| 43 |
+
`collect_trajectory(env, task, policy)` runs a pluggable policy
|
| 44 |
+
(ScriptedPolicy for tests; OpenRouterPolicy stub; mini-swe-agent/SWE-agent
|
| 45 |
+
adoption is the documented upgrade) through `FeatureDeletionEnv` to
|
| 46 |
+
`_grade()`. Its env-grounded trajectories are ALSO the tree-of-work's seed
|
| 47 |
+
nodes — fixing the seed/oracle disjointness as a byproduct. `admit()` routes
|
| 48 |
+
typed signal: clean full pass → SFT; clean near-miss → DPO candidate;
|
| 49 |
+
guard-broken/hacked → quarantine (never raw negative gradient).
|
| 50 |
+
4. **One canonical trajectory IR** (`datagen/trajectory.py`): `ToolCall` (whose
|
| 51 |
+
`canonical_form()` is the v1 divergence-gate action algebra, replacing the
|
| 52 |
+
whitespace stub), `CanonicalTrajectory`, adapters from Claude Code traces
|
| 53 |
+
(explicitly UNGRADED — demoted to flat/SFT uses), and `to_policy_row()` —
|
| 54 |
+
the ONE policy-visible serializer, unit-tested to never emit
|
| 55 |
+
`golden_diff`/`deleted_symbols` (sentinel test).
|
| 56 |
+
5. **One reconciled dataset contract** (`pipeline/s3_contract.py`, supersedes
|
| 57 |
+
design-F1's and design-F2's divergent layouts): `runs/<id>/{tasks,
|
| 58 |
+
tasks_full(RESTRICTED), traj, corpus_sft, corpus_dpo, holdout, quarantine}`
|
| 59 |
+
+ `RunManifest` (counts, cost, budget, `parent_run_id` lineage, status) +
|
| 60 |
+
dataset card. Policy-safe task rows carry `golden_diff_sha256`, never the
|
| 61 |
+
diff. DiLoCo rendezvous and `wm_tuples/` are deliberately OUT (separate
|
| 62 |
+
concern; ablation-gated respectively).
|
| 63 |
+
6. **Cross-generation dedup** (`pipeline/dedup.py`): stable-hash MinHash over
|
| 64 |
+
word 5-shingles; a run can dedup against the prior generation's signature
|
| 65 |
+
file (flywheel-collapse mitigation). datasketch/LSH is the upgrade path.
|
| 66 |
+
7. **The local stage-driver** (`pipeline/build_corpus.py`): holdout-split FIRST
|
| 67 |
+
(held-out tasks never rolled out), rollouts under a budget ceiling
|
| 68 |
+
(partial-marking), typed routing, dedup, write-once-per-run idempotency.
|
| 69 |
+
|
| 70 |
+
## Fidelity corrections shipped with this ADR (deepread findings, all verified)
|
| 71 |
+
|
| 72 |
+
- **V1:** "SDPO is mathematically the same as Composer's mechanism" corrected
|
| 73 |
+
in `opsd.py` + `COMPOSER_RECIPE_MAPPING.md` — Cursor cites SDPO/OPSD as
|
| 74 |
+
*background*; our channel is a third, blog-inspired design (turn-localized
|
| 75 |
+
hint splice, live stop-grad teacher, no EMA).
|
| 76 |
+
- **V5:** fabricated numbers struck/tagged: 69.3%/Terminal-Bench parity (no
|
| 77 |
+
primary source), "24 other generators" (back-formed), "85% post-training
|
| 78 |
+
compute" (community speculation) — `research/01`, mapping doc, `research/06`,
|
| 79 |
+
`research/09`.
|
| 80 |
+
- **V7:** Streaming DiLoCo citation fixed (`diloco/__init__.py`): 2501.18512 =
|
| 81 |
+
Douillard et al.; Eager Updates = Kale et al. 2502.12996.
|
| 82 |
+
- **V11:** `teacher_replay.py` cost docstring relabeled ($0.98 = 50-state
|
| 83 |
+
synthetic trace; real sessions ~2 OOM more).
|
| 84 |
+
- **V13:** `kl_in_reward.py` "verl's only reverse-KL option" → default/
|
| 85 |
+
recommended (verl also ships a k3-family option).
|
| 86 |
+
|
| 87 |
+
## What is deliberately NOT in Stage 0
|
| 88 |
+
|
| 89 |
+
- AWS orchestration (Glue/EMR/Batch/Bedrock-batch/Step Functions) — Stage 4,
|
| 90 |
+
only after local runs are routine (finding D-9).
|
| 91 |
+
- Tree depth>1 — gated on a `Sandbox.fork()` spike + a measured divergence-gate
|
| 92 |
+
firing rate (findings D-3/D-4). Depth-1 multi-candidate rollouts need no fork.
|
| 93 |
+
- World-model `wm_tuples/` emission — gated on the P4 ablation being scheduled
|
| 94 |
+
(finding D-14; CWM evidence is mid-training, not RL-time aux head — V6).
|
| 95 |
+
- Secrets/PII scrub at trace ingest (finding V9) — REQUIRED before any raw
|
| 96 |
+
Claude Code session is uploaded to shared storage; tracked as the next
|
| 97 |
+
pipeline item. Local-only runs are unaffected.
|
| 98 |
+
|
| 99 |
+
## Acceptance gate
|
| 100 |
+
|
| 101 |
+
- [x] `repo_gate`: 53 tests (license tiers, decontamination, gate verdicts).
|
| 102 |
+
- [x] `swesmith_adapter`: 18 tests (patch INVERSION semantics, reverse round-trip,
|
| 103 |
+
strategy provenance, image conventions).
|
| 104 |
+
- [x] `trajectory` + `rollout_harness`: 13 tests (IR round-trips, SENTINEL leak
|
| 105 |
+
guard, env-grounded episode to grade 1.0 / guard-broken / near-miss,
|
| 106 |
+
admission routing).
|
| 107 |
+
- [x] `pipeline`: 12 tests (layout, manifest+budget, leak guard at the writer,
|
| 108 |
+
MinHash within-run + cross-generation, build_corpus e2e with holdout
|
| 109 |
+
exclusion + budget stop + write-once).
|
| 110 |
+
- [x] Full suite green: 511 passed / 66 skipped.
|
| 111 |
+
- [ ] Live swesmith synthesis on a real pointed-at repo (needs Docker+Linux) —
|
| 112 |
+
the documented `[~]` gate, same shape as ADR-010's Docker e2e.
|
| 113 |
+
|
| 114 |
+
## More Information
|
| 115 |
+
|
| 116 |
+
- `research/deepread/13-synthesis-architecture.md` — the architecture this implements.
|
| 117 |
+
- `research/deepread/12-verified-findings.md` — the verified finding ledger (V1–V15).
|
| 118 |
+
- `research/deepread/02-swe-task-synthesis.md` — the SWE-smith/R2E-Gym/SWE-Gym deep-read.
|
| 119 |
+
- ADR-010 (the substrate-inversion base this extends), ADR-002 (trace source).
|
|
@@ -83,6 +83,14 @@ aws = [
|
|
| 83 |
"boto3>=1.34",
|
| 84 |
"sagemaker>=2.200,<3",
|
| 85 |
]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 86 |
# Replaysim dataset normalization (per ADR-004)
|
| 87 |
#
|
| 88 |
# NOTE: data-juicer is intentionally NOT pinned as an extra. The package
|
|
|
|
| 83 |
"boto3>=1.34",
|
| 84 |
"sagemaker>=2.200,<3",
|
| 85 |
]
|
| 86 |
+
# SWE-smith task-synthesis engine (deepread finding V4 buy-vs-build verdict):
|
| 87 |
+
# the swesmith toolkit builds env images from arbitrary GitHub repos and
|
| 88 |
+
# synthesizes bugs (PR Mirror = this repo's gold-patch-reversion mechanic).
|
| 89 |
+
# LIVE synthesis needs Docker on Linux (the toolkit does not support macOS/
|
| 90 |
+
# Windows officially); the SwesmithAdapter itself needs nothing beyond core.
|
| 91 |
+
swesmith = [
|
| 92 |
+
"swesmith>=0.1",
|
| 93 |
+
]
|
| 94 |
# Replaysim dataset normalization (per ADR-004)
|
| 95 |
#
|
| 96 |
# NOTE: data-juicer is intentionally NOT pinned as an extra. The package
|
|
@@ -11,7 +11,7 @@
|
|
| 11 |
> The targeted-textual-feedback method is correctly described, but this file does **not** cite the three self-distillation papers Cursor cites in footnote 1 (OPSD `arXiv:2601.18734`, SDPO `arXiv:2601.20802`, Self-Distillation Continual Learning `arXiv:2601.19897`). The mapping document does.
|
| 12 |
|
| 13 |
## Overview
|
| 14 |
-
Cursor's Composer 2.5 is an advanced agentic coding model that powers the Cursor IDE. Released in mid-May 2026, it represents a massive leap in agentic capabilities, particularly for long-running, multi-file software engineering tasks. While the base weights are Moonshot AI's open-source **Kimi K2.5** model,
|
| 15 |
|
| 16 |
The resulting model is highly optimized for the exact constraints and tools of the Cursor environment (file edits, terminal usage, LSP interaction). Composer 2.5 is praised for having fewer "false-start" tool calls, avoiding prompt-baiting, and demonstrating a much calmer, more effective collaboration loop than its predecessors.
|
| 17 |
|
|
@@ -60,8 +60,8 @@ During post-training, Cursor employs **Sharded Muon** and **Dual Mesh HSDP (Hybr
|
|
| 60 |
## Performance Characteristics
|
| 61 |
Cursor claims Composer 2.5 achieves a Pareto-optimal tradeoff between intelligence and inference cost compared to frontier models (Opus 4.5/4.6, GPT-5.4/5.5).
|
| 62 |
|
| 63 |
-
* **Intelligence Improvements**: On Cursor's internal *CursorBench* (which tests sweeping, multi-file edits with ambiguous prompts), Composer 2.5
|
| 64 |
-
* **Frontier Parity**:
|
| 65 |
* **Cost Efficiency**:
|
| 66 |
* Standard Tier: $0.50 per 1M input / $2.50 per 1M output tokens.
|
| 67 |
* Fast Tier: $3.00 per 1M input / $15.00 per 1M output tokens.
|
|
|
|
| 11 |
> The targeted-textual-feedback method is correctly described, but this file does **not** cite the three self-distillation papers Cursor cites in footnote 1 (OPSD `arXiv:2601.18734`, SDPO `arXiv:2601.20802`, Self-Distillation Continual Learning `arXiv:2601.19897`). The mapping document does.
|
| 12 |
|
| 13 |
## Overview
|
| 14 |
+
Cursor's Composer 2.5 is an advanced agentic coding model that powers the Cursor IDE. Released in mid-May 2026, it represents a massive leap in agentic capabilities, particularly for long-running, multi-file software engineering tasks. While the base weights are Moonshot AI's open-source **Kimi K2.5** model, a large share of the compute budget went to Cursor's proprietary post-training/RL pipeline (the widely-circulated "85%" figure is community speculation, in NO primary source — deepread finding V5).
|
| 15 |
|
| 16 |
The resulting model is highly optimized for the exact constraints and tools of the Cursor environment (file edits, terminal usage, LSP interaction). Composer 2.5 is praised for having fewer "false-start" tool calls, avoiding prompt-baiting, and demonstrating a much calmer, more effective collaboration loop than its predecessors.
|
| 17 |
|
|
|
|
| 60 |
## Performance Characteristics
|
| 61 |
Cursor claims Composer 2.5 achieves a Pareto-optimal tradeoff between intelligence and inference cost compared to frontier models (Opus 4.5/4.6, GPT-5.4/5.5).
|
| 62 |
|
| 63 |
+
* **Intelligence Improvements**: On Cursor's internal *CursorBench* (which tests sweeping, multi-file edits with ambiguous prompts), Composer 2.5's score is NOT in any primary source (the circulating 69.3% figure appears in neither the 2.5 blog nor the Composer 2 techreport — deepread finding V5; the techreport's Table 1 gives Composer 2 = 61.3 CursorBench). Treat all 2.5 benchmark numbers as unverified.
|
| 64 |
+
* **Frontier Parity**: Claims of Terminal-Bench 2.0 / SWE-bench Multilingual parity circulate in secondary commentary only; neither primary source contains benchmark numbers for 2.5 (deepread finding V5).
|
| 65 |
* **Cost Efficiency**:
|
| 66 |
* Standard Tier: $0.50 per 1M input / $2.50 per 1M output tokens.
|
| 67 |
* Fast Tier: $3.00 per 1M input / $15.00 per 1M output tokens.
|
|
@@ -327,7 +327,7 @@ Feature-Deletion is **embarrassingly parallel and CPU-bound** — no GPU in the
|
|
| 327 |
|
| 328 |
1. **Deletion-target selection heuristic** — blog silent (`research/09` §1 "NO CHANGE"). We propose coverage-selectivity (§5 Path B); Cursor's actual heuristic is unknown.
|
| 329 |
2. **Deleter model vs. program** — blog implies an agent deletes ("asked to delete code… such that the codebase remains functional"); we default to *programmatic* deletion (cheaper, deterministic, no second model). An LLM-deleter is a v0.2 escalation.
|
| 330 |
-
3. **The other
|
| 331 |
4. **"Agentic monitoring tools" internals** — unspecified; our §3c monitor is a best-effort programmatic stand-in.
|
| 332 |
5. **Composer2.pdf (arXiv:2603.24477)** — flagged by `research/09` action-item #1 as the likely home of data-mix % and generator inventory; **not yet extracted**. Recommend a follow-up pull before scaling the generator suite.
|
| 333 |
|
|
|
|
| 327 |
|
| 328 |
1. **Deletion-target selection heuristic** — blog silent (`research/09` §1 "NO CHANGE"). We propose coverage-selectivity (§5 Path B); Cursor's actual heuristic is unknown.
|
| 329 |
2. **Deleter model vs. program** — blog implies an agent deletes ("asked to delete code… such that the codebase remains functional"); we default to *programmatic* deletion (cheaper, deterministic, no second model). An LLM-deleter is a v0.2 escalation.
|
| 330 |
+
3. **The other generators (count UNKNOWN)** — Feature Deletion is "one synthetic approach… a range of approaches"; the rest are unnamed and uncounted (the old "~24" was a back-formation from the 25x task multiplier — deepread finding V5). Out of scope here; this brief delivers the one named generator.
|
| 331 |
4. **"Agentic monitoring tools" internals** — unspecified; our §3c monitor is a best-effort programmatic stand-in.
|
| 332 |
5. **Composer2.pdf (arXiv:2603.24477)** — flagged by `research/09` action-item #1 as the likely home of data-mix % and generator inventory; **not yet extracted**. Recommend a follow-up pull before scaling the generator suite.
|
| 333 |
|
|
@@ -20,7 +20,7 @@ The **2.5 blog body is byte-for-byte unchanged** from what the mapping doc captu
|
|
| 20 |
|
| 21 |
**DELTAS (not in / under-stated in COMPOSER_RECIPE_MAPPING.md):**
|
| 22 |
|
| 23 |
-
- **[DELTA — new emphasis]** The phrase *"we both **select for** and **create** harder tasks **dynamically throughout the run**"* is a **dynamic curriculum / online task-selection** signal. The mapping doc captured "Feature Deletion +
|
| 24 |
- **[DELTA — new authoritative source for CPT data mix]** The Composer 2 technical-report blog states the CPT data mix explicitly: *"continued pretraining on a data mix that **emphasizes code** to deepen the base model's coding knowledge"* and *"We find that **reducing pretraining loss improves downstream RL performance**, with better base knowledge reliably translating into a better agent."* The mapping doc marked "continued pretraining on heavily code-weighted data" as `[BLOG-VERIFIED]` from the 2.5 Muon section — but the **causal claim (CPT loss ↓ ⇒ RL performance ↑)** is new and is the stated *justification* for doing CPT at all. Relevant to our "skip CPT, start from Qwen3-Coder" decision: Cursor's own evidence says base-knowledge quality gates RL ceiling, which strengthens the case for starting from an already-code-tuned base.
|
| 25 |
- **[DELTA — new artifact]** There is now a **full Composer 2 arXiv technical report: [arXiv:2603.24477](https://arxiv.org/abs/2603.24477)** and a downloadable PDF at **`https://cursor.com/resources/Composer2.pdf`** (authored by Sasha Rush et al.). The report explicitly *"covers... ablations on the training recipe, our approach to agent behavior shaping, and the design of our evaluation suite."* The mapping doc cited only the blog stub and never the arXiv ID/PDF. **This PDF is the most likely place to resolve the data-mix weighting %, the RL algorithm name, and the hint-generation mechanism — none of which are in either blog.** → Recommend a dedicated follow-up extraction of Composer2.pdf.
|
| 26 |
- **[CONFIRM — "Anyrun"]** Mapping doc flagged "Anyrun" as possibly not Cursor-sourced. **Confirmed real:** the Composer 2 report blog says *"**Anyrun**, our internal compute platform for running hundreds of thousands of sandboxed coding environments."* It is a Composer-**2** artifact (carried into 2.5), correctly attributed. Resolves the mapping doc's open flag.
|
|
|
|
| 20 |
|
| 21 |
**DELTAS (not in / under-stated in COMPOSER_RECIPE_MAPPING.md):**
|
| 22 |
|
| 23 |
+
- **[DELTA — new emphasis]** The phrase *"we both **select for** and **create** harder tasks **dynamically throughout the run**"* is a **dynamic curriculum / online task-selection** signal. The mapping doc captured "Feature Deletion + other unnamed generators" (its old "24" count was a back-formation — deepread finding V5) but did **not** flag that task difficulty is filtered *online* (the model "begins to get most training problems correct," so hard tasks are up-weighted live). This is a data-*mix*/curriculum detail with direct replication impact: our generator suite needs a difficulty filter / pass-rate gate, not just a static task bank.
|
| 24 |
- **[DELTA — new authoritative source for CPT data mix]** The Composer 2 technical-report blog states the CPT data mix explicitly: *"continued pretraining on a data mix that **emphasizes code** to deepen the base model's coding knowledge"* and *"We find that **reducing pretraining loss improves downstream RL performance**, with better base knowledge reliably translating into a better agent."* The mapping doc marked "continued pretraining on heavily code-weighted data" as `[BLOG-VERIFIED]` from the 2.5 Muon section — but the **causal claim (CPT loss ↓ ⇒ RL performance ↑)** is new and is the stated *justification* for doing CPT at all. Relevant to our "skip CPT, start from Qwen3-Coder" decision: Cursor's own evidence says base-knowledge quality gates RL ceiling, which strengthens the case for starting from an already-code-tuned base.
|
| 25 |
- **[DELTA — new artifact]** There is now a **full Composer 2 arXiv technical report: [arXiv:2603.24477](https://arxiv.org/abs/2603.24477)** and a downloadable PDF at **`https://cursor.com/resources/Composer2.pdf`** (authored by Sasha Rush et al.). The report explicitly *"covers... ablations on the training recipe, our approach to agent behavior shaping, and the design of our evaluation suite."* The mapping doc cited only the blog stub and never the arXiv ID/PDF. **This PDF is the most likely place to resolve the data-mix weighting %, the RL algorithm name, and the hint-generation mechanism — none of which are in either blog.** → Recommend a dedicated follow-up extraction of Composer2.pdf.
|
| 26 |
- **[CONFIRM — "Anyrun"]** Mapping doc flagged "Anyrun" as possibly not Cursor-sourced. **Confirmed real:** the Composer 2 report blog says *"**Anyrun**, our internal compute platform for running hundreds of thousands of sandboxed coding environments."* It is a Composer-**2** artifact (carried into 2.5), correctly attributed. Resolves the mapping doc's open flag.
|
|
@@ -0,0 +1,224 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2304.06767] RAFT: Reward rAnked FineTuning for Generative Foundation Model
|
| 3 |
+
Alignment'
|
| 4 |
+
id: 230406767-raft-reward-ranked-finetuning-for-generative-foundation-model-alignmen
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:31:18.566124Z'
|
| 8 |
+
source: https://arxiv.org/abs/2304.06767
|
| 9 |
+
source_domain: arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:31:18.565918Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2304.06767] RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
|
| 20 |
+
Computer Science > Machine Learning
|
| 21 |
+
arXiv:2304.06767
|
| 22 |
+
(cs)
|
| 23 |
+
[Submitted on 13 Apr 2023 (
|
| 24 |
+
v1
|
| 25 |
+
), last revised 1 Dec 2023 (this version, v4)]
|
| 26 |
+
Title:
|
| 27 |
+
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
|
| 28 |
+
Authors:
|
| 29 |
+
Hanze Dong
|
| 30 |
+
,
|
| 31 |
+
Wei Xiong
|
| 32 |
+
,
|
| 33 |
+
Deepanshu Goyal
|
| 34 |
+
,
|
| 35 |
+
Yihan Zhang
|
| 36 |
+
,
|
| 37 |
+
Winnie Chow
|
| 38 |
+
,
|
| 39 |
+
Rui Pan
|
| 40 |
+
,
|
| 41 |
+
Shizhe Diao
|
| 42 |
+
,
|
| 43 |
+
Jipeng Zhang
|
| 44 |
+
,
|
| 45 |
+
Kashun Shum
|
| 46 |
+
,
|
| 47 |
+
Tong Zhang
|
| 48 |
+
View a PDF of the paper titled RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment, by Hanze Dong and 9 other authors
|
| 49 |
+
View PDF
|
| 50 |
+
HTML (experimental)
|
| 51 |
+
Abstract:
|
| 52 |
+
Generative foundation models are susceptible to implicit biases that can arise from extensive unsupervised training data. Such biases can produce suboptimal samples, skewed outcomes, and unfairness, with potentially serious consequences. Consequently, aligning these models with human ethics and preferences is an essential step toward ensuring their responsible and effective deployment in real-world applications. Prior research has primarily employed Reinforcement Learning from Human Feedback (RLHF) to address this problem, where generative models are fine-tuned with RL algorithms guided by a human-feedback-informed reward model. However, the inefficiencies and instabilities associated with RL algorithms frequently present substantial obstacles to the successful alignment, necessitating the development of a more robust and streamlined approach. To this end, we introduce a new framework, Reward rAnked FineTuning (RAFT), designed to align generative models effectively. Utilizing a reward model and a sufficient number of samples, our approach selects the high-quality samples, discarding those that exhibit undesired behavior, and subsequently enhancing the model by fine-tuning on these filtered samples. Our studies show that RAFT can effectively improve the model performance in both reward learning and other automated metrics in both large language models and diffusion models.
|
| 53 |
+
Comments:
|
| 54 |
+
29 pages, 12 figures, Published in Transactions on Machine Learning Research (TMLR)
|
| 55 |
+
Subjects:
|
| 56 |
+
Machine Learning (cs.LG)
|
| 57 |
+
; Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
|
| 58 |
+
Cite as:
|
| 59 |
+
arXiv:2304.06767
|
| 60 |
+
[cs.LG]
|
| 61 |
+
(or
|
| 62 |
+
arXiv:2304.06767v4
|
| 63 |
+
[cs.LG]
|
| 64 |
+
for this version)
|
| 65 |
+
https://doi.org/10.48550/arXiv.2304.06767
|
| 66 |
+
Focus to learn more
|
| 67 |
+
arXiv-issued DOI via DataCite
|
| 68 |
+
Submission history
|
| 69 |
+
From: Hanze Dong [
|
| 70 |
+
view email
|
| 71 |
+
]
|
| 72 |
+
[v1]
|
| 73 |
+
Thu, 13 Apr 2023 18:22:40 UTC (62,967 KB)
|
| 74 |
+
[v2]
|
| 75 |
+
Thu, 25 May 2023 06:27:31 UTC (42,022 KB)
|
| 76 |
+
[v3]
|
| 77 |
+
Wed, 30 Aug 2023 01:25:29 UTC (33,955 KB)
|
| 78 |
+
[v4]
|
| 79 |
+
Fri, 1 Dec 2023 14:28:06 UTC (34,049 KB)
|
| 80 |
+
Full-text links:
|
| 81 |
+
Access Paper:
|
| 82 |
+
View a PDF of the paper titled RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment, by Hanze Dong and 9 other authors
|
| 83 |
+
View PDF
|
| 84 |
+
HTML (experimental)
|
| 85 |
+
TeX Source
|
| 86 |
+
view license
|
| 87 |
+
Current browse context:
|
| 88 |
+
cs.LG
|
| 89 |
+
< prev
|
| 90 |
+
|
|
| 91 |
+
next >
|
| 92 |
+
new
|
| 93 |
+
|
|
| 94 |
+
recent
|
| 95 |
+
|
|
| 96 |
+
2023-04
|
| 97 |
+
Change to browse by:
|
| 98 |
+
cs
|
| 99 |
+
cs.AI
|
| 100 |
+
cs.CL
|
| 101 |
+
cs.CV
|
| 102 |
+
stat
|
| 103 |
+
stat.ML
|
| 104 |
+
References & Citations
|
| 105 |
+
NASA ADS
|
| 106 |
+
Google Scholar
|
| 107 |
+
Semantic Scholar
|
| 108 |
+
export BibTeX citation
|
| 109 |
+
Loading...
|
| 110 |
+
BibTeX formatted citation
|
| 111 |
+
×
|
| 112 |
+
loading...
|
| 113 |
+
Data provided by:
|
| 114 |
+
Bookmark
|
| 115 |
+
Bibliographic Tools
|
| 116 |
+
Bibliographic and Citation Tools
|
| 117 |
+
Bibliographic Explorer Toggle
|
| 118 |
+
Bibliographic Explorer
|
| 119 |
+
(
|
| 120 |
+
What is the Explorer?
|
| 121 |
+
)
|
| 122 |
+
Connected Papers Toggle
|
| 123 |
+
Connected Papers
|
| 124 |
+
(
|
| 125 |
+
What is Connected Papers?
|
| 126 |
+
)
|
| 127 |
+
Litmaps Toggle
|
| 128 |
+
Litmaps
|
| 129 |
+
(
|
| 130 |
+
What is Litmaps?
|
| 131 |
+
)
|
| 132 |
+
scite.ai Toggle
|
| 133 |
+
scite Smart Citations
|
| 134 |
+
(
|
| 135 |
+
What are Smart Citations?
|
| 136 |
+
)
|
| 137 |
+
Code, Data, Media
|
| 138 |
+
Code, Data and Media Associated with this Article
|
| 139 |
+
alphaXiv Toggle
|
| 140 |
+
alphaXiv
|
| 141 |
+
(
|
| 142 |
+
What is alphaXiv?
|
| 143 |
+
)
|
| 144 |
+
Links to Code Toggle
|
| 145 |
+
CatalyzeX Code Finder for Papers
|
| 146 |
+
(
|
| 147 |
+
What is CatalyzeX?
|
| 148 |
+
)
|
| 149 |
+
DagsHub Toggle
|
| 150 |
+
DagsHub
|
| 151 |
+
(
|
| 152 |
+
What is DagsHub?
|
| 153 |
+
)
|
| 154 |
+
GotitPub Toggle
|
| 155 |
+
Gotit.pub
|
| 156 |
+
(
|
| 157 |
+
What is GotitPub?
|
| 158 |
+
)
|
| 159 |
+
Huggingface Toggle
|
| 160 |
+
Hugging Face
|
| 161 |
+
(
|
| 162 |
+
What is Huggingface?
|
| 163 |
+
)
|
| 164 |
+
Links to Code Toggle
|
| 165 |
+
Papers with Code
|
| 166 |
+
(
|
| 167 |
+
What is Papers with Code?
|
| 168 |
+
)
|
| 169 |
+
ScienceCast Toggle
|
| 170 |
+
ScienceCast
|
| 171 |
+
(
|
| 172 |
+
What is ScienceCast?
|
| 173 |
+
)
|
| 174 |
+
Demos
|
| 175 |
+
Demos
|
| 176 |
+
Replicate Toggle
|
| 177 |
+
Replicate
|
| 178 |
+
(
|
| 179 |
+
What is Replicate?
|
| 180 |
+
)
|
| 181 |
+
Spaces Toggle
|
| 182 |
+
Hugging Face Spaces
|
| 183 |
+
(
|
| 184 |
+
What is Spaces?
|
| 185 |
+
)
|
| 186 |
+
Spaces Toggle
|
| 187 |
+
TXYZ.AI
|
| 188 |
+
(
|
| 189 |
+
What is TXYZ.AI?
|
| 190 |
+
)
|
| 191 |
+
Related Papers
|
| 192 |
+
Recommenders and Search Tools
|
| 193 |
+
Link to Influence Flower
|
| 194 |
+
Influence Flower
|
| 195 |
+
(
|
| 196 |
+
What are Influence Flowers?
|
| 197 |
+
)
|
| 198 |
+
Core recommender toggle
|
| 199 |
+
CORE Recommender
|
| 200 |
+
(
|
| 201 |
+
What is CORE?
|
| 202 |
+
)
|
| 203 |
+
IArxiv recommender toggle
|
| 204 |
+
IArxiv Recommender
|
| 205 |
+
(
|
| 206 |
+
What is IArxiv?
|
| 207 |
+
)
|
| 208 |
+
Author
|
| 209 |
+
Venue
|
| 210 |
+
Institution
|
| 211 |
+
Topic
|
| 212 |
+
About arXivLabs
|
| 213 |
+
arXivLabs: experimental projects with community collaborators
|
| 214 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 215 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 216 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 217 |
+
Learn more about arXivLabs
|
| 218 |
+
.
|
| 219 |
+
Which authors of this paper are endorsers?
|
| 220 |
+
|
|
| 221 |
+
Disable MathJax
|
| 222 |
+
(
|
| 223 |
+
What is MathJax?
|
| 224 |
+
)
|
|
@@ -0,0 +1,2735 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2305.10601] Tree of Thoughts: Deliberate Problem Solving with Large Language
|
| 3 |
+
Models'
|
| 4 |
+
id: 230510601-tree-of-thoughts-deliberate-problem-solving-with-large-language-models-2
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:41:12.876142Z'
|
| 8 |
+
source: https://ar5iv.labs.arxiv.org/html/2305.10601
|
| 9 |
+
source_domain: ar5iv.labs.arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:41:12.875985Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2305.10601] Tree of Thoughts: Deliberate Problem Solving with Large Language Models
|
| 20 |
+
Tree of Thoughts: Deliberate Problem Solving
|
| 21 |
+
with Large Language Models
|
| 22 |
+
Shunyu Yao
|
| 23 |
+
Princeton University
|
| 24 |
+
Dian Yu
|
| 25 |
+
Google DeepMind
|
| 26 |
+
Jeffrey Zhao
|
| 27 |
+
Google DeepMind
|
| 28 |
+
Izhak Shafran
|
| 29 |
+
Google DeepMind
|
| 30 |
+
Thomas L. Griffiths
|
| 31 |
+
Princeton University
|
| 32 |
+
Yuan Cao
|
| 33 |
+
Google DeepMind
|
| 34 |
+
Karthik Narasimhan
|
| 35 |
+
Princeton University
|
| 36 |
+
Abstract
|
| 37 |
+
Language models are increasingly being deployed for general problem solving across a wide range of tasks, but are still confined to token-level, left-to-right decision-making processes during inference. This means they can fall short in tasks that require exploration, strategic lookahead, or where initial decisions play a pivotal role.
|
| 38 |
+
To surmount these challenges, we introduce a new framework for language model inference, “Tree of Thoughts” (ToT), which generalizes over the popular “Chain of Thought” approach to prompting language models, and enables exploration over coherent units of text (“thoughts”) that serve as intermediate steps toward problem solving.
|
| 39 |
+
ToT allows LMs to perform deliberate decision making by considering multiple different reasoning paths and self-evaluating choices to decide the next course of action, as well as looking ahead or backtracking when necessary to make global choices.
|
| 40 |
+
Our experiments show that ToT significantly enhances language models’ problem-solving abilities on three novel tasks requiring non-trivial planning or search: Game of 24, Creative Writing, and Mini Crosswords.
|
| 41 |
+
For instance, in Game of 24, while GPT-4 with chain-of-thought prompting only solved 4% of tasks, our method achieved a success rate of 74%. Code repo with all prompts:
|
| 42 |
+
https://github.com/princeton-nlp/tree-of-thought-llm
|
| 43 |
+
.
|
| 44 |
+
1
|
| 45 |
+
Introduction
|
| 46 |
+
Originally designed to generate text, scaled-up versions of language models (LMs) such as GPT
|
| 47 |
+
[
|
| 48 |
+
25
|
| 49 |
+
,
|
| 50 |
+
26
|
| 51 |
+
,
|
| 52 |
+
1
|
| 53 |
+
,
|
| 54 |
+
23
|
| 55 |
+
]
|
| 56 |
+
and PaLM
|
| 57 |
+
[
|
| 58 |
+
5
|
| 59 |
+
]
|
| 60 |
+
have been shown to be increasingly capable of performing an ever wider range of tasks requiring mathematical, symbolic, commonsense, and knowledge reasoning. It is perhaps surprising that underlying all this progress is still the original autoregressive mechanism for generating text, which makes token-level decisions one by one and in a left-to-right fashion.
|
| 61 |
+
Is such a simple mechanism sufficient for a LM to be built toward a general problem solver?
|
| 62 |
+
If not, what problems would challenge the current paradigm, and what should be alternative mechanisms?
|
| 63 |
+
The literature on human cognition provides some clues to answer these questions.
|
| 64 |
+
Research on “dual process” models suggests that people have two modes in which they engage with decisions – a fast, automatic, unconscious mode (“System 1”) and a slow, deliberate, conscious mode (“System 2”)
|
| 65 |
+
[
|
| 66 |
+
30
|
| 67 |
+
,
|
| 68 |
+
31
|
| 69 |
+
,
|
| 70 |
+
16
|
| 71 |
+
,
|
| 72 |
+
15
|
| 73 |
+
]
|
| 74 |
+
.
|
| 75 |
+
These two modes have previously been connected to a variety of mathematical models used in machine learning. For example, research on reinforcement learning in humans and other animals has explored the circumstances under which they engage in associative “model free” learning or more deliberative “model based” planning
|
| 76 |
+
[
|
| 77 |
+
7
|
| 78 |
+
]
|
| 79 |
+
.
|
| 80 |
+
The simple associative token-level choices of LMs are also reminiscent of “System 1”, and thus might benefit from augmentation by a more deliberate “System 2” planning process that (1) maintains and explores diverse alternatives for current choices instead of just picking one, and (2) evaluates its current status and actively looks ahead or backtracks to make more global decisions.
|
| 81 |
+
To design such a planning process, we return to the origins of artificial intelligence (and cognitive science), drawing inspiration from the planning processes explored by Newell, Shaw, and Simon starting in the 1950s
|
| 82 |
+
[
|
| 83 |
+
21
|
| 84 |
+
,
|
| 85 |
+
22
|
| 86 |
+
]
|
| 87 |
+
. Newell and colleagues characterized
|
| 88 |
+
problem solving
|
| 89 |
+
[
|
| 90 |
+
21
|
| 91 |
+
]
|
| 92 |
+
as search through a combinatorial problem space, represented as a tree. We thus propose the Tree of Thoughts (ToT) framework for general problem solving with language models. As Figure
|
| 93 |
+
1
|
| 94 |
+
illustrates, while existing methods (detailed below) sample continuous language sequences for problem solving, ToT actively maintains a tree of thoughts, where each
|
| 95 |
+
thought
|
| 96 |
+
is a coherent language sequence that serves as an intermediate step toward problem solving (Table
|
| 97 |
+
1
|
| 98 |
+
). Such a high-level semantic unit allows the LM to self-evaluate the progress different intermediate thoughts make towards solving the problem through a deliberate reasoning process that is also instantiated in language (Figures
|
| 99 |
+
2
|
| 100 |
+
,
|
| 101 |
+
4
|
| 102 |
+
,
|
| 103 |
+
6
|
| 104 |
+
). This implementation of search heuristics via LM self-evaluation and deliberation is novel, as previous search heuristics are either programmed or learned. Finally, we combine this language-based capability to generate and evaluate diverse thoughts with search algorithms, such as breadth-first search (BFS) or depth-first search (DFS), which allow systematic exploration of the tree of thoughts with lookahead and backtracking.
|
| 105 |
+
Empirically, we propose three new problems that challenge existing LM inference methods even with the state-of-the-art language model, GPT-4
|
| 106 |
+
[
|
| 107 |
+
23
|
| 108 |
+
]
|
| 109 |
+
: Game of 24, Creative Writing, and Crosswords (Table
|
| 110 |
+
1
|
| 111 |
+
).
|
| 112 |
+
These tasks require deductive, mathematical, commonsense, lexical reasoning abilities, and a way to incorporate systematic planning or search.
|
| 113 |
+
We show ToT obtains superior results on all three tasks by being general and flexible enough to support different levels of thoughts, different ways to generate and evaluate thoughts, and different search algorithms that adapt to the nature of different problems. We also analyze how such choices affect model performances via systematic ablations and discuss future directions to better train and use LMs.
|
| 114 |
+
Figure 1:
|
| 115 |
+
Schematic illustrating various approaches to problem solving with LLMs. Each rectangle box represents a
|
| 116 |
+
thought
|
| 117 |
+
, which is a coherent language sequence that serves as an intermediate step toward problem solving. See concrete examples of how thoughts are generated, evaluated, and searched in Figures
|
| 118 |
+
2
|
| 119 |
+
,
|
| 120 |
+
4
|
| 121 |
+
,
|
| 122 |
+
6
|
| 123 |
+
.
|
| 124 |
+
2
|
| 125 |
+
Background
|
| 126 |
+
We first formalize some existing methods that use large language models for problem-solving, which our approach is inspired by and later compared with.
|
| 127 |
+
We use
|
| 128 |
+
p
|
| 129 |
+
θ
|
| 130 |
+
subscript
|
| 131 |
+
𝑝
|
| 132 |
+
𝜃
|
| 133 |
+
p_{\theta}
|
| 134 |
+
to denote a pre-trained LM with parameters
|
| 135 |
+
θ
|
| 136 |
+
𝜃
|
| 137 |
+
\theta
|
| 138 |
+
, and
|
| 139 |
+
lowercase letters
|
| 140 |
+
x
|
| 141 |
+
,
|
| 142 |
+
y
|
| 143 |
+
,
|
| 144 |
+
z
|
| 145 |
+
,
|
| 146 |
+
s
|
| 147 |
+
,
|
| 148 |
+
⋯
|
| 149 |
+
𝑥
|
| 150 |
+
𝑦
|
| 151 |
+
𝑧
|
| 152 |
+
𝑠
|
| 153 |
+
⋯
|
| 154 |
+
x,y,z,s,\cdots
|
| 155 |
+
to denote a language sequence
|
| 156 |
+
, i.e.
|
| 157 |
+
x
|
| 158 |
+
=
|
| 159 |
+
(
|
| 160 |
+
x
|
| 161 |
+
|
| 162 |
+
[
|
| 163 |
+
1
|
| 164 |
+
]
|
| 165 |
+
,
|
| 166 |
+
⋯
|
| 167 |
+
,
|
| 168 |
+
x
|
| 169 |
+
|
| 170 |
+
[
|
| 171 |
+
n
|
| 172 |
+
]
|
| 173 |
+
)
|
| 174 |
+
𝑥
|
| 175 |
+
𝑥
|
| 176 |
+
delimited-[]
|
| 177 |
+
1
|
| 178 |
+
⋯
|
| 179 |
+
𝑥
|
| 180 |
+
delimited-[]
|
| 181 |
+
𝑛
|
| 182 |
+
x=(x[1],\cdots,x[n])
|
| 183 |
+
where each
|
| 184 |
+
x
|
| 185 |
+
|
| 186 |
+
[
|
| 187 |
+
i
|
| 188 |
+
]
|
| 189 |
+
𝑥
|
| 190 |
+
delimited-[]
|
| 191 |
+
𝑖
|
| 192 |
+
x[i]
|
| 193 |
+
is a token, so that
|
| 194 |
+
p
|
| 195 |
+
θ
|
| 196 |
+
|
| 197 |
+
(
|
| 198 |
+
x
|
| 199 |
+
)
|
| 200 |
+
=
|
| 201 |
+
∏
|
| 202 |
+
i
|
| 203 |
+
=
|
| 204 |
+
1
|
| 205 |
+
n
|
| 206 |
+
p
|
| 207 |
+
θ
|
| 208 |
+
|
| 209 |
+
(
|
| 210 |
+
x
|
| 211 |
+
|
| 212 |
+
[
|
| 213 |
+
i
|
| 214 |
+
]
|
| 215 |
+
|
|
| 216 |
+
x
|
| 217 |
+
|
| 218 |
+
[
|
| 219 |
+
1
|
| 220 |
+
|
| 221 |
+
…
|
| 222 |
+
|
| 223 |
+
i
|
| 224 |
+
]
|
| 225 |
+
)
|
| 226 |
+
subscript
|
| 227 |
+
𝑝
|
| 228 |
+
𝜃
|
| 229 |
+
𝑥
|
| 230 |
+
superscript
|
| 231 |
+
subscript
|
| 232 |
+
product
|
| 233 |
+
𝑖
|
| 234 |
+
1
|
| 235 |
+
𝑛
|
| 236 |
+
subscript
|
| 237 |
+
𝑝
|
| 238 |
+
𝜃
|
| 239 |
+
conditional
|
| 240 |
+
𝑥
|
| 241 |
+
delimited-[]
|
| 242 |
+
𝑖
|
| 243 |
+
𝑥
|
| 244 |
+
delimited-[]
|
| 245 |
+
1
|
| 246 |
+
…
|
| 247 |
+
𝑖
|
| 248 |
+
p_{\theta}(x)=\prod_{i=1}^{n}p_{\theta}(x[i]|x[1...i])
|
| 249 |
+
. We use uppercase letters
|
| 250 |
+
S
|
| 251 |
+
,
|
| 252 |
+
⋯
|
| 253 |
+
𝑆
|
| 254 |
+
⋯
|
| 255 |
+
S,\cdots
|
| 256 |
+
to denote a collection of language sequences.
|
| 257 |
+
Input-output (IO) prompting
|
| 258 |
+
is the most common way to turn a problem input
|
| 259 |
+
x
|
| 260 |
+
𝑥
|
| 261 |
+
x
|
| 262 |
+
into output
|
| 263 |
+
y
|
| 264 |
+
𝑦
|
| 265 |
+
y
|
| 266 |
+
with LM:
|
| 267 |
+
y
|
| 268 |
+
∼
|
| 269 |
+
p
|
| 270 |
+
θ
|
| 271 |
+
|
| 272 |
+
(
|
| 273 |
+
y
|
| 274 |
+
|
|
| 275 |
+
prompt
|
| 276 |
+
I
|
| 277 |
+
|
| 278 |
+
O
|
| 279 |
+
|
| 280 |
+
(
|
| 281 |
+
x
|
| 282 |
+
)
|
| 283 |
+
)
|
| 284 |
+
similar-to
|
| 285 |
+
𝑦
|
| 286 |
+
subscript
|
| 287 |
+
𝑝
|
| 288 |
+
𝜃
|
| 289 |
+
conditional
|
| 290 |
+
𝑦
|
| 291 |
+
subscript
|
| 292 |
+
prompt
|
| 293 |
+
𝐼
|
| 294 |
+
𝑂
|
| 295 |
+
𝑥
|
| 296 |
+
y\sim p_{\theta}(y|\texttt{prompt}_{{IO}}(x))
|
| 297 |
+
, where
|
| 298 |
+
prompt
|
| 299 |
+
I
|
| 300 |
+
|
| 301 |
+
O
|
| 302 |
+
|
| 303 |
+
(
|
| 304 |
+
x
|
| 305 |
+
)
|
| 306 |
+
subscript
|
| 307 |
+
prompt
|
| 308 |
+
𝐼
|
| 309 |
+
𝑂
|
| 310 |
+
𝑥
|
| 311 |
+
\texttt{prompt}_{IO}(x)
|
| 312 |
+
wraps input
|
| 313 |
+
x
|
| 314 |
+
𝑥
|
| 315 |
+
x
|
| 316 |
+
with task instructions and/or few-shot input-output examples. For simplicity, let us denote
|
| 317 |
+
p
|
| 318 |
+
θ
|
| 319 |
+
prompt
|
| 320 |
+
|
| 321 |
+
(
|
| 322 |
+
output
|
| 323 |
+
∣
|
| 324 |
+
input
|
| 325 |
+
)
|
| 326 |
+
=
|
| 327 |
+
p
|
| 328 |
+
θ
|
| 329 |
+
|
| 330 |
+
(
|
| 331 |
+
output
|
| 332 |
+
∣
|
| 333 |
+
prompt
|
| 334 |
+
|
| 335 |
+
(
|
| 336 |
+
input
|
| 337 |
+
)
|
| 338 |
+
)
|
| 339 |
+
superscript
|
| 340 |
+
subscript
|
| 341 |
+
𝑝
|
| 342 |
+
𝜃
|
| 343 |
+
prompt
|
| 344 |
+
conditional
|
| 345 |
+
output
|
| 346 |
+
input
|
| 347 |
+
subscript
|
| 348 |
+
𝑝
|
| 349 |
+
𝜃
|
| 350 |
+
conditional
|
| 351 |
+
output
|
| 352 |
+
prompt
|
| 353 |
+
input
|
| 354 |
+
p_{\theta}^{{\rm prompt}}(\texttt{output}\mid\texttt{input})=p_{\theta}(\texttt{output}\mid\texttt{prompt}(\texttt{input}))
|
| 355 |
+
, so that IO prompting can be formulated as
|
| 356 |
+
y
|
| 357 |
+
∼
|
| 358 |
+
p
|
| 359 |
+
θ
|
| 360 |
+
I
|
| 361 |
+
|
| 362 |
+
O
|
| 363 |
+
|
| 364 |
+
(
|
| 365 |
+
y
|
| 366 |
+
|
|
| 367 |
+
x
|
| 368 |
+
)
|
| 369 |
+
similar-to
|
| 370 |
+
𝑦
|
| 371 |
+
superscript
|
| 372 |
+
subscript
|
| 373 |
+
𝑝
|
| 374 |
+
𝜃
|
| 375 |
+
𝐼
|
| 376 |
+
𝑂
|
| 377 |
+
conditional
|
| 378 |
+
𝑦
|
| 379 |
+
𝑥
|
| 380 |
+
y\sim p_{\theta}^{IO}(y|x)
|
| 381 |
+
.
|
| 382 |
+
Chain-of-thought (CoT) prompting
|
| 383 |
+
[
|
| 384 |
+
38
|
| 385 |
+
]
|
| 386 |
+
was proposed to address cases where the mapping of input
|
| 387 |
+
x
|
| 388 |
+
𝑥
|
| 389 |
+
x
|
| 390 |
+
to output
|
| 391 |
+
y
|
| 392 |
+
𝑦
|
| 393 |
+
y
|
| 394 |
+
is non-trivial (e.g. when
|
| 395 |
+
x
|
| 396 |
+
𝑥
|
| 397 |
+
x
|
| 398 |
+
is a math question and
|
| 399 |
+
y
|
| 400 |
+
𝑦
|
| 401 |
+
y
|
| 402 |
+
is the final numerical answer). The key idea is to introduce a chain of
|
| 403 |
+
thoughts
|
| 404 |
+
z
|
| 405 |
+
1
|
| 406 |
+
,
|
| 407 |
+
⋯
|
| 408 |
+
,
|
| 409 |
+
z
|
| 410 |
+
n
|
| 411 |
+
subscript
|
| 412 |
+
𝑧
|
| 413 |
+
1
|
| 414 |
+
⋯
|
| 415 |
+
subscript
|
| 416 |
+
𝑧
|
| 417 |
+
𝑛
|
| 418 |
+
z_{1},\cdots,z_{n}
|
| 419 |
+
to bridge
|
| 420 |
+
x
|
| 421 |
+
𝑥
|
| 422 |
+
x
|
| 423 |
+
and
|
| 424 |
+
y
|
| 425 |
+
𝑦
|
| 426 |
+
y
|
| 427 |
+
, where each
|
| 428 |
+
z
|
| 429 |
+
i
|
| 430 |
+
subscript
|
| 431 |
+
𝑧
|
| 432 |
+
𝑖
|
| 433 |
+
z_{i}
|
| 434 |
+
is a coherent language sequence that serves as a meaningful intermediate step toward problem solving (e.g.
|
| 435 |
+
z
|
| 436 |
+
i
|
| 437 |
+
subscript
|
| 438 |
+
𝑧
|
| 439 |
+
𝑖
|
| 440 |
+
z_{i}
|
| 441 |
+
could be an intermediate equation for math QA). To solve problems with CoT, each thought
|
| 442 |
+
z
|
| 443 |
+
i
|
| 444 |
+
∼
|
| 445 |
+
p
|
| 446 |
+
θ
|
| 447 |
+
C
|
| 448 |
+
|
| 449 |
+
o
|
| 450 |
+
|
| 451 |
+
T
|
| 452 |
+
|
| 453 |
+
(
|
| 454 |
+
z
|
| 455 |
+
i
|
| 456 |
+
∣
|
| 457 |
+
x
|
| 458 |
+
,
|
| 459 |
+
z
|
| 460 |
+
1
|
| 461 |
+
|
| 462 |
+
⋯
|
| 463 |
+
|
| 464 |
+
i
|
| 465 |
+
−
|
| 466 |
+
1
|
| 467 |
+
)
|
| 468 |
+
similar-to
|
| 469 |
+
subscript
|
| 470 |
+
𝑧
|
| 471 |
+
𝑖
|
| 472 |
+
superscript
|
| 473 |
+
subscript
|
| 474 |
+
𝑝
|
| 475 |
+
𝜃
|
| 476 |
+
𝐶
|
| 477 |
+
𝑜
|
| 478 |
+
𝑇
|
| 479 |
+
conditional
|
| 480 |
+
subscript
|
| 481 |
+
𝑧
|
| 482 |
+
𝑖
|
| 483 |
+
𝑥
|
| 484 |
+
subscript
|
| 485 |
+
𝑧
|
| 486 |
+
1
|
| 487 |
+
⋯
|
| 488 |
+
𝑖
|
| 489 |
+
1
|
| 490 |
+
z_{i}\sim p_{\theta}^{CoT}(z_{i}\mid x,z_{1\cdots i-1})
|
| 491 |
+
is sampled sequentially, then the output
|
| 492 |
+
y
|
| 493 |
+
∼
|
| 494 |
+
p
|
| 495 |
+
θ
|
| 496 |
+
C
|
| 497 |
+
|
| 498 |
+
o
|
| 499 |
+
|
| 500 |
+
T
|
| 501 |
+
|
| 502 |
+
(
|
| 503 |
+
y
|
| 504 |
+
|
|
| 505 |
+
x
|
| 506 |
+
,
|
| 507 |
+
z
|
| 508 |
+
1
|
| 509 |
+
|
| 510 |
+
⋯
|
| 511 |
+
|
| 512 |
+
n
|
| 513 |
+
)
|
| 514 |
+
similar-to
|
| 515 |
+
𝑦
|
| 516 |
+
superscript
|
| 517 |
+
subscript
|
| 518 |
+
𝑝
|
| 519 |
+
𝜃
|
| 520 |
+
𝐶
|
| 521 |
+
𝑜
|
| 522 |
+
𝑇
|
| 523 |
+
conditional
|
| 524 |
+
𝑦
|
| 525 |
+
𝑥
|
| 526 |
+
subscript
|
| 527 |
+
𝑧
|
| 528 |
+
1
|
| 529 |
+
⋯
|
| 530 |
+
𝑛
|
| 531 |
+
y\sim p_{\theta}^{CoT}(y|x,z_{1\cdots n})
|
| 532 |
+
. In practice,
|
| 533 |
+
[
|
| 534 |
+
z
|
| 535 |
+
1
|
| 536 |
+
|
| 537 |
+
⋯
|
| 538 |
+
|
| 539 |
+
n
|
| 540 |
+
,
|
| 541 |
+
y
|
| 542 |
+
]
|
| 543 |
+
��
|
| 544 |
+
p
|
| 545 |
+
θ
|
| 546 |
+
C
|
| 547 |
+
|
| 548 |
+
o
|
| 549 |
+
|
| 550 |
+
T
|
| 551 |
+
|
| 552 |
+
(
|
| 553 |
+
z
|
| 554 |
+
1
|
| 555 |
+
|
| 556 |
+
⋯
|
| 557 |
+
|
| 558 |
+
n
|
| 559 |
+
,
|
| 560 |
+
y
|
| 561 |
+
|
|
| 562 |
+
x
|
| 563 |
+
)
|
| 564 |
+
similar-to
|
| 565 |
+
subscript
|
| 566 |
+
𝑧
|
| 567 |
+
1
|
| 568 |
+
⋯
|
| 569 |
+
𝑛
|
| 570 |
+
𝑦
|
| 571 |
+
superscript
|
| 572 |
+
subscript
|
| 573 |
+
𝑝
|
| 574 |
+
𝜃
|
| 575 |
+
𝐶
|
| 576 |
+
𝑜
|
| 577 |
+
𝑇
|
| 578 |
+
subscript
|
| 579 |
+
𝑧
|
| 580 |
+
1
|
| 581 |
+
⋯
|
| 582 |
+
𝑛
|
| 583 |
+
conditional
|
| 584 |
+
𝑦
|
| 585 |
+
𝑥
|
| 586 |
+
[z_{1\cdots n},y]\sim p_{\theta}^{CoT}(z_{1\cdots n},y|x)
|
| 587 |
+
is sampled as a continuous language sequence, and the
|
| 588 |
+
decomposition
|
| 589 |
+
of thoughts (e.g. is each
|
| 590 |
+
z
|
| 591 |
+
i
|
| 592 |
+
subscript
|
| 593 |
+
𝑧
|
| 594 |
+
𝑖
|
| 595 |
+
z_{i}
|
| 596 |
+
a phrase, a sentence, or a paragraph) is left ambiguous.
|
| 597 |
+
Self-consistency with CoT (CoT-SC)
|
| 598 |
+
[
|
| 599 |
+
36
|
| 600 |
+
]
|
| 601 |
+
is an ensemble approach that samples
|
| 602 |
+
k
|
| 603 |
+
𝑘
|
| 604 |
+
k
|
| 605 |
+
i.i.d. chains of thought:
|
| 606 |
+
[
|
| 607 |
+
z
|
| 608 |
+
1
|
| 609 |
+
|
| 610 |
+
⋯
|
| 611 |
+
|
| 612 |
+
n
|
| 613 |
+
(
|
| 614 |
+
i
|
| 615 |
+
)
|
| 616 |
+
,
|
| 617 |
+
y
|
| 618 |
+
(
|
| 619 |
+
i
|
| 620 |
+
)
|
| 621 |
+
]
|
| 622 |
+
∼
|
| 623 |
+
p
|
| 624 |
+
θ
|
| 625 |
+
C
|
| 626 |
+
|
| 627 |
+
o
|
| 628 |
+
|
| 629 |
+
T
|
| 630 |
+
|
| 631 |
+
(
|
| 632 |
+
z
|
| 633 |
+
1
|
| 634 |
+
|
| 635 |
+
⋯
|
| 636 |
+
|
| 637 |
+
n
|
| 638 |
+
,
|
| 639 |
+
y
|
| 640 |
+
|
|
| 641 |
+
x
|
| 642 |
+
)
|
| 643 |
+
|
| 644 |
+
(
|
| 645 |
+
i
|
| 646 |
+
=
|
| 647 |
+
1
|
| 648 |
+
|
| 649 |
+
⋯
|
| 650 |
+
|
| 651 |
+
k
|
| 652 |
+
)
|
| 653 |
+
similar-to
|
| 654 |
+
subscript
|
| 655 |
+
superscript
|
| 656 |
+
𝑧
|
| 657 |
+
𝑖
|
| 658 |
+
1
|
| 659 |
+
⋯
|
| 660 |
+
𝑛
|
| 661 |
+
superscript
|
| 662 |
+
𝑦
|
| 663 |
+
𝑖
|
| 664 |
+
superscript
|
| 665 |
+
subscript
|
| 666 |
+
𝑝
|
| 667 |
+
𝜃
|
| 668 |
+
𝐶
|
| 669 |
+
𝑜
|
| 670 |
+
𝑇
|
| 671 |
+
subscript
|
| 672 |
+
𝑧
|
| 673 |
+
1
|
| 674 |
+
⋯
|
| 675 |
+
𝑛
|
| 676 |
+
conditional
|
| 677 |
+
𝑦
|
| 678 |
+
𝑥
|
| 679 |
+
𝑖
|
| 680 |
+
1
|
| 681 |
+
⋯
|
| 682 |
+
𝑘
|
| 683 |
+
[z^{(i)}_{1\cdots n},y^{(i)}]\sim p_{\theta}^{CoT}(z_{1\cdots n},y|x)\ (i=1\cdots k)
|
| 684 |
+
, then returns the most frequent output:
|
| 685 |
+
arg
|
| 686 |
+
|
| 687 |
+
max
|
| 688 |
+
y
|
| 689 |
+
|
| 690 |
+
#
|
| 691 |
+
|
| 692 |
+
{
|
| 693 |
+
i
|
| 694 |
+
∣
|
| 695 |
+
y
|
| 696 |
+
(
|
| 697 |
+
i
|
| 698 |
+
)
|
| 699 |
+
=
|
| 700 |
+
y
|
| 701 |
+
}
|
| 702 |
+
subscript
|
| 703 |
+
𝑦
|
| 704 |
+
#
|
| 705 |
+
conditional-set
|
| 706 |
+
𝑖
|
| 707 |
+
superscript
|
| 708 |
+
𝑦
|
| 709 |
+
𝑖
|
| 710 |
+
𝑦
|
| 711 |
+
\arg\max_{y}\#\{i\mid y^{(i)}=y\}
|
| 712 |
+
. CoT-SC improves upon CoT, because there are generally different thought processes for the same problem (e.g. different ways to prove the same theorem), and the output decision can be more faithful by exploring a richer set of thoughts. However, within each chain there is no local exploration of different thought steps, and the “most frequent” heuristic only applies when the output space is limited (e.g. multi-choice QA).
|
| 713 |
+
3
|
| 714 |
+
Tree of Thoughts: Deliberate Problem Solving with LM
|
| 715 |
+
A genuine problem-solving process involves the repeated use of available information to initiate exploration, which discloses, in turn, more information until a way to attain the solution is finally discovered.——
|
| 716 |
+
Newell et al. [
|
| 717 |
+
21
|
| 718 |
+
]
|
| 719 |
+
Research on human problem-solving suggests that people search through a combinatorial problem-space – a tree where the nodes represent partial solutions, and the branches correspond to operators that modify them
|
| 720 |
+
[
|
| 721 |
+
21
|
| 722 |
+
,
|
| 723 |
+
22
|
| 724 |
+
]
|
| 725 |
+
. Which branch to take is determined by heuristics that help to navigate the problem-space and guide the problem-solver towards a solution. This perspective highlights two key shortcomings of existing approaches that use LMs to solve general problems: 1) Locally, they do not explore
|
| 726 |
+
different
|
| 727 |
+
continuations within a thought process – the branches of the tree. 2) Globally, they do not incorporate any type of planning, lookahead, or backtracking to help evaluate these different options – the kind of heuristic-guided search that seems characteristic of human problem-solving.
|
| 728 |
+
To address these shortcomings, we introduce
|
| 729 |
+
Tree of Thoughts (ToT)
|
| 730 |
+
, a paradigm that allows LMs to explore multiple reasoning paths over thoughts (Figure
|
| 731 |
+
1
|
| 732 |
+
(c)). ToT frames any problem as a search over a tree, where each node is a
|
| 733 |
+
state
|
| 734 |
+
s
|
| 735 |
+
=
|
| 736 |
+
[
|
| 737 |
+
x
|
| 738 |
+
,
|
| 739 |
+
z
|
| 740 |
+
1
|
| 741 |
+
|
| 742 |
+
⋯
|
| 743 |
+
|
| 744 |
+
i
|
| 745 |
+
]
|
| 746 |
+
𝑠
|
| 747 |
+
𝑥
|
| 748 |
+
subscript
|
| 749 |
+
𝑧
|
| 750 |
+
1
|
| 751 |
+
⋯
|
| 752 |
+
𝑖
|
| 753 |
+
s=[x,z_{1\cdots i}]
|
| 754 |
+
representing a partial solution with the input and the sequence of thoughts so far. A specific instantiation of ToT involves answering four questions: 1. How to
|
| 755 |
+
decompose
|
| 756 |
+
the intermediate process into thought steps; 2. How to
|
| 757 |
+
generate
|
| 758 |
+
potential thoughts from each state; 3. How to heuristically
|
| 759 |
+
evaluate
|
| 760 |
+
states; 4. What
|
| 761 |
+
search
|
| 762 |
+
algorithm to use.
|
| 763 |
+
1. Thought decomposition.
|
| 764 |
+
While CoT samples thoughts coherently without explicit decomposition, ToT leverages problem properties to design and decompose intermediate thought steps. As Table
|
| 765 |
+
1
|
| 766 |
+
shows, depending on different problems, a thought could be a couple of words (Crosswords), a line of equation (Game of 24), or a whole paragraph of writing plan (Creative Writing). In general, a thought should be “small” enough so that LMs can generate promising and diverse samples (e.g. generating a whole book is usually too “big” to be coherent), yet “big” enough so that LMs can evaluate its prospect toward problem solving (e.g. generating one token is usually too “small” to evaluate).
|
| 767 |
+
2. Thought generator
|
| 768 |
+
G
|
| 769 |
+
|
| 770 |
+
(
|
| 771 |
+
p
|
| 772 |
+
θ
|
| 773 |
+
,
|
| 774 |
+
s
|
| 775 |
+
,
|
| 776 |
+
k
|
| 777 |
+
)
|
| 778 |
+
𝐺
|
| 779 |
+
subscript
|
| 780 |
+
𝑝
|
| 781 |
+
𝜃
|
| 782 |
+
𝑠
|
| 783 |
+
𝑘
|
| 784 |
+
G(p_{\theta},s,k)
|
| 785 |
+
.
|
| 786 |
+
Given a tree state
|
| 787 |
+
s
|
| 788 |
+
=
|
| 789 |
+
[
|
| 790 |
+
x
|
| 791 |
+
,
|
| 792 |
+
z
|
| 793 |
+
1
|
| 794 |
+
|
| 795 |
+
⋯
|
| 796 |
+
|
| 797 |
+
i
|
| 798 |
+
]
|
| 799 |
+
𝑠
|
| 800 |
+
𝑥
|
| 801 |
+
subscript
|
| 802 |
+
𝑧
|
| 803 |
+
1
|
| 804 |
+
⋯
|
| 805 |
+
𝑖
|
| 806 |
+
s=[x,z_{1\cdots i}]
|
| 807 |
+
, we consider two strategies to generate
|
| 808 |
+
k
|
| 809 |
+
𝑘
|
| 810 |
+
k
|
| 811 |
+
candidates for the next thought step:
|
| 812 |
+
(a)
|
| 813 |
+
Sample
|
| 814 |
+
i.i.d. thoughts from a CoT prompt (Creative Writing, Figure
|
| 815 |
+
4
|
| 816 |
+
):
|
| 817 |
+
z
|
| 818 |
+
(
|
| 819 |
+
j
|
| 820 |
+
)
|
| 821 |
+
∼
|
| 822 |
+
p
|
| 823 |
+
θ
|
| 824 |
+
C
|
| 825 |
+
|
| 826 |
+
o
|
| 827 |
+
|
| 828 |
+
T
|
| 829 |
+
|
| 830 |
+
(
|
| 831 |
+
z
|
| 832 |
+
i
|
| 833 |
+
+
|
| 834 |
+
1
|
| 835 |
+
|
|
| 836 |
+
s
|
| 837 |
+
)
|
| 838 |
+
=
|
| 839 |
+
p
|
| 840 |
+
θ
|
| 841 |
+
C
|
| 842 |
+
|
| 843 |
+
o
|
| 844 |
+
|
| 845 |
+
T
|
| 846 |
+
|
| 847 |
+
(
|
| 848 |
+
z
|
| 849 |
+
i
|
| 850 |
+
+
|
| 851 |
+
1
|
| 852 |
+
|
|
| 853 |
+
x
|
| 854 |
+
,
|
| 855 |
+
z
|
| 856 |
+
1
|
| 857 |
+
|
| 858 |
+
⋯
|
| 859 |
+
|
| 860 |
+
i
|
| 861 |
+
)
|
| 862 |
+
|
| 863 |
+
(
|
| 864 |
+
j
|
| 865 |
+
=
|
| 866 |
+
1
|
| 867 |
+
|
| 868 |
+
⋯
|
| 869 |
+
|
| 870 |
+
k
|
| 871 |
+
)
|
| 872 |
+
similar-to
|
| 873 |
+
superscript
|
| 874 |
+
𝑧
|
| 875 |
+
𝑗
|
| 876 |
+
superscript
|
| 877 |
+
subscript
|
| 878 |
+
𝑝
|
| 879 |
+
𝜃
|
| 880 |
+
𝐶
|
| 881 |
+
𝑜
|
| 882 |
+
𝑇
|
| 883 |
+
conditional
|
| 884 |
+
subscript
|
| 885 |
+
𝑧
|
| 886 |
+
𝑖
|
| 887 |
+
1
|
| 888 |
+
𝑠
|
| 889 |
+
superscript
|
| 890 |
+
subscript
|
| 891 |
+
𝑝
|
| 892 |
+
𝜃
|
| 893 |
+
𝐶
|
| 894 |
+
𝑜
|
| 895 |
+
𝑇
|
| 896 |
+
conditional
|
| 897 |
+
subscript
|
| 898 |
+
𝑧
|
| 899 |
+
𝑖
|
| 900 |
+
1
|
| 901 |
+
𝑥
|
| 902 |
+
subscript
|
| 903 |
+
𝑧
|
| 904 |
+
1
|
| 905 |
+
⋯
|
| 906 |
+
𝑖
|
| 907 |
+
𝑗
|
| 908 |
+
1
|
| 909 |
+
⋯
|
| 910 |
+
𝑘
|
| 911 |
+
z^{(j)}\sim p_{\theta}^{CoT}(z_{i+1}|s)=p_{\theta}^{CoT}(z_{i+1}|x,z_{1\cdots i})\ (j=1\cdots k)
|
| 912 |
+
. This works better when the thought space is rich (e.g. each thought is a paragraph), and i.i.d. samples lead to diversity;
|
| 913 |
+
(b)
|
| 914 |
+
Propose
|
| 915 |
+
thoughts sequentially using a “propose prompt” (Game of 24, Figure
|
| 916 |
+
2
|
| 917 |
+
; Crosswords, Figure
|
| 918 |
+
6
|
| 919 |
+
):
|
| 920 |
+
[
|
| 921 |
+
z
|
| 922 |
+
(
|
| 923 |
+
1
|
| 924 |
+
)
|
| 925 |
+
,
|
| 926 |
+
⋯
|
| 927 |
+
,
|
| 928 |
+
z
|
| 929 |
+
(
|
| 930 |
+
k
|
| 931 |
+
)
|
| 932 |
+
]
|
| 933 |
+
∼
|
| 934 |
+
p
|
| 935 |
+
θ
|
| 936 |
+
p
|
| 937 |
+
|
| 938 |
+
r
|
| 939 |
+
|
| 940 |
+
o
|
| 941 |
+
|
| 942 |
+
p
|
| 943 |
+
|
| 944 |
+
o
|
| 945 |
+
|
| 946 |
+
s
|
| 947 |
+
|
| 948 |
+
e
|
| 949 |
+
|
| 950 |
+
(
|
| 951 |
+
z
|
| 952 |
+
i
|
| 953 |
+
+
|
| 954 |
+
1
|
| 955 |
+
(
|
| 956 |
+
1
|
| 957 |
+
|
| 958 |
+
⋯
|
| 959 |
+
|
| 960 |
+
k
|
| 961 |
+
)
|
| 962 |
+
∣
|
| 963 |
+
s
|
| 964 |
+
)
|
| 965 |
+
similar-to
|
| 966 |
+
superscript
|
| 967 |
+
𝑧
|
| 968 |
+
1
|
| 969 |
+
⋯
|
| 970 |
+
superscript
|
| 971 |
+
𝑧
|
| 972 |
+
𝑘
|
| 973 |
+
superscript
|
| 974 |
+
subscript
|
| 975 |
+
𝑝
|
| 976 |
+
𝜃
|
| 977 |
+
𝑝
|
| 978 |
+
𝑟
|
| 979 |
+
𝑜
|
| 980 |
+
𝑝
|
| 981 |
+
𝑜
|
| 982 |
+
𝑠
|
| 983 |
+
𝑒
|
| 984 |
+
conditional
|
| 985 |
+
superscript
|
| 986 |
+
subscript
|
| 987 |
+
𝑧
|
| 988 |
+
𝑖
|
| 989 |
+
1
|
| 990 |
+
1
|
| 991 |
+
⋯
|
| 992 |
+
𝑘
|
| 993 |
+
𝑠
|
| 994 |
+
[z^{(1)},\cdots,z^{(k)}]\sim p_{\theta}^{propose}(z_{i+1}^{(1\cdots k)}\mid s)
|
| 995 |
+
. This works better when the thought space is more constrained (e.g. each thought is just a word or a line), so proposing different thoughts in the same context avoids duplication.
|
| 996 |
+
3. State evaluator
|
| 997 |
+
V
|
| 998 |
+
|
| 999 |
+
(
|
| 1000 |
+
p
|
| 1001 |
+
θ
|
| 1002 |
+
,
|
| 1003 |
+
S
|
| 1004 |
+
)
|
| 1005 |
+
𝑉
|
| 1006 |
+
subscript
|
| 1007 |
+
𝑝
|
| 1008 |
+
𝜃
|
| 1009 |
+
𝑆
|
| 1010 |
+
V(p_{\theta},S)
|
| 1011 |
+
.
|
| 1012 |
+
Given a frontier of different states, the state evaluator evaluates the progress they make towards solving the problem, serving as a
|
| 1013 |
+
heuristic
|
| 1014 |
+
for the search algorithm to determine which states to keep exploring and in which order. While heuristics are a standard approach to solving search problems, they are typically either programmed (e.g. DeepBlue
|
| 1015 |
+
[
|
| 1016 |
+
3
|
| 1017 |
+
]
|
| 1018 |
+
) or learned (e.g. AlphaGo
|
| 1019 |
+
[
|
| 1020 |
+
29
|
| 1021 |
+
]
|
| 1022 |
+
). We propose a third alternative, by using the LM to deliberately reason about states. When applicable, such a deliberate heuristic can be more flexible than programmed rules, and more sample-efficient than learned models.
|
| 1023 |
+
Similar to the thought generator, we consider two strategies to evaluate states either independently or together:
|
| 1024 |
+
(a)
|
| 1025 |
+
Value
|
| 1026 |
+
each state independently:
|
| 1027 |
+
V
|
| 1028 |
+
|
| 1029 |
+
(
|
| 1030 |
+
p
|
| 1031 |
+
θ
|
| 1032 |
+
,
|
| 1033 |
+
S
|
| 1034 |
+
)
|
| 1035 |
+
|
| 1036 |
+
(
|
| 1037 |
+
s
|
| 1038 |
+
)
|
| 1039 |
+
∼
|
| 1040 |
+
p
|
| 1041 |
+
θ
|
| 1042 |
+
v
|
| 1043 |
+
|
| 1044 |
+
a
|
| 1045 |
+
|
| 1046 |
+
l
|
| 1047 |
+
|
| 1048 |
+
u
|
| 1049 |
+
|
| 1050 |
+
e
|
| 1051 |
+
|
| 1052 |
+
(
|
| 1053 |
+
v
|
| 1054 |
+
|
|
| 1055 |
+
s
|
| 1056 |
+
)
|
| 1057 |
+
|
| 1058 |
+
∀
|
| 1059 |
+
s
|
| 1060 |
+
∈
|
| 1061 |
+
S
|
| 1062 |
+
similar-to
|
| 1063 |
+
𝑉
|
| 1064 |
+
subscript
|
| 1065 |
+
𝑝
|
| 1066 |
+
𝜃
|
| 1067 |
+
𝑆
|
| 1068 |
+
𝑠
|
| 1069 |
+
superscript
|
| 1070 |
+
subscript
|
| 1071 |
+
𝑝
|
| 1072 |
+
𝜃
|
| 1073 |
+
𝑣
|
| 1074 |
+
𝑎
|
| 1075 |
+
𝑙
|
| 1076 |
+
𝑢
|
| 1077 |
+
𝑒
|
| 1078 |
+
conditional
|
| 1079 |
+
𝑣
|
| 1080 |
+
𝑠
|
| 1081 |
+
for-all
|
| 1082 |
+
𝑠
|
| 1083 |
+
𝑆
|
| 1084 |
+
V(p_{\theta},S)(s)\sim p_{\theta}^{value}(v|s)\ \forall s\in S
|
| 1085 |
+
, where a value prompt reasons about the state
|
| 1086 |
+
s
|
| 1087 |
+
𝑠
|
| 1088 |
+
s
|
| 1089 |
+
to generate a scalar value
|
| 1090 |
+
v
|
| 1091 |
+
𝑣
|
| 1092 |
+
v
|
| 1093 |
+
(e.g. 1-10) or a classification (e.g. sure/likely/impossible) that could be heuristically turned into a value. The basis of such evaluative reasoning can vary across problems and thought steps. In this work, we explore evaluation via few
|
| 1094 |
+
lookahead
|
| 1095 |
+
simulations (e.g. quickly confirm that 5, 5, 14 can reach 24 via 5 + 5 + 14, or “hot_l” can mean “inn” via filling “e” in “_”) plus commonsense (e.g. 1 2 3 are too small to reach 24, or no word can start with “tzxc”). While the former might promote “good” states, the latter could help eliminate “bad” states. Such valuations do not need to be perfect, and only need to be approximately helpful for decision making.
|
| 1096 |
+
(b)
|
| 1097 |
+
Vote
|
| 1098 |
+
across states:
|
| 1099 |
+
V
|
| 1100 |
+
|
| 1101 |
+
(
|
| 1102 |
+
p
|
| 1103 |
+
θ
|
| 1104 |
+
,
|
| 1105 |
+
S
|
| 1106 |
+
)
|
| 1107 |
+
|
| 1108 |
+
(
|
| 1109 |
+
s
|
| 1110 |
+
)
|
| 1111 |
+
=
|
| 1112 |
+
𝟙
|
| 1113 |
+
|
| 1114 |
+
[
|
| 1115 |
+
s
|
| 1116 |
+
=
|
| 1117 |
+
s
|
| 1118 |
+
∗
|
| 1119 |
+
]
|
| 1120 |
+
𝑉
|
| 1121 |
+
subscript
|
| 1122 |
+
𝑝
|
| 1123 |
+
𝜃
|
| 1124 |
+
𝑆
|
| 1125 |
+
𝑠
|
| 1126 |
+
1
|
| 1127 |
+
delimited-[]
|
| 1128 |
+
𝑠
|
| 1129 |
+
superscript
|
| 1130 |
+
𝑠
|
| 1131 |
+
V(p_{\theta},S)(s)=\mathds{1}[s=s^{*}]
|
| 1132 |
+
, where a “good” state
|
| 1133 |
+
s
|
| 1134 |
+
∗
|
| 1135 |
+
∼
|
| 1136 |
+
p
|
| 1137 |
+
θ
|
| 1138 |
+
v
|
| 1139 |
+
|
| 1140 |
+
o
|
| 1141 |
+
|
| 1142 |
+
t
|
| 1143 |
+
|
| 1144 |
+
e
|
| 1145 |
+
|
| 1146 |
+
(
|
| 1147 |
+
s
|
| 1148 |
+
∗
|
| 1149 |
+
|
|
| 1150 |
+
S
|
| 1151 |
+
)
|
| 1152 |
+
similar-to
|
| 1153 |
+
superscript
|
| 1154 |
+
𝑠
|
| 1155 |
+
superscript
|
| 1156 |
+
subscript
|
| 1157 |
+
𝑝
|
| 1158 |
+
𝜃
|
| 1159 |
+
𝑣
|
| 1160 |
+
𝑜
|
| 1161 |
+
𝑡
|
| 1162 |
+
𝑒
|
| 1163 |
+
conditional
|
| 1164 |
+
superscript
|
| 1165 |
+
𝑠
|
| 1166 |
+
𝑆
|
| 1167 |
+
s^{*}\sim p_{\theta}^{vote}(s^{*}|S)
|
| 1168 |
+
is voted out based on deliberately comparing different states in
|
| 1169 |
+
S
|
| 1170 |
+
𝑆
|
| 1171 |
+
S
|
| 1172 |
+
in a vote prompt.
|
| 1173 |
+
When problem success is harder to directly value (e.g. passage coherency), it is natural to to instead compare different partial solutions and vote for the most promising one. This is similar in spirit to a “step-wise” self-consistency strategy, i.e. cast “which state to explore” as a multi-choice QA, and use LM samples to vote for it.
|
| 1174 |
+
For both strategies, we could prompt the LM multiple times to aggregate the value or vote results to trade time/resource/cost for more faithful/robust heuristics.
|
| 1175 |
+
Algorithm 1
|
| 1176 |
+
ToT-BFS(
|
| 1177 |
+
x
|
| 1178 |
+
,
|
| 1179 |
+
p
|
| 1180 |
+
θ
|
| 1181 |
+
,
|
| 1182 |
+
G
|
| 1183 |
+
,
|
| 1184 |
+
k
|
| 1185 |
+
,
|
| 1186 |
+
V
|
| 1187 |
+
,
|
| 1188 |
+
T
|
| 1189 |
+
,
|
| 1190 |
+
b
|
| 1191 |
+
𝑥
|
| 1192 |
+
subscript
|
| 1193 |
+
𝑝
|
| 1194 |
+
𝜃
|
| 1195 |
+
𝐺
|
| 1196 |
+
𝑘
|
| 1197 |
+
𝑉
|
| 1198 |
+
𝑇
|
| 1199 |
+
𝑏
|
| 1200 |
+
x,p_{\theta},G,k,V,T,b
|
| 1201 |
+
)
|
| 1202 |
+
Input
|
| 1203 |
+
x
|
| 1204 |
+
𝑥
|
| 1205 |
+
x
|
| 1206 |
+
, LM
|
| 1207 |
+
p
|
| 1208 |
+
θ
|
| 1209 |
+
subscript
|
| 1210 |
+
𝑝
|
| 1211 |
+
𝜃
|
| 1212 |
+
p_{\theta}
|
| 1213 |
+
, thought generator
|
| 1214 |
+
G
|
| 1215 |
+
|
| 1216 |
+
(
|
| 1217 |
+
)
|
| 1218 |
+
𝐺
|
| 1219 |
+
G()
|
| 1220 |
+
& size limit
|
| 1221 |
+
k
|
| 1222 |
+
𝑘
|
| 1223 |
+
k
|
| 1224 |
+
, states evaluator
|
| 1225 |
+
V
|
| 1226 |
+
|
| 1227 |
+
(
|
| 1228 |
+
)
|
| 1229 |
+
𝑉
|
| 1230 |
+
V()
|
| 1231 |
+
, step limit
|
| 1232 |
+
T
|
| 1233 |
+
𝑇
|
| 1234 |
+
T
|
| 1235 |
+
, breadth limit
|
| 1236 |
+
b
|
| 1237 |
+
𝑏
|
| 1238 |
+
b
|
| 1239 |
+
.
|
| 1240 |
+
S
|
| 1241 |
+
0
|
| 1242 |
+
←
|
| 1243 |
+
{
|
| 1244 |
+
x
|
| 1245 |
+
}
|
| 1246 |
+
←
|
| 1247 |
+
subscript
|
| 1248 |
+
𝑆
|
| 1249 |
+
0
|
| 1250 |
+
𝑥
|
| 1251 |
+
S_{0}\leftarrow\{x\}
|
| 1252 |
+
for
|
| 1253 |
+
t
|
| 1254 |
+
=
|
| 1255 |
+
1
|
| 1256 |
+
,
|
| 1257 |
+
⋯
|
| 1258 |
+
,
|
| 1259 |
+
T
|
| 1260 |
+
𝑡
|
| 1261 |
+
1
|
| 1262 |
+
⋯
|
| 1263 |
+
𝑇
|
| 1264 |
+
t=1,\cdots,T
|
| 1265 |
+
do
|
| 1266 |
+
S
|
| 1267 |
+
t
|
| 1268 |
+
′
|
| 1269 |
+
←
|
| 1270 |
+
{
|
| 1271 |
+
[
|
| 1272 |
+
s
|
| 1273 |
+
,
|
| 1274 |
+
z
|
| 1275 |
+
]
|
| 1276 |
+
∣
|
| 1277 |
+
s
|
| 1278 |
+
∈
|
| 1279 |
+
S
|
| 1280 |
+
t
|
| 1281 |
+
−
|
| 1282 |
+
1
|
| 1283 |
+
,
|
| 1284 |
+
z
|
| 1285 |
+
t
|
| 1286 |
+
∈
|
| 1287 |
+
G
|
| 1288 |
+
|
| 1289 |
+
(
|
| 1290 |
+
p
|
| 1291 |
+
θ
|
| 1292 |
+
,
|
| 1293 |
+
s
|
| 1294 |
+
,
|
| 1295 |
+
k
|
| 1296 |
+
)
|
| 1297 |
+
}
|
| 1298 |
+
←
|
| 1299 |
+
subscript
|
| 1300 |
+
superscript
|
| 1301 |
+
𝑆
|
| 1302 |
+
′
|
| 1303 |
+
𝑡
|
| 1304 |
+
conditional-set
|
| 1305 |
+
𝑠
|
| 1306 |
+
𝑧
|
| 1307 |
+
formulae-sequence
|
| 1308 |
+
𝑠
|
| 1309 |
+
subscript
|
| 1310 |
+
𝑆
|
| 1311 |
+
𝑡
|
| 1312 |
+
1
|
| 1313 |
+
subscript
|
| 1314 |
+
𝑧
|
| 1315 |
+
𝑡
|
| 1316 |
+
G
|
| 1317 |
+
subscript
|
| 1318 |
+
𝑝
|
| 1319 |
+
𝜃
|
| 1320 |
+
𝑠
|
| 1321 |
+
𝑘
|
| 1322 |
+
S^{\prime}_{t}\leftarrow\{[s,z]\mid s\in S_{t-1},z_{t}\in{\color[rgb]{0,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,0}\pgfsys@color@gray@stroke{0}\pgfsys@color@gray@fill{0}\mathrm{G}}(p_{\theta},s,k)\}
|
| 1323 |
+
V
|
| 1324 |
+
t
|
| 1325 |
+
←
|
| 1326 |
+
V
|
| 1327 |
+
|
| 1328 |
+
(
|
| 1329 |
+
p
|
| 1330 |
+
θ
|
| 1331 |
+
,
|
| 1332 |
+
S
|
| 1333 |
+
t
|
| 1334 |
+
′
|
| 1335 |
+
)
|
| 1336 |
+
←
|
| 1337 |
+
subscript
|
| 1338 |
+
𝑉
|
| 1339 |
+
𝑡
|
| 1340 |
+
𝑉
|
| 1341 |
+
subscript
|
| 1342 |
+
𝑝
|
| 1343 |
+
𝜃
|
| 1344 |
+
subscript
|
| 1345 |
+
superscript
|
| 1346 |
+
𝑆
|
| 1347 |
+
′
|
| 1348 |
+
𝑡
|
| 1349 |
+
V_{t}\leftarrow V(p_{\theta},S^{\prime}_{t})
|
| 1350 |
+
S
|
| 1351 |
+
t
|
| 1352 |
+
←
|
| 1353 |
+
arg
|
| 1354 |
+
|
| 1355 |
+
max
|
| 1356 |
+
S
|
| 1357 |
+
⊂
|
| 1358 |
+
S
|
| 1359 |
+
t
|
| 1360 |
+
′
|
| 1361 |
+
,
|
| 1362 |
+
|
|
| 1363 |
+
S
|
| 1364 |
+
|
|
| 1365 |
+
=
|
| 1366 |
+
b
|
| 1367 |
+
|
| 1368 |
+
∑
|
| 1369 |
+
s
|
| 1370 |
+
∈
|
| 1371 |
+
S
|
| 1372 |
+
V
|
| 1373 |
+
t
|
| 1374 |
+
|
| 1375 |
+
(
|
| 1376 |
+
s
|
| 1377 |
+
)
|
| 1378 |
+
←
|
| 1379 |
+
subscript
|
| 1380 |
+
𝑆
|
| 1381 |
+
𝑡
|
| 1382 |
+
subscript
|
| 1383 |
+
formulae-sequence
|
| 1384 |
+
𝑆
|
| 1385 |
+
subscript
|
| 1386 |
+
superscript
|
| 1387 |
+
𝑆
|
| 1388 |
+
′
|
| 1389 |
+
𝑡
|
| 1390 |
+
𝑆
|
| 1391 |
+
𝑏
|
| 1392 |
+
subscript
|
| 1393 |
+
𝑠
|
| 1394 |
+
𝑆
|
| 1395 |
+
subscript
|
| 1396 |
+
𝑉
|
| 1397 |
+
𝑡
|
| 1398 |
+
𝑠
|
| 1399 |
+
S_{t}\leftarrow\arg\max_{S\subset S^{\prime}_{t},|S|=b}\sum_{s\in S}V_{t}(s)
|
| 1400 |
+
end
|
| 1401 |
+
for
|
| 1402 |
+
return
|
| 1403 |
+
G
|
| 1404 |
+
|
| 1405 |
+
(
|
| 1406 |
+
p
|
| 1407 |
+
θ
|
| 1408 |
+
,
|
| 1409 |
+
arg
|
| 1410 |
+
|
| 1411 |
+
max
|
| 1412 |
+
s
|
| 1413 |
+
∈
|
| 1414 |
+
S
|
| 1415 |
+
T
|
| 1416 |
+
|
| 1417 |
+
V
|
| 1418 |
+
T
|
| 1419 |
+
|
| 1420 |
+
(
|
| 1421 |
+
s
|
| 1422 |
+
)
|
| 1423 |
+
,
|
| 1424 |
+
1
|
| 1425 |
+
)
|
| 1426 |
+
𝐺
|
| 1427 |
+
subscript
|
| 1428 |
+
𝑝
|
| 1429 |
+
𝜃
|
| 1430 |
+
subscript
|
| 1431 |
+
𝑠
|
| 1432 |
+
subscript
|
| 1433 |
+
𝑆
|
| 1434 |
+
𝑇
|
| 1435 |
+
subscript
|
| 1436 |
+
𝑉
|
| 1437 |
+
𝑇
|
| 1438 |
+
𝑠
|
| 1439 |
+
1
|
| 1440 |
+
G(p_{\theta},\arg\max_{s\in S_{T}}V_{T}(s),1)
|
| 1441 |
+
Algorithm 2
|
| 1442 |
+
ToT-DFS(
|
| 1443 |
+
s
|
| 1444 |
+
,
|
| 1445 |
+
t
|
| 1446 |
+
,
|
| 1447 |
+
p
|
| 1448 |
+
θ
|
| 1449 |
+
,
|
| 1450 |
+
G
|
| 1451 |
+
,
|
| 1452 |
+
k
|
| 1453 |
+
,
|
| 1454 |
+
V
|
| 1455 |
+
,
|
| 1456 |
+
T
|
| 1457 |
+
,
|
| 1458 |
+
v
|
| 1459 |
+
t
|
| 1460 |
+
|
| 1461 |
+
h
|
| 1462 |
+
𝑠
|
| 1463 |
+
𝑡
|
| 1464 |
+
subscript
|
| 1465 |
+
𝑝
|
| 1466 |
+
𝜃
|
| 1467 |
+
𝐺
|
| 1468 |
+
𝑘
|
| 1469 |
+
𝑉
|
| 1470 |
+
𝑇
|
| 1471 |
+
subscript
|
| 1472 |
+
𝑣
|
| 1473 |
+
𝑡
|
| 1474 |
+
ℎ
|
| 1475 |
+
s,t,p_{\theta},G,k,V,T,v_{\small th}
|
| 1476 |
+
)
|
| 1477 |
+
Current state
|
| 1478 |
+
s
|
| 1479 |
+
𝑠
|
| 1480 |
+
s
|
| 1481 |
+
, step
|
| 1482 |
+
t
|
| 1483 |
+
𝑡
|
| 1484 |
+
t
|
| 1485 |
+
, LM
|
| 1486 |
+
p
|
| 1487 |
+
θ
|
| 1488 |
+
subscript
|
| 1489 |
+
𝑝
|
| 1490 |
+
𝜃
|
| 1491 |
+
p_{\theta}
|
| 1492 |
+
, thought generator
|
| 1493 |
+
G
|
| 1494 |
+
|
| 1495 |
+
(
|
| 1496 |
+
)
|
| 1497 |
+
𝐺
|
| 1498 |
+
G()
|
| 1499 |
+
and size limit
|
| 1500 |
+
k
|
| 1501 |
+
𝑘
|
| 1502 |
+
k
|
| 1503 |
+
, states evaluator
|
| 1504 |
+
V
|
| 1505 |
+
|
| 1506 |
+
(
|
| 1507 |
+
)
|
| 1508 |
+
𝑉
|
| 1509 |
+
V()
|
| 1510 |
+
, step limit
|
| 1511 |
+
T
|
| 1512 |
+
𝑇
|
| 1513 |
+
T
|
| 1514 |
+
, threshold
|
| 1515 |
+
v
|
| 1516 |
+
t
|
| 1517 |
+
|
| 1518 |
+
h
|
| 1519 |
+
subscript
|
| 1520 |
+
𝑣
|
| 1521 |
+
𝑡
|
| 1522 |
+
ℎ
|
| 1523 |
+
v_{\small th}
|
| 1524 |
+
if
|
| 1525 |
+
t
|
| 1526 |
+
>
|
| 1527 |
+
T
|
| 1528 |
+
𝑡
|
| 1529 |
+
𝑇
|
| 1530 |
+
t>T
|
| 1531 |
+
then
|
| 1532 |
+
record output
|
| 1533 |
+
G
|
| 1534 |
+
|
| 1535 |
+
(
|
| 1536 |
+
p
|
| 1537 |
+
θ
|
| 1538 |
+
,
|
| 1539 |
+
s
|
| 1540 |
+
,
|
| 1541 |
+
1
|
| 1542 |
+
)
|
| 1543 |
+
𝐺
|
| 1544 |
+
subscript
|
| 1545 |
+
𝑝
|
| 1546 |
+
𝜃
|
| 1547 |
+
𝑠
|
| 1548 |
+
1
|
| 1549 |
+
G(p_{\theta},s,1)
|
| 1550 |
+
end
|
| 1551 |
+
if
|
| 1552 |
+
for
|
| 1553 |
+
s
|
| 1554 |
+
′
|
| 1555 |
+
∈
|
| 1556 |
+
G
|
| 1557 |
+
|
| 1558 |
+
(
|
| 1559 |
+
p
|
| 1560 |
+
θ
|
| 1561 |
+
,
|
| 1562 |
+
s
|
| 1563 |
+
,
|
| 1564 |
+
k
|
| 1565 |
+
)
|
| 1566 |
+
superscript
|
| 1567 |
+
𝑠
|
| 1568 |
+
′
|
| 1569 |
+
𝐺
|
| 1570 |
+
subscript
|
| 1571 |
+
𝑝
|
| 1572 |
+
𝜃
|
| 1573 |
+
𝑠
|
| 1574 |
+
𝑘
|
| 1575 |
+
s^{\prime}\in G(p_{\theta},s,k)
|
| 1576 |
+
do
|
| 1577 |
+
▷
|
| 1578 |
+
▷
|
| 1579 |
+
\triangleright
|
| 1580 |
+
sorted candidates
|
| 1581 |
+
if
|
| 1582 |
+
V
|
| 1583 |
+
|
| 1584 |
+
(
|
| 1585 |
+
p
|
| 1586 |
+
θ
|
| 1587 |
+
,
|
| 1588 |
+
{
|
| 1589 |
+
s
|
| 1590 |
+
′
|
| 1591 |
+
}
|
| 1592 |
+
)
|
| 1593 |
+
|
| 1594 |
+
(
|
| 1595 |
+
s
|
| 1596 |
+
)
|
| 1597 |
+
>
|
| 1598 |
+
v
|
| 1599 |
+
t
|
| 1600 |
+
|
| 1601 |
+
h
|
| 1602 |
+
|
| 1603 |
+
r
|
| 1604 |
+
|
| 1605 |
+
e
|
| 1606 |
+
|
| 1607 |
+
s
|
| 1608 |
+
𝑉
|
| 1609 |
+
subscript
|
| 1610 |
+
𝑝
|
| 1611 |
+
𝜃
|
| 1612 |
+
superscript
|
| 1613 |
+
𝑠
|
| 1614 |
+
′
|
| 1615 |
+
𝑠
|
| 1616 |
+
subscript
|
| 1617 |
+
𝑣
|
| 1618 |
+
𝑡
|
| 1619 |
+
ℎ
|
| 1620 |
+
𝑟
|
| 1621 |
+
𝑒
|
| 1622 |
+
𝑠
|
| 1623 |
+
V(p_{\theta},\{s^{\prime}\})(s)>v_{\small thres}
|
| 1624 |
+
then
|
| 1625 |
+
▷
|
| 1626 |
+
▷
|
| 1627 |
+
\triangleright
|
| 1628 |
+
pruning
|
| 1629 |
+
DFS
|
| 1630 |
+
(
|
| 1631 |
+
s
|
| 1632 |
+
′
|
| 1633 |
+
,
|
| 1634 |
+
t
|
| 1635 |
+
+
|
| 1636 |
+
1
|
| 1637 |
+
)
|
| 1638 |
+
superscript
|
| 1639 |
+
𝑠
|
| 1640 |
+
′
|
| 1641 |
+
𝑡
|
| 1642 |
+
1
|
| 1643 |
+
(s^{\prime},t+1)
|
| 1644 |
+
end
|
| 1645 |
+
if
|
| 1646 |
+
end
|
| 1647 |
+
for
|
| 1648 |
+
4. Search algorithm.
|
| 1649 |
+
Finally, within the ToT framework, one can plug and play different search algorithms depending on the tree structure. We explore two relatively simple search algorithms and leave more advanced ones (e.g. A*
|
| 1650 |
+
[
|
| 1651 |
+
11
|
| 1652 |
+
]
|
| 1653 |
+
, MCTS
|
| 1654 |
+
[
|
| 1655 |
+
2
|
| 1656 |
+
]
|
| 1657 |
+
) for future work:
|
| 1658 |
+
(a)
|
| 1659 |
+
Breadth-first search (BFS)
|
| 1660 |
+
(Algorithm
|
| 1661 |
+
1
|
| 1662 |
+
) maintains a set of the
|
| 1663 |
+
b
|
| 1664 |
+
𝑏
|
| 1665 |
+
b
|
| 1666 |
+
most promising states per step. This is used for Game of 24 and Creative Writing where the tree depth is limit (
|
| 1667 |
+
T
|
| 1668 |
+
≤
|
| 1669 |
+
3
|
| 1670 |
+
𝑇
|
| 1671 |
+
3
|
| 1672 |
+
T\leq 3
|
| 1673 |
+
), and initial thought steps can be evaluated and pruned to a small set (
|
| 1674 |
+
b
|
| 1675 |
+
≤
|
| 1676 |
+
5
|
| 1677 |
+
𝑏
|
| 1678 |
+
5
|
| 1679 |
+
b\leq 5
|
| 1680 |
+
).
|
| 1681 |
+
(b)
|
| 1682 |
+
Depth-first search (DFS)
|
| 1683 |
+
(Algorithm
|
| 1684 |
+
2
|
| 1685 |
+
) explores the most promising state first, until the final output is reached (
|
| 1686 |
+
t
|
| 1687 |
+
>
|
| 1688 |
+
T
|
| 1689 |
+
𝑡
|
| 1690 |
+
𝑇
|
| 1691 |
+
t>T
|
| 1692 |
+
), or the state evaluator deems it impossible to solve the problem from the current
|
| 1693 |
+
s
|
| 1694 |
+
𝑠
|
| 1695 |
+
s
|
| 1696 |
+
(
|
| 1697 |
+
V
|
| 1698 |
+
|
| 1699 |
+
(
|
| 1700 |
+
p
|
| 1701 |
+
θ
|
| 1702 |
+
,
|
| 1703 |
+
{
|
| 1704 |
+
s
|
| 1705 |
+
}
|
| 1706 |
+
)
|
| 1707 |
+
|
| 1708 |
+
(
|
| 1709 |
+
s
|
| 1710 |
+
)
|
| 1711 |
+
≤
|
| 1712 |
+
v
|
| 1713 |
+
t
|
| 1714 |
+
|
| 1715 |
+
h
|
| 1716 |
+
𝑉
|
| 1717 |
+
subscript
|
| 1718 |
+
𝑝
|
| 1719 |
+
𝜃
|
| 1720 |
+
𝑠
|
| 1721 |
+
𝑠
|
| 1722 |
+
subscript
|
| 1723 |
+
𝑣
|
| 1724 |
+
𝑡
|
| 1725 |
+
ℎ
|
| 1726 |
+
V(p_{\theta},\{s\})(s)\leq v_{th}
|
| 1727 |
+
for a value threshold
|
| 1728 |
+
v
|
| 1729 |
+
t
|
| 1730 |
+
|
| 1731 |
+
h
|
| 1732 |
+
subscript
|
| 1733 |
+
𝑣
|
| 1734 |
+
𝑡
|
| 1735 |
+
ℎ
|
| 1736 |
+
v_{th}
|
| 1737 |
+
). In the latter case, the subtree from
|
| 1738 |
+
s
|
| 1739 |
+
𝑠
|
| 1740 |
+
s
|
| 1741 |
+
is
|
| 1742 |
+
pruned
|
| 1743 |
+
to trade exploration for exploitation. In both cases, DFS
|
| 1744 |
+
backtracks
|
| 1745 |
+
to the parent state of
|
| 1746 |
+
s
|
| 1747 |
+
𝑠
|
| 1748 |
+
s
|
| 1749 |
+
to continue exploration.
|
| 1750 |
+
Conceptually, ToT has several benefits as a method for general problem-solving with LMs: (1)
|
| 1751 |
+
Generality.
|
| 1752 |
+
IO, CoT, CoT-SC, and self-refinement can be seen as special cases of ToT (i.e. trees of limited depth and breadth; Figure
|
| 1753 |
+
1
|
| 1754 |
+
). (2)
|
| 1755 |
+
Modularity.
|
| 1756 |
+
The base LM, as well as the thought decomposition, generation, evaluation, and search procedures can all be varied independently. (3)
|
| 1757 |
+
Adaptability
|
| 1758 |
+
. Different problem properties, LM capabilities, and resource constraints can be accommodated. (4)
|
| 1759 |
+
Convenience.
|
| 1760 |
+
No extra training is needed, just a pre-trained LM is sufficient. The next section will show how these conceptual benefits translate to strong empirical performance in different problems.
|
| 1761 |
+
4
|
| 1762 |
+
Experiments
|
| 1763 |
+
Game of 24
|
| 1764 |
+
Creative Writing
|
| 1765 |
+
5x5 Crosswords
|
| 1766 |
+
Input
|
| 1767 |
+
4 numbers
|
| 1768 |
+
(4 9 10 13)
|
| 1769 |
+
4 random sentences
|
| 1770 |
+
10 clues
|
| 1771 |
+
(h1. presented;..)
|
| 1772 |
+
Output
|
| 1773 |
+
An equation to reach 24
|
| 1774 |
+
(13-9)*(10-4)=24
|
| 1775 |
+
A passage of 4 paragraphs ending in the 4 sentences
|
| 1776 |
+
5x5 letters:
|
| 1777 |
+
SHOWN; WIRRA; AVAIL; …
|
| 1778 |
+
Thoughts
|
| 1779 |
+
3 intermediate equations
|
| 1780 |
+
(13-9=4 (left 4,4,10); 10-4=6 (left 4,6); 4*6=24)
|
| 1781 |
+
A short writing plan
|
| 1782 |
+
(1. Introduce a book that connects…)
|
| 1783 |
+
Words to fill in for clues:
|
| 1784 |
+
(h1. shown; v5. naled; …)
|
| 1785 |
+
#ToT steps
|
| 1786 |
+
3
|
| 1787 |
+
1
|
| 1788 |
+
5-10 (variable)
|
| 1789 |
+
Table 1:
|
| 1790 |
+
Task overview. Input, output, thought examples are in blue.
|
| 1791 |
+
We propose three tasks that are hard even when sampling from the state-of-the-art language model, GPT-4
|
| 1792 |
+
[
|
| 1793 |
+
23
|
| 1794 |
+
]
|
| 1795 |
+
, using standard IO prompting or chain-of-thought (CoT) prompting. We show how deliberate search in trees of thoughts (ToT) produces better results, and more importantly, interesting and promising new ways to use language models to solve problems requiring search or planning.
|
| 1796 |
+
Unless otherwise stated, we perform experiments using a Chat Completion mode GPT-4
|
| 1797 |
+
1
|
| 1798 |
+
1
|
| 1799 |
+
1
|
| 1800 |
+
Experiments were done between May 5-16, 2023.
|
| 1801 |
+
with a sampling temperature of 0.7.
|
| 1802 |
+
4.1
|
| 1803 |
+
Game of 24
|
| 1804 |
+
Game of 24 is a mathematical reasoning challenge, where the goal is to use 4 numbers and basic arithmetic operations (+-*/) to obtain 24.
|
| 1805 |
+
For example, given input “4 9 10 13”, a solution output could be “(10 - 4) * (13 - 9) = 24”.
|
| 1806 |
+
Figure 2:
|
| 1807 |
+
ToT in a game of 24. The LM is prompted for (a) thought generation and (b) valuation.
|
| 1808 |
+
Method
|
| 1809 |
+
Success
|
| 1810 |
+
IO prompt
|
| 1811 |
+
7.3%
|
| 1812 |
+
CoT prompt
|
| 1813 |
+
4.0%
|
| 1814 |
+
CoT-SC
|
| 1815 |
+
(k=100)
|
| 1816 |
+
9.0%
|
| 1817 |
+
ToT (ours)
|
| 1818 |
+
(b=1)
|
| 1819 |
+
45%
|
| 1820 |
+
ToT (ours)
|
| 1821 |
+
(b=5)
|
| 1822 |
+
74%
|
| 1823 |
+
IO + Refine
|
| 1824 |
+
(k=10)
|
| 1825 |
+
27%
|
| 1826 |
+
IO
|
| 1827 |
+
(best of 100)
|
| 1828 |
+
33%
|
| 1829 |
+
CoT
|
| 1830 |
+
(best of 100)
|
| 1831 |
+
49%
|
| 1832 |
+
Table 2:
|
| 1833 |
+
Game of 24 Results.
|
| 1834 |
+
Figure 3:
|
| 1835 |
+
Game of 24 (a) scale analysis & (b) error analysis.
|
| 1836 |
+
Task Setup.
|
| 1837 |
+
We scrape data from
|
| 1838 |
+
4nums.com
|
| 1839 |
+
, which has 1,362 games that are sorted from easy to hard by human solving time, and use a subset of relatively hard games indexed 901-1,000 for testing. For each task, we consider the output as success if it is a valid equation that equals 24 and uses the input numbers each exactly once. We report the success rate across 100 games as the metric.
|
| 1840 |
+
Baselines.
|
| 1841 |
+
We use a standard input-output (IO) prompt with 5 in-context examples. For chain-of-thought (CoT) prompting, we augment each input-output pair with 3 intermediate equations, each operating on two remaining numbers. For example, given input “4 9 10 13”, the thoughts could be “13 - 9 = 4 (left: 4 4 10); 10 - 4 = 6 (left: 4 6); 4 * 6 = 24 (left: 24)”. For each game, we sample IO and CoT prompting for 100 times for average performance.
|
| 1842 |
+
We also consider a CoT self-consistency baseline, which takes the majority output from 100 CoT samples, and an iterative-refine approach on top of an IO sample for at most
|
| 1843 |
+
10
|
| 1844 |
+
10
|
| 1845 |
+
10
|
| 1846 |
+
iterations. At each iteration, the LM is conditioned on all previous history to “reflect on your mistakes and generate a refined answer” if the output is incorrect. Note that it uses groundtruth feedback signals about equation correctness.
|
| 1847 |
+
ToT Setup.
|
| 1848 |
+
To frame Game of 24 into ToT, it is natural to decompose the thoughts into 3 steps, each an intermediate equation. As shown in Figure
|
| 1849 |
+
2
|
| 1850 |
+
(a), at each tree node, we exact the remaining numbers and prompt the LM to propose some possible next steps.
|
| 1851 |
+
The same “propose prompt” is used for all 3 thought steps, though it only has one example with 4 input numbers.
|
| 1852 |
+
We perform a breadth-first search (BFS) in ToT, where at each step we keep the best
|
| 1853 |
+
b
|
| 1854 |
+
=
|
| 1855 |
+
5
|
| 1856 |
+
𝑏
|
| 1857 |
+
5
|
| 1858 |
+
b=5
|
| 1859 |
+
candidates.
|
| 1860 |
+
To perform deliberate BFS in ToT, as shown in Figure
|
| 1861 |
+
2
|
| 1862 |
+
(b), we prompt LM to evaluate each thought candidate as “sure/maybe/impossible” with regard to reaching 24. The aim is to promote correct partial solutions that can be verdicted within few lookahead trials, and eliminate impossible partial solutions based on “too big/small” commonsense, and keep the rest “maybe”. We sample values
|
| 1863 |
+
3
|
| 1864 |
+
3
|
| 1865 |
+
3
|
| 1866 |
+
times for each thought.
|
| 1867 |
+
Results.
|
| 1868 |
+
As shown in Table
|
| 1869 |
+
3
|
| 1870 |
+
, IO, CoT, and CoT-SC prompting methods perform badly on the task, achieving only 7.3%, 4.0%, and 9.0% success rates. In contrast, ToT with a breadth of
|
| 1871 |
+
b
|
| 1872 |
+
=
|
| 1873 |
+
1
|
| 1874 |
+
𝑏
|
| 1875 |
+
1
|
| 1876 |
+
b=1
|
| 1877 |
+
already achieves a success rate of
|
| 1878 |
+
45
|
| 1879 |
+
%
|
| 1880 |
+
percent
|
| 1881 |
+
45
|
| 1882 |
+
45\%
|
| 1883 |
+
, while
|
| 1884 |
+
b
|
| 1885 |
+
=
|
| 1886 |
+
5
|
| 1887 |
+
𝑏
|
| 1888 |
+
5
|
| 1889 |
+
b=5
|
| 1890 |
+
achieves
|
| 1891 |
+
74
|
| 1892 |
+
%
|
| 1893 |
+
percent
|
| 1894 |
+
74
|
| 1895 |
+
74\%
|
| 1896 |
+
.
|
| 1897 |
+
We also consider an oracle setup for IO/CoT, by calculating the success rate using best of
|
| 1898 |
+
k
|
| 1899 |
+
𝑘
|
| 1900 |
+
k
|
| 1901 |
+
samples
|
| 1902 |
+
(
|
| 1903 |
+
1
|
| 1904 |
+
≤
|
| 1905 |
+
k
|
| 1906 |
+
≤
|
| 1907 |
+
100
|
| 1908 |
+
)
|
| 1909 |
+
1
|
| 1910 |
+
𝑘
|
| 1911 |
+
100
|
| 1912 |
+
(1\leq k\leq 100)
|
| 1913 |
+
. To compare IO/CoT (best of k) with ToT, we consider calculating the tree nodes visited per task in ToT across
|
| 1914 |
+
b
|
| 1915 |
+
=
|
| 1916 |
+
1
|
| 1917 |
+
|
| 1918 |
+
⋯
|
| 1919 |
+
|
| 1920 |
+
5
|
| 1921 |
+
𝑏
|
| 1922 |
+
1
|
| 1923 |
+
⋯
|
| 1924 |
+
5
|
| 1925 |
+
b=1\cdots 5
|
| 1926 |
+
, and map the 5 success rates in Figure
|
| 1927 |
+
3
|
| 1928 |
+
(a), treating IO/CoT (best of
|
| 1929 |
+
k
|
| 1930 |
+
𝑘
|
| 1931 |
+
k
|
| 1932 |
+
) as visiting
|
| 1933 |
+
k
|
| 1934 |
+
𝑘
|
| 1935 |
+
k
|
| 1936 |
+
nodes in a bandit. Not surprisingly, CoT scales better than IO, and best of 100 CoT samples achieve a success rate of
|
| 1937 |
+
49
|
| 1938 |
+
%
|
| 1939 |
+
percent
|
| 1940 |
+
49
|
| 1941 |
+
49\%
|
| 1942 |
+
, but still much worse than exploring more nodes in ToT (
|
| 1943 |
+
b
|
| 1944 |
+
>
|
| 1945 |
+
1
|
| 1946 |
+
𝑏
|
| 1947 |
+
1
|
| 1948 |
+
b>1
|
| 1949 |
+
).
|
| 1950 |
+
Error analysis.
|
| 1951 |
+
Figure
|
| 1952 |
+
3
|
| 1953 |
+
(b) breaks down at which step CoT and ToT samples fail the task, i.e. the thought (in CoT) or all
|
| 1954 |
+
b
|
| 1955 |
+
𝑏
|
| 1956 |
+
b
|
| 1957 |
+
thoughts (in ToT) are invalid or impossible to reach 24. Notably, around 60% of CoT samples already failed the task after generating the first step, or equivalently, the first three words (e.g. “
|
| 1958 |
+
4
|
| 1959 |
+
+
|
| 1960 |
+
9
|
| 1961 |
+
4
|
| 1962 |
+
9
|
| 1963 |
+
4+9
|
| 1964 |
+
”). This highlights the issues with direct left-to-right decoding.
|
| 1965 |
+
4.2
|
| 1966 |
+
Creative writing
|
| 1967 |
+
Next, we invent a creative writing task where the input is 4 random sentences and the output should be a coherent passage with 4 paragraphs that end in the 4 input sentences respectively.
|
| 1968 |
+
Such a task is open-ended and exploratory, and challenges creative thinking as well as high-level planning.
|
| 1969 |
+
Task setup.
|
| 1970 |
+
We sample random sentences from
|
| 1971 |
+
randomwordgenerator.com
|
| 1972 |
+
to form 100 inputs, and there is no groundtruth passage for each input constraint. As we find that GPT-4 can follow the input constraints most of the time, we focus on evaluating passage coherency in two ways: using a GPT-4 zero-shot prompt to provide a 1-10 scalar score, or using human judgments to compare pairs of outputs from different methods. For the former, we sample 5 scores and average them for each task output, and we find these 5 scores usually consistent, with a standard deviation of around
|
| 1973 |
+
0.56
|
| 1974 |
+
0.56
|
| 1975 |
+
0.56
|
| 1976 |
+
on average across outputs. For the latter, we employ a subset of the authors in a blind study to compare the coherency of CoT vs. ToT generated passage pairs, where the order of passages is random flipped over 100 inputs.
|
| 1977 |
+
Baselines.
|
| 1978 |
+
Given the creative nature of the task, both IO and CoT prompts are zero-shot. While the former prompts the LM to directly generate a coherent passage given input constraints, the latter prompts the LM to first make a brief plan then write the passage, i.e. the plan serves as the intermediate thought step. We generate 10 IO and CoT samples per task.
|
| 1979 |
+
We also consider an iterative-refine (
|
| 1980 |
+
k
|
| 1981 |
+
≤
|
| 1982 |
+
5
|
| 1983 |
+
𝑘
|
| 1984 |
+
5
|
| 1985 |
+
k\leq 5
|
| 1986 |
+
) method on top of a random IO sample for each task, where the LM is conditioned on input constraints and the last generated passage to decide if the passage is already “perfectly coherent”, and if not generate a refined one.
|
| 1987 |
+
ToT setup.
|
| 1988 |
+
We build a ToT with depth 2 (and only 1 intermediate thought step) — the LM first generates
|
| 1989 |
+
k
|
| 1990 |
+
=
|
| 1991 |
+
5
|
| 1992 |
+
𝑘
|
| 1993 |
+
5
|
| 1994 |
+
k=5
|
| 1995 |
+
plans and votes for the best one (Figure
|
| 1996 |
+
4
|
| 1997 |
+
), then similarly generate
|
| 1998 |
+
k
|
| 1999 |
+
=
|
| 2000 |
+
5
|
| 2001 |
+
𝑘
|
| 2002 |
+
5
|
| 2003 |
+
k=5
|
| 2004 |
+
passages based on the best plan then vote for the best one. Here the breadth limit
|
| 2005 |
+
b
|
| 2006 |
+
=
|
| 2007 |
+
1
|
| 2008 |
+
𝑏
|
| 2009 |
+
1
|
| 2010 |
+
b=1
|
| 2011 |
+
, as only one choice is kept per step. A simple zero-shot vote prompt (“analyze choices below, then conclude which is most promising for the instruction”) is used to sample 5 votes at both steps.
|
| 2012 |
+
Results.
|
| 2013 |
+
Figure
|
| 2014 |
+
5
|
| 2015 |
+
(a) shows average GPT-4 scores across 100 tasks, where ToT (7.56) is deemed to generate more coherent passages than IO (6.19) and CoT (6.93) on average. While such an automatic metric might be noisy, Figure
|
| 2016 |
+
5
|
| 2017 |
+
(b) confirms the finding by showing that humans prefer ToT over CoT in 41 out of 100 passage pairs, while only prefer CoT over ToT in 21 (other 38 pairs are found “similarly coherent”). Lastly, iterative-refine is more effective on this natural language task, where it improves IO coherency score from 6.19 to 7.67, and ToT coherency score from 7.56 to 7.91.
|
| 2018 |
+
We believe it could be thought of as a third approach to thought generation in the ToT framework, where new thoughts can arise from refining old thoughts instead of i.i.d. or sequentially generated.
|
| 2019 |
+
Figure 4:
|
| 2020 |
+
A step of deliberate search in a randomly picked Creative Writing task. Given the input, the LM samples 5 different plans, then votes 5 times to decide which plan is best. The majority choice is used to consequently write the output passage with the same sample-vote procedure.
|
| 2021 |
+
Figure 5:
|
| 2022 |
+
Creative Writing results.
|
| 2023 |
+
Method
|
| 2024 |
+
Success Rate (%)
|
| 2025 |
+
Letter
|
| 2026 |
+
Word
|
| 2027 |
+
Game
|
| 2028 |
+
IO
|
| 2029 |
+
38.7
|
| 2030 |
+
14
|
| 2031 |
+
0
|
| 2032 |
+
CoT
|
| 2033 |
+
40.6
|
| 2034 |
+
15.6
|
| 2035 |
+
1
|
| 2036 |
+
ToT (ours)
|
| 2037 |
+
78
|
| 2038 |
+
60
|
| 2039 |
+
20
|
| 2040 |
+
+best state
|
| 2041 |
+
82.4
|
| 2042 |
+
67.5
|
| 2043 |
+
35
|
| 2044 |
+
-prune
|
| 2045 |
+
65.4
|
| 2046 |
+
41.5
|
| 2047 |
+
5
|
| 2048 |
+
-backtrack
|
| 2049 |
+
54.6
|
| 2050 |
+
20
|
| 2051 |
+
5
|
| 2052 |
+
Table 3:
|
| 2053 |
+
Mini Crosswords results.
|
| 2054 |
+
4.3
|
| 2055 |
+
Mini crosswords
|
| 2056 |
+
Figure 6:
|
| 2057 |
+
In Mini Crosswords, (a) how thoughts are proposed and aggregated in a priority queue for depth-first search (DFS), and (b) how a state is evaluated based on the possibility of filling in each remaining word clue, and pruned if any remaining clue is deemed not possible to fill by the LM. Then DFS backtracks to the parent state and explore the next promising thought for clue.
|
| 2058 |
+
In Game of 24 and Creative Writing, ToT is relatively shallow — at most 3 thought steps are needed to reach the final output. Here we explore
|
| 2059 |
+
5
|
| 2060 |
+
×
|
| 2061 |
+
5
|
| 2062 |
+
5
|
| 2063 |
+
5
|
| 2064 |
+
5\times 5
|
| 2065 |
+
mini crosswords as a harder search problem involving natural language. Again, the goal is not just to solve the task, as more general crosswords can be readily solved with specialized NLP pipelines
|
| 2066 |
+
[
|
| 2067 |
+
34
|
| 2068 |
+
]
|
| 2069 |
+
that leverages large-scale retrieval instead of LM. Rather, we aim to explore the limit of LM as a general problem solver that explores its own thoughts and guides its own exploration with deliberate reasoning as heuristics.
|
| 2070 |
+
Task setup.
|
| 2071 |
+
We scrape data from
|
| 2072 |
+
GooBix
|
| 2073 |
+
, which contains 156 games of
|
| 2074 |
+
5
|
| 2075 |
+
×
|
| 2076 |
+
5
|
| 2077 |
+
5
|
| 2078 |
+
5
|
| 2079 |
+
5\times 5
|
| 2080 |
+
mini crosswords. As we observe adjacent games contain similar clues, we use 20 games with indices
|
| 2081 |
+
1
|
| 2082 |
+
,
|
| 2083 |
+
6
|
| 2084 |
+
,
|
| 2085 |
+
⋯
|
| 2086 |
+
,
|
| 2087 |
+
91
|
| 2088 |
+
,
|
| 2089 |
+
96
|
| 2090 |
+
1
|
| 2091 |
+
6
|
| 2092 |
+
⋯
|
| 2093 |
+
91
|
| 2094 |
+
96
|
| 2095 |
+
1,6,\cdots,91,96
|
| 2096 |
+
for testing, and games
|
| 2097 |
+
136
|
| 2098 |
+
,
|
| 2099 |
+
141
|
| 2100 |
+
,
|
| 2101 |
+
146
|
| 2102 |
+
,
|
| 2103 |
+
151
|
| 2104 |
+
,
|
| 2105 |
+
156
|
| 2106 |
+
136
|
| 2107 |
+
141
|
| 2108 |
+
146
|
| 2109 |
+
151
|
| 2110 |
+
156
|
| 2111 |
+
136,141,146,151,156
|
| 2112 |
+
for prompting.
|
| 2113 |
+
For each task, the input describes the 5 horizontal clues and 5 vertical clues, and the output should be a board of
|
| 2114 |
+
5
|
| 2115 |
+
×
|
| 2116 |
+
5
|
| 2117 |
+
=
|
| 2118 |
+
25
|
| 2119 |
+
5
|
| 2120 |
+
5
|
| 2121 |
+
25
|
| 2122 |
+
5\times 5=25
|
| 2123 |
+
letters to solve the crosswords. For evaluation, we consider three levels of success: the portion of correct letters (25 per game), words (10 per game), and games.
|
| 2124 |
+
Baselines.
|
| 2125 |
+
We provide 5 example input-output pairs in the IO prompt, and in the CoT prompt additionally include intermediate words in the order h1..5 then v1..5. We run each prompt for 10 samples and average the results.
|
| 2126 |
+
ToT setup.
|
| 2127 |
+
We leverage a depth-first search (Algorithm
|
| 2128 |
+
2
|
| 2129 |
+
) that keeps exploring the most promising subsequent word clue until the state is no longer promising, then backtrack to the parent state to explore alternative thoughts.
|
| 2130 |
+
To make search tractable, subsequent thoughts are constrained not to change any filled words or letters, so that the ToT has at most 10 intermediate steps.
|
| 2131 |
+
For thought generation, at each state we translate all existing thoughts (e.g. “h2.motor; h1.tasks” for the state in Figure
|
| 2132 |
+
6
|
| 2133 |
+
(a)) into letter constraints for remaining clues (e.g. “v1.To heap: tm___;…”) and prompt a proposal prompt
|
| 2134 |
+
5
|
| 2135 |
+
5
|
| 2136 |
+
5
|
| 2137 |
+
times to come up with candidates for where and what to fill in the next word. Importantly, we also prompt the LM to give a confidence level for different thoughts, and aggregate these across proposals to obtain a sorted list of next thoughts to explore (Figure
|
| 2138 |
+
6
|
| 2139 |
+
(a)).
|
| 2140 |
+
For state evaluations, we similarly translate each state into letter constraints for remaining clues, then evaluate for each clue if it is possible to fill given the constraints. If any remaining clue is deemed “impossible” to fill in (e.g. “v1. To heap: tm_s_”), then the exploration of the state’s subtree is pruned and DFS backtracks to its parent to explore the next promising thought. We limit DFS search steps to 100, and simply render the deepest explored state (the first explored one if multiple) into the final output.
|
| 2141 |
+
Results.
|
| 2142 |
+
As shown in Table
|
| 2143 |
+
5
|
| 2144 |
+
, IO and CoT prompting methods perform poorly with a word-level success rate less than
|
| 2145 |
+
16
|
| 2146 |
+
%
|
| 2147 |
+
percent
|
| 2148 |
+
16
|
| 2149 |
+
16\%
|
| 2150 |
+
, while ToT significantly improves all metrics, achieving a word-level success rate of
|
| 2151 |
+
60
|
| 2152 |
+
%
|
| 2153 |
+
percent
|
| 2154 |
+
60
|
| 2155 |
+
60\%
|
| 2156 |
+
and solving 4 out of 20 games. Such an improvement is not surprising, given IO and CoT lack mechanisms to try different clues, make changes to decisions, or backtrack.
|
| 2157 |
+
Oracle and ablation studies.
|
| 2158 |
+
When outputting from the oracle best DFS state (instead of the heuristically determined best state) per task, ToT performance is even higher and actually solves 7/20 games (Table
|
| 2159 |
+
5
|
| 2160 |
+
, “+best state”), indicating our simple output heuristics can be readily improved. Interestingly, sometimes when the crosswords game is actually solved, the state evaluator might still deem some words as “impossible” and prune — possibly because
|
| 2161 |
+
5
|
| 2162 |
+
×
|
| 2163 |
+
5
|
| 2164 |
+
5
|
| 2165 |
+
5
|
| 2166 |
+
5\times 5
|
| 2167 |
+
crosswords by design have some rare or obselete words that GPT-4 cannot recognize
|
| 2168 |
+
2
|
| 2169 |
+
2
|
| 2170 |
+
2
|
| 2171 |
+
For example, “agend” is an obsolete form of “agendum”, but GPT-4 deems it a typo for “agenda”. External retrieval
|
| 2172 |
+
or web interaction
|
| 2173 |
+
could augment LM for problem solving under knowledge uncertainty.
|
| 2174 |
+
.
|
| 2175 |
+
Given the state evaluation as a pruning heuristic is imperfect, we also explore ablating the pruning, and find the performance generally worse (Table
|
| 2176 |
+
5
|
| 2177 |
+
, “-prune”). However, it could actually find the correct solution for 4/20 games (though only outputting 1 via heuristic), 3 of which are games ToT+pruning cannot solve within 100 steps. Thus, better heuristics for DFS pruning are critical for problem solving in this case.
|
| 2178 |
+
Lastly, we confirm the importance of backtracking by running an ablation that keeps filling the most promising clue for at most 20 steps, allowing overwrites. This is similar to a “greedy” BFS search with breadth limit of
|
| 2179 |
+
b
|
| 2180 |
+
=
|
| 2181 |
+
1
|
| 2182 |
+
𝑏
|
| 2183 |
+
1
|
| 2184 |
+
b=1
|
| 2185 |
+
, and performs poorly with a word level success of only
|
| 2186 |
+
20
|
| 2187 |
+
%
|
| 2188 |
+
percent
|
| 2189 |
+
20
|
| 2190 |
+
20\%
|
| 2191 |
+
(Table
|
| 2192 |
+
5
|
| 2193 |
+
, “-backtrack”).
|
| 2194 |
+
5
|
| 2195 |
+
Related Work
|
| 2196 |
+
Planning and decision making.
|
| 2197 |
+
Smart planning and decision making are critical to achieving predefined goals. As they are trained on vast amount of world knowledge and human examples,
|
| 2198 |
+
LMs are known to have already absorbed rich commonsense that makes it possible to propose reasonable plans conditioned on problem setting and environmental states
|
| 2199 |
+
[
|
| 2200 |
+
12
|
| 2201 |
+
,
|
| 2202 |
+
42
|
| 2203 |
+
,
|
| 2204 |
+
37
|
| 2205 |
+
,
|
| 2206 |
+
13
|
| 2207 |
+
,
|
| 2208 |
+
35
|
| 2209 |
+
,
|
| 2210 |
+
41
|
| 2211 |
+
,
|
| 2212 |
+
40
|
| 2213 |
+
]
|
| 2214 |
+
. Our proposed ToT approach extends existing planning formulations by considering multiple potentially feasible plans simultaneously at each problem-solving step, and proceeding with the most promising ones. The integration between thought sampling and value feedback organically integrates planning and decision-making mechanisms, enabling effective search inside a solution tree. On the other hand, traditional decision-making procedures usually require training dedicated reward and policy models as in reinforcement learning (for example CHAI
|
| 2215 |
+
[
|
| 2216 |
+
33
|
| 2217 |
+
]
|
| 2218 |
+
), whereas we use the LM itself to provide the value estimates for decision making.
|
| 2219 |
+
RAP
|
| 2220 |
+
[
|
| 2221 |
+
9
|
| 2222 |
+
]
|
| 2223 |
+
is a concurrent work that treats language model reasoning as planning with its internal world model, and proposes a MCTS-based method similar to ToT. However, its tasks are simpler than ours, and its framework lacks the modularity to incorporate different tree search algorithms.
|
| 2224 |
+
Self-reflection.
|
| 2225 |
+
Using LLMs to assess the viability of their own predictions is becoming an increasingly important procedure in problem solving.
|
| 2226 |
+
[
|
| 2227 |
+
28
|
| 2228 |
+
,
|
| 2229 |
+
20
|
| 2230 |
+
,
|
| 2231 |
+
24
|
| 2232 |
+
]
|
| 2233 |
+
introduced the “self-reflection” mechanism, in which LMs provide feedback to their generation candidates.
|
| 2234 |
+
[
|
| 2235 |
+
4
|
| 2236 |
+
]
|
| 2237 |
+
improves LMs code generation accuracy by injecting feedback messages generated by the LM itself based on its code execution results. Similarly,
|
| 2238 |
+
[
|
| 2239 |
+
17
|
| 2240 |
+
]
|
| 2241 |
+
also introduces “critic” or review steps over the actions and states, deciding the next action to take in solving computer operation tasks. Another recent work very relevant to ours is “self-eval guided decoding”
|
| 2242 |
+
[
|
| 2243 |
+
39
|
| 2244 |
+
]
|
| 2245 |
+
. Similar to our method, self-eval decoding also follows a tree-search procedure with leaves sampled from stochastic beam search decoding, which are then evaluated by LLM itself with carefully prepared self-eval prompts. Their approach however, uses the PAL formulation
|
| 2246 |
+
[
|
| 2247 |
+
8
|
| 2248 |
+
]
|
| 2249 |
+
which represents thoughts as codes, which makes it difficult to tackle challenging tasks like creative writing which we consider in this paper. Our Tree-of-Thought formulation is thus more versatile and handles challenging tasks on which GPT-4 only achieves very low accuracy with standard prompts.
|
| 2250 |
+
Program-guided LLM generation.
|
| 2251 |
+
Our proposal is also related to recent advancements that organize LM’s behavior with systematic procedures
|
| 2252 |
+
[
|
| 2253 |
+
14
|
| 2254 |
+
,
|
| 2255 |
+
44
|
| 2256 |
+
,
|
| 2257 |
+
6
|
| 2258 |
+
,
|
| 2259 |
+
43
|
| 2260 |
+
]
|
| 2261 |
+
or symbolic program guidance. For example,
|
| 2262 |
+
Schlag et al. [
|
| 2263 |
+
27
|
| 2264 |
+
]
|
| 2265 |
+
embeds LMs in an algorithmic search procedure to help solve problems like question answering step-by-step, in which the search trees are expanded by relevant paragraphs that might provide answers. This approach however differs from ours in that trees are expanded by sampling external paragraphs instead of the LM’s own thoughts, and there is no reflection or voting steps. Another approach, LLM+P
|
| 2266 |
+
[
|
| 2267 |
+
18
|
| 2268 |
+
]
|
| 2269 |
+
, goes one step further and delegates the actual planning process to a classical planner.
|
| 2270 |
+
Classical search methods.
|
| 2271 |
+
Last but not least, our approach can be treated as a modern rendition of classical search methods for problem solving. For example it can be considered as a heuristic search algorithm like A*
|
| 2272 |
+
[
|
| 2273 |
+
10
|
| 2274 |
+
]
|
| 2275 |
+
, in which the heuristic at each search node is provided by the LM’s self-assessment. From this perspective, our method is also related to NeuroLogic A*esque decoding
|
| 2276 |
+
[
|
| 2277 |
+
19
|
| 2278 |
+
]
|
| 2279 |
+
, which is inspired by A* search but introduces look-ahead heuristics that are efficient for LMs to improve the beam-search or top-k sampling decoding. This method however is constrained to sentence generation tasks, whereas our framework are designed for complex, multi-step problem solving guarded by value feedback.
|
| 2280 |
+
6
|
| 2281 |
+
Discussion
|
| 2282 |
+
Limitations and future directions.
|
| 2283 |
+
Deliberate search such as ToT might not be necessary for many existing tasks that GPT-4 already excels at (see Appendix
|
| 2284 |
+
B.1
|
| 2285 |
+
), and as an initial step this work only explores three relatively simple tasks that challenges GPT-4 (see Appendix
|
| 2286 |
+
B.2
|
| 2287 |
+
for some GPT-3.5 experiment results) and calls of better search and planning abilities incorporated with LMs. However, as we begin to deploy LMs for more real-world decision making applications (e.g. coding, data analysis, robotics, etc.), more complex tasks could emerge and present new opportunities to study these research questions. Also, search methods like ToT requires more resources (e.g. GPT-4 API cost) than sampling methods in order to improve task performances, but the modular flexibility of ToT allows users to customize such performance-cost tradeoffs, and ongoing open-source efforts
|
| 2288 |
+
[
|
| 2289 |
+
32
|
| 2290 |
+
]
|
| 2291 |
+
should readily reduce such costs in the near future. More details about cost and efficiency are in Appendix
|
| 2292 |
+
B.3
|
| 2293 |
+
. Lastly, this work focuses on using an off-the-shelf LM, and fine-tuning LMs using a ToT-style high-level counterfactual decision making (e.g. deliberating over potential choices for the next paragraph, instead of predicting the next token) might present opportunities to enhance the problem-solving capabilities of LMs.
|
| 2294 |
+
Conclusion.
|
| 2295 |
+
The associative “System 1” of LMs can be beneficially augmented by a “System 2” based on searching a tree of possible paths to the solution to a problem. The Tree of Thoughts framework provides a way to translate classical insights about problem-solving into actionable methods for contemporary LMs. At the same time, LMs address a weakness of these classical methods, providing a way to solve complex problems that are not easily formalized, such as creative writing. We see this intersection of LMs with classical approaches to AI as an exciting direction.
|
| 2296 |
+
Broader Impact
|
| 2297 |
+
ToT is a framework that empowers LMs to more autonomously and intelligently make decisions and solve problems. While current tasks are limited to reasoning and search problems, future applications involving interaction with external environments or humans could bring potential danger, e.g. facilitating harmful uses of LMs. On the other hand, ToT also improves the interpretability of model decisions and the opportunity for human alignment, as the resulting representations are readable, high-level language reasoning instead of implicit, low-level token values.
|
| 2298 |
+
Acknowledgements
|
| 2299 |
+
SY and KN acknowledge support from an Oracle Collaborative Research award and the National Science Foundation under Grant No. 2239363. Any opinions, findings, conclusions, or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation. SY is also supported by the Harold W. Dodds Fellowship from Princeton.
|
| 2300 |
+
References
|
| 2301 |
+
Brown et al. [2020]
|
| 2302 |
+
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal,
|
| 2303 |
+
A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al.
|
| 2304 |
+
Language models are few-shot learners.
|
| 2305 |
+
Advances in neural information processing systems
|
| 2306 |
+
,
|
| 2307 |
+
33:1877–1901, 2020.
|
| 2308 |
+
Browne et al. [2012]
|
| 2309 |
+
C. Browne, E. J. Powley, D. Whitehouse, S. M. M. Lucas, P. I. Cowling,
|
| 2310 |
+
P. Rohlfshagen, S. Tavener, D. P. Liebana, S. Samothrakis, and S. Colton.
|
| 2311 |
+
A survey of monte carlo tree search methods.
|
| 2312 |
+
IEEE Transactions on Computational Intelligence and AI in
|
| 2313 |
+
Games
|
| 2314 |
+
, 4:1–43, 2012.
|
| 2315 |
+
Campbell et al. [2002]
|
| 2316 |
+
M. Campbell, A. J. Hoane Jr, and F.-h. Hsu.
|
| 2317 |
+
Deep blue.
|
| 2318 |
+
Artificial intelligence
|
| 2319 |
+
, 134(1-2):57–83,
|
| 2320 |
+
2002.
|
| 2321 |
+
Chen et al. [2023]
|
| 2322 |
+
X. Chen, M. Lin, N. Schärli, and D. Zhou.
|
| 2323 |
+
Teaching large language models to self-debug, 2023.
|
| 2324 |
+
Chowdhery et al. [2022]
|
| 2325 |
+
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham,
|
| 2326 |
+
H. W. Chung, C. Sutton, S. Gehrmann, et al.
|
| 2327 |
+
Palm: Scaling language modeling with pathways.
|
| 2328 |
+
arXiv preprint arXiv:2204.02311
|
| 2329 |
+
, 2022.
|
| 2330 |
+
Creswell and Shanahan [2022]
|
| 2331 |
+
A. Creswell and M. Shanahan.
|
| 2332 |
+
Faithful reasoning using large language models.
|
| 2333 |
+
arXiv preprint arXiv:2208.14271
|
| 2334 |
+
, 2022.
|
| 2335 |
+
Daw et al. [2005]
|
| 2336 |
+
N. D. Daw, Y. Niv, and P. Dayan.
|
| 2337 |
+
Uncertainty-based competition between prefrontal and dorsolateral
|
| 2338 |
+
striatal systems for behavioral control.
|
| 2339 |
+
Nature neuroscience
|
| 2340 |
+
, 8(12):1704–1711,
|
| 2341 |
+
2005.
|
| 2342 |
+
Gao et al. [2023]
|
| 2343 |
+
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig.
|
| 2344 |
+
Pal: Program-aided language models, 2023.
|
| 2345 |
+
Hao et al. [2023]
|
| 2346 |
+
S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu.
|
| 2347 |
+
Reasoning with language model is planning with world model.
|
| 2348 |
+
arXiv preprint arXiv:2305.14992
|
| 2349 |
+
, 2023.
|
| 2350 |
+
Hart et al. [1968a]
|
| 2351 |
+
P. E. Hart, N. J. Nilsson, and B. Raphael.
|
| 2352 |
+
A formal basis for the heuristic determination of minimum cost paths.
|
| 2353 |
+
IEEE Transactions on Systems Science and Cybernetics
|
| 2354 |
+
,
|
| 2355 |
+
4(2):100–107, 1968a.
|
| 2356 |
+
doi:
|
| 2357 |
+
10.1109/TSSC.1968.300136
|
| 2358 |
+
.
|
| 2359 |
+
Hart et al. [1968b]
|
| 2360 |
+
P. E. Hart, N. J. Nilsson, and B. Raphael.
|
| 2361 |
+
A formal basis for the heuristic determination of minimum cost paths.
|
| 2362 |
+
IEEE transactions on Systems Science and Cybernetics
|
| 2363 |
+
,
|
| 2364 |
+
4(2):100–107, 1968b.
|
| 2365 |
+
Huang et al. [2022a]
|
| 2366 |
+
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch.
|
| 2367 |
+
Language models as zero-shot planners: Extracting actionable
|
| 2368 |
+
knowledge for embodied agents, 2022a.
|
| 2369 |
+
Huang et al. [2022b]
|
| 2370 |
+
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson,
|
| 2371 |
+
I. Mordatch, Y. Chebotar, et al.
|
| 2372 |
+
Inner monologue: Embodied reasoning through planning with language
|
| 2373 |
+
models.
|
| 2374 |
+
arXiv preprint arXiv:2207.05608
|
| 2375 |
+
, 2022b.
|
| 2376 |
+
Jung et al. [2022]
|
| 2377 |
+
J. Jung, L. Qin, S. Welleck, F. Brahman, C. Bhagavatula, R. L. Bras, and
|
| 2378 |
+
Y. Choi.
|
| 2379 |
+
Maieutic prompting: Logically consistent reasoning with recursive
|
| 2380 |
+
explanations.
|
| 2381 |
+
arXiv preprint arXiv:2205.11822
|
| 2382 |
+
, 2022.
|
| 2383 |
+
Kahneman [2011]
|
| 2384 |
+
D. Kahneman.
|
| 2385 |
+
Thinking, fast and slow
|
| 2386 |
+
.
|
| 2387 |
+
Macmillan, 2011.
|
| 2388 |
+
Kahneman et al. [2002]
|
| 2389 |
+
D. Kahneman, S. Frederick, et al.
|
| 2390 |
+
Representativeness revisited: Attribute substitution in intuitive
|
| 2391 |
+
judgment.
|
| 2392 |
+
Heuristics and biases: The psychology of intuitive judgment
|
| 2393 |
+
,
|
| 2394 |
+
49(49-81):74, 2002.
|
| 2395 |
+
Kim et al. [2023]
|
| 2396 |
+
G. Kim, P. Baldi, and S. McAleer.
|
| 2397 |
+
Language models can solve computer tasks, 2023.
|
| 2398 |
+
Liu et al. [2023]
|
| 2399 |
+
B. Liu, Y. Jiang, X. Zhang, Q. Liu, S. Zhang, J. Biswas, and P. Stone.
|
| 2400 |
+
Llm+p: Empowering large language models with optimal planning
|
| 2401 |
+
proficiency, 2023.
|
| 2402 |
+
Lu et al. [2021]
|
| 2403 |
+
X. Lu, S. Welleck, P. West, L. Jiang, J. Kasai, D. Khashabi, R. L. Bras,
|
| 2404 |
+
L. Qin, Y. Yu, R. Zellers, N. A. Smith, and Y. Choi.
|
| 2405 |
+
Neurologic a*esque decoding: Constrained text generation with
|
| 2406 |
+
lookahead heuristics.
|
| 2407 |
+
In
|
| 2408 |
+
North American Chapter of the Association for Computational
|
| 2409 |
+
Linguistics
|
| 2410 |
+
, 2021.
|
| 2411 |
+
Madaan et al. [2023]
|
| 2412 |
+
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon,
|
| 2413 |
+
N. Dziri, S. Prabhumoye, Y. Yang, S. Welleck, B. P. Majumder, S. Gupta,
|
| 2414 |
+
A. Yazdanbakhsh, and P. Clark.
|
| 2415 |
+
Self-refine: Iterative refinement with self-feedback, 2023.
|
| 2416 |
+
Newell et al. [1959]
|
| 2417 |
+
A. Newell, J. C. Shaw, and H. A. Simon.
|
| 2418 |
+
Report on a general problem solving program.
|
| 2419 |
+
In
|
| 2420 |
+
IFIP congress
|
| 2421 |
+
, volume 256, page 64. Pittsburgh, PA, 1959.
|
| 2422 |
+
Newell et al. [1972]
|
| 2423 |
+
A. Newell, H. A. Simon, et al.
|
| 2424 |
+
Human problem solving
|
| 2425 |
+
.
|
| 2426 |
+
Prentice-Hall, 1972.
|
| 2427 |
+
OpenAI [2023]
|
| 2428 |
+
OpenAI.
|
| 2429 |
+
Gpt-4 technical report.
|
| 2430 |
+
ArXiv
|
| 2431 |
+
, abs/2303.08774, 2023.
|
| 2432 |
+
Paul et al. [2023]
|
| 2433 |
+
D. Paul, M. Ismayilzada, M. Peyrard, B. Borges, A. Bosselut, R. West, and
|
| 2434 |
+
B. Faltings.
|
| 2435 |
+
Refiner: Reasoning feedback on intermediate representations, 2023.
|
| 2436 |
+
Radford et al. [2018]
|
| 2437 |
+
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al.
|
| 2438 |
+
Improving language understanding by generative pre-training.
|
| 2439 |
+
OpenAI blog
|
| 2440 |
+
, 2018.
|
| 2441 |
+
Radford et al. [2019]
|
| 2442 |
+
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al.
|
| 2443 |
+
Language models are unsupervised multitask learners.
|
| 2444 |
+
OpenAI blog
|
| 2445 |
+
, 1(8):9, 2019.
|
| 2446 |
+
Schlag et al. [2023]
|
| 2447 |
+
I. Schlag, S. Sukhbaatar, A. Celikyilmaz, W. tau Yih, J. Weston,
|
| 2448 |
+
J. Schmidhuber, and X. Li.
|
| 2449 |
+
Large language model programs, 2023.
|
| 2450 |
+
Shinn et al. [2023]
|
| 2451 |
+
N. Shinn, B. Labash, and A. Gopinath.
|
| 2452 |
+
Reflexion: an autonomous agent with dynamic memory and
|
| 2453 |
+
self-reflection, 2023.
|
| 2454 |
+
Silver et al. [2017]
|
| 2455 |
+
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez,
|
| 2456 |
+
T. Hubert, L. Baker, M. Lai, A. Bolton, et al.
|
| 2457 |
+
Mastering the game of go without human knowledge.
|
| 2458 |
+
nature
|
| 2459 |
+
, 550(7676):354–359, 2017.
|
| 2460 |
+
Sloman [1996]
|
| 2461 |
+
S. A. Sloman.
|
| 2462 |
+
The empirical case for two systems of reasoning.
|
| 2463 |
+
Psychological bulletin
|
| 2464 |
+
, 119(1):3, 1996.
|
| 2465 |
+
Stanovich [1999]
|
| 2466 |
+
K. E. Stanovich.
|
| 2467 |
+
Who is rational? Studies of individual differences in
|
| 2468 |
+
reasoning
|
| 2469 |
+
.
|
| 2470 |
+
Psychology Press, 1999.
|
| 2471 |
+
Touvron et al. [2023]
|
| 2472 |
+
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix,
|
| 2473 |
+
B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al.
|
| 2474 |
+
Llama: Open and efficient foundation language models.
|
| 2475 |
+
arXiv preprint arXiv:2302.13971
|
| 2476 |
+
, 2023.
|
| 2477 |
+
Verma et al. [2022]
|
| 2478 |
+
S. Verma, J. Fu, S. Yang, and S. Levine.
|
| 2479 |
+
Chai: A chatbot ai for task-oriented dialogue with offline
|
| 2480 |
+
reinforcement learning.
|
| 2481 |
+
In
|
| 2482 |
+
Proceedings of the 2022 Conference of the North American
|
| 2483 |
+
Chapter of the Association for Computational Linguistics: Human Language
|
| 2484 |
+
Technologies
|
| 2485 |
+
, pages 4471–4491, 2022.
|
| 2486 |
+
Wallace et al. [2022]
|
| 2487 |
+
E. Wallace, N. Tomlin, A. Xu, K. Yang, E. Pathak, M. Ginsberg, and D. Klein.
|
| 2488 |
+
Automated crossword solving.
|
| 2489 |
+
arXiv preprint arXiv:2205.09665
|
| 2490 |
+
, 2022.
|
| 2491 |
+
Wang et al. [2023a]
|
| 2492 |
+
L. Wang, W. Xu, Y. Lan, Z. Hu, Y. Lan, R. K.-W. Lee, and E.-P. Lim.
|
| 2493 |
+
Plan-and-solve prompting: Improving zero-shot chain-of-thought
|
| 2494 |
+
reasoning by large language models, 2023a.
|
| 2495 |
+
Wang et al. [2022]
|
| 2496 |
+
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, and D. Zhou.
|
| 2497 |
+
Self-consistency improves chain of thought reasoning in language
|
| 2498 |
+
models.
|
| 2499 |
+
arXiv preprint arXiv:2203.11171
|
| 2500 |
+
, 2022.
|
| 2501 |
+
Wang et al. [2023b]
|
| 2502 |
+
Z. Wang, S. Cai, A. Liu, X. Ma, and Y. Liang.
|
| 2503 |
+
Describe, explain, plan and select: Interactive planning with large
|
| 2504 |
+
language models enables open-world multi-task agents, 2023b.
|
| 2505 |
+
Wei et al. [2022]
|
| 2506 |
+
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou.
|
| 2507 |
+
Chain of thought prompting elicits reasoning in large language
|
| 2508 |
+
models.
|
| 2509 |
+
arXiv preprint arXiv:2201.11903
|
| 2510 |
+
, 2022.
|
| 2511 |
+
Xie et al. [2023]
|
| 2512 |
+
Y. Xie, K. Kawaguchi, Y. Zhao, X. Zhao, M.-Y. Kan, J. He, and Q. Xie.
|
| 2513 |
+
Decomposition enhances reasoning via self-evaluation guided decoding,
|
| 2514 |
+
2023.
|
| 2515 |
+
Yang et al. [2023]
|
| 2516 |
+
S. Yang, O. Nachum, Y. Du, J. Wei, P. Abbeel, and D. Schuurmans.
|
| 2517 |
+
Foundation models for decision making: Problems, methods, and
|
| 2518 |
+
opportunities, 2023.
|
| 2519 |
+
Yao et al. [2022]
|
| 2520 |
+
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao.
|
| 2521 |
+
ReAct: Synergizing reasoning and acting in language models.
|
| 2522 |
+
arXiv preprint arXiv:2210.03629
|
| 2523 |
+
, 2022.
|
| 2524 |
+
Zhang et al. [2023]
|
| 2525 |
+
S. Zhang, Z. Chen, Y. Shen, M. Ding, J. B. Tenenbaum, and C. Gan.
|
| 2526 |
+
Planning with large language models for code generation.
|
| 2527 |
+
In
|
| 2528 |
+
The Eleventh International Conference on Learning
|
| 2529 |
+
Representations
|
| 2530 |
+
, 2023.
|
| 2531 |
+
URL
|
| 2532 |
+
https://openreview.net/forum?id=Lr8cOOtYbfL
|
| 2533 |
+
.
|
| 2534 |
+
Zhou et al. [2022]
|
| 2535 |
+
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans,
|
| 2536 |
+
C. Cui, O. Bousquet, Q. Le, et al.
|
| 2537 |
+
Least-to-most prompting enables complex reasoning in large language
|
| 2538 |
+
models.
|
| 2539 |
+
arXiv preprint arXiv:2205.10625
|
| 2540 |
+
, 2022.
|
| 2541 |
+
Zhu et al. [2022]
|
| 2542 |
+
X. Zhu, J. Wang, L. Zhang, Y. Zhang, R. Gan, J. Zhang, and Y. Yang.
|
| 2543 |
+
Solving math word problem via cooperative reasoning induced language
|
| 2544 |
+
models.
|
| 2545 |
+
arXiv preprint arXiv:2210.16257
|
| 2546 |
+
, 2022.
|
| 2547 |
+
Appendix A
|
| 2548 |
+
Code, Prompts, Trajectories
|
| 2549 |
+
All code is available at
|
| 2550 |
+
https://github.com/princeton-nlp/tree-of-thought-llm
|
| 2551 |
+
.
|
| 2552 |
+
All prompts are available at
|
| 2553 |
+
https://github.com/princeton-nlp/tree-of-thought-llm/tree/master/src/tot/prompts
|
| 2554 |
+
.
|
| 2555 |
+
Trajectories are available at
|
| 2556 |
+
https://github.com/princeton-nlp/tree-of-thought-llm/tree/master/logs
|
| 2557 |
+
.
|
| 2558 |
+
Appendix B
|
| 2559 |
+
Additional Experiment Results
|
| 2560 |
+
Given the motivation of exploring and extending the capability frontier of language models, our experiments in the main paper have focused on a setup with the state-of-the-art language model (GPT-4), and three hard tasks invented to challenge it. Here, we report additional experiments with weaker LLM or easier tasks, and discuss cost and efficiency.
|
| 2561 |
+
GSM8K
|
| 2562 |
+
StrategyQA
|
| 2563 |
+
IO
|
| 2564 |
+
51
|
| 2565 |
+
73
|
| 2566 |
+
CoT
|
| 2567 |
+
86
|
| 2568 |
+
82
|
| 2569 |
+
ToT
|
| 2570 |
+
90
|
| 2571 |
+
83
|
| 2572 |
+
Table 4:
|
| 2573 |
+
New tasks with
|
| 2574 |
+
zero-shot ToT and GPT-4.
|
| 2575 |
+
GPT-4
|
| 2576 |
+
GPT-3.5
|
| 2577 |
+
IO
|
| 2578 |
+
7.3%
|
| 2579 |
+
6%
|
| 2580 |
+
CoT
|
| 2581 |
+
4.0%
|
| 2582 |
+
3%
|
| 2583 |
+
ToT
|
| 2584 |
+
74%
|
| 2585 |
+
19%
|
| 2586 |
+
Table 5:
|
| 2587 |
+
Game of 24 with
|
| 2588 |
+
GPT-4 vs GPT-3.5.
|
| 2589 |
+
GPT-4
|
| 2590 |
+
GPT-3.5
|
| 2591 |
+
IO
|
| 2592 |
+
6.19
|
| 2593 |
+
4.47
|
| 2594 |
+
CoT
|
| 2595 |
+
6.93
|
| 2596 |
+
5.16
|
| 2597 |
+
ToT
|
| 2598 |
+
7.56
|
| 2599 |
+
6.62
|
| 2600 |
+
Table 6:
|
| 2601 |
+
Creative Writing with
|
| 2602 |
+
GPT-4 vs. GPT-3.5.
|
| 2603 |
+
B.1
|
| 2604 |
+
Extension to new tasks (GSM8k, StrategyQA) with zero-shot ToT
|
| 2605 |
+
While more common NLP tasks might be too easy for GPT-4 and do not require ToT (which is why we considered harder new tasks), we believe applying ToT to new tasks could be straightforward. For example, we implemented a simple and generic zero-shot ToT-BFS similar to creative writing (sample 5 problem solving strategies then vote for the best one; then sample 5 solutions based on the best strategy then vote for the best one) for GSM8K and StrategyQA with few extra lines of code:
|
| 2606 |
+
# define the answer format of new tasks
|
| 2607 |
+
gsm8k_format = ‘"the answer is n" where n is a number’
|
| 2608 |
+
strategyqa_format = ‘either "the answer is yes" or "the answer is no"’
|
| 2609 |
+
|
| 2610 |
+
# define zero-shot io prompting
|
| 2611 |
+
standard_prompt = ‘Answer the following question with {format}: {input}’
|
| 2612 |
+
|
| 2613 |
+
# define thought format for zero-shot cot and zero-shot tot
|
| 2614 |
+
cot_prompt = ‘‘‘Answer the following question: {input}
|
| 2615 |
+
|
| 2616 |
+
Make a strategy then write. Your output should be of the following format:
|
| 2617 |
+
|
| 2618 |
+
Strategy:
|
| 2619 |
+
Your strategy about how to answer the question.
|
| 2620 |
+
|
| 2621 |
+
Answer:
|
| 2622 |
+
Your answer to the question. It should end with {format}.
|
| 2623 |
+
’’’
|
| 2624 |
+
|
| 2625 |
+
# define zero-shot voting used for zero-shot tot
|
| 2626 |
+
vote_prompt = ‘‘‘Given an instruction and several choices,
|
| 2627 |
+
decide which choice is most promising.
|
| 2628 |
+
Analyze each choice in detail, then conclude in the last line
|
| 2629 |
+
"The best choice is {s}", where s the integer id of the choice.
|
| 2630 |
+
’’’
|
| 2631 |
+
We evaluated on a subset of 100 random GSM8K test and StrategyQA dev questions. As shown in Table
|
| 2632 |
+
B
|
| 2633 |
+
and as expected, ToT improves over CoT on both tasks (but only slightly, given GPT-4 + CoT is already very good on such tasks, and StrategyQA’s bottleneck is external knowledge, not reasoning). Considering computational costs, it is more suitable to try smaller LLMs + ToT for traditional NLP tasks, or GPT-4 + ToT for hard tasks that challenge GPT-4 + CoT’s reasoning.
|
| 2634 |
+
B.2
|
| 2635 |
+
Extension to new LMs (GPT-3.5)
|
| 2636 |
+
To understand how ToT works with other LLMs, we also ran GPT-3.5-turbo for Creative Writing (Table
|
| 2637 |
+
B
|
| 2638 |
+
) and Game of 24 (Table
|
| 2639 |
+
B
|
| 2640 |
+
).
|
| 2641 |
+
On both tasks, “ToT
|
| 2642 |
+
>
|
| 2643 |
+
>
|
| 2644 |
+
CoT
|
| 2645 |
+
>
|
| 2646 |
+
>
|
| 2647 |
+
IO” remains true for GPT-3.5.
|
| 2648 |
+
On Creative Writing, we find GPT-3.5+ToT outperform GPT-4+IO, and similar to GPT-4+CoT, which suggests ToT could also work well on weaker language models.
|
| 2649 |
+
On Game of 24 (we changed 1-shot proposal prompt to 3-shot to make it work), GPT-3.5+ToT’s 19% is far worse than GPT-4+ToT’s 74%. To further understand the importance of generation vs. evaluation, we ran GPT-4 generation + GPT-3.5 evaluation (64%) and GPT-3.5 generation + GPT-4 evaluation (31%). This suggests the game’s bottleneck is thought generation, and different generation/evaluation language models might attain decent results while reducing costs.
|
| 2650 |
+
B.3
|
| 2651 |
+
Cost and efficiency
|
| 2652 |
+
Running ToT requires significantly more computations than IO or CoT prompting. For example, in Game of 24 (Table
|
| 2653 |
+
7
|
| 2654 |
+
below), solving a problem with ToT requires 5.5k completion tokens, close to 100 CoT trials (6.7k tokens). But the performance of ToT is better than best of 100 independent CoT trials.
|
| 2655 |
+
Game of 24
|
| 2656 |
+
Generate/Prompt tokens
|
| 2657 |
+
Cost per case
|
| 2658 |
+
Success
|
| 2659 |
+
IO (best of 100)
|
| 2660 |
+
1.8k / 1.0k
|
| 2661 |
+
$0.13
|
| 2662 |
+
33%
|
| 2663 |
+
CoT (best of 100)
|
| 2664 |
+
6.7k / 2.2k
|
| 2665 |
+
$0.47
|
| 2666 |
+
49%
|
| 2667 |
+
ToT
|
| 2668 |
+
5.5k / 1.4k
|
| 2669 |
+
$0.74
|
| 2670 |
+
74%
|
| 2671 |
+
Table 7:
|
| 2672 |
+
Cost analysis on Game of 24.
|
| 2673 |
+
On Creative Writing (Table
|
| 2674 |
+
8
|
| 2675 |
+
below), we found ToT takes around 5x completion tokens and money cost, which is intuitive as
|
| 2676 |
+
b
|
| 2677 |
+
=
|
| 2678 |
+
5
|
| 2679 |
+
𝑏
|
| 2680 |
+
5
|
| 2681 |
+
b=5
|
| 2682 |
+
and most tokens are generated passages.
|
| 2683 |
+
Creative Writing
|
| 2684 |
+
Generate/Prompt tokens
|
| 2685 |
+
Cost per case
|
| 2686 |
+
IO
|
| 2687 |
+
0.9k / 0.4k
|
| 2688 |
+
$0.06
|
| 2689 |
+
CoT
|
| 2690 |
+
0.9k / 0.4k
|
| 2691 |
+
$0.07
|
| 2692 |
+
ToT
|
| 2693 |
+
4k / 2.9k
|
| 2694 |
+
$0.32
|
| 2695 |
+
Table 8:
|
| 2696 |
+
Cost analysis on Game of 24.
|
| 2697 |
+
So completing Game of 24 and Creative Writing’s main ToT experiments cost around
|
| 2698 |
+
0.74
|
| 2699 |
+
×
|
| 2700 |
+
100
|
| 2701 |
+
+
|
| 2702 |
+
0.32
|
| 2703 |
+
×
|
| 2704 |
+
100
|
| 2705 |
+
=
|
| 2706 |
+
106
|
| 2707 |
+
0.74
|
| 2708 |
+
100
|
| 2709 |
+
0.32
|
| 2710 |
+
100
|
| 2711 |
+
106
|
| 2712 |
+
0.74\times 100+0.32\times 100=106
|
| 2713 |
+
dollars. Crosswords’ DFS experiments should be also within
|
| 2714 |
+
100
|
| 2715 |
+
100
|
| 2716 |
+
100
|
| 2717 |
+
dollars. In general, cost and efficiency of ToT highly depend on the prompts and search algorithms used, and could require 5-100 times more generated tokens than CoT. Some actionable insights:
|
| 2718 |
+
•
|
| 2719 |
+
We recommend using ToT on tasks requiring deliberate reasoning, on which CoT struggles.
|
| 2720 |
+
•
|
| 2721 |
+
Flexibility of ToT allows some performance-cost tradeoff, e.g., change beam size or vote number in BFS, few-shot vs. zero-shot prompting, GPT-3.5 vs. GPT-4, etc. One could configure the setup based on some resource constraints or performance goal.
|
| 2722 |
+
•
|
| 2723 |
+
There is much space for improving efficiency, e.g., BFS could early stop when solution is found, or trim down beam size to when some thoughts are ”impossible”.
|
| 2724 |
+
•
|
| 2725 |
+
We believe that more computation is indeed required in order for the model to achieve stronger intelligence, and this should not become a blocking issue as in the long run, (open-source) LMs will become much cheaper and more efficient. It is also a great direction how to better train/finetune LMs for thought generation and/or evaluation.
|
| 2726 |
+
◄
|
| 2727 |
+
Feeling
|
| 2728 |
+
lucky?
|
| 2729 |
+
Conversion
|
| 2730 |
+
report
|
| 2731 |
+
Report
|
| 2732 |
+
an issue
|
| 2733 |
+
View original
|
| 2734 |
+
on arXiv
|
| 2735 |
+
►
|
|
@@ -0,0 +1,213 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2305.10601] Tree of Thoughts: Deliberate Problem Solving with Large Language
|
| 3 |
+
Models'
|
| 4 |
+
id: 230510601-tree-of-thoughts-deliberate-problem-solving-with-large-language-models
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:39:55.627654Z'
|
| 8 |
+
source: https://arxiv.org/abs/2305.10601
|
| 9 |
+
source_domain: arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:39:55.627506Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2305.10601] Tree of Thoughts: Deliberate Problem Solving with Large Language Models
|
| 20 |
+
Computer Science > Computation and Language
|
| 21 |
+
arXiv:2305.10601
|
| 22 |
+
(cs)
|
| 23 |
+
[Submitted on 17 May 2023 (
|
| 24 |
+
v1
|
| 25 |
+
), last revised 3 Dec 2023 (this version, v2)]
|
| 26 |
+
Title:
|
| 27 |
+
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
|
| 28 |
+
Authors:
|
| 29 |
+
Shunyu Yao
|
| 30 |
+
,
|
| 31 |
+
Dian Yu
|
| 32 |
+
,
|
| 33 |
+
Jeffrey Zhao
|
| 34 |
+
,
|
| 35 |
+
Izhak Shafran
|
| 36 |
+
,
|
| 37 |
+
Thomas L. Griffiths
|
| 38 |
+
,
|
| 39 |
+
Yuan Cao
|
| 40 |
+
,
|
| 41 |
+
Karthik Narasimhan
|
| 42 |
+
View a PDF of the paper titled Tree of Thoughts: Deliberate Problem Solving with Large Language Models, by Shunyu Yao and 6 other authors
|
| 43 |
+
View PDF
|
| 44 |
+
HTML (experimental)
|
| 45 |
+
Abstract:
|
| 46 |
+
Language models are increasingly being deployed for general problem solving across a wide range of tasks, but are still confined to token-level, left-to-right decision-making processes during inference. This means they can fall short in tasks that require exploration, strategic lookahead, or where initial decisions play a pivotal role. To surmount these challenges, we introduce a new framework for language model inference, Tree of Thoughts (ToT), which generalizes over the popular Chain of Thought approach to prompting language models, and enables exploration over coherent units of text (thoughts) that serve as intermediate steps toward problem solving. ToT allows LMs to perform deliberate decision making by considering multiple different reasoning paths and self-evaluating choices to decide the next course of action, as well as looking ahead or backtracking when necessary to make global choices. Our experiments show that ToT significantly enhances language models' problem-solving abilities on three novel tasks requiring non-trivial planning or search: Game of 24, Creative Writing, and Mini Crosswords. For instance, in Game of 24, while GPT-4 with chain-of-thought prompting only solved 4% of tasks, our method achieved a success rate of 74%. Code repo with all prompts:
|
| 47 |
+
this https URL
|
| 48 |
+
.
|
| 49 |
+
Comments:
|
| 50 |
+
NeurIPS 2023 camera ready version. Code repo with all prompts:
|
| 51 |
+
this https URL
|
| 52 |
+
Subjects:
|
| 53 |
+
Computation and Language (cs.CL)
|
| 54 |
+
; Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
|
| 55 |
+
Cite as:
|
| 56 |
+
arXiv:2305.10601
|
| 57 |
+
[cs.CL]
|
| 58 |
+
(or
|
| 59 |
+
arXiv:2305.10601v2
|
| 60 |
+
[cs.CL]
|
| 61 |
+
for this version)
|
| 62 |
+
https://doi.org/10.48550/arXiv.2305.10601
|
| 63 |
+
Focus to learn more
|
| 64 |
+
arXiv-issued DOI via DataCite
|
| 65 |
+
Submission history
|
| 66 |
+
From: Shunyu Yao [
|
| 67 |
+
view email
|
| 68 |
+
]
|
| 69 |
+
[v1]
|
| 70 |
+
Wed, 17 May 2023 23:16:17 UTC (609 KB)
|
| 71 |
+
[v2]
|
| 72 |
+
Sun, 3 Dec 2023 22:50:35 UTC (623 KB)
|
| 73 |
+
Full-text links:
|
| 74 |
+
Access Paper:
|
| 75 |
+
View a PDF of the paper titled Tree of Thoughts: Deliberate Problem Solving with Large Language Models, by Shunyu Yao and 6 other authors
|
| 76 |
+
View PDF
|
| 77 |
+
HTML (experimental)
|
| 78 |
+
TeX Source
|
| 79 |
+
view license
|
| 80 |
+
Current browse context:
|
| 81 |
+
cs.CL
|
| 82 |
+
< prev
|
| 83 |
+
|
|
| 84 |
+
next >
|
| 85 |
+
new
|
| 86 |
+
|
|
| 87 |
+
recent
|
| 88 |
+
|
|
| 89 |
+
2023-05
|
| 90 |
+
Change to browse by:
|
| 91 |
+
cs
|
| 92 |
+
cs.AI
|
| 93 |
+
cs.LG
|
| 94 |
+
References & Citations
|
| 95 |
+
NASA ADS
|
| 96 |
+
Google Scholar
|
| 97 |
+
Semantic Scholar
|
| 98 |
+
5 blog links
|
| 99 |
+
(
|
| 100 |
+
what is this?
|
| 101 |
+
)
|
| 102 |
+
export BibTeX citation
|
| 103 |
+
Loading...
|
| 104 |
+
BibTeX formatted citation
|
| 105 |
+
×
|
| 106 |
+
loading...
|
| 107 |
+
Data provided by:
|
| 108 |
+
Bookmark
|
| 109 |
+
Bibliographic Tools
|
| 110 |
+
Bibliographic and Citation Tools
|
| 111 |
+
Bibliographic Explorer Toggle
|
| 112 |
+
Bibliographic Explorer
|
| 113 |
+
(
|
| 114 |
+
What is the Explorer?
|
| 115 |
+
)
|
| 116 |
+
Connected Papers Toggle
|
| 117 |
+
Connected Papers
|
| 118 |
+
(
|
| 119 |
+
What is Connected Papers?
|
| 120 |
+
)
|
| 121 |
+
Litmaps Toggle
|
| 122 |
+
Litmaps
|
| 123 |
+
(
|
| 124 |
+
What is Litmaps?
|
| 125 |
+
)
|
| 126 |
+
scite.ai Toggle
|
| 127 |
+
scite Smart Citations
|
| 128 |
+
(
|
| 129 |
+
What are Smart Citations?
|
| 130 |
+
)
|
| 131 |
+
Code, Data, Media
|
| 132 |
+
Code, Data and Media Associated with this Article
|
| 133 |
+
alphaXiv Toggle
|
| 134 |
+
alphaXiv
|
| 135 |
+
(
|
| 136 |
+
What is alphaXiv?
|
| 137 |
+
)
|
| 138 |
+
Links to Code Toggle
|
| 139 |
+
CatalyzeX Code Finder for Papers
|
| 140 |
+
(
|
| 141 |
+
What is CatalyzeX?
|
| 142 |
+
)
|
| 143 |
+
DagsHub Toggle
|
| 144 |
+
DagsHub
|
| 145 |
+
(
|
| 146 |
+
What is DagsHub?
|
| 147 |
+
)
|
| 148 |
+
GotitPub Toggle
|
| 149 |
+
Gotit.pub
|
| 150 |
+
(
|
| 151 |
+
What is GotitPub?
|
| 152 |
+
)
|
| 153 |
+
Huggingface Toggle
|
| 154 |
+
Hugging Face
|
| 155 |
+
(
|
| 156 |
+
What is Huggingface?
|
| 157 |
+
)
|
| 158 |
+
Links to Code Toggle
|
| 159 |
+
Papers with Code
|
| 160 |
+
(
|
| 161 |
+
What is Papers with Code?
|
| 162 |
+
)
|
| 163 |
+
ScienceCast Toggle
|
| 164 |
+
ScienceCast
|
| 165 |
+
(
|
| 166 |
+
What is ScienceCast?
|
| 167 |
+
)
|
| 168 |
+
Demos
|
| 169 |
+
Demos
|
| 170 |
+
Replicate Toggle
|
| 171 |
+
Replicate
|
| 172 |
+
(
|
| 173 |
+
What is Replicate?
|
| 174 |
+
)
|
| 175 |
+
Spaces Toggle
|
| 176 |
+
Hugging Face Spaces
|
| 177 |
+
(
|
| 178 |
+
What is Spaces?
|
| 179 |
+
)
|
| 180 |
+
Spaces Toggle
|
| 181 |
+
TXYZ.AI
|
| 182 |
+
(
|
| 183 |
+
What is TXYZ.AI?
|
| 184 |
+
)
|
| 185 |
+
Related Papers
|
| 186 |
+
Recommenders and Search Tools
|
| 187 |
+
Link to Influence Flower
|
| 188 |
+
Influence Flower
|
| 189 |
+
(
|
| 190 |
+
What are Influence Flowers?
|
| 191 |
+
)
|
| 192 |
+
Core recommender toggle
|
| 193 |
+
CORE Recommender
|
| 194 |
+
(
|
| 195 |
+
What is CORE?
|
| 196 |
+
)
|
| 197 |
+
Author
|
| 198 |
+
Venue
|
| 199 |
+
Institution
|
| 200 |
+
Topic
|
| 201 |
+
About arXivLabs
|
| 202 |
+
arXivLabs: experimental projects with community collaborators
|
| 203 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 204 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 205 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 206 |
+
Learn more about arXivLabs
|
| 207 |
+
.
|
| 208 |
+
Which authors of this paper are endorsers?
|
| 209 |
+
|
|
| 210 |
+
Disable MathJax
|
| 211 |
+
(
|
| 212 |
+
What is MathJax?
|
| 213 |
+
)
|
|
@@ -0,0 +1,4095 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2310.04406] Language Agent Tree Search Unifies Reasoning Acting and Planning
|
| 3 |
+
in Language Models'
|
| 4 |
+
id: 231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la-2
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:40:44.608943Z'
|
| 8 |
+
source: https://ar5iv.labs.arxiv.org/html/2310.04406
|
| 9 |
+
source_domain: ar5iv.labs.arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:40:44.608803Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2310.04406] Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
|
| 20 |
+
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
|
| 21 |
+
Andy Zhou
|
| 22 |
+
University of Illinois at Urbana-Champaign
|
| 23 |
+
AI@UIUC
|
| 24 |
+
Kai Yan
|
| 25 |
+
University of Illinois at Urbana-Champaign
|
| 26 |
+
Michal Shlapentokh-Rothman
|
| 27 |
+
University of Illinois at Urbana-Champaign
|
| 28 |
+
Haohan Wang
|
| 29 |
+
University of Illinois at Urbana-Champaign
|
| 30 |
+
Yu-Xiong Wang
|
| 31 |
+
University of Illinois at Urbana-Champaign
|
| 32 |
+
Abstract
|
| 33 |
+
While large language models (LLMs) have demonstrated impressive performance on a range of decision-making tasks, they rely on simple acting processes and fall short of broad deployment as autonomous agents. We introduce LATS (Language Agent Tree Search), a general framework that synergizes the capabilities of LLMs in planning, acting, and reasoning. Drawing inspiration from Monte Carlo tree search commonly used in model-based reinforcement learning, LATS employs LLMs as agents, value functions, and optimizers, repurposing their latent strengths for enhanced decision-making. What is crucial in this method is the use of an environment for external feedback, which offers a more deliberate and adaptive problem-solving mechanism that moves beyond the limitations of existing techniques. Our experimental evaluation across diverse domains, such as programming, HotPotQA, and WebShop, illustrates the applicability of LATS for decision-making while maintaining competitive reasoning performance. In particular, LATS achieves 94.4% for programming on HumanEval with GPT-4 and an average score of 75.9 for web browsing on WebShop with GPT-3.5, demonstrating the effectiveness and generality of our method.
|
| 34 |
+
1
|
| 35 |
+
Introduction
|
| 36 |
+
General autonomous agents capable of reasoning and decision-making in a variety of environments
|
| 37 |
+
(Wooldridge & Jennings,
|
| 38 |
+
1995
|
| 39 |
+
)
|
| 40 |
+
have been of longstanding interest in the field of artificial intelligence. While this has traditionally been studied in reinforcement learning, the recent rise of large language models (LLMs)
|
| 41 |
+
(Brown et al.,
|
| 42 |
+
2020
|
| 43 |
+
; Chowdhery et al.,
|
| 44 |
+
2022
|
| 45 |
+
; Touvron et al.,
|
| 46 |
+
2023
|
| 47 |
+
; OpenAI,
|
| 48 |
+
2023
|
| 49 |
+
)
|
| 50 |
+
with strong reasoning and general adaptability offers an alternative paradigm. Not only have LLMs excelled on standard NLP tasks such as text summarization
|
| 51 |
+
(Nallapati et al.,
|
| 52 |
+
2016
|
| 53 |
+
)
|
| 54 |
+
or natural language inference
|
| 55 |
+
(Bowman et al.,
|
| 56 |
+
2015
|
| 57 |
+
)
|
| 58 |
+
, but they have been adapted to an increasingly diverse set of tasks that often require advanced common-sense reasoning or quantitative skills
|
| 59 |
+
(Cobbe et al.,
|
| 60 |
+
2021
|
| 61 |
+
; Saparov & He,
|
| 62 |
+
2022
|
| 63 |
+
)
|
| 64 |
+
. LLMs are also capable of performing in complex environments that involve knowledge and reasoning, such as web navigation
|
| 65 |
+
(Yao et al.,
|
| 66 |
+
2022
|
| 67 |
+
; Deng et al.,
|
| 68 |
+
2023
|
| 69 |
+
)
|
| 70 |
+
, tool-use
|
| 71 |
+
(Schick et al.,
|
| 72 |
+
2023
|
| 73 |
+
)
|
| 74 |
+
, or open-ended games
|
| 75 |
+
(Fan et al.,
|
| 76 |
+
2022
|
| 77 |
+
)
|
| 78 |
+
.
|
| 79 |
+
Figure 1:
|
| 80 |
+
An overview of LATS. LATS uses an external environment and self-reflection to improve reasoning and decision-making.
|
| 81 |
+
Reasoning and acting abilities have also been improved by prompting techniques that augment LLMs with feedback or observations from an external environment
|
| 82 |
+
(Yao et al.,
|
| 83 |
+
2023b
|
| 84 |
+
; Gao et al.,
|
| 85 |
+
2022
|
| 86 |
+
; Shinn et al.,
|
| 87 |
+
2023
|
| 88 |
+
)
|
| 89 |
+
. This eliminates the need to rely entirely on the base abilities of the Language Model (LM), enhancing it through external tools or semantic feedback. Despite this strength, these methods are reflexive and fall short of humans’ deliberate and thoughtful decision-making characteristics to solve problems
|
| 90 |
+
(Sloman,
|
| 91 |
+
1996
|
| 92 |
+
; Evans,
|
| 93 |
+
2010
|
| 94 |
+
)
|
| 95 |
+
. In particular, such methods fail to consider multiple reasoning paths or to plan ahead. Recent search-guided LLM works
|
| 96 |
+
(Xie et al.,
|
| 97 |
+
2023
|
| 98 |
+
; Yao et al.,
|
| 99 |
+
2023a
|
| 100 |
+
; Hao et al.,
|
| 101 |
+
2023
|
| 102 |
+
)
|
| 103 |
+
address this issue by searching over multiple reasoning chains. While these methods enable planning, these methods operate in isolation and do not incorporate external feedback that can improve reasoning.
|
| 104 |
+
To help address these issues, we propose LATS (Language Agent Tree Search), a general framework for decision-making and reasoning with language models. LATS unifies LM planning, acting, and reasoning strategies by expanding ReAct
|
| 105 |
+
(Yao et al.,
|
| 106 |
+
2023b
|
| 107 |
+
)
|
| 108 |
+
into a search over a combinatorial space of possible reasoning and acting steps. We adapt Monte Carlo tree search (MCTS) from model-based reinforcement learning
|
| 109 |
+
(Silver et al.,
|
| 110 |
+
2017
|
| 111 |
+
; Anthony et al.,
|
| 112 |
+
2017
|
| 113 |
+
; Jiang et al.,
|
| 114 |
+
2018
|
| 115 |
+
)
|
| 116 |
+
to language agents, repurposing a pretrained LLM as an agent, value function, and optimizer. Utilizing the strong natural language understanding and in-context learning ability of modern LMs, we use text as an interface between each component of the framework, allowing LATS to adapt planning to environmental conditions without additional training. To the best of our knowledge,
|
| 117 |
+
LATS is the first framework that combines reasoning, acting, and planning to enhance LLMs
|
| 118 |
+
. Notably, LATS doubles the performance of GPT-3.5 on HotPotQA
|
| 119 |
+
(Yang et al.,
|
| 120 |
+
2018
|
| 121 |
+
)
|
| 122 |
+
over ReAct
|
| 123 |
+
(Yao et al.,
|
| 124 |
+
2023b
|
| 125 |
+
)
|
| 126 |
+
and raises the average score by
|
| 127 |
+
22.1
|
| 128 |
+
22.1
|
| 129 |
+
22.1
|
| 130 |
+
on WebShop
|
| 131 |
+
(Yao et al.,
|
| 132 |
+
2022
|
| 133 |
+
)
|
| 134 |
+
. When used with GPT-4, LATS achieves a
|
| 135 |
+
94.4
|
| 136 |
+
94.4
|
| 137 |
+
94.4
|
| 138 |
+
Pass@1 rate for programming on HumanEval
|
| 139 |
+
(Chen et al.,
|
| 140 |
+
2021
|
| 141 |
+
)
|
| 142 |
+
, setting the state of the art. To summarize, our
|
| 143 |
+
contributions
|
| 144 |
+
are the following:
|
| 145 |
+
•
|
| 146 |
+
We introduce an LM-based Monte Carlo tree search variant to deliberately construct the best trajectory from sampled actions, enabling more flexible and adaptive problem-solving compared to reflexive prompting methods. This is guided by heuristics from the LM.
|
| 147 |
+
•
|
| 148 |
+
By integrating external feedback and self-reflection, LATS enhances model sensibility and enables agents to learn from experience, surpassing reasoning-based search methods.
|
| 149 |
+
•
|
| 150 |
+
Through experiments across diverse domains like programming, interactive QA, and web navigation, we demonstrate the versatility of LATS in harnessing LLMs for autonomous reasoning and decision-making.
|
| 151 |
+
2
|
| 152 |
+
Related Work
|
| 153 |
+
Approach
|
| 154 |
+
Reasoning
|
| 155 |
+
Acting
|
| 156 |
+
Planning
|
| 157 |
+
Self
|
| 158 |
+
External
|
| 159 |
+
Reflection
|
| 160 |
+
Memory
|
| 161 |
+
CoT
|
| 162 |
+
(Wei et al.,
|
| 163 |
+
2022
|
| 164 |
+
)
|
| 165 |
+
✓
|
| 166 |
+
×
|
| 167 |
+
\times
|
| 168 |
+
×
|
| 169 |
+
\times
|
| 170 |
+
×
|
| 171 |
+
\times
|
| 172 |
+
×
|
| 173 |
+
\times
|
| 174 |
+
ReAct
|
| 175 |
+
(Yao et al.,
|
| 176 |
+
2023b
|
| 177 |
+
)
|
| 178 |
+
✓
|
| 179 |
+
✓
|
| 180 |
+
×
|
| 181 |
+
\times
|
| 182 |
+
×
|
| 183 |
+
\times
|
| 184 |
+
×
|
| 185 |
+
\times
|
| 186 |
+
ToT
|
| 187 |
+
(Yao et al.,
|
| 188 |
+
2023a
|
| 189 |
+
)
|
| 190 |
+
✓
|
| 191 |
+
×
|
| 192 |
+
\times
|
| 193 |
+
✓
|
| 194 |
+
✓
|
| 195 |
+
✓
|
| 196 |
+
RAP
|
| 197 |
+
(Hao et al.,
|
| 198 |
+
2023
|
| 199 |
+
)
|
| 200 |
+
✓
|
| 201 |
+
×
|
| 202 |
+
\times
|
| 203 |
+
✓
|
| 204 |
+
×
|
| 205 |
+
\times
|
| 206 |
+
✓
|
| 207 |
+
Self-Refine
|
| 208 |
+
(Madaan et al.,
|
| 209 |
+
2023
|
| 210 |
+
)
|
| 211 |
+
✓
|
| 212 |
+
×
|
| 213 |
+
\times
|
| 214 |
+
×
|
| 215 |
+
\times
|
| 216 |
+
✓
|
| 217 |
+
×
|
| 218 |
+
\times
|
| 219 |
+
Beam Search
|
| 220 |
+
(Xie et al.,
|
| 221 |
+
2023
|
| 222 |
+
)
|
| 223 |
+
✓
|
| 224 |
+
×
|
| 225 |
+
\times
|
| 226 |
+
×
|
| 227 |
+
\times
|
| 228 |
+
✓
|
| 229 |
+
×
|
| 230 |
+
\times
|
| 231 |
+
Reflexion
|
| 232 |
+
(Shinn et al.,
|
| 233 |
+
2023
|
| 234 |
+
)
|
| 235 |
+
✓
|
| 236 |
+
✓
|
| 237 |
+
×
|
| 238 |
+
\times
|
| 239 |
+
✓
|
| 240 |
+
✓
|
| 241 |
+
LATS (Ours)
|
| 242 |
+
✓
|
| 243 |
+
✓
|
| 244 |
+
✓
|
| 245 |
+
✓
|
| 246 |
+
✓
|
| 247 |
+
Table 1:
|
| 248 |
+
A summary of related work on reasoning, acting, and planning. LATS is the first work incorporating designs from all three domains, allowing use in all corresponding tasks. We refer to planning as the use of a search algorithm, self-reflection as the use of LM-generated feedback, and external memory as storaging past text context for future updates of solution.
|
| 249 |
+
a) Tree-of-Thoughts
|
| 250 |
+
b) Reasoning via Planning
|
| 251 |
+
c) Language Agent Tree Search
|
| 252 |
+
Figure 2:
|
| 253 |
+
An overview of the differences between LATS and recently proposed LM search algorithms ToT
|
| 254 |
+
(Yao et al.,
|
| 255 |
+
2023a
|
| 256 |
+
)
|
| 257 |
+
and RAP
|
| 258 |
+
(Hao et al.,
|
| 259 |
+
2023
|
| 260 |
+
)
|
| 261 |
+
. LATS leverages environmental feedback and self-reflection to further adapt search and improve performance.
|
| 262 |
+
LLMs for reasoning.
|
| 263 |
+
For LLMs, reasoning typically involves decomposing complex inputs into sequential intermediate steps towards a final answer
|
| 264 |
+
(Cobbe et al.,
|
| 265 |
+
2021
|
| 266 |
+
)
|
| 267 |
+
, demonstrated with Chain-of-Thought (CoT) prompting
|
| 268 |
+
(Wei et al.,
|
| 269 |
+
2022
|
| 270 |
+
)
|
| 271 |
+
and its variants
|
| 272 |
+
(Wei et al.,
|
| 273 |
+
2022
|
| 274 |
+
; Kojima et al.,
|
| 275 |
+
2022
|
| 276 |
+
; Wang et al.,
|
| 277 |
+
2022
|
| 278 |
+
)
|
| 279 |
+
. However, these methods, which create chains autoregressively in a single step, often suffer from error propagation as the number of steps increases
|
| 280 |
+
(Guo et al.,
|
| 281 |
+
2018
|
| 282 |
+
; Chen et al.,
|
| 283 |
+
2022b
|
| 284 |
+
)
|
| 285 |
+
due to compound errors. Various advancements aim to mitigate this issue; some approaches, such as Self-Consistency
|
| 286 |
+
(Wang et al.,
|
| 287 |
+
2022
|
| 288 |
+
)
|
| 289 |
+
, employ majority voting over sampled chains, while others focus on multi-step decomposition, such as least-to-most prompting
|
| 290 |
+
(Zhou et al.,
|
| 291 |
+
2022
|
| 292 |
+
)
|
| 293 |
+
, or use of external tools such as a scratchpad
|
| 294 |
+
(Nye et al.,
|
| 295 |
+
2021
|
| 296 |
+
)
|
| 297 |
+
or compiler
|
| 298 |
+
(Gao et al.,
|
| 299 |
+
2022
|
| 300 |
+
)
|
| 301 |
+
. Recently, CoT has been improved with search algorithms
|
| 302 |
+
(Yao et al.,
|
| 303 |
+
2023a
|
| 304 |
+
; Hao et al.,
|
| 305 |
+
2023
|
| 306 |
+
; Besta et al.,
|
| 307 |
+
2023
|
| 308 |
+
)
|
| 309 |
+
that can sample trajectories more effectively. Tree-of-thought (ToT) prompting
|
| 310 |
+
(Yao et al.,
|
| 311 |
+
2023a
|
| 312 |
+
)
|
| 313 |
+
uses DFS or BFS-based search guided by an LM-generated heuristic while Reasoning via Planning (RAP)
|
| 314 |
+
(Hao et al.,
|
| 315 |
+
2023
|
| 316 |
+
)
|
| 317 |
+
uses MCTS with rollouts simulated by the LM. However, they rely solely on LM internal knowledge and cannot adapt to useful external feedback.
|
| 318 |
+
LLMs for acting.
|
| 319 |
+
The strong reasoning and common-sense abilities of LLMs have also been adapted for decision-making or acting tasks as a policy model in interactive environments. In the realm of robotics LLMs have been employed as high-level controllers of control policies
|
| 320 |
+
(Ahn et al.,
|
| 321 |
+
2022
|
| 322 |
+
; Huang et al.,
|
| 323 |
+
2022
|
| 324 |
+
; Driess et al.,
|
| 325 |
+
2023
|
| 326 |
+
)
|
| 327 |
+
. Similar work
|
| 328 |
+
(Baker et al.,
|
| 329 |
+
2022
|
| 330 |
+
; Wang et al.,
|
| 331 |
+
2023
|
| 332 |
+
; Zhu et al.,
|
| 333 |
+
2023
|
| 334 |
+
)
|
| 335 |
+
has also adapted LLM agents to complex multimodal games such as Minecraft
|
| 336 |
+
(Guss et al.,
|
| 337 |
+
2019
|
| 338 |
+
; Fan et al.,
|
| 339 |
+
2022
|
| 340 |
+
)
|
| 341 |
+
. LLMs are particularly useful in text-based environments
|
| 342 |
+
(Liu et al.,
|
| 343 |
+
2018
|
| 344 |
+
; Shridhar et al.,
|
| 345 |
+
2020
|
| 346 |
+
; Liu et al.,
|
| 347 |
+
2023
|
| 348 |
+
)
|
| 349 |
+
, where acting-based prompting techniques such as ReAct
|
| 350 |
+
(Yao et al.,
|
| 351 |
+
2023b
|
| 352 |
+
)
|
| 353 |
+
have seen success. Similar to CoT, ReAct is limited by its simplicity and cannot effectively adapt to environment conditions. Many extensions have been proposed to address this, including Self-refine
|
| 354 |
+
(Madaan et al.,
|
| 355 |
+
2023
|
| 356 |
+
)
|
| 357 |
+
and Reflexion
|
| 358 |
+
(Shinn et al.,
|
| 359 |
+
2023
|
| 360 |
+
; Yao et al.,
|
| 361 |
+
2023c
|
| 362 |
+
)
|
| 363 |
+
, which uses self-reflection to enhance reasoning and decision-making, and AdaPlanner
|
| 364 |
+
(Sun et al.,
|
| 365 |
+
2023
|
| 366 |
+
)
|
| 367 |
+
, which incorporates both positive and negative environmental feedback. However these methods focus on refining an individual plan or trajectory and do not consider alternative choices at each step. In addition, recent work
|
| 368 |
+
(Huang et al.,
|
| 369 |
+
2023
|
| 370 |
+
)
|
| 371 |
+
has suggested LLMs cannot self-correct their internal reasoning, making it critical to use external feedback. Alternatively to pure decision-making environments, the reasoning and practical abilities of LLMs have been enhanced by access to external tools, such as APIs, search engines, calculators, or other models
|
| 372 |
+
(Schick et al.,
|
| 373 |
+
2023
|
| 374 |
+
; Shen et al.,
|
| 375 |
+
2023
|
| 376 |
+
; Surís et al.,
|
| 377 |
+
2023
|
| 378 |
+
)
|
| 379 |
+
. Contrary to reasoning-based approaches, these methods have not been improved with planning, limiting their effectiveness. We summarize them in Tab.
|
| 380 |
+
1
|
| 381 |
+
.
|
| 382 |
+
Tree-based search.
|
| 383 |
+
Tree-based search, where multiple branches of outcomes are explored during search, is widely used in many planning algorithms
|
| 384 |
+
(Świechowski et al.,
|
| 385 |
+
2023
|
| 386 |
+
; LaValle et al.,
|
| 387 |
+
2001
|
| 388 |
+
)
|
| 389 |
+
and Reinforcement Learning (RL)
|
| 390 |
+
(Hafner et al.,
|
| 391 |
+
2019
|
| 392 |
+
; Du et al.,
|
| 393 |
+
2023
|
| 394 |
+
; Wu et al.,
|
| 395 |
+
2023
|
| 396 |
+
)
|
| 397 |
+
algorithms for its good exploration-exploitation trade-off. Though tree-based search requires an environment model that can expand from arbitrary state
|
| 398 |
+
(Vodopivec et al.,
|
| 399 |
+
2017
|
| 400 |
+
)
|
| 401 |
+
, which often requires extra training in RL
|
| 402 |
+
(Hafner et al.,
|
| 403 |
+
2023
|
| 404 |
+
)
|
| 405 |
+
, such problem does not exist for LM tasks as we can conveniently backup to any state by setting the input to be the context and corresponding previous output by the LM. Thus, we work on the tree-based framework and use MCTS
|
| 406 |
+
(Świechowski et al.,
|
| 407 |
+
2023
|
| 408 |
+
)
|
| 409 |
+
to fully release the potential of LMs, while avoiding the cost of training a value function over language descriptions by leveraging the in-context learning
|
| 410 |
+
(Brown et al.,
|
| 411 |
+
2020
|
| 412 |
+
)
|
| 413 |
+
abilities of LLMs.
|
| 414 |
+
3
|
| 415 |
+
Preliminaries
|
| 416 |
+
3.1
|
| 417 |
+
Problem Setting and Prompting
|
| 418 |
+
Before describing LATS, we first define our problem and outline a few established methods that leverage large language models for reasoning or decision-making. In LM reasoning or decision making, we are given an input
|
| 419 |
+
x
|
| 420 |
+
𝑥
|
| 421 |
+
x
|
| 422 |
+
in natural language and a pretrained language model
|
| 423 |
+
p
|
| 424 |
+
θ
|
| 425 |
+
|
| 426 |
+
(
|
| 427 |
+
x
|
| 428 |
+
)
|
| 429 |
+
subscript
|
| 430 |
+
𝑝
|
| 431 |
+
𝜃
|
| 432 |
+
𝑥
|
| 433 |
+
p_{\theta}(x)
|
| 434 |
+
parameterized by
|
| 435 |
+
θ
|
| 436 |
+
𝜃
|
| 437 |
+
\theta
|
| 438 |
+
; our goal is to generate a final output
|
| 439 |
+
y
|
| 440 |
+
∼
|
| 441 |
+
p
|
| 442 |
+
θ
|
| 443 |
+
|
| 444 |
+
(
|
| 445 |
+
x
|
| 446 |
+
)
|
| 447 |
+
similar-to
|
| 448 |
+
𝑦
|
| 449 |
+
subscript
|
| 450 |
+
𝑝
|
| 451 |
+
𝜃
|
| 452 |
+
𝑥
|
| 453 |
+
y\sim p_{\theta}(x)
|
| 454 |
+
corresponding to the answer (reasoning) or completes the task (decision-making). Both
|
| 455 |
+
x
|
| 456 |
+
𝑥
|
| 457 |
+
x
|
| 458 |
+
and
|
| 459 |
+
y
|
| 460 |
+
𝑦
|
| 461 |
+
y
|
| 462 |
+
are language
|
| 463 |
+
sequences
|
| 464 |
+
, which are comprised of a list of
|
| 465 |
+
tokens
|
| 466 |
+
(the basic elements of natural language, often words), denoted as
|
| 467 |
+
x
|
| 468 |
+
=
|
| 469 |
+
(
|
| 470 |
+
x
|
| 471 |
+
|
| 472 |
+
[
|
| 473 |
+
1
|
| 474 |
+
]
|
| 475 |
+
,
|
| 476 |
+
…
|
| 477 |
+
,
|
| 478 |
+
x
|
| 479 |
+
|
| 480 |
+
[
|
| 481 |
+
n
|
| 482 |
+
]
|
| 483 |
+
)
|
| 484 |
+
𝑥
|
| 485 |
+
𝑥
|
| 486 |
+
delimited-[]
|
| 487 |
+
1
|
| 488 |
+
…
|
| 489 |
+
𝑥
|
| 490 |
+
delimited-[]
|
| 491 |
+
𝑛
|
| 492 |
+
x=(x[1],\dots,x[n])
|
| 493 |
+
and
|
| 494 |
+
y
|
| 495 |
+
=
|
| 496 |
+
(
|
| 497 |
+
y
|
| 498 |
+
|
| 499 |
+
[
|
| 500 |
+
1
|
| 501 |
+
]
|
| 502 |
+
,
|
| 503 |
+
…
|
| 504 |
+
,
|
| 505 |
+
y
|
| 506 |
+
|
| 507 |
+
[
|
| 508 |
+
n
|
| 509 |
+
]
|
| 510 |
+
)
|
| 511 |
+
𝑦
|
| 512 |
+
𝑦
|
| 513 |
+
delimited-[]
|
| 514 |
+
1
|
| 515 |
+
…
|
| 516 |
+
𝑦
|
| 517 |
+
delimited-[]
|
| 518 |
+
𝑛
|
| 519 |
+
y=(y[1],\dots,y[n])
|
| 520 |
+
. The LM decodes text autoregressively, i.e., without other inputs, the probability for an LM to generate a sequence
|
| 521 |
+
x
|
| 522 |
+
𝑥
|
| 523 |
+
x
|
| 524 |
+
is given by
|
| 525 |
+
p
|
| 526 |
+
θ
|
| 527 |
+
|
| 528 |
+
(
|
| 529 |
+
x
|
| 530 |
+
)
|
| 531 |
+
=
|
| 532 |
+
∏
|
| 533 |
+
i
|
| 534 |
+
=
|
| 535 |
+
1
|
| 536 |
+
n
|
| 537 |
+
p
|
| 538 |
+
θ
|
| 539 |
+
|
| 540 |
+
(
|
| 541 |
+
x
|
| 542 |
+
|
| 543 |
+
[
|
| 544 |
+
i
|
| 545 |
+
]
|
| 546 |
+
|
|
| 547 |
+
x
|
| 548 |
+
|
| 549 |
+
[
|
| 550 |
+
1
|
| 551 |
+
|
| 552 |
+
…
|
| 553 |
+
|
| 554 |
+
i
|
| 555 |
+
−
|
| 556 |
+
1
|
| 557 |
+
]
|
| 558 |
+
)
|
| 559 |
+
subscript
|
| 560 |
+
𝑝
|
| 561 |
+
𝜃
|
| 562 |
+
𝑥
|
| 563 |
+
superscript
|
| 564 |
+
subscript
|
| 565 |
+
product
|
| 566 |
+
𝑖
|
| 567 |
+
1
|
| 568 |
+
𝑛
|
| 569 |
+
subscript
|
| 570 |
+
𝑝
|
| 571 |
+
𝜃
|
| 572 |
+
conditional
|
| 573 |
+
𝑥
|
| 574 |
+
delimited-[]
|
| 575 |
+
𝑖
|
| 576 |
+
𝑥
|
| 577 |
+
delimited-[]
|
| 578 |
+
1
|
| 579 |
+
…
|
| 580 |
+
𝑖
|
| 581 |
+
1
|
| 582 |
+
p_{\theta}(x)=\prod_{i=1}^{n}p_{\theta}(x[i]|x[1\dots i-1])
|
| 583 |
+
. Usually, to improve the LM,
|
| 584 |
+
prompts
|
| 585 |
+
are provided along with the input
|
| 586 |
+
x
|
| 587 |
+
𝑥
|
| 588 |
+
x
|
| 589 |
+
, which are specific instructions or few-shot input-output examples. We denote the generic process where an input
|
| 590 |
+
x
|
| 591 |
+
𝑥
|
| 592 |
+
x
|
| 593 |
+
is transformed into an output
|
| 594 |
+
y
|
| 595 |
+
𝑦
|
| 596 |
+
y
|
| 597 |
+
by LM:
|
| 598 |
+
y
|
| 599 |
+
∼
|
| 600 |
+
p
|
| 601 |
+
θ
|
| 602 |
+
|
| 603 |
+
(
|
| 604 |
+
y
|
| 605 |
+
|
|
| 606 |
+
prompt
|
| 607 |
+
I
|
| 608 |
+
|
| 609 |
+
O
|
| 610 |
+
|
| 611 |
+
(
|
| 612 |
+
x
|
| 613 |
+
)
|
| 614 |
+
)
|
| 615 |
+
similar-to
|
| 616 |
+
𝑦
|
| 617 |
+
subscript
|
| 618 |
+
𝑝
|
| 619 |
+
𝜃
|
| 620 |
+
conditional
|
| 621 |
+
𝑦
|
| 622 |
+
subscript
|
| 623 |
+
prompt
|
| 624 |
+
𝐼
|
| 625 |
+
𝑂
|
| 626 |
+
𝑥
|
| 627 |
+
y\sim p_{\theta}(y|\texttt{prompt}_{IO}(x))
|
| 628 |
+
, where
|
| 629 |
+
prompt
|
| 630 |
+
I
|
| 631 |
+
|
| 632 |
+
O
|
| 633 |
+
|
| 634 |
+
(
|
| 635 |
+
x
|
| 636 |
+
)
|
| 637 |
+
subscript
|
| 638 |
+
prompt
|
| 639 |
+
𝐼
|
| 640 |
+
𝑂
|
| 641 |
+
𝑥
|
| 642 |
+
\texttt{prompt}_{IO}(x)
|
| 643 |
+
denotes the input
|
| 644 |
+
x
|
| 645 |
+
𝑥
|
| 646 |
+
x
|
| 647 |
+
.
|
| 648 |
+
Chain-of-thought (CoT) Prompting
|
| 649 |
+
(Wei et al.,
|
| 650 |
+
2022
|
| 651 |
+
)
|
| 652 |
+
was introduced to cater to scenarios where direct mapping from
|
| 653 |
+
x
|
| 654 |
+
𝑥
|
| 655 |
+
x
|
| 656 |
+
to
|
| 657 |
+
y
|
| 658 |
+
𝑦
|
| 659 |
+
y
|
| 660 |
+
is intricate, such as when
|
| 661 |
+
x
|
| 662 |
+
𝑥
|
| 663 |
+
x
|
| 664 |
+
is from a mathematical query or challenging question. This method hinges on creating
|
| 665 |
+
thoughts
|
| 666 |
+
z
|
| 667 |
+
1
|
| 668 |
+
,
|
| 669 |
+
…
|
| 670 |
+
,
|
| 671 |
+
z
|
| 672 |
+
n
|
| 673 |
+
subscript
|
| 674 |
+
𝑧
|
| 675 |
+
1
|
| 676 |
+
…
|
| 677 |
+
subscript
|
| 678 |
+
𝑧
|
| 679 |
+
𝑛
|
| 680 |
+
z_{1},\dots,z_{n}
|
| 681 |
+
that act as stepping stones between
|
| 682 |
+
x
|
| 683 |
+
𝑥
|
| 684 |
+
x
|
| 685 |
+
and
|
| 686 |
+
y
|
| 687 |
+
𝑦
|
| 688 |
+
y
|
| 689 |
+
; each thought
|
| 690 |
+
z
|
| 691 |
+
i
|
| 692 |
+
subscript
|
| 693 |
+
𝑧
|
| 694 |
+
𝑖
|
| 695 |
+
z_{i}
|
| 696 |
+
is a language sequence. To employ CoT prompting, thoughts are extracted sequentially as
|
| 697 |
+
z
|
| 698 |
+
i
|
| 699 |
+
∼
|
| 700 |
+
p
|
| 701 |
+
θ
|
| 702 |
+
C
|
| 703 |
+
|
| 704 |
+
o
|
| 705 |
+
|
| 706 |
+
T
|
| 707 |
+
|
| 708 |
+
(
|
| 709 |
+
z
|
| 710 |
+
i
|
| 711 |
+
|
|
| 712 |
+
x
|
| 713 |
+
,
|
| 714 |
+
z
|
| 715 |
+
1
|
| 716 |
+
|
| 717 |
+
⋯
|
| 718 |
+
|
| 719 |
+
i
|
| 720 |
+
−
|
| 721 |
+
1
|
| 722 |
+
)
|
| 723 |
+
similar-to
|
| 724 |
+
subscript
|
| 725 |
+
𝑧
|
| 726 |
+
𝑖
|
| 727 |
+
superscript
|
| 728 |
+
subscript
|
| 729 |
+
𝑝
|
| 730 |
+
𝜃
|
| 731 |
+
𝐶
|
| 732 |
+
𝑜
|
| 733 |
+
𝑇
|
| 734 |
+
conditional
|
| 735 |
+
subscript
|
| 736 |
+
𝑧
|
| 737 |
+
𝑖
|
| 738 |
+
𝑥
|
| 739 |
+
subscript
|
| 740 |
+
𝑧
|
| 741 |
+
1
|
| 742 |
+
⋯
|
| 743 |
+
𝑖
|
| 744 |
+
1
|
| 745 |
+
z_{i}\sim p_{\theta}^{CoT}(z_{i}|x,z_{1\cdots i-1})
|
| 746 |
+
, with the final output being
|
| 747 |
+
y
|
| 748 |
+
∼
|
| 749 |
+
p
|
| 750 |
+
θ
|
| 751 |
+
C
|
| 752 |
+
|
| 753 |
+
o
|
| 754 |
+
|
| 755 |
+
T
|
| 756 |
+
|
| 757 |
+
(
|
| 758 |
+
y
|
| 759 |
+
|
|
| 760 |
+
x
|
| 761 |
+
,
|
| 762 |
+
z
|
| 763 |
+
1
|
| 764 |
+
|
| 765 |
+
⋯
|
| 766 |
+
|
| 767 |
+
n
|
| 768 |
+
)
|
| 769 |
+
similar-to
|
| 770 |
+
𝑦
|
| 771 |
+
superscript
|
| 772 |
+
subscript
|
| 773 |
+
𝑝
|
| 774 |
+
𝜃
|
| 775 |
+
𝐶
|
| 776 |
+
𝑜
|
| 777 |
+
𝑇
|
| 778 |
+
conditional
|
| 779 |
+
𝑦
|
| 780 |
+
𝑥
|
| 781 |
+
subscript
|
| 782 |
+
𝑧
|
| 783 |
+
1
|
| 784 |
+
⋯
|
| 785 |
+
𝑛
|
| 786 |
+
y\sim p_{\theta}^{CoT}(y|x,z_{1\cdots n})
|
| 787 |
+
.
|
| 788 |
+
Tree-of-thought (ToT) Prompting
|
| 789 |
+
(Yao et al.,
|
| 790 |
+
2023a
|
| 791 |
+
)
|
| 792 |
+
extends CoT prompting by exploring multiple reasoning paths over thoughts. It frames problems as a search over a tree where each node
|
| 793 |
+
s
|
| 794 |
+
=
|
| 795 |
+
[
|
| 796 |
+
x
|
| 797 |
+
,
|
| 798 |
+
z
|
| 799 |
+
1
|
| 800 |
+
⋅
|
| 801 |
+
i
|
| 802 |
+
]
|
| 803 |
+
𝑠
|
| 804 |
+
𝑥
|
| 805 |
+
subscript
|
| 806 |
+
𝑧
|
| 807 |
+
⋅
|
| 808 |
+
1
|
| 809 |
+
𝑖
|
| 810 |
+
s=[x,z_{1\cdot i}]
|
| 811 |
+
represents a partial solution state comprising the original input
|
| 812 |
+
x
|
| 813 |
+
𝑥
|
| 814 |
+
x
|
| 815 |
+
and thought sequence
|
| 816 |
+
z
|
| 817 |
+
1
|
| 818 |
+
|
| 819 |
+
⋯
|
| 820 |
+
|
| 821 |
+
i
|
| 822 |
+
subscript
|
| 823 |
+
𝑧
|
| 824 |
+
1
|
| 825 |
+
⋯
|
| 826 |
+
𝑖
|
| 827 |
+
z_{1\cdots i}
|
| 828 |
+
. Thoughts
|
| 829 |
+
z
|
| 830 |
+
i
|
| 831 |
+
subscript
|
| 832 |
+
𝑧
|
| 833 |
+
𝑖
|
| 834 |
+
z_{i}
|
| 835 |
+
are generated by proposal or sampling with CoT
|
| 836 |
+
z
|
| 837 |
+
i
|
| 838 |
+
∼
|
| 839 |
+
p
|
| 840 |
+
θ
|
| 841 |
+
C
|
| 842 |
+
|
| 843 |
+
o
|
| 844 |
+
|
| 845 |
+
T
|
| 846 |
+
|
| 847 |
+
(
|
| 848 |
+
z
|
| 849 |
+
i
|
| 850 |
+
|
|
| 851 |
+
x
|
| 852 |
+
,
|
| 853 |
+
z
|
| 854 |
+
1
|
| 855 |
+
|
| 856 |
+
⋯
|
| 857 |
+
|
| 858 |
+
i
|
| 859 |
+
−
|
| 860 |
+
1
|
| 861 |
+
)
|
| 862 |
+
similar-to
|
| 863 |
+
subscript
|
| 864 |
+
𝑧
|
| 865 |
+
𝑖
|
| 866 |
+
superscript
|
| 867 |
+
subscript
|
| 868 |
+
𝑝
|
| 869 |
+
𝜃
|
| 870 |
+
𝐶
|
| 871 |
+
𝑜
|
| 872 |
+
𝑇
|
| 873 |
+
conditional
|
| 874 |
+
subscript
|
| 875 |
+
𝑧
|
| 876 |
+
𝑖
|
| 877 |
+
𝑥
|
| 878 |
+
subscript
|
| 879 |
+
𝑧
|
| 880 |
+
1
|
| 881 |
+
⋯
|
| 882 |
+
𝑖
|
| 883 |
+
1
|
| 884 |
+
z_{i}\sim p_{\theta}^{CoT}(z_{i}|x,z_{1\cdots i-1})
|
| 885 |
+
. Deliberate search algorithms like breadth-first or depth-first search are used to systematically explore the tree, guided by heuristics based on language model evaluations
|
| 886 |
+
V
|
| 887 |
+
|
| 888 |
+
(
|
| 889 |
+
s
|
| 890 |
+
)
|
| 891 |
+
𝑉
|
| 892 |
+
𝑠
|
| 893 |
+
V(s)
|
| 894 |
+
of each state.
|
| 895 |
+
Reasoning via Planning
|
| 896 |
+
(RAP)
|
| 897 |
+
(Hao et al.,
|
| 898 |
+
2023
|
| 899 |
+
)
|
| 900 |
+
is similar to ToT, except that MCTS is used over DFS or BFS. Heuristics are designed from an LM, such as the likelihood or confidence of an action, and the LM is used as a world model to predict subsequent states during the simulation step.
|
| 901 |
+
ReAct
|
| 902 |
+
(Yao et al.,
|
| 903 |
+
2023b
|
| 904 |
+
)
|
| 905 |
+
extends language models to tasks where the mapping from
|
| 906 |
+
x
|
| 907 |
+
𝑥
|
| 908 |
+
x
|
| 909 |
+
to
|
| 910 |
+
y
|
| 911 |
+
𝑦
|
| 912 |
+
y
|
| 913 |
+
is enhanced by or requires interactions with an external environment, such as a game or API. This technique constructs an action space
|
| 914 |
+
A
|
| 915 |
+
^
|
| 916 |
+
=
|
| 917 |
+
A
|
| 918 |
+
∪
|
| 919 |
+
Z
|
| 920 |
+
^
|
| 921 |
+
𝐴
|
| 922 |
+
𝐴
|
| 923 |
+
𝑍
|
| 924 |
+
\hat{A}=A\cup Z
|
| 925 |
+
that adds permissible actions
|
| 926 |
+
a
|
| 927 |
+
𝑎
|
| 928 |
+
a
|
| 929 |
+
to the reasoning traces
|
| 930 |
+
z
|
| 931 |
+
𝑧
|
| 932 |
+
z
|
| 933 |
+
from CoT. Observations
|
| 934 |
+
o
|
| 935 |
+
𝑜
|
| 936 |
+
o
|
| 937 |
+
from the environment are used to improve both reasoning and acting. To solve problems with ReAct, after each observation, actions are generated from
|
| 938 |
+
p
|
| 939 |
+
θ
|
| 940 |
+
subscript
|
| 941 |
+
𝑝
|
| 942 |
+
𝜃
|
| 943 |
+
p_{\theta}
|
| 944 |
+
sequentially as
|
| 945 |
+
a
|
| 946 |
+
i
|
| 947 |
+
∼
|
| 948 |
+
p
|
| 949 |
+
θ
|
| 950 |
+
R
|
| 951 |
+
|
| 952 |
+
e
|
| 953 |
+
|
| 954 |
+
A
|
| 955 |
+
|
| 956 |
+
c
|
| 957 |
+
|
| 958 |
+
t
|
| 959 |
+
|
| 960 |
+
(
|
| 961 |
+
a
|
| 962 |
+
i
|
| 963 |
+
|
|
| 964 |
+
x
|
| 965 |
+
,
|
| 966 |
+
o
|
| 967 |
+
1
|
| 968 |
+
|
| 969 |
+
⋯
|
| 970 |
+
|
| 971 |
+
i
|
| 972 |
+
−
|
| 973 |
+
1
|
| 974 |
+
,
|
| 975 |
+
a
|
| 976 |
+
1
|
| 977 |
+
|
| 978 |
+
⋯
|
| 979 |
+
|
| 980 |
+
i
|
| 981 |
+
−
|
| 982 |
+
1
|
| 983 |
+
)
|
| 984 |
+
similar-to
|
| 985 |
+
subscript
|
| 986 |
+
𝑎
|
| 987 |
+
𝑖
|
| 988 |
+
superscript
|
| 989 |
+
subscript
|
| 990 |
+
𝑝
|
| 991 |
+
𝜃
|
| 992 |
+
𝑅
|
| 993 |
+
𝑒
|
| 994 |
+
𝐴
|
| 995 |
+
𝑐
|
| 996 |
+
𝑡
|
| 997 |
+
conditional
|
| 998 |
+
subscript
|
| 999 |
+
𝑎
|
| 1000 |
+
𝑖
|
| 1001 |
+
𝑥
|
| 1002 |
+
subscript
|
| 1003 |
+
𝑜
|
| 1004 |
+
1
|
| 1005 |
+
⋯
|
| 1006 |
+
𝑖
|
| 1007 |
+
1
|
| 1008 |
+
subscript
|
| 1009 |
+
𝑎
|
| 1010 |
+
1
|
| 1011 |
+
⋯
|
| 1012 |
+
𝑖
|
| 1013 |
+
1
|
| 1014 |
+
a_{i}\sim p_{\theta}^{ReAct}(a_{i}|x,o_{1\cdots i-1},a_{1\cdots i-1})
|
| 1015 |
+
, with the final output being
|
| 1016 |
+
y
|
| 1017 |
+
∼
|
| 1018 |
+
p
|
| 1019 |
+
θ
|
| 1020 |
+
R
|
| 1021 |
+
|
| 1022 |
+
e
|
| 1023 |
+
|
| 1024 |
+
A
|
| 1025 |
+
|
| 1026 |
+
c
|
| 1027 |
+
|
| 1028 |
+
t
|
| 1029 |
+
|
| 1030 |
+
(
|
| 1031 |
+
y
|
| 1032 |
+
|
|
| 1033 |
+
x
|
| 1034 |
+
,
|
| 1035 |
+
o
|
| 1036 |
+
1
|
| 1037 |
+
|
| 1038 |
+
⋯
|
| 1039 |
+
|
| 1040 |
+
n
|
| 1041 |
+
,
|
| 1042 |
+
a
|
| 1043 |
+
1
|
| 1044 |
+
|
| 1045 |
+
⋯
|
| 1046 |
+
|
| 1047 |
+
n
|
| 1048 |
+
)
|
| 1049 |
+
similar-to
|
| 1050 |
+
𝑦
|
| 1051 |
+
superscript
|
| 1052 |
+
subscript
|
| 1053 |
+
𝑝
|
| 1054 |
+
𝜃
|
| 1055 |
+
𝑅
|
| 1056 |
+
𝑒
|
| 1057 |
+
𝐴
|
| 1058 |
+
𝑐
|
| 1059 |
+
𝑡
|
| 1060 |
+
conditional
|
| 1061 |
+
𝑦
|
| 1062 |
+
𝑥
|
| 1063 |
+
subscript
|
| 1064 |
+
𝑜
|
| 1065 |
+
1
|
| 1066 |
+
⋯
|
| 1067 |
+
𝑛
|
| 1068 |
+
subscript
|
| 1069 |
+
𝑎
|
| 1070 |
+
1
|
| 1071 |
+
⋯
|
| 1072 |
+
𝑛
|
| 1073 |
+
y\sim p_{\theta}^{ReAct}(y~{}|~{}x,o_{1\cdots n},a_{1\cdots n})
|
| 1074 |
+
.
|
| 1075 |
+
While the previously described prompting techniques improve LM performance on reasoning tasks, they falter on difficult tasks that involve multifaceted decision-making due to several shortcomings: 1)
|
| 1076 |
+
Flexibility
|
| 1077 |
+
: Base prompting methods (CoT or ReAct) autoregressively sample from the LM, neglecting potential alternative continuations from specific states. 2)
|
| 1078 |
+
Sensibility
|
| 1079 |
+
: Reasoning-based methods (CoT, RAP, or ToT) rely solely on the internal representations of the LM and cannot consider external observations. This dependency risks fact hallucination and error propagation while setting a performance ceiling. 3)
|
| 1080 |
+
Adaptability
|
| 1081 |
+
: Current planning frameworks (RAP or ToT) use simple search algorithms such as BFS or cannot leverage environmental feedback to improve planning. Additionally, the agent is static and cannot reuse previous experience or learn from trial and error. While RAP also adopts MCTS, it is constrained to tasks where the LM can become a world model and accurately predict states. These shortcomings limit the ability of LMs to be deployed as general problem-solving agents and form the motivation for LATS.
|
| 1082 |
+
3.2
|
| 1083 |
+
Monte-Carlo Tree Search (MCTS)
|
| 1084 |
+
Monte-Carlo Tree Search (MCTS) is a heuristic search algorithm that is proved successful on many decision-making environments such as Atari
|
| 1085 |
+
(Ye et al.,
|
| 1086 |
+
2021
|
| 1087 |
+
)
|
| 1088 |
+
and Go
|
| 1089 |
+
(Silver et al.,
|
| 1090 |
+
2016
|
| 1091 |
+
)
|
| 1092 |
+
. MCTS builds a decision tree where every node in the tree is a state and edge is an action. MCTS runs for
|
| 1093 |
+
k
|
| 1094 |
+
𝑘
|
| 1095 |
+
k
|
| 1096 |
+
episodes; for each episode, it starts from the root (i.e., initial state) and iteratively conducts two steps to expand the tree: 1)
|
| 1097 |
+
Expansion
|
| 1098 |
+
, where multiple children states
|
| 1099 |
+
s
|
| 1100 |
+
𝑠
|
| 1101 |
+
s
|
| 1102 |
+
are explored from the current parent state
|
| 1103 |
+
p
|
| 1104 |
+
𝑝
|
| 1105 |
+
p
|
| 1106 |
+
by sampling
|
| 1107 |
+
n
|
| 1108 |
+
𝑛
|
| 1109 |
+
n
|
| 1110 |
+
actions, and 2)
|
| 1111 |
+
Selection
|
| 1112 |
+
, where the children with the highest UCT
|
| 1113 |
+
(Upper Confidence bounds applied to Trees)
|
| 1114 |
+
(Kocsis & Szepesvári,
|
| 1115 |
+
2006
|
| 1116 |
+
)
|
| 1117 |
+
value is selected by the next iteration. The UCT of a child state
|
| 1118 |
+
s
|
| 1119 |
+
𝑠
|
| 1120 |
+
s
|
| 1121 |
+
is calculated as follows:
|
| 1122 |
+
U
|
| 1123 |
+
|
| 1124 |
+
C
|
| 1125 |
+
|
| 1126 |
+
T
|
| 1127 |
+
|
| 1128 |
+
(
|
| 1129 |
+
s
|
| 1130 |
+
)
|
| 1131 |
+
=
|
| 1132 |
+
V
|
| 1133 |
+
|
| 1134 |
+
(
|
| 1135 |
+
s
|
| 1136 |
+
)
|
| 1137 |
+
+
|
| 1138 |
+
w
|
| 1139 |
+
|
| 1140 |
+
ln
|
| 1141 |
+
|
| 1142 |
+
N
|
| 1143 |
+
|
| 1144 |
+
(
|
| 1145 |
+
p
|
| 1146 |
+
)
|
| 1147 |
+
N
|
| 1148 |
+
|
| 1149 |
+
(
|
| 1150 |
+
s
|
| 1151 |
+
)
|
| 1152 |
+
,
|
| 1153 |
+
𝑈
|
| 1154 |
+
𝐶
|
| 1155 |
+
𝑇
|
| 1156 |
+
𝑠
|
| 1157 |
+
𝑉
|
| 1158 |
+
𝑠
|
| 1159 |
+
𝑤
|
| 1160 |
+
𝑁
|
| 1161 |
+
𝑝
|
| 1162 |
+
𝑁
|
| 1163 |
+
𝑠
|
| 1164 |
+
UCT(s)=V(s)+w\sqrt{\frac{\ln N(p)}{N(s)}},
|
| 1165 |
+
(1)
|
| 1166 |
+
where
|
| 1167 |
+
N
|
| 1168 |
+
|
| 1169 |
+
(
|
| 1170 |
+
s
|
| 1171 |
+
)
|
| 1172 |
+
𝑁
|
| 1173 |
+
𝑠
|
| 1174 |
+
N(s)
|
| 1175 |
+
is the number of visits to a node
|
| 1176 |
+
s
|
| 1177 |
+
𝑠
|
| 1178 |
+
s
|
| 1179 |
+
,
|
| 1180 |
+
V
|
| 1181 |
+
|
| 1182 |
+
(
|
| 1183 |
+
s
|
| 1184 |
+
)
|
| 1185 |
+
𝑉
|
| 1186 |
+
𝑠
|
| 1187 |
+
V(s)
|
| 1188 |
+
is the value function (expected return) from the subtree of
|
| 1189 |
+
s
|
| 1190 |
+
𝑠
|
| 1191 |
+
s
|
| 1192 |
+
,
|
| 1193 |
+
w
|
| 1194 |
+
𝑤
|
| 1195 |
+
w
|
| 1196 |
+
is the exploration weight, and
|
| 1197 |
+
p
|
| 1198 |
+
𝑝
|
| 1199 |
+
p
|
| 1200 |
+
is the parent node of
|
| 1201 |
+
s
|
| 1202 |
+
𝑠
|
| 1203 |
+
s
|
| 1204 |
+
. The child node with the highest UCT value is selected for expansion in the next iteration. When the end of an episode is reached, a
|
| 1205 |
+
backpropagation
|
| 1206 |
+
is carried out: the return
|
| 1207 |
+
r
|
| 1208 |
+
𝑟
|
| 1209 |
+
r
|
| 1210 |
+
is used for updating every
|
| 1211 |
+
V
|
| 1212 |
+
|
| 1213 |
+
(
|
| 1214 |
+
s
|
| 1215 |
+
)
|
| 1216 |
+
𝑉
|
| 1217 |
+
𝑠
|
| 1218 |
+
V(s)
|
| 1219 |
+
along the path
|
| 1220 |
+
with the formula
|
| 1221 |
+
V
|
| 1222 |
+
|
| 1223 |
+
(
|
| 1224 |
+
s
|
| 1225 |
+
)
|
| 1226 |
+
=
|
| 1227 |
+
V
|
| 1228 |
+
old
|
| 1229 |
+
|
| 1230 |
+
(
|
| 1231 |
+
s
|
| 1232 |
+
)
|
| 1233 |
+
|
| 1234 |
+
(
|
| 1235 |
+
N
|
| 1236 |
+
|
| 1237 |
+
(
|
| 1238 |
+
s
|
| 1239 |
+
)
|
| 1240 |
+
−
|
| 1241 |
+
1
|
| 1242 |
+
)
|
| 1243 |
+
+
|
| 1244 |
+
r
|
| 1245 |
+
N
|
| 1246 |
+
|
| 1247 |
+
(
|
| 1248 |
+
s
|
| 1249 |
+
)
|
| 1250 |
+
𝑉
|
| 1251 |
+
𝑠
|
| 1252 |
+
subscript
|
| 1253 |
+
𝑉
|
| 1254 |
+
old
|
| 1255 |
+
𝑠
|
| 1256 |
+
𝑁
|
| 1257 |
+
𝑠
|
| 1258 |
+
1
|
| 1259 |
+
𝑟
|
| 1260 |
+
𝑁
|
| 1261 |
+
𝑠
|
| 1262 |
+
V(s)=\frac{V_{\text{old}}(s)(N(s)-1)+r}{N(s)}
|
| 1263 |
+
, where
|
| 1264 |
+
V
|
| 1265 |
+
old
|
| 1266 |
+
|
| 1267 |
+
(
|
| 1268 |
+
s
|
| 1269 |
+
)
|
| 1270 |
+
subscript
|
| 1271 |
+
𝑉
|
| 1272 |
+
old
|
| 1273 |
+
𝑠
|
| 1274 |
+
V_{\text{old}}(s)
|
| 1275 |
+
is the old value function. Normally, the major shortcoming of MCTS is that it requires an environment model to undo previous steps and form a searching tree, which is often a strong assumption. However, such a limitation does not exist for LMs, as we can conveniently reset to any step by simply copy-pasting historical text input. Such a special property is the key motivation of our work.
|
| 1276 |
+
4
|
| 1277 |
+
Unifying Planning, Reasoning, and Acting
|
| 1278 |
+
4.1
|
| 1279 |
+
LM Agent
|
| 1280 |
+
LATS supports sequential reasoning or decision-making tasks on the basis of ReAct. At time step
|
| 1281 |
+
t
|
| 1282 |
+
𝑡
|
| 1283 |
+
t
|
| 1284 |
+
, an agent receives an observation
|
| 1285 |
+
o
|
| 1286 |
+
t
|
| 1287 |
+
∈
|
| 1288 |
+
O
|
| 1289 |
+
subscript
|
| 1290 |
+
𝑜
|
| 1291 |
+
𝑡
|
| 1292 |
+
𝑂
|
| 1293 |
+
o_{t}\in O
|
| 1294 |
+
from the environment and takes an action
|
| 1295 |
+
a
|
| 1296 |
+
t
|
| 1297 |
+
∈
|
| 1298 |
+
A
|
| 1299 |
+
subscript
|
| 1300 |
+
𝑎
|
| 1301 |
+
𝑡
|
| 1302 |
+
𝐴
|
| 1303 |
+
a_{t}\in A
|
| 1304 |
+
following some policy
|
| 1305 |
+
π
|
| 1306 |
+
|
| 1307 |
+
(
|
| 1308 |
+
a
|
| 1309 |
+
t
|
| 1310 |
+
|
|
| 1311 |
+
x
|
| 1312 |
+
,
|
| 1313 |
+
o
|
| 1314 |
+
1
|
| 1315 |
+
|
| 1316 |
+
⋯
|
| 1317 |
+
|
| 1318 |
+
i
|
| 1319 |
+
−
|
| 1320 |
+
1
|
| 1321 |
+
,
|
| 1322 |
+
a
|
| 1323 |
+
1
|
| 1324 |
+
|
| 1325 |
+
⋯
|
| 1326 |
+
|
| 1327 |
+
i
|
| 1328 |
+
−
|
| 1329 |
+
1
|
| 1330 |
+
)
|
| 1331 |
+
𝜋
|
| 1332 |
+
conditional
|
| 1333 |
+
subscript
|
| 1334 |
+
𝑎
|
| 1335 |
+
𝑡
|
| 1336 |
+
𝑥
|
| 1337 |
+
subscript
|
| 1338 |
+
𝑜
|
| 1339 |
+
1
|
| 1340 |
+
⋯
|
| 1341 |
+
𝑖
|
| 1342 |
+
1
|
| 1343 |
+
subscript
|
| 1344 |
+
𝑎
|
| 1345 |
+
1
|
| 1346 |
+
⋯
|
| 1347 |
+
𝑖
|
| 1348 |
+
1
|
| 1349 |
+
\pi(a_{t}|x,o_{1\cdots i-1},a_{1\cdots i-1})
|
| 1350 |
+
, where
|
| 1351 |
+
x
|
| 1352 |
+
𝑥
|
| 1353 |
+
x
|
| 1354 |
+
consists of the task instruction and a number of few-shot examples. We initialize the agent with
|
| 1355 |
+
p
|
| 1356 |
+
θ
|
| 1357 |
+
subscript
|
| 1358 |
+
𝑝
|
| 1359 |
+
𝜃
|
| 1360 |
+
p_{\theta}
|
| 1361 |
+
to leverage the useful language representations of an LM as a base decision-maker. We follow the ReAct instantiation in which the action space
|
| 1362 |
+
A
|
| 1363 |
+
^
|
| 1364 |
+
=
|
| 1365 |
+
A
|
| 1366 |
+
∪
|
| 1367 |
+
Z
|
| 1368 |
+
^
|
| 1369 |
+
𝐴
|
| 1370 |
+
𝐴
|
| 1371 |
+
𝑍
|
| 1372 |
+
\hat{A}=A\cup Z
|
| 1373 |
+
consists of both the space of permissible actions
|
| 1374 |
+
A
|
| 1375 |
+
𝐴
|
| 1376 |
+
A
|
| 1377 |
+
and language space of reasoning traces
|
| 1378 |
+
Z
|
| 1379 |
+
𝑍
|
| 1380 |
+
Z
|
| 1381 |
+
. Actions directly affect the environment and result in observation, while thoughts are used to formalize decisions by organizing information, planning future actions, or injecting internal knowledge. The exact instantiation of the action space depends on the particular environment; for decision-making tasks actions might consist of commands on a website while for reasoning tasks the action space might be limited to a few external tools or APIs.
|
| 1382 |
+
Instead of greedily decoding one trajectory or solution, we sample
|
| 1383 |
+
n
|
| 1384 |
+
𝑛
|
| 1385 |
+
n
|
| 1386 |
+
actions from
|
| 1387 |
+
p
|
| 1388 |
+
θ
|
| 1389 |
+
subscript
|
| 1390 |
+
𝑝
|
| 1391 |
+
𝜃
|
| 1392 |
+
p_{\theta}
|
| 1393 |
+
using the current state. This is based on the intuition that for complex decision-making tasks, there is likely to be a range of potential trajectories or reasoning paths that are correct
|
| 1394 |
+
(Evans,
|
| 1395 |
+
2010
|
| 1396 |
+
)
|
| 1397 |
+
. Sampling a diverse set of candidates at each step mitigates the stochastic nature of LM text generation and enables greater exploration in both the decision-making and reasoning space. We wrap
|
| 1398 |
+
p
|
| 1399 |
+
θ
|
| 1400 |
+
subscript
|
| 1401 |
+
𝑝
|
| 1402 |
+
𝜃
|
| 1403 |
+
p_{\theta}
|
| 1404 |
+
within our proposed search algorithm to deliberately construct the best trajectory from sampled actions.
|
| 1405 |
+
4.2
|
| 1406 |
+
LATS
|
| 1407 |
+
Figure 3:
|
| 1408 |
+
An overview of the six operations of LATS. A node is
|
| 1409 |
+
selected
|
| 1410 |
+
,
|
| 1411 |
+
expanded
|
| 1412 |
+
,
|
| 1413 |
+
evaluated
|
| 1414 |
+
, then
|
| 1415 |
+
simulated
|
| 1416 |
+
until a terminal node is reached, then the resulting value is
|
| 1417 |
+
backpropagated
|
| 1418 |
+
. If the trajectory fails, a
|
| 1419 |
+
reflection
|
| 1420 |
+
is generated and used as additional context for future trials. These operations are performed in succession until the budget is reached or task is successful.
|
| 1421 |
+
The main component of LATS is a search algorithm that controls the overall problem-solving process with deliberate planning. To find the most promising trajectory and systemically balance exploration with exploitation, we adopt a variant of Monte Carlo Tree Search (MCTS) that frames decision-making as a tree search, in which each node
|
| 1422 |
+
s
|
| 1423 |
+
=
|
| 1424 |
+
[
|
| 1425 |
+
x
|
| 1426 |
+
,
|
| 1427 |
+
a
|
| 1428 |
+
1
|
| 1429 |
+
|
| 1430 |
+
⋯
|
| 1431 |
+
|
| 1432 |
+
i
|
| 1433 |
+
,
|
| 1434 |
+
o
|
| 1435 |
+
1
|
| 1436 |
+
|
| 1437 |
+
⋯
|
| 1438 |
+
|
| 1439 |
+
i
|
| 1440 |
+
]
|
| 1441 |
+
𝑠
|
| 1442 |
+
𝑥
|
| 1443 |
+
subscript
|
| 1444 |
+
𝑎
|
| 1445 |
+
1
|
| 1446 |
+
⋯
|
| 1447 |
+
𝑖
|
| 1448 |
+
subscript
|
| 1449 |
+
𝑜
|
| 1450 |
+
1
|
| 1451 |
+
⋯
|
| 1452 |
+
𝑖
|
| 1453 |
+
s=[x,a_{1\cdots i},o_{1\cdots i}]
|
| 1454 |
+
represents a state comprising the original input
|
| 1455 |
+
x
|
| 1456 |
+
𝑥
|
| 1457 |
+
x
|
| 1458 |
+
, action sequence
|
| 1459 |
+
a
|
| 1460 |
+
1
|
| 1461 |
+
⋅
|
| 1462 |
+
i
|
| 1463 |
+
subscript
|
| 1464 |
+
𝑎
|
| 1465 |
+
⋅
|
| 1466 |
+
1
|
| 1467 |
+
𝑖
|
| 1468 |
+
a_{1\cdot i}
|
| 1469 |
+
, and observation sequence
|
| 1470 |
+
o
|
| 1471 |
+
1
|
| 1472 |
+
⋅
|
| 1473 |
+
i
|
| 1474 |
+
subscript
|
| 1475 |
+
𝑜
|
| 1476 |
+
⋅
|
| 1477 |
+
1
|
| 1478 |
+
𝑖
|
| 1479 |
+
o_{1\cdot i}
|
| 1480 |
+
.
|
| 1481 |
+
To adapt MCTS for language agents, LATS repurposes
|
| 1482 |
+
p
|
| 1483 |
+
θ
|
| 1484 |
+
subscript
|
| 1485 |
+
𝑝
|
| 1486 |
+
𝜃
|
| 1487 |
+
p_{\theta}
|
| 1488 |
+
as an agent, state evaluator, and feedback generator, leveraging the useful language priors of modern LMs to facilitate planning. While standard MCTS and RAP
|
| 1489 |
+
Hao et al. (
|
| 1490 |
+
2023
|
| 1491 |
+
)
|
| 1492 |
+
rely on internal dynamics models to facilitate simulation, LATS is model-free and uses environment interaction. LATS consists of a series of operations,
|
| 1493 |
+
selection, expansion, evaluation, simulation, backpropagation, and reflection
|
| 1494 |
+
, performed in succession until the task is successfully completed or a computational limit is reached. The full psuedocode of LATS can be found in Sec.
|
| 1495 |
+
A
|
| 1496 |
+
in the Appendix.
|
| 1497 |
+
Selection.
|
| 1498 |
+
In the first operation, the algorithm identifies a segment of the current tree most suitable for subsequent expansion. Starting from the root node, denoted as the initial state
|
| 1499 |
+
s
|
| 1500 |
+
0
|
| 1501 |
+
subscript
|
| 1502 |
+
𝑠
|
| 1503 |
+
0
|
| 1504 |
+
s_{0}
|
| 1505 |
+
, a child node is selected at each tree level until a leaf node is reached. To balance exploration and exploitation, we use the UCT algorithm as shown in Eq.
|
| 1506 |
+
1
|
| 1507 |
+
.
|
| 1508 |
+
Expansion.
|
| 1509 |
+
After selecting a node, the second operation expands the tree by sampling
|
| 1510 |
+
n
|
| 1511 |
+
𝑛
|
| 1512 |
+
n
|
| 1513 |
+
actions from
|
| 1514 |
+
p
|
| 1515 |
+
θ
|
| 1516 |
+
subscript
|
| 1517 |
+
𝑝
|
| 1518 |
+
𝜃
|
| 1519 |
+
p_{\theta}
|
| 1520 |
+
, as described in the prior section. The environment receives each action and returns corresponding feedback as an observation. This results in
|
| 1521 |
+
n
|
| 1522 |
+
𝑛
|
| 1523 |
+
n
|
| 1524 |
+
new child nodes added to the tree. This tree is stored in an external long-term memory structure.
|
| 1525 |
+
Evaluation.
|
| 1526 |
+
The third operation assigns a scalar value to each new child node to be used for selection and backpropagation. This value effectively quantifies the agent’s progress in task completion, serving as a heuristic to steer the search algorithm towards the most promising regions of the tree. Following
|
| 1527 |
+
Yao et al. (
|
| 1528 |
+
2023a
|
| 1529 |
+
)
|
| 1530 |
+
we repurpose
|
| 1531 |
+
p
|
| 1532 |
+
θ
|
| 1533 |
+
subscript
|
| 1534 |
+
𝑝
|
| 1535 |
+
𝜃
|
| 1536 |
+
p_{\theta}
|
| 1537 |
+
into a value function by prompting it to reason about a given state. To obtain a scalar value, we instruct
|
| 1538 |
+
p
|
| 1539 |
+
θ
|
| 1540 |
+
subscript
|
| 1541 |
+
𝑝
|
| 1542 |
+
𝜃
|
| 1543 |
+
p_{\theta}
|
| 1544 |
+
to end its reasoning trace with a score indicating the correctness of the trajectory. This method offers enhanced flexibility over programmed heuristics
|
| 1545 |
+
(Campbell et al.,
|
| 1546 |
+
2002
|
| 1547 |
+
)
|
| 1548 |
+
and greater efficiency than learned heuristics
|
| 1549 |
+
(Silver et al.,
|
| 1550 |
+
2017
|
| 1551 |
+
)
|
| 1552 |
+
.
|
| 1553 |
+
Simulation.
|
| 1554 |
+
The fourth operation expands the currently selected node until a terminal state is reached. At each depth level we sample and evaluate nodes with the same operations, but prioritize nodes of highest value. Reaching a terminal state provides objective feedback on the correctness of a trajectory. If the task is completed successfully, then LATS terminates the search. If the solution is partially successful or unsuccessful, then we perform two additional operations as described below.
|
| 1555 |
+
Backpropagation.
|
| 1556 |
+
This operation updates the values of the tree based on the outcome of a trajectory. For each node
|
| 1557 |
+
s
|
| 1558 |
+
0
|
| 1559 |
+
,
|
| 1560 |
+
s
|
| 1561 |
+
1
|
| 1562 |
+
,
|
| 1563 |
+
…
|
| 1564 |
+
,
|
| 1565 |
+
s
|
| 1566 |
+
n
|
| 1567 |
+
subscript
|
| 1568 |
+
𝑠
|
| 1569 |
+
0
|
| 1570 |
+
subscript
|
| 1571 |
+
𝑠
|
| 1572 |
+
1
|
| 1573 |
+
…
|
| 1574 |
+
subscript
|
| 1575 |
+
𝑠
|
| 1576 |
+
𝑛
|
| 1577 |
+
s_{0},s_{1},\dots,s_{n}
|
| 1578 |
+
in the trajectory from root (initial state
|
| 1579 |
+
s
|
| 1580 |
+
0
|
| 1581 |
+
subscript
|
| 1582 |
+
𝑠
|
| 1583 |
+
0
|
| 1584 |
+
s_{0}
|
| 1585 |
+
) of the searching tree to leaf (terminal state
|
| 1586 |
+
s
|
| 1587 |
+
n
|
| 1588 |
+
subscript
|
| 1589 |
+
𝑠
|
| 1590 |
+
𝑛
|
| 1591 |
+
s_{n}
|
| 1592 |
+
), its value is updated to reflect the outcome of the simulation by
|
| 1593 |
+
N
|
| 1594 |
+
|
| 1595 |
+
(
|
| 1596 |
+
s
|
| 1597 |
+
i
|
| 1598 |
+
)
|
| 1599 |
+
=
|
| 1600 |
+
N
|
| 1601 |
+
old
|
| 1602 |
+
|
| 1603 |
+
(
|
| 1604 |
+
s
|
| 1605 |
+
i
|
| 1606 |
+
)
|
| 1607 |
+
+
|
| 1608 |
+
1
|
| 1609 |
+
𝑁
|
| 1610 |
+
subscript
|
| 1611 |
+
𝑠
|
| 1612 |
+
𝑖
|
| 1613 |
+
subscript
|
| 1614 |
+
𝑁
|
| 1615 |
+
old
|
| 1616 |
+
subscript
|
| 1617 |
+
𝑠
|
| 1618 |
+
𝑖
|
| 1619 |
+
1
|
| 1620 |
+
N(s_{i})=N_{\text{old}}(s_{i})+1
|
| 1621 |
+
and
|
| 1622 |
+
V
|
| 1623 |
+
|
| 1624 |
+
(
|
| 1625 |
+
s
|
| 1626 |
+
i
|
| 1627 |
+
)
|
| 1628 |
+
=
|
| 1629 |
+
r
|
| 1630 |
+
+
|
| 1631 |
+
N
|
| 1632 |
+
old
|
| 1633 |
+
|
| 1634 |
+
(
|
| 1635 |
+
s
|
| 1636 |
+
i
|
| 1637 |
+
)
|
| 1638 |
+
|
| 1639 |
+
V
|
| 1640 |
+
old
|
| 1641 |
+
|
| 1642 |
+
(
|
| 1643 |
+
s
|
| 1644 |
+
i
|
| 1645 |
+
)
|
| 1646 |
+
N
|
| 1647 |
+
|
| 1648 |
+
(
|
| 1649 |
+
s
|
| 1650 |
+
i
|
| 1651 |
+
)
|
| 1652 |
+
𝑉
|
| 1653 |
+
subscript
|
| 1654 |
+
𝑠
|
| 1655 |
+
𝑖
|
| 1656 |
+
𝑟
|
| 1657 |
+
subscript
|
| 1658 |
+
𝑁
|
| 1659 |
+
old
|
| 1660 |
+
subscript
|
| 1661 |
+
𝑠
|
| 1662 |
+
𝑖
|
| 1663 |
+
subscript
|
| 1664 |
+
𝑉
|
| 1665 |
+
old
|
| 1666 |
+
subscript
|
| 1667 |
+
𝑠
|
| 1668 |
+
𝑖
|
| 1669 |
+
𝑁
|
| 1670 |
+
subscript
|
| 1671 |
+
𝑠
|
| 1672 |
+
𝑖
|
| 1673 |
+
V(s_{i})=\frac{r+N_{\text{old}}(s_{i})V_{\text{old}}(s_{i})}{N(s_{i})}
|
| 1674 |
+
, where
|
| 1675 |
+
r
|
| 1676 |
+
𝑟
|
| 1677 |
+
r
|
| 1678 |
+
is the return and
|
| 1679 |
+
N
|
| 1680 |
+
old
|
| 1681 |
+
,
|
| 1682 |
+
V
|
| 1683 |
+
old
|
| 1684 |
+
subscript
|
| 1685 |
+
𝑁
|
| 1686 |
+
old
|
| 1687 |
+
subscript
|
| 1688 |
+
𝑉
|
| 1689 |
+
old
|
| 1690 |
+
N_{\text{old}},V_{\text{old}}
|
| 1691 |
+
are the old number of visits and value function. These updated values are used in the UCT formula (Eq.
|
| 1692 |
+
1
|
| 1693 |
+
) to guide the selection of the next node for exploration.
|
| 1694 |
+
Reflection.
|
| 1695 |
+
In addition to the environmental feedback, we also leverage
|
| 1696 |
+
self-reflection
|
| 1697 |
+
to further refine the decision-making process
|
| 1698 |
+
(Shinn et al.,
|
| 1699 |
+
2023
|
| 1700 |
+
; Madaan et al.,
|
| 1701 |
+
2023
|
| 1702 |
+
)
|
| 1703 |
+
. Upon encountering an unsuccessful terminal node,
|
| 1704 |
+
p
|
| 1705 |
+
θ
|
| 1706 |
+
subscript
|
| 1707 |
+
𝑝
|
| 1708 |
+
𝜃
|
| 1709 |
+
p_{\theta}
|
| 1710 |
+
is prompted with the trajectory and final reward to provide a verbal self-reflection that summarizes the errors in the reasoning or acting process and proposes superior alternatives. We store both failed trajectories and corresponding reflections in the memory. In subsequent iterations, these are integrated as additional context to the agent and value function, refining both through in-context learning. This imparts a semantic gradient signal more useful than a scalar value, enabling the agent to learn from trial and error without the cost of expensive optimization processes such as reinforcement learning.
|
| 1711 |
+
Conceptually, LATS has the following advantages as a general framework for reasoning and decision-making with LM agents.
|
| 1712 |
+
(1)
|
| 1713 |
+
Generality
|
| 1714 |
+
: LATS supports both reasoning and decision-making tasks by defining a shared space of thoughts and actions. (2)
|
| 1715 |
+
Deliberate
|
| 1716 |
+
: The use of MCTS and LM value function ensures a principled search that selects options with high value while exploring promising alternatives. (3)
|
| 1717 |
+
Adaptability
|
| 1718 |
+
: LATS is designed around the use of external feedback through observations and self-reflection, enabling greater adaptation during problem-solving. (4)
|
| 1719 |
+
Flexibility
|
| 1720 |
+
: LATS can accommodate different scenarios, environments, and resource stipulations by modifying state design and tree dimensions. (5)
|
| 1721 |
+
Modularity
|
| 1722 |
+
: The base LM agent, reflection generator, and value function can be independently altered and adapted to individual LM properties.
|
| 1723 |
+
5
|
| 1724 |
+
Experiments
|
| 1725 |
+
To demonstrate the general applicability of LATS, we evaluate our method on a variety of decision-making domains that requires both reasoning and acting ability: programming
|
| 1726 |
+
(Chen et al.,
|
| 1727 |
+
2021
|
| 1728 |
+
; Austin et al.,
|
| 1729 |
+
2021
|
| 1730 |
+
)
|
| 1731 |
+
, HotPotQA
|
| 1732 |
+
(Yang et al.,
|
| 1733 |
+
2018
|
| 1734 |
+
)
|
| 1735 |
+
, and WebShop
|
| 1736 |
+
(Yao et al.,
|
| 1737 |
+
2022
|
| 1738 |
+
)
|
| 1739 |
+
.
|
| 1740 |
+
5.1
|
| 1741 |
+
HotPotQA
|
| 1742 |
+
For a task that can be approached with both reasoning-based and acting-based strategies, we consider HotPotQA
|
| 1743 |
+
(Yang et al.,
|
| 1744 |
+
2018
|
| 1745 |
+
)
|
| 1746 |
+
, a multi-hop question-answering benchmark that requires retrieval over two or more Wikipedia passages. For the action space, in addition to LM thoughts we follow the setup from
|
| 1747 |
+
Yao et al. (
|
| 1748 |
+
2023b
|
| 1749 |
+
)
|
| 1750 |
+
, which provides the agent with API calls to search and lookup information. The output of these API calls and self-generated reflections form the observation space. We use a subset of 100 questions and three few-shot examples for each method. For ToT, we use DFS as the base search algorithm and scoring with the LM as the heuristic. For all methods that involve sampling, including LATS, we sample
|
| 1751 |
+
k
|
| 1752 |
+
=
|
| 1753 |
+
50
|
| 1754 |
+
𝑘
|
| 1755 |
+
50
|
| 1756 |
+
k=50
|
| 1757 |
+
trajectories. More details and prompts can be found in Sec.
|
| 1758 |
+
D
|
| 1759 |
+
and Sec.
|
| 1760 |
+
E
|
| 1761 |
+
in the Appendix.
|
| 1762 |
+
We evaluate internal reasoning strategies by removing actions and observations from the context, corresponding to CoT
|
| 1763 |
+
(Wei et al.,
|
| 1764 |
+
2022
|
| 1765 |
+
)
|
| 1766 |
+
and its variants, CoT-SC
|
| 1767 |
+
(Wang et al.,
|
| 1768 |
+
2022
|
| 1769 |
+
)
|
| 1770 |
+
, ToT
|
| 1771 |
+
(Yao et al.,
|
| 1772 |
+
2023a
|
| 1773 |
+
)
|
| 1774 |
+
, and RAP
|
| 1775 |
+
(Hao et al.,
|
| 1776 |
+
2023
|
| 1777 |
+
)
|
| 1778 |
+
. These methods rely solely on the agent’s existing knowledge to answer the question. We also consider acting-based methods ReAct, Reflexion, and LATS, which augment the agent with the interactive API environment and primarily evaluate its information retrieval abilities. While LATS is designed for scenarios where external feedback can enhance reasoning, we also implement a reasoning-only version with CoT as the base prompt. We also combine internal and external reasoning in LATS by first prompting with a CoT-based prompt, then switching to a ReAct-based prompt upon failure. This is closer to how humans might approach this task, by using tools to lookup additional information only when the answer is not already known.
|
| 1779 |
+
Prompt Method
|
| 1780 |
+
HotpotQA (EM)
|
| 1781 |
+
I/O
|
| 1782 |
+
0.32
|
| 1783 |
+
CoT
|
| 1784 |
+
(Wei et al.,
|
| 1785 |
+
2022
|
| 1786 |
+
)
|
| 1787 |
+
0.34
|
| 1788 |
+
CoT - SC
|
| 1789 |
+
(Wang et al.,
|
| 1790 |
+
2022
|
| 1791 |
+
)
|
| 1792 |
+
0.38
|
| 1793 |
+
ToT
|
| 1794 |
+
(Yao et al.,
|
| 1795 |
+
2023a
|
| 1796 |
+
)
|
| 1797 |
+
0.55
|
| 1798 |
+
RAP
|
| 1799 |
+
(Hao et al.,
|
| 1800 |
+
2023
|
| 1801 |
+
)
|
| 1802 |
+
0.60
|
| 1803 |
+
RAP (n = 10)
|
| 1804 |
+
0.60
|
| 1805 |
+
LATS (CoT)
|
| 1806 |
+
0.60
|
| 1807 |
+
Prompt Method
|
| 1808 |
+
HotpotQA (EM)
|
| 1809 |
+
ReAct
|
| 1810 |
+
(Yao et al.,
|
| 1811 |
+
2023b
|
| 1812 |
+
)
|
| 1813 |
+
0.32
|
| 1814 |
+
ReAct (best of k)
|
| 1815 |
+
0.38
|
| 1816 |
+
Reflexion
|
| 1817 |
+
(Shinn et al.,
|
| 1818 |
+
2023
|
| 1819 |
+
)
|
| 1820 |
+
0.51
|
| 1821 |
+
LATS
|
| 1822 |
+
0.61
|
| 1823 |
+
LATS (n = 3)
|
| 1824 |
+
0.56
|
| 1825 |
+
LATS (n = 10)
|
| 1826 |
+
0.64
|
| 1827 |
+
LATS (CoT + ReAct)
|
| 1828 |
+
0.71
|
| 1829 |
+
Table 2:
|
| 1830 |
+
GPT-3.5 reasoning-based prompting (left) and acting-based prompting (right) results on HotpotQA. LATS achieves the highest exact match (EM) for acting and is competitive on reasoning. Unless otherwise specified, we sample
|
| 1831 |
+
n
|
| 1832 |
+
=
|
| 1833 |
+
5
|
| 1834 |
+
𝑛
|
| 1835 |
+
5
|
| 1836 |
+
n=5
|
| 1837 |
+
nodes during expansion and
|
| 1838 |
+
k
|
| 1839 |
+
=
|
| 1840 |
+
50
|
| 1841 |
+
𝑘
|
| 1842 |
+
50
|
| 1843 |
+
k=50
|
| 1844 |
+
trajectories.
|
| 1845 |
+
Results.
|
| 1846 |
+
We observe in Tab.
|
| 1847 |
+
2
|
| 1848 |
+
that both internal reasoning and external retrieval strategies perform well on HotPotQA. Due to their large-scale training corpus, modern LLMs already encode factual knowledge and can often directly answer the question correctly. While CoT can slightly enhance performance on questions requiring reasoning, larger gains are observed with search methods ToT and RAP, which can sample and explore more outputs. We observe similar results for acting-based methods. LATS surpasses ReAct, even when sampling the same number of trajectories, by expanding more nodes with principled search (see Fig.
|
| 1849 |
+
5
|
| 1850 |
+
in Appendix
|
| 1851 |
+
D
|
| 1852 |
+
for a qualitative sample). This is demonstrated when modifying
|
| 1853 |
+
n
|
| 1854 |
+
𝑛
|
| 1855 |
+
n
|
| 1856 |
+
, the number of nodes expanded during each iteration. Increasing
|
| 1857 |
+
n
|
| 1858 |
+
𝑛
|
| 1859 |
+
n
|
| 1860 |
+
can consistently improve performance, although at greater computational and inference costs. LATS is also competitive to RAP on internal reasoning but performs worse than acting. Combining internal and external reasoning in LATS results in the highest performance, indicating the importance of external feedback in augmenting reasoning even in tasks the base LM can already perform.
|
| 1861 |
+
5.2
|
| 1862 |
+
Programming
|
| 1863 |
+
Prompt Method
|
| 1864 |
+
Model
|
| 1865 |
+
Pass@1
|
| 1866 |
+
CoT
|
| 1867 |
+
(Wei et al.,
|
| 1868 |
+
2022
|
| 1869 |
+
)
|
| 1870 |
+
GPT-3.5
|
| 1871 |
+
46.9
|
| 1872 |
+
ReAct
|
| 1873 |
+
(Yao et al.,
|
| 1874 |
+
2023b
|
| 1875 |
+
)
|
| 1876 |
+
GPT-3.5
|
| 1877 |
+
56.9
|
| 1878 |
+
Reflexion
|
| 1879 |
+
(Shinn et al.,
|
| 1880 |
+
2023
|
| 1881 |
+
)
|
| 1882 |
+
GPT-3.5
|
| 1883 |
+
68.1
|
| 1884 |
+
ToT
|
| 1885 |
+
(Yao et al.,
|
| 1886 |
+
2023a
|
| 1887 |
+
)
|
| 1888 |
+
GPT-3.5
|
| 1889 |
+
54.4
|
| 1890 |
+
RAP
|
| 1891 |
+
(Hao et al.,
|
| 1892 |
+
2023
|
| 1893 |
+
)
|
| 1894 |
+
GPT-3.5
|
| 1895 |
+
63.1
|
| 1896 |
+
LATS (Ours)
|
| 1897 |
+
GPT-3.5
|
| 1898 |
+
83.8
|
| 1899 |
+
I/O
|
| 1900 |
+
GPT-4
|
| 1901 |
+
80.1
|
| 1902 |
+
Reflexion
|
| 1903 |
+
GPT-4
|
| 1904 |
+
91.0
|
| 1905 |
+
LATS
|
| 1906 |
+
GPT-4
|
| 1907 |
+
94.4
|
| 1908 |
+
Prompt Method
|
| 1909 |
+
Pass@1
|
| 1910 |
+
CoT
|
| 1911 |
+
(Wei et al.,
|
| 1912 |
+
2022
|
| 1913 |
+
)
|
| 1914 |
+
54.9
|
| 1915 |
+
ReAct
|
| 1916 |
+
(Wei et al.,
|
| 1917 |
+
2022
|
| 1918 |
+
)
|
| 1919 |
+
67.0
|
| 1920 |
+
Reflexion
|
| 1921 |
+
(Shinn et al.,
|
| 1922 |
+
2023
|
| 1923 |
+
)
|
| 1924 |
+
70.0
|
| 1925 |
+
ToT
|
| 1926 |
+
(Yao et al.,
|
| 1927 |
+
2023a
|
| 1928 |
+
)
|
| 1929 |
+
65.8
|
| 1930 |
+
RAP
|
| 1931 |
+
(Hao et al.,
|
| 1932 |
+
2023
|
| 1933 |
+
)
|
| 1934 |
+
71.4
|
| 1935 |
+
LATS (Ours)
|
| 1936 |
+
81.1
|
| 1937 |
+
Table 3:
|
| 1938 |
+
GPT-3.5 and GPT-4 Pass@1 accuracy on HumanEval
|
| 1939 |
+
(Chen et al.,
|
| 1940 |
+
2021
|
| 1941 |
+
)
|
| 1942 |
+
and MBPP
|
| 1943 |
+
(Austin et al.,
|
| 1944 |
+
2021
|
| 1945 |
+
)
|
| 1946 |
+
. Prompting with LATS achieves the highest performance. We sample 5 solutions during expansion for
|
| 1947 |
+
8
|
| 1948 |
+
iterations.
|
| 1949 |
+
To demonstrate the importance of external observations for complex reasoning tasks, we evaluate the baselines and LATS on programming with Humaneval
|
| 1950 |
+
(Chen et al.,
|
| 1951 |
+
2021
|
| 1952 |
+
)
|
| 1953 |
+
and MBPP
|
| 1954 |
+
(Austin et al.,
|
| 1955 |
+
2021
|
| 1956 |
+
)
|
| 1957 |
+
. Both datasets measure the correctness of synthesized programs in Python from natural language docstrings. We use individual solutions as the action space and test suite and compiler feedback as the external observation. We follow
|
| 1958 |
+
Chen et al. (
|
| 1959 |
+
2022a
|
| 1960 |
+
)
|
| 1961 |
+
and use an LLM to generate a synthetic test suite of syntactically valid “assert” statements for each question. For each step, the solution is evaluated on this test suite, and the results including successful and failed tests and compiler output, are added to the context as an observation. We use the same test suite for Reflexion.
|
| 1962 |
+
For this task, the reasoning and acting baselines share an action space, but acting methods are able to incorporate observations as additional context. For LATS, since each action corresponds to a complete solution, we skip the simulation step of LATS and directly use the percentage of passed tests as the backpropagated reward. We use
|
| 1963 |
+
k
|
| 1964 |
+
=
|
| 1965 |
+
8
|
| 1966 |
+
𝑘
|
| 1967 |
+
8
|
| 1968 |
+
k=8
|
| 1969 |
+
iterations, set the number of generated tests at
|
| 1970 |
+
4
|
| 1971 |
+
4
|
| 1972 |
+
4
|
| 1973 |
+
, and sample
|
| 1974 |
+
n
|
| 1975 |
+
=
|
| 1976 |
+
5
|
| 1977 |
+
𝑛
|
| 1978 |
+
5
|
| 1979 |
+
n=5
|
| 1980 |
+
solutions during expansion. After the search is completed, we select the solution with the highest value and evaluate it on the real test suite for the pass@1 accuracy evaluation. More details and prompts can be found in Sec.
|
| 1981 |
+
D
|
| 1982 |
+
and Sec.
|
| 1983 |
+
F
|
| 1984 |
+
in the Appendix.
|
| 1985 |
+
Results.
|
| 1986 |
+
We find in Tab
|
| 1987 |
+
3
|
| 1988 |
+
that both search and semantic feedback are crucial for better performance. Despite not using observations, ToT and RAP are competitive with Reflexion. LATS has the highest performance on both datasets. Since RAP uses a similar search algorithm as LATS, this reveals the importance of external feedback for difficult reasoning tasks such as programming. With GPT-4, using LATS sets the state of the art for HumanEval, showing LATS can be used with more advanced LLMs for higher performance.
|
| 1989 |
+
5.3
|
| 1990 |
+
Webshop
|
| 1991 |
+
For a complex decision-making environment with practical applications, we consider WebShop
|
| 1992 |
+
(Yao et al.,
|
| 1993 |
+
2022
|
| 1994 |
+
)
|
| 1995 |
+
, an online shopping environment composed of a website with 1.18M real-world products and 12k human instructions. Agents must navigate a website through a variety of commands to purchase an item matching a user specification. We use the preconstructed action space of search and click commands and browser feedback and reflections for the observation. The performance is gauged using two metrics: an average score, reflecting the percentage of user-specified attributes met by the selected product, and a success rate, indicating the frequency with which the chosen product fulfills all given conditions. We compare against acting-based prompting methods and RL-based approaches. We evaluate on 50 instructions, expand
|
| 1996 |
+
n
|
| 1997 |
+
=
|
| 1998 |
+
5
|
| 1999 |
+
𝑛
|
| 2000 |
+
5
|
| 2001 |
+
n=5
|
| 2002 |
+
children for LATS, and set
|
| 2003 |
+
k
|
| 2004 |
+
=
|
| 2005 |
+
30
|
| 2006 |
+
𝑘
|
| 2007 |
+
30
|
| 2008 |
+
k=30
|
| 2009 |
+
for LATS, ReAct best of
|
| 2010 |
+
k
|
| 2011 |
+
𝑘
|
| 2012 |
+
k
|
| 2013 |
+
, and Reflexion. More details and prompts are in Appendix
|
| 2014 |
+
D
|
| 2015 |
+
and
|
| 2016 |
+
G
|
| 2017 |
+
.
|
| 2018 |
+
Results.
|
| 2019 |
+
We find in Tab.
|
| 2020 |
+
5
|
| 2021 |
+
that GPT-3.5 with ReAct is competitive to imitation learning, and can exceed reinforcement learning techniques with stronger prompting strategies. Sampling
|
| 2022 |
+
k
|
| 2023 |
+
=
|
| 2024 |
+
30
|
| 2025 |
+
𝑘
|
| 2026 |
+
30
|
| 2027 |
+
k=30
|
| 2028 |
+
trajectories with ReAct and Reflexion results in a similar performance, suggesting the semantic feedback is not as helpful in complex environments like WebShop. Indeed like in
|
| 2029 |
+
Shinn et al. (
|
| 2030 |
+
2023
|
| 2031 |
+
)
|
| 2032 |
+
, we find that generated reflections are often generic and do not provide useful feedback, resulting in a tendency for the agent to become stuck in local minima. However, using LATS indeed results in a noticeable improvement, indicating a more effective exploration for the same number of iterations.
|
| 2033 |
+
5.4
|
| 2034 |
+
Additional Observations
|
| 2035 |
+
Method
|
| 2036 |
+
Score
|
| 2037 |
+
SR
|
| 2038 |
+
ReAct
|
| 2039 |
+
(Yao et al.,
|
| 2040 |
+
2023b
|
| 2041 |
+
)
|
| 2042 |
+
53.8
|
| 2043 |
+
28.0
|
| 2044 |
+
ReAct (best of k)
|
| 2045 |
+
59.1
|
| 2046 |
+
32.0
|
| 2047 |
+
Reflexion
|
| 2048 |
+
(Shinn et al.,
|
| 2049 |
+
2023
|
| 2050 |
+
)
|
| 2051 |
+
64.2
|
| 2052 |
+
35.0
|
| 2053 |
+
LATS
|
| 2054 |
+
75.9
|
| 2055 |
+
38.0
|
| 2056 |
+
IL
|
| 2057 |
+
59.9
|
| 2058 |
+
29.1
|
| 2059 |
+
IL+RL
|
| 2060 |
+
62.4
|
| 2061 |
+
28.7
|
| 2062 |
+
Fine-tuning
|
| 2063 |
+
(Furuta et al.,
|
| 2064 |
+
2023
|
| 2065 |
+
)
|
| 2066 |
+
67.5
|
| 2067 |
+
45.0
|
| 2068 |
+
Expert
|
| 2069 |
+
82.1
|
| 2070 |
+
59.6
|
| 2071 |
+
Table 4:
|
| 2072 |
+
Score and success rate (SR) on Webshop. Table is separated into prompting, RL-based training, and human performance. For the same number of iterations, LATS improves both score and success rate, and surpasses RL-based training. IL/IL+RL taken from
|
| 2073 |
+
Yao et al. (
|
| 2074 |
+
2022
|
| 2075 |
+
)
|
| 2076 |
+
.
|
| 2077 |
+
Prompt Method
|
| 2078 |
+
HotPotQA (EM)
|
| 2079 |
+
ToT (ReAct)
|
| 2080 |
+
0.39
|
| 2081 |
+
RAP (ReAct)
|
| 2082 |
+
0.54
|
| 2083 |
+
LATS (No LM Heuristic)
|
| 2084 |
+
0.37
|
| 2085 |
+
LATS (DFS)
|
| 2086 |
+
0.42
|
| 2087 |
+
LATS (No Reflection)
|
| 2088 |
+
0.56
|
| 2089 |
+
LATS
|
| 2090 |
+
0.61
|
| 2091 |
+
Table 5:
|
| 2092 |
+
Ablation results on LATS and baseline variants in HotPotQA; we use ReAct as the base prompt and sample
|
| 2093 |
+
n
|
| 2094 |
+
=
|
| 2095 |
+
5
|
| 2096 |
+
𝑛
|
| 2097 |
+
5
|
| 2098 |
+
n=5
|
| 2099 |
+
children and
|
| 2100 |
+
k
|
| 2101 |
+
=
|
| 2102 |
+
50
|
| 2103 |
+
𝑘
|
| 2104 |
+
50
|
| 2105 |
+
k=50
|
| 2106 |
+
maximum trajectories. LATS requires every component and operation for optimal performance.
|
| 2107 |
+
We also conduct additional experiments on HotPotQA to demonstrate the effect of each component of LATS. We also design a version of ToT and RAP with ReAct prompt and can handle external observations. We use HotPotQA as our setup incorporates both reasoning (through thoughts) and acting (through API calls); the results are shown in Tab.
|
| 2108 |
+
5
|
| 2109 |
+
. More ablations for token consumption on HotPotQA are in Tab.
|
| 2110 |
+
7
|
| 2111 |
+
in Appendix
|
| 2112 |
+
C
|
| 2113 |
+
. Note that baselines generally perform worse than the reasoning-only setting of HotPotQA, which indicates that the acting-based setting is more challenging and adaption of search algorithms to decision-making scenarios is non-trivial.
|
| 2114 |
+
Self-reflection.
|
| 2115 |
+
We use self-reflection to provide additional semantic signals for the agent. We observe a
|
| 2116 |
+
0.05
|
| 2117 |
+
0.05
|
| 2118 |
+
0.05
|
| 2119 |
+
performance drop when removed from LATS, suggesting this is useful. This is a smaller gain Reflexion
|
| 2120 |
+
(Shinn et al.,
|
| 2121 |
+
2023
|
| 2122 |
+
)
|
| 2123 |
+
observes over ReAct
|
| 2124 |
+
(Yao et al.,
|
| 2125 |
+
2023b
|
| 2126 |
+
)
|
| 2127 |
+
as shown in Tab.
|
| 2128 |
+
2
|
| 2129 |
+
, suggesting overlap between the types of questions where there is an improvement with self-reflection and search. This variant outperforms RAP-ReAct, reflecting our improvements to MCTS.
|
| 2130 |
+
Search Algorithm.
|
| 2131 |
+
MCTS is a more principled search algorithm than variants like A* or DFS search and the basis for observed performance gains. We observe the effects of using DFS, and incorporate the LM-based heuristic used in ToT
|
| 2132 |
+
(Yao et al.,
|
| 2133 |
+
2023a
|
| 2134 |
+
)
|
| 2135 |
+
in which branches with low values are pruned. This removes the selection and backpropagation operations, and we observe a
|
| 2136 |
+
0.08
|
| 2137 |
+
0.08
|
| 2138 |
+
0.08
|
| 2139 |
+
drop in performance when sampling the same number of nodes, but outperforms ToT-ReAct.
|
| 2140 |
+
6
|
| 2141 |
+
Conclusion
|
| 2142 |
+
In this work, we introduce Language Agent Tree Search (LATS), the first framework to unify planning, acting, and reasoning for enhanced LLM problem solving. By deliberately constructing trajectories with search algorithms, incorporating external feedback, and enabling agents to learn from experience, LATS addresses key limitations of prior prompting techniques. Our evaluations demonstrate the ability of LATS to harness LLM capabilities for a variety of decision-making tasks while keeping its reasoning ability without additional training. The proposed synergies between search, interaction, and reflection offer a versatile approach to autonomous decision-making, highlighting the potential of LLMs as generalist agents. A full discussion of the limitations and broader impacts is in Appendix
|
| 2143 |
+
B
|
| 2144 |
+
.
|
| 2145 |
+
References
|
| 2146 |
+
Ahn et al. (2022)
|
| 2147 |
+
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan, and Andy Zeng.
|
| 2148 |
+
Do as i can, not as i say: Grounding language in robotic affordances.
|
| 2149 |
+
arXiv:2204.01691
|
| 2150 |
+
, 2022.
|
| 2151 |
+
Anthony et al. (2017)
|
| 2152 |
+
T. Anthony, Z. Tian, and D. Barber.
|
| 2153 |
+
Thinking fast and slow with deep learning and tree search.
|
| 2154 |
+
In
|
| 2155 |
+
NIPS
|
| 2156 |
+
, 2017.
|
| 2157 |
+
Austin et al. (2021)
|
| 2158 |
+
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al.
|
| 2159 |
+
Program synthesis with large language models.
|
| 2160 |
+
arXiv:2108.07732
|
| 2161 |
+
, 2021.
|
| 2162 |
+
Baker et al. (2022)
|
| 2163 |
+
Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune.
|
| 2164 |
+
Video pretraining (vpt): Learning to act by watching unlabeled online videos.
|
| 2165 |
+
arXiv:2206.11795
|
| 2166 |
+
, 2022.
|
| 2167 |
+
Besta et al. (2023)
|
| 2168 |
+
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Michal Podstawski, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler.
|
| 2169 |
+
Graph of thoughts: Solving elaborate problems with large language models.
|
| 2170 |
+
arXiv:2308.09687
|
| 2171 |
+
, 2023.
|
| 2172 |
+
Bowman et al. (2015)
|
| 2173 |
+
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning.
|
| 2174 |
+
A large annotated corpus for learning natural language inference.
|
| 2175 |
+
In
|
| 2176 |
+
EMNLP
|
| 2177 |
+
, 2015.
|
| 2178 |
+
Brown et al. (2020)
|
| 2179 |
+
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei.
|
| 2180 |
+
Language models are few-shot learners.
|
| 2181 |
+
In
|
| 2182 |
+
NeurIPS
|
| 2183 |
+
, 2020.
|
| 2184 |
+
Campbell et al. (2002)
|
| 2185 |
+
Murray Campbell, A Joseph Hoane Jr, and Feng-hsiung Hsu.
|
| 2186 |
+
Deep blue.
|
| 2187 |
+
Artificial intelligence
|
| 2188 |
+
, 2002.
|
| 2189 |
+
Chen et al. (2022a)
|
| 2190 |
+
Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen.
|
| 2191 |
+
Codet: Code generation with generated tests.
|
| 2192 |
+
arXiv:2207.10397
|
| 2193 |
+
, 2022a.
|
| 2194 |
+
Chen et al. (2021)
|
| 2195 |
+
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al.
|
| 2196 |
+
Evaluating large language models trained on code.
|
| 2197 |
+
arXiv:2107.03374
|
| 2198 |
+
, 2021.
|
| 2199 |
+
Chen et al. (2022b)
|
| 2200 |
+
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen.
|
| 2201 |
+
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks.
|
| 2202 |
+
arXiv preprint arXiv:2211.12588
|
| 2203 |
+
, 2022b.
|
| 2204 |
+
Chowdhery et al. (2022)
|
| 2205 |
+
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al.
|
| 2206 |
+
Palm: Scaling language modeling with pathways.
|
| 2207 |
+
arXiv:2204.02311
|
| 2208 |
+
, 2022.
|
| 2209 |
+
Cobbe et al. (2021)
|
| 2210 |
+
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al.
|
| 2211 |
+
Training verifiers to solve math word problems.
|
| 2212 |
+
arXiv:2110.14168
|
| 2213 |
+
, 2021.
|
| 2214 |
+
Deng et al. (2023)
|
| 2215 |
+
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su.
|
| 2216 |
+
Mind2web: Towards a generalist agent for the web.
|
| 2217 |
+
arXiv:2306.06070
|
| 2218 |
+
, 2023.
|
| 2219 |
+
Driess et al. (2023)
|
| 2220 |
+
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence.
|
| 2221 |
+
Palm-e: An embodied multimodal language model.
|
| 2222 |
+
arXiv:2303.03378
|
| 2223 |
+
, 2023.
|
| 2224 |
+
Du et al. (2023)
|
| 2225 |
+
Yilun Du, Mengjiao Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Joshua B. Tenenbaum, Dale Schuurmans, and Pieter Abbeel.
|
| 2226 |
+
Learning universal policies via text-guided video generation.
|
| 2227 |
+
arXiv:2302.00111
|
| 2228 |
+
, 2023.
|
| 2229 |
+
Evans (2010)
|
| 2230 |
+
Jonathan St BT Evans.
|
| 2231 |
+
Intuition and reasoning: A dual-process perspective.
|
| 2232 |
+
Psychological Inquiry
|
| 2233 |
+
, 2010.
|
| 2234 |
+
Fan et al. (2022)
|
| 2235 |
+
Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar.
|
| 2236 |
+
Minedojo: Building open-ended embodied agents with internet-scale knowledge.
|
| 2237 |
+
In
|
| 2238 |
+
NeurIPS Datasets and Benchmarks Track
|
| 2239 |
+
, 2022.
|
| 2240 |
+
Furuta et al. (2023)
|
| 2241 |
+
Hiroki Furuta, Ofir Nachum, Kuang-Huei Lee, Yutaka Matsuo, Shixiang Shane Gu, and Izzeddin Gur.
|
| 2242 |
+
Multimodal web navigation with instruction-finetuned foundation models.
|
| 2243 |
+
arXiv preprint arXiv:2305.11854
|
| 2244 |
+
, 2023.
|
| 2245 |
+
Gao et al. (2022)
|
| 2246 |
+
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig.
|
| 2247 |
+
Pal: Program-aided language models.
|
| 2248 |
+
arXiv preprint arXiv:2211.10435
|
| 2249 |
+
, 2022.
|
| 2250 |
+
Guo et al. (2018)
|
| 2251 |
+
Jiaxian Guo, Sidi Lu, Han Cai, Weinan Zhang, Yong Yu, and Jun Wang.
|
| 2252 |
+
Long text generation via adversarial training with leaked information.
|
| 2253 |
+
AAAI
|
| 2254 |
+
, 2018.
|
| 2255 |
+
Guss et al. (2019)
|
| 2256 |
+
William H. Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela Veloso, and Ruslan Salakhutdinov.
|
| 2257 |
+
Minerl: A large-scale dataset of minecraft demonstrations.
|
| 2258 |
+
In
|
| 2259 |
+
IJCAI
|
| 2260 |
+
, 2019.
|
| 2261 |
+
Hafner et al. (2019)
|
| 2262 |
+
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson.
|
| 2263 |
+
Learning latent dynamics for planning from pixels.
|
| 2264 |
+
In
|
| 2265 |
+
ICML
|
| 2266 |
+
, 2019.
|
| 2267 |
+
Hafner et al. (2023)
|
| 2268 |
+
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap.
|
| 2269 |
+
Mastering diverse domains through world models.
|
| 2270 |
+
arXiv:2301.04104
|
| 2271 |
+
, 2023.
|
| 2272 |
+
Hao et al. (2023)
|
| 2273 |
+
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu.
|
| 2274 |
+
Reasoning with language model is planning with world model.
|
| 2275 |
+
arXiv:2305.14992
|
| 2276 |
+
, 2023.
|
| 2277 |
+
Huang et al. (2023)
|
| 2278 |
+
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou.
|
| 2279 |
+
Large language models cannot self-correct reasoning yet.
|
| 2280 |
+
arXiv:2310.01798
|
| 2281 |
+
, 2023.
|
| 2282 |
+
Huang et al. (2022)
|
| 2283 |
+
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al.
|
| 2284 |
+
Inner monologue: Embodied reasoning through planning with language models.
|
| 2285 |
+
arXiv:2207.05608
|
| 2286 |
+
, 2022.
|
| 2287 |
+
Jiang et al. (2018)
|
| 2288 |
+
D. Jiang, E. Ekwedike, and H. Liu.
|
| 2289 |
+
Feedback-based tree search for reinforcement learning.
|
| 2290 |
+
In
|
| 2291 |
+
ICML
|
| 2292 |
+
, 2018.
|
| 2293 |
+
Kocsis & Szepesvári (2006)
|
| 2294 |
+
Levente Kocsis and Csaba Szepesvári.
|
| 2295 |
+
Bandit based monte-carlo planning.
|
| 2296 |
+
In
|
| 2297 |
+
ECML
|
| 2298 |
+
, 2006.
|
| 2299 |
+
Kojima et al. (2022)
|
| 2300 |
+
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa.
|
| 2301 |
+
Large language models are zero-shot reasoners.
|
| 2302 |
+
arXiv:2205.11916
|
| 2303 |
+
, 2022.
|
| 2304 |
+
LaValle et al. (2001)
|
| 2305 |
+
Steven M LaValle, James J Kuffner, BR Donald, et al.
|
| 2306 |
+
Rapidly-exploring random trees: Progress and prospects.
|
| 2307 |
+
Algorithmic and computational robotics: new directions
|
| 2308 |
+
, 2001.
|
| 2309 |
+
Liu et al. (2018)
|
| 2310 |
+
Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang.
|
| 2311 |
+
Reinforcement learning on web interfaces using workflow-guided exploration.
|
| 2312 |
+
In
|
| 2313 |
+
ICLR
|
| 2314 |
+
, 2018.
|
| 2315 |
+
Liu et al. (2023)
|
| 2316 |
+
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang.
|
| 2317 |
+
Agentbench: Evaluating llms as agents.
|
| 2318 |
+
arXiv:2308.03688
|
| 2319 |
+
, 2023.
|
| 2320 |
+
Madaan et al. (2023)
|
| 2321 |
+
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark.
|
| 2322 |
+
Self-refine: Iterative refinement with self-feedback.
|
| 2323 |
+
arXiv:2303.17651
|
| 2324 |
+
, 2023.
|
| 2325 |
+
Nallapati et al. (2016)
|
| 2326 |
+
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Caglar Gulcehre, and Bing Xiang.
|
| 2327 |
+
Abstractive text summarization using sequence-to-sequence rnns and beyond.
|
| 2328 |
+
In
|
| 2329 |
+
SIGNLL
|
| 2330 |
+
, 2016.
|
| 2331 |
+
Nye et al. (2021)
|
| 2332 |
+
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al.
|
| 2333 |
+
Show your work: Scratchpads for intermediate computation with language models.
|
| 2334 |
+
arXiv:2112.00114
|
| 2335 |
+
, 2021.
|
| 2336 |
+
OpenAI (2023)
|
| 2337 |
+
OpenAI.
|
| 2338 |
+
Gpt-4 technical report.
|
| 2339 |
+
arXiv:2303.08774
|
| 2340 |
+
, 2023.
|
| 2341 |
+
Saparov & He (2022)
|
| 2342 |
+
Abulhair Saparov and He He.
|
| 2343 |
+
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought.
|
| 2344 |
+
arXiv:2210.01240
|
| 2345 |
+
, 2022.
|
| 2346 |
+
Schick et al. (2023)
|
| 2347 |
+
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom.
|
| 2348 |
+
Toolformer: Language models can teach themselves to use tools.
|
| 2349 |
+
arXiv:2302.04761
|
| 2350 |
+
, 2023.
|
| 2351 |
+
Shen et al. (2023)
|
| 2352 |
+
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang.
|
| 2353 |
+
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface.
|
| 2354 |
+
arXiv:2303.17580
|
| 2355 |
+
, 2023.
|
| 2356 |
+
Shinn et al. (2023)
|
| 2357 |
+
Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao.
|
| 2358 |
+
Reflexion: Language agents with verbal reinforcement learning.
|
| 2359 |
+
arXiv:2303.11366
|
| 2360 |
+
, 2023.
|
| 2361 |
+
Shridhar et al. (2020)
|
| 2362 |
+
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht.
|
| 2363 |
+
Alfworld: Aligning text and embodied environments for interactive learning.
|
| 2364 |
+
arXiv:2010.03768
|
| 2365 |
+
, 2020.
|
| 2366 |
+
Silver et al. (2016)
|
| 2367 |
+
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al.
|
| 2368 |
+
Mastering the game of go with deep neural networks and tree search.
|
| 2369 |
+
nature
|
| 2370 |
+
, 2016.
|
| 2371 |
+
Silver et al. (2017)
|
| 2372 |
+
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy P. Lillicrap, Fan Hui, L. Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis.
|
| 2373 |
+
Mastering the game of go without human knowledge.
|
| 2374 |
+
Nature
|
| 2375 |
+
, 2017.
|
| 2376 |
+
Sloman (1996)
|
| 2377 |
+
Steven A. Sloman.
|
| 2378 |
+
The empirical case for two systems of reasoning.
|
| 2379 |
+
Psychological Bulletin
|
| 2380 |
+
, 1996.
|
| 2381 |
+
Sun et al. (2023)
|
| 2382 |
+
Haotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai, and Chao Zhang.
|
| 2383 |
+
Adaplanner: Adaptive planning from feedback with language models.
|
| 2384 |
+
arXiv:2305.16653
|
| 2385 |
+
, 2023.
|
| 2386 |
+
Surís et al. (2023)
|
| 2387 |
+
Dídac Surís, Sachit Menon, and Carl Vondrick.
|
| 2388 |
+
Vipergpt: Visual inference via python execution for reasoning.
|
| 2389 |
+
arXiv preprint arXiv:2303.08128
|
| 2390 |
+
, 2023.
|
| 2391 |
+
Świechowski et al. (2023)
|
| 2392 |
+
Maciej Świechowski, Konrad Godlewski, Bartosz Sawicki, and Jacek Mańdziuk.
|
| 2393 |
+
Monte carlo tree search: A review of recent modifications and applications.
|
| 2394 |
+
Artificial Intelligence Review
|
| 2395 |
+
, 2023.
|
| 2396 |
+
Touvron et al. (2023)
|
| 2397 |
+
Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Cantón Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony S. Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel M. Kloumann, A. V. Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, R. Subramanian, Xia Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zhengxu Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and
|
| 2398 |
+
Thomas Scialom.
|
| 2399 |
+
Llama 2: Open foundation and fine-tuned chat models.
|
| 2400 |
+
arXiv:2307.09288
|
| 2401 |
+
, 2023.
|
| 2402 |
+
Vodopivec et al. (2017)
|
| 2403 |
+
Tom Vodopivec, Spyridon Samothrakis, and Branko Ster.
|
| 2404 |
+
On monte carlo tree search and reinforcement learning.
|
| 2405 |
+
Journal of Artificial Intelligence Research
|
| 2406 |
+
, 2017.
|
| 2407 |
+
Wang et al. (2023)
|
| 2408 |
+
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar.
|
| 2409 |
+
Voyager: An open-ended embodied agent with large language models.
|
| 2410 |
+
arXiv:2305.16291
|
| 2411 |
+
, 2023.
|
| 2412 |
+
Wang et al. (2022)
|
| 2413 |
+
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou.
|
| 2414 |
+
Self-consistency improves chain of thought reasoning in language models.
|
| 2415 |
+
arXiv:2203.11171
|
| 2416 |
+
, 2022.
|
| 2417 |
+
Wei et al. (2022)
|
| 2418 |
+
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou.
|
| 2419 |
+
Chain of thought prompting elicits reasoning in large language models.
|
| 2420 |
+
arXiv:2201.11903
|
| 2421 |
+
, 2022.
|
| 2422 |
+
Wooldridge & Jennings (1995)
|
| 2423 |
+
Michael Wooldridge and Nicholas R Jennings.
|
| 2424 |
+
Intelligent agents: Theory and practice.
|
| 2425 |
+
The knowledge engineering review
|
| 2426 |
+
, 1995.
|
| 2427 |
+
Wu et al. (2023)
|
| 2428 |
+
Philipp Wu, Alejandro Escontrela, Danijar Hafner, Pieter Abbeel, and Ken Goldberg.
|
| 2429 |
+
Daydreamer: World models for physical robot learning.
|
| 2430 |
+
In
|
| 2431 |
+
CoRL
|
| 2432 |
+
. PMLR, 2023.
|
| 2433 |
+
Xie et al. (2023)
|
| 2434 |
+
Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao, Min-Yen Kan, Junxian He, and Qizhe Xie.
|
| 2435 |
+
Decomposition enhances reasoning via self-evaluation guided decoding.
|
| 2436 |
+
arXiv:2305.00633
|
| 2437 |
+
, 2023.
|
| 2438 |
+
Yang et al. (2018)
|
| 2439 |
+
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning.
|
| 2440 |
+
Hotpotqa: A dataset for diverse, explainable multi-hop question answering.
|
| 2441 |
+
arXiv:1809.09600
|
| 2442 |
+
, 2018.
|
| 2443 |
+
Yao et al. (2022)
|
| 2444 |
+
Shunyu Yao, Howard Chen, John Yang, and Karthik R Narasimhan.
|
| 2445 |
+
Webshop: Towards scalable real-world web interaction with grounded language agents.
|
| 2446 |
+
In
|
| 2447 |
+
NeurIPS
|
| 2448 |
+
, 2022.
|
| 2449 |
+
Yao et al. (2023a)
|
| 2450 |
+
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan.
|
| 2451 |
+
Tree of thoughts: Deliberate problem solving with large language models.
|
| 2452 |
+
arXiv:2305.10601
|
| 2453 |
+
, 2023a.
|
| 2454 |
+
Yao et al. (2023b)
|
| 2455 |
+
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao.
|
| 2456 |
+
ReAct: Synergizing reasoning and acting in language models.
|
| 2457 |
+
In
|
| 2458 |
+
ICLR
|
| 2459 |
+
, 2023b.
|
| 2460 |
+
Yao et al. (2023c)
|
| 2461 |
+
Weiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu, Yihao Feng, Le Xue, Rithesh Murthy, Zeyuan Chen, Jianguo Zhang, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese.
|
| 2462 |
+
Retroformer: Retrospective large language agents with policy gradient optimization.
|
| 2463 |
+
arXiv preprint arXiv:2308.02151
|
| 2464 |
+
, 2023c.
|
| 2465 |
+
Ye et al. (2021)
|
| 2466 |
+
Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao.
|
| 2467 |
+
Mastering atari games with limited data.
|
| 2468 |
+
In
|
| 2469 |
+
NeurIPS
|
| 2470 |
+
, 2021.
|
| 2471 |
+
Zhou et al. (2022)
|
| 2472 |
+
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi.
|
| 2473 |
+
Least-to-most prompting enables complex reasoning in large language models.
|
| 2474 |
+
arXiv:2205.10625
|
| 2475 |
+
, 2022.
|
| 2476 |
+
Zhu et al. (2023)
|
| 2477 |
+
Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, Yu Qiao, Zhaoxiang Zhang, and Jifeng Dai.
|
| 2478 |
+
Ghost in the minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory.
|
| 2479 |
+
arXiv:2305.17144
|
| 2480 |
+
, 2023.
|
| 2481 |
+
7
|
| 2482 |
+
Appendix
|
| 2483 |
+
The appendix is organized as follows. First in Sec.
|
| 2484 |
+
A
|
| 2485 |
+
, we show the pseudocode of our proposed algorithm, LATS; then in Sec.
|
| 2486 |
+
B
|
| 2487 |
+
, we provide further discussion of our method and its limitations, future direction and broader impact; then in Sec.
|
| 2488 |
+
C
|
| 2489 |
+
we provide additional experimental results; then in Sec.
|
| 2490 |
+
D
|
| 2491 |
+
, we specify the environment details in our experiments; finally, we list our prompts used for the three environments in Sec.
|
| 2492 |
+
E
|
| 2493 |
+
(HotPotQA), Sec.
|
| 2494 |
+
F
|
| 2495 |
+
(Programming) and Sec.
|
| 2496 |
+
G
|
| 2497 |
+
(Webshop) respectively.
|
| 2498 |
+
Appendix A
|
| 2499 |
+
LATS Pseudocode
|
| 2500 |
+
Alg.
|
| 2501 |
+
1
|
| 2502 |
+
shows the pseudocode of our algorithm LATS. Nodes are stored explicitly in the memory. Unless otherwise specified, in all experiments we use
|
| 2503 |
+
n
|
| 2504 |
+
=
|
| 2505 |
+
5
|
| 2506 |
+
𝑛
|
| 2507 |
+
5
|
| 2508 |
+
n=5
|
| 2509 |
+
and
|
| 2510 |
+
w
|
| 2511 |
+
=
|
| 2512 |
+
1
|
| 2513 |
+
𝑤
|
| 2514 |
+
1
|
| 2515 |
+
w=1
|
| 2516 |
+
.
|
| 2517 |
+
Algorithm 1
|
| 2518 |
+
LATS
|
| 2519 |
+
|
| 2520 |
+
(
|
| 2521 |
+
S
|
| 2522 |
+
0
|
| 2523 |
+
,
|
| 2524 |
+
p
|
| 2525 |
+
θ
|
| 2526 |
+
,
|
| 2527 |
+
p
|
| 2528 |
+
V
|
| 2529 |
+
,
|
| 2530 |
+
p
|
| 2531 |
+
ref
|
| 2532 |
+
,
|
| 2533 |
+
d
|
| 2534 |
+
,
|
| 2535 |
+
k
|
| 2536 |
+
,
|
| 2537 |
+
n
|
| 2538 |
+
,
|
| 2539 |
+
w
|
| 2540 |
+
)
|
| 2541 |
+
LATS
|
| 2542 |
+
subscript
|
| 2543 |
+
𝑆
|
| 2544 |
+
0
|
| 2545 |
+
subscript
|
| 2546 |
+
𝑝
|
| 2547 |
+
𝜃
|
| 2548 |
+
subscript
|
| 2549 |
+
𝑝
|
| 2550 |
+
𝑉
|
| 2551 |
+
subscript
|
| 2552 |
+
𝑝
|
| 2553 |
+
ref
|
| 2554 |
+
𝑑
|
| 2555 |
+
𝑘
|
| 2556 |
+
𝑛
|
| 2557 |
+
𝑤
|
| 2558 |
+
\operatorname{LATS}(S_{0},p_{\theta},{p_{V}},p_{\text{ref}},d,k,n,w)
|
| 2559 |
+
Initial state
|
| 2560 |
+
s
|
| 2561 |
+
1
|
| 2562 |
+
subscript
|
| 2563 |
+
𝑠
|
| 2564 |
+
1
|
| 2565 |
+
s_{1}
|
| 2566 |
+
, action generator
|
| 2567 |
+
p
|
| 2568 |
+
θ
|
| 2569 |
+
subscript
|
| 2570 |
+
𝑝
|
| 2571 |
+
𝜃
|
| 2572 |
+
p_{\theta}
|
| 2573 |
+
, value function
|
| 2574 |
+
p
|
| 2575 |
+
V
|
| 2576 |
+
subscript
|
| 2577 |
+
𝑝
|
| 2578 |
+
𝑉
|
| 2579 |
+
p_{V}
|
| 2580 |
+
, reflection generator
|
| 2581 |
+
p
|
| 2582 |
+
ref
|
| 2583 |
+
subscript
|
| 2584 |
+
𝑝
|
| 2585 |
+
ref
|
| 2586 |
+
p_{\text{ref}}
|
| 2587 |
+
, number of generated actions
|
| 2588 |
+
n
|
| 2589 |
+
𝑛
|
| 2590 |
+
n
|
| 2591 |
+
, depth limit
|
| 2592 |
+
L
|
| 2593 |
+
𝐿
|
| 2594 |
+
L
|
| 2595 |
+
, number of roll-outs
|
| 2596 |
+
K
|
| 2597 |
+
𝐾
|
| 2598 |
+
K
|
| 2599 |
+
, context
|
| 2600 |
+
c
|
| 2601 |
+
𝑐
|
| 2602 |
+
c
|
| 2603 |
+
, and exploration weight
|
| 2604 |
+
w
|
| 2605 |
+
𝑤
|
| 2606 |
+
w
|
| 2607 |
+
Initialize action space
|
| 2608 |
+
A
|
| 2609 |
+
𝐴
|
| 2610 |
+
A
|
| 2611 |
+
, observation space
|
| 2612 |
+
O
|
| 2613 |
+
𝑂
|
| 2614 |
+
O
|
| 2615 |
+
Initialize the state-action value function
|
| 2616 |
+
p
|
| 2617 |
+
V
|
| 2618 |
+
:
|
| 2619 |
+
S
|
| 2620 |
+
×
|
| 2621 |
+
A
|
| 2622 |
+
↦
|
| 2623 |
+
ℝ
|
| 2624 |
+
:
|
| 2625 |
+
subscript
|
| 2626 |
+
𝑝
|
| 2627 |
+
𝑉
|
| 2628 |
+
maps-to
|
| 2629 |
+
𝑆
|
| 2630 |
+
𝐴
|
| 2631 |
+
ℝ
|
| 2632 |
+
{p_{V}}:S\times A\mapsto\mathbb{R}
|
| 2633 |
+
and visit counter
|
| 2634 |
+
N
|
| 2635 |
+
:
|
| 2636 |
+
S
|
| 2637 |
+
↦
|
| 2638 |
+
ℕ
|
| 2639 |
+
:
|
| 2640 |
+
𝑁
|
| 2641 |
+
maps-to
|
| 2642 |
+
𝑆
|
| 2643 |
+
ℕ
|
| 2644 |
+
{N}:S\mapsto\mathbb{N}
|
| 2645 |
+
to zero
|
| 2646 |
+
for
|
| 2647 |
+
k
|
| 2648 |
+
←
|
| 2649 |
+
0
|
| 2650 |
+
,
|
| 2651 |
+
…
|
| 2652 |
+
,
|
| 2653 |
+
K
|
| 2654 |
+
−
|
| 2655 |
+
1
|
| 2656 |
+
←
|
| 2657 |
+
𝑘
|
| 2658 |
+
0
|
| 2659 |
+
…
|
| 2660 |
+
𝐾
|
| 2661 |
+
1
|
| 2662 |
+
k\leftarrow 0,\dots,K-1
|
| 2663 |
+
do
|
| 2664 |
+
for
|
| 2665 |
+
t
|
| 2666 |
+
←
|
| 2667 |
+
0
|
| 2668 |
+
,
|
| 2669 |
+
…
|
| 2670 |
+
,
|
| 2671 |
+
L
|
| 2672 |
+
−
|
| 2673 |
+
1
|
| 2674 |
+
←
|
| 2675 |
+
𝑡
|
| 2676 |
+
0
|
| 2677 |
+
…
|
| 2678 |
+
𝐿
|
| 2679 |
+
1
|
| 2680 |
+
t\leftarrow 0,\dots,L-1
|
| 2681 |
+
do
|
| 2682 |
+
if
|
| 2683 |
+
s
|
| 2684 |
+
t
|
| 2685 |
+
subscript
|
| 2686 |
+
𝑠
|
| 2687 |
+
𝑡
|
| 2688 |
+
s_{t}
|
| 2689 |
+
not terminal
|
| 2690 |
+
then
|
| 2691 |
+
▷
|
| 2692 |
+
▷
|
| 2693 |
+
\triangleright
|
| 2694 |
+
Expansion & Simulation
|
| 2695 |
+
for
|
| 2696 |
+
i
|
| 2697 |
+
←
|
| 2698 |
+
1
|
| 2699 |
+
,
|
| 2700 |
+
…
|
| 2701 |
+
,
|
| 2702 |
+
n
|
| 2703 |
+
←
|
| 2704 |
+
𝑖
|
| 2705 |
+
1
|
| 2706 |
+
…
|
| 2707 |
+
𝑛
|
| 2708 |
+
i\leftarrow 1,\dots,n
|
| 2709 |
+
do
|
| 2710 |
+
Sample
|
| 2711 |
+
a
|
| 2712 |
+
t
|
| 2713 |
+
(
|
| 2714 |
+
i
|
| 2715 |
+
)
|
| 2716 |
+
∼
|
| 2717 |
+
p
|
| 2718 |
+
θ
|
| 2719 |
+
|
| 2720 |
+
(
|
| 2721 |
+
a
|
| 2722 |
+
∣
|
| 2723 |
+
s
|
| 2724 |
+
t
|
| 2725 |
+
)
|
| 2726 |
+
similar-to
|
| 2727 |
+
superscript
|
| 2728 |
+
subscript
|
| 2729 |
+
𝑎
|
| 2730 |
+
𝑡
|
| 2731 |
+
𝑖
|
| 2732 |
+
subscript
|
| 2733 |
+
𝑝
|
| 2734 |
+
𝜃
|
| 2735 |
+
conditional
|
| 2736 |
+
𝑎
|
| 2737 |
+
subscript
|
| 2738 |
+
𝑠
|
| 2739 |
+
𝑡
|
| 2740 |
+
a_{t}^{(i)}\sim p_{\theta}(a\mid s_{t})
|
| 2741 |
+
Get
|
| 2742 |
+
o
|
| 2743 |
+
t
|
| 2744 |
+
(
|
| 2745 |
+
i
|
| 2746 |
+
)
|
| 2747 |
+
superscript
|
| 2748 |
+
subscript
|
| 2749 |
+
𝑜
|
| 2750 |
+
𝑡
|
| 2751 |
+
𝑖
|
| 2752 |
+
o_{t}^{(i)}
|
| 2753 |
+
from environment,
|
| 2754 |
+
s
|
| 2755 |
+
t
|
| 2756 |
+
+
|
| 2757 |
+
1
|
| 2758 |
+
(
|
| 2759 |
+
i
|
| 2760 |
+
)
|
| 2761 |
+
←
|
| 2762 |
+
(
|
| 2763 |
+
c
|
| 2764 |
+
t
|
| 2765 |
+
(
|
| 2766 |
+
i
|
| 2767 |
+
)
|
| 2768 |
+
,
|
| 2769 |
+
o
|
| 2770 |
+
t
|
| 2771 |
+
(
|
| 2772 |
+
i
|
| 2773 |
+
)
|
| 2774 |
+
,
|
| 2775 |
+
a
|
| 2776 |
+
t
|
| 2777 |
+
(
|
| 2778 |
+
i
|
| 2779 |
+
)
|
| 2780 |
+
)
|
| 2781 |
+
←
|
| 2782 |
+
superscript
|
| 2783 |
+
subscript
|
| 2784 |
+
𝑠
|
| 2785 |
+
𝑡
|
| 2786 |
+
1
|
| 2787 |
+
𝑖
|
| 2788 |
+
superscript
|
| 2789 |
+
subscript
|
| 2790 |
+
𝑐
|
| 2791 |
+
𝑡
|
| 2792 |
+
𝑖
|
| 2793 |
+
superscript
|
| 2794 |
+
subscript
|
| 2795 |
+
𝑜
|
| 2796 |
+
𝑡
|
| 2797 |
+
𝑖
|
| 2798 |
+
superscript
|
| 2799 |
+
subscript
|
| 2800 |
+
𝑎
|
| 2801 |
+
𝑡
|
| 2802 |
+
𝑖
|
| 2803 |
+
s_{t+1}^{(i)}\leftarrow(c_{t}^{(i)},o_{t}^{(i)},a_{t}^{(i)})
|
| 2804 |
+
,
|
| 2805 |
+
c
|
| 2806 |
+
t
|
| 2807 |
+
+
|
| 2808 |
+
1
|
| 2809 |
+
(
|
| 2810 |
+
i
|
| 2811 |
+
)
|
| 2812 |
+
←
|
| 2813 |
+
(
|
| 2814 |
+
o
|
| 2815 |
+
t
|
| 2816 |
+
(
|
| 2817 |
+
i
|
| 2818 |
+
)
|
| 2819 |
+
,
|
| 2820 |
+
a
|
| 2821 |
+
t
|
| 2822 |
+
(
|
| 2823 |
+
i
|
| 2824 |
+
)
|
| 2825 |
+
)
|
| 2826 |
+
←
|
| 2827 |
+
superscript
|
| 2828 |
+
subscript
|
| 2829 |
+
𝑐
|
| 2830 |
+
𝑡
|
| 2831 |
+
1
|
| 2832 |
+
𝑖
|
| 2833 |
+
superscript
|
| 2834 |
+
subscript
|
| 2835 |
+
𝑜
|
| 2836 |
+
𝑡
|
| 2837 |
+
𝑖
|
| 2838 |
+
superscript
|
| 2839 |
+
subscript
|
| 2840 |
+
𝑎
|
| 2841 |
+
𝑡
|
| 2842 |
+
𝑖
|
| 2843 |
+
c_{t+1}^{(i)}\leftarrow(o_{t}^{(i)},a_{t}^{(i)})
|
| 2844 |
+
Evaluate
|
| 2845 |
+
V
|
| 2846 |
+
t
|
| 2847 |
+
(
|
| 2848 |
+
i
|
| 2849 |
+
)
|
| 2850 |
+
∼
|
| 2851 |
+
p
|
| 2852 |
+
V
|
| 2853 |
+
|
| 2854 |
+
(
|
| 2855 |
+
s
|
| 2856 |
+
t
|
| 2857 |
+
(
|
| 2858 |
+
i
|
| 2859 |
+
)
|
| 2860 |
+
)
|
| 2861 |
+
similar-to
|
| 2862 |
+
superscript
|
| 2863 |
+
subscript
|
| 2864 |
+
𝑉
|
| 2865 |
+
𝑡
|
| 2866 |
+
𝑖
|
| 2867 |
+
subscript
|
| 2868 |
+
𝑝
|
| 2869 |
+
𝑉
|
| 2870 |
+
superscript
|
| 2871 |
+
subscript
|
| 2872 |
+
𝑠
|
| 2873 |
+
𝑡
|
| 2874 |
+
𝑖
|
| 2875 |
+
{V}_{t}^{(i)}\sim{p_{V}}(s_{t}^{(i)})
|
| 2876 |
+
▷
|
| 2877 |
+
▷
|
| 2878 |
+
\triangleright
|
| 2879 |
+
Evaluation
|
| 2880 |
+
V
|
| 2881 |
+
|
| 2882 |
+
(
|
| 2883 |
+
s
|
| 2884 |
+
t
|
| 2885 |
+
)
|
| 2886 |
+
←
|
| 2887 |
+
V
|
| 2888 |
+
t
|
| 2889 |
+
(
|
| 2890 |
+
i
|
| 2891 |
+
)
|
| 2892 |
+
←
|
| 2893 |
+
𝑉
|
| 2894 |
+
subscript
|
| 2895 |
+
𝑠
|
| 2896 |
+
𝑡
|
| 2897 |
+
superscript
|
| 2898 |
+
subscript
|
| 2899 |
+
𝑉
|
| 2900 |
+
𝑡
|
| 2901 |
+
𝑖
|
| 2902 |
+
{V}(s_{t})\leftarrow{V}_{t}^{(i)}
|
| 2903 |
+
Add
|
| 2904 |
+
s
|
| 2905 |
+
t
|
| 2906 |
+
(
|
| 2907 |
+
i
|
| 2908 |
+
)
|
| 2909 |
+
superscript
|
| 2910 |
+
subscript
|
| 2911 |
+
𝑠
|
| 2912 |
+
𝑡
|
| 2913 |
+
𝑖
|
| 2914 |
+
s_{t}^{(i)}
|
| 2915 |
+
to children
|
| 2916 |
+
end
|
| 2917 |
+
for
|
| 2918 |
+
end
|
| 2919 |
+
if
|
| 2920 |
+
if
|
| 2921 |
+
s
|
| 2922 |
+
t
|
| 2923 |
+
subscript
|
| 2924 |
+
𝑠
|
| 2925 |
+
𝑡
|
| 2926 |
+
s_{t}
|
| 2927 |
+
is terminal
|
| 2928 |
+
then
|
| 2929 |
+
▷
|
| 2930 |
+
▷
|
| 2931 |
+
\triangleright
|
| 2932 |
+
Reflection
|
| 2933 |
+
Get
|
| 2934 |
+
r
|
| 2935 |
+
𝑟
|
| 2936 |
+
r
|
| 2937 |
+
from environment
|
| 2938 |
+
if
|
| 2939 |
+
r
|
| 2940 |
+
𝑟
|
| 2941 |
+
r
|
| 2942 |
+
not success
|
| 2943 |
+
then
|
| 2944 |
+
reflection
|
| 2945 |
+
←
|
| 2946 |
+
p
|
| 2947 |
+
ref
|
| 2948 |
+
|
| 2949 |
+
(
|
| 2950 |
+
c
|
| 2951 |
+
t
|
| 2952 |
+
)
|
| 2953 |
+
←
|
| 2954 |
+
reflection
|
| 2955 |
+
subscript
|
| 2956 |
+
𝑝
|
| 2957 |
+
ref
|
| 2958 |
+
subscript
|
| 2959 |
+
𝑐
|
| 2960 |
+
𝑡
|
| 2961 |
+
\text{reflection}\leftarrow p_{\text{ref}}(c_{t})
|
| 2962 |
+
c
|
| 2963 |
+
←
|
| 2964 |
+
reflection
|
| 2965 |
+
←
|
| 2966 |
+
𝑐
|
| 2967 |
+
reflection
|
| 2968 |
+
c\leftarrow\text{reflection}
|
| 2969 |
+
end
|
| 2970 |
+
if
|
| 2971 |
+
end
|
| 2972 |
+
if
|
| 2973 |
+
a
|
| 2974 |
+
t
|
| 2975 |
+
←
|
| 2976 |
+
arg
|
| 2977 |
+
|
| 2978 |
+
max
|
| 2979 |
+
a
|
| 2980 |
+
∈
|
| 2981 |
+
e
|
| 2982 |
+
|
| 2983 |
+
(
|
| 2984 |
+
s
|
| 2985 |
+
t
|
| 2986 |
+
)
|
| 2987 |
+
|
| 2988 |
+
[
|
| 2989 |
+
V
|
| 2990 |
+
|
| 2991 |
+
(
|
| 2992 |
+
s
|
| 2993 |
+
t
|
| 2994 |
+
)
|
| 2995 |
+
+
|
| 2996 |
+
w
|
| 2997 |
+
|
| 2998 |
+
ln
|
| 2999 |
+
|
| 3000 |
+
N
|
| 3001 |
+
|
| 3002 |
+
(
|
| 3003 |
+
s
|
| 3004 |
+
t
|
| 3005 |
+
−
|
| 3006 |
+
1
|
| 3007 |
+
)
|
| 3008 |
+
N
|
| 3009 |
+
|
| 3010 |
+
(
|
| 3011 |
+
s
|
| 3012 |
+
t
|
| 3013 |
+
)
|
| 3014 |
+
]
|
| 3015 |
+
←
|
| 3016 |
+
subscript
|
| 3017 |
+
𝑎
|
| 3018 |
+
𝑡
|
| 3019 |
+
subscript
|
| 3020 |
+
𝑎
|
| 3021 |
+
𝑒
|
| 3022 |
+
subscript
|
| 3023 |
+
𝑠
|
| 3024 |
+
𝑡
|
| 3025 |
+
𝑉
|
| 3026 |
+
subscript
|
| 3027 |
+
𝑠
|
| 3028 |
+
𝑡
|
| 3029 |
+
𝑤
|
| 3030 |
+
𝑁
|
| 3031 |
+
subscript
|
| 3032 |
+
𝑠
|
| 3033 |
+
𝑡
|
| 3034 |
+
1
|
| 3035 |
+
𝑁
|
| 3036 |
+
subscript
|
| 3037 |
+
𝑠
|
| 3038 |
+
𝑡
|
| 3039 |
+
a_{t}\leftarrow\arg\max_{a\in e(s_{t})}\left[{V(s_{t})}+w\sqrt{\frac{\ln{N}(s_{t-1})}{{N}(s_{t})}}\right]
|
| 3040 |
+
▷
|
| 3041 |
+
▷
|
| 3042 |
+
\triangleright
|
| 3043 |
+
Selection
|
| 3044 |
+
N
|
| 3045 |
+
|
| 3046 |
+
(
|
| 3047 |
+
s
|
| 3048 |
+
t
|
| 3049 |
+
+
|
| 3050 |
+
1
|
| 3051 |
+
)
|
| 3052 |
+
←
|
| 3053 |
+
N
|
| 3054 |
+
|
| 3055 |
+
(
|
| 3056 |
+
s
|
| 3057 |
+
t
|
| 3058 |
+
+
|
| 3059 |
+
1
|
| 3060 |
+
)
|
| 3061 |
+
+
|
| 3062 |
+
1
|
| 3063 |
+
←
|
| 3064 |
+
𝑁
|
| 3065 |
+
subscript
|
| 3066 |
+
𝑠
|
| 3067 |
+
𝑡
|
| 3068 |
+
1
|
| 3069 |
+
𝑁
|
| 3070 |
+
subscript
|
| 3071 |
+
𝑠
|
| 3072 |
+
𝑡
|
| 3073 |
+
1
|
| 3074 |
+
1
|
| 3075 |
+
{N}(s_{t+1})\leftarrow{N}(s_{t+1})+1
|
| 3076 |
+
if
|
| 3077 |
+
a
|
| 3078 |
+
t
|
| 3079 |
+
subscript
|
| 3080 |
+
𝑎
|
| 3081 |
+
𝑡
|
| 3082 |
+
a_{t}
|
| 3083 |
+
is an output action
|
| 3084 |
+
then
|
| 3085 |
+
break
|
| 3086 |
+
end
|
| 3087 |
+
for
|
| 3088 |
+
T
|
| 3089 |
+
←
|
| 3090 |
+
←
|
| 3091 |
+
𝑇
|
| 3092 |
+
absent
|
| 3093 |
+
T\leftarrow
|
| 3094 |
+
the actual number of steps
|
| 3095 |
+
for
|
| 3096 |
+
t
|
| 3097 |
+
←
|
| 3098 |
+
T
|
| 3099 |
+
−
|
| 3100 |
+
1
|
| 3101 |
+
,
|
| 3102 |
+
…
|
| 3103 |
+
,
|
| 3104 |
+
0
|
| 3105 |
+
←
|
| 3106 |
+
𝑡
|
| 3107 |
+
𝑇
|
| 3108 |
+
1
|
| 3109 |
+
…
|
| 3110 |
+
0
|
| 3111 |
+
t\leftarrow T-1,\dots,0
|
| 3112 |
+
do
|
| 3113 |
+
▷
|
| 3114 |
+
▷
|
| 3115 |
+
\triangleright
|
| 3116 |
+
Backpropagation
|
| 3117 |
+
V
|
| 3118 |
+
|
| 3119 |
+
(
|
| 3120 |
+
s
|
| 3121 |
+
t
|
| 3122 |
+
)
|
| 3123 |
+
←
|
| 3124 |
+
V
|
| 3125 |
+
|
| 3126 |
+
(
|
| 3127 |
+
s
|
| 3128 |
+
t
|
| 3129 |
+
)
|
| 3130 |
+
|
| 3131 |
+
(
|
| 3132 |
+
N
|
| 3133 |
+
|
| 3134 |
+
(
|
| 3135 |
+
s
|
| 3136 |
+
t
|
| 3137 |
+
)
|
| 3138 |
+
−
|
| 3139 |
+
1
|
| 3140 |
+
)
|
| 3141 |
+
+
|
| 3142 |
+
r
|
| 3143 |
+
N
|
| 3144 |
+
|
| 3145 |
+
(
|
| 3146 |
+
s
|
| 3147 |
+
t
|
| 3148 |
+
)
|
| 3149 |
+
←
|
| 3150 |
+
𝑉
|
| 3151 |
+
subscript
|
| 3152 |
+
𝑠
|
| 3153 |
+
𝑡
|
| 3154 |
+
𝑉
|
| 3155 |
+
subscript
|
| 3156 |
+
𝑠
|
| 3157 |
+
𝑡
|
| 3158 |
+
𝑁
|
| 3159 |
+
subscript
|
| 3160 |
+
𝑠
|
| 3161 |
+
𝑡
|
| 3162 |
+
1
|
| 3163 |
+
𝑟
|
| 3164 |
+
𝑁
|
| 3165 |
+
subscript
|
| 3166 |
+
𝑠
|
| 3167 |
+
𝑡
|
| 3168 |
+
V(s_{t})\leftarrow\frac{V(s_{t})(N(s_{t})-1)+r}{N(s_{t})}
|
| 3169 |
+
end
|
| 3170 |
+
for
|
| 3171 |
+
end
|
| 3172 |
+
for
|
| 3173 |
+
Appendix B
|
| 3174 |
+
Discussion
|
| 3175 |
+
Limitations.
|
| 3176 |
+
Although LATS can improve reasoning and decision-making, this arrives at a higher computational cost relative to simpler prompting methods like ReAct or Reflexion. The search process takes more time than standard prompting or simpler techniques, and requires greater inference costs. While such an issue is mitigated by the fact that the number of nodes
|
| 3177 |
+
n
|
| 3178 |
+
𝑛
|
| 3179 |
+
n
|
| 3180 |
+
expanded at every step provides a natural trade-off between performance and efficiency (setting
|
| 3181 |
+
n
|
| 3182 |
+
=
|
| 3183 |
+
1
|
| 3184 |
+
𝑛
|
| 3185 |
+
1
|
| 3186 |
+
n=1
|
| 3187 |
+
makes the method as effecient as ReAct with multiple trials or CoT-SC), in practice we recommend using LATS for difficult tasks like programming or for situations where performance is prioritized over efficiency. We hope that continued advancements in LLMs will reduce costs and increase the practicality of LATS.
|
| 3188 |
+
Additionally, the benchmarks we use in this paper are relatively simple and focused on decision-making, compared to the complexity of real-world interactive environments. In addition, some environments might not easily support rollbacks to previous states. However, the design of LATS is flexible and can be adjusted to various resource constraints. Using planning-based prompting methods like LATS in environments like Minecraft
|
| 3189 |
+
(Fan et al.,
|
| 3190 |
+
2022
|
| 3191 |
+
)
|
| 3192 |
+
and more reasoning benchmarks would be interesting avenues for future work.
|
| 3193 |
+
Broader impact.
|
| 3194 |
+
LATS is a framework that enhances LLM performance through interactions with an environment. This improvement in autonomous decision-making may facilitate harmful uses of LLMs. Alternatively, LATS enhances interpretability and the potential for greater alignment, as it generates understandable, high-level linguistic reasoning and actions through several rounds of decision-making and reflection, rather than relying on implicit, low-level token values.
|
| 3195 |
+
Appendix C
|
| 3196 |
+
Ablations
|
| 3197 |
+
Prompt Method
|
| 3198 |
+
HotpotQA (EM)
|
| 3199 |
+
LATS (w=0.5)
|
| 3200 |
+
0.55
|
| 3201 |
+
LATS (w=2.0)
|
| 3202 |
+
0.61
|
| 3203 |
+
LATS (d=4)
|
| 3204 |
+
0.58
|
| 3205 |
+
LATS (CoT)
|
| 3206 |
+
0.60
|
| 3207 |
+
LATS (No LM Heuristic)
|
| 3208 |
+
0.37
|
| 3209 |
+
LATS
|
| 3210 |
+
0.61
|
| 3211 |
+
Table 6:
|
| 3212 |
+
Ablation results on LATS and baseline variants in HotPotQA measured by Exact Match (EM). We test different depth
|
| 3213 |
+
d
|
| 3214 |
+
𝑑
|
| 3215 |
+
d
|
| 3216 |
+
, exploration factor
|
| 3217 |
+
w
|
| 3218 |
+
𝑤
|
| 3219 |
+
w
|
| 3220 |
+
, and versions of LATS using CoT and without the LM value function. We sample
|
| 3221 |
+
n
|
| 3222 |
+
=
|
| 3223 |
+
5
|
| 3224 |
+
𝑛
|
| 3225 |
+
5
|
| 3226 |
+
n=5
|
| 3227 |
+
and
|
| 3228 |
+
k
|
| 3229 |
+
=
|
| 3230 |
+
50
|
| 3231 |
+
𝑘
|
| 3232 |
+
50
|
| 3233 |
+
k=50
|
| 3234 |
+
trajectories.
|
| 3235 |
+
Figure 4:
|
| 3236 |
+
Performance over successive iterations on HumanEval with GPT-3.5.
|
| 3237 |
+
In this section, we ablate various designs of LATS. Experiments are conducted on HotPotQA with a maximum of
|
| 3238 |
+
k
|
| 3239 |
+
=
|
| 3240 |
+
50
|
| 3241 |
+
𝑘
|
| 3242 |
+
50
|
| 3243 |
+
k=50
|
| 3244 |
+
trajectories and sampling size of
|
| 3245 |
+
n
|
| 3246 |
+
=
|
| 3247 |
+
5
|
| 3248 |
+
𝑛
|
| 3249 |
+
5
|
| 3250 |
+
n=5
|
| 3251 |
+
and HumanEval with a maximum of
|
| 3252 |
+
k
|
| 3253 |
+
=
|
| 3254 |
+
8
|
| 3255 |
+
𝑘
|
| 3256 |
+
8
|
| 3257 |
+
k=8
|
| 3258 |
+
trajectories and sampling size of
|
| 3259 |
+
n
|
| 3260 |
+
=
|
| 3261 |
+
5
|
| 3262 |
+
𝑛
|
| 3263 |
+
5
|
| 3264 |
+
n=5
|
| 3265 |
+
. The result for HotPotQA is shown in Tab.
|
| 3266 |
+
5
|
| 3267 |
+
and HumanEval in Fig.
|
| 3268 |
+
4
|
| 3269 |
+
.
|
| 3270 |
+
Exploration weight.
|
| 3271 |
+
We find that there is lower performance on HotPotQA when the exploration weight
|
| 3272 |
+
w
|
| 3273 |
+
𝑤
|
| 3274 |
+
w
|
| 3275 |
+
in the selection formula is decreased to
|
| 3276 |
+
0.5
|
| 3277 |
+
0.5
|
| 3278 |
+
0.5
|
| 3279 |
+
, suggesting that this reduces the effectiveness of the search. Increasing
|
| 3280 |
+
w
|
| 3281 |
+
𝑤
|
| 3282 |
+
w
|
| 3283 |
+
to
|
| 3284 |
+
2.0
|
| 3285 |
+
2.0
|
| 3286 |
+
2.0
|
| 3287 |
+
does not lead to a performance improvement, but we tend to observe faster convergence. The optimal setting depends on the particular environment and complexity of the state space.
|
| 3288 |
+
Depth.
|
| 3289 |
+
In our main experiments we use a maximum depth of
|
| 3290 |
+
d
|
| 3291 |
+
=
|
| 3292 |
+
7
|
| 3293 |
+
𝑑
|
| 3294 |
+
7
|
| 3295 |
+
d=7
|
| 3296 |
+
on HotPotQA for all methods, following previous work
|
| 3297 |
+
(Yao et al.,
|
| 3298 |
+
2023b
|
| 3299 |
+
)
|
| 3300 |
+
. We ablate the effect on LATS after reducing it to
|
| 3301 |
+
d
|
| 3302 |
+
=
|
| 3303 |
+
4
|
| 3304 |
+
𝑑
|
| 3305 |
+
4
|
| 3306 |
+
d=4
|
| 3307 |
+
. This results in only a slight drop in performance. We find that most questions can be answered within four steps, and using a greater number of steps tends to force the agent into local minima and rarely improves success.
|
| 3308 |
+
LM value function.
|
| 3309 |
+
The LM value function scores states based on expected future reward. Without this heuristic, the only signal to guide search would be from environment rewards for completed trajectories, which are scarce and often binary. When we remove the evaluation operation, we observe a dramatic
|
| 3310 |
+
0.24
|
| 3311 |
+
0.24
|
| 3312 |
+
0.24
|
| 3313 |
+
drop in performance.
|
| 3314 |
+
Performance over time.
|
| 3315 |
+
To see the effects of increasing the number of trajectories sampled, we change
|
| 3316 |
+
k
|
| 3317 |
+
𝑘
|
| 3318 |
+
k
|
| 3319 |
+
to different values. We conduct this experiment on HumanEval, which has a more noticeable difference due to sampling less trajectories. The results are shown in Fig.
|
| 3320 |
+
4
|
| 3321 |
+
, in which LATS scales better with more iterations than Reflexion.
|
| 3322 |
+
Sample complexity and Token cost.
|
| 3323 |
+
One possible concern of LATS is that the tree-structured search might consume much more tokens than existing methods. To further study the computational cost of LATS compared to prior methods, we examine the sample complexity (i.e. asymptotic token cost) of all methods considered in this paper, and count the average number of nodes expanded by our method and other tree-structured methods (ToT and RAP) upon successful search on HotPotQA. We present the results in Tab.
|
| 3324 |
+
7
|
| 3325 |
+
; the result shows that our method has the same sample complexity as other tree-based search methods, and has less average number of nodes expanded upon success, which indicates less token cost. The token cost gap will be even larger when taking failed trajectories into account, since our method has higher success rate and reaches computational budget limit less often.
|
| 3326 |
+
Method
|
| 3327 |
+
Performance (
|
| 3328 |
+
↑
|
| 3329 |
+
↑
|
| 3330 |
+
\uparrow
|
| 3331 |
+
)
|
| 3332 |
+
Sample complexity (
|
| 3333 |
+
↓
|
| 3334 |
+
↓
|
| 3335 |
+
\downarrow
|
| 3336 |
+
)
|
| 3337 |
+
Avg. #nodes upon success (
|
| 3338 |
+
↓
|
| 3339 |
+
↓
|
| 3340 |
+
\downarrow
|
| 3341 |
+
)
|
| 3342 |
+
ReAct (Best
|
| 3343 |
+
k
|
| 3344 |
+
=
|
| 3345 |
+
250
|
| 3346 |
+
𝑘
|
| 3347 |
+
250
|
| 3348 |
+
k=250
|
| 3349 |
+
)
|
| 3350 |
+
0.42
|
| 3351 |
+
0.42
|
| 3352 |
+
0.42
|
| 3353 |
+
O
|
| 3354 |
+
|
| 3355 |
+
(
|
| 3356 |
+
k
|
| 3357 |
+
)
|
| 3358 |
+
𝑂
|
| 3359 |
+
𝑘
|
| 3360 |
+
O(k)
|
| 3361 |
+
N/A
|
| 3362 |
+
CoT-SC (
|
| 3363 |
+
n
|
| 3364 |
+
=
|
| 3365 |
+
1
|
| 3366 |
+
,
|
| 3367 |
+
k
|
| 3368 |
+
=
|
| 3369 |
+
250
|
| 3370 |
+
formulae-sequence
|
| 3371 |
+
𝑛
|
| 3372 |
+
1
|
| 3373 |
+
𝑘
|
| 3374 |
+
250
|
| 3375 |
+
n=1,k=250
|
| 3376 |
+
)
|
| 3377 |
+
0.40
|
| 3378 |
+
0.40
|
| 3379 |
+
0.40
|
| 3380 |
+
O
|
| 3381 |
+
|
| 3382 |
+
(
|
| 3383 |
+
k
|
| 3384 |
+
)
|
| 3385 |
+
𝑂
|
| 3386 |
+
𝑘
|
| 3387 |
+
O(k)
|
| 3388 |
+
N/A
|
| 3389 |
+
LATS (
|
| 3390 |
+
n
|
| 3391 |
+
=
|
| 3392 |
+
1
|
| 3393 |
+
,
|
| 3394 |
+
k
|
| 3395 |
+
=
|
| 3396 |
+
50
|
| 3397 |
+
formulae-sequence
|
| 3398 |
+
𝑛
|
| 3399 |
+
1
|
| 3400 |
+
𝑘
|
| 3401 |
+
50
|
| 3402 |
+
n=1,k=50
|
| 3403 |
+
)
|
| 3404 |
+
0.48
|
| 3405 |
+
0.48
|
| 3406 |
+
0.48
|
| 3407 |
+
O
|
| 3408 |
+
|
| 3409 |
+
(
|
| 3410 |
+
k
|
| 3411 |
+
)
|
| 3412 |
+
𝑂
|
| 3413 |
+
𝑘
|
| 3414 |
+
O(k)
|
| 3415 |
+
N/A
|
| 3416 |
+
ToT (ReAct)
|
| 3417 |
+
0.49
|
| 3418 |
+
0.49
|
| 3419 |
+
0.49
|
| 3420 |
+
O
|
| 3421 |
+
|
| 3422 |
+
(
|
| 3423 |
+
k
|
| 3424 |
+
|
| 3425 |
+
n
|
| 3426 |
+
)
|
| 3427 |
+
𝑂
|
| 3428 |
+
𝑘
|
| 3429 |
+
𝑛
|
| 3430 |
+
O(kn)
|
| 3431 |
+
84.05
|
| 3432 |
+
84.05
|
| 3433 |
+
84.05
|
| 3434 |
+
RAP (ReAct)
|
| 3435 |
+
0.54
|
| 3436 |
+
0.54
|
| 3437 |
+
0.54
|
| 3438 |
+
O
|
| 3439 |
+
|
| 3440 |
+
(
|
| 3441 |
+
k
|
| 3442 |
+
|
| 3443 |
+
n
|
| 3444 |
+
)
|
| 3445 |
+
𝑂
|
| 3446 |
+
𝑘
|
| 3447 |
+
𝑛
|
| 3448 |
+
O(kn)
|
| 3449 |
+
70.60
|
| 3450 |
+
70.60
|
| 3451 |
+
70.60
|
| 3452 |
+
LATS (
|
| 3453 |
+
n
|
| 3454 |
+
=
|
| 3455 |
+
5
|
| 3456 |
+
,
|
| 3457 |
+
k
|
| 3458 |
+
=
|
| 3459 |
+
50
|
| 3460 |
+
formulae-sequence
|
| 3461 |
+
𝑛
|
| 3462 |
+
5
|
| 3463 |
+
𝑘
|
| 3464 |
+
50
|
| 3465 |
+
n=5,k=50
|
| 3466 |
+
)
|
| 3467 |
+
0.61
|
| 3468 |
+
0.61
|
| 3469 |
+
0.61
|
| 3470 |
+
O
|
| 3471 |
+
|
| 3472 |
+
(
|
| 3473 |
+
k
|
| 3474 |
+
|
| 3475 |
+
n
|
| 3476 |
+
)
|
| 3477 |
+
𝑂
|
| 3478 |
+
𝑘
|
| 3479 |
+
𝑛
|
| 3480 |
+
O(kn)
|
| 3481 |
+
66.65
|
| 3482 |
+
66.65
|
| 3483 |
+
66.65
|
| 3484 |
+
Table 7:
|
| 3485 |
+
The performance, sample complexity of different methods and average number of nodes expanded upon success by methods with tree-based search.
|
| 3486 |
+
n
|
| 3487 |
+
𝑛
|
| 3488 |
+
n
|
| 3489 |
+
is the number of children nodes expanded at every step and
|
| 3490 |
+
k
|
| 3491 |
+
𝑘
|
| 3492 |
+
k
|
| 3493 |
+
is the number of trajectories. Our method has the same sample complexity as other methods with tree-based search and expands less nodes upon success, which indicates lower token cost.
|
| 3494 |
+
Appendix D
|
| 3495 |
+
Environment Details
|
| 3496 |
+
D.1
|
| 3497 |
+
HotPotQA
|
| 3498 |
+
Figure 5:
|
| 3499 |
+
Example trajectories on HotPotQA for ReAct (left) and LATS (right). LATS can sample more actions and avoid failure from previous mistakes by evaluating states with an LM to guide the search toward promising areas of the tree.
|
| 3500 |
+
HotPotQA
|
| 3501 |
+
(Yang et al.,
|
| 3502 |
+
2018
|
| 3503 |
+
)
|
| 3504 |
+
is a question-answering dataset that requires reasoning over multiple supporting documents to answer questions. It contains 113k Wikipedia-based question-answer pairs crafted by crowdworkers to be diverse, multi-hop, and explainable. Questions cover a range of types like entities, locations, dates, and comparison of shared properties between two entities. Crowdworkers also provide supporting facts from the documents that justify the answer. We use the HotPotQA benchmark setting with all the Wikipedia paragraphs to test retrieval. We use a randomly selected subset of 100 questions for our experiments and a maximum depth limit of 6. Fig.
|
| 3505 |
+
5
|
| 3506 |
+
illustrates how ReAct and LATS work on an example task of HotPotQA, and gives a qualitative example on how LATS outperforms ReAct on the task.
|
| 3507 |
+
Action Space.
|
| 3508 |
+
We adopt the Wikipedia web API proposed in
|
| 3509 |
+
Yao et al. (
|
| 3510 |
+
2023b
|
| 3511 |
+
)
|
| 3512 |
+
, with three types of actions to support interactive information retrieval:
|
| 3513 |
+
(1)
|
| 3514 |
+
search
|
| 3515 |
+
[
|
| 3516 |
+
entity
|
| 3517 |
+
], which returns the first 5 sentences from the corresponding
|
| 3518 |
+
entity
|
| 3519 |
+
wiki page if it exists, or else suggests top-5 similar entities from the Wikipedia search engine,
|
| 3520 |
+
(2)
|
| 3521 |
+
lookup
|
| 3522 |
+
[
|
| 3523 |
+
string
|
| 3524 |
+
], which returns the next sentence in the page containing
|
| 3525 |
+
string
|
| 3526 |
+
,
|
| 3527 |
+
(3)
|
| 3528 |
+
finish
|
| 3529 |
+
[
|
| 3530 |
+
answer
|
| 3531 |
+
], which finishes the current task with
|
| 3532 |
+
answer
|
| 3533 |
+
.
|
| 3534 |
+
These API calls and free-form thoughts form the action space for this environment.
|
| 3535 |
+
D.2
|
| 3536 |
+
Programming
|
| 3537 |
+
The HumanEval dataset
|
| 3538 |
+
(Chen et al.,
|
| 3539 |
+
2021
|
| 3540 |
+
)
|
| 3541 |
+
is a collection of 164 handwritten programming problems introduced to evaluate the functional correctness of models for synthesizing programs from natural language descriptions. Each problem includes a function signature, docstring description, reference implementation, and multiple unit tests, with an average of 7.7 tests per problem. The programming tasks assess comprehension of natural language, reasoning, algorithms, and basic mathematics, at a difficulty level comparable to simple software interview questions. Pass rates are evaluated with the pass@k metric, where k samples are generated per problem and a problem is considered solved if any sample passes all tests. We use all 164 problems for our experiments and a maximum depth limit of 8.
|
| 3542 |
+
The Mostly Basic Programming Problems (MBPP)
|
| 3543 |
+
Austin et al. (
|
| 3544 |
+
2021
|
| 3545 |
+
)
|
| 3546 |
+
benchmark contains 974 short Python functions designed to evaluate program synthesis techniques. The dataset was constructed by crowdsourcing from workers with basic Python knowledge. Each data point consists of a natural language description of a programming task, a reference solution implementation, and three test cases for functional correctness. The natural language prompts are typically short, one-sentence descriptions. Solutions cover common programming constructs including mathematical operations, list processing, string manipulation, and usage of the Python standard library. On average, solutions are 6.8 lines of code. The dataset is also supplemented with an additional set of 426 problems that were manually verified for unambiguous specifications, standard function signatures, and accurate test cases. We use a randomly selected subset of 397 problems for our experiments.
|
| 3547 |
+
D.3
|
| 3548 |
+
WebShop
|
| 3549 |
+
WebShop
|
| 3550 |
+
(Yao et al.,
|
| 3551 |
+
2022
|
| 3552 |
+
)
|
| 3553 |
+
is an interactive web-based environment designed to evaluate agents on grounded language understanding and decision-making. It simulates an e-commerce shopping task by providing agents with over 1 million real-world products scraped from Amazon, spanning 5 categories and 113 subcategories. These products contain rich linguistic information, with an average text length of 262 words and a vocabulary size of 224k. In addition, there are over 800k unique product options available for customization. The environment renders webpages in two modes: HTML mode provides pixel-level observations with interactive elements, while simple mode converts the raw HTML into a structured text observation more amenable for training agents. The action space consists of query searches and button clicks, which transition between 4 page types: search, results, item and item-detail. Instructions are crowdsourced natural language specifying product attributes and options, with a total of 12k collected. Automatic rewards are computed by comparing the product purchased by the agent against the attributes and options specified in the instruction, using both lexical matching and semantic similarity metrics.
|
| 3554 |
+
Type
|
| 3555 |
+
Argument
|
| 3556 |
+
State
|
| 3557 |
+
→
|
| 3558 |
+
→
|
| 3559 |
+
\rightarrow
|
| 3560 |
+
Next State
|
| 3561 |
+
search
|
| 3562 |
+
[
|
| 3563 |
+
Query
|
| 3564 |
+
]
|
| 3565 |
+
Search
|
| 3566 |
+
→
|
| 3567 |
+
→
|
| 3568 |
+
\rightarrow
|
| 3569 |
+
Results
|
| 3570 |
+
choose
|
| 3571 |
+
Back to search
|
| 3572 |
+
∗
|
| 3573 |
+
*
|
| 3574 |
+
→
|
| 3575 |
+
→
|
| 3576 |
+
\rightarrow
|
| 3577 |
+
Search
|
| 3578 |
+
choose
|
| 3579 |
+
Prev/Next page
|
| 3580 |
+
Results
|
| 3581 |
+
→
|
| 3582 |
+
→
|
| 3583 |
+
\rightarrow
|
| 3584 |
+
Results
|
| 3585 |
+
choose
|
| 3586 |
+
[
|
| 3587 |
+
Product title
|
| 3588 |
+
]
|
| 3589 |
+
Results
|
| 3590 |
+
→
|
| 3591 |
+
→
|
| 3592 |
+
\rightarrow
|
| 3593 |
+
Item
|
| 3594 |
+
choose
|
| 3595 |
+
[
|
| 3596 |
+
Option
|
| 3597 |
+
]
|
| 3598 |
+
Item
|
| 3599 |
+
→
|
| 3600 |
+
→
|
| 3601 |
+
\rightarrow
|
| 3602 |
+
Item
|
| 3603 |
+
choose
|
| 3604 |
+
Desc/Overview
|
| 3605 |
+
Item
|
| 3606 |
+
→
|
| 3607 |
+
→
|
| 3608 |
+
\rightarrow
|
| 3609 |
+
Item-Detail
|
| 3610 |
+
choose
|
| 3611 |
+
Previous
|
| 3612 |
+
Item-Detail
|
| 3613 |
+
→
|
| 3614 |
+
→
|
| 3615 |
+
\rightarrow
|
| 3616 |
+
Item
|
| 3617 |
+
choose
|
| 3618 |
+
Buy
|
| 3619 |
+
Item
|
| 3620 |
+
→
|
| 3621 |
+
→
|
| 3622 |
+
\rightarrow
|
| 3623 |
+
Episode End
|
| 3624 |
+
Table 8:
|
| 3625 |
+
Action space of webshop.
|
| 3626 |
+
There are two evaluation metrics used in WebShop: (1)
|
| 3627 |
+
Task Score
|
| 3628 |
+
: defined as
|
| 3629 |
+
(
|
| 3630 |
+
100
|
| 3631 |
+
×
|
| 3632 |
+
avg. reward
|
| 3633 |
+
)
|
| 3634 |
+
100
|
| 3635 |
+
avg. reward
|
| 3636 |
+
(100\times\text{avg. reward})
|
| 3637 |
+
, which captures the average reward obtained across episodes; and (2)
|
| 3638 |
+
Success Rate (SR)
|
| 3639 |
+
defined as the portion of instructions where
|
| 3640 |
+
r
|
| 3641 |
+
=
|
| 3642 |
+
1
|
| 3643 |
+
𝑟
|
| 3644 |
+
1
|
| 3645 |
+
r=1
|
| 3646 |
+
. The reward is calculated based on the number of attributes satisfied by the selected item. We use 50 environments for our experiments and a maximum depth limit of 15.
|
| 3647 |
+
Appendix E
|
| 3648 |
+
HotPotQA Prompts
|
| 3649 |
+
E.1
|
| 3650 |
+
Base Acting Prompt
|
| 3651 |
+
Solve a question answering task with interleaving Thought, Action, Observation steps. Thought can reason about the current situation, and Action can be three types:
|
| 3652 |
+
(1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
|
| 3653 |
+
(2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
|
| 3654 |
+
(3) Finish[answer], which returns the answer and finishes the task.
|
| 3655 |
+
After each observation, provide the next Thought and next Action.
|
| 3656 |
+
Here are some examples:
|
| 3657 |
+
Question: Which magazine was started first Arthur’s Magazine or First for Women?
|
| 3658 |
+
Thought 1: I need to search Arthur’s Magazine and First for Women, and find which was started first.
|
| 3659 |
+
Action 1: Search[Arthur’s Magazine]
|
| 3660 |
+
Observation 1: Arthur’s Magazine (18440̆0e20̆0800̆0931846) was an American literary periodical published in Philadelphia in the 19th century. Edited by Timothy Shay Arthur, it featured work by Edgar A. Poe, J.H. Ingraham, Sarah Josepha Hale, Thomas G. Spear, and others.[1][2] In May 1846 it was merged into Godey’s Lady’s Book.[3]
|
| 3661 |
+
Thought 2: Arthur’s Magazine was started in 1844. I need to search First for Women next.
|
| 3662 |
+
Action 2: Search[First for Women]
|
| 3663 |
+
Observation 2: First for Women is a woman’s magazine published by Bauer Media Group in the USA.[1] The magazine was started in 1989. It is based in Englewood Cliffs, New Jersey.[2] In 2011 the circulation of the magazine was 1,310,696 copies.[3]
|
| 3664 |
+
Thought 3: First for Women was started in 1989. 1844 (Arthur’s Magazine) ¡ 1989 (First for Women), so Arthur’s Magazine was started first.
|
| 3665 |
+
Action 3: Finish[Arthur’s Magazine]
|
| 3666 |
+
(examples)
|
| 3667 |
+
You have attempted to answer the following question before and failed. The following reflection(s) give a plan to avoid failing to answer the question in the same way you did previously. Use them to improve your strategy of correctly answering the given question.
|
| 3668 |
+
(trajectories)
|
| 3669 |
+
(input)
|
| 3670 |
+
E.2
|
| 3671 |
+
Base Reasoning Prompt
|
| 3672 |
+
Solve a question answering task by having a Thought, then Finish with your answer. Thought can reason about the current situation. Finish[answer] returns the answer and finishes the task. You will be given context that you should use to help you answer the question. Start your response with either Action or an indexed Thought
|
| 3673 |
+
Here are some examples:
|
| 3674 |
+
Question: What is the elevation range for the area that the eastern sector of the Colorado orogeny extends into?
|
| 3675 |
+
Let’s think step by step.
|
| 3676 |
+
Thought 1: The eastern sector of Colorado orogeny extends into the High Plains.
|
| 3677 |
+
Thought 2: High Plains rise in elevation from around 1,800 to 7,000 ft
|
| 3678 |
+
Thought 3: The answer is 1,800 to 7,000 ft.
|
| 3679 |
+
Action: Finish[1,800 to 7,000 ft]
|
| 3680 |
+
(examples)
|
| 3681 |
+
Previous trial:
|
| 3682 |
+
(trajectories)
|
| 3683 |
+
(input)
|
| 3684 |
+
E.3
|
| 3685 |
+
Value Function Prompt
|
| 3686 |
+
Analyze the trajectories of a solution to a question answering task. The trajectories are labeled by environmental observations about the situation, thoughts that can reason about the current situation and actions that can be three types:
|
| 3687 |
+
(1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
|
| 3688 |
+
(2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
|
| 3689 |
+
(3) Finish[answer], which returns the answer and finishes the task.
|
| 3690 |
+
Given a question and a trajectory, evaluate its correctness and provide your reasoning and analysis in detail. Focus on the latest thought, action, and observation. Incomplete trajectories can be correct if the thoughts and actions so far are correct, even if the answer is not found yet. Do not generate additional thoughts or actions. Then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
|
| 3691 |
+
Question: Which magazine was started first Arthur’s Magazine or First for Women?
|
| 3692 |
+
Thought 1: I need to search Arthur’s Magazine and First for Women, and find which was started first.
|
| 3693 |
+
Action 1: Search[Arthur’s Magazine]
|
| 3694 |
+
Observation 1: Arthur’s Magazine (18440̆0e20̆0800̆0931846) was an American literary periodical published in Philadelphia in the 19th century. Edited by Timothy Shay Arthur, it featured work by Edgar A. Poe, J.H. Ingraham, Sarah Josepha Hale, Thomas G. Spear, and others.[1][2] In May 1846 it was merged into Godey’s Lady’s Book.[3]
|
| 3695 |
+
This trajectory is correct as it is reasonable to search for the first magazine provided in the question. It is also better to have simple searches corresponding to a single entity, making this the best action.
|
| 3696 |
+
Thus the correctness score is 10
|
| 3697 |
+
(other examples)
|
| 3698 |
+
(failed trajectories)
|
| 3699 |
+
(context)
|
| 3700 |
+
E.4
|
| 3701 |
+
Reflection Prompt
|
| 3702 |
+
Analyze the trajectories of a solution to a question answering task. The trajectories are labeled by environmental observations about the situation, thoughts that can reason about the current situation and actions that can be three types:
|
| 3703 |
+
(1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
|
| 3704 |
+
(2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
|
| 3705 |
+
(3) Finish[answer], which returns the answer and finishes the task.
|
| 3706 |
+
Given a question and a trajectory, evaluate its correctness and provide your reasoning and analysis in detail. Focus on the latest thought, action, and observation. Incomplete trajectories can be correct if the thoughts and actions so far are correct, even if the answer is not found yet. Do not generate additional thoughts or actions. Then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
|
| 3707 |
+
Question: Which magazine was started first Arthur’s Magazine or First for Women?
|
| 3708 |
+
Thought 1: I need to search Arthur’s Magazine and First for Women, and find which was started first.
|
| 3709 |
+
Action 1: Search[Arthur’s Magazine]
|
| 3710 |
+
Observation 1: Arthur’s Magazine (18440̆0e20̆0800̆0931846) was an American literary periodical published in Philadelphia in the 19th century. Edited by Timothy Shay Arthur, it featured work by Edgar A. Poe, J.H. Ingraham, Sarah Josepha Hale, Thomas G. Spear, and others.[1][2] In May 1846 it was merged into Godey’s Lady’s Book.[3]
|
| 3711 |
+
This trajectory is correct as it is reasonable to search for the first magazine provided in the question. It is also better to have simple searches corresponding to a single entity, making this the best action.
|
| 3712 |
+
Thus the correctness score is 10
|
| 3713 |
+
(other examples)
|
| 3714 |
+
(failed trajectories)
|
| 3715 |
+
(context)
|
| 3716 |
+
Appendix F
|
| 3717 |
+
Programming Prompts
|
| 3718 |
+
F.1
|
| 3719 |
+
HumanEval function implementation example
|
| 3720 |
+
Sample function signature:
|
| 3721 |
+
⬇
|
| 3722 |
+
def
|
| 3723 |
+
minSubArraySum
|
| 3724 |
+
(
|
| 3725 |
+
nums
|
| 3726 |
+
):
|
| 3727 |
+
Given
|
| 3728 |
+
an
|
| 3729 |
+
array
|
| 3730 |
+
of
|
| 3731 |
+
integers
|
| 3732 |
+
nums
|
| 3733 |
+
,
|
| 3734 |
+
find
|
| 3735 |
+
the
|
| 3736 |
+
minimum
|
| 3737 |
+
sum
|
| 3738 |
+
of
|
| 3739 |
+
any
|
| 3740 |
+
non
|
| 3741 |
+
-
|
| 3742 |
+
empty
|
| 3743 |
+
sub
|
| 3744 |
+
-
|
| 3745 |
+
array
|
| 3746 |
+
of
|
| 3747 |
+
nums
|
| 3748 |
+
.
|
| 3749 |
+
Example
|
| 3750 |
+
minSubArraySum
|
| 3751 |
+
([2,
|
| 3752 |
+
3,
|
| 3753 |
+
4,
|
| 3754 |
+
1,
|
| 3755 |
+
2,
|
| 3756 |
+
4])
|
| 3757 |
+
==
|
| 3758 |
+
1
|
| 3759 |
+
minSubArraySum
|
| 3760 |
+
([-1,
|
| 3761 |
+
-2,
|
| 3762 |
+
-3])
|
| 3763 |
+
==
|
| 3764 |
+
-6
|
| 3765 |
+
Sample function body implementation:
|
| 3766 |
+
⬇
|
| 3767 |
+
min_sum
|
| 3768 |
+
=
|
| 3769 |
+
float
|
| 3770 |
+
(’
|
| 3771 |
+
inf
|
| 3772 |
+
’)
|
| 3773 |
+
for
|
| 3774 |
+
i
|
| 3775 |
+
in
|
| 3776 |
+
range
|
| 3777 |
+
(
|
| 3778 |
+
len
|
| 3779 |
+
(
|
| 3780 |
+
nums
|
| 3781 |
+
)):
|
| 3782 |
+
current_sum
|
| 3783 |
+
=
|
| 3784 |
+
0
|
| 3785 |
+
for
|
| 3786 |
+
j
|
| 3787 |
+
in
|
| 3788 |
+
range
|
| 3789 |
+
(
|
| 3790 |
+
i
|
| 3791 |
+
,
|
| 3792 |
+
len
|
| 3793 |
+
(
|
| 3794 |
+
nums
|
| 3795 |
+
)):
|
| 3796 |
+
current_sum
|
| 3797 |
+
+=
|
| 3798 |
+
nums
|
| 3799 |
+
[
|
| 3800 |
+
j
|
| 3801 |
+
]
|
| 3802 |
+
if
|
| 3803 |
+
current_sum
|
| 3804 |
+
<
|
| 3805 |
+
min_sum
|
| 3806 |
+
:
|
| 3807 |
+
min_sum
|
| 3808 |
+
=
|
| 3809 |
+
current_sum
|
| 3810 |
+
return
|
| 3811 |
+
min_sum
|
| 3812 |
+
F.2
|
| 3813 |
+
Base Acting/Reasoning Prompt
|
| 3814 |
+
You are an AI Python assistant. You will be given your previous implementation of a function, a series of unit tests results, and your self-reflection on your previous implementation. Write your full implementation (restate the function signature).
|
| 3815 |
+
Example 1:
|
| 3816 |
+
[previous impl]:
|
| 3817 |
+
⬇
|
| 3818 |
+
def
|
| 3819 |
+
add
|
| 3820 |
+
(
|
| 3821 |
+
a
|
| 3822 |
+
:
|
| 3823 |
+
int
|
| 3824 |
+
,
|
| 3825 |
+
b
|
| 3826 |
+
:
|
| 3827 |
+
int
|
| 3828 |
+
)
|
| 3829 |
+
->
|
| 3830 |
+
int
|
| 3831 |
+
:
|
| 3832 |
+
”””
|
| 3833 |
+
Given
|
| 3834 |
+
integers
|
| 3835 |
+
a
|
| 3836 |
+
and
|
| 3837 |
+
b
|
| 3838 |
+
,
|
| 3839 |
+
return
|
| 3840 |
+
the
|
| 3841 |
+
total
|
| 3842 |
+
value
|
| 3843 |
+
of
|
| 3844 |
+
a
|
| 3845 |
+
and
|
| 3846 |
+
b
|
| 3847 |
+
.
|
| 3848 |
+
”””
|
| 3849 |
+
return
|
| 3850 |
+
a
|
| 3851 |
+
-
|
| 3852 |
+
b
|
| 3853 |
+
[unit test results from previous impl]:
|
| 3854 |
+
Tested passed:
|
| 3855 |
+
Tests failed:
|
| 3856 |
+
assert add(1, 2) == 3 # output: -1
|
| 3857 |
+
assert add(1, 2) == 4 # output: -1
|
| 3858 |
+
[reflection on previous impl]:
|
| 3859 |
+
The implementation failed the test cases where the input integers are 1 and 2. The issue arises because the code does not add the two integers together, but instead subtracts the second integer from the first. To fix this issue, we should change the operator from ‘-‘ to ‘+‘ in the return statement. This will ensure that the function returns the correct output for the given input.
|
| 3860 |
+
[improved impl]:
|
| 3861 |
+
⬇
|
| 3862 |
+
def
|
| 3863 |
+
add
|
| 3864 |
+
(
|
| 3865 |
+
a
|
| 3866 |
+
:
|
| 3867 |
+
int
|
| 3868 |
+
,
|
| 3869 |
+
b
|
| 3870 |
+
:
|
| 3871 |
+
int
|
| 3872 |
+
)
|
| 3873 |
+
->
|
| 3874 |
+
int
|
| 3875 |
+
:
|
| 3876 |
+
”””
|
| 3877 |
+
Given
|
| 3878 |
+
integers
|
| 3879 |
+
a
|
| 3880 |
+
and
|
| 3881 |
+
b
|
| 3882 |
+
,
|
| 3883 |
+
return
|
| 3884 |
+
the
|
| 3885 |
+
total
|
| 3886 |
+
value
|
| 3887 |
+
of
|
| 3888 |
+
a
|
| 3889 |
+
and
|
| 3890 |
+
b
|
| 3891 |
+
.
|
| 3892 |
+
”””
|
| 3893 |
+
return
|
| 3894 |
+
a
|
| 3895 |
+
+
|
| 3896 |
+
b
|
| 3897 |
+
F.3
|
| 3898 |
+
Reflection Prompt
|
| 3899 |
+
You are a Python programming assistant. You will be given a function implementation and a series of unit test results. Your goal is to write a few sentences to explain why your implementation is wrong as indicated by the tests. You will need this as guidance when you try again later. Only provide the few sentence description in your answer, not the implementation. You will be given a few examples by the user.
|
| 3900 |
+
Example 1:
|
| 3901 |
+
[previous impl]:
|
| 3902 |
+
⬇
|
| 3903 |
+
def
|
| 3904 |
+
add
|
| 3905 |
+
(
|
| 3906 |
+
a
|
| 3907 |
+
:
|
| 3908 |
+
int
|
| 3909 |
+
,
|
| 3910 |
+
b
|
| 3911 |
+
:
|
| 3912 |
+
int
|
| 3913 |
+
)
|
| 3914 |
+
->
|
| 3915 |
+
int
|
| 3916 |
+
:
|
| 3917 |
+
”””
|
| 3918 |
+
Given
|
| 3919 |
+
integers
|
| 3920 |
+
a
|
| 3921 |
+
and
|
| 3922 |
+
b
|
| 3923 |
+
,
|
| 3924 |
+
return
|
| 3925 |
+
the
|
| 3926 |
+
total
|
| 3927 |
+
value
|
| 3928 |
+
of
|
| 3929 |
+
a
|
| 3930 |
+
and
|
| 3931 |
+
b
|
| 3932 |
+
.
|
| 3933 |
+
”””
|
| 3934 |
+
return
|
| 3935 |
+
a
|
| 3936 |
+
-
|
| 3937 |
+
b
|
| 3938 |
+
[unit test results from previous impl]:
|
| 3939 |
+
Tested passed:
|
| 3940 |
+
Tests failed:
|
| 3941 |
+
assert add(1, 2) == 3 # output: -1
|
| 3942 |
+
assert add(1, 2) == 4 # output: -1
|
| 3943 |
+
[reflection on previous impl]:
|
| 3944 |
+
The implementation failed the test cases where the input integers are 1 and 2. The issue arises because the code does not add the two integers together, but instead subtracts the second integer from the first. To fix this issue, we should change the operator from ‘-‘ to ‘+‘ in the return statement. This will ensure that the function returns the correct output for the given input.
|
| 3945 |
+
F.4
|
| 3946 |
+
Test Case Generation Prompt
|
| 3947 |
+
You are an AI coding assistant that can write unique, diverse, and intuitive unit tests for functions given the signature and docstring.
|
| 3948 |
+
(input)
|
| 3949 |
+
Appendix G
|
| 3950 |
+
WebShop Prompts
|
| 3951 |
+
G.1
|
| 3952 |
+
Acting Prompt
|
| 3953 |
+
Webshop
|
| 3954 |
+
Instruction:
|
| 3955 |
+
i would like a 3 ounce bottle of bright citrus deodorant for sensitive skin, and price lower than 50.00 dollars
|
| 3956 |
+
[Search]
|
| 3957 |
+
Action: search[3 ounce bright citrus deodorant sensitive skin]
|
| 3958 |
+
Observation:
|
| 3959 |
+
[Back to Search]
|
| 3960 |
+
Page 1 (Total results: 50)
|
| 3961 |
+
[Next
|
| 3962 |
+
]
|
| 3963 |
+
[B078GWRC1J]
|
| 3964 |
+
Bright Citrus Deodorant by Earth Mama — Natural and Safe for Sensitive Skin, Pregnancy and Breastfeeding, Contains Organic Calendula 3-Ounce
|
| 3965 |
+
$10.99
|
| 3966 |
+
[B078GTKVXY]
|
| 3967 |
+
Ginger Fresh Deodorant by Earth Mama — Natural and Safe for Sensitive Skin, Pregnancy and Breastfeeding, Contains Organic Calendula 3-Ounce
|
| 3968 |
+
$10.99
|
| 3969 |
+
[B08KBVJ4XN]
|
| 3970 |
+
Barrel and Oak - Aluminum-Free Deodorant, Deodorant for Men, Essential Oil-Based Scent, 24-Hour Odor Protection, Cedar & Patchouli Blend, Gentle on Sensitive Skin (Mountain Sage, 2.7 oz, 2-Pack)
|
| 3971 |
+
$15.95
|
| 3972 |
+
Action: think[B078GWRC1J and B078GTKVXY are bright citrus deodorant less then 50 dollars. I can check B078GWRC1J first.]
|
| 3973 |
+
Observation: OK.
|
| 3974 |
+
Action: click[B078GWRC1J]
|
| 3975 |
+
Observation:
|
| 3976 |
+
[Back to Search]
|
| 3977 |
+
[Prev]
|
| 3978 |
+
scent [assorted scents][bright citrus][calming lavender][ginger fresh][simply non-scents]
|
| 3979 |
+
size [travel set (4-pack)][3 ounce (pack of 1)][3-ounce (2-pack)]
|
| 3980 |
+
Bright Citrus Deodorant by Earth Mama — Natural and Safe for Sensitive Skin, Pregnancy and Breastfeeding, Contains Organic Calendula 3-Ounce
|
| 3981 |
+
Price: $10.99
|
| 3982 |
+
Rating: N.A.
|
| 3983 |
+
[Description]
|
| 3984 |
+
[Features]
|
| 3985 |
+
[Reviews]
|
| 3986 |
+
[Buy Now]
|
| 3987 |
+
Action: think[For 3 ounce bottle of bright citrus deodorant for sensitive skin, the item has options ’bright citrus’ and ’3 ounce (pack of 1)’ and seems good to buy.]
|
| 3988 |
+
Observation: OK.
|
| 3989 |
+
Action: click[bright citrus]
|
| 3990 |
+
Observation: You have clicked bright citrus.
|
| 3991 |
+
Action: click[3 ounce (pack of 1)]
|
| 3992 |
+
Observation: You have clicked 3 ounce (pack of 1).
|
| 3993 |
+
Action: click[Buy Now]
|
| 3994 |
+
G.2
|
| 3995 |
+
Value Function Prompt
|
| 3996 |
+
Given an item to purchase and a trajectory that aims to buy an item that exactly matches the specification, which corresponds to the ideal score of 1.0, analyze the following trajectory, then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
|
| 3997 |
+
Here are some examples
|
| 3998 |
+
Webshop
|
| 3999 |
+
Instruction:
|
| 4000 |
+
i am looking for dairy free and apple variety pack of chips, and price lower than 30.00 dollars
|
| 4001 |
+
[Search]
|
| 4002 |
+
Action: search[dairy free and apple variety pack of chips]
|
| 4003 |
+
Observation:
|
| 4004 |
+
[Back to Search]
|
| 4005 |
+
Page 1 (Total results: 50)
|
| 4006 |
+
[Next]
|
| 4007 |
+
[B07HRFJWP8]
|
| 4008 |
+
Enjoy Life Foods Soft Baked Ovals, Breakfast Bars, Nut Free Bars, Soy Free, Dairy Free, Non GMO, Gluten Free, Vegan, Variety Pack, 4 Boxes (20 Bars Total)
|
| 4009 |
+
$100.0
|
| 4010 |
+
[B01KMHY5PG]
|
| 4011 |
+
Enjoy Life Soft Baked Chewy Bars, Variety Pack, Nut Free Bars, Soy Free, Dairy Free, Gluten Free, 6 Boxes (30 Total Bars)
|
| 4012 |
+
$21.49
|
| 4013 |
+
[B008D2X8C4]
|
| 4014 |
+
Enjoy Life Lentil Chips Variety Pack, Dairy Free Chips, Soy Free, Nut Free, Non GMO, Vegan, Gluten Free, 24 Bags (0.8 oz)
|
| 4015 |
+
$100.0
|
| 4016 |
+
Action: think[The first two items are not what I want. The third item is the dairy free and apple variety pack of chips I want.]
|
| 4017 |
+
Observation: OK.
|
| 4018 |
+
Action: click[B008D2X8C4]
|
| 4019 |
+
Observation:
|
| 4020 |
+
[Back to Search]
|
| 4021 |
+
[Prev]
|
| 4022 |
+
flavor name [dill and sour cream][garlic & parmesan][light sea salt][margherita pizza][thai chili lime][variety pack]
|
| 4023 |
+
size [0.8 ounce (pack of 24)][4 ounce (pack of 12)]
|
| 4024 |
+
Price: $100.0
|
| 4025 |
+
Rating: N.A.
|
| 4026 |
+
[Description]
|
| 4027 |
+
[Features]
|
| 4028 |
+
[Reviews]
|
| 4029 |
+
[Buy Now]
|
| 4030 |
+
Action: think[The item has options ’variety pack’ and ’0.8 ounce (pack of 24)’ and seems good to buy.]
|
| 4031 |
+
Observation: OK.
|
| 4032 |
+
Action: click[variety pack]
|
| 4033 |
+
Observation: You have clicked variety pack.
|
| 4034 |
+
Action: click[0.8 ounce (pack of 24)]
|
| 4035 |
+
Observation: You have clicked 0.8 ounce (pack of 24).
|
| 4036 |
+
Action: click[Buy Now]
|
| 4037 |
+
Reflection: In this attempt, I was unsuccessful. I accidentally bought a product that was $100, which is more than my budget of $30. Either way, the initial search results were not good. Next time, I will do search[”variety pack of chips”] and then check if the results meet the dairy free and the $30 budget constraints. I will continue to refine my searches so that I can find more products.
|
| 4038 |
+
Thus the correctness score is 5
|
| 4039 |
+
(input)
|
| 4040 |
+
G.3
|
| 4041 |
+
Reflection Prompt
|
| 4042 |
+
Given an item to purchase and a trajectory that aims to buy an item that exactly matches the specification, which corresponds to the ideal score of 1.0, analyze the following trajectory, then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
|
| 4043 |
+
Here are some examples
|
| 4044 |
+
Webshop
|
| 4045 |
+
Instruction:
|
| 4046 |
+
i am looking for dairy free and apple variety pack of chips, and price lower than 30.00 dollars
|
| 4047 |
+
[Search]
|
| 4048 |
+
Action: search[dairy free and apple variety pack of chips]
|
| 4049 |
+
Observation:
|
| 4050 |
+
[Back to Search]
|
| 4051 |
+
Page 1 (Total results: 50)
|
| 4052 |
+
[Next]
|
| 4053 |
+
[B07HRFJWP8]
|
| 4054 |
+
Enjoy Life Foods Soft Baked Ovals, Breakfast Bars, Nut Free Bars, Soy Free, Dairy Free, Non GMO, Gluten Free, Vegan, Variety Pack, 4 Boxes (20 Bars Total)
|
| 4055 |
+
$100.0
|
| 4056 |
+
[B01KMHY5PG]
|
| 4057 |
+
Enjoy Life Soft Baked Chewy Bars, Variety Pack, Nut Free Bars, Soy Free, Dairy Free, Gluten Free, 6 Boxes (30 Total Bars)
|
| 4058 |
+
$21.49
|
| 4059 |
+
[B008D2X8C4]
|
| 4060 |
+
Enjoy Life Lentil Chips Variety Pack, Dairy Free Chips, Soy Free, Nut Free, Non GMO, Vegan, Gluten Free, 24 Bags (0.8 oz)
|
| 4061 |
+
$100.0
|
| 4062 |
+
Action: think[The first two items are not what I want. The third item is the dairy free and apple variety pack of chips I want.]
|
| 4063 |
+
Observation: OK.
|
| 4064 |
+
Action: click[B008D2X8C4]
|
| 4065 |
+
Observation:
|
| 4066 |
+
[Back to Search]
|
| 4067 |
+
[Prev]
|
| 4068 |
+
flavor name [dill and sour cream][garlic & parmesan][light sea salt][margherita pizza][thai chili lime][variety pack]
|
| 4069 |
+
size [0.8 ounce (pack of 24)][4 ounce (pack of 12)]
|
| 4070 |
+
Price: $100.0
|
| 4071 |
+
Rating: N.A.
|
| 4072 |
+
[Description]
|
| 4073 |
+
[Features]
|
| 4074 |
+
[Reviews]
|
| 4075 |
+
[Buy Now]
|
| 4076 |
+
Action: think[The item has options ’variety pack’ and ’0.8 ounce (pack of 24)’ and seems good to buy.]
|
| 4077 |
+
Observation: OK.
|
| 4078 |
+
Action: click[variety pack]
|
| 4079 |
+
Observation: You have clicked variety pack.
|
| 4080 |
+
Action: click[0.8 ounce (pack of 24)]
|
| 4081 |
+
Observation: You have clicked 0.8 ounce (pack of 24).
|
| 4082 |
+
Action: click[Buy Now]
|
| 4083 |
+
Reflection: In this attempt, I was unsuccessful. I accidentally bought a product that was $100, which is more than my budget of $30. Either way, the initial search results were not good. Next time, I will do search[”variety pack of chips”] and then check if the results meet the dairy free and the $30 budget constraints. I will continue to refine my searches so that I can find more products.
|
| 4084 |
+
(input)
|
| 4085 |
+
Reflection:
|
| 4086 |
+
◄
|
| 4087 |
+
Feeling
|
| 4088 |
+
lucky?
|
| 4089 |
+
Conversion
|
| 4090 |
+
report
|
| 4091 |
+
Report
|
| 4092 |
+
an issue
|
| 4093 |
+
View original
|
| 4094 |
+
on arXiv
|
| 4095 |
+
►
|
|
@@ -0,0 +1,4095 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2310.04406] Language Agent Tree Search Unifies Reasoning Acting and Planning
|
| 3 |
+
in Language Models'
|
| 4 |
+
id: 231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la-3
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:40:52.405072Z'
|
| 8 |
+
source: https://ar5iv.labs.arxiv.org/html/2310.04406
|
| 9 |
+
source_domain: ar5iv.labs.arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:40:52.404928Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2310.04406] Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
|
| 20 |
+
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
|
| 21 |
+
Andy Zhou
|
| 22 |
+
University of Illinois at Urbana-Champaign
|
| 23 |
+
AI@UIUC
|
| 24 |
+
Kai Yan
|
| 25 |
+
University of Illinois at Urbana-Champaign
|
| 26 |
+
Michal Shlapentokh-Rothman
|
| 27 |
+
University of Illinois at Urbana-Champaign
|
| 28 |
+
Haohan Wang
|
| 29 |
+
University of Illinois at Urbana-Champaign
|
| 30 |
+
Yu-Xiong Wang
|
| 31 |
+
University of Illinois at Urbana-Champaign
|
| 32 |
+
Abstract
|
| 33 |
+
While large language models (LLMs) have demonstrated impressive performance on a range of decision-making tasks, they rely on simple acting processes and fall short of broad deployment as autonomous agents. We introduce LATS (Language Agent Tree Search), a general framework that synergizes the capabilities of LLMs in planning, acting, and reasoning. Drawing inspiration from Monte Carlo tree search commonly used in model-based reinforcement learning, LATS employs LLMs as agents, value functions, and optimizers, repurposing their latent strengths for enhanced decision-making. What is crucial in this method is the use of an environment for external feedback, which offers a more deliberate and adaptive problem-solving mechanism that moves beyond the limitations of existing techniques. Our experimental evaluation across diverse domains, such as programming, HotPotQA, and WebShop, illustrates the applicability of LATS for decision-making while maintaining competitive reasoning performance. In particular, LATS achieves 94.4% for programming on HumanEval with GPT-4 and an average score of 75.9 for web browsing on WebShop with GPT-3.5, demonstrating the effectiveness and generality of our method.
|
| 34 |
+
1
|
| 35 |
+
Introduction
|
| 36 |
+
General autonomous agents capable of reasoning and decision-making in a variety of environments
|
| 37 |
+
(Wooldridge & Jennings,
|
| 38 |
+
1995
|
| 39 |
+
)
|
| 40 |
+
have been of longstanding interest in the field of artificial intelligence. While this has traditionally been studied in reinforcement learning, the recent rise of large language models (LLMs)
|
| 41 |
+
(Brown et al.,
|
| 42 |
+
2020
|
| 43 |
+
; Chowdhery et al.,
|
| 44 |
+
2022
|
| 45 |
+
; Touvron et al.,
|
| 46 |
+
2023
|
| 47 |
+
; OpenAI,
|
| 48 |
+
2023
|
| 49 |
+
)
|
| 50 |
+
with strong reasoning and general adaptability offers an alternative paradigm. Not only have LLMs excelled on standard NLP tasks such as text summarization
|
| 51 |
+
(Nallapati et al.,
|
| 52 |
+
2016
|
| 53 |
+
)
|
| 54 |
+
or natural language inference
|
| 55 |
+
(Bowman et al.,
|
| 56 |
+
2015
|
| 57 |
+
)
|
| 58 |
+
, but they have been adapted to an increasingly diverse set of tasks that often require advanced common-sense reasoning or quantitative skills
|
| 59 |
+
(Cobbe et al.,
|
| 60 |
+
2021
|
| 61 |
+
; Saparov & He,
|
| 62 |
+
2022
|
| 63 |
+
)
|
| 64 |
+
. LLMs are also capable of performing in complex environments that involve knowledge and reasoning, such as web navigation
|
| 65 |
+
(Yao et al.,
|
| 66 |
+
2022
|
| 67 |
+
; Deng et al.,
|
| 68 |
+
2023
|
| 69 |
+
)
|
| 70 |
+
, tool-use
|
| 71 |
+
(Schick et al.,
|
| 72 |
+
2023
|
| 73 |
+
)
|
| 74 |
+
, or open-ended games
|
| 75 |
+
(Fan et al.,
|
| 76 |
+
2022
|
| 77 |
+
)
|
| 78 |
+
.
|
| 79 |
+
Figure 1:
|
| 80 |
+
An overview of LATS. LATS uses an external environment and self-reflection to improve reasoning and decision-making.
|
| 81 |
+
Reasoning and acting abilities have also been improved by prompting techniques that augment LLMs with feedback or observations from an external environment
|
| 82 |
+
(Yao et al.,
|
| 83 |
+
2023b
|
| 84 |
+
; Gao et al.,
|
| 85 |
+
2022
|
| 86 |
+
; Shinn et al.,
|
| 87 |
+
2023
|
| 88 |
+
)
|
| 89 |
+
. This eliminates the need to rely entirely on the base abilities of the Language Model (LM), enhancing it through external tools or semantic feedback. Despite this strength, these methods are reflexive and fall short of humans’ deliberate and thoughtful decision-making characteristics to solve problems
|
| 90 |
+
(Sloman,
|
| 91 |
+
1996
|
| 92 |
+
; Evans,
|
| 93 |
+
2010
|
| 94 |
+
)
|
| 95 |
+
. In particular, such methods fail to consider multiple reasoning paths or to plan ahead. Recent search-guided LLM works
|
| 96 |
+
(Xie et al.,
|
| 97 |
+
2023
|
| 98 |
+
; Yao et al.,
|
| 99 |
+
2023a
|
| 100 |
+
; Hao et al.,
|
| 101 |
+
2023
|
| 102 |
+
)
|
| 103 |
+
address this issue by searching over multiple reasoning chains. While these methods enable planning, these methods operate in isolation and do not incorporate external feedback that can improve reasoning.
|
| 104 |
+
To help address these issues, we propose LATS (Language Agent Tree Search), a general framework for decision-making and reasoning with language models. LATS unifies LM planning, acting, and reasoning strategies by expanding ReAct
|
| 105 |
+
(Yao et al.,
|
| 106 |
+
2023b
|
| 107 |
+
)
|
| 108 |
+
into a search over a combinatorial space of possible reasoning and acting steps. We adapt Monte Carlo tree search (MCTS) from model-based reinforcement learning
|
| 109 |
+
(Silver et al.,
|
| 110 |
+
2017
|
| 111 |
+
; Anthony et al.,
|
| 112 |
+
2017
|
| 113 |
+
; Jiang et al.,
|
| 114 |
+
2018
|
| 115 |
+
)
|
| 116 |
+
to language agents, repurposing a pretrained LLM as an agent, value function, and optimizer. Utilizing the strong natural language understanding and in-context learning ability of modern LMs, we use text as an interface between each component of the framework, allowing LATS to adapt planning to environmental conditions without additional training. To the best of our knowledge,
|
| 117 |
+
LATS is the first framework that combines reasoning, acting, and planning to enhance LLMs
|
| 118 |
+
. Notably, LATS doubles the performance of GPT-3.5 on HotPotQA
|
| 119 |
+
(Yang et al.,
|
| 120 |
+
2018
|
| 121 |
+
)
|
| 122 |
+
over ReAct
|
| 123 |
+
(Yao et al.,
|
| 124 |
+
2023b
|
| 125 |
+
)
|
| 126 |
+
and raises the average score by
|
| 127 |
+
22.1
|
| 128 |
+
22.1
|
| 129 |
+
22.1
|
| 130 |
+
on WebShop
|
| 131 |
+
(Yao et al.,
|
| 132 |
+
2022
|
| 133 |
+
)
|
| 134 |
+
. When used with GPT-4, LATS achieves a
|
| 135 |
+
94.4
|
| 136 |
+
94.4
|
| 137 |
+
94.4
|
| 138 |
+
Pass@1 rate for programming on HumanEval
|
| 139 |
+
(Chen et al.,
|
| 140 |
+
2021
|
| 141 |
+
)
|
| 142 |
+
, setting the state of the art. To summarize, our
|
| 143 |
+
contributions
|
| 144 |
+
are the following:
|
| 145 |
+
•
|
| 146 |
+
We introduce an LM-based Monte Carlo tree search variant to deliberately construct the best trajectory from sampled actions, enabling more flexible and adaptive problem-solving compared to reflexive prompting methods. This is guided by heuristics from the LM.
|
| 147 |
+
•
|
| 148 |
+
By integrating external feedback and self-reflection, LATS enhances model sensibility and enables agents to learn from experience, surpassing reasoning-based search methods.
|
| 149 |
+
•
|
| 150 |
+
Through experiments across diverse domains like programming, interactive QA, and web navigation, we demonstrate the versatility of LATS in harnessing LLMs for autonomous reasoning and decision-making.
|
| 151 |
+
2
|
| 152 |
+
Related Work
|
| 153 |
+
Approach
|
| 154 |
+
Reasoning
|
| 155 |
+
Acting
|
| 156 |
+
Planning
|
| 157 |
+
Self
|
| 158 |
+
External
|
| 159 |
+
Reflection
|
| 160 |
+
Memory
|
| 161 |
+
CoT
|
| 162 |
+
(Wei et al.,
|
| 163 |
+
2022
|
| 164 |
+
)
|
| 165 |
+
✓
|
| 166 |
+
×
|
| 167 |
+
\times
|
| 168 |
+
×
|
| 169 |
+
\times
|
| 170 |
+
×
|
| 171 |
+
\times
|
| 172 |
+
×
|
| 173 |
+
\times
|
| 174 |
+
ReAct
|
| 175 |
+
(Yao et al.,
|
| 176 |
+
2023b
|
| 177 |
+
)
|
| 178 |
+
✓
|
| 179 |
+
✓
|
| 180 |
+
×
|
| 181 |
+
\times
|
| 182 |
+
×
|
| 183 |
+
\times
|
| 184 |
+
×
|
| 185 |
+
\times
|
| 186 |
+
ToT
|
| 187 |
+
(Yao et al.,
|
| 188 |
+
2023a
|
| 189 |
+
)
|
| 190 |
+
✓
|
| 191 |
+
×
|
| 192 |
+
\times
|
| 193 |
+
✓
|
| 194 |
+
✓
|
| 195 |
+
✓
|
| 196 |
+
RAP
|
| 197 |
+
(Hao et al.,
|
| 198 |
+
2023
|
| 199 |
+
)
|
| 200 |
+
✓
|
| 201 |
+
×
|
| 202 |
+
\times
|
| 203 |
+
✓
|
| 204 |
+
×
|
| 205 |
+
\times
|
| 206 |
+
✓
|
| 207 |
+
Self-Refine
|
| 208 |
+
(Madaan et al.,
|
| 209 |
+
2023
|
| 210 |
+
)
|
| 211 |
+
✓
|
| 212 |
+
×
|
| 213 |
+
\times
|
| 214 |
+
×
|
| 215 |
+
\times
|
| 216 |
+
✓
|
| 217 |
+
×
|
| 218 |
+
\times
|
| 219 |
+
Beam Search
|
| 220 |
+
(Xie et al.,
|
| 221 |
+
2023
|
| 222 |
+
)
|
| 223 |
+
✓
|
| 224 |
+
×
|
| 225 |
+
\times
|
| 226 |
+
×
|
| 227 |
+
\times
|
| 228 |
+
✓
|
| 229 |
+
×
|
| 230 |
+
\times
|
| 231 |
+
Reflexion
|
| 232 |
+
(Shinn et al.,
|
| 233 |
+
2023
|
| 234 |
+
)
|
| 235 |
+
✓
|
| 236 |
+
✓
|
| 237 |
+
×
|
| 238 |
+
\times
|
| 239 |
+
✓
|
| 240 |
+
✓
|
| 241 |
+
LATS (Ours)
|
| 242 |
+
✓
|
| 243 |
+
✓
|
| 244 |
+
✓
|
| 245 |
+
✓
|
| 246 |
+
✓
|
| 247 |
+
Table 1:
|
| 248 |
+
A summary of related work on reasoning, acting, and planning. LATS is the first work incorporating designs from all three domains, allowing use in all corresponding tasks. We refer to planning as the use of a search algorithm, self-reflection as the use of LM-generated feedback, and external memory as storaging past text context for future updates of solution.
|
| 249 |
+
a) Tree-of-Thoughts
|
| 250 |
+
b) Reasoning via Planning
|
| 251 |
+
c) Language Agent Tree Search
|
| 252 |
+
Figure 2:
|
| 253 |
+
An overview of the differences between LATS and recently proposed LM search algorithms ToT
|
| 254 |
+
(Yao et al.,
|
| 255 |
+
2023a
|
| 256 |
+
)
|
| 257 |
+
and RAP
|
| 258 |
+
(Hao et al.,
|
| 259 |
+
2023
|
| 260 |
+
)
|
| 261 |
+
. LATS leverages environmental feedback and self-reflection to further adapt search and improve performance.
|
| 262 |
+
LLMs for reasoning.
|
| 263 |
+
For LLMs, reasoning typically involves decomposing complex inputs into sequential intermediate steps towards a final answer
|
| 264 |
+
(Cobbe et al.,
|
| 265 |
+
2021
|
| 266 |
+
)
|
| 267 |
+
, demonstrated with Chain-of-Thought (CoT) prompting
|
| 268 |
+
(Wei et al.,
|
| 269 |
+
2022
|
| 270 |
+
)
|
| 271 |
+
and its variants
|
| 272 |
+
(Wei et al.,
|
| 273 |
+
2022
|
| 274 |
+
; Kojima et al.,
|
| 275 |
+
2022
|
| 276 |
+
; Wang et al.,
|
| 277 |
+
2022
|
| 278 |
+
)
|
| 279 |
+
. However, these methods, which create chains autoregressively in a single step, often suffer from error propagation as the number of steps increases
|
| 280 |
+
(Guo et al.,
|
| 281 |
+
2018
|
| 282 |
+
; Chen et al.,
|
| 283 |
+
2022b
|
| 284 |
+
)
|
| 285 |
+
due to compound errors. Various advancements aim to mitigate this issue; some approaches, such as Self-Consistency
|
| 286 |
+
(Wang et al.,
|
| 287 |
+
2022
|
| 288 |
+
)
|
| 289 |
+
, employ majority voting over sampled chains, while others focus on multi-step decomposition, such as least-to-most prompting
|
| 290 |
+
(Zhou et al.,
|
| 291 |
+
2022
|
| 292 |
+
)
|
| 293 |
+
, or use of external tools such as a scratchpad
|
| 294 |
+
(Nye et al.,
|
| 295 |
+
2021
|
| 296 |
+
)
|
| 297 |
+
or compiler
|
| 298 |
+
(Gao et al.,
|
| 299 |
+
2022
|
| 300 |
+
)
|
| 301 |
+
. Recently, CoT has been improved with search algorithms
|
| 302 |
+
(Yao et al.,
|
| 303 |
+
2023a
|
| 304 |
+
; Hao et al.,
|
| 305 |
+
2023
|
| 306 |
+
; Besta et al.,
|
| 307 |
+
2023
|
| 308 |
+
)
|
| 309 |
+
that can sample trajectories more effectively. Tree-of-thought (ToT) prompting
|
| 310 |
+
(Yao et al.,
|
| 311 |
+
2023a
|
| 312 |
+
)
|
| 313 |
+
uses DFS or BFS-based search guided by an LM-generated heuristic while Reasoning via Planning (RAP)
|
| 314 |
+
(Hao et al.,
|
| 315 |
+
2023
|
| 316 |
+
)
|
| 317 |
+
uses MCTS with rollouts simulated by the LM. However, they rely solely on LM internal knowledge and cannot adapt to useful external feedback.
|
| 318 |
+
LLMs for acting.
|
| 319 |
+
The strong reasoning and common-sense abilities of LLMs have also been adapted for decision-making or acting tasks as a policy model in interactive environments. In the realm of robotics LLMs have been employed as high-level controllers of control policies
|
| 320 |
+
(Ahn et al.,
|
| 321 |
+
2022
|
| 322 |
+
; Huang et al.,
|
| 323 |
+
2022
|
| 324 |
+
; Driess et al.,
|
| 325 |
+
2023
|
| 326 |
+
)
|
| 327 |
+
. Similar work
|
| 328 |
+
(Baker et al.,
|
| 329 |
+
2022
|
| 330 |
+
; Wang et al.,
|
| 331 |
+
2023
|
| 332 |
+
; Zhu et al.,
|
| 333 |
+
2023
|
| 334 |
+
)
|
| 335 |
+
has also adapted LLM agents to complex multimodal games such as Minecraft
|
| 336 |
+
(Guss et al.,
|
| 337 |
+
2019
|
| 338 |
+
; Fan et al.,
|
| 339 |
+
2022
|
| 340 |
+
)
|
| 341 |
+
. LLMs are particularly useful in text-based environments
|
| 342 |
+
(Liu et al.,
|
| 343 |
+
2018
|
| 344 |
+
; Shridhar et al.,
|
| 345 |
+
2020
|
| 346 |
+
; Liu et al.,
|
| 347 |
+
2023
|
| 348 |
+
)
|
| 349 |
+
, where acting-based prompting techniques such as ReAct
|
| 350 |
+
(Yao et al.,
|
| 351 |
+
2023b
|
| 352 |
+
)
|
| 353 |
+
have seen success. Similar to CoT, ReAct is limited by its simplicity and cannot effectively adapt to environment conditions. Many extensions have been proposed to address this, including Self-refine
|
| 354 |
+
(Madaan et al.,
|
| 355 |
+
2023
|
| 356 |
+
)
|
| 357 |
+
and Reflexion
|
| 358 |
+
(Shinn et al.,
|
| 359 |
+
2023
|
| 360 |
+
; Yao et al.,
|
| 361 |
+
2023c
|
| 362 |
+
)
|
| 363 |
+
, which uses self-reflection to enhance reasoning and decision-making, and AdaPlanner
|
| 364 |
+
(Sun et al.,
|
| 365 |
+
2023
|
| 366 |
+
)
|
| 367 |
+
, which incorporates both positive and negative environmental feedback. However these methods focus on refining an individual plan or trajectory and do not consider alternative choices at each step. In addition, recent work
|
| 368 |
+
(Huang et al.,
|
| 369 |
+
2023
|
| 370 |
+
)
|
| 371 |
+
has suggested LLMs cannot self-correct their internal reasoning, making it critical to use external feedback. Alternatively to pure decision-making environments, the reasoning and practical abilities of LLMs have been enhanced by access to external tools, such as APIs, search engines, calculators, or other models
|
| 372 |
+
(Schick et al.,
|
| 373 |
+
2023
|
| 374 |
+
; Shen et al.,
|
| 375 |
+
2023
|
| 376 |
+
; Surís et al.,
|
| 377 |
+
2023
|
| 378 |
+
)
|
| 379 |
+
. Contrary to reasoning-based approaches, these methods have not been improved with planning, limiting their effectiveness. We summarize them in Tab.
|
| 380 |
+
1
|
| 381 |
+
.
|
| 382 |
+
Tree-based search.
|
| 383 |
+
Tree-based search, where multiple branches of outcomes are explored during search, is widely used in many planning algorithms
|
| 384 |
+
(Świechowski et al.,
|
| 385 |
+
2023
|
| 386 |
+
; LaValle et al.,
|
| 387 |
+
2001
|
| 388 |
+
)
|
| 389 |
+
and Reinforcement Learning (RL)
|
| 390 |
+
(Hafner et al.,
|
| 391 |
+
2019
|
| 392 |
+
; Du et al.,
|
| 393 |
+
2023
|
| 394 |
+
; Wu et al.,
|
| 395 |
+
2023
|
| 396 |
+
)
|
| 397 |
+
algorithms for its good exploration-exploitation trade-off. Though tree-based search requires an environment model that can expand from arbitrary state
|
| 398 |
+
(Vodopivec et al.,
|
| 399 |
+
2017
|
| 400 |
+
)
|
| 401 |
+
, which often requires extra training in RL
|
| 402 |
+
(Hafner et al.,
|
| 403 |
+
2023
|
| 404 |
+
)
|
| 405 |
+
, such problem does not exist for LM tasks as we can conveniently backup to any state by setting the input to be the context and corresponding previous output by the LM. Thus, we work on the tree-based framework and use MCTS
|
| 406 |
+
(Świechowski et al.,
|
| 407 |
+
2023
|
| 408 |
+
)
|
| 409 |
+
to fully release the potential of LMs, while avoiding the cost of training a value function over language descriptions by leveraging the in-context learning
|
| 410 |
+
(Brown et al.,
|
| 411 |
+
2020
|
| 412 |
+
)
|
| 413 |
+
abilities of LLMs.
|
| 414 |
+
3
|
| 415 |
+
Preliminaries
|
| 416 |
+
3.1
|
| 417 |
+
Problem Setting and Prompting
|
| 418 |
+
Before describing LATS, we first define our problem and outline a few established methods that leverage large language models for reasoning or decision-making. In LM reasoning or decision making, we are given an input
|
| 419 |
+
x
|
| 420 |
+
𝑥
|
| 421 |
+
x
|
| 422 |
+
in natural language and a pretrained language model
|
| 423 |
+
p
|
| 424 |
+
θ
|
| 425 |
+
|
| 426 |
+
(
|
| 427 |
+
x
|
| 428 |
+
)
|
| 429 |
+
subscript
|
| 430 |
+
𝑝
|
| 431 |
+
𝜃
|
| 432 |
+
𝑥
|
| 433 |
+
p_{\theta}(x)
|
| 434 |
+
parameterized by
|
| 435 |
+
θ
|
| 436 |
+
𝜃
|
| 437 |
+
\theta
|
| 438 |
+
; our goal is to generate a final output
|
| 439 |
+
y
|
| 440 |
+
∼
|
| 441 |
+
p
|
| 442 |
+
θ
|
| 443 |
+
|
| 444 |
+
(
|
| 445 |
+
x
|
| 446 |
+
)
|
| 447 |
+
similar-to
|
| 448 |
+
𝑦
|
| 449 |
+
subscript
|
| 450 |
+
𝑝
|
| 451 |
+
𝜃
|
| 452 |
+
𝑥
|
| 453 |
+
y\sim p_{\theta}(x)
|
| 454 |
+
corresponding to the answer (reasoning) or completes the task (decision-making). Both
|
| 455 |
+
x
|
| 456 |
+
𝑥
|
| 457 |
+
x
|
| 458 |
+
and
|
| 459 |
+
y
|
| 460 |
+
𝑦
|
| 461 |
+
y
|
| 462 |
+
are language
|
| 463 |
+
sequences
|
| 464 |
+
, which are comprised of a list of
|
| 465 |
+
tokens
|
| 466 |
+
(the basic elements of natural language, often words), denoted as
|
| 467 |
+
x
|
| 468 |
+
=
|
| 469 |
+
(
|
| 470 |
+
x
|
| 471 |
+
|
| 472 |
+
[
|
| 473 |
+
1
|
| 474 |
+
]
|
| 475 |
+
,
|
| 476 |
+
…
|
| 477 |
+
,
|
| 478 |
+
x
|
| 479 |
+
|
| 480 |
+
[
|
| 481 |
+
n
|
| 482 |
+
]
|
| 483 |
+
)
|
| 484 |
+
𝑥
|
| 485 |
+
𝑥
|
| 486 |
+
delimited-[]
|
| 487 |
+
1
|
| 488 |
+
…
|
| 489 |
+
𝑥
|
| 490 |
+
delimited-[]
|
| 491 |
+
𝑛
|
| 492 |
+
x=(x[1],\dots,x[n])
|
| 493 |
+
and
|
| 494 |
+
y
|
| 495 |
+
=
|
| 496 |
+
(
|
| 497 |
+
y
|
| 498 |
+
|
| 499 |
+
[
|
| 500 |
+
1
|
| 501 |
+
]
|
| 502 |
+
,
|
| 503 |
+
…
|
| 504 |
+
,
|
| 505 |
+
y
|
| 506 |
+
|
| 507 |
+
[
|
| 508 |
+
n
|
| 509 |
+
]
|
| 510 |
+
)
|
| 511 |
+
𝑦
|
| 512 |
+
𝑦
|
| 513 |
+
delimited-[]
|
| 514 |
+
1
|
| 515 |
+
…
|
| 516 |
+
𝑦
|
| 517 |
+
delimited-[]
|
| 518 |
+
𝑛
|
| 519 |
+
y=(y[1],\dots,y[n])
|
| 520 |
+
. The LM decodes text autoregressively, i.e., without other inputs, the probability for an LM to generate a sequence
|
| 521 |
+
x
|
| 522 |
+
𝑥
|
| 523 |
+
x
|
| 524 |
+
is given by
|
| 525 |
+
p
|
| 526 |
+
θ
|
| 527 |
+
|
| 528 |
+
(
|
| 529 |
+
x
|
| 530 |
+
)
|
| 531 |
+
=
|
| 532 |
+
∏
|
| 533 |
+
i
|
| 534 |
+
=
|
| 535 |
+
1
|
| 536 |
+
n
|
| 537 |
+
p
|
| 538 |
+
θ
|
| 539 |
+
|
| 540 |
+
(
|
| 541 |
+
x
|
| 542 |
+
|
| 543 |
+
[
|
| 544 |
+
i
|
| 545 |
+
]
|
| 546 |
+
|
|
| 547 |
+
x
|
| 548 |
+
|
| 549 |
+
[
|
| 550 |
+
1
|
| 551 |
+
|
| 552 |
+
…
|
| 553 |
+
|
| 554 |
+
i
|
| 555 |
+
−
|
| 556 |
+
1
|
| 557 |
+
]
|
| 558 |
+
)
|
| 559 |
+
subscript
|
| 560 |
+
𝑝
|
| 561 |
+
𝜃
|
| 562 |
+
𝑥
|
| 563 |
+
superscript
|
| 564 |
+
subscript
|
| 565 |
+
product
|
| 566 |
+
𝑖
|
| 567 |
+
1
|
| 568 |
+
𝑛
|
| 569 |
+
subscript
|
| 570 |
+
𝑝
|
| 571 |
+
𝜃
|
| 572 |
+
conditional
|
| 573 |
+
𝑥
|
| 574 |
+
delimited-[]
|
| 575 |
+
𝑖
|
| 576 |
+
𝑥
|
| 577 |
+
delimited-[]
|
| 578 |
+
1
|
| 579 |
+
…
|
| 580 |
+
𝑖
|
| 581 |
+
1
|
| 582 |
+
p_{\theta}(x)=\prod_{i=1}^{n}p_{\theta}(x[i]|x[1\dots i-1])
|
| 583 |
+
. Usually, to improve the LM,
|
| 584 |
+
prompts
|
| 585 |
+
are provided along with the input
|
| 586 |
+
x
|
| 587 |
+
𝑥
|
| 588 |
+
x
|
| 589 |
+
, which are specific instructions or few-shot input-output examples. We denote the generic process where an input
|
| 590 |
+
x
|
| 591 |
+
𝑥
|
| 592 |
+
x
|
| 593 |
+
is transformed into an output
|
| 594 |
+
y
|
| 595 |
+
𝑦
|
| 596 |
+
y
|
| 597 |
+
by LM:
|
| 598 |
+
y
|
| 599 |
+
∼
|
| 600 |
+
p
|
| 601 |
+
θ
|
| 602 |
+
|
| 603 |
+
(
|
| 604 |
+
y
|
| 605 |
+
|
|
| 606 |
+
prompt
|
| 607 |
+
I
|
| 608 |
+
|
| 609 |
+
O
|
| 610 |
+
|
| 611 |
+
(
|
| 612 |
+
x
|
| 613 |
+
)
|
| 614 |
+
)
|
| 615 |
+
similar-to
|
| 616 |
+
𝑦
|
| 617 |
+
subscript
|
| 618 |
+
𝑝
|
| 619 |
+
𝜃
|
| 620 |
+
conditional
|
| 621 |
+
𝑦
|
| 622 |
+
subscript
|
| 623 |
+
prompt
|
| 624 |
+
𝐼
|
| 625 |
+
𝑂
|
| 626 |
+
𝑥
|
| 627 |
+
y\sim p_{\theta}(y|\texttt{prompt}_{IO}(x))
|
| 628 |
+
, where
|
| 629 |
+
prompt
|
| 630 |
+
I
|
| 631 |
+
|
| 632 |
+
O
|
| 633 |
+
|
| 634 |
+
(
|
| 635 |
+
x
|
| 636 |
+
)
|
| 637 |
+
subscript
|
| 638 |
+
prompt
|
| 639 |
+
𝐼
|
| 640 |
+
𝑂
|
| 641 |
+
𝑥
|
| 642 |
+
\texttt{prompt}_{IO}(x)
|
| 643 |
+
denotes the input
|
| 644 |
+
x
|
| 645 |
+
𝑥
|
| 646 |
+
x
|
| 647 |
+
.
|
| 648 |
+
Chain-of-thought (CoT) Prompting
|
| 649 |
+
(Wei et al.,
|
| 650 |
+
2022
|
| 651 |
+
)
|
| 652 |
+
was introduced to cater to scenarios where direct mapping from
|
| 653 |
+
x
|
| 654 |
+
𝑥
|
| 655 |
+
x
|
| 656 |
+
to
|
| 657 |
+
y
|
| 658 |
+
𝑦
|
| 659 |
+
y
|
| 660 |
+
is intricate, such as when
|
| 661 |
+
x
|
| 662 |
+
𝑥
|
| 663 |
+
x
|
| 664 |
+
is from a mathematical query or challenging question. This method hinges on creating
|
| 665 |
+
thoughts
|
| 666 |
+
z
|
| 667 |
+
1
|
| 668 |
+
,
|
| 669 |
+
…
|
| 670 |
+
,
|
| 671 |
+
z
|
| 672 |
+
n
|
| 673 |
+
subscript
|
| 674 |
+
𝑧
|
| 675 |
+
1
|
| 676 |
+
…
|
| 677 |
+
subscript
|
| 678 |
+
𝑧
|
| 679 |
+
𝑛
|
| 680 |
+
z_{1},\dots,z_{n}
|
| 681 |
+
that act as stepping stones between
|
| 682 |
+
x
|
| 683 |
+
𝑥
|
| 684 |
+
x
|
| 685 |
+
and
|
| 686 |
+
y
|
| 687 |
+
𝑦
|
| 688 |
+
y
|
| 689 |
+
; each thought
|
| 690 |
+
z
|
| 691 |
+
i
|
| 692 |
+
subscript
|
| 693 |
+
𝑧
|
| 694 |
+
𝑖
|
| 695 |
+
z_{i}
|
| 696 |
+
is a language sequence. To employ CoT prompting, thoughts are extracted sequentially as
|
| 697 |
+
z
|
| 698 |
+
i
|
| 699 |
+
∼
|
| 700 |
+
p
|
| 701 |
+
θ
|
| 702 |
+
C
|
| 703 |
+
|
| 704 |
+
o
|
| 705 |
+
|
| 706 |
+
T
|
| 707 |
+
|
| 708 |
+
(
|
| 709 |
+
z
|
| 710 |
+
i
|
| 711 |
+
|
|
| 712 |
+
x
|
| 713 |
+
,
|
| 714 |
+
z
|
| 715 |
+
1
|
| 716 |
+
|
| 717 |
+
⋯
|
| 718 |
+
|
| 719 |
+
i
|
| 720 |
+
−
|
| 721 |
+
1
|
| 722 |
+
)
|
| 723 |
+
similar-to
|
| 724 |
+
subscript
|
| 725 |
+
𝑧
|
| 726 |
+
𝑖
|
| 727 |
+
superscript
|
| 728 |
+
subscript
|
| 729 |
+
𝑝
|
| 730 |
+
𝜃
|
| 731 |
+
𝐶
|
| 732 |
+
𝑜
|
| 733 |
+
𝑇
|
| 734 |
+
conditional
|
| 735 |
+
subscript
|
| 736 |
+
𝑧
|
| 737 |
+
𝑖
|
| 738 |
+
𝑥
|
| 739 |
+
subscript
|
| 740 |
+
𝑧
|
| 741 |
+
1
|
| 742 |
+
⋯
|
| 743 |
+
𝑖
|
| 744 |
+
1
|
| 745 |
+
z_{i}\sim p_{\theta}^{CoT}(z_{i}|x,z_{1\cdots i-1})
|
| 746 |
+
, with the final output being
|
| 747 |
+
y
|
| 748 |
+
∼
|
| 749 |
+
p
|
| 750 |
+
θ
|
| 751 |
+
C
|
| 752 |
+
|
| 753 |
+
o
|
| 754 |
+
|
| 755 |
+
T
|
| 756 |
+
|
| 757 |
+
(
|
| 758 |
+
y
|
| 759 |
+
|
|
| 760 |
+
x
|
| 761 |
+
,
|
| 762 |
+
z
|
| 763 |
+
1
|
| 764 |
+
|
| 765 |
+
⋯
|
| 766 |
+
|
| 767 |
+
n
|
| 768 |
+
)
|
| 769 |
+
similar-to
|
| 770 |
+
𝑦
|
| 771 |
+
superscript
|
| 772 |
+
subscript
|
| 773 |
+
𝑝
|
| 774 |
+
𝜃
|
| 775 |
+
𝐶
|
| 776 |
+
𝑜
|
| 777 |
+
𝑇
|
| 778 |
+
conditional
|
| 779 |
+
𝑦
|
| 780 |
+
𝑥
|
| 781 |
+
subscript
|
| 782 |
+
𝑧
|
| 783 |
+
1
|
| 784 |
+
⋯
|
| 785 |
+
𝑛
|
| 786 |
+
y\sim p_{\theta}^{CoT}(y|x,z_{1\cdots n})
|
| 787 |
+
.
|
| 788 |
+
Tree-of-thought (ToT) Prompting
|
| 789 |
+
(Yao et al.,
|
| 790 |
+
2023a
|
| 791 |
+
)
|
| 792 |
+
extends CoT prompting by exploring multiple reasoning paths over thoughts. It frames problems as a search over a tree where each node
|
| 793 |
+
s
|
| 794 |
+
=
|
| 795 |
+
[
|
| 796 |
+
x
|
| 797 |
+
,
|
| 798 |
+
z
|
| 799 |
+
1
|
| 800 |
+
⋅
|
| 801 |
+
i
|
| 802 |
+
]
|
| 803 |
+
𝑠
|
| 804 |
+
𝑥
|
| 805 |
+
subscript
|
| 806 |
+
𝑧
|
| 807 |
+
⋅
|
| 808 |
+
1
|
| 809 |
+
𝑖
|
| 810 |
+
s=[x,z_{1\cdot i}]
|
| 811 |
+
represents a partial solution state comprising the original input
|
| 812 |
+
x
|
| 813 |
+
𝑥
|
| 814 |
+
x
|
| 815 |
+
and thought sequence
|
| 816 |
+
z
|
| 817 |
+
1
|
| 818 |
+
|
| 819 |
+
⋯
|
| 820 |
+
|
| 821 |
+
i
|
| 822 |
+
subscript
|
| 823 |
+
𝑧
|
| 824 |
+
1
|
| 825 |
+
⋯
|
| 826 |
+
𝑖
|
| 827 |
+
z_{1\cdots i}
|
| 828 |
+
. Thoughts
|
| 829 |
+
z
|
| 830 |
+
i
|
| 831 |
+
subscript
|
| 832 |
+
𝑧
|
| 833 |
+
𝑖
|
| 834 |
+
z_{i}
|
| 835 |
+
are generated by proposal or sampling with CoT
|
| 836 |
+
z
|
| 837 |
+
i
|
| 838 |
+
∼
|
| 839 |
+
p
|
| 840 |
+
θ
|
| 841 |
+
C
|
| 842 |
+
|
| 843 |
+
o
|
| 844 |
+
|
| 845 |
+
T
|
| 846 |
+
|
| 847 |
+
(
|
| 848 |
+
z
|
| 849 |
+
i
|
| 850 |
+
|
|
| 851 |
+
x
|
| 852 |
+
,
|
| 853 |
+
z
|
| 854 |
+
1
|
| 855 |
+
|
| 856 |
+
⋯
|
| 857 |
+
|
| 858 |
+
i
|
| 859 |
+
−
|
| 860 |
+
1
|
| 861 |
+
)
|
| 862 |
+
similar-to
|
| 863 |
+
subscript
|
| 864 |
+
𝑧
|
| 865 |
+
𝑖
|
| 866 |
+
superscript
|
| 867 |
+
subscript
|
| 868 |
+
𝑝
|
| 869 |
+
𝜃
|
| 870 |
+
𝐶
|
| 871 |
+
𝑜
|
| 872 |
+
𝑇
|
| 873 |
+
conditional
|
| 874 |
+
subscript
|
| 875 |
+
𝑧
|
| 876 |
+
𝑖
|
| 877 |
+
𝑥
|
| 878 |
+
subscript
|
| 879 |
+
𝑧
|
| 880 |
+
1
|
| 881 |
+
⋯
|
| 882 |
+
𝑖
|
| 883 |
+
1
|
| 884 |
+
z_{i}\sim p_{\theta}^{CoT}(z_{i}|x,z_{1\cdots i-1})
|
| 885 |
+
. Deliberate search algorithms like breadth-first or depth-first search are used to systematically explore the tree, guided by heuristics based on language model evaluations
|
| 886 |
+
V
|
| 887 |
+
|
| 888 |
+
(
|
| 889 |
+
s
|
| 890 |
+
)
|
| 891 |
+
𝑉
|
| 892 |
+
𝑠
|
| 893 |
+
V(s)
|
| 894 |
+
of each state.
|
| 895 |
+
Reasoning via Planning
|
| 896 |
+
(RAP)
|
| 897 |
+
(Hao et al.,
|
| 898 |
+
2023
|
| 899 |
+
)
|
| 900 |
+
is similar to ToT, except that MCTS is used over DFS or BFS. Heuristics are designed from an LM, such as the likelihood or confidence of an action, and the LM is used as a world model to predict subsequent states during the simulation step.
|
| 901 |
+
ReAct
|
| 902 |
+
(Yao et al.,
|
| 903 |
+
2023b
|
| 904 |
+
)
|
| 905 |
+
extends language models to tasks where the mapping from
|
| 906 |
+
x
|
| 907 |
+
𝑥
|
| 908 |
+
x
|
| 909 |
+
to
|
| 910 |
+
y
|
| 911 |
+
𝑦
|
| 912 |
+
y
|
| 913 |
+
is enhanced by or requires interactions with an external environment, such as a game or API. This technique constructs an action space
|
| 914 |
+
A
|
| 915 |
+
^
|
| 916 |
+
=
|
| 917 |
+
A
|
| 918 |
+
∪
|
| 919 |
+
Z
|
| 920 |
+
^
|
| 921 |
+
𝐴
|
| 922 |
+
𝐴
|
| 923 |
+
𝑍
|
| 924 |
+
\hat{A}=A\cup Z
|
| 925 |
+
that adds permissible actions
|
| 926 |
+
a
|
| 927 |
+
𝑎
|
| 928 |
+
a
|
| 929 |
+
to the reasoning traces
|
| 930 |
+
z
|
| 931 |
+
𝑧
|
| 932 |
+
z
|
| 933 |
+
from CoT. Observations
|
| 934 |
+
o
|
| 935 |
+
𝑜
|
| 936 |
+
o
|
| 937 |
+
from the environment are used to improve both reasoning and acting. To solve problems with ReAct, after each observation, actions are generated from
|
| 938 |
+
p
|
| 939 |
+
θ
|
| 940 |
+
subscript
|
| 941 |
+
𝑝
|
| 942 |
+
𝜃
|
| 943 |
+
p_{\theta}
|
| 944 |
+
sequentially as
|
| 945 |
+
a
|
| 946 |
+
i
|
| 947 |
+
∼
|
| 948 |
+
p
|
| 949 |
+
θ
|
| 950 |
+
R
|
| 951 |
+
|
| 952 |
+
e
|
| 953 |
+
|
| 954 |
+
A
|
| 955 |
+
|
| 956 |
+
c
|
| 957 |
+
|
| 958 |
+
t
|
| 959 |
+
|
| 960 |
+
(
|
| 961 |
+
a
|
| 962 |
+
i
|
| 963 |
+
|
|
| 964 |
+
x
|
| 965 |
+
,
|
| 966 |
+
o
|
| 967 |
+
1
|
| 968 |
+
|
| 969 |
+
⋯
|
| 970 |
+
|
| 971 |
+
i
|
| 972 |
+
−
|
| 973 |
+
1
|
| 974 |
+
,
|
| 975 |
+
a
|
| 976 |
+
1
|
| 977 |
+
|
| 978 |
+
⋯
|
| 979 |
+
|
| 980 |
+
i
|
| 981 |
+
−
|
| 982 |
+
1
|
| 983 |
+
)
|
| 984 |
+
similar-to
|
| 985 |
+
subscript
|
| 986 |
+
𝑎
|
| 987 |
+
𝑖
|
| 988 |
+
superscript
|
| 989 |
+
subscript
|
| 990 |
+
𝑝
|
| 991 |
+
𝜃
|
| 992 |
+
𝑅
|
| 993 |
+
𝑒
|
| 994 |
+
𝐴
|
| 995 |
+
𝑐
|
| 996 |
+
𝑡
|
| 997 |
+
conditional
|
| 998 |
+
subscript
|
| 999 |
+
𝑎
|
| 1000 |
+
𝑖
|
| 1001 |
+
𝑥
|
| 1002 |
+
subscript
|
| 1003 |
+
𝑜
|
| 1004 |
+
1
|
| 1005 |
+
⋯
|
| 1006 |
+
𝑖
|
| 1007 |
+
1
|
| 1008 |
+
subscript
|
| 1009 |
+
𝑎
|
| 1010 |
+
1
|
| 1011 |
+
⋯
|
| 1012 |
+
𝑖
|
| 1013 |
+
1
|
| 1014 |
+
a_{i}\sim p_{\theta}^{ReAct}(a_{i}|x,o_{1\cdots i-1},a_{1\cdots i-1})
|
| 1015 |
+
, with the final output being
|
| 1016 |
+
y
|
| 1017 |
+
∼
|
| 1018 |
+
p
|
| 1019 |
+
θ
|
| 1020 |
+
R
|
| 1021 |
+
|
| 1022 |
+
e
|
| 1023 |
+
|
| 1024 |
+
A
|
| 1025 |
+
|
| 1026 |
+
c
|
| 1027 |
+
|
| 1028 |
+
t
|
| 1029 |
+
|
| 1030 |
+
(
|
| 1031 |
+
y
|
| 1032 |
+
|
|
| 1033 |
+
x
|
| 1034 |
+
,
|
| 1035 |
+
o
|
| 1036 |
+
1
|
| 1037 |
+
|
| 1038 |
+
⋯
|
| 1039 |
+
|
| 1040 |
+
n
|
| 1041 |
+
,
|
| 1042 |
+
a
|
| 1043 |
+
1
|
| 1044 |
+
|
| 1045 |
+
⋯
|
| 1046 |
+
|
| 1047 |
+
n
|
| 1048 |
+
)
|
| 1049 |
+
similar-to
|
| 1050 |
+
𝑦
|
| 1051 |
+
superscript
|
| 1052 |
+
subscript
|
| 1053 |
+
𝑝
|
| 1054 |
+
𝜃
|
| 1055 |
+
𝑅
|
| 1056 |
+
𝑒
|
| 1057 |
+
𝐴
|
| 1058 |
+
𝑐
|
| 1059 |
+
𝑡
|
| 1060 |
+
conditional
|
| 1061 |
+
𝑦
|
| 1062 |
+
𝑥
|
| 1063 |
+
subscript
|
| 1064 |
+
𝑜
|
| 1065 |
+
1
|
| 1066 |
+
⋯
|
| 1067 |
+
𝑛
|
| 1068 |
+
subscript
|
| 1069 |
+
𝑎
|
| 1070 |
+
1
|
| 1071 |
+
⋯
|
| 1072 |
+
𝑛
|
| 1073 |
+
y\sim p_{\theta}^{ReAct}(y~{}|~{}x,o_{1\cdots n},a_{1\cdots n})
|
| 1074 |
+
.
|
| 1075 |
+
While the previously described prompting techniques improve LM performance on reasoning tasks, they falter on difficult tasks that involve multifaceted decision-making due to several shortcomings: 1)
|
| 1076 |
+
Flexibility
|
| 1077 |
+
: Base prompting methods (CoT or ReAct) autoregressively sample from the LM, neglecting potential alternative continuations from specific states. 2)
|
| 1078 |
+
Sensibility
|
| 1079 |
+
: Reasoning-based methods (CoT, RAP, or ToT) rely solely on the internal representations of the LM and cannot consider external observations. This dependency risks fact hallucination and error propagation while setting a performance ceiling. 3)
|
| 1080 |
+
Adaptability
|
| 1081 |
+
: Current planning frameworks (RAP or ToT) use simple search algorithms such as BFS or cannot leverage environmental feedback to improve planning. Additionally, the agent is static and cannot reuse previous experience or learn from trial and error. While RAP also adopts MCTS, it is constrained to tasks where the LM can become a world model and accurately predict states. These shortcomings limit the ability of LMs to be deployed as general problem-solving agents and form the motivation for LATS.
|
| 1082 |
+
3.2
|
| 1083 |
+
Monte-Carlo Tree Search (MCTS)
|
| 1084 |
+
Monte-Carlo Tree Search (MCTS) is a heuristic search algorithm that is proved successful on many decision-making environments such as Atari
|
| 1085 |
+
(Ye et al.,
|
| 1086 |
+
2021
|
| 1087 |
+
)
|
| 1088 |
+
and Go
|
| 1089 |
+
(Silver et al.,
|
| 1090 |
+
2016
|
| 1091 |
+
)
|
| 1092 |
+
. MCTS builds a decision tree where every node in the tree is a state and edge is an action. MCTS runs for
|
| 1093 |
+
k
|
| 1094 |
+
𝑘
|
| 1095 |
+
k
|
| 1096 |
+
episodes; for each episode, it starts from the root (i.e., initial state) and iteratively conducts two steps to expand the tree: 1)
|
| 1097 |
+
Expansion
|
| 1098 |
+
, where multiple children states
|
| 1099 |
+
s
|
| 1100 |
+
𝑠
|
| 1101 |
+
s
|
| 1102 |
+
are explored from the current parent state
|
| 1103 |
+
p
|
| 1104 |
+
𝑝
|
| 1105 |
+
p
|
| 1106 |
+
by sampling
|
| 1107 |
+
n
|
| 1108 |
+
𝑛
|
| 1109 |
+
n
|
| 1110 |
+
actions, and 2)
|
| 1111 |
+
Selection
|
| 1112 |
+
, where the children with the highest UCT
|
| 1113 |
+
(Upper Confidence bounds applied to Trees)
|
| 1114 |
+
(Kocsis & Szepesvári,
|
| 1115 |
+
2006
|
| 1116 |
+
)
|
| 1117 |
+
value is selected by the next iteration. The UCT of a child state
|
| 1118 |
+
s
|
| 1119 |
+
𝑠
|
| 1120 |
+
s
|
| 1121 |
+
is calculated as follows:
|
| 1122 |
+
U
|
| 1123 |
+
|
| 1124 |
+
C
|
| 1125 |
+
|
| 1126 |
+
T
|
| 1127 |
+
|
| 1128 |
+
(
|
| 1129 |
+
s
|
| 1130 |
+
)
|
| 1131 |
+
=
|
| 1132 |
+
V
|
| 1133 |
+
|
| 1134 |
+
(
|
| 1135 |
+
s
|
| 1136 |
+
)
|
| 1137 |
+
+
|
| 1138 |
+
w
|
| 1139 |
+
|
| 1140 |
+
ln
|
| 1141 |
+
|
| 1142 |
+
N
|
| 1143 |
+
|
| 1144 |
+
(
|
| 1145 |
+
p
|
| 1146 |
+
)
|
| 1147 |
+
N
|
| 1148 |
+
|
| 1149 |
+
(
|
| 1150 |
+
s
|
| 1151 |
+
)
|
| 1152 |
+
,
|
| 1153 |
+
𝑈
|
| 1154 |
+
𝐶
|
| 1155 |
+
𝑇
|
| 1156 |
+
𝑠
|
| 1157 |
+
𝑉
|
| 1158 |
+
𝑠
|
| 1159 |
+
𝑤
|
| 1160 |
+
𝑁
|
| 1161 |
+
𝑝
|
| 1162 |
+
𝑁
|
| 1163 |
+
𝑠
|
| 1164 |
+
UCT(s)=V(s)+w\sqrt{\frac{\ln N(p)}{N(s)}},
|
| 1165 |
+
(1)
|
| 1166 |
+
where
|
| 1167 |
+
N
|
| 1168 |
+
|
| 1169 |
+
(
|
| 1170 |
+
s
|
| 1171 |
+
)
|
| 1172 |
+
𝑁
|
| 1173 |
+
𝑠
|
| 1174 |
+
N(s)
|
| 1175 |
+
is the number of visits to a node
|
| 1176 |
+
s
|
| 1177 |
+
𝑠
|
| 1178 |
+
s
|
| 1179 |
+
,
|
| 1180 |
+
V
|
| 1181 |
+
|
| 1182 |
+
(
|
| 1183 |
+
s
|
| 1184 |
+
)
|
| 1185 |
+
𝑉
|
| 1186 |
+
𝑠
|
| 1187 |
+
V(s)
|
| 1188 |
+
is the value function (expected return) from the subtree of
|
| 1189 |
+
s
|
| 1190 |
+
𝑠
|
| 1191 |
+
s
|
| 1192 |
+
,
|
| 1193 |
+
w
|
| 1194 |
+
𝑤
|
| 1195 |
+
w
|
| 1196 |
+
is the exploration weight, and
|
| 1197 |
+
p
|
| 1198 |
+
𝑝
|
| 1199 |
+
p
|
| 1200 |
+
is the parent node of
|
| 1201 |
+
s
|
| 1202 |
+
𝑠
|
| 1203 |
+
s
|
| 1204 |
+
. The child node with the highest UCT value is selected for expansion in the next iteration. When the end of an episode is reached, a
|
| 1205 |
+
backpropagation
|
| 1206 |
+
is carried out: the return
|
| 1207 |
+
r
|
| 1208 |
+
𝑟
|
| 1209 |
+
r
|
| 1210 |
+
is used for updating every
|
| 1211 |
+
V
|
| 1212 |
+
|
| 1213 |
+
(
|
| 1214 |
+
s
|
| 1215 |
+
)
|
| 1216 |
+
𝑉
|
| 1217 |
+
𝑠
|
| 1218 |
+
V(s)
|
| 1219 |
+
along the path
|
| 1220 |
+
with the formula
|
| 1221 |
+
V
|
| 1222 |
+
|
| 1223 |
+
(
|
| 1224 |
+
s
|
| 1225 |
+
)
|
| 1226 |
+
=
|
| 1227 |
+
V
|
| 1228 |
+
old
|
| 1229 |
+
|
| 1230 |
+
(
|
| 1231 |
+
s
|
| 1232 |
+
)
|
| 1233 |
+
|
| 1234 |
+
(
|
| 1235 |
+
N
|
| 1236 |
+
|
| 1237 |
+
(
|
| 1238 |
+
s
|
| 1239 |
+
)
|
| 1240 |
+
−
|
| 1241 |
+
1
|
| 1242 |
+
)
|
| 1243 |
+
+
|
| 1244 |
+
r
|
| 1245 |
+
N
|
| 1246 |
+
|
| 1247 |
+
(
|
| 1248 |
+
s
|
| 1249 |
+
)
|
| 1250 |
+
𝑉
|
| 1251 |
+
𝑠
|
| 1252 |
+
subscript
|
| 1253 |
+
𝑉
|
| 1254 |
+
old
|
| 1255 |
+
𝑠
|
| 1256 |
+
𝑁
|
| 1257 |
+
𝑠
|
| 1258 |
+
1
|
| 1259 |
+
𝑟
|
| 1260 |
+
𝑁
|
| 1261 |
+
𝑠
|
| 1262 |
+
V(s)=\frac{V_{\text{old}}(s)(N(s)-1)+r}{N(s)}
|
| 1263 |
+
, where
|
| 1264 |
+
V
|
| 1265 |
+
old
|
| 1266 |
+
|
| 1267 |
+
(
|
| 1268 |
+
s
|
| 1269 |
+
)
|
| 1270 |
+
subscript
|
| 1271 |
+
𝑉
|
| 1272 |
+
old
|
| 1273 |
+
𝑠
|
| 1274 |
+
V_{\text{old}}(s)
|
| 1275 |
+
is the old value function. Normally, the major shortcoming of MCTS is that it requires an environment model to undo previous steps and form a searching tree, which is often a strong assumption. However, such a limitation does not exist for LMs, as we can conveniently reset to any step by simply copy-pasting historical text input. Such a special property is the key motivation of our work.
|
| 1276 |
+
4
|
| 1277 |
+
Unifying Planning, Reasoning, and Acting
|
| 1278 |
+
4.1
|
| 1279 |
+
LM Agent
|
| 1280 |
+
LATS supports sequential reasoning or decision-making tasks on the basis of ReAct. At time step
|
| 1281 |
+
t
|
| 1282 |
+
𝑡
|
| 1283 |
+
t
|
| 1284 |
+
, an agent receives an observation
|
| 1285 |
+
o
|
| 1286 |
+
t
|
| 1287 |
+
∈
|
| 1288 |
+
O
|
| 1289 |
+
subscript
|
| 1290 |
+
𝑜
|
| 1291 |
+
𝑡
|
| 1292 |
+
𝑂
|
| 1293 |
+
o_{t}\in O
|
| 1294 |
+
from the environment and takes an action
|
| 1295 |
+
a
|
| 1296 |
+
t
|
| 1297 |
+
∈
|
| 1298 |
+
A
|
| 1299 |
+
subscript
|
| 1300 |
+
𝑎
|
| 1301 |
+
𝑡
|
| 1302 |
+
𝐴
|
| 1303 |
+
a_{t}\in A
|
| 1304 |
+
following some policy
|
| 1305 |
+
π
|
| 1306 |
+
|
| 1307 |
+
(
|
| 1308 |
+
a
|
| 1309 |
+
t
|
| 1310 |
+
|
|
| 1311 |
+
x
|
| 1312 |
+
,
|
| 1313 |
+
o
|
| 1314 |
+
1
|
| 1315 |
+
|
| 1316 |
+
⋯
|
| 1317 |
+
|
| 1318 |
+
i
|
| 1319 |
+
−
|
| 1320 |
+
1
|
| 1321 |
+
,
|
| 1322 |
+
a
|
| 1323 |
+
1
|
| 1324 |
+
|
| 1325 |
+
⋯
|
| 1326 |
+
|
| 1327 |
+
i
|
| 1328 |
+
−
|
| 1329 |
+
1
|
| 1330 |
+
)
|
| 1331 |
+
𝜋
|
| 1332 |
+
conditional
|
| 1333 |
+
subscript
|
| 1334 |
+
𝑎
|
| 1335 |
+
𝑡
|
| 1336 |
+
𝑥
|
| 1337 |
+
subscript
|
| 1338 |
+
𝑜
|
| 1339 |
+
1
|
| 1340 |
+
⋯
|
| 1341 |
+
𝑖
|
| 1342 |
+
1
|
| 1343 |
+
subscript
|
| 1344 |
+
𝑎
|
| 1345 |
+
1
|
| 1346 |
+
⋯
|
| 1347 |
+
𝑖
|
| 1348 |
+
1
|
| 1349 |
+
\pi(a_{t}|x,o_{1\cdots i-1},a_{1\cdots i-1})
|
| 1350 |
+
, where
|
| 1351 |
+
x
|
| 1352 |
+
𝑥
|
| 1353 |
+
x
|
| 1354 |
+
consists of the task instruction and a number of few-shot examples. We initialize the agent with
|
| 1355 |
+
p
|
| 1356 |
+
θ
|
| 1357 |
+
subscript
|
| 1358 |
+
𝑝
|
| 1359 |
+
𝜃
|
| 1360 |
+
p_{\theta}
|
| 1361 |
+
to leverage the useful language representations of an LM as a base decision-maker. We follow the ReAct instantiation in which the action space
|
| 1362 |
+
A
|
| 1363 |
+
^
|
| 1364 |
+
=
|
| 1365 |
+
A
|
| 1366 |
+
∪
|
| 1367 |
+
Z
|
| 1368 |
+
^
|
| 1369 |
+
𝐴
|
| 1370 |
+
𝐴
|
| 1371 |
+
𝑍
|
| 1372 |
+
\hat{A}=A\cup Z
|
| 1373 |
+
consists of both the space of permissible actions
|
| 1374 |
+
A
|
| 1375 |
+
𝐴
|
| 1376 |
+
A
|
| 1377 |
+
and language space of reasoning traces
|
| 1378 |
+
Z
|
| 1379 |
+
𝑍
|
| 1380 |
+
Z
|
| 1381 |
+
. Actions directly affect the environment and result in observation, while thoughts are used to formalize decisions by organizing information, planning future actions, or injecting internal knowledge. The exact instantiation of the action space depends on the particular environment; for decision-making tasks actions might consist of commands on a website while for reasoning tasks the action space might be limited to a few external tools or APIs.
|
| 1382 |
+
Instead of greedily decoding one trajectory or solution, we sample
|
| 1383 |
+
n
|
| 1384 |
+
𝑛
|
| 1385 |
+
n
|
| 1386 |
+
actions from
|
| 1387 |
+
p
|
| 1388 |
+
θ
|
| 1389 |
+
subscript
|
| 1390 |
+
𝑝
|
| 1391 |
+
𝜃
|
| 1392 |
+
p_{\theta}
|
| 1393 |
+
using the current state. This is based on the intuition that for complex decision-making tasks, there is likely to be a range of potential trajectories or reasoning paths that are correct
|
| 1394 |
+
(Evans,
|
| 1395 |
+
2010
|
| 1396 |
+
)
|
| 1397 |
+
. Sampling a diverse set of candidates at each step mitigates the stochastic nature of LM text generation and enables greater exploration in both the decision-making and reasoning space. We wrap
|
| 1398 |
+
p
|
| 1399 |
+
θ
|
| 1400 |
+
subscript
|
| 1401 |
+
𝑝
|
| 1402 |
+
𝜃
|
| 1403 |
+
p_{\theta}
|
| 1404 |
+
within our proposed search algorithm to deliberately construct the best trajectory from sampled actions.
|
| 1405 |
+
4.2
|
| 1406 |
+
LATS
|
| 1407 |
+
Figure 3:
|
| 1408 |
+
An overview of the six operations of LATS. A node is
|
| 1409 |
+
selected
|
| 1410 |
+
,
|
| 1411 |
+
expanded
|
| 1412 |
+
,
|
| 1413 |
+
evaluated
|
| 1414 |
+
, then
|
| 1415 |
+
simulated
|
| 1416 |
+
until a terminal node is reached, then the resulting value is
|
| 1417 |
+
backpropagated
|
| 1418 |
+
. If the trajectory fails, a
|
| 1419 |
+
reflection
|
| 1420 |
+
is generated and used as additional context for future trials. These operations are performed in succession until the budget is reached or task is successful.
|
| 1421 |
+
The main component of LATS is a search algorithm that controls the overall problem-solving process with deliberate planning. To find the most promising trajectory and systemically balance exploration with exploitation, we adopt a variant of Monte Carlo Tree Search (MCTS) that frames decision-making as a tree search, in which each node
|
| 1422 |
+
s
|
| 1423 |
+
=
|
| 1424 |
+
[
|
| 1425 |
+
x
|
| 1426 |
+
,
|
| 1427 |
+
a
|
| 1428 |
+
1
|
| 1429 |
+
|
| 1430 |
+
⋯
|
| 1431 |
+
|
| 1432 |
+
i
|
| 1433 |
+
,
|
| 1434 |
+
o
|
| 1435 |
+
1
|
| 1436 |
+
|
| 1437 |
+
⋯
|
| 1438 |
+
|
| 1439 |
+
i
|
| 1440 |
+
]
|
| 1441 |
+
𝑠
|
| 1442 |
+
𝑥
|
| 1443 |
+
subscript
|
| 1444 |
+
𝑎
|
| 1445 |
+
1
|
| 1446 |
+
⋯
|
| 1447 |
+
𝑖
|
| 1448 |
+
subscript
|
| 1449 |
+
𝑜
|
| 1450 |
+
1
|
| 1451 |
+
⋯
|
| 1452 |
+
𝑖
|
| 1453 |
+
s=[x,a_{1\cdots i},o_{1\cdots i}]
|
| 1454 |
+
represents a state comprising the original input
|
| 1455 |
+
x
|
| 1456 |
+
𝑥
|
| 1457 |
+
x
|
| 1458 |
+
, action sequence
|
| 1459 |
+
a
|
| 1460 |
+
1
|
| 1461 |
+
⋅
|
| 1462 |
+
i
|
| 1463 |
+
subscript
|
| 1464 |
+
𝑎
|
| 1465 |
+
⋅
|
| 1466 |
+
1
|
| 1467 |
+
𝑖
|
| 1468 |
+
a_{1\cdot i}
|
| 1469 |
+
, and observation sequence
|
| 1470 |
+
o
|
| 1471 |
+
1
|
| 1472 |
+
⋅
|
| 1473 |
+
i
|
| 1474 |
+
subscript
|
| 1475 |
+
𝑜
|
| 1476 |
+
⋅
|
| 1477 |
+
1
|
| 1478 |
+
𝑖
|
| 1479 |
+
o_{1\cdot i}
|
| 1480 |
+
.
|
| 1481 |
+
To adapt MCTS for language agents, LATS repurposes
|
| 1482 |
+
p
|
| 1483 |
+
θ
|
| 1484 |
+
subscript
|
| 1485 |
+
𝑝
|
| 1486 |
+
𝜃
|
| 1487 |
+
p_{\theta}
|
| 1488 |
+
as an agent, state evaluator, and feedback generator, leveraging the useful language priors of modern LMs to facilitate planning. While standard MCTS and RAP
|
| 1489 |
+
Hao et al. (
|
| 1490 |
+
2023
|
| 1491 |
+
)
|
| 1492 |
+
rely on internal dynamics models to facilitate simulation, LATS is model-free and uses environment interaction. LATS consists of a series of operations,
|
| 1493 |
+
selection, expansion, evaluation, simulation, backpropagation, and reflection
|
| 1494 |
+
, performed in succession until the task is successfully completed or a computational limit is reached. The full psuedocode of LATS can be found in Sec.
|
| 1495 |
+
A
|
| 1496 |
+
in the Appendix.
|
| 1497 |
+
Selection.
|
| 1498 |
+
In the first operation, the algorithm identifies a segment of the current tree most suitable for subsequent expansion. Starting from the root node, denoted as the initial state
|
| 1499 |
+
s
|
| 1500 |
+
0
|
| 1501 |
+
subscript
|
| 1502 |
+
𝑠
|
| 1503 |
+
0
|
| 1504 |
+
s_{0}
|
| 1505 |
+
, a child node is selected at each tree level until a leaf node is reached. To balance exploration and exploitation, we use the UCT algorithm as shown in Eq.
|
| 1506 |
+
1
|
| 1507 |
+
.
|
| 1508 |
+
Expansion.
|
| 1509 |
+
After selecting a node, the second operation expands the tree by sampling
|
| 1510 |
+
n
|
| 1511 |
+
𝑛
|
| 1512 |
+
n
|
| 1513 |
+
actions from
|
| 1514 |
+
p
|
| 1515 |
+
θ
|
| 1516 |
+
subscript
|
| 1517 |
+
𝑝
|
| 1518 |
+
𝜃
|
| 1519 |
+
p_{\theta}
|
| 1520 |
+
, as described in the prior section. The environment receives each action and returns corresponding feedback as an observation. This results in
|
| 1521 |
+
n
|
| 1522 |
+
𝑛
|
| 1523 |
+
n
|
| 1524 |
+
new child nodes added to the tree. This tree is stored in an external long-term memory structure.
|
| 1525 |
+
Evaluation.
|
| 1526 |
+
The third operation assigns a scalar value to each new child node to be used for selection and backpropagation. This value effectively quantifies the agent’s progress in task completion, serving as a heuristic to steer the search algorithm towards the most promising regions of the tree. Following
|
| 1527 |
+
Yao et al. (
|
| 1528 |
+
2023a
|
| 1529 |
+
)
|
| 1530 |
+
we repurpose
|
| 1531 |
+
p
|
| 1532 |
+
θ
|
| 1533 |
+
subscript
|
| 1534 |
+
𝑝
|
| 1535 |
+
𝜃
|
| 1536 |
+
p_{\theta}
|
| 1537 |
+
into a value function by prompting it to reason about a given state. To obtain a scalar value, we instruct
|
| 1538 |
+
p
|
| 1539 |
+
θ
|
| 1540 |
+
subscript
|
| 1541 |
+
𝑝
|
| 1542 |
+
𝜃
|
| 1543 |
+
p_{\theta}
|
| 1544 |
+
to end its reasoning trace with a score indicating the correctness of the trajectory. This method offers enhanced flexibility over programmed heuristics
|
| 1545 |
+
(Campbell et al.,
|
| 1546 |
+
2002
|
| 1547 |
+
)
|
| 1548 |
+
and greater efficiency than learned heuristics
|
| 1549 |
+
(Silver et al.,
|
| 1550 |
+
2017
|
| 1551 |
+
)
|
| 1552 |
+
.
|
| 1553 |
+
Simulation.
|
| 1554 |
+
The fourth operation expands the currently selected node until a terminal state is reached. At each depth level we sample and evaluate nodes with the same operations, but prioritize nodes of highest value. Reaching a terminal state provides objective feedback on the correctness of a trajectory. If the task is completed successfully, then LATS terminates the search. If the solution is partially successful or unsuccessful, then we perform two additional operations as described below.
|
| 1555 |
+
Backpropagation.
|
| 1556 |
+
This operation updates the values of the tree based on the outcome of a trajectory. For each node
|
| 1557 |
+
s
|
| 1558 |
+
0
|
| 1559 |
+
,
|
| 1560 |
+
s
|
| 1561 |
+
1
|
| 1562 |
+
,
|
| 1563 |
+
…
|
| 1564 |
+
,
|
| 1565 |
+
s
|
| 1566 |
+
n
|
| 1567 |
+
subscript
|
| 1568 |
+
𝑠
|
| 1569 |
+
0
|
| 1570 |
+
subscript
|
| 1571 |
+
𝑠
|
| 1572 |
+
1
|
| 1573 |
+
…
|
| 1574 |
+
subscript
|
| 1575 |
+
𝑠
|
| 1576 |
+
𝑛
|
| 1577 |
+
s_{0},s_{1},\dots,s_{n}
|
| 1578 |
+
in the trajectory from root (initial state
|
| 1579 |
+
s
|
| 1580 |
+
0
|
| 1581 |
+
subscript
|
| 1582 |
+
𝑠
|
| 1583 |
+
0
|
| 1584 |
+
s_{0}
|
| 1585 |
+
) of the searching tree to leaf (terminal state
|
| 1586 |
+
s
|
| 1587 |
+
n
|
| 1588 |
+
subscript
|
| 1589 |
+
𝑠
|
| 1590 |
+
𝑛
|
| 1591 |
+
s_{n}
|
| 1592 |
+
), its value is updated to reflect the outcome of the simulation by
|
| 1593 |
+
N
|
| 1594 |
+
|
| 1595 |
+
(
|
| 1596 |
+
s
|
| 1597 |
+
i
|
| 1598 |
+
)
|
| 1599 |
+
=
|
| 1600 |
+
N
|
| 1601 |
+
old
|
| 1602 |
+
|
| 1603 |
+
(
|
| 1604 |
+
s
|
| 1605 |
+
i
|
| 1606 |
+
)
|
| 1607 |
+
+
|
| 1608 |
+
1
|
| 1609 |
+
𝑁
|
| 1610 |
+
subscript
|
| 1611 |
+
𝑠
|
| 1612 |
+
𝑖
|
| 1613 |
+
subscript
|
| 1614 |
+
𝑁
|
| 1615 |
+
old
|
| 1616 |
+
subscript
|
| 1617 |
+
𝑠
|
| 1618 |
+
𝑖
|
| 1619 |
+
1
|
| 1620 |
+
N(s_{i})=N_{\text{old}}(s_{i})+1
|
| 1621 |
+
and
|
| 1622 |
+
V
|
| 1623 |
+
|
| 1624 |
+
(
|
| 1625 |
+
s
|
| 1626 |
+
i
|
| 1627 |
+
)
|
| 1628 |
+
=
|
| 1629 |
+
r
|
| 1630 |
+
+
|
| 1631 |
+
N
|
| 1632 |
+
old
|
| 1633 |
+
|
| 1634 |
+
(
|
| 1635 |
+
s
|
| 1636 |
+
i
|
| 1637 |
+
)
|
| 1638 |
+
|
| 1639 |
+
V
|
| 1640 |
+
old
|
| 1641 |
+
|
| 1642 |
+
(
|
| 1643 |
+
s
|
| 1644 |
+
i
|
| 1645 |
+
)
|
| 1646 |
+
N
|
| 1647 |
+
|
| 1648 |
+
(
|
| 1649 |
+
s
|
| 1650 |
+
i
|
| 1651 |
+
)
|
| 1652 |
+
𝑉
|
| 1653 |
+
subscript
|
| 1654 |
+
𝑠
|
| 1655 |
+
𝑖
|
| 1656 |
+
𝑟
|
| 1657 |
+
subscript
|
| 1658 |
+
𝑁
|
| 1659 |
+
old
|
| 1660 |
+
subscript
|
| 1661 |
+
𝑠
|
| 1662 |
+
𝑖
|
| 1663 |
+
subscript
|
| 1664 |
+
𝑉
|
| 1665 |
+
old
|
| 1666 |
+
subscript
|
| 1667 |
+
𝑠
|
| 1668 |
+
𝑖
|
| 1669 |
+
𝑁
|
| 1670 |
+
subscript
|
| 1671 |
+
𝑠
|
| 1672 |
+
𝑖
|
| 1673 |
+
V(s_{i})=\frac{r+N_{\text{old}}(s_{i})V_{\text{old}}(s_{i})}{N(s_{i})}
|
| 1674 |
+
, where
|
| 1675 |
+
r
|
| 1676 |
+
𝑟
|
| 1677 |
+
r
|
| 1678 |
+
is the return and
|
| 1679 |
+
N
|
| 1680 |
+
old
|
| 1681 |
+
,
|
| 1682 |
+
V
|
| 1683 |
+
old
|
| 1684 |
+
subscript
|
| 1685 |
+
𝑁
|
| 1686 |
+
old
|
| 1687 |
+
subscript
|
| 1688 |
+
𝑉
|
| 1689 |
+
old
|
| 1690 |
+
N_{\text{old}},V_{\text{old}}
|
| 1691 |
+
are the old number of visits and value function. These updated values are used in the UCT formula (Eq.
|
| 1692 |
+
1
|
| 1693 |
+
) to guide the selection of the next node for exploration.
|
| 1694 |
+
Reflection.
|
| 1695 |
+
In addition to the environmental feedback, we also leverage
|
| 1696 |
+
self-reflection
|
| 1697 |
+
to further refine the decision-making process
|
| 1698 |
+
(Shinn et al.,
|
| 1699 |
+
2023
|
| 1700 |
+
; Madaan et al.,
|
| 1701 |
+
2023
|
| 1702 |
+
)
|
| 1703 |
+
. Upon encountering an unsuccessful terminal node,
|
| 1704 |
+
p
|
| 1705 |
+
θ
|
| 1706 |
+
subscript
|
| 1707 |
+
𝑝
|
| 1708 |
+
𝜃
|
| 1709 |
+
p_{\theta}
|
| 1710 |
+
is prompted with the trajectory and final reward to provide a verbal self-reflection that summarizes the errors in the reasoning or acting process and proposes superior alternatives. We store both failed trajectories and corresponding reflections in the memory. In subsequent iterations, these are integrated as additional context to the agent and value function, refining both through in-context learning. This imparts a semantic gradient signal more useful than a scalar value, enabling the agent to learn from trial and error without the cost of expensive optimization processes such as reinforcement learning.
|
| 1711 |
+
Conceptually, LATS has the following advantages as a general framework for reasoning and decision-making with LM agents.
|
| 1712 |
+
(1)
|
| 1713 |
+
Generality
|
| 1714 |
+
: LATS supports both reasoning and decision-making tasks by defining a shared space of thoughts and actions. (2)
|
| 1715 |
+
Deliberate
|
| 1716 |
+
: The use of MCTS and LM value function ensures a principled search that selects options with high value while exploring promising alternatives. (3)
|
| 1717 |
+
Adaptability
|
| 1718 |
+
: LATS is designed around the use of external feedback through observations and self-reflection, enabling greater adaptation during problem-solving. (4)
|
| 1719 |
+
Flexibility
|
| 1720 |
+
: LATS can accommodate different scenarios, environments, and resource stipulations by modifying state design and tree dimensions. (5)
|
| 1721 |
+
Modularity
|
| 1722 |
+
: The base LM agent, reflection generator, and value function can be independently altered and adapted to individual LM properties.
|
| 1723 |
+
5
|
| 1724 |
+
Experiments
|
| 1725 |
+
To demonstrate the general applicability of LATS, we evaluate our method on a variety of decision-making domains that requires both reasoning and acting ability: programming
|
| 1726 |
+
(Chen et al.,
|
| 1727 |
+
2021
|
| 1728 |
+
; Austin et al.,
|
| 1729 |
+
2021
|
| 1730 |
+
)
|
| 1731 |
+
, HotPotQA
|
| 1732 |
+
(Yang et al.,
|
| 1733 |
+
2018
|
| 1734 |
+
)
|
| 1735 |
+
, and WebShop
|
| 1736 |
+
(Yao et al.,
|
| 1737 |
+
2022
|
| 1738 |
+
)
|
| 1739 |
+
.
|
| 1740 |
+
5.1
|
| 1741 |
+
HotPotQA
|
| 1742 |
+
For a task that can be approached with both reasoning-based and acting-based strategies, we consider HotPotQA
|
| 1743 |
+
(Yang et al.,
|
| 1744 |
+
2018
|
| 1745 |
+
)
|
| 1746 |
+
, a multi-hop question-answering benchmark that requires retrieval over two or more Wikipedia passages. For the action space, in addition to LM thoughts we follow the setup from
|
| 1747 |
+
Yao et al. (
|
| 1748 |
+
2023b
|
| 1749 |
+
)
|
| 1750 |
+
, which provides the agent with API calls to search and lookup information. The output of these API calls and self-generated reflections form the observation space. We use a subset of 100 questions and three few-shot examples for each method. For ToT, we use DFS as the base search algorithm and scoring with the LM as the heuristic. For all methods that involve sampling, including LATS, we sample
|
| 1751 |
+
k
|
| 1752 |
+
=
|
| 1753 |
+
50
|
| 1754 |
+
𝑘
|
| 1755 |
+
50
|
| 1756 |
+
k=50
|
| 1757 |
+
trajectories. More details and prompts can be found in Sec.
|
| 1758 |
+
D
|
| 1759 |
+
and Sec.
|
| 1760 |
+
E
|
| 1761 |
+
in the Appendix.
|
| 1762 |
+
We evaluate internal reasoning strategies by removing actions and observations from the context, corresponding to CoT
|
| 1763 |
+
(Wei et al.,
|
| 1764 |
+
2022
|
| 1765 |
+
)
|
| 1766 |
+
and its variants, CoT-SC
|
| 1767 |
+
(Wang et al.,
|
| 1768 |
+
2022
|
| 1769 |
+
)
|
| 1770 |
+
, ToT
|
| 1771 |
+
(Yao et al.,
|
| 1772 |
+
2023a
|
| 1773 |
+
)
|
| 1774 |
+
, and RAP
|
| 1775 |
+
(Hao et al.,
|
| 1776 |
+
2023
|
| 1777 |
+
)
|
| 1778 |
+
. These methods rely solely on the agent’s existing knowledge to answer the question. We also consider acting-based methods ReAct, Reflexion, and LATS, which augment the agent with the interactive API environment and primarily evaluate its information retrieval abilities. While LATS is designed for scenarios where external feedback can enhance reasoning, we also implement a reasoning-only version with CoT as the base prompt. We also combine internal and external reasoning in LATS by first prompting with a CoT-based prompt, then switching to a ReAct-based prompt upon failure. This is closer to how humans might approach this task, by using tools to lookup additional information only when the answer is not already known.
|
| 1779 |
+
Prompt Method
|
| 1780 |
+
HotpotQA (EM)
|
| 1781 |
+
I/O
|
| 1782 |
+
0.32
|
| 1783 |
+
CoT
|
| 1784 |
+
(Wei et al.,
|
| 1785 |
+
2022
|
| 1786 |
+
)
|
| 1787 |
+
0.34
|
| 1788 |
+
CoT - SC
|
| 1789 |
+
(Wang et al.,
|
| 1790 |
+
2022
|
| 1791 |
+
)
|
| 1792 |
+
0.38
|
| 1793 |
+
ToT
|
| 1794 |
+
(Yao et al.,
|
| 1795 |
+
2023a
|
| 1796 |
+
)
|
| 1797 |
+
0.55
|
| 1798 |
+
RAP
|
| 1799 |
+
(Hao et al.,
|
| 1800 |
+
2023
|
| 1801 |
+
)
|
| 1802 |
+
0.60
|
| 1803 |
+
RAP (n = 10)
|
| 1804 |
+
0.60
|
| 1805 |
+
LATS (CoT)
|
| 1806 |
+
0.60
|
| 1807 |
+
Prompt Method
|
| 1808 |
+
HotpotQA (EM)
|
| 1809 |
+
ReAct
|
| 1810 |
+
(Yao et al.,
|
| 1811 |
+
2023b
|
| 1812 |
+
)
|
| 1813 |
+
0.32
|
| 1814 |
+
ReAct (best of k)
|
| 1815 |
+
0.38
|
| 1816 |
+
Reflexion
|
| 1817 |
+
(Shinn et al.,
|
| 1818 |
+
2023
|
| 1819 |
+
)
|
| 1820 |
+
0.51
|
| 1821 |
+
LATS
|
| 1822 |
+
0.61
|
| 1823 |
+
LATS (n = 3)
|
| 1824 |
+
0.56
|
| 1825 |
+
LATS (n = 10)
|
| 1826 |
+
0.64
|
| 1827 |
+
LATS (CoT + ReAct)
|
| 1828 |
+
0.71
|
| 1829 |
+
Table 2:
|
| 1830 |
+
GPT-3.5 reasoning-based prompting (left) and acting-based prompting (right) results on HotpotQA. LATS achieves the highest exact match (EM) for acting and is competitive on reasoning. Unless otherwise specified, we sample
|
| 1831 |
+
n
|
| 1832 |
+
=
|
| 1833 |
+
5
|
| 1834 |
+
𝑛
|
| 1835 |
+
5
|
| 1836 |
+
n=5
|
| 1837 |
+
nodes during expansion and
|
| 1838 |
+
k
|
| 1839 |
+
=
|
| 1840 |
+
50
|
| 1841 |
+
𝑘
|
| 1842 |
+
50
|
| 1843 |
+
k=50
|
| 1844 |
+
trajectories.
|
| 1845 |
+
Results.
|
| 1846 |
+
We observe in Tab.
|
| 1847 |
+
2
|
| 1848 |
+
that both internal reasoning and external retrieval strategies perform well on HotPotQA. Due to their large-scale training corpus, modern LLMs already encode factual knowledge and can often directly answer the question correctly. While CoT can slightly enhance performance on questions requiring reasoning, larger gains are observed with search methods ToT and RAP, which can sample and explore more outputs. We observe similar results for acting-based methods. LATS surpasses ReAct, even when sampling the same number of trajectories, by expanding more nodes with principled search (see Fig.
|
| 1849 |
+
5
|
| 1850 |
+
in Appendix
|
| 1851 |
+
D
|
| 1852 |
+
for a qualitative sample). This is demonstrated when modifying
|
| 1853 |
+
n
|
| 1854 |
+
𝑛
|
| 1855 |
+
n
|
| 1856 |
+
, the number of nodes expanded during each iteration. Increasing
|
| 1857 |
+
n
|
| 1858 |
+
𝑛
|
| 1859 |
+
n
|
| 1860 |
+
can consistently improve performance, although at greater computational and inference costs. LATS is also competitive to RAP on internal reasoning but performs worse than acting. Combining internal and external reasoning in LATS results in the highest performance, indicating the importance of external feedback in augmenting reasoning even in tasks the base LM can already perform.
|
| 1861 |
+
5.2
|
| 1862 |
+
Programming
|
| 1863 |
+
Prompt Method
|
| 1864 |
+
Model
|
| 1865 |
+
Pass@1
|
| 1866 |
+
CoT
|
| 1867 |
+
(Wei et al.,
|
| 1868 |
+
2022
|
| 1869 |
+
)
|
| 1870 |
+
GPT-3.5
|
| 1871 |
+
46.9
|
| 1872 |
+
ReAct
|
| 1873 |
+
(Yao et al.,
|
| 1874 |
+
2023b
|
| 1875 |
+
)
|
| 1876 |
+
GPT-3.5
|
| 1877 |
+
56.9
|
| 1878 |
+
Reflexion
|
| 1879 |
+
(Shinn et al.,
|
| 1880 |
+
2023
|
| 1881 |
+
)
|
| 1882 |
+
GPT-3.5
|
| 1883 |
+
68.1
|
| 1884 |
+
ToT
|
| 1885 |
+
(Yao et al.,
|
| 1886 |
+
2023a
|
| 1887 |
+
)
|
| 1888 |
+
GPT-3.5
|
| 1889 |
+
54.4
|
| 1890 |
+
RAP
|
| 1891 |
+
(Hao et al.,
|
| 1892 |
+
2023
|
| 1893 |
+
)
|
| 1894 |
+
GPT-3.5
|
| 1895 |
+
63.1
|
| 1896 |
+
LATS (Ours)
|
| 1897 |
+
GPT-3.5
|
| 1898 |
+
83.8
|
| 1899 |
+
I/O
|
| 1900 |
+
GPT-4
|
| 1901 |
+
80.1
|
| 1902 |
+
Reflexion
|
| 1903 |
+
GPT-4
|
| 1904 |
+
91.0
|
| 1905 |
+
LATS
|
| 1906 |
+
GPT-4
|
| 1907 |
+
94.4
|
| 1908 |
+
Prompt Method
|
| 1909 |
+
Pass@1
|
| 1910 |
+
CoT
|
| 1911 |
+
(Wei et al.,
|
| 1912 |
+
2022
|
| 1913 |
+
)
|
| 1914 |
+
54.9
|
| 1915 |
+
ReAct
|
| 1916 |
+
(Wei et al.,
|
| 1917 |
+
2022
|
| 1918 |
+
)
|
| 1919 |
+
67.0
|
| 1920 |
+
Reflexion
|
| 1921 |
+
(Shinn et al.,
|
| 1922 |
+
2023
|
| 1923 |
+
)
|
| 1924 |
+
70.0
|
| 1925 |
+
ToT
|
| 1926 |
+
(Yao et al.,
|
| 1927 |
+
2023a
|
| 1928 |
+
)
|
| 1929 |
+
65.8
|
| 1930 |
+
RAP
|
| 1931 |
+
(Hao et al.,
|
| 1932 |
+
2023
|
| 1933 |
+
)
|
| 1934 |
+
71.4
|
| 1935 |
+
LATS (Ours)
|
| 1936 |
+
81.1
|
| 1937 |
+
Table 3:
|
| 1938 |
+
GPT-3.5 and GPT-4 Pass@1 accuracy on HumanEval
|
| 1939 |
+
(Chen et al.,
|
| 1940 |
+
2021
|
| 1941 |
+
)
|
| 1942 |
+
and MBPP
|
| 1943 |
+
(Austin et al.,
|
| 1944 |
+
2021
|
| 1945 |
+
)
|
| 1946 |
+
. Prompting with LATS achieves the highest performance. We sample 5 solutions during expansion for
|
| 1947 |
+
8
|
| 1948 |
+
iterations.
|
| 1949 |
+
To demonstrate the importance of external observations for complex reasoning tasks, we evaluate the baselines and LATS on programming with Humaneval
|
| 1950 |
+
(Chen et al.,
|
| 1951 |
+
2021
|
| 1952 |
+
)
|
| 1953 |
+
and MBPP
|
| 1954 |
+
(Austin et al.,
|
| 1955 |
+
2021
|
| 1956 |
+
)
|
| 1957 |
+
. Both datasets measure the correctness of synthesized programs in Python from natural language docstrings. We use individual solutions as the action space and test suite and compiler feedback as the external observation. We follow
|
| 1958 |
+
Chen et al. (
|
| 1959 |
+
2022a
|
| 1960 |
+
)
|
| 1961 |
+
and use an LLM to generate a synthetic test suite of syntactically valid “assert” statements for each question. For each step, the solution is evaluated on this test suite, and the results including successful and failed tests and compiler output, are added to the context as an observation. We use the same test suite for Reflexion.
|
| 1962 |
+
For this task, the reasoning and acting baselines share an action space, but acting methods are able to incorporate observations as additional context. For LATS, since each action corresponds to a complete solution, we skip the simulation step of LATS and directly use the percentage of passed tests as the backpropagated reward. We use
|
| 1963 |
+
k
|
| 1964 |
+
=
|
| 1965 |
+
8
|
| 1966 |
+
𝑘
|
| 1967 |
+
8
|
| 1968 |
+
k=8
|
| 1969 |
+
iterations, set the number of generated tests at
|
| 1970 |
+
4
|
| 1971 |
+
4
|
| 1972 |
+
4
|
| 1973 |
+
, and sample
|
| 1974 |
+
n
|
| 1975 |
+
=
|
| 1976 |
+
5
|
| 1977 |
+
𝑛
|
| 1978 |
+
5
|
| 1979 |
+
n=5
|
| 1980 |
+
solutions during expansion. After the search is completed, we select the solution with the highest value and evaluate it on the real test suite for the pass@1 accuracy evaluation. More details and prompts can be found in Sec.
|
| 1981 |
+
D
|
| 1982 |
+
and Sec.
|
| 1983 |
+
F
|
| 1984 |
+
in the Appendix.
|
| 1985 |
+
Results.
|
| 1986 |
+
We find in Tab
|
| 1987 |
+
3
|
| 1988 |
+
that both search and semantic feedback are crucial for better performance. Despite not using observations, ToT and RAP are competitive with Reflexion. LATS has the highest performance on both datasets. Since RAP uses a similar search algorithm as LATS, this reveals the importance of external feedback for difficult reasoning tasks such as programming. With GPT-4, using LATS sets the state of the art for HumanEval, showing LATS can be used with more advanced LLMs for higher performance.
|
| 1989 |
+
5.3
|
| 1990 |
+
Webshop
|
| 1991 |
+
For a complex decision-making environment with practical applications, we consider WebShop
|
| 1992 |
+
(Yao et al.,
|
| 1993 |
+
2022
|
| 1994 |
+
)
|
| 1995 |
+
, an online shopping environment composed of a website with 1.18M real-world products and 12k human instructions. Agents must navigate a website through a variety of commands to purchase an item matching a user specification. We use the preconstructed action space of search and click commands and browser feedback and reflections for the observation. The performance is gauged using two metrics: an average score, reflecting the percentage of user-specified attributes met by the selected product, and a success rate, indicating the frequency with which the chosen product fulfills all given conditions. We compare against acting-based prompting methods and RL-based approaches. We evaluate on 50 instructions, expand
|
| 1996 |
+
n
|
| 1997 |
+
=
|
| 1998 |
+
5
|
| 1999 |
+
𝑛
|
| 2000 |
+
5
|
| 2001 |
+
n=5
|
| 2002 |
+
children for LATS, and set
|
| 2003 |
+
k
|
| 2004 |
+
=
|
| 2005 |
+
30
|
| 2006 |
+
𝑘
|
| 2007 |
+
30
|
| 2008 |
+
k=30
|
| 2009 |
+
for LATS, ReAct best of
|
| 2010 |
+
k
|
| 2011 |
+
𝑘
|
| 2012 |
+
k
|
| 2013 |
+
, and Reflexion. More details and prompts are in Appendix
|
| 2014 |
+
D
|
| 2015 |
+
and
|
| 2016 |
+
G
|
| 2017 |
+
.
|
| 2018 |
+
Results.
|
| 2019 |
+
We find in Tab.
|
| 2020 |
+
5
|
| 2021 |
+
that GPT-3.5 with ReAct is competitive to imitation learning, and can exceed reinforcement learning techniques with stronger prompting strategies. Sampling
|
| 2022 |
+
k
|
| 2023 |
+
=
|
| 2024 |
+
30
|
| 2025 |
+
𝑘
|
| 2026 |
+
30
|
| 2027 |
+
k=30
|
| 2028 |
+
trajectories with ReAct and Reflexion results in a similar performance, suggesting the semantic feedback is not as helpful in complex environments like WebShop. Indeed like in
|
| 2029 |
+
Shinn et al. (
|
| 2030 |
+
2023
|
| 2031 |
+
)
|
| 2032 |
+
, we find that generated reflections are often generic and do not provide useful feedback, resulting in a tendency for the agent to become stuck in local minima. However, using LATS indeed results in a noticeable improvement, indicating a more effective exploration for the same number of iterations.
|
| 2033 |
+
5.4
|
| 2034 |
+
Additional Observations
|
| 2035 |
+
Method
|
| 2036 |
+
Score
|
| 2037 |
+
SR
|
| 2038 |
+
ReAct
|
| 2039 |
+
(Yao et al.,
|
| 2040 |
+
2023b
|
| 2041 |
+
)
|
| 2042 |
+
53.8
|
| 2043 |
+
28.0
|
| 2044 |
+
ReAct (best of k)
|
| 2045 |
+
59.1
|
| 2046 |
+
32.0
|
| 2047 |
+
Reflexion
|
| 2048 |
+
(Shinn et al.,
|
| 2049 |
+
2023
|
| 2050 |
+
)
|
| 2051 |
+
64.2
|
| 2052 |
+
35.0
|
| 2053 |
+
LATS
|
| 2054 |
+
75.9
|
| 2055 |
+
38.0
|
| 2056 |
+
IL
|
| 2057 |
+
59.9
|
| 2058 |
+
29.1
|
| 2059 |
+
IL+RL
|
| 2060 |
+
62.4
|
| 2061 |
+
28.7
|
| 2062 |
+
Fine-tuning
|
| 2063 |
+
(Furuta et al.,
|
| 2064 |
+
2023
|
| 2065 |
+
)
|
| 2066 |
+
67.5
|
| 2067 |
+
45.0
|
| 2068 |
+
Expert
|
| 2069 |
+
82.1
|
| 2070 |
+
59.6
|
| 2071 |
+
Table 4:
|
| 2072 |
+
Score and success rate (SR) on Webshop. Table is separated into prompting, RL-based training, and human performance. For the same number of iterations, LATS improves both score and success rate, and surpasses RL-based training. IL/IL+RL taken from
|
| 2073 |
+
Yao et al. (
|
| 2074 |
+
2022
|
| 2075 |
+
)
|
| 2076 |
+
.
|
| 2077 |
+
Prompt Method
|
| 2078 |
+
HotPotQA (EM)
|
| 2079 |
+
ToT (ReAct)
|
| 2080 |
+
0.39
|
| 2081 |
+
RAP (ReAct)
|
| 2082 |
+
0.54
|
| 2083 |
+
LATS (No LM Heuristic)
|
| 2084 |
+
0.37
|
| 2085 |
+
LATS (DFS)
|
| 2086 |
+
0.42
|
| 2087 |
+
LATS (No Reflection)
|
| 2088 |
+
0.56
|
| 2089 |
+
LATS
|
| 2090 |
+
0.61
|
| 2091 |
+
Table 5:
|
| 2092 |
+
Ablation results on LATS and baseline variants in HotPotQA; we use ReAct as the base prompt and sample
|
| 2093 |
+
n
|
| 2094 |
+
=
|
| 2095 |
+
5
|
| 2096 |
+
𝑛
|
| 2097 |
+
5
|
| 2098 |
+
n=5
|
| 2099 |
+
children and
|
| 2100 |
+
k
|
| 2101 |
+
=
|
| 2102 |
+
50
|
| 2103 |
+
𝑘
|
| 2104 |
+
50
|
| 2105 |
+
k=50
|
| 2106 |
+
maximum trajectories. LATS requires every component and operation for optimal performance.
|
| 2107 |
+
We also conduct additional experiments on HotPotQA to demonstrate the effect of each component of LATS. We also design a version of ToT and RAP with ReAct prompt and can handle external observations. We use HotPotQA as our setup incorporates both reasoning (through thoughts) and acting (through API calls); the results are shown in Tab.
|
| 2108 |
+
5
|
| 2109 |
+
. More ablations for token consumption on HotPotQA are in Tab.
|
| 2110 |
+
7
|
| 2111 |
+
in Appendix
|
| 2112 |
+
C
|
| 2113 |
+
. Note that baselines generally perform worse than the reasoning-only setting of HotPotQA, which indicates that the acting-based setting is more challenging and adaption of search algorithms to decision-making scenarios is non-trivial.
|
| 2114 |
+
Self-reflection.
|
| 2115 |
+
We use self-reflection to provide additional semantic signals for the agent. We observe a
|
| 2116 |
+
0.05
|
| 2117 |
+
0.05
|
| 2118 |
+
0.05
|
| 2119 |
+
performance drop when removed from LATS, suggesting this is useful. This is a smaller gain Reflexion
|
| 2120 |
+
(Shinn et al.,
|
| 2121 |
+
2023
|
| 2122 |
+
)
|
| 2123 |
+
observes over ReAct
|
| 2124 |
+
(Yao et al.,
|
| 2125 |
+
2023b
|
| 2126 |
+
)
|
| 2127 |
+
as shown in Tab.
|
| 2128 |
+
2
|
| 2129 |
+
, suggesting overlap between the types of questions where there is an improvement with self-reflection and search. This variant outperforms RAP-ReAct, reflecting our improvements to MCTS.
|
| 2130 |
+
Search Algorithm.
|
| 2131 |
+
MCTS is a more principled search algorithm than variants like A* or DFS search and the basis for observed performance gains. We observe the effects of using DFS, and incorporate the LM-based heuristic used in ToT
|
| 2132 |
+
(Yao et al.,
|
| 2133 |
+
2023a
|
| 2134 |
+
)
|
| 2135 |
+
in which branches with low values are pruned. This removes the selection and backpropagation operations, and we observe a
|
| 2136 |
+
0.08
|
| 2137 |
+
0.08
|
| 2138 |
+
0.08
|
| 2139 |
+
drop in performance when sampling the same number of nodes, but outperforms ToT-ReAct.
|
| 2140 |
+
6
|
| 2141 |
+
Conclusion
|
| 2142 |
+
In this work, we introduce Language Agent Tree Search (LATS), the first framework to unify planning, acting, and reasoning for enhanced LLM problem solving. By deliberately constructing trajectories with search algorithms, incorporating external feedback, and enabling agents to learn from experience, LATS addresses key limitations of prior prompting techniques. Our evaluations demonstrate the ability of LATS to harness LLM capabilities for a variety of decision-making tasks while keeping its reasoning ability without additional training. The proposed synergies between search, interaction, and reflection offer a versatile approach to autonomous decision-making, highlighting the potential of LLMs as generalist agents. A full discussion of the limitations and broader impacts is in Appendix
|
| 2143 |
+
B
|
| 2144 |
+
.
|
| 2145 |
+
References
|
| 2146 |
+
Ahn et al. (2022)
|
| 2147 |
+
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan, and Andy Zeng.
|
| 2148 |
+
Do as i can, not as i say: Grounding language in robotic affordances.
|
| 2149 |
+
arXiv:2204.01691
|
| 2150 |
+
, 2022.
|
| 2151 |
+
Anthony et al. (2017)
|
| 2152 |
+
T. Anthony, Z. Tian, and D. Barber.
|
| 2153 |
+
Thinking fast and slow with deep learning and tree search.
|
| 2154 |
+
In
|
| 2155 |
+
NIPS
|
| 2156 |
+
, 2017.
|
| 2157 |
+
Austin et al. (2021)
|
| 2158 |
+
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al.
|
| 2159 |
+
Program synthesis with large language models.
|
| 2160 |
+
arXiv:2108.07732
|
| 2161 |
+
, 2021.
|
| 2162 |
+
Baker et al. (2022)
|
| 2163 |
+
Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune.
|
| 2164 |
+
Video pretraining (vpt): Learning to act by watching unlabeled online videos.
|
| 2165 |
+
arXiv:2206.11795
|
| 2166 |
+
, 2022.
|
| 2167 |
+
Besta et al. (2023)
|
| 2168 |
+
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Michal Podstawski, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler.
|
| 2169 |
+
Graph of thoughts: Solving elaborate problems with large language models.
|
| 2170 |
+
arXiv:2308.09687
|
| 2171 |
+
, 2023.
|
| 2172 |
+
Bowman et al. (2015)
|
| 2173 |
+
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning.
|
| 2174 |
+
A large annotated corpus for learning natural language inference.
|
| 2175 |
+
In
|
| 2176 |
+
EMNLP
|
| 2177 |
+
, 2015.
|
| 2178 |
+
Brown et al. (2020)
|
| 2179 |
+
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei.
|
| 2180 |
+
Language models are few-shot learners.
|
| 2181 |
+
In
|
| 2182 |
+
NeurIPS
|
| 2183 |
+
, 2020.
|
| 2184 |
+
Campbell et al. (2002)
|
| 2185 |
+
Murray Campbell, A Joseph Hoane Jr, and Feng-hsiung Hsu.
|
| 2186 |
+
Deep blue.
|
| 2187 |
+
Artificial intelligence
|
| 2188 |
+
, 2002.
|
| 2189 |
+
Chen et al. (2022a)
|
| 2190 |
+
Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen.
|
| 2191 |
+
Codet: Code generation with generated tests.
|
| 2192 |
+
arXiv:2207.10397
|
| 2193 |
+
, 2022a.
|
| 2194 |
+
Chen et al. (2021)
|
| 2195 |
+
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al.
|
| 2196 |
+
Evaluating large language models trained on code.
|
| 2197 |
+
arXiv:2107.03374
|
| 2198 |
+
, 2021.
|
| 2199 |
+
Chen et al. (2022b)
|
| 2200 |
+
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen.
|
| 2201 |
+
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks.
|
| 2202 |
+
arXiv preprint arXiv:2211.12588
|
| 2203 |
+
, 2022b.
|
| 2204 |
+
Chowdhery et al. (2022)
|
| 2205 |
+
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al.
|
| 2206 |
+
Palm: Scaling language modeling with pathways.
|
| 2207 |
+
arXiv:2204.02311
|
| 2208 |
+
, 2022.
|
| 2209 |
+
Cobbe et al. (2021)
|
| 2210 |
+
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al.
|
| 2211 |
+
Training verifiers to solve math word problems.
|
| 2212 |
+
arXiv:2110.14168
|
| 2213 |
+
, 2021.
|
| 2214 |
+
Deng et al. (2023)
|
| 2215 |
+
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su.
|
| 2216 |
+
Mind2web: Towards a generalist agent for the web.
|
| 2217 |
+
arXiv:2306.06070
|
| 2218 |
+
, 2023.
|
| 2219 |
+
Driess et al. (2023)
|
| 2220 |
+
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence.
|
| 2221 |
+
Palm-e: An embodied multimodal language model.
|
| 2222 |
+
arXiv:2303.03378
|
| 2223 |
+
, 2023.
|
| 2224 |
+
Du et al. (2023)
|
| 2225 |
+
Yilun Du, Mengjiao Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Joshua B. Tenenbaum, Dale Schuurmans, and Pieter Abbeel.
|
| 2226 |
+
Learning universal policies via text-guided video generation.
|
| 2227 |
+
arXiv:2302.00111
|
| 2228 |
+
, 2023.
|
| 2229 |
+
Evans (2010)
|
| 2230 |
+
Jonathan St BT Evans.
|
| 2231 |
+
Intuition and reasoning: A dual-process perspective.
|
| 2232 |
+
Psychological Inquiry
|
| 2233 |
+
, 2010.
|
| 2234 |
+
Fan et al. (2022)
|
| 2235 |
+
Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar.
|
| 2236 |
+
Minedojo: Building open-ended embodied agents with internet-scale knowledge.
|
| 2237 |
+
In
|
| 2238 |
+
NeurIPS Datasets and Benchmarks Track
|
| 2239 |
+
, 2022.
|
| 2240 |
+
Furuta et al. (2023)
|
| 2241 |
+
Hiroki Furuta, Ofir Nachum, Kuang-Huei Lee, Yutaka Matsuo, Shixiang Shane Gu, and Izzeddin Gur.
|
| 2242 |
+
Multimodal web navigation with instruction-finetuned foundation models.
|
| 2243 |
+
arXiv preprint arXiv:2305.11854
|
| 2244 |
+
, 2023.
|
| 2245 |
+
Gao et al. (2022)
|
| 2246 |
+
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig.
|
| 2247 |
+
Pal: Program-aided language models.
|
| 2248 |
+
arXiv preprint arXiv:2211.10435
|
| 2249 |
+
, 2022.
|
| 2250 |
+
Guo et al. (2018)
|
| 2251 |
+
Jiaxian Guo, Sidi Lu, Han Cai, Weinan Zhang, Yong Yu, and Jun Wang.
|
| 2252 |
+
Long text generation via adversarial training with leaked information.
|
| 2253 |
+
AAAI
|
| 2254 |
+
, 2018.
|
| 2255 |
+
Guss et al. (2019)
|
| 2256 |
+
William H. Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela Veloso, and Ruslan Salakhutdinov.
|
| 2257 |
+
Minerl: A large-scale dataset of minecraft demonstrations.
|
| 2258 |
+
In
|
| 2259 |
+
IJCAI
|
| 2260 |
+
, 2019.
|
| 2261 |
+
Hafner et al. (2019)
|
| 2262 |
+
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson.
|
| 2263 |
+
Learning latent dynamics for planning from pixels.
|
| 2264 |
+
In
|
| 2265 |
+
ICML
|
| 2266 |
+
, 2019.
|
| 2267 |
+
Hafner et al. (2023)
|
| 2268 |
+
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap.
|
| 2269 |
+
Mastering diverse domains through world models.
|
| 2270 |
+
arXiv:2301.04104
|
| 2271 |
+
, 2023.
|
| 2272 |
+
Hao et al. (2023)
|
| 2273 |
+
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu.
|
| 2274 |
+
Reasoning with language model is planning with world model.
|
| 2275 |
+
arXiv:2305.14992
|
| 2276 |
+
, 2023.
|
| 2277 |
+
Huang et al. (2023)
|
| 2278 |
+
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou.
|
| 2279 |
+
Large language models cannot self-correct reasoning yet.
|
| 2280 |
+
arXiv:2310.01798
|
| 2281 |
+
, 2023.
|
| 2282 |
+
Huang et al. (2022)
|
| 2283 |
+
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al.
|
| 2284 |
+
Inner monologue: Embodied reasoning through planning with language models.
|
| 2285 |
+
arXiv:2207.05608
|
| 2286 |
+
, 2022.
|
| 2287 |
+
Jiang et al. (2018)
|
| 2288 |
+
D. Jiang, E. Ekwedike, and H. Liu.
|
| 2289 |
+
Feedback-based tree search for reinforcement learning.
|
| 2290 |
+
In
|
| 2291 |
+
ICML
|
| 2292 |
+
, 2018.
|
| 2293 |
+
Kocsis & Szepesvári (2006)
|
| 2294 |
+
Levente Kocsis and Csaba Szepesvári.
|
| 2295 |
+
Bandit based monte-carlo planning.
|
| 2296 |
+
In
|
| 2297 |
+
ECML
|
| 2298 |
+
, 2006.
|
| 2299 |
+
Kojima et al. (2022)
|
| 2300 |
+
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa.
|
| 2301 |
+
Large language models are zero-shot reasoners.
|
| 2302 |
+
arXiv:2205.11916
|
| 2303 |
+
, 2022.
|
| 2304 |
+
LaValle et al. (2001)
|
| 2305 |
+
Steven M LaValle, James J Kuffner, BR Donald, et al.
|
| 2306 |
+
Rapidly-exploring random trees: Progress and prospects.
|
| 2307 |
+
Algorithmic and computational robotics: new directions
|
| 2308 |
+
, 2001.
|
| 2309 |
+
Liu et al. (2018)
|
| 2310 |
+
Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang.
|
| 2311 |
+
Reinforcement learning on web interfaces using workflow-guided exploration.
|
| 2312 |
+
In
|
| 2313 |
+
ICLR
|
| 2314 |
+
, 2018.
|
| 2315 |
+
Liu et al. (2023)
|
| 2316 |
+
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang.
|
| 2317 |
+
Agentbench: Evaluating llms as agents.
|
| 2318 |
+
arXiv:2308.03688
|
| 2319 |
+
, 2023.
|
| 2320 |
+
Madaan et al. (2023)
|
| 2321 |
+
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark.
|
| 2322 |
+
Self-refine: Iterative refinement with self-feedback.
|
| 2323 |
+
arXiv:2303.17651
|
| 2324 |
+
, 2023.
|
| 2325 |
+
Nallapati et al. (2016)
|
| 2326 |
+
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Caglar Gulcehre, and Bing Xiang.
|
| 2327 |
+
Abstractive text summarization using sequence-to-sequence rnns and beyond.
|
| 2328 |
+
In
|
| 2329 |
+
SIGNLL
|
| 2330 |
+
, 2016.
|
| 2331 |
+
Nye et al. (2021)
|
| 2332 |
+
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al.
|
| 2333 |
+
Show your work: Scratchpads for intermediate computation with language models.
|
| 2334 |
+
arXiv:2112.00114
|
| 2335 |
+
, 2021.
|
| 2336 |
+
OpenAI (2023)
|
| 2337 |
+
OpenAI.
|
| 2338 |
+
Gpt-4 technical report.
|
| 2339 |
+
arXiv:2303.08774
|
| 2340 |
+
, 2023.
|
| 2341 |
+
Saparov & He (2022)
|
| 2342 |
+
Abulhair Saparov and He He.
|
| 2343 |
+
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought.
|
| 2344 |
+
arXiv:2210.01240
|
| 2345 |
+
, 2022.
|
| 2346 |
+
Schick et al. (2023)
|
| 2347 |
+
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom.
|
| 2348 |
+
Toolformer: Language models can teach themselves to use tools.
|
| 2349 |
+
arXiv:2302.04761
|
| 2350 |
+
, 2023.
|
| 2351 |
+
Shen et al. (2023)
|
| 2352 |
+
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang.
|
| 2353 |
+
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface.
|
| 2354 |
+
arXiv:2303.17580
|
| 2355 |
+
, 2023.
|
| 2356 |
+
Shinn et al. (2023)
|
| 2357 |
+
Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao.
|
| 2358 |
+
Reflexion: Language agents with verbal reinforcement learning.
|
| 2359 |
+
arXiv:2303.11366
|
| 2360 |
+
, 2023.
|
| 2361 |
+
Shridhar et al. (2020)
|
| 2362 |
+
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht.
|
| 2363 |
+
Alfworld: Aligning text and embodied environments for interactive learning.
|
| 2364 |
+
arXiv:2010.03768
|
| 2365 |
+
, 2020.
|
| 2366 |
+
Silver et al. (2016)
|
| 2367 |
+
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al.
|
| 2368 |
+
Mastering the game of go with deep neural networks and tree search.
|
| 2369 |
+
nature
|
| 2370 |
+
, 2016.
|
| 2371 |
+
Silver et al. (2017)
|
| 2372 |
+
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy P. Lillicrap, Fan Hui, L. Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis.
|
| 2373 |
+
Mastering the game of go without human knowledge.
|
| 2374 |
+
Nature
|
| 2375 |
+
, 2017.
|
| 2376 |
+
Sloman (1996)
|
| 2377 |
+
Steven A. Sloman.
|
| 2378 |
+
The empirical case for two systems of reasoning.
|
| 2379 |
+
Psychological Bulletin
|
| 2380 |
+
, 1996.
|
| 2381 |
+
Sun et al. (2023)
|
| 2382 |
+
Haotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai, and Chao Zhang.
|
| 2383 |
+
Adaplanner: Adaptive planning from feedback with language models.
|
| 2384 |
+
arXiv:2305.16653
|
| 2385 |
+
, 2023.
|
| 2386 |
+
Surís et al. (2023)
|
| 2387 |
+
Dídac Surís, Sachit Menon, and Carl Vondrick.
|
| 2388 |
+
Vipergpt: Visual inference via python execution for reasoning.
|
| 2389 |
+
arXiv preprint arXiv:2303.08128
|
| 2390 |
+
, 2023.
|
| 2391 |
+
Świechowski et al. (2023)
|
| 2392 |
+
Maciej Świechowski, Konrad Godlewski, Bartosz Sawicki, and Jacek Mańdziuk.
|
| 2393 |
+
Monte carlo tree search: A review of recent modifications and applications.
|
| 2394 |
+
Artificial Intelligence Review
|
| 2395 |
+
, 2023.
|
| 2396 |
+
Touvron et al. (2023)
|
| 2397 |
+
Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Cantón Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony S. Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel M. Kloumann, A. V. Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, R. Subramanian, Xia Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zhengxu Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and
|
| 2398 |
+
Thomas Scialom.
|
| 2399 |
+
Llama 2: Open foundation and fine-tuned chat models.
|
| 2400 |
+
arXiv:2307.09288
|
| 2401 |
+
, 2023.
|
| 2402 |
+
Vodopivec et al. (2017)
|
| 2403 |
+
Tom Vodopivec, Spyridon Samothrakis, and Branko Ster.
|
| 2404 |
+
On monte carlo tree search and reinforcement learning.
|
| 2405 |
+
Journal of Artificial Intelligence Research
|
| 2406 |
+
, 2017.
|
| 2407 |
+
Wang et al. (2023)
|
| 2408 |
+
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar.
|
| 2409 |
+
Voyager: An open-ended embodied agent with large language models.
|
| 2410 |
+
arXiv:2305.16291
|
| 2411 |
+
, 2023.
|
| 2412 |
+
Wang et al. (2022)
|
| 2413 |
+
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou.
|
| 2414 |
+
Self-consistency improves chain of thought reasoning in language models.
|
| 2415 |
+
arXiv:2203.11171
|
| 2416 |
+
, 2022.
|
| 2417 |
+
Wei et al. (2022)
|
| 2418 |
+
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou.
|
| 2419 |
+
Chain of thought prompting elicits reasoning in large language models.
|
| 2420 |
+
arXiv:2201.11903
|
| 2421 |
+
, 2022.
|
| 2422 |
+
Wooldridge & Jennings (1995)
|
| 2423 |
+
Michael Wooldridge and Nicholas R Jennings.
|
| 2424 |
+
Intelligent agents: Theory and practice.
|
| 2425 |
+
The knowledge engineering review
|
| 2426 |
+
, 1995.
|
| 2427 |
+
Wu et al. (2023)
|
| 2428 |
+
Philipp Wu, Alejandro Escontrela, Danijar Hafner, Pieter Abbeel, and Ken Goldberg.
|
| 2429 |
+
Daydreamer: World models for physical robot learning.
|
| 2430 |
+
In
|
| 2431 |
+
CoRL
|
| 2432 |
+
. PMLR, 2023.
|
| 2433 |
+
Xie et al. (2023)
|
| 2434 |
+
Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao, Min-Yen Kan, Junxian He, and Qizhe Xie.
|
| 2435 |
+
Decomposition enhances reasoning via self-evaluation guided decoding.
|
| 2436 |
+
arXiv:2305.00633
|
| 2437 |
+
, 2023.
|
| 2438 |
+
Yang et al. (2018)
|
| 2439 |
+
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning.
|
| 2440 |
+
Hotpotqa: A dataset for diverse, explainable multi-hop question answering.
|
| 2441 |
+
arXiv:1809.09600
|
| 2442 |
+
, 2018.
|
| 2443 |
+
Yao et al. (2022)
|
| 2444 |
+
Shunyu Yao, Howard Chen, John Yang, and Karthik R Narasimhan.
|
| 2445 |
+
Webshop: Towards scalable real-world web interaction with grounded language agents.
|
| 2446 |
+
In
|
| 2447 |
+
NeurIPS
|
| 2448 |
+
, 2022.
|
| 2449 |
+
Yao et al. (2023a)
|
| 2450 |
+
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan.
|
| 2451 |
+
Tree of thoughts: Deliberate problem solving with large language models.
|
| 2452 |
+
arXiv:2305.10601
|
| 2453 |
+
, 2023a.
|
| 2454 |
+
Yao et al. (2023b)
|
| 2455 |
+
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao.
|
| 2456 |
+
ReAct: Synergizing reasoning and acting in language models.
|
| 2457 |
+
In
|
| 2458 |
+
ICLR
|
| 2459 |
+
, 2023b.
|
| 2460 |
+
Yao et al. (2023c)
|
| 2461 |
+
Weiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu, Yihao Feng, Le Xue, Rithesh Murthy, Zeyuan Chen, Jianguo Zhang, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese.
|
| 2462 |
+
Retroformer: Retrospective large language agents with policy gradient optimization.
|
| 2463 |
+
arXiv preprint arXiv:2308.02151
|
| 2464 |
+
, 2023c.
|
| 2465 |
+
Ye et al. (2021)
|
| 2466 |
+
Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao.
|
| 2467 |
+
Mastering atari games with limited data.
|
| 2468 |
+
In
|
| 2469 |
+
NeurIPS
|
| 2470 |
+
, 2021.
|
| 2471 |
+
Zhou et al. (2022)
|
| 2472 |
+
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi.
|
| 2473 |
+
Least-to-most prompting enables complex reasoning in large language models.
|
| 2474 |
+
arXiv:2205.10625
|
| 2475 |
+
, 2022.
|
| 2476 |
+
Zhu et al. (2023)
|
| 2477 |
+
Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, Yu Qiao, Zhaoxiang Zhang, and Jifeng Dai.
|
| 2478 |
+
Ghost in the minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory.
|
| 2479 |
+
arXiv:2305.17144
|
| 2480 |
+
, 2023.
|
| 2481 |
+
7
|
| 2482 |
+
Appendix
|
| 2483 |
+
The appendix is organized as follows. First in Sec.
|
| 2484 |
+
A
|
| 2485 |
+
, we show the pseudocode of our proposed algorithm, LATS; then in Sec.
|
| 2486 |
+
B
|
| 2487 |
+
, we provide further discussion of our method and its limitations, future direction and broader impact; then in Sec.
|
| 2488 |
+
C
|
| 2489 |
+
we provide additional experimental results; then in Sec.
|
| 2490 |
+
D
|
| 2491 |
+
, we specify the environment details in our experiments; finally, we list our prompts used for the three environments in Sec.
|
| 2492 |
+
E
|
| 2493 |
+
(HotPotQA), Sec.
|
| 2494 |
+
F
|
| 2495 |
+
(Programming) and Sec.
|
| 2496 |
+
G
|
| 2497 |
+
(Webshop) respectively.
|
| 2498 |
+
Appendix A
|
| 2499 |
+
LATS Pseudocode
|
| 2500 |
+
Alg.
|
| 2501 |
+
1
|
| 2502 |
+
shows the pseudocode of our algorithm LATS. Nodes are stored explicitly in the memory. Unless otherwise specified, in all experiments we use
|
| 2503 |
+
n
|
| 2504 |
+
=
|
| 2505 |
+
5
|
| 2506 |
+
𝑛
|
| 2507 |
+
5
|
| 2508 |
+
n=5
|
| 2509 |
+
and
|
| 2510 |
+
w
|
| 2511 |
+
=
|
| 2512 |
+
1
|
| 2513 |
+
𝑤
|
| 2514 |
+
1
|
| 2515 |
+
w=1
|
| 2516 |
+
.
|
| 2517 |
+
Algorithm 1
|
| 2518 |
+
LATS
|
| 2519 |
+
|
| 2520 |
+
(
|
| 2521 |
+
S
|
| 2522 |
+
0
|
| 2523 |
+
,
|
| 2524 |
+
p
|
| 2525 |
+
θ
|
| 2526 |
+
,
|
| 2527 |
+
p
|
| 2528 |
+
V
|
| 2529 |
+
,
|
| 2530 |
+
p
|
| 2531 |
+
ref
|
| 2532 |
+
,
|
| 2533 |
+
d
|
| 2534 |
+
,
|
| 2535 |
+
k
|
| 2536 |
+
,
|
| 2537 |
+
n
|
| 2538 |
+
,
|
| 2539 |
+
w
|
| 2540 |
+
)
|
| 2541 |
+
LATS
|
| 2542 |
+
subscript
|
| 2543 |
+
𝑆
|
| 2544 |
+
0
|
| 2545 |
+
subscript
|
| 2546 |
+
𝑝
|
| 2547 |
+
𝜃
|
| 2548 |
+
subscript
|
| 2549 |
+
𝑝
|
| 2550 |
+
𝑉
|
| 2551 |
+
subscript
|
| 2552 |
+
𝑝
|
| 2553 |
+
ref
|
| 2554 |
+
𝑑
|
| 2555 |
+
𝑘
|
| 2556 |
+
𝑛
|
| 2557 |
+
𝑤
|
| 2558 |
+
\operatorname{LATS}(S_{0},p_{\theta},{p_{V}},p_{\text{ref}},d,k,n,w)
|
| 2559 |
+
Initial state
|
| 2560 |
+
s
|
| 2561 |
+
1
|
| 2562 |
+
subscript
|
| 2563 |
+
𝑠
|
| 2564 |
+
1
|
| 2565 |
+
s_{1}
|
| 2566 |
+
, action generator
|
| 2567 |
+
p
|
| 2568 |
+
θ
|
| 2569 |
+
subscript
|
| 2570 |
+
𝑝
|
| 2571 |
+
𝜃
|
| 2572 |
+
p_{\theta}
|
| 2573 |
+
, value function
|
| 2574 |
+
p
|
| 2575 |
+
V
|
| 2576 |
+
subscript
|
| 2577 |
+
𝑝
|
| 2578 |
+
𝑉
|
| 2579 |
+
p_{V}
|
| 2580 |
+
, reflection generator
|
| 2581 |
+
p
|
| 2582 |
+
ref
|
| 2583 |
+
subscript
|
| 2584 |
+
𝑝
|
| 2585 |
+
ref
|
| 2586 |
+
p_{\text{ref}}
|
| 2587 |
+
, number of generated actions
|
| 2588 |
+
n
|
| 2589 |
+
𝑛
|
| 2590 |
+
n
|
| 2591 |
+
, depth limit
|
| 2592 |
+
L
|
| 2593 |
+
𝐿
|
| 2594 |
+
L
|
| 2595 |
+
, number of roll-outs
|
| 2596 |
+
K
|
| 2597 |
+
𝐾
|
| 2598 |
+
K
|
| 2599 |
+
, context
|
| 2600 |
+
c
|
| 2601 |
+
𝑐
|
| 2602 |
+
c
|
| 2603 |
+
, and exploration weight
|
| 2604 |
+
w
|
| 2605 |
+
𝑤
|
| 2606 |
+
w
|
| 2607 |
+
Initialize action space
|
| 2608 |
+
A
|
| 2609 |
+
𝐴
|
| 2610 |
+
A
|
| 2611 |
+
, observation space
|
| 2612 |
+
O
|
| 2613 |
+
𝑂
|
| 2614 |
+
O
|
| 2615 |
+
Initialize the state-action value function
|
| 2616 |
+
p
|
| 2617 |
+
V
|
| 2618 |
+
:
|
| 2619 |
+
S
|
| 2620 |
+
×
|
| 2621 |
+
A
|
| 2622 |
+
↦
|
| 2623 |
+
ℝ
|
| 2624 |
+
:
|
| 2625 |
+
subscript
|
| 2626 |
+
𝑝
|
| 2627 |
+
𝑉
|
| 2628 |
+
maps-to
|
| 2629 |
+
𝑆
|
| 2630 |
+
𝐴
|
| 2631 |
+
ℝ
|
| 2632 |
+
{p_{V}}:S\times A\mapsto\mathbb{R}
|
| 2633 |
+
and visit counter
|
| 2634 |
+
N
|
| 2635 |
+
:
|
| 2636 |
+
S
|
| 2637 |
+
↦
|
| 2638 |
+
ℕ
|
| 2639 |
+
:
|
| 2640 |
+
𝑁
|
| 2641 |
+
maps-to
|
| 2642 |
+
𝑆
|
| 2643 |
+
ℕ
|
| 2644 |
+
{N}:S\mapsto\mathbb{N}
|
| 2645 |
+
to zero
|
| 2646 |
+
for
|
| 2647 |
+
k
|
| 2648 |
+
←
|
| 2649 |
+
0
|
| 2650 |
+
,
|
| 2651 |
+
…
|
| 2652 |
+
,
|
| 2653 |
+
K
|
| 2654 |
+
−
|
| 2655 |
+
1
|
| 2656 |
+
←
|
| 2657 |
+
𝑘
|
| 2658 |
+
0
|
| 2659 |
+
…
|
| 2660 |
+
𝐾
|
| 2661 |
+
1
|
| 2662 |
+
k\leftarrow 0,\dots,K-1
|
| 2663 |
+
do
|
| 2664 |
+
for
|
| 2665 |
+
t
|
| 2666 |
+
←
|
| 2667 |
+
0
|
| 2668 |
+
,
|
| 2669 |
+
…
|
| 2670 |
+
,
|
| 2671 |
+
L
|
| 2672 |
+
−
|
| 2673 |
+
1
|
| 2674 |
+
←
|
| 2675 |
+
𝑡
|
| 2676 |
+
0
|
| 2677 |
+
…
|
| 2678 |
+
𝐿
|
| 2679 |
+
1
|
| 2680 |
+
t\leftarrow 0,\dots,L-1
|
| 2681 |
+
do
|
| 2682 |
+
if
|
| 2683 |
+
s
|
| 2684 |
+
t
|
| 2685 |
+
subscript
|
| 2686 |
+
𝑠
|
| 2687 |
+
𝑡
|
| 2688 |
+
s_{t}
|
| 2689 |
+
not terminal
|
| 2690 |
+
then
|
| 2691 |
+
▷
|
| 2692 |
+
▷
|
| 2693 |
+
\triangleright
|
| 2694 |
+
Expansion & Simulation
|
| 2695 |
+
for
|
| 2696 |
+
i
|
| 2697 |
+
←
|
| 2698 |
+
1
|
| 2699 |
+
,
|
| 2700 |
+
…
|
| 2701 |
+
,
|
| 2702 |
+
n
|
| 2703 |
+
←
|
| 2704 |
+
𝑖
|
| 2705 |
+
1
|
| 2706 |
+
…
|
| 2707 |
+
𝑛
|
| 2708 |
+
i\leftarrow 1,\dots,n
|
| 2709 |
+
do
|
| 2710 |
+
Sample
|
| 2711 |
+
a
|
| 2712 |
+
t
|
| 2713 |
+
(
|
| 2714 |
+
i
|
| 2715 |
+
)
|
| 2716 |
+
∼
|
| 2717 |
+
p
|
| 2718 |
+
θ
|
| 2719 |
+
|
| 2720 |
+
(
|
| 2721 |
+
a
|
| 2722 |
+
∣
|
| 2723 |
+
s
|
| 2724 |
+
t
|
| 2725 |
+
)
|
| 2726 |
+
similar-to
|
| 2727 |
+
superscript
|
| 2728 |
+
subscript
|
| 2729 |
+
𝑎
|
| 2730 |
+
𝑡
|
| 2731 |
+
𝑖
|
| 2732 |
+
subscript
|
| 2733 |
+
𝑝
|
| 2734 |
+
𝜃
|
| 2735 |
+
conditional
|
| 2736 |
+
𝑎
|
| 2737 |
+
subscript
|
| 2738 |
+
𝑠
|
| 2739 |
+
𝑡
|
| 2740 |
+
a_{t}^{(i)}\sim p_{\theta}(a\mid s_{t})
|
| 2741 |
+
Get
|
| 2742 |
+
o
|
| 2743 |
+
t
|
| 2744 |
+
(
|
| 2745 |
+
i
|
| 2746 |
+
)
|
| 2747 |
+
superscript
|
| 2748 |
+
subscript
|
| 2749 |
+
𝑜
|
| 2750 |
+
𝑡
|
| 2751 |
+
𝑖
|
| 2752 |
+
o_{t}^{(i)}
|
| 2753 |
+
from environment,
|
| 2754 |
+
s
|
| 2755 |
+
t
|
| 2756 |
+
+
|
| 2757 |
+
1
|
| 2758 |
+
(
|
| 2759 |
+
i
|
| 2760 |
+
)
|
| 2761 |
+
←
|
| 2762 |
+
(
|
| 2763 |
+
c
|
| 2764 |
+
t
|
| 2765 |
+
(
|
| 2766 |
+
i
|
| 2767 |
+
)
|
| 2768 |
+
,
|
| 2769 |
+
o
|
| 2770 |
+
t
|
| 2771 |
+
(
|
| 2772 |
+
i
|
| 2773 |
+
)
|
| 2774 |
+
,
|
| 2775 |
+
a
|
| 2776 |
+
t
|
| 2777 |
+
(
|
| 2778 |
+
i
|
| 2779 |
+
)
|
| 2780 |
+
)
|
| 2781 |
+
←
|
| 2782 |
+
superscript
|
| 2783 |
+
subscript
|
| 2784 |
+
𝑠
|
| 2785 |
+
𝑡
|
| 2786 |
+
1
|
| 2787 |
+
𝑖
|
| 2788 |
+
superscript
|
| 2789 |
+
subscript
|
| 2790 |
+
𝑐
|
| 2791 |
+
𝑡
|
| 2792 |
+
𝑖
|
| 2793 |
+
superscript
|
| 2794 |
+
subscript
|
| 2795 |
+
𝑜
|
| 2796 |
+
𝑡
|
| 2797 |
+
𝑖
|
| 2798 |
+
superscript
|
| 2799 |
+
subscript
|
| 2800 |
+
𝑎
|
| 2801 |
+
𝑡
|
| 2802 |
+
𝑖
|
| 2803 |
+
s_{t+1}^{(i)}\leftarrow(c_{t}^{(i)},o_{t}^{(i)},a_{t}^{(i)})
|
| 2804 |
+
,
|
| 2805 |
+
c
|
| 2806 |
+
t
|
| 2807 |
+
+
|
| 2808 |
+
1
|
| 2809 |
+
(
|
| 2810 |
+
i
|
| 2811 |
+
)
|
| 2812 |
+
←
|
| 2813 |
+
(
|
| 2814 |
+
o
|
| 2815 |
+
t
|
| 2816 |
+
(
|
| 2817 |
+
i
|
| 2818 |
+
)
|
| 2819 |
+
,
|
| 2820 |
+
a
|
| 2821 |
+
t
|
| 2822 |
+
(
|
| 2823 |
+
i
|
| 2824 |
+
)
|
| 2825 |
+
)
|
| 2826 |
+
←
|
| 2827 |
+
superscript
|
| 2828 |
+
subscript
|
| 2829 |
+
𝑐
|
| 2830 |
+
𝑡
|
| 2831 |
+
1
|
| 2832 |
+
𝑖
|
| 2833 |
+
superscript
|
| 2834 |
+
subscript
|
| 2835 |
+
𝑜
|
| 2836 |
+
𝑡
|
| 2837 |
+
𝑖
|
| 2838 |
+
superscript
|
| 2839 |
+
subscript
|
| 2840 |
+
𝑎
|
| 2841 |
+
𝑡
|
| 2842 |
+
𝑖
|
| 2843 |
+
c_{t+1}^{(i)}\leftarrow(o_{t}^{(i)},a_{t}^{(i)})
|
| 2844 |
+
Evaluate
|
| 2845 |
+
V
|
| 2846 |
+
t
|
| 2847 |
+
(
|
| 2848 |
+
i
|
| 2849 |
+
)
|
| 2850 |
+
∼
|
| 2851 |
+
p
|
| 2852 |
+
V
|
| 2853 |
+
|
| 2854 |
+
(
|
| 2855 |
+
s
|
| 2856 |
+
t
|
| 2857 |
+
(
|
| 2858 |
+
i
|
| 2859 |
+
)
|
| 2860 |
+
)
|
| 2861 |
+
similar-to
|
| 2862 |
+
superscript
|
| 2863 |
+
subscript
|
| 2864 |
+
𝑉
|
| 2865 |
+
𝑡
|
| 2866 |
+
𝑖
|
| 2867 |
+
subscript
|
| 2868 |
+
𝑝
|
| 2869 |
+
𝑉
|
| 2870 |
+
superscript
|
| 2871 |
+
subscript
|
| 2872 |
+
𝑠
|
| 2873 |
+
𝑡
|
| 2874 |
+
𝑖
|
| 2875 |
+
{V}_{t}^{(i)}\sim{p_{V}}(s_{t}^{(i)})
|
| 2876 |
+
▷
|
| 2877 |
+
▷
|
| 2878 |
+
\triangleright
|
| 2879 |
+
Evaluation
|
| 2880 |
+
V
|
| 2881 |
+
|
| 2882 |
+
(
|
| 2883 |
+
s
|
| 2884 |
+
t
|
| 2885 |
+
)
|
| 2886 |
+
←
|
| 2887 |
+
V
|
| 2888 |
+
t
|
| 2889 |
+
(
|
| 2890 |
+
i
|
| 2891 |
+
)
|
| 2892 |
+
←
|
| 2893 |
+
𝑉
|
| 2894 |
+
subscript
|
| 2895 |
+
𝑠
|
| 2896 |
+
𝑡
|
| 2897 |
+
superscript
|
| 2898 |
+
subscript
|
| 2899 |
+
𝑉
|
| 2900 |
+
𝑡
|
| 2901 |
+
𝑖
|
| 2902 |
+
{V}(s_{t})\leftarrow{V}_{t}^{(i)}
|
| 2903 |
+
Add
|
| 2904 |
+
s
|
| 2905 |
+
t
|
| 2906 |
+
(
|
| 2907 |
+
i
|
| 2908 |
+
)
|
| 2909 |
+
superscript
|
| 2910 |
+
subscript
|
| 2911 |
+
𝑠
|
| 2912 |
+
𝑡
|
| 2913 |
+
𝑖
|
| 2914 |
+
s_{t}^{(i)}
|
| 2915 |
+
to children
|
| 2916 |
+
end
|
| 2917 |
+
for
|
| 2918 |
+
end
|
| 2919 |
+
if
|
| 2920 |
+
if
|
| 2921 |
+
s
|
| 2922 |
+
t
|
| 2923 |
+
subscript
|
| 2924 |
+
𝑠
|
| 2925 |
+
𝑡
|
| 2926 |
+
s_{t}
|
| 2927 |
+
is terminal
|
| 2928 |
+
then
|
| 2929 |
+
▷
|
| 2930 |
+
▷
|
| 2931 |
+
\triangleright
|
| 2932 |
+
Reflection
|
| 2933 |
+
Get
|
| 2934 |
+
r
|
| 2935 |
+
𝑟
|
| 2936 |
+
r
|
| 2937 |
+
from environment
|
| 2938 |
+
if
|
| 2939 |
+
r
|
| 2940 |
+
𝑟
|
| 2941 |
+
r
|
| 2942 |
+
not success
|
| 2943 |
+
then
|
| 2944 |
+
reflection
|
| 2945 |
+
←
|
| 2946 |
+
p
|
| 2947 |
+
ref
|
| 2948 |
+
|
| 2949 |
+
(
|
| 2950 |
+
c
|
| 2951 |
+
t
|
| 2952 |
+
)
|
| 2953 |
+
←
|
| 2954 |
+
reflection
|
| 2955 |
+
subscript
|
| 2956 |
+
𝑝
|
| 2957 |
+
ref
|
| 2958 |
+
subscript
|
| 2959 |
+
𝑐
|
| 2960 |
+
𝑡
|
| 2961 |
+
\text{reflection}\leftarrow p_{\text{ref}}(c_{t})
|
| 2962 |
+
c
|
| 2963 |
+
←
|
| 2964 |
+
reflection
|
| 2965 |
+
←
|
| 2966 |
+
𝑐
|
| 2967 |
+
reflection
|
| 2968 |
+
c\leftarrow\text{reflection}
|
| 2969 |
+
end
|
| 2970 |
+
if
|
| 2971 |
+
end
|
| 2972 |
+
if
|
| 2973 |
+
a
|
| 2974 |
+
t
|
| 2975 |
+
←
|
| 2976 |
+
arg
|
| 2977 |
+
|
| 2978 |
+
max
|
| 2979 |
+
a
|
| 2980 |
+
∈
|
| 2981 |
+
e
|
| 2982 |
+
|
| 2983 |
+
(
|
| 2984 |
+
s
|
| 2985 |
+
t
|
| 2986 |
+
)
|
| 2987 |
+
|
| 2988 |
+
[
|
| 2989 |
+
V
|
| 2990 |
+
|
| 2991 |
+
(
|
| 2992 |
+
s
|
| 2993 |
+
t
|
| 2994 |
+
)
|
| 2995 |
+
+
|
| 2996 |
+
w
|
| 2997 |
+
|
| 2998 |
+
ln
|
| 2999 |
+
|
| 3000 |
+
N
|
| 3001 |
+
|
| 3002 |
+
(
|
| 3003 |
+
s
|
| 3004 |
+
t
|
| 3005 |
+
−
|
| 3006 |
+
1
|
| 3007 |
+
)
|
| 3008 |
+
N
|
| 3009 |
+
|
| 3010 |
+
(
|
| 3011 |
+
s
|
| 3012 |
+
t
|
| 3013 |
+
)
|
| 3014 |
+
]
|
| 3015 |
+
←
|
| 3016 |
+
subscript
|
| 3017 |
+
𝑎
|
| 3018 |
+
𝑡
|
| 3019 |
+
subscript
|
| 3020 |
+
𝑎
|
| 3021 |
+
𝑒
|
| 3022 |
+
subscript
|
| 3023 |
+
𝑠
|
| 3024 |
+
𝑡
|
| 3025 |
+
𝑉
|
| 3026 |
+
subscript
|
| 3027 |
+
𝑠
|
| 3028 |
+
𝑡
|
| 3029 |
+
𝑤
|
| 3030 |
+
𝑁
|
| 3031 |
+
subscript
|
| 3032 |
+
𝑠
|
| 3033 |
+
𝑡
|
| 3034 |
+
1
|
| 3035 |
+
𝑁
|
| 3036 |
+
subscript
|
| 3037 |
+
𝑠
|
| 3038 |
+
𝑡
|
| 3039 |
+
a_{t}\leftarrow\arg\max_{a\in e(s_{t})}\left[{V(s_{t})}+w\sqrt{\frac{\ln{N}(s_{t-1})}{{N}(s_{t})}}\right]
|
| 3040 |
+
▷
|
| 3041 |
+
▷
|
| 3042 |
+
\triangleright
|
| 3043 |
+
Selection
|
| 3044 |
+
N
|
| 3045 |
+
|
| 3046 |
+
(
|
| 3047 |
+
s
|
| 3048 |
+
t
|
| 3049 |
+
+
|
| 3050 |
+
1
|
| 3051 |
+
)
|
| 3052 |
+
←
|
| 3053 |
+
N
|
| 3054 |
+
|
| 3055 |
+
(
|
| 3056 |
+
s
|
| 3057 |
+
t
|
| 3058 |
+
+
|
| 3059 |
+
1
|
| 3060 |
+
)
|
| 3061 |
+
+
|
| 3062 |
+
1
|
| 3063 |
+
←
|
| 3064 |
+
𝑁
|
| 3065 |
+
subscript
|
| 3066 |
+
𝑠
|
| 3067 |
+
𝑡
|
| 3068 |
+
1
|
| 3069 |
+
𝑁
|
| 3070 |
+
subscript
|
| 3071 |
+
𝑠
|
| 3072 |
+
𝑡
|
| 3073 |
+
1
|
| 3074 |
+
1
|
| 3075 |
+
{N}(s_{t+1})\leftarrow{N}(s_{t+1})+1
|
| 3076 |
+
if
|
| 3077 |
+
a
|
| 3078 |
+
t
|
| 3079 |
+
subscript
|
| 3080 |
+
𝑎
|
| 3081 |
+
𝑡
|
| 3082 |
+
a_{t}
|
| 3083 |
+
is an output action
|
| 3084 |
+
then
|
| 3085 |
+
break
|
| 3086 |
+
end
|
| 3087 |
+
for
|
| 3088 |
+
T
|
| 3089 |
+
←
|
| 3090 |
+
←
|
| 3091 |
+
𝑇
|
| 3092 |
+
absent
|
| 3093 |
+
T\leftarrow
|
| 3094 |
+
the actual number of steps
|
| 3095 |
+
for
|
| 3096 |
+
t
|
| 3097 |
+
←
|
| 3098 |
+
T
|
| 3099 |
+
−
|
| 3100 |
+
1
|
| 3101 |
+
,
|
| 3102 |
+
…
|
| 3103 |
+
,
|
| 3104 |
+
0
|
| 3105 |
+
←
|
| 3106 |
+
𝑡
|
| 3107 |
+
𝑇
|
| 3108 |
+
1
|
| 3109 |
+
…
|
| 3110 |
+
0
|
| 3111 |
+
t\leftarrow T-1,\dots,0
|
| 3112 |
+
do
|
| 3113 |
+
▷
|
| 3114 |
+
▷
|
| 3115 |
+
\triangleright
|
| 3116 |
+
Backpropagation
|
| 3117 |
+
V
|
| 3118 |
+
|
| 3119 |
+
(
|
| 3120 |
+
s
|
| 3121 |
+
t
|
| 3122 |
+
)
|
| 3123 |
+
←
|
| 3124 |
+
V
|
| 3125 |
+
|
| 3126 |
+
(
|
| 3127 |
+
s
|
| 3128 |
+
t
|
| 3129 |
+
)
|
| 3130 |
+
|
| 3131 |
+
(
|
| 3132 |
+
N
|
| 3133 |
+
|
| 3134 |
+
(
|
| 3135 |
+
s
|
| 3136 |
+
t
|
| 3137 |
+
)
|
| 3138 |
+
−
|
| 3139 |
+
1
|
| 3140 |
+
)
|
| 3141 |
+
+
|
| 3142 |
+
r
|
| 3143 |
+
N
|
| 3144 |
+
|
| 3145 |
+
(
|
| 3146 |
+
s
|
| 3147 |
+
t
|
| 3148 |
+
)
|
| 3149 |
+
←
|
| 3150 |
+
𝑉
|
| 3151 |
+
subscript
|
| 3152 |
+
𝑠
|
| 3153 |
+
𝑡
|
| 3154 |
+
𝑉
|
| 3155 |
+
subscript
|
| 3156 |
+
𝑠
|
| 3157 |
+
𝑡
|
| 3158 |
+
𝑁
|
| 3159 |
+
subscript
|
| 3160 |
+
𝑠
|
| 3161 |
+
𝑡
|
| 3162 |
+
1
|
| 3163 |
+
𝑟
|
| 3164 |
+
𝑁
|
| 3165 |
+
subscript
|
| 3166 |
+
𝑠
|
| 3167 |
+
𝑡
|
| 3168 |
+
V(s_{t})\leftarrow\frac{V(s_{t})(N(s_{t})-1)+r}{N(s_{t})}
|
| 3169 |
+
end
|
| 3170 |
+
for
|
| 3171 |
+
end
|
| 3172 |
+
for
|
| 3173 |
+
Appendix B
|
| 3174 |
+
Discussion
|
| 3175 |
+
Limitations.
|
| 3176 |
+
Although LATS can improve reasoning and decision-making, this arrives at a higher computational cost relative to simpler prompting methods like ReAct or Reflexion. The search process takes more time than standard prompting or simpler techniques, and requires greater inference costs. While such an issue is mitigated by the fact that the number of nodes
|
| 3177 |
+
n
|
| 3178 |
+
𝑛
|
| 3179 |
+
n
|
| 3180 |
+
expanded at every step provides a natural trade-off between performance and efficiency (setting
|
| 3181 |
+
n
|
| 3182 |
+
=
|
| 3183 |
+
1
|
| 3184 |
+
𝑛
|
| 3185 |
+
1
|
| 3186 |
+
n=1
|
| 3187 |
+
makes the method as effecient as ReAct with multiple trials or CoT-SC), in practice we recommend using LATS for difficult tasks like programming or for situations where performance is prioritized over efficiency. We hope that continued advancements in LLMs will reduce costs and increase the practicality of LATS.
|
| 3188 |
+
Additionally, the benchmarks we use in this paper are relatively simple and focused on decision-making, compared to the complexity of real-world interactive environments. In addition, some environments might not easily support rollbacks to previous states. However, the design of LATS is flexible and can be adjusted to various resource constraints. Using planning-based prompting methods like LATS in environments like Minecraft
|
| 3189 |
+
(Fan et al.,
|
| 3190 |
+
2022
|
| 3191 |
+
)
|
| 3192 |
+
and more reasoning benchmarks would be interesting avenues for future work.
|
| 3193 |
+
Broader impact.
|
| 3194 |
+
LATS is a framework that enhances LLM performance through interactions with an environment. This improvement in autonomous decision-making may facilitate harmful uses of LLMs. Alternatively, LATS enhances interpretability and the potential for greater alignment, as it generates understandable, high-level linguistic reasoning and actions through several rounds of decision-making and reflection, rather than relying on implicit, low-level token values.
|
| 3195 |
+
Appendix C
|
| 3196 |
+
Ablations
|
| 3197 |
+
Prompt Method
|
| 3198 |
+
HotpotQA (EM)
|
| 3199 |
+
LATS (w=0.5)
|
| 3200 |
+
0.55
|
| 3201 |
+
LATS (w=2.0)
|
| 3202 |
+
0.61
|
| 3203 |
+
LATS (d=4)
|
| 3204 |
+
0.58
|
| 3205 |
+
LATS (CoT)
|
| 3206 |
+
0.60
|
| 3207 |
+
LATS (No LM Heuristic)
|
| 3208 |
+
0.37
|
| 3209 |
+
LATS
|
| 3210 |
+
0.61
|
| 3211 |
+
Table 6:
|
| 3212 |
+
Ablation results on LATS and baseline variants in HotPotQA measured by Exact Match (EM). We test different depth
|
| 3213 |
+
d
|
| 3214 |
+
𝑑
|
| 3215 |
+
d
|
| 3216 |
+
, exploration factor
|
| 3217 |
+
w
|
| 3218 |
+
𝑤
|
| 3219 |
+
w
|
| 3220 |
+
, and versions of LATS using CoT and without the LM value function. We sample
|
| 3221 |
+
n
|
| 3222 |
+
=
|
| 3223 |
+
5
|
| 3224 |
+
𝑛
|
| 3225 |
+
5
|
| 3226 |
+
n=5
|
| 3227 |
+
and
|
| 3228 |
+
k
|
| 3229 |
+
=
|
| 3230 |
+
50
|
| 3231 |
+
𝑘
|
| 3232 |
+
50
|
| 3233 |
+
k=50
|
| 3234 |
+
trajectories.
|
| 3235 |
+
Figure 4:
|
| 3236 |
+
Performance over successive iterations on HumanEval with GPT-3.5.
|
| 3237 |
+
In this section, we ablate various designs of LATS. Experiments are conducted on HotPotQA with a maximum of
|
| 3238 |
+
k
|
| 3239 |
+
=
|
| 3240 |
+
50
|
| 3241 |
+
𝑘
|
| 3242 |
+
50
|
| 3243 |
+
k=50
|
| 3244 |
+
trajectories and sampling size of
|
| 3245 |
+
n
|
| 3246 |
+
=
|
| 3247 |
+
5
|
| 3248 |
+
𝑛
|
| 3249 |
+
5
|
| 3250 |
+
n=5
|
| 3251 |
+
and HumanEval with a maximum of
|
| 3252 |
+
k
|
| 3253 |
+
=
|
| 3254 |
+
8
|
| 3255 |
+
𝑘
|
| 3256 |
+
8
|
| 3257 |
+
k=8
|
| 3258 |
+
trajectories and sampling size of
|
| 3259 |
+
n
|
| 3260 |
+
=
|
| 3261 |
+
5
|
| 3262 |
+
𝑛
|
| 3263 |
+
5
|
| 3264 |
+
n=5
|
| 3265 |
+
. The result for HotPotQA is shown in Tab.
|
| 3266 |
+
5
|
| 3267 |
+
and HumanEval in Fig.
|
| 3268 |
+
4
|
| 3269 |
+
.
|
| 3270 |
+
Exploration weight.
|
| 3271 |
+
We find that there is lower performance on HotPotQA when the exploration weight
|
| 3272 |
+
w
|
| 3273 |
+
𝑤
|
| 3274 |
+
w
|
| 3275 |
+
in the selection formula is decreased to
|
| 3276 |
+
0.5
|
| 3277 |
+
0.5
|
| 3278 |
+
0.5
|
| 3279 |
+
, suggesting that this reduces the effectiveness of the search. Increasing
|
| 3280 |
+
w
|
| 3281 |
+
𝑤
|
| 3282 |
+
w
|
| 3283 |
+
to
|
| 3284 |
+
2.0
|
| 3285 |
+
2.0
|
| 3286 |
+
2.0
|
| 3287 |
+
does not lead to a performance improvement, but we tend to observe faster convergence. The optimal setting depends on the particular environment and complexity of the state space.
|
| 3288 |
+
Depth.
|
| 3289 |
+
In our main experiments we use a maximum depth of
|
| 3290 |
+
d
|
| 3291 |
+
=
|
| 3292 |
+
7
|
| 3293 |
+
𝑑
|
| 3294 |
+
7
|
| 3295 |
+
d=7
|
| 3296 |
+
on HotPotQA for all methods, following previous work
|
| 3297 |
+
(Yao et al.,
|
| 3298 |
+
2023b
|
| 3299 |
+
)
|
| 3300 |
+
. We ablate the effect on LATS after reducing it to
|
| 3301 |
+
d
|
| 3302 |
+
=
|
| 3303 |
+
4
|
| 3304 |
+
𝑑
|
| 3305 |
+
4
|
| 3306 |
+
d=4
|
| 3307 |
+
. This results in only a slight drop in performance. We find that most questions can be answered within four steps, and using a greater number of steps tends to force the agent into local minima and rarely improves success.
|
| 3308 |
+
LM value function.
|
| 3309 |
+
The LM value function scores states based on expected future reward. Without this heuristic, the only signal to guide search would be from environment rewards for completed trajectories, which are scarce and often binary. When we remove the evaluation operation, we observe a dramatic
|
| 3310 |
+
0.24
|
| 3311 |
+
0.24
|
| 3312 |
+
0.24
|
| 3313 |
+
drop in performance.
|
| 3314 |
+
Performance over time.
|
| 3315 |
+
To see the effects of increasing the number of trajectories sampled, we change
|
| 3316 |
+
k
|
| 3317 |
+
𝑘
|
| 3318 |
+
k
|
| 3319 |
+
to different values. We conduct this experiment on HumanEval, which has a more noticeable difference due to sampling less trajectories. The results are shown in Fig.
|
| 3320 |
+
4
|
| 3321 |
+
, in which LATS scales better with more iterations than Reflexion.
|
| 3322 |
+
Sample complexity and Token cost.
|
| 3323 |
+
One possible concern of LATS is that the tree-structured search might consume much more tokens than existing methods. To further study the computational cost of LATS compared to prior methods, we examine the sample complexity (i.e. asymptotic token cost) of all methods considered in this paper, and count the average number of nodes expanded by our method and other tree-structured methods (ToT and RAP) upon successful search on HotPotQA. We present the results in Tab.
|
| 3324 |
+
7
|
| 3325 |
+
; the result shows that our method has the same sample complexity as other tree-based search methods, and has less average number of nodes expanded upon success, which indicates less token cost. The token cost gap will be even larger when taking failed trajectories into account, since our method has higher success rate and reaches computational budget limit less often.
|
| 3326 |
+
Method
|
| 3327 |
+
Performance (
|
| 3328 |
+
↑
|
| 3329 |
+
↑
|
| 3330 |
+
\uparrow
|
| 3331 |
+
)
|
| 3332 |
+
Sample complexity (
|
| 3333 |
+
↓
|
| 3334 |
+
↓
|
| 3335 |
+
\downarrow
|
| 3336 |
+
)
|
| 3337 |
+
Avg. #nodes upon success (
|
| 3338 |
+
↓
|
| 3339 |
+
↓
|
| 3340 |
+
\downarrow
|
| 3341 |
+
)
|
| 3342 |
+
ReAct (Best
|
| 3343 |
+
k
|
| 3344 |
+
=
|
| 3345 |
+
250
|
| 3346 |
+
𝑘
|
| 3347 |
+
250
|
| 3348 |
+
k=250
|
| 3349 |
+
)
|
| 3350 |
+
0.42
|
| 3351 |
+
0.42
|
| 3352 |
+
0.42
|
| 3353 |
+
O
|
| 3354 |
+
|
| 3355 |
+
(
|
| 3356 |
+
k
|
| 3357 |
+
)
|
| 3358 |
+
𝑂
|
| 3359 |
+
𝑘
|
| 3360 |
+
O(k)
|
| 3361 |
+
N/A
|
| 3362 |
+
CoT-SC (
|
| 3363 |
+
n
|
| 3364 |
+
=
|
| 3365 |
+
1
|
| 3366 |
+
,
|
| 3367 |
+
k
|
| 3368 |
+
=
|
| 3369 |
+
250
|
| 3370 |
+
formulae-sequence
|
| 3371 |
+
𝑛
|
| 3372 |
+
1
|
| 3373 |
+
𝑘
|
| 3374 |
+
250
|
| 3375 |
+
n=1,k=250
|
| 3376 |
+
)
|
| 3377 |
+
0.40
|
| 3378 |
+
0.40
|
| 3379 |
+
0.40
|
| 3380 |
+
O
|
| 3381 |
+
|
| 3382 |
+
(
|
| 3383 |
+
k
|
| 3384 |
+
)
|
| 3385 |
+
𝑂
|
| 3386 |
+
𝑘
|
| 3387 |
+
O(k)
|
| 3388 |
+
N/A
|
| 3389 |
+
LATS (
|
| 3390 |
+
n
|
| 3391 |
+
=
|
| 3392 |
+
1
|
| 3393 |
+
,
|
| 3394 |
+
k
|
| 3395 |
+
=
|
| 3396 |
+
50
|
| 3397 |
+
formulae-sequence
|
| 3398 |
+
𝑛
|
| 3399 |
+
1
|
| 3400 |
+
𝑘
|
| 3401 |
+
50
|
| 3402 |
+
n=1,k=50
|
| 3403 |
+
)
|
| 3404 |
+
0.48
|
| 3405 |
+
0.48
|
| 3406 |
+
0.48
|
| 3407 |
+
O
|
| 3408 |
+
|
| 3409 |
+
(
|
| 3410 |
+
k
|
| 3411 |
+
)
|
| 3412 |
+
𝑂
|
| 3413 |
+
𝑘
|
| 3414 |
+
O(k)
|
| 3415 |
+
N/A
|
| 3416 |
+
ToT (ReAct)
|
| 3417 |
+
0.49
|
| 3418 |
+
0.49
|
| 3419 |
+
0.49
|
| 3420 |
+
O
|
| 3421 |
+
|
| 3422 |
+
(
|
| 3423 |
+
k
|
| 3424 |
+
|
| 3425 |
+
n
|
| 3426 |
+
)
|
| 3427 |
+
𝑂
|
| 3428 |
+
𝑘
|
| 3429 |
+
𝑛
|
| 3430 |
+
O(kn)
|
| 3431 |
+
84.05
|
| 3432 |
+
84.05
|
| 3433 |
+
84.05
|
| 3434 |
+
RAP (ReAct)
|
| 3435 |
+
0.54
|
| 3436 |
+
0.54
|
| 3437 |
+
0.54
|
| 3438 |
+
O
|
| 3439 |
+
|
| 3440 |
+
(
|
| 3441 |
+
k
|
| 3442 |
+
|
| 3443 |
+
n
|
| 3444 |
+
)
|
| 3445 |
+
𝑂
|
| 3446 |
+
𝑘
|
| 3447 |
+
𝑛
|
| 3448 |
+
O(kn)
|
| 3449 |
+
70.60
|
| 3450 |
+
70.60
|
| 3451 |
+
70.60
|
| 3452 |
+
LATS (
|
| 3453 |
+
n
|
| 3454 |
+
=
|
| 3455 |
+
5
|
| 3456 |
+
,
|
| 3457 |
+
k
|
| 3458 |
+
=
|
| 3459 |
+
50
|
| 3460 |
+
formulae-sequence
|
| 3461 |
+
𝑛
|
| 3462 |
+
5
|
| 3463 |
+
𝑘
|
| 3464 |
+
50
|
| 3465 |
+
n=5,k=50
|
| 3466 |
+
)
|
| 3467 |
+
0.61
|
| 3468 |
+
0.61
|
| 3469 |
+
0.61
|
| 3470 |
+
O
|
| 3471 |
+
|
| 3472 |
+
(
|
| 3473 |
+
k
|
| 3474 |
+
|
| 3475 |
+
n
|
| 3476 |
+
)
|
| 3477 |
+
𝑂
|
| 3478 |
+
𝑘
|
| 3479 |
+
𝑛
|
| 3480 |
+
O(kn)
|
| 3481 |
+
66.65
|
| 3482 |
+
66.65
|
| 3483 |
+
66.65
|
| 3484 |
+
Table 7:
|
| 3485 |
+
The performance, sample complexity of different methods and average number of nodes expanded upon success by methods with tree-based search.
|
| 3486 |
+
n
|
| 3487 |
+
𝑛
|
| 3488 |
+
n
|
| 3489 |
+
is the number of children nodes expanded at every step and
|
| 3490 |
+
k
|
| 3491 |
+
𝑘
|
| 3492 |
+
k
|
| 3493 |
+
is the number of trajectories. Our method has the same sample complexity as other methods with tree-based search and expands less nodes upon success, which indicates lower token cost.
|
| 3494 |
+
Appendix D
|
| 3495 |
+
Environment Details
|
| 3496 |
+
D.1
|
| 3497 |
+
HotPotQA
|
| 3498 |
+
Figure 5:
|
| 3499 |
+
Example trajectories on HotPotQA for ReAct (left) and LATS (right). LATS can sample more actions and avoid failure from previous mistakes by evaluating states with an LM to guide the search toward promising areas of the tree.
|
| 3500 |
+
HotPotQA
|
| 3501 |
+
(Yang et al.,
|
| 3502 |
+
2018
|
| 3503 |
+
)
|
| 3504 |
+
is a question-answering dataset that requires reasoning over multiple supporting documents to answer questions. It contains 113k Wikipedia-based question-answer pairs crafted by crowdworkers to be diverse, multi-hop, and explainable. Questions cover a range of types like entities, locations, dates, and comparison of shared properties between two entities. Crowdworkers also provide supporting facts from the documents that justify the answer. We use the HotPotQA benchmark setting with all the Wikipedia paragraphs to test retrieval. We use a randomly selected subset of 100 questions for our experiments and a maximum depth limit of 6. Fig.
|
| 3505 |
+
5
|
| 3506 |
+
illustrates how ReAct and LATS work on an example task of HotPotQA, and gives a qualitative example on how LATS outperforms ReAct on the task.
|
| 3507 |
+
Action Space.
|
| 3508 |
+
We adopt the Wikipedia web API proposed in
|
| 3509 |
+
Yao et al. (
|
| 3510 |
+
2023b
|
| 3511 |
+
)
|
| 3512 |
+
, with three types of actions to support interactive information retrieval:
|
| 3513 |
+
(1)
|
| 3514 |
+
search
|
| 3515 |
+
[
|
| 3516 |
+
entity
|
| 3517 |
+
], which returns the first 5 sentences from the corresponding
|
| 3518 |
+
entity
|
| 3519 |
+
wiki page if it exists, or else suggests top-5 similar entities from the Wikipedia search engine,
|
| 3520 |
+
(2)
|
| 3521 |
+
lookup
|
| 3522 |
+
[
|
| 3523 |
+
string
|
| 3524 |
+
], which returns the next sentence in the page containing
|
| 3525 |
+
string
|
| 3526 |
+
,
|
| 3527 |
+
(3)
|
| 3528 |
+
finish
|
| 3529 |
+
[
|
| 3530 |
+
answer
|
| 3531 |
+
], which finishes the current task with
|
| 3532 |
+
answer
|
| 3533 |
+
.
|
| 3534 |
+
These API calls and free-form thoughts form the action space for this environment.
|
| 3535 |
+
D.2
|
| 3536 |
+
Programming
|
| 3537 |
+
The HumanEval dataset
|
| 3538 |
+
(Chen et al.,
|
| 3539 |
+
2021
|
| 3540 |
+
)
|
| 3541 |
+
is a collection of 164 handwritten programming problems introduced to evaluate the functional correctness of models for synthesizing programs from natural language descriptions. Each problem includes a function signature, docstring description, reference implementation, and multiple unit tests, with an average of 7.7 tests per problem. The programming tasks assess comprehension of natural language, reasoning, algorithms, and basic mathematics, at a difficulty level comparable to simple software interview questions. Pass rates are evaluated with the pass@k metric, where k samples are generated per problem and a problem is considered solved if any sample passes all tests. We use all 164 problems for our experiments and a maximum depth limit of 8.
|
| 3542 |
+
The Mostly Basic Programming Problems (MBPP)
|
| 3543 |
+
Austin et al. (
|
| 3544 |
+
2021
|
| 3545 |
+
)
|
| 3546 |
+
benchmark contains 974 short Python functions designed to evaluate program synthesis techniques. The dataset was constructed by crowdsourcing from workers with basic Python knowledge. Each data point consists of a natural language description of a programming task, a reference solution implementation, and three test cases for functional correctness. The natural language prompts are typically short, one-sentence descriptions. Solutions cover common programming constructs including mathematical operations, list processing, string manipulation, and usage of the Python standard library. On average, solutions are 6.8 lines of code. The dataset is also supplemented with an additional set of 426 problems that were manually verified for unambiguous specifications, standard function signatures, and accurate test cases. We use a randomly selected subset of 397 problems for our experiments.
|
| 3547 |
+
D.3
|
| 3548 |
+
WebShop
|
| 3549 |
+
WebShop
|
| 3550 |
+
(Yao et al.,
|
| 3551 |
+
2022
|
| 3552 |
+
)
|
| 3553 |
+
is an interactive web-based environment designed to evaluate agents on grounded language understanding and decision-making. It simulates an e-commerce shopping task by providing agents with over 1 million real-world products scraped from Amazon, spanning 5 categories and 113 subcategories. These products contain rich linguistic information, with an average text length of 262 words and a vocabulary size of 224k. In addition, there are over 800k unique product options available for customization. The environment renders webpages in two modes: HTML mode provides pixel-level observations with interactive elements, while simple mode converts the raw HTML into a structured text observation more amenable for training agents. The action space consists of query searches and button clicks, which transition between 4 page types: search, results, item and item-detail. Instructions are crowdsourced natural language specifying product attributes and options, with a total of 12k collected. Automatic rewards are computed by comparing the product purchased by the agent against the attributes and options specified in the instruction, using both lexical matching and semantic similarity metrics.
|
| 3554 |
+
Type
|
| 3555 |
+
Argument
|
| 3556 |
+
State
|
| 3557 |
+
→
|
| 3558 |
+
→
|
| 3559 |
+
\rightarrow
|
| 3560 |
+
Next State
|
| 3561 |
+
search
|
| 3562 |
+
[
|
| 3563 |
+
Query
|
| 3564 |
+
]
|
| 3565 |
+
Search
|
| 3566 |
+
→
|
| 3567 |
+
→
|
| 3568 |
+
\rightarrow
|
| 3569 |
+
Results
|
| 3570 |
+
choose
|
| 3571 |
+
Back to search
|
| 3572 |
+
∗
|
| 3573 |
+
*
|
| 3574 |
+
→
|
| 3575 |
+
→
|
| 3576 |
+
\rightarrow
|
| 3577 |
+
Search
|
| 3578 |
+
choose
|
| 3579 |
+
Prev/Next page
|
| 3580 |
+
Results
|
| 3581 |
+
→
|
| 3582 |
+
→
|
| 3583 |
+
\rightarrow
|
| 3584 |
+
Results
|
| 3585 |
+
choose
|
| 3586 |
+
[
|
| 3587 |
+
Product title
|
| 3588 |
+
]
|
| 3589 |
+
Results
|
| 3590 |
+
→
|
| 3591 |
+
→
|
| 3592 |
+
\rightarrow
|
| 3593 |
+
Item
|
| 3594 |
+
choose
|
| 3595 |
+
[
|
| 3596 |
+
Option
|
| 3597 |
+
]
|
| 3598 |
+
Item
|
| 3599 |
+
→
|
| 3600 |
+
→
|
| 3601 |
+
\rightarrow
|
| 3602 |
+
Item
|
| 3603 |
+
choose
|
| 3604 |
+
Desc/Overview
|
| 3605 |
+
Item
|
| 3606 |
+
→
|
| 3607 |
+
→
|
| 3608 |
+
\rightarrow
|
| 3609 |
+
Item-Detail
|
| 3610 |
+
choose
|
| 3611 |
+
Previous
|
| 3612 |
+
Item-Detail
|
| 3613 |
+
→
|
| 3614 |
+
→
|
| 3615 |
+
\rightarrow
|
| 3616 |
+
Item
|
| 3617 |
+
choose
|
| 3618 |
+
Buy
|
| 3619 |
+
Item
|
| 3620 |
+
→
|
| 3621 |
+
→
|
| 3622 |
+
\rightarrow
|
| 3623 |
+
Episode End
|
| 3624 |
+
Table 8:
|
| 3625 |
+
Action space of webshop.
|
| 3626 |
+
There are two evaluation metrics used in WebShop: (1)
|
| 3627 |
+
Task Score
|
| 3628 |
+
: defined as
|
| 3629 |
+
(
|
| 3630 |
+
100
|
| 3631 |
+
×
|
| 3632 |
+
avg. reward
|
| 3633 |
+
)
|
| 3634 |
+
100
|
| 3635 |
+
avg. reward
|
| 3636 |
+
(100\times\text{avg. reward})
|
| 3637 |
+
, which captures the average reward obtained across episodes; and (2)
|
| 3638 |
+
Success Rate (SR)
|
| 3639 |
+
defined as the portion of instructions where
|
| 3640 |
+
r
|
| 3641 |
+
=
|
| 3642 |
+
1
|
| 3643 |
+
𝑟
|
| 3644 |
+
1
|
| 3645 |
+
r=1
|
| 3646 |
+
. The reward is calculated based on the number of attributes satisfied by the selected item. We use 50 environments for our experiments and a maximum depth limit of 15.
|
| 3647 |
+
Appendix E
|
| 3648 |
+
HotPotQA Prompts
|
| 3649 |
+
E.1
|
| 3650 |
+
Base Acting Prompt
|
| 3651 |
+
Solve a question answering task with interleaving Thought, Action, Observation steps. Thought can reason about the current situation, and Action can be three types:
|
| 3652 |
+
(1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
|
| 3653 |
+
(2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
|
| 3654 |
+
(3) Finish[answer], which returns the answer and finishes the task.
|
| 3655 |
+
After each observation, provide the next Thought and next Action.
|
| 3656 |
+
Here are some examples:
|
| 3657 |
+
Question: Which magazine was started first Arthur’s Magazine or First for Women?
|
| 3658 |
+
Thought 1: I need to search Arthur’s Magazine and First for Women, and find which was started first.
|
| 3659 |
+
Action 1: Search[Arthur’s Magazine]
|
| 3660 |
+
Observation 1: Arthur’s Magazine (18440̆0e20̆0800̆0931846) was an American literary periodical published in Philadelphia in the 19th century. Edited by Timothy Shay Arthur, it featured work by Edgar A. Poe, J.H. Ingraham, Sarah Josepha Hale, Thomas G. Spear, and others.[1][2] In May 1846 it was merged into Godey’s Lady’s Book.[3]
|
| 3661 |
+
Thought 2: Arthur’s Magazine was started in 1844. I need to search First for Women next.
|
| 3662 |
+
Action 2: Search[First for Women]
|
| 3663 |
+
Observation 2: First for Women is a woman’s magazine published by Bauer Media Group in the USA.[1] The magazine was started in 1989. It is based in Englewood Cliffs, New Jersey.[2] In 2011 the circulation of the magazine was 1,310,696 copies.[3]
|
| 3664 |
+
Thought 3: First for Women was started in 1989. 1844 (Arthur’s Magazine) ¡ 1989 (First for Women), so Arthur’s Magazine was started first.
|
| 3665 |
+
Action 3: Finish[Arthur’s Magazine]
|
| 3666 |
+
(examples)
|
| 3667 |
+
You have attempted to answer the following question before and failed. The following reflection(s) give a plan to avoid failing to answer the question in the same way you did previously. Use them to improve your strategy of correctly answering the given question.
|
| 3668 |
+
(trajectories)
|
| 3669 |
+
(input)
|
| 3670 |
+
E.2
|
| 3671 |
+
Base Reasoning Prompt
|
| 3672 |
+
Solve a question answering task by having a Thought, then Finish with your answer. Thought can reason about the current situation. Finish[answer] returns the answer and finishes the task. You will be given context that you should use to help you answer the question. Start your response with either Action or an indexed Thought
|
| 3673 |
+
Here are some examples:
|
| 3674 |
+
Question: What is the elevation range for the area that the eastern sector of the Colorado orogeny extends into?
|
| 3675 |
+
Let’s think step by step.
|
| 3676 |
+
Thought 1: The eastern sector of Colorado orogeny extends into the High Plains.
|
| 3677 |
+
Thought 2: High Plains rise in elevation from around 1,800 to 7,000 ft
|
| 3678 |
+
Thought 3: The answer is 1,800 to 7,000 ft.
|
| 3679 |
+
Action: Finish[1,800 to 7,000 ft]
|
| 3680 |
+
(examples)
|
| 3681 |
+
Previous trial:
|
| 3682 |
+
(trajectories)
|
| 3683 |
+
(input)
|
| 3684 |
+
E.3
|
| 3685 |
+
Value Function Prompt
|
| 3686 |
+
Analyze the trajectories of a solution to a question answering task. The trajectories are labeled by environmental observations about the situation, thoughts that can reason about the current situation and actions that can be three types:
|
| 3687 |
+
(1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
|
| 3688 |
+
(2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
|
| 3689 |
+
(3) Finish[answer], which returns the answer and finishes the task.
|
| 3690 |
+
Given a question and a trajectory, evaluate its correctness and provide your reasoning and analysis in detail. Focus on the latest thought, action, and observation. Incomplete trajectories can be correct if the thoughts and actions so far are correct, even if the answer is not found yet. Do not generate additional thoughts or actions. Then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
|
| 3691 |
+
Question: Which magazine was started first Arthur’s Magazine or First for Women?
|
| 3692 |
+
Thought 1: I need to search Arthur’s Magazine and First for Women, and find which was started first.
|
| 3693 |
+
Action 1: Search[Arthur’s Magazine]
|
| 3694 |
+
Observation 1: Arthur’s Magazine (18440̆0e20̆0800̆0931846) was an American literary periodical published in Philadelphia in the 19th century. Edited by Timothy Shay Arthur, it featured work by Edgar A. Poe, J.H. Ingraham, Sarah Josepha Hale, Thomas G. Spear, and others.[1][2] In May 1846 it was merged into Godey’s Lady’s Book.[3]
|
| 3695 |
+
This trajectory is correct as it is reasonable to search for the first magazine provided in the question. It is also better to have simple searches corresponding to a single entity, making this the best action.
|
| 3696 |
+
Thus the correctness score is 10
|
| 3697 |
+
(other examples)
|
| 3698 |
+
(failed trajectories)
|
| 3699 |
+
(context)
|
| 3700 |
+
E.4
|
| 3701 |
+
Reflection Prompt
|
| 3702 |
+
Analyze the trajectories of a solution to a question answering task. The trajectories are labeled by environmental observations about the situation, thoughts that can reason about the current situation and actions that can be three types:
|
| 3703 |
+
(1) Search[entity], which searches the exact entity on Wikipedia and returns the first paragraph if it exists. If not, it will return some similar entities to search.
|
| 3704 |
+
(2) Lookup[keyword], which returns the next sentence containing keyword in the current passage.
|
| 3705 |
+
(3) Finish[answer], which returns the answer and finishes the task.
|
| 3706 |
+
Given a question and a trajectory, evaluate its correctness and provide your reasoning and analysis in detail. Focus on the latest thought, action, and observation. Incomplete trajectories can be correct if the thoughts and actions so far are correct, even if the answer is not found yet. Do not generate additional thoughts or actions. Then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
|
| 3707 |
+
Question: Which magazine was started first Arthur’s Magazine or First for Women?
|
| 3708 |
+
Thought 1: I need to search Arthur’s Magazine and First for Women, and find which was started first.
|
| 3709 |
+
Action 1: Search[Arthur’s Magazine]
|
| 3710 |
+
Observation 1: Arthur’s Magazine (18440̆0e20̆0800̆0931846) was an American literary periodical published in Philadelphia in the 19th century. Edited by Timothy Shay Arthur, it featured work by Edgar A. Poe, J.H. Ingraham, Sarah Josepha Hale, Thomas G. Spear, and others.[1][2] In May 1846 it was merged into Godey’s Lady’s Book.[3]
|
| 3711 |
+
This trajectory is correct as it is reasonable to search for the first magazine provided in the question. It is also better to have simple searches corresponding to a single entity, making this the best action.
|
| 3712 |
+
Thus the correctness score is 10
|
| 3713 |
+
(other examples)
|
| 3714 |
+
(failed trajectories)
|
| 3715 |
+
(context)
|
| 3716 |
+
Appendix F
|
| 3717 |
+
Programming Prompts
|
| 3718 |
+
F.1
|
| 3719 |
+
HumanEval function implementation example
|
| 3720 |
+
Sample function signature:
|
| 3721 |
+
⬇
|
| 3722 |
+
def
|
| 3723 |
+
minSubArraySum
|
| 3724 |
+
(
|
| 3725 |
+
nums
|
| 3726 |
+
):
|
| 3727 |
+
Given
|
| 3728 |
+
an
|
| 3729 |
+
array
|
| 3730 |
+
of
|
| 3731 |
+
integers
|
| 3732 |
+
nums
|
| 3733 |
+
,
|
| 3734 |
+
find
|
| 3735 |
+
the
|
| 3736 |
+
minimum
|
| 3737 |
+
sum
|
| 3738 |
+
of
|
| 3739 |
+
any
|
| 3740 |
+
non
|
| 3741 |
+
-
|
| 3742 |
+
empty
|
| 3743 |
+
sub
|
| 3744 |
+
-
|
| 3745 |
+
array
|
| 3746 |
+
of
|
| 3747 |
+
nums
|
| 3748 |
+
.
|
| 3749 |
+
Example
|
| 3750 |
+
minSubArraySum
|
| 3751 |
+
([2,
|
| 3752 |
+
3,
|
| 3753 |
+
4,
|
| 3754 |
+
1,
|
| 3755 |
+
2,
|
| 3756 |
+
4])
|
| 3757 |
+
==
|
| 3758 |
+
1
|
| 3759 |
+
minSubArraySum
|
| 3760 |
+
([-1,
|
| 3761 |
+
-2,
|
| 3762 |
+
-3])
|
| 3763 |
+
==
|
| 3764 |
+
-6
|
| 3765 |
+
Sample function body implementation:
|
| 3766 |
+
⬇
|
| 3767 |
+
min_sum
|
| 3768 |
+
=
|
| 3769 |
+
float
|
| 3770 |
+
(’
|
| 3771 |
+
inf
|
| 3772 |
+
’)
|
| 3773 |
+
for
|
| 3774 |
+
i
|
| 3775 |
+
in
|
| 3776 |
+
range
|
| 3777 |
+
(
|
| 3778 |
+
len
|
| 3779 |
+
(
|
| 3780 |
+
nums
|
| 3781 |
+
)):
|
| 3782 |
+
current_sum
|
| 3783 |
+
=
|
| 3784 |
+
0
|
| 3785 |
+
for
|
| 3786 |
+
j
|
| 3787 |
+
in
|
| 3788 |
+
range
|
| 3789 |
+
(
|
| 3790 |
+
i
|
| 3791 |
+
,
|
| 3792 |
+
len
|
| 3793 |
+
(
|
| 3794 |
+
nums
|
| 3795 |
+
)):
|
| 3796 |
+
current_sum
|
| 3797 |
+
+=
|
| 3798 |
+
nums
|
| 3799 |
+
[
|
| 3800 |
+
j
|
| 3801 |
+
]
|
| 3802 |
+
if
|
| 3803 |
+
current_sum
|
| 3804 |
+
<
|
| 3805 |
+
min_sum
|
| 3806 |
+
:
|
| 3807 |
+
min_sum
|
| 3808 |
+
=
|
| 3809 |
+
current_sum
|
| 3810 |
+
return
|
| 3811 |
+
min_sum
|
| 3812 |
+
F.2
|
| 3813 |
+
Base Acting/Reasoning Prompt
|
| 3814 |
+
You are an AI Python assistant. You will be given your previous implementation of a function, a series of unit tests results, and your self-reflection on your previous implementation. Write your full implementation (restate the function signature).
|
| 3815 |
+
Example 1:
|
| 3816 |
+
[previous impl]:
|
| 3817 |
+
⬇
|
| 3818 |
+
def
|
| 3819 |
+
add
|
| 3820 |
+
(
|
| 3821 |
+
a
|
| 3822 |
+
:
|
| 3823 |
+
int
|
| 3824 |
+
,
|
| 3825 |
+
b
|
| 3826 |
+
:
|
| 3827 |
+
int
|
| 3828 |
+
)
|
| 3829 |
+
->
|
| 3830 |
+
int
|
| 3831 |
+
:
|
| 3832 |
+
”””
|
| 3833 |
+
Given
|
| 3834 |
+
integers
|
| 3835 |
+
a
|
| 3836 |
+
and
|
| 3837 |
+
b
|
| 3838 |
+
,
|
| 3839 |
+
return
|
| 3840 |
+
the
|
| 3841 |
+
total
|
| 3842 |
+
value
|
| 3843 |
+
of
|
| 3844 |
+
a
|
| 3845 |
+
and
|
| 3846 |
+
b
|
| 3847 |
+
.
|
| 3848 |
+
”””
|
| 3849 |
+
return
|
| 3850 |
+
a
|
| 3851 |
+
-
|
| 3852 |
+
b
|
| 3853 |
+
[unit test results from previous impl]:
|
| 3854 |
+
Tested passed:
|
| 3855 |
+
Tests failed:
|
| 3856 |
+
assert add(1, 2) == 3 # output: -1
|
| 3857 |
+
assert add(1, 2) == 4 # output: -1
|
| 3858 |
+
[reflection on previous impl]:
|
| 3859 |
+
The implementation failed the test cases where the input integers are 1 and 2. The issue arises because the code does not add the two integers together, but instead subtracts the second integer from the first. To fix this issue, we should change the operator from ‘-‘ to ‘+‘ in the return statement. This will ensure that the function returns the correct output for the given input.
|
| 3860 |
+
[improved impl]:
|
| 3861 |
+
⬇
|
| 3862 |
+
def
|
| 3863 |
+
add
|
| 3864 |
+
(
|
| 3865 |
+
a
|
| 3866 |
+
:
|
| 3867 |
+
int
|
| 3868 |
+
,
|
| 3869 |
+
b
|
| 3870 |
+
:
|
| 3871 |
+
int
|
| 3872 |
+
)
|
| 3873 |
+
->
|
| 3874 |
+
int
|
| 3875 |
+
:
|
| 3876 |
+
”””
|
| 3877 |
+
Given
|
| 3878 |
+
integers
|
| 3879 |
+
a
|
| 3880 |
+
and
|
| 3881 |
+
b
|
| 3882 |
+
,
|
| 3883 |
+
return
|
| 3884 |
+
the
|
| 3885 |
+
total
|
| 3886 |
+
value
|
| 3887 |
+
of
|
| 3888 |
+
a
|
| 3889 |
+
and
|
| 3890 |
+
b
|
| 3891 |
+
.
|
| 3892 |
+
”””
|
| 3893 |
+
return
|
| 3894 |
+
a
|
| 3895 |
+
+
|
| 3896 |
+
b
|
| 3897 |
+
F.3
|
| 3898 |
+
Reflection Prompt
|
| 3899 |
+
You are a Python programming assistant. You will be given a function implementation and a series of unit test results. Your goal is to write a few sentences to explain why your implementation is wrong as indicated by the tests. You will need this as guidance when you try again later. Only provide the few sentence description in your answer, not the implementation. You will be given a few examples by the user.
|
| 3900 |
+
Example 1:
|
| 3901 |
+
[previous impl]:
|
| 3902 |
+
⬇
|
| 3903 |
+
def
|
| 3904 |
+
add
|
| 3905 |
+
(
|
| 3906 |
+
a
|
| 3907 |
+
:
|
| 3908 |
+
int
|
| 3909 |
+
,
|
| 3910 |
+
b
|
| 3911 |
+
:
|
| 3912 |
+
int
|
| 3913 |
+
)
|
| 3914 |
+
->
|
| 3915 |
+
int
|
| 3916 |
+
:
|
| 3917 |
+
”””
|
| 3918 |
+
Given
|
| 3919 |
+
integers
|
| 3920 |
+
a
|
| 3921 |
+
and
|
| 3922 |
+
b
|
| 3923 |
+
,
|
| 3924 |
+
return
|
| 3925 |
+
the
|
| 3926 |
+
total
|
| 3927 |
+
value
|
| 3928 |
+
of
|
| 3929 |
+
a
|
| 3930 |
+
and
|
| 3931 |
+
b
|
| 3932 |
+
.
|
| 3933 |
+
”””
|
| 3934 |
+
return
|
| 3935 |
+
a
|
| 3936 |
+
-
|
| 3937 |
+
b
|
| 3938 |
+
[unit test results from previous impl]:
|
| 3939 |
+
Tested passed:
|
| 3940 |
+
Tests failed:
|
| 3941 |
+
assert add(1, 2) == 3 # output: -1
|
| 3942 |
+
assert add(1, 2) == 4 # output: -1
|
| 3943 |
+
[reflection on previous impl]:
|
| 3944 |
+
The implementation failed the test cases where the input integers are 1 and 2. The issue arises because the code does not add the two integers together, but instead subtracts the second integer from the first. To fix this issue, we should change the operator from ‘-‘ to ‘+‘ in the return statement. This will ensure that the function returns the correct output for the given input.
|
| 3945 |
+
F.4
|
| 3946 |
+
Test Case Generation Prompt
|
| 3947 |
+
You are an AI coding assistant that can write unique, diverse, and intuitive unit tests for functions given the signature and docstring.
|
| 3948 |
+
(input)
|
| 3949 |
+
Appendix G
|
| 3950 |
+
WebShop Prompts
|
| 3951 |
+
G.1
|
| 3952 |
+
Acting Prompt
|
| 3953 |
+
Webshop
|
| 3954 |
+
Instruction:
|
| 3955 |
+
i would like a 3 ounce bottle of bright citrus deodorant for sensitive skin, and price lower than 50.00 dollars
|
| 3956 |
+
[Search]
|
| 3957 |
+
Action: search[3 ounce bright citrus deodorant sensitive skin]
|
| 3958 |
+
Observation:
|
| 3959 |
+
[Back to Search]
|
| 3960 |
+
Page 1 (Total results: 50)
|
| 3961 |
+
[Next
|
| 3962 |
+
]
|
| 3963 |
+
[B078GWRC1J]
|
| 3964 |
+
Bright Citrus Deodorant by Earth Mama — Natural and Safe for Sensitive Skin, Pregnancy and Breastfeeding, Contains Organic Calendula 3-Ounce
|
| 3965 |
+
$10.99
|
| 3966 |
+
[B078GTKVXY]
|
| 3967 |
+
Ginger Fresh Deodorant by Earth Mama — Natural and Safe for Sensitive Skin, Pregnancy and Breastfeeding, Contains Organic Calendula 3-Ounce
|
| 3968 |
+
$10.99
|
| 3969 |
+
[B08KBVJ4XN]
|
| 3970 |
+
Barrel and Oak - Aluminum-Free Deodorant, Deodorant for Men, Essential Oil-Based Scent, 24-Hour Odor Protection, Cedar & Patchouli Blend, Gentle on Sensitive Skin (Mountain Sage, 2.7 oz, 2-Pack)
|
| 3971 |
+
$15.95
|
| 3972 |
+
Action: think[B078GWRC1J and B078GTKVXY are bright citrus deodorant less then 50 dollars. I can check B078GWRC1J first.]
|
| 3973 |
+
Observation: OK.
|
| 3974 |
+
Action: click[B078GWRC1J]
|
| 3975 |
+
Observation:
|
| 3976 |
+
[Back to Search]
|
| 3977 |
+
[Prev]
|
| 3978 |
+
scent [assorted scents][bright citrus][calming lavender][ginger fresh][simply non-scents]
|
| 3979 |
+
size [travel set (4-pack)][3 ounce (pack of 1)][3-ounce (2-pack)]
|
| 3980 |
+
Bright Citrus Deodorant by Earth Mama — Natural and Safe for Sensitive Skin, Pregnancy and Breastfeeding, Contains Organic Calendula 3-Ounce
|
| 3981 |
+
Price: $10.99
|
| 3982 |
+
Rating: N.A.
|
| 3983 |
+
[Description]
|
| 3984 |
+
[Features]
|
| 3985 |
+
[Reviews]
|
| 3986 |
+
[Buy Now]
|
| 3987 |
+
Action: think[For 3 ounce bottle of bright citrus deodorant for sensitive skin, the item has options ’bright citrus’ and ’3 ounce (pack of 1)’ and seems good to buy.]
|
| 3988 |
+
Observation: OK.
|
| 3989 |
+
Action: click[bright citrus]
|
| 3990 |
+
Observation: You have clicked bright citrus.
|
| 3991 |
+
Action: click[3 ounce (pack of 1)]
|
| 3992 |
+
Observation: You have clicked 3 ounce (pack of 1).
|
| 3993 |
+
Action: click[Buy Now]
|
| 3994 |
+
G.2
|
| 3995 |
+
Value Function Prompt
|
| 3996 |
+
Given an item to purchase and a trajectory that aims to buy an item that exactly matches the specification, which corresponds to the ideal score of 1.0, analyze the following trajectory, then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
|
| 3997 |
+
Here are some examples
|
| 3998 |
+
Webshop
|
| 3999 |
+
Instruction:
|
| 4000 |
+
i am looking for dairy free and apple variety pack of chips, and price lower than 30.00 dollars
|
| 4001 |
+
[Search]
|
| 4002 |
+
Action: search[dairy free and apple variety pack of chips]
|
| 4003 |
+
Observation:
|
| 4004 |
+
[Back to Search]
|
| 4005 |
+
Page 1 (Total results: 50)
|
| 4006 |
+
[Next]
|
| 4007 |
+
[B07HRFJWP8]
|
| 4008 |
+
Enjoy Life Foods Soft Baked Ovals, Breakfast Bars, Nut Free Bars, Soy Free, Dairy Free, Non GMO, Gluten Free, Vegan, Variety Pack, 4 Boxes (20 Bars Total)
|
| 4009 |
+
$100.0
|
| 4010 |
+
[B01KMHY5PG]
|
| 4011 |
+
Enjoy Life Soft Baked Chewy Bars, Variety Pack, Nut Free Bars, Soy Free, Dairy Free, Gluten Free, 6 Boxes (30 Total Bars)
|
| 4012 |
+
$21.49
|
| 4013 |
+
[B008D2X8C4]
|
| 4014 |
+
Enjoy Life Lentil Chips Variety Pack, Dairy Free Chips, Soy Free, Nut Free, Non GMO, Vegan, Gluten Free, 24 Bags (0.8 oz)
|
| 4015 |
+
$100.0
|
| 4016 |
+
Action: think[The first two items are not what I want. The third item is the dairy free and apple variety pack of chips I want.]
|
| 4017 |
+
Observation: OK.
|
| 4018 |
+
Action: click[B008D2X8C4]
|
| 4019 |
+
Observation:
|
| 4020 |
+
[Back to Search]
|
| 4021 |
+
[Prev]
|
| 4022 |
+
flavor name [dill and sour cream][garlic & parmesan][light sea salt][margherita pizza][thai chili lime][variety pack]
|
| 4023 |
+
size [0.8 ounce (pack of 24)][4 ounce (pack of 12)]
|
| 4024 |
+
Price: $100.0
|
| 4025 |
+
Rating: N.A.
|
| 4026 |
+
[Description]
|
| 4027 |
+
[Features]
|
| 4028 |
+
[Reviews]
|
| 4029 |
+
[Buy Now]
|
| 4030 |
+
Action: think[The item has options ’variety pack’ and ’0.8 ounce (pack of 24)’ and seems good to buy.]
|
| 4031 |
+
Observation: OK.
|
| 4032 |
+
Action: click[variety pack]
|
| 4033 |
+
Observation: You have clicked variety pack.
|
| 4034 |
+
Action: click[0.8 ounce (pack of 24)]
|
| 4035 |
+
Observation: You have clicked 0.8 ounce (pack of 24).
|
| 4036 |
+
Action: click[Buy Now]
|
| 4037 |
+
Reflection: In this attempt, I was unsuccessful. I accidentally bought a product that was $100, which is more than my budget of $30. Either way, the initial search results were not good. Next time, I will do search[”variety pack of chips”] and then check if the results meet the dairy free and the $30 budget constraints. I will continue to refine my searches so that I can find more products.
|
| 4038 |
+
Thus the correctness score is 5
|
| 4039 |
+
(input)
|
| 4040 |
+
G.3
|
| 4041 |
+
Reflection Prompt
|
| 4042 |
+
Given an item to purchase and a trajectory that aims to buy an item that exactly matches the specification, which corresponds to the ideal score of 1.0, analyze the following trajectory, then at the last line conclude ”Thus the correctness score is s”, where s is an integer from 1 to 10.
|
| 4043 |
+
Here are some examples
|
| 4044 |
+
Webshop
|
| 4045 |
+
Instruction:
|
| 4046 |
+
i am looking for dairy free and apple variety pack of chips, and price lower than 30.00 dollars
|
| 4047 |
+
[Search]
|
| 4048 |
+
Action: search[dairy free and apple variety pack of chips]
|
| 4049 |
+
Observation:
|
| 4050 |
+
[Back to Search]
|
| 4051 |
+
Page 1 (Total results: 50)
|
| 4052 |
+
[Next]
|
| 4053 |
+
[B07HRFJWP8]
|
| 4054 |
+
Enjoy Life Foods Soft Baked Ovals, Breakfast Bars, Nut Free Bars, Soy Free, Dairy Free, Non GMO, Gluten Free, Vegan, Variety Pack, 4 Boxes (20 Bars Total)
|
| 4055 |
+
$100.0
|
| 4056 |
+
[B01KMHY5PG]
|
| 4057 |
+
Enjoy Life Soft Baked Chewy Bars, Variety Pack, Nut Free Bars, Soy Free, Dairy Free, Gluten Free, 6 Boxes (30 Total Bars)
|
| 4058 |
+
$21.49
|
| 4059 |
+
[B008D2X8C4]
|
| 4060 |
+
Enjoy Life Lentil Chips Variety Pack, Dairy Free Chips, Soy Free, Nut Free, Non GMO, Vegan, Gluten Free, 24 Bags (0.8 oz)
|
| 4061 |
+
$100.0
|
| 4062 |
+
Action: think[The first two items are not what I want. The third item is the dairy free and apple variety pack of chips I want.]
|
| 4063 |
+
Observation: OK.
|
| 4064 |
+
Action: click[B008D2X8C4]
|
| 4065 |
+
Observation:
|
| 4066 |
+
[Back to Search]
|
| 4067 |
+
[Prev]
|
| 4068 |
+
flavor name [dill and sour cream][garlic & parmesan][light sea salt][margherita pizza][thai chili lime][variety pack]
|
| 4069 |
+
size [0.8 ounce (pack of 24)][4 ounce (pack of 12)]
|
| 4070 |
+
Price: $100.0
|
| 4071 |
+
Rating: N.A.
|
| 4072 |
+
[Description]
|
| 4073 |
+
[Features]
|
| 4074 |
+
[Reviews]
|
| 4075 |
+
[Buy Now]
|
| 4076 |
+
Action: think[The item has options ’variety pack’ and ’0.8 ounce (pack of 24)’ and seems good to buy.]
|
| 4077 |
+
Observation: OK.
|
| 4078 |
+
Action: click[variety pack]
|
| 4079 |
+
Observation: You have clicked variety pack.
|
| 4080 |
+
Action: click[0.8 ounce (pack of 24)]
|
| 4081 |
+
Observation: You have clicked 0.8 ounce (pack of 24).
|
| 4082 |
+
Action: click[Buy Now]
|
| 4083 |
+
Reflection: In this attempt, I was unsuccessful. I accidentally bought a product that was $100, which is more than my budget of $30. Either way, the initial search results were not good. Next time, I will do search[”variety pack of chips”] and then check if the results meet the dairy free and the $30 budget constraints. I will continue to refine my searches so that I can find more products.
|
| 4084 |
+
(input)
|
| 4085 |
+
Reflection:
|
| 4086 |
+
◄
|
| 4087 |
+
Feeling
|
| 4088 |
+
lucky?
|
| 4089 |
+
Conversion
|
| 4090 |
+
report
|
| 4091 |
+
Report
|
| 4092 |
+
an issue
|
| 4093 |
+
View original
|
| 4094 |
+
on arXiv
|
| 4095 |
+
►
|
|
@@ -0,0 +1,202 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2310.04406] Language Agent Tree Search Unifies Reasoning Acting and Planning
|
| 3 |
+
in Language Models'
|
| 4 |
+
id: 231004406-language-agent-tree-search-unifies-reasoning-acting-and-planning-in-la
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:39:54.848871Z'
|
| 8 |
+
source: https://arxiv.org/abs/2310.04406
|
| 9 |
+
source_domain: arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:39:54.848723Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2310.04406] Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
|
| 20 |
+
Computer Science > Artificial Intelligence
|
| 21 |
+
arXiv:2310.04406
|
| 22 |
+
(cs)
|
| 23 |
+
[Submitted on 6 Oct 2023 (
|
| 24 |
+
v1
|
| 25 |
+
), last revised 6 Jun 2024 (this version, v3)]
|
| 26 |
+
Title:
|
| 27 |
+
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
|
| 28 |
+
Authors:
|
| 29 |
+
Andy Zhou
|
| 30 |
+
,
|
| 31 |
+
Kai Yan
|
| 32 |
+
,
|
| 33 |
+
Michal Shlapentokh-Rothman
|
| 34 |
+
,
|
| 35 |
+
Haohan Wang
|
| 36 |
+
,
|
| 37 |
+
Yu-Xiong Wang
|
| 38 |
+
View a PDF of the paper titled Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models, by Andy Zhou and 4 other authors
|
| 39 |
+
View PDF
|
| 40 |
+
HTML (experimental)
|
| 41 |
+
Abstract:
|
| 42 |
+
While language models (LMs) have shown potential across a range of decision-making tasks, their reliance on simple acting processes limits their broad deployment as autonomous agents. In this paper, we introduce Language Agent Tree Search (LATS) -- the first general framework that synergizes the capabilities of LMs in reasoning, acting, and planning. By leveraging the in-context learning ability of LMs, we integrate Monte Carlo Tree Search into LATS to enable LMs as agents, along with LM-powered value functions and self-reflections for proficient exploration and enhanced decision-making. A key feature of our approach is the incorporation of an environment for external feedback, which offers a more deliberate and adaptive problem-solving mechanism that surpasses the constraints of existing techniques. Our experimental evaluation across diverse domains, including programming, interactive question-answering (QA), web navigation, and math, validates the effectiveness and generality of LATS in decision-making while maintaining competitive or improved reasoning performance. Notably, LATS achieves state-of-the-art pass@1 accuracy (92.7%) for programming on HumanEval with GPT-4 and demonstrates gradient-free performance (average score of 75.9) comparable to gradient-based fine-tuning for web navigation on WebShop with GPT-3.5. Code can be found at
|
| 43 |
+
this https URL
|
| 44 |
+
Comments:
|
| 45 |
+
Code at
|
| 46 |
+
this https URL
|
| 47 |
+
Subjects:
|
| 48 |
+
Artificial Intelligence (cs.AI)
|
| 49 |
+
; Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
|
| 50 |
+
Cite as:
|
| 51 |
+
arXiv:2310.04406
|
| 52 |
+
[cs.AI]
|
| 53 |
+
(or
|
| 54 |
+
arXiv:2310.04406v3
|
| 55 |
+
[cs.AI]
|
| 56 |
+
for this version)
|
| 57 |
+
https://doi.org/10.48550/arXiv.2310.04406
|
| 58 |
+
Focus to learn more
|
| 59 |
+
arXiv-issued DOI via DataCite
|
| 60 |
+
Submission history
|
| 61 |
+
From: Andy Zhou [
|
| 62 |
+
view email
|
| 63 |
+
]
|
| 64 |
+
[v1]
|
| 65 |
+
Fri, 6 Oct 2023 17:55:11 UTC (371 KB)
|
| 66 |
+
[v2]
|
| 67 |
+
Tue, 5 Dec 2023 05:25:55 UTC (465 KB)
|
| 68 |
+
[v3]
|
| 69 |
+
Thu, 6 Jun 2024 02:51:17 UTC (960 KB)
|
| 70 |
+
Full-text links:
|
| 71 |
+
Access Paper:
|
| 72 |
+
View a PDF of the paper titled Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models, by Andy Zhou and 4 other authors
|
| 73 |
+
View PDF
|
| 74 |
+
HTML (experimental)
|
| 75 |
+
TeX Source
|
| 76 |
+
view license
|
| 77 |
+
Current browse context:
|
| 78 |
+
cs.AI
|
| 79 |
+
< prev
|
| 80 |
+
|
|
| 81 |
+
next >
|
| 82 |
+
new
|
| 83 |
+
|
|
| 84 |
+
recent
|
| 85 |
+
|
|
| 86 |
+
2023-10
|
| 87 |
+
Change to browse by:
|
| 88 |
+
cs
|
| 89 |
+
cs.CL
|
| 90 |
+
cs.CV
|
| 91 |
+
cs.LG
|
| 92 |
+
References & Citations
|
| 93 |
+
NASA ADS
|
| 94 |
+
Google Scholar
|
| 95 |
+
Semantic Scholar
|
| 96 |
+
export BibTeX citation
|
| 97 |
+
Loading...
|
| 98 |
+
BibTeX formatted citation
|
| 99 |
+
×
|
| 100 |
+
loading...
|
| 101 |
+
Data provided by:
|
| 102 |
+
Bookmark
|
| 103 |
+
Bibliographic Tools
|
| 104 |
+
Bibliographic and Citation Tools
|
| 105 |
+
Bibliographic Explorer Toggle
|
| 106 |
+
Bibliographic Explorer
|
| 107 |
+
(
|
| 108 |
+
What is the Explorer?
|
| 109 |
+
)
|
| 110 |
+
Connected Papers Toggle
|
| 111 |
+
Connected Papers
|
| 112 |
+
(
|
| 113 |
+
What is Connected Papers?
|
| 114 |
+
)
|
| 115 |
+
Litmaps Toggle
|
| 116 |
+
Litmaps
|
| 117 |
+
(
|
| 118 |
+
What is Litmaps?
|
| 119 |
+
)
|
| 120 |
+
scite.ai Toggle
|
| 121 |
+
scite Smart Citations
|
| 122 |
+
(
|
| 123 |
+
What are Smart Citations?
|
| 124 |
+
)
|
| 125 |
+
Code, Data, Media
|
| 126 |
+
Code, Data and Media Associated with this Article
|
| 127 |
+
alphaXiv Toggle
|
| 128 |
+
alphaXiv
|
| 129 |
+
(
|
| 130 |
+
What is alphaXiv?
|
| 131 |
+
)
|
| 132 |
+
Links to Code Toggle
|
| 133 |
+
CatalyzeX Code Finder for Papers
|
| 134 |
+
(
|
| 135 |
+
What is CatalyzeX?
|
| 136 |
+
)
|
| 137 |
+
DagsHub Toggle
|
| 138 |
+
DagsHub
|
| 139 |
+
(
|
| 140 |
+
What is DagsHub?
|
| 141 |
+
)
|
| 142 |
+
GotitPub Toggle
|
| 143 |
+
Gotit.pub
|
| 144 |
+
(
|
| 145 |
+
What is GotitPub?
|
| 146 |
+
)
|
| 147 |
+
Huggingface Toggle
|
| 148 |
+
Hugging Face
|
| 149 |
+
(
|
| 150 |
+
What is Huggingface?
|
| 151 |
+
)
|
| 152 |
+
ScienceCast Toggle
|
| 153 |
+
ScienceCast
|
| 154 |
+
(
|
| 155 |
+
What is ScienceCast?
|
| 156 |
+
)
|
| 157 |
+
Demos
|
| 158 |
+
Demos
|
| 159 |
+
Replicate Toggle
|
| 160 |
+
Replicate
|
| 161 |
+
(
|
| 162 |
+
What is Replicate?
|
| 163 |
+
)
|
| 164 |
+
Spaces Toggle
|
| 165 |
+
Hugging Face Spaces
|
| 166 |
+
(
|
| 167 |
+
What is Spaces?
|
| 168 |
+
)
|
| 169 |
+
Spaces Toggle
|
| 170 |
+
TXYZ.AI
|
| 171 |
+
(
|
| 172 |
+
What is TXYZ.AI?
|
| 173 |
+
)
|
| 174 |
+
Related Papers
|
| 175 |
+
Recommenders and Search Tools
|
| 176 |
+
Link to Influence Flower
|
| 177 |
+
Influence Flower
|
| 178 |
+
(
|
| 179 |
+
What are Influence Flowers?
|
| 180 |
+
)
|
| 181 |
+
Core recommender toggle
|
| 182 |
+
CORE Recommender
|
| 183 |
+
(
|
| 184 |
+
What is CORE?
|
| 185 |
+
)
|
| 186 |
+
Author
|
| 187 |
+
Venue
|
| 188 |
+
Institution
|
| 189 |
+
Topic
|
| 190 |
+
About arXivLabs
|
| 191 |
+
arXivLabs: experimental projects with community collaborators
|
| 192 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 193 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 194 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 195 |
+
Learn more about arXivLabs
|
| 196 |
+
.
|
| 197 |
+
Which authors of this paper are endorsers?
|
| 198 |
+
|
|
| 199 |
+
Disable MathJax
|
| 200 |
+
(
|
| 201 |
+
What is MathJax?
|
| 202 |
+
)
|
|
@@ -0,0 +1,203 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2310.06770] SWE-bench: Can Language Models Resolve Real-World GitHub Issues?'
|
| 3 |
+
id: 231006770-swe-bench-can-language-models-resolve-real-world-github-issues
|
| 4 |
+
tags:
|
| 5 |
+
- deepread
|
| 6 |
+
created: '2026-06-10T00:23:35.577828Z'
|
| 7 |
+
source: https://arxiv.org/abs/2310.06770
|
| 8 |
+
source_domain: arxiv.org
|
| 9 |
+
fetched_at: '2026-06-10T00:23:35.577638Z'
|
| 10 |
+
fetch_provider: builtin
|
| 11 |
+
status: draft
|
| 12 |
+
type: note
|
| 13 |
+
tier: institutional
|
| 14 |
+
content_type: paper
|
| 15 |
+
deprecated: false
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
[2310.06770] SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
|
| 19 |
+
Computer Science > Computation and Language
|
| 20 |
+
arXiv:2310.06770
|
| 21 |
+
(cs)
|
| 22 |
+
[Submitted on 10 Oct 2023 (
|
| 23 |
+
v1
|
| 24 |
+
), last revised 11 Nov 2024 (this version, v3)]
|
| 25 |
+
Title:
|
| 26 |
+
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
|
| 27 |
+
Authors:
|
| 28 |
+
Carlos E. Jimenez
|
| 29 |
+
,
|
| 30 |
+
John Yang
|
| 31 |
+
,
|
| 32 |
+
Alexander Wettig
|
| 33 |
+
,
|
| 34 |
+
Shunyu Yao
|
| 35 |
+
,
|
| 36 |
+
Kexin Pei
|
| 37 |
+
,
|
| 38 |
+
Ofir Press
|
| 39 |
+
,
|
| 40 |
+
Karthik Narasimhan
|
| 41 |
+
View a PDF of the paper titled SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, by Carlos E. Jimenez and 6 other authors
|
| 42 |
+
View PDF
|
| 43 |
+
Abstract:
|
| 44 |
+
Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models. To this end, we introduce SWE-bench, an evaluation framework consisting of $2,294$ software engineering problems drawn from real GitHub issues and corresponding pull requests across $12$ popular Python repositories. Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks. Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere $1.96$% of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous.
|
| 45 |
+
Comments:
|
| 46 |
+
Data, code, and leaderboard are available at
|
| 47 |
+
this https URL
|
| 48 |
+
ICLR 2024,
|
| 49 |
+
this https URL
|
| 50 |
+
Subjects:
|
| 51 |
+
Computation and Language (cs.CL)
|
| 52 |
+
; Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
|
| 53 |
+
Cite as:
|
| 54 |
+
arXiv:2310.06770
|
| 55 |
+
[cs.CL]
|
| 56 |
+
(or
|
| 57 |
+
arXiv:2310.06770v3
|
| 58 |
+
[cs.CL]
|
| 59 |
+
for this version)
|
| 60 |
+
https://doi.org/10.48550/arXiv.2310.06770
|
| 61 |
+
Focus to learn more
|
| 62 |
+
arXiv-issued DOI via DataCite
|
| 63 |
+
Submission history
|
| 64 |
+
From: Carlos E. Jimenez [
|
| 65 |
+
view email
|
| 66 |
+
]
|
| 67 |
+
[v1]
|
| 68 |
+
Tue, 10 Oct 2023 16:47:29 UTC (2,003 KB)
|
| 69 |
+
[v2]
|
| 70 |
+
Fri, 5 Apr 2024 18:16:29 UTC (2,258 KB)
|
| 71 |
+
[v3]
|
| 72 |
+
Mon, 11 Nov 2024 23:05:04 UTC (2,398 KB)
|
| 73 |
+
Full-text links:
|
| 74 |
+
Access Paper:
|
| 75 |
+
View a PDF of the paper titled SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, by Carlos E. Jimenez and 6 other authors
|
| 76 |
+
View PDF
|
| 77 |
+
TeX Source
|
| 78 |
+
view license
|
| 79 |
+
Current browse context:
|
| 80 |
+
cs.CL
|
| 81 |
+
< prev
|
| 82 |
+
|
|
| 83 |
+
next >
|
| 84 |
+
new
|
| 85 |
+
|
|
| 86 |
+
recent
|
| 87 |
+
|
|
| 88 |
+
2023-10
|
| 89 |
+
Change to browse by:
|
| 90 |
+
cs
|
| 91 |
+
cs.AI
|
| 92 |
+
cs.SE
|
| 93 |
+
References & Citations
|
| 94 |
+
NASA ADS
|
| 95 |
+
Google Scholar
|
| 96 |
+
Semantic Scholar
|
| 97 |
+
export BibTeX citation
|
| 98 |
+
Loading...
|
| 99 |
+
BibTeX formatted citation
|
| 100 |
+
×
|
| 101 |
+
loading...
|
| 102 |
+
Data provided by:
|
| 103 |
+
Bookmark
|
| 104 |
+
Bibliographic Tools
|
| 105 |
+
Bibliographic and Citation Tools
|
| 106 |
+
Bibliographic Explorer Toggle
|
| 107 |
+
Bibliographic Explorer
|
| 108 |
+
(
|
| 109 |
+
What is the Explorer?
|
| 110 |
+
)
|
| 111 |
+
Connected Papers Toggle
|
| 112 |
+
Connected Papers
|
| 113 |
+
(
|
| 114 |
+
What is Connected Papers?
|
| 115 |
+
)
|
| 116 |
+
Litmaps Toggle
|
| 117 |
+
Litmaps
|
| 118 |
+
(
|
| 119 |
+
What is Litmaps?
|
| 120 |
+
)
|
| 121 |
+
scite.ai Toggle
|
| 122 |
+
scite Smart Citations
|
| 123 |
+
(
|
| 124 |
+
What are Smart Citations?
|
| 125 |
+
)
|
| 126 |
+
Code, Data, Media
|
| 127 |
+
Code, Data and Media Associated with this Article
|
| 128 |
+
alphaXiv Toggle
|
| 129 |
+
alphaXiv
|
| 130 |
+
(
|
| 131 |
+
What is alphaXiv?
|
| 132 |
+
)
|
| 133 |
+
Links to Code Toggle
|
| 134 |
+
CatalyzeX Code Finder for Papers
|
| 135 |
+
(
|
| 136 |
+
What is CatalyzeX?
|
| 137 |
+
)
|
| 138 |
+
DagsHub Toggle
|
| 139 |
+
DagsHub
|
| 140 |
+
(
|
| 141 |
+
What is DagsHub?
|
| 142 |
+
)
|
| 143 |
+
GotitPub Toggle
|
| 144 |
+
Gotit.pub
|
| 145 |
+
(
|
| 146 |
+
What is GotitPub?
|
| 147 |
+
)
|
| 148 |
+
Huggingface Toggle
|
| 149 |
+
Hugging Face
|
| 150 |
+
(
|
| 151 |
+
What is Huggingface?
|
| 152 |
+
)
|
| 153 |
+
ScienceCast Toggle
|
| 154 |
+
ScienceCast
|
| 155 |
+
(
|
| 156 |
+
What is ScienceCast?
|
| 157 |
+
)
|
| 158 |
+
Demos
|
| 159 |
+
Demos
|
| 160 |
+
Replicate Toggle
|
| 161 |
+
Replicate
|
| 162 |
+
(
|
| 163 |
+
What is Replicate?
|
| 164 |
+
)
|
| 165 |
+
Spaces Toggle
|
| 166 |
+
Hugging Face Spaces
|
| 167 |
+
(
|
| 168 |
+
What is Spaces?
|
| 169 |
+
)
|
| 170 |
+
Spaces Toggle
|
| 171 |
+
TXYZ.AI
|
| 172 |
+
(
|
| 173 |
+
What is TXYZ.AI?
|
| 174 |
+
)
|
| 175 |
+
Related Papers
|
| 176 |
+
Recommenders and Search Tools
|
| 177 |
+
Link to Influence Flower
|
| 178 |
+
Influence Flower
|
| 179 |
+
(
|
| 180 |
+
What are Influence Flowers?
|
| 181 |
+
)
|
| 182 |
+
Core recommender toggle
|
| 183 |
+
CORE Recommender
|
| 184 |
+
(
|
| 185 |
+
What is CORE?
|
| 186 |
+
)
|
| 187 |
+
Author
|
| 188 |
+
Venue
|
| 189 |
+
Institution
|
| 190 |
+
Topic
|
| 191 |
+
About arXivLabs
|
| 192 |
+
arXivLabs: experimental projects with community collaborators
|
| 193 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 194 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 195 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 196 |
+
Learn more about arXivLabs
|
| 197 |
+
.
|
| 198 |
+
Which authors of this paper are endorsers?
|
| 199 |
+
|
|
| 200 |
+
Disable MathJax
|
| 201 |
+
(
|
| 202 |
+
What is MathJax?
|
| 203 |
+
)
|
|
@@ -0,0 +1,208 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2311.08105] DiLoCo: Distributed Low-Communication Training of Language Models'
|
| 3 |
+
id: 231108105-diloco-distributed-low-communication-training-of-language-models
|
| 4 |
+
tags:
|
| 5 |
+
- deepread
|
| 6 |
+
created: '2026-06-10T00:30:20.411067Z'
|
| 7 |
+
source: https://arxiv.org/abs/2311.08105
|
| 8 |
+
source_domain: arxiv.org
|
| 9 |
+
fetched_at: '2026-06-10T00:30:20.410923Z'
|
| 10 |
+
fetch_provider: builtin
|
| 11 |
+
status: draft
|
| 12 |
+
type: note
|
| 13 |
+
tier: institutional
|
| 14 |
+
content_type: paper
|
| 15 |
+
deprecated: false
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
[2311.08105] DiLoCo: Distributed Low-Communication Training of Language Models
|
| 19 |
+
Computer Science > Machine Learning
|
| 20 |
+
arXiv:2311.08105
|
| 21 |
+
(cs)
|
| 22 |
+
[Submitted on 14 Nov 2023 (
|
| 23 |
+
v1
|
| 24 |
+
), last revised 23 Sep 2024 (this version, v3)]
|
| 25 |
+
Title:
|
| 26 |
+
DiLoCo: Distributed Low-Communication Training of Language Models
|
| 27 |
+
Authors:
|
| 28 |
+
Arthur Douillard
|
| 29 |
+
,
|
| 30 |
+
Qixuan Feng
|
| 31 |
+
,
|
| 32 |
+
Andrei A. Rusu
|
| 33 |
+
,
|
| 34 |
+
Rachita Chhaparia
|
| 35 |
+
,
|
| 36 |
+
Yani Donchev
|
| 37 |
+
,
|
| 38 |
+
Adhiguna Kuncoro
|
| 39 |
+
,
|
| 40 |
+
Marc'Aurelio Ranzato
|
| 41 |
+
,
|
| 42 |
+
Arthur Szlam
|
| 43 |
+
,
|
| 44 |
+
Jiajun Shen
|
| 45 |
+
View a PDF of the paper titled DiLoCo: Distributed Low-Communication Training of Language Models, by Arthur Douillard and 8 other authors
|
| 46 |
+
View PDF
|
| 47 |
+
HTML (experimental)
|
| 48 |
+
Abstract:
|
| 49 |
+
Large language models (LLM) have become a critical component in many applications of machine learning. However, standard approaches to training LLM require a large number of tightly interconnected accelerators, with devices exchanging gradients and other intermediate states at each optimization step. While it is difficult to build and maintain a single computing cluster hosting many accelerators, it might be easier to find several computing clusters each hosting a smaller number of devices. In this work, we propose a distributed optimization algorithm, Distributed Low-Communication (DiLoCo), that enables training of language models on islands of devices that are poorly connected. The approach is a variant of federated averaging, where the number of inner steps is large, the inner optimizer is AdamW, and the outer optimizer is Nesterov momentum. On the widely used C4 dataset, we show that DiLoCo on 8 workers performs as well as fully synchronous optimization while communicating 500 times less. DiLoCo exhibits great robustness to the data distribution of each worker. It is also robust to resources becoming unavailable over time, and vice versa, it can seamlessly leverage resources that become available during training.
|
| 50 |
+
Subjects:
|
| 51 |
+
Machine Learning (cs.LG)
|
| 52 |
+
; Computation and Language (cs.CL)
|
| 53 |
+
Cite as:
|
| 54 |
+
arXiv:2311.08105
|
| 55 |
+
[cs.LG]
|
| 56 |
+
(or
|
| 57 |
+
arXiv:2311.08105v3
|
| 58 |
+
[cs.LG]
|
| 59 |
+
for this version)
|
| 60 |
+
https://doi.org/10.48550/arXiv.2311.08105
|
| 61 |
+
Focus to learn more
|
| 62 |
+
arXiv-issued DOI via DataCite
|
| 63 |
+
Submission history
|
| 64 |
+
From: Arthur Douillard [
|
| 65 |
+
view email
|
| 66 |
+
]
|
| 67 |
+
[v1]
|
| 68 |
+
Tue, 14 Nov 2023 12:05:45 UTC (1,609 KB)
|
| 69 |
+
[v2]
|
| 70 |
+
Sat, 2 Dec 2023 14:10:14 UTC (1,610 KB)
|
| 71 |
+
[v3]
|
| 72 |
+
Mon, 23 Sep 2024 10:41:27 UTC (1,610 KB)
|
| 73 |
+
Full-text links:
|
| 74 |
+
Access Paper:
|
| 75 |
+
View a PDF of the paper titled DiLoCo: Distributed Low-Communication Training of Language Models, by Arthur Douillard and 8 other authors
|
| 76 |
+
View PDF
|
| 77 |
+
HTML (experimental)
|
| 78 |
+
TeX Source
|
| 79 |
+
view license
|
| 80 |
+
Current browse context:
|
| 81 |
+
cs.LG
|
| 82 |
+
< prev
|
| 83 |
+
|
|
| 84 |
+
next >
|
| 85 |
+
new
|
| 86 |
+
|
|
| 87 |
+
recent
|
| 88 |
+
|
|
| 89 |
+
2023-11
|
| 90 |
+
Change to browse by:
|
| 91 |
+
cs
|
| 92 |
+
cs.CL
|
| 93 |
+
References & Citations
|
| 94 |
+
NASA ADS
|
| 95 |
+
Google Scholar
|
| 96 |
+
Semantic Scholar
|
| 97 |
+
export BibTeX citation
|
| 98 |
+
Loading...
|
| 99 |
+
BibTeX formatted citation
|
| 100 |
+
×
|
| 101 |
+
loading...
|
| 102 |
+
Data provided by:
|
| 103 |
+
Bookmark
|
| 104 |
+
Bibliographic Tools
|
| 105 |
+
Bibliographic and Citation Tools
|
| 106 |
+
Bibliographic Explorer Toggle
|
| 107 |
+
Bibliographic Explorer
|
| 108 |
+
(
|
| 109 |
+
What is the Explorer?
|
| 110 |
+
)
|
| 111 |
+
Connected Papers Toggle
|
| 112 |
+
Connected Papers
|
| 113 |
+
(
|
| 114 |
+
What is Connected Papers?
|
| 115 |
+
)
|
| 116 |
+
Litmaps Toggle
|
| 117 |
+
Litmaps
|
| 118 |
+
(
|
| 119 |
+
What is Litmaps?
|
| 120 |
+
)
|
| 121 |
+
scite.ai Toggle
|
| 122 |
+
scite Smart Citations
|
| 123 |
+
(
|
| 124 |
+
What are Smart Citations?
|
| 125 |
+
)
|
| 126 |
+
Code, Data, Media
|
| 127 |
+
Code, Data and Media Associated with this Article
|
| 128 |
+
alphaXiv Toggle
|
| 129 |
+
alphaXiv
|
| 130 |
+
(
|
| 131 |
+
What is alphaXiv?
|
| 132 |
+
)
|
| 133 |
+
Links to Code Toggle
|
| 134 |
+
CatalyzeX Code Finder for Papers
|
| 135 |
+
(
|
| 136 |
+
What is CatalyzeX?
|
| 137 |
+
)
|
| 138 |
+
DagsHub Toggle
|
| 139 |
+
DagsHub
|
| 140 |
+
(
|
| 141 |
+
What is DagsHub?
|
| 142 |
+
)
|
| 143 |
+
GotitPub Toggle
|
| 144 |
+
Gotit.pub
|
| 145 |
+
(
|
| 146 |
+
What is GotitPub?
|
| 147 |
+
)
|
| 148 |
+
Huggingface Toggle
|
| 149 |
+
Hugging Face
|
| 150 |
+
(
|
| 151 |
+
What is Huggingface?
|
| 152 |
+
)
|
| 153 |
+
ScienceCast Toggle
|
| 154 |
+
ScienceCast
|
| 155 |
+
(
|
| 156 |
+
What is ScienceCast?
|
| 157 |
+
)
|
| 158 |
+
Demos
|
| 159 |
+
Demos
|
| 160 |
+
Replicate Toggle
|
| 161 |
+
Replicate
|
| 162 |
+
(
|
| 163 |
+
What is Replicate?
|
| 164 |
+
)
|
| 165 |
+
Spaces Toggle
|
| 166 |
+
Hugging Face Spaces
|
| 167 |
+
(
|
| 168 |
+
What is Spaces?
|
| 169 |
+
)
|
| 170 |
+
Spaces Toggle
|
| 171 |
+
TXYZ.AI
|
| 172 |
+
(
|
| 173 |
+
What is TXYZ.AI?
|
| 174 |
+
)
|
| 175 |
+
Related Papers
|
| 176 |
+
Recommenders and Search Tools
|
| 177 |
+
Link to Influence Flower
|
| 178 |
+
Influence Flower
|
| 179 |
+
(
|
| 180 |
+
What are Influence Flowers?
|
| 181 |
+
)
|
| 182 |
+
Core recommender toggle
|
| 183 |
+
CORE Recommender
|
| 184 |
+
(
|
| 185 |
+
What is CORE?
|
| 186 |
+
)
|
| 187 |
+
IArxiv recommender toggle
|
| 188 |
+
IArxiv Recommender
|
| 189 |
+
(
|
| 190 |
+
What is IArxiv?
|
| 191 |
+
)
|
| 192 |
+
Author
|
| 193 |
+
Venue
|
| 194 |
+
Institution
|
| 195 |
+
Topic
|
| 196 |
+
About arXivLabs
|
| 197 |
+
arXivLabs: experimental projects with community collaborators
|
| 198 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 199 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 200 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 201 |
+
Learn more about arXivLabs
|
| 202 |
+
.
|
| 203 |
+
Which authors of this paper are endorsers?
|
| 204 |
+
|
|
| 205 |
+
Disable MathJax
|
| 206 |
+
(
|
| 207 |
+
What is MathJax?
|
| 208 |
+
)
|
|
@@ -0,0 +1,204 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2311.08516] LLMs cannot find reasoning errors, but can correct them given
|
| 3 |
+
the error location'
|
| 4 |
+
id: 231108516-llms-cannot-find-reasoning-errors-but-can-correct-them-given-the-error
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:40:16.980357Z'
|
| 8 |
+
source: https://arxiv.org/abs/2311.08516
|
| 9 |
+
source_domain: arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:40:16.980220Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2311.08516] LLMs cannot find reasoning errors, but can correct them given the error location
|
| 20 |
+
Computer Science > Artificial Intelligence
|
| 21 |
+
arXiv:2311.08516
|
| 22 |
+
(cs)
|
| 23 |
+
[Submitted on 14 Nov 2023 (
|
| 24 |
+
v1
|
| 25 |
+
), last revised 4 Jun 2024 (this version, v3)]
|
| 26 |
+
Title:
|
| 27 |
+
LLMs cannot find reasoning errors, but can correct them given the error location
|
| 28 |
+
Authors:
|
| 29 |
+
Gladys Tyen
|
| 30 |
+
,
|
| 31 |
+
Hassan Mansoor
|
| 32 |
+
,
|
| 33 |
+
Victor Cărbune
|
| 34 |
+
,
|
| 35 |
+
Peter Chen
|
| 36 |
+
,
|
| 37 |
+
Tony Mak
|
| 38 |
+
View a PDF of the paper titled LLMs cannot find reasoning errors, but can correct them given the error location, by Gladys Tyen and 4 other authors
|
| 39 |
+
View PDF
|
| 40 |
+
HTML (experimental)
|
| 41 |
+
Abstract:
|
| 42 |
+
While self-correction has shown promise in improving LLM outputs in terms of style and quality (e.g. Chen et al., 2023b; Madaan et al., 2023), recent attempts to self-correct logical or reasoning errors often cause correct answers to become incorrect, resulting in worse performances overall (Huang et al., 2023). In this paper, we show that poor self-correction performance stems from LLMs' inability to find logical mistakes, rather than their ability to correct a known mistake. Firstly, we benchmark several state-of-the-art LLMs on their mistake-finding ability and demonstrate that they generally struggle with the task, even in highly objective, unambiguous cases. Secondly, we test the correction abilities of LLMs -- separately from mistake finding -- using a backtracking setup that feeds ground truth mistake location information to the model. We show that this boosts downstream task performance across our 5 reasoning tasks, indicating that LLMs' correction abilities are robust. Finally, we show that it is possible to obtain mistake location information without ground truth labels or in-domain training data. We train a small classifier with out-of-domain data, which exhibits stronger mistake-finding performance than prompting a large model. We release our dataset of LLM-generated logical mistakes, BIG-Bench Mistake, to enable further research into locating LLM reasoning mistakes.
|
| 43 |
+
Comments:
|
| 44 |
+
ACL 2024 Findings
|
| 45 |
+
Subjects:
|
| 46 |
+
Artificial Intelligence (cs.AI)
|
| 47 |
+
; Computation and Language (cs.CL); Machine Learning (cs.LG)
|
| 48 |
+
Cite as:
|
| 49 |
+
arXiv:2311.08516
|
| 50 |
+
[cs.AI]
|
| 51 |
+
(or
|
| 52 |
+
arXiv:2311.08516v3
|
| 53 |
+
[cs.AI]
|
| 54 |
+
for this version)
|
| 55 |
+
https://doi.org/10.48550/arXiv.2311.08516
|
| 56 |
+
Focus to learn more
|
| 57 |
+
arXiv-issued DOI via DataCite
|
| 58 |
+
Submission history
|
| 59 |
+
From: Gladys Tyen [
|
| 60 |
+
view email
|
| 61 |
+
]
|
| 62 |
+
[v1]
|
| 63 |
+
Tue, 14 Nov 2023 20:12:38 UTC (7,191 KB)
|
| 64 |
+
[v2]
|
| 65 |
+
Tue, 9 Jan 2024 03:32:32 UTC (7,191 KB)
|
| 66 |
+
[v3]
|
| 67 |
+
Tue, 4 Jun 2024 10:25:13 UTC (7,319 KB)
|
| 68 |
+
Full-text links:
|
| 69 |
+
Access Paper:
|
| 70 |
+
View a PDF of the paper titled LLMs cannot find reasoning errors, but can correct them given the error location, by Gladys Tyen and 4 other authors
|
| 71 |
+
View PDF
|
| 72 |
+
HTML (experimental)
|
| 73 |
+
TeX Source
|
| 74 |
+
view license
|
| 75 |
+
Current browse context:
|
| 76 |
+
cs.AI
|
| 77 |
+
< prev
|
| 78 |
+
|
|
| 79 |
+
next >
|
| 80 |
+
new
|
| 81 |
+
|
|
| 82 |
+
recent
|
| 83 |
+
|
|
| 84 |
+
2023-11
|
| 85 |
+
Change to browse by:
|
| 86 |
+
cs
|
| 87 |
+
cs.CL
|
| 88 |
+
cs.LG
|
| 89 |
+
References & Citations
|
| 90 |
+
NASA ADS
|
| 91 |
+
Google Scholar
|
| 92 |
+
Semantic Scholar
|
| 93 |
+
export BibTeX citation
|
| 94 |
+
Loading...
|
| 95 |
+
BibTeX formatted citation
|
| 96 |
+
×
|
| 97 |
+
loading...
|
| 98 |
+
Data provided by:
|
| 99 |
+
Bookmark
|
| 100 |
+
Bibliographic Tools
|
| 101 |
+
Bibliographic and Citation Tools
|
| 102 |
+
Bibliographic Explorer Toggle
|
| 103 |
+
Bibliographic Explorer
|
| 104 |
+
(
|
| 105 |
+
What is the Explorer?
|
| 106 |
+
)
|
| 107 |
+
Connected Papers Toggle
|
| 108 |
+
Connected Papers
|
| 109 |
+
(
|
| 110 |
+
What is Connected Papers?
|
| 111 |
+
)
|
| 112 |
+
Litmaps Toggle
|
| 113 |
+
Litmaps
|
| 114 |
+
(
|
| 115 |
+
What is Litmaps?
|
| 116 |
+
)
|
| 117 |
+
scite.ai Toggle
|
| 118 |
+
scite Smart Citations
|
| 119 |
+
(
|
| 120 |
+
What are Smart Citations?
|
| 121 |
+
)
|
| 122 |
+
Code, Data, Media
|
| 123 |
+
Code, Data and Media Associated with this Article
|
| 124 |
+
alphaXiv Toggle
|
| 125 |
+
alphaXiv
|
| 126 |
+
(
|
| 127 |
+
What is alphaXiv?
|
| 128 |
+
)
|
| 129 |
+
Links to Code Toggle
|
| 130 |
+
CatalyzeX Code Finder for Papers
|
| 131 |
+
(
|
| 132 |
+
What is CatalyzeX?
|
| 133 |
+
)
|
| 134 |
+
DagsHub Toggle
|
| 135 |
+
DagsHub
|
| 136 |
+
(
|
| 137 |
+
What is DagsHub?
|
| 138 |
+
)
|
| 139 |
+
GotitPub Toggle
|
| 140 |
+
Gotit.pub
|
| 141 |
+
(
|
| 142 |
+
What is GotitPub?
|
| 143 |
+
)
|
| 144 |
+
Huggingface Toggle
|
| 145 |
+
Hugging Face
|
| 146 |
+
(
|
| 147 |
+
What is Huggingface?
|
| 148 |
+
)
|
| 149 |
+
Links to Code Toggle
|
| 150 |
+
Papers with Code
|
| 151 |
+
(
|
| 152 |
+
What is Papers with Code?
|
| 153 |
+
)
|
| 154 |
+
ScienceCast Toggle
|
| 155 |
+
ScienceCast
|
| 156 |
+
(
|
| 157 |
+
What is ScienceCast?
|
| 158 |
+
)
|
| 159 |
+
Demos
|
| 160 |
+
Demos
|
| 161 |
+
Replicate Toggle
|
| 162 |
+
Replicate
|
| 163 |
+
(
|
| 164 |
+
What is Replicate?
|
| 165 |
+
)
|
| 166 |
+
Spaces Toggle
|
| 167 |
+
Hugging Face Spaces
|
| 168 |
+
(
|
| 169 |
+
What is Spaces?
|
| 170 |
+
)
|
| 171 |
+
Spaces Toggle
|
| 172 |
+
TXYZ.AI
|
| 173 |
+
(
|
| 174 |
+
What is TXYZ.AI?
|
| 175 |
+
)
|
| 176 |
+
Related Papers
|
| 177 |
+
Recommenders and Search Tools
|
| 178 |
+
Link to Influence Flower
|
| 179 |
+
Influence Flower
|
| 180 |
+
(
|
| 181 |
+
What are Influence Flowers?
|
| 182 |
+
)
|
| 183 |
+
Core recommender toggle
|
| 184 |
+
CORE Recommender
|
| 185 |
+
(
|
| 186 |
+
What is CORE?
|
| 187 |
+
)
|
| 188 |
+
Author
|
| 189 |
+
Venue
|
| 190 |
+
Institution
|
| 191 |
+
Topic
|
| 192 |
+
About arXivLabs
|
| 193 |
+
arXivLabs: experimental projects with community collaborators
|
| 194 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 195 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 196 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 197 |
+
Learn more about arXivLabs
|
| 198 |
+
.
|
| 199 |
+
Which authors of this paper are endorsers?
|
| 200 |
+
|
|
| 201 |
+
Disable MathJax
|
| 202 |
+
(
|
| 203 |
+
What is MathJax?
|
| 204 |
+
)
|
|
@@ -0,0 +1,203 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2312.09152] Evaluating Augmented Reality Communication: How Can We Teach
|
| 3 |
+
Procedural Skill in AR?'
|
| 4 |
+
id: 231209152-evaluating-augmented-reality-communication-how-can-we-teach-procedural
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:40:10.692238Z'
|
| 8 |
+
source: https://arxiv.org/abs/2312.09152
|
| 9 |
+
source_domain: arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:40:10.692096Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2312.09152] Evaluating Augmented Reality Communication: How Can We Teach Procedural Skill in AR?
|
| 20 |
+
Computer Science > Human-Computer Interaction
|
| 21 |
+
arXiv:2312.09152
|
| 22 |
+
(cs)
|
| 23 |
+
[Submitted on 14 Dec 2023]
|
| 24 |
+
Title:
|
| 25 |
+
Evaluating Augmented Reality Communication: How Can We Teach Procedural Skill in AR?
|
| 26 |
+
Authors:
|
| 27 |
+
Manuel Rebol
|
| 28 |
+
,
|
| 29 |
+
Krzysztof Pietroszek
|
| 30 |
+
,
|
| 31 |
+
Neal Sikka
|
| 32 |
+
,
|
| 33 |
+
Claudia Ranniger
|
| 34 |
+
,
|
| 35 |
+
Colton Hood
|
| 36 |
+
,
|
| 37 |
+
Adam Rutenberg
|
| 38 |
+
,
|
| 39 |
+
Puja Sasankan
|
| 40 |
+
,
|
| 41 |
+
Christian Gütl
|
| 42 |
+
View a PDF of the paper titled Evaluating Augmented Reality Communication: How Can We Teach Procedural Skill in AR?, by Manuel Rebol and 7 other authors
|
| 43 |
+
View PDF
|
| 44 |
+
HTML (experimental)
|
| 45 |
+
Abstract:
|
| 46 |
+
Augmented reality (AR) has great potential for use in healthcare applications, especially remote medical training and supervision. In this paper, we analyze the usage of an AR communication system to teach a medical procedure, the placement of a central venous catheter (CVC) under ultrasound guidance. We examine various AR communication and collaboration components, including gestural communication, volumetric information, annotations, augmented objects, and augmented screens. We compare how teaching in AR differs from teaching through videoconferencing-based communication. Our results include a detailed medical training steps analysis in which we compare how verbal and visual communication differs between video and AR training. We identify procedural steps in which medical experts give visual instructions utilizing AR components. We examine the change in AR usage and interaction over time and recognize patterns between users. Moreover, AR design recommendations are given based on post-training interviews.
|
| 47 |
+
Comments:
|
| 48 |
+
this https URL
|
| 49 |
+
Subjects:
|
| 50 |
+
Human-Computer Interaction (cs.HC)
|
| 51 |
+
Cite as:
|
| 52 |
+
arXiv:2312.09152
|
| 53 |
+
[cs.HC]
|
| 54 |
+
(or
|
| 55 |
+
arXiv:2312.09152v1
|
| 56 |
+
[cs.HC]
|
| 57 |
+
for this version)
|
| 58 |
+
https://doi.org/10.48550/arXiv.2312.09152
|
| 59 |
+
Focus to learn more
|
| 60 |
+
arXiv-issued DOI via DataCite
|
| 61 |
+
Journal reference:
|
| 62 |
+
Proceedings of the 29th ACM Symposium on Virtual Reality Software and Technology (VRST 2023)
|
| 63 |
+
Related DOI
|
| 64 |
+
:
|
| 65 |
+
https://doi.org/10.1145/3611659.3615685
|
| 66 |
+
Focus to learn more
|
| 67 |
+
DOI(s) linking to related resources
|
| 68 |
+
Submission history
|
| 69 |
+
From: Manuel Rebol [
|
| 70 |
+
view email
|
| 71 |
+
]
|
| 72 |
+
[v1]
|
| 73 |
+
Thu, 14 Dec 2023 17:22:22 UTC (2,671 KB)
|
| 74 |
+
Full-text links:
|
| 75 |
+
Access Paper:
|
| 76 |
+
View a PDF of the paper titled Evaluating Augmented Reality Communication: How Can We Teach Procedural Skill in AR?, by Manuel Rebol and 7 other authors
|
| 77 |
+
View PDF
|
| 78 |
+
HTML (experimental)
|
| 79 |
+
TeX Source
|
| 80 |
+
view license
|
| 81 |
+
Current browse context:
|
| 82 |
+
cs.HC
|
| 83 |
+
< prev
|
| 84 |
+
|
|
| 85 |
+
next >
|
| 86 |
+
new
|
| 87 |
+
|
|
| 88 |
+
recent
|
| 89 |
+
|
|
| 90 |
+
2023-12
|
| 91 |
+
Change to browse by:
|
| 92 |
+
cs
|
| 93 |
+
References & Citations
|
| 94 |
+
NASA ADS
|
| 95 |
+
Google Scholar
|
| 96 |
+
Semantic Scholar
|
| 97 |
+
export BibTeX citation
|
| 98 |
+
Loading...
|
| 99 |
+
BibTeX formatted citation
|
| 100 |
+
×
|
| 101 |
+
loading...
|
| 102 |
+
Data provided by:
|
| 103 |
+
Bookmark
|
| 104 |
+
Bibliographic Tools
|
| 105 |
+
Bibliographic and Citation Tools
|
| 106 |
+
Bibliographic Explorer Toggle
|
| 107 |
+
Bibliographic Explorer
|
| 108 |
+
(
|
| 109 |
+
What is the Explorer?
|
| 110 |
+
)
|
| 111 |
+
Connected Papers Toggle
|
| 112 |
+
Connected Papers
|
| 113 |
+
(
|
| 114 |
+
What is Connected Papers?
|
| 115 |
+
)
|
| 116 |
+
Litmaps Toggle
|
| 117 |
+
Litmaps
|
| 118 |
+
(
|
| 119 |
+
What is Litmaps?
|
| 120 |
+
)
|
| 121 |
+
scite.ai Toggle
|
| 122 |
+
scite Smart Citations
|
| 123 |
+
(
|
| 124 |
+
What are Smart Citations?
|
| 125 |
+
)
|
| 126 |
+
Code, Data, Media
|
| 127 |
+
Code, Data and Media Associated with this Article
|
| 128 |
+
alphaXiv Toggle
|
| 129 |
+
alphaXiv
|
| 130 |
+
(
|
| 131 |
+
What is alphaXiv?
|
| 132 |
+
)
|
| 133 |
+
Links to Code Toggle
|
| 134 |
+
CatalyzeX Code Finder for Papers
|
| 135 |
+
(
|
| 136 |
+
What is CatalyzeX?
|
| 137 |
+
)
|
| 138 |
+
DagsHub Toggle
|
| 139 |
+
DagsHub
|
| 140 |
+
(
|
| 141 |
+
What is DagsHub?
|
| 142 |
+
)
|
| 143 |
+
GotitPub Toggle
|
| 144 |
+
Gotit.pub
|
| 145 |
+
(
|
| 146 |
+
What is GotitPub?
|
| 147 |
+
)
|
| 148 |
+
Huggingface Toggle
|
| 149 |
+
Hugging Face
|
| 150 |
+
(
|
| 151 |
+
What is Huggingface?
|
| 152 |
+
)
|
| 153 |
+
ScienceCast Toggle
|
| 154 |
+
ScienceCast
|
| 155 |
+
(
|
| 156 |
+
What is ScienceCast?
|
| 157 |
+
)
|
| 158 |
+
Demos
|
| 159 |
+
Demos
|
| 160 |
+
Replicate Toggle
|
| 161 |
+
Replicate
|
| 162 |
+
(
|
| 163 |
+
What is Replicate?
|
| 164 |
+
)
|
| 165 |
+
Spaces Toggle
|
| 166 |
+
Hugging Face Spaces
|
| 167 |
+
(
|
| 168 |
+
What is Spaces?
|
| 169 |
+
)
|
| 170 |
+
Spaces Toggle
|
| 171 |
+
TXYZ.AI
|
| 172 |
+
(
|
| 173 |
+
What is TXYZ.AI?
|
| 174 |
+
)
|
| 175 |
+
Related Papers
|
| 176 |
+
Recommenders and Search Tools
|
| 177 |
+
Link to Influence Flower
|
| 178 |
+
Influence Flower
|
| 179 |
+
(
|
| 180 |
+
What are Influence Flowers?
|
| 181 |
+
)
|
| 182 |
+
Core recommender toggle
|
| 183 |
+
CORE Recommender
|
| 184 |
+
(
|
| 185 |
+
What is CORE?
|
| 186 |
+
)
|
| 187 |
+
Author
|
| 188 |
+
Venue
|
| 189 |
+
Institution
|
| 190 |
+
Topic
|
| 191 |
+
About arXivLabs
|
| 192 |
+
arXivLabs: experimental projects with community collaborators
|
| 193 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 194 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 195 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 196 |
+
Learn more about arXivLabs
|
| 197 |
+
.
|
| 198 |
+
Which authors of this paper are endorsers?
|
| 199 |
+
|
|
| 200 |
+
Disable MathJax
|
| 201 |
+
(
|
| 202 |
+
What is MathJax?
|
| 203 |
+
)
|
|
@@ -0,0 +1,203 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2402.01817] LLMs Can''t Plan, But Can Help Planning in LLM-Modulo Frameworks'
|
| 3 |
+
id: 240201817-llms-cant-plan-but-can-help-planning-in-llm-modulo-frameworks
|
| 4 |
+
tags:
|
| 5 |
+
- deepread
|
| 6 |
+
created: '2026-06-10T00:40:16.134883Z'
|
| 7 |
+
source: https://arxiv.org/abs/2402.01817
|
| 8 |
+
source_domain: arxiv.org
|
| 9 |
+
fetched_at: '2026-06-10T00:40:16.134739Z'
|
| 10 |
+
fetch_provider: builtin
|
| 11 |
+
status: draft
|
| 12 |
+
type: note
|
| 13 |
+
tier: institutional
|
| 14 |
+
content_type: paper
|
| 15 |
+
deprecated: false
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
[2402.01817] LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
|
| 19 |
+
Computer Science > Artificial Intelligence
|
| 20 |
+
arXiv:2402.01817
|
| 21 |
+
(cs)
|
| 22 |
+
[Submitted on 2 Feb 2024 (
|
| 23 |
+
v1
|
| 24 |
+
), last revised 12 Jun 2024 (this version, v3)]
|
| 25 |
+
Title:
|
| 26 |
+
LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
|
| 27 |
+
Authors:
|
| 28 |
+
Subbarao Kambhampati
|
| 29 |
+
,
|
| 30 |
+
Karthik Valmeekam
|
| 31 |
+
,
|
| 32 |
+
Lin Guan
|
| 33 |
+
,
|
| 34 |
+
Mudit Verma
|
| 35 |
+
,
|
| 36 |
+
Kaya Stechly
|
| 37 |
+
,
|
| 38 |
+
Siddhant Bhambri
|
| 39 |
+
,
|
| 40 |
+
Lucas Saldyt
|
| 41 |
+
,
|
| 42 |
+
Anil Murthy
|
| 43 |
+
View a PDF of the paper titled LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks, by Subbarao Kambhampati and 7 other authors
|
| 44 |
+
View PDF
|
| 45 |
+
HTML (experimental)
|
| 46 |
+
Abstract:
|
| 47 |
+
There is considerable confusion about the role of Large Language Models (LLMs) in planning and reasoning tasks. On one side are over-optimistic claims that LLMs can indeed do these tasks with just the right prompting or self-verification strategies. On the other side are perhaps over-pessimistic claims that all that LLMs are good for in planning/reasoning tasks are as mere translators of the problem specification from one syntactic format to another, and ship the problem off to external symbolic solvers. In this position paper, we take the view that both these extremes are misguided. We argue that auto-regressive LLMs cannot, by themselves, do planning or self-verification (which is after all a form of reasoning), and shed some light on the reasons for misunderstandings in the literature. We will also argue that LLMs should be viewed as universal approximate knowledge sources that have much more meaningful roles to play in planning/reasoning tasks beyond simple front-end/back-end format translators. We present a vision of {\bf LLM-Modulo Frameworks} that combine the strengths of LLMs with external model-based verifiers in a tighter bi-directional interaction regime. We will show how the models driving the external verifiers themselves can be acquired with the help of LLMs. We will also argue that rather than simply pipelining LLMs and symbolic components, this LLM-Modulo Framework provides a better neuro-symbolic approach that offers tighter integration between LLMs and symbolic components, and allows extending the scope of model-based planning/reasoning regimes towards more flexible knowledge, problem and preference specifications.
|
| 48 |
+
Subjects:
|
| 49 |
+
Artificial Intelligence (cs.AI)
|
| 50 |
+
; Machine Learning (cs.LG)
|
| 51 |
+
Cite as:
|
| 52 |
+
arXiv:2402.01817
|
| 53 |
+
[cs.AI]
|
| 54 |
+
(or
|
| 55 |
+
arXiv:2402.01817v3
|
| 56 |
+
[cs.AI]
|
| 57 |
+
for this version)
|
| 58 |
+
https://doi.org/10.48550/arXiv.2402.01817
|
| 59 |
+
Focus to learn more
|
| 60 |
+
arXiv-issued DOI via DataCite
|
| 61 |
+
Journal reference:
|
| 62 |
+
Proceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024
|
| 63 |
+
Submission history
|
| 64 |
+
From: Subbarao Kambhampati [
|
| 65 |
+
view email
|
| 66 |
+
]
|
| 67 |
+
[v1]
|
| 68 |
+
Fri, 2 Feb 2024 14:43:18 UTC (4,551 KB)
|
| 69 |
+
[v2]
|
| 70 |
+
Tue, 6 Feb 2024 01:29:37 UTC (4,552 KB)
|
| 71 |
+
[v3]
|
| 72 |
+
Wed, 12 Jun 2024 01:13:11 UTC (6,405 KB)
|
| 73 |
+
Full-text links:
|
| 74 |
+
Access Paper:
|
| 75 |
+
View a PDF of the paper titled LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks, by Subbarao Kambhampati and 7 other authors
|
| 76 |
+
View PDF
|
| 77 |
+
HTML (experimental)
|
| 78 |
+
TeX Source
|
| 79 |
+
view license
|
| 80 |
+
Current browse context:
|
| 81 |
+
cs.AI
|
| 82 |
+
< prev
|
| 83 |
+
|
|
| 84 |
+
next >
|
| 85 |
+
new
|
| 86 |
+
|
|
| 87 |
+
recent
|
| 88 |
+
|
|
| 89 |
+
2024-02
|
| 90 |
+
Change to browse by:
|
| 91 |
+
cs
|
| 92 |
+
cs.LG
|
| 93 |
+
References & Citations
|
| 94 |
+
NASA ADS
|
| 95 |
+
Google Scholar
|
| 96 |
+
Semantic Scholar
|
| 97 |
+
export BibTeX citation
|
| 98 |
+
Loading...
|
| 99 |
+
BibTeX formatted citation
|
| 100 |
+
×
|
| 101 |
+
loading...
|
| 102 |
+
Data provided by:
|
| 103 |
+
Bookmark
|
| 104 |
+
Bibliographic Tools
|
| 105 |
+
Bibliographic and Citation Tools
|
| 106 |
+
Bibliographic Explorer Toggle
|
| 107 |
+
Bibliographic Explorer
|
| 108 |
+
(
|
| 109 |
+
What is the Explorer?
|
| 110 |
+
)
|
| 111 |
+
Connected Papers Toggle
|
| 112 |
+
Connected Papers
|
| 113 |
+
(
|
| 114 |
+
What is Connected Papers?
|
| 115 |
+
)
|
| 116 |
+
Litmaps Toggle
|
| 117 |
+
Litmaps
|
| 118 |
+
(
|
| 119 |
+
What is Litmaps?
|
| 120 |
+
)
|
| 121 |
+
scite.ai Toggle
|
| 122 |
+
scite Smart Citations
|
| 123 |
+
(
|
| 124 |
+
What are Smart Citations?
|
| 125 |
+
)
|
| 126 |
+
Code, Data, Media
|
| 127 |
+
Code, Data and Media Associated with this Article
|
| 128 |
+
alphaXiv Toggle
|
| 129 |
+
alphaXiv
|
| 130 |
+
(
|
| 131 |
+
What is alphaXiv?
|
| 132 |
+
)
|
| 133 |
+
Links to Code Toggle
|
| 134 |
+
CatalyzeX Code Finder for Papers
|
| 135 |
+
(
|
| 136 |
+
What is CatalyzeX?
|
| 137 |
+
)
|
| 138 |
+
DagsHub Toggle
|
| 139 |
+
DagsHub
|
| 140 |
+
(
|
| 141 |
+
What is DagsHub?
|
| 142 |
+
)
|
| 143 |
+
GotitPub Toggle
|
| 144 |
+
Gotit.pub
|
| 145 |
+
(
|
| 146 |
+
What is GotitPub?
|
| 147 |
+
)
|
| 148 |
+
Huggingface Toggle
|
| 149 |
+
Hugging Face
|
| 150 |
+
(
|
| 151 |
+
What is Huggingface?
|
| 152 |
+
)
|
| 153 |
+
ScienceCast Toggle
|
| 154 |
+
ScienceCast
|
| 155 |
+
(
|
| 156 |
+
What is ScienceCast?
|
| 157 |
+
)
|
| 158 |
+
Demos
|
| 159 |
+
Demos
|
| 160 |
+
Replicate Toggle
|
| 161 |
+
Replicate
|
| 162 |
+
(
|
| 163 |
+
What is Replicate?
|
| 164 |
+
)
|
| 165 |
+
Spaces Toggle
|
| 166 |
+
Hugging Face Spaces
|
| 167 |
+
(
|
| 168 |
+
What is Spaces?
|
| 169 |
+
)
|
| 170 |
+
Spaces Toggle
|
| 171 |
+
TXYZ.AI
|
| 172 |
+
(
|
| 173 |
+
What is TXYZ.AI?
|
| 174 |
+
)
|
| 175 |
+
Related Papers
|
| 176 |
+
Recommenders and Search Tools
|
| 177 |
+
Link to Influence Flower
|
| 178 |
+
Influence Flower
|
| 179 |
+
(
|
| 180 |
+
What are Influence Flowers?
|
| 181 |
+
)
|
| 182 |
+
Core recommender toggle
|
| 183 |
+
CORE Recommender
|
| 184 |
+
(
|
| 185 |
+
What is CORE?
|
| 186 |
+
)
|
| 187 |
+
Author
|
| 188 |
+
Venue
|
| 189 |
+
Institution
|
| 190 |
+
Topic
|
| 191 |
+
About arXivLabs
|
| 192 |
+
arXivLabs: experimental projects with community collaborators
|
| 193 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 194 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 195 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 196 |
+
Learn more about arXivLabs
|
| 197 |
+
.
|
| 198 |
+
Which authors of this paper are endorsers?
|
| 199 |
+
|
|
| 200 |
+
Disable MathJax
|
| 201 |
+
(
|
| 202 |
+
What is MathJax?
|
| 203 |
+
)
|
|
@@ -0,0 +1,214 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2402.03300] DeepSeekMath: Pushing the Limits of Mathematical Reasoning in
|
| 3 |
+
Open Language Models'
|
| 4 |
+
id: 240203300-deepseekmath-pushing-the-limits-of-mathematical-reasoning-in-open-lang
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-09T23:28:29.232007Z'
|
| 8 |
+
source: https://arxiv.org/abs/2402.03300
|
| 9 |
+
source_domain: arxiv.org
|
| 10 |
+
fetched_at: '2026-06-09T23:28:29.231807Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2402.03300] DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
|
| 20 |
+
Computer Science > Computation and Language
|
| 21 |
+
arXiv:2402.03300
|
| 22 |
+
(cs)
|
| 23 |
+
[Submitted on 5 Feb 2024 (
|
| 24 |
+
v1
|
| 25 |
+
), last revised 27 Apr 2024 (this version, v3)]
|
| 26 |
+
Title:
|
| 27 |
+
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
|
| 28 |
+
Authors:
|
| 29 |
+
Zhihong Shao
|
| 30 |
+
,
|
| 31 |
+
Peiyi Wang
|
| 32 |
+
,
|
| 33 |
+
Qihao Zhu
|
| 34 |
+
,
|
| 35 |
+
Runxin Xu
|
| 36 |
+
,
|
| 37 |
+
Junxiao Song
|
| 38 |
+
,
|
| 39 |
+
Xiao Bi
|
| 40 |
+
,
|
| 41 |
+
Haowei Zhang
|
| 42 |
+
,
|
| 43 |
+
Mingchuan Zhang
|
| 44 |
+
,
|
| 45 |
+
Y.K. Li
|
| 46 |
+
,
|
| 47 |
+
Y. Wu
|
| 48 |
+
,
|
| 49 |
+
Daya Guo
|
| 50 |
+
View a PDF of the paper titled DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, by Zhihong Shao and 10 other authors
|
| 51 |
+
View PDF
|
| 52 |
+
HTML (experimental)
|
| 53 |
+
Abstract:
|
| 54 |
+
Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pre-training DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens sourced from Common Crawl, together with natural language and code data. DeepSeekMath 7B has achieved an impressive score of 51.7% on the competition-level MATH benchmark without relying on external toolkits and voting techniques, approaching the performance level of Gemini-Ultra and GPT-4. Self-consistency over 64 samples from DeepSeekMath 7B achieves 60.9% on MATH. The mathematical reasoning capability of DeepSeekMath is attributed to two key factors: First, we harness the significant potential of publicly available web data through a meticulously engineered data selection pipeline. Second, we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.
|
| 55 |
+
Subjects:
|
| 56 |
+
Computation and Language (cs.CL)
|
| 57 |
+
; Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
|
| 58 |
+
Cite as:
|
| 59 |
+
arXiv:2402.03300
|
| 60 |
+
[cs.CL]
|
| 61 |
+
(or
|
| 62 |
+
arXiv:2402.03300v3
|
| 63 |
+
[cs.CL]
|
| 64 |
+
for this version)
|
| 65 |
+
https://doi.org/10.48550/arXiv.2402.03300
|
| 66 |
+
Focus to learn more
|
| 67 |
+
arXiv-issued DOI via DataCite
|
| 68 |
+
Submission history
|
| 69 |
+
From: Zhihong Shao [
|
| 70 |
+
view email
|
| 71 |
+
]
|
| 72 |
+
[v1]
|
| 73 |
+
Mon, 5 Feb 2024 18:55:32 UTC (3,417 KB)
|
| 74 |
+
[v2]
|
| 75 |
+
Tue, 6 Feb 2024 18:39:38 UTC (3,417 KB)
|
| 76 |
+
[v3]
|
| 77 |
+
Sat, 27 Apr 2024 15:25:53 UTC (3,417 KB)
|
| 78 |
+
Full-text links:
|
| 79 |
+
Access Paper:
|
| 80 |
+
View a PDF of the paper titled DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, by Zhihong Shao and 10 other authors
|
| 81 |
+
View PDF
|
| 82 |
+
HTML (experimental)
|
| 83 |
+
TeX Source
|
| 84 |
+
view license
|
| 85 |
+
Current browse context:
|
| 86 |
+
cs.CL
|
| 87 |
+
< prev
|
| 88 |
+
|
|
| 89 |
+
next >
|
| 90 |
+
new
|
| 91 |
+
|
|
| 92 |
+
recent
|
| 93 |
+
|
|
| 94 |
+
2024-02
|
| 95 |
+
Change to browse by:
|
| 96 |
+
cs
|
| 97 |
+
cs.AI
|
| 98 |
+
cs.LG
|
| 99 |
+
References & Citations
|
| 100 |
+
NASA ADS
|
| 101 |
+
Google Scholar
|
| 102 |
+
Semantic Scholar
|
| 103 |
+
export BibTeX citation
|
| 104 |
+
Loading...
|
| 105 |
+
BibTeX formatted citation
|
| 106 |
+
×
|
| 107 |
+
loading...
|
| 108 |
+
Data provided by:
|
| 109 |
+
Bookmark
|
| 110 |
+
Bibliographic Tools
|
| 111 |
+
Bibliographic and Citation Tools
|
| 112 |
+
Bibliographic Explorer Toggle
|
| 113 |
+
Bibliographic Explorer
|
| 114 |
+
(
|
| 115 |
+
What is the Explorer?
|
| 116 |
+
)
|
| 117 |
+
Connected Papers Toggle
|
| 118 |
+
Connected Papers
|
| 119 |
+
(
|
| 120 |
+
What is Connected Papers?
|
| 121 |
+
)
|
| 122 |
+
Litmaps Toggle
|
| 123 |
+
Litmaps
|
| 124 |
+
(
|
| 125 |
+
What is Litmaps?
|
| 126 |
+
)
|
| 127 |
+
scite.ai Toggle
|
| 128 |
+
scite Smart Citations
|
| 129 |
+
(
|
| 130 |
+
What are Smart Citations?
|
| 131 |
+
)
|
| 132 |
+
Code, Data, Media
|
| 133 |
+
Code, Data and Media Associated with this Article
|
| 134 |
+
alphaXiv Toggle
|
| 135 |
+
alphaXiv
|
| 136 |
+
(
|
| 137 |
+
What is alphaXiv?
|
| 138 |
+
)
|
| 139 |
+
Links to Code Toggle
|
| 140 |
+
CatalyzeX Code Finder for Papers
|
| 141 |
+
(
|
| 142 |
+
What is CatalyzeX?
|
| 143 |
+
)
|
| 144 |
+
DagsHub Toggle
|
| 145 |
+
DagsHub
|
| 146 |
+
(
|
| 147 |
+
What is DagsHub?
|
| 148 |
+
)
|
| 149 |
+
GotitPub Toggle
|
| 150 |
+
Gotit.pub
|
| 151 |
+
(
|
| 152 |
+
What is GotitPub?
|
| 153 |
+
)
|
| 154 |
+
Huggingface Toggle
|
| 155 |
+
Hugging Face
|
| 156 |
+
(
|
| 157 |
+
What is Huggingface?
|
| 158 |
+
)
|
| 159 |
+
Links to Code Toggle
|
| 160 |
+
Papers with Code
|
| 161 |
+
(
|
| 162 |
+
What is Papers with Code?
|
| 163 |
+
)
|
| 164 |
+
ScienceCast Toggle
|
| 165 |
+
ScienceCast
|
| 166 |
+
(
|
| 167 |
+
What is ScienceCast?
|
| 168 |
+
)
|
| 169 |
+
Demos
|
| 170 |
+
Demos
|
| 171 |
+
Replicate Toggle
|
| 172 |
+
Replicate
|
| 173 |
+
(
|
| 174 |
+
What is Replicate?
|
| 175 |
+
)
|
| 176 |
+
Spaces Toggle
|
| 177 |
+
Hugging Face Spaces
|
| 178 |
+
(
|
| 179 |
+
What is Spaces?
|
| 180 |
+
)
|
| 181 |
+
Spaces Toggle
|
| 182 |
+
TXYZ.AI
|
| 183 |
+
(
|
| 184 |
+
What is TXYZ.AI?
|
| 185 |
+
)
|
| 186 |
+
Related Papers
|
| 187 |
+
Recommenders and Search Tools
|
| 188 |
+
Link to Influence Flower
|
| 189 |
+
Influence Flower
|
| 190 |
+
(
|
| 191 |
+
What are Influence Flowers?
|
| 192 |
+
)
|
| 193 |
+
Core recommender toggle
|
| 194 |
+
CORE Recommender
|
| 195 |
+
(
|
| 196 |
+
What is CORE?
|
| 197 |
+
)
|
| 198 |
+
Author
|
| 199 |
+
Venue
|
| 200 |
+
Institution
|
| 201 |
+
Topic
|
| 202 |
+
About arXivLabs
|
| 203 |
+
arXivLabs: experimental projects with community collaborators
|
| 204 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 205 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 206 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 207 |
+
Learn more about arXivLabs
|
| 208 |
+
.
|
| 209 |
+
Which authors of this paper are endorsers?
|
| 210 |
+
|
|
| 211 |
+
Disable MathJax
|
| 212 |
+
(
|
| 213 |
+
What is MathJax?
|
| 214 |
+
)
|
|
@@ -0,0 +1,228 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2404.11018] Many-Shot In-Context Learning'
|
| 3 |
+
id: 240411018-many-shot-in-context-learning
|
| 4 |
+
tags:
|
| 5 |
+
- deepread
|
| 6 |
+
created: '2026-06-10T00:40:15.011649Z'
|
| 7 |
+
source: https://arxiv.org/abs/2404.11018
|
| 8 |
+
source_domain: arxiv.org
|
| 9 |
+
fetched_at: '2026-06-10T00:40:15.011513Z'
|
| 10 |
+
fetch_provider: builtin
|
| 11 |
+
status: draft
|
| 12 |
+
type: note
|
| 13 |
+
tier: institutional
|
| 14 |
+
content_type: paper
|
| 15 |
+
deprecated: false
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
[2404.11018] Many-Shot In-Context Learning
|
| 19 |
+
Computer Science > Machine Learning
|
| 20 |
+
arXiv:2404.11018
|
| 21 |
+
(cs)
|
| 22 |
+
[Submitted on 17 Apr 2024 (
|
| 23 |
+
v1
|
| 24 |
+
), last revised 17 Oct 2024 (this version, v3)]
|
| 25 |
+
Title:
|
| 26 |
+
Many-Shot In-Context Learning
|
| 27 |
+
Authors:
|
| 28 |
+
Rishabh Agarwal
|
| 29 |
+
,
|
| 30 |
+
Avi Singh
|
| 31 |
+
,
|
| 32 |
+
Lei M. Zhang
|
| 33 |
+
,
|
| 34 |
+
Bernd Bohnet
|
| 35 |
+
,
|
| 36 |
+
Luis Rosias
|
| 37 |
+
,
|
| 38 |
+
Stephanie Chan
|
| 39 |
+
,
|
| 40 |
+
Biao Zhang
|
| 41 |
+
,
|
| 42 |
+
Ankesh Anand
|
| 43 |
+
,
|
| 44 |
+
Zaheer Abbas
|
| 45 |
+
,
|
| 46 |
+
Azade Nova
|
| 47 |
+
,
|
| 48 |
+
John D. Co-Reyes
|
| 49 |
+
,
|
| 50 |
+
Eric Chu
|
| 51 |
+
,
|
| 52 |
+
Feryal Behbahani
|
| 53 |
+
,
|
| 54 |
+
Aleksandra Faust
|
| 55 |
+
,
|
| 56 |
+
Hugo Larochelle
|
| 57 |
+
View a PDF of the paper titled Many-Shot In-Context Learning, by Rishabh Agarwal and 13 other authors
|
| 58 |
+
View PDF
|
| 59 |
+
HTML (experimental)
|
| 60 |
+
Abstract:
|
| 61 |
+
Large language models (LLMs) excel at few-shot in-context learning (ICL) -- learning from a few examples provided in context at inference, without any weight updates. Newly expanded context windows allow us to investigate ICL with hundreds or thousands of examples -- the many-shot regime. Going from few-shot to many-shot, we observe significant performance gains across a wide variety of generative and discriminative tasks. While promising, many-shot ICL can be bottlenecked by the available amount of human-generated examples. To mitigate this limitation, we explore two new settings: Reinforced and Unsupervised ICL. Reinforced ICL uses model-generated chain-of-thought rationales in place of human examples. Unsupervised ICL removes rationales from the prompt altogether, and prompts the model only with domain-specific questions. We find that both Reinforced and Unsupervised ICL can be quite effective in the many-shot regime, particularly on complex reasoning tasks. Finally, we demonstrate that, unlike few-shot learning, many-shot learning is effective at overriding pretraining biases, can learn high-dimensional functions with numerical inputs, and performs comparably to fine-tuning. We also find that inference cost increases linearly in the many-shot regime, and frontier LLMs benefit from many-shot ICL to varying degrees. Our analysis also reveals the limitations of next-token prediction loss as an indicator of downstream ICL performance.
|
| 62 |
+
Comments:
|
| 63 |
+
NeurIPS (Spotlight)
|
| 64 |
+
Subjects:
|
| 65 |
+
Machine Learning (cs.LG)
|
| 66 |
+
; Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
|
| 67 |
+
Cite as:
|
| 68 |
+
arXiv:2404.11018
|
| 69 |
+
[cs.LG]
|
| 70 |
+
(or
|
| 71 |
+
arXiv:2404.11018v3
|
| 72 |
+
[cs.LG]
|
| 73 |
+
for this version)
|
| 74 |
+
https://doi.org/10.48550/arXiv.2404.11018
|
| 75 |
+
Focus to learn more
|
| 76 |
+
arXiv-issued DOI via DataCite
|
| 77 |
+
Submission history
|
| 78 |
+
From: Rishabh Agarwal [
|
| 79 |
+
view email
|
| 80 |
+
]
|
| 81 |
+
[v1]
|
| 82 |
+
Wed, 17 Apr 2024 02:49:26 UTC (327 KB)
|
| 83 |
+
[v2]
|
| 84 |
+
Wed, 22 May 2024 17:06:10 UTC (370 KB)
|
| 85 |
+
[v3]
|
| 86 |
+
Thu, 17 Oct 2024 17:45:09 UTC (414 KB)
|
| 87 |
+
Full-text links:
|
| 88 |
+
Access Paper:
|
| 89 |
+
View a PDF of the paper titled Many-Shot In-Context Learning, by Rishabh Agarwal and 13 other authors
|
| 90 |
+
View PDF
|
| 91 |
+
HTML (experimental)
|
| 92 |
+
TeX Source
|
| 93 |
+
view license
|
| 94 |
+
Current browse context:
|
| 95 |
+
cs.LG
|
| 96 |
+
< prev
|
| 97 |
+
|
|
| 98 |
+
next >
|
| 99 |
+
new
|
| 100 |
+
|
|
| 101 |
+
recent
|
| 102 |
+
|
|
| 103 |
+
2024-04
|
| 104 |
+
Change to browse by:
|
| 105 |
+
cs
|
| 106 |
+
cs.AI
|
| 107 |
+
cs.CL
|
| 108 |
+
References & Citations
|
| 109 |
+
NASA ADS
|
| 110 |
+
Google Scholar
|
| 111 |
+
Semantic Scholar
|
| 112 |
+
export BibTeX citation
|
| 113 |
+
Loading...
|
| 114 |
+
BibTeX formatted citation
|
| 115 |
+
×
|
| 116 |
+
loading...
|
| 117 |
+
Data provided by:
|
| 118 |
+
Bookmark
|
| 119 |
+
Bibliographic Tools
|
| 120 |
+
Bibliographic and Citation Tools
|
| 121 |
+
Bibliographic Explorer Toggle
|
| 122 |
+
Bibliographic Explorer
|
| 123 |
+
(
|
| 124 |
+
What is the Explorer?
|
| 125 |
+
)
|
| 126 |
+
Connected Papers Toggle
|
| 127 |
+
Connected Papers
|
| 128 |
+
(
|
| 129 |
+
What is Connected Papers?
|
| 130 |
+
)
|
| 131 |
+
Litmaps Toggle
|
| 132 |
+
Litmaps
|
| 133 |
+
(
|
| 134 |
+
What is Litmaps?
|
| 135 |
+
)
|
| 136 |
+
scite.ai Toggle
|
| 137 |
+
scite Smart Citations
|
| 138 |
+
(
|
| 139 |
+
What are Smart Citations?
|
| 140 |
+
)
|
| 141 |
+
Code, Data, Media
|
| 142 |
+
Code, Data and Media Associated with this Article
|
| 143 |
+
alphaXiv Toggle
|
| 144 |
+
alphaXiv
|
| 145 |
+
(
|
| 146 |
+
What is alphaXiv?
|
| 147 |
+
)
|
| 148 |
+
Links to Code Toggle
|
| 149 |
+
CatalyzeX Code Finder for Papers
|
| 150 |
+
(
|
| 151 |
+
What is CatalyzeX?
|
| 152 |
+
)
|
| 153 |
+
DagsHub Toggle
|
| 154 |
+
DagsHub
|
| 155 |
+
(
|
| 156 |
+
What is DagsHub?
|
| 157 |
+
)
|
| 158 |
+
GotitPub Toggle
|
| 159 |
+
Gotit.pub
|
| 160 |
+
(
|
| 161 |
+
What is GotitPub?
|
| 162 |
+
)
|
| 163 |
+
Huggingface Toggle
|
| 164 |
+
Hugging Face
|
| 165 |
+
(
|
| 166 |
+
What is Huggingface?
|
| 167 |
+
)
|
| 168 |
+
Links to Code Toggle
|
| 169 |
+
Papers with Code
|
| 170 |
+
(
|
| 171 |
+
What is Papers with Code?
|
| 172 |
+
)
|
| 173 |
+
ScienceCast Toggle
|
| 174 |
+
ScienceCast
|
| 175 |
+
(
|
| 176 |
+
What is ScienceCast?
|
| 177 |
+
)
|
| 178 |
+
Demos
|
| 179 |
+
Demos
|
| 180 |
+
Replicate Toggle
|
| 181 |
+
Replicate
|
| 182 |
+
(
|
| 183 |
+
What is Replicate?
|
| 184 |
+
)
|
| 185 |
+
Spaces Toggle
|
| 186 |
+
Hugging Face Spaces
|
| 187 |
+
(
|
| 188 |
+
What is Spaces?
|
| 189 |
+
)
|
| 190 |
+
Spaces Toggle
|
| 191 |
+
TXYZ.AI
|
| 192 |
+
(
|
| 193 |
+
What is TXYZ.AI?
|
| 194 |
+
)
|
| 195 |
+
Related Papers
|
| 196 |
+
Recommenders and Search Tools
|
| 197 |
+
Link to Influence Flower
|
| 198 |
+
Influence Flower
|
| 199 |
+
(
|
| 200 |
+
What are Influence Flowers?
|
| 201 |
+
)
|
| 202 |
+
Core recommender toggle
|
| 203 |
+
CORE Recommender
|
| 204 |
+
(
|
| 205 |
+
What is CORE?
|
| 206 |
+
)
|
| 207 |
+
IArxiv recommender toggle
|
| 208 |
+
IArxiv Recommender
|
| 209 |
+
(
|
| 210 |
+
What is IArxiv?
|
| 211 |
+
)
|
| 212 |
+
Author
|
| 213 |
+
Venue
|
| 214 |
+
Institution
|
| 215 |
+
Topic
|
| 216 |
+
About arXivLabs
|
| 217 |
+
arXivLabs: experimental projects with community collaborators
|
| 218 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 219 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 220 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 221 |
+
Learn more about arXivLabs
|
| 222 |
+
.
|
| 223 |
+
Which authors of this paper are endorsers?
|
| 224 |
+
|
|
| 225 |
+
Disable MathJax
|
| 226 |
+
(
|
| 227 |
+
What is MathJax?
|
| 228 |
+
)
|
|
@@ -0,0 +1,197 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2406.12543] Phase-controlled heat modulation with Aharonov-Bohm interferometers'
|
| 3 |
+
id: 240612543-phase-controlled-heat-modulation-with-aharonov-bohm-interferometers
|
| 4 |
+
tags:
|
| 5 |
+
- deepread
|
| 6 |
+
created: '2026-06-10T00:40:09.876451Z'
|
| 7 |
+
source: https://arxiv.org/abs/2406.12543
|
| 8 |
+
source_domain: arxiv.org
|
| 9 |
+
fetched_at: '2026-06-10T00:40:09.876309Z'
|
| 10 |
+
fetch_provider: builtin
|
| 11 |
+
status: draft
|
| 12 |
+
type: note
|
| 13 |
+
tier: institutional
|
| 14 |
+
content_type: paper
|
| 15 |
+
deprecated: false
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
[2406.12543] Phase-controlled heat modulation with Aharonov-Bohm interferometers
|
| 19 |
+
Condensed Matter > Mesoscale and Nanoscale Physics
|
| 20 |
+
arXiv:2406.12543
|
| 21 |
+
(cond-mat)
|
| 22 |
+
[Submitted on 18 Jun 2024]
|
| 23 |
+
Title:
|
| 24 |
+
Phase-controlled heat modulation with Aharonov-Bohm interferometers
|
| 25 |
+
Authors:
|
| 26 |
+
Sun-Yong Hwang
|
| 27 |
+
,
|
| 28 |
+
Björn Sothmann
|
| 29 |
+
,
|
| 30 |
+
Rosa López
|
| 31 |
+
View a PDF of the paper titled Phase-controlled heat modulation with Aharonov-Bohm interferometers, by Sun-Yong Hwang and 2 other authors
|
| 32 |
+
View PDF
|
| 33 |
+
HTML (experimental)
|
| 34 |
+
Abstract:
|
| 35 |
+
A heat modulator is proposed based on a voltage-biased Aharonov-Bohm interferometer. Once an electrical bias is applied, Peltier effects give rise to a flow of heat that can be modulated by a magnetic flux. We determine the corresponding temperature changes using a simple thermal model. Our calculations demonstrate that the modulated temperature difference can be as large as 80 mK at base temperature about 600 mK with relative temperature variations reaching 10\%. Our model also predicts, quite generally, the emergence of spin-polarized heat flows without any ferromagnetic contacts, if Rashba spin-orbit interaction is combined with the applied magnetic flux, which potentially paves the way towards caloritronic information processing.
|
| 36 |
+
Comments:
|
| 37 |
+
8 pages, 4 figures
|
| 38 |
+
Subjects:
|
| 39 |
+
Mesoscale and Nanoscale Physics (cond-mat.mes-hall)
|
| 40 |
+
Cite as:
|
| 41 |
+
arXiv:2406.12543
|
| 42 |
+
[cond-mat.mes-hall]
|
| 43 |
+
(or
|
| 44 |
+
arXiv:2406.12543v1
|
| 45 |
+
[cond-mat.mes-hall]
|
| 46 |
+
for this version)
|
| 47 |
+
https://doi.org/10.48550/arXiv.2406.12543
|
| 48 |
+
Focus to learn more
|
| 49 |
+
arXiv-issued DOI via DataCite
|
| 50 |
+
Journal reference:
|
| 51 |
+
Phys. Rev. Research 6, 013215 (2024)
|
| 52 |
+
Related DOI
|
| 53 |
+
:
|
| 54 |
+
https://doi.org/10.1103/PhysRevResearch.6.013215
|
| 55 |
+
Focus to learn more
|
| 56 |
+
DOI(s) linking to related resources
|
| 57 |
+
Submission history
|
| 58 |
+
From: Sun-Yong Hwang [
|
| 59 |
+
view email
|
| 60 |
+
]
|
| 61 |
+
[v1]
|
| 62 |
+
Tue, 18 Jun 2024 12:22:44 UTC (1,894 KB)
|
| 63 |
+
Full-text links:
|
| 64 |
+
Access Paper:
|
| 65 |
+
View a PDF of the paper titled Phase-controlled heat modulation with Aharonov-Bohm interferometers, by Sun-Yong Hwang and 2 other authors
|
| 66 |
+
View PDF
|
| 67 |
+
HTML (experimental)
|
| 68 |
+
TeX Source
|
| 69 |
+
view license
|
| 70 |
+
Current browse context:
|
| 71 |
+
cond-mat.mes-hall
|
| 72 |
+
< prev
|
| 73 |
+
|
|
| 74 |
+
next >
|
| 75 |
+
new
|
| 76 |
+
|
|
| 77 |
+
recent
|
| 78 |
+
|
|
| 79 |
+
2024-06
|
| 80 |
+
Change to browse by:
|
| 81 |
+
cond-mat
|
| 82 |
+
References & Citations
|
| 83 |
+
NASA ADS
|
| 84 |
+
Google Scholar
|
| 85 |
+
Semantic Scholar
|
| 86 |
+
export BibTeX citation
|
| 87 |
+
Loading...
|
| 88 |
+
BibTeX formatted citation
|
| 89 |
+
×
|
| 90 |
+
loading...
|
| 91 |
+
Data provided by:
|
| 92 |
+
Bookmark
|
| 93 |
+
Bibliographic Tools
|
| 94 |
+
Bibliographic and Citation Tools
|
| 95 |
+
Bibliographic Explorer Toggle
|
| 96 |
+
Bibliographic Explorer
|
| 97 |
+
(
|
| 98 |
+
What is the Explorer?
|
| 99 |
+
)
|
| 100 |
+
Connected Papers Toggle
|
| 101 |
+
Connected Papers
|
| 102 |
+
(
|
| 103 |
+
What is Connected Papers?
|
| 104 |
+
)
|
| 105 |
+
Litmaps Toggle
|
| 106 |
+
Litmaps
|
| 107 |
+
(
|
| 108 |
+
What is Litmaps?
|
| 109 |
+
)
|
| 110 |
+
scite.ai Toggle
|
| 111 |
+
scite Smart Citations
|
| 112 |
+
(
|
| 113 |
+
What are Smart Citations?
|
| 114 |
+
)
|
| 115 |
+
Code, Data, Media
|
| 116 |
+
Code, Data and Media Associated with this Article
|
| 117 |
+
alphaXiv Toggle
|
| 118 |
+
alphaXiv
|
| 119 |
+
(
|
| 120 |
+
What is alphaXiv?
|
| 121 |
+
)
|
| 122 |
+
Links to Code Toggle
|
| 123 |
+
CatalyzeX Code Finder for Papers
|
| 124 |
+
(
|
| 125 |
+
What is CatalyzeX?
|
| 126 |
+
)
|
| 127 |
+
DagsHub Toggle
|
| 128 |
+
DagsHub
|
| 129 |
+
(
|
| 130 |
+
What is DagsHub?
|
| 131 |
+
)
|
| 132 |
+
GotitPub Toggle
|
| 133 |
+
Gotit.pub
|
| 134 |
+
(
|
| 135 |
+
What is GotitPub?
|
| 136 |
+
)
|
| 137 |
+
Huggingface Toggle
|
| 138 |
+
Hugging Face
|
| 139 |
+
(
|
| 140 |
+
What is Huggingface?
|
| 141 |
+
)
|
| 142 |
+
ScienceCast Toggle
|
| 143 |
+
ScienceCast
|
| 144 |
+
(
|
| 145 |
+
What is ScienceCast?
|
| 146 |
+
)
|
| 147 |
+
Demos
|
| 148 |
+
Demos
|
| 149 |
+
Replicate Toggle
|
| 150 |
+
Replicate
|
| 151 |
+
(
|
| 152 |
+
What is Replicate?
|
| 153 |
+
)
|
| 154 |
+
Spaces Toggle
|
| 155 |
+
Hugging Face Spaces
|
| 156 |
+
(
|
| 157 |
+
What is Spaces?
|
| 158 |
+
)
|
| 159 |
+
Spaces Toggle
|
| 160 |
+
TXYZ.AI
|
| 161 |
+
(
|
| 162 |
+
What is TXYZ.AI?
|
| 163 |
+
)
|
| 164 |
+
Related Papers
|
| 165 |
+
Recommenders and Search Tools
|
| 166 |
+
Link to Influence Flower
|
| 167 |
+
Influence Flower
|
| 168 |
+
(
|
| 169 |
+
What are Influence Flowers?
|
| 170 |
+
)
|
| 171 |
+
Core recommender toggle
|
| 172 |
+
CORE Recommender
|
| 173 |
+
(
|
| 174 |
+
What is CORE?
|
| 175 |
+
)
|
| 176 |
+
IArxiv recommender toggle
|
| 177 |
+
IArxiv Recommender
|
| 178 |
+
(
|
| 179 |
+
What is IArxiv?
|
| 180 |
+
)
|
| 181 |
+
Author
|
| 182 |
+
Venue
|
| 183 |
+
Institution
|
| 184 |
+
Topic
|
| 185 |
+
About arXivLabs
|
| 186 |
+
arXivLabs: experimental projects with community collaborators
|
| 187 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 188 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 189 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 190 |
+
Learn more about arXivLabs
|
| 191 |
+
.
|
| 192 |
+
Which authors of this paper are endorsers?
|
| 193 |
+
|
|
| 194 |
+
Disable MathJax
|
| 195 |
+
(
|
| 196 |
+
What is MathJax?
|
| 197 |
+
)
|
|
@@ -0,0 +1,2384 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2408.06195] Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers'
|
| 3 |
+
id: 240806195-mutual-reasoning-makes-smaller-llms-stronger-problem-solvers-2
|
| 4 |
+
tags:
|
| 5 |
+
- deepread
|
| 6 |
+
created: '2026-06-10T00:40:45.751799Z'
|
| 7 |
+
source: https://ar5iv.labs.arxiv.org/html/2408.06195
|
| 8 |
+
source_domain: ar5iv.labs.arxiv.org
|
| 9 |
+
fetched_at: '2026-06-10T00:40:45.751621Z'
|
| 10 |
+
fetch_provider: builtin
|
| 11 |
+
status: draft
|
| 12 |
+
type: note
|
| 13 |
+
tier: institutional
|
| 14 |
+
content_type: paper
|
| 15 |
+
deprecated: false
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
[2408.06195] Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
|
| 19 |
+
Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
|
| 20 |
+
Zhenting Qi
|
| 21 |
+
∗‡†
|
| 22 |
+
Mingyuan Ma
|
| 23 |
+
∗‡†
|
| 24 |
+
Jiahang Xu
|
| 25 |
+
∗‡
|
| 26 |
+
Li Lyna Zhang
|
| 27 |
+
‡⋄
|
| 28 |
+
Fan Yang
|
| 29 |
+
‡
|
| 30 |
+
Mao Yang
|
| 31 |
+
‡
|
| 32 |
+
‡
|
| 33 |
+
Microsoft Research Asia
|
| 34 |
+
†
|
| 35 |
+
Harvard University
|
| 36 |
+
Abstract
|
| 37 |
+
This paper introduces rStar, a self-play mutual reasoning approach that significantly improves reasoning capabilities of small language models (SLMs) without fine-tuning or superior models. rStar decouples reasoning into a self-play mutual generation-discrimination process. First, a target SLM augments the Monte Carlo Tree Search (MCTS) with
|
| 38 |
+
a rich set of human-like reasoning actions
|
| 39 |
+
to construct higher quality reasoning trajectories. Next, another SLM, with capabilities similar to the target SLM, acts as a discriminator to verify each trajectory generated by the target SLM. The mutually agreed reasoning trajectories are considered
|
| 40 |
+
mutual consistent
|
| 41 |
+
, thus are more likely to be correct. Extensive experiments across five SLMs demonstrate rStar can effectively solve diverse reasoning problems, including GSM8K, GSM-Hard, MATH, SVAMP, and StrategyQA. Remarkably, rStar boosts GSM8K accuracy from 12.51% to 63.91% for LLaMA2-7B, from 36.46% to 81.88% for Mistral-7B, from 74.53% to 91.13% for LLaMA3-8B-Instruct. Code will be available at
|
| 42 |
+
here
|
| 43 |
+
.
|
| 44 |
+
$*$
|
| 45 |
+
$*$
|
| 46 |
+
footnotetext:
|
| 47 |
+
Equal contribution. Zhenting Qi and Mingyuan Ma did the work during an internship at MSRA
|
| 48 |
+
$\diamond$
|
| 49 |
+
$\diamond$
|
| 50 |
+
footnotetext:
|
| 51 |
+
Corresponding author: lzhani@microsoft.com
|
| 52 |
+
1
|
| 53 |
+
Introduction
|
| 54 |
+
Despite their success, large language models (LLMs) face significant challenges in complex reasoning
|
| 55 |
+
(Valmeekam et al.,
|
| 56 |
+
2022
|
| 57 |
+
; Weng et al.,
|
| 58 |
+
2023
|
| 59 |
+
)
|
| 60 |
+
. For example, state of the art models like Mistral-7B
|
| 61 |
+
(Jiang et al.,
|
| 62 |
+
2023
|
| 63 |
+
)
|
| 64 |
+
can only achieve 36.5% accuracy on the GSM8K dataset, even with techniques like Chain-of-Throught (CoT)
|
| 65 |
+
(Wei et al.,
|
| 66 |
+
2022
|
| 67 |
+
)
|
| 68 |
+
. Although fine-tuning is shown to be an effective way to improve reasoning capability, most LLMs rely on fine-tuning data distilled or synthesized by
|
| 69 |
+
superior
|
| 70 |
+
models like GPT-4
|
| 71 |
+
(Wang et al.,
|
| 72 |
+
2024a
|
| 73 |
+
; Gou et al.,
|
| 74 |
+
2023
|
| 75 |
+
)
|
| 76 |
+
. Meanwhile, the community has been actively working on a complimentary and yet more challenging approach: Reasoning improvements
|
| 77 |
+
without
|
| 78 |
+
a superior teacher LLM.
|
| 79 |
+
Figure 1:
|
| 80 |
+
With 32 rounds of inference, rStar makes SLMs highly capable problem-solvers, matching or even surpassing the reasoning performance achieved after domain-specialized SFT.
|
| 81 |
+
A promising paradigm to improve reasoning without superior models is to leverage the knowledge within LLMs themselves
|
| 82 |
+
(Wang et al.,
|
| 83 |
+
2023
|
| 84 |
+
; Hao et al.,
|
| 85 |
+
2023
|
| 86 |
+
; Madaan et al.,
|
| 87 |
+
2024
|
| 88 |
+
)
|
| 89 |
+
. For example, RAP
|
| 90 |
+
(Hao et al.,
|
| 91 |
+
2023
|
| 92 |
+
)
|
| 93 |
+
adopts a self-exploration solution to iteratively improve LLM’s reasoning performance through self-rewarded feedback. Unfortunately, study suggests that this paradigm often suffers from two fundamental issues.
|
| 94 |
+
First, LLMs often struggle to effectively explore the solution space during reasoning. The self-exploration often traps in a solution space with low-quality reasoning steps even after many attempts. For example, our experiments reveal that after 32 rounds of self-exploration with RAP
|
| 95 |
+
(Hao et al.,
|
| 96 |
+
2023
|
| 97 |
+
)
|
| 98 |
+
, only 24% of the trajectories generated by LLaMA2-7B on GSM8K are correct.
|
| 99 |
+
Second, even the self-exploration can find high quality reasoning steps, it is difficult for SLMs to tell which reasoning steps are of higher quality or determine which final answers are correct, thus it is hard to effectively guide the self-exploration. Our study shows that a naïve reward-based self-exploration guidance can lead to results no better than random guesses (see Appendix
|
| 100 |
+
A.1
|
| 101 |
+
).
|
| 102 |
+
A more troublesome fact is that the above two issues are more pronounced in the smaller version of LLMs, i.e.,
|
| 103 |
+
SLM
|
| 104 |
+
s, due to their weaker capabilities. For instance, while GPT-4 can improve by self-refining its output
|
| 105 |
+
(Madaan et al.,
|
| 106 |
+
2024
|
| 107 |
+
; Wu et al.,
|
| 108 |
+
2024
|
| 109 |
+
; Zhou et al.,
|
| 110 |
+
2024
|
| 111 |
+
)
|
| 112 |
+
, the approaches are less effective in SLMs and may even lead to worse performance
|
| 113 |
+
(Forsman,
|
| 114 |
+
2024
|
| 115 |
+
)
|
| 116 |
+
. This significantly hinders the adoption of neural language models.
|
| 117 |
+
This paper introduces
|
| 118 |
+
S
|
| 119 |
+
elf-play mu
|
| 120 |
+
T
|
| 121 |
+
u
|
| 122 |
+
A
|
| 123 |
+
l
|
| 124 |
+
R
|
| 125 |
+
easoning
|
| 126 |
+
(rStar), a novel approach that boosts SLMs’ reasoning capability during inference without fine-tuning or superior models. To address the aforementioned challenges, rStar decouples reasoning into a self-play mutual generation-discrimination process as illustrated in Fig.
|
| 127 |
+
2
|
| 128 |
+
.
|
| 129 |
+
Specifically, rStar is unique in the following approaches. First, although relying on a conventional Monte Carlo Tree Search (MCTS) for SLMs to self-generate reasoning steps, rStar advocates
|
| 130 |
+
a richer set of reasoning actions
|
| 131 |
+
in the self-exploration. The new proposed actions simulate human reasoning behaviors given the current reasoning state, such as decomposing and searching for a specific reasoning step, proposing a new sub-question, or rephrasing the given question. This enables SLMs to generate high-quality candidate reasoning trajectories during self-exploration.
|
| 132 |
+
Second, to effectively guide the exploration among the generated reasoning trajectories, rStar augments the MCTS process with a new discrimination process called
|
| 133 |
+
mutual consistency
|
| 134 |
+
. In particular, rStar employs a second SLM with the similar capability, acting as a discriminator to provide unsupervised feedback on each candidate reasoning trajectory generated by MCTS. To improve the accuracy of the feedback, rStar hints the second SLM with sampled partial reasoning trajectories, asking it to complete the remaining reasoning steps. And rStar deems the mutually agreed reasoning trajectories of higher quality. Mutual consistency mirrors the common human practice in the absence of supervision, where agreement among peers (i.e., two SLMs) on derived answers suggests a higher likelihood of correctness.
|
| 135 |
+
As a result, mutual consistency offers more effective reasoning across diverse tasks than other approaches like self-consistency
|
| 136 |
+
(Wang et al.,
|
| 137 |
+
2023
|
| 138 |
+
)
|
| 139 |
+
and avoids the risk of overfitting when training a reward model
|
| 140 |
+
(Chen et al.,
|
| 141 |
+
2024a
|
| 142 |
+
; Wang et al.,
|
| 143 |
+
2024b
|
| 144 |
+
)
|
| 145 |
+
.
|
| 146 |
+
Figure 2:
|
| 147 |
+
Our self-play mutual reasoning is a generation-discrimination process: (1) a self-generator augments the target SLM to generate candidate reasoning trajectories using MCTS; (2) the discriminator uses another SLM to provide unsupervised feedback on each trajectory based on partial hints; (3) based on this feedback, the target SLM decides a final reasoning trajectory as the solution.
|
| 148 |
+
Extensive experiments across five SLMs and five diverse reasoning tasks demonstrate the effectiveness of rStar. With just 32 rounds of MCTS inference, rStar significantly enhances SLMs’ reasoning capabilities, matching or even surpassing the accuracy achieved after fine-tuning. For example, rStar boosts GSM8K accuracy from 12.51% to 63.91% for LLaMA2-7B, from 36.46% to 81.88% for Mistral, and from 47.23% to 85.52% for LLaMA3-8B. Furthermore, we conduct comprehensive experiments to verify rStar’s superiority over state-of-the-art baselines, including single-round inference techniques like few-shot CoT, multi-round prompting approaches such as self-consistency, and self-improvement techniques such as RAP, ToT, self-evaluation and self-verification.
|
| 149 |
+
2
|
| 150 |
+
Related Work
|
| 151 |
+
Prompting Language Models to Reason
|
| 152 |
+
.
|
| 153 |
+
Prompting-based methods, such as Chain-of-Thought
|
| 154 |
+
(Wei et al.,
|
| 155 |
+
2022
|
| 156 |
+
)
|
| 157 |
+
, focus on designing instructions and pipelines to enhance LLMs’ reasoning performance during inference. Recent advances include planning
|
| 158 |
+
(Hao et al.,
|
| 159 |
+
2023
|
| 160 |
+
; Ding et al.,
|
| 161 |
+
2023
|
| 162 |
+
)
|
| 163 |
+
, problem decomposition
|
| 164 |
+
(Zhou et al.,
|
| 165 |
+
2022
|
| 166 |
+
; Khot et al.,
|
| 167 |
+
2022
|
| 168 |
+
; Hao et al.,
|
| 169 |
+
2023
|
| 170 |
+
)
|
| 171 |
+
, abstraction
|
| 172 |
+
(Zheng et al.,
|
| 173 |
+
2023
|
| 174 |
+
)
|
| 175 |
+
, programming
|
| 176 |
+
(Chen et al.,
|
| 177 |
+
2022
|
| 178 |
+
; Zhou et al.,
|
| 179 |
+
2023
|
| 180 |
+
)
|
| 181 |
+
.
|
| 182 |
+
These methods aim to improve single-round inference performance and are orthogonal to ours.
|
| 183 |
+
LLM Self-improvement
|
| 184 |
+
. Recently, research on the self-improvement of LLMs has rapidly increased.
|
| 185 |
+
Fine-tuning based methods
|
| 186 |
+
(Chen et al.,
|
| 187 |
+
2024b
|
| 188 |
+
;
|
| 189 |
+
a
|
| 190 |
+
)
|
| 191 |
+
leverage the capabilities of a well-pretrained LLM to synthesize data and progressively enhance its performance. Advanced prompting techniques, such as self-verification
|
| 192 |
+
(Gero et al.,
|
| 193 |
+
2023
|
| 194 |
+
; Zhou et al.,
|
| 195 |
+
2023
|
| 196 |
+
)
|
| 197 |
+
, and RAP
|
| 198 |
+
(Hao et al.,
|
| 199 |
+
2023
|
| 200 |
+
)
|
| 201 |
+
, improve performance through iterative self-exploring based on self-diagnosed feedback at inference time.
|
| 202 |
+
However, as illustrated in previous section, the achieved performance often depend on the LLM’s inherent capabilities, and for SLMs, their weaker instruction-following ability and unreliable self-rewarding can mislead self-improvement.
|
| 203 |
+
Sampling Reasoning Paths
|
| 204 |
+
. Recent works
|
| 205 |
+
(Brown et al.,
|
| 206 |
+
2024
|
| 207 |
+
; Li et al.,
|
| 208 |
+
2024
|
| 209 |
+
; Snell et al.,
|
| 210 |
+
2024
|
| 211 |
+
)
|
| 212 |
+
on mathematical reasoning have shown that sampling diverse reasoning paths can significantly enhance performance compared to greedy one-time decoding. Self-Consistency
|
| 213 |
+
(Wang et al.,
|
| 214 |
+
2023
|
| 215 |
+
)
|
| 216 |
+
sample a complete CoT path each time. Tree-search approaches
|
| 217 |
+
(Yao et al.,
|
| 218 |
+
2024
|
| 219 |
+
; Hao et al.,
|
| 220 |
+
2023
|
| 221 |
+
; Zhang et al.,
|
| 222 |
+
2024
|
| 223 |
+
)
|
| 224 |
+
, like MCTS, further improve the performance by breaking down tasks and sampling simpler, individual intermediate reasoning steps. However, most approaches have limited action spaces. For example, RAP
|
| 225 |
+
(Hao et al.,
|
| 226 |
+
2023
|
| 227 |
+
)
|
| 228 |
+
decomposes only subproblems, while AlphaMath
|
| 229 |
+
(Chen et al.,
|
| 230 |
+
2024a
|
| 231 |
+
)
|
| 232 |
+
searches only for one CoT step, limiting effectiveness in generating better trajectories.
|
| 233 |
+
Answer Verification
|
| 234 |
+
. To select correct reasoning trajectories, majority voting
|
| 235 |
+
(Wang et al.,
|
| 236 |
+
2023
|
| 237 |
+
)
|
| 238 |
+
is a widely-used approach. To improve accuracy, some works train value or rewards model for verification
|
| 239 |
+
(Wang et al.,
|
| 240 |
+
2024b
|
| 241 |
+
; Chen et al.,
|
| 242 |
+
2024a
|
| 243 |
+
)
|
| 244 |
+
, but these require additional annotations and have risks in overfitting to specific tasks. Self-verification
|
| 245 |
+
(Weng et al.,
|
| 246 |
+
2023
|
| 247 |
+
)
|
| 248 |
+
leverages LLM capabilities for backward self-verification. Nevertheless, its effectiveness hinges on its inherent ability to reason effectively. Recent studies have shown that LLM struggles to evaluate itself and rectify its initial responses without any external feedbacks
|
| 249 |
+
(Huang et al.,
|
| 250 |
+
2023
|
| 251 |
+
; Feng et al.,
|
| 252 |
+
2023
|
| 253 |
+
)
|
| 254 |
+
.
|
| 255 |
+
3
|
| 256 |
+
Methodology
|
| 257 |
+
3.1
|
| 258 |
+
Overview
|
| 259 |
+
Problem Formulation
|
| 260 |
+
. To solve a reasoning problem by SLMs, we formulate it as a multi-step reasoning generation task, which breaks
|
| 261 |
+
the problem into simpler sub-tasks. This is more effective than traditional CoT-based reasoning
|
| 262 |
+
(Wei et al.,
|
| 263 |
+
2022
|
| 264 |
+
; Wang et al.,
|
| 265 |
+
2023
|
| 266 |
+
)
|
| 267 |
+
, as it is much easier for SLMs to correctly generate one step than complete reasoning steps in a single inference. We leverage the Monte-Carlo Tree Search (MCTS) algorithm
|
| 268 |
+
(Kocsis & Szepesvári,
|
| 269 |
+
2006
|
| 270 |
+
)
|
| 271 |
+
to augment the target SLM for self-generating multi-step reasoning solutions.
|
| 272 |
+
Formally, for a given problem
|
| 273 |
+
x
|
| 274 |
+
𝑥
|
| 275 |
+
x
|
| 276 |
+
and a target SLM
|
| 277 |
+
M
|
| 278 |
+
𝑀
|
| 279 |
+
M
|
| 280 |
+
, the MCTS augments
|
| 281 |
+
M
|
| 282 |
+
𝑀
|
| 283 |
+
M
|
| 284 |
+
to incrementally build a search tree
|
| 285 |
+
𝒯
|
| 286 |
+
𝒯
|
| 287 |
+
\mathcal{T}
|
| 288 |
+
. As illustrated in Fig.
|
| 289 |
+
3
|
| 290 |
+
, the root node represents the question
|
| 291 |
+
x
|
| 292 |
+
𝑥
|
| 293 |
+
x
|
| 294 |
+
, an edge represents an action
|
| 295 |
+
a
|
| 296 |
+
𝑎
|
| 297 |
+
a
|
| 298 |
+
, each child node is an intermediate step
|
| 299 |
+
s
|
| 300 |
+
𝑠
|
| 301 |
+
s
|
| 302 |
+
generated by
|
| 303 |
+
M
|
| 304 |
+
𝑀
|
| 305 |
+
M
|
| 306 |
+
under the corresponding action. A path from the root node to a leaf node (denoted as
|
| 307 |
+
s
|
| 308 |
+
d
|
| 309 |
+
subscript
|
| 310 |
+
𝑠
|
| 311 |
+
𝑑
|
| 312 |
+
s_{d}
|
| 313 |
+
, also called a terminal node) constitutes a candidate solution trajectory
|
| 314 |
+
𝐭
|
| 315 |
+
=
|
| 316 |
+
x
|
| 317 |
+
⊕
|
| 318 |
+
s
|
| 319 |
+
1
|
| 320 |
+
⊕
|
| 321 |
+
s
|
| 322 |
+
2
|
| 323 |
+
⊕
|
| 324 |
+
…
|
| 325 |
+
⊕
|
| 326 |
+
s
|
| 327 |
+
d
|
| 328 |
+
𝐭
|
| 329 |
+
direct-sum
|
| 330 |
+
𝑥
|
| 331 |
+
subscript
|
| 332 |
+
𝑠
|
| 333 |
+
1
|
| 334 |
+
subscript
|
| 335 |
+
𝑠
|
| 336 |
+
2
|
| 337 |
+
…
|
| 338 |
+
subscript
|
| 339 |
+
𝑠
|
| 340 |
+
𝑑
|
| 341 |
+
\mathbf{t}=x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{d}
|
| 342 |
+
. From the search tree
|
| 343 |
+
𝒯
|
| 344 |
+
𝒯
|
| 345 |
+
\mathcal{T}
|
| 346 |
+
, we can extract a set of solution trajectories
|
| 347 |
+
𝕋
|
| 348 |
+
=
|
| 349 |
+
{
|
| 350 |
+
𝐭
|
| 351 |
+
1
|
| 352 |
+
,
|
| 353 |
+
𝐭
|
| 354 |
+
2
|
| 355 |
+
,
|
| 356 |
+
…
|
| 357 |
+
,
|
| 358 |
+
𝐭
|
| 359 |
+
n
|
| 360 |
+
}
|
| 361 |
+
|
| 362 |
+
(
|
| 363 |
+
n
|
| 364 |
+
≥
|
| 365 |
+
1
|
| 366 |
+
)
|
| 367 |
+
𝕋
|
| 368 |
+
superscript
|
| 369 |
+
𝐭
|
| 370 |
+
1
|
| 371 |
+
superscript
|
| 372 |
+
𝐭
|
| 373 |
+
2
|
| 374 |
+
…
|
| 375 |
+
superscript
|
| 376 |
+
𝐭
|
| 377 |
+
𝑛
|
| 378 |
+
𝑛
|
| 379 |
+
1
|
| 380 |
+
\mathbb{T}=\{\mathbf{t}^{1},\mathbf{t}^{2},...,\mathbf{t}^{n}\}(n\geq 1)
|
| 381 |
+
. Our goal is to find the trajectories that can achieve the correct answer for the given question.
|
| 382 |
+
Challenges in SLM Self-Improvement
|
| 383 |
+
. MCTS allows an SLM to explore and evaluate multiple potential solutions. Ideally, by balancing exploration of new possibilities with the exploitation of high-reward actions, the SLM can gradually refine its reasoning steps to generate a final correct reasoning trajectory. However, due to the limited capabilities in SLMs, traditional MCTS yields minimal improvement. First, the vast solution space makes it challenging for SLMs to generate effective solutions. Existing MCTS-based methods
|
| 384 |
+
(Hao et al.,
|
| 385 |
+
2023
|
| 386 |
+
; Kang et al.,
|
| 387 |
+
2024
|
| 388 |
+
)
|
| 389 |
+
that use single actions limit diversity and struggle to generalize across tasks. Approaches like self-consistency
|
| 390 |
+
(Wang et al.,
|
| 391 |
+
2023
|
| 392 |
+
)
|
| 393 |
+
use random sampling ensure diversity, SLMs often produce poor-quality solutions, requiring many attempts to find a correct solution, thereby increasing inference costs.
|
| 394 |
+
Second, it’s challenging to accurately reward each action. Without ground truth labels, it’s difficult to verify the correctness for each intermediate step
|
| 395 |
+
s
|
| 396 |
+
i
|
| 397 |
+
subscript
|
| 398 |
+
𝑠
|
| 399 |
+
𝑖
|
| 400 |
+
s_{i}
|
| 401 |
+
and the final answer in
|
| 402 |
+
s
|
| 403 |
+
d
|
| 404 |
+
subscript
|
| 405 |
+
𝑠
|
| 406 |
+
𝑑
|
| 407 |
+
s_{d}
|
| 408 |
+
. Majority voting in self-consistency requires most traces to be correct, which is often not the case for SLMs. Methods like RAP
|
| 409 |
+
(Hao et al.,
|
| 410 |
+
2023
|
| 411 |
+
)
|
| 412 |
+
use self-rewarding, but our study shows SLMs perform near-random self-rewarding (Appendix
|
| 413 |
+
A.1
|
| 414 |
+
). Training a reward model, as in M
|
| 415 |
+
∗
|
| 416 |
+
(Kang et al.,
|
| 417 |
+
2024
|
| 418 |
+
)
|
| 419 |
+
, can address this challenge but faces difficulties in collecting training data and generalizing across various tasks.
|
| 420 |
+
Overview
|
| 421 |
+
.
|
| 422 |
+
To address these challenges, this section introduces our methodology, rStar, which decomposes reasoning into solution generation and mutual verification in Fig.
|
| 423 |
+
2
|
| 424 |
+
. To tackle the first challenge, we introduce a richer set of human-like reasoning actions that allows for thorough space exploration across diverse reasoning tasks. To address the second challenge, we design an SLM-tailored reward function to evaluate intermediate steps, avoiding reliance on their often unreliable self-evaluations. Moreover, we use another SLM as a discriminator to augment the MCTS process, mutually verifying the correctness of each trajectory with the generator SLM.
|
| 425 |
+
3.2
|
| 426 |
+
Self-generating Reasoning Trajectory with MCTS Rollout
|
| 427 |
+
Figure 3:
|
| 428 |
+
An example to illustrate the process of self-generator. Highlighted nodes from top to bottom constitute a complete reasoning trace.
|
| 429 |
+
Given a question, MCTS augments the target SLM to explore a rich, human-like reasoning action space and generate the next steps based on the current state.
|
| 430 |
+
A Rich Set of Human-like Reasoning Actions
|
| 431 |
+
. At the core of MCTS generation lies the action space, which defines the scope of tree exploration. Most MCTS-based methods use a single action type to build the tree. For instance, in RAP, the action is to propose the next sub-question, whereas in AlphaMath
|
| 432 |
+
(Chen et al.,
|
| 433 |
+
2024a
|
| 434 |
+
)
|
| 435 |
+
and MindStar
|
| 436 |
+
(Kang et al.,
|
| 437 |
+
2024
|
| 438 |
+
)
|
| 439 |
+
, the action is to generate the next reasoning step.
|
| 440 |
+
However, relying on a single action type can easily lead to ineffective space exploration.
|
| 441 |
+
To address this, we revisit how humans approach reasoning.
|
| 442 |
+
Different people solve problems using diverse actions: some break into sub-questions, others solve it directly, and some might rephrase the problem to focus on key conditions. Moreover, people adjust their approach based on current states, choosing different actions as needed. Inspired by this human reasoning process, we introduce a richer set of 5 actions to maximize the SLM’s potential for correctly solving complex reasoning problems.
|
| 443 |
+
⋄
|
| 444 |
+
⋄
|
| 445 |
+
\diamond
|
| 446 |
+
A1
|
| 447 |
+
: Propose an one-step thought
|
| 448 |
+
. This action prompts the LLM to generate the next one-step thought for a given question, by considering the existing reasoning steps. Unlike the CoT, which generates complete thoughts, this approach simplifies the reasoning process and allows the LLM to perform better decision making
|
| 449 |
+
(Yao et al.,
|
| 450 |
+
2024
|
| 451 |
+
; Besta et al.,
|
| 452 |
+
2024
|
| 453 |
+
)
|
| 454 |
+
.
|
| 455 |
+
⋄
|
| 456 |
+
⋄
|
| 457 |
+
\diamond
|
| 458 |
+
A2
|
| 459 |
+
: Propose the remaining thought steps.
|
| 460 |
+
Instead of generating only one step thought per state, this action aligns with standard CoT, enabling “fast thinking” to solve simple question in fewer steps. Given the already generated reasoning steps, it prompts the LLM to directly produce the remaining steps until reaching the final answer.
|
| 461 |
+
⋄
|
| 462 |
+
⋄
|
| 463 |
+
\diamond
|
| 464 |
+
A3
|
| 465 |
+
: Propose next sub-question along with its answer.
|
| 466 |
+
This action is inspired by
|
| 467 |
+
least-to-most prompting
|
| 468 |
+
(Zhou et al.,
|
| 469 |
+
2022
|
| 470 |
+
)
|
| 471 |
+
, which breaks down a complex problem into a series of simpler sub-questions and solves them sequentially. Following RAP’s implementation, we prompt the LLM to ask and then answer the next sub-question.
|
| 472 |
+
⋄
|
| 473 |
+
⋄
|
| 474 |
+
\diamond
|
| 475 |
+
A4
|
| 476 |
+
: Answer the sub-question again.
|
| 477 |
+
Considering that a sub-question might not be answered correctly by
|
| 478 |
+
A3
|
| 479 |
+
, we propose this action to re-answer it. To improve accuracy, this action prompts the LLM to use few-shot CoT. Note that the original answer generated by
|
| 480 |
+
A3
|
| 481 |
+
did not use a CoT-like prompt but instead followed the least-to-most problem decomposition prompt
|
| 482 |
+
(Zhou et al.,
|
| 483 |
+
2022
|
| 484 |
+
)
|
| 485 |
+
.
|
| 486 |
+
⋄
|
| 487 |
+
⋄
|
| 488 |
+
\diamond
|
| 489 |
+
A5
|
| 490 |
+
: Rephrase the question/sub-question.
|
| 491 |
+
When analyzing incorrect cases, we found that many of them are due the LLM misunderstanding the question. For example, it might miss a specific condition provided in the question. Therefore, we propose a new action to rephrase the question more simply. Specifically, we prompt the LLM to clearly list all conditions given in the problem statement.
|
| 492 |
+
Table 1:
|
| 493 |
+
Ablation study on the effectiveness of our rich action space: we evaluate LLaMA3-8B on 200 sampled GSM8K questions.
|
| 494 |
+
Action Space
|
| 495 |
+
Accuracy
|
| 496 |
+
A
|
| 497 |
+
3
|
| 498 |
+
subscript
|
| 499 |
+
𝐴
|
| 500 |
+
3
|
| 501 |
+
A_{3}
|
| 502 |
+
(i.e., RAP)
|
| 503 |
+
70.5
|
| 504 |
+
A
|
| 505 |
+
3
|
| 506 |
+
subscript
|
| 507 |
+
𝐴
|
| 508 |
+
3
|
| 509 |
+
A_{3}
|
| 510 |
+
+
|
| 511 |
+
A
|
| 512 |
+
5
|
| 513 |
+
subscript
|
| 514 |
+
𝐴
|
| 515 |
+
5
|
| 516 |
+
A_{5}
|
| 517 |
+
72.5
|
| 518 |
+
A
|
| 519 |
+
3
|
| 520 |
+
subscript
|
| 521 |
+
𝐴
|
| 522 |
+
3
|
| 523 |
+
A_{3}
|
| 524 |
+
+
|
| 525 |
+
A
|
| 526 |
+
4
|
| 527 |
+
subscript
|
| 528 |
+
𝐴
|
| 529 |
+
4
|
| 530 |
+
A_{4}
|
| 531 |
+
+
|
| 532 |
+
A
|
| 533 |
+
5
|
| 534 |
+
subscript
|
| 535 |
+
𝐴
|
| 536 |
+
5
|
| 537 |
+
A_{5}
|
| 538 |
+
73.5
|
| 539 |
+
A
|
| 540 |
+
2
|
| 541 |
+
subscript
|
| 542 |
+
𝐴
|
| 543 |
+
2
|
| 544 |
+
A_{2}
|
| 545 |
+
+
|
| 546 |
+
A
|
| 547 |
+
3
|
| 548 |
+
subscript
|
| 549 |
+
𝐴
|
| 550 |
+
3
|
| 551 |
+
A_{3}
|
| 552 |
+
+
|
| 553 |
+
A
|
| 554 |
+
4
|
| 555 |
+
subscript
|
| 556 |
+
𝐴
|
| 557 |
+
4
|
| 558 |
+
A_{4}
|
| 559 |
+
+
|
| 560 |
+
A
|
| 561 |
+
5
|
| 562 |
+
subscript
|
| 563 |
+
𝐴
|
| 564 |
+
5
|
| 565 |
+
A_{5}
|
| 566 |
+
74.0
|
| 567 |
+
All (
|
| 568 |
+
A
|
| 569 |
+
1
|
| 570 |
+
subscript
|
| 571 |
+
𝐴
|
| 572 |
+
1
|
| 573 |
+
A_{1}
|
| 574 |
+
+
|
| 575 |
+
A
|
| 576 |
+
2
|
| 577 |
+
subscript
|
| 578 |
+
𝐴
|
| 579 |
+
2
|
| 580 |
+
A_{2}
|
| 581 |
+
+
|
| 582 |
+
A
|
| 583 |
+
3
|
| 584 |
+
subscript
|
| 585 |
+
𝐴
|
| 586 |
+
3
|
| 587 |
+
A_{3}
|
| 588 |
+
+
|
| 589 |
+
A
|
| 590 |
+
4
|
| 591 |
+
subscript
|
| 592 |
+
𝐴
|
| 593 |
+
4
|
| 594 |
+
A_{4}
|
| 595 |
+
+
|
| 596 |
+
A
|
| 597 |
+
5
|
| 598 |
+
subscript
|
| 599 |
+
𝐴
|
| 600 |
+
5
|
| 601 |
+
A_{5}
|
| 602 |
+
)
|
| 603 |
+
75.0
|
| 604 |
+
The above 5 actions define a highly diverse action space
|
| 605 |
+
{
|
| 606 |
+
A
|
| 607 |
+
1
|
| 608 |
+
,
|
| 609 |
+
A
|
| 610 |
+
2
|
| 611 |
+
,
|
| 612 |
+
A
|
| 613 |
+
3
|
| 614 |
+
,
|
| 615 |
+
A
|
| 616 |
+
4
|
| 617 |
+
,
|
| 618 |
+
A
|
| 619 |
+
5
|
| 620 |
+
}
|
| 621 |
+
subscript
|
| 622 |
+
𝐴
|
| 623 |
+
1
|
| 624 |
+
subscript
|
| 625 |
+
𝐴
|
| 626 |
+
2
|
| 627 |
+
subscript
|
| 628 |
+
𝐴
|
| 629 |
+
3
|
| 630 |
+
subscript
|
| 631 |
+
𝐴
|
| 632 |
+
4
|
| 633 |
+
subscript
|
| 634 |
+
𝐴
|
| 635 |
+
5
|
| 636 |
+
\{A_{1},A_{2},A_{3},A_{4},A_{5}\}
|
| 637 |
+
.
|
| 638 |
+
At each step
|
| 639 |
+
i
|
| 640 |
+
𝑖
|
| 641 |
+
i
|
| 642 |
+
, MCTS selects an action
|
| 643 |
+
a
|
| 644 |
+
i
|
| 645 |
+
subscript
|
| 646 |
+
𝑎
|
| 647 |
+
𝑖
|
| 648 |
+
a_{i}
|
| 649 |
+
from this space. We then use this action
|
| 650 |
+
a
|
| 651 |
+
i
|
| 652 |
+
subscript
|
| 653 |
+
𝑎
|
| 654 |
+
𝑖
|
| 655 |
+
a_{i}
|
| 656 |
+
to prompt the LLM to generate the next reasoning step
|
| 657 |
+
s
|
| 658 |
+
i
|
| 659 |
+
subscript
|
| 660 |
+
𝑠
|
| 661 |
+
𝑖
|
| 662 |
+
s_{i}
|
| 663 |
+
, based on the current state, which is the previous generated trajectory
|
| 664 |
+
x
|
| 665 |
+
⊕
|
| 666 |
+
s
|
| 667 |
+
1
|
| 668 |
+
⊕
|
| 669 |
+
s
|
| 670 |
+
2
|
| 671 |
+
⊕
|
| 672 |
+
…
|
| 673 |
+
⊕
|
| 674 |
+
s
|
| 675 |
+
i
|
| 676 |
+
−
|
| 677 |
+
1
|
| 678 |
+
direct-sum
|
| 679 |
+
𝑥
|
| 680 |
+
subscript
|
| 681 |
+
𝑠
|
| 682 |
+
1
|
| 683 |
+
subscript
|
| 684 |
+
𝑠
|
| 685 |
+
2
|
| 686 |
+
…
|
| 687 |
+
subscript
|
| 688 |
+
𝑠
|
| 689 |
+
𝑖
|
| 690 |
+
1
|
| 691 |
+
x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{i-1}
|
| 692 |
+
. Note that certain actions require orders. For example,
|
| 693 |
+
A4
|
| 694 |
+
can only happen after
|
| 695 |
+
A3
|
| 696 |
+
, and
|
| 697 |
+
A5
|
| 698 |
+
can only happen after the root question. As shown in Table
|
| 699 |
+
1
|
| 700 |
+
, each action plays a crucial role in improving the final reasoning accuracy.
|
| 701 |
+
Reward Function.
|
| 702 |
+
Another critical component in MCTS is the reward function, which evaluates the value of each action and directs the tree expansion.
|
| 703 |
+
We design a simple yet effective reward function for SLMs. First, we exclude self-rewarding techniques for any intermediate nodes due to the limited capabilities of SLMs. Second, to ensure generalization across different reasoning tasks, we avoid introducing external supervision (e.g., tools or trained value models). Our approach draws inspiration from AlphaGo
|
| 704 |
+
(Silver et al.,
|
| 705 |
+
2017
|
| 706 |
+
)
|
| 707 |
+
, where we score each intermediate node based on its contribution to the final correct answer. Consequently, actions that frequently lead to correct answers receive higher rewards, making them more likely to be selected in future MCTS tree expansions.
|
| 708 |
+
We define
|
| 709 |
+
Q
|
| 710 |
+
|
| 711 |
+
(
|
| 712 |
+
s
|
| 713 |
+
,
|
| 714 |
+
a
|
| 715 |
+
)
|
| 716 |
+
𝑄
|
| 717 |
+
𝑠
|
| 718 |
+
𝑎
|
| 719 |
+
Q(s,a)
|
| 720 |
+
as the reward value for node
|
| 721 |
+
s
|
| 722 |
+
𝑠
|
| 723 |
+
s
|
| 724 |
+
generated under action
|
| 725 |
+
a
|
| 726 |
+
𝑎
|
| 727 |
+
a
|
| 728 |
+
.
|
| 729 |
+
Initially, all unexplored nodes are assigned
|
| 730 |
+
Q
|
| 731 |
+
|
| 732 |
+
(
|
| 733 |
+
s
|
| 734 |
+
i
|
| 735 |
+
,
|
| 736 |
+
a
|
| 737 |
+
i
|
| 738 |
+
)
|
| 739 |
+
=
|
| 740 |
+
0
|
| 741 |
+
𝑄
|
| 742 |
+
subscript
|
| 743 |
+
𝑠
|
| 744 |
+
𝑖
|
| 745 |
+
subscript
|
| 746 |
+
𝑎
|
| 747 |
+
𝑖
|
| 748 |
+
0
|
| 749 |
+
Q(s_{i},a_{i})=0
|
| 750 |
+
, leading to random tree expansions. Upon reaching the first terminal node
|
| 751 |
+
n
|
| 752 |
+
d
|
| 753 |
+
subscript
|
| 754 |
+
𝑛
|
| 755 |
+
𝑑
|
| 756 |
+
n_{d}
|
| 757 |
+
, we compute a reward score
|
| 758 |
+
Q
|
| 759 |
+
|
| 760 |
+
(
|
| 761 |
+
s
|
| 762 |
+
d
|
| 763 |
+
,
|
| 764 |
+
a
|
| 765 |
+
d
|
| 766 |
+
)
|
| 767 |
+
𝑄
|
| 768 |
+
subscript
|
| 769 |
+
𝑠
|
| 770 |
+
𝑑
|
| 771 |
+
subscript
|
| 772 |
+
𝑎
|
| 773 |
+
𝑑
|
| 774 |
+
Q(s_{d},a_{d})
|
| 775 |
+
based on whether it reaches the correct answer.
|
| 776 |
+
This score is then back-propagated to each intermediate node along the trajectory
|
| 777 |
+
𝐭
|
| 778 |
+
=
|
| 779 |
+
x
|
| 780 |
+
⊕
|
| 781 |
+
s
|
| 782 |
+
1
|
| 783 |
+
⊕
|
| 784 |
+
s
|
| 785 |
+
2
|
| 786 |
+
⊕
|
| 787 |
+
…
|
| 788 |
+
⊕
|
| 789 |
+
s
|
| 790 |
+
d
|
| 791 |
+
𝐭
|
| 792 |
+
direct-sum
|
| 793 |
+
𝑥
|
| 794 |
+
subscript
|
| 795 |
+
𝑠
|
| 796 |
+
1
|
| 797 |
+
subscript
|
| 798 |
+
𝑠
|
| 799 |
+
2
|
| 800 |
+
…
|
| 801 |
+
subscript
|
| 802 |
+
𝑠
|
| 803 |
+
𝑑
|
| 804 |
+
\mathbf{t}=x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{d}
|
| 805 |
+
. Specifically, for each
|
| 806 |
+
s
|
| 807 |
+
i
|
| 808 |
+
subscript
|
| 809 |
+
𝑠
|
| 810 |
+
𝑖
|
| 811 |
+
s_{i}
|
| 812 |
+
(for
|
| 813 |
+
i
|
| 814 |
+
=
|
| 815 |
+
1
|
| 816 |
+
,
|
| 817 |
+
2
|
| 818 |
+
,
|
| 819 |
+
…
|
| 820 |
+
,
|
| 821 |
+
d
|
| 822 |
+
−
|
| 823 |
+
1
|
| 824 |
+
𝑖
|
| 825 |
+
1
|
| 826 |
+
2
|
| 827 |
+
…
|
| 828 |
+
𝑑
|
| 829 |
+
1
|
| 830 |
+
i=1,2,...,d-1
|
| 831 |
+
), its
|
| 832 |
+
Q
|
| 833 |
+
𝑄
|
| 834 |
+
Q
|
| 835 |
+
value is updated as follows:
|
| 836 |
+
Q
|
| 837 |
+
|
| 838 |
+
(
|
| 839 |
+
s
|
| 840 |
+
i
|
| 841 |
+
,
|
| 842 |
+
a
|
| 843 |
+
i
|
| 844 |
+
)
|
| 845 |
+
=
|
| 846 |
+
Q
|
| 847 |
+
|
| 848 |
+
(
|
| 849 |
+
s
|
| 850 |
+
i
|
| 851 |
+
,
|
| 852 |
+
a
|
| 853 |
+
i
|
| 854 |
+
)
|
| 855 |
+
+
|
| 856 |
+
Q
|
| 857 |
+
|
| 858 |
+
(
|
| 859 |
+
s
|
| 860 |
+
d
|
| 861 |
+
,
|
| 862 |
+
a
|
| 863 |
+
d
|
| 864 |
+
)
|
| 865 |
+
𝑄
|
| 866 |
+
subscript
|
| 867 |
+
𝑠
|
| 868 |
+
𝑖
|
| 869 |
+
subscript
|
| 870 |
+
𝑎
|
| 871 |
+
𝑖
|
| 872 |
+
𝑄
|
| 873 |
+
subscript
|
| 874 |
+
𝑠
|
| 875 |
+
𝑖
|
| 876 |
+
subscript
|
| 877 |
+
𝑎
|
| 878 |
+
𝑖
|
| 879 |
+
𝑄
|
| 880 |
+
subscript
|
| 881 |
+
𝑠
|
| 882 |
+
𝑑
|
| 883 |
+
subscript
|
| 884 |
+
𝑎
|
| 885 |
+
𝑑
|
| 886 |
+
Q(s_{i},a_{i})=Q(s_{i},a_{i})+Q(s_{d},a_{d})
|
| 887 |
+
. To compute the
|
| 888 |
+
Q
|
| 889 |
+
|
| 890 |
+
(
|
| 891 |
+
s
|
| 892 |
+
d
|
| 893 |
+
,
|
| 894 |
+
a
|
| 895 |
+
d
|
| 896 |
+
)
|
| 897 |
+
𝑄
|
| 898 |
+
subscript
|
| 899 |
+
𝑠
|
| 900 |
+
𝑑
|
| 901 |
+
subscript
|
| 902 |
+
𝑎
|
| 903 |
+
𝑑
|
| 904 |
+
Q(s_{d},a_{d})
|
| 905 |
+
for the terminal node, we use the likelihood (confidence) of self-consistency majority voting as the reward value.
|
| 906 |
+
Figure 4:
|
| 907 |
+
The prompt example for mutual reasoning consistency.
|
| 908 |
+
Solution Generation with MCTS Rollout
|
| 909 |
+
. We now describe how our MCTS generates candidate reasoning trajectories. Starting from the initial root node
|
| 910 |
+
s
|
| 911 |
+
0
|
| 912 |
+
subscript
|
| 913 |
+
𝑠
|
| 914 |
+
0
|
| 915 |
+
s_{0}
|
| 916 |
+
, we perform multiple searches consisting of
|
| 917 |
+
selection
|
| 918 |
+
,
|
| 919 |
+
expansion
|
| 920 |
+
,
|
| 921 |
+
simulations
|
| 922 |
+
and
|
| 923 |
+
back-propagation
|
| 924 |
+
. Specifically, the simulation is performed using the default
|
| 925 |
+
rollout
|
| 926 |
+
policy, and to achieve more accurate reward estimation, we perform multiple rollouts. To balance the exploration and exploitation, we use the well-known Upper Confidence Bounds applied to Trees (UCT)
|
| 927 |
+
(Kocsis & Szepesvári,
|
| 928 |
+
2006
|
| 929 |
+
)
|
| 930 |
+
to select each node. This selection process is mathematically represented as:
|
| 931 |
+
UCT
|
| 932 |
+
|
| 933 |
+
(
|
| 934 |
+
s
|
| 935 |
+
,
|
| 936 |
+
a
|
| 937 |
+
)
|
| 938 |
+
=
|
| 939 |
+
Q
|
| 940 |
+
|
| 941 |
+
(
|
| 942 |
+
s
|
| 943 |
+
,
|
| 944 |
+
a
|
| 945 |
+
)
|
| 946 |
+
N
|
| 947 |
+
|
| 948 |
+
(
|
| 949 |
+
s
|
| 950 |
+
,
|
| 951 |
+
a
|
| 952 |
+
)
|
| 953 |
+
+
|
| 954 |
+
c
|
| 955 |
+
|
| 956 |
+
ln
|
| 957 |
+
|
| 958 |
+
N
|
| 959 |
+
p
|
| 960 |
+
|
| 961 |
+
a
|
| 962 |
+
|
| 963 |
+
r
|
| 964 |
+
|
| 965 |
+
e
|
| 966 |
+
|
| 967 |
+
n
|
| 968 |
+
|
| 969 |
+
t
|
| 970 |
+
|
| 971 |
+
(
|
| 972 |
+
s
|
| 973 |
+
)
|
| 974 |
+
N
|
| 975 |
+
|
| 976 |
+
(
|
| 977 |
+
s
|
| 978 |
+
,
|
| 979 |
+
a
|
| 980 |
+
)
|
| 981 |
+
.
|
| 982 |
+
UCT
|
| 983 |
+
𝑠
|
| 984 |
+
𝑎
|
| 985 |
+
𝑄
|
| 986 |
+
𝑠
|
| 987 |
+
𝑎
|
| 988 |
+
𝑁
|
| 989 |
+
𝑠
|
| 990 |
+
𝑎
|
| 991 |
+
𝑐
|
| 992 |
+
subscript
|
| 993 |
+
𝑁
|
| 994 |
+
𝑝
|
| 995 |
+
𝑎
|
| 996 |
+
𝑟
|
| 997 |
+
𝑒
|
| 998 |
+
𝑛
|
| 999 |
+
𝑡
|
| 1000 |
+
𝑠
|
| 1001 |
+
𝑁
|
| 1002 |
+
𝑠
|
| 1003 |
+
𝑎
|
| 1004 |
+
\text{UCT}(s,a)=\frac{Q(s,a)}{N(s,a)}+c\sqrt{\frac{\ln N_{parent}(s)}{N(s,a)}}.
|
| 1005 |
+
where
|
| 1006 |
+
N
|
| 1007 |
+
|
| 1008 |
+
(
|
| 1009 |
+
s
|
| 1010 |
+
,
|
| 1011 |
+
a
|
| 1012 |
+
)
|
| 1013 |
+
𝑁
|
| 1014 |
+
𝑠
|
| 1015 |
+
𝑎
|
| 1016 |
+
N(s,a)
|
| 1017 |
+
is the number of times node
|
| 1018 |
+
s
|
| 1019 |
+
𝑠
|
| 1020 |
+
s
|
| 1021 |
+
has been visited in previous iterations, and
|
| 1022 |
+
N
|
| 1023 |
+
p
|
| 1024 |
+
|
| 1025 |
+
a
|
| 1026 |
+
|
| 1027 |
+
r
|
| 1028 |
+
|
| 1029 |
+
e
|
| 1030 |
+
|
| 1031 |
+
n
|
| 1032 |
+
|
| 1033 |
+
t
|
| 1034 |
+
|
| 1035 |
+
(
|
| 1036 |
+
s
|
| 1037 |
+
)
|
| 1038 |
+
subscript
|
| 1039 |
+
𝑁
|
| 1040 |
+
𝑝
|
| 1041 |
+
𝑎
|
| 1042 |
+
𝑟
|
| 1043 |
+
𝑒
|
| 1044 |
+
𝑛
|
| 1045 |
+
𝑡
|
| 1046 |
+
𝑠
|
| 1047 |
+
N_{parent}(s)
|
| 1048 |
+
represents the visiting count of the parent node of
|
| 1049 |
+
s
|
| 1050 |
+
𝑠
|
| 1051 |
+
s
|
| 1052 |
+
.
|
| 1053 |
+
Q
|
| 1054 |
+
|
| 1055 |
+
(
|
| 1056 |
+
s
|
| 1057 |
+
,
|
| 1058 |
+
a
|
| 1059 |
+
)
|
| 1060 |
+
𝑄
|
| 1061 |
+
𝑠
|
| 1062 |
+
𝑎
|
| 1063 |
+
Q(s,a)
|
| 1064 |
+
is the estimated reward value and will be updated through back-propagation.
|
| 1065 |
+
c
|
| 1066 |
+
𝑐
|
| 1067 |
+
c
|
| 1068 |
+
is a constant that balances exploitation and exploration.
|
| 1069 |
+
Once the search reaches a terminal node, either a terminal state or a predetermined maximum tree depth
|
| 1070 |
+
d
|
| 1071 |
+
𝑑
|
| 1072 |
+
d
|
| 1073 |
+
, we obtain a trajectory from the root to terminal node. We collect all trajectories from the rollout iterations as candidate solutions. The next section explains how we verify each of them.
|
| 1074 |
+
3.3
|
| 1075 |
+
Reasoning Trajectory Selection with Mutual Consistency
|
| 1076 |
+
In traditional MCTS, typically only one trajectory is selected as the final solution based on a specific metric, such as choosing the path with the highest reward from the rollout iterations. Unfortunately, after trying various existing methods, we found it challenging to define a single metric that reliably selects the trajectory containing the correct answer.
|
| 1077 |
+
Therefore, we collect all trajectories and propose mutual reasoning consistency for answer selection.
|
| 1078 |
+
Mutual Reasoning Consistency by Discriminator SLM
|
| 1079 |
+
2
|
| 1080 |
+
. As shown in Fig.
|
| 1081 |
+
2
|
| 1082 |
+
, in addition to the target SLM
|
| 1083 |
+
M
|
| 1084 |
+
𝑀
|
| 1085 |
+
M
|
| 1086 |
+
, we introduce another SLM
|
| 1087 |
+
M
|
| 1088 |
+
^
|
| 1089 |
+
^
|
| 1090 |
+
𝑀
|
| 1091 |
+
\hat{M}
|
| 1092 |
+
to serve as a discriminator, providing external unsupervised feedback for each candidate trajectory.
|
| 1093 |
+
Specifically, for
|
| 1094 |
+
𝐭
|
| 1095 |
+
=
|
| 1096 |
+
x
|
| 1097 |
+
⊕
|
| 1098 |
+
s
|
| 1099 |
+
1
|
| 1100 |
+
⊕
|
| 1101 |
+
s
|
| 1102 |
+
2
|
| 1103 |
+
⊕
|
| 1104 |
+
…
|
| 1105 |
+
⊕
|
| 1106 |
+
s
|
| 1107 |
+
d
|
| 1108 |
+
𝐭
|
| 1109 |
+
direct-sum
|
| 1110 |
+
𝑥
|
| 1111 |
+
subscript
|
| 1112 |
+
𝑠
|
| 1113 |
+
1
|
| 1114 |
+
subscript
|
| 1115 |
+
𝑠
|
| 1116 |
+
2
|
| 1117 |
+
…
|
| 1118 |
+
subscript
|
| 1119 |
+
𝑠
|
| 1120 |
+
𝑑
|
| 1121 |
+
\mathbf{t}=x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{d}
|
| 1122 |
+
, we mask the reasoning steps starting from a randomly sampled step
|
| 1123 |
+
i
|
| 1124 |
+
𝑖
|
| 1125 |
+
i
|
| 1126 |
+
(
|
| 1127 |
+
i
|
| 1128 |
+
<
|
| 1129 |
+
d
|
| 1130 |
+
𝑖
|
| 1131 |
+
𝑑
|
| 1132 |
+
i<d
|
| 1133 |
+
). We then provide the earlier reasoning trajectory
|
| 1134 |
+
𝐭
|
| 1135 |
+
=
|
| 1136 |
+
x
|
| 1137 |
+
⊕
|
| 1138 |
+
s
|
| 1139 |
+
1
|
| 1140 |
+
⊕
|
| 1141 |
+
s
|
| 1142 |
+
2
|
| 1143 |
+
⊕
|
| 1144 |
+
…
|
| 1145 |
+
⊕
|
| 1146 |
+
s
|
| 1147 |
+
i
|
| 1148 |
+
−
|
| 1149 |
+
1
|
| 1150 |
+
𝐭
|
| 1151 |
+
direct-sum
|
| 1152 |
+
𝑥
|
| 1153 |
+
subscript
|
| 1154 |
+
𝑠
|
| 1155 |
+
1
|
| 1156 |
+
subscript
|
| 1157 |
+
𝑠
|
| 1158 |
+
2
|
| 1159 |
+
…
|
| 1160 |
+
subscript
|
| 1161 |
+
𝑠
|
| 1162 |
+
𝑖
|
| 1163 |
+
1
|
| 1164 |
+
\mathbf{t}=x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{i-1}
|
| 1165 |
+
as a prompt to
|
| 1166 |
+
M
|
| 1167 |
+
^
|
| 1168 |
+
^
|
| 1169 |
+
𝑀
|
| 1170 |
+
\hat{M}
|
| 1171 |
+
to complete the remaining steps for the question. Due to the provision of the earlier
|
| 1172 |
+
i
|
| 1173 |
+
−
|
| 1174 |
+
1
|
| 1175 |
+
𝑖
|
| 1176 |
+
1
|
| 1177 |
+
i-1
|
| 1178 |
+
reasoning steps as a hint, we reduce the difficulty, thereby increasing the likelihood that SLM
|
| 1179 |
+
M
|
| 1180 |
+
^
|
| 1181 |
+
^
|
| 1182 |
+
𝑀
|
| 1183 |
+
\hat{M}
|
| 1184 |
+
can provide the correct answer.
|
| 1185 |
+
As shown in Fig.
|
| 1186 |
+
4
|
| 1187 |
+
, we compare whether the answer completed by
|
| 1188 |
+
M
|
| 1189 |
+
^
|
| 1190 |
+
^
|
| 1191 |
+
𝑀
|
| 1192 |
+
\hat{M}
|
| 1193 |
+
matches the original trajectory
|
| 1194 |
+
𝐭
|
| 1195 |
+
𝐭
|
| 1196 |
+
\mathbf{t}
|
| 1197 |
+
. If they are consistent, we consider
|
| 1198 |
+
t
|
| 1199 |
+
𝑡
|
| 1200 |
+
t
|
| 1201 |
+
as an validate trajectory for final selection.
|
| 1202 |
+
We provide an intuitive explanation to illustrate the rational behind our approach. Consider students solving a problem without a teacher’s feedback. A student (SLM
|
| 1203 |
+
1
|
| 1204 |
+
) unsure of their solution might ask a peer (SLM
|
| 1205 |
+
2
|
| 1206 |
+
) to review their reasoning. If the peer, given the same initial steps, arrives at the same answer, the student gains confidence in their solution. This peer verification process reflects the mutual reasoning consistency we aim to achieve.
|
| 1207 |
+
Final Trajectory Selection by SLM
|
| 1208 |
+
1
|
| 1209 |
+
. After applying mutual reasoning consistency to all candidate trajectories, we return to the target SLM
|
| 1210 |
+
M
|
| 1211 |
+
𝑀
|
| 1212 |
+
M
|
| 1213 |
+
to select the final trajectory from the validated ones. We compute each trajectory’s final score by multiplying its reward with the terminal node’s confidence score achieved from rollouts. The trajectory with the highest final score is chosen as the solution.
|
| 1214 |
+
Table 2:
|
| 1215 |
+
rStar greatly improves reasoning accuracy across various SLMs and tasks. rStar (generator@maj): uses majority voting for answer verification to show the MCTS generator’s effectiveness.
|
| 1216 |
+
Method
|
| 1217 |
+
LLaMA2-7B
|
| 1218 |
+
Mistral-7B
|
| 1219 |
+
LLaMA3-8B
|
| 1220 |
+
LLaMA3-8B-Instruct
|
| 1221 |
+
Phi3-mini-4k
|
| 1222 |
+
GSM8K
|
| 1223 |
+
Zero-shot CoT
|
| 1224 |
+
1.44
|
| 1225 |
+
17.89
|
| 1226 |
+
22.66
|
| 1227 |
+
68.38
|
| 1228 |
+
20.17
|
| 1229 |
+
Few-shot CoT
|
| 1230 |
+
12.51
|
| 1231 |
+
36.46
|
| 1232 |
+
47.23
|
| 1233 |
+
74.53
|
| 1234 |
+
83.45
|
| 1235 |
+
SC@maj8
|
| 1236 |
+
15.31
|
| 1237 |
+
42.91
|
| 1238 |
+
54.21
|
| 1239 |
+
78.39
|
| 1240 |
+
86.35
|
| 1241 |
+
SC@maj64
|
| 1242 |
+
20.77
|
| 1243 |
+
52.84
|
| 1244 |
+
64.37
|
| 1245 |
+
83.24
|
| 1246 |
+
88.02
|
| 1247 |
+
SC@maj128
|
| 1248 |
+
23.05
|
| 1249 |
+
57.25
|
| 1250 |
+
67.55
|
| 1251 |
+
84.69
|
| 1252 |
+
88.68
|
| 1253 |
+
ToT
|
| 1254 |
+
12.96
|
| 1255 |
+
38.89
|
| 1256 |
+
36.01
|
| 1257 |
+
69.07
|
| 1258 |
+
79.68
|
| 1259 |
+
RAP
|
| 1260 |
+
24.34
|
| 1261 |
+
56.25
|
| 1262 |
+
57.99
|
| 1263 |
+
80.59
|
| 1264 |
+
81.88
|
| 1265 |
+
rStar (generator @maj)
|
| 1266 |
+
27.22
|
| 1267 |
+
64.59
|
| 1268 |
+
74.38
|
| 1269 |
+
88.70
|
| 1270 |
+
90.44
|
| 1271 |
+
rStar
|
| 1272 |
+
63.91
|
| 1273 |
+
81.88
|
| 1274 |
+
85.52
|
| 1275 |
+
91.13
|
| 1276 |
+
90.67
|
| 1277 |
+
GSM-Hard
|
| 1278 |
+
Zero-shot CoT
|
| 1279 |
+
0.83
|
| 1280 |
+
5.16
|
| 1281 |
+
6.44
|
| 1282 |
+
14.94
|
| 1283 |
+
33.73
|
| 1284 |
+
Few-shot CoT
|
| 1285 |
+
3.71
|
| 1286 |
+
13.57
|
| 1287 |
+
13.80
|
| 1288 |
+
25.63
|
| 1289 |
+
40.63
|
| 1290 |
+
SC@maj8
|
| 1291 |
+
4.39
|
| 1292 |
+
17.36
|
| 1293 |
+
18.20
|
| 1294 |
+
28.51
|
| 1295 |
+
42.00
|
| 1296 |
+
SC@maj64
|
| 1297 |
+
6.52
|
| 1298 |
+
22.59
|
| 1299 |
+
23.73
|
| 1300 |
+
30.33
|
| 1301 |
+
44.80
|
| 1302 |
+
SC@maj128
|
| 1303 |
+
6.89
|
| 1304 |
+
25.01
|
| 1305 |
+
25.47
|
| 1306 |
+
31.16
|
| 1307 |
+
45.56
|
| 1308 |
+
ToT
|
| 1309 |
+
2.35
|
| 1310 |
+
11.47
|
| 1311 |
+
10.61
|
| 1312 |
+
19.64
|
| 1313 |
+
32.68
|
| 1314 |
+
RAP
|
| 1315 |
+
7.28
|
| 1316 |
+
22.52
|
| 1317 |
+
18.95
|
| 1318 |
+
29.64
|
| 1319 |
+
40.94
|
| 1320 |
+
rStar (generator @maj)
|
| 1321 |
+
8.64
|
| 1322 |
+
29.26
|
| 1323 |
+
26.76
|
| 1324 |
+
33.35
|
| 1325 |
+
46.55
|
| 1326 |
+
rStar
|
| 1327 |
+
18.57
|
| 1328 |
+
37.91
|
| 1329 |
+
32.97
|
| 1330 |
+
37.53
|
| 1331 |
+
46.55
|
| 1332 |
+
SVAMP
|
| 1333 |
+
Zero-shot CoT
|
| 1334 |
+
8.90
|
| 1335 |
+
26.10
|
| 1336 |
+
40.20
|
| 1337 |
+
70.90
|
| 1338 |
+
84.70
|
| 1339 |
+
Few-shot CoT
|
| 1340 |
+
48.10
|
| 1341 |
+
72.80
|
| 1342 |
+
76.90
|
| 1343 |
+
89.20
|
| 1344 |
+
92.80
|
| 1345 |
+
SC@maj8
|
| 1346 |
+
49.90
|
| 1347 |
+
74.60
|
| 1348 |
+
79.10
|
| 1349 |
+
89.20
|
| 1350 |
+
93.50
|
| 1351 |
+
SC@maj64
|
| 1352 |
+
54.10
|
| 1353 |
+
76.70
|
| 1354 |
+
80.70
|
| 1355 |
+
90.50
|
| 1356 |
+
93.30
|
| 1357 |
+
SC@maj128
|
| 1358 |
+
54.50
|
| 1359 |
+
76.60
|
| 1360 |
+
80.80
|
| 1361 |
+
90.60
|
| 1362 |
+
93.70
|
| 1363 |
+
ToT
|
| 1364 |
+
33.40
|
| 1365 |
+
56.30
|
| 1366 |
+
62.20
|
| 1367 |
+
79.80
|
| 1368 |
+
84.90
|
| 1369 |
+
RAP
|
| 1370 |
+
41.00
|
| 1371 |
+
71.80
|
| 1372 |
+
73.10
|
| 1373 |
+
85.70
|
| 1374 |
+
91.50
|
| 1375 |
+
rStar (generator @maj)
|
| 1376 |
+
60.30
|
| 1377 |
+
83.10
|
| 1378 |
+
86.20
|
| 1379 |
+
91.89
|
| 1380 |
+
93.80
|
| 1381 |
+
rStar
|
| 1382 |
+
74.90
|
| 1383 |
+
86.40
|
| 1384 |
+
90.00
|
| 1385 |
+
94.29
|
| 1386 |
+
94.10
|
| 1387 |
+
StrategyQA
|
| 1388 |
+
Zero-shot CoT
|
| 1389 |
+
52.67
|
| 1390 |
+
57.20
|
| 1391 |
+
41.48
|
| 1392 |
+
57.21
|
| 1393 |
+
54.68
|
| 1394 |
+
Few-shot CoT
|
| 1395 |
+
58.82
|
| 1396 |
+
65.65
|
| 1397 |
+
64.05
|
| 1398 |
+
68.41
|
| 1399 |
+
63.61
|
| 1400 |
+
SC@maj8
|
| 1401 |
+
59.10
|
| 1402 |
+
65.50
|
| 1403 |
+
63.76
|
| 1404 |
+
68.26
|
| 1405 |
+
64.34
|
| 1406 |
+
SC@maj64
|
| 1407 |
+
58.51
|
| 1408 |
+
63.61
|
| 1409 |
+
63.46
|
| 1410 |
+
67.39
|
| 1411 |
+
62.74
|
| 1412 |
+
SC@maj128
|
| 1413 |
+
58.37
|
| 1414 |
+
62.01
|
| 1415 |
+
63.31
|
| 1416 |
+
66.67
|
| 1417 |
+
59.53
|
| 1418 |
+
ToT
|
| 1419 |
+
45.27
|
| 1420 |
+
55.75
|
| 1421 |
+
57.64
|
| 1422 |
+
60.41
|
| 1423 |
+
40.47
|
| 1424 |
+
RAP
|
| 1425 |
+
59.68
|
| 1426 |
+
64.48
|
| 1427 |
+
63.32
|
| 1428 |
+
68.71
|
| 1429 |
+
60.26
|
| 1430 |
+
rStar (generator @maj)
|
| 1431 |
+
61.57
|
| 1432 |
+
69.43
|
| 1433 |
+
65.50
|
| 1434 |
+
71.47
|
| 1435 |
+
65.50
|
| 1436 |
+
rStar
|
| 1437 |
+
67.25
|
| 1438 |
+
70.31
|
| 1439 |
+
67.69
|
| 1440 |
+
71.57
|
| 1441 |
+
67.25
|
| 1442 |
+
4
|
| 1443 |
+
Experiments
|
| 1444 |
+
4.1
|
| 1445 |
+
Setup
|
| 1446 |
+
Models and Datasets
|
| 1447 |
+
. rStar is a general approach applicable to various LLMs and reasoning tasks. We evaluate 5 SLMs: Phi3-mini (3.8B)
|
| 1448 |
+
(Abdin et al.,
|
| 1449 |
+
2024
|
| 1450 |
+
)
|
| 1451 |
+
, LLaMA2-7B, Mistral-7B
|
| 1452 |
+
(Jiang et al.,
|
| 1453 |
+
2023
|
| 1454 |
+
)
|
| 1455 |
+
, LLaMA3-8B, and LLaMA3-8B-Instruct
|
| 1456 |
+
(Meta,
|
| 1457 |
+
2024
|
| 1458 |
+
)
|
| 1459 |
+
. We test across 5 reasoning tasks, including 4 mathematical tasks (GSM8K
|
| 1460 |
+
(Cobbe et al.,
|
| 1461 |
+
2021
|
| 1462 |
+
)
|
| 1463 |
+
, GSM-Hard
|
| 1464 |
+
(Gao et al.,
|
| 1465 |
+
2022
|
| 1466 |
+
)
|
| 1467 |
+
, MATH
|
| 1468 |
+
(Hendrycks et al.,
|
| 1469 |
+
2021
|
| 1470 |
+
)
|
| 1471 |
+
, SVAMP
|
| 1472 |
+
(Patel et al.,
|
| 1473 |
+
2021
|
| 1474 |
+
)
|
| 1475 |
+
) and one commonsense reasoning task (StrategyQA
|
| 1476 |
+
(Geva et al.,
|
| 1477 |
+
2021
|
| 1478 |
+
)
|
| 1479 |
+
).
|
| 1480 |
+
Implementation Details.
|
| 1481 |
+
In the trajectory self-generation stage, we augment each target SLM with our MCTS, performing 32 rollouts. Except for MATH, where we set the depth
|
| 1482 |
+
d
|
| 1483 |
+
𝑑
|
| 1484 |
+
d
|
| 1485 |
+
to 8, all other tasks have a
|
| 1486 |
+
d
|
| 1487 |
+
𝑑
|
| 1488 |
+
d
|
| 1489 |
+
=5. Actions
|
| 1490 |
+
A
|
| 1491 |
+
1
|
| 1492 |
+
subscript
|
| 1493 |
+
𝐴
|
| 1494 |
+
1
|
| 1495 |
+
A_{1}
|
| 1496 |
+
and
|
| 1497 |
+
A
|
| 1498 |
+
3
|
| 1499 |
+
subscript
|
| 1500 |
+
𝐴
|
| 1501 |
+
3
|
| 1502 |
+
A_{3}
|
| 1503 |
+
have a maximum of 5 nodes per depth, while the other actions have a default node count of 1. In the trajectory discrimination stage, we use Phi3-mini-4k as the discriminator, which has only 3.8B parameters, for effective inference. Moreover, the discriminator performs inference in a parallelized manner, making the verification process highly efficient.
|
| 1504 |
+
Notably, when Phi3 is the target SLM, it performs self-discrimination. For a given trajectory, we randomly split it between 20% and 80% of its steps, providing the first half of the steps as input to the discriminator SLM, which then completes the remaining steps.
|
| 1505 |
+
Detailed prompts are available in appendix
|
| 1506 |
+
A.3
|
| 1507 |
+
.
|
| 1508 |
+
4.2
|
| 1509 |
+
Main Results
|
| 1510 |
+
Baselines
|
| 1511 |
+
. We compare rStar against three strong baseline types:
|
| 1512 |
+
(i)
|
| 1513 |
+
single-round CoT prompting
|
| 1514 |
+
, including zero-shot CoT
|
| 1515 |
+
(Kojima et al.,
|
| 1516 |
+
2022
|
| 1517 |
+
)
|
| 1518 |
+
and few-shot CoT
|
| 1519 |
+
(Wei et al.,
|
| 1520 |
+
2022
|
| 1521 |
+
)
|
| 1522 |
+
;
|
| 1523 |
+
(ii)
|
| 1524 |
+
multi-round CoT prompting
|
| 1525 |
+
using the widely adopted self-consistency (SC) method
|
| 1526 |
+
(Wang et al.,
|
| 1527 |
+
2023
|
| 1528 |
+
)
|
| 1529 |
+
. We sample answers 8, 64, and 128 times, employing majority voting for answer selection; and
|
| 1530 |
+
(iii)
|
| 1531 |
+
multi-round self-improving approaches
|
| 1532 |
+
. For this, we select ToT
|
| 1533 |
+
(Yao et al.,
|
| 1534 |
+
2024
|
| 1535 |
+
)
|
| 1536 |
+
and RAP
|
| 1537 |
+
(Hao et al.,
|
| 1538 |
+
2023
|
| 1539 |
+
)
|
| 1540 |
+
as baselines, using BFS and MCTS for tree search, respectively. Note that the action in ToT corresponds to our action
|
| 1541 |
+
A
|
| 1542 |
+
1
|
| 1543 |
+
subscript
|
| 1544 |
+
𝐴
|
| 1545 |
+
1
|
| 1546 |
+
A_{1}
|
| 1547 |
+
, while RAP corresponds to our action
|
| 1548 |
+
A
|
| 1549 |
+
3
|
| 1550 |
+
subscript
|
| 1551 |
+
𝐴
|
| 1552 |
+
3
|
| 1553 |
+
A_{3}
|
| 1554 |
+
. For the answer selection, we follow their original implementations.
|
| 1555 |
+
Results on diverse reasoning benchmarks
|
| 1556 |
+
. We start by evaluating the effectiveness of rStar on general reasoning benchmarks. Table
|
| 1557 |
+
2
|
| 1558 |
+
compares its accuracy with state-of-the-art baselines on diverse SLMs and reasoning datasets. To demonstrate the effectiveness of our generator, we also provide the accuracy of rStar (gen. @maj), which do not apply our discriminator and use majority voting for answer verification. We highlight three key observations:
|
| 1559 |
+
(1)
|
| 1560 |
+
SLMs empowered with rStar demonstrate highly capable problem-solving abilities. For example, LLaMA2-7B originally had an accuracy of only 12.51% on GSM8K using few-shot CoT. However, with improvements from rStar, its accuracy increased to 63.91%, nearly matching the accuracy achieved with fine-tuning as shown in Fig.
|
| 1561 |
+
1
|
| 1562 |
+
. Similarly, Mistral with rStar can even outperform fine-tuned MetaMath by +4.18%. This improvement shows that SLMs already have strong reasoning capabilities but need guidance to generate and select the correct solutions.
|
| 1563 |
+
(2)
|
| 1564 |
+
rStar consistently improves the reasoning accuracy of various evaluated SLMs across different tasks to a state-of-the-art level. In contrast, none of the baseline approaches consistently perform well across all four benchmarks. For example, while SC excels in three mathematical tasks, it is less effective on the logical reasoning task of StrategyQA. Specifically, SC with more sampling can even lower the score on StrategyQA. RAP performs better than SC on StrategyQA but falls short compared to SC on most mathematical reasoning tasks.
|
| 1565 |
+
(3)
|
| 1566 |
+
Even without our proposed discriminator for reasoning trajectory verification, our MCTS generator demonstrates greater effectiveness in improving reasoning accuracy for SLMs compared to existing multi-round inference baselines. For example, rStar (generator @maj) achieves up to 2.88%-16.39% higher accuracy than RAP, 10.60%- 38.37% higher accuracy than ToT, and 1.69% - 7.34% higher accuracy than SC on the GSM8K dataset.
|
| 1567 |
+
Table 3:
|
| 1568 |
+
Reasoning performance comparison on the challenging MATH-500 dataset. Due to the extensive LaTeX syntax in the dataset, which is challenging for pre-trained LLMs in instruction following, we evaluate only on LLaMA3-8B-instruct and Phi3-Mini-4k-Instruct.
|
| 1569 |
+
Method
|
| 1570 |
+
LLaMA3-8b-Instruct
|
| 1571 |
+
Phi3-mini-4k
|
| 1572 |
+
Zeroshot CoT
|
| 1573 |
+
5.80
|
| 1574 |
+
3.60
|
| 1575 |
+
Fewshot CoT
|
| 1576 |
+
17.80
|
| 1577 |
+
32.20
|
| 1578 |
+
SC@maj8
|
| 1579 |
+
30.00
|
| 1580 |
+
40.40
|
| 1581 |
+
SC@maj64
|
| 1582 |
+
33.00
|
| 1583 |
+
45.20
|
| 1584 |
+
SC@maj128
|
| 1585 |
+
33.80
|
| 1586 |
+
45.60
|
| 1587 |
+
ToT
|
| 1588 |
+
13.60
|
| 1589 |
+
18.20
|
| 1590 |
+
RAP
|
| 1591 |
+
18.80
|
| 1592 |
+
27.80
|
| 1593 |
+
rStar (generator @maj)
|
| 1594 |
+
38.30
|
| 1595 |
+
48.40
|
| 1596 |
+
rStar
|
| 1597 |
+
42.94
|
| 1598 |
+
48.60
|
| 1599 |
+
Figure 5:
|
| 1600 |
+
Performance comparison on the GSM8K dataset under different number of rollouts. rStar can significantly improve reasoning accuracy with just 2 rollouts.
|
| 1601 |
+
Results on challenging mathematical dataset
|
| 1602 |
+
. We also evaluate the effectiveness of rStar on more challenging mathematical datasets. In particular, we select the GSM-Hard and MATH datasets. Following
|
| 1603 |
+
(Wang et al.,
|
| 1604 |
+
2024b
|
| 1605 |
+
; Lightman et al.,
|
| 1606 |
+
2023
|
| 1607 |
+
)
|
| 1608 |
+
, we use MATH-500, a subset of representative problems from the MATH dataset, to speedup the evaluation. As shown in Table
|
| 1609 |
+
2
|
| 1610 |
+
and Table
|
| 1611 |
+
3
|
| 1612 |
+
, rStar
|
| 1613 |
+
is capable of significantly improve the reasoning accuracy of SLMs on these challenging mathematical datasets.
|
| 1614 |
+
Remarkably, when compared to SOTA baselines, we observe a significant improvements of up to +12.9% and +9.14% on GSM-Hard and MATH-500, respectively.
|
| 1615 |
+
4.3
|
| 1616 |
+
Ablation Study
|
| 1617 |
+
Table 4:
|
| 1618 |
+
Ablation study on the effectiveness of our MCTS generator. Ours+self-eval: we apply self-evaluation to prompt model for self-rewarding each intermediate action in our generator.
|
| 1619 |
+
Generator
|
| 1620 |
+
LLaMA3-8B
|
| 1621 |
+
LLaMA3-8B-Instruct
|
| 1622 |
+
GSM8K
|
| 1623 |
+
StrategyQA
|
| 1624 |
+
GSM8K
|
| 1625 |
+
StrategyQA
|
| 1626 |
+
Answer verification
|
| 1627 |
+
Answer verification
|
| 1628 |
+
Answer verification
|
| 1629 |
+
Answer verification
|
| 1630 |
+
Maj
|
| 1631 |
+
Ours
|
| 1632 |
+
Maj
|
| 1633 |
+
Ours
|
| 1634 |
+
Maj
|
| 1635 |
+
Ours
|
| 1636 |
+
Maj
|
| 1637 |
+
Ours
|
| 1638 |
+
RAP
|
| 1639 |
+
56.56
|
| 1640 |
+
57.31
|
| 1641 |
+
62.30
|
| 1642 |
+
64.63
|
| 1643 |
+
81.35
|
| 1644 |
+
84.69
|
| 1645 |
+
69.43
|
| 1646 |
+
70.60
|
| 1647 |
+
SC (@128)
|
| 1648 |
+
67.55
|
| 1649 |
+
85.06
|
| 1650 |
+
63.31
|
| 1651 |
+
65.65
|
| 1652 |
+
84.69
|
| 1653 |
+
89.99
|
| 1654 |
+
66.67
|
| 1655 |
+
68.56
|
| 1656 |
+
Ours+Self-eval
|
| 1657 |
+
70.28
|
| 1658 |
+
82.18
|
| 1659 |
+
65.07
|
| 1660 |
+
66.22
|
| 1661 |
+
88.07
|
| 1662 |
+
89.92
|
| 1663 |
+
69.28
|
| 1664 |
+
69.43
|
| 1665 |
+
Ours
|
| 1666 |
+
74.38
|
| 1667 |
+
85.52
|
| 1668 |
+
65.50
|
| 1669 |
+
67.69
|
| 1670 |
+
88.70
|
| 1671 |
+
91.13
|
| 1672 |
+
71.47
|
| 1673 |
+
71.57
|
| 1674 |
+
Table 5:
|
| 1675 |
+
Ablation study on discriminator effectiveness. We evaluate accuracy on GSM8K.
|
| 1676 |
+
Left
|
| 1677 |
+
: Our discriminator consistently outperforms others in verifying solution trajectories generated by different generators.
|
| 1678 |
+
Right
|
| 1679 |
+
: The ablation study on the choice of discriminator model.
|
| 1680 |
+
Discriminator
|
| 1681 |
+
LLaMA3-8B
|
| 1682 |
+
LLaMA3-8B-Instruct
|
| 1683 |
+
Generator
|
| 1684 |
+
Generator
|
| 1685 |
+
SC
|
| 1686 |
+
Ours
|
| 1687 |
+
SC
|
| 1688 |
+
Ours
|
| 1689 |
+
Maj
|
| 1690 |
+
67.55
|
| 1691 |
+
74.38
|
| 1692 |
+
84.69
|
| 1693 |
+
88.70
|
| 1694 |
+
Self-verification
|
| 1695 |
+
74.00
|
| 1696 |
+
75.52
|
| 1697 |
+
83.02
|
| 1698 |
+
86.63
|
| 1699 |
+
Ours
|
| 1700 |
+
85.06
|
| 1701 |
+
85.52
|
| 1702 |
+
89.99
|
| 1703 |
+
91.13
|
| 1704 |
+
Model
|
| 1705 |
+
Discriminator SLM
|
| 1706 |
+
Accuracy
|
| 1707 |
+
LLaMA3-8B-Instruct
|
| 1708 |
+
Maj
|
| 1709 |
+
88.70
|
| 1710 |
+
LLaMA3-8B-Instruct
|
| 1711 |
+
88.78
|
| 1712 |
+
LLaMA3.1-8B-Instruct
|
| 1713 |
+
89.52
|
| 1714 |
+
Phi3-Mini-Instruct
|
| 1715 |
+
91.13
|
| 1716 |
+
GPT-4 (2024-05-01)
|
| 1717 |
+
92.57
|
| 1718 |
+
Effectiveness under different rollouts
|
| 1719 |
+
. rStar uses a rollout policy for MCTS tree expansion. More rollouts generate more candidate solution trajectories but increase inference cost. In Fig.
|
| 1720 |
+
5
|
| 1721 |
+
, we compare the accuracy of SC, RAP, and our rStar across different rollouts on GSM8K. For SC, we sample solutions based on each number of rollouts and use majority voting to select the answer. We highlight two key observations:
|
| 1722 |
+
(1)
|
| 1723 |
+
Even with just 2 rollouts, rStar significantly improves reasoning accuracy for SLMs, demonstrating its effectiveness;
|
| 1724 |
+
(2)
|
| 1725 |
+
Both rStar and SC benefit from more rollouts, whereas RAP tends to saturate and even decline after 4 rollouts on LLaMA3-8B-Instruct. One reason is that the single-type action space in RAP limits the effective MCTS exploration.
|
| 1726 |
+
The effectiveness of MCTS generator
|
| 1727 |
+
. We compare our MCTS generator with three baselines: (i) the MCTS generator used in RAP; (ii) SC with 128 randomly sampled solutions; and (iii) our generator with Self-evaluation, a popular technique that self-evaluates the reward score for each action. Baseline (iii) specifically evaluates the effectiveness of our reward function.
|
| 1728 |
+
To isolate the impact of answer verification methods, each generator is evaluated under both majority voting and our discriminator for trajectory selection. As shown in Table
|
| 1729 |
+
4
|
| 1730 |
+
, our generator consistently outperforms the baseline generators across different answer verification methods. More, we demonstrate the effectiveness of our SLM-tailored reward function, as self-evaluation reduces our generator’s accuracy.
|
| 1731 |
+
The effectiveness of discriminator
|
| 1732 |
+
. We setup two experiments for evaluation. First, we compare our discrimination approach with two baselines: the majority voting and self-verification
|
| 1733 |
+
(Weng et al.,
|
| 1734 |
+
2023
|
| 1735 |
+
)
|
| 1736 |
+
. Specifically, we follow the key idea in
|
| 1737 |
+
Weng et al. (
|
| 1738 |
+
2023
|
| 1739 |
+
)
|
| 1740 |
+
to prompt the SLM (i.e., generator SLM) to self-verify the correctness of each trajectory. To demonstrate the generalization ability of our discriminator, we used candidate solutions from different generators for evaluation.
|
| 1741 |
+
As shown in Table
|
| 1742 |
+
5
|
| 1743 |
+
(Left), our discriminator significantly improves reasoning accuracy when performing answer verification on trajectories generated by different generators. Similar to the previous self-evaluation experiment, self-verification on SLMs is ineffective.
|
| 1744 |
+
Second, we study the impact of discriminator model selection. Our current discriminator models are all Phi3-Mini-Instruct. We tested various LLMs, both stronger and weaker, as discriminators for LLaMA3-8B-Instruct. As shown in Table
|
| 1745 |
+
5
|
| 1746 |
+
(Right), the choice of discriminator model generally does not affect the effectiveness of our mutual reasoning consistency for answer verification. Notably, using the powerful GPT-4 as the discriminator only slightly improves performance (91.13% to 92.57%), demonstrating that mutual reasoning consistency can effectively verify answers using SLMs.
|
| 1747 |
+
5
|
| 1748 |
+
Conclusion
|
| 1749 |
+
In this work, we present rStar, a generator-discriminator self-play approach that significantly grow the reasoning capabilities for SLMs at the inference time. Our approach reveals that SLMs, such as LLaMA2-7B, already exhibit strong reasoning capabilities prior to domain specialized supervised fine-tuning. rStar achieves state-of-the-art performance across five SLMs and five diverse reasoning tasks, substantially outperforming existing multi-round prompting and self-improvement approaches. Furthermore, we conduct extensive ablation studies and analysis, contributing to the development of more advanced SLM self-improved reasoning.
|
| 1750 |
+
References
|
| 1751 |
+
Abdin et al. (2024)
|
| 1752 |
+
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah,
|
| 1753 |
+
Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl,
|
| 1754 |
+
et al.
|
| 1755 |
+
Phi-3 technical report: A highly capable language model locally on
|
| 1756 |
+
your phone.
|
| 1757 |
+
arXiv preprint arXiv:2404.14219
|
| 1758 |
+
, 2024.
|
| 1759 |
+
Besta et al. (2024)
|
| 1760 |
+
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal
|
| 1761 |
+
Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert
|
| 1762 |
+
Niewiadomski, Piotr Nyczyk, et al.
|
| 1763 |
+
Graph of thoughts: Solving elaborate problems with large language
|
| 1764 |
+
models.
|
| 1765 |
+
In
|
| 1766 |
+
Proceedings of the AAAI Conference on Artificial
|
| 1767 |
+
Intelligence
|
| 1768 |
+
, volume 38, pp. 17682–17690, 2024.
|
| 1769 |
+
Brown et al. (2024)
|
| 1770 |
+
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le,
|
| 1771 |
+
Christopher Ré, and Azalia Mirhoseini.
|
| 1772 |
+
Large language monkeys: Scaling inference compute with repeated
|
| 1773 |
+
sampling.
|
| 1774 |
+
arXiv preprint arXiv:2407.21787
|
| 1775 |
+
, 2024.
|
| 1776 |
+
Chen et al. (2024a)
|
| 1777 |
+
Guoxin Chen, Minpeng Liao, Chengxi Li, and Kai Fan.
|
| 1778 |
+
Alphamath almost zero: process supervision without process,
|
| 1779 |
+
2024a.
|
| 1780 |
+
Chen et al. (2022)
|
| 1781 |
+
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen.
|
| 1782 |
+
Program of thoughts prompting: Disentangling computation from
|
| 1783 |
+
reasoning for numerical reasoning tasks.
|
| 1784 |
+
arXiv preprint arXiv:2211.12588
|
| 1785 |
+
, 2022.
|
| 1786 |
+
Chen et al. (2024b)
|
| 1787 |
+
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu.
|
| 1788 |
+
Self-play fine-tuning converts weak language models to strong
|
| 1789 |
+
language models.
|
| 1790 |
+
arXiv preprint arXiv:2401.01335
|
| 1791 |
+
, 2024b.
|
| 1792 |
+
Cobbe et al. (2021)
|
| 1793 |
+
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz
|
| 1794 |
+
Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano,
|
| 1795 |
+
et al.
|
| 1796 |
+
Training verifiers to solve math word problems.
|
| 1797 |
+
arXiv preprint arXiv:2110.14168
|
| 1798 |
+
, 2021.
|
| 1799 |
+
Ding et al. (2023)
|
| 1800 |
+
Ruomeng Ding, Chaoyun Zhang, Lu Wang, Yong Xu, Minghua Ma, Wei Zhang, Si Qin,
|
| 1801 |
+
Saravan Rajmohan, Qingwei Lin, and Dongmei Zhang.
|
| 1802 |
+
Everything of thoughts: Defying the law of penrose triangle for
|
| 1803 |
+
thought generation.
|
| 1804 |
+
arXiv preprint arXiv:2311.04254
|
| 1805 |
+
, 2023.
|
| 1806 |
+
Feng et al. (2023)
|
| 1807 |
+
Xidong Feng, Ziyu Wan, Muning Wen, Ying Wen, Weinan Zhang, and Jun Wang.
|
| 1808 |
+
Alphazero-like tree-search can guide large language model decoding
|
| 1809 |
+
and training.
|
| 1810 |
+
arXiv preprint arXiv:2309.17179
|
| 1811 |
+
, 2023.
|
| 1812 |
+
Forsman (2024)
|
| 1813 |
+
Anton Forsman.
|
| 1814 |
+
Analyzing the performance of self-refine on different large language
|
| 1815 |
+
models.
|
| 1816 |
+
2024.
|
| 1817 |
+
URL
|
| 1818 |
+
https://github.com/anforsm/self-refine/blob/main/report.pdf
|
| 1819 |
+
.
|
| 1820 |
+
Gao et al. (2022)
|
| 1821 |
+
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie
|
| 1822 |
+
Callan, and Graham Neubig.
|
| 1823 |
+
Pal: Program-aided language models.
|
| 1824 |
+
arXiv preprint arXiv:2211.10435
|
| 1825 |
+
, 2022.
|
| 1826 |
+
Gero et al. (2023)
|
| 1827 |
+
Zelalem Gero, Chandan Singh, Hao Cheng, Tristan Naumann, Michel Galley,
|
| 1828 |
+
Jianfeng Gao, and Hoifung Poon.
|
| 1829 |
+
Self-verification improves few-shot clinical information extraction.
|
| 1830 |
+
arXiv preprint arXiv:2306.00024
|
| 1831 |
+
, 2023.
|
| 1832 |
+
Geva et al. (2021)
|
| 1833 |
+
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan
|
| 1834 |
+
Berant.
|
| 1835 |
+
Did aristotle use a laptop? a question answering benchmark with
|
| 1836 |
+
implicit reasoning strategies.
|
| 1837 |
+
Transactions of the Association for Computational Linguistics
|
| 1838 |
+
,
|
| 1839 |
+
9:346–361, 2021.
|
| 1840 |
+
URL
|
| 1841 |
+
https://huggingface.co/datasets/ChilleD/StrategyQA
|
| 1842 |
+
.
|
| 1843 |
+
Gou et al. (2023)
|
| 1844 |
+
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yujiu Yang, Minlie Huang, Nan Duan,
|
| 1845 |
+
Weizhu Chen, et al.
|
| 1846 |
+
Tora: A tool-integrated reasoning agent for mathematical problem
|
| 1847 |
+
solving.
|
| 1848 |
+
arXiv preprint arXiv:2309.17452
|
| 1849 |
+
, 2023.
|
| 1850 |
+
Hao et al. (2023)
|
| 1851 |
+
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and
|
| 1852 |
+
Zhiting Hu.
|
| 1853 |
+
Reasoning with language model is planning with world model.
|
| 1854 |
+
arXiv preprint arXiv:2305.14992
|
| 1855 |
+
, 2023.
|
| 1856 |
+
Hendrycks et al. (2021)
|
| 1857 |
+
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric
|
| 1858 |
+
Tang, Dawn Song, and Jacob Steinhardt.
|
| 1859 |
+
Measuring mathematical problem solving with the math dataset.
|
| 1860 |
+
arXiv preprint arXiv:2103.03874
|
| 1861 |
+
, 2021.
|
| 1862 |
+
Huang et al. (2023)
|
| 1863 |
+
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu,
|
| 1864 |
+
Xinying Song, and Denny Zhou.
|
| 1865 |
+
Large language models cannot self-correct reasoning yet.
|
| 1866 |
+
arXiv preprint arXiv:2310.01798
|
| 1867 |
+
, 2023.
|
| 1868 |
+
Jiang et al. (2023)
|
| 1869 |
+
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford,
|
| 1870 |
+
Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel,
|
| 1871 |
+
Guillaume Lample, Lucile Saulnier, et al.
|
| 1872 |
+
Mistral 7b.
|
| 1873 |
+
arXiv preprint arXiv:2310.06825
|
| 1874 |
+
, 2023.
|
| 1875 |
+
Kang et al. (2024)
|
| 1876 |
+
Jikun Kang, Xin Zhe Li, Xi Chen, Amirreza Kazemi, and Boxing Chen.
|
| 1877 |
+
Mindstar: Enhancing math reasoning in pre-trained llms at inference
|
| 1878 |
+
time.
|
| 1879 |
+
arXiv preprint arXiv:2405.16265
|
| 1880 |
+
, 2024.
|
| 1881 |
+
Khot et al. (2022)
|
| 1882 |
+
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter
|
| 1883 |
+
Clark, and Ashish Sabharwal.
|
| 1884 |
+
Decomposed prompting: A modular approach for solving complex tasks.
|
| 1885 |
+
arXiv preprint arXiv:2210.02406
|
| 1886 |
+
, 2022.
|
| 1887 |
+
Kocsis & Szepesvári (2006)
|
| 1888 |
+
Levente Kocsis and Csaba Szepesvári.
|
| 1889 |
+
Bandit based monte-carlo planning.
|
| 1890 |
+
volume 2006, pp. 282–293, 09 2006.
|
| 1891 |
+
ISBN 978-3-540-45375-8.
|
| 1892 |
+
doi:
|
| 1893 |
+
10.1007/11871842_29
|
| 1894 |
+
.
|
| 1895 |
+
Kojima et al. (2022)
|
| 1896 |
+
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke
|
| 1897 |
+
Iwasawa.
|
| 1898 |
+
Large language models are zero-shot reasoners.
|
| 1899 |
+
Advances in neural information processing systems
|
| 1900 |
+
,
|
| 1901 |
+
35:22199–22213, 2022.
|
| 1902 |
+
Li et al. (2024)
|
| 1903 |
+
Chen Li, Weiqi Wang, Jingcheng Hu, Yixuan Wei, Nanning Zheng, Han Hu, Zheng
|
| 1904 |
+
Zhang, and Houwen Peng.
|
| 1905 |
+
Common 7b language models already possess strong math capabilities.
|
| 1906 |
+
arXiv preprint arXiv:2403.04706
|
| 1907 |
+
, 2024.
|
| 1908 |
+
Lightman et al. (2023)
|
| 1909 |
+
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy
|
| 1910 |
+
Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe.
|
| 1911 |
+
Let’s verify step by step.
|
| 1912 |
+
arXiv preprint arXiv:2305.20050
|
| 1913 |
+
, 2023.
|
| 1914 |
+
Madaan et al. (2024)
|
| 1915 |
+
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah
|
| 1916 |
+
Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al.
|
| 1917 |
+
Self-refine: Iterative refinement with self-feedback.
|
| 1918 |
+
Advances in Neural Information Processing Systems
|
| 1919 |
+
, 36, 2024.
|
| 1920 |
+
Meta (2024)
|
| 1921 |
+
Meta.
|
| 1922 |
+
Introducing meta llama3: The most capable openly available llm to
|
| 1923 |
+
date, 2024.
|
| 1924 |
+
URL
|
| 1925 |
+
https://ai.meta.com/blog/meta-llama-3/
|
| 1926 |
+
.
|
| 1927 |
+
Patel et al. (2021)
|
| 1928 |
+
Arkil Patel, Satwik Bhattamishra, and Navin Goyal.
|
| 1929 |
+
Are nlp models really able to solve simple math word problems?
|
| 1930 |
+
In
|
| 1931 |
+
Proceedings of the 2021 Conference of the North American
|
| 1932 |
+
Chapter of the Association for Computational Linguistics: Human Language
|
| 1933 |
+
Technologies
|
| 1934 |
+
, pp. 2080–2094, 2021.
|
| 1935 |
+
Roy & Roth (2015)
|
| 1936 |
+
Subhro Roy and Dan Roth.
|
| 1937 |
+
Solving General Arithmetic Word Problems.
|
| 1938 |
+
In
|
| 1939 |
+
Proc. of the Conference on Empirical Methods in Natural
|
| 1940 |
+
Language Processing (EMNLP)
|
| 1941 |
+
, 2015.
|
| 1942 |
+
URL
|
| 1943 |
+
http://cogcomp.org/papers/arithmetic.pdf
|
| 1944 |
+
.
|
| 1945 |
+
Silver et al. (2017)
|
| 1946 |
+
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew
|
| 1947 |
+
Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore
|
| 1948 |
+
Graepel, et al.
|
| 1949 |
+
Mastering chess and shogi by self-play with a general reinforcement
|
| 1950 |
+
learning algorithm.
|
| 1951 |
+
arXiv preprint arXiv:1712.01815
|
| 1952 |
+
, 2017.
|
| 1953 |
+
Snell et al. (2024)
|
| 1954 |
+
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar.
|
| 1955 |
+
Scaling llm test-time compute optimally can be more effective than
|
| 1956 |
+
scaling model parameters, 2024.
|
| 1957 |
+
URL
|
| 1958 |
+
https://arxiv.org/abs/2408.03314
|
| 1959 |
+
.
|
| 1960 |
+
Valmeekam et al. (2022)
|
| 1961 |
+
Karthik Valmeekam, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati.
|
| 1962 |
+
Large language models still can’t plan (a benchmark for LLMs on
|
| 1963 |
+
planning and reasoning about change).
|
| 1964 |
+
In
|
| 1965 |
+
NeurIPS 2022 Foundation Models for Decision Making
|
| 1966 |
+
Workshop
|
| 1967 |
+
, 2022.
|
| 1968 |
+
URL
|
| 1969 |
+
https://openreview.net/forum?id=wUU-7XTL5XO
|
| 1970 |
+
.
|
| 1971 |
+
Wang et al. (2024a)
|
| 1972 |
+
Ke Wang, Houxing Ren, Aojun Zhou, Zimu Lu, Sichun Luo, Weikang Shi, Renrui
|
| 1973 |
+
Zhang, Linqi Song, Mingjie Zhan, and Hongsheng Li.
|
| 1974 |
+
Mathcoder: Seamless code integration in LLMs for enhanced
|
| 1975 |
+
mathematical reasoning.
|
| 1976 |
+
In
|
| 1977 |
+
The Twelfth International Conference on Learning
|
| 1978 |
+
Representations
|
| 1979 |
+
, 2024a.
|
| 1980 |
+
URL
|
| 1981 |
+
https://openreview.net/forum?id=z8TW0ttBPp
|
| 1982 |
+
.
|
| 1983 |
+
Wang et al. (2024b)
|
| 1984 |
+
Peiyi Wang, Lei Li, Zhihong Shao, R. X. Xu, Damai Dai, Yifei Li, Deli Chen,
|
| 1985 |
+
Y. Wu, and Zhifang Sui.
|
| 1986 |
+
Math-shepherd: Verify and reinforce llms step-by-step without human
|
| 1987 |
+
annotations, 2024b.
|
| 1988 |
+
Wang et al. (2023)
|
| 1989 |
+
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang,
|
| 1990 |
+
Aakanksha Chowdhery, and Denny Zhou.
|
| 1991 |
+
Self-consistency improves chain of thought reasoning in language
|
| 1992 |
+
models.
|
| 1993 |
+
In
|
| 1994 |
+
The Eleventh International Conference on Learning
|
| 1995 |
+
Representations
|
| 1996 |
+
, 2023.
|
| 1997 |
+
URL
|
| 1998 |
+
https://openreview.net/forum?id=1PL1NIMMrw
|
| 1999 |
+
.
|
| 2000 |
+
Wei et al. (2022)
|
| 2001 |
+
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V
|
| 2002 |
+
Le, Denny Zhou, et al.
|
| 2003 |
+
Chain-of-thought prompting elicits reasoning in large language
|
| 2004 |
+
models.
|
| 2005 |
+
Advances in Neural Information Processing Systems
|
| 2006 |
+
,
|
| 2007 |
+
35:24824–24837, 2022.
|
| 2008 |
+
Weng et al. (2023)
|
| 2009 |
+
Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun,
|
| 2010 |
+
Kang Liu, and Jun Zhao.
|
| 2011 |
+
Large language models are better reasoners with self-verification.
|
| 2012 |
+
2023.
|
| 2013 |
+
Wu et al. (2024)
|
| 2014 |
+
Zhenyu Wu, Qingkai Zeng, Zhihan Zhang, Zhaoxuan Tan, Chao Shen, and Meng Jiang.
|
| 2015 |
+
Large language models can self-correct with minimal effort.
|
| 2016 |
+
arXiv preprint arXiv:2405.14092
|
| 2017 |
+
, 2024.
|
| 2018 |
+
Yao et al. (2024)
|
| 2019 |
+
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and
|
| 2020 |
+
Karthik Narasimhan.
|
| 2021 |
+
Tree of thoughts: Deliberate problem solving with large language
|
| 2022 |
+
models.
|
| 2023 |
+
Advances in Neural Information Processing Systems
|
| 2024 |
+
, 36, 2024.
|
| 2025 |
+
Zhang et al. (2024)
|
| 2026 |
+
Di Zhang, Jiatong Li, Xiaoshui Huang, Dongzhan Zhou, Yuqiang Li, and Wanli
|
| 2027 |
+
Ouyang.
|
| 2028 |
+
Accessing gpt-4 level mathematical olympiad solutions via monte carlo
|
| 2029 |
+
tree self-refine with llama-3 8b.
|
| 2030 |
+
arXiv preprint arXiv:2406.07394
|
| 2031 |
+
, 2024.
|
| 2032 |
+
Zheng et al. (2023)
|
| 2033 |
+
Huaixiu Steven Zheng, Swaroop Mishra, Xinyun Chen, Heng-Tze Cheng, Ed H Chi,
|
| 2034 |
+
Quoc V Le, and Denny Zhou.
|
| 2035 |
+
Take a step back: Evoking reasoning via abstraction in large language
|
| 2036 |
+
models.
|
| 2037 |
+
arXiv preprint arXiv:2310.06117
|
| 2038 |
+
, 2023.
|
| 2039 |
+
Zhou et al. (2023)
|
| 2040 |
+
Aojun Zhou, Ke Wang, Zimu Lu, Weikang Shi, Sichun Luo, Zipeng Qin, Shaoqing Lu,
|
| 2041 |
+
Anya Jia, Linqi Song, Mingjie Zhan, et al.
|
| 2042 |
+
Solving challenging math word problems using gpt-4 code interpreter
|
| 2043 |
+
with code-based self-verification.
|
| 2044 |
+
arXiv preprint arXiv:2308.07921
|
| 2045 |
+
, 2023.
|
| 2046 |
+
Zhou et al. (2022)
|
| 2047 |
+
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi
|
| 2048 |
+
Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, et al.
|
| 2049 |
+
Least-to-most prompting enables complex reasoning in large language
|
| 2050 |
+
models.
|
| 2051 |
+
In
|
| 2052 |
+
The Eleventh International Conference on Learning
|
| 2053 |
+
Representations
|
| 2054 |
+
, 2022.
|
| 2055 |
+
Zhou et al. (2024)
|
| 2056 |
+
Pei Zhou, Jay Pujara, Xiang Ren, Xinyun Chen, Heng-Tze Cheng, Quoc V Le, Ed H
|
| 2057 |
+
Chi, Denny Zhou, Swaroop Mishra, and Huaixiu Steven Zheng.
|
| 2058 |
+
Self-discover: Large language models self-compose reasoning
|
| 2059 |
+
structures.
|
| 2060 |
+
arXiv preprint arXiv:2402.03620
|
| 2061 |
+
, 2024.
|
| 2062 |
+
Appendix A
|
| 2063 |
+
Appendix
|
| 2064 |
+
A.1
|
| 2065 |
+
Experiments to evaluate the self-rewarding in SLMs
|
| 2066 |
+
Table 6:
|
| 2067 |
+
Analysis on the effectiveness of SLMs’ self-rewarding. The original
|
| 2068 |
+
r
|
| 2069 |
+
1
|
| 2070 |
+
subscript
|
| 2071 |
+
𝑟
|
| 2072 |
+
1
|
| 2073 |
+
r_{1}
|
| 2074 |
+
is a self-evaluation of the helpfulness of the new proposed subquestion, while
|
| 2075 |
+
r
|
| 2076 |
+
2
|
| 2077 |
+
subscript
|
| 2078 |
+
𝑟
|
| 2079 |
+
2
|
| 2080 |
+
r_{2}
|
| 2081 |
+
measures the confidence in answering the subquestion through self-consistency majority voting. Results show that replacing the self-evaluated
|
| 2082 |
+
r
|
| 2083 |
+
1
|
| 2084 |
+
subscript
|
| 2085 |
+
𝑟
|
| 2086 |
+
1
|
| 2087 |
+
r_{1}
|
| 2088 |
+
to random values does not significantly impact the final reasoning performance.
|
| 2089 |
+
Method
|
| 2090 |
+
LLaMA2-7B
|
| 2091 |
+
Mistral
|
| 2092 |
+
GSM8K
|
| 2093 |
+
RAP
|
| 2094 |
+
24.34
|
| 2095 |
+
56.25
|
| 2096 |
+
RAP + random
|
| 2097 |
+
r
|
| 2098 |
+
1
|
| 2099 |
+
subscript
|
| 2100 |
+
𝑟
|
| 2101 |
+
1
|
| 2102 |
+
r_{1}
|
| 2103 |
+
22.90
|
| 2104 |
+
55.50
|
| 2105 |
+
RAP + random
|
| 2106 |
+
r
|
| 2107 |
+
2
|
| 2108 |
+
subscript
|
| 2109 |
+
𝑟
|
| 2110 |
+
2
|
| 2111 |
+
r_{2}
|
| 2112 |
+
22.67
|
| 2113 |
+
49.66
|
| 2114 |
+
Multiarith
|
| 2115 |
+
RAP
|
| 2116 |
+
57.22
|
| 2117 |
+
91.11
|
| 2118 |
+
RAP + random
|
| 2119 |
+
r
|
| 2120 |
+
1
|
| 2121 |
+
subscript
|
| 2122 |
+
𝑟
|
| 2123 |
+
1
|
| 2124 |
+
r_{1}
|
| 2125 |
+
52.78
|
| 2126 |
+
90.56
|
| 2127 |
+
RAP + random
|
| 2128 |
+
r
|
| 2129 |
+
2
|
| 2130 |
+
subscript
|
| 2131 |
+
𝑟
|
| 2132 |
+
2
|
| 2133 |
+
r_{2}
|
| 2134 |
+
47.22
|
| 2135 |
+
81.11
|
| 2136 |
+
Ablation study on self-rewarding in RAP
|
| 2137 |
+
. RAP rewards both intermediate and terminal nodes. For each node generated by its action, it combines two scores,
|
| 2138 |
+
r
|
| 2139 |
+
1
|
| 2140 |
+
subscript
|
| 2141 |
+
𝑟
|
| 2142 |
+
1
|
| 2143 |
+
r_{1}
|
| 2144 |
+
and
|
| 2145 |
+
r
|
| 2146 |
+
2
|
| 2147 |
+
subscript
|
| 2148 |
+
𝑟
|
| 2149 |
+
2
|
| 2150 |
+
r_{2}
|
| 2151 |
+
, to determine the final reward score. Formally,
|
| 2152 |
+
r
|
| 2153 |
+
=
|
| 2154 |
+
r
|
| 2155 |
+
1
|
| 2156 |
+
×
|
| 2157 |
+
r
|
| 2158 |
+
2
|
| 2159 |
+
𝑟
|
| 2160 |
+
subscript
|
| 2161 |
+
𝑟
|
| 2162 |
+
1
|
| 2163 |
+
subscript
|
| 2164 |
+
𝑟
|
| 2165 |
+
2
|
| 2166 |
+
r=r_{1}\times r_{2}
|
| 2167 |
+
.
|
| 2168 |
+
r
|
| 2169 |
+
1
|
| 2170 |
+
subscript
|
| 2171 |
+
𝑟
|
| 2172 |
+
1
|
| 2173 |
+
r_{1}
|
| 2174 |
+
is a self-evaluation score that evaluates the LLM’s own estimation of the helpfulness of the current node. Specifically, it prompts the LLM with the question "
|
| 2175 |
+
Is the new question useful
|
| 2176 |
+
?".
|
| 2177 |
+
r
|
| 2178 |
+
2
|
| 2179 |
+
subscript
|
| 2180 |
+
𝑟
|
| 2181 |
+
2
|
| 2182 |
+
r_{2}
|
| 2183 |
+
is the confidence of correctly answering the proposed new question, measured by self-consistency majority voting.
|
| 2184 |
+
To evaluate the effectiveness of self-rewarding in RAP, we replace
|
| 2185 |
+
r
|
| 2186 |
+
1
|
| 2187 |
+
subscript
|
| 2188 |
+
𝑟
|
| 2189 |
+
1
|
| 2190 |
+
r_{1}
|
| 2191 |
+
and
|
| 2192 |
+
r
|
| 2193 |
+
2
|
| 2194 |
+
subscript
|
| 2195 |
+
𝑟
|
| 2196 |
+
2
|
| 2197 |
+
r_{2}
|
| 2198 |
+
with random values sampled from (0,1)and re-run RAP on LLaMA2-7B and Mistral-7B. We select a challenging dataset, GSM8K and an easy mathematical reasoning dataset, Multiarith
|
| 2199 |
+
(Roy & Roth,
|
| 2200 |
+
2015
|
| 2201 |
+
)
|
| 2202 |
+
, for evaluation.
|
| 2203 |
+
Table
|
| 2204 |
+
6
|
| 2205 |
+
compares the results with original RAP. We can see that replacing
|
| 2206 |
+
r
|
| 2207 |
+
1
|
| 2208 |
+
subscript
|
| 2209 |
+
𝑟
|
| 2210 |
+
1
|
| 2211 |
+
r_{1}
|
| 2212 |
+
with random values has minimal impact on RAP’s performance across different SLMs and datasets. However, replacing
|
| 2213 |
+
r
|
| 2214 |
+
2
|
| 2215 |
+
subscript
|
| 2216 |
+
𝑟
|
| 2217 |
+
2
|
| 2218 |
+
r_{2}
|
| 2219 |
+
with random values result in a noticeable drop in accuracy on Mistral and Multiarith. This indicates that self-evaluation
|
| 2220 |
+
r
|
| 2221 |
+
1
|
| 2222 |
+
subscript
|
| 2223 |
+
𝑟
|
| 2224 |
+
1
|
| 2225 |
+
r_{1}
|
| 2226 |
+
has minimal effect, suggesting that LLaMA2-7B and Mistral are essentially performing near-random self-evaluations.
|
| 2227 |
+
A.2
|
| 2228 |
+
Discussions
|
| 2229 |
+
Discussions on the importance of generator and discriminator
|
| 2230 |
+
. In our experiments, we found that on certain SLMs, the discriminator yields more significant improvement than the generator. For instance, on LLaMA2-7B, rStar (generator @maj) can improves accuracy by +4.17% on GSM8K, while our discriminator can further boosts accuracy by +36.69%. However, both the generator and discriminator are crucial to final performance. The generator must effectively produce the correct solution first, and the discriminator then provide reasonable feedback to select the correct solution.
|
| 2231 |
+
The importance of the generator and discriminator varies based on the SLM’s solution generation effectiveness. For LLaMA2-7B, the proportion of correct solution trajectories is low, leading to a lower majority voting score. In this case, the discriminator is crucial for identifying the correct solution, significantly boosting accuracy. Conversely, for stronger models like LLaMA3-8B-instruct, our generator produces a higher proportion of correct solutions. Therefore, improving the generator results in greater accuracy improvements.
|
| 2232 |
+
Table 7:
|
| 2233 |
+
Inference costs of rStar on GSM8K. We show the average number of inferences and generated tokens required to solve a question.
|
| 2234 |
+
LLaMA2-7B
|
| 2235 |
+
Mistral
|
| 2236 |
+
Avg. calls
|
| 2237 |
+
166.81
|
| 2238 |
+
148.90
|
| 2239 |
+
Avg. generated tokens
|
| 2240 |
+
367.1k
|
| 2241 |
+
348.6k
|
| 2242 |
+
Inference cost
|
| 2243 |
+
. rStar grows SLMs reasoning capabilities at inference time. The primary inference cost arises from our MCTS self-generator. Table
|
| 2244 |
+
7
|
| 2245 |
+
shows the average number of inferences and tokens generated for solving a GSM8K question after 32 rollouts.
|
| 2246 |
+
On LLaMA2-7B and Mistral, this averages 166 and 148 model calls to solve a question, respectively. Currently, completing the 32 rollouts for the entire GSM8K test set takes about 4.5 days on a single A100 GPU per model. These costs can be significantly reduced by distributing tasks across multiple GPUs or batching model calls within each rollout.
|
| 2247 |
+
A.3
|
| 2248 |
+
Prompt examples in MTCS
|
| 2249 |
+
Action 1. Propose an one-step thought.
|
| 2250 |
+
### Instruction:
|
| 2251 |
+
There are 15 trees in the grove. Grove workers will plant trees in the grove today. After they are done, there will be 21 trees. How many trees did the grove workers plant today?
|
| 2252 |
+
### Response:
|
| 2253 |
+
Let’s think step by step.
|
| 2254 |
+
Step 1: Identify the initial number of trees. The problem states there are 15 trees in the grove.
|
| 2255 |
+
Step 2: Identify the final number of trees. The problem states there will be 21 trees after the workers are done planting.
|
| 2256 |
+
Step 3: Subtract the initial number of trees from the final number of trees to find out how many trees were planted.
|
| 2257 |
+
Step 4: Therefore, the grove workers planted 21 (final number of trees) - 15 (initial number of trees) = 6 trees today.
|
| 2258 |
+
Step 5: The answer is 6.
|
| 2259 |
+
### Instruction:
|
| 2260 |
+
If there are 3 cars in the parking lot and 2 more cars arrive, how many cars are in the parking lot?
|
| 2261 |
+
### Response:
|
| 2262 |
+
Let’s think step by step.
|
| 2263 |
+
Step 1: Start with the number of cars that are already in the parking lot, which is 3 cars.
|
| 2264 |
+
Step 2: Add the number of cars that arrive, which is 2 cars.
|
| 2265 |
+
Step 3: Add the numbers together. there are 3 cars + 2 cars = 5 cars in the parking lot.
|
| 2266 |
+
Step 4: The answer is 5.
|
| 2267 |
+
### Instruction:
|
| 2268 |
+
Leah had 32 chocolates and her sister had 42. If they ate 35, how many pieces do they have left in total?
|
| 2269 |
+
### Response:
|
| 2270 |
+
Let’s think step by step.
|
| 2271 |
+
Step 1: Add the number of chocolates Leah and her sister had initially. Leah had 32 chocolates and her sister had 42 chocolates. So, they had 32 + 42 = 74 chocolates in total.
|
| 2272 |
+
Step 2: Subtract the number of chocolates they ate from the total number of chocolates they had. They ate 35 chocolates. So, they have 74 - 35 = 39 chocolates left.
|
| 2273 |
+
Step 3: The answer is 39.
|
| 2274 |
+
### Instruction:
|
| 2275 |
+
Jason had 20 lollipops. He gave Denny some lollipops. Now Jason has 12 lollipops. How many lollipops did Jason give to Denny?
|
| 2276 |
+
### Response:
|
| 2277 |
+
Let’s think step by step.
|
| 2278 |
+
Step 1: Identify the initial number of lollipops. The problem states Jason had 20 lollipops.
|
| 2279 |
+
Step 2: Identify the final number of lollipops. The problem states Jason now has 12 lollipops.
|
| 2280 |
+
Step 3: Subtract the final number of lollipops from the initial number to find out how many lollipops Jason gave to Denny.
|
| 2281 |
+
Step 4: Therefore, Jason gave Denny 20 (initial number of lollipops) - 12 (final number of lollipops) = 8 lollipops.
|
| 2282 |
+
Step 5: The answer is 8.
|
| 2283 |
+
###
|
| 2284 |
+
Instruction:
|
| 2285 |
+
{user question}
|
| 2286 |
+
###
|
| 2287 |
+
Response:
|
| 2288 |
+
Let’s think step by step.
|
| 2289 |
+
Action 2: Propose the remaining thought steps /A4: Answer the sub-question again.
|
| 2290 |
+
### Instruction:
|
| 2291 |
+
There are 15 trees in the grove. Grove workers will plant trees in the grove today. After they are done, there will be 21 trees. How many trees did the grove workers plant today?
|
| 2292 |
+
### Response:
|
| 2293 |
+
Let’s think step by step. There are 15 trees originally. Then there were 21 trees after some more were planted. So there must have been 21 - 15 = 6. The answer is: 6.
|
| 2294 |
+
### Instruction:
|
| 2295 |
+
If there are 3 cars in the parking lot and 2 more cars arrive, how many cars are in the parking lot?
|
| 2296 |
+
### Response:
|
| 2297 |
+
Let’s think step by step. There are originally 3 cars. 2 more cars arrive. 3 + 2 = 5. The answer is: 5.
|
| 2298 |
+
### Instruction:
|
| 2299 |
+
Leah had 32 chocolates and her sister had 42. If they ate 35, how many pieces do they have left in total?
|
| 2300 |
+
### Response:
|
| 2301 |
+
Let’s think step by step. Originally, Leah had 32 chocolates. Her sister had 42. So in total they had 32 + 42 = 74. After eating 35, they had 74 - 35 = 39. The answer is: 39.
|
| 2302 |
+
### Instruction:
|
| 2303 |
+
Jason had 20 lollipops. He gave Denny some lollipops. Now Jason has 12 lollipops. How many lollipops did Jason give to Denny?
|
| 2304 |
+
### Response:
|
| 2305 |
+
Let’s think step by step. Jason started with 20 lollipops. Then he had 12 after giving some to Denny. So he gave Denny 20 - 12 = 8. The answer is: 8.
|
| 2306 |
+
### Instruction:
|
| 2307 |
+
Shawn has five toys. For Christmas, he got two toys each from his mom and dad. How many toys does he have now?
|
| 2308 |
+
### Response:
|
| 2309 |
+
Let’s think step by step. Shawn started with 5 toys. If he got 2 toys each from his mom and dad, then that is 4 more toys. 5 + 4 = 9. The answer is: 9.
|
| 2310 |
+
### Instruction:
|
| 2311 |
+
There were nine computers in the server room. Five more computers were installed each day, from monday to thursday. How many computers are now in the server room?
|
| 2312 |
+
### Response:
|
| 2313 |
+
Let’s think step by step. There were originally 9 computers. For each of 4 days, 5 more computers were added. So 5 * 4 = 20 computers were added. 9 + 20 is 29. The answer is: 29.
|
| 2314 |
+
### Instruction:
|
| 2315 |
+
Michael had 58 golf balls. On tuesday, he lost 23 golf balls. On wednesday, he lost 2 more. How many golf balls did he have at the end of wednesday?
|
| 2316 |
+
### Response:
|
| 2317 |
+
Let’s think step by step. Michael started with 58 golf balls. After losing 23 on tuesday, he had 58 - 23 = 35. After losing 2 more, he had 35 - 2 = 33 golf balls. The answer is: 33.
|
| 2318 |
+
### Instruction:
|
| 2319 |
+
Olivia has $23. She bought five bagels for $3 each. How much money does she have left?
|
| 2320 |
+
### Response:
|
| 2321 |
+
Let’s think step by step. Olivia had 23 dollars. 5 bagels for 3 dollars each will be 5 x 3 = 15 dollars. So she has 23 - 15 dollars left. 23 - 15 is 8. The answer is: 8.
|
| 2322 |
+
###
|
| 2323 |
+
Instruction:
|
| 2324 |
+
{user question}
|
| 2325 |
+
###
|
| 2326 |
+
Response
|
| 2327 |
+
:
|
| 2328 |
+
Action 3: Propose next sub-question along with its answer.
|
| 2329 |
+
Given a question, please decompose it into sub-questions. For each sub-question, please answer it in a complete sentence, ending with "The answer is <a numeric answer>". When the original question is answerable, please start the subquestion with "Now we can answer the question: <original question>".
|
| 2330 |
+
Question 1: Four years ago, Kody was only half as old as Mohamed. If Mohamed is currently twice as 30 years old, how old is Kody?
|
| 2331 |
+
Question 1.1: How old is Mohamed currently?
|
| 2332 |
+
Answer 1.1: Mohamed is twice as old as 30 years, which means he is 30 * 2 = 60 years old.
|
| 2333 |
+
Question 1.2: What was Kody’s age four years ago, given that it was half of Mohamed’s age at that time?
|
| 2334 |
+
Answer 1.2: Four years ago, Mohamed was 60 - 4 = 56 years old, so Kody was half of that, which is 56 / 2 = 28 years old.
|
| 2335 |
+
Question 1.3: Now we can answer the question: How old is Kody?
|
| 2336 |
+
Answer 1.3: Kody is currently 28 + 4 = 32 years old. The answer is 32.
|
| 2337 |
+
Question 2: On a moonless night, three fireflies danced in the evening breeze. They were joined by four less than a dozen more fireflies before two of the fireflies flew away. How many fireflies remained?
|
| 2338 |
+
Question 2.1: How many fireflies joined?
|
| 2339 |
+
Answer 2.1: The fireflies were joined by four less than a dozen more fireflies, which are 12 - 4 = 8 fireflies. The answer is 8.
|
| 2340 |
+
Question 2.2: Now we can answer the question: How many fireflies remained?
|
| 2341 |
+
Answer 2.2: Three fireflies were dancing originally. They were joined by 8 fireflies before two of them flew away. So there were 3 + 8 - 2 = 9 remaining. The answer is 9.
|
| 2342 |
+
Question 3: Ali has four $10 bills and six $20 bills that he saved after working for Mr. James on his farm. Ali gives her sister half of the total money he has and uses 3/5 of the remaining amount of money to buy dinner. Calculate the amount of money he has after buying the dinner.
|
| 2343 |
+
Question 3.1: How much money does Ali have after giving half of his total money to his sister?
|
| 2344 |
+
Answer 3.1: Ali initially has four $10 bills and six $20 bills, totaling 4 * 10 + 6 * 20 = 160 dollars. Giving half of this to his sister leaves him with 160 / 2 = 80 dollars. The answer is 80.
|
| 2345 |
+
Question 3.2: How much money does Ali spend on dinner?
|
| 2346 |
+
Answer 3.2: Ali uses 3/5 of his remaining money, which is 80 dollars, to buy dinner. Therefore, he spends 80 * 3/5 = 48 dollars on dinner. The answer is 48.
|
| 2347 |
+
Question 3.3: Now we can answer the question: How much money does Ali have after buying the dinner?
|
| 2348 |
+
Answer 3.3: After buying the dinner, Ali has 80 - 48 = 32 dollars left. The answer is 32.
|
| 2349 |
+
Question 4: A car is driving through a tunnel with many turns. After a while, the car must travel through a ring that requires a total of 4 right-hand turns. After the 1st turn, it travels 5 meters. After the 2nd turn, it travels 8 meters. After the 3rd turn, it travels a little further and at the 4th turn, it immediately exits the tunnel. If the car has driven a total of 23 meters around the ring, how far did it have to travel after the 3rd turn?
|
| 2350 |
+
Question 4.1: How far did the car travel except for the 3rd turn?
|
| 2351 |
+
Answer 4.1: It travels 5 meters after the 1st, 8 meters after the 2nd, and 0 meters after the 4th turn. It’s a total of 5 + 8 + 0 = 13 meters. The answer is 13.
|
| 2352 |
+
Question 4.2: Now we can answer the question: How far did the car have to travel after the 3rd turn?
|
| 2353 |
+
Answer 4.2: The car has driven a total of 23 meters around the ring. It travels 13 meters except for the 3rd turn. So it has to travel 23 - 13 = 10 meters after the 3rd turn. The answer is 10.
|
| 2354 |
+
Question 5: {user question}
|
| 2355 |
+
Action 5: Rephrase the question/sub-question.
|
| 2356 |
+
You are an AI assistant to help me rephrase questions by splitting the question context into conditions. In your rephrased question, remember to fully express the information in the original question.
|
| 2357 |
+
Original Question: Olivia has $23. She bought five bagels for $3 each. How much money does she have left?
|
| 2358 |
+
Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Olivia starts with $23. Condition 2: She buys five bagels, each costing $3. Question: How much money does Olivia have remaining after her purchase?
|
| 2359 |
+
Original Question: Michael had 58 golf balls. On Tuesday, he lost 23 golf balls. On Wednesday, he lost 2 more. How many golf balls did he have at the end of Wednesday?
|
| 2360 |
+
Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Michael initially has 58 golf balls. Condition 2: On Tuesday, he loses 23 golf balls. Condition 3: On Wednesday, he loses 2 additional golf balls. Question: What is the total number of golf balls Michael has left at the end of Wednesday?
|
| 2361 |
+
Original Question: Angelo and Melanie want to plan how many hours over the next week they should study together for their test next week. They have 2 chapters of their textbook to study and 4 worksheets to memorize. They figure out that they should dedicate 3 hours to each chapter of their textbook and 1.5 hours for each worksheet. If they plan to study no more than 4 hours each day, how many days should they plan to study total over the next week if they take a 10-minute break every hour, include 3 10-minute snack breaks each day, and 30 minutes for lunch each day?
|
| 2362 |
+
Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Angelo and Melanie need to study 2 textbook chapters and 4 worksheets. Condition 2: They allocate 3 hours per textbook chapter and 1.5 hours per worksheet. Condition 3: Their daily study limit is 4 hours, with a 10-minute break every hour, three 10-minute snack breaks, and a 30-minute lunch break each day. Question: Over the next week, for how many days should they plan to study to cover all their materials?
|
| 2363 |
+
Original Question: Leah had 32 chocolates and her sister had 42. If they ate 35, how many pieces do they have left in total?
|
| 2364 |
+
Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Leah has 32 chocolates. Condition 2: Her sister has 42 chocolates. Condition 3: Together, they consume 35 chocolates. Question: How many chocolates remain between them after they have eaten some?
|
| 2365 |
+
Original Question: There were nine computers in the server room. Five more computers were installed each day, from Monday to Thursday. How many computers are now in the server room?
|
| 2366 |
+
Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Initially, there are nine computers in the server room. Condition 2: Each day, from Monday to Thursday, five additional computers are installed. Question: What is the total number of computers in the server room after these installations?
|
| 2367 |
+
Original Question: Jason had 20 lollipops. He gave Denny some lollipops. Now Jason has 12 lollipops. How many lollipops did Jason give to Denny?
|
| 2368 |
+
Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Jason starts with 20 lollipops. Condition 2: After giving some lollipops to Denny, Jason has 12 lollipops left. Question: How many lollipops did Jason give to Denny?
|
| 2369 |
+
Original Question: Sam bought a dozen boxes, each with 30 highlighter pens inside, for $10 each box. He rearranged five of these boxes into packages of six highlighters each and sold them for $3 per package. He sold the rest of the highlighters separately at the rate of three pens for $2. How much profit did he make in total, in dollars?
|
| 2370 |
+
Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Sam purchases a dozen boxes of highlighters, with each box containing 30 pens, at $10 per box. Condition 2: He repackages five boxes into packages of six highlighters, selling each package for $3. Condition 3: He sells the remaining highlighters at a rate of three for $2. Question: What is Sam’s total profit from these transactions?
|
| 2371 |
+
Original Question: There are 15 trees in the grove. Grove workers will plant trees in the grove today. After they are done, there will be 21 trees. How many trees did the grove workers plant today?
|
| 2372 |
+
Rephrased Question: Given a list of conditions, please answer the question. Condition 1: Initially, there are 15 trees in the grove. Condition 2: Grove workers will add more trees to the grove today. Condition 3: After planting, the total number of trees in the grove will increase to 21. Question: How many trees did the grove workers plant today?
|
| 2373 |
+
Original Question: {user question}
|
| 2374 |
+
Rephrased Question:
|
| 2375 |
+
◄
|
| 2376 |
+
Feeling
|
| 2377 |
+
lucky?
|
| 2378 |
+
Conversion
|
| 2379 |
+
report
|
| 2380 |
+
Report
|
| 2381 |
+
an issue
|
| 2382 |
+
View original
|
| 2383 |
+
on arXiv
|
| 2384 |
+
►
|
|
@@ -0,0 +1,191 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2408.06195] Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers'
|
| 3 |
+
id: 240806195-mutual-reasoning-makes-smaller-llms-stronger-problem-solvers
|
| 4 |
+
tags:
|
| 5 |
+
- deepread
|
| 6 |
+
created: '2026-06-10T00:39:56.384617Z'
|
| 7 |
+
source: https://arxiv.org/abs/2408.06195
|
| 8 |
+
source_domain: arxiv.org
|
| 9 |
+
fetched_at: '2026-06-10T00:39:56.384488Z'
|
| 10 |
+
fetch_provider: builtin
|
| 11 |
+
status: draft
|
| 12 |
+
type: note
|
| 13 |
+
tier: institutional
|
| 14 |
+
content_type: paper
|
| 15 |
+
deprecated: false
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
[2408.06195] Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
|
| 19 |
+
Computer Science > Computation and Language
|
| 20 |
+
arXiv:2408.06195
|
| 21 |
+
(cs)
|
| 22 |
+
[Submitted on 12 Aug 2024]
|
| 23 |
+
Title:
|
| 24 |
+
Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
|
| 25 |
+
Authors:
|
| 26 |
+
Zhenting Qi
|
| 27 |
+
,
|
| 28 |
+
Mingyuan Ma
|
| 29 |
+
,
|
| 30 |
+
Jiahang Xu
|
| 31 |
+
,
|
| 32 |
+
Li Lyna Zhang
|
| 33 |
+
,
|
| 34 |
+
Fan Yang
|
| 35 |
+
,
|
| 36 |
+
Mao Yang
|
| 37 |
+
View a PDF of the paper titled Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers, by Zhenting Qi and 5 other authors
|
| 38 |
+
View PDF
|
| 39 |
+
HTML (experimental)
|
| 40 |
+
Abstract:
|
| 41 |
+
This paper introduces rStar, a self-play mutual reasoning approach that significantly improves reasoning capabilities of small language models (SLMs) without fine-tuning or superior models. rStar decouples reasoning into a self-play mutual generation-discrimination process. First, a target SLM augments the Monte Carlo Tree Search (MCTS) with a rich set of human-like reasoning actions to construct higher quality reasoning trajectories. Next, another SLM, with capabilities similar to the target SLM, acts as a discriminator to verify each trajectory generated by the target SLM. The mutually agreed reasoning trajectories are considered mutual consistent, thus are more likely to be correct. Extensive experiments across five SLMs demonstrate rStar can effectively solve diverse reasoning problems, including GSM8K, GSM-Hard, MATH, SVAMP, and StrategyQA. Remarkably, rStar boosts GSM8K accuracy from 12.51% to 63.91% for LLaMA2-7B, from 36.46% to 81.88% for Mistral-7B, from 74.53% to 91.13% for LLaMA3-8B-Instruct. Code will be available at
|
| 42 |
+
this https URL
|
| 43 |
+
.
|
| 44 |
+
Subjects:
|
| 45 |
+
Computation and Language (cs.CL)
|
| 46 |
+
Cite as:
|
| 47 |
+
arXiv:2408.06195
|
| 48 |
+
[cs.CL]
|
| 49 |
+
(or
|
| 50 |
+
arXiv:2408.06195v1
|
| 51 |
+
[cs.CL]
|
| 52 |
+
for this version)
|
| 53 |
+
https://doi.org/10.48550/arXiv.2408.06195
|
| 54 |
+
Focus to learn more
|
| 55 |
+
arXiv-issued DOI via DataCite
|
| 56 |
+
Submission history
|
| 57 |
+
From: Li Lyna Zhang [
|
| 58 |
+
view email
|
| 59 |
+
]
|
| 60 |
+
[v1]
|
| 61 |
+
Mon, 12 Aug 2024 14:42:13 UTC (1,140 KB)
|
| 62 |
+
Full-text links:
|
| 63 |
+
Access Paper:
|
| 64 |
+
View a PDF of the paper titled Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers, by Zhenting Qi and 5 other authors
|
| 65 |
+
View PDF
|
| 66 |
+
HTML (experimental)
|
| 67 |
+
TeX Source
|
| 68 |
+
view license
|
| 69 |
+
Current browse context:
|
| 70 |
+
cs.CL
|
| 71 |
+
< prev
|
| 72 |
+
|
|
| 73 |
+
next >
|
| 74 |
+
new
|
| 75 |
+
|
|
| 76 |
+
recent
|
| 77 |
+
|
|
| 78 |
+
2024-08
|
| 79 |
+
Change to browse by:
|
| 80 |
+
cs
|
| 81 |
+
References & Citations
|
| 82 |
+
NASA ADS
|
| 83 |
+
Google Scholar
|
| 84 |
+
Semantic Scholar
|
| 85 |
+
export BibTeX citation
|
| 86 |
+
Loading...
|
| 87 |
+
BibTeX formatted citation
|
| 88 |
+
×
|
| 89 |
+
loading...
|
| 90 |
+
Data provided by:
|
| 91 |
+
Bookmark
|
| 92 |
+
Bibliographic Tools
|
| 93 |
+
Bibliographic and Citation Tools
|
| 94 |
+
Bibliographic Explorer Toggle
|
| 95 |
+
Bibliographic Explorer
|
| 96 |
+
(
|
| 97 |
+
What is the Explorer?
|
| 98 |
+
)
|
| 99 |
+
Connected Papers Toggle
|
| 100 |
+
Connected Papers
|
| 101 |
+
(
|
| 102 |
+
What is Connected Papers?
|
| 103 |
+
)
|
| 104 |
+
Litmaps Toggle
|
| 105 |
+
Litmaps
|
| 106 |
+
(
|
| 107 |
+
What is Litmaps?
|
| 108 |
+
)
|
| 109 |
+
scite.ai Toggle
|
| 110 |
+
scite Smart Citations
|
| 111 |
+
(
|
| 112 |
+
What are Smart Citations?
|
| 113 |
+
)
|
| 114 |
+
Code, Data, Media
|
| 115 |
+
Code, Data and Media Associated with this Article
|
| 116 |
+
alphaXiv Toggle
|
| 117 |
+
alphaXiv
|
| 118 |
+
(
|
| 119 |
+
What is alphaXiv?
|
| 120 |
+
)
|
| 121 |
+
Links to Code Toggle
|
| 122 |
+
CatalyzeX Code Finder for Papers
|
| 123 |
+
(
|
| 124 |
+
What is CatalyzeX?
|
| 125 |
+
)
|
| 126 |
+
DagsHub Toggle
|
| 127 |
+
DagsHub
|
| 128 |
+
(
|
| 129 |
+
What is DagsHub?
|
| 130 |
+
)
|
| 131 |
+
GotitPub Toggle
|
| 132 |
+
Gotit.pub
|
| 133 |
+
(
|
| 134 |
+
What is GotitPub?
|
| 135 |
+
)
|
| 136 |
+
Huggingface Toggle
|
| 137 |
+
Hugging Face
|
| 138 |
+
(
|
| 139 |
+
What is Huggingface?
|
| 140 |
+
)
|
| 141 |
+
ScienceCast Toggle
|
| 142 |
+
ScienceCast
|
| 143 |
+
(
|
| 144 |
+
What is ScienceCast?
|
| 145 |
+
)
|
| 146 |
+
Demos
|
| 147 |
+
Demos
|
| 148 |
+
Replicate Toggle
|
| 149 |
+
Replicate
|
| 150 |
+
(
|
| 151 |
+
What is Replicate?
|
| 152 |
+
)
|
| 153 |
+
Spaces Toggle
|
| 154 |
+
Hugging Face Spaces
|
| 155 |
+
(
|
| 156 |
+
What is Spaces?
|
| 157 |
+
)
|
| 158 |
+
Spaces Toggle
|
| 159 |
+
TXYZ.AI
|
| 160 |
+
(
|
| 161 |
+
What is TXYZ.AI?
|
| 162 |
+
)
|
| 163 |
+
Related Papers
|
| 164 |
+
Recommenders and Search Tools
|
| 165 |
+
Link to Influence Flower
|
| 166 |
+
Influence Flower
|
| 167 |
+
(
|
| 168 |
+
What are Influence Flowers?
|
| 169 |
+
)
|
| 170 |
+
Core recommender toggle
|
| 171 |
+
CORE Recommender
|
| 172 |
+
(
|
| 173 |
+
What is CORE?
|
| 174 |
+
)
|
| 175 |
+
Author
|
| 176 |
+
Venue
|
| 177 |
+
Institution
|
| 178 |
+
Topic
|
| 179 |
+
About arXivLabs
|
| 180 |
+
arXivLabs: experimental projects with community collaborators
|
| 181 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 182 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 183 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 184 |
+
Learn more about arXivLabs
|
| 185 |
+
.
|
| 186 |
+
Which authors of this paper are endorsers?
|
| 187 |
+
|
|
| 188 |
+
Disable MathJax
|
| 189 |
+
(
|
| 190 |
+
What is MathJax?
|
| 191 |
+
)
|
|
@@ -0,0 +1,2144 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2410.20285] SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search
|
| 3 |
+
and Iterative Refinement'
|
| 4 |
+
id: 241020285-swe-search-enhancing-software-agents-with-monte-carlo-tree-search-and-2
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:41:20.614811Z'
|
| 8 |
+
source: https://ar5iv.labs.arxiv.org/html/2410.20285
|
| 9 |
+
source_domain: ar5iv.labs.arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:41:20.614654Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2410.20285] SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement
|
| 20 |
+
SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement
|
| 21 |
+
Antonis Antoniades
|
| 22 |
+
1∗
|
| 23 |
+
, Albert Örwall
|
| 24 |
+
2
|
| 25 |
+
,
|
| 26 |
+
Kexun Zhang
|
| 27 |
+
3
|
| 28 |
+
,
|
| 29 |
+
Yuxi Xie
|
| 30 |
+
4
|
| 31 |
+
, Anirudh Goyal
|
| 32 |
+
5
|
| 33 |
+
, William Wang
|
| 34 |
+
1
|
| 35 |
+
1
|
| 36 |
+
University of California, Santa Barbara,
|
| 37 |
+
2
|
| 38 |
+
Moatless AI,
|
| 39 |
+
3
|
| 40 |
+
Carnegie Mellon University,
|
| 41 |
+
4
|
| 42 |
+
National University of Singapore,
|
| 43 |
+
5
|
| 44 |
+
Mila
|
| 45 |
+
Denotes equal contribution.
|
| 46 |
+
Correspondence to:
|
| 47 |
+
antonis@ucsb.edu
|
| 48 |
+
,
|
| 49 |
+
albert@moatless.ai
|
| 50 |
+
.
|
| 51 |
+
Code:
|
| 52 |
+
github.com/aorwall/moatless-tree-search
|
| 53 |
+
Abstract
|
| 54 |
+
Software engineers operating in complex and dynamic environments must continuously adapt to evolving requirements, learn iteratively from experience, and reconsider their approaches based on new insights. However, current large language model (LLM)-based software agents often rely on rigid processes and tend to repeat ineffective actions without the capacity to evaluate their performance or adapt their strategies over time. To address these challenges, we propose SWE-Search, a multi-agent framework that integrates Monte Carlo Tree Search (MCTS) with a self-improvement mechanism to enhance software agents’ performance on repository-level software tasks. SWE-Search extends traditional MCTS by incorporating a hybrid value function that leverages LLMs for both numerical value estimation and qualitative evaluation. This enables self-feedback loops where agents iteratively refine their strategies based on both quantitative numerical evaluations and qualitative natural language assessments of pursued trajectories. The framework includes a SWE-Agent for adaptive exploration, a Value Agent for iterative feedback, and a Discriminator Agent that facilitates multi-agent debate for collaborative decision-making. Applied to the SWE-bench benchmark, our approach demonstrates a 23% relative improvement in performance across five models compared to standard open-source agents without MCTS. Our analysis reveals how performance scales with increased search depth and identifies key factors that facilitate effective self-evaluation in software agents. This work highlights the potential of self-evaluation driven search techniques to enhance agent reasoning and planning in complex, dynamic software engineering environments.
|
| 55 |
+
1
|
| 56 |
+
Introduction
|
| 57 |
+
Software engineering is a complex and iterative process involving exploration, problem-solving, and decision-making under uncertainty. Tasks such as debugging, feature development, and code refactoring require continuous assessment of different approaches, frequent backtracking, and the incorporation of new information. While machine learning has made progress in automating parts of this workflow
|
| 58 |
+
(Li et al.,
|
| 59 |
+
2022
|
| 60 |
+
; OpenAI et al.,
|
| 61 |
+
2024
|
| 62 |
+
; Ouyang et al.,
|
| 63 |
+
2022
|
| 64 |
+
; Yang et al.,
|
| 65 |
+
2024b
|
| 66 |
+
)
|
| 67 |
+
, replicating the adaptive and strategic behavior of human engineers remains a significant challenge. This is due to the inherently non-linear and iterative nature of software engineering, where engineers dynamically explore various solutions, refine strategies based on feedback, and collaborate to identify the most effective path forward. Current large language model (LLM)-based software agents
|
| 68 |
+
(Xia et al.,
|
| 69 |
+
2024
|
| 70 |
+
; Zhang et al.,
|
| 71 |
+
2024d
|
| 72 |
+
)
|
| 73 |
+
, while powerful, often struggle with complex, long-horizon tasks that require adaptive strategies and flexible reassessment over time. These agents can become trapped in repetitive patterns, limiting their effectiveness in tackling more intricate software engineering problems.
|
| 74 |
+
To address these challenges, we introduce
|
| 75 |
+
SWE-Search
|
| 76 |
+
, a multi-agent system that replicates the adaptability, iterative learning, and collaborative decision-making of human engineers. SWE-Search is designed to address three critical needs in software engineering:
|
| 77 |
+
Flexible Exploration and Adaptation
|
| 78 |
+
: Engineering problems often require exploring multiple approaches and adapting strategies based on evolving information
|
| 79 |
+
(Li et al.,
|
| 80 |
+
2022
|
| 81 |
+
)
|
| 82 |
+
. SWE-Search’s SWE-Agent operates in a flexible state space, allowing it to fluidly transition between actions such as planning, searching, and editing. This design mirrors the way engineers backtrack and adjust their approach dynamically, ensuring the agent can revise its course when faced with new challenges or information, and points towards the direction of more general, open-ended systems
|
| 83 |
+
(Wang et al.,
|
| 84 |
+
2023
|
| 85 |
+
; Ma et al.,
|
| 86 |
+
2024a
|
| 87 |
+
; Lu et al.,
|
| 88 |
+
2024b
|
| 89 |
+
; Faldor et al.,
|
| 90 |
+
2024
|
| 91 |
+
; Hu et al.,
|
| 92 |
+
2024
|
| 93 |
+
; Lu et al.,
|
| 94 |
+
2024a
|
| 95 |
+
)
|
| 96 |
+
.
|
| 97 |
+
Iterative Learning through Feedback
|
| 98 |
+
: Effective engineering relies heavily on continuous testing and refinement. To replicate this, SWE-Search integrates a Monte Carlo Tree Search (MCTS)
|
| 99 |
+
(Silver et al.,
|
| 100 |
+
2016b
|
| 101 |
+
)
|
| 102 |
+
planning module paired with a Value Agent. The MCTS module balances exploration and exploitation to guide the agent through complex solution spaces. The Value Agent augments this process by providing both utility estimates and qualitative feedback, allowing the agent to iteratively improve its decision-making based on past experiences, similar to how engineers refine their work through feedback and debugging.
|
| 103 |
+
Collaborative Decision-Making
|
| 104 |
+
: Complex problems often benefit from diverse perspectives
|
| 105 |
+
(Khan et al.,
|
| 106 |
+
2024
|
| 107 |
+
; Amayuelas et al.,
|
| 108 |
+
2024
|
| 109 |
+
; Du et al.,
|
| 110 |
+
2023
|
| 111 |
+
; Zhang et al.,
|
| 112 |
+
2024c
|
| 113 |
+
)
|
| 114 |
+
. In SWE-Search, once a set of potential solutions is generated, the Discriminator Agent facilitates a multi-agent debate. Each agent advocates for different solutions by presenting arguments, which are critically evaluated by a judge agent. This process mirrors real-world engineering collaboration, where teams deliberate to refine and select the most robust solutions.
|
| 115 |
+
The architecture of SWE-Search is designed to automate software engineering tasks through these adaptive, feedback-driven, and collaborative processes. The SWE-Agent serves as the system’s problem solver, operating in a dynamic environment where it can backtrack and adapt its actions as necessary. The MCTS Planning Module efficiently guides exploration and exploitation, ensuring that the agent balances the need for innovation with the need to focus on promising solutions. The Value Agent provides continual feedback, offering both quantitative assessments and qualitative insights, helping the agent refine its strategy iteratively. Finally, the Discriminator Agent ensures that the final decision is rigorously vetted through a multi-agent debate, simulating the collaborative decision-making processes commonly found in engineering teams.
|
| 116 |
+
We evaluate SWE-Search on the SWE-bench benchmark, a comprehensive dataset from real-world open-source repositories. SWE-bench tests agents’ ability to resolve software issues by generating code patches that fix failing tests. SWE-Search demonstrates a
|
| 117 |
+
23
|
| 118 |
+
%
|
| 119 |
+
percent
|
| 120 |
+
23
|
| 121 |
+
23\%
|
| 122 |
+
relative performance improvement across five models compared to standard open-source agents, highlighting the effectiveness of strategic search and iterative self-evaluation. Through detailed analysis, we explore how performance scales with increased search depth and identify key factors that enhance self-assessment in software agents. Our work demonstrates the potential of MCTS and iterative learning to improve agent reasoning and planning in dynamic, complex domains like software engineering, introducing a new paradigm for autonomous software development.
|
| 123 |
+
2
|
| 124 |
+
Related Work
|
| 125 |
+
Search methods
|
| 126 |
+
Various search approaches have been applied to Large Language Models (LLMs) to facilitate System 2
|
| 127 |
+
(Kahneman,
|
| 128 |
+
2011
|
| 129 |
+
; Saha et al.,
|
| 130 |
+
2024
|
| 131 |
+
; Pan et al.,
|
| 132 |
+
2023
|
| 133 |
+
; Bounsi et al.,
|
| 134 |
+
2024
|
| 135 |
+
)
|
| 136 |
+
thinking in non-linear reasoning structures. A critical feature of these approaches is their ability to backtrack. Unlike greedy processes
|
| 137 |
+
(Black,
|
| 138 |
+
2005
|
| 139 |
+
)
|
| 140 |
+
, search algorithms explore multiple branches at each step, potentially escaping paths that lead to dead ends. These methods differ in their strategies for exploring and memorizing possible choices, and in their heuristics for switching between them. Breadth-first search
|
| 141 |
+
(Moore,
|
| 142 |
+
1959
|
| 143 |
+
)
|
| 144 |
+
maintains all possible search paths, incurring significant memory and computational costs. Depth-first search
|
| 145 |
+
(Cormen et al.,
|
| 146 |
+
2009
|
| 147 |
+
)
|
| 148 |
+
, in contrast, prioritizes the most promising path in a more greedy manner. When applied to LLMs, these methods demonstrate a trade-off between diversity and quality in text generation
|
| 149 |
+
(Yao et al.,
|
| 150 |
+
2023
|
| 151 |
+
)
|
| 152 |
+
. The A
|
| 153 |
+
∗
|
| 154 |
+
algorithm
|
| 155 |
+
(Hart et al.,
|
| 156 |
+
1968
|
| 157 |
+
)
|
| 158 |
+
combines aspects of breadth-first and greedy search to find optimal solutions using a predetermined evaluation function. In this work, we adopt Monte Carlo Tree Search (MCTS)
|
| 159 |
+
(Silver et al.,
|
| 160 |
+
2016b
|
| 161 |
+
)
|
| 162 |
+
, an advanced search algorithm that conducts statistical tree search without requiring dedicated evaluation heuristics for each state. MCTS has achieved impressive results in complex strategy games
|
| 163 |
+
(Silver et al.,
|
| 164 |
+
2016a
|
| 165 |
+
)
|
| 166 |
+
, protein folding
|
| 167 |
+
(Jumper et al.,
|
| 168 |
+
2021
|
| 169 |
+
)
|
| 170 |
+
, and algorithm discovery
|
| 171 |
+
(Fawzi et al.,
|
| 172 |
+
2022
|
| 173 |
+
)
|
| 174 |
+
.
|
| 175 |
+
Software Agents
|
| 176 |
+
Software agents are designed to perform autonomous actions within large codebases. Given a repository-level task, these agents typically locate relevant files and code segments before implementing necessary changes. We focus on the SWE-bench task
|
| 177 |
+
(Jimenez et al.,
|
| 178 |
+
2024
|
| 179 |
+
)
|
| 180 |
+
, which involves resolving real-world GitHub issues. Among the agents with disclosed technical details on SWE-bench,
|
| 181 |
+
Yang et al. (
|
| 182 |
+
2024b
|
| 183 |
+
)
|
| 184 |
+
introduced the concept of agent-computer interfaces with SWE-agent. OpenDevin
|
| 185 |
+
(Wang et al.,
|
| 186 |
+
2024b
|
| 187 |
+
)
|
| 188 |
+
presents a collection of community-driven agents, including CodeAct
|
| 189 |
+
(Wang et al.,
|
| 190 |
+
2024a
|
| 191 |
+
)
|
| 192 |
+
. The Agentless approach demonstrated competitive performance using a simple two-step process of localization and repair. AutoCodeRover
|
| 193 |
+
(Zhang et al.,
|
| 194 |
+
2024d
|
| 195 |
+
)
|
| 196 |
+
incorporated advanced code tools such as abstract syntax trees and spectrum-based fault localization. The Alibaba Lingma Agent
|
| 197 |
+
(Ma et al.,
|
| 198 |
+
2024b
|
| 199 |
+
)
|
| 200 |
+
introduced a search-based approach for repository exploration, followed by a structured editing phase. While effective, it constitutes a more hand-designed solution specifically designed to interface with the search functionality of their agent.
|
| 201 |
+
3
|
| 202 |
+
Methodology
|
| 203 |
+
SWE-Search is a multi-agent system designed to tackle complex software engineering tasks by integrating dynamic planning, value estimation, and deliberative decision-making. The core motivation behind this method is to emulate the sophisticated, iterative workflows of human software engineers, where exploration, planning, and collaboration are crucial to solving intricate problems. By leveraging the strengths of Monte Carlo Tree Search (MCTS) for planning, a Value Agent for utility estimation and feedback, and a Discriminator Agent for final decision-making through debate, SWE-Search provides a comprehensive, adaptive framework capable of navigating and solving real-world software engineering challenges.
|
| 204 |
+
SWE-Search consists of four primary components that work in synergy:
|
| 205 |
+
SWE-Search Framework and Action Agent
|
| 206 |
+
: Building on the moatless-tools framework
|
| 207 |
+
(Örwall,
|
| 208 |
+
2024
|
| 209 |
+
)
|
| 210 |
+
, SWE-Search operates in a dynamic code environment with a flexible state-space and a git-like commit tree structure. This design facilitates efficient backtracking to previous states, enabling the Action Agent to explore diverse solution trajectories. The adaptable state-space enhances the system’s ability to exploit the MCTS module effectively.
|
| 211 |
+
Search Algorithm
|
| 212 |
+
: The core of SWE-Search’s exploration strategy is based on a Monte Carlo Tree Search (MCTS) which uses a heuristic-based selection process similar to AlphaZero
|
| 213 |
+
(Silver et al.,
|
| 214 |
+
2016a
|
| 215 |
+
)
|
| 216 |
+
, specifically tailored for software engineering tasks. This modified MCTS algorithm effectively balances exploration and exploitation, helping the agent explore a diverse set of solutions and converge quickly on the most promising strategies.
|
| 217 |
+
Value (Function) Agent
|
| 218 |
+
: To approximate the utility of each observation, we employ an LLM-based value function, which in addition to outputting a value, also generates an explanation in natural language. This explanation can be leveraged to improve subsequent actions from parent nodes, enabling iterative self-improvement of the search process.
|
| 219 |
+
Discriminator Agent
|
| 220 |
+
: In the final stage of SWE-Search, the Discriminator Agent evaluates the solutions generated by the search process. Inspired by multi-agent debate frameworks
|
| 221 |
+
Du et al. (
|
| 222 |
+
2023
|
| 223 |
+
); Khan et al. (
|
| 224 |
+
2024
|
| 225 |
+
); Amayuelas et al. (
|
| 226 |
+
2024
|
| 227 |
+
)
|
| 228 |
+
, this agent engages in a structured debate, where multiple agents argue for or against the proposed solutions. The debate process not only surfaces diverse perspectives but also leads to a more rigorously justified final decision.
|
| 229 |
+
This system architecture combines the strengths of dynamic action selection, strategic planning, and collaborative deliberation, creating a comprehensive tool capable of handling the complexity and iterative nature of software engineering tasks.
|
| 230 |
+
3.1
|
| 231 |
+
Problem Formulation
|
| 232 |
+
Figure 1:
|
| 233 |
+
SWE-Search Overview.
|
| 234 |
+
Tree search.
|
| 235 |
+
Each state is represented as a node from which the agent can expand from, and each corresponding action is presented as an edge.
|
| 236 |
+
Evaluation.
|
| 237 |
+
Uses all relevant context including trajectory information, file context, and executed tests, to provide a quantitative value estimation and qualitative explanation in natural language.
|
| 238 |
+
Expansion.
|
| 239 |
+
Nodes can be expanded using value function feedback from future actions.
|
| 240 |
+
The task of the SWE agent can be formalized as a tuple
|
| 241 |
+
ℳ
|
| 242 |
+
=
|
| 243 |
+
(
|
| 244 |
+
𝒮
|
| 245 |
+
,
|
| 246 |
+
𝒞
|
| 247 |
+
,
|
| 248 |
+
𝒜
|
| 249 |
+
,
|
| 250 |
+
𝒱
|
| 251 |
+
,
|
| 252 |
+
𝒫
|
| 253 |
+
,
|
| 254 |
+
p
|
| 255 |
+
0
|
| 256 |
+
,
|
| 257 |
+
ρ
|
| 258 |
+
)
|
| 259 |
+
ℳ
|
| 260 |
+
𝒮
|
| 261 |
+
𝒞
|
| 262 |
+
𝒜
|
| 263 |
+
𝒱
|
| 264 |
+
𝒫
|
| 265 |
+
subscript
|
| 266 |
+
𝑝
|
| 267 |
+
0
|
| 268 |
+
𝜌
|
| 269 |
+
\mathcal{M}=(\mathcal{S},\mathcal{C},\mathcal{A},\mathcal{V},\mathcal{P},p_{0},\rho)
|
| 270 |
+
. Here,
|
| 271 |
+
𝒮
|
| 272 |
+
𝒮
|
| 273 |
+
\mathcal{S}
|
| 274 |
+
represents the state space, encompassing all possible states such as the current context of the files the agent is working on and the overall status of the codebase. The context space, denoted as
|
| 275 |
+
𝒞
|
| 276 |
+
𝒞
|
| 277 |
+
\mathcal{C}
|
| 278 |
+
, includes metadata about the repository and the initial problem description. The value function
|
| 279 |
+
𝒱
|
| 280 |
+
𝒱
|
| 281 |
+
\mathcal{V}
|
| 282 |
+
assigns a utility score to each state-action pair
|
| 283 |
+
O
|
| 284 |
+
|
| 285 |
+
(
|
| 286 |
+
a
|
| 287 |
+
,
|
| 288 |
+
t
|
| 289 |
+
)
|
| 290 |
+
𝑂
|
| 291 |
+
𝑎
|
| 292 |
+
𝑡
|
| 293 |
+
O(a,t)
|
| 294 |
+
, guiding the agent’s decisions.
|
| 295 |
+
The environment’s dynamics are defined by a context-dependent transition function
|
| 296 |
+
𝒫
|
| 297 |
+
:
|
| 298 |
+
𝒮
|
| 299 |
+
×
|
| 300 |
+
𝒜
|
| 301 |
+
×
|
| 302 |
+
𝒞
|
| 303 |
+
→
|
| 304 |
+
Δ
|
| 305 |
+
|
| 306 |
+
(
|
| 307 |
+
𝒮
|
| 308 |
+
)
|
| 309 |
+
:
|
| 310 |
+
𝒫
|
| 311 |
+
→
|
| 312 |
+
𝒮
|
| 313 |
+
𝒜
|
| 314 |
+
𝒞
|
| 315 |
+
Δ
|
| 316 |
+
𝒮
|
| 317 |
+
\mathcal{P}:\mathcal{S}\times\mathcal{A}\times\mathcal{C}\rightarrow\Delta(\mathcal{S})
|
| 318 |
+
, which models the evolution of the repository’s state after each action. The initial state distribution,
|
| 319 |
+
p
|
| 320 |
+
0
|
| 321 |
+
:
|
| 322 |
+
𝒞
|
| 323 |
+
→
|
| 324 |
+
Δ
|
| 325 |
+
|
| 326 |
+
(
|
| 327 |
+
𝒮
|
| 328 |
+
)
|
| 329 |
+
:
|
| 330 |
+
subscript
|
| 331 |
+
𝑝
|
| 332 |
+
0
|
| 333 |
+
→
|
| 334 |
+
𝒞
|
| 335 |
+
Δ
|
| 336 |
+
���
|
| 337 |
+
p_{0}:\mathcal{C}\rightarrow\Delta(\mathcal{S})
|
| 338 |
+
, specifies how the initial state depends on the given context, while
|
| 339 |
+
ρ
|
| 340 |
+
∈
|
| 341 |
+
Δ
|
| 342 |
+
|
| 343 |
+
(
|
| 344 |
+
𝒞
|
| 345 |
+
)
|
| 346 |
+
𝜌
|
| 347 |
+
Δ
|
| 348 |
+
𝒞
|
| 349 |
+
\rho\in\Delta(\mathcal{C})
|
| 350 |
+
defines the distribution over contexts.
|
| 351 |
+
Given an initial context
|
| 352 |
+
c
|
| 353 |
+
∼
|
| 354 |
+
ρ
|
| 355 |
+
similar-to
|
| 356 |
+
𝑐
|
| 357 |
+
𝜌
|
| 358 |
+
c\sim\rho
|
| 359 |
+
and an initial state
|
| 360 |
+
s
|
| 361 |
+
0
|
| 362 |
+
∼
|
| 363 |
+
p
|
| 364 |
+
0
|
| 365 |
+
(
|
| 366 |
+
⋅
|
| 367 |
+
∣
|
| 368 |
+
c
|
| 369 |
+
)
|
| 370 |
+
s_{0}\sim p_{0}(\cdot\mid c)
|
| 371 |
+
, the SWE agent executes its policy
|
| 372 |
+
π
|
| 373 |
+
:
|
| 374 |
+
𝒮
|
| 375 |
+
×
|
| 376 |
+
𝒞
|
| 377 |
+
→
|
| 378 |
+
Δ
|
| 379 |
+
|
| 380 |
+
(
|
| 381 |
+
𝒜
|
| 382 |
+
)
|
| 383 |
+
:
|
| 384 |
+
𝜋
|
| 385 |
+
→
|
| 386 |
+
𝒮
|
| 387 |
+
𝒞
|
| 388 |
+
Δ
|
| 389 |
+
𝒜
|
| 390 |
+
\pi:\mathcal{S}\times\mathcal{C}\rightarrow\Delta(\mathcal{A})
|
| 391 |
+
, which selects actions based on the current state and context. At each time step
|
| 392 |
+
t
|
| 393 |
+
𝑡
|
| 394 |
+
t
|
| 395 |
+
, the agent takes an action
|
| 396 |
+
a
|
| 397 |
+
t
|
| 398 |
+
∼
|
| 399 |
+
π
|
| 400 |
+
|
| 401 |
+
(
|
| 402 |
+
s
|
| 403 |
+
t
|
| 404 |
+
,
|
| 405 |
+
c
|
| 406 |
+
)
|
| 407 |
+
similar-to
|
| 408 |
+
subscript
|
| 409 |
+
𝑎
|
| 410 |
+
𝑡
|
| 411 |
+
𝜋
|
| 412 |
+
subscript
|
| 413 |
+
𝑠
|
| 414 |
+
𝑡
|
| 415 |
+
𝑐
|
| 416 |
+
a_{t}\sim\pi(s_{t},c)
|
| 417 |
+
and receives a corresponding reward
|
| 418 |
+
ℛ
|
| 419 |
+
|
| 420 |
+
(
|
| 421 |
+
s
|
| 422 |
+
t
|
| 423 |
+
,
|
| 424 |
+
a
|
| 425 |
+
t
|
| 426 |
+
,
|
| 427 |
+
c
|
| 428 |
+
)
|
| 429 |
+
ℛ
|
| 430 |
+
subscript
|
| 431 |
+
𝑠
|
| 432 |
+
𝑡
|
| 433 |
+
subscript
|
| 434 |
+
𝑎
|
| 435 |
+
𝑡
|
| 436 |
+
𝑐
|
| 437 |
+
\mathcal{R}(s_{t},a_{t},c)
|
| 438 |
+
. The environment then transitions to a new state
|
| 439 |
+
s
|
| 440 |
+
t
|
| 441 |
+
+
|
| 442 |
+
1
|
| 443 |
+
∼
|
| 444 |
+
𝒫
|
| 445 |
+
(
|
| 446 |
+
⋅
|
| 447 |
+
∣
|
| 448 |
+
s
|
| 449 |
+
t
|
| 450 |
+
,
|
| 451 |
+
a
|
| 452 |
+
t
|
| 453 |
+
,
|
| 454 |
+
c
|
| 455 |
+
)
|
| 456 |
+
s_{t+1}\sim\mathcal{P}(\cdot\mid s_{t},a_{t},c)
|
| 457 |
+
, and the agent continues to observe this updated state. Over time, this process generates a trajectory
|
| 458 |
+
τ
|
| 459 |
+
:=
|
| 460 |
+
{
|
| 461 |
+
s
|
| 462 |
+
t
|
| 463 |
+
,
|
| 464 |
+
a
|
| 465 |
+
t
|
| 466 |
+
,
|
| 467 |
+
r
|
| 468 |
+
t
|
| 469 |
+
}
|
| 470 |
+
t
|
| 471 |
+
=
|
| 472 |
+
0
|
| 473 |
+
T
|
| 474 |
+
assign
|
| 475 |
+
𝜏
|
| 476 |
+
superscript
|
| 477 |
+
subscript
|
| 478 |
+
subscript
|
| 479 |
+
𝑠
|
| 480 |
+
𝑡
|
| 481 |
+
subscript
|
| 482 |
+
𝑎
|
| 483 |
+
𝑡
|
| 484 |
+
subscript
|
| 485 |
+
𝑟
|
| 486 |
+
𝑡
|
| 487 |
+
𝑡
|
| 488 |
+
0
|
| 489 |
+
𝑇
|
| 490 |
+
\tau:=\{s_{t},a_{t},r_{t}\}_{t=0}^{T}
|
| 491 |
+
as the agent interacts with the environment.
|
| 492 |
+
The agent’s objective is to maximize the cumulative reward over the trajectory, which is captured by the value function
|
| 493 |
+
v
|
| 494 |
+
|
| 495 |
+
(
|
| 496 |
+
s
|
| 497 |
+
t
|
| 498 |
+
,
|
| 499 |
+
a
|
| 500 |
+
t
|
| 501 |
+
,
|
| 502 |
+
{
|
| 503 |
+
s
|
| 504 |
+
i
|
| 505 |
+
}
|
| 506 |
+
i
|
| 507 |
+
=
|
| 508 |
+
0
|
| 509 |
+
t
|
| 510 |
+
−
|
| 511 |
+
1
|
| 512 |
+
,
|
| 513 |
+
{
|
| 514 |
+
a
|
| 515 |
+
i
|
| 516 |
+
}
|
| 517 |
+
i
|
| 518 |
+
=
|
| 519 |
+
0
|
| 520 |
+
t
|
| 521 |
+
−
|
| 522 |
+
1
|
| 523 |
+
)
|
| 524 |
+
𝑣
|
| 525 |
+
subscript
|
| 526 |
+
𝑠
|
| 527 |
+
𝑡
|
| 528 |
+
subscript
|
| 529 |
+
𝑎
|
| 530 |
+
𝑡
|
| 531 |
+
superscript
|
| 532 |
+
subscript
|
| 533 |
+
subscript
|
| 534 |
+
𝑠
|
| 535 |
+
𝑖
|
| 536 |
+
𝑖
|
| 537 |
+
0
|
| 538 |
+
𝑡
|
| 539 |
+
1
|
| 540 |
+
superscript
|
| 541 |
+
subscript
|
| 542 |
+
subscript
|
| 543 |
+
𝑎
|
| 544 |
+
𝑖
|
| 545 |
+
𝑖
|
| 546 |
+
0
|
| 547 |
+
𝑡
|
| 548 |
+
1
|
| 549 |
+
v(s_{t},a_{t},\{s_{i}\}_{i=0}^{t-1},\{a_{i}\}_{i=0}^{t-1})
|
| 550 |
+
. This value function depends not only on the current state and action but also on the history of previous states and actions, which deviates from the assumptions of a Markovian process. Formally, the agent seeks to maximize the expected cumulative reward, defined as:
|
| 551 |
+
max
|
| 552 |
+
π
|
| 553 |
+
|
| 554 |
+
V
|
| 555 |
+
T
|
| 556 |
+
|
| 557 |
+
(
|
| 558 |
+
ρ
|
| 559 |
+
)
|
| 560 |
+
=
|
| 561 |
+
max
|
| 562 |
+
π
|
| 563 |
+
|
| 564 |
+
𝔼
|
| 565 |
+
|
| 566 |
+
τ
|
| 567 |
+
|
| 568 |
+
[
|
| 569 |
+
∑
|
| 570 |
+
t
|
| 571 |
+
=
|
| 572 |
+
0
|
| 573 |
+
T
|
| 574 |
+
ℛ
|
| 575 |
+
|
| 576 |
+
(
|
| 577 |
+
s
|
| 578 |
+
t
|
| 579 |
+
,
|
| 580 |
+
a
|
| 581 |
+
t
|
| 582 |
+
,
|
| 583 |
+
c
|
| 584 |
+
)
|
| 585 |
+
∣
|
| 586 |
+
c
|
| 587 |
+
∼
|
| 588 |
+
ρ
|
| 589 |
+
;
|
| 590 |
+
π
|
| 591 |
+
]
|
| 592 |
+
subscript
|
| 593 |
+
𝜋
|
| 594 |
+
superscript
|
| 595 |
+
𝑉
|
| 596 |
+
𝑇
|
| 597 |
+
𝜌
|
| 598 |
+
subscript
|
| 599 |
+
𝜋
|
| 600 |
+
𝔼
|
| 601 |
+
𝜏
|
| 602 |
+
delimited-[]
|
| 603 |
+
similar-to
|
| 604 |
+
conditional
|
| 605 |
+
superscript
|
| 606 |
+
subscript
|
| 607 |
+
𝑡
|
| 608 |
+
0
|
| 609 |
+
𝑇
|
| 610 |
+
ℛ
|
| 611 |
+
subscript
|
| 612 |
+
𝑠
|
| 613 |
+
𝑡
|
| 614 |
+
subscript
|
| 615 |
+
𝑎
|
| 616 |
+
𝑡
|
| 617 |
+
𝑐
|
| 618 |
+
𝑐
|
| 619 |
+
𝜌
|
| 620 |
+
𝜋
|
| 621 |
+
\max_{\pi}V^{T}(\rho)=\max_{\pi}\mathbb{E}{\tau}\left[\sum_{t=0}^{T}\mathcal{R}(s_{t},a_{t},c)\mid c\sim\rho;\pi\right]
|
| 622 |
+
.
|
| 623 |
+
This optimization captures the agent’s (in-context) process, as it adjusts its policy
|
| 624 |
+
π
|
| 625 |
+
𝜋
|
| 626 |
+
\pi
|
| 627 |
+
to achieve the highest expected return across multiple trajectories, considering both current and historical information.
|
| 628 |
+
3.2
|
| 629 |
+
SWE-Search Framework and Action Agent
|
| 630 |
+
The SWE-Search Action Agent builds on the moatless-tools framework
|
| 631 |
+
(Örwall,
|
| 632 |
+
2024
|
| 633 |
+
)
|
| 634 |
+
. Its action space,
|
| 635 |
+
𝒜
|
| 636 |
+
𝒜
|
| 637 |
+
\mathcal{A}
|
| 638 |
+
, is organized as a two-tier hierarchy, comprising both action types and their corresponding specific actions. Formally, this can be expressed as
|
| 639 |
+
𝒜
|
| 640 |
+
=
|
| 641 |
+
(
|
| 642 |
+
t
|
| 643 |
+
,
|
| 644 |
+
a
|
| 645 |
+
)
|
| 646 |
+
∣
|
| 647 |
+
t
|
| 648 |
+
∈
|
| 649 |
+
𝒯
|
| 650 |
+
,
|
| 651 |
+
a
|
| 652 |
+
∈
|
| 653 |
+
𝒜
|
| 654 |
+
t
|
| 655 |
+
formulae-sequence
|
| 656 |
+
𝒜
|
| 657 |
+
conditional
|
| 658 |
+
𝑡
|
| 659 |
+
𝑎
|
| 660 |
+
𝑡
|
| 661 |
+
𝒯
|
| 662 |
+
𝑎
|
| 663 |
+
subscript
|
| 664 |
+
𝒜
|
| 665 |
+
𝑡
|
| 666 |
+
\mathcal{A}={(t,a)\mid t\in\mathcal{T},a\in\mathcal{A}_{t}}
|
| 667 |
+
, where
|
| 668 |
+
𝒯
|
| 669 |
+
𝒯
|
| 670 |
+
\mathcal{T}
|
| 671 |
+
represents the set of action types (e.g.,
|
| 672 |
+
Search
|
| 673 |
+
,
|
| 674 |
+
Plan
|
| 675 |
+
,
|
| 676 |
+
Edit
|
| 677 |
+
), and
|
| 678 |
+
𝒜
|
| 679 |
+
t
|
| 680 |
+
subscript
|
| 681 |
+
𝒜
|
| 682 |
+
𝑡
|
| 683 |
+
\mathcal{A}_{t}
|
| 684 |
+
is the set of possible actions corresponding to each type
|
| 685 |
+
t
|
| 686 |
+
𝑡
|
| 687 |
+
t
|
| 688 |
+
. These actions range from tool invocations and code modifications to the generation of structured text. To enhance the agent’s effectiveness in search-driven tasks, we introduced the following modifications:
|
| 689 |
+
One key modification we implemented is the expansion of the
|
| 690 |
+
Plan
|
| 691 |
+
state, allowing it to transition flexibly to any other state, rather than being limited to transitioning only to
|
| 692 |
+
Edit
|
| 693 |
+
. This change is motivated by the need to enable more dynamic and adaptive problem-solving behaviors within the agent. In the context of software engineering, rigid state transitions can be overly restrictive, forcing the agent into predetermined pathways that may not always align with the complexities of real-world scenarios. For instance, during code modification tasks, an agent might recognize mid-process that further planning, additional searches, or different types of analysis are necessary before proceeding with edits. Restricting transitions only to editing would artificially constrain the agent, potentially leading it to suboptimal actions or causing it to become stuck in unproductive loops. By allowing transitions to any state, we empower the agent to adapt to new information as it arises (
|
| 694 |
+
Fig.
|
| 695 |
+
2
|
| 696 |
+
), exploring a wider variety of trajectories. This enhanced flexibility reflects the iterative and often non-linear nature of real software engineering workflows, where engineers frequently revisit planning, testing, and research phases before committing to edits.
|
| 697 |
+
Second, the agent is empowered to execute any tests within the codebase at its discretion, as well as to create and implement new tests. The results of these tests are incorporated into both the value function and the agent’s subsequent decision-making process. It is crucial to highlight that the tests required to resolve a given instance (i.e., fail-to-pass tests) are not explicitly revealed to the agent. However, the agent can leverage any pre-existing tests within the repository, simulating the behavior of a real-world software engineer.
|
| 698 |
+
1
|
| 699 |
+
1
|
| 700 |
+
1
|
| 701 |
+
This approach aligns with the practices of other SWE agents, and has been validated by the authors of SWE-bench, who confirmed its legitimacy as long as the fail-to-pass tests remain concealed from the model.
|
| 702 |
+
3.3
|
| 703 |
+
Value (Function) Agent
|
| 704 |
+
The role of the Value Agent extends beyond simply estimating the expected utility of a given state-action pair
|
| 705 |
+
O
|
| 706 |
+
n
|
| 707 |
+
|
| 708 |
+
(
|
| 709 |
+
s
|
| 710 |
+
n
|
| 711 |
+
,
|
| 712 |
+
a
|
| 713 |
+
n
|
| 714 |
+
)
|
| 715 |
+
subscript
|
| 716 |
+
𝑂
|
| 717 |
+
𝑛
|
| 718 |
+
subscript
|
| 719 |
+
𝑠
|
| 720 |
+
𝑛
|
| 721 |
+
subscript
|
| 722 |
+
𝑎
|
| 723 |
+
𝑛
|
| 724 |
+
O_{n}(s_{n},a_{n})
|
| 725 |
+
. In addition to calculating the value
|
| 726 |
+
v
|
| 727 |
+
n
|
| 728 |
+
subscript
|
| 729 |
+
𝑣
|
| 730 |
+
𝑛
|
| 731 |
+
v_{n}
|
| 732 |
+
, the Value Agent generates a written explanation, denoted as
|
| 733 |
+
ε
|
| 734 |
+
𝜀
|
| 735 |
+
\varepsilon
|
| 736 |
+
. This explanation serves a dual purpose: it provides transparency into the decision-making process and functions as feedback for the Action Agent, which can leverage this explanation when re-expanding from the parent node of
|
| 737 |
+
O
|
| 738 |
+
n
|
| 739 |
+
subscript
|
| 740 |
+
𝑂
|
| 741 |
+
𝑛
|
| 742 |
+
O_{n}
|
| 743 |
+
(see
|
| 744 |
+
Figure
|
| 745 |
+
1
|
| 746 |
+
,
|
| 747 |
+
hindsight feedback
|
| 748 |
+
). This approach enables the system to iteratively refine its decision-making process, mirroring how a human software engineer continuously re-evaluates their approach based on new information to improve their problem-solving strategy.
|
| 749 |
+
The input to the value function consists of all state-action pairs up to and including the current state being evaluated, alongside specific instructions on how to assess the state. This allows the Value Agent to contextualize the decision within the trajectory, accounting for the sequence of actions and states leading up to the present. The final output of the value function can be formalized as:
|
| 750 |
+
(
|
| 751 |
+
v
|
| 752 |
+
t
|
| 753 |
+
,
|
| 754 |
+
ε
|
| 755 |
+
t
|
| 756 |
+
)
|
| 757 |
+
=
|
| 758 |
+
V
|
| 759 |
+
|
| 760 |
+
(
|
| 761 |
+
s
|
| 762 |
+
t
|
| 763 |
+
,
|
| 764 |
+
a
|
| 765 |
+
t
|
| 766 |
+
,
|
| 767 |
+
{
|
| 768 |
+
s
|
| 769 |
+
i
|
| 770 |
+
}
|
| 771 |
+
i
|
| 772 |
+
=
|
| 773 |
+
0
|
| 774 |
+
|
| 775 |
+
…
|
| 776 |
+
|
| 777 |
+
t
|
| 778 |
+
−
|
| 779 |
+
1
|
| 780 |
+
,
|
| 781 |
+
{
|
| 782 |
+
a
|
| 783 |
+
i
|
| 784 |
+
}
|
| 785 |
+
i
|
| 786 |
+
=
|
| 787 |
+
0
|
| 788 |
+
|
| 789 |
+
…
|
| 790 |
+
|
| 791 |
+
t
|
| 792 |
+
−
|
| 793 |
+
1
|
| 794 |
+
)
|
| 795 |
+
subscript
|
| 796 |
+
𝑣
|
| 797 |
+
𝑡
|
| 798 |
+
subscript
|
| 799 |
+
𝜀
|
| 800 |
+
𝑡
|
| 801 |
+
𝑉
|
| 802 |
+
subscript
|
| 803 |
+
𝑠
|
| 804 |
+
𝑡
|
| 805 |
+
subscript
|
| 806 |
+
𝑎
|
| 807 |
+
𝑡
|
| 808 |
+
subscript
|
| 809 |
+
subscript
|
| 810 |
+
𝑠
|
| 811 |
+
𝑖
|
| 812 |
+
𝑖
|
| 813 |
+
0
|
| 814 |
+
…
|
| 815 |
+
𝑡
|
| 816 |
+
1
|
| 817 |
+
subscript
|
| 818 |
+
subscript
|
| 819 |
+
𝑎
|
| 820 |
+
𝑖
|
| 821 |
+
𝑖
|
| 822 |
+
0
|
| 823 |
+
…
|
| 824 |
+
𝑡
|
| 825 |
+
1
|
| 826 |
+
(v_{t},\varepsilon_{t})=V(s_{t},a_{t},\{s_{i}\}_{i=0{\ldots}t-1},\{a_{i}\}_{i=0{\ldots}t-1})
|
| 827 |
+
(1)
|
| 828 |
+
Here,
|
| 829 |
+
v
|
| 830 |
+
t
|
| 831 |
+
subscript
|
| 832 |
+
𝑣
|
| 833 |
+
𝑡
|
| 834 |
+
v_{t}
|
| 835 |
+
represents the expected utility of the current state-action pair, while
|
| 836 |
+
ε
|
| 837 |
+
t
|
| 838 |
+
subscript
|
| 839 |
+
𝜀
|
| 840 |
+
𝑡
|
| 841 |
+
\varepsilon_{t}
|
| 842 |
+
is the accompanying explanation.
|
| 843 |
+
In practice, the Value Agent is tasked with analyzing the entire trajectory leading up to the current state-action pair, providing not only the required utility estimate
|
| 844 |
+
v
|
| 845 |
+
t
|
| 846 |
+
subscript
|
| 847 |
+
𝑣
|
| 848 |
+
𝑡
|
| 849 |
+
v_{t}
|
| 850 |
+
, but also a detailed explanation
|
| 851 |
+
ε
|
| 852 |
+
t
|
| 853 |
+
subscript
|
| 854 |
+
𝜀
|
| 855 |
+
𝑡
|
| 856 |
+
\varepsilon_{t}
|
| 857 |
+
. This explanation is critical for the agent’s overall performance, as it offers insight into the reasoning behind utility estimates, which in turn informs the Action Agent’s future decisions. We have observed that one of the key factors driving the effectiveness of the Value Agent lies in the clarity and specificity of these explanations. A well-articulated explanation can illuminate the strengths and limitations of different state types (e.g.,
|
| 858 |
+
Search
|
| 859 |
+
,
|
| 860 |
+
Edit
|
| 861 |
+
,
|
| 862 |
+
Plan
|
| 863 |
+
), helping the Action Agent better understand which types of states are more promising or risky to pursue.
|
| 864 |
+
By providing detailed feedback on the potential utility of different actions and contextualizing them within the broader trajectory, the Value Agent enables more informed and strategic decision-making by the Action Agent. This integration of both quantitative and qualitative feedback leads to improved performance and more adaptive behavior throughout the task (
|
| 865 |
+
Fig.
|
| 866 |
+
4
|
| 867 |
+
a
|
| 868 |
+
).
|
| 869 |
+
3.4
|
| 870 |
+
Search Algorithm
|
| 871 |
+
Our search tree is structured with nodes representing states
|
| 872 |
+
𝒮
|
| 873 |
+
t
|
| 874 |
+
subscript
|
| 875 |
+
𝒮
|
| 876 |
+
𝑡
|
| 877 |
+
\mathcal{S}_{t}
|
| 878 |
+
and edges representing actions
|
| 879 |
+
𝒜
|
| 880 |
+
t
|
| 881 |
+
subscript
|
| 882 |
+
𝒜
|
| 883 |
+
𝑡
|
| 884 |
+
\mathcal{A}_{t}
|
| 885 |
+
. The search algorithm employed is a modified Monte Carlo Tree Search (MCTS), specifically adapted for the tasks of the SWE-Agent. Unlike prior approaches for web agents that utilize language models in the selection process
|
| 886 |
+
Koh et al. (
|
| 887 |
+
2024
|
| 888 |
+
); Zhang et al. (
|
| 889 |
+
2024b
|
| 890 |
+
)
|
| 891 |
+
, we deliberately choose not to rely on language models for node selection. Instead, we adopt a more straightforward heuristic-based selection function, similar to the approach used in AlphaZero
|
| 892 |
+
Silver et al. (
|
| 893 |
+
2016a
|
| 894 |
+
;
|
| 895 |
+
2018
|
| 896 |
+
)
|
| 897 |
+
. This decision is driven by the need for interpretability, efficiency, and the focus on tasks where heuristic-based exploration suffices to guide the agent effectively through complex software engineering environments.
|
| 898 |
+
At the core of our algorithm is a modified Upper Confidence Bound for Trees (UCT) selection criterion
|
| 899 |
+
Kocsis & Szepesvári (
|
| 900 |
+
2006
|
| 901 |
+
)
|
| 902 |
+
, which determines the next node to expand. This criterion balances exploitation of known high-reward actions with exploration of less-visited states. We introduce additional terms to encourage strategic exploration early in the search process, and to penalize over-exploration at later stages when convergence on the optimal solution is desired. The modified UCT function is expressed as:
|
| 903 |
+
U
|
| 904 |
+
|
| 905 |
+
C
|
| 906 |
+
|
| 907 |
+
T
|
| 908 |
+
|
| 909 |
+
(
|
| 910 |
+
s
|
| 911 |
+
,
|
| 912 |
+
a
|
| 913 |
+
)
|
| 914 |
+
=
|
| 915 |
+
e
|
| 916 |
+
|
| 917 |
+
x
|
| 918 |
+
|
| 919 |
+
p
|
| 920 |
+
|
| 921 |
+
l
|
| 922 |
+
|
| 923 |
+
o
|
| 924 |
+
|
| 925 |
+
i
|
| 926 |
+
|
| 927 |
+
t
|
| 928 |
+
|
| 929 |
+
a
|
| 930 |
+
|
| 931 |
+
t
|
| 932 |
+
|
| 933 |
+
i
|
| 934 |
+
|
| 935 |
+
o
|
| 936 |
+
|
| 937 |
+
n
|
| 938 |
+
+
|
| 939 |
+
e
|
| 940 |
+
|
| 941 |
+
x
|
| 942 |
+
|
| 943 |
+
p
|
| 944 |
+
|
| 945 |
+
l
|
| 946 |
+
|
| 947 |
+
o
|
| 948 |
+
|
| 949 |
+
r
|
| 950 |
+
|
| 951 |
+
a
|
| 952 |
+
|
| 953 |
+
t
|
| 954 |
+
|
| 955 |
+
i
|
| 956 |
+
|
| 957 |
+
o
|
| 958 |
+
|
| 959 |
+
n
|
| 960 |
+
+
|
| 961 |
+
e
|
| 962 |
+
|
| 963 |
+
a
|
| 964 |
+
|
| 965 |
+
r
|
| 966 |
+
|
| 967 |
+
l
|
| 968 |
+
|
| 969 |
+
y
|
| 970 |
+
|
| 971 |
+
_
|
| 972 |
+
|
| 973 |
+
d
|
| 974 |
+
|
| 975 |
+
e
|
| 976 |
+
|
| 977 |
+
p
|
| 978 |
+
|
| 979 |
+
t
|
| 980 |
+
|
| 981 |
+
h
|
| 982 |
+
|
| 983 |
+
_
|
| 984 |
+
|
| 985 |
+
b
|
| 986 |
+
|
| 987 |
+
o
|
| 988 |
+
|
| 989 |
+
n
|
| 990 |
+
|
| 991 |
+
u
|
| 992 |
+
|
| 993 |
+
s
|
| 994 |
+
−
|
| 995 |
+
l
|
| 996 |
+
|
| 997 |
+
a
|
| 998 |
+
|
| 999 |
+
t
|
| 1000 |
+
|
| 1001 |
+
e
|
| 1002 |
+
|
| 1003 |
+
_
|
| 1004 |
+
|
| 1005 |
+
d
|
| 1006 |
+
|
| 1007 |
+
e
|
| 1008 |
+
|
| 1009 |
+
p
|
| 1010 |
+
|
| 1011 |
+
t
|
| 1012 |
+
|
| 1013 |
+
h
|
| 1014 |
+
|
| 1015 |
+
_
|
| 1016 |
+
|
| 1017 |
+
p
|
| 1018 |
+
|
| 1019 |
+
e
|
| 1020 |
+
|
| 1021 |
+
n
|
| 1022 |
+
|
| 1023 |
+
a
|
| 1024 |
+
|
| 1025 |
+
l
|
| 1026 |
+
|
| 1027 |
+
t
|
| 1028 |
+
|
| 1029 |
+
y
|
| 1030 |
+
𝑈
|
| 1031 |
+
𝐶
|
| 1032 |
+
𝑇
|
| 1033 |
+
𝑠
|
| 1034 |
+
𝑎
|
| 1035 |
+
𝑒
|
| 1036 |
+
𝑥
|
| 1037 |
+
𝑝
|
| 1038 |
+
𝑙
|
| 1039 |
+
𝑜
|
| 1040 |
+
𝑖
|
| 1041 |
+
𝑡
|
| 1042 |
+
𝑎
|
| 1043 |
+
𝑡
|
| 1044 |
+
𝑖
|
| 1045 |
+
𝑜
|
| 1046 |
+
𝑛
|
| 1047 |
+
𝑒
|
| 1048 |
+
𝑥
|
| 1049 |
+
𝑝
|
| 1050 |
+
𝑙
|
| 1051 |
+
𝑜
|
| 1052 |
+
𝑟
|
| 1053 |
+
𝑎
|
| 1054 |
+
𝑡
|
| 1055 |
+
𝑖
|
| 1056 |
+
𝑜
|
| 1057 |
+
𝑛
|
| 1058 |
+
𝑒
|
| 1059 |
+
𝑎
|
| 1060 |
+
𝑟
|
| 1061 |
+
𝑙
|
| 1062 |
+
𝑦
|
| 1063 |
+
_
|
| 1064 |
+
𝑑
|
| 1065 |
+
𝑒
|
| 1066 |
+
𝑝
|
| 1067 |
+
𝑡
|
| 1068 |
+
ℎ
|
| 1069 |
+
_
|
| 1070 |
+
𝑏
|
| 1071 |
+
𝑜
|
| 1072 |
+
𝑛
|
| 1073 |
+
𝑢
|
| 1074 |
+
𝑠
|
| 1075 |
+
𝑙
|
| 1076 |
+
𝑎
|
| 1077 |
+
𝑡
|
| 1078 |
+
𝑒
|
| 1079 |
+
_
|
| 1080 |
+
𝑑
|
| 1081 |
+
𝑒
|
| 1082 |
+
𝑝
|
| 1083 |
+
𝑡
|
| 1084 |
+
ℎ
|
| 1085 |
+
_
|
| 1086 |
+
𝑝
|
| 1087 |
+
𝑒
|
| 1088 |
+
𝑛
|
| 1089 |
+
𝑎
|
| 1090 |
+
𝑙
|
| 1091 |
+
𝑡
|
| 1092 |
+
𝑦
|
| 1093 |
+
UCT(s,a)=exploitation+exploration+early\_depth\_bonus-late\_depth\_penalty
|
| 1094 |
+
(2)
|
| 1095 |
+
This can be expressed more formally as:
|
| 1096 |
+
U
|
| 1097 |
+
|
| 1098 |
+
C
|
| 1099 |
+
|
| 1100 |
+
T
|
| 1101 |
+
|
| 1102 |
+
(
|
| 1103 |
+
s
|
| 1104 |
+
,
|
| 1105 |
+
a
|
| 1106 |
+
)
|
| 1107 |
+
=
|
| 1108 |
+
V
|
| 1109 |
+
|
| 1110 |
+
(
|
| 1111 |
+
s
|
| 1112 |
+
,
|
| 1113 |
+
a
|
| 1114 |
+
)
|
| 1115 |
+
+
|
| 1116 |
+
C
|
| 1117 |
+
|
| 1118 |
+
ln
|
| 1119 |
+
|
| 1120 |
+
N
|
| 1121 |
+
|
| 1122 |
+
(
|
| 1123 |
+
s
|
| 1124 |
+
)
|
| 1125 |
+
N
|
| 1126 |
+
|
| 1127 |
+
(
|
| 1128 |
+
s
|
| 1129 |
+
,
|
| 1130 |
+
a
|
| 1131 |
+
)
|
| 1132 |
+
+
|
| 1133 |
+
α
|
| 1134 |
+
|
| 1135 |
+
e
|
| 1136 |
+
−
|
| 1137 |
+
β
|
| 1138 |
+
|
| 1139 |
+
(
|
| 1140 |
+
d
|
| 1141 |
+
−
|
| 1142 |
+
1
|
| 1143 |
+
)
|
| 1144 |
+
−
|
| 1145 |
+
γ
|
| 1146 |
+
|
| 1147 |
+
d
|
| 1148 |
+
𝑈
|
| 1149 |
+
𝐶
|
| 1150 |
+
𝑇
|
| 1151 |
+
𝑠
|
| 1152 |
+
𝑎
|
| 1153 |
+
𝑉
|
| 1154 |
+
𝑠
|
| 1155 |
+
𝑎
|
| 1156 |
+
𝐶
|
| 1157 |
+
𝑁
|
| 1158 |
+
𝑠
|
| 1159 |
+
𝑁
|
| 1160 |
+
𝑠
|
| 1161 |
+
𝑎
|
| 1162 |
+
𝛼
|
| 1163 |
+
superscript
|
| 1164 |
+
𝑒
|
| 1165 |
+
𝛽
|
| 1166 |
+
𝑑
|
| 1167 |
+
1
|
| 1168 |
+
𝛾
|
| 1169 |
+
𝑑
|
| 1170 |
+
UCT(s,a)=V(s,a)+C\sqrt{\frac{\ln N(s)}{N(s,a)}}+\alpha e^{-\beta(d-1)}-\gamma\sqrt{d}
|
| 1171 |
+
(3)
|
| 1172 |
+
V
|
| 1173 |
+
|
| 1174 |
+
(
|
| 1175 |
+
s
|
| 1176 |
+
,
|
| 1177 |
+
a
|
| 1178 |
+
)
|
| 1179 |
+
𝑉
|
| 1180 |
+
𝑠
|
| 1181 |
+
𝑎
|
| 1182 |
+
V(s,a)
|
| 1183 |
+
is the value estimate of the state-action pair
|
| 1184 |
+
,
|
| 1185 |
+
N
|
| 1186 |
+
|
| 1187 |
+
(
|
| 1188 |
+
s
|
| 1189 |
+
,
|
| 1190 |
+
a
|
| 1191 |
+
)
|
| 1192 |
+
𝑁
|
| 1193 |
+
𝑠
|
| 1194 |
+
𝑎
|
| 1195 |
+
N(s,a)
|
| 1196 |
+
is the number of times the state-action pair
|
| 1197 |
+
(
|
| 1198 |
+
s
|
| 1199 |
+
,
|
| 1200 |
+
a
|
| 1201 |
+
)
|
| 1202 |
+
𝑠
|
| 1203 |
+
𝑎
|
| 1204 |
+
(s,a)
|
| 1205 |
+
has been visited,
|
| 1206 |
+
N
|
| 1207 |
+
|
| 1208 |
+
(
|
| 1209 |
+
s
|
| 1210 |
+
)
|
| 1211 |
+
𝑁
|
| 1212 |
+
𝑠
|
| 1213 |
+
N(s)
|
| 1214 |
+
is the visit count of state
|
| 1215 |
+
s
|
| 1216 |
+
𝑠
|
| 1217 |
+
s
|
| 1218 |
+
,
|
| 1219 |
+
d
|
| 1220 |
+
𝑑
|
| 1221 |
+
d
|
| 1222 |
+
is the depth of the node in the search tree, and
|
| 1223 |
+
C
|
| 1224 |
+
𝐶
|
| 1225 |
+
C
|
| 1226 |
+
,
|
| 1227 |
+
α
|
| 1228 |
+
𝛼
|
| 1229 |
+
\alpha
|
| 1230 |
+
,
|
| 1231 |
+
β
|
| 1232 |
+
𝛽
|
| 1233 |
+
\beta
|
| 1234 |
+
, and
|
| 1235 |
+
γ
|
| 1236 |
+
𝛾
|
| 1237 |
+
\gamma
|
| 1238 |
+
are constants that control the balance between exploration, exploitation, and depth-dependent rewards and penalties.
|
| 1239 |
+
This formulation is inspired by the way software engineers explore potential solutions to a task. In practice, an engineer’s search process can be broken down into the following key phases, which our algorithm mirrors:
|
| 1240 |
+
Early Exploration
|
| 1241 |
+
: Initially, an engineer explores a wide variety of potential approaches to fully understand the problem and identify promising strategies. This is encouraged in our algorithm by the
|
| 1242 |
+
e
|
| 1243 |
+
|
| 1244 |
+
a
|
| 1245 |
+
|
| 1246 |
+
r
|
| 1247 |
+
|
| 1248 |
+
l
|
| 1249 |
+
|
| 1250 |
+
y
|
| 1251 |
+
|
| 1252 |
+
_
|
| 1253 |
+
|
| 1254 |
+
d
|
| 1255 |
+
|
| 1256 |
+
e
|
| 1257 |
+
|
| 1258 |
+
p
|
| 1259 |
+
|
| 1260 |
+
t
|
| 1261 |
+
|
| 1262 |
+
h
|
| 1263 |
+
|
| 1264 |
+
_
|
| 1265 |
+
|
| 1266 |
+
b
|
| 1267 |
+
|
| 1268 |
+
o
|
| 1269 |
+
|
| 1270 |
+
n
|
| 1271 |
+
|
| 1272 |
+
u
|
| 1273 |
+
|
| 1274 |
+
s
|
| 1275 |
+
𝑒
|
| 1276 |
+
𝑎
|
| 1277 |
+
𝑟
|
| 1278 |
+
𝑙
|
| 1279 |
+
𝑦
|
| 1280 |
+
_
|
| 1281 |
+
𝑑
|
| 1282 |
+
𝑒
|
| 1283 |
+
𝑝
|
| 1284 |
+
𝑡
|
| 1285 |
+
ℎ
|
| 1286 |
+
_
|
| 1287 |
+
𝑏
|
| 1288 |
+
𝑜
|
| 1289 |
+
𝑛
|
| 1290 |
+
𝑢
|
| 1291 |
+
𝑠
|
| 1292 |
+
early\_depth\_bonus
|
| 1293 |
+
, represented by the term
|
| 1294 |
+
α
|
| 1295 |
+
|
| 1296 |
+
e
|
| 1297 |
+
−
|
| 1298 |
+
β
|
| 1299 |
+
|
| 1300 |
+
(
|
| 1301 |
+
d
|
| 1302 |
+
−
|
| 1303 |
+
1
|
| 1304 |
+
)
|
| 1305 |
+
𝛼
|
| 1306 |
+
superscript
|
| 1307 |
+
𝑒
|
| 1308 |
+
𝛽
|
| 1309 |
+
𝑑
|
| 1310 |
+
1
|
| 1311 |
+
\alpha e^{-\beta(d-1)}
|
| 1312 |
+
, which rewards exploration at shallow depths, simulating the early phases of wide exploration.
|
| 1313 |
+
Convergence and Exploitation
|
| 1314 |
+
: As the engineer gains more information and narrows down the options, the focus shifts to exploiting the most effective solution paths. This transition is handled by the standard UCT exploitation term
|
| 1315 |
+
V
|
| 1316 |
+
|
| 1317 |
+
(
|
| 1318 |
+
s
|
| 1319 |
+
,
|
| 1320 |
+
a
|
| 1321 |
+
)
|
| 1322 |
+
𝑉
|
| 1323 |
+
𝑠
|
| 1324 |
+
𝑎
|
| 1325 |
+
V(s,a)
|
| 1326 |
+
and is further reinforced by the
|
| 1327 |
+
l
|
| 1328 |
+
|
| 1329 |
+
a
|
| 1330 |
+
|
| 1331 |
+
t
|
| 1332 |
+
|
| 1333 |
+
e
|
| 1334 |
+
|
| 1335 |
+
_
|
| 1336 |
+
|
| 1337 |
+
d
|
| 1338 |
+
|
| 1339 |
+
e
|
| 1340 |
+
|
| 1341 |
+
p
|
| 1342 |
+
|
| 1343 |
+
t
|
| 1344 |
+
|
| 1345 |
+
h
|
| 1346 |
+
|
| 1347 |
+
_
|
| 1348 |
+
|
| 1349 |
+
p
|
| 1350 |
+
|
| 1351 |
+
e
|
| 1352 |
+
|
| 1353 |
+
n
|
| 1354 |
+
|
| 1355 |
+
a
|
| 1356 |
+
|
| 1357 |
+
l
|
| 1358 |
+
|
| 1359 |
+
t
|
| 1360 |
+
|
| 1361 |
+
y
|
| 1362 |
+
𝑙
|
| 1363 |
+
𝑎
|
| 1364 |
+
𝑡
|
| 1365 |
+
𝑒
|
| 1366 |
+
_
|
| 1367 |
+
𝑑
|
| 1368 |
+
𝑒
|
| 1369 |
+
𝑝
|
| 1370 |
+
𝑡
|
| 1371 |
+
ℎ
|
| 1372 |
+
_
|
| 1373 |
+
𝑝
|
| 1374 |
+
𝑒
|
| 1375 |
+
𝑛
|
| 1376 |
+
𝑎
|
| 1377 |
+
𝑙
|
| 1378 |
+
𝑡
|
| 1379 |
+
𝑦
|
| 1380 |
+
late\_depth\_penalty
|
| 1381 |
+
(
|
| 1382 |
+
−
|
| 1383 |
+
γ
|
| 1384 |
+
|
| 1385 |
+
d
|
| 1386 |
+
𝛾
|
| 1387 |
+
𝑑
|
| 1388 |
+
-\gamma\sqrt{d}
|
| 1389 |
+
), which discourages over-exploration as the agent delves deeper into the search tree.
|
| 1390 |
+
Quick Abandonment of Poor Strategies
|
| 1391 |
+
: Software engineers are also adept at abandoning poor strategies when new information indicates that a particular approach is not viable. We capture this behavior by implementing a simple heuristic rule that abandons nodes associated with consecutive low rewards, ensuring that the agent does not waste resources on unproductive trajectories.
|
| 1392 |
+
At each step, the node with the highest UCT value is selected for expansion, formalized as:
|
| 1393 |
+
s
|
| 1394 |
+
∗
|
| 1395 |
+
=
|
| 1396 |
+
arg
|
| 1397 |
+
|
| 1398 |
+
max
|
| 1399 |
+
(
|
| 1400 |
+
s
|
| 1401 |
+
,
|
| 1402 |
+
a
|
| 1403 |
+
)
|
| 1404 |
+
|
| 1405 |
+
U
|
| 1406 |
+
|
| 1407 |
+
C
|
| 1408 |
+
|
| 1409 |
+
T
|
| 1410 |
+
|
| 1411 |
+
(
|
| 1412 |
+
s
|
| 1413 |
+
,
|
| 1414 |
+
a
|
| 1415 |
+
)
|
| 1416 |
+
superscript
|
| 1417 |
+
𝑠
|
| 1418 |
+
subscript
|
| 1419 |
+
arg
|
| 1420 |
+
max
|
| 1421 |
+
𝑠
|
| 1422 |
+
𝑎
|
| 1423 |
+
𝑈
|
| 1424 |
+
𝐶
|
| 1425 |
+
𝑇
|
| 1426 |
+
𝑠
|
| 1427 |
+
𝑎
|
| 1428 |
+
s^{*}=\operatorname*{arg\,max}_{(s,a)}UCT(s,a)
|
| 1429 |
+
(4)
|
| 1430 |
+
This approach effectively mimics the decision-making process of a software engineer, who balances exploration of potential strategies with a focus on converging towards the optimal solution, while remaining flexible enough to backtrack when necessary. By incorporating heuristic feedback and depth-based adjustments, the algorithm avoids getting stuck in unproductive paths and enhances the agent’s ability to identify high-reward strategies with minimal computational overhead
|
| 1431 |
+
Appendix
|
| 1432 |
+
6
|
| 1433 |
+
.
|
| 1434 |
+
3.4.1
|
| 1435 |
+
Discriminator Agent
|
| 1436 |
+
The final stage of SWE-Search involves the Discriminator Agent, whose role is to evaluate the candidate solutions generated by the search process and select the one most likely to resolve the issue at hand. This module accepts up to five final solutions produced by the search and engages in a multi-agent debate to determine the most promising option. Drawing inspiration from recent work on persuasive multi-agent debates
|
| 1437 |
+
(Khan et al.,
|
| 1438 |
+
2024
|
| 1439 |
+
; Amayuelas et al.,
|
| 1440 |
+
2024
|
| 1441 |
+
)
|
| 1442 |
+
, the Discriminator leverages the collective reasoning of multiple agents to ensure a more robust final selection. Configuration and hyperparameter details can be found in
|
| 1443 |
+
Table
|
| 1444 |
+
2
|
| 1445 |
+
.
|
| 1446 |
+
In this stage, agents are presented with the original problem statement and candidate solutions. They engage in a structured debate to determine the most effective solution, supporting their choices with logical reasoning and evidence from the search process. This debate encourages a thorough exploration of trade-offs between solutions, potentially uncovering strengths or weaknesses not evident during individual searches. Finally, a judge agent evaluates the arguments and selects the solution deemed most likely to resolve the issue. This process simulates the collaborative decision-making in software engineering teams, where diverse perspectives lead to a more thorough evaluation of candidate solutions, ultimately increasing the likelihood of identifying the most optimal outcome.
|
| 1447 |
+
The discriminator process not only enhances the robustness of the final solution but also adds transparency, as the reasoning behind the choice is clearly articulated and evaluated. This ensures that the selected solution is well-reasoned and thoroughly vetted before implementation.
|
| 1448 |
+
Figure 2:
|
| 1449 |
+
Hindsight feedback error correction.
|
| 1450 |
+
Instance sympy__sympy-15678, SWE-Search with Qwen2.5-72B-Instruct. Initially, the Action Agent performs edits and runs tests, which pass. It prematurely concludes the search. Without actually knowing the proposed solution does not resolve the issue, the Value Agent identifies potentially missed tests and assigns a low reward. Upon re-expansion using the Value Agent’s feedback, new tests fail, prompting the Action Agent to make additional edits, which result in a preferred solution which ultimately resolves the issue.
|
| 1451 |
+
4
|
| 1452 |
+
Experiments
|
| 1453 |
+
Benchmark
|
| 1454 |
+
For our experiments, we utilize SWE-bench Lite, a curated subset of the official SWE-bench, containing 300 instances. This dataset is specifically designed to be self-contained and focuses primarily on evaluating functional bug fixes, providing a controlled environment to assess the performance of our system.
|
| 1455 |
+
Evaluation Metrics
|
| 1456 |
+
We use two metrics: resolve rate (
|
| 1457 |
+
Pass@1
|
| 1458 |
+
) and
|
| 1459 |
+
Pass@5
|
| 1460 |
+
. Resolve rate is the percentage of issues successfully resolved, measuring overall effectiveness. Pass@5 is the percentage of issues where a correct solution is found within five attempts. This allows us to assess the efficiency of the search in identifying successful bug fixes within a limited number of iterations.
|
| 1461 |
+
Baselines
|
| 1462 |
+
Software agents leverage diverse tools, architectures, and models, leading to variability in their performance on subsets of the SWE-bench Lite dataset
|
| 1463 |
+
(Zhang et al.,
|
| 1464 |
+
2024a
|
| 1465 |
+
)
|
| 1466 |
+
. For comparison, we build upon the moatless-tools framework
|
| 1467 |
+
(Örwall,
|
| 1468 |
+
2024
|
| 1469 |
+
)
|
| 1470 |
+
, a high-performing open-source agent commonly used in research settings
|
| 1471 |
+
(Chowdhury et al.,
|
| 1472 |
+
2024
|
| 1473 |
+
)
|
| 1474 |
+
. To isolate the impact of our search approach, we adapt moatless-tools as our baseline, referred to as moatless-adapted. This allows us to fairly compare the performance of SWE-Search against moatless-adapted across various models, including two closed-source models (GPT-4o, GPT-4o-mini) and three open-source models (Qwen2.5-72B-Instruct
|
| 1475 |
+
(Yang et al.,
|
| 1476 |
+
2024a
|
| 1477 |
+
)
|
| 1478 |
+
, Llama-3.1-70B-Instruct
|
| 1479 |
+
(Dubey et al.,
|
| 1480 |
+
2024
|
| 1481 |
+
)
|
| 1482 |
+
, and DeepSeek-V2.5
|
| 1483 |
+
(DeepSeek-AI et al.,
|
| 1484 |
+
2024
|
| 1485 |
+
)
|
| 1486 |
+
). We also reference official moatless-tools GPT-4o results on SWE-bench Lite to ensure a fair and consistent comparison.
|
| 1487 |
+
Implementation Details
|
| 1488 |
+
For consistency, we use identical prompts across all models. In SWE-Search, we limit each node to a maximum of three expansions and cap the total search iterations at 100. Further details on model hyperparameters can be found in
|
| 1489 |
+
Appendix,
|
| 1490 |
+
2
|
| 1491 |
+
.
|
| 1492 |
+
Table 1:
|
| 1493 |
+
Resolve Rate Comparison, SWE-bench Lite
|
| 1494 |
+
Model
|
| 1495 |
+
Moatless-v1
|
| 1496 |
+
Moatless-adapted
|
| 1497 |
+
SWE-Search
|
| 1498 |
+
%
|
| 1499 |
+
Δ
|
| 1500 |
+
Δ
|
| 1501 |
+
\Delta
|
| 1502 |
+
GPT-4o
|
| 1503 |
+
24.3
|
| 1504 |
+
25.7
|
| 1505 |
+
31.0
|
| 1506 |
+
+17
|
| 1507 |
+
GPT-4o-mini
|
| 1508 |
+
–
|
| 1509 |
+
13.0
|
| 1510 |
+
17.0
|
| 1511 |
+
+24
|
| 1512 |
+
Qwen-2.5-72b-Instruct
|
| 1513 |
+
–
|
| 1514 |
+
18.0
|
| 1515 |
+
24.7
|
| 1516 |
+
+27
|
| 1517 |
+
Deepseek-V2.5
|
| 1518 |
+
–
|
| 1519 |
+
16.3
|
| 1520 |
+
21.0
|
| 1521 |
+
+22
|
| 1522 |
+
Llama-3.1-70b-Instruct
|
| 1523 |
+
–
|
| 1524 |
+
13.6
|
| 1525 |
+
17.7
|
| 1526 |
+
+23
|
| 1527 |
+
Mean %
|
| 1528 |
+
Δ
|
| 1529 |
+
Δ
|
| 1530 |
+
\Delta
|
| 1531 |
+
+23
|
| 1532 |
+
4.1
|
| 1533 |
+
Experimental Results
|
| 1534 |
+
4.1.1
|
| 1535 |
+
SWE-Search Surpasses all Corresponding Base Agents and Enables Smaller, Open Source Models to Approach GPT-4o
|
| 1536 |
+
On average, SWE-Search outperforms the baseline agent across all five models, achieving a 23% relative improvement
|
| 1537 |
+
(Table
|
| 1538 |
+
1
|
| 1539 |
+
)
|
| 1540 |
+
. Notably, SWE-Search with Qwen-2.5-72B-Instruct exceeds the performance of GPT-4o using the original Moatless-v1 framework, and closely matches its performance when compared with the Moatless-adapted agent, with only a slight difference (
|
| 1541 |
+
Δ
|
| 1542 |
+
=
|
| 1543 |
+
−
|
| 1544 |
+
1
|
| 1545 |
+
%
|
| 1546 |
+
Δ
|
| 1547 |
+
percent
|
| 1548 |
+
1
|
| 1549 |
+
\Delta=-1\%
|
| 1550 |
+
). Interestingly, all five models demonstrate significant improvement when utilizing the proposed approach, with consistent gains across different models.
|
| 1551 |
+
4.1.2
|
| 1552 |
+
Search Enables Agents to Make Better Use of More Flexibility
|
| 1553 |
+
To prevent goal divergence, most agents, including moatless-tools, rely on strict transition rules, where state transitions follow predetermined sequences (e.g., Search
|
| 1554 |
+
→
|
| 1555 |
+
→
|
| 1556 |
+
\rightarrow
|
| 1557 |
+
Identify, Plan
|
| 1558 |
+
→
|
| 1559 |
+
→
|
| 1560 |
+
\rightarrow
|
| 1561 |
+
Edit). In moatless-adapted, we introduce a more flexible transition logic that allows a Plan state to transition into any other state type. This added flexibility has both advantages and drawbacks. On the positive side, it enables the agent to autonomously correct its trajectory without external feedback, particularly when the necessary adjustments span only a limited portion of the task. However, this increased flexibility also introduces the risk of the agent becoming trapped in infinite loops. Without a high-level control mechanism to detect and mitigate these situations, the agent may fail to recover from such loops. This trade-off is evident in the modest performance difference between Moatless-v1 and moatless-adapted, with a slight performance improvement of only 1.4% (
|
| 1562 |
+
Table
|
| 1563 |
+
1
|
| 1564 |
+
).
|
| 1565 |
+
4.1.3
|
| 1566 |
+
Impact of Hindsight Feedback on Agent Performance
|
| 1567 |
+
One key advantage of utilizing LLMs as general value functions is their dual ability to provide both quantitative value estimates and qualitative assessments in natural language. These qualitative insights can significantly enhance the agent’s action generation and search process by offering detailed feedback on potential errors or overlooked aspects of the task. In practice, feedback was also crucial in eliciting diversity in the actions taken by the agent, as without it, the agent would often take very similar actions when re-expanding from a parent node.
|
| 1568 |
+
As shown in
|
| 1569 |
+
Figure
|
| 1570 |
+
2
|
| 1571 |
+
, this mechanism plays a critical role in improving the agent’s performance. During the initial expansion, the agent prematurely concludes that the task is complete. However, the value function correctly identifies gaps in the test coverage, specifically in addressing potential corner cases, and assigns a low reward. This feedback prompts the agent to re-expand the parent state, leading to the introduction of new tests, which subsequently fail. The agent then performs a series of edits (summarized in the figure for brevity), ultimately resolving the task correctly. Empirically, we observe that the instances unresolved by moatless-adapted but successfully solved by SWE-Search are often attributed to this search-and-feedback loop, where iterative feedback drives the agent toward a correct solution.
|
| 1572 |
+
4.2
|
| 1573 |
+
Importance of Comprehensive State Information for Value Function Performance
|
| 1574 |
+
Model
|
| 1575 |
+
Pass@1
|
| 1576 |
+
Pass@5
|
| 1577 |
+
GPT-4o
|
| 1578 |
+
31.0
|
| 1579 |
+
34.0
|
| 1580 |
+
GPT-4o-mini
|
| 1581 |
+
17.0
|
| 1582 |
+
22.3
|
| 1583 |
+
Qwen-2.5-72b-Instruct
|
| 1584 |
+
24.7
|
| 1585 |
+
25.7
|
| 1586 |
+
Deepseek-V2.5
|
| 1587 |
+
21.0
|
| 1588 |
+
23.3
|
| 1589 |
+
Llama-3.1-70b-Instruct
|
| 1590 |
+
21.0
|
| 1591 |
+
22.3
|
| 1592 |
+
Figure 3:
|
| 1593 |
+
SWE-bench SWE-Search results
|
| 1594 |
+
The effectiveness of SWE-Search hinges on the value function’s ability to accurately differentiate between desirable and undesirable states, and to provide actionable feedback that drives improvement. However, our experiments revealed that the value function sometimes failed to recognize critical decision points in the search tree. It frequently misinterpreted the purpose of certain actions, leading to the undervaluation of effective strategies by assigning low rewards. As shown in
|
| 1595 |
+
Figure
|
| 1596 |
+
4
|
| 1597 |
+
a
|
| 1598 |
+
, before the introduction of state-specific value prompts, the agent consistently assigned low rewards even when the Action Agent correctly identified the need for additional context, such as locating relevant files. This issue persisted despite the agent successfully identifying the files later. By implementing state-specific prompts across core state clusters (Searching, Planning, Editing), the value function became significantly more adept at interpreting the intent behind actions and evaluating their outcomes within each state. For further details on experiments distinguishing between effective and ineffective states, refer to
|
| 1599 |
+
Appendix
|
| 1600 |
+
8
|
| 1601 |
+
.
|
| 1602 |
+
Figure 4:
|
| 1603 |
+
(a) Importance of state-specific value prompts.
|
| 1604 |
+
On the left and right are the respective Value Agents’ outputs with and without state-specific prompts. While the action in both cases is effective in finding the right file, the non-state-specific scenario does not recognize this and assigns a low reward. On the contrary, the state-specific prompt correctly assigns a high reward to this state.
|
| 1605 |
+
(b) Performance scaling with search depth across different language models.
|
| 1606 |
+
The graph shows the number of issues resolved as a function of the number of transitions (search iterations) for all models used.
|
| 1607 |
+
Scaling SWE agents with Inference-time Compute
|
| 1608 |
+
The success of large language models (LLMs) has traditionally been attributed to the expansion of training data and model size, i.e., training-time compute
|
| 1609 |
+
(Wei et al.,
|
| 1610 |
+
2022
|
| 1611 |
+
; Chung et al.,
|
| 1612 |
+
2022
|
| 1613 |
+
)
|
| 1614 |
+
. Recently, researchers have started exploring how different methods scale with inference-time
|
| 1615 |
+
(OpenAI,
|
| 1616 |
+
2024
|
| 1617 |
+
; Snell et al.,
|
| 1618 |
+
2024
|
| 1619 |
+
; Dubey et al.,
|
| 1620 |
+
2024
|
| 1621 |
+
)
|
| 1622 |
+
. Here, we study the performance of software engineering agents through increased inference-time compute. As shown in
|
| 1623 |
+
Figure
|
| 1624 |
+
4
|
| 1625 |
+
b
|
| 1626 |
+
, increasing search iterations leads to a consistent rise in the number of resolved issues. To ensure experimental feasibility across the 300 instances in the SWE-bench Lite dataset, we applied conservative parameters (maximum iterations
|
| 1627 |
+
=
|
| 1628 |
+
100
|
| 1629 |
+
absent
|
| 1630 |
+
100
|
| 1631 |
+
=100
|
| 1632 |
+
, maximum expansions per node
|
| 1633 |
+
=
|
| 1634 |
+
3
|
| 1635 |
+
absent
|
| 1636 |
+
3
|
| 1637 |
+
=3
|
| 1638 |
+
). Approaches like SWE-Search enable the allocation of greater resources to specific challenges, such as addressing critical software vulnerabilities
|
| 1639 |
+
(Rigaki et al.,
|
| 1640 |
+
2024
|
| 1641 |
+
; Fang et al.,
|
| 1642 |
+
2024
|
| 1643 |
+
)
|
| 1644 |
+
, offering a scalable solution to complex tasks.
|
| 1645 |
+
Figure 5:
|
| 1646 |
+
(a) Value Function vs. Discriminator Comparison.
|
| 1647 |
+
Comparison of value function vs. discriminator ability to discern the final solution that resolved the issue when there is one. The discriminator performs better across all models except GPT-4o-mini. DeepSeek-V2.5 had the smallest disparity between the two methods, suggesting an ability to act as a well-calibrated value function.
|
| 1648 |
+
(b) Model-Specific Issue Resolution.
|
| 1649 |
+
Venn diagram of resolved issues by model. Each model can solve a handful of unique instances.
|
| 1650 |
+
Convergence of Value Function and Discriminator to Right Solution
|
| 1651 |
+
The search process can yield multiple proposed solutions. Ideally, the mean trajectory value of the the proposed solution that resolves the issue will always be the highest, which would yields the ideal performance of the agent
|
| 1652 |
+
(Table
|
| 1653 |
+
3
|
| 1654 |
+
)
|
| 1655 |
+
. In practice, the value function successfully converged on the correct solution 73% of the time on average across the five models. The discriminator module performed even better, increasing the proportion of correct solutions selected to 84%. While in typical large action spaces, Monte Carlo Tree Search (MCTS) is run for thousands of iterations
|
| 1656 |
+
(Silver et al.,
|
| 1657 |
+
2016b
|
| 1658 |
+
)
|
| 1659 |
+
, the value function’s success rate remains impressive given the computational constraints. However, SWE-Search could further benefit from enhanced methods for identifying the correct solutions more consistently, allowing it to fully reach its potential.
|
| 1660 |
+
Different Models can Resolve Vastly Different Issue Subsets
|
| 1661 |
+
When comparing the resolved instances across the five models, we observed significant diversity in the subsets of issues each model successfully solved. As shown in
|
| 1662 |
+
Figure
|
| 1663 |
+
5
|
| 1664 |
+
, each model managed to resolve at least one unique instance. Notably, a surprising number of issues (33) were solved by other models but not by GPT-4o. This suggests that model diversity could play an important role, at least in the short term, in enhancing the performance of SWE-agents.
|
| 1665 |
+
5
|
| 1666 |
+
Discussion and Conclusion
|
| 1667 |
+
In this paper, we introduced SWE-Search, a general framework that integrates Monte Carlo Tree Search (MCTS) and qualitative feedback to enhance the performance of software engineering agents. The proposed approach demonstrated improvements over different baseline models, highlighting the potential of search-based methods in software engineering tasks.
|
| 1668 |
+
One of the key advantages of search-based approaches, as demonstrated in our work, is their ability to scale performance with increased inference-time compute. This flexibility allows the system to adapt to problems that require higher computational resources, such as discovering software vulnerabilities or even generating large codebases from scratch. Future research should focus on two main directions: (a) investigating how search agents scale with computational resources, and (b) expanding the application of software agent search to a broader range of complex use cases.
|
| 1669 |
+
Given that search techniques like MCTS closely resemble the problem-solving processes of human software engineers, we expect these methods to become increasingly prevalent in agent-driven systems. As the nature of software engineering tasks evolves, system architectures will need to become more fluid and adaptable, fully leveraging the potential of search-based techniques. This evolution will likely lead to the development of larger, more general agentic systems capable of tackling a wide array of software engineering challenges.
|
| 1670 |
+
References
|
| 1671 |
+
Amayuelas et al. (2024)
|
| 1672 |
+
Alfonso Amayuelas, Xianjun Yang, Antonis Antoniades, Wenyue Hua, Liangming Pan, and William Wang.
|
| 1673 |
+
Multiagent collaboration attack: Investigating adversarial attacks in large language model collaborations via debate, 2024.
|
| 1674 |
+
URL
|
| 1675 |
+
https://arxiv.org/abs/2406.14711
|
| 1676 |
+
.
|
| 1677 |
+
Black (2005)
|
| 1678 |
+
Paul E. Black.
|
| 1679 |
+
greedy algorithm, feb 2005.
|
| 1680 |
+
URL
|
| 1681 |
+
https://www.nist.gov/dads/HTML/greedyalgo.html
|
| 1682 |
+
.
|
| 1683 |
+
Accessed: TODAY.
|
| 1684 |
+
Bounsi et al. (2024)
|
| 1685 |
+
Wilfried Bounsi, Borja Ibarz, Andrew Dudzik, Jessica B. Hamrick, Larisa Markeeva, Alex Vitvitskyi, Razvan Pascanu, and Petar Veličković.
|
| 1686 |
+
Transformers meet neural algorithmic reasoners, 2024.
|
| 1687 |
+
URL
|
| 1688 |
+
https://arxiv.org/abs/2406.09308
|
| 1689 |
+
.
|
| 1690 |
+
Chowdhury et al. (2024)
|
| 1691 |
+
Neil Chowdhury, James Aung, Chan Jun Shern, Oliver Jaffe, Dane Sherburn, Giulio Starace, Evan Mays, Rachel Dias, Marwan Aljubeh, Mia Glaese, Carlos E. Jimenez, John Yang, Kevin Liu, and Aleksander Madry.
|
| 1692 |
+
Introducing SWE-bench verified, August 2024.
|
| 1693 |
+
URL
|
| 1694 |
+
https://openai.com/research/introducing-swe-bench-verified
|
| 1695 |
+
.
|
| 1696 |
+
OpenAI Blog.
|
| 1697 |
+
Chung et al. (2022)
|
| 1698 |
+
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei.
|
| 1699 |
+
Scaling instruction-finetuned language models, 2022.
|
| 1700 |
+
URL
|
| 1701 |
+
https://arxiv.org/abs/2210.11416
|
| 1702 |
+
.
|
| 1703 |
+
Cormen et al. (2009)
|
| 1704 |
+
Thomas H. Cormen, Charles E. Leiserson, Ronald L. Rivest, and Clifford Stein.
|
| 1705 |
+
Introduction to Algorithms, Third Edition
|
| 1706 |
+
.
|
| 1707 |
+
The MIT Press, 3rd edition, 2009.
|
| 1708 |
+
ISBN 0262033844.
|
| 1709 |
+
DeepSeek-AI et al. (2024)
|
| 1710 |
+
DeepSeek-AI, Qihao Zhu, Daya Guo, Zhihong Shao, Dejian Yang, Peiyi Wang, Runxin Xu, Y. Wu, Yukun Li, Huazuo Gao, Shirong Ma, Wangding Zeng, Xiao Bi, Zihui Gu, Hanwei Xu, Damai Dai, Kai Dong, Liyue Zhang, Yishi Piao, Zhibin Gou, Zhenda Xie, Zhewen Hao, Bingxuan Wang, Junxiao Song, Deli Chen, Xin Xie, Kang Guan, Yuxiang You, Aixin Liu, Qiushi Du, Wenjun Gao, Xuan Lu, Qinyu Chen, Yaohui Wang, Chengqi Deng, Jiashi Li, Chenggang Zhao, Chong Ruan, Fuli Luo, and Wenfeng Liang.
|
| 1711 |
+
Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence, 2024.
|
| 1712 |
+
URL
|
| 1713 |
+
https://arxiv.org/abs/2406.11931
|
| 1714 |
+
.
|
| 1715 |
+
Du et al. (2023)
|
| 1716 |
+
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch.
|
| 1717 |
+
Improving factuality and reasoning in language models through multiagent debate, 2023.
|
| 1718 |
+
URL
|
| 1719 |
+
https://arxiv.org/abs/2305.14325
|
| 1720 |
+
.
|
| 1721 |
+
Dubey et al. (2024)
|
| 1722 |
+
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov,
|
| 1723 |
+
Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Olivier
|
| 1724 |
+
Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vladan Petrovic, Weiwei Chu,
|
| 1725 |
+
Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaoqing Ellen Tan, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aaron Grattafiori, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alex Vaughan, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Franco, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl
|
| 1726 |
+
Parker, Carly Burton, Catalina Mejia, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, Danny Wyatt, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Firat Ozgenel, Francesco Caggioni, Francisco Guzmán, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Govind Thattai, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Irina-Elena Veliche, Itai Gat, Jake Weissman, James
|
| 1727 |
+
Geboski, James Kohli, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Karthik Prasad, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kun Huang, Kunal Chawla, Kushal Lakhotia, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Maria Tsimpoukelli, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini
|
| 1728 |
+
Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikolay Pavlovich Laptev, Ning Dong, Ning Zhang, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Rohan Maheswari, Russ Howes, Ruty Rinott, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta,
|
| 1729 |
+
Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Kohler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vítor Albiero, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaofang Wang, Xiaojian Wu, Xiaolan Wang, Xide Xia, Xilun Wu, Xinbo Gao, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yuchen Hao, Yundi Qian, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, and Zhiwei Zhao.
|
| 1730 |
+
The llama 3 herd of models, 2024.
|
| 1731 |
+
URL
|
| 1732 |
+
https://arxiv.org/abs/2407.21783
|
| 1733 |
+
.
|
| 1734 |
+
Faldor et al. (2024)
|
| 1735 |
+
Maxence Faldor, Jenny Zhang, Antoine Cully, and Jeff Clune.
|
| 1736 |
+
Omni-epic: Open-endedness via models of human notions of interestingness with environments programmed in code, 2024.
|
| 1737 |
+
URL
|
| 1738 |
+
https://arxiv.org/abs/2405.15568
|
| 1739 |
+
.
|
| 1740 |
+
Fang et al. (2024)
|
| 1741 |
+
Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang.
|
| 1742 |
+
Teams of llm agents can exploit zero-day vulnerabilities, 2024.
|
| 1743 |
+
URL
|
| 1744 |
+
https://arxiv.org/abs/2406.01637
|
| 1745 |
+
.
|
| 1746 |
+
Fawzi et al. (2022)
|
| 1747 |
+
A. Fawzi, M. Balog, A. Huang, T. Hubert, B. Romera-Paredes, M. Barekatain, A. Novikov, F. J. R. Ruiz, J. Schrittwieser, G. Swirszcz, D. Silver, D. Hassabis, and P. Kohli.
|
| 1748 |
+
Discovering faster matrix multiplication algorithms with reinforcement learning.
|
| 1749 |
+
Nature
|
| 1750 |
+
, 610(7930):47–53, 2022.
|
| 1751 |
+
doi:
|
| 1752 |
+
10.1038/s41586-022-05172-4
|
| 1753 |
+
.
|
| 1754 |
+
Hart et al. (1968)
|
| 1755 |
+
Peter E. Hart, Nils J. Nilsson, and Bertram Raphael.
|
| 1756 |
+
A formal basis for the heuristic determination of minimum cost paths.
|
| 1757 |
+
IEEE Trans. Syst. Sci. Cybern.
|
| 1758 |
+
, 4(2):100–107, 1968.
|
| 1759 |
+
doi:
|
| 1760 |
+
10.1109/TSSC.1968.300136
|
| 1761 |
+
.
|
| 1762 |
+
URL
|
| 1763 |
+
https://doi.org/10.1109/TSSC.1968.300136
|
| 1764 |
+
.
|
| 1765 |
+
Hu et al. (2024)
|
| 1766 |
+
Shengran Hu, Cong Lu, and Jeff Clune.
|
| 1767 |
+
Automated design of agentic systems, 2024.
|
| 1768 |
+
URL
|
| 1769 |
+
https://arxiv.org/abs/2408.08435
|
| 1770 |
+
.
|
| 1771 |
+
Jimenez et al. (2024)
|
| 1772 |
+
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.
|
| 1773 |
+
Swe-bench: Can language models resolve real-world github issues?, 2024.
|
| 1774 |
+
URL
|
| 1775 |
+
https://arxiv.org/abs/2310.06770
|
| 1776 |
+
.
|
| 1777 |
+
Jumper et al. (2021)
|
| 1778 |
+
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera-Paredes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Petersen, D. Reiman, E. Clancy, M. Zielinski, M. Steinegger, M. Pacholska, T. Berghammer, S. Bodenstein, D. Silver, O. Vinyals, A. W. Senior, K. Kavukcuoglu, P. Kohli, and D. Hassabis.
|
| 1779 |
+
Highly accurate protein structure prediction with AlphaFold.
|
| 1780 |
+
Nature
|
| 1781 |
+
, 596(7873):583–589, 2021.
|
| 1782 |
+
doi:
|
| 1783 |
+
10.1038/s41586-021-03819-2
|
| 1784 |
+
.
|
| 1785 |
+
Kahneman (2011)
|
| 1786 |
+
Daniel Kahneman.
|
| 1787 |
+
Thinking, fast and slow
|
| 1788 |
+
.
|
| 1789 |
+
Farrar, Straus and Giroux, New York, NY, US, 2011.
|
| 1790 |
+
ISBN 978-0-374-27563-1.
|
| 1791 |
+
Khan et al. (2024)
|
| 1792 |
+
Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, and Ethan Perez.
|
| 1793 |
+
Debating with more persuasive llms leads to more truthful answers, 2024.
|
| 1794 |
+
URL
|
| 1795 |
+
https://arxiv.org/abs/2402.06782
|
| 1796 |
+
.
|
| 1797 |
+
Kocsis & Szepesvári (2006)
|
| 1798 |
+
Levente Kocsis and Csaba Szepesvári.
|
| 1799 |
+
Bandit based monte-carlo planning.
|
| 1800 |
+
In Johannes Fürnkranz, Tobias Scheffer, and Myra Spiliopoulou (eds.),
|
| 1801 |
+
Machine Learning: ECML 2006
|
| 1802 |
+
, pp. 282–293, Berlin, Heidelberg, 2006. Springer Berlin Heidelberg.
|
| 1803 |
+
ISBN 978-3-540-46056-5.
|
| 1804 |
+
Koh et al. (2024)
|
| 1805 |
+
Jing Yu Koh, Stephen McAleer, Daniel Fried, and Ruslan Salakhutdinov.
|
| 1806 |
+
Tree search for language model agents, 2024.
|
| 1807 |
+
URL
|
| 1808 |
+
https://arxiv.org/abs/2407.01476
|
| 1809 |
+
.
|
| 1810 |
+
Li et al. (2022)
|
| 1811 |
+
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals.
|
| 1812 |
+
Competition-level code generation with alphacode.
|
| 1813 |
+
Science
|
| 1814 |
+
, 378(6624):1092–1097, 2022.
|
| 1815 |
+
doi:
|
| 1816 |
+
10.1126/science.abq1158
|
| 1817 |
+
.
|
| 1818 |
+
URL
|
| 1819 |
+
https://www.science.org/doi/abs/10.1126/science.abq1158
|
| 1820 |
+
.
|
| 1821 |
+
Lu et al. (2024a)
|
| 1822 |
+
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha.
|
| 1823 |
+
The ai scientist: Towards fully automated open-ended scientific discovery, 2024a.
|
| 1824 |
+
URL
|
| 1825 |
+
https://arxiv.org/abs/2408.06292
|
| 1826 |
+
.
|
| 1827 |
+
Lu et al. (2024b)
|
| 1828 |
+
Cong Lu, Shengran Hu, and Jeff Clune.
|
| 1829 |
+
Intelligent go-explore: Standing on the shoulders of giant foundation models, 2024b.
|
| 1830 |
+
URL
|
| 1831 |
+
https://arxiv.org/abs/2405.15143
|
| 1832 |
+
.
|
| 1833 |
+
Ma et al. (2024a)
|
| 1834 |
+
Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar.
|
| 1835 |
+
Eureka: Human-level reward design via coding large language models, 2024a.
|
| 1836 |
+
URL
|
| 1837 |
+
https://arxiv.org/abs/2310.12931
|
| 1838 |
+
.
|
| 1839 |
+
Ma et al. (2024b)
|
| 1840 |
+
Yingwei Ma, Qingping Yang, Rongyu Cao, Binhua Li, Fei Huang, and Yongbin Li.
|
| 1841 |
+
How to understand whole software repository?, 2024b.
|
| 1842 |
+
URL
|
| 1843 |
+
https://arxiv.org/abs/2406.01422
|
| 1844 |
+
.
|
| 1845 |
+
Moore (1959)
|
| 1846 |
+
E.F. Moore.
|
| 1847 |
+
The Shortest Path Through a Maze
|
| 1848 |
+
.
|
| 1849 |
+
Bell Telephone System. Technical publications. monograph. Bell Telephone System., 1959.
|
| 1850 |
+
URL
|
| 1851 |
+
https://books.google.com/books?id=IVZBHAAACAAJ
|
| 1852 |
+
.
|
| 1853 |
+
OpenAI (2024)
|
| 1854 |
+
OpenAI.
|
| 1855 |
+
OpenAI o1 System Card, September 2024.
|
| 1856 |
+
URL
|
| 1857 |
+
https://openai.com/research/o1-system-card
|
| 1858 |
+
.
|
| 1859 |
+
Online report.
|
| 1860 |
+
OpenAI et al. (2024)
|
| 1861 |
+
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo
|
| 1862 |
+
Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan
|
| 1863 |
+
Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez,
|
| 1864 |
+
Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia
|
| 1865 |
+
Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph.
|
| 1866 |
+
Gpt-4 technical report, 2024.
|
| 1867 |
+
URL
|
| 1868 |
+
https://arxiv.org/abs/2303.08774
|
| 1869 |
+
.
|
| 1870 |
+
Örwall (2024)
|
| 1871 |
+
Albert Örwall.
|
| 1872 |
+
Moatless tools, jun 2024.
|
| 1873 |
+
URL
|
| 1874 |
+
https://github.com/aorwall/moatless-tools
|
| 1875 |
+
.
|
| 1876 |
+
Accessed: 2024-07-16.
|
| 1877 |
+
Ouyang et al. (2022)
|
| 1878 |
+
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe.
|
| 1879 |
+
Training language models to follow instructions with human feedback, 2022.
|
| 1880 |
+
URL
|
| 1881 |
+
https://arxiv.org/abs/2203.02155
|
| 1882 |
+
.
|
| 1883 |
+
Pan et al. (2023)
|
| 1884 |
+
Liangming Pan, Alon Albalak, Xinyi Wang, and William Yang Wang.
|
| 1885 |
+
Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning, 2023.
|
| 1886 |
+
URL
|
| 1887 |
+
https://arxiv.org/abs/2305.12295
|
| 1888 |
+
.
|
| 1889 |
+
Rigaki et al. (2024)
|
| 1890 |
+
Maria Rigaki, Carlos Catania, and Sebastian Garcia.
|
| 1891 |
+
Hackphyr: A local fine-tuned llm agent for network security environments, 2024.
|
| 1892 |
+
URL
|
| 1893 |
+
https://arxiv.org/abs/2409.11276
|
| 1894 |
+
.
|
| 1895 |
+
Saha et al. (2024)
|
| 1896 |
+
Swarnadeep Saha, Archiki Prasad, Justin Chih-Yao Chen, Peter Hase, Elias Stengel-Eskin, and Mohit Bansal.
|
| 1897 |
+
System-1.x: Learning to balance fast and slow planning with language models, 2024.
|
| 1898 |
+
URL
|
| 1899 |
+
https://arxiv.org/abs/2407.14414
|
| 1900 |
+
.
|
| 1901 |
+
Silver et al. (2016a)
|
| 1902 |
+
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis.
|
| 1903 |
+
Mastering the game of go with deep neural networks and tree search.
|
| 1904 |
+
Nature
|
| 1905 |
+
, 529(7587):484–489, 1 2016a.
|
| 1906 |
+
ISSN 1476-4687.
|
| 1907 |
+
doi:
|
| 1908 |
+
10.1038/nature16961
|
| 1909 |
+
.
|
| 1910 |
+
URL
|
| 1911 |
+
https://doi.org/10.1038/nature16961
|
| 1912 |
+
.
|
| 1913 |
+
Silver et al. (2016b)
|
| 1914 |
+
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis.
|
| 1915 |
+
Mastering the game of go with deep neural networks and tree search.
|
| 1916 |
+
Nat.
|
| 1917 |
+
, 529(7587):484–489, 2016b.
|
| 1918 |
+
doi:
|
| 1919 |
+
10.1038/NATURE16961
|
| 1920 |
+
.
|
| 1921 |
+
URL
|
| 1922 |
+
https://doi.org/10.1038/nature16961
|
| 1923 |
+
.
|
| 1924 |
+
Silver et al. (2018)
|
| 1925 |
+
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis.
|
| 1926 |
+
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.
|
| 1927 |
+
Science
|
| 1928 |
+
, 362(6419):1140–1144, 2018.
|
| 1929 |
+
doi:
|
| 1930 |
+
10.1126/science.aar6404
|
| 1931 |
+
.
|
| 1932 |
+
URL
|
| 1933 |
+
https://www.science.org/doi/abs/10.1126/science.aar6404
|
| 1934 |
+
.
|
| 1935 |
+
Snell et al. (2024)
|
| 1936 |
+
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar.
|
| 1937 |
+
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024.
|
| 1938 |
+
URL
|
| 1939 |
+
https://arxiv.org/abs/2408.03314
|
| 1940 |
+
.
|
| 1941 |
+
Wang et al. (2023)
|
| 1942 |
+
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar.
|
| 1943 |
+
Voyager: An open-ended embodied agent with large language models, 2023.
|
| 1944 |
+
URL
|
| 1945 |
+
https://arxiv.org/abs/2305.16291
|
| 1946 |
+
.
|
| 1947 |
+
Wang et al. (2024a)
|
| 1948 |
+
Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji.
|
| 1949 |
+
Executable code actions elicit better llm agents, 2024a.
|
| 1950 |
+
URL
|
| 1951 |
+
https://arxiv.org/abs/2402.01030
|
| 1952 |
+
.
|
| 1953 |
+
Wang et al. (2024b)
|
| 1954 |
+
Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, Robert Brennan, Hao Peng, Heng Ji, and Graham Neubig.
|
| 1955 |
+
Opendevin: An open platform for ai software developers as generalist agents, 2024b.
|
| 1956 |
+
URL
|
| 1957 |
+
https://arxiv.org/abs/2407.16741
|
| 1958 |
+
.
|
| 1959 |
+
Wei et al. (2022)
|
| 1960 |
+
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al.
|
| 1961 |
+
Emergent abilities of large language models.
|
| 1962 |
+
arXiv preprint arXiv:2206.07682
|
| 1963 |
+
, 2022.
|
| 1964 |
+
Xia et al. (2024)
|
| 1965 |
+
Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang.
|
| 1966 |
+
Agentless: Demystifying llm-based software engineering agents, 2024.
|
| 1967 |
+
URL
|
| 1968 |
+
https://arxiv.org/abs/2407.01489
|
| 1969 |
+
.
|
| 1970 |
+
Yang et al. (2024a)
|
| 1971 |
+
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Xuejing Liu, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, Zhifang Guo, and Zhihao Fan.
|
| 1972 |
+
Qwen2 technical report, 2024a.
|
| 1973 |
+
URL
|
| 1974 |
+
https://arxiv.org/abs/2407.10671
|
| 1975 |
+
.
|
| 1976 |
+
Yang et al. (2024b)
|
| 1977 |
+
John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press.
|
| 1978 |
+
Swe-agent: Agent-computer interfaces enable automated software engineering, 2024b.
|
| 1979 |
+
URL
|
| 1980 |
+
https://arxiv.org/abs/2405.15793
|
| 1981 |
+
.
|
| 1982 |
+
Yao et al. (2023)
|
| 1983 |
+
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan.
|
| 1984 |
+
Tree of thoughts: Deliberate problem solving with large language models.
|
| 1985 |
+
In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.),
|
| 1986 |
+
Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023
|
| 1987 |
+
, 2023.
|
| 1988 |
+
URL
|
| 1989 |
+
http://papers.nips.cc/paper_files/paper/2023/hash/271db9922b8d1f4dd7aaef84ed5ac703-Abstract-Conference.html
|
| 1990 |
+
.
|
| 1991 |
+
Zhang et al. (2024a)
|
| 1992 |
+
Kexun Zhang, Weiran Yao, Zuxin Liu, Yihao Feng, Zhiwei Liu, Rithesh Murthy, Tian Lan, Lei Li, Renze Lou, Jiacheng Xu, Bo Pang, Yingbo Zhou, Shelby Heinecke, Silvio Savarese, Huan Wang, and Caiming Xiong.
|
| 1993 |
+
Diversity empowers intelligence: Integrating expertise of software engineering agents, 2024a.
|
| 1994 |
+
URL
|
| 1995 |
+
https://arxiv.org/abs/2408.07060
|
| 1996 |
+
.
|
| 1997 |
+
Zhang et al. (2024b)
|
| 1998 |
+
Yao Zhang, Zijian Ma, Yunpu Ma, Zhen Han, Yu Wu, and Volker Tresp.
|
| 1999 |
+
Webpilot: A versatile and autonomous multi-agent system for web task execution with strategic exploration, 2024b.
|
| 2000 |
+
URL
|
| 2001 |
+
https://arxiv.org/abs/2408.15978
|
| 2002 |
+
.
|
| 2003 |
+
Zhang et al. (2024c)
|
| 2004 |
+
Yiqun Zhang, Xiaocui Yang, Shi Feng, Daling Wang, Yifei Zhang, and Kaisong Song.
|
| 2005 |
+
Can llms beat humans in debating? a dynamic multi-agent framework for competitive debate, 2024c.
|
| 2006 |
+
URL
|
| 2007 |
+
https://arxiv.org/abs/2408.04472
|
| 2008 |
+
.
|
| 2009 |
+
Zhang et al. (2024d)
|
| 2010 |
+
Yuntong Zhang, Haifeng Ruan, Zhiyu Fan, and Abhik Roychoudhury.
|
| 2011 |
+
Autocoderover: Autonomous program improvement, 2024d.
|
| 2012 |
+
URL
|
| 2013 |
+
https://arxiv.org/abs/2404.05427
|
| 2014 |
+
.
|
| 2015 |
+
Appendix A
|
| 2016 |
+
Reproducibility
|
| 2017 |
+
All models and data used in our work are publicly available. We additionally provide hyperparameter details in
|
| 2018 |
+
Appendix
|
| 2019 |
+
2
|
| 2020 |
+
. The code will be released as a public repository upon publication.
|
| 2021 |
+
Appendix B
|
| 2022 |
+
Additional Implementation Details
|
| 2023 |
+
Moatless-adapted is an extended version of the moatless-tools library with support for a tree structure, the ability to revert to earlier versions of the codebase, and the capability to run tests.
|
| 2024 |
+
The standard implementation of moatless-tools is based on a finite state machine structure where a state holds information about file context and properties set in the configuration or from previous states. It can then transition to a new state when an action is executed. The request that initiates the action is created by an LLM. This follows a linear structure where one state can transition to another state. In moatless-adapted, this model is extended so that a state can expand by using actions to create more states. The connections between states are then represented in a tree structure with nodes.
|
| 2025 |
+
Each state has a file context associated with it. This file context will be included in the prompt sent to an LLM. To limit the size of the prompt, files are divided into ”spans,” where a span could be, for example, a section of code (e.g., imports), a class, or a function. These are identified by span IDs. Thus, the LLM sees a limited part of the code at a time but can request more context by searching for or adding files and spans. The file context therefore changes over time, and a specific state of file context is linked to a specific state.
|
| 2026 |
+
In the standard implementation of moatless-tools, changes to the codebase are made linearly, and each change is saved directly to the file system. In moatless-adapted, however, there is a need to be able to revert to earlier states and thus return to a previous version of the codebase. To handle this, the code is stored in a git repository where each change is committed, and each state has a reference to a commit as well as the current patch of the diff from the initial commit that existed before starting. This way, one can go back to an earlier state by specifying the state ID, and the commit that was current at that time will be checked out.
|
| 2027 |
+
The test files present in the file context are run each time the Plan state is initiated, and the test results are provided to the state. The tests are then run in Docker images built via the SWE-bench library. To use this approach in a benchmark where a larger number of instances should be able to run simultaneously, a solution is used where these images are run as pods in a Kubernetes cluster. Moatless-tools communicates with the testbed by applying patches and running commands via an API. When a new instance starts, a pod is created which is then reset at each run, applying the current patch and running tests according to the test command specified in the SWE-bench library. It’s important to add here that the agent is not aware of the
|
| 2028 |
+
PASS_TO_PASS
|
| 2029 |
+
or
|
| 2030 |
+
FAIL_TO_PASS
|
| 2031 |
+
tests in the SWE-bench harness, but only knows how to run the tests. This corresponds to a real engineering environment where each project can have its own test commands.
|
| 2032 |
+
Appendix C
|
| 2033 |
+
MCTS Hyperparameters
|
| 2034 |
+
The Monte Carlo Tree Search (MCTS) algorithm used in this study employs several hyperparameters.
|
| 2035 |
+
Table 2:
|
| 2036 |
+
MCTS Hyperparameters
|
| 2037 |
+
Hyperparameter
|
| 2038 |
+
Description
|
| 2039 |
+
Default
|
| 2040 |
+
c_param
|
| 2041 |
+
UCT exploration parameter
|
| 2042 |
+
1.41
|
| 2043 |
+
max_expansions
|
| 2044 |
+
Max children per node
|
| 2045 |
+
5
|
| 2046 |
+
max_iterations
|
| 2047 |
+
Max MCTS iterations
|
| 2048 |
+
100
|
| 2049 |
+
provide_feedback
|
| 2050 |
+
Enable feedback
|
| 2051 |
+
True
|
| 2052 |
+
best_first
|
| 2053 |
+
Use best-first strategy
|
| 2054 |
+
True
|
| 2055 |
+
value_function_temperature
|
| 2056 |
+
Value function temperature
|
| 2057 |
+
0.2
|
| 2058 |
+
max_depth
|
| 2059 |
+
Max tree depth
|
| 2060 |
+
20
|
| 2061 |
+
UCT Score Calculation Parameters
|
| 2062 |
+
exploration_weight
|
| 2063 |
+
UCT exploration weight
|
| 2064 |
+
1.0
|
| 2065 |
+
depth_weight
|
| 2066 |
+
Depth penalty weight
|
| 2067 |
+
0.8
|
| 2068 |
+
depth_bonus_factor
|
| 2069 |
+
Depth bonus factor
|
| 2070 |
+
200.0
|
| 2071 |
+
high_value_threshold
|
| 2072 |
+
High-value node threshold
|
| 2073 |
+
55.0
|
| 2074 |
+
low_value_threshold
|
| 2075 |
+
Low-value node threshold
|
| 2076 |
+
50.0
|
| 2077 |
+
very_high_value_threshold
|
| 2078 |
+
Very high-value threshold
|
| 2079 |
+
75.0
|
| 2080 |
+
high_value_leaf_bonus_constant
|
| 2081 |
+
High-value leaf bonus
|
| 2082 |
+
20.0
|
| 2083 |
+
high_value_bad_children_bonus_constant
|
| 2084 |
+
High-value bad children bonus
|
| 2085 |
+
20.0
|
| 2086 |
+
high_value_child_penalty_constant
|
| 2087 |
+
High-value child penalty
|
| 2088 |
+
5.0
|
| 2089 |
+
Action Model Parameters
|
| 2090 |
+
action_model_temperature
|
| 2091 |
+
Action model temperature
|
| 2092 |
+
0.2
|
| 2093 |
+
Discriminator Parameters
|
| 2094 |
+
number_of_agents
|
| 2095 |
+
Number of Discriminator Agents
|
| 2096 |
+
5
|
| 2097 |
+
number_of_round
|
| 2098 |
+
Number of debate rounds
|
| 2099 |
+
3
|
| 2100 |
+
discriminator_temperature
|
| 2101 |
+
Discriminator temperature
|
| 2102 |
+
1.0
|
| 2103 |
+
These hyperparameters can be adjusted to fine-tune the MCTS algorithm’s performance for specific problem domains or computational constraints. The values listed here are the defaults as defined in the
|
| 2104 |
+
TreeSearchSettings
|
| 2105 |
+
class and the MCTS implementation.
|
| 2106 |
+
Appendix D
|
| 2107 |
+
Ability of MCTS to Escape Unproductive Loops vs. Baseline
|
| 2108 |
+
Figure 6:
|
| 2109 |
+
Avoiding Repetitive Actions, django__django__10914.
|
| 2110 |
+
We found that the base agent can often get stuck performing repetitive actions
|
| 2111 |
+
that do not bring it closer to solving the issue, and which commonly lead to unresolvable dead-ends. In this example, the base agent
|
| 2112 |
+
was stuck implementing wrong tests which continuously returned errors. In contrast, when this happens in
|
| 2113 |
+
SWE-Search, the Value Agent recognizes this, terminating these trajectories quickly,
|
| 2114 |
+
as happens in Node 73 (orange).
|
| 2115 |
+
Appendix E
|
| 2116 |
+
Model Instance Resolution Uniqueness
|
| 2117 |
+
To understand the complementary strengths of different models in resolving software issues, we analyzed how unique their resolved issue subsets where. Figure
|
| 2118 |
+
7
|
| 2119 |
+
illustrates the resolution patterns for each model across five of the codebases in SWE-bench-lite.
|
| 2120 |
+
Figure 7:
|
| 2121 |
+
Unique Issue Resolution Patterns Across Models and Libraries.
|
| 2122 |
+
Each column represents a different Python reposiroty, and each row within a column represents a specific issue. Colored blocks indicate successful resolution by the corresponding model (see legend). White spaces denote unresolved issues. This visualization highlights the diverse problem-solving capabilities of different models across various software domains, demonstrating that no single model dominates across all issues and libraries.
|
| 2123 |
+
Appendix F
|
| 2124 |
+
Ability of Value Function to Discern Successful Trajectories
|
| 2125 |
+
Before implementing SWE-Search, we conducted a general study across many models to evaluate the models’ ability to differentiate states which led to resolved vs. unresolved issues. Figure
|
| 2126 |
+
8
|
| 2127 |
+
shows the results of this study. We found that in general, models assigned higher rewards to states which eventually led to resolved issues. Of particular interest was the Deepseek model, which seemed to identify critical errors in trajectories effectively. This was also observed in the final agent (see Fig.
|
| 2128 |
+
5
|
| 2129 |
+
a).
|
| 2130 |
+
Figure 8:
|
| 2131 |
+
Average State Reward Comparison Across Models.
|
| 2132 |
+
This graph compares the average state rewards assigned by different language models for resolved (green) and unresolved (red) issues. Error bars indicate standard deviation. Most models consistently assign higher rewards to states leading to resolved issues, with the exception of the. The ’Average’ column represents the mean across all models, demonstrating a clear distinction between resolved and unresolved states.
|
| 2133 |
+
Appendix G
|
| 2134 |
+
Value Function Prompts
|
| 2135 |
+
◄
|
| 2136 |
+
Feeling
|
| 2137 |
+
lucky?
|
| 2138 |
+
Conversion
|
| 2139 |
+
report
|
| 2140 |
+
Report
|
| 2141 |
+
an issue
|
| 2142 |
+
View original
|
| 2143 |
+
on arXiv
|
| 2144 |
+
►
|
|
@@ -0,0 +1,200 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2412.21139] Training Software Engineering Agents and Verifiers with SWE-Gym'
|
| 3 |
+
id: 241221139-training-software-engineering-agents-and-verifiers-with-swe-gym
|
| 4 |
+
tags:
|
| 5 |
+
- deepread
|
| 6 |
+
created: '2026-06-10T00:22:57.639430Z'
|
| 7 |
+
source: https://arxiv.org/abs/2412.21139
|
| 8 |
+
source_domain: arxiv.org
|
| 9 |
+
fetched_at: '2026-06-10T00:22:57.639309Z'
|
| 10 |
+
fetch_provider: builtin
|
| 11 |
+
status: draft
|
| 12 |
+
type: note
|
| 13 |
+
tier: institutional
|
| 14 |
+
content_type: paper
|
| 15 |
+
deprecated: false
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
[2412.21139] Training Software Engineering Agents and Verifiers with SWE-Gym
|
| 19 |
+
Computer Science > Software Engineering
|
| 20 |
+
arXiv:2412.21139
|
| 21 |
+
(cs)
|
| 22 |
+
[Submitted on 30 Dec 2024 (
|
| 23 |
+
v1
|
| 24 |
+
), last revised 6 Jun 2025 (this version, v2)]
|
| 25 |
+
Title:
|
| 26 |
+
Training Software Engineering Agents and Verifiers with SWE-Gym
|
| 27 |
+
Authors:
|
| 28 |
+
Jiayi Pan
|
| 29 |
+
,
|
| 30 |
+
Xingyao Wang
|
| 31 |
+
,
|
| 32 |
+
Graham Neubig
|
| 33 |
+
,
|
| 34 |
+
Navdeep Jaitly
|
| 35 |
+
,
|
| 36 |
+
Heng Ji
|
| 37 |
+
,
|
| 38 |
+
Alane Suhr
|
| 39 |
+
,
|
| 40 |
+
Yizhe Zhang
|
| 41 |
+
View a PDF of the paper titled Training Software Engineering Agents and Verifiers with SWE-Gym, by Jiayi Pan and 6 other authors
|
| 42 |
+
View PDF
|
| 43 |
+
HTML (experimental)
|
| 44 |
+
Abstract:
|
| 45 |
+
We present SWE-Gym, the first environment for training real-world software engineering (SWE) agents. SWE-Gym contains 2,438 real-world Python task instances, each comprising a codebase with an executable runtime environment, unit tests, and a task specified in natural language. We use SWE-Gym to train language model based SWE agents, achieving up to 19% absolute gains in resolve rate on the popular SWE-Bench Verified and Lite test sets. We also experiment with inference-time scaling through verifiers trained on agent trajectories sampled from SWE-Gym. When combined with our fine-tuned SWE agents, we achieve 32.0% and 26.0% on SWE-Bench Verified and Lite, respectively, reflecting a new state-of-the-art for open-weight SWE agents. To facilitate further research, we publicly release SWE-Gym, models, and agent trajectories.
|
| 46 |
+
Comments:
|
| 47 |
+
Accepted at ICML 2025. Code at
|
| 48 |
+
this https URL
|
| 49 |
+
Subjects:
|
| 50 |
+
Software Engineering (cs.SE)
|
| 51 |
+
; Computation and Language (cs.CL)
|
| 52 |
+
Cite as:
|
| 53 |
+
arXiv:2412.21139
|
| 54 |
+
[cs.SE]
|
| 55 |
+
(or
|
| 56 |
+
arXiv:2412.21139v2
|
| 57 |
+
[cs.SE]
|
| 58 |
+
for this version)
|
| 59 |
+
https://doi.org/10.48550/arXiv.2412.21139
|
| 60 |
+
Focus to learn more
|
| 61 |
+
arXiv-issued DOI via DataCite
|
| 62 |
+
Submission history
|
| 63 |
+
From: Jiayi Pan [
|
| 64 |
+
view email
|
| 65 |
+
]
|
| 66 |
+
[v1]
|
| 67 |
+
Mon, 30 Dec 2024 18:15:39 UTC (156 KB)
|
| 68 |
+
[v2]
|
| 69 |
+
Fri, 6 Jun 2025 07:53:20 UTC (295 KB)
|
| 70 |
+
Full-text links:
|
| 71 |
+
Access Paper:
|
| 72 |
+
View a PDF of the paper titled Training Software Engineering Agents and Verifiers with SWE-Gym, by Jiayi Pan and 6 other authors
|
| 73 |
+
View PDF
|
| 74 |
+
HTML (experimental)
|
| 75 |
+
TeX Source
|
| 76 |
+
view license
|
| 77 |
+
Current browse context:
|
| 78 |
+
cs.SE
|
| 79 |
+
< prev
|
| 80 |
+
|
|
| 81 |
+
next >
|
| 82 |
+
new
|
| 83 |
+
|
|
| 84 |
+
recent
|
| 85 |
+
|
|
| 86 |
+
2024-12
|
| 87 |
+
Change to browse by:
|
| 88 |
+
cs
|
| 89 |
+
cs.CL
|
| 90 |
+
References & Citations
|
| 91 |
+
NASA ADS
|
| 92 |
+
Google Scholar
|
| 93 |
+
Semantic Scholar
|
| 94 |
+
export BibTeX citation
|
| 95 |
+
Loading...
|
| 96 |
+
BibTeX formatted citation
|
| 97 |
+
×
|
| 98 |
+
loading...
|
| 99 |
+
Data provided by:
|
| 100 |
+
Bookmark
|
| 101 |
+
Bibliographic Tools
|
| 102 |
+
Bibliographic and Citation Tools
|
| 103 |
+
Bibliographic Explorer Toggle
|
| 104 |
+
Bibliographic Explorer
|
| 105 |
+
(
|
| 106 |
+
What is the Explorer?
|
| 107 |
+
)
|
| 108 |
+
Connected Papers Toggle
|
| 109 |
+
Connected Papers
|
| 110 |
+
(
|
| 111 |
+
What is Connected Papers?
|
| 112 |
+
)
|
| 113 |
+
Litmaps Toggle
|
| 114 |
+
Litmaps
|
| 115 |
+
(
|
| 116 |
+
What is Litmaps?
|
| 117 |
+
)
|
| 118 |
+
scite.ai Toggle
|
| 119 |
+
scite Smart Citations
|
| 120 |
+
(
|
| 121 |
+
What are Smart Citations?
|
| 122 |
+
)
|
| 123 |
+
Code, Data, Media
|
| 124 |
+
Code, Data and Media Associated with this Article
|
| 125 |
+
alphaXiv Toggle
|
| 126 |
+
alphaXiv
|
| 127 |
+
(
|
| 128 |
+
What is alphaXiv?
|
| 129 |
+
)
|
| 130 |
+
Links to Code Toggle
|
| 131 |
+
CatalyzeX Code Finder for Papers
|
| 132 |
+
(
|
| 133 |
+
What is CatalyzeX?
|
| 134 |
+
)
|
| 135 |
+
DagsHub Toggle
|
| 136 |
+
DagsHub
|
| 137 |
+
(
|
| 138 |
+
What is DagsHub?
|
| 139 |
+
)
|
| 140 |
+
GotitPub Toggle
|
| 141 |
+
Gotit.pub
|
| 142 |
+
(
|
| 143 |
+
What is GotitPub?
|
| 144 |
+
)
|
| 145 |
+
Huggingface Toggle
|
| 146 |
+
Hugging Face
|
| 147 |
+
(
|
| 148 |
+
What is Huggingface?
|
| 149 |
+
)
|
| 150 |
+
ScienceCast Toggle
|
| 151 |
+
ScienceCast
|
| 152 |
+
(
|
| 153 |
+
What is ScienceCast?
|
| 154 |
+
)
|
| 155 |
+
Demos
|
| 156 |
+
Demos
|
| 157 |
+
Replicate Toggle
|
| 158 |
+
Replicate
|
| 159 |
+
(
|
| 160 |
+
What is Replicate?
|
| 161 |
+
)
|
| 162 |
+
Spaces Toggle
|
| 163 |
+
Hugging Face Spaces
|
| 164 |
+
(
|
| 165 |
+
What is Spaces?
|
| 166 |
+
)
|
| 167 |
+
Spaces Toggle
|
| 168 |
+
TXYZ.AI
|
| 169 |
+
(
|
| 170 |
+
What is TXYZ.AI?
|
| 171 |
+
)
|
| 172 |
+
Related Papers
|
| 173 |
+
Recommenders and Search Tools
|
| 174 |
+
Link to Influence Flower
|
| 175 |
+
Influence Flower
|
| 176 |
+
(
|
| 177 |
+
What are Influence Flowers?
|
| 178 |
+
)
|
| 179 |
+
Core recommender toggle
|
| 180 |
+
CORE Recommender
|
| 181 |
+
(
|
| 182 |
+
What is CORE?
|
| 183 |
+
)
|
| 184 |
+
Author
|
| 185 |
+
Venue
|
| 186 |
+
Institution
|
| 187 |
+
Topic
|
| 188 |
+
About arXivLabs
|
| 189 |
+
arXivLabs: experimental projects with community collaborators
|
| 190 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 191 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 192 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 193 |
+
Learn more about arXivLabs
|
| 194 |
+
.
|
| 195 |
+
Which authors of this paper are endorsers?
|
| 196 |
+
|
|
| 197 |
+
Disable MathJax
|
| 198 |
+
(
|
| 199 |
+
What is MathJax?
|
| 200 |
+
)
|
|
@@ -0,0 +1,196 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2501.04519] rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved
|
| 3 |
+
Deep Thinking'
|
| 4 |
+
id: 250104519-rstar-math-small-llms-can-master-math-reasoning-with-self-evolved-deep
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:40:01.597011Z'
|
| 8 |
+
source: https://arxiv.org/abs/2501.04519
|
| 9 |
+
source_domain: arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:40:01.596873Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2501.04519] rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
|
| 20 |
+
Computer Science > Computation and Language
|
| 21 |
+
arXiv:2501.04519
|
| 22 |
+
(cs)
|
| 23 |
+
[Submitted on 8 Jan 2025]
|
| 24 |
+
Title:
|
| 25 |
+
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
|
| 26 |
+
Authors:
|
| 27 |
+
Xinyu Guan
|
| 28 |
+
,
|
| 29 |
+
Li Lyna Zhang
|
| 30 |
+
,
|
| 31 |
+
Yifei Liu
|
| 32 |
+
,
|
| 33 |
+
Ning Shang
|
| 34 |
+
,
|
| 35 |
+
Youran Sun
|
| 36 |
+
,
|
| 37 |
+
Yi Zhu
|
| 38 |
+
,
|
| 39 |
+
Fan Yang
|
| 40 |
+
,
|
| 41 |
+
Mao Yang
|
| 42 |
+
View a PDF of the paper titled rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking, by Xinyu Guan and 7 other authors
|
| 43 |
+
View PDF
|
| 44 |
+
HTML (experimental)
|
| 45 |
+
Abstract:
|
| 46 |
+
We present rStar-Math to demonstrate that small language models (SLMs) can rival or even surpass the math reasoning capability of OpenAI o1, without distillation from superior models. rStar-Math achieves this by exercising "deep thinking" through Monte Carlo Tree Search (MCTS), where a math policy SLM performs test-time search guided by an SLM-based process reward model. rStar-Math introduces three innovations to tackle the challenges in training the two SLMs: (1) a novel code-augmented CoT data sythesis method, which performs extensive MCTS rollouts to generate step-by-step verified reasoning trajectories used to train the policy SLM; (2) a novel process reward model training method that avoids naïve step-level score annotation, yielding a more effective process preference model (PPM); (3) a self-evolution recipe in which the policy SLM and PPM are built from scratch and iteratively evolved to improve reasoning capabilities. Through 4 rounds of self-evolution with millions of synthesized solutions for 747k math problems, rStar-Math boosts SLMs' math reasoning to state-of-the-art levels. On the MATH benchmark, it improves Qwen2.5-Math-7B from 58.8% to 90.0% and Phi3-mini-3.8B from 41.4% to 86.4%, surpassing o1-preview by +4.5% and +0.9%. On the USA Math Olympiad (AIME), rStar-Math solves an average of 53.3% (8/15) of problems, ranking among the top 20% the brightest high school math students. Code and data will be available at
|
| 47 |
+
this https URL
|
| 48 |
+
.
|
| 49 |
+
Subjects:
|
| 50 |
+
Computation and Language (cs.CL)
|
| 51 |
+
Cite as:
|
| 52 |
+
arXiv:2501.04519
|
| 53 |
+
[cs.CL]
|
| 54 |
+
(or
|
| 55 |
+
arXiv:2501.04519v1
|
| 56 |
+
[cs.CL]
|
| 57 |
+
for this version)
|
| 58 |
+
https://doi.org/10.48550/arXiv.2501.04519
|
| 59 |
+
Focus to learn more
|
| 60 |
+
arXiv-issued DOI via DataCite
|
| 61 |
+
Submission history
|
| 62 |
+
From: Li Lyna Zhang [
|
| 63 |
+
view email
|
| 64 |
+
]
|
| 65 |
+
[v1]
|
| 66 |
+
Wed, 8 Jan 2025 14:12:57 UTC (632 KB)
|
| 67 |
+
Full-text links:
|
| 68 |
+
Access Paper:
|
| 69 |
+
View a PDF of the paper titled rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking, by Xinyu Guan and 7 other authors
|
| 70 |
+
View PDF
|
| 71 |
+
HTML (experimental)
|
| 72 |
+
TeX Source
|
| 73 |
+
view license
|
| 74 |
+
Current browse context:
|
| 75 |
+
cs.CL
|
| 76 |
+
< prev
|
| 77 |
+
|
|
| 78 |
+
next >
|
| 79 |
+
new
|
| 80 |
+
|
|
| 81 |
+
recent
|
| 82 |
+
|
|
| 83 |
+
2025-01
|
| 84 |
+
Change to browse by:
|
| 85 |
+
cs
|
| 86 |
+
References & Citations
|
| 87 |
+
NASA ADS
|
| 88 |
+
Google Scholar
|
| 89 |
+
Semantic Scholar
|
| 90 |
+
export BibTeX citation
|
| 91 |
+
Loading...
|
| 92 |
+
BibTeX formatted citation
|
| 93 |
+
×
|
| 94 |
+
loading...
|
| 95 |
+
Data provided by:
|
| 96 |
+
Bookmark
|
| 97 |
+
Bibliographic Tools
|
| 98 |
+
Bibliographic and Citation Tools
|
| 99 |
+
Bibliographic Explorer Toggle
|
| 100 |
+
Bibliographic Explorer
|
| 101 |
+
(
|
| 102 |
+
What is the Explorer?
|
| 103 |
+
)
|
| 104 |
+
Connected Papers Toggle
|
| 105 |
+
Connected Papers
|
| 106 |
+
(
|
| 107 |
+
What is Connected Papers?
|
| 108 |
+
)
|
| 109 |
+
Litmaps Toggle
|
| 110 |
+
Litmaps
|
| 111 |
+
(
|
| 112 |
+
What is Litmaps?
|
| 113 |
+
)
|
| 114 |
+
scite.ai Toggle
|
| 115 |
+
scite Smart Citations
|
| 116 |
+
(
|
| 117 |
+
What are Smart Citations?
|
| 118 |
+
)
|
| 119 |
+
Code, Data, Media
|
| 120 |
+
Code, Data and Media Associated with this Article
|
| 121 |
+
alphaXiv Toggle
|
| 122 |
+
alphaXiv
|
| 123 |
+
(
|
| 124 |
+
What is alphaXiv?
|
| 125 |
+
)
|
| 126 |
+
Links to Code Toggle
|
| 127 |
+
CatalyzeX Code Finder for Papers
|
| 128 |
+
(
|
| 129 |
+
What is CatalyzeX?
|
| 130 |
+
)
|
| 131 |
+
DagsHub Toggle
|
| 132 |
+
DagsHub
|
| 133 |
+
(
|
| 134 |
+
What is DagsHub?
|
| 135 |
+
)
|
| 136 |
+
GotitPub Toggle
|
| 137 |
+
Gotit.pub
|
| 138 |
+
(
|
| 139 |
+
What is GotitPub?
|
| 140 |
+
)
|
| 141 |
+
Huggingface Toggle
|
| 142 |
+
Hugging Face
|
| 143 |
+
(
|
| 144 |
+
What is Huggingface?
|
| 145 |
+
)
|
| 146 |
+
ScienceCast Toggle
|
| 147 |
+
ScienceCast
|
| 148 |
+
(
|
| 149 |
+
What is ScienceCast?
|
| 150 |
+
)
|
| 151 |
+
Demos
|
| 152 |
+
Demos
|
| 153 |
+
Replicate Toggle
|
| 154 |
+
Replicate
|
| 155 |
+
(
|
| 156 |
+
What is Replicate?
|
| 157 |
+
)
|
| 158 |
+
Spaces Toggle
|
| 159 |
+
Hugging Face Spaces
|
| 160 |
+
(
|
| 161 |
+
What is Spaces?
|
| 162 |
+
)
|
| 163 |
+
Spaces Toggle
|
| 164 |
+
TXYZ.AI
|
| 165 |
+
(
|
| 166 |
+
What is TXYZ.AI?
|
| 167 |
+
)
|
| 168 |
+
Related Papers
|
| 169 |
+
Recommenders and Search Tools
|
| 170 |
+
Link to Influence Flower
|
| 171 |
+
Influence Flower
|
| 172 |
+
(
|
| 173 |
+
What are Influence Flowers?
|
| 174 |
+
)
|
| 175 |
+
Core recommender toggle
|
| 176 |
+
CORE Recommender
|
| 177 |
+
(
|
| 178 |
+
What is CORE?
|
| 179 |
+
)
|
| 180 |
+
Author
|
| 181 |
+
Venue
|
| 182 |
+
Institution
|
| 183 |
+
Topic
|
| 184 |
+
About arXivLabs
|
| 185 |
+
arXivLabs: experimental projects with community collaborators
|
| 186 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 187 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 188 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 189 |
+
Learn more about arXivLabs
|
| 190 |
+
.
|
| 191 |
+
Which authors of this paper are endorsers?
|
| 192 |
+
|
|
| 193 |
+
Disable MathJax
|
| 194 |
+
(
|
| 195 |
+
What is MathJax?
|
| 196 |
+
)
|
|
@@ -0,0 +1,3557 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2501.04519] \sysname: Small LLMs Can Master Math Reasoning with Self-Evolved
|
| 3 |
+
Deep Thinking'
|
| 4 |
+
id: 250104519-sysname-small-llms-can-master-math-reasoning-with-self-evolved-deep-th
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:40:46.873514Z'
|
| 8 |
+
source: https://ar5iv.labs.arxiv.org/html/2501.04519
|
| 9 |
+
source_domain: ar5iv.labs.arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:40:46.873327Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2501.04519] \sysname: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
|
| 20 |
+
\sysname
|
| 21 |
+
: Small LLMs Can Master Math Reasoning
|
| 22 |
+
with Self-Evolved Deep Thinking
|
| 23 |
+
Xinyu Guan
|
| 24 |
+
∗
|
| 25 |
+
Li Lyna Zhang
|
| 26 |
+
∗⋄
|
| 27 |
+
Yifei Liu
|
| 28 |
+
Ning Shang Youran Sun Yi Zhu Fan Yang Mao Yang
|
| 29 |
+
Microsoft Research Asia
|
| 30 |
+
Abstract
|
| 31 |
+
We present
|
| 32 |
+
\sysname
|
| 33 |
+
to demonstrate that small language models (SLMs) can rival or even surpass the math reasoning capability of OpenAI o1, without distillation from superior models.
|
| 34 |
+
\sysname
|
| 35 |
+
achieves this by exercising “deep thinking” through Monte Carlo Tree Search (MCTS), where a math
|
| 36 |
+
policy SLM
|
| 37 |
+
performs test-time search guided by an SLM-based
|
| 38 |
+
process reward model
|
| 39 |
+
.
|
| 40 |
+
\sysname
|
| 41 |
+
introduces three innovations to tackle the challenges in training the two SLMs:
|
| 42 |
+
(1)
|
| 43 |
+
a novel code-augmented CoT data sythesis method, which performs extensive MCTS rollouts to generate
|
| 44 |
+
step-by-step verified reasoning trajectories
|
| 45 |
+
used to train the policy SLM;
|
| 46 |
+
(2)
|
| 47 |
+
a novel process reward model training method that avoids naïve step-level score annotation, yielding a more effective
|
| 48 |
+
process preference model (PPM)
|
| 49 |
+
;
|
| 50 |
+
(3)
|
| 51 |
+
a
|
| 52 |
+
self-evolution recipe
|
| 53 |
+
in which the policy SLM and PPM are built from scratch and iteratively evolved to improve reasoning capabilities.
|
| 54 |
+
Through 4 rounds of self-evolution with millions of synthesized solutions for 747k math problems,
|
| 55 |
+
\sysname
|
| 56 |
+
boosts SLMs’ math reasoning to state-of-the-art levels. On the MATH benchmark, it improves Qwen2.5-Math-7B from 58.8% to 90.0% and Phi3-mini-3.8B from 41.4% to 86.4%, surpassing o1-preview by +4.5% and +0.9%. On the USA Math Olympiad (AIME),
|
| 57 |
+
\sysname
|
| 58 |
+
solves an average of 53.3% (8/15) of problems, ranking among the top 20% the brightest high school math students. Code and data will be available at
|
| 59 |
+
https://github.com/microsoft/rStar
|
| 60 |
+
.
|
| 61 |
+
Task
|
| 62 |
+
(pass@1 Acc)
|
| 63 |
+
rStar-Math
|
| 64 |
+
(Qwen-7B)
|
| 65 |
+
rStar-Math
|
| 66 |
+
(Qwen-1.5B)
|
| 67 |
+
rStar-Math
|
| 68 |
+
(Phi3-mini)
|
| 69 |
+
OpenAI
|
| 70 |
+
o1-preview
|
| 71 |
+
OpenAI
|
| 72 |
+
o1-mini
|
| 73 |
+
QWQ
|
| 74 |
+
32B-preview
|
| 75 |
+
GPT-4o
|
| 76 |
+
DeepSeek-V3
|
| 77 |
+
MATH
|
| 78 |
+
90.0
|
| 79 |
+
88.6
|
| 80 |
+
86.4
|
| 81 |
+
85.5
|
| 82 |
+
90.0
|
| 83 |
+
90.6
|
| 84 |
+
76.6
|
| 85 |
+
90.2
|
| 86 |
+
AIME 2024
|
| 87 |
+
53.3
|
| 88 |
+
46.7
|
| 89 |
+
43.3
|
| 90 |
+
44.6
|
| 91 |
+
56.7
|
| 92 |
+
50.0
|
| 93 |
+
9.3
|
| 94 |
+
39.2
|
| 95 |
+
Olympiad Bench
|
| 96 |
+
65.6
|
| 97 |
+
64.6
|
| 98 |
+
60.3
|
| 99 |
+
-
|
| 100 |
+
65.3
|
| 101 |
+
61.2
|
| 102 |
+
43.3
|
| 103 |
+
55.4
|
| 104 |
+
College Math
|
| 105 |
+
60.5
|
| 106 |
+
59.3
|
| 107 |
+
59.1
|
| 108 |
+
-
|
| 109 |
+
57.8
|
| 110 |
+
55.8
|
| 111 |
+
48.5
|
| 112 |
+
58.9
|
| 113 |
+
Omni-Math
|
| 114 |
+
50.5
|
| 115 |
+
48.5
|
| 116 |
+
46.0
|
| 117 |
+
52.5
|
| 118 |
+
60.5
|
| 119 |
+
49.6
|
| 120 |
+
30.5
|
| 121 |
+
35.9
|
| 122 |
+
Table 1:
|
| 123 |
+
\sysname
|
| 124 |
+
enables frontier math reasoning in SLMs via deep thinking over 64 trajectories.
|
| 125 |
+
$*$
|
| 126 |
+
$*$
|
| 127 |
+
footnotetext:
|
| 128 |
+
Equal contribution.
|
| 129 |
+
$\diamond$
|
| 130 |
+
$\diamond$
|
| 131 |
+
footnotetext:
|
| 132 |
+
Project leader; correspondence to lzhani@microsoft.com
|
| 133 |
+
$\S$
|
| 134 |
+
$\S$
|
| 135 |
+
footnotetext:
|
| 136 |
+
Xinyu Guan and Youran Sun did this work during the internship at MSRA. Xinyu Guan (2001gxy@gmail.com) is with Peking University, Youran Sun is with Tsinghua University.
|
| 137 |
+
1
|
| 138 |
+
Introduction
|
| 139 |
+
Recent studies have demonstrated that large language models (LLMs) are capable of tackling mathematical problems
|
| 140 |
+
(Team,
|
| 141 |
+
2024a
|
| 142 |
+
; Yang et al.,
|
| 143 |
+
2024
|
| 144 |
+
; OpenAI,
|
| 145 |
+
2024
|
| 146 |
+
; Liu et al.,
|
| 147 |
+
2024
|
| 148 |
+
)
|
| 149 |
+
. However, the conventional approach of having LLMs generate complete solutions in a single inference – akin to System 1 thinking
|
| 150 |
+
(Daniel,
|
| 151 |
+
2011
|
| 152 |
+
)
|
| 153 |
+
– often yields fast but error-prone results
|
| 154 |
+
(Valmeekam et al.,
|
| 155 |
+
2023
|
| 156 |
+
; OpenAI,
|
| 157 |
+
2023
|
| 158 |
+
)
|
| 159 |
+
. In response, test-time compute scaling
|
| 160 |
+
(Snell et al.,
|
| 161 |
+
2024
|
| 162 |
+
; Qi et al.,
|
| 163 |
+
2024
|
| 164 |
+
)
|
| 165 |
+
suggests a paradigm shift toward a System 2-style thinking, which emulates human reasoning through a slower and deeper thought process. In this paradigm, an LLM serves as a policy model to generate multiple math reasoning steps, which are then evaluated by another LLM acting as a reward model
|
| 166 |
+
(OpenAI,
|
| 167 |
+
2024
|
| 168 |
+
)
|
| 169 |
+
. The steps and solutions deemed more likely to be correct are selected. The process repeats iteratively and ultimately derives the final answer.
|
| 170 |
+
In the test-time compute paradigm, the key is to train a powerful policy model that generates promising solution steps and a reliable reward model that accurately evaluates them, both of which depend on
|
| 171 |
+
high-quality
|
| 172 |
+
training data. Unfortunately, it is well-known that off-the-shelf high-quality math reasoning data is scarce, and synthesizing high-quality math data faces fundamental challenges.
|
| 173 |
+
For the policy model, it is challenging to distinguish erroneous reasoning steps from the correct ones, complicating the elimination of low-quality data. It is worth noting that in math reasoning, a correct final answer does not ensure the correctness of the entire reasoning trace
|
| 174 |
+
(Lanham et al.,
|
| 175 |
+
2023
|
| 176 |
+
)
|
| 177 |
+
. Incorrect intermediate steps significantly decrease data quality.
|
| 178 |
+
As for the reward model, process reward modeling (PRM) shows a great potential by providing fine-grained feedback on intermediate steps
|
| 179 |
+
(Lightman et al.,
|
| 180 |
+
2023
|
| 181 |
+
)
|
| 182 |
+
. However, the training data is even scarcer in this regard: accurate step-by-step feedback requires intense human labeling efforts and is impractical to scale, while those automatic annotation attempts show limited gains due to noisy reward scores
|
| 183 |
+
(Luo et al.,
|
| 184 |
+
2024
|
| 185 |
+
; Wang et al.,
|
| 186 |
+
2024c
|
| 187 |
+
; Chen et al.,
|
| 188 |
+
2024
|
| 189 |
+
)
|
| 190 |
+
.
|
| 191 |
+
Due to the above challenges, existing distill-based data synthesis approaches to training policy models, e.g., scaling up GPT4-distilled CoT data
|
| 192 |
+
(Tang et al.,
|
| 193 |
+
2024
|
| 194 |
+
; Huang et al.,
|
| 195 |
+
2024
|
| 196 |
+
)
|
| 197 |
+
, have shown diminishing returns and cannot exceed the capability of their teacher model; meanwhile, as of today, training reliable PRMs for math reasoning remains an open question.
|
| 198 |
+
Figure 1:
|
| 199 |
+
The overview of
|
| 200 |
+
\sysname
|
| 201 |
+
.
|
| 202 |
+
In this work, we introduce
|
| 203 |
+
\sysname
|
| 204 |
+
, a self-evolvable System 2-style reasoning approach that achieves the state-of-the-art math reasoning, rivaling and sometimes even surpassing OpenAI o1 on challenging math competition benchmarks with a model size as small as 7 billion. Unlike solutions relying on superior LLMs for data synthesis,
|
| 205 |
+
\sysname
|
| 206 |
+
leverages smaller language models (SLMs) with Monte Carlo Tree Search (MCTS) to establish a self-evolutionary process, iteratively generating higher-quality training data. To achieve self-evolution,
|
| 207 |
+
\sysname
|
| 208 |
+
introduces three key innovations.
|
| 209 |
+
First, a novel code-augmented CoT data synthesis method, which performs
|
| 210 |
+
extensive
|
| 211 |
+
MCTS rollouts to generate
|
| 212 |
+
step-by-step verified reasoning trajectories
|
| 213 |
+
with
|
| 214 |
+
self-annotated MCTS Q-values
|
| 215 |
+
. Specifically, math problem-solving is decomposed into multi-step generation within MCTS. At each step, the SLM serving as the policy model samples candidate nodes, each generating a one-step CoT and the corresponding Python code. To verify the generation quality, only nodes with successful Python code execution are retained, thus mitigating errors in intermediate steps. Moreover, extensive MCTS rollouts automatically assign a Q-value to each intermediate step based on its contribution: steps contributing to more trajectories that lead to the correct answer are given higher Q-values and considered higher quality. This ensures that the reasoning trajectories generated by SLMs consist of correct, high-quality intermediate steps.
|
| 216 |
+
Second, a novel method that trains an SLM acting as a
|
| 217 |
+
process preference model
|
| 218 |
+
, i.e., a PPM to implement the desired PRM, that reliably predicts a reward label for each math reasoning step. The PPM leverages the fact that, although Q-values are still not precise enough to score each reasoning step despite using extensive MCTS rollouts, the Q-values can reliably distinguish positive (correct) steps from negative (irrelevant/incorrect) ones. Thus the training method constructs preference pairs for each step based on Q-values and uses a pairwise ranking loss
|
| 219 |
+
(Ouyang et al.,
|
| 220 |
+
2022
|
| 221 |
+
)
|
| 222 |
+
to optimize PPM’s score prediction for each reasoning step, achieving reliable labeling. This approach avoids conventional methods that directly use Q-values as reward labels
|
| 223 |
+
(Luo et al.,
|
| 224 |
+
2024
|
| 225 |
+
; Chen et al.,
|
| 226 |
+
2024
|
| 227 |
+
)
|
| 228 |
+
, which are inherently noisy and imprecise in stepwise reward assignment.
|
| 229 |
+
Finally, a four-round self-evolution recipe that progressively builds both a frontier policy model and PPM from scratch. We begin by curating a dataset of 747k math word problems from publicly available sources. In each round, we use the latest policy model and PPM to perform MCTS, generating increasingly high-quality training data using the above two methods to train a stronger policy model and PPM for next round. Each round achieves progressive refinement: (1) a stronger policy SLM, (2) a more reliable PPM, (3) generating better reasoning trajectories via PPM-augmented MCTS, and (4) improving training data coverage to tackle more challenging and even competition-level math problems.
|
| 230 |
+
Extensive experiments across four SLMs (1.5B-7B) and seven math reasoning tasks demonstrate the effectiveness of
|
| 231 |
+
\sysname
|
| 232 |
+
. Remarkably,
|
| 233 |
+
\sysname
|
| 234 |
+
improves all four SLMs, matching or even surpassing OpenAI o1 on challenging math benchmarks. On MATH benchmark, with 8 search trajectories,
|
| 235 |
+
\sysname
|
| 236 |
+
boosts Qwen2.5-Math-7B from 58.8% to 89.4% and Qwen2.5-Math-1.5B from 51.2% to 87.8%. With 64 trajectories, the scores rise to 90% and 88.4%, outperforming o1-preview by 4.5% and 2.6% and matching o1-mini’s 90%. On the Olympiad-level AIME 2024,
|
| 237 |
+
\sysname
|
| 238 |
+
solves on average 53.3% (8/15) of the problems, exceeding o1-preview by 8.7% and all other open-sourced LLMs. We further conduct comprehensive experiments to verify the superiority of step-by-step verified reasoning trajectories over state-of-the-art data synthesis baselines, as well as the PPM’s effectiveness compared to outcome reward models and Q value-based PRMs. Finally, we present key findings from
|
| 239 |
+
\sysname
|
| 240 |
+
deep thinking, including the intrinsic self-reflection capability and PPM’s preference for theorem-applications intermediate steps.
|
| 241 |
+
2
|
| 242 |
+
Related Works
|
| 243 |
+
Math Data Synthesis
|
| 244 |
+
. Advancements in LLM math reasoning have largely relied on curating high-quality CoT data, with most leading approaches being GPT-distilled, using frontier models like GPT-4 for synthesis
|
| 245 |
+
(Wang et al.,
|
| 246 |
+
2024b
|
| 247 |
+
; Gou et al.,
|
| 248 |
+
2023
|
| 249 |
+
; Luo et al.,
|
| 250 |
+
2023
|
| 251 |
+
)
|
| 252 |
+
. Notable works include NuminaMath
|
| 253 |
+
(Jia LI and Polu,
|
| 254 |
+
2024a
|
| 255 |
+
)
|
| 256 |
+
and
|
| 257 |
+
MetaMath
|
| 258 |
+
(Yu et al.,
|
| 259 |
+
2023b
|
| 260 |
+
)
|
| 261 |
+
. While effective, this limits reasoning to the capabilities of the teacher LLM.
|
| 262 |
+
Hard problems that the teacher LLM cannot solve are excluded in the training set.
|
| 263 |
+
Even solvable problems may contain error-prone intermediate steps, which are hard to detect. Although rejection sampling methods
|
| 264 |
+
(Yuan et al.,
|
| 265 |
+
2023
|
| 266 |
+
; Brown et al.,
|
| 267 |
+
2024
|
| 268 |
+
)
|
| 269 |
+
can improve data quality,
|
| 270 |
+
they do not guarantee correct intermediate steps. As a result, scaling up CoT data has diminishing returns, with gains nearing saturation—e.g., OpenMathInstruct-2
|
| 271 |
+
(Toshniwal et al.,
|
| 272 |
+
2024
|
| 273 |
+
)
|
| 274 |
+
only sees a 3.9% boost on MATH despite an 8× increase in dataset size.
|
| 275 |
+
Scaling Test-time Compute
|
| 276 |
+
has introduced new scaling laws, allowing LLMs to improve performance across by generating multiple samples and using reward models for best-solution selection
|
| 277 |
+
(Snell et al.,
|
| 278 |
+
2024
|
| 279 |
+
; Wu et al.,
|
| 280 |
+
2024
|
| 281 |
+
; Brown et al.,
|
| 282 |
+
2024
|
| 283 |
+
)
|
| 284 |
+
. Various test-time search methods have been proposed
|
| 285 |
+
(Kang et al.,
|
| 286 |
+
2024
|
| 287 |
+
; Wang et al.,
|
| 288 |
+
2024a
|
| 289 |
+
)
|
| 290 |
+
, including random sampling
|
| 291 |
+
(Wang et al.,
|
| 292 |
+
2023
|
| 293 |
+
)
|
| 294 |
+
and tree-search methods
|
| 295 |
+
(Yao et al.,
|
| 296 |
+
2024
|
| 297 |
+
; Hao et al.,
|
| 298 |
+
2023
|
| 299 |
+
; Zhang et al.,
|
| 300 |
+
2024b
|
| 301 |
+
; Qi et al.,
|
| 302 |
+
2024
|
| 303 |
+
)
|
| 304 |
+
like MCTS. However, open-source methods for scaling test-time computation have shown limited gains in math reasoning, often due to policy LLM or reward model limitations.
|
| 305 |
+
\sysname
|
| 306 |
+
addresses this by iteratively evolving the policy LLM and reward model, achieving System 2 mathematical reasoning performance comparable to OpenAI o1
|
| 307 |
+
(OpenAI,
|
| 308 |
+
2024
|
| 309 |
+
)
|
| 310 |
+
.
|
| 311 |
+
Reward Models
|
| 312 |
+
are crucial for effective System 2 reasoning but are challenging to obtain. Recent works include LLM-as-a-Judge for verification
|
| 313 |
+
(Zheng et al.,
|
| 314 |
+
2023
|
| 315 |
+
; Qi et al.,
|
| 316 |
+
2024
|
| 317 |
+
)
|
| 318 |
+
and specialized reward models like Outcome Reward Model
|
| 319 |
+
(Yang et al.,
|
| 320 |
+
2024
|
| 321 |
+
; Yu et al.,
|
| 322 |
+
2023a
|
| 323 |
+
)
|
| 324 |
+
and Process Reward Model (PRM)
|
| 325 |
+
(Lightman et al.,
|
| 326 |
+
2024
|
| 327 |
+
)
|
| 328 |
+
. While PRMs offer promising dense, step-level reward signals for
|
| 329 |
+
complex reasoning
|
| 330 |
+
(Luo et al.,
|
| 331 |
+
2024
|
| 332 |
+
; Wang et al.,
|
| 333 |
+
2024c
|
| 334 |
+
)
|
| 335 |
+
, collecting step-level annotations remains an obstacle. While
|
| 336 |
+
Kang et al. (
|
| 337 |
+
2024
|
| 338 |
+
); Wang et al. (
|
| 339 |
+
2024a
|
| 340 |
+
)
|
| 341 |
+
rely on costly human-annotated datasets like PRM800k
|
| 342 |
+
(Lightman et al.,
|
| 343 |
+
2024
|
| 344 |
+
)
|
| 345 |
+
,
|
| 346 |
+
recent approaches
|
| 347 |
+
(Wang et al.,
|
| 348 |
+
2024c
|
| 349 |
+
; Luo et al.,
|
| 350 |
+
2024
|
| 351 |
+
)
|
| 352 |
+
explore automated annotation via Monte Carlo Sampling or MCTS. However, they struggle to generate precise reward scores, which limits performance gains.
|
| 353 |
+
\sysname
|
| 354 |
+
introduces a novel process preference reward (PPM) that eliminates the need for accurate step-level reward score annotation.
|
| 355 |
+
3
|
| 356 |
+
Methodology
|
| 357 |
+
3.1
|
| 358 |
+
Design Choices
|
| 359 |
+
MCTS for Effective System 2 Reasoning
|
| 360 |
+
.
|
| 361 |
+
We aim to train a math policy SLM and a process reward model (PRM), and integrating both within Monte Carlo Tree Search (MCTS) for System 2 deep thinking. MCTS is chosen for two key reasons. First, it breaks down complex math problems into simpler single-step generation tasks, reducing the difficulty for the policy SLM compared to other System 2 methods like Best-of-N
|
| 362 |
+
(Brown et al.,
|
| 363 |
+
2024
|
| 364 |
+
)
|
| 365 |
+
or self-consistency
|
| 366 |
+
(Wang et al.,
|
| 367 |
+
2023
|
| 368 |
+
)
|
| 369 |
+
, which require generating full solutions in one inference.
|
| 370 |
+
Second, the step-by-step generation in MCTS naturally yields step-level training data for both models. Standard MCTS rollout automatically assign Q-value to each step based on its contribution to the final correct answer, obviating the need for human-generated step-level annotations for process reward model training.
|
| 371 |
+
Ideally, advanced LLMs such as GPT-4 could be integrated within MCTS to generate training data. However, this approach faces two key challenges. First, even these powerful models struggle to consistently solve difficult problems, such as Olympiad-level mathematics. Consequently, the resulting training data would primarily consist of simpler solvable problems, limiting its diversity and quality. Second, annotating per-step Q-values demands extensive MCTS rollouts; insufficient tree exploration can lead to spurious Q-value assignments, such as overestimating suboptimal steps. Given that each rollout involves multiple single-step generations and these models are computationally expensive, increasing rollouts significantly raises inference costs.
|
| 372 |
+
Overview
|
| 373 |
+
. To this end, we explore using two 7B SLMs (a policy SLM and a PRM) to generate higher-quality training data, with their smaller size allowing for extensive MCTS rollouts on accessible hardware (e.g., 4
|
| 374 |
+
×
|
| 375 |
+
\times
|
| 376 |
+
40GB A100 GPUs). However, self-generating data presents greater challenges for SLMs, due to their weaker capabilities.
|
| 377 |
+
SLMs frequently fail to generate correct solutions, and even when the final answer is correct, the intermediate steps are often flawed or of poor quality. Moreover, SLMs solve fewer challenging problems compared to advanced models like GPT-4.
|
| 378 |
+
This section introduces our methodology, as illustrated in Fig.
|
| 379 |
+
1
|
| 380 |
+
. To mitigate errors and low-quality intermediate steps, we introduce a code-augmented CoT synthetic method, which performs extensive MCTS rollouts to generate step-by-step verified reasoning trajectories, annotated with Q-values. To further improve SLM performance on challenging problems, we introduce a four-round self-evolution recipe. In each round, both the policy SLM and the reward model are updated to stronger versions, progressively tackling more difficult problems and generating higher-quality training data. Finally, we present a novel process reward model training approach that eliminates the need for precise per-step reward annotations, yielding the more
|
| 381 |
+
effective process preference model (PPM).
|
| 382 |
+
3.2
|
| 383 |
+
Step-by-Step Verified Reasoning Trajectory
|
| 384 |
+
We start by introducing our method for generating step-by-step verified reasoning trajectories with per-step Q-value annotations. Given a problem
|
| 385 |
+
x
|
| 386 |
+
x
|
| 387 |
+
and a policy model
|
| 388 |
+
M
|
| 389 |
+
M
|
| 390 |
+
, we run the standard MCTS to incrementally construct a search tree for step-by-step solution exploration. As shown in Fig.
|
| 391 |
+
1
|
| 392 |
+
(a),
|
| 393 |
+
the root node represents question
|
| 394 |
+
x
|
| 395 |
+
x
|
| 396 |
+
, while child nodes correspond to intermediate steps
|
| 397 |
+
s
|
| 398 |
+
s
|
| 399 |
+
generated by
|
| 400 |
+
M
|
| 401 |
+
M
|
| 402 |
+
. A root-to-leaf path ending at terminal node
|
| 403 |
+
s
|
| 404 |
+
d
|
| 405 |
+
s_{d}
|
| 406 |
+
forms a trajectory
|
| 407 |
+
𝐭
|
| 408 |
+
=
|
| 409 |
+
x
|
| 410 |
+
⊕
|
| 411 |
+
s
|
| 412 |
+
1
|
| 413 |
+
⊕
|
| 414 |
+
s
|
| 415 |
+
2
|
| 416 |
+
⊕
|
| 417 |
+
…
|
| 418 |
+
⊕
|
| 419 |
+
s
|
| 420 |
+
d
|
| 421 |
+
\mathbf{t}=x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{d}
|
| 422 |
+
, with each step
|
| 423 |
+
s
|
| 424 |
+
i
|
| 425 |
+
s_{i}
|
| 426 |
+
assigned a Q-value
|
| 427 |
+
Q
|
| 428 |
+
|
| 429 |
+
(
|
| 430 |
+
s
|
| 431 |
+
i
|
| 432 |
+
)
|
| 433 |
+
Q(s_{i})
|
| 434 |
+
.
|
| 435 |
+
From the search tree
|
| 436 |
+
𝒯
|
| 437 |
+
\mathcal{T}
|
| 438 |
+
, we extract solution trajectories
|
| 439 |
+
𝕋
|
| 440 |
+
=
|
| 441 |
+
{
|
| 442 |
+
𝐭
|
| 443 |
+
1
|
| 444 |
+
,
|
| 445 |
+
𝐭
|
| 446 |
+
2
|
| 447 |
+
,
|
| 448 |
+
…
|
| 449 |
+
,
|
| 450 |
+
𝐭
|
| 451 |
+
n
|
| 452 |
+
}
|
| 453 |
+
|
| 454 |
+
(
|
| 455 |
+
n
|
| 456 |
+
≥
|
| 457 |
+
1
|
| 458 |
+
)
|
| 459 |
+
\mathbb{T}=\{\mathbf{t}^{1},\mathbf{t}^{2},...,\mathbf{t}^{n}\}(n\geq 1)
|
| 460 |
+
. Our goal is to select high-quality trajectories from
|
| 461 |
+
𝒯
|
| 462 |
+
\mathcal{T}
|
| 463 |
+
to construct the training set. For this purpose, we introduce code-augmented CoT synthesis method to filter out low-quality generations and perform extensive rollouts to improve the reliability of Q-value accuracy.
|
| 464 |
+
Code-augmented CoT Generation
|
| 465 |
+
. Prior MCTS approaches primarily generate natural language (NL) CoTs
|
| 466 |
+
(Qi et al.,
|
| 467 |
+
2024
|
| 468 |
+
; Zhang et al.,
|
| 469 |
+
2024a
|
| 470 |
+
)
|
| 471 |
+
. However, LLMs often suffer from hallucination, producing incorrect or irrelevant steps yet still arrive at the correct answer by chance
|
| 472 |
+
(Lanham et al.,
|
| 473 |
+
2023
|
| 474 |
+
)
|
| 475 |
+
. These flawed steps are challenging to detect and eliminate. To address this, we propose a novel code execution augmented CoT. As shown in Fig.
|
| 476 |
+
2
|
| 477 |
+
, the policy model generates a one-step NL CoT alongside its corresponding Python code, where the NL CoT is embedded as a Python comment. Only generations with successfully executed Python code are retained as valid candidates.
|
| 478 |
+
Figure 2:
|
| 479 |
+
An example of Code-augmented CoT.
|
| 480 |
+
Specifically, starting from the initial root node
|
| 481 |
+
x
|
| 482 |
+
x
|
| 483 |
+
, we perform multiple MCTS iterations through
|
| 484 |
+
selection
|
| 485 |
+
,
|
| 486 |
+
expansion
|
| 487 |
+
,
|
| 488 |
+
rollout
|
| 489 |
+
, and
|
| 490 |
+
back-propagation
|
| 491 |
+
. At step
|
| 492 |
+
i
|
| 493 |
+
i
|
| 494 |
+
, we collect the latest reasoning trajectory
|
| 495 |
+
x
|
| 496 |
+
⊕
|
| 497 |
+
s
|
| 498 |
+
1
|
| 499 |
+
⊕
|
| 500 |
+
s
|
| 501 |
+
2
|
| 502 |
+
⊕
|
| 503 |
+
…
|
| 504 |
+
⊕
|
| 505 |
+
s
|
| 506 |
+
i
|
| 507 |
+
−
|
| 508 |
+
1
|
| 509 |
+
x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{i-1}
|
| 510 |
+
as the current state. Based on this state, we prompt (see Appendix
|
| 511 |
+
A.3
|
| 512 |
+
) the policy model to generate
|
| 513 |
+
n
|
| 514 |
+
n
|
| 515 |
+
candidates
|
| 516 |
+
s
|
| 517 |
+
i
|
| 518 |
+
,
|
| 519 |
+
0
|
| 520 |
+
,
|
| 521 |
+
…
|
| 522 |
+
,
|
| 523 |
+
s
|
| 524 |
+
i
|
| 525 |
+
,
|
| 526 |
+
n
|
| 527 |
+
−
|
| 528 |
+
1
|
| 529 |
+
s_{i,0},...,s_{i,n-1}
|
| 530 |
+
for step
|
| 531 |
+
i
|
| 532 |
+
i
|
| 533 |
+
. Python code execution is then employed to filter valid nodes. As shown in Fig.
|
| 534 |
+
2
|
| 535 |
+
, each generation
|
| 536 |
+
s
|
| 537 |
+
i
|
| 538 |
+
,
|
| 539 |
+
j
|
| 540 |
+
s_{i,j}
|
| 541 |
+
is concatenated with the code from all previous steps, forming
|
| 542 |
+
s
|
| 543 |
+
1
|
| 544 |
+
⊕
|
| 545 |
+
s
|
| 546 |
+
2
|
| 547 |
+
⊕
|
| 548 |
+
…
|
| 549 |
+
⊕
|
| 550 |
+
s
|
| 551 |
+
i
|
| 552 |
+
−
|
| 553 |
+
1
|
| 554 |
+
⊕
|
| 555 |
+
s
|
| 556 |
+
i
|
| 557 |
+
,
|
| 558 |
+
j
|
| 559 |
+
s_{1}\oplus s_{2}\oplus...\oplus s_{i-1}\oplus s_{i,j}
|
| 560 |
+
. Candidates that execute successfully are retained as valid nodes and scored by the PPM, which assigns a Q-value
|
| 561 |
+
q
|
| 562 |
+
|
| 563 |
+
(
|
| 564 |
+
s
|
| 565 |
+
i
|
| 566 |
+
)
|
| 567 |
+
q(s_{i})
|
| 568 |
+
.
|
| 569 |
+
Then, we use the well-known Upper Confidence bounds for Trees (UCT)
|
| 570 |
+
(Kocsis and Szepesvári,
|
| 571 |
+
2006
|
| 572 |
+
)
|
| 573 |
+
to select the best node among the
|
| 574 |
+
n
|
| 575 |
+
n
|
| 576 |
+
candidates. This selection process is mathematically represented as:
|
| 577 |
+
UCT
|
| 578 |
+
|
| 579 |
+
(
|
| 580 |
+
s
|
| 581 |
+
)
|
| 582 |
+
=
|
| 583 |
+
Q
|
| 584 |
+
|
| 585 |
+
(
|
| 586 |
+
s
|
| 587 |
+
)
|
| 588 |
+
+
|
| 589 |
+
c
|
| 590 |
+
|
| 591 |
+
ln
|
| 592 |
+
|
| 593 |
+
N
|
| 594 |
+
p
|
| 595 |
+
|
| 596 |
+
a
|
| 597 |
+
|
| 598 |
+
r
|
| 599 |
+
|
| 600 |
+
e
|
| 601 |
+
|
| 602 |
+
n
|
| 603 |
+
|
| 604 |
+
t
|
| 605 |
+
|
| 606 |
+
(
|
| 607 |
+
s
|
| 608 |
+
)
|
| 609 |
+
N
|
| 610 |
+
|
| 611 |
+
(
|
| 612 |
+
s
|
| 613 |
+
)
|
| 614 |
+
;
|
| 615 |
+
where
|
| 616 |
+
Q
|
| 617 |
+
|
| 618 |
+
(
|
| 619 |
+
s
|
| 620 |
+
)
|
| 621 |
+
=
|
| 622 |
+
q
|
| 623 |
+
|
| 624 |
+
(
|
| 625 |
+
s
|
| 626 |
+
)
|
| 627 |
+
N
|
| 628 |
+
|
| 629 |
+
(
|
| 630 |
+
s
|
| 631 |
+
)
|
| 632 |
+
\displaystyle\text{UCT}(s)=Q(s)+c\sqrt{\frac{\ln N_{parent}(s)}{N(s)}};\quad\text{where}\quad Q(s)=\frac{q(s)}{N(s)}
|
| 633 |
+
(1)
|
| 634 |
+
where
|
| 635 |
+
N
|
| 636 |
+
|
| 637 |
+
(
|
| 638 |
+
s
|
| 639 |
+
)
|
| 640 |
+
N(s)
|
| 641 |
+
denotes the number of visits to node
|
| 642 |
+
s
|
| 643 |
+
s
|
| 644 |
+
, and
|
| 645 |
+
N
|
| 646 |
+
parent
|
| 647 |
+
|
| 648 |
+
(
|
| 649 |
+
s
|
| 650 |
+
)
|
| 651 |
+
N_{\text{parent}}(s)
|
| 652 |
+
is the visit count of
|
| 653 |
+
s
|
| 654 |
+
s
|
| 655 |
+
’s parent node. The predicted reward
|
| 656 |
+
q
|
| 657 |
+
|
| 658 |
+
(
|
| 659 |
+
s
|
| 660 |
+
)
|
| 661 |
+
q(s)
|
| 662 |
+
is provided by the PPM and will be updated through back-propagation.
|
| 663 |
+
c
|
| 664 |
+
c
|
| 665 |
+
is a constant that balances exploitation and exploration.
|
| 666 |
+
Extensive Rollouts for Q-value Annotation
|
| 667 |
+
. Accurate Q-value
|
| 668 |
+
Q
|
| 669 |
+
|
| 670 |
+
(
|
| 671 |
+
s
|
| 672 |
+
)
|
| 673 |
+
Q(s)
|
| 674 |
+
annotation in Eq.
|
| 675 |
+
1
|
| 676 |
+
is crucial for guiding MCTS node selection towards correct problem-solving paths and identifying high-quality steps within trajectories.
|
| 677 |
+
To improve Q-value reliability, we draw inspiration from Go players, who retrospectively evaluate the reward of each move based on game outcomes. Although initial estimates may be imprecise, repeated gameplay refines these evaluations over time. Similarly, in each rollout, we update the Q-value of each step based on its contribution to achieving the correct final answer. After extensive MCTS rollouts, steps consistently leading to correct answers achieve higher Q-values, occasional successes yield moderate Q-values, and consistently incorrect steps receive low Q-values. Specifically, we introduce two self-annotation methods to obtain these step-level Q-values. Fig.
|
| 678 |
+
1
|
| 679 |
+
(c) shows the detailed setting in the four rounds of self-evolution.
|
| 680 |
+
Terminal-guided annotation
|
| 681 |
+
. During the first two rounds, when the PPM is unavailable or insufficiently accurate, we use terminal-guided annotation. Formally, let
|
| 682 |
+
q
|
| 683 |
+
|
| 684 |
+
(
|
| 685 |
+
s
|
| 686 |
+
i
|
| 687 |
+
)
|
| 688 |
+
k
|
| 689 |
+
q(s_{i})^{k}
|
| 690 |
+
denote the q value for step
|
| 691 |
+
s
|
| 692 |
+
i
|
| 693 |
+
s_{i}
|
| 694 |
+
after back-propagation in the
|
| 695 |
+
k
|
| 696 |
+
t
|
| 697 |
+
|
| 698 |
+
h
|
| 699 |
+
k^{th}
|
| 700 |
+
rollout. Following AlphaGo
|
| 701 |
+
(Silver et al.,
|
| 702 |
+
2017
|
| 703 |
+
)
|
| 704 |
+
and rStar
|
| 705 |
+
(Qi et al.,
|
| 706 |
+
2024
|
| 707 |
+
)
|
| 708 |
+
, we score each intermediate node based on its contribution to the final correct answer:
|
| 709 |
+
q
|
| 710 |
+
|
| 711 |
+
(
|
| 712 |
+
s
|
| 713 |
+
i
|
| 714 |
+
)
|
| 715 |
+
k
|
| 716 |
+
=
|
| 717 |
+
q
|
| 718 |
+
|
| 719 |
+
(
|
| 720 |
+
s
|
| 721 |
+
i
|
| 722 |
+
)
|
| 723 |
+
k
|
| 724 |
+
−
|
| 725 |
+
1
|
| 726 |
+
+
|
| 727 |
+
q
|
| 728 |
+
|
| 729 |
+
(
|
| 730 |
+
s
|
| 731 |
+
d
|
| 732 |
+
)
|
| 733 |
+
k
|
| 734 |
+
;
|
| 735 |
+
\displaystyle q(s_{i})^{k}=q(s_{i})^{k-1}+q(s_{d})^{k};
|
| 736 |
+
(2)
|
| 737 |
+
where the initial q value
|
| 738 |
+
q
|
| 739 |
+
|
| 740 |
+
(
|
| 741 |
+
s
|
| 742 |
+
i
|
| 743 |
+
)
|
| 744 |
+
0
|
| 745 |
+
=
|
| 746 |
+
0
|
| 747 |
+
q(s_{i})^{0}=0
|
| 748 |
+
in the first rollout. If this step frequently leads to a correct answer, its
|
| 749 |
+
q
|
| 750 |
+
q
|
| 751 |
+
value will increase; otherwise, it decreases. Terminal nodes are scored as
|
| 752 |
+
q
|
| 753 |
+
|
| 754 |
+
(
|
| 755 |
+
s
|
| 756 |
+
d
|
| 757 |
+
)
|
| 758 |
+
=
|
| 759 |
+
1
|
| 760 |
+
q(s_{d})=1
|
| 761 |
+
for correct answers and
|
| 762 |
+
q
|
| 763 |
+
|
| 764 |
+
(
|
| 765 |
+
s
|
| 766 |
+
d
|
| 767 |
+
)
|
| 768 |
+
=
|
| 769 |
+
−
|
| 770 |
+
1
|
| 771 |
+
q(s_{d})=-1
|
| 772 |
+
otherwise, as shown in Fig.
|
| 773 |
+
1
|
| 774 |
+
.
|
| 775 |
+
PRM-augmented annotation
|
| 776 |
+
. Starting from the third round, we use PPM to score each step for more effective generation. Compared to terminal-guided annotation, which requires multiple rollouts for a meaningful
|
| 777 |
+
q
|
| 778 |
+
q
|
| 779 |
+
value, PPM directly predicts a non-zero initial
|
| 780 |
+
q
|
| 781 |
+
q
|
| 782 |
+
value.
|
| 783 |
+
PPM-augmented MCTS also helps the policy model to generate higher-quality steps, guiding solutions towards correct paths. Formally, for step
|
| 784 |
+
s
|
| 785 |
+
i
|
| 786 |
+
s_{i}
|
| 787 |
+
, PPM predicts an initial
|
| 788 |
+
q
|
| 789 |
+
|
| 790 |
+
(
|
| 791 |
+
s
|
| 792 |
+
i
|
| 793 |
+
)
|
| 794 |
+
0
|
| 795 |
+
q(s_{i})^{0}
|
| 796 |
+
value based on the partial trajectory:
|
| 797 |
+
q
|
| 798 |
+
|
| 799 |
+
(
|
| 800 |
+
s
|
| 801 |
+
i
|
| 802 |
+
)
|
| 803 |
+
0
|
| 804 |
+
=
|
| 805 |
+
P
|
| 806 |
+
|
| 807 |
+
P
|
| 808 |
+
|
| 809 |
+
M
|
| 810 |
+
|
| 811 |
+
(
|
| 812 |
+
x
|
| 813 |
+
⊕
|
| 814 |
+
s
|
| 815 |
+
1
|
| 816 |
+
⊕
|
| 817 |
+
s
|
| 818 |
+
2
|
| 819 |
+
⊕
|
| 820 |
+
…
|
| 821 |
+
⊕
|
| 822 |
+
s
|
| 823 |
+
i
|
| 824 |
+
−
|
| 825 |
+
1
|
| 826 |
+
⊕
|
| 827 |
+
s
|
| 828 |
+
i
|
| 829 |
+
)
|
| 830 |
+
\displaystyle q(s_{i})^{0}=PPM(x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{i-1}\oplus s_{i})
|
| 831 |
+
(3)
|
| 832 |
+
This
|
| 833 |
+
q
|
| 834 |
+
q
|
| 835 |
+
value will be updated based on terminal node’s
|
| 836 |
+
q
|
| 837 |
+
|
| 838 |
+
(
|
| 839 |
+
s
|
| 840 |
+
d
|
| 841 |
+
)
|
| 842 |
+
q(s_{d})
|
| 843 |
+
value through MCTS
|
| 844 |
+
back-propagation
|
| 845 |
+
in Eq.
|
| 846 |
+
2
|
| 847 |
+
.
|
| 848 |
+
For terminal node
|
| 849 |
+
s
|
| 850 |
+
d
|
| 851 |
+
s_{d}
|
| 852 |
+
, we do not use PRM for scoring during training data generation. Instead, we assign a more accurate score based on ground truth labels as terminal-guided rewarding.
|
| 853 |
+
3.3
|
| 854 |
+
Process Preference Model
|
| 855 |
+
Process reward models, which provide granular step-level reward signals, is highly desirable for solving challenging math problems. However, obtaining high-quality step-level training data remains an open challenge. Existing methods rely on human annotations
|
| 856 |
+
(Lightman et al.,
|
| 857 |
+
2023
|
| 858 |
+
)
|
| 859 |
+
or MCTS-generated scores
|
| 860 |
+
(Zhang et al.,
|
| 861 |
+
2024a
|
| 862 |
+
; Chen et al.,
|
| 863 |
+
2024
|
| 864 |
+
)
|
| 865 |
+
to assign a score for each step. These scores then serve as training targets, with methods such as MSE loss
|
| 866 |
+
(Chen et al.,
|
| 867 |
+
2024
|
| 868 |
+
)
|
| 869 |
+
or pointwise loss
|
| 870 |
+
(Wang et al.,
|
| 871 |
+
2024c
|
| 872 |
+
; Luo et al.,
|
| 873 |
+
2024
|
| 874 |
+
; Zhang et al.,
|
| 875 |
+
2024a
|
| 876 |
+
)
|
| 877 |
+
used to minimize the difference between predicted and labeled scores.
|
| 878 |
+
As a result, the precision of these annotated step-level reward scores directly determines the effectiveness of the resulting process reward model.
|
| 879 |
+
Unfortunately, precise per-step scoring remains a unsolved challenge. Although our extensive MCTS rollouts improve the reliability of Q-values, precisely evaluating fine-grained step quality presents a major obstacle. For instance, among a set of correct steps, it is difficult to rank them as best, second-best, or average and then assign precise scores. Similarly, among incorrect steps, differentiating the worst from moderately poor steps poses analogous challenges. Even expert human annotation struggles with consistency, particularly at scale, leading to inherent noise in training labels.
|
| 880 |
+
We introduce a novel training method that trains a process preference model (PPM) by constructing step-level positive-negative preference pairs. As shown in Fig.
|
| 881 |
+
1
|
| 882 |
+
(b), instead of using Q-values as direct reward labels, we use them to select steps from MCTS tree for preference pair construction. For each step, we select two candidates with the highest Q-values as positive steps and two with the lowest as negative steps. Critically, the selected positive steps must lead to a correct final answer, while negative steps must lead to incorrect answers. For intermediate steps (except the final answer step), the positive and negative pairs share the same preceding steps. For the final answer step, where identical reasoning trajectories rarely yield different final answers, we relax this restriction.
|
| 883 |
+
We select two correct trajectories with the highest average Q-values as positive examples and two incorrect trajectories with the lowest average Q-values as negative examples. Following
|
| 884 |
+
(Ouyang et al.,
|
| 885 |
+
2022
|
| 886 |
+
)
|
| 887 |
+
, we define our loss function using the standard Bradley-Terry model with a pairwise ranking loss:
|
| 888 |
+
ℒ
|
| 889 |
+
p
|
| 890 |
+
|
| 891 |
+
p
|
| 892 |
+
|
| 893 |
+
m
|
| 894 |
+
|
| 895 |
+
(
|
| 896 |
+
θ
|
| 897 |
+
)
|
| 898 |
+
=
|
| 899 |
+
−
|
| 900 |
+
1
|
| 901 |
+
2
|
| 902 |
+
×
|
| 903 |
+
2
|
| 904 |
+
|
| 905 |
+
E
|
| 906 |
+
(
|
| 907 |
+
x
|
| 908 |
+
,
|
| 909 |
+
y
|
| 910 |
+
i
|
| 911 |
+
p
|
| 912 |
+
|
| 913 |
+
o
|
| 914 |
+
|
| 915 |
+
s
|
| 916 |
+
,
|
| 917 |
+
y
|
| 918 |
+
i
|
| 919 |
+
n
|
| 920 |
+
|
| 921 |
+
e
|
| 922 |
+
|
| 923 |
+
g
|
| 924 |
+
∈
|
| 925 |
+
𝔻
|
| 926 |
+
)
|
| 927 |
+
|
| 928 |
+
[
|
| 929 |
+
l
|
| 930 |
+
|
| 931 |
+
o
|
| 932 |
+
|
| 933 |
+
g
|
| 934 |
+
|
| 935 |
+
(
|
| 936 |
+
σ
|
| 937 |
+
|
| 938 |
+
(
|
| 939 |
+
r
|
| 940 |
+
θ
|
| 941 |
+
|
| 942 |
+
(
|
| 943 |
+
x
|
| 944 |
+
,
|
| 945 |
+
y
|
| 946 |
+
i
|
| 947 |
+
p
|
| 948 |
+
|
| 949 |
+
o
|
| 950 |
+
|
| 951 |
+
s
|
| 952 |
+
)
|
| 953 |
+
−
|
| 954 |
+
r
|
| 955 |
+
θ
|
| 956 |
+
|
| 957 |
+
(
|
| 958 |
+
x
|
| 959 |
+
,
|
| 960 |
+
y
|
| 961 |
+
i
|
| 962 |
+
n
|
| 963 |
+
|
| 964 |
+
e
|
| 965 |
+
|
| 966 |
+
g
|
| 967 |
+
)
|
| 968 |
+
)
|
| 969 |
+
)
|
| 970 |
+
]
|
| 971 |
+
\displaystyle\mathcal{L}_{ppm}(\theta)=-\frac{1}{2\times 2}E_{(x,y_{i}^{pos},y_{i}^{neg}\in\mathbb{D})}[log(\sigma(r_{\theta}(x,y_{i}^{pos})-r_{\theta}(x,y_{i}^{neg})))]
|
| 972 |
+
(4)
|
| 973 |
+
when
|
| 974 |
+
i
|
| 975 |
+
is not final answer step
|
| 976 |
+
,
|
| 977 |
+
y
|
| 978 |
+
i
|
| 979 |
+
p
|
| 980 |
+
|
| 981 |
+
o
|
| 982 |
+
|
| 983 |
+
s
|
| 984 |
+
=
|
| 985 |
+
s
|
| 986 |
+
1
|
| 987 |
+
⊕
|
| 988 |
+
…
|
| 989 |
+
⊕
|
| 990 |
+
s
|
| 991 |
+
i
|
| 992 |
+
−
|
| 993 |
+
1
|
| 994 |
+
⊕
|
| 995 |
+
s
|
| 996 |
+
i
|
| 997 |
+
p
|
| 998 |
+
|
| 999 |
+
o
|
| 1000 |
+
|
| 1001 |
+
s
|
| 1002 |
+
;
|
| 1003 |
+
y
|
| 1004 |
+
i
|
| 1005 |
+
n
|
| 1006 |
+
|
| 1007 |
+
e
|
| 1008 |
+
|
| 1009 |
+
g
|
| 1010 |
+
=
|
| 1011 |
+
s
|
| 1012 |
+
1
|
| 1013 |
+
⊕
|
| 1014 |
+
…
|
| 1015 |
+
⊕
|
| 1016 |
+
s
|
| 1017 |
+
i
|
| 1018 |
+
−
|
| 1019 |
+
1
|
| 1020 |
+
⊕
|
| 1021 |
+
s
|
| 1022 |
+
i
|
| 1023 |
+
n
|
| 1024 |
+
|
| 1025 |
+
e
|
| 1026 |
+
|
| 1027 |
+
g
|
| 1028 |
+
\displaystyle\text{when $i$ is not final answer step},y_{i}^{pos}=s_{1}\oplus...\oplus s_{i-1}\oplus s_{i}^{pos};y_{i}^{neg}=s_{1}\oplus...\oplus s_{i-1}\oplus s_{i}^{neg}\vskip-4.30554pt
|
| 1029 |
+
(5)
|
| 1030 |
+
Here,
|
| 1031 |
+
r
|
| 1032 |
+
θ
|
| 1033 |
+
|
| 1034 |
+
(
|
| 1035 |
+
x
|
| 1036 |
+
,
|
| 1037 |
+
y
|
| 1038 |
+
i
|
| 1039 |
+
)
|
| 1040 |
+
r_{\theta}(x,y_{i})
|
| 1041 |
+
denotes the output of the PPM, where
|
| 1042 |
+
x
|
| 1043 |
+
x
|
| 1044 |
+
is the problem and
|
| 1045 |
+
y
|
| 1046 |
+
y
|
| 1047 |
+
is the trajectory from the first step to the
|
| 1048 |
+
i
|
| 1049 |
+
t
|
| 1050 |
+
|
| 1051 |
+
h
|
| 1052 |
+
i^{th}
|
| 1053 |
+
step.
|
| 1054 |
+
3.4
|
| 1055 |
+
Self-Evolved Deep Thinking
|
| 1056 |
+
3.4.1
|
| 1057 |
+
Training with Step-by-Step Verified Reasoning Trajectory
|
| 1058 |
+
Math Problems Collection
|
| 1059 |
+
. We collect a large dataset of 747k math word problems with final answer ground-truth labels, primarily from NuminaMath
|
| 1060 |
+
(Jia LI and Polu,
|
| 1061 |
+
2024a
|
| 1062 |
+
)
|
| 1063 |
+
and MetaMath
|
| 1064 |
+
(Yu et al.,
|
| 1065 |
+
2023b
|
| 1066 |
+
)
|
| 1067 |
+
. Notably, only competition-level problems (e.g., Olympiads and AIME/AMC) from NuminaMath are included, as we observe that grade-school-level problems do not significantly improve LLM complex math reasoning. To augment the limited competition-level problems, we follow
|
| 1068 |
+
(Li et al.,
|
| 1069 |
+
2024
|
| 1070 |
+
)
|
| 1071 |
+
and use GPT-4 to synthesize new problems based on the seed problems in 7.5k MATH train set and 3.6k AMC-AIME training split. However, GPT-4 often generated unsolvable problems or incorrect solutions for challenging seed problems. To filter these, we prompt GPT-4 to generate 10 solutions per problem, retaining only those with at least 3 consistent solutions.
|
| 1072 |
+
Reasoning Trajectories Collection
|
| 1073 |
+
. Instead of using the original solutions in the 747k math dataset, we conduct extensive MCTS rollouts (Sec.
|
| 1074 |
+
3.2
|
| 1075 |
+
) to generate higher-quality step-by-step verified reasoning trajectories. In each self-evolution round, we perform 16 rollouts per math problem, which leads to 16 reasoning trajectories. Problems are then categories by difficulty based on the correct ratio of the generated trajectories:
|
| 1076 |
+
easy
|
| 1077 |
+
(all solutions are correct),
|
| 1078 |
+
medium
|
| 1079 |
+
(a mix of correct and incorrect solutions) and
|
| 1080 |
+
hard
|
| 1081 |
+
(all solutions are incorrect). For
|
| 1082 |
+
hard
|
| 1083 |
+
problems with no correct trajectories, an additional MCTS with 16 rollouts is performed. After that, all step-by-step trajectories and their annotated Q-values are collected and filtered to train the policy SLM and process preference model.
|
| 1084 |
+
Supervised Fine-tuning the Policy SLM
|
| 1085 |
+
. Through extensive experiments, we find that selecting high-quality reasoning trajectories is the key for fine-tuning a frontier math LLM. While methods such as GPT-distillation and Best-of-N can include low-quality or erroneous intermediate steps, a more effective approach ensures that every step in the trajectory is of high quality. To achieve this, we use per-step Q-values to select optimal trajectories from MCTS rollouts. Specifically, for each math problem, we select the top-2 trajectories with the highest average Q-values among those leading to correct answers as SFT training data.
|
| 1086 |
+
Training PPM
|
| 1087 |
+
. The PPM is initialized from the fine-tuned policy model, with its next-token prediction head replaced by a scalar-value head consisting of a linear layer and a tanh function to constrain outputs to the range [-1, 1]. We filter out math problems where all solution trajectories are fully correct or incorrect. For problems with mixed outcomes, we select two positive and two negative examples for each step based on Q-values, which are used as preference pairs for training data.
|
| 1088 |
+
3.4.2
|
| 1089 |
+
Recipe for Self-Evolution
|
| 1090 |
+
Table 2:
|
| 1091 |
+
Percentage of the 747k math problems correctly solved in each round. Only problems have correct solutions are included in the training set. The first round uses DeepSeek-Coder-Instruct as the policy LLM, while later rounds use our fine-tuned 7B policy SLM.
|
| 1092 |
+
#
|
| 1093 |
+
models in MCTS
|
| 1094 |
+
GSM-level
|
| 1095 |
+
MATH-level
|
| 1096 |
+
Olympiad-level
|
| 1097 |
+
All
|
| 1098 |
+
Round 1
|
| 1099 |
+
DeepSeek-Coder-V2-Instruct
|
| 1100 |
+
96.61%
|
| 1101 |
+
67.36%
|
| 1102 |
+
20.99%
|
| 1103 |
+
60.17%
|
| 1104 |
+
Round 2
|
| 1105 |
+
policy SLM-r1
|
| 1106 |
+
97.88%
|
| 1107 |
+
67.40%
|
| 1108 |
+
56.04%
|
| 1109 |
+
66.60%
|
| 1110 |
+
Round 3
|
| 1111 |
+
policy SLM-r2, PPM-r2
|
| 1112 |
+
98.15%
|
| 1113 |
+
88.69%
|
| 1114 |
+
62.16%
|
| 1115 |
+
77.86%
|
| 1116 |
+
Round 4
|
| 1117 |
+
policy SLM-r3, PPM-r3
|
| 1118 |
+
98.15%
|
| 1119 |
+
94.53%
|
| 1120 |
+
80.58%
|
| 1121 |
+
90.25%
|
| 1122 |
+
Table 3:
|
| 1123 |
+
Pass@1 accuracy of the resulting policy SLM in each round, showing continuous improvement until surpassing the bootstrap model.
|
| 1124 |
+
Round#
|
| 1125 |
+
MATH
|
| 1126 |
+
AIME 2024
|
| 1127 |
+
AMC 2023
|
| 1128 |
+
Olympiad Bench
|
| 1129 |
+
College Math
|
| 1130 |
+
GSM8K
|
| 1131 |
+
GaokaoEn 2023
|
| 1132 |
+
DeepSeek-Coder-V2-Instruct
|
| 1133 |
+
(bootstrap model)
|
| 1134 |
+
75.3
|
| 1135 |
+
13.3
|
| 1136 |
+
57.5
|
| 1137 |
+
37.6
|
| 1138 |
+
46.2
|
| 1139 |
+
94.9
|
| 1140 |
+
64.7
|
| 1141 |
+
Base (Qwen2.5-Math-7B)
|
| 1142 |
+
58.8
|
| 1143 |
+
0.0
|
| 1144 |
+
22.5
|
| 1145 |
+
21.8
|
| 1146 |
+
41.6
|
| 1147 |
+
91.6
|
| 1148 |
+
51.7
|
| 1149 |
+
\hdashline
|
| 1150 |
+
policy SLM-r1
|
| 1151 |
+
69.6
|
| 1152 |
+
3.3
|
| 1153 |
+
30.0
|
| 1154 |
+
34.7
|
| 1155 |
+
44.5
|
| 1156 |
+
88.4
|
| 1157 |
+
57.4
|
| 1158 |
+
policy SLM-r2
|
| 1159 |
+
73.6
|
| 1160 |
+
10.0
|
| 1161 |
+
35.0
|
| 1162 |
+
39.0
|
| 1163 |
+
45.7
|
| 1164 |
+
89.1
|
| 1165 |
+
59.7
|
| 1166 |
+
policy SLM-r3
|
| 1167 |
+
75.8
|
| 1168 |
+
16.7
|
| 1169 |
+
45.0
|
| 1170 |
+
44.1
|
| 1171 |
+
49.6
|
| 1172 |
+
89.3
|
| 1173 |
+
62.8
|
| 1174 |
+
policy SLM-r4
|
| 1175 |
+
78.4
|
| 1176 |
+
26.7
|
| 1177 |
+
47.5
|
| 1178 |
+
47.1
|
| 1179 |
+
52.5
|
| 1180 |
+
89.7
|
| 1181 |
+
65.7
|
| 1182 |
+
Table 4:
|
| 1183 |
+
The quality of PPM consistently improves across rounds. The policy model has been fixed with policy SLM-r1 for a fair comparison.
|
| 1184 |
+
Round#
|
| 1185 |
+
MATH
|
| 1186 |
+
AIME 2024
|
| 1187 |
+
AMC 2023
|
| 1188 |
+
Olympiad Bench
|
| 1189 |
+
College Math
|
| 1190 |
+
GSM8K
|
| 1191 |
+
GaokaoEn 2023
|
| 1192 |
+
PPM-r1
|
| 1193 |
+
75.2
|
| 1194 |
+
10.0
|
| 1195 |
+
57.5
|
| 1196 |
+
35.7
|
| 1197 |
+
45.4
|
| 1198 |
+
90.9
|
| 1199 |
+
60.3
|
| 1200 |
+
PPM-r2
|
| 1201 |
+
84.1
|
| 1202 |
+
26.7
|
| 1203 |
+
75.0
|
| 1204 |
+
52.7
|
| 1205 |
+
54.2
|
| 1206 |
+
93.3
|
| 1207 |
+
73.0
|
| 1208 |
+
PPM-r3
|
| 1209 |
+
85.2
|
| 1210 |
+
33.3
|
| 1211 |
+
77.5
|
| 1212 |
+
59.5
|
| 1213 |
+
55.6
|
| 1214 |
+
93.9
|
| 1215 |
+
76.6
|
| 1216 |
+
PPM-r4
|
| 1217 |
+
87.0
|
| 1218 |
+
43.3
|
| 1219 |
+
77.5
|
| 1220 |
+
61.5
|
| 1221 |
+
56.8
|
| 1222 |
+
94.2
|
| 1223 |
+
77.8
|
| 1224 |
+
Due to the weaker capabilities of SLMs, we perform four rounds of MCTS deep thinking to progressively generate higher-quality data and expand the training set with more challenging math problems. Each round uses MCTS to generate step-by-step verified reasoning trajectories, which are then used to train the new policy SLM and PPM. The new models are then applied in next round to generate higher-quality training data. Fig.
|
| 1225 |
+
1
|
| 1226 |
+
(c) and Table
|
| 1227 |
+
2
|
| 1228 |
+
detail the models used for data generation in each round, along with the identifiers of the trained policy model and PPM. Next, we outline the details and specific improvements targeted in each round.
|
| 1229 |
+
Round 1: Bootstrapping an initial strong policy SLM-r1
|
| 1230 |
+
. To enable SLMs to self-generate reasonably good training data, we perform a bootstrap round to fine-tune an initial strong policy model, denoted as SLM-r1.
|
| 1231 |
+
As shown in Table
|
| 1232 |
+
2
|
| 1233 |
+
, we run MCTS with DeepSeek-Coder-V2-Instruct (236B) to collect the SFT data. With no available reward model in this round, we use terminal-guided annotation for Q-values and limit MCTS to 8 rollouts for efficiency. For correct solutions, the top-2 trajectories with the highest average Q-values are selected as SFT data. We also train PPM-r1, but the limited rollouts yields unreliable Q-values, affecting the effectiveness of PPM-r1 ( Table
|
| 1234 |
+
4
|
| 1235 |
+
).
|
| 1236 |
+
Round 2: Training a reliable PPM-r2
|
| 1237 |
+
. In this round, with the policy model updated to the 7B SLM-r1, we conduct extensive MCTS rollouts for more reliable Q-value annotation and train the first reliable reward model, PPM-r2. Specifically, we perform 16 MCTS rollouts per problem. The resulting step-by-step verified reasoning trajectories show significant improvements in both quality and Q-value precision. As shown in Table
|
| 1238 |
+
4
|
| 1239 |
+
, PPM-r2 is notably more effective than in the bootstrap round. Moreover, the policy SLM-r2 also continues to improve as expected (Table
|
| 1240 |
+
3
|
| 1241 |
+
).
|
| 1242 |
+
Round 3: PPM-augmented MCTS to significantly improve data quality
|
| 1243 |
+
. With the reliable PPM-r2, we perform PPM-augmented MCTS in this round to generate data, leading to significantly higher-quality trajectories that cover more math and Olympiad-level problems in the training set (Table
|
| 1244 |
+
2
|
| 1245 |
+
). The generated reasoning trajectories and self-annotated Q-values are then used to train the new policy SLM-r3 and PPM-r3, both of which show significant improvements.
|
| 1246 |
+
Round 4: Solving challenging math problems
|
| 1247 |
+
. After the third round, while grade school and MATH problems achieve high success rates, only 62.16% of Olympiad-level problems are included in the training set. This is
|
| 1248 |
+
NOT
|
| 1249 |
+
solely due to weak reasoning abilities in our SLMs, as many Olympiad problems remain unsolved by GPT-4 or o1. To improve coverage, we adopt a straightforward strategy. For unsolved problems after 16 MCTS rollouts, we perform an additional 64 rollouts, and if needed, increase to 128. We also conduct multiple MCTS tree expansions with different random seeds. This boosts the success rate of Olympiad-level problems to 80.58%.
|
| 1250 |
+
After four rounds of self-evolution, 90.25% of the 747k math problems are successfully covered into the training set, as shown in Table
|
| 1251 |
+
2
|
| 1252 |
+
. Among the remaining unsolved problems, a significant portion consists of synthetic questions. We manually review a random sample of 20 problems and find that 19 are incorrectly labeled with wrong answers. Based on this, we conclude that the remaining unsolved problems are of low quality and thus terminate the self-evolution at round 4.
|
| 1253 |
+
4
|
| 1254 |
+
Evaluation
|
| 1255 |
+
4.1
|
| 1256 |
+
Setup
|
| 1257 |
+
Evaluation Datasets
|
| 1258 |
+
. We evaluate
|
| 1259 |
+
\sysname
|
| 1260 |
+
on diverse mathematical benchmarks. In addition to the widely-used GSM8K
|
| 1261 |
+
(Cobbe et al.,
|
| 1262 |
+
2021
|
| 1263 |
+
)
|
| 1264 |
+
, we include challenging benchmarks from multiple domains:
|
| 1265 |
+
(i)
|
| 1266 |
+
competition and Olympiad-level benchmarks, such as MATH-500
|
| 1267 |
+
(Lightman et al.,
|
| 1268 |
+
2023
|
| 1269 |
+
)
|
| 1270 |
+
, AIME 2024
|
| 1271 |
+
(AI-MO,
|
| 1272 |
+
2024a
|
| 1273 |
+
)
|
| 1274 |
+
, AMC 2023
|
| 1275 |
+
(AI-MO,
|
| 1276 |
+
2024b
|
| 1277 |
+
)
|
| 1278 |
+
and Olympiad Bench
|
| 1279 |
+
(He et al.,
|
| 1280 |
+
2024
|
| 1281 |
+
)
|
| 1282 |
+
. Specifically, AIME is the exams designed to challenge the brightest high school math students in American, with the 2024 dataset comprising 30 problems from AIME I and II exams;
|
| 1283 |
+
(ii)
|
| 1284 |
+
college-level math problems from College Math
|
| 1285 |
+
(Tang et al.,
|
| 1286 |
+
2024
|
| 1287 |
+
)
|
| 1288 |
+
and
|
| 1289 |
+
(iii)
|
| 1290 |
+
out-of-domain math benchmark: GaoKao (Chinese
|
| 1291 |
+
College Entrance Exam) En 2023
|
| 1292 |
+
(Liao et al.,
|
| 1293 |
+
2024
|
| 1294 |
+
)
|
| 1295 |
+
.
|
| 1296 |
+
Base Models and Setup
|
| 1297 |
+
.
|
| 1298 |
+
\sysname
|
| 1299 |
+
is a general approach applicable to various LLMs. To show its effectiveness and generalizability, we use SLMs of different sizes as the base policy models:
|
| 1300 |
+
Qwen2.5-Math-1.5B
|
| 1301 |
+
(Qwen,
|
| 1302 |
+
2024b
|
| 1303 |
+
)
|
| 1304 |
+
, Phi3-mini-Instruct (3B)
|
| 1305 |
+
(Microsoft,
|
| 1306 |
+
2024
|
| 1307 |
+
; Abdin et al.,
|
| 1308 |
+
2024
|
| 1309 |
+
)
|
| 1310 |
+
, Qwen2-Math-7B
|
| 1311 |
+
(Qwen,
|
| 1312 |
+
2024a
|
| 1313 |
+
)
|
| 1314 |
+
and Qwen2.5-Math-7B
|
| 1315 |
+
(Qwen,
|
| 1316 |
+
2024c
|
| 1317 |
+
)
|
| 1318 |
+
. Among these, Phi3-mini-Instruct is a general-purpose SLM without specialization in math reasoning.
|
| 1319 |
+
Due to limited GPU resources, we performed 4 rounds of self-evolution exclusively on Qwen2.5-Math-7B, yielding 4 evolved policy SLMs (Table
|
| 1320 |
+
3
|
| 1321 |
+
) and 4 PPMs (Table
|
| 1322 |
+
4
|
| 1323 |
+
). For the other 3 policy LLMs, we fine-tune them using step-by-step verified trajectories generated from Qwen2.5-Math-7B’s 4th round. The final PPM from this round is then used as the reward model for the 3 policy SLMs.
|
| 1324 |
+
Baselines
|
| 1325 |
+
.
|
| 1326 |
+
\sysname
|
| 1327 |
+
is a System 2 method. We compare it against three strong baselines representing both System 1 and System 2 approaches:
|
| 1328 |
+
(i)
|
| 1329 |
+
Frontier LLMs
|
| 1330 |
+
, including GPT-4o, the latest Claude, OpenAI o1-preview and o1-mini.
|
| 1331 |
+
We measure their accuracy on AMC 2023, Olympiad Bench, College Math, Gaokao and GSM8K, with accuracy numbers for other benchmarks are taken from public technical reports
|
| 1332 |
+
(Team,
|
| 1333 |
+
2024a
|
| 1334 |
+
)
|
| 1335 |
+
.
|
| 1336 |
+
(ii)
|
| 1337 |
+
Open-sourced superior reasoning models
|
| 1338 |
+
, including DeepSeek-Coder-v2-Instruct, Mathstral
|
| 1339 |
+
(Team,
|
| 1340 |
+
2024b
|
| 1341 |
+
)
|
| 1342 |
+
, NuminaMath-72B
|
| 1343 |
+
(Jia LI and Polu,
|
| 1344 |
+
2024a
|
| 1345 |
+
)
|
| 1346 |
+
, and LLaMA3.1
|
| 1347 |
+
(Dubey et al.,
|
| 1348 |
+
2024
|
| 1349 |
+
)
|
| 1350 |
+
, which represent the current mainstream System 1 approaches for improving LLM math reasoning.
|
| 1351 |
+
(iii)
|
| 1352 |
+
Both System 1 and System 2 performance of the base models trained from the original models teams
|
| 1353 |
+
, including Instruct versions (e.g., Qwen2.5-Math-7B-Instruct) and Best-of-N (e.g., Qwen2.5-Math-72B-Instruct+Qwen2.5-Math-RM-72B). Notably, the reward model used for the three Qwen base models is a 72B ORM, significantly larger than our 7B PPM.
|
| 1354 |
+
Evaluation Metric
|
| 1355 |
+
. We report Pass@1 accuracy for all baselines. For System 2 baselines, we use default evaluation settings, such as default thinking time for o1-mini and o1-preview. For Qwen models with Best-of-N, we re-evaluate MATH-500, AIME/AMC accuracy; other benchmarks results are from their technical reports. For a fair comparison,
|
| 1356 |
+
\sysname
|
| 1357 |
+
run MCTS to generate the same number of solutions as Qwen. Specifically, for AIME/AMC, we generate 16 trajectories for AIME/AMC and 8 for other benchmarks, using PPM to select the best solution. We also report performance with increased test-time computation using 64 trajectories, denoted as
|
| 1358 |
+
\sysname
|
| 1359 |
+
64
|
| 1360 |
+
.
|
| 1361 |
+
Table 5:
|
| 1362 |
+
The results of
|
| 1363 |
+
\sysname
|
| 1364 |
+
and other frontier LLMs on the most challenging math benchmarks.
|
| 1365 |
+
\sysname
|
| 1366 |
+
64
|
| 1367 |
+
shows the Pass@1 accuracy achieved when sampling 64 trajectories.
|
| 1368 |
+
Competition and College Level
|
| 1369 |
+
OOD
|
| 1370 |
+
Model
|
| 1371 |
+
Method
|
| 1372 |
+
MATH
|
| 1373 |
+
AIME
|
| 1374 |
+
2024
|
| 1375 |
+
AMC
|
| 1376 |
+
2023
|
| 1377 |
+
Olympiad
|
| 1378 |
+
Bench
|
| 1379 |
+
College
|
| 1380 |
+
Math
|
| 1381 |
+
GSM8K
|
| 1382 |
+
Gaokao
|
| 1383 |
+
En 2023
|
| 1384 |
+
Frontier LLMs
|
| 1385 |
+
GPT-4o
|
| 1386 |
+
System 1
|
| 1387 |
+
76.6
|
| 1388 |
+
9.3
|
| 1389 |
+
47.5
|
| 1390 |
+
43.3
|
| 1391 |
+
48.5
|
| 1392 |
+
92.9
|
| 1393 |
+
67.5
|
| 1394 |
+
Claude3.5-Sonnet
|
| 1395 |
+
System 1
|
| 1396 |
+
78.3
|
| 1397 |
+
16.0
|
| 1398 |
+
-
|
| 1399 |
+
-
|
| 1400 |
+
-
|
| 1401 |
+
96.4
|
| 1402 |
+
-
|
| 1403 |
+
GPT-o1-preview
|
| 1404 |
+
-
|
| 1405 |
+
85.5
|
| 1406 |
+
44.6
|
| 1407 |
+
90.0
|
| 1408 |
+
-
|
| 1409 |
+
-
|
| 1410 |
+
-
|
| 1411 |
+
-
|
| 1412 |
+
GPT-o1-mini
|
| 1413 |
+
-
|
| 1414 |
+
90.0
|
| 1415 |
+
56.7
|
| 1416 |
+
95.0
|
| 1417 |
+
65.3
|
| 1418 |
+
57.8
|
| 1419 |
+
94.8
|
| 1420 |
+
78.4
|
| 1421 |
+
Open-Sourced Reasoning LLMs
|
| 1422 |
+
DeepSeek-Coder-V2-Instruct
|
| 1423 |
+
System 1
|
| 1424 |
+
75.3
|
| 1425 |
+
13.3
|
| 1426 |
+
57.5
|
| 1427 |
+
37.6
|
| 1428 |
+
46.2
|
| 1429 |
+
94.9
|
| 1430 |
+
64.7
|
| 1431 |
+
Mathstral-7B-v0.1
|
| 1432 |
+
System 1
|
| 1433 |
+
57.8
|
| 1434 |
+
0.0
|
| 1435 |
+
37.5
|
| 1436 |
+
21.5
|
| 1437 |
+
33.7
|
| 1438 |
+
84.9
|
| 1439 |
+
46.0
|
| 1440 |
+
NuminaMath-72B-CoT
|
| 1441 |
+
System 1
|
| 1442 |
+
64.0
|
| 1443 |
+
3.3
|
| 1444 |
+
70.0
|
| 1445 |
+
32.6
|
| 1446 |
+
39.7
|
| 1447 |
+
90.8
|
| 1448 |
+
58.4
|
| 1449 |
+
LLaMA3.1-8B-Instruct
|
| 1450 |
+
System 1
|
| 1451 |
+
51.4
|
| 1452 |
+
6.7
|
| 1453 |
+
25.0
|
| 1454 |
+
15.4
|
| 1455 |
+
33.8
|
| 1456 |
+
76.6
|
| 1457 |
+
38.4
|
| 1458 |
+
LLaMA3.1-70B-Instruct
|
| 1459 |
+
System 1
|
| 1460 |
+
65.4
|
| 1461 |
+
23.3
|
| 1462 |
+
50.0
|
| 1463 |
+
27.7
|
| 1464 |
+
42.5
|
| 1465 |
+
94.1
|
| 1466 |
+
54.0
|
| 1467 |
+
Qwen2.5-Math-72B-Instruct
|
| 1468 |
+
System 1
|
| 1469 |
+
85.6
|
| 1470 |
+
30.0
|
| 1471 |
+
70.0
|
| 1472 |
+
49.0
|
| 1473 |
+
49.5
|
| 1474 |
+
95.9
|
| 1475 |
+
71.9
|
| 1476 |
+
Qwen2.5-Math-72B-Instruct+72B ORM
|
| 1477 |
+
System 2
|
| 1478 |
+
85.8
|
| 1479 |
+
36.7
|
| 1480 |
+
72.5
|
| 1481 |
+
54.5
|
| 1482 |
+
50.6
|
| 1483 |
+
96.4
|
| 1484 |
+
76.9
|
| 1485 |
+
General Base Model: Phi3-mini-Instruct (3.8B)
|
| 1486 |
+
Phi3-mini-Instruct (base model)
|
| 1487 |
+
System 1
|
| 1488 |
+
41.4
|
| 1489 |
+
3.33
|
| 1490 |
+
7.5
|
| 1491 |
+
12.3
|
| 1492 |
+
33.1
|
| 1493 |
+
85.7
|
| 1494 |
+
37.1
|
| 1495 |
+
\sysname
|
| 1496 |
+
(3.8B SLM+7B PPM)
|
| 1497 |
+
System 2
|
| 1498 |
+
85.4
|
| 1499 |
+
40.0
|
| 1500 |
+
77.5
|
| 1501 |
+
59.3
|
| 1502 |
+
58.0
|
| 1503 |
+
94.5
|
| 1504 |
+
77.1
|
| 1505 |
+
\sysname
|
| 1506 |
+
64
|
| 1507 |
+
(3.8B SLM+7B PPM)
|
| 1508 |
+
System 2
|
| 1509 |
+
86.4
|
| 1510 |
+
43.3
|
| 1511 |
+
80.0
|
| 1512 |
+
60.3
|
| 1513 |
+
59.1
|
| 1514 |
+
94.7
|
| 1515 |
+
77.7
|
| 1516 |
+
Math-Specialized Base Model: Qwen2.5-Math-1.5B
|
| 1517 |
+
Qwen2.5-Math-1.5B (base model)
|
| 1518 |
+
System 1
|
| 1519 |
+
51.2
|
| 1520 |
+
0.0
|
| 1521 |
+
22.5
|
| 1522 |
+
16.7
|
| 1523 |
+
38.4
|
| 1524 |
+
74.6
|
| 1525 |
+
46.5
|
| 1526 |
+
Qwen2.5-Math-1.5B-Instruct
|
| 1527 |
+
System 1
|
| 1528 |
+
60.0
|
| 1529 |
+
10.0
|
| 1530 |
+
60.0
|
| 1531 |
+
38.1
|
| 1532 |
+
47.7
|
| 1533 |
+
84.8
|
| 1534 |
+
65.5
|
| 1535 |
+
Qwen2.5-Math-1.5B-Instruct+72B ORM
|
| 1536 |
+
System 2
|
| 1537 |
+
83.4
|
| 1538 |
+
20.0
|
| 1539 |
+
72.5
|
| 1540 |
+
47.3
|
| 1541 |
+
50.2
|
| 1542 |
+
94.1
|
| 1543 |
+
73.0
|
| 1544 |
+
\sysname
|
| 1545 |
+
(1.5B SLM+7B PPM)
|
| 1546 |
+
System 2
|
| 1547 |
+
87.8
|
| 1548 |
+
46.7
|
| 1549 |
+
80.0
|
| 1550 |
+
63.5
|
| 1551 |
+
59.0
|
| 1552 |
+
94.3
|
| 1553 |
+
77.7
|
| 1554 |
+
\sysname
|
| 1555 |
+
64
|
| 1556 |
+
(1.5B SLM+7B PPM)
|
| 1557 |
+
System 2
|
| 1558 |
+
88.6
|
| 1559 |
+
46.7
|
| 1560 |
+
85.0
|
| 1561 |
+
64.6
|
| 1562 |
+
59.3
|
| 1563 |
+
94.8
|
| 1564 |
+
79.5
|
| 1565 |
+
Math-Specialized Base Model: Qwen2-Math-7B
|
| 1566 |
+
Qwen2-Math-7B (base model)
|
| 1567 |
+
System 1
|
| 1568 |
+
53.4
|
| 1569 |
+
3.3
|
| 1570 |
+
25.0
|
| 1571 |
+
17.3
|
| 1572 |
+
39.4
|
| 1573 |
+
80.4
|
| 1574 |
+
47.3
|
| 1575 |
+
Qwen2-Math-7B-Instruct
|
| 1576 |
+
System 1
|
| 1577 |
+
73.2
|
| 1578 |
+
13.3
|
| 1579 |
+
62.5
|
| 1580 |
+
38.2
|
| 1581 |
+
45.9
|
| 1582 |
+
89.9
|
| 1583 |
+
62.1
|
| 1584 |
+
Qwen2-Math-7B-Instruct+72B ORM
|
| 1585 |
+
System 2
|
| 1586 |
+
83.4
|
| 1587 |
+
23.3
|
| 1588 |
+
62.5
|
| 1589 |
+
47.6
|
| 1590 |
+
47.9
|
| 1591 |
+
95.1
|
| 1592 |
+
71.9
|
| 1593 |
+
\sysname
|
| 1594 |
+
(7B SLM+7B PPM)
|
| 1595 |
+
System 2
|
| 1596 |
+
88.2
|
| 1597 |
+
43.3
|
| 1598 |
+
80.0
|
| 1599 |
+
63.1
|
| 1600 |
+
58.4
|
| 1601 |
+
94.6
|
| 1602 |
+
78.2
|
| 1603 |
+
\sysname
|
| 1604 |
+
64
|
| 1605 |
+
(7B SLM+7B PPM)
|
| 1606 |
+
System 2
|
| 1607 |
+
88.6
|
| 1608 |
+
46.7
|
| 1609 |
+
85.0
|
| 1610 |
+
63.4
|
| 1611 |
+
59.3
|
| 1612 |
+
94.8
|
| 1613 |
+
79.2
|
| 1614 |
+
Math-Specialized Base Model: Qwen2.5-Math-7B
|
| 1615 |
+
Qwen2.5-Math-7B (base model)
|
| 1616 |
+
System 1
|
| 1617 |
+
58.8
|
| 1618 |
+
0.0
|
| 1619 |
+
22.5
|
| 1620 |
+
21.8
|
| 1621 |
+
41.6
|
| 1622 |
+
91.6
|
| 1623 |
+
51.7
|
| 1624 |
+
Qwen2.5-Math-7B-Instruct
|
| 1625 |
+
System 1
|
| 1626 |
+
82.6
|
| 1627 |
+
6.0
|
| 1628 |
+
62.5
|
| 1629 |
+
41.6
|
| 1630 |
+
46.8
|
| 1631 |
+
95.2
|
| 1632 |
+
66.8
|
| 1633 |
+
Qwen2.5-Math-7B-Instruct+72B ORM
|
| 1634 |
+
System 2
|
| 1635 |
+
88.4
|
| 1636 |
+
26.7
|
| 1637 |
+
75.0
|
| 1638 |
+
49.9
|
| 1639 |
+
49.6
|
| 1640 |
+
97.9
|
| 1641 |
+
75.1
|
| 1642 |
+
\sysname
|
| 1643 |
+
(7B SLM+7B PPM)
|
| 1644 |
+
System 2
|
| 1645 |
+
89.4
|
| 1646 |
+
50.0
|
| 1647 |
+
87.5
|
| 1648 |
+
65.3
|
| 1649 |
+
59.0
|
| 1650 |
+
95.0
|
| 1651 |
+
80.5
|
| 1652 |
+
\sysname
|
| 1653 |
+
64
|
| 1654 |
+
(7B SLM+7B PPM)
|
| 1655 |
+
System 2
|
| 1656 |
+
90.0
|
| 1657 |
+
53.3
|
| 1658 |
+
87.5
|
| 1659 |
+
65.6
|
| 1660 |
+
60.5
|
| 1661 |
+
95.2
|
| 1662 |
+
81.3
|
| 1663 |
+
4.2
|
| 1664 |
+
Main Results
|
| 1665 |
+
Results on diverse challenging math benchmarks
|
| 1666 |
+
. Table
|
| 1667 |
+
5
|
| 1668 |
+
shows the results of
|
| 1669 |
+
\sysname
|
| 1670 |
+
with comparing to state-of-the-art reasoning models. We highlight three key observations:
|
| 1671 |
+
(1)
|
| 1672 |
+
\sysname
|
| 1673 |
+
significantly improves SLMs math reasoning capabilities, achieving performance comparable to or surpassing OpenAI o1 with substantially smaller model size (1.5B-7B). For example, Qwen2.5-Math-7B, originally at 58.8% accuracy on MATH, improved dramatically to 90.0% with
|
| 1674 |
+
\sysname
|
| 1675 |
+
, outperforming o1-preview and Claude 3.5 Sonnet while matching o1-mini. On the College Math benchmark,
|
| 1676 |
+
\sysname
|
| 1677 |
+
exceeds o1-mini by 2.7%. On AIME 2024,
|
| 1678 |
+
\sysname
|
| 1679 |
+
scored 53.3%, ranking just below o1-mini, with the 7B model solving 8/15 problems in both AIME I and II, placing in the top 20% of the brightest high school math students.
|
| 1680 |
+
Notably, 8 of the unsolved problems were geometry-based, requiring visual understanding, a capability
|
| 1681 |
+
\sysname
|
| 1682 |
+
currently does not support.
|
| 1683 |
+
(2)
|
| 1684 |
+
Despite using smaller policy models (1.5B-7B) and reward models (7B),
|
| 1685 |
+
\sysname
|
| 1686 |
+
significantly outperforms state-of-the-art System 2 baselines. Compared to Qwen Best-of-N baselines, which use the same base models (Qwen2-Math-7B, Qwen2.5-Math-1.5B/7B) but a 10
|
| 1687 |
+
×
|
| 1688 |
+
\times
|
| 1689 |
+
larger reward model (Qwen2.5-Math-RM-72B),
|
| 1690 |
+
\sysname
|
| 1691 |
+
consistently improves the reasoning accuracy of all base models to state-of-the-art levels. Even against Best-of-N with a 10
|
| 1692 |
+
×
|
| 1693 |
+
\times
|
| 1694 |
+
larger Qwen2.5-Math-72B-Instruct policy model,
|
| 1695 |
+
\sysname
|
| 1696 |
+
surpasses it on all benchmarks except GSM8K, using the same number of sampled solutions.
|
| 1697 |
+
(3)
|
| 1698 |
+
Beyond well-known benchmarks like MATH, GSM8K, and AIME, which may risk over-optimization,
|
| 1699 |
+
\sysname
|
| 1700 |
+
shows strong generalizability on other challenging math benchmarks, including Olympiad Bench, College Math, and the Chinese College Entrance Math Exam (Gaokao), setting new state-of-the-art scores. As discussed in Sec.
|
| 1701 |
+
3.4
|
| 1702 |
+
, our training set is primarily sourced from public datasets, with no specific optimizations for these benchmarks.
|
| 1703 |
+
Figure 3:
|
| 1704 |
+
Reasoning performance under scaling up the test-time compute.
|
| 1705 |
+
Scaling up test-time computation
|
| 1706 |
+
.
|
| 1707 |
+
\sysname
|
| 1708 |
+
uses MCTS to augment the policy model, searching solutions guided by the PPM. By increasing test-time computation, it explores more trajectories, potentially improving performance.
|
| 1709 |
+
In Fig.
|
| 1710 |
+
3
|
| 1711 |
+
, we show the impact of test-time compute scaling by comparing the accuracy of the official Qwen Best-of-N across different numbers of sampled trajectories on four challenging math benchmarks. Sampling only one trajectory corresponds to the policy LLM’s Pass@1 accuracy, indicating a fallback to System 1 reasoning. We highlight two key observations:
|
| 1712 |
+
(1)
|
| 1713 |
+
With only 4 trajectories,
|
| 1714 |
+
\sysname
|
| 1715 |
+
significantly outperforms Best-of-N baselines, exceeding o1-preview and approaching o1-mini, demonstrating its effectiveness.
|
| 1716 |
+
(2)
|
| 1717 |
+
Scaling test-time compute improves reasoning accuracy across all benchmarks, though with varying trends. On Math, AIME, and Olympiad Bench,
|
| 1718 |
+
\sysname
|
| 1719 |
+
shows saturation or slow improvement at 64 trajectories, while on College Math, performance continues to improve steadily.
|
| 1720 |
+
4.3
|
| 1721 |
+
Ablation Study and Analysis
|
| 1722 |
+
We ablate the effectiveness of our three innovations. For System 2-style inference, Pass@1 accuracy is measured with 16 trajectories for AIME and AMC, and 8 for other benchmarks.
|
| 1723 |
+
Table 6:
|
| 1724 |
+
The continuously improved math reasoning capabilities through
|
| 1725 |
+
\sysname
|
| 1726 |
+
self-evolved deep thinking. Starting from round 2, the 7B base model powered by
|
| 1727 |
+
\sysname
|
| 1728 |
+
surpasses GPT-4o.
|
| 1729 |
+
Round#
|
| 1730 |
+
MATH
|
| 1731 |
+
AIME 2024
|
| 1732 |
+
AMC 2023
|
| 1733 |
+
Olympiad Bench
|
| 1734 |
+
College Math
|
| 1735 |
+
GSM8K
|
| 1736 |
+
GaokaoEn 2023
|
| 1737 |
+
GPT-4o
|
| 1738 |
+
76.6
|
| 1739 |
+
9.3
|
| 1740 |
+
47.5
|
| 1741 |
+
43.3
|
| 1742 |
+
48.5
|
| 1743 |
+
92.9
|
| 1744 |
+
67.5
|
| 1745 |
+
Base 7B model
|
| 1746 |
+
58.8
|
| 1747 |
+
0.0
|
| 1748 |
+
22.5
|
| 1749 |
+
21.8
|
| 1750 |
+
41.6
|
| 1751 |
+
91.6
|
| 1752 |
+
51.7
|
| 1753 |
+
\sysname
|
| 1754 |
+
Round 1
|
| 1755 |
+
75.2
|
| 1756 |
+
10.0
|
| 1757 |
+
57.5
|
| 1758 |
+
35.7
|
| 1759 |
+
45.4
|
| 1760 |
+
90.9
|
| 1761 |
+
60.3
|
| 1762 |
+
\sysname
|
| 1763 |
+
Round 2
|
| 1764 |
+
86.6
|
| 1765 |
+
43.3
|
| 1766 |
+
75.0
|
| 1767 |
+
59.4
|
| 1768 |
+
55.6
|
| 1769 |
+
94.0
|
| 1770 |
+
76.4
|
| 1771 |
+
\sysname
|
| 1772 |
+
Round 3
|
| 1773 |
+
87.0
|
| 1774 |
+
46.7
|
| 1775 |
+
80.0
|
| 1776 |
+
61.6
|
| 1777 |
+
56.5
|
| 1778 |
+
94.2
|
| 1779 |
+
77.1
|
| 1780 |
+
\sysname
|
| 1781 |
+
Round 4
|
| 1782 |
+
89.4
|
| 1783 |
+
50.0
|
| 1784 |
+
87.5
|
| 1785 |
+
65.3
|
| 1786 |
+
59.0
|
| 1787 |
+
95.0
|
| 1788 |
+
80.5
|
| 1789 |
+
The effectiveness of self-evolution
|
| 1790 |
+
. The impressive results in Table
|
| 1791 |
+
5
|
| 1792 |
+
are achieved after 4 rounds of
|
| 1793 |
+
\sysname
|
| 1794 |
+
self-evolved deep thinking. Table
|
| 1795 |
+
6
|
| 1796 |
+
shows the math reasoning performance in each round, demonstrating a continuous improvement in accuracy.
|
| 1797 |
+
In round 1, the main improvement comes from applying SFT to the base model. Round 2 brings a significant boost with the application of a stronger PPM in MCTS, which unlocks the full potential of System 2 deep reasoning. Notably, starting from round 2,
|
| 1798 |
+
\sysname
|
| 1799 |
+
outperforms GPT-4o. Rounds 3 and 4 show further improvements, driven by stronger System 2 reasoning through better policy SLMs and PPMs.
|
| 1800 |
+
The effectiveness of step-by-step verified reasoning trajectory
|
| 1801 |
+
.
|
| 1802 |
+
\sysname
|
| 1803 |
+
generates step-by-step verified reasoning trajectories, which eliminate error intermediate steps and further expand training set with more challenging problems. To evaluate its effectiveness, we use the data generated from round 4 as SFT training data and compare it against
|
| 1804 |
+
three strong baselines:
|
| 1805 |
+
(i)
|
| 1806 |
+
GPT-distillation, which includes open-sourced CoT solutions synthesized using GPT-4, such as MetaMath
|
| 1807 |
+
(Yu et al.,
|
| 1808 |
+
2023b
|
| 1809 |
+
)
|
| 1810 |
+
, NuminaMath-CoT
|
| 1811 |
+
(Jia LI and Polu,
|
| 1812 |
+
2024b
|
| 1813 |
+
)
|
| 1814 |
+
;
|
| 1815 |
+
(ii)
|
| 1816 |
+
Random sampling from self-generation,
|
| 1817 |
+
which use the same policy model (i.e., policy SLM-r3) to randomly generate trajectories;
|
| 1818 |
+
(iii)
|
| 1819 |
+
Rejection sampling, where 32 trajectories are randomly sampled from the policy model, with high-quality solutions ranked by our trained ORM (appendix
|
| 1820 |
+
A.1
|
| 1821 |
+
). For fairness, we select two correct trajectories for each math problem in baseline (ii) and (iii). All SFT experiments use the same training recipe.
|
| 1822 |
+
Table 7:
|
| 1823 |
+
Ablation study on the effectiveness of our step-by-step verified reasoning trajectories as the SFT dataset. We report the SFT accuracy of Qwen2.5-Math-7B fine-tuned with different datasets.
|
| 1824 |
+
Dataset
|
| 1825 |
+
MATH
|
| 1826 |
+
AIME
|
| 1827 |
+
AMC
|
| 1828 |
+
Olympiad Bench
|
| 1829 |
+
College Math
|
| 1830 |
+
GSM8K
|
| 1831 |
+
GaokaoEn 2023
|
| 1832 |
+
GPT-4o
|
| 1833 |
+
-
|
| 1834 |
+
76.6
|
| 1835 |
+
9.3
|
| 1836 |
+
47.5
|
| 1837 |
+
43.3
|
| 1838 |
+
48.5
|
| 1839 |
+
92.9
|
| 1840 |
+
67.5
|
| 1841 |
+
GPT4-distillation
|
| 1842 |
+
(Open-sourced)
|
| 1843 |
+
MetaMath
|
| 1844 |
+
55.2
|
| 1845 |
+
3.33
|
| 1846 |
+
32.5
|
| 1847 |
+
19.1
|
| 1848 |
+
39.2
|
| 1849 |
+
85.1
|
| 1850 |
+
43.6
|
| 1851 |
+
NuminaMath-CoT
|
| 1852 |
+
69.6
|
| 1853 |
+
10.0
|
| 1854 |
+
50.0
|
| 1855 |
+
37.2
|
| 1856 |
+
43.4
|
| 1857 |
+
89.8
|
| 1858 |
+
59.5
|
| 1859 |
+
Self-generation
|
| 1860 |
+
by policy SLM-r3
|
| 1861 |
+
Random sample
|
| 1862 |
+
72.4
|
| 1863 |
+
10.0
|
| 1864 |
+
45.0
|
| 1865 |
+
41.0
|
| 1866 |
+
48.0
|
| 1867 |
+
87.5
|
| 1868 |
+
57.1
|
| 1869 |
+
Rejection sampling
|
| 1870 |
+
73.4
|
| 1871 |
+
13.3
|
| 1872 |
+
47.5
|
| 1873 |
+
44.7
|
| 1874 |
+
50.8
|
| 1875 |
+
89.3
|
| 1876 |
+
61.7
|
| 1877 |
+
Step-by-step verified (ours)
|
| 1878 |
+
78.4
|
| 1879 |
+
26.7
|
| 1880 |
+
47.5
|
| 1881 |
+
47.1
|
| 1882 |
+
52.5
|
| 1883 |
+
89.7
|
| 1884 |
+
65.7
|
| 1885 |
+
Table
|
| 1886 |
+
7
|
| 1887 |
+
shows the math reasoning accuracy of Qwen2.5-Math-7B fine-tuned on different datasets. We highlight two observations:
|
| 1888 |
+
(i)
|
| 1889 |
+
Fine-tuning with our step-by-step verified trajectories significantly outperforms all other baselines. This is primarily due to our PPM-augmented MCTS for code-augmented CoT synthesis, which provides denser verification during math solution generation. It proves more effective than both random sampling, which lacks verification, and rejection sampling, where ORM provides only sparse verification.
|
| 1890 |
+
(ii)
|
| 1891 |
+
Even randomly sampled code-augmented CoT solutions from our SLM yields comparable or better performance than GPT-4 synthesized NuminaMath and MetaMath datasets.
|
| 1892 |
+
This indicates that our policy SLMs, after rounds of self-evolution, can generate high-quality math solutions. These results demonstrates the huge potential of our method to self-generate higher-quality reasoning data without relying on advanced LLM distillation.
|
| 1893 |
+
The effectiveness of PPM
|
| 1894 |
+
. We train both a strong ORM and Q-value score-based PRM (PQM) for comparison. To ensure a fair evaluation, we use the highest-quality training data: the step-by-step verified trajectories generated in round 4, with selected math problems matching those used for PPM training. Similar to PPM, we use step-level Q-values as to select positive and negative trajectories for each math problem.
|
| 1895 |
+
The ORM is trained using a pairwise ranking loss
|
| 1896 |
+
(Ouyang et al.,
|
| 1897 |
+
2022
|
| 1898 |
+
)
|
| 1899 |
+
, while the PQM follows
|
| 1900 |
+
(Chen et al.,
|
| 1901 |
+
2024
|
| 1902 |
+
; Zhang et al.,
|
| 1903 |
+
2024a
|
| 1904 |
+
)
|
| 1905 |
+
to use Q-values as reward labels and optimize with MSE loss. Detailed training settings are provided in Appendix
|
| 1906 |
+
A.1
|
| 1907 |
+
.
|
| 1908 |
+
Table 8:
|
| 1909 |
+
Ablation study on the reward model. Process reward models (PQM and PPM) outperform ORM, with PPM pushing the frontier of math reasoning capabilities.
|
| 1910 |
+
RM
|
| 1911 |
+
Inference
|
| 1912 |
+
MATH
|
| 1913 |
+
AIME
|
| 1914 |
+
AMC
|
| 1915 |
+
Olympiad Bench
|
| 1916 |
+
College Math
|
| 1917 |
+
GSM8K
|
| 1918 |
+
GaokaoEn
|
| 1919 |
+
o1-mini
|
| 1920 |
+
-
|
| 1921 |
+
90.0
|
| 1922 |
+
56.7
|
| 1923 |
+
95.0
|
| 1924 |
+
65.3
|
| 1925 |
+
55.6
|
| 1926 |
+
94.8
|
| 1927 |
+
78.6
|
| 1928 |
+
ORM
|
| 1929 |
+
Best-of-N
|
| 1930 |
+
82.6
|
| 1931 |
+
26.7
|
| 1932 |
+
65.0
|
| 1933 |
+
55.1
|
| 1934 |
+
55.5
|
| 1935 |
+
92.3
|
| 1936 |
+
72.5
|
| 1937 |
+
PQM
|
| 1938 |
+
MCTS
|
| 1939 |
+
88.2
|
| 1940 |
+
46.7
|
| 1941 |
+
85.0
|
| 1942 |
+
62.9
|
| 1943 |
+
57.6
|
| 1944 |
+
94.6
|
| 1945 |
+
79.5
|
| 1946 |
+
PPM
|
| 1947 |
+
MCTS
|
| 1948 |
+
89.4
|
| 1949 |
+
50.0
|
| 1950 |
+
87.5
|
| 1951 |
+
65.3
|
| 1952 |
+
59.0
|
| 1953 |
+
95.0
|
| 1954 |
+
80.5
|
| 1955 |
+
Table
|
| 1956 |
+
8
|
| 1957 |
+
compares the performance of ORM, PQM, and PPM for System 2 reasoning using our final round policy model. ORM provides reward signals only at the end of problem solving, so we use the Best-of-N method, while PRM and PPM leverage MCTS-driven search. As shown in Table
|
| 1958 |
+
8
|
| 1959 |
+
, both PQM and PPM outperform ORM by providing denser step-level reward signals, leading to higher accuracy on complex math reasoning tasks. However, PQM struggles on more challenging benchmarks, such as MATH and Olympiad Bench, due to the inherent imprecision of Q-values.
|
| 1960 |
+
In contrast, PPM constructs step-level preference data for training, enabling our 7B policy model to achieve comparable or superior performance to o1-mini across all benchmarks.
|
| 1961 |
+
5
|
| 1962 |
+
Findings and Discussions
|
| 1963 |
+
Figure 4:
|
| 1964 |
+
An example of intrinsic self-reflection during
|
| 1965 |
+
\sysname
|
| 1966 |
+
deep thinking.
|
| 1967 |
+
The emergence of intrinsic self-reflection capability
|
| 1968 |
+
. A key breakthrough in OpenAI o1 is its intrinsic self-reflection capability. When the model makes an error, it recognizes the mistake and can self-correct with a correct answer
|
| 1969 |
+
(Noam Brown and Lightman,
|
| 1970 |
+
2024
|
| 1971 |
+
)
|
| 1972 |
+
. Yet it has consistently
|
| 1973 |
+
been found to be largely ineffective in open-sourced LLMs. The community has actively explored various approaches, including self-correction
|
| 1974 |
+
(Huang et al.,
|
| 1975 |
+
2023
|
| 1976 |
+
; Kumar et al.,
|
| 1977 |
+
2024
|
| 1978 |
+
)
|
| 1979 |
+
, self-reflection
|
| 1980 |
+
(Renze and Guven,
|
| 1981 |
+
2024
|
| 1982 |
+
; Shinn et al.,
|
| 1983 |
+
2024
|
| 1984 |
+
)
|
| 1985 |
+
, to explicitly train or prompt LLMs to develop such capability.
|
| 1986 |
+
In our experiments, we unexpectedly observe that our MCTS-driven deep thinking exhibits self-reflection during problem-solving. As shown in Fig.
|
| 1987 |
+
4
|
| 1988 |
+
, the model initially formalizes an equation using
|
| 1989 |
+
SymPy
|
| 1990 |
+
in the first three steps, which would lead to an incorrect answer (left branch). Interestingly, in the fourth step (right branch), the policy model recognizes the low quality of its earlier steps and refrains from continuing along the initial problem-solving path. Instead, it backtracks and resolves the problem using a new, simpler approach, ultimately arriving at the correct answer. An additional example of self-correction is provided in Appendix
|
| 1991 |
+
A.2
|
| 1992 |
+
. Notably, no self-reflection training data or prompt was included, suggesting that advanced System 2 reasoning can foster intrinsic self-reflection.
|
| 1993 |
+
Figure 5:
|
| 1994 |
+
Pass@1 accuracy of policy models and their accuracy after applying System 2 reasoning with various reward models, shows that reward models primarily determine the final performance.
|
| 1995 |
+
PPM shapes the reasoning boundary in System 2 deep thinking
|
| 1996 |
+
. Both the policy and reward models are crucial for System 2 deep reasoning. Our experiments show that once the policy model attains a reasonably strong capability level,
|
| 1997 |
+
(see Appendix
|
| 1998 |
+
A.1
|
| 1999 |
+
), the PPM becomes the key determinant of the upper performance limit.
|
| 2000 |
+
Fig.
|
| 2001 |
+
5
|
| 2002 |
+
summarizes the accuracy of policy models of different sizes, as well as the improvements achieved with reward models. Despite variations in Pass@1 accuracy due to differences in training strategies, datasets, and model scales, the reward model proves to be the dominant factor in System 2 reasoning. For instance, although the SFT accuracy of
|
| 2003 |
+
\sysname
|
| 2004 |
+
-7B is lower than Qwen2.5-Math-72B-Instruct, pairing it with our 7B PPM allows
|
| 2005 |
+
\sysname
|
| 2006 |
+
to outperform the 72B policy model with Qwen 72B ORM. Moreover, despite varying Pass@1 accuracy across our three policy SLM sizes, the final reasoning accuracy converges after applying the PPM.
|
| 2007 |
+
PPM spots theorem-application steps
|
| 2008 |
+
. When solving challenging math problems, identifying and applying relevant theorems or key conclusions often form the cornerstone of successful problem-solving
|
| 2009 |
+
(Xin et al.,
|
| 2010 |
+
2024
|
| 2011 |
+
)
|
| 2012 |
+
. In our experiments, we find that during
|
| 2013 |
+
\sysname
|
| 2014 |
+
problem-solving, our PPM effectively identifies critical theorem-application intermediate steps within policy model’s deep thinking process. These steps are predicted with high reward scores, guiding the policy model to generate the correct solution. Appendix
|
| 2015 |
+
A.2
|
| 2016 |
+
provides examples where the PPM successfully identifies key theorems such as Fermat’s little theorem
|
| 2017 |
+
(Weisstein,
|
| 2018 |
+
a
|
| 2019 |
+
)
|
| 2020 |
+
, Vieta’s formulas
|
| 2021 |
+
(Weisstein,
|
| 2022 |
+
b
|
| 2023 |
+
)
|
| 2024 |
+
, the AM-GM inequality
|
| 2025 |
+
(
|
| 2026 |
+
amg,
|
| 2027 |
+
)
|
| 2028 |
+
, the Pythagorean theorem
|
| 2029 |
+
(
|
| 2030 |
+
pyt,
|
| 2031 |
+
)
|
| 2032 |
+
, and the Shoelace Theorem
|
| 2033 |
+
(
|
| 2034 |
+
sho,
|
| 2035 |
+
)
|
| 2036 |
+
, etc.
|
| 2037 |
+
Generalization discussions
|
| 2038 |
+
.
|
| 2039 |
+
\sysname
|
| 2040 |
+
offers a general methodology for improving LLM reasoning applicable to various domains. First,
|
| 2041 |
+
\sysname
|
| 2042 |
+
can generalize to more challenging math tasks, such as theorem proving, though its current focus is on word problems due to dataset limitations. Nonetheless,
|
| 2043 |
+
\sysname
|
| 2044 |
+
demonstrates the potential to prove mathematical statements. As shown in Appendix
|
| 2045 |
+
A.2
|
| 2046 |
+
, it successfully proves an Olympiad-level problem involving Fermat’s Little Theorem, providing a step-by-step correct proof through its deep reasoning process. Second,
|
| 2047 |
+
\sysname
|
| 2048 |
+
can generalize to other domains, such as code and commonsense reasoning. Notably, synthesizing step-by-step verified training trajectories for general reasoning requires a mechanism to provide feedback on whether a given trajectory reaches the desired output at the end of MCTS rollout. For instance, in code reasoning, this could involve designing extensive test cases; in general reasoning, feedback could be obtained through human labeling or mutual verification with another LLM
|
| 2049 |
+
(Qi et al.,
|
| 2050 |
+
2024
|
| 2051 |
+
)
|
| 2052 |
+
.
|
| 2053 |
+
6
|
| 2054 |
+
Conclusion
|
| 2055 |
+
In this work, we present
|
| 2056 |
+
\sysname
|
| 2057 |
+
, a self-evolved System 2 deep thinking approach that significantly boosts the math reasoning capabilities of small LLMs, achieving state-of-the-art OpenAI o1-level performance. Our approach demonstrates that SLMs can self-generate high-quality training data for frontier-level math reasoning. Extensive experiments across four different-sized SLMs and challenging math benchmarks demonstrate the superiority of
|
| 2058 |
+
\sysname
|
| 2059 |
+
, with achieving leading results while outperforming existing math reasoning LLMs and Best-of-N baselines. We also reveal key findings, including the emergence of self-reflection and the effectiveness of the PPM in identifying critical intermediate steps, such as theorem-application steps. Finally,
|
| 2060 |
+
\sysname
|
| 2061 |
+
can achieve further improvements by collecting more challenging math problems, we leave this as future work.
|
| 2062 |
+
Acknowledgement
|
| 2063 |
+
In the early stages of this work, we faced significant challenges due to limited GPU resources and restricted access to the GPT-4 API. We are deeply grateful to Qiufeng Yin and Chengmin Chi for their assistance in collecting math problems and providing GPT-4 resources for new math problem synthesis. Special thanks go to my colleagues, Lingxiao Ma, Ying Cao, Baotong Lu, Jing Liu, Jiahang Xu, Chengruidong Zhang, Siyuan Wang, Gaokai Zhang, Yujian Li, and Yang Wang, for generously sharing their GPU quotas with us.
|
| 2064 |
+
References
|
| 2065 |
+
[1]
|
| 2066 |
+
Inequality of arithmetic and geometric means.
|
| 2067 |
+
URL
|
| 2068 |
+
https://artofproblemsolving.com/wiki/index.php/AM-GM_Inequality
|
| 2069 |
+
.
|
| 2070 |
+
[2]
|
| 2071 |
+
Pythagorean theorem.
|
| 2072 |
+
URL
|
| 2073 |
+
https://en.wikipedia.org/wiki/Pythagorean_theorem
|
| 2074 |
+
.
|
| 2075 |
+
[3]
|
| 2076 |
+
Shoelace theorem.
|
| 2077 |
+
URL
|
| 2078 |
+
https://artofproblemsolving.com/wiki/index.php/Shoelace_Theorem
|
| 2079 |
+
.
|
| 2080 |
+
Abdin et al. [2024]
|
| 2081 |
+
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah,
|
| 2082 |
+
Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl,
|
| 2083 |
+
et al.
|
| 2084 |
+
Phi-3 technical report: A highly capable language model locally on
|
| 2085 |
+
your phone.
|
| 2086 |
+
arXiv preprint arXiv:2404.14219
|
| 2087 |
+
, 2024.
|
| 2088 |
+
AI-MO [2024a]
|
| 2089 |
+
AI-MO.
|
| 2090 |
+
Aime 2024, 2024a.
|
| 2091 |
+
URL
|
| 2092 |
+
https://huggingface.co/datasets/AI-MO/aimo-validation-aime
|
| 2093 |
+
.
|
| 2094 |
+
AI-MO [2024b]
|
| 2095 |
+
AI-MO.
|
| 2096 |
+
Amc 2023, 2024b.
|
| 2097 |
+
URL
|
| 2098 |
+
https://huggingface.co/datasets/AI-MO/aimo-validation-amc
|
| 2099 |
+
.
|
| 2100 |
+
Brown et al. [2024]
|
| 2101 |
+
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le,
|
| 2102 |
+
Christopher Ré, and Azalia Mirhoseini.
|
| 2103 |
+
Large language monkeys: Scaling inference compute with repeated
|
| 2104 |
+
sampling.
|
| 2105 |
+
arXiv preprint arXiv:2407.21787
|
| 2106 |
+
, 2024.
|
| 2107 |
+
Chen et al. [2024]
|
| 2108 |
+
Guoxin Chen, Minpeng Liao, Chengxi Li, and Kai Fan.
|
| 2109 |
+
Alphamath almost zero: process supervision without process, 2024.
|
| 2110 |
+
Cobbe et al. [2021]
|
| 2111 |
+
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz
|
| 2112 |
+
Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano,
|
| 2113 |
+
et al.
|
| 2114 |
+
Training verifiers to solve math word problems.
|
| 2115 |
+
arXiv preprint arXiv:2110.14168
|
| 2116 |
+
, 2021.
|
| 2117 |
+
Daniel [2011]
|
| 2118 |
+
Kahneman Daniel.
|
| 2119 |
+
Thinking, fast and slow.
|
| 2120 |
+
Macmillan
|
| 2121 |
+
, 2011.
|
| 2122 |
+
Dubey et al. [2024]
|
| 2123 |
+
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad
|
| 2124 |
+
Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan,
|
| 2125 |
+
et al.
|
| 2126 |
+
The llama 3 herd of models.
|
| 2127 |
+
arXiv preprint arXiv:2407.21783
|
| 2128 |
+
, 2024.
|
| 2129 |
+
Gou et al. [2023]
|
| 2130 |
+
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yujiu Yang, Minlie Huang, Nan Duan,
|
| 2131 |
+
Weizhu Chen, et al.
|
| 2132 |
+
Tora: A tool-integrated reasoning agent for mathematical problem
|
| 2133 |
+
solving.
|
| 2134 |
+
arXiv preprint arXiv:2309.17452
|
| 2135 |
+
, 2023.
|
| 2136 |
+
Hao et al. [2023]
|
| 2137 |
+
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and
|
| 2138 |
+
Zhiting Hu.
|
| 2139 |
+
Reasoning with language model is planning with world model.
|
| 2140 |
+
arXiv preprint arXiv:2305.14992
|
| 2141 |
+
, 2023.
|
| 2142 |
+
He et al. [2024]
|
| 2143 |
+
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen,
|
| 2144 |
+
Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, et al.
|
| 2145 |
+
Olympiadbench: A challenging benchmark for promoting agi with
|
| 2146 |
+
olympiad-level bilingual multimodal scientific problems.
|
| 2147 |
+
arXiv preprint arXiv:2402.14008
|
| 2148 |
+
, 2024.
|
| 2149 |
+
Huang et al. [2023]
|
| 2150 |
+
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu,
|
| 2151 |
+
Xinying Song, and Denny Zhou.
|
| 2152 |
+
Large language models cannot self-correct reasoning yet.
|
| 2153 |
+
arXiv preprint arXiv:2310.01798
|
| 2154 |
+
, 2023.
|
| 2155 |
+
Huang et al. [2024]
|
| 2156 |
+
Zhen Huang, Haoyang Zou, Xuefeng Li, Yixiu Liu, Yuxiang Zheng, Ethan Chern,
|
| 2157 |
+
Shijie Xia, Yiwei Qin, Weizhe Yuan, and Pengfei Liu.
|
| 2158 |
+
O1 replication journey – part 2: Surpassing o1-preview through
|
| 2159 |
+
simple distillation big progress or bitter lesson?
|
| 2160 |
+
Github
|
| 2161 |
+
, 2024.
|
| 2162 |
+
URL
|
| 2163 |
+
https://github.com/GAIR-NLP/O1-Journey
|
| 2164 |
+
.
|
| 2165 |
+
Jia LI and Polu [2024a]
|
| 2166 |
+
Lewis Tunstall Ben Lipkin Roman Soletskyi Shengyi Costa Huang Kashif Rasul
|
| 2167 |
+
Longhui Yu Albert Jiang Ziju Shen Zihan Qin Bin Dong Li Zhou Yann Fleureau
|
| 2168 |
+
Guillaume Lample Jia LI, Edward Beeching and Stanislas Polu.
|
| 2169 |
+
Numinamath.
|
| 2170 |
+
[https://github.com/project-numina/aimo-progress-prize](https://github.com/project-numina/aimo-progress-prize/blob/main/report/numina_dataset.pdf)
|
| 2171 |
+
,
|
| 2172 |
+
2024a.
|
| 2173 |
+
Jia LI and Polu [2024b]
|
| 2174 |
+
Lewis Tunstall Ben Lipkin Roman Soletskyi Shengyi Costa Huang Kashif Rasul
|
| 2175 |
+
Longhui Yu Albert Jiang Ziju Shen Zihan Qin Bin Dong Li Zhou Yann Fleureau
|
| 2176 |
+
Guillaume Lample Jia LI, Edward Beeching and Stanislas Polu.
|
| 2177 |
+
Numinamath cot, 2024b.
|
| 2178 |
+
URL
|
| 2179 |
+
https://huggingface.co/datasets/AI-MO/NuminaMath-CoT
|
| 2180 |
+
.
|
| 2181 |
+
Kang et al. [2024]
|
| 2182 |
+
Jikun Kang, Xin Zhe Li, Xi Chen, Amirreza Kazemi, and Boxing Chen.
|
| 2183 |
+
Mindstar: Enhancing math reasoning in pre-trained llms at inference
|
| 2184 |
+
time.
|
| 2185 |
+
arXiv preprint arXiv:2405.16265
|
| 2186 |
+
, 2024.
|
| 2187 |
+
Kocsis and Szepesvári [2006]
|
| 2188 |
+
Levente Kocsis and Csaba Szepesvári.
|
| 2189 |
+
Bandit based monte-carlo planning.
|
| 2190 |
+
volume 2006, pages 282–293, 09 2006.
|
| 2191 |
+
ISBN 978-3-540-45375-8.
|
| 2192 |
+
doi:
|
| 2193 |
+
10.1007/11871842_29
|
| 2194 |
+
.
|
| 2195 |
+
Kumar et al. [2024]
|
| 2196 |
+
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi
|
| 2197 |
+
Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, et al.
|
| 2198 |
+
Training language models to self-correct via reinforcement learning.
|
| 2199 |
+
arXiv preprint arXiv:2409.12917
|
| 2200 |
+
, 2024.
|
| 2201 |
+
Lanham et al. [2023]
|
| 2202 |
+
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison,
|
| 2203 |
+
Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion,
|
| 2204 |
+
et al.
|
| 2205 |
+
Measuring faithfulness in chain-of-thought reasoning.
|
| 2206 |
+
arXiv preprint arXiv:2307.13702
|
| 2207 |
+
, 2023.
|
| 2208 |
+
Li et al. [2024]
|
| 2209 |
+
Chen Li, Weiqi Wang, Jingcheng Hu, Yixuan Wei, Nanning Zheng, Han Hu, Zheng
|
| 2210 |
+
Zhang, and Houwen Peng.
|
| 2211 |
+
Common 7b language models already possess strong math capabilities.
|
| 2212 |
+
arXiv preprint arXiv:2403.04706
|
| 2213 |
+
, 2024.
|
| 2214 |
+
Liao et al. [2024]
|
| 2215 |
+
Minpeng Liao, Wei Luo, Chengxi Li, Jing Wu, and Kai Fan.
|
| 2216 |
+
Mario: Math reasoning with code interpreter output–a reproducible
|
| 2217 |
+
pipeline.
|
| 2218 |
+
arXiv preprint arXiv:2401.08190
|
| 2219 |
+
, 2024.
|
| 2220 |
+
Lightman et al. [2023]
|
| 2221 |
+
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy
|
| 2222 |
+
Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe.
|
| 2223 |
+
Let’s verify step by step.
|
| 2224 |
+
arXiv preprint arXiv:2305.20050
|
| 2225 |
+
, 2023.
|
| 2226 |
+
Lightman et al. [2024]
|
| 2227 |
+
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker,
|
| 2228 |
+
Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe.
|
| 2229 |
+
Let’s verify step by step.
|
| 2230 |
+
In
|
| 2231 |
+
The Twelfth International Conference on Learning
|
| 2232 |
+
Representations
|
| 2233 |
+
, 2024.
|
| 2234 |
+
URL
|
| 2235 |
+
https://openreview.net/forum?id=v8L0pN6EOi
|
| 2236 |
+
.
|
| 2237 |
+
Liu et al. [2024]
|
| 2238 |
+
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang
|
| 2239 |
+
Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al.
|
| 2240 |
+
Deepseek-v3 technical report.
|
| 2241 |
+
arXiv preprint arXiv:2412.19437
|
| 2242 |
+
, 2024.
|
| 2243 |
+
Luo et al. [2023]
|
| 2244 |
+
Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo
|
| 2245 |
+
Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang.
|
| 2246 |
+
Wizardmath: Empowering mathematical reasoning for large language
|
| 2247 |
+
models via reinforced evol-instruct.
|
| 2248 |
+
arXiv preprint arXiv:2308.09583
|
| 2249 |
+
, 2023.
|
| 2250 |
+
Luo et al. [2024]
|
| 2251 |
+
Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Phatale, Harsh Lara, Yunxuan
|
| 2252 |
+
Li, Lei Shu, Yun Zhu, Lei Meng, Jiao Sun, et al.
|
| 2253 |
+
Improve mathematical reasoning in language models by automated
|
| 2254 |
+
process supervision.
|
| 2255 |
+
arXiv preprint arXiv:2406.06592
|
| 2256 |
+
, 2024.
|
| 2257 |
+
Microsoft [2024]
|
| 2258 |
+
Microsoft.
|
| 2259 |
+
Phi-3-mini-4k-instruct, 2024.
|
| 2260 |
+
URL
|
| 2261 |
+
https://huggingface.co/microsoft/Phi-3-mini-4k-instruct
|
| 2262 |
+
.
|
| 2263 |
+
Noam Brown and Lightman [2024]
|
| 2264 |
+
Ilge Akkaya Noam Brown and Hunter Lightman.
|
| 2265 |
+
Openai’s noam brown, ilge akkaya and hunter lightman on o1 and
|
| 2266 |
+
teaching llms to reason better, 2024.
|
| 2267 |
+
URL
|
| 2268 |
+
https://www.youtube.com/watch?v=jPluSXJpdrA
|
| 2269 |
+
.
|
| 2270 |
+
OpenAI [2023]
|
| 2271 |
+
OpenAI.
|
| 2272 |
+
Gpt-4 technical report.
|
| 2273 |
+
2023.
|
| 2274 |
+
OpenAI [2024]
|
| 2275 |
+
OpenAI.
|
| 2276 |
+
Openai o1 system card.
|
| 2277 |
+
preprint
|
| 2278 |
+
, 2024.
|
| 2279 |
+
Ouyang et al. [2022]
|
| 2280 |
+
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela
|
| 2281 |
+
Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al.
|
| 2282 |
+
Training language models to follow instructions with human feedback.
|
| 2283 |
+
Advances in Neural Information Processing Systems
|
| 2284 |
+
,
|
| 2285 |
+
35:27730–27744, 2022.
|
| 2286 |
+
Qi et al. [2024]
|
| 2287 |
+
Zhenting Qi, Mingyuan Ma, Jiahang Xu, Li Lyna Zhang, Fan Yang, and Mao Yang.
|
| 2288 |
+
Mutual reasoning makes smaller llms stronger problem-solvers.
|
| 2289 |
+
arXiv preprint arXiv:2408.06195
|
| 2290 |
+
, 2024.
|
| 2291 |
+
Qwen [2024a]
|
| 2292 |
+
Qwen.
|
| 2293 |
+
Qwen2-math-7b, 2024a.
|
| 2294 |
+
URL
|
| 2295 |
+
https://huggingface.co/Qwen/Qwen2-Math-7B
|
| 2296 |
+
.
|
| 2297 |
+
Qwen [2024b]
|
| 2298 |
+
Qwen.
|
| 2299 |
+
Qwen2.5-math-1.5b, 2024b.
|
| 2300 |
+
URL
|
| 2301 |
+
https://huggingface.co/Qwen/Qwen2.5-Math-1.5B
|
| 2302 |
+
.
|
| 2303 |
+
Qwen [2024c]
|
| 2304 |
+
Qwen.
|
| 2305 |
+
Qwen2.5-math-7b, 2024c.
|
| 2306 |
+
URL
|
| 2307 |
+
https://huggingface.co/Qwen/Qwen2.5-Math-7B
|
| 2308 |
+
.
|
| 2309 |
+
Renze and Guven [2024]
|
| 2310 |
+
Matthew Renze and Erhan Guven.
|
| 2311 |
+
Self-reflection in llm agents: Effects on problem-solving
|
| 2312 |
+
performance.
|
| 2313 |
+
arXiv preprint arXiv:2405.06682
|
| 2314 |
+
, 2024.
|
| 2315 |
+
Shinn et al. [2024]
|
| 2316 |
+
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu
|
| 2317 |
+
Yao.
|
| 2318 |
+
Reflexion: Language agents with verbal reinforcement learning.
|
| 2319 |
+
Advances in Neural Information Processing Systems
|
| 2320 |
+
, 36, 2024.
|
| 2321 |
+
Silver et al. [2017]
|
| 2322 |
+
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew
|
| 2323 |
+
Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore
|
| 2324 |
+
Graepel, et al.
|
| 2325 |
+
Mastering chess and shogi by self-play with a general reinforcement
|
| 2326 |
+
learning algorithm.
|
| 2327 |
+
arXiv preprint arXiv:1712.01815
|
| 2328 |
+
, 2017.
|
| 2329 |
+
Snell et al. [2024]
|
| 2330 |
+
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar.
|
| 2331 |
+
Scaling llm test-time compute optimally can be more effective than
|
| 2332 |
+
scaling model parameters.
|
| 2333 |
+
arXiv preprint arXiv:2408.03314
|
| 2334 |
+
, 2024.
|
| 2335 |
+
Tang et al. [2024]
|
| 2336 |
+
Zhengyang Tang, Xingxing Zhang, Benyou Wan, and Furu Wei.
|
| 2337 |
+
Mathscale: Scaling instruction tuning for mathematical reasoning.
|
| 2338 |
+
arXiv preprint arXiv:2403.02884
|
| 2339 |
+
, 2024.
|
| 2340 |
+
Team [2024a]
|
| 2341 |
+
Qwen Team.
|
| 2342 |
+
Qwq: Reflect deeply on the boundaries of the unknown, November
|
| 2343 |
+
2024a.
|
| 2344 |
+
URL
|
| 2345 |
+
https://qwenlm.github.io/blog/qwq-32b-preview/
|
| 2346 |
+
.
|
| 2347 |
+
Team [2024b]
|
| 2348 |
+
The Mistral AI Team.
|
| 2349 |
+
Mathstral-7b-v0.1, 2024b.
|
| 2350 |
+
URL
|
| 2351 |
+
https://huggingface.co/mistralai/Mathstral-7B-v0.1
|
| 2352 |
+
.
|
| 2353 |
+
Toshniwal et al. [2024]
|
| 2354 |
+
Shubham Toshniwal, Wei Du, Ivan Moshkov, Branislav Kisacanin, Alexan
|
| 2355 |
+
Ayrapetyan, and Igor Gitman.
|
| 2356 |
+
Openmathinstruct-2: Accelerating ai for math with massive open-source
|
| 2357 |
+
instruction data.
|
| 2358 |
+
arXiv preprint arXiv:2410.01560
|
| 2359 |
+
, 2024.
|
| 2360 |
+
Valmeekam et al. [2023]
|
| 2361 |
+
Karthik Valmeekam, Sarath Sreedharan, Matthew Marquez, Alberto Olmo, and
|
| 2362 |
+
Subbarao Kambhampati.
|
| 2363 |
+
On the planning abilities of large language models (a critical
|
| 2364 |
+
investigation with a proposed benchmark).
|
| 2365 |
+
arXiv preprint arXiv:2302.06706
|
| 2366 |
+
, 2023.
|
| 2367 |
+
Wang et al. [2024a]
|
| 2368 |
+
Chaojie Wang, Yanchen Deng, Zhiyi Lv, Shuicheng Yan, and An Bo.
|
| 2369 |
+
Q*: Improving multi-step reasoning for llms with deliberative
|
| 2370 |
+
planning, 2024a.
|
| 2371 |
+
Wang et al. [2024b]
|
| 2372 |
+
Ke Wang, Houxing Ren, Aojun Zhou, Zimu Lu, Sichun Luo, Weikang Shi, Renrui
|
| 2373 |
+
Zhang, Linqi Song, Mingjie Zhan, and Hongsheng Li.
|
| 2374 |
+
Mathcoder: Seamless code integration in LLMs for enhanced
|
| 2375 |
+
mathematical reasoning.
|
| 2376 |
+
In
|
| 2377 |
+
The Twelfth International Conference on Learning
|
| 2378 |
+
Representations
|
| 2379 |
+
, 2024b.
|
| 2380 |
+
URL
|
| 2381 |
+
https://openreview.net/forum?id=z8TW0ttBPp
|
| 2382 |
+
.
|
| 2383 |
+
Wang et al. [2024c]
|
| 2384 |
+
Peiyi Wang, Lei Li, Zhihong Shao, R. X. Xu, Damai Dai, Yifei Li, Deli Chen,
|
| 2385 |
+
Y. Wu, and Zhifang Sui.
|
| 2386 |
+
Math-shepherd: Verify and reinforce llms step-by-step without human
|
| 2387 |
+
annotations, 2024c.
|
| 2388 |
+
Wang et al. [2023]
|
| 2389 |
+
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang,
|
| 2390 |
+
Aakanksha Chowdhery, and Denny Zhou.
|
| 2391 |
+
Self-consistency improves chain of thought reasoning in language
|
| 2392 |
+
models.
|
| 2393 |
+
In
|
| 2394 |
+
The Eleventh International Conference on Learning
|
| 2395 |
+
Representations
|
| 2396 |
+
, 2023.
|
| 2397 |
+
URL
|
| 2398 |
+
https://openreview.net/forum?id=1PL1NIMMrw
|
| 2399 |
+
.
|
| 2400 |
+
Weisstein [a]
|
| 2401 |
+
Eric W. Weisstein.
|
| 2402 |
+
Fermat’s little theorem, a.
|
| 2403 |
+
URL
|
| 2404 |
+
https://mathworld.wolfram.com/FermatsLittleTheorem.html
|
| 2405 |
+
.
|
| 2406 |
+
Weisstein [b]
|
| 2407 |
+
Eric W. Weisstein.
|
| 2408 |
+
Vieta’s formulas, from mathworld—a wolfram web resource,
|
| 2409 |
+
b.
|
| 2410 |
+
URL
|
| 2411 |
+
http://mathworld.wolfram.com/Tree.html
|
| 2412 |
+
.
|
| 2413 |
+
Wu et al. [2024]
|
| 2414 |
+
Yangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck, and Yiming Yang.
|
| 2415 |
+
An empirical analysis of compute-optimal inference for
|
| 2416 |
+
problem-solving with language models.
|
| 2417 |
+
arXiv preprint arXiv:2408.00724
|
| 2418 |
+
, 2024.
|
| 2419 |
+
Xin et al. [2024]
|
| 2420 |
+
Huajian Xin, Daya Guo, Zhihong Shao, Zhizhou Ren, Qihao Zhu, Bo Liu, Chong
|
| 2421 |
+
Ruan, Wenda Li, and Xiaodan Liang.
|
| 2422 |
+
Deepseek-prover: Advancing theorem proving in llms through
|
| 2423 |
+
large-scale synthetic data.
|
| 2424 |
+
arXiv preprint arXiv:2405.14333
|
| 2425 |
+
, 2024.
|
| 2426 |
+
Yang et al. [2024]
|
| 2427 |
+
An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li,
|
| 2428 |
+
Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, et al.
|
| 2429 |
+
Qwen2. 5-math technical report: Toward mathematical expert model via
|
| 2430 |
+
self-improvement.
|
| 2431 |
+
arXiv preprint arXiv:2409.12122
|
| 2432 |
+
, 2024.
|
| 2433 |
+
Yao et al. [2024]
|
| 2434 |
+
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and
|
| 2435 |
+
Karthik Narasimhan.
|
| 2436 |
+
Tree of thoughts: Deliberate problem solving with large language
|
| 2437 |
+
models.
|
| 2438 |
+
Advances in Neural Information Processing Systems
|
| 2439 |
+
, 36, 2024.
|
| 2440 |
+
Yu et al. [2023a]
|
| 2441 |
+
Fei Yu, Anningzhe Gao, and Benyou Wang.
|
| 2442 |
+
Outcome-supervised verifiers for planning in mathematical reasoning.
|
| 2443 |
+
arXiv preprint arXiv:2311.09724
|
| 2444 |
+
, 2023a.
|
| 2445 |
+
Yu et al. [2023b]
|
| 2446 |
+
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang,
|
| 2447 |
+
James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu.
|
| 2448 |
+
Metamath: Bootstrap your own mathematical questions for large
|
| 2449 |
+
language models.
|
| 2450 |
+
arXiv preprint arXiv:2309.12284
|
| 2451 |
+
, 2023b.
|
| 2452 |
+
Yuan et al. [2023]
|
| 2453 |
+
Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Keming Lu, Chuanqi Tan,
|
| 2454 |
+
Chang Zhou, and Jingren Zhou.
|
| 2455 |
+
Scaling relationship on learning mathematical reasoning with large
|
| 2456 |
+
language models.
|
| 2457 |
+
arXiv preprint arXiv:2308.01825
|
| 2458 |
+
, 2023.
|
| 2459 |
+
Zhang et al. [2024a]
|
| 2460 |
+
Dan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue, Yuxiao Dong, and Jie Tang.
|
| 2461 |
+
Rest-mcts*: Llm self-training via process reward guided tree search.
|
| 2462 |
+
arXiv preprint arXiv:2406.03816
|
| 2463 |
+
, 2024a.
|
| 2464 |
+
Zhang et al. [2024b]
|
| 2465 |
+
Di Zhang, Jiatong Li, Xiaoshui Huang, Dongzhan Zhou, Yuqiang Li, and Wanli
|
| 2466 |
+
Ouyang.
|
| 2467 |
+
Accessing gpt-4 level mathematical olympiad solutions via monte carlo
|
| 2468 |
+
tree self-refine with llama-3 8b.
|
| 2469 |
+
arXiv preprint arXiv:2406.07394
|
| 2470 |
+
, 2024b.
|
| 2471 |
+
Zheng et al. [2023]
|
| 2472 |
+
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao
|
| 2473 |
+
Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E.
|
| 2474 |
+
Gonzalez, and Ion Stoica.
|
| 2475 |
+
Judging LLM-as-a-judge with MT-bench and chatbot arena.
|
| 2476 |
+
In
|
| 2477 |
+
Thirty-seventh Conference on Neural Information Processing
|
| 2478 |
+
Systems Datasets and Benchmarks Track
|
| 2479 |
+
, 2023.
|
| 2480 |
+
Appendix A
|
| 2481 |
+
Appendix
|
| 2482 |
+
A.1
|
| 2483 |
+
Additional Experiments and Details
|
| 2484 |
+
Data Generation Details
|
| 2485 |
+
. As detailed in Sec.
|
| 2486 |
+
3.4
|
| 2487 |
+
, each round starts by self-generating step-by-step verified trajectories for 747k math word problems. The maximum tree depth
|
| 2488 |
+
d
|
| 2489 |
+
d
|
| 2490 |
+
is set to 16, with 16 MCTS rollouts conducted per problem by default. At each step, we allow to explore 8 candidate nodes, and the constant
|
| 2491 |
+
c
|
| 2492 |
+
c
|
| 2493 |
+
in Eq.
|
| 2494 |
+
1
|
| 2495 |
+
is set to 2 to promote greater exploration. In the bootstrap round, due to the large size of the initial policy model (236B), we used smaller parameters: 8 rollouts and 5 candidate nodes per step. To improve the accuracy of solving challenging problems in round 4, we increase the number of candidate nodes to 16 and conduct 2 MCTS tree expansions per problem using different random seeds. Detailed prompts are available in Appendix
|
| 2496 |
+
A.3
|
| 2497 |
+
.
|
| 2498 |
+
Training Details
|
| 2499 |
+
. In each round, we collect step-by-step verified trajectories to fine-tune the policy LLM and train the PPM. To reduce noise
|
| 2500 |
+
in synthetic math problems (e.g., incorrect ground-truth answers labeled by GPT-4), we remove synthetic problems with trajectories achieving less than 50% accuracy. Based on our extensive experiments, the policy LLM is fine-tuned from the initial base model in each round, rather than training incrementally on the model from the previous round.
|
| 2501 |
+
All policy SLMs are trained for 2 epochs with a sequence length of 4096 tokens and a batch size of 128. We use AdamW optimizer with a linear learning rate scheduler, setting the initial learning rate to 7e-6 for Qwen models, and a cosine scheduler with an initial learning rate of 5e-6 for Phi3-mini-Instruct.
|
| 2502 |
+
The PPM is trained for 1 epoch with a batch size of 512 and an initial learning rate of 7e-6.
|
| 2503 |
+
Training the ORM and PQM
|
| 2504 |
+
. The Outcome Reward Model (ORM) and the Q-value-based Process Reward Model (PQM) share the same model architecture and training parameters with our PPM. To train the ORM, we collect trajectories from math problems containing both correct and incorrect solutions. Specifically, the two trajectories with the highest average Q-values are selected as positive examples, while the two with the lowest are chosen as negative examples. Following Qwen2.5-Math
|
| 2505 |
+
(Yang et al.,
|
| 2506 |
+
2024
|
| 2507 |
+
)
|
| 2508 |
+
, we adopt the pairwise ranking loss
|
| 2509 |
+
(Ouyang et al.,
|
| 2510 |
+
2022
|
| 2511 |
+
)
|
| 2512 |
+
to optimize the ORM. To train the PQM, we follow
|
| 2513 |
+
Chen et al. (
|
| 2514 |
+
2024
|
| 2515 |
+
)
|
| 2516 |
+
to use step-level Q-values as reward labels. Let
|
| 2517 |
+
𝐱
|
| 2518 |
+
=
|
| 2519 |
+
x
|
| 2520 |
+
⊕
|
| 2521 |
+
s
|
| 2522 |
+
1
|
| 2523 |
+
⊕
|
| 2524 |
+
s
|
| 2525 |
+
2
|
| 2526 |
+
⊕
|
| 2527 |
+
…
|
| 2528 |
+
⊕
|
| 2529 |
+
s
|
| 2530 |
+
d
|
| 2531 |
+
\mathbf{x}=x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{d}
|
| 2532 |
+
be the trajectory, with annotated Q-values
|
| 2533 |
+
𝐐
|
| 2534 |
+
=
|
| 2535 |
+
(
|
| 2536 |
+
Q
|
| 2537 |
+
|
| 2538 |
+
(
|
| 2539 |
+
s
|
| 2540 |
+
1
|
| 2541 |
+
)
|
| 2542 |
+
,
|
| 2543 |
+
Q
|
| 2544 |
+
|
| 2545 |
+
(
|
| 2546 |
+
s
|
| 2547 |
+
1
|
| 2548 |
+
)
|
| 2549 |
+
,
|
| 2550 |
+
…
|
| 2551 |
+
,
|
| 2552 |
+
Q
|
| 2553 |
+
|
| 2554 |
+
(
|
| 2555 |
+
s
|
| 2556 |
+
d
|
| 2557 |
+
)
|
| 2558 |
+
)
|
| 2559 |
+
\mathbf{Q}=(Q(s_{1}),Q(s_{1}),...,Q(s_{d}))
|
| 2560 |
+
and predicted Q-values
|
| 2561 |
+
𝐐
|
| 2562 |
+
′
|
| 2563 |
+
=
|
| 2564 |
+
(
|
| 2565 |
+
Q
|
| 2566 |
+
′
|
| 2567 |
+
|
| 2568 |
+
(
|
| 2569 |
+
s
|
| 2570 |
+
1
|
| 2571 |
+
)
|
| 2572 |
+
,
|
| 2573 |
+
Q
|
| 2574 |
+
′
|
| 2575 |
+
|
| 2576 |
+
(
|
| 2577 |
+
s
|
| 2578 |
+
1
|
| 2579 |
+
)
|
| 2580 |
+
,
|
| 2581 |
+
…
|
| 2582 |
+
,
|
| 2583 |
+
Q
|
| 2584 |
+
′
|
| 2585 |
+
|
| 2586 |
+
(
|
| 2587 |
+
s
|
| 2588 |
+
d
|
| 2589 |
+
)
|
| 2590 |
+
)
|
| 2591 |
+
\mathbf{Q^{\prime}}=(Q^{\prime}(s_{1}),Q^{\prime}(s_{1}),...,Q^{\prime}(s_{d}))
|
| 2592 |
+
for each step. To stabilize PQM training, we treat each trajectory as a single training sample and predict Q-values for all steps simultaneously, rather than splitting it into individual per-step samples. Specifically, to predict the Q-value
|
| 2593 |
+
Q
|
| 2594 |
+
′
|
| 2595 |
+
|
| 2596 |
+
(
|
| 2597 |
+
s
|
| 2598 |
+
i
|
| 2599 |
+
)
|
| 2600 |
+
Q^{\prime}(s_{i})
|
| 2601 |
+
for step
|
| 2602 |
+
s
|
| 2603 |
+
i
|
| 2604 |
+
s_{i}
|
| 2605 |
+
, PQM takes the trajectory from the question up to step
|
| 2606 |
+
s
|
| 2607 |
+
i
|
| 2608 |
+
s_{i}
|
| 2609 |
+
(i.e.,
|
| 2610 |
+
x
|
| 2611 |
+
⊕
|
| 2612 |
+
s
|
| 2613 |
+
1
|
| 2614 |
+
⊕
|
| 2615 |
+
s
|
| 2616 |
+
2
|
| 2617 |
+
⊕
|
| 2618 |
+
…
|
| 2619 |
+
⊕
|
| 2620 |
+
s
|
| 2621 |
+
i
|
| 2622 |
+
x\oplus s_{1}\oplus s_{2}\oplus...\oplus s_{i}
|
| 2623 |
+
) as input and outputs a value between -1 and 1. We use a mean squared error (MSE) loss for PQM training:
|
| 2624 |
+
ℒ
|
| 2625 |
+
p
|
| 2626 |
+
|
| 2627 |
+
r
|
| 2628 |
+
|
| 2629 |
+
m
|
| 2630 |
+
|
| 2631 |
+
(
|
| 2632 |
+
𝐱
|
| 2633 |
+
)
|
| 2634 |
+
=
|
| 2635 |
+
‖
|
| 2636 |
+
𝐐
|
| 2637 |
+
−
|
| 2638 |
+
𝐐
|
| 2639 |
+
′
|
| 2640 |
+
‖
|
| 2641 |
+
𝟐
|
| 2642 |
+
\mathcal{L}_{prm}(\bf{x})=\|\bf{Q}-\bf{Q^{\prime}}\|^{2}
|
| 2643 |
+
(6)
|
| 2644 |
+
Self-evolution Inference Costs.
|
| 2645 |
+
In the initial bootstrap round, we use DeepSeek-Coder-v2-Instruct (236B) as the policy model, using 10 nodes of 8×80GB H100 GPUs with 8 MCTS rollouts. This required approximately two weeks to finish the data generation. For rounds 2–4, using our fine-tuned 7B SLM as the policy model, data generation was performed on 15 nodes of 4×40GB A100 GPUs,
|
| 2646 |
+
with each round completed in three days. In the final round, to include more challenging problems, we increased the number of MCTS rollouts to 64, extending the data generation time to one week.
|
| 2647 |
+
Table 9:
|
| 2648 |
+
Inference costs of
|
| 2649 |
+
\sysname
|
| 2650 |
+
. We show the average number of generated tokens required to generate a trajectory for a given question.
|
| 2651 |
+
MATH
|
| 2652 |
+
AIME 2024
|
| 2653 |
+
AMC 2023
|
| 2654 |
+
Olympiad Bench
|
| 2655 |
+
College Math
|
| 2656 |
+
GSM8K
|
| 2657 |
+
GaokaoEn 2023
|
| 2658 |
+
5453
|
| 2659 |
+
15693
|
| 2660 |
+
14544
|
| 2661 |
+
7889
|
| 2662 |
+
4503
|
| 2663 |
+
3299
|
| 2664 |
+
6375
|
| 2665 |
+
Inference Setting
|
| 2666 |
+
. In our evaluation, we run multiple MCTS to generate candidate solution trajectories. For each problem, we generate 32 candidate nodes at each step and use the PPM to score each node. Since the PPM effectively provides step-level quality evaluations, we limit MCTS to just 4 rollouts per step to update the Q-values. After completing MCTS, the trajectory with the highest PPM score is selected as the final answer. Table
|
| 2667 |
+
9
|
| 2668 |
+
presents the average number of tokens generated to produce a trajectory in MCTS.
|
| 2669 |
+
Table 10:
|
| 2670 |
+
Pass@1 (greedy) accuracy of our fine-tuned policy models for Phi3-mini, Qwen2.5-Math-1.5B, Qwen2-Math-7B and Qwen2.5-Math-7B.
|
| 2671 |
+
Model
|
| 2672 |
+
MATH
|
| 2673 |
+
AIME 2024
|
| 2674 |
+
AMC 2023
|
| 2675 |
+
Olympiad Bench
|
| 2676 |
+
College Math
|
| 2677 |
+
GSM8K
|
| 2678 |
+
GaokaoEn 2023
|
| 2679 |
+
General Base Model: Phi3-mini-Instruct (3.8B)
|
| 2680 |
+
Phi3-mini-Instruct
|
| 2681 |
+
41.4
|
| 2682 |
+
3.33
|
| 2683 |
+
7.5
|
| 2684 |
+
12.3
|
| 2685 |
+
33.1
|
| 2686 |
+
85.7
|
| 2687 |
+
37.1
|
| 2688 |
+
Our policy model
|
| 2689 |
+
68.0
|
| 2690 |
+
10.0
|
| 2691 |
+
37.5
|
| 2692 |
+
36.6
|
| 2693 |
+
48.7
|
| 2694 |
+
87.9
|
| 2695 |
+
53.2
|
| 2696 |
+
Math-Specialized Base Model: Qwen2.5-Math-1.5B
|
| 2697 |
+
Qwen2.5-Math-1.5B
|
| 2698 |
+
51.2
|
| 2699 |
+
0.0
|
| 2700 |
+
22.5
|
| 2701 |
+
16.7
|
| 2702 |
+
38.4
|
| 2703 |
+
74.6
|
| 2704 |
+
46.5
|
| 2705 |
+
Qwen2.5-Math-1.5B-Instruct
|
| 2706 |
+
60.0
|
| 2707 |
+
10.0
|
| 2708 |
+
60.0
|
| 2709 |
+
38.1
|
| 2710 |
+
47.7
|
| 2711 |
+
84.8
|
| 2712 |
+
65.5
|
| 2713 |
+
Our policy model
|
| 2714 |
+
74.8
|
| 2715 |
+
13.3
|
| 2716 |
+
47.5
|
| 2717 |
+
42.5
|
| 2718 |
+
50.1
|
| 2719 |
+
83.1
|
| 2720 |
+
58.7
|
| 2721 |
+
Math-Specialized Base Model: Qwen2-Math-7B
|
| 2722 |
+
Qwen2-Math-7B
|
| 2723 |
+
53.4
|
| 2724 |
+
3.3
|
| 2725 |
+
25.0
|
| 2726 |
+
17.3
|
| 2727 |
+
39.4
|
| 2728 |
+
80.4
|
| 2729 |
+
47.3
|
| 2730 |
+
Qwen2-Math-7B-Instruct
|
| 2731 |
+
73.2
|
| 2732 |
+
13.3
|
| 2733 |
+
62.5
|
| 2734 |
+
38.2
|
| 2735 |
+
45.9
|
| 2736 |
+
89.9
|
| 2737 |
+
62.1
|
| 2738 |
+
Our policy model
|
| 2739 |
+
73.8
|
| 2740 |
+
16.7
|
| 2741 |
+
45.0
|
| 2742 |
+
43.9
|
| 2743 |
+
52.0
|
| 2744 |
+
88.3
|
| 2745 |
+
65.2
|
| 2746 |
+
Math-Specialized Base Model: Qwen2.5-Math-7B
|
| 2747 |
+
Qwen2.5-Math-7B
|
| 2748 |
+
58.8
|
| 2749 |
+
0.0
|
| 2750 |
+
22.5
|
| 2751 |
+
21.8
|
| 2752 |
+
41.6
|
| 2753 |
+
91.6
|
| 2754 |
+
51.7
|
| 2755 |
+
Qwen2.5-Math-7B-Instruct
|
| 2756 |
+
82.6
|
| 2757 |
+
6.0
|
| 2758 |
+
62.5
|
| 2759 |
+
41.6
|
| 2760 |
+
46.8
|
| 2761 |
+
95.2
|
| 2762 |
+
66.8
|
| 2763 |
+
Our policy model
|
| 2764 |
+
78.4
|
| 2765 |
+
26.7
|
| 2766 |
+
47.5
|
| 2767 |
+
47.1
|
| 2768 |
+
52.5
|
| 2769 |
+
89.7
|
| 2770 |
+
65.7
|
| 2771 |
+
Figure 6:
|
| 2772 |
+
Pass@N accuracy with random sampling from different policy models. Compared to the official Qwen instruct version, our policy model exhibits a stronger ability to sample correct solutions.
|
| 2773 |
+
Figure 7:
|
| 2774 |
+
Pass@N accuracy with PPM-augmented MCTS. Under the same PPM guidance, the four policy models of varying sizes demonstrate convergent capabilities in sampling correct solutions.
|
| 2775 |
+
Pass@N.
|
| 2776 |
+
Table
|
| 2777 |
+
10
|
| 2778 |
+
compares the math reasoning performance of our policy models with the instruct versions developed by the original model team. Our policy models do not consistently outperform the instruct versions. For example, on the Qwen2.5-Math-7B base model, Qwen2.5-Math-7B-Instruct achieves 4.2% higher accuracy on the MATH benchmark. However, in System 2 deep thinking paradigm, the pass@1 accuracy alone does not fully reflect the reasoning capabilities for the policy model. To provide a more comprehensive evaluation, Fig.
|
| 2779 |
+
6
|
| 2780 |
+
and Fig.
|
| 2781 |
+
7
|
| 2782 |
+
present the pass@N accuracy. In this metric, the policy model generates
|
| 2783 |
+
N
|
| 2784 |
+
N
|
| 2785 |
+
solutions under two settings: random sampling (Fig.
|
| 2786 |
+
6
|
| 2787 |
+
) and PPM-augmented MCTS deep thinking (Fig.
|
| 2788 |
+
7
|
| 2789 |
+
). A problem is considered solved if one solution is correct.
|
| 2790 |
+
As shown in Figure
|
| 2791 |
+
6
|
| 2792 |
+
, our policy model achieves comparable Pass@64 accuracy with Qwen2.5-Math-7B-Instruct on MATH and significantly outperforms it on other benchmarks.
|
| 2793 |
+
Figure
|
| 2794 |
+
7
|
| 2795 |
+
further compares the pass@N accuracy of our four policy models (different sizes) after MCTS deep thinking under the same PPM guidance. We can observe that after generating 64 trajectories, the pass@N accuracy of different policy models becomes comparable.
|
| 2796 |
+
A.2
|
| 2797 |
+
Examples of Reasoning Trajectories
|
| 2798 |
+
Self-correction Example
|
| 2799 |
+
Question: The coordinates of a parallelogram are (5, 3), (6, 8), (7, 4) and
|
| 2800 |
+
(
|
| 2801 |
+
x
|
| 2802 |
+
,
|
| 2803 |
+
y
|
| 2804 |
+
)
|
| 2805 |
+
(x,y)
|
| 2806 |
+
and
|
| 2807 |
+
x
|
| 2808 |
+
>
|
| 2809 |
+
7
|
| 2810 |
+
x>7
|
| 2811 |
+
. What is the value of
|
| 2812 |
+
x
|
| 2813 |
+
+
|
| 2814 |
+
y
|
| 2815 |
+
x+y
|
| 2816 |
+
Our trajectory
|
| 2817 |
+
:
|
| 2818 |
+
<code>
|
| 2819 |
+
# Import necessary libraries
|
| 2820 |
+
⬇
|
| 2821 |
+
import
|
| 2822 |
+
numpy
|
| 2823 |
+
as
|
| 2824 |
+
np
|
| 2825 |
+
<end_of_step>
|
| 2826 |
+
# Define the coordinates of the parallelogram
|
| 2827 |
+
⬇
|
| 2828 |
+
A
|
| 2829 |
+
=
|
| 2830 |
+
np
|
| 2831 |
+
.
|
| 2832 |
+
array
|
| 2833 |
+
([5,
|
| 2834 |
+
3])
|
| 2835 |
+
B
|
| 2836 |
+
=
|
| 2837 |
+
np
|
| 2838 |
+
.
|
| 2839 |
+
array
|
| 2840 |
+
([6,
|
| 2841 |
+
8])
|
| 2842 |
+
C
|
| 2843 |
+
=
|
| 2844 |
+
np
|
| 2845 |
+
.
|
| 2846 |
+
array
|
| 2847 |
+
([7,
|
| 2848 |
+
4])
|
| 2849 |
+
<end_of_step>
|
| 2850 |
+
# Calculate the vector differences to find the missing coordinate D
|
| 2851 |
+
# In a parallelogram, the vector difference between two adjacent sides is equal to the vector difference between the other two adjacent sides
|
| 2852 |
+
# We can use the formula: D = A + (C - B)
|
| 2853 |
+
⬇
|
| 2854 |
+
D
|
| 2855 |
+
=
|
| 2856 |
+
A
|
| 2857 |
+
+
|
| 2858 |
+
(
|
| 2859 |
+
C
|
| 2860 |
+
-
|
| 2861 |
+
B
|
| 2862 |
+
)
|
| 2863 |
+
<end_of_step>
|
| 2864 |
+
# Since the coordinates provided are (5, 3), (6, 8), (7, 4) and (x, y), we can infer that the vector from B to C should be the same as the vector from A to the missing point (x, y)
|
| 2865 |
+
# This means: (x - 5, y - 3) = (1, -4) because the vector from B to C is (7 - 6, 4 - 8) = (1, -4)
|
| 2866 |
+
# Solve for x and y
|
| 2867 |
+
⬇
|
| 2868 |
+
x
|
| 2869 |
+
=
|
| 2870 |
+
5
|
| 2871 |
+
+
|
| 2872 |
+
1
|
| 2873 |
+
y
|
| 2874 |
+
=
|
| 2875 |
+
3
|
| 2876 |
+
+
|
| 2877 |
+
(-4)
|
| 2878 |
+
<end_of_step>
|
| 2879 |
+
# Verify that the point (x, y) exists and satisfies the condition x > 7
|
| 2880 |
+
# Since x = 6 is not greater than 7, we need to use the correct formula that maintains the parallelogram property and x > 7
|
| 2881 |
+
# Correcting the formula to find the other possible D [
|
| 2882 |
+
Self-correction!
|
| 2883 |
+
]
|
| 2884 |
+
⬇
|
| 2885 |
+
D
|
| 2886 |
+
=
|
| 2887 |
+
C
|
| 2888 |
+
+
|
| 2889 |
+
(
|
| 2890 |
+
B
|
| 2891 |
+
-
|
| 2892 |
+
A
|
| 2893 |
+
)
|
| 2894 |
+
<end_of_step>
|
| 2895 |
+
# Solve for x and y again
|
| 2896 |
+
⬇
|
| 2897 |
+
x
|
| 2898 |
+
=
|
| 2899 |
+
7
|
| 2900 |
+
+
|
| 2901 |
+
(6
|
| 2902 |
+
-
|
| 2903 |
+
5)
|
| 2904 |
+
y
|
| 2905 |
+
=
|
| 2906 |
+
4
|
| 2907 |
+
+
|
| 2908 |
+
(8
|
| 2909 |
+
-
|
| 2910 |
+
3)
|
| 2911 |
+
<end_of_step>
|
| 2912 |
+
# Verify that the point (x, y) exists and satisfies the condition x > 7
|
| 2913 |
+
# Since x = 8 is greater than 7, which satisfies the condition given, we can now find the sum of x and y
|
| 2914 |
+
⬇
|
| 2915 |
+
x_plus_y
|
| 2916 |
+
=
|
| 2917 |
+
x
|
| 2918 |
+
+
|
| 2919 |
+
y
|
| 2920 |
+
<end_of_step>
|
| 2921 |
+
# Now print the final answer
|
| 2922 |
+
⬇
|
| 2923 |
+
print
|
| 2924 |
+
(
|
| 2925 |
+
x_plus_y
|
| 2926 |
+
)
|
| 2927 |
+
# This should print the value of x + y
|
| 2928 |
+
<end_of_code>
|
| 2929 |
+
<output>17<end_of_output>
|
| 2930 |
+
<answer>The value of x + y is
|
| 2931 |
+
17
|
| 2932 |
+
<end_of_answer>
|
| 2933 |
+
Math statement proof using Fermat’s little theorem
|
| 2934 |
+
Question: If ( pt+a ) is a perfect square for some value of ( t ), then ( a
|
| 2935 |
+
p
|
| 2936 |
+
−
|
| 2937 |
+
1
|
| 2938 |
+
2
|
| 2939 |
+
{}^{\frac{p-1}{2}}
|
| 2940 |
+
- 1 ) is divisible by ( p ) (where ( p ) is a prime number).
|
| 2941 |
+
Our trajectory for proving
|
| 2942 |
+
:
|
| 2943 |
+
<code>
|
| 2944 |
+
# Step 1: Define the variables
|
| 2945 |
+
⬇
|
| 2946 |
+
from
|
| 2947 |
+
sympy
|
| 2948 |
+
import
|
| 2949 |
+
symbols
|
| 2950 |
+
,
|
| 2951 |
+
isprime
|
| 2952 |
+
,
|
| 2953 |
+
sqrt
|
| 2954 |
+
,
|
| 2955 |
+
mod_inverse
|
| 2956 |
+
# Define the variables
|
| 2957 |
+
⬇
|
| 2958 |
+
p
|
| 2959 |
+
,
|
| 2960 |
+
t
|
| 2961 |
+
,
|
| 2962 |
+
a
|
| 2963 |
+
=
|
| 2964 |
+
symbols
|
| 2965 |
+
(’
|
| 2966 |
+
p
|
| 2967 |
+
t
|
| 2968 |
+
a
|
| 2969 |
+
’)\\
|
| 2970 |
+
<end_of_step>
|
| 2971 |
+
# Step 2: Assume p*t + a is a perfect square
|
| 2972 |
+
# This means there exists an integer k such that p*t + a = k
|
| 2973 |
+
2
|
| 2974 |
+
# Rearrange to find a in terms of p, t, and k
|
| 2975 |
+
# a = k
|
| 2976 |
+
2
|
| 2977 |
+
- p*t
|
| 2978 |
+
<end_of_step>
|
| 2979 |
+
#
|
| 2980 |
+
Step 3: Use Fermat’s Little Theorem
|
| 2981 |
+
# Fermat’s Little Theorem states that if p is a prime and a is an integer not divisible by p,
|
| 2982 |
+
then a
|
| 2983 |
+
(p-1)
|
| 2984 |
+
≡
|
| 2985 |
+
\equiv
|
| 2986 |
+
1 (mod p)
|
| 2987 |
+
# We need to show that a
|
| 2988 |
+
((p-1)/2)
|
| 2989 |
+
- 1 is divisible by p
|
| 2990 |
+
<end_of_step>
|
| 2991 |
+
# Step 4: Substitute a = k
|
| 2992 |
+
2
|
| 2993 |
+
- p*t into Fermat’s Little Theorem
|
| 2994 |
+
# a
|
| 2995 |
+
((p-1)/2)
|
| 2996 |
+
≡
|
| 2997 |
+
\equiv
|
| 2998 |
+
(k
|
| 2999 |
+
2
|
| 3000 |
+
- p*t)
|
| 3001 |
+
((p-1)/2)
|
| 3002 |
+
(mod p)
|
| 3003 |
+
# Since p*t is divisible by p, (k
|
| 3004 |
+
2
|
| 3005 |
+
- p*t)
|
| 3006 |
+
((p-1)/2)
|
| 3007 |
+
≡
|
| 3008 |
+
\equiv
|
| 3009 |
+
k
|
| 3010 |
+
(p-1)
|
| 3011 |
+
(mod p)
|
| 3012 |
+
# By Fermat’s Little Theorem, k
|
| 3013 |
+
(p-1)
|
| 3014 |
+
≡
|
| 3015 |
+
\equiv
|
| 3016 |
+
1 (mod p)
|
| 3017 |
+
<end_of_step>
|
| 3018 |
+
# Step 5: Therefore, a
|
| 3019 |
+
((p-1)/2)
|
| 3020 |
+
≡
|
| 3021 |
+
\equiv
|
| 3022 |
+
1 (mod p)
|
| 3023 |
+
# This implies a
|
| 3024 |
+
((p-1)/2)
|
| 3025 |
+
- 1
|
| 3026 |
+
≡
|
| 3027 |
+
\equiv
|
| 3028 |
+
0 (mod p)
|
| 3029 |
+
# Hence, a
|
| 3030 |
+
((p-1)/2)
|
| 3031 |
+
- 1 is divisible by p
|
| 3032 |
+
<end_of_step>
|
| 3033 |
+
# Now print the final answer
|
| 3034 |
+
print("The expression a
|
| 3035 |
+
((p-1)/2)
|
| 3036 |
+
- 1 is divisible by p.")
|
| 3037 |
+
<end_of_code>
|
| 3038 |
+
<output>
|
| 3039 |
+
The expression
|
| 3040 |
+
a
|
| 3041 |
+
(
|
| 3042 |
+
(
|
| 3043 |
+
p
|
| 3044 |
+
−
|
| 3045 |
+
1
|
| 3046 |
+
)
|
| 3047 |
+
/
|
| 3048 |
+
2
|
| 3049 |
+
)
|
| 3050 |
+
a^{((p-1)/2)}
|
| 3051 |
+
- 1 is divisible by p.
|
| 3052 |
+
The statement is true.
|
| 3053 |
+
<end_of_output>
|
| 3054 |
+
Problem-solving with application of the Vieta’s formulas and the AM-GM inequality
|
| 3055 |
+
Question: Let
|
| 3056 |
+
a
|
| 3057 |
+
a
|
| 3058 |
+
be a positive real number such that all the roots of
|
| 3059 |
+
x
|
| 3060 |
+
3
|
| 3061 |
+
+
|
| 3062 |
+
a
|
| 3063 |
+
|
| 3064 |
+
x
|
| 3065 |
+
2
|
| 3066 |
+
+
|
| 3067 |
+
a
|
| 3068 |
+
|
| 3069 |
+
x
|
| 3070 |
+
+
|
| 3071 |
+
1
|
| 3072 |
+
=
|
| 3073 |
+
0
|
| 3074 |
+
x^{3}+ax^{2}+ax+1=0
|
| 3075 |
+
are real. Find the smallest possible value of
|
| 3076 |
+
a
|
| 3077 |
+
.
|
| 3078 |
+
a.
|
| 3079 |
+
Our trajectory for proving
|
| 3080 |
+
:
|
| 3081 |
+
<code>
|
| 3082 |
+
# Step 1: Define the polynomial equation
|
| 3083 |
+
from sympy import symbols, solve, Eq
|
| 3084 |
+
# Define the variable and the parameter
|
| 3085 |
+
x, a = symbols(’x a’)
|
| 3086 |
+
<end_of_step>
|
| 3087 |
+
[PPM Score: -0.46430]
|
| 3088 |
+
# Step 2: Define the polynomial equation
|
| 3089 |
+
polynomial = x**3 + a*x**2 + a*x + 1
|
| 3090 |
+
<end_of_step>
|
| 3091 |
+
[PPM Score: 0.53937]
|
| 3092 |
+
#
|
| 3093 |
+
Step 3: Use Vieta’s formulas to relate the coefficients to the roots
|
| 3094 |
+
# Since the polynomial is monic (leading coefficient is 1), the sum of the roots (taken one at a time) is -a
|
| 3095 |
+
# The sum of the products of the roots taken two at a time is a
|
| 3096 |
+
# The product of the roots is -1
|
| 3097 |
+
# Since all roots are real, we can denote them as r1, r2, r3
|
| 3098 |
+
# From Vieta’s formulas:
|
| 3099 |
+
# r1 + r2 + r3 = -a
|
| 3100 |
+
# r1*r2 + r2*r3 + r3*r1 = a
|
| 3101 |
+
# r1*r2*r3 = -1
|
| 3102 |
+
<end_of_step>
|
| 3103 |
+
[PPM Score:
|
| 3104 |
+
0.9989]
|
| 3105 |
+
#
|
| 3106 |
+
Step 4: Use the AM-GM inequality to find the minimum value of a
|
| 3107 |
+
# By AM-GM, (r1*r2 + r2*r3 + r3*r1) / 3 >= (r1*r2*r3)
|
| 3108 |
+
(2/3)
|
| 3109 |
+
# Substituting the known values:
|
| 3110 |
+
# a / 3 >= (-1)
|
| 3111 |
+
(2/3)
|
| 3112 |
+
# a / 3 >= 1
|
| 3113 |
+
# a >= 3
|
| 3114 |
+
<end_of_step>
|
| 3115 |
+
[PPM Score:
|
| 3116 |
+
0.9999]
|
| 3117 |
+
# Step 5: Verify if a = 3 is achievable
|
| 3118 |
+
# If a = 3, then r1 + r2 + r3 = -3 and r1*r2 + r2*r3 + r3*r1 = 3
|
| 3119 |
+
# We need to check if there exist real roots that satisfy these conditions
|
| 3120 |
+
# Consider the case where r1 = r2 = r3 = -1
|
| 3121 |
+
# Then r1 + r2 + r3 = -3 and r1*r2 + r2*r3 + r3*r1 = 3, and r1*r2*r3 = -1
|
| 3122 |
+
# Thus, a = 3 is achievable with real roots
|
| 3123 |
+
<end_of_step>
|
| 3124 |
+
[PPM Score: 1.0]
|
| 3125 |
+
# Step 6: Print the final answer
|
| 3126 |
+
print(3)
|
| 3127 |
+
<end_of_code>
|
| 3128 |
+
<output>3<end_of_output>
|
| 3129 |
+
<answer>The smallest possible value of a is
|
| 3130 |
+
3
|
| 3131 |
+
.
|
| 3132 |
+
<end_of_answer>
|
| 3133 |
+
Problem-solving with application the Shoelace Theorem
|
| 3134 |
+
Question: In the diagram,
|
| 3135 |
+
D
|
| 3136 |
+
D
|
| 3137 |
+
and
|
| 3138 |
+
E
|
| 3139 |
+
E
|
| 3140 |
+
are the midpoints of
|
| 3141 |
+
A
|
| 3142 |
+
|
| 3143 |
+
B
|
| 3144 |
+
¯
|
| 3145 |
+
\overline{AB}
|
| 3146 |
+
and
|
| 3147 |
+
B
|
| 3148 |
+
|
| 3149 |
+
C
|
| 3150 |
+
¯
|
| 3151 |
+
\overline{BC}
|
| 3152 |
+
respectively. Determine the area of quadrilateral
|
| 3153 |
+
D
|
| 3154 |
+
|
| 3155 |
+
B
|
| 3156 |
+
|
| 3157 |
+
E
|
| 3158 |
+
|
| 3159 |
+
F
|
| 3160 |
+
DBEF
|
| 3161 |
+
. [asy]
|
| 3162 |
+
size(180); defaultpen(linewidth(.7pt)+fontsize(10pt));
|
| 3163 |
+
pair A, B, C, D, E, F;
|
| 3164 |
+
A=(0,6);
|
| 3165 |
+
B=(0,0);
|
| 3166 |
+
C=(8,0);
|
| 3167 |
+
D=(0,3);
|
| 3168 |
+
E=(4,0);
|
| 3169 |
+
F=(8/3,2);
|
| 3170 |
+
draw(E–A–C–D);
|
| 3171 |
+
draw((-1,0)–(10,0), EndArrow);
|
| 3172 |
+
draw((0,-1)–(0,8), EndArrow);
|
| 3173 |
+
label("
|
| 3174 |
+
A
|
| 3175 |
+
|
| 3176 |
+
(
|
| 3177 |
+
0
|
| 3178 |
+
,
|
| 3179 |
+
6
|
| 3180 |
+
)
|
| 3181 |
+
A(0,6)
|
| 3182 |
+
", A, W);
|
| 3183 |
+
label("
|
| 3184 |
+
B
|
| 3185 |
+
|
| 3186 |
+
(
|
| 3187 |
+
0
|
| 3188 |
+
,
|
| 3189 |
+
0
|
| 3190 |
+
)
|
| 3191 |
+
B(0,0)
|
| 3192 |
+
", B, SW);
|
| 3193 |
+
label("
|
| 3194 |
+
C
|
| 3195 |
+
|
| 3196 |
+
(
|
| 3197 |
+
8
|
| 3198 |
+
,
|
| 3199 |
+
0
|
| 3200 |
+
)
|
| 3201 |
+
C(8,0)
|
| 3202 |
+
", C, S);
|
| 3203 |
+
label("
|
| 3204 |
+
D
|
| 3205 |
+
D
|
| 3206 |
+
", D, W);
|
| 3207 |
+
label("
|
| 3208 |
+
E
|
| 3209 |
+
E
|
| 3210 |
+
", E, S);
|
| 3211 |
+
label("
|
| 3212 |
+
F
|
| 3213 |
+
F
|
| 3214 |
+
", F, SW);
|
| 3215 |
+
label("
|
| 3216 |
+
x
|
| 3217 |
+
x
|
| 3218 |
+
", (10,0), dir(0));
|
| 3219 |
+
label("
|
| 3220 |
+
y
|
| 3221 |
+
y
|
| 3222 |
+
", (0,8), dir(90));
|
| 3223 |
+
[/asy]
|
| 3224 |
+
Our trajectory for proving
|
| 3225 |
+
:
|
| 3226 |
+
<code>
|
| 3227 |
+
# Step 1: Define the coordinates of points A, B, C, D, and E
|
| 3228 |
+
⬇
|
| 3229 |
+
A
|
| 3230 |
+
=
|
| 3231 |
+
(0,
|
| 3232 |
+
6)
|
| 3233 |
+
B
|
| 3234 |
+
=
|
| 3235 |
+
(0,
|
| 3236 |
+
0)
|
| 3237 |
+
C
|
| 3238 |
+
=
|
| 3239 |
+
(8,
|
| 3240 |
+
0)
|
| 3241 |
+
D
|
| 3242 |
+
=
|
| 3243 |
+
(0,
|
| 3244 |
+
3)
|
| 3245 |
+
#
|
| 3246 |
+
Midpoint
|
| 3247 |
+
of
|
| 3248 |
+
AB
|
| 3249 |
+
E
|
| 3250 |
+
=
|
| 3251 |
+
(4,
|
| 3252 |
+
0)
|
| 3253 |
+
#
|
| 3254 |
+
Midpoint
|
| 3255 |
+
of
|
| 3256 |
+
BC
|
| 3257 |
+
<end_of_step>
|
| 3258 |
+
[PPM Score: -0.76639]
|
| 3259 |
+
# Step 2: Define the coordinates of point F using the section formula (known to be (8/3, 2))
|
| 3260 |
+
⬇
|
| 3261 |
+
F
|
| 3262 |
+
=
|
| 3263 |
+
(8/3,
|
| 3264 |
+
2)
|
| 3265 |
+
<end_of_step>
|
| 3266 |
+
[PPM Score: 0.8429]
|
| 3267 |
+
#
|
| 3268 |
+
Step 3: Use the Shoelace Theorem to find the area of quadrilateral DBEF
|
| 3269 |
+
# Shoelace Theorem formula: Area = 0.5 * |x1y2 + x2y3 + x3y4 + x4y1 - (y1x2 + y2x3 + y3x4 + y4x1)|
|
| 3270 |
+
# Coordinates of D, B, E, F are (0, 3), (0, 0), (4, 0), (8/3, 2) respectively
|
| 3271 |
+
⬇
|
| 3272 |
+
x1
|
| 3273 |
+
,
|
| 3274 |
+
y1
|
| 3275 |
+
=
|
| 3276 |
+
D
|
| 3277 |
+
x2
|
| 3278 |
+
,
|
| 3279 |
+
y2
|
| 3280 |
+
=
|
| 3281 |
+
B
|
| 3282 |
+
x3
|
| 3283 |
+
,
|
| 3284 |
+
y3
|
| 3285 |
+
=
|
| 3286 |
+
E
|
| 3287 |
+
x4
|
| 3288 |
+
,
|
| 3289 |
+
y4
|
| 3290 |
+
=
|
| 3291 |
+
F
|
| 3292 |
+
area
|
| 3293 |
+
=
|
| 3294 |
+
0.5
|
| 3295 |
+
*
|
| 3296 |
+
abs
|
| 3297 |
+
(
|
| 3298 |
+
x1
|
| 3299 |
+
*
|
| 3300 |
+
y2
|
| 3301 |
+
+
|
| 3302 |
+
x2
|
| 3303 |
+
*
|
| 3304 |
+
y3
|
| 3305 |
+
+
|
| 3306 |
+
x3
|
| 3307 |
+
*
|
| 3308 |
+
y4
|
| 3309 |
+
+
|
| 3310 |
+
x4
|
| 3311 |
+
*
|
| 3312 |
+
y1
|
| 3313 |
+
-
|
| 3314 |
+
(
|
| 3315 |
+
y1
|
| 3316 |
+
*
|
| 3317 |
+
x2
|
| 3318 |
+
+
|
| 3319 |
+
y2
|
| 3320 |
+
*
|
| 3321 |
+
x3
|
| 3322 |
+
+
|
| 3323 |
+
y3
|
| 3324 |
+
*
|
| 3325 |
+
x4
|
| 3326 |
+
+
|
| 3327 |
+
y4
|
| 3328 |
+
*
|
| 3329 |
+
x1
|
| 3330 |
+
))
|
| 3331 |
+
<end_of_step>
|
| 3332 |
+
[PPM Score:
|
| 3333 |
+
0.9972]
|
| 3334 |
+
# Step 4: Print the area of quadrilateral DBEF
|
| 3335 |
+
⬇
|
| 3336 |
+
print
|
| 3337 |
+
(
|
| 3338 |
+
area
|
| 3339 |
+
)
|
| 3340 |
+
<end_of_code>
|
| 3341 |
+
<output>8.0<end_of_output>
|
| 3342 |
+
<answer>
|
| 3343 |
+
8
|
| 3344 |
+
<end_of_answer>
|
| 3345 |
+
A.3
|
| 3346 |
+
Prompt Examples
|
| 3347 |
+
MCTS Prompt for DeepSeek-V2-Coder-Instruct (bootstrap round)
|
| 3348 |
+
You are a powerful agent with broad math knowledge and great python programming skills. You need to use python interpreter to do accurate calculation on math equations.
|
| 3349 |
+
!!! Remember:
|
| 3350 |
+
1. Use code solve the problem step by step. The solution should include three parts: <code>, <output>, and <answer>.
|
| 3351 |
+
2. All calculations should be done in python code. Provide concise reasoning and thinking in the comments of the code.
|
| 3352 |
+
3. The most related python packages include ‘math‘, ‘sympy‘, ‘scipy‘, and ‘numpy‘.
|
| 3353 |
+
4. Please use the following template:
|
| 3354 |
+
Question: the input question
|
| 3355 |
+
<code>Construct the code step by step. Use <end_of_step> to indicate the end of each step. Ensure your code can execute correctly(excluding <end_of_step>) and print the answer. Avoid undefined variables (NameError), unimported packages, or formatting errors (SyntaxError, TypeError). In the last step of the code, print the final answer and add a comment: Now print the final answer.<end_of_code>
|
| 3356 |
+
<output>Execute the code in using the Python interpreter and display the printed results.<end_of_output>
|
| 3357 |
+
<answer>The concise answer without verbose context, put your final answer’s numerical part (without unit, only focus on the numerical part if it’s a choice question) in
|
| 3358 |
+
boxed.<end_of_answer> Now! It’s your turn.
|
| 3359 |
+
Question:
|
| 3360 |
+
{input}
|
| 3361 |
+
The following are 2 demonstration examples:
|
| 3362 |
+
Question: Terrell usually lifts two 20-pound weights 12 times. If he uses two 15-pound weights instead, how many times must Terrell lift them in order to lift the same total weight?
|
| 3363 |
+
<code>
|
| 3364 |
+
# Step 1: Calculate the total weight lifted with two 20-pound weights
|
| 3365 |
+
total_weight_20 = 2 * 20 * 12
|
| 3366 |
+
<end_of_step>
|
| 3367 |
+
# Step 2: Calculate the weight lifted per repetition with two 15-pound weights
|
| 3368 |
+
weight_per_rep_15 = 2 * 15
|
| 3369 |
+
<end_of_step>
|
| 3370 |
+
# Step 3: Calculate the number of repetitions needed to lift the same total weight with two 15-pound weights
|
| 3371 |
+
reps_needed = total_weight_20 / weight_per_rep_15
|
| 3372 |
+
<end_of_step>
|
| 3373 |
+
# Now print the final answer
|
| 3374 |
+
print(reps_needed)
|
| 3375 |
+
<end_of_code>
|
| 3376 |
+
<output>16.0 <end_of_output> <answer>From the result, we can see that Terrell must lift the 15-pound weights
|
| 3377 |
+
boxed16 times to lift the same total weight.
|
| 3378 |
+
<end_of_answer>,
|
| 3379 |
+
Question: Find the value of
|
| 3380 |
+
x
|
| 3381 |
+
x
|
| 3382 |
+
that satisfies
|
| 3383 |
+
3
|
| 3384 |
+
|
| 3385 |
+
x
|
| 3386 |
+
+
|
| 3387 |
+
5
|
| 3388 |
+
6
|
| 3389 |
+
|
| 3390 |
+
x
|
| 3391 |
+
+
|
| 3392 |
+
5
|
| 3393 |
+
=
|
| 3394 |
+
5
|
| 3395 |
+
3
|
| 3396 |
+
\frac{\sqrt{3x+5}}{\sqrt{6x+5}}=\frac{\sqrt{5}}{3}
|
| 3397 |
+
. Express your answer as a common fraction.
|
| 3398 |
+
<code>
|
| 3399 |
+
from sympy import symbols, Eq, solve, sqrt
|
| 3400 |
+
# Define the variable x
|
| 3401 |
+
x = symbols(’x’)
|
| 3402 |
+
<end_of_step>
|
| 3403 |
+
# Define the equation
|
| 3404 |
+
equation = Eq(sqrt(3*x + 5) / sqrt(6*x + 5), sqrt(5) / 3)
|
| 3405 |
+
<end_of_step>
|
| 3406 |
+
# Solve the equation for x
|
| 3407 |
+
solution = solve(equation, x)
|
| 3408 |
+
<end_of_step>
|
| 3409 |
+
# Now print the final answer
|
| 3410 |
+
print(solution)
|
| 3411 |
+
<end_of_code>
|
| 3412 |
+
<output>[20/3]
|
| 3413 |
+
<end_of_output> <answer>From the result, we can see that the value of x is
|
| 3414 |
+
20
|
| 3415 |
+
3
|
| 3416 |
+
<end_of_answer>
|
| 3417 |
+
MCTS Prompt for our policy SLM
|
| 3418 |
+
MCTS prompt for next step generation
|
| 3419 |
+
:
|
| 3420 |
+
<|user|>:
|
| 3421 |
+
{input question}
|
| 3422 |
+
<|assistant|>: Let’s think step by step and solve the problem with code.
|
| 3423 |
+
{trajectory of previous steps}
|
| 3424 |
+
SFT Prompt Example
|
| 3425 |
+
:
|
| 3426 |
+
<|user|>:
|
| 3427 |
+
How many numbers can you get by multiplying two or more distinct members of the set
|
| 3428 |
+
1
|
| 3429 |
+
,
|
| 3430 |
+
2
|
| 3431 |
+
,
|
| 3432 |
+
3
|
| 3433 |
+
,
|
| 3434 |
+
5
|
| 3435 |
+
,
|
| 3436 |
+
11
|
| 3437 |
+
\\
|
| 3438 |
+
{1,2,3,5,11\\
|
| 3439 |
+
}
|
| 3440 |
+
together?
|
| 3441 |
+
<|assistant|>: Let’s think step by step and solve the problem with code.
|
| 3442 |
+
<code>
|
| 3443 |
+
# Step 1: Create the set of numbers
|
| 3444 |
+
⬇
|
| 3445 |
+
numbers
|
| 3446 |
+
=
|
| 3447 |
+
{1,
|
| 3448 |
+
2,
|
| 3449 |
+
3,
|
| 3450 |
+
5,
|
| 3451 |
+
11}
|
| 3452 |
+
<end_of_step>
|
| 3453 |
+
# Step 2: Define a function to generate all possible products of two or more distinct members of the set
|
| 3454 |
+
⬇
|
| 3455 |
+
from
|
| 3456 |
+
itertools
|
| 3457 |
+
import
|
| 3458 |
+
combinations
|
| 3459 |
+
def
|
| 3460 |
+
generate_products
|
| 3461 |
+
(
|
| 3462 |
+
numbers
|
| 3463 |
+
):
|
| 3464 |
+
products
|
| 3465 |
+
=
|
| 3466 |
+
set
|
| 3467 |
+
()
|
| 3468 |
+
for
|
| 3469 |
+
r
|
| 3470 |
+
in
|
| 3471 |
+
range
|
| 3472 |
+
(2,
|
| 3473 |
+
len
|
| 3474 |
+
(
|
| 3475 |
+
numbers
|
| 3476 |
+
)
|
| 3477 |
+
+
|
| 3478 |
+
1):
|
| 3479 |
+
#
|
| 3480 |
+
For
|
| 3481 |
+
combinations
|
| 3482 |
+
of
|
| 3483 |
+
2,
|
| 3484 |
+
3,
|
| 3485 |
+
4,
|
| 3486 |
+
and
|
| 3487 |
+
5
|
| 3488 |
+
elements
|
| 3489 |
+
for
|
| 3490 |
+
combo
|
| 3491 |
+
in
|
| 3492 |
+
combinations
|
| 3493 |
+
(
|
| 3494 |
+
numbers
|
| 3495 |
+
,
|
| 3496 |
+
r
|
| 3497 |
+
):
|
| 3498 |
+
product
|
| 3499 |
+
=
|
| 3500 |
+
1
|
| 3501 |
+
for
|
| 3502 |
+
num
|
| 3503 |
+
in
|
| 3504 |
+
combo
|
| 3505 |
+
:
|
| 3506 |
+
product
|
| 3507 |
+
*=
|
| 3508 |
+
num
|
| 3509 |
+
products
|
| 3510 |
+
.
|
| 3511 |
+
add
|
| 3512 |
+
(
|
| 3513 |
+
product
|
| 3514 |
+
)
|
| 3515 |
+
return
|
| 3516 |
+
products
|
| 3517 |
+
<end_of_step>
|
| 3518 |
+
# Step 3: Calculate the products
|
| 3519 |
+
⬇
|
| 3520 |
+
unique_products
|
| 3521 |
+
=
|
| 3522 |
+
generate_products
|
| 3523 |
+
(
|
| 3524 |
+
numbers
|
| 3525 |
+
)
|
| 3526 |
+
<end_of_step>
|
| 3527 |
+
# Step 4: Calculate the number of unique products
|
| 3528 |
+
⬇
|
| 3529 |
+
num_unique_products
|
| 3530 |
+
=
|
| 3531 |
+
len
|
| 3532 |
+
(
|
| 3533 |
+
unique_products
|
| 3534 |
+
)
|
| 3535 |
+
<end_of_step>
|
| 3536 |
+
# Now print the final answer
|
| 3537 |
+
⬇
|
| 3538 |
+
print
|
| 3539 |
+
(
|
| 3540 |
+
num_unique_products
|
| 3541 |
+
)
|
| 3542 |
+
<end_of_code>
|
| 3543 |
+
<output>15<end_of_output>
|
| 3544 |
+
<answer>From the result, we can see that there are
|
| 3545 |
+
15
|
| 3546 |
+
unique products.
|
| 3547 |
+
<end_of_answer>
|
| 3548 |
+
◄
|
| 3549 |
+
Feeling
|
| 3550 |
+
lucky?
|
| 3551 |
+
Conversion
|
| 3552 |
+
report
|
| 3553 |
+
Report
|
| 3554 |
+
an issue
|
| 3555 |
+
View original
|
| 3556 |
+
on arXiv
|
| 3557 |
+
►
|
|
@@ -0,0 +1,203 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2501.09136] Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG'
|
| 3 |
+
id: 250109136-agentic-retrieval-augmented-generation-a-survey-on-agentic-rag
|
| 4 |
+
tags:
|
| 5 |
+
- deepread
|
| 6 |
+
created: '2026-06-10T00:24:47.557837Z'
|
| 7 |
+
source: https://arxiv.org/abs/2501.09136
|
| 8 |
+
source_domain: arxiv.org
|
| 9 |
+
fetched_at: '2026-06-10T00:24:47.557707Z'
|
| 10 |
+
fetch_provider: builtin
|
| 11 |
+
status: draft
|
| 12 |
+
type: note
|
| 13 |
+
tier: institutional
|
| 14 |
+
content_type: paper
|
| 15 |
+
deprecated: false
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
[2501.09136] Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
|
| 19 |
+
Computer Science > Artificial Intelligence
|
| 20 |
+
arXiv:2501.09136
|
| 21 |
+
(cs)
|
| 22 |
+
[Submitted on 15 Jan 2025 (
|
| 23 |
+
v1
|
| 24 |
+
), last revised 1 Apr 2026 (this version, v4)]
|
| 25 |
+
Title:
|
| 26 |
+
Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
|
| 27 |
+
Authors:
|
| 28 |
+
Aditi Singh
|
| 29 |
+
,
|
| 30 |
+
Abul Ehtesham
|
| 31 |
+
,
|
| 32 |
+
Saket Kumar
|
| 33 |
+
,
|
| 34 |
+
Tala Talaei Khoei
|
| 35 |
+
,
|
| 36 |
+
Athanasios V. Vasilakos
|
| 37 |
+
View a PDF of the paper titled Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG, by Aditi Singh and 4 other authors
|
| 38 |
+
View PDF
|
| 39 |
+
HTML (experimental)
|
| 40 |
+
Abstract:
|
| 41 |
+
Large Language Models (LLMs) have advanced artificial intelligence by enabling human-like text generation and natural language understanding. However, their reliance on static training data limits their ability to respond to dynamic, real-time queries, resulting in outdated or inaccurate outputs. Retrieval-Augmented Generation (RAG) has emerged as a solution, enhancing LLMs by integrating real-time data retrieval to provide contextually relevant and up-to-date responses. Despite its promise, traditional RAG systems are constrained by static workflows and lack the adaptability required for multi-step reasoning and complex task management. Agentic Retrieval-Augmented Generation (Agentic RAG) transcends these limitations by embedding autonomous AI agents into the RAG pipeline. These agents leverage agentic design patterns reflection, planning, tool use, and multi-agent collaboration to dynamically manage retrieval strategies, iteratively refine contextual understanding, and adapt workflows through operational structures ranging from sequential steps to adaptive collaboration. This integration enables Agentic RAG systems to deliver flexibility, scalability, and context-awareness across diverse applications. This paper presents an analytical survey of Agentic RAG systems. It traces the evolution of RAG paradigms, introduces a principled taxonomy of Agentic RAG architectures based on agent cardinality, control structure, autonomy, and knowledge representation, and provides a comparative analysis of design trade-offs across existing frameworks. The survey examines applications in healthcare, finance, education, and enterprise document processing, and distills practical lessons for system designers and practitioners. Finally, it identifies key open research challenges related to evaluation, coordination, memory management, efficiency, and governance, outlining directions for future research.
|
| 42 |
+
Subjects:
|
| 43 |
+
Artificial Intelligence (cs.AI)
|
| 44 |
+
; Computation and Language (cs.CL); Information Retrieval (cs.IR)
|
| 45 |
+
Cite as:
|
| 46 |
+
arXiv:2501.09136
|
| 47 |
+
[cs.AI]
|
| 48 |
+
(or
|
| 49 |
+
arXiv:2501.09136v4
|
| 50 |
+
[cs.AI]
|
| 51 |
+
for this version)
|
| 52 |
+
https://doi.org/10.48550/arXiv.2501.09136
|
| 53 |
+
Focus to learn more
|
| 54 |
+
arXiv-issued DOI via DataCite
|
| 55 |
+
Submission history
|
| 56 |
+
From: Abul Ehtesham [
|
| 57 |
+
view email
|
| 58 |
+
]
|
| 59 |
+
[v1]
|
| 60 |
+
Wed, 15 Jan 2025 20:40:25 UTC (20,962 KB)
|
| 61 |
+
[v2]
|
| 62 |
+
Mon, 3 Feb 2025 04:01:36 UTC (22,453 KB)
|
| 63 |
+
[v3]
|
| 64 |
+
Tue, 4 Feb 2025 04:48:00 UTC (22,430 KB)
|
| 65 |
+
[v4]
|
| 66 |
+
Wed, 1 Apr 2026 15:51:06 UTC (13,996 KB)
|
| 67 |
+
Full-text links:
|
| 68 |
+
Access Paper:
|
| 69 |
+
View a PDF of the paper titled Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG, by Aditi Singh and 4 other authors
|
| 70 |
+
View PDF
|
| 71 |
+
HTML (experimental)
|
| 72 |
+
TeX Source
|
| 73 |
+
view license
|
| 74 |
+
Current browse context:
|
| 75 |
+
cs.AI
|
| 76 |
+
< prev
|
| 77 |
+
|
|
| 78 |
+
next >
|
| 79 |
+
new
|
| 80 |
+
|
|
| 81 |
+
recent
|
| 82 |
+
|
|
| 83 |
+
2025-01
|
| 84 |
+
Change to browse by:
|
| 85 |
+
cs
|
| 86 |
+
cs.CL
|
| 87 |
+
cs.IR
|
| 88 |
+
References & Citations
|
| 89 |
+
NASA ADS
|
| 90 |
+
Google Scholar
|
| 91 |
+
Semantic Scholar
|
| 92 |
+
export BibTeX citation
|
| 93 |
+
Loading...
|
| 94 |
+
BibTeX formatted citation
|
| 95 |
+
×
|
| 96 |
+
loading...
|
| 97 |
+
Data provided by:
|
| 98 |
+
Bookmark
|
| 99 |
+
Bibliographic Tools
|
| 100 |
+
Bibliographic and Citation Tools
|
| 101 |
+
Bibliographic Explorer Toggle
|
| 102 |
+
Bibliographic Explorer
|
| 103 |
+
(
|
| 104 |
+
What is the Explorer?
|
| 105 |
+
)
|
| 106 |
+
Connected Papers Toggle
|
| 107 |
+
Connected Papers
|
| 108 |
+
(
|
| 109 |
+
What is Connected Papers?
|
| 110 |
+
)
|
| 111 |
+
Litmaps Toggle
|
| 112 |
+
Litmaps
|
| 113 |
+
(
|
| 114 |
+
What is Litmaps?
|
| 115 |
+
)
|
| 116 |
+
scite.ai Toggle
|
| 117 |
+
scite Smart Citations
|
| 118 |
+
(
|
| 119 |
+
What are Smart Citations?
|
| 120 |
+
)
|
| 121 |
+
Code, Data, Media
|
| 122 |
+
Code, Data and Media Associated with this Article
|
| 123 |
+
alphaXiv Toggle
|
| 124 |
+
alphaXiv
|
| 125 |
+
(
|
| 126 |
+
What is alphaXiv?
|
| 127 |
+
)
|
| 128 |
+
Links to Code Toggle
|
| 129 |
+
CatalyzeX Code Finder for Papers
|
| 130 |
+
(
|
| 131 |
+
What is CatalyzeX?
|
| 132 |
+
)
|
| 133 |
+
DagsHub Toggle
|
| 134 |
+
DagsHub
|
| 135 |
+
(
|
| 136 |
+
What is DagsHub?
|
| 137 |
+
)
|
| 138 |
+
GotitPub Toggle
|
| 139 |
+
Gotit.pub
|
| 140 |
+
(
|
| 141 |
+
What is GotitPub?
|
| 142 |
+
)
|
| 143 |
+
Huggingface Toggle
|
| 144 |
+
Hugging Face
|
| 145 |
+
(
|
| 146 |
+
What is Huggingface?
|
| 147 |
+
)
|
| 148 |
+
Links to Code Toggle
|
| 149 |
+
Papers with Code
|
| 150 |
+
(
|
| 151 |
+
What is Papers with Code?
|
| 152 |
+
)
|
| 153 |
+
ScienceCast Toggle
|
| 154 |
+
ScienceCast
|
| 155 |
+
(
|
| 156 |
+
What is ScienceCast?
|
| 157 |
+
)
|
| 158 |
+
Demos
|
| 159 |
+
Demos
|
| 160 |
+
Replicate Toggle
|
| 161 |
+
Replicate
|
| 162 |
+
(
|
| 163 |
+
What is Replicate?
|
| 164 |
+
)
|
| 165 |
+
Spaces Toggle
|
| 166 |
+
Hugging Face Spaces
|
| 167 |
+
(
|
| 168 |
+
What is Spaces?
|
| 169 |
+
)
|
| 170 |
+
Spaces Toggle
|
| 171 |
+
TXYZ.AI
|
| 172 |
+
(
|
| 173 |
+
What is TXYZ.AI?
|
| 174 |
+
)
|
| 175 |
+
Related Papers
|
| 176 |
+
Recommenders and Search Tools
|
| 177 |
+
Link to Influence Flower
|
| 178 |
+
Influence Flower
|
| 179 |
+
(
|
| 180 |
+
What are Influence Flowers?
|
| 181 |
+
)
|
| 182 |
+
Core recommender toggle
|
| 183 |
+
CORE Recommender
|
| 184 |
+
(
|
| 185 |
+
What is CORE?
|
| 186 |
+
)
|
| 187 |
+
Author
|
| 188 |
+
Venue
|
| 189 |
+
Institution
|
| 190 |
+
Topic
|
| 191 |
+
About arXivLabs
|
| 192 |
+
arXivLabs: experimental projects with community collaborators
|
| 193 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 194 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 195 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 196 |
+
Learn more about arXivLabs
|
| 197 |
+
.
|
| 198 |
+
Which authors of this paper are endorsers?
|
| 199 |
+
|
|
| 200 |
+
Disable MathJax
|
| 201 |
+
(
|
| 202 |
+
What is MathJax?
|
| 203 |
+
)
|
|
@@ -0,0 +1,196 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2501.09891] Evolving Deeper LLM Thinking'
|
| 3 |
+
id: 250109891-evolving-deeper-llm-thinking
|
| 4 |
+
tags:
|
| 5 |
+
- deepread
|
| 6 |
+
created: '2026-06-10T00:24:58.674469Z'
|
| 7 |
+
source: https://arxiv.org/abs/2501.09891
|
| 8 |
+
source_domain: arxiv.org
|
| 9 |
+
fetched_at: '2026-06-10T00:24:58.674344Z'
|
| 10 |
+
fetch_provider: builtin
|
| 11 |
+
status: draft
|
| 12 |
+
type: note
|
| 13 |
+
tier: institutional
|
| 14 |
+
content_type: paper
|
| 15 |
+
deprecated: false
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
[2501.09891] Evolving Deeper LLM Thinking
|
| 19 |
+
Computer Science > Artificial Intelligence
|
| 20 |
+
arXiv:2501.09891
|
| 21 |
+
(cs)
|
| 22 |
+
[Submitted on 17 Jan 2025]
|
| 23 |
+
Title:
|
| 24 |
+
Evolving Deeper LLM Thinking
|
| 25 |
+
Authors:
|
| 26 |
+
Kuang-Huei Lee
|
| 27 |
+
,
|
| 28 |
+
Ian Fischer
|
| 29 |
+
,
|
| 30 |
+
Yueh-Hua Wu
|
| 31 |
+
,
|
| 32 |
+
Dave Marwood
|
| 33 |
+
,
|
| 34 |
+
Shumeet Baluja
|
| 35 |
+
,
|
| 36 |
+
Dale Schuurmans
|
| 37 |
+
,
|
| 38 |
+
Xinyun Chen
|
| 39 |
+
View a PDF of the paper titled Evolving Deeper LLM Thinking, by Kuang-Huei Lee and 6 other authors
|
| 40 |
+
View PDF
|
| 41 |
+
HTML (experimental)
|
| 42 |
+
Abstract:
|
| 43 |
+
We explore an evolutionary search strategy for scaling inference time compute in Large Language Models. The proposed approach, Mind Evolution, uses a language model to generate, recombine and refine candidate responses. The proposed approach avoids the need to formalize the underlying inference problem whenever a solution evaluator is available. Controlling for inference cost, we find that Mind Evolution significantly outperforms other inference strategies such as Best-of-N and Sequential Revision in natural language planning tasks. In the TravelPlanner and Natural Plan benchmarks, Mind Evolution solves more than 98% of the problem instances using Gemini 1.5 Pro without the use of a formal solver.
|
| 44 |
+
Subjects:
|
| 45 |
+
Artificial Intelligence (cs.AI)
|
| 46 |
+
Cite as:
|
| 47 |
+
arXiv:2501.09891
|
| 48 |
+
[cs.AI]
|
| 49 |
+
(or
|
| 50 |
+
arXiv:2501.09891v1
|
| 51 |
+
[cs.AI]
|
| 52 |
+
for this version)
|
| 53 |
+
https://doi.org/10.48550/arXiv.2501.09891
|
| 54 |
+
Focus to learn more
|
| 55 |
+
arXiv-issued DOI via DataCite
|
| 56 |
+
Submission history
|
| 57 |
+
From: Dale Schuurmans [
|
| 58 |
+
view email
|
| 59 |
+
]
|
| 60 |
+
[v1]
|
| 61 |
+
Fri, 17 Jan 2025 00:41:44 UTC (3,183 KB)
|
| 62 |
+
Full-text links:
|
| 63 |
+
Access Paper:
|
| 64 |
+
View a PDF of the paper titled Evolving Deeper LLM Thinking, by Kuang-Huei Lee and 6 other authors
|
| 65 |
+
View PDF
|
| 66 |
+
HTML (experimental)
|
| 67 |
+
TeX Source
|
| 68 |
+
view license
|
| 69 |
+
Current browse context:
|
| 70 |
+
cs.AI
|
| 71 |
+
< prev
|
| 72 |
+
|
|
| 73 |
+
next >
|
| 74 |
+
new
|
| 75 |
+
|
|
| 76 |
+
recent
|
| 77 |
+
|
|
| 78 |
+
2025-01
|
| 79 |
+
Change to browse by:
|
| 80 |
+
cs
|
| 81 |
+
References & Citations
|
| 82 |
+
NASA ADS
|
| 83 |
+
Google Scholar
|
| 84 |
+
Semantic Scholar
|
| 85 |
+
export BibTeX citation
|
| 86 |
+
Loading...
|
| 87 |
+
BibTeX formatted citation
|
| 88 |
+
×
|
| 89 |
+
loading...
|
| 90 |
+
Data provided by:
|
| 91 |
+
Bookmark
|
| 92 |
+
Bibliographic Tools
|
| 93 |
+
Bibliographic and Citation Tools
|
| 94 |
+
Bibliographic Explorer Toggle
|
| 95 |
+
Bibliographic Explorer
|
| 96 |
+
(
|
| 97 |
+
What is the Explorer?
|
| 98 |
+
)
|
| 99 |
+
Connected Papers Toggle
|
| 100 |
+
Connected Papers
|
| 101 |
+
(
|
| 102 |
+
What is Connected Papers?
|
| 103 |
+
)
|
| 104 |
+
Litmaps Toggle
|
| 105 |
+
Litmaps
|
| 106 |
+
(
|
| 107 |
+
What is Litmaps?
|
| 108 |
+
)
|
| 109 |
+
scite.ai Toggle
|
| 110 |
+
scite Smart Citations
|
| 111 |
+
(
|
| 112 |
+
What are Smart Citations?
|
| 113 |
+
)
|
| 114 |
+
Code, Data, Media
|
| 115 |
+
Code, Data and Media Associated with this Article
|
| 116 |
+
alphaXiv Toggle
|
| 117 |
+
alphaXiv
|
| 118 |
+
(
|
| 119 |
+
What is alphaXiv?
|
| 120 |
+
)
|
| 121 |
+
Links to Code Toggle
|
| 122 |
+
CatalyzeX Code Finder for Papers
|
| 123 |
+
(
|
| 124 |
+
What is CatalyzeX?
|
| 125 |
+
)
|
| 126 |
+
DagsHub Toggle
|
| 127 |
+
DagsHub
|
| 128 |
+
(
|
| 129 |
+
What is DagsHub?
|
| 130 |
+
)
|
| 131 |
+
GotitPub Toggle
|
| 132 |
+
Gotit.pub
|
| 133 |
+
(
|
| 134 |
+
What is GotitPub?
|
| 135 |
+
)
|
| 136 |
+
Huggingface Toggle
|
| 137 |
+
Hugging Face
|
| 138 |
+
(
|
| 139 |
+
What is Huggingface?
|
| 140 |
+
)
|
| 141 |
+
Links to Code Toggle
|
| 142 |
+
Papers with Code
|
| 143 |
+
(
|
| 144 |
+
What is Papers with Code?
|
| 145 |
+
)
|
| 146 |
+
ScienceCast Toggle
|
| 147 |
+
ScienceCast
|
| 148 |
+
(
|
| 149 |
+
What is ScienceCast?
|
| 150 |
+
)
|
| 151 |
+
Demos
|
| 152 |
+
Demos
|
| 153 |
+
Replicate Toggle
|
| 154 |
+
Replicate
|
| 155 |
+
(
|
| 156 |
+
What is Replicate?
|
| 157 |
+
)
|
| 158 |
+
Spaces Toggle
|
| 159 |
+
Hugging Face Spaces
|
| 160 |
+
(
|
| 161 |
+
What is Spaces?
|
| 162 |
+
)
|
| 163 |
+
Spaces Toggle
|
| 164 |
+
TXYZ.AI
|
| 165 |
+
(
|
| 166 |
+
What is TXYZ.AI?
|
| 167 |
+
)
|
| 168 |
+
Related Papers
|
| 169 |
+
Recommenders and Search Tools
|
| 170 |
+
Link to Influence Flower
|
| 171 |
+
Influence Flower
|
| 172 |
+
(
|
| 173 |
+
What are Influence Flowers?
|
| 174 |
+
)
|
| 175 |
+
Core recommender toggle
|
| 176 |
+
CORE Recommender
|
| 177 |
+
(
|
| 178 |
+
What is CORE?
|
| 179 |
+
)
|
| 180 |
+
Author
|
| 181 |
+
Venue
|
| 182 |
+
Institution
|
| 183 |
+
Topic
|
| 184 |
+
About arXivLabs
|
| 185 |
+
arXivLabs: experimental projects with community collaborators
|
| 186 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 187 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 188 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 189 |
+
Learn more about arXivLabs
|
| 190 |
+
.
|
| 191 |
+
Which authors of this paper are endorsers?
|
| 192 |
+
|
|
| 193 |
+
Disable MathJax
|
| 194 |
+
(
|
| 195 |
+
What is MathJax?
|
| 196 |
+
)
|
|
@@ -0,0 +1,386 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2501.12599] Kimi k1.5: Scaling Reinforcement Learning with LLMs'
|
| 3 |
+
id: 250112599-kimi-k15-scaling-reinforcement-learning-with-llms
|
| 4 |
+
tags:
|
| 5 |
+
- deepread
|
| 6 |
+
created: '2026-06-10T00:24:53.655467Z'
|
| 7 |
+
source: https://arxiv.org/abs/2501.12599
|
| 8 |
+
source_domain: arxiv.org
|
| 9 |
+
fetched_at: '2026-06-10T00:24:53.655188Z'
|
| 10 |
+
fetch_provider: builtin
|
| 11 |
+
status: draft
|
| 12 |
+
type: note
|
| 13 |
+
tier: institutional
|
| 14 |
+
content_type: paper
|
| 15 |
+
deprecated: false
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
[2501.12599] Kimi k1.5: Scaling Reinforcement Learning with LLMs
|
| 19 |
+
Computer Science > Artificial Intelligence
|
| 20 |
+
arXiv:2501.12599
|
| 21 |
+
(cs)
|
| 22 |
+
[Submitted on 22 Jan 2025 (
|
| 23 |
+
v1
|
| 24 |
+
), last revised 3 Jun 2025 (this version, v4)]
|
| 25 |
+
Title:
|
| 26 |
+
Kimi k1.5: Scaling Reinforcement Learning with LLMs
|
| 27 |
+
Authors:
|
| 28 |
+
Kimi Team
|
| 29 |
+
,
|
| 30 |
+
Angang Du
|
| 31 |
+
,
|
| 32 |
+
Bofei Gao
|
| 33 |
+
,
|
| 34 |
+
Bowei Xing
|
| 35 |
+
,
|
| 36 |
+
Changjiu Jiang
|
| 37 |
+
,
|
| 38 |
+
Cheng Chen
|
| 39 |
+
,
|
| 40 |
+
Cheng Li
|
| 41 |
+
,
|
| 42 |
+
Chenjun Xiao
|
| 43 |
+
,
|
| 44 |
+
Chenzhuang Du
|
| 45 |
+
,
|
| 46 |
+
Chonghua Liao
|
| 47 |
+
,
|
| 48 |
+
Chuning Tang
|
| 49 |
+
,
|
| 50 |
+
Congcong Wang
|
| 51 |
+
,
|
| 52 |
+
Dehao Zhang
|
| 53 |
+
,
|
| 54 |
+
Enming Yuan
|
| 55 |
+
,
|
| 56 |
+
Enzhe Lu
|
| 57 |
+
,
|
| 58 |
+
Fengxiang Tang
|
| 59 |
+
,
|
| 60 |
+
Flood Sung
|
| 61 |
+
,
|
| 62 |
+
Guangda Wei
|
| 63 |
+
,
|
| 64 |
+
Guokun Lai
|
| 65 |
+
,
|
| 66 |
+
Haiqing Guo
|
| 67 |
+
,
|
| 68 |
+
Han Zhu
|
| 69 |
+
,
|
| 70 |
+
Hao Ding
|
| 71 |
+
,
|
| 72 |
+
Hao Hu
|
| 73 |
+
,
|
| 74 |
+
Hao Yang
|
| 75 |
+
,
|
| 76 |
+
Hao Zhang
|
| 77 |
+
,
|
| 78 |
+
Haotian Yao
|
| 79 |
+
,
|
| 80 |
+
Haotian Zhao
|
| 81 |
+
,
|
| 82 |
+
Haoyu Lu
|
| 83 |
+
,
|
| 84 |
+
Haoze Li
|
| 85 |
+
,
|
| 86 |
+
Haozhen Yu
|
| 87 |
+
,
|
| 88 |
+
Hongcheng Gao
|
| 89 |
+
,
|
| 90 |
+
Huabin Zheng
|
| 91 |
+
,
|
| 92 |
+
Huan Yuan
|
| 93 |
+
,
|
| 94 |
+
Jia Chen
|
| 95 |
+
,
|
| 96 |
+
Jianhang Guo
|
| 97 |
+
,
|
| 98 |
+
Jianlin Su
|
| 99 |
+
,
|
| 100 |
+
Jianzhou Wang
|
| 101 |
+
,
|
| 102 |
+
Jie Zhao
|
| 103 |
+
,
|
| 104 |
+
Jin Zhang
|
| 105 |
+
,
|
| 106 |
+
Jingyuan Liu
|
| 107 |
+
,
|
| 108 |
+
Junjie Yan
|
| 109 |
+
,
|
| 110 |
+
Junyan Wu
|
| 111 |
+
,
|
| 112 |
+
Lidong Shi
|
| 113 |
+
,
|
| 114 |
+
Ling Ye
|
| 115 |
+
,
|
| 116 |
+
Longhui Yu
|
| 117 |
+
,
|
| 118 |
+
Mengnan Dong
|
| 119 |
+
,
|
| 120 |
+
Neo Zhang
|
| 121 |
+
,
|
| 122 |
+
Ningchen Ma
|
| 123 |
+
,
|
| 124 |
+
Qiwei Pan
|
| 125 |
+
,
|
| 126 |
+
Qucheng Gong
|
| 127 |
+
,
|
| 128 |
+
Shaowei Liu
|
| 129 |
+
,
|
| 130 |
+
Shengling Ma
|
| 131 |
+
,
|
| 132 |
+
Shupeng Wei
|
| 133 |
+
,
|
| 134 |
+
Sihan Cao
|
| 135 |
+
,
|
| 136 |
+
Siying Huang
|
| 137 |
+
,
|
| 138 |
+
Tao Jiang
|
| 139 |
+
,
|
| 140 |
+
Weihao Gao
|
| 141 |
+
,
|
| 142 |
+
Weimin Xiong
|
| 143 |
+
,
|
| 144 |
+
Weiran He
|
| 145 |
+
,
|
| 146 |
+
Weixiao Huang
|
| 147 |
+
,
|
| 148 |
+
Weixin Xu
|
| 149 |
+
,
|
| 150 |
+
Wenhao Wu
|
| 151 |
+
,
|
| 152 |
+
Wenyang He
|
| 153 |
+
,
|
| 154 |
+
Xianghui Wei
|
| 155 |
+
,
|
| 156 |
+
Xianqing Jia
|
| 157 |
+
,
|
| 158 |
+
Xingzhe Wu
|
| 159 |
+
,
|
| 160 |
+
Xinran Xu
|
| 161 |
+
,
|
| 162 |
+
Xinxing Zu
|
| 163 |
+
,
|
| 164 |
+
Xinyu Zhou
|
| 165 |
+
,
|
| 166 |
+
Xuehai Pan
|
| 167 |
+
,
|
| 168 |
+
Y. Charles
|
| 169 |
+
,
|
| 170 |
+
Yang Li
|
| 171 |
+
,
|
| 172 |
+
Yangyang Hu
|
| 173 |
+
,
|
| 174 |
+
Yangyang Liu
|
| 175 |
+
,
|
| 176 |
+
Yanru Chen
|
| 177 |
+
,
|
| 178 |
+
Yejie Wang
|
| 179 |
+
,
|
| 180 |
+
Yibo Liu
|
| 181 |
+
,
|
| 182 |
+
Yidao Qin
|
| 183 |
+
,
|
| 184 |
+
Yifeng Liu
|
| 185 |
+
,
|
| 186 |
+
Ying Yang
|
| 187 |
+
,
|
| 188 |
+
Yiping Bao
|
| 189 |
+
,
|
| 190 |
+
Yulun Du
|
| 191 |
+
,
|
| 192 |
+
Yuxin Wu
|
| 193 |
+
,
|
| 194 |
+
Yuzhi Wang
|
| 195 |
+
,
|
| 196 |
+
Zaida Zhou
|
| 197 |
+
,
|
| 198 |
+
Zhaoji Wang
|
| 199 |
+
,
|
| 200 |
+
Zhaowei Li
|
| 201 |
+
,
|
| 202 |
+
Zhen Zhu
|
| 203 |
+
,
|
| 204 |
+
Zheng Zhang
|
| 205 |
+
,
|
| 206 |
+
Zhexu Wang
|
| 207 |
+
,
|
| 208 |
+
Zhilin Yang
|
| 209 |
+
,
|
| 210 |
+
Zhiqi Huang
|
| 211 |
+
,
|
| 212 |
+
Zihao Huang
|
| 213 |
+
,
|
| 214 |
+
Ziyao Xu
|
| 215 |
+
,
|
| 216 |
+
Zonghan Yang
|
| 217 |
+
,
|
| 218 |
+
Zongyu Lin
|
| 219 |
+
View a PDF of the paper titled Kimi k1.5: Scaling Reinforcement Learning with LLMs, by Kimi Team and 95 other authors
|
| 220 |
+
View PDF
|
| 221 |
+
HTML (experimental)
|
| 222 |
+
Abstract:
|
| 223 |
+
Language model pretraining with next token prediction has proved effective for scaling compute but is limited to the amount of available training data. Scaling reinforcement learning (RL) unlocks a new axis for the continued improvement of artificial intelligence, with the promise that large language models (LLMs) can scale their training data by learning to explore with rewards. However, prior published work has not produced competitive results. In light of this, we report on the training practice of Kimi k1.5, our latest multi-modal LLM trained with RL, including its RL training techniques, multi-modal data recipes, and infrastructure optimization. Long context scaling and improved policy optimization methods are key ingredients of our approach, which establishes a simplistic, effective RL framework without relying on more complex techniques such as Monte Carlo tree search, value functions, and process reward models. Notably, our system achieves state-of-the-art reasoning performance across multiple benchmarks and modalities -- e.g., 77.5 on AIME, 96.2 on MATH 500, 94-th percentile on Codeforces, 74.9 on MathVista -- matching OpenAI's o1. Moreover, we present effective long2short methods that use long-CoT techniques to improve short-CoT models, yielding state-of-the-art short-CoT reasoning results -- e.g., 60.8 on AIME, 94.6 on MATH500, 47.3 on LiveCodeBench -- outperforming existing short-CoT models such as GPT-4o and Claude Sonnet 3.5 by a large margin (up to +550%).
|
| 224 |
+
Comments:
|
| 225 |
+
25 pages
|
| 226 |
+
Subjects:
|
| 227 |
+
Artificial Intelligence (cs.AI)
|
| 228 |
+
; Machine Learning (cs.LG)
|
| 229 |
+
Cite as:
|
| 230 |
+
arXiv:2501.12599
|
| 231 |
+
[cs.AI]
|
| 232 |
+
(or
|
| 233 |
+
arXiv:2501.12599v4
|
| 234 |
+
[cs.AI]
|
| 235 |
+
for this version)
|
| 236 |
+
https://doi.org/10.48550/arXiv.2501.12599
|
| 237 |
+
Focus to learn more
|
| 238 |
+
arXiv-issued DOI via DataCite
|
| 239 |
+
Submission history
|
| 240 |
+
From: Flood Sung [
|
| 241 |
+
view email
|
| 242 |
+
]
|
| 243 |
+
[v1]
|
| 244 |
+
Wed, 22 Jan 2025 02:48:14 UTC (614 KB)
|
| 245 |
+
[v2]
|
| 246 |
+
Wed, 5 Mar 2025 02:16:32 UTC (614 KB)
|
| 247 |
+
[v3]
|
| 248 |
+
Wed, 28 May 2025 03:57:30 UTC (614 KB)
|
| 249 |
+
[v4]
|
| 250 |
+
Tue, 3 Jun 2025 02:14:54 UTC (603 KB)
|
| 251 |
+
Full-text links:
|
| 252 |
+
Access Paper:
|
| 253 |
+
View a PDF of the paper titled Kimi k1.5: Scaling Reinforcement Learning with LLMs, by Kimi Team and 95 other authors
|
| 254 |
+
View PDF
|
| 255 |
+
HTML (experimental)
|
| 256 |
+
TeX Source
|
| 257 |
+
view license
|
| 258 |
+
Current browse context:
|
| 259 |
+
cs.AI
|
| 260 |
+
< prev
|
| 261 |
+
|
|
| 262 |
+
next >
|
| 263 |
+
new
|
| 264 |
+
|
|
| 265 |
+
recent
|
| 266 |
+
|
|
| 267 |
+
2025-01
|
| 268 |
+
Change to browse by:
|
| 269 |
+
cs
|
| 270 |
+
cs.LG
|
| 271 |
+
References & Citations
|
| 272 |
+
NASA ADS
|
| 273 |
+
Google Scholar
|
| 274 |
+
Semantic Scholar
|
| 275 |
+
export BibTeX citation
|
| 276 |
+
Loading...
|
| 277 |
+
BibTeX formatted citation
|
| 278 |
+
×
|
| 279 |
+
loading...
|
| 280 |
+
Data provided by:
|
| 281 |
+
Bookmark
|
| 282 |
+
Bibliographic Tools
|
| 283 |
+
Bibliographic and Citation Tools
|
| 284 |
+
Bibliographic Explorer Toggle
|
| 285 |
+
Bibliographic Explorer
|
| 286 |
+
(
|
| 287 |
+
What is the Explorer?
|
| 288 |
+
)
|
| 289 |
+
Connected Papers Toggle
|
| 290 |
+
Connected Papers
|
| 291 |
+
(
|
| 292 |
+
What is Connected Papers?
|
| 293 |
+
)
|
| 294 |
+
Litmaps Toggle
|
| 295 |
+
Litmaps
|
| 296 |
+
(
|
| 297 |
+
What is Litmaps?
|
| 298 |
+
)
|
| 299 |
+
scite.ai Toggle
|
| 300 |
+
scite Smart Citations
|
| 301 |
+
(
|
| 302 |
+
What are Smart Citations?
|
| 303 |
+
)
|
| 304 |
+
Code, Data, Media
|
| 305 |
+
Code, Data and Media Associated with this Article
|
| 306 |
+
alphaXiv Toggle
|
| 307 |
+
alphaXiv
|
| 308 |
+
(
|
| 309 |
+
What is alphaXiv?
|
| 310 |
+
)
|
| 311 |
+
Links to Code Toggle
|
| 312 |
+
CatalyzeX Code Finder for Papers
|
| 313 |
+
(
|
| 314 |
+
What is CatalyzeX?
|
| 315 |
+
)
|
| 316 |
+
DagsHub Toggle
|
| 317 |
+
DagsHub
|
| 318 |
+
(
|
| 319 |
+
What is DagsHub?
|
| 320 |
+
)
|
| 321 |
+
GotitPub Toggle
|
| 322 |
+
Gotit.pub
|
| 323 |
+
(
|
| 324 |
+
What is GotitPub?
|
| 325 |
+
)
|
| 326 |
+
Huggingface Toggle
|
| 327 |
+
Hugging Face
|
| 328 |
+
(
|
| 329 |
+
What is Huggingface?
|
| 330 |
+
)
|
| 331 |
+
Links to Code Toggle
|
| 332 |
+
Papers with Code
|
| 333 |
+
(
|
| 334 |
+
What is Papers with Code?
|
| 335 |
+
)
|
| 336 |
+
ScienceCast Toggle
|
| 337 |
+
ScienceCast
|
| 338 |
+
(
|
| 339 |
+
What is ScienceCast?
|
| 340 |
+
)
|
| 341 |
+
Demos
|
| 342 |
+
Demos
|
| 343 |
+
Replicate Toggle
|
| 344 |
+
Replicate
|
| 345 |
+
(
|
| 346 |
+
What is Replicate?
|
| 347 |
+
)
|
| 348 |
+
Spaces Toggle
|
| 349 |
+
Hugging Face Spaces
|
| 350 |
+
(
|
| 351 |
+
What is Spaces?
|
| 352 |
+
)
|
| 353 |
+
Spaces Toggle
|
| 354 |
+
TXYZ.AI
|
| 355 |
+
(
|
| 356 |
+
What is TXYZ.AI?
|
| 357 |
+
)
|
| 358 |
+
Related Papers
|
| 359 |
+
Recommenders and Search Tools
|
| 360 |
+
Link to Influence Flower
|
| 361 |
+
Influence Flower
|
| 362 |
+
(
|
| 363 |
+
What are Influence Flowers?
|
| 364 |
+
)
|
| 365 |
+
Core recommender toggle
|
| 366 |
+
CORE Recommender
|
| 367 |
+
(
|
| 368 |
+
What is CORE?
|
| 369 |
+
)
|
| 370 |
+
Author
|
| 371 |
+
Venue
|
| 372 |
+
Institution
|
| 373 |
+
Topic
|
| 374 |
+
About arXivLabs
|
| 375 |
+
arXivLabs: experimental projects with community collaborators
|
| 376 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 377 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 378 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 379 |
+
Learn more about arXivLabs
|
| 380 |
+
.
|
| 381 |
+
Which authors of this paper are endorsers?
|
| 382 |
+
|
|
| 383 |
+
Disable MathJax
|
| 384 |
+
(
|
| 385 |
+
What is MathJax?
|
| 386 |
+
)
|
|
@@ -0,0 +1,206 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2501.18512] Streaming DiLoCo with overlapping communication: Towards a Distributed
|
| 3 |
+
Free Lunch'
|
| 4 |
+
id: 250118512-streaming-diloco-with-overlapping-communication-towards-a-distributed
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:30:21.211856Z'
|
| 8 |
+
source: https://arxiv.org/abs/2501.18512
|
| 9 |
+
source_domain: arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:30:21.211709Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2501.18512] Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
|
| 20 |
+
Computer Science > Computation and Language
|
| 21 |
+
arXiv:2501.18512
|
| 22 |
+
(cs)
|
| 23 |
+
[Submitted on 30 Jan 2025]
|
| 24 |
+
Title:
|
| 25 |
+
Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
|
| 26 |
+
Authors:
|
| 27 |
+
Arthur Douillard
|
| 28 |
+
,
|
| 29 |
+
Yanislav Donchev
|
| 30 |
+
,
|
| 31 |
+
Keith Rush
|
| 32 |
+
,
|
| 33 |
+
Satyen Kale
|
| 34 |
+
,
|
| 35 |
+
Zachary Charles
|
| 36 |
+
,
|
| 37 |
+
Zachary Garrett
|
| 38 |
+
,
|
| 39 |
+
Gabriel Teston
|
| 40 |
+
,
|
| 41 |
+
Dave Lacey
|
| 42 |
+
,
|
| 43 |
+
Ross McIlroy
|
| 44 |
+
,
|
| 45 |
+
Jiajun Shen
|
| 46 |
+
,
|
| 47 |
+
Alexandre Ramé
|
| 48 |
+
,
|
| 49 |
+
Arthur Szlam
|
| 50 |
+
,
|
| 51 |
+
Marc'Aurelio Ranzato
|
| 52 |
+
,
|
| 53 |
+
Paul Barham
|
| 54 |
+
View a PDF of the paper titled Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch, by Arthur Douillard and Yanislav Donchev and Keith Rush and Satyen Kale and Zachary Charles and Zachary Garrett and Gabriel Teston and Dave Lacey and Ross McIlroy and Jiajun Shen and Alexandre Ram\'e and Arthur Szlam and Marc'Aurelio Ranzato and Paul Barham
|
| 55 |
+
View PDF
|
| 56 |
+
HTML (experimental)
|
| 57 |
+
Abstract:
|
| 58 |
+
Training of large language models (LLMs) is typically distributed across a large number of accelerators to reduce training time. Since internal states and parameter gradients need to be exchanged at each and every single gradient step, all devices need to be co-located using low-latency high-bandwidth communication links to support the required high volume of exchanged bits. Recently, distributed algorithms like DiLoCo have relaxed such co-location constraint: accelerators can be grouped into ``workers'', where synchronizations between workers only occur infrequently. This in turn means that workers can afford being connected by lower bandwidth communication links without affecting learning quality. However, in these methods, communication across workers still requires the same peak bandwidth as before, as the synchronizations require all parameters to be exchanged across all workers. In this paper, we improve DiLoCo in three ways. First, we synchronize only subsets of parameters in sequence, rather than all at once, which greatly reduces peak bandwidth. Second, we allow workers to continue training while synchronizing, which decreases wall clock time. Third, we quantize the data exchanged by workers, which further reduces bandwidth across workers. By properly combining these modifications, we show experimentally that we can distribute training of billion-scale parameters and reach similar quality as before, but reducing required bandwidth by two orders of magnitude.
|
| 59 |
+
Subjects:
|
| 60 |
+
Computation and Language (cs.CL)
|
| 61 |
+
Cite as:
|
| 62 |
+
arXiv:2501.18512
|
| 63 |
+
[cs.CL]
|
| 64 |
+
(or
|
| 65 |
+
arXiv:2501.18512v1
|
| 66 |
+
[cs.CL]
|
| 67 |
+
for this version)
|
| 68 |
+
https://doi.org/10.48550/arXiv.2501.18512
|
| 69 |
+
Focus to learn more
|
| 70 |
+
arXiv-issued DOI via DataCite
|
| 71 |
+
Submission history
|
| 72 |
+
From: Arthur Douillard [
|
| 73 |
+
view email
|
| 74 |
+
]
|
| 75 |
+
[v1]
|
| 76 |
+
Thu, 30 Jan 2025 17:23:50 UTC (3,278 KB)
|
| 77 |
+
Full-text links:
|
| 78 |
+
Access Paper:
|
| 79 |
+
View a PDF of the paper titled Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch, by Arthur Douillard and Yanislav Donchev and Keith Rush and Satyen Kale and Zachary Charles and Zachary Garrett and Gabriel Teston and Dave Lacey and Ross McIlroy and Jiajun Shen and Alexandre Ram\'e and Arthur Szlam and Marc'Aurelio Ranzato and Paul Barham
|
| 80 |
+
View PDF
|
| 81 |
+
HTML (experimental)
|
| 82 |
+
TeX Source
|
| 83 |
+
view license
|
| 84 |
+
Current browse context:
|
| 85 |
+
cs.CL
|
| 86 |
+
< prev
|
| 87 |
+
|
|
| 88 |
+
next >
|
| 89 |
+
new
|
| 90 |
+
|
|
| 91 |
+
recent
|
| 92 |
+
|
|
| 93 |
+
2025-01
|
| 94 |
+
Change to browse by:
|
| 95 |
+
cs
|
| 96 |
+
References & Citations
|
| 97 |
+
NASA ADS
|
| 98 |
+
Google Scholar
|
| 99 |
+
Semantic Scholar
|
| 100 |
+
export BibTeX citation
|
| 101 |
+
Loading...
|
| 102 |
+
BibTeX formatted citation
|
| 103 |
+
×
|
| 104 |
+
loading...
|
| 105 |
+
Data provided by:
|
| 106 |
+
Bookmark
|
| 107 |
+
Bibliographic Tools
|
| 108 |
+
Bibliographic and Citation Tools
|
| 109 |
+
Bibliographic Explorer Toggle
|
| 110 |
+
Bibliographic Explorer
|
| 111 |
+
(
|
| 112 |
+
What is the Explorer?
|
| 113 |
+
)
|
| 114 |
+
Connected Papers Toggle
|
| 115 |
+
Connected Papers
|
| 116 |
+
(
|
| 117 |
+
What is Connected Papers?
|
| 118 |
+
)
|
| 119 |
+
Litmaps Toggle
|
| 120 |
+
Litmaps
|
| 121 |
+
(
|
| 122 |
+
What is Litmaps?
|
| 123 |
+
)
|
| 124 |
+
scite.ai Toggle
|
| 125 |
+
scite Smart Citations
|
| 126 |
+
(
|
| 127 |
+
What are Smart Citations?
|
| 128 |
+
)
|
| 129 |
+
Code, Data, Media
|
| 130 |
+
Code, Data and Media Associated with this Article
|
| 131 |
+
alphaXiv Toggle
|
| 132 |
+
alphaXiv
|
| 133 |
+
(
|
| 134 |
+
What is alphaXiv?
|
| 135 |
+
)
|
| 136 |
+
Links to Code Toggle
|
| 137 |
+
CatalyzeX Code Finder for Papers
|
| 138 |
+
(
|
| 139 |
+
What is CatalyzeX?
|
| 140 |
+
)
|
| 141 |
+
DagsHub Toggle
|
| 142 |
+
DagsHub
|
| 143 |
+
(
|
| 144 |
+
What is DagsHub?
|
| 145 |
+
)
|
| 146 |
+
GotitPub Toggle
|
| 147 |
+
Gotit.pub
|
| 148 |
+
(
|
| 149 |
+
What is GotitPub?
|
| 150 |
+
)
|
| 151 |
+
Huggingface Toggle
|
| 152 |
+
Hugging Face
|
| 153 |
+
(
|
| 154 |
+
What is Huggingface?
|
| 155 |
+
)
|
| 156 |
+
ScienceCast Toggle
|
| 157 |
+
ScienceCast
|
| 158 |
+
(
|
| 159 |
+
What is ScienceCast?
|
| 160 |
+
)
|
| 161 |
+
Demos
|
| 162 |
+
Demos
|
| 163 |
+
Replicate Toggle
|
| 164 |
+
Replicate
|
| 165 |
+
(
|
| 166 |
+
What is Replicate?
|
| 167 |
+
)
|
| 168 |
+
Spaces Toggle
|
| 169 |
+
Hugging Face Spaces
|
| 170 |
+
(
|
| 171 |
+
What is Spaces?
|
| 172 |
+
)
|
| 173 |
+
Spaces Toggle
|
| 174 |
+
TXYZ.AI
|
| 175 |
+
(
|
| 176 |
+
What is TXYZ.AI?
|
| 177 |
+
)
|
| 178 |
+
Related Papers
|
| 179 |
+
Recommenders and Search Tools
|
| 180 |
+
Link to Influence Flower
|
| 181 |
+
Influence Flower
|
| 182 |
+
(
|
| 183 |
+
What are Influence Flowers?
|
| 184 |
+
)
|
| 185 |
+
Core recommender toggle
|
| 186 |
+
CORE Recommender
|
| 187 |
+
(
|
| 188 |
+
What is CORE?
|
| 189 |
+
)
|
| 190 |
+
Author
|
| 191 |
+
Venue
|
| 192 |
+
Institution
|
| 193 |
+
Topic
|
| 194 |
+
About arXivLabs
|
| 195 |
+
arXivLabs: experimental projects with community collaborators
|
| 196 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 197 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 198 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 199 |
+
Learn more about arXivLabs
|
| 200 |
+
.
|
| 201 |
+
Which authors of this paper are endorsers?
|
| 202 |
+
|
|
| 203 |
+
Disable MathJax
|
| 204 |
+
(
|
| 205 |
+
What is MathJax?
|
| 206 |
+
)
|
|
@@ -0,0 +1,179 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2501.18639] A Comprehensive Survey of the Lean 4 Theorem Prover: Architecture,
|
| 3 |
+
Applications, and Advances'
|
| 4 |
+
id: 250118639-a-comprehensive-survey-of-the-lean-4-theorem-prover-architecture-appli
|
| 5 |
+
tags:
|
| 6 |
+
- deepread
|
| 7 |
+
created: '2026-06-10T00:25:14.929249Z'
|
| 8 |
+
source: https://arxiv.org/abs/2501.18639
|
| 9 |
+
source_domain: arxiv.org
|
| 10 |
+
fetched_at: '2026-06-10T00:25:14.929112Z'
|
| 11 |
+
fetch_provider: builtin
|
| 12 |
+
status: draft
|
| 13 |
+
type: note
|
| 14 |
+
tier: institutional
|
| 15 |
+
content_type: paper
|
| 16 |
+
deprecated: false
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
[2501.18639] A Comprehensive Survey of the Lean 4 Theorem Prover: Architecture, Applications, and Advances
|
| 20 |
+
Computer Science > Logic in Computer Science
|
| 21 |
+
arXiv:2501.18639
|
| 22 |
+
(cs)
|
| 23 |
+
[Submitted on 28 Jan 2025]
|
| 24 |
+
Title:
|
| 25 |
+
A Comprehensive Survey of the Lean 4 Theorem Prover: Architecture, Applications, and Advances
|
| 26 |
+
Authors:
|
| 27 |
+
Xichen Tang
|
| 28 |
+
View a PDF of the paper titled A Comprehensive Survey of the Lean 4 Theorem Prover: Architecture, Applications, and Advances, by Xichen Tang
|
| 29 |
+
View PDF
|
| 30 |
+
Abstract:
|
| 31 |
+
This comprehensive survey examines Lean 4, a state-of-the-art interactive theorem prover and functional programming language. We analyze its architectural design, type system, metaprogramming capabilities, and practical applications in formal verification and mathematics. Through detailed comparisons with other proof assistants and extensive case studies, we demonstrate Lean 4's unique advantages in proof automation, performance, and usability. The paper also explores recent developments in its ecosystem, including libraries, tools, and educational applications, providing insights into its growing impact on formal methods and mathematical formalization.
|
| 32 |
+
Subjects:
|
| 33 |
+
Logic in Computer Science (cs.LO)
|
| 34 |
+
; Programming Languages (cs.PL)
|
| 35 |
+
Cite as:
|
| 36 |
+
arXiv:2501.18639
|
| 37 |
+
[cs.LO]
|
| 38 |
+
(or
|
| 39 |
+
arXiv:2501.18639v1
|
| 40 |
+
[cs.LO]
|
| 41 |
+
for this version)
|
| 42 |
+
https://doi.org/10.48550/arXiv.2501.18639
|
| 43 |
+
Focus to learn more
|
| 44 |
+
arXiv-issued DOI via DataCite
|
| 45 |
+
Submission history
|
| 46 |
+
From: Xichen Tang [
|
| 47 |
+
view email
|
| 48 |
+
]
|
| 49 |
+
[v1]
|
| 50 |
+
Tue, 28 Jan 2025 17:15:54 UTC (2,729 KB)
|
| 51 |
+
Full-text links:
|
| 52 |
+
Access Paper:
|
| 53 |
+
View a PDF of the paper titled A Comprehensive Survey of the Lean 4 Theorem Prover: Architecture, Applications, and Advances, by Xichen Tang
|
| 54 |
+
View PDF
|
| 55 |
+
view license
|
| 56 |
+
Current browse context:
|
| 57 |
+
cs.LO
|
| 58 |
+
< prev
|
| 59 |
+
|
|
| 60 |
+
next >
|
| 61 |
+
new
|
| 62 |
+
|
|
| 63 |
+
recent
|
| 64 |
+
|
|
| 65 |
+
2025-01
|
| 66 |
+
Change to browse by:
|
| 67 |
+
cs
|
| 68 |
+
cs.PL
|
| 69 |
+
References & Citations
|
| 70 |
+
NASA ADS
|
| 71 |
+
Google Scholar
|
| 72 |
+
Semantic Scholar
|
| 73 |
+
export BibTeX citation
|
| 74 |
+
Loading...
|
| 75 |
+
BibTeX formatted citation
|
| 76 |
+
×
|
| 77 |
+
loading...
|
| 78 |
+
Data provided by:
|
| 79 |
+
Bookmark
|
| 80 |
+
Bibliographic Tools
|
| 81 |
+
Bibliographic and Citation Tools
|
| 82 |
+
Bibliographic Explorer Toggle
|
| 83 |
+
Bibliographic Explorer
|
| 84 |
+
(
|
| 85 |
+
What is the Explorer?
|
| 86 |
+
)
|
| 87 |
+
Connected Papers Toggle
|
| 88 |
+
Connected Papers
|
| 89 |
+
(
|
| 90 |
+
What is Connected Papers?
|
| 91 |
+
)
|
| 92 |
+
Litmaps Toggle
|
| 93 |
+
Litmaps
|
| 94 |
+
(
|
| 95 |
+
What is Litmaps?
|
| 96 |
+
)
|
| 97 |
+
scite.ai Toggle
|
| 98 |
+
scite Smart Citations
|
| 99 |
+
(
|
| 100 |
+
What are Smart Citations?
|
| 101 |
+
)
|
| 102 |
+
Code, Data, Media
|
| 103 |
+
Code, Data and Media Associated with this Article
|
| 104 |
+
alphaXiv Toggle
|
| 105 |
+
alphaXiv
|
| 106 |
+
(
|
| 107 |
+
What is alphaXiv?
|
| 108 |
+
)
|
| 109 |
+
Links to Code Toggle
|
| 110 |
+
CatalyzeX Code Finder for Papers
|
| 111 |
+
(
|
| 112 |
+
What is CatalyzeX?
|
| 113 |
+
)
|
| 114 |
+
DagsHub Toggle
|
| 115 |
+
DagsHub
|
| 116 |
+
(
|
| 117 |
+
What is DagsHub?
|
| 118 |
+
)
|
| 119 |
+
GotitPub Toggle
|
| 120 |
+
Gotit.pub
|
| 121 |
+
(
|
| 122 |
+
What is GotitPub?
|
| 123 |
+
)
|
| 124 |
+
Huggingface Toggle
|
| 125 |
+
Hugging Face
|
| 126 |
+
(
|
| 127 |
+
What is Huggingface?
|
| 128 |
+
)
|
| 129 |
+
ScienceCast Toggle
|
| 130 |
+
ScienceCast
|
| 131 |
+
(
|
| 132 |
+
What is ScienceCast?
|
| 133 |
+
)
|
| 134 |
+
Demos
|
| 135 |
+
Demos
|
| 136 |
+
Replicate Toggle
|
| 137 |
+
Replicate
|
| 138 |
+
(
|
| 139 |
+
What is Replicate?
|
| 140 |
+
)
|
| 141 |
+
Spaces Toggle
|
| 142 |
+
Hugging Face Spaces
|
| 143 |
+
(
|
| 144 |
+
What is Spaces?
|
| 145 |
+
)
|
| 146 |
+
Spaces Toggle
|
| 147 |
+
TXYZ.AI
|
| 148 |
+
(
|
| 149 |
+
What is TXYZ.AI?
|
| 150 |
+
)
|
| 151 |
+
Related Papers
|
| 152 |
+
Recommenders and Search Tools
|
| 153 |
+
Link to Influence Flower
|
| 154 |
+
Influence Flower
|
| 155 |
+
(
|
| 156 |
+
What are Influence Flowers?
|
| 157 |
+
)
|
| 158 |
+
Core recommender toggle
|
| 159 |
+
CORE Recommender
|
| 160 |
+
(
|
| 161 |
+
What is CORE?
|
| 162 |
+
)
|
| 163 |
+
Author
|
| 164 |
+
Venue
|
| 165 |
+
Institution
|
| 166 |
+
Topic
|
| 167 |
+
About arXivLabs
|
| 168 |
+
arXivLabs: experimental projects with community collaborators
|
| 169 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 170 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 171 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 172 |
+
Learn more about arXivLabs
|
| 173 |
+
.
|
| 174 |
+
Which authors of this paper are endorsers?
|
| 175 |
+
|
|
| 176 |
+
Disable MathJax
|
| 177 |
+
(
|
| 178 |
+
What is MathJax?
|
| 179 |
+
)
|
|
@@ -0,0 +1,183 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: '[2502.02047] AmaSQuAD: A Benchmark for Amharic Extractive Question Answering'
|
| 3 |
+
id: 250202047-amasquad-a-benchmark-for-amharic-extractive-question-answering
|
| 4 |
+
tags:
|
| 5 |
+
- deepread
|
| 6 |
+
created: '2026-06-10T00:24:12.341628Z'
|
| 7 |
+
source: https://arxiv.org/abs/2502.02047
|
| 8 |
+
source_domain: arxiv.org
|
| 9 |
+
fetched_at: '2026-06-10T00:24:12.341471Z'
|
| 10 |
+
fetch_provider: builtin
|
| 11 |
+
status: draft
|
| 12 |
+
type: note
|
| 13 |
+
tier: institutional
|
| 14 |
+
content_type: paper
|
| 15 |
+
deprecated: false
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
[2502.02047] AmaSQuAD: A Benchmark for Amharic Extractive Question Answering
|
| 19 |
+
Computer Science > Computation and Language
|
| 20 |
+
arXiv:2502.02047
|
| 21 |
+
(cs)
|
| 22 |
+
[Submitted on 4 Feb 2025]
|
| 23 |
+
Title:
|
| 24 |
+
AmaSQuAD: A Benchmark for Amharic Extractive Question Answering
|
| 25 |
+
Authors:
|
| 26 |
+
Nebiyou Daniel Hailemariam
|
| 27 |
+
,
|
| 28 |
+
Blessed Guda
|
| 29 |
+
,
|
| 30 |
+
Tsegazeab Tefferi
|
| 31 |
+
View a PDF of the paper titled AmaSQuAD: A Benchmark for Amharic Extractive Question Answering, by Nebiyou Daniel Hailemariam and 2 other authors
|
| 32 |
+
View PDF
|
| 33 |
+
HTML (experimental)
|
| 34 |
+
Abstract:
|
| 35 |
+
This research presents a novel framework for translating extractive question-answering datasets into low-resource languages, as demonstrated by the creation of the AmaSQuAD dataset, a translation of SQuAD 2.0 into Amharic. The methodology addresses challenges related to misalignment between translated questions and answers, as well as the presence of multiple answer instances in the translated context. For this purpose, we used cosine similarity utilizing embeddings from a fine-tuned BERT-based model for Amharic and Longest Common Subsequence (LCS). Additionally, we fine-tune the XLM-R model on the AmaSQuAD synthetic dataset for Amharic Question-Answering. The results show an improvement in baseline performance, with the fine-tuned model achieving an increase in the F1 score from 36.55% to 44.41% and 50.01% to 57.5% on the AmaSQuAD development dataset. Moreover, the model demonstrates improvement on the human-curated AmQA dataset, increasing the F1 score from 67.80% to 68.80% and the exact match score from 52.50% to 52.66%.The AmaSQuAD dataset is publicly available Datasets
|
| 36 |
+
Subjects:
|
| 37 |
+
Computation and Language (cs.CL)
|
| 38 |
+
Cite as:
|
| 39 |
+
arXiv:2502.02047
|
| 40 |
+
[cs.CL]
|
| 41 |
+
(or
|
| 42 |
+
arXiv:2502.02047v1
|
| 43 |
+
[cs.CL]
|
| 44 |
+
for this version)
|
| 45 |
+
https://doi.org/10.48550/arXiv.2502.02047
|
| 46 |
+
Focus to learn more
|
| 47 |
+
arXiv-issued DOI via DataCite
|
| 48 |
+
Submission history
|
| 49 |
+
From: Blessed Guda [
|
| 50 |
+
view email
|
| 51 |
+
]
|
| 52 |
+
[v1]
|
| 53 |
+
Tue, 4 Feb 2025 06:27:39 UTC (778 KB)
|
| 54 |
+
Full-text links:
|
| 55 |
+
Access Paper:
|
| 56 |
+
View a PDF of the paper titled AmaSQuAD: A Benchmark for Amharic Extractive Question Answering, by Nebiyou Daniel Hailemariam and 2 other authors
|
| 57 |
+
View PDF
|
| 58 |
+
HTML (experimental)
|
| 59 |
+
TeX Source
|
| 60 |
+
view license
|
| 61 |
+
Current browse context:
|
| 62 |
+
cs.CL
|
| 63 |
+
< prev
|
| 64 |
+
|
|
| 65 |
+
next >
|
| 66 |
+
new
|
| 67 |
+
|
|
| 68 |
+
recent
|
| 69 |
+
|
|
| 70 |
+
2025-02
|
| 71 |
+
Change to browse by:
|
| 72 |
+
cs
|
| 73 |
+
References & Citations
|
| 74 |
+
NASA ADS
|
| 75 |
+
Google Scholar
|
| 76 |
+
Semantic Scholar
|
| 77 |
+
export BibTeX citation
|
| 78 |
+
Loading...
|
| 79 |
+
BibTeX formatted citation
|
| 80 |
+
×
|
| 81 |
+
loading...
|
| 82 |
+
Data provided by:
|
| 83 |
+
Bookmark
|
| 84 |
+
Bibliographic Tools
|
| 85 |
+
Bibliographic and Citation Tools
|
| 86 |
+
Bibliographic Explorer Toggle
|
| 87 |
+
Bibliographic Explorer
|
| 88 |
+
(
|
| 89 |
+
What is the Explorer?
|
| 90 |
+
)
|
| 91 |
+
Connected Papers Toggle
|
| 92 |
+
Connected Papers
|
| 93 |
+
(
|
| 94 |
+
What is Connected Papers?
|
| 95 |
+
)
|
| 96 |
+
Litmaps Toggle
|
| 97 |
+
Litmaps
|
| 98 |
+
(
|
| 99 |
+
What is Litmaps?
|
| 100 |
+
)
|
| 101 |
+
scite.ai Toggle
|
| 102 |
+
scite Smart Citations
|
| 103 |
+
(
|
| 104 |
+
What are Smart Citations?
|
| 105 |
+
)
|
| 106 |
+
Code, Data, Media
|
| 107 |
+
Code, Data and Media Associated with this Article
|
| 108 |
+
alphaXiv Toggle
|
| 109 |
+
alphaXiv
|
| 110 |
+
(
|
| 111 |
+
What is alphaXiv?
|
| 112 |
+
)
|
| 113 |
+
Links to Code Toggle
|
| 114 |
+
CatalyzeX Code Finder for Papers
|
| 115 |
+
(
|
| 116 |
+
What is CatalyzeX?
|
| 117 |
+
)
|
| 118 |
+
DagsHub Toggle
|
| 119 |
+
DagsHub
|
| 120 |
+
(
|
| 121 |
+
What is DagsHub?
|
| 122 |
+
)
|
| 123 |
+
GotitPub Toggle
|
| 124 |
+
Gotit.pub
|
| 125 |
+
(
|
| 126 |
+
What is GotitPub?
|
| 127 |
+
)
|
| 128 |
+
Huggingface Toggle
|
| 129 |
+
Hugging Face
|
| 130 |
+
(
|
| 131 |
+
What is Huggingface?
|
| 132 |
+
)
|
| 133 |
+
ScienceCast Toggle
|
| 134 |
+
ScienceCast
|
| 135 |
+
(
|
| 136 |
+
What is ScienceCast?
|
| 137 |
+
)
|
| 138 |
+
Demos
|
| 139 |
+
Demos
|
| 140 |
+
Replicate Toggle
|
| 141 |
+
Replicate
|
| 142 |
+
(
|
| 143 |
+
What is Replicate?
|
| 144 |
+
)
|
| 145 |
+
Spaces Toggle
|
| 146 |
+
Hugging Face Spaces
|
| 147 |
+
(
|
| 148 |
+
What is Spaces?
|
| 149 |
+
)
|
| 150 |
+
Spaces Toggle
|
| 151 |
+
TXYZ.AI
|
| 152 |
+
(
|
| 153 |
+
What is TXYZ.AI?
|
| 154 |
+
)
|
| 155 |
+
Related Papers
|
| 156 |
+
Recommenders and Search Tools
|
| 157 |
+
Link to Influence Flower
|
| 158 |
+
Influence Flower
|
| 159 |
+
(
|
| 160 |
+
What are Influence Flowers?
|
| 161 |
+
)
|
| 162 |
+
Core recommender toggle
|
| 163 |
+
CORE Recommender
|
| 164 |
+
(
|
| 165 |
+
What is CORE?
|
| 166 |
+
)
|
| 167 |
+
Author
|
| 168 |
+
Venue
|
| 169 |
+
Institution
|
| 170 |
+
Topic
|
| 171 |
+
About arXivLabs
|
| 172 |
+
arXivLabs: experimental projects with community collaborators
|
| 173 |
+
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
|
| 174 |
+
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
|
| 175 |
+
Have an idea for a project that will add value for arXiv's community?
|
| 176 |
+
Learn more about arXivLabs
|
| 177 |
+
.
|
| 178 |
+
Which authors of this paper are endorsers?
|
| 179 |
+
|
|
| 180 |
+
Disable MathJax
|
| 181 |
+
(
|
| 182 |
+
What is MathJax?
|
| 183 |
+
)
|