seonglae commited on
Commit
77adcae
·
1 Parent(s): 664c6e5

links: point the paper button at arXiv rather than the workshop OpenReview

Browse files

The page is now a NeurIPS paper with its own arXiv record, 2608.23670. The
variable keeps its name; renaming it is a separate change.

Files changed (3) hide show
  1. app/src/pages/index.astro +1 -1
  2. index.html +12 -12
  3. index.html.gz +0 -0
app/src/pages/index.astro CHANGED
@@ -144,7 +144,7 @@ const citationAuthorsText = authorNames.join(", ");
144
  // Published at ICML 2026 AIWILD workshop. Use the canonical OpenReview entry
145
  // rather than an auto-generated @misc (which left a dangling comma when no DOI).
146
  const venue = "Second Workshop on Agents in the Wild: Safety, Security, and Beyond";
147
- const openreviewUrl = "https://openreview.net/forum?id=1cprFkvFT0";
148
  const fullTitle = "Automata from Agent Traces: Failure and Next-Step Prediction";
149
  const authorsBib = authorNames.join(" and ");
150
  const doi = (ArticleMod as any)?.frontmatter?.doi
 
144
  // Published at ICML 2026 AIWILD workshop. Use the canonical OpenReview entry
145
  // rather than an auto-generated @misc (which left a dangling comma when no DOI).
146
  const venue = "Second Workshop on Agents in the Wild: Safety, Security, and Beyond";
147
+ const openreviewUrl = "https://arxiv.org/abs/2608.23670";
148
  const fullTitle = "Automata from Agent Traces: Failure and Next-Step Prediction";
149
  const authorsBib = authorNames.join(" and ");
150
  const doi = (ArticleMod as any)?.frontmatter?.doi
index.html CHANGED
@@ -9,7 +9,7 @@
9
  document.documentElement.setAttribute("data-theme", theme);
10
  } catch {}
11
  })();
12
- </script><link rel="stylesheet" href="/_astro/index.DpTQb45s.css"><script type="module" src="/_astro/hoisted.DzT9C6fG.js"></script></head> <body> <button id="theme-toggle" aria-label="Toggle color theme" data-astro-cid-x3pjskd3> <svg class="icon light" width="20" height="20" viewBox="0 0 24 24" aria-hidden="true" focusable="false" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" data-astro-cid-x3pjskd3> <circle cx="12" cy="12" r="5" data-astro-cid-x3pjskd3></circle> <line x1="12" y1="1" x2="12" y2="4" data-astro-cid-x3pjskd3></line> <line x1="12" y1="20" x2="12" y2="23" data-astro-cid-x3pjskd3></line> <line x1="1" y1="12" x2="4" y2="12" data-astro-cid-x3pjskd3></line> <line x1="20" y1="12" x2="23" y2="12" data-astro-cid-x3pjskd3></line> <line x1="4.22" y1="4.22" x2="6.34" y2="6.34" data-astro-cid-x3pjskd3></line> <line x1="17.66" y1="17.66" x2="19.78" y2="19.78" data-astro-cid-x3pjskd3></line> <line x1="4.22" y1="19.78" x2="6.34" y2="17.66" data-astro-cid-x3pjskd3></line> <line x1="17.66" y1="6.34" x2="19.78" y2="4.22" data-astro-cid-x3pjskd3></line> </svg> <svg class="icon dark" width="20" height="20" viewBox="0 0 24 24" aria-hidden="true" focusable="false" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" data-astro-cid-x3pjskd3> <path d="M21 12.79A9 9 0 1 1 11.21 3 7 7 0 0 0 21 12.79z" data-astro-cid-x3pjskd3></path> </svg> </button> <section class="hero" data-astro-cid-bbe6dxrz> <h1 class="hero-title" data-astro-cid-bbe6dxrz>Automata from Agent Traces</h1> <div class="hero-banner" data-astro-cid-bbe6dxrz> <figure class="html-embed"><div class="html-embed__card is-frameless"><div id="frag-ii2xmcaz23o"><!-- Interactive FSM Hero: flat dashboard-style graph with live trace simulation -->
13
  <div class="fsm-hero">
14
  <div class="hero-pills"></div>
15
  <div class="hero-graph"></div>
@@ -642,7 +642,7 @@
642
  else ensureD3(bootstrap);
643
  })();
644
  </script>
645
- </div></div></figure> <p class="hero-desc" data-astro-cid-bbe6dxrz>One automaton, built from traces, serves as a memory, prediction, and monitoring prior</p> </div> </section> <header class="meta" aria-label="Article meta information" data-astro-cid-bbe6dxrz> <div class="meta-container" data-astro-cid-bbe6dxrz> <div class="meta-container-cell" data-astro-cid-bbe6dxrz> <h3 data-astro-cid-bbe6dxrz>Authors</h3> <ul class="authors" data-astro-cid-bbe6dxrz> <li data-astro-cid-bbe6dxrz> <a href="https://seongland.com" data-astro-cid-bbe6dxrz>Seonglae Cho</a><sup data-astro-cid-bbe6dxrz>1</sup><span data-astro-cid-bbe6dxrz>, </span> </li><li data-astro-cid-bbe6dxrz> Franklin Cardenoso Fernandez<sup data-astro-cid-bbe6dxrz>1, 3</sup><span data-astro-cid-bbe6dxrz>, </span> </li><li data-astro-cid-bbe6dxrz> Umar Mohammed<sup data-astro-cid-bbe6dxrz>1</sup><span data-astro-cid-bbe6dxrz>, </span> </li><li data-astro-cid-bbe6dxrz> <a href="https://981526092.github.io/zekunwu.github.io/" data-astro-cid-bbe6dxrz>Zekun Wu</a><sup data-astro-cid-bbe6dxrz>2</sup><span data-astro-cid-bbe6dxrz>, </span> </li><li data-astro-cid-bbe6dxrz> <a href="https://kleytoncosta.com" data-astro-cid-bbe6dxrz>Kleyton Da Costa</a><sup data-astro-cid-bbe6dxrz>2</sup><span data-astro-cid-bbe6dxrz>, </span> </li><li data-astro-cid-bbe6dxrz> Ilham Wicaksono<sup data-astro-cid-bbe6dxrz>1</sup><span data-astro-cid-bbe6dxrz>, </span> </li><li data-astro-cid-bbe6dxrz> Adriano Koshiyama<sup data-astro-cid-bbe6dxrz>2</sup> </li> </ul> </div> <div class="meta-container-cell meta-container-cell--affiliations" data-astro-cid-bbe6dxrz> <h3 data-astro-cid-bbe6dxrz>Affiliations</h3> <ol class="affiliations" data-astro-cid-bbe6dxrz> <li value="1" data-astro-cid-bbe6dxrz> <a href="https://www.holisticai.com" target="_blank" rel="noopener noreferrer" data-astro-cid-bbe6dxrz> Holistic AI </a> </li><li value="2" data-astro-cid-bbe6dxrz> <a href="https://www.ucl.ac.uk" target="_blank" rel="noopener noreferrer" data-astro-cid-bbe6dxrz> University College London </a> </li><li value="3" data-astro-cid-bbe6dxrz> <a href="https://www.puc-rio.br" target="_blank" rel="noopener noreferrer" data-astro-cid-bbe6dxrz> PUC-Rio </a> </li> </ol> </div> <div class="meta-container-cell meta-container-cell--published" data-astro-cid-bbe6dxrz> <h3 data-astro-cid-bbe6dxrz>Published</h3> <p data-astro-cid-bbe6dxrz>Jun. 30, 2026</p> </div> <div class="meta-container-cell meta-container-cell--links" data-astro-cid-bbe6dxrz> <h3 data-astro-cid-bbe6dxrz>Links</h3> <p class="hero-links" data-astro-cid-bbe6dxrz> <a class="button lk-paper" href="https://openreview.net/forum?id=1cprFkvFT0" target="_blank" rel="noopener noreferrer" data-astro-cid-bbe6dxrz>Paper</a> <a class="button lk-demo" href="https://seongland.com/article/asg/browser" target="_blank" rel="noopener noreferrer" data-astro-cid-bbe6dxrz>Demo</a> <a class="button lk-dash" href="https://seongland.com/article/asg/browser?tab=overview" target="_blank" rel="noopener noreferrer" data-astro-cid-bbe6dxrz>Dashboard</a> </p> </div> <!-- {doi && (
646
  <div class="meta-container-cell">
647
  <h3>DOI</h3>
648
  <p><a href={`https://doi.org/${doi}`} target="_blank" rel="noopener noreferrer">{doi}</a></p>
@@ -1357,7 +1357,7 @@
1357
  <p>Yang and colleagues put LLM-based agents to work resolving GitHub issues <span class="" id="citation--yang2024sweagent--yang2025swesmith--1">(<a href="#bib-yang2024sweagent" id="refctx-bib-yang2024sweagent-1">Yang et al., 2024</a>, 2025<a href="#bib-yang2025swesmith" id="refctx-bib-yang2025swesmith-1">)</a></span>. Others have them navigating websites <span class="" id="citation--zhou2024webarena--2">(<a href="#bib-zhou2024webarena" id="refctx-bib-zhou2024webarena-1">Zhou et al., 2024</a>)</span>, operating desktops <span class="" id="citation--xie2024osworld--3">(<a href="https://openreview.net/forum?id=tN61DTr4Ed" id="refctx-bib-xie2024osworld-1" data-ref-id="bib-xie2024osworld" target="_blank" rel="noopener noreferrer">Xie et al., 2024</a>)</span>, driving mobile interfaces <span class="" id="citation--lu2024guiodyssey--4">(<a href="#bib-lu2024guiodyssey" id="refctx-bib-lu2024guiodyssey-1">Lu et al., 2025</a>)</span>, managing customer service interactions <span class="" id="citation--yao2024taubench--5">(<a href="#bib-yao2024taubench" id="refctx-bib-yao2024taubench-1">Yao et al., 2024</a>)</span>, orchestrating multi-agent pipelines <span class="" id="citation--wu2023autogen--hong2023metagpt--6">(<a href="#bib-hong2023metagpt" id="refctx-bib-hong2023metagpt-1">Hong et al., 2024</a>; <a href="#bib-wu2023autogen" id="refctx-bib-wu2023autogen-1">Wu et al., 2023</a>)</span>, and attributing the blame when one of those pipelines breaks down <span class="" id="citation--yang2025whoandwhen--7">(<a href="#bib-yang2025whoandwhen" id="refctx-bib-yang2025whoandwhen-1">Zhang et al., 2025</a>)</span>. Following the ReAct pattern <span class="" id="citation--yao2023react--8">(<a href="#bib-yao2023react" id="refctx-bib-yao2023react-1">Yao et al., 2023</a>)</span>, they interleave chain-of-thought reasoning <span class="" id="citation--wei2022cot--9">(<a href="#bib-wei2022cot" id="refctx-bib-wei2022cot-1">Wei et al., 2022</a>)</span> with tool calls — and every run produces an execution trace of tool calls, natural language, and environment feedback.</p>
1358
  <p>This article recovers that machine. From nothing but the agent’s own execution traces — no labels, no task descriptions — we extract a compact <a href="https://texonom.com/6ccd60d415e74d07a615b9714ce83510">finite-state machine</a>, or FSM, of 7 to 43 states. That machine does the two things an operator actually needs: predict what the agent will do next, and catch a failing run before it wastes the compute.</p>
1359
  <p>The whole article follows one thread: heterogeneous agent traces pass through one deterministic abstraction <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ϕ</mi></mrow><annotation encoding="application/x-tex">\phi</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em"></span><span class="mord mathnormal">ϕ</span></span></span></span>, pile up into a prefix tree, and collapse in a single classical merge into one automaton. That one object — not four bespoke pipelines — is then read four different ways: as workflow memory, a next-step predictor, a failure detector, and a runtime monitor.</p>
1360
- <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">How the machine is built, and how it is used</figcaption><div class="html-embed__card"><div id="frag-0ks97jpz74sa"><!-- System / data-flow diagram. LEFT->RIGHT: how the machine is BUILT (construction,
1361
  offline): raw traces -> phi abstraction -> prefix tree -> one merge -> automaton.
1362
  RIGHT: how the machine is USED at runtime by four consumers (what each reads from it).
1363
  No result metrics — this shows data flow and execution flow, not numbers.
@@ -1579,7 +1579,7 @@
1579
  <p>A coding agent cycles through <code>search</code>, <code>edit</code>, <code>execute</code>. A customer service agent alternates between database queries and user communication. This structure emerges from the interaction between the system prompt, the available tools, and the task distribution — yet it’s <strong>nowhere written down</strong>.</p>
1580
  <p>Understanding this latent structure matters the moment you deploy: safety auditing has to verify that an agent visits the right states and avoids the attack chains that agent security benchmarks enumerate <span class="" id="citation--zhang2025asb--10">(<a href="#bib-zhang2025asb" id="refctx-bib-zhang2025asb-1">H. Zhang et al., 2025</a>)</span>, debugging means locating the bottleneck states where agents get stuck, and production monitoring flags behavioral drift before it costs anything.</p>
1581
  <p>Yet current agent analysis works at the level of a single trace <span class="" id="citation--wu2025agentgraph--11">(<a href="#bib-wu2025agentgraph" id="refctx-bib-wu2025agentgraph-1">Z. Wu et al., 2025</a>)</span>. Benchmarks such as AgentBench and AgentBoard score whether a run succeeded <span class="" id="citation--liu2024agentbench--ma2024agentboard--12">(<a href="#bib-liu2024agentbench" id="refctx-bib-liu2024agentbench-1">Liu et al., 2024</a>; <a href="#bib-ma2024agentboard" id="refctx-bib-ma2024agentboard-1">Ma et al., 2024</a>)</span>, and sandboxes such as ToolEmu probe what a run risks <span class="" id="citation--ruan2024toolemu--13">(<a href="#bib-ruan2024toolemu" id="refctx-bib-ruan2024toolemu-1">Ruan et al., 2024</a>)</span>. Neither offers a structural model of the behavior that links one run to the next.</p>
1582
- <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">A trace, read symbol by symbol, drives a finite-state machine</figcaption><div class="html-embed__card"><div id="frag-snjdga0pck"><!-- Tape -> FSM: a trace tape feeds a read-head that drives the FSM state (mini of the dashboard landing) -->
1583
  <div class="tape-fsm"></div>
1584
  <style>
1585
  .tape-fsm { position: relative; width: 100%; font-family: inherit; }
@@ -1987,7 +1987,7 @@
1987
  <p>Consider a coding agent trace from SWE-agent <span class="" id="citation--yang2024sweagent--4">(<a href="#bib-yang2024sweagent" id="refctx-bib-yang2024sweagent-1">Yang et al., 2024</a>)</span> with 47 messages. The raw trace contains system prompts, file contents, error messages, and tool invocations. After extraction, the activity sequence is:</p>
1988
  <p style="text-align:center; line-height:2.2;"><p><code>init</code> → <code>user</code> → <code>search</code> → <code>user</code> → <code>edit</code> → <code>user</code> → <code>execute</code> → <code>user</code> → <code>edit</code> → <code>user</code> → <code>submit</code></p></p>
1989
  <p>From 47 messages and thousands of tokens, we get the 11 symbols the FSM will model — drawn from an alphabet of 24 possible activities.</p>
1990
- <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Trace Explorer</figcaption><div class="html-embed__card"><div id="frag-ng6hfzn9hv"><!-- Trace Explorer: Interactive trace-to-symbol mapper -->
1991
  <div class="trace-explorer"></div>
1992
  <style>
1993
  .trace-explorer { position: relative; width: 100%; font-family: inherit; }
@@ -2285,7 +2285,7 @@
2285
  <div class="note note--neutral" data-astro-cid-qg6lmfty> <!-- When there's no title, emoji is above content -->
2286
  <div class="note__layout" data-astro-cid-qg6lmfty> <div class="note__body" data-astro-cid-qg6lmfty> <div class="note__content" data-astro-cid-qg6lmfty> <p>Structural equivalence is exactly the <a href="https://texonom.com/ab9c3c96247d8389af2a8156427385b4">Myhill-Nerode</a> equivalence on the observed prefix language, and by that theorem the quotient is the <strong>unique minimal <a href="https://texonom.com/f08438340ef447f89eade6aac52f2cbb">deterministic finite automaton (DFA)</a></strong>, so no smaller automaton can reproduce the observed behavior <span class="" id="citation--hopcroft2006automata--1">(<a href="#bib-hopcroft2006automata" id="refctx-bib-hopcroft2006automata-1">Hopcroft et al., 2006</a>)</span>.</p> </div> </div> </div> </div>
2287
  <p>The SWE-agent prefix tree with 59,510 states collapses to just <strong>25 states</strong> — a 2,380<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em"></span><span class="mord">×</span></span></span></span> compression — and the resulting FSM still replays held-out traces at 0.999 fitness.</p>
2288
- <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Prefix Tree to FSM Collapse</figcaption><div class="html-embed__card"><div id="frag-ynuommwq6b"><!-- Prefix Tree Collapse: Animated tree → FSM visualization -->
2289
  <div class="prefix-collapse"></div>
2290
  <style>
2291
  .prefix-collapse { position: relative; width: 100%; min-height: 480px; }
@@ -2687,7 +2687,7 @@
2687
  <h2 id="the-resulting-fsm"><a href="#the-resulting-fsm">The Resulting FSM</a></h2>
2688
  <p>The FSM <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="script">M</mi><mo>=</mo><mo stretchy="false">(</mo><mi>Q</mi><mo separator="true">,</mo><mi mathvariant="script">A</mi><mo separator="true">,</mo><mi>δ</mi><mo separator="true">,</mo><msub><mi>q</mi><mn>0</mn></msub><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">\mathcal{M} = (Q, \mathcal{A}, \delta, q_0)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathcal">M</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mopen">(</span><span class="mord mathnormal">Q</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathcal">A</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.03785em">δ</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.03588em">q</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em"><span style="top:-2.55em;margin-left:-0.0359em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">0</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mclose">)</span></span></span></span> encodes the agent’s <strong>behavioral topology</strong>. Recurring patterns become loops. And the state count tracks the number of distinct behavioral modes.</p>
2689
  <p>In the tau2-bench retail and telecom customer service agents, a tool-call loop (<code>assistant:tool_call</code> to <code>tool:text</code>) dominates execution, with the conversational path through <code>assistant:text</code> as a separate branch. We find that in a coding agent the <code>search</code>, <code>edit</code>, <code>execute</code> cycle accounts for most of the trace.</p>
2690
- <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Explorable FSM Graphs</figcaption><div class="html-embed__card"><div id="frag-tiuhx71aiwk"><!-- FSM Force Graph: Explorable FSM - browser-style warm rendering -->
2691
  <div class="fsm-force-graph"></div>
2692
  <style>
2693
  .fsm-force-graph { position: relative; width: 100%; min-height: 500px; }
@@ -3105,7 +3105,7 @@
3105
  <p>We compare against nine baselines: RPNI <span class="" id="citation--oncina1992rpni--1">(<a href="#bib-oncina1992rpni" id="refctx-bib-oncina1992rpni-1">Oncina &amp; Garcı́a, 1992</a>)</span>, EDSM <span class="" id="citation--lang1998edsm--2">(<a href="#bib-lang1998edsm" id="refctx-bib-lang1998edsm-1">Lang et al., 1998</a>)</span>, Alergia <span class="" id="citation--carrasco1994alergia--3">(<a href="#bib-carrasco1994alergia" id="refctx-bib-carrasco1994alergia-1">Carrasco &amp; Oncina, 1994</a>)</span> and k-Tails <span class="" id="citation--biermann1972ktails--4">(<a href="#bib-biermann1972ktails" id="refctx-bib-biermann1972ktails-1">Biermann &amp; Feldman, 1972</a>)</span> from <a href="https://texonom.com/37bc3c96247d803b8156ee3fcfdd4556">automata learning</a>, run through AALpy <span class="" id="citation--muskardin2022aalpy--5">(<a href="#bib-muskardin2022aalpy" id="refctx-bib-muskardin2022aalpy-1">Muškardin et al., 2022</a>)</span>; the HMM <span class="" id="citation--rabiner1989hmm--6">(<a href="#bib-rabiner1989hmm" id="refctx-bib-rabiner1989hmm-1">Rabiner, 1989</a>)</span>; the Alpha, Inductive and Heuristic miners from process mining <span class="" id="citation--vanderaalst2016process--7">(<a href="#bib-vanderaalst2016process" id="refctx-bib-vanderaalst2016process-1">van der Aalst, 2016</a>)</span>, run through PM4Py <span class="" id="citation--berti2019pm4py--8">(<a href="https://arxiv.org/abs/1905.06169" id="refctx-bib-berti2019pm4py-1" data-ref-id="bib-berti2019pm4py" target="_blank" rel="noopener noreferrer">Berti et al., 2019</a>)</span>; and AWM, agent workflow extraction <span class="" id="citation--wang2024agent_workflow_memory--9">(<a href="#bib-wang2024agent_workflow_memory" id="refctx-bib-wang2024agent_workflow_memory-1">Wang et al., 2024</a>)</span>. All receive only the same positive training sequences, no failure labels.</p>
3106
  <h2 id="how-much-smaller"><a href="#how-much-smaller">How Much Smaller</a></h2>
3107
  <p>Our FSMs achieve <strong>15 to 3,036<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em"></span><span class="mord">×</span></span></span></span> compression</strong> over RPNI while replaying held-out traces at fitness of at least 0.997. The ratio grows with trace length and branching: 15<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em"></span><span class="mord">×</span></span></span></span> on WebArena (short web traces, where RPNI succeeds) up to 3,036<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em"></span><span class="mord">×</span></span></span></span> on GUI-Odyssey <span class="" id="citation--lu2024guiodyssey--10">(<a href="#bib-lu2024guiodyssey" id="refctx-bib-lu2024guiodyssey-1">Lu et al., 2025</a>)</span> (long, repetitive mobile-GUI traces, where RPNI’s prefix tree explodes to 21,255 states against our 7).</p>
3108
- <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Compression Ratios</figcaption><div class="html-embed__card"><div id="frag-fofxgorisq9"><!-- Compression: log-scale dumbbell of ours vs RPNI state counts across 12 datasets -->
3109
  <div class="compression-bars"></div>
3110
  <style>
3111
  .compression-bars { position: relative; width: 100%; min-height: 440px; }
@@ -3294,7 +3294,7 @@
3294
  <p>Compare all eight methods interactively in the <a href="https://seongland.com/article/asg/browser?tab=baselines">baselines view</a>, or watch structure stabilize in the <a href="https://seongland.com/article/asg/browser?tab=convergence">convergence view</a>.</p>
3295
  <h2 id="convergence-and-stability"><a href="#convergence-and-stability">Convergence and Stability</a></h2>
3296
  <p>Replay fitness reaches its plateau well before the training set is exhausted: on SWE-agent it’s already at 0.985 within 1% of the training traces and settles at 0.996 by 10%, and the state count keeps inching up as rare command patterns appear. Structured tool-call domains converge fastest: SWE-smith holds 0.9996 from the first 1% of data. Open web and delegation traces take longer, with Mind2Web <span class="" id="citation--deng2024mind2web--11">(<a href="#bib-deng2024mind2web" id="refctx-bib-deng2024mind2web-1">Deng et al., 2023</a>)</span> needing 5% and Who&amp;When <span class="" id="citation--yang2025whoandwhen--12">(<a href="#bib-yang2025whoandwhen" id="refctx-bib-yang2025whoandwhen-1">Zhang et al., 2025</a>)</span> 10% of their traces to clear 0.95 fitness.</p>
3297
- <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Convergence Curves</figcaption><div class="html-embed__card"><div id="frag-to55k18jekr"><!-- Convergence Curves: Small-multiples sparkline cards with convergence gauges -->
3298
  <div class="convergence-curves"></div>
3299
  <style>
3300
  .convergence-curves { position: relative; width: 100%; }
@@ -3551,7 +3551,7 @@
3551
  </script>
3552
  </div></div><figcaption class="html-embed__desc" style="text-align:left">Replay fitness as training traces accumulate, for the four datasets with incremental-convergence runs. The dashed marker shows where each first reaches 0.95 fitness.</figcaption></figure> </div>
3553
  <h2 id="baselines-at-a-glance"><a href="#baselines-at-a-glance">Baselines at a Glance</a></h2>
3554
- <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Method Comparison</figcaption><div class="html-embed__card"><div id="frag-57zux2b7lsc"><!-- Baseline Heatmap: Interactive comparison table of methods x datasets -->
3555
  <div class="baseline-heatmap"></div>
3556
  <style>
3557
  .baseline-heatmap { position: relative; width: 100%; }
@@ -3903,7 +3903,7 @@
3903
  <div class="note note--neutral" data-astro-cid-qg6lmfty> <!-- When there's no title, emoji is above content -->
3904
  <div class="note__layout" data-astro-cid-qg6lmfty> <div class="note__body" data-astro-cid-qg6lmfty> <div class="note__content" data-astro-cid-qg6lmfty> <p>Raw fitness is useless here (AUROC near 0.50): successful and failed traces both replay perfectly. But the signal is in the <strong>per-state decomposition</strong> and in <em>surprise</em>: failing traces take low-probability transitions under the FSM.</p> </div> </div> </div> </div>
3905
  <p>Failure prediction scales with machine size: more states give a finer map of where a run can go wrong. The 43-state telecom agent tops out at <strong>0.941</strong>; WebArena <span class="" id="citation--zhou2024webarena--11">(<a href="#bib-zhou2024webarena" id="refctx-bib-zhou2024webarena-2">Zhou et al., 2024</a>)</span> (0.903) and AgentNet <span class="" id="citation--wang2025opencua--12">(<a href="#bib-wang2025opencua" id="refctx-bib-wang2025opencua-1">X. Wang et al., 2025</a>)</span> (0.890) follow; SWE-agent <span class="" id="citation--yang2024sweagent--13">(<a href="#bib-yang2024sweagent" id="refctx-bib-yang2024sweagent-2">Yang et al., 2024</a>)</span> — with 25 states — reaches 0.799. ATBench, the only safety-labeled benchmark, reaches 0.894 (0.864 ± 0.024 under repeated CV). We see across all eight real-trace datasets the CV standard deviation stays in 0.012 to 0.031 — so these aren’t single-split artifacts.</p>
3906
- <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Failure Prediction</figcaption><div class="html-embed__card"><div id="frag-d7wkry9uuyn"><!-- Failure Features: Dumbbell chart with CV + Holdout AUROC and top features -->
3907
  <div class="failure-features"></div>
3908
  <style>
3909
  .failure-features { position: relative; width: 100%; min-height: 420px; }
@@ -4168,7 +4168,7 @@
4168
  author={Seonglae Cho and Franklin Cardenoso Fernandez and Umar Mohammed and Zekun Wu and Kleyton Da Costa and Ilham Wicaksono and Adriano Koshiyama},
4169
  booktitle={Second Workshop on Agents in the Wild: Safety, Security, and Beyond},
4170
  year={2026},
4171
- url={https://openreview.net/forum?id=1cprFkvFT0},
4172
  howpublished={\url{https://seongland.com/article/asg}}
4173
  }</pre> </section> <section class="reuse-block"> <h3>Reuse</h3> <p>Diagrams and text are licensed under <a href="https://creativecommons.org/licenses/by/4.0/" target="_blank" rel="noopener noreferrer">CC-BY 4.0</a>.
4174
  </p> </section> <section class="references-block"> </section> <div class="template-credit"> <p>
 
9
  document.documentElement.setAttribute("data-theme", theme);
10
  } catch {}
11
  })();
12
+ </script><link rel="stylesheet" href="/_astro/index.DpTQb45s.css"><script type="module" src="/_astro/hoisted.DzT9C6fG.js"></script></head> <body> <button id="theme-toggle" aria-label="Toggle color theme" data-astro-cid-x3pjskd3> <svg class="icon light" width="20" height="20" viewBox="0 0 24 24" aria-hidden="true" focusable="false" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" data-astro-cid-x3pjskd3> <circle cx="12" cy="12" r="5" data-astro-cid-x3pjskd3></circle> <line x1="12" y1="1" x2="12" y2="4" data-astro-cid-x3pjskd3></line> <line x1="12" y1="20" x2="12" y2="23" data-astro-cid-x3pjskd3></line> <line x1="1" y1="12" x2="4" y2="12" data-astro-cid-x3pjskd3></line> <line x1="20" y1="12" x2="23" y2="12" data-astro-cid-x3pjskd3></line> <line x1="4.22" y1="4.22" x2="6.34" y2="6.34" data-astro-cid-x3pjskd3></line> <line x1="17.66" y1="17.66" x2="19.78" y2="19.78" data-astro-cid-x3pjskd3></line> <line x1="4.22" y1="19.78" x2="6.34" y2="17.66" data-astro-cid-x3pjskd3></line> <line x1="17.66" y1="6.34" x2="19.78" y2="4.22" data-astro-cid-x3pjskd3></line> </svg> <svg class="icon dark" width="20" height="20" viewBox="0 0 24 24" aria-hidden="true" focusable="false" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" data-astro-cid-x3pjskd3> <path d="M21 12.79A9 9 0 1 1 11.21 3 7 7 0 0 0 21 12.79z" data-astro-cid-x3pjskd3></path> </svg> </button> <section class="hero" data-astro-cid-bbe6dxrz> <h1 class="hero-title" data-astro-cid-bbe6dxrz>Automata from Agent Traces</h1> <div class="hero-banner" data-astro-cid-bbe6dxrz> <figure class="html-embed"><div class="html-embed__card is-frameless"><div id="frag-hsvhw4lir2a"><!-- Interactive FSM Hero: flat dashboard-style graph with live trace simulation -->
13
  <div class="fsm-hero">
14
  <div class="hero-pills"></div>
15
  <div class="hero-graph"></div>
 
642
  else ensureD3(bootstrap);
643
  })();
644
  </script>
645
+ </div></div></figure> <p class="hero-desc" data-astro-cid-bbe6dxrz>One automaton, built from traces, serves as a memory, prediction, and monitoring prior</p> </div> </section> <header class="meta" aria-label="Article meta information" data-astro-cid-bbe6dxrz> <div class="meta-container" data-astro-cid-bbe6dxrz> <div class="meta-container-cell" data-astro-cid-bbe6dxrz> <h3 data-astro-cid-bbe6dxrz>Authors</h3> <ul class="authors" data-astro-cid-bbe6dxrz> <li data-astro-cid-bbe6dxrz> <a href="https://seongland.com" data-astro-cid-bbe6dxrz>Seonglae Cho</a><sup data-astro-cid-bbe6dxrz>1</sup><span data-astro-cid-bbe6dxrz>, </span> </li><li data-astro-cid-bbe6dxrz> Franklin Cardenoso Fernandez<sup data-astro-cid-bbe6dxrz>1, 3</sup><span data-astro-cid-bbe6dxrz>, </span> </li><li data-astro-cid-bbe6dxrz> Umar Mohammed<sup data-astro-cid-bbe6dxrz>1</sup><span data-astro-cid-bbe6dxrz>, </span> </li><li data-astro-cid-bbe6dxrz> <a href="https://981526092.github.io/zekunwu.github.io/" data-astro-cid-bbe6dxrz>Zekun Wu</a><sup data-astro-cid-bbe6dxrz>2</sup><span data-astro-cid-bbe6dxrz>, </span> </li><li data-astro-cid-bbe6dxrz> <a href="https://kleytoncosta.com" data-astro-cid-bbe6dxrz>Kleyton Da Costa</a><sup data-astro-cid-bbe6dxrz>2</sup><span data-astro-cid-bbe6dxrz>, </span> </li><li data-astro-cid-bbe6dxrz> Ilham Wicaksono<sup data-astro-cid-bbe6dxrz>1</sup><span data-astro-cid-bbe6dxrz>, </span> </li><li data-astro-cid-bbe6dxrz> Adriano Koshiyama<sup data-astro-cid-bbe6dxrz>2</sup> </li> </ul> </div> <div class="meta-container-cell meta-container-cell--affiliations" data-astro-cid-bbe6dxrz> <h3 data-astro-cid-bbe6dxrz>Affiliations</h3> <ol class="affiliations" data-astro-cid-bbe6dxrz> <li value="1" data-astro-cid-bbe6dxrz> <a href="https://www.holisticai.com" target="_blank" rel="noopener noreferrer" data-astro-cid-bbe6dxrz> Holistic AI </a> </li><li value="2" data-astro-cid-bbe6dxrz> <a href="https://www.ucl.ac.uk" target="_blank" rel="noopener noreferrer" data-astro-cid-bbe6dxrz> University College London </a> </li><li value="3" data-astro-cid-bbe6dxrz> <a href="https://www.puc-rio.br" target="_blank" rel="noopener noreferrer" data-astro-cid-bbe6dxrz> PUC-Rio </a> </li> </ol> </div> <div class="meta-container-cell meta-container-cell--published" data-astro-cid-bbe6dxrz> <h3 data-astro-cid-bbe6dxrz>Published</h3> <p data-astro-cid-bbe6dxrz>Jun. 30, 2026</p> </div> <div class="meta-container-cell meta-container-cell--links" data-astro-cid-bbe6dxrz> <h3 data-astro-cid-bbe6dxrz>Links</h3> <p class="hero-links" data-astro-cid-bbe6dxrz> <a class="button lk-paper" href="https://arxiv.org/abs/2608.23670" target="_blank" rel="noopener noreferrer" data-astro-cid-bbe6dxrz>Paper</a> <a class="button lk-demo" href="https://seongland.com/article/asg/browser" target="_blank" rel="noopener noreferrer" data-astro-cid-bbe6dxrz>Demo</a> <a class="button lk-dash" href="https://seongland.com/article/asg/browser?tab=overview" target="_blank" rel="noopener noreferrer" data-astro-cid-bbe6dxrz>Dashboard</a> </p> </div> <!-- {doi && (
646
  <div class="meta-container-cell">
647
  <h3>DOI</h3>
648
  <p><a href={`https://doi.org/${doi}`} target="_blank" rel="noopener noreferrer">{doi}</a></p>
 
1357
  <p>Yang and colleagues put LLM-based agents to work resolving GitHub issues <span class="" id="citation--yang2024sweagent--yang2025swesmith--1">(<a href="#bib-yang2024sweagent" id="refctx-bib-yang2024sweagent-1">Yang et al., 2024</a>, 2025<a href="#bib-yang2025swesmith" id="refctx-bib-yang2025swesmith-1">)</a></span>. Others have them navigating websites <span class="" id="citation--zhou2024webarena--2">(<a href="#bib-zhou2024webarena" id="refctx-bib-zhou2024webarena-1">Zhou et al., 2024</a>)</span>, operating desktops <span class="" id="citation--xie2024osworld--3">(<a href="https://openreview.net/forum?id=tN61DTr4Ed" id="refctx-bib-xie2024osworld-1" data-ref-id="bib-xie2024osworld" target="_blank" rel="noopener noreferrer">Xie et al., 2024</a>)</span>, driving mobile interfaces <span class="" id="citation--lu2024guiodyssey--4">(<a href="#bib-lu2024guiodyssey" id="refctx-bib-lu2024guiodyssey-1">Lu et al., 2025</a>)</span>, managing customer service interactions <span class="" id="citation--yao2024taubench--5">(<a href="#bib-yao2024taubench" id="refctx-bib-yao2024taubench-1">Yao et al., 2024</a>)</span>, orchestrating multi-agent pipelines <span class="" id="citation--wu2023autogen--hong2023metagpt--6">(<a href="#bib-hong2023metagpt" id="refctx-bib-hong2023metagpt-1">Hong et al., 2024</a>; <a href="#bib-wu2023autogen" id="refctx-bib-wu2023autogen-1">Wu et al., 2023</a>)</span>, and attributing the blame when one of those pipelines breaks down <span class="" id="citation--yang2025whoandwhen--7">(<a href="#bib-yang2025whoandwhen" id="refctx-bib-yang2025whoandwhen-1">Zhang et al., 2025</a>)</span>. Following the ReAct pattern <span class="" id="citation--yao2023react--8">(<a href="#bib-yao2023react" id="refctx-bib-yao2023react-1">Yao et al., 2023</a>)</span>, they interleave chain-of-thought reasoning <span class="" id="citation--wei2022cot--9">(<a href="#bib-wei2022cot" id="refctx-bib-wei2022cot-1">Wei et al., 2022</a>)</span> with tool calls — and every run produces an execution trace of tool calls, natural language, and environment feedback.</p>
1358
  <p>This article recovers that machine. From nothing but the agent’s own execution traces — no labels, no task descriptions — we extract a compact <a href="https://texonom.com/6ccd60d415e74d07a615b9714ce83510">finite-state machine</a>, or FSM, of 7 to 43 states. That machine does the two things an operator actually needs: predict what the agent will do next, and catch a failing run before it wastes the compute.</p>
1359
  <p>The whole article follows one thread: heterogeneous agent traces pass through one deterministic abstraction <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ϕ</mi></mrow><annotation encoding="application/x-tex">\phi</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em"></span><span class="mord mathnormal">ϕ</span></span></span></span>, pile up into a prefix tree, and collapse in a single classical merge into one automaton. That one object — not four bespoke pipelines — is then read four different ways: as workflow memory, a next-step predictor, a failure detector, and a runtime monitor.</p>
1360
+ <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">How the machine is built, and how it is used</figcaption><div class="html-embed__card"><div id="frag-1qhaojmay9y"><!-- System / data-flow diagram. LEFT->RIGHT: how the machine is BUILT (construction,
1361
  offline): raw traces -> phi abstraction -> prefix tree -> one merge -> automaton.
1362
  RIGHT: how the machine is USED at runtime by four consumers (what each reads from it).
1363
  No result metrics — this shows data flow and execution flow, not numbers.
 
1579
  <p>A coding agent cycles through <code>search</code>, <code>edit</code>, <code>execute</code>. A customer service agent alternates between database queries and user communication. This structure emerges from the interaction between the system prompt, the available tools, and the task distribution — yet it’s <strong>nowhere written down</strong>.</p>
1580
  <p>Understanding this latent structure matters the moment you deploy: safety auditing has to verify that an agent visits the right states and avoids the attack chains that agent security benchmarks enumerate <span class="" id="citation--zhang2025asb--10">(<a href="#bib-zhang2025asb" id="refctx-bib-zhang2025asb-1">H. Zhang et al., 2025</a>)</span>, debugging means locating the bottleneck states where agents get stuck, and production monitoring flags behavioral drift before it costs anything.</p>
1581
  <p>Yet current agent analysis works at the level of a single trace <span class="" id="citation--wu2025agentgraph--11">(<a href="#bib-wu2025agentgraph" id="refctx-bib-wu2025agentgraph-1">Z. Wu et al., 2025</a>)</span>. Benchmarks such as AgentBench and AgentBoard score whether a run succeeded <span class="" id="citation--liu2024agentbench--ma2024agentboard--12">(<a href="#bib-liu2024agentbench" id="refctx-bib-liu2024agentbench-1">Liu et al., 2024</a>; <a href="#bib-ma2024agentboard" id="refctx-bib-ma2024agentboard-1">Ma et al., 2024</a>)</span>, and sandboxes such as ToolEmu probe what a run risks <span class="" id="citation--ruan2024toolemu--13">(<a href="#bib-ruan2024toolemu" id="refctx-bib-ruan2024toolemu-1">Ruan et al., 2024</a>)</span>. Neither offers a structural model of the behavior that links one run to the next.</p>
1582
+ <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">A trace, read symbol by symbol, drives a finite-state machine</figcaption><div class="html-embed__card"><div id="frag-heb6hyv9u7v"><!-- Tape -> FSM: a trace tape feeds a read-head that drives the FSM state (mini of the dashboard landing) -->
1583
  <div class="tape-fsm"></div>
1584
  <style>
1585
  .tape-fsm { position: relative; width: 100%; font-family: inherit; }
 
1987
  <p>Consider a coding agent trace from SWE-agent <span class="" id="citation--yang2024sweagent--4">(<a href="#bib-yang2024sweagent" id="refctx-bib-yang2024sweagent-1">Yang et al., 2024</a>)</span> with 47 messages. The raw trace contains system prompts, file contents, error messages, and tool invocations. After extraction, the activity sequence is:</p>
1988
  <p style="text-align:center; line-height:2.2;"><p><code>init</code> → <code>user</code> → <code>search</code> → <code>user</code> → <code>edit</code> → <code>user</code> → <code>execute</code> → <code>user</code> → <code>edit</code> → <code>user</code> → <code>submit</code></p></p>
1989
  <p>From 47 messages and thousands of tokens, we get the 11 symbols the FSM will model — drawn from an alphabet of 24 possible activities.</p>
1990
+ <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Trace Explorer</figcaption><div class="html-embed__card"><div id="frag-1js91fhlzu8"><!-- Trace Explorer: Interactive trace-to-symbol mapper -->
1991
  <div class="trace-explorer"></div>
1992
  <style>
1993
  .trace-explorer { position: relative; width: 100%; font-family: inherit; }
 
2285
  <div class="note note--neutral" data-astro-cid-qg6lmfty> <!-- When there's no title, emoji is above content -->
2286
  <div class="note__layout" data-astro-cid-qg6lmfty> <div class="note__body" data-astro-cid-qg6lmfty> <div class="note__content" data-astro-cid-qg6lmfty> <p>Structural equivalence is exactly the <a href="https://texonom.com/ab9c3c96247d8389af2a8156427385b4">Myhill-Nerode</a> equivalence on the observed prefix language, and by that theorem the quotient is the <strong>unique minimal <a href="https://texonom.com/f08438340ef447f89eade6aac52f2cbb">deterministic finite automaton (DFA)</a></strong>, so no smaller automaton can reproduce the observed behavior <span class="" id="citation--hopcroft2006automata--1">(<a href="#bib-hopcroft2006automata" id="refctx-bib-hopcroft2006automata-1">Hopcroft et al., 2006</a>)</span>.</p> </div> </div> </div> </div>
2287
  <p>The SWE-agent prefix tree with 59,510 states collapses to just <strong>25 states</strong> — a 2,380<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em"></span><span class="mord">×</span></span></span></span> compression — and the resulting FSM still replays held-out traces at 0.999 fitness.</p>
2288
+ <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Prefix Tree to FSM Collapse</figcaption><div class="html-embed__card"><div id="frag-qmk1rbd7hyb"><!-- Prefix Tree Collapse: Animated tree → FSM visualization -->
2289
  <div class="prefix-collapse"></div>
2290
  <style>
2291
  .prefix-collapse { position: relative; width: 100%; min-height: 480px; }
 
2687
  <h2 id="the-resulting-fsm"><a href="#the-resulting-fsm">The Resulting FSM</a></h2>
2688
  <p>The FSM <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="script">M</mi><mo>=</mo><mo stretchy="false">(</mo><mi>Q</mi><mo separator="true">,</mo><mi mathvariant="script">A</mi><mo separator="true">,</mo><mi>δ</mi><mo separator="true">,</mo><msub><mi>q</mi><mn>0</mn></msub><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">\mathcal{M} = (Q, \mathcal{A}, \delta, q_0)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em"></span><span class="mord mathcal">M</span><span class="mspace" style="margin-right:0.2778em"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em"></span><span class="mopen">(</span><span class="mord mathnormal">Q</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathcal">A</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord mathnormal" style="margin-right:0.03785em">δ</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.03588em">q</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em"><span style="top:-2.55em;margin-left:-0.0359em;margin-right:0.05em"><span class="pstrut" style="height:2.7em"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">0</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em"><span></span></span></span></span></span></span><span class="mclose">)</span></span></span></span> encodes the agent’s <strong>behavioral topology</strong>. Recurring patterns become loops. And the state count tracks the number of distinct behavioral modes.</p>
2689
  <p>In the tau2-bench retail and telecom customer service agents, a tool-call loop (<code>assistant:tool_call</code> to <code>tool:text</code>) dominates execution, with the conversational path through <code>assistant:text</code> as a separate branch. We find that in a coding agent the <code>search</code>, <code>edit</code>, <code>execute</code> cycle accounts for most of the trace.</p>
2690
+ <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Explorable FSM Graphs</figcaption><div class="html-embed__card"><div id="frag-gaz7qnjakj7"><!-- FSM Force Graph: Explorable FSM - browser-style warm rendering -->
2691
  <div class="fsm-force-graph"></div>
2692
  <style>
2693
  .fsm-force-graph { position: relative; width: 100%; min-height: 500px; }
 
3105
  <p>We compare against nine baselines: RPNI <span class="" id="citation--oncina1992rpni--1">(<a href="#bib-oncina1992rpni" id="refctx-bib-oncina1992rpni-1">Oncina &amp; Garcı́a, 1992</a>)</span>, EDSM <span class="" id="citation--lang1998edsm--2">(<a href="#bib-lang1998edsm" id="refctx-bib-lang1998edsm-1">Lang et al., 1998</a>)</span>, Alergia <span class="" id="citation--carrasco1994alergia--3">(<a href="#bib-carrasco1994alergia" id="refctx-bib-carrasco1994alergia-1">Carrasco &amp; Oncina, 1994</a>)</span> and k-Tails <span class="" id="citation--biermann1972ktails--4">(<a href="#bib-biermann1972ktails" id="refctx-bib-biermann1972ktails-1">Biermann &amp; Feldman, 1972</a>)</span> from <a href="https://texonom.com/37bc3c96247d803b8156ee3fcfdd4556">automata learning</a>, run through AALpy <span class="" id="citation--muskardin2022aalpy--5">(<a href="#bib-muskardin2022aalpy" id="refctx-bib-muskardin2022aalpy-1">Muškardin et al., 2022</a>)</span>; the HMM <span class="" id="citation--rabiner1989hmm--6">(<a href="#bib-rabiner1989hmm" id="refctx-bib-rabiner1989hmm-1">Rabiner, 1989</a>)</span>; the Alpha, Inductive and Heuristic miners from process mining <span class="" id="citation--vanderaalst2016process--7">(<a href="#bib-vanderaalst2016process" id="refctx-bib-vanderaalst2016process-1">van der Aalst, 2016</a>)</span>, run through PM4Py <span class="" id="citation--berti2019pm4py--8">(<a href="https://arxiv.org/abs/1905.06169" id="refctx-bib-berti2019pm4py-1" data-ref-id="bib-berti2019pm4py" target="_blank" rel="noopener noreferrer">Berti et al., 2019</a>)</span>; and AWM, agent workflow extraction <span class="" id="citation--wang2024agent_workflow_memory--9">(<a href="#bib-wang2024agent_workflow_memory" id="refctx-bib-wang2024agent_workflow_memory-1">Wang et al., 2024</a>)</span>. All receive only the same positive training sequences, no failure labels.</p>
3106
  <h2 id="how-much-smaller"><a href="#how-much-smaller">How Much Smaller</a></h2>
3107
  <p>Our FSMs achieve <strong>15 to 3,036<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em"></span><span class="mord">×</span></span></span></span> compression</strong> over RPNI while replaying held-out traces at fitness of at least 0.997. The ratio grows with trace length and branching: 15<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em"></span><span class="mord">×</span></span></span></span> on WebArena (short web traces, where RPNI succeeds) up to 3,036<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em"></span><span class="mord">×</span></span></span></span> on GUI-Odyssey <span class="" id="citation--lu2024guiodyssey--10">(<a href="#bib-lu2024guiodyssey" id="refctx-bib-lu2024guiodyssey-1">Lu et al., 2025</a>)</span> (long, repetitive mobile-GUI traces, where RPNI’s prefix tree explodes to 21,255 states against our 7).</p>
3108
+ <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Compression Ratios</figcaption><div class="html-embed__card"><div id="frag-0f6dvhcb8vn5"><!-- Compression: log-scale dumbbell of ours vs RPNI state counts across 12 datasets -->
3109
  <div class="compression-bars"></div>
3110
  <style>
3111
  .compression-bars { position: relative; width: 100%; min-height: 440px; }
 
3294
  <p>Compare all eight methods interactively in the <a href="https://seongland.com/article/asg/browser?tab=baselines">baselines view</a>, or watch structure stabilize in the <a href="https://seongland.com/article/asg/browser?tab=convergence">convergence view</a>.</p>
3295
  <h2 id="convergence-and-stability"><a href="#convergence-and-stability">Convergence and Stability</a></h2>
3296
  <p>Replay fitness reaches its plateau well before the training set is exhausted: on SWE-agent it’s already at 0.985 within 1% of the training traces and settles at 0.996 by 10%, and the state count keeps inching up as rare command patterns appear. Structured tool-call domains converge fastest: SWE-smith holds 0.9996 from the first 1% of data. Open web and delegation traces take longer, with Mind2Web <span class="" id="citation--deng2024mind2web--11">(<a href="#bib-deng2024mind2web" id="refctx-bib-deng2024mind2web-1">Deng et al., 2023</a>)</span> needing 5% and Who&amp;When <span class="" id="citation--yang2025whoandwhen--12">(<a href="#bib-yang2025whoandwhen" id="refctx-bib-yang2025whoandwhen-1">Zhang et al., 2025</a>)</span> 10% of their traces to clear 0.95 fitness.</p>
3297
+ <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Convergence Curves</figcaption><div class="html-embed__card"><div id="frag-ujn7qtudt98"><!-- Convergence Curves: Small-multiples sparkline cards with convergence gauges -->
3298
  <div class="convergence-curves"></div>
3299
  <style>
3300
  .convergence-curves { position: relative; width: 100%; }
 
3551
  </script>
3552
  </div></div><figcaption class="html-embed__desc" style="text-align:left">Replay fitness as training traces accumulate, for the four datasets with incremental-convergence runs. The dashed marker shows where each first reaches 0.95 fitness.</figcaption></figure> </div>
3553
  <h2 id="baselines-at-a-glance"><a href="#baselines-at-a-glance">Baselines at a Glance</a></h2>
3554
+ <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Method Comparison</figcaption><div class="html-embed__card"><div id="frag-daq4o60t5z7"><!-- Baseline Heatmap: Interactive comparison table of methods x datasets -->
3555
  <div class="baseline-heatmap"></div>
3556
  <style>
3557
  .baseline-heatmap { position: relative; width: 100%; }
 
3903
  <div class="note note--neutral" data-astro-cid-qg6lmfty> <!-- When there's no title, emoji is above content -->
3904
  <div class="note__layout" data-astro-cid-qg6lmfty> <div class="note__body" data-astro-cid-qg6lmfty> <div class="note__content" data-astro-cid-qg6lmfty> <p>Raw fitness is useless here (AUROC near 0.50): successful and failed traces both replay perfectly. But the signal is in the <strong>per-state decomposition</strong> and in <em>surprise</em>: failing traces take low-probability transitions under the FSM.</p> </div> </div> </div> </div>
3905
  <p>Failure prediction scales with machine size: more states give a finer map of where a run can go wrong. The 43-state telecom agent tops out at <strong>0.941</strong>; WebArena <span class="" id="citation--zhou2024webarena--11">(<a href="#bib-zhou2024webarena" id="refctx-bib-zhou2024webarena-2">Zhou et al., 2024</a>)</span> (0.903) and AgentNet <span class="" id="citation--wang2025opencua--12">(<a href="#bib-wang2025opencua" id="refctx-bib-wang2025opencua-1">X. Wang et al., 2025</a>)</span> (0.890) follow; SWE-agent <span class="" id="citation--yang2024sweagent--13">(<a href="#bib-yang2024sweagent" id="refctx-bib-yang2024sweagent-2">Yang et al., 2024</a>)</span> — with 25 states — reaches 0.799. ATBench, the only safety-labeled benchmark, reaches 0.894 (0.864 ± 0.024 under repeated CV). We see across all eight real-trace datasets the CV standard deviation stays in 0.012 to 0.031 — so these aren’t single-split artifacts.</p>
3906
+ <div class="wide"> <figure class="html-embed"><figcaption class="html-embed__title" style="text-align:left">Failure Prediction</figcaption><div class="html-embed__card"><div id="frag-71atzyto7i6"><!-- Failure Features: Dumbbell chart with CV + Holdout AUROC and top features -->
3907
  <div class="failure-features"></div>
3908
  <style>
3909
  .failure-features { position: relative; width: 100%; min-height: 420px; }
 
4168
  author={Seonglae Cho and Franklin Cardenoso Fernandez and Umar Mohammed and Zekun Wu and Kleyton Da Costa and Ilham Wicaksono and Adriano Koshiyama},
4169
  booktitle={Second Workshop on Agents in the Wild: Safety, Security, and Beyond},
4170
  year={2026},
4171
+ url={https://arxiv.org/abs/2608.23670},
4172
  howpublished={\url{https://seongland.com/article/asg}}
4173
  }</pre> </section> <section class="reuse-block"> <h3>Reuse</h3> <p>Diagrams and text are licensed under <a href="https://creativecommons.org/licenses/by/4.0/" target="_blank" rel="noopener noreferrer">CC-BY 4.0</a>.
4174
  </p> </section> <section class="references-block"> </section> <div class="template-credit"> <p>
index.html.gz CHANGED
Binary files a/index.html.gz and b/index.html.gz differ