Spaces:
Sleeping
Sleeping
File size: 33,757 Bytes
116524e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 | # ACE Architecture
> Architecture for ACE, a pipeline-based framework for building self-improving AI agents. Roles are backed by PydanticAI agents; the pipeline engine handles composition and concurrency.
For full code examples and API reference, see [ACE_REFERENCE.md](ACE_REFERENCE.md).
For design decisions and rejected alternatives, see [ACE_DECISIONS.md](ACE_DECISIONS.md).
For the pipeline engine, see [PIPELINE_DESIGN.md](PIPELINE_DESIGN.md).
---
## Overview
ACE (Agentic Context Engine) builds AI agents that learn from their own executions. It combines:
- A **pipeline engine** (`pipeline/`) with typed step contracts, concurrent execution, and structured error handling
- **Roles** (Agent, Reflector, SkillManager) backed by PydanticAI agents for structured LLM interactions
- A **Skillbook** β an evolving knowledge base of strategies that agents read from and learning loops write to
- **Integration steps** for external frameworks (browser-use, LangChain, Claude Code, Anthropic SDK)
- **Observability** via Logfire auto-instrumentation of all PydanticAI agent calls
The LLM interaction layer uses PydanticAI exclusively. Three legacy hand-rolled LLM clients (LiteLLM, Instructor, ClaudeCode) were replaced β PydanticAI handles structured output, retries with error feedback, and multi-provider support as maintained infrastructure. The pipeline engine and skillbook/learning loop are untouched.
| Kept (core IP) | Replaced (commodity plumbing) |
|---|---|
| Pipeline engine (`requires`/`provides`, `async_boundary`, `max_workers`) | LLM client abstraction (3 implementations β PydanticAI agents) |
| Skillbook & learning loop (Reflect β Update β Apply) | Structured output parsing + retries (β PydanticAI native validation) |
| Step composition (`learning_tail`, pipeline nesting) | RR iteration loop, code extraction, budget tracking (~2,500 lines β PydanticAI agent + tools) |
| Domain-specific prompts | Sub-agent call management (CallBudget β `UsageLimits`) |
---
## Naming
| Legacy | Current | What it does |
|---|---|---|
| `OfflineACE` | `TraceAnalyser` | Analyse pre-recorded traces β evolve a skillbook |
| `OnlineACE` | `ACE` | Live execution β feedback β learning loop |
| `ACEBase` | `ACERunner` | Shared runner infrastructure (composition, not inheritance from Pipeline) |
| `ACEStepResult` | Removed β use `SampleResult` from the pipeline engine | Unified result type |
---
## Architecture Layers
The framework separates concerns into four layers:
| Layer | Location | Responsibility | Example |
|-------|----------|----------------|---------|
| **Protocols** | `ace/protocols/` | Interface contracts | `ReflectorLike.reflect()` |
| **Roles** | `ace/implementations/` | Business logic (LLM calls) | `Reflector`, `RRStep` |
| **Steps** | `ace/steps/` | Context plumbing (extract β call role β put back) | `ReflectStep` |
| **Runners** | `ace/runners/` | Orchestration (sample loop, epoch management) | `ACELiteLLM` |
**Protocols** define what a role must look like. **Roles** implement the logic. **Steps** adapt between the pipeline's context-based data flow and the role's parameter-based API. **Runners** compose steps into pipelines and iterate over inputs.
Roles are interchangeable anywhere their protocol is expected β both `Reflector` (simple single-pass) and `RRStep` (recursive multi-iteration) satisfy `ReflectorLike`. The runner and pipeline don't know or care which one is in use.
---
## Core Concepts
### Sample
The input unit for ACE. A question with optional context and ground truth:
```python
@dataclass
class Sample:
question: str
context: str = ""
ground_truth: str | None = None
metadata: dict = field(default_factory=dict)
id: str | None = None
```
### ACESample β protocol for step access
Steps access `ctx.sample.question` uniformly. A `Protocol` makes this duck typing explicit and type-safe. `Sample` satisfies it structurally β no inheritance required.
### SkillbookView β read-only projection
The `Skillbook` is mutable β steps add, update, and remove skills. Placing it directly on a `frozen=True` context would allow mutation through the reference, breaking the immutability guarantee.
`SkillbookView` wraps a `Skillbook` and exposes only read methods (`as_prompt()`, `get_skill()`, `skills()`, `stats()`). Write methods don't exist on the class β calling them raises `AttributeError` at runtime and a type error at check time.
**Enforcement:**
- **Type checker** β mypy/pyright flags `ctx.skillbook.add_skill(...)` because `SkillbookView` has no such method.
- **Runtime** β `AttributeError` if someone calls a write method anyway.
- **Convention** β the underlying `_sb` is underscore-prefixed. Accessing it is a deliberate violation.
Steps that only **read** the skillbook (ReflectStep) access `ctx.skillbook` β the view. Steps that **write** the skillbook (AgentStep, UpdateStep, DeduplicateStep, CheckpointStep) receive the real `Skillbook` via constructor injection and use `self.skillbook`. `AgentStep` bumps `used_count`; `UpdateStep` invokes the agentic SkillManager whose tools apply ADD / UPDATE / REMOVE / TAG directly.
### ACEStepContext β immutable step-to-step data
Subclass of the pipeline engine's `StepContext`. Carries all step-to-step data for the ACE pipeline. The pipeline engine only knows about `sample` and `metadata`; all ACE-specific fields live here.
Key fields:
| Field | Type | Source |
|---|---|---|
| `mode` | `Literal["online", "offline"]` | `"online"` (default) β reserved for downstream steps |
| `sample` | `ACESample \| None` | Set by runner's `_build_context()` |
| `skillbook` | `SkillbookView \| None` | Read-only projection of the real Skillbook |
| `trace` | `object \| None` | Raw execution record β any type, no enforced schema |
| `agent_output` | `AgentOutput \| None` | Produced by `AgentStep` |
| `reflections` | `tuple[ReflectorOutput, ...]` | Produced by `ReflectStep` / `RRStep` |
| `skill_manager_output` | `UpdateBatch \| None` | Produced by `UpdateStep` (audit log of mutations the SM already applied) |
| `injected_skill_ids` | `tuple[str, ...]` | Produced by `AgentStep` β skill IDs rendered into the agent prompt; downstream attribution scope |
| `epoch`, `total_epochs` | `int` | Runner bookkeeping |
| `step_index`, `total_steps` | `int` | Runner bookkeeping |
| `global_sample_index` | `int` | Runner bookkeeping (used by interval steps) |
The `trace` field holds the raw execution record from any external system β a browser-use `AgentHistoryList`, a LangChain result dict, a Claude Code transcript, or any arbitrary Python object. The Reflector receives the raw trace and is responsible for making sense of it.
The `reflections` field is a tuple. In single-trace mode, it's a 1-tuple. In batch mode, it holds one `ReflectorOutput` per trace. Downstream steps iterate uniformly β no special-casing.
### Context vs constructor injection
| | On the context | Injected via constructor |
|---|---|---|
| **Nature** | Step-to-step data + read-only dependencies | Mutable shared state |
| **Lifetime** | Per-sample (born in `_build_context`, dies after pipeline) | Per-runner (created once, shared across samples) |
| **Immutable?** | Yes β frozen fields, read-only views | No β mutable by design |
| **Examples** | `agent_output`, `reflections`, `skillbook` (view) | `skillbook` (real), `environment`, `dedup_manager` |
| **Validated by engine?** | Yes β `requires`/`provides` | No β runtime error if missing |
---
## Protocols
Steps depend on protocols, not concrete classes. Each protocol defines the minimal interface a step needs. Concrete implementations satisfy them structurally β no inheritance required.
| Protocol | Method | Used by | Satisfied by |
|---|---|---|---|
| `AgentLike` | `generate(question, context, skillbook, reflection, **kwargs) β AgentOutput` | `AgentStep` | `Agent` |
| `ReflectorLike` | `reflect(question, agent_output, skillbook, ground_truth, feedback, **kwargs) β ReflectorOutput` | `ReflectStep` | `Reflector`, `RRStep` |
| `SkillManagerLike` | `update_skills(reflections, skillbook, question_context, progress, **kwargs) β SkillManagerOutput` | `UpdateStep` | `SkillManager` |
| `DeduplicationManagerLike` | `get_similarity_report(skillbook) β str \| None` | `DeduplicateStep` | `DeduplicationManager` |
Roles take a model string directly (e.g. `Agent("gpt-4o-mini")`). Internally each role creates a PydanticAI agent that handles structured output natively β no separate LLM client protocol is needed.
**Why protocols, not ABC:** Protocols use structural typing (duck typing checked by mypy). A class satisfies a protocol if it has the right methods β no `class Agent(AgentLike)` inheritance needed. Users can pass any object with a matching method, mocks satisfy protocols without ceremony, and steps are decoupled from implementations at the type level.
---
## Roles (Implementations)
Concrete LLM-based implementations of the role protocols. Live in `ace/implementations/` β fully self-contained.
| Class | Protocol | Method | What it does |
|---|---|---|---|
| `Agent` | `AgentLike` | `generate()` | Produces answers using the current skillbook of strategies |
| `Reflector` | `ReflectorLike` | `reflect()` | Single-pass analysis of agent outputs to extract lessons |
| `RRStep` | `ReflectorLike` + `StepProtocol` | `reflect()` / `__call__()` | Recursive multi-iteration reflection via PydanticAI agent with tools |
| `SkillManager` | `SkillManagerLike` | `update_skills()` | Transforms reflections into actionable skillbook updates |
All three share the same constructor pattern: `__init__(self, model: str, *, prompt_template=..., max_retries=3)`. The `model` parameter is resolved via `resolve_model()` to a PydanticAI agent.
`RRStep` is both a `StepProtocol[ACEStepContext]` (composable in any pipeline) and `ReflectorLike` (usable as a drop-in reflector). It is a subclass of `RecursiveAgent` with `execute_code` and `recurse` tools, plus two-tier compaction and depth-based recursion. See [RR_DESIGN.md](RR_DESIGN.md) for the full Recursive Reflector architecture.
---
## Steps
Reusable step implementations in `ace/steps/`. Each satisfies `StepProtocol[ACEStepContext]`. Each step does exactly one thing.
**Design principle: steps are stateless.** A step's `__call__` is a pure function of its constructor arguments and the incoming `ACEStepContext`. No internal counters, no accumulated state between invocations. Run-scoped information (like a global sample index for interval logic) comes from the context.
### Step summary
| Step | Requires | Provides | Side effects | `max_workers` |
|---|---|---|---|---|
| **AgentStep** | `sample`, `skillbook` | `agent_output` | None | 1 |
| **EvaluateStep** | `sample`, `agent_output` | `trace` | None | 1 |
| **ReflectStep** | `trace`, `skillbook` | `reflections` | None | 3; `async_boundary = True` |
| **UpdateStep** | `reflections`, `skillbook` | `skill_manager_output` | Agentic SkillManager mutates skillbook directly via ADD / UPDATE / REMOVE / TAG tools; output is an audit log | 1 |
| **DeduplicateStep** | `global_sample_index` | β | Consolidates similar skills | 1 |
| **CheckpointStep** | `global_sample_index` | β | Saves skillbook to disk | 1 |
| **LoadTracesStep** | `sample` | `trace` | None | 1 |
| **PersistStep** | `skillbook` | β | Writes skillbook to external file | 1 |
| **ExportSkillbookMarkdownStep** | `skillbook` | β | Exports skillbook as markdown | 1 |
**Requires vs Injected:** `Requires` lists context fields (validated by the pipeline engine at construction time). The `skillbook` on the context is a `SkillbookView` (read-only). Steps that **write** to the skillbook receive the real `Skillbook` via constructor injection.
**`trace` as the universal learning input:** The learning tail's entry point (ReflectStep) requires only `trace` and `skillbook`. In the standard ACE pipeline, `EvaluateStep` bundles structured fields into a `trace` dict. In TraceAnalyser, `_build_context` places the raw trace directly. In integrations, the execute step provides `trace` from its framework's native output. The learning tail is agnostic to trace format.
Steps with empty `provides` are pure side-effect steps β they mutate shared state (skillbook) or write to external systems (disk) but add no new fields to the context.
---
## Runners
### Class hierarchy
```
ACERunner (shared infrastructure: epoch loop, delegates to Pipeline.run())
βββ TraceAnalyser β [Reflect β Update β Apply]
βββ ACE β [Agent β Evaluate β Reflect β Update β Apply]
βββ BrowserUse β [BrowserExecute β BrowserToTrace β learning_tail]
βββ LangChain β [LangChainExecute β LangChainToTrace β learning_tail]
βββ ClaudeCode β [ClaudeCodeExecute β ClaudeCodeToTrace β learning_tail]
βββ OpenClaw (script) β [LoadTraces β OpenClawToTrace β learning_tail]
ACELiteLLM (standalone convenience wrapper β not an ACERunner subclass)
βββ ask() β direct Agent call, no pipeline
βββ learn() β delegates to lazy-init ACE runner
βββ learn_from_traces() β delegates to lazy-init TraceAnalyser
βββ learn_from_feedback()β runs learning_tail from last ask()
RRStep (RecursiveAgent subclass β composable iterative step)
βββ __call__() β StepProtocol entry; usable in any runner's pipeline
βββ reflect() β ReflectorLike entry; drop-in reflector for runners
βββ _run_reflection() β PydanticAI agent with execute_code and recurse tools
```
All runners compose a `Pipeline` rather than extending it.
### ACERunner β shared base
Encapsulates everything runners have in common: the epoch loop and Iterable validation. Per-sample iteration, error handling, background execution, and checkpoints are all delegated to `Pipeline.run()`.
Subclasses only override `run()` (public signature) and `_build_context()` (input mapping).
**Responsibilities:**
| Concern | Owner |
|---|---|
| Epoch loop + Iterable validation | `ACERunner._run()` |
| Per-sample iteration + error isolation | `Pipeline.run()` |
| Foreground/background split | `Pipeline.run()` (via `async_boundary`) |
| Concurrent workers | `Pipeline.run(workers=N)` |
| Checkpoints | `CheckpointStep` (in the pipeline) |
| Background drain | `ACERunner.wait_for_background()` β `Pipeline.wait_for_background()` |
| Skillbook I/O | `save(path)` on the runner |
Each sample is independent β no state persists across samples. The skillbook is the only cross-sample coupling.
**Eventual consistency:** `SkillbookView` is a thin delegation wrapper, not a snapshot β it reads from the live `Skillbook` at call time. When background learning is active, concurrent samples may observe partially-updated skillbook state. This is by design: steps see a best-effort view rather than a point-in-time snapshot. The trade-off is acceptable because (1) the skillbook is LLM prompt context where a few missing or extra skills have negligible impact, (2) serialising reads would eliminate the concurrency benefit, and (3) write steps already run with `max_workers = 1`.
### TraceAnalyser
Analyses pre-recorded traces without executing an agent. Runs the learning tail only. Accepts raw trace objects of any type.
**When to use:** You have execution logs from an external system and want to build or refine a skillbook from historical data. Multi-epoch mode re-processes all traces with the evolving skillbook.
**Pipeline:**
```
[ReflectStep] β [UpdateStep] (SkillManager mutates the skillbook directly)
```
No AgentStep, no EvaluateStep. The trace already contains the agent's output and the evaluation feedback.
**Multi-epoch semantics:** Each epoch re-processes all traces with the current skillbook. Early epochs extract obvious patterns; later epochs refine and consolidate.
### ACE
The full live adaptive pipeline. An agent executes, the reflector analyses, the skill manager updates. Optionally evaluates against a `TaskEnvironment` for feedback-driven learning.
**When to use:** Building a new agent, or running closed-loop learning where the agent improves in real time.
**Pipeline:**
```
[AgentStep] β [EvaluateStep] β [ReflectStep] β [UpdateStep] (SkillManager mutates the skillbook directly)
```
A single class handles both single-pass (`epochs=1`) and multi-epoch batch training (`epochs > 1`). The `environment` is optional β when provided, `EvaluateStep` generates feedback. When omitted, the Reflector learns from ground-truth comparison or the agent's reasoning alone.
### ACELiteLLM β standalone convenience wrapper
`ACELiteLLM` is not an `ACERunner` subclass. It wraps two different runners (`ACE` and `TraceAnalyser`) and exposes a fundamentally different API:
| Method | What it does |
|---|---|
| `ask(question, context)` | Direct Agent call β no pipeline. Stores interaction for `learn_from_feedback()` |
| `learn(samples, environment, epochs)` | Delegates to lazy-init ACE runner |
| `learn_from_traces(traces, epochs)` | Delegates to lazy-init TraceAnalyser |
| `learn_from_feedback(feedback, ground_truth)` | Manual single-shot learning from last `ask()` call |
Runners are cached and invalidated on `load()` (new skillbook object means stale references).
### Factory methods
All runners provide a `from_roles` factory that takes pre-built role instances. Integration runners also provide `from_model()` that auto-builds PydanticAI-backed roles from a model string.
**Common parameters on `from_roles`:**
| Parameter | Default | Description |
|---|---|---|
| `skillbook` | `Skillbook()` | Starting skillbook |
| `dedup_manager` | `None` | Appends a `DeduplicateStep` |
| `dedup_interval` | `10` | Deduplication frequency |
| `checkpoint_dir` | `None` | Appends a `CheckpointStep` |
| `checkpoint_interval` | `10` | Checkpoint frequency |
| `extra_steps` | `None` | Additional steps appended after the learning tail |
### `learning_tail()` β reusable learning steps
Every integration assembles the same `[Reflect β Update β Apply]` suffix. `learning_tail()` returns this standard step list, with optional dedup and checkpoint steps. If the provided reflector already exposes `provides = {'reflections'}` (e.g. `RRStep`), it's inserted directly instead of being wrapped in `ReflectStep`.
---
## Integration Pattern
External frameworks integrate via composable pipeline steps in `ace/integrations/`. Each integration provides:
1. **Result type** β an integration-specific dataclass (e.g. `BrowserResult`, `ClaudeCodeResult`)
2. **Execute step** β INJECT skillbook context + EXECUTE the framework, writes to `ctx.trace`
3. **ToTrace step** β converts the integration-specific result into the standardised trace dict
### Execute β Convert β Learn
```
Standard ACE: [Agent β Evaluate] β [Reflect β Update β Apply]
β°ββ execute (built-in) βββ― β°ββββββββ learn (shared) βββββββ―
provides: trace (dict) ββββββββββββββββββΊ requires: trace
Browser-use: [BrowserExecute] β [BrowserToTrace] β [Reflect β Update β Apply]
β°ββ execute βββββ― β°ββ convert βββ― β°ββββββββ learn (shared) βββββββ―
provides: trace rewrites trace requires: trace
(BrowserResult) (BrowserResult β dict)
TraceAnalyser: [_build_context] β [Reflect β Update β Apply]
β°ββ sets ctx.trace (raw object) ββββββββ― β°ββββββββ learn (shared) βββββββ―
```
The standardised trace dict keys match what `ReflectStep` expects: `question`, `reasoning`, `answer`, `skill_ids`, `feedback`, `ground_truth`.
### Result types
| Integration | Result type | Key fields |
|---|---|---|
| Browser-use | `BrowserResult` | `task`, `success`, `output`, `error`, `steps_count`, `duration_seconds`, `cited_skill_ids`, `chronological_steps`, `raw_history` |
| Claude Code | `ClaudeCodeResult` | `task`, `success`, `output`, `execution_trace`, `returncode`, `error` |
| Claude SDK | `ClaudeSDKResult` | `task`, `success`, `output`, `error`, `model`, `stop_reason`, `input_tokens`, `output_tokens`, `tool_calls`, `cited_skill_ids` |
| LangChain | `LangChainResult` | `task`, `output`, `result_type`, `success`, `error`, `intermediate_steps`, `messages`, `raw_result` |
### Why two steps instead of one
Splitting execute from trace conversion gives independent testability, reusability (execute step usable standalone), and separation of concerns (framework interaction vs trace formatting).
### Live vs offline
| | Integration Runner | TraceAnalyser |
|---|---|---|
| When | Live execution | Post-hoc analysis |
| Agent | Framework runs it | Already ran |
| Feedback | Generated live | Baked into trace |
| Use case | Production deployment | Historical batch learning, debugging |
Both update the same skillbook. A common workflow: TraceAnalyser builds an initial skillbook from historical data, then an integration runner refines it during live deployment.
> **MCP Server** is a different pattern. It does not add pipeline steps β it's a thin async layer over `ACELiteLLM` that exposes ACE as an MCP tool provider. See [MCP Server docs](../integrations/mcp.md).
---
## Configuration & Providers
### Principles
1. **API keys never appear in ACE APIs** β no `api_key` parameter anywhere. Keys are resolved from the environment by LiteLLM at call time.
2. **Per-role model selection** β Agent, Reflector, and SkillManager can each use different models.
3. **Validate before running** β `ace setup` and `validate_connection()` make a tiny LLM call to verify auth before writing code.
### Config types
- **`ModelConfig`** β which model to use for a role (model string, temperature, max_tokens). No secrets.
- **`ACEModelConfig`** β model selection per ACE role. Serialises to/from `ace.toml` (committable, no secrets).
### Construction paths
| Constructor | Input | Use case |
|---|---|---|
| `ACELiteLLM.from_setup()` | `ace.toml` + `.env` | Teams, CI, guided setup |
| `ACELiteLLM.from_config(config)` | `ACEModelConfig` object | Per-role model selection in code |
| `ACELiteLLM.from_model("gpt-4o")` | Model string | Quick start, single model |
| `ACELiteLLM("gpt-4o-mini", ...)` | Model string + overrides | Full control |
### CLI
| Command | What it does |
|---|---|
| `ace setup` | Interactive wizard: model name, API key, validate, assign per-role models. Saves `.env` + `ace.toml`. |
| `ace models [query]` | Search LiteLLM's model registry (2,600+ models). Filter by `--provider`. |
| `ace providers` | List providers with env var names and key status. |
| `ace validate <model>` | Test a model connection with a tiny LLM call. |
### File layout
| File | Secrets? | Committable? | Purpose |
|---|---|---|---|
| `.env` | Yes | No (gitignored) | API keys only |
| `ace.toml` | No | Yes | Model names + parameters per role |
### Key resolution flow
```
API Key: .env β os.environ β LiteLLM reads OPENAI_API_KEY / ANTHROPIC_API_KEY / etc.
Model: ace.toml β ACEModelConfig.for_role("agent") β resolve_model(model) β PydanticAI agent
```
### Provider resolution
ACE model strings follow LiteLLM convention (`provider/model`). The resolver in `ace/providers/pydantic_ai.py` routes them through three paths:
1. **PydanticAI-native prefix** β strings like `openai:gpt-4o` pass through unchanged
2. **LiteLLM prefix β native provider** β when the first path segment matches a PydanticAI native provider, `/` is rewritten to `:` (e.g. `bedrock/model` β `bedrock:model`)
3. **Fallback** β everything else is prefixed with `litellm:` for the proxy provider
```
LiteLLM string β PydanticAI string
βββββββββββββββββββββββββββββββββββββββββββββββββ ββββββββββββββββββββββββββββββββββββββββββ
gpt-4o-mini β litellm:gpt-4o-mini
bedrock/eu.anthropic.claude-haiku-4-5-v1:0 β bedrock:eu.anthropic.claude-haiku-4-5-v1:0
groq/llama-3.1-70b-versatile β groq:llama-3.1-70b-versatile
openrouter/anthropic/claude-3.5-sonnet β openrouter:anthropic/claude-3.5-sonnet
anthropic/claude-3-5-sonnet-20241022 β anthropic:claude-3-5-sonnet-20241022
ollama/llama3 β litellm:ollama/llama3
together_ai/meta-llama/Llama-3-70b β litellm:together_ai/meta-llama/Llama-3-70b
```
Mapped LiteLLM prefixes: `anthropic`, `azure`, `azure_ai`, `bedrock`, `cohere`, `deepseek`, `groq`, `mistral`, `openrouter`, `vertex_ai`. All others fall through to `litellm:`.
Native providers are faster (no proxy hop) and use the provider's own API key env vars directly. Install with extras: `uv add "pydantic-ai-slim[anthropic,openai,bedrock]"`.
---
## Deduplication
Skill deduplication subsystem in `ace/deduplication/` β fully self-contained.
| Class | Role |
|---|---|
| `SimilarityDetector` | Computes embeddings, detects similar pairs via cosine similarity |
| `DeduplicationManager` | Coordinates detection and consolidation |
**Embedding providers:** LiteLLM (remote) or sentence-transformers (local, lazy-loaded).
**Consolidation operations:**
| Operation | Effect |
|---|---|
| `MergeOp` | Combine skills β accumulate counters, soft-delete others |
| `DeleteOp` | Soft-delete a redundant skill |
| `KeepOp` | Store a similarity decision so the pair is not flagged again |
| `UpdateOp` | Refine content to differentiate, clear embedding |
**Pipeline integration:** Deduplication runs as a separate `DeduplicateStep` at a configurable interval, not inside the SkillManager. Appended by factory methods when a `DeduplicationManagerLike` is provided.
---
## Observability
PydanticAI has first-class Logfire integration. One call auto-instruments everything:
```python
logfire.configure()
logfire.instrument_pydantic_ai()
```
This automatically captures agent runs, tool calls, model requests, structured output validation, and sub-agent delegation β the entire RR execution appears as a structured trace. No custom span-building code.
**Pipeline integration:** Logfire is OpenTelemetry-based, so pipeline-level spans coexist. Steps that don't use PydanticAI can use `logfire.span()` / `logfire.info()` directly.
**Setup:** `ace/observability/configure_logfire()` auto-instruments all PydanticAI agents. Opt-in via `ACELiteLLM(logfire=True)`. Config is purely env-based (`LOGFIRE_TOKEN`), no changes to `ace.toml`.
---
## Concurrency
Both TraceAnalyser and ACE inherit async capabilities from the pipeline engine. No custom async machinery is needed.
### ReflectStep as async boundary
`ReflectStep.async_boundary = True` means everything before it (Agent, Evaluate) runs in the foreground, and everything from ReflectStep onwards runs in a background thread pool:
```
sample 1: [AgentStep] [EvaluateStep] ββfireβββΊ [ReflectStep] [UpdateStep]
sample 2: [AgentStep] [EvaluateStep] ββfireβββΊ ...
β
async_boundary
```
### Concurrency knobs
| Knob | Where | Effect |
|---|---|---|
| `ReflectStep.max_workers = 3` | Step class attribute | Up to 3 reflections in parallel |
| `UpdateStep.max_workers = 1` | Step class attribute | Serialises skill manager LLM calls AND skillbook writes (SM tools mutate in place) |
| `wait_for_background(timeout)` | Runner method | Blocks until background threads drain |
### Cancellation
ACE inherits cancellation from the pipeline engine. Pass a `CancellationToken` to `run()`. The pipeline checks it before each foreground step. Within a step, PydanticAI's async runtime handles cancellation natively. The token flows via `contextvars.ContextVar` β no parameter changes needed across layers. See [PIPELINE_DESIGN.md Β§ Cancellation](PIPELINE_DESIGN.md#cancellation).
---
## Error Handling
Follows the pipeline engine's error model without additions.
- **Per-sample isolation:** A failing sample does not abort the run. The exception is recorded in `SampleResult.error` and `SampleResult.failed_at`.
- **Background failures:** Captured and attached to `SampleResult` by the pipeline engine.
- **No retry logic in the runner.** Retries are the responsibility of individual steps (e.g., PydanticAI's built-in retry with error feedback).
---
## Directory Structure
```
ace/
__init__.py β Public API re-exports
core/
context.py β ACEStepContext, SkillbookView, ACESample
insight_source.py β TraceIdentity, TraceReference, InsightSource
outputs.py β AgentOutput, ReflectorOutput, SkillManagerOutput
skillbook.py β Skill, Skillbook, SimilarityDecision
environments.py β Sample, TaskEnvironment, SimpleEnvironment
protocols/ β Role protocols (one file per protocol)
agent.py, reflector.py, skill_manager.py, deduplication.py
implementations/ β PydanticAI-backed role implementations
agent.py, reflector.py, skill_manager.py, helpers.py, prompts.py
steps/ β Pipeline steps (one file per class)
__init__.py β learning_tail() helper
agent.py, evaluate.py, reflect.py, update.py,
apply.py, deduplicate.py, checkpoint.py,
load_traces.py, persist.py, export_markdown.py, observability.py
runners/ β Runner classes
base.py β ACERunner
trace_analyser.py, ace.py, browser_use.py, langchain.py,
claude_code.py, litellm.py
integrations/ β Integration steps (execute + result + converter)
browser_use.py, langchain.py, claude_code.py, claude_sdk.py
openclaw/ β OpenClaw trace converter
mcp/ β Optional MCP server
providers/ β PydanticAI model resolution
pydantic_ai.py, config.py, registry.py
deduplication/ β Skill deduplication subsystem
detector.py, manager.py, operations.py, prompts.py
rr/ β Recursive Reflector (PydanticAI agent)
observability/ β Logfire configuration
```
### Key modules
| Module | Contents |
|---|---|
| `ace/core/` | `ACEStepContext`, `SkillbookView`, `Skillbook`, `AgentOutput`, `ReflectorOutput`, `InsightSource` |
| `ace/protocols/` | `AgentLike`, `ReflectorLike`, `SkillManagerLike` protocols |
| `ace/implementations/` | PydanticAI-backed `Agent`, `Reflector`, `SkillManager` |
| `ace/steps/` | All pipeline steps + `learning_tail()` |
| `ace/runners/` | `ACERunner`, `TraceAnalyser`, `ACE`, `BrowserUse`, `LangChain`, `ClaudeCode`, `ACELiteLLM` |
| `ace/providers/` | `resolve_model`, `ACEModelConfig`, `validate_connection` |
| `ace/steps/rr_step.py` | `RRStep` (RecursiveAgent subclass), `RRConfig`, `TraceSandbox` |
| `ace/core/recursive_agent.py` | `RecursiveAgent`, `AgenticConfig`, `AgenticDeps`, compaction, recursion, `usage_callback` hook |
| `ace/core/metered_model.py` | `MeteredModel` β pydantic-ai `WrapperModel` that fires the `usage_callback` once per request |
| `ace/integrations/` | Execute steps, result types, ToTrace converters; MCP server |
| `ace/deduplication/` | Dedup subsystem (detector, manager, operations) |
| `ace/observability/` | Logfire configuration (`configure_logfire()`) |
---
## Future Directions
Issues acknowledged but deferred.
**Streaming / lazy iteration:** `_run()` eagerly materializes the full iterable before passing to `Pipeline.run()`. True streaming would require the pipeline to accept an iterator. Deliberate simplification β revisit if memory pressure from large single-pass runs becomes real.
**Builder API for custom pipelines:** The current API offers two extremes: factory methods that hide the pipeline, and manual construction that requires understanding step contracts. A builder could bridge this gap, but `learning_tail()` covers the most common customisation (custom execute step + standard learning). Worth pursuing when users hit friction with manual wiring.
**Skillbook rollback and versioning:** Currently the skillbook is mutated in place with no undo. A lightweight versioning mechanism (snapshotting at epoch boundaries, `rollback(to_version)`) would enable automatic revert when metrics degrade. Deferred because checkpoints cover the common recovery scenario.
**LiteLLM proxy base URL support:** Users running a LiteLLM proxy may need `api_base` configuration. Deferred because all current users connect directly to providers.
|