|
Download README.md from observability/agent-observability: direct link, hf CLI and curl.
- Browser
- Download file 7.98 kB
-
https://huggingface.co/spaces/observability/agent-observability/resolve/main/README.md
- Command line
-
hf download hf://spaces/observability/agent-observability/README.md
-
curl -L -o README.md https://huggingface.co/spaces/observability/agent-observability/resolve/main/README.md
7.98 kB
| title: Agent Observability | |
| emoji: π€ | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: static | |
| pinned: false | |
| # Agent Observability | |
| ### Trace, debug and evaluate advanced AI agents | |
| **Agent Observability** is an interactive Hugging Face Space focused on the telemetry required to understand how AI agents behave across planning, tools, memory, verification and long-horizon execution. | |
| A modern agent may perform dozens or hundreds of actions before reaching a result. | |
| Observability makes that behavior inspectable. | |
| > **If an agent can act, you need to know what it did, why it did it and what happened next.** | |
| --- | |
| # What Is Agent Observability? | |
| **Agent observability** is the practice of capturing and correlating the execution signals produced by an AI agent. | |
| These signals may include: | |
| - goals | |
| - plans | |
| - agent steps | |
| - model calls | |
| - tool calls | |
| - tool results | |
| - memory reads | |
| - memory writes | |
| - retries | |
| - replanning | |
| - verification | |
| - permission checks | |
| - human approvals | |
| - cost | |
| - latency | |
| - failures | |
| A useful agent trace might look like: | |
| ```text | |
| Goal | |
| β | |
| Plan | |
| β | |
| Agent Step | |
| β | |
| Tool Call | |
| β | |
| Observation | |
| β | |
| Memory Update | |
| β | |
| Verification | |
| β | |
| Replan | |
| β | |
| Next Step | |
| ``` | |
| --- | |
| # Why Agent Observability Matters | |
| Agents introduce runtime complexity beyond simple model calls. | |
| Without observability, it may be impossible to answer: | |
| - Why did the agent choose this tool? | |
| - Which memory influenced the decision? | |
| - Where did the plan change? | |
| - Why did the agent retry? | |
| - Which step caused the failure? | |
| - Did the agent recover? | |
| - Did the agent violate a permission boundary? | |
| - How much did each step cost? | |
| - Which action required human approval? | |
| - Did the final result pass verification? | |
| --- | |
| # Core Agent Signals | |
| ## Goal | |
| Capture the original objective. | |
| Useful fields: | |
| - goal | |
| - constraints | |
| - success criteria | |
| - deadline | |
| - budget | |
| ## Plan | |
| Capture: | |
| - task decomposition | |
| - plan version | |
| - dependencies | |
| - priority | |
| - replans | |
| ## Agent Step | |
| Capture: | |
| - step number | |
| - step type | |
| - action | |
| - status | |
| - duration | |
| ## Tool Call | |
| Capture: | |
| - tool | |
| - arguments | |
| - result | |
| - error | |
| - retry | |
| - permissions | |
| ## Memory | |
| Capture: | |
| - read | |
| - write | |
| - update | |
| - provenance | |
| - freshness | |
| - conflict | |
| ## Verification | |
| Capture: | |
| - verifier | |
| - evidence | |
| - result | |
| - failure reason | |
| - next action | |
| --- | |
| # Long-Horizon Observability | |
| Long-running agents require persistent telemetry. | |
| Relevant signals include: | |
| - total steps | |
| - total duration | |
| - total cost | |
| - retries | |
| - replans | |
| - checkpoint events | |
| - repeated actions | |
| - memory usage | |
| - tool failures | |
| - verification failures | |
| - human interventions | |
| --- | |
| # Plan Drift | |
| Plan drift occurs when execution gradually diverges from the intended plan. | |
| Observability can help detect: | |
| - unexpected branches | |
| - repeated replanning | |
| - skipped dependencies | |
| - growing step count | |
| - inactive tasks | |
| --- | |
| # Goal Drift | |
| Goal drift occurs when the agent begins optimizing for a different objective. | |
| Useful checks include: | |
| ```text | |
| Original Goal | |
| β | |
| Current Plan | |
| β | |
| Current Action | |
| β | |
| Compare | |
| β | |
| Aligned / Drifted | |
| ``` | |
| --- | |
| # Tool Observability | |
| Agent tool calls should include: | |
| - tool name | |
| - tool version | |
| - input arguments | |
| - execution time | |
| - output | |
| - error | |
| - retry count | |
| - permission decision | |
| --- | |
| # Memory Observability | |
| Persistent agents require visibility into: | |
| - memory read | |
| - memory write | |
| - memory update | |
| - memory deletion | |
| - retrieved memory | |
| - score | |
| - source | |
| - timestamp | |
| - confidence | |
| --- | |
| # Recovery Observability | |
| Recovery events may include: | |
| - retry | |
| - backoff | |
| - alternate tool | |
| - alternate model | |
| - checkpoint restore | |
| - replan | |
| - human escalation | |
| - abort | |
| --- | |
| # Human Oversight | |
| Important events may require human review. | |
| Capture: | |
| - approval requested | |
| - approver | |
| - decision | |
| - timestamp | |
| - reason | |
| - result | |
| --- | |
| # Agent Trace Example | |
| ```text | |
| trace_id: agent-001 | |
| 1. Goal received | |
| 2. Plan created | |
| 3. Browser tool selected | |
| 4. Browser tool called | |
| 5. Tool returned error | |
| 6. Retry triggered | |
| 7. Alternate tool selected | |
| 8. Tool succeeded | |
| 9. Memory updated | |
| 10. Verifier passed | |
| 11. Final response generated | |
| ``` | |
| --- | |
| # Agent Metrics | |
| Useful metrics include: | |
| - task success rate | |
| - average steps per task | |
| - tool success rate | |
| - retry rate | |
| - recovery rate | |
| - replan rate | |
| - verification pass rate | |
| - human intervention rate | |
| - memory retrieval precision | |
| - average cost per task | |
| - average latency | |
| - repeated-action rate | |
| --- | |
| # Agent Failure Modes | |
| ## Hidden Tool Failure | |
| A tool fails but the error is not propagated. | |
| ## Silent Memory Drift | |
| Stale state influences later decisions. | |
| ## Retry Loop | |
| The agent repeats the same failed action. | |
| ## Excessive Replanning | |
| The plan changes too frequently. | |
| ## Goal Drift | |
| Execution no longer matches the original objective. | |
| ## Missing Verification | |
| A wrong intermediate result is accepted. | |
| ## Permission Blind Spot | |
| A high-impact action occurs without a visible authorization decision. | |
| --- | |
| # Agent Observability Architecture | |
| ```text | |
| AGENT | |
| β | |
| Instrumentation | |
| β | |
| βββββββββββββββββββββββββββ | |
| β Steps β Tools β Memory β | |
| β Plans β Evals β Costs β | |
| βββββββββββββββββββββββββββ | |
| β | |
| Trace | |
| β | |
| Correlation | |
| β | |
| βββββββββββββββββββββββββββ | |
| β Search β Dashboards β | |
| β Alerts β Evaluation β | |
| βββββββββββββββββββββββββββ | |
| β | |
| Insight | |
| ``` | |
| --- | |
| # Interactive Explorer | |
| The included `index.html` lets users inspect: | |
| - Goal tracing | |
| - Plan tracing | |
| - Step tracing | |
| - Tool tracing | |
| - Memory tracing | |
| - Verification tracing | |
| - Recovery tracing | |
| - Human approval tracing | |
| - Cost tracing | |
| - Long-horizon drift detection | |
| Each topic includes: | |
| - recommended fields | |
| - why it matters | |
| - failure modes | |
| - operational metrics | |
| --- | |
| # SEO & GEO Topic Map | |
| This Space is structured around: | |
| - Agent Observability | |
| - AI Agent Observability | |
| - agent tracing | |
| - AI agent tracing | |
| - agent telemetry | |
| - long-horizon agent observability | |
| - tool tracing | |
| - memory tracing | |
| - planning observability | |
| - AI agent debugging | |
| - AI agent monitoring | |
| - agent runtime observability | |
| - agent verification | |
| - agent recovery | |
| - goal drift | |
| - plan drift | |
| - agent cost monitoring | |
| --- | |
| # GEO Entity Relationships | |
| ```text | |
| Agent Observability | |
| OBSERVES β Goals | |
| OBSERVES β Plans | |
| OBSERVES β Agent Steps | |
| OBSERVES β Tools | |
| OBSERVES β Memory | |
| OBSERVES β Verification | |
| OBSERVES β Recovery | |
| TRACKS β Cost | |
| TRACKS β Latency | |
| DETECTS β Goal Drift | |
| DETECTS β Plan Drift | |
| SUPPORTS β Debugging | |
| SUPPORTS β Evaluation | |
| SUPPORTS β Reliability | |
| ``` | |
| --- | |
| # Collaboration & Partnerships | |
| **Agent Observability** is open to collaboration with companies, research teams, universities and open-source projects working on agent infrastructure and observability. | |
| Relevant areas include: | |
| - AI agents | |
| - agent runtimes | |
| - tracing | |
| - telemetry | |
| - tool observability | |
| - memory observability | |
| - long-horizon agents | |
| - verification | |
| - evaluation | |
| - orchestration | |
| - human oversight | |
| Possible collaboration formats include: | |
| - joint Hugging Face Spaces | |
| - trace visualizations | |
| - framework integrations | |
| - benchmark projects | |
| - technical demos | |
| - agent runtime integrations | |
| - clearly disclosed partnerships and sponsorships | |
| ## Collaboration Contact | |
| **agenten@magenta.de** | |
| --- | |
| # Independence | |
| **Agent Observability** is an independent Hugging Face Space. | |
| It is not an official project of Hugging Face, any AI laboratory, observability vendor, agent framework or technology company. | |
| --- | |
| # Long-Term Vision | |
| The goal is to make agent execution understandable, inspectable and debuggable from the first goal to the final action. | |
| > **Observe every step. Correlate every action. Understand the whole agent.** | |