|
Download README.md from observability/observability-explorer: direct link, hf CLI and curl.
- Browser
- Download file 7.21 kB
-
https://huggingface.co/spaces/observability/observability-explorer/resolve/main/README.md
- Command line
-
hf download hf://spaces/observability/observability-explorer/README.md
-
curl -L -o README.md https://huggingface.co/spaces/observability/observability-explorer/resolve/main/README.md
7.21 kB
| title: Observability Explorer | |
| emoji: ποΈ | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: static | |
| pinned: false | |
| # Observability Explorer | |
| ### Explore the signals, layers and workflows behind AI Observability | |
| **Observability Explorer** is an interactive Hugging Face Space for understanding how modern AI systems can be traced, measured, debugged and improved. | |
| The Space covers observability across: | |
| - LLMs | |
| - AI agents | |
| - tool use | |
| - retrieval | |
| - memory | |
| - inference | |
| - orchestration | |
| - multi-agent systems | |
| - evaluation | |
| - verification | |
| - cost | |
| - production AI infrastructure | |
| > **Monitoring tells you that something changed. Observability helps you understand why.** | |
| --- | |
| # What Is AI Observability? | |
| **AI observability** is the practice of collecting and interpreting runtime signals so developers and operators can understand: | |
| - what happened | |
| - in what order | |
| - which component acted | |
| - what data was used | |
| - what model was selected | |
| - which tool was called | |
| - whether the result was verified | |
| - how long the execution took | |
| - how much it cost | |
| - why the system failed | |
| A modern AI execution may involve many connected components: | |
| ```text | |
| User Request | |
| β | |
| Router | |
| β | |
| Model | |
| β | |
| Retriever | |
| β | |
| Tool | |
| β | |
| Agent Step | |
| β | |
| Memory | |
| β | |
| Verifier | |
| β | |
| Final Response | |
| ``` | |
| Observability links these stages into one interpretable execution story. | |
| --- | |
| # Core Observability Signals | |
| ## Traces | |
| Traces connect the full path of a request. | |
| Useful for: | |
| - end-to-end debugging | |
| - latency analysis | |
| - root-cause analysis | |
| - workflow inspection | |
| ## Spans | |
| Spans represent individual operations inside a trace. | |
| Examples: | |
| - model call | |
| - retrieval | |
| - tool execution | |
| - memory read | |
| - memory write | |
| - verification | |
| - routing decision | |
| ## Logs | |
| Logs capture discrete events such as: | |
| - tool failed | |
| - retry triggered | |
| - fallback model used | |
| - permission denied | |
| - checkpoint restored | |
| ## Metrics | |
| Metrics summarize numerical signals such as: | |
| - latency | |
| - error rate | |
| - token usage | |
| - cost | |
| - tool success rate | |
| - verifier pass rate | |
| ## Events | |
| Events capture meaningful state changes: | |
| - agent started | |
| - plan updated | |
| - memory changed | |
| - approval requested | |
| - task completed | |
| --- | |
| # AI Observability Layers | |
| ```text | |
| Application | |
| β | |
| AI Workflow | |
| β | |
| Models / Agents / Tools | |
| β | |
| Instrumentation | |
| β | |
| Traces / Logs / Metrics / Events | |
| β | |
| Evaluation / Verification | |
| β | |
| Dashboards / Search / Alerts | |
| β | |
| Debugging / Improvement | |
| ``` | |
| --- | |
| # LLM Observability | |
| LLM observability tracks: | |
| - model | |
| - provider | |
| - prompt | |
| - prompt version | |
| - input tokens | |
| - output tokens | |
| - latency | |
| - cost | |
| - structured output | |
| - errors | |
| - fallback behavior | |
| --- | |
| # Agent Observability | |
| Agent observability extends beyond model calls. | |
| A useful agent trace may include: | |
| ```text | |
| Goal | |
| β | |
| Plan | |
| β | |
| Agent Step | |
| β | |
| Tool Call | |
| β | |
| Observation | |
| β | |
| Memory Update | |
| β | |
| Verification | |
| β | |
| Replan | |
| ``` | |
| Relevant signals include: | |
| - step count | |
| - tool calls | |
| - retries | |
| - memory access | |
| - recovery | |
| - replanning | |
| - human intervention | |
| - task success | |
| --- | |
| # Retrieval Observability | |
| Retrieval systems should expose: | |
| - query | |
| - retrieved documents | |
| - reranker | |
| - selected chunks | |
| - relevance scores | |
| - retrieval latency | |
| This helps distinguish: | |
| > Did the model fail, or did retrieval provide weak context? | |
| --- | |
| # Tool Observability | |
| Tool telemetry can include: | |
| - tool name | |
| - input arguments | |
| - execution time | |
| - result | |
| - error | |
| - retry | |
| - permissions | |
| - risk class | |
| --- | |
| # Memory Observability | |
| Persistent agents need visibility into: | |
| - memory reads | |
| - memory writes | |
| - updates | |
| - deletions | |
| - provenance | |
| - freshness | |
| - conflicts | |
| --- | |
| # Routing Observability | |
| Model routing should capture: | |
| - available models | |
| - selected model | |
| - selection reason | |
| - cost estimate | |
| - latency estimate | |
| - fallback behavior | |
| --- | |
| # Cost Observability | |
| Useful cost dimensions include: | |
| - model cost | |
| - tool cost | |
| - retrieval cost | |
| - embedding cost | |
| - retry cost | |
| - total task cost | |
| --- | |
| # Evaluation + Observability | |
| Evaluation answers: | |
| > How good is the system? | |
| Observability answers: | |
| > What happened during execution? | |
| Together: | |
| ```text | |
| Execution | |
| β | |
| Observability | |
| β | |
| Evaluation | |
| β | |
| Diagnosis | |
| β | |
| Improvement | |
| ``` | |
| --- | |
| # Long-Horizon Observability | |
| Long-running agents may require: | |
| - persistent trace IDs | |
| - checkpoints | |
| - step history | |
| - memory telemetry | |
| - budget tracking | |
| - repeated-action detection | |
| - plan drift detection | |
| - goal drift detection | |
| --- | |
| # Failure Modes | |
| ## Missing Trace Coverage | |
| Important components are not instrumented. | |
| ## Broken Correlation | |
| Events cannot be connected into one execution path. | |
| ## Excessive Logging | |
| Too much telemetry creates noise and cost. | |
| ## Sensitive Data Leakage | |
| Prompts or outputs expose private data. | |
| ## Missing Tool Visibility | |
| External actions cannot be reconstructed. | |
| ## Missing Memory Visibility | |
| Persistent state changes remain hidden. | |
| --- | |
| # Interactive Explorer | |
| The included `index.html` lets users explore: | |
| - traces | |
| - spans | |
| - logs | |
| - metrics | |
| - events | |
| - LLM observability | |
| - agent observability | |
| - retrieval observability | |
| - tool observability | |
| - memory observability | |
| - cost observability | |
| - evaluation-aware observability | |
| Each topic includes: | |
| - definition | |
| - what to capture | |
| - why it matters | |
| - typical failure modes | |
| - relevance to production AI | |
| --- | |
| # SEO & GEO Topic Map | |
| This Space is structured around: | |
| - AI Observability | |
| - Observability Explorer | |
| - LLM Observability | |
| - Agent Observability | |
| - AI tracing | |
| - LLM tracing | |
| - agent tracing | |
| - AI telemetry | |
| - AI monitoring | |
| - model monitoring | |
| - prompt observability | |
| - tool observability | |
| - memory observability | |
| - retrieval observability | |
| - inference observability | |
| - AI cost monitoring | |
| - AI debugging | |
| - AI evaluation | |
| - AI validation | |
| - AI verification | |
| - production AI systems | |
| --- | |
| # GEO Entity Relationships | |
| ```text | |
| AI Observability | |
| USES β Traces | |
| USES β Spans | |
| USES β Logs | |
| USES β Metrics | |
| USES β Events | |
| OBSERVES β Models | |
| OBSERVES β Agents | |
| OBSERVES β Tools | |
| OBSERVES β Memory | |
| OBSERVES β Retrieval | |
| SUPPORTS β Evaluation | |
| SUPPORTS β Verification | |
| SUPPORTS β Reliability | |
| ENABLES β Debugging | |
| ``` | |
| --- | |
| # Collaboration & Partnerships | |
| **Observability Explorer** is open to collaboration with companies, research teams, universities and open-source projects working on AI observability and production AI infrastructure. | |
| Relevant areas include: | |
| - tracing | |
| - telemetry | |
| - LLM observability | |
| - agent observability | |
| - prompt tracking | |
| - retrieval tracing | |
| - memory tracing | |
| - model routing | |
| - evaluation | |
| - verification | |
| - cost monitoring | |
| - production AI | |
| Possible collaboration formats include: | |
| - joint Hugging Face Spaces | |
| - framework integrations | |
| - trace visualizations | |
| - technical demos | |
| - benchmark projects | |
| - ecosystem maps | |
| - open-source integrations | |
| - clearly disclosed partnerships and sponsorships | |
| ## Collaboration Contact | |
| **agenten@magenta.de** | |
| --- | |
| # Independence | |
| **Observability Explorer** is an independent Hugging Face Space. | |
| It is not an official project of Hugging Face, any AI laboratory, model provider, observability vendor, agent framework or technology company. | |
| --- | |
| # Long-Term Vision | |
| The goal is to make AI observability easier to understand and apply across increasingly complex AI systems. | |
| > **Trace. Measure. Understand. Improve.** | |