File size: 13,673 Bytes
c0af099 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 | # Integrating DefenseClaw with Agent Canvas
[DefenseClaw](https://github.com/cisco-ai-defense/defenseclaw) is a security governance layer for agentic AI runtimes β it scans skills and MCP servers before they run, inspects LLM traffic at runtime, and produces durable audit evidence. This guide explains how to run DefenseClaw alongside the [OpenHands Agent Server](https://github.com/OpenHands/software-agent-sdk/tree/main/openhands-agent-server) that powers Agent Canvas, without making any code-level changes to either project.
> **Status:** DefenseClaw is purpose-built around the OpenClaw runtime and its TypeScript plugin hooks. The integration described here targets the lowest-friction overlap points β skill injection, LLM proxying, CLI scanning, and audit export β that work without modifying Agent Canvas or DefenseClaw source code. [Future work](#future-work-code-level-extensions) describes deeper hooks that would require code changes.
---
## How the Two Systems Fit Together
```mermaid
flowchart TD
UI["Agent Canvas (browser)"]
AS["OpenHands Agent Server\nlocalhost:18000"]
GP["DefenseClaw Guardrail Proxy\nlocalhost:4000"]
LLM["LLM Provider"]
GW["DefenseClaw Gateway Sidecar\nlocalhost:18970"]
CLI["DefenseClaw CLI / TUI"]
UI -->|HTTP| AS
AS -->|LLM API calls| GP
GP -->|forwarded request| LLM
GW <-->|REST API| AS
CLI <-->|REST API| GW
style GW fill:#fff3cd,stroke:#856404
style CLI fill:#fff3cd,stroke:#856404
style GP fill:#f8d7da,stroke:#842029
```
**Shared concepts:**
| Agent Canvas / Agent Server | DefenseClaw equivalent |
|---|---|
| Skills (`.agents/skills/`) | Skills (scanned by `cisco-ai-skill-scanner` + CodeGuard) |
| MCP servers | MCP servers (scanned by `cisco-ai-mcp-scanner`) |
| LLM settings (`base_url`) | Guardrail proxy upstream target |
| Workspace files (generated code) | CodeGuard scan surface |
| Agent Server hooks | Potential enforcement point (future work) |
---
## Prerequisites
| Component | Version |
|---|---|
| Agent Canvas / Agent Server | Current `main` |
| Python | 3.10+ |
| Go | 1.26.2+ (for DefenseClaw gateway) |
| DefenseClaw | Latest release |
---
## Installation
### 1. Install and initialise DefenseClaw
```bash
# Install from the release script
curl -LsSf https://raw.githubusercontent.com/cisco-ai-defense/defenseclaw/main/scripts/install.sh | bash
# Initialise config and enable the guardrail proxy
defenseclaw init --enable-guardrail
```
Verify the installation:
```bash
defenseclaw doctor
```
Start the Go gateway sidecar (keep this running alongside the Agent Server):
```bash
defenseclaw-gateway start
```
### 2. Start Agent Canvas
Follow the standard [Agent Canvas quickstart](../README.md). The integration steps below assume the Agent Server is reachable at `http://localhost:18000`.
---
## Integration Points
### A. Load the CodeGuard Skill
DefenseClaw ships a ready-made OpenHands skill β `skills/codeguard/SKILL.md` β that teaches the agent the CodeGuard security rules. When the skill is active, the agent writes code that avoids the patterns DefenseClaw blocks at scan time (hardcoded secrets, `os.system()`, string-interpolated SQL, weak crypto, path traversal, etc.).
**Install the skill into a user or project skill directory:**
```bash
# User-level (applies to all Agent Server conversations on this machine)
mkdir -p ~/.agents/skills/codeguard
curl -fsSL https://raw.githubusercontent.com/cisco-ai-defense/defenseclaw/main/skills/codeguard/SKILL.md \
-o ~/.agents/skills/codeguard/SKILL.md
# Project-level (checked in alongside your project, only affects that workspace)
mkdir -p .agents/skills/codeguard
curl -fsSL https://raw.githubusercontent.com/cisco-ai-defense/defenseclaw/main/skills/codeguard/SKILL.md \
-o .agents/skills/codeguard/SKILL.md
```
The Agent Server loads skills from these directories automatically at conversation start. No restart of the server is required for user-level skills; project-level skills are loaded when the conversation workspace is opened.
**What this achieves:** The agent's system prompt is augmented with the full CodeGuard rule set. Code it generates will pre-emptively avoid the patterns that the downstream `defenseclaw codeguard scan` would flag.
---
### B. Route LLM Traffic Through the Guardrail Proxy
The DefenseClaw guardrail proxy runs on `localhost:4000` and acts as an OpenAI-compatible reverse proxy. Pointing the Agent Server's LLM calls through it causes every prompt and completion to be inspected β in observe mode (log only) or action mode (block on policy violations).
**Configure the LLM base URL in Agent Canvas:**
Open the Agent Canvas settings panel β select your active backend β under **LLM settings**, set **Base URL** to:
```
http://localhost:4000
```
Leave the model name and API key as-is. The proxy reads the original `Authorization` / `x-api-key` header, forwards the request to the real provider, and injects its own `X-DC-Target-URL` routing header β the agent code and Agent Server require no changes.
**Via environment variable (server-side):**
If you configure your Agent Server through environment variables, set the LLM base URL before starting it:
```bash
# Example using OpenAI; set model and key as normal, only base_url changes
export OH_LLM__BASE_URL="http://localhost:4000"
npm run dev
```
> Consult the Agent Server [settings schema](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-agent-server/openhands/agent_server/settings_router.py) for the exact environment variable name used in your deployment.
**Start the guardrail in observe mode (safe default) or action mode:**
```bash
# Observe β log findings, never block (recommended while tuning)
defenseclaw setup guardrail --mode observe --restart
# Action β block prompts and responses that match policies
defenseclaw setup guardrail --mode action --restart
```
**Supported providers:**
The DefenseClaw proxy handles Anthropic (`api.anthropic.com`), OpenAI (`api.openai.com`), OpenRouter, Azure OpenAI, Gemini, Ollama, and Bedrock. Provider detection is automatic based on the target URL.
---
### C. Scan Skills Before Loading
Before installing a skill from the marketplace or an external source into the Agent Server, use the DefenseClaw CLI to vet it:
```bash
# Scan a locally downloaded skill directory
defenseclaw skill scan path/to/skill-directory
# Scan an installed skill by name (requires the skill to be registered in the DefenseClaw inventory)
defenseclaw skill scan my-skill-name
# List all skills currently visible to DefenseClaw
defenseclaw skill list
```
The scanner applies `cisco-ai-skill-scanner` rules plus CodeGuard static analysis and emits a verdict (`PASS`, `WARN`, `BLOCK`) with per-finding details. HIGH and CRITICAL findings block skill use in action mode.
**Workflow recommendation:** Add `defenseclaw skill scan <skill-dir>` as a pre-commit or CI step in repositories that ship skills for Agent Canvas.
---
### D. Scan Agent-Generated Code
After an agent conversation produces code in the workspace, run CodeGuard on the output before committing:
```bash
# Scan an entire workspace directory
defenseclaw codeguard scan /path/to/workspace
# Scan a single file
defenseclaw codeguard scan /path/to/workspace/src/auth.py
# Output as JSON (useful in CI pipelines)
defenseclaw codeguard scan /path/to/workspace --json
```
CodeGuard checks for hardcoded secrets, dangerous command execution, SQL injection, unsafe deserialization, weak cryptography, SSRF-prone network calls, and path traversal β covering Python, JavaScript, TypeScript, Go, Java, Ruby, and PHP.
**Zero-friction CI gate example (GitHub Actions):**
```yaml
- name: Scan agent-generated code
run: |
defenseclaw codeguard scan ${{ github.workspace }} --json \
| python3 -c "
import sys, json
findings = json.load(sys.stdin)
criticals = [f for f in findings if f.get('severity') in ('HIGH','CRITICAL')]
if criticals:
for f in criticals:
print(f'::error file={f[\"file\"]},line={f[\"line\"]}::{f[\"rule\"]}: {f[\"message\"]}')
sys.exit(1)
"
```
---
### E. Monitor via the DefenseClaw TUI and Audit Store
All scan results, guardrail decisions, tool-call inspections, and policy verdicts are written to DefenseClaw's SQLite audit store. The TUI gives a live operator view:
```bash
defenseclaw tui
```
The TUI panels cover:
- **Alerts** β recent HIGH/CRITICAL findings and blocked events
- **Scans** β historical scan results per skill/file
- **Tools** β tool-call verdicts from the inspection engine
- **Policy** β current block/allow lists
**Export to external systems:**
| Target | Setup |
|---|---|
| OTLP (Prometheus/Grafana/Honeycomb) | `defenseclaw setup observability --otlp-endpoint http://collector:4317` |
| Splunk HEC | `defenseclaw setup splunk --hec-url http://splunk:8088 --hec-token $TOKEN` |
| Slack / PagerDuty / Webex | `defenseclaw setup notifications --slack-webhook $SLACK_URL` |
| Local Splunk bundle (Docker) | `defenseclaw setup splunk --logs --accept-splunk-license` |
---
## Integration Summary
| Goal | Mechanism | Config change? | Code change? |
|---|---|---|---|
| Agent writes secure code by default | CodeGuard skill in `.agents/skills/` | Drop-in file | No |
| Inspect all LLM prompts and responses | Guardrail proxy at `localhost:4000` | Set `base_url` | No |
| Vet skills before loading | `defenseclaw skill scan` in CI/workflow | None | No |
| Scan agent-generated code | `defenseclaw codeguard scan <workspace>` | None | No |
| Audit trail and alerting | DefenseClaw TUI, OTLP, Splunk, webhooks | DefenseClaw config | No |
---
## Future Work: Code-Level Extensions
The following integrations would require changes to Agent Canvas, the Agent Server, or DefenseClaw, but would significantly deepen the security posture.
### 1. Native `SecurityAnalyzer` hook
The OpenHands SDK exposes a [`SecurityAnalyzer`](https://docs.openhands.dev/sdk/arch/security.md) interface. A custom implementation could call DefenseClaw's `/api/v1/inspect/tool` endpoint before every tool invocation β mirroring the inspection the OpenClaw TypeScript plugin performs. This would gate bash commands, file writes, and other tool calls through DefenseClaw's four-stage inspection pipeline (regex, Cisco AI Defense cloud rules, LLM judge, OPA policy) before they execute.
```python
# Sketch β not yet implemented
class DefenseClawSecurityAnalyzer(SecurityAnalyzer):
async def analyze(self, action: Action) -> ActionSecurityRisk:
resp = await httpx.post(
"http://localhost:18970/api/v1/inspect/tool",
json={"tool": action.tool_name, "args": action.args},
headers={"X-DefenseClaw-Client": "agent-server"},
)
if resp.json()["action"] == "block":
return ActionSecurityRisk.HIGH
return ActionSecurityRisk.LOW
```
### 2. Skill install pipeline integration
The Agent Server's `skills_service.py` (`service_install_skill`) runs skill validation during install. A pre-install hook that calls `defenseclaw skill scan` and fails the install on HIGH/CRITICAL findings would enforce a mandatory scan gate β no skill reaches the agent without passing DefenseClaw's scanner. This change would live in `openhands-agent-server`.
### 3. Hooks integration
The Agent Server loads `.openhands/hooks.json` from the workspace. An `on_conversation_end` hook that runs `defenseclaw codeguard scan <workspace>` and writes findings to a structured report file would give per-session security evidence without manual operator intervention.
### 4. Agent Canvas security dashboard
A dedicated panel in the Agent Canvas UI that queries DefenseClaw's gateway REST API (`GET /alerts`, `GET /enforce/blocked`) would surface guardrail findings inline with the conversation view β correlating blocked prompts or tool calls with the agent turn that triggered them.
### 5. Agent Server β DefenseClaw audit bridge
The Agent Server supports outgoing webhooks (`WebhookSpec`). A webhook handler that forwards conversation events to `POST /audit/event` on the DefenseClaw gateway would allow DefenseClaw's audit store to record Agent Server conversation lifecycle events (start, tool invocation, finish) alongside its own security findings β building a single correlated audit trail.
### 6. Skill registry alignment
DefenseClaw's registry system (`defenseclaw registry add`) ingests external skill/MCP catalogs from ClawHub, Smithery, skills.sh, HTTP YAML, and Git sources. Aligning the Agent Server's marketplace skill catalog with the DefenseClaw registry would allow `defenseclaw skill scan all` to exhaustively vet the entire available catalog, not just individually installed skills.
---
## References
- [DefenseClaw GitHub](https://github.com/cisco-ai-defense/defenseclaw)
- [DefenseClaw Quick Start](https://github.com/cisco-ai-defense/defenseclaw/blob/main/docs/QUICKSTART.md)
- [DefenseClaw API Reference](https://github.com/cisco-ai-defense/defenseclaw/blob/main/docs/API.md)
- [DefenseClaw Guardrail Architecture](https://github.com/cisco-ai-defense/defenseclaw/blob/main/docs/GUARDRAIL.md)
- [DefenseClaw CodeGuard Skill](https://github.com/cisco-ai-defense/defenseclaw/blob/main/skills/codeguard/SKILL.md)
- [OpenHands Agent Server](https://github.com/OpenHands/software-agent-sdk/tree/main/openhands-agent-server)
- [OpenHands SDK Security Analyzer](https://docs.openhands.dev/sdk/arch/security.md)
- [Agent Canvas Self-Hosting](./SELF_HOSTING.md)
---
_This document was created by an AI agent (OpenHands) on behalf of the user._
|