File size: 13,673 Bytes
c0af099
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
# Integrating DefenseClaw with Agent Canvas

[DefenseClaw](https://github.com/cisco-ai-defense/defenseclaw) is a security governance layer for agentic AI runtimes β€” it scans skills and MCP servers before they run, inspects LLM traffic at runtime, and produces durable audit evidence. This guide explains how to run DefenseClaw alongside the [OpenHands Agent Server](https://github.com/OpenHands/software-agent-sdk/tree/main/openhands-agent-server) that powers Agent Canvas, without making any code-level changes to either project.

> **Status:** DefenseClaw is purpose-built around the OpenClaw runtime and its TypeScript plugin hooks. The integration described here targets the lowest-friction overlap points β€” skill injection, LLM proxying, CLI scanning, and audit export β€” that work without modifying Agent Canvas or DefenseClaw source code. [Future work](#future-work-code-level-extensions) describes deeper hooks that would require code changes.

---

## How the Two Systems Fit Together

```mermaid
flowchart TD
    UI["Agent Canvas (browser)"]
    AS["OpenHands Agent Server\nlocalhost:18000"]
    GP["DefenseClaw Guardrail Proxy\nlocalhost:4000"]
    LLM["LLM Provider"]
    GW["DefenseClaw Gateway Sidecar\nlocalhost:18970"]
    CLI["DefenseClaw CLI / TUI"]

    UI -->|HTTP| AS
    AS -->|LLM API calls| GP
    GP -->|forwarded request| LLM
    GW <-->|REST API| AS
    CLI <-->|REST API| GW

    style GW fill:#fff3cd,stroke:#856404
    style CLI fill:#fff3cd,stroke:#856404
    style GP fill:#f8d7da,stroke:#842029
```

**Shared concepts:**

| Agent Canvas / Agent Server | DefenseClaw equivalent |
|---|---|
| Skills (`.agents/skills/`) | Skills (scanned by `cisco-ai-skill-scanner` + CodeGuard) |
| MCP servers | MCP servers (scanned by `cisco-ai-mcp-scanner`) |
| LLM settings (`base_url`) | Guardrail proxy upstream target |
| Workspace files (generated code) | CodeGuard scan surface |
| Agent Server hooks | Potential enforcement point (future work) |

---

## Prerequisites

| Component | Version |
|---|---|
| Agent Canvas / Agent Server | Current `main` |
| Python | 3.10+ |
| Go | 1.26.2+ (for DefenseClaw gateway) |
| DefenseClaw | Latest release |

---

## Installation

### 1. Install and initialise DefenseClaw

```bash
# Install from the release script
curl -LsSf https://raw.githubusercontent.com/cisco-ai-defense/defenseclaw/main/scripts/install.sh | bash

# Initialise config and enable the guardrail proxy
defenseclaw init --enable-guardrail
```

Verify the installation:

```bash
defenseclaw doctor
```

Start the Go gateway sidecar (keep this running alongside the Agent Server):

```bash
defenseclaw-gateway start
```

### 2. Start Agent Canvas

Follow the standard [Agent Canvas quickstart](../README.md). The integration steps below assume the Agent Server is reachable at `http://localhost:18000`.

---

## Integration Points

### A. Load the CodeGuard Skill

DefenseClaw ships a ready-made OpenHands skill β€” `skills/codeguard/SKILL.md` β€” that teaches the agent the CodeGuard security rules. When the skill is active, the agent writes code that avoids the patterns DefenseClaw blocks at scan time (hardcoded secrets, `os.system()`, string-interpolated SQL, weak crypto, path traversal, etc.).

**Install the skill into a user or project skill directory:**

```bash
# User-level (applies to all Agent Server conversations on this machine)
mkdir -p ~/.agents/skills/codeguard
curl -fsSL https://raw.githubusercontent.com/cisco-ai-defense/defenseclaw/main/skills/codeguard/SKILL.md \
  -o ~/.agents/skills/codeguard/SKILL.md

# Project-level (checked in alongside your project, only affects that workspace)
mkdir -p .agents/skills/codeguard
curl -fsSL https://raw.githubusercontent.com/cisco-ai-defense/defenseclaw/main/skills/codeguard/SKILL.md \
  -o .agents/skills/codeguard/SKILL.md
```

The Agent Server loads skills from these directories automatically at conversation start. No restart of the server is required for user-level skills; project-level skills are loaded when the conversation workspace is opened.

**What this achieves:** The agent's system prompt is augmented with the full CodeGuard rule set. Code it generates will pre-emptively avoid the patterns that the downstream `defenseclaw codeguard scan` would flag.

---

### B. Route LLM Traffic Through the Guardrail Proxy

The DefenseClaw guardrail proxy runs on `localhost:4000` and acts as an OpenAI-compatible reverse proxy. Pointing the Agent Server's LLM calls through it causes every prompt and completion to be inspected β€” in observe mode (log only) or action mode (block on policy violations).

**Configure the LLM base URL in Agent Canvas:**

Open the Agent Canvas settings panel β†’ select your active backend β†’ under **LLM settings**, set **Base URL** to:

```
http://localhost:4000
```

Leave the model name and API key as-is. The proxy reads the original `Authorization` / `x-api-key` header, forwards the request to the real provider, and injects its own `X-DC-Target-URL` routing header β€” the agent code and Agent Server require no changes.

**Via environment variable (server-side):**

If you configure your Agent Server through environment variables, set the LLM base URL before starting it:

```bash
# Example using OpenAI; set model and key as normal, only base_url changes
export OH_LLM__BASE_URL="http://localhost:4000"
npm run dev
```

> Consult the Agent Server [settings schema](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-agent-server/openhands/agent_server/settings_router.py) for the exact environment variable name used in your deployment.

**Start the guardrail in observe mode (safe default) or action mode:**

```bash
# Observe β€” log findings, never block (recommended while tuning)
defenseclaw setup guardrail --mode observe --restart

# Action β€” block prompts and responses that match policies
defenseclaw setup guardrail --mode action --restart
```

**Supported providers:**

The DefenseClaw proxy handles Anthropic (`api.anthropic.com`), OpenAI (`api.openai.com`), OpenRouter, Azure OpenAI, Gemini, Ollama, and Bedrock. Provider detection is automatic based on the target URL.

---

### C. Scan Skills Before Loading

Before installing a skill from the marketplace or an external source into the Agent Server, use the DefenseClaw CLI to vet it:

```bash
# Scan a locally downloaded skill directory
defenseclaw skill scan path/to/skill-directory

# Scan an installed skill by name (requires the skill to be registered in the DefenseClaw inventory)
defenseclaw skill scan my-skill-name

# List all skills currently visible to DefenseClaw
defenseclaw skill list
```

The scanner applies `cisco-ai-skill-scanner` rules plus CodeGuard static analysis and emits a verdict (`PASS`, `WARN`, `BLOCK`) with per-finding details. HIGH and CRITICAL findings block skill use in action mode.

**Workflow recommendation:** Add `defenseclaw skill scan <skill-dir>` as a pre-commit or CI step in repositories that ship skills for Agent Canvas.

---

### D. Scan Agent-Generated Code

After an agent conversation produces code in the workspace, run CodeGuard on the output before committing:

```bash
# Scan an entire workspace directory
defenseclaw codeguard scan /path/to/workspace

# Scan a single file
defenseclaw codeguard scan /path/to/workspace/src/auth.py

# Output as JSON (useful in CI pipelines)
defenseclaw codeguard scan /path/to/workspace --json
```

CodeGuard checks for hardcoded secrets, dangerous command execution, SQL injection, unsafe deserialization, weak cryptography, SSRF-prone network calls, and path traversal β€” covering Python, JavaScript, TypeScript, Go, Java, Ruby, and PHP.

**Zero-friction CI gate example (GitHub Actions):**

```yaml
- name: Scan agent-generated code
  run: |
    defenseclaw codeguard scan ${{ github.workspace }} --json \
      | python3 -c "
    import sys, json
    findings = json.load(sys.stdin)
    criticals = [f for f in findings if f.get('severity') in ('HIGH','CRITICAL')]
    if criticals:
        for f in criticals:
            print(f'::error file={f[\"file\"]},line={f[\"line\"]}::{f[\"rule\"]}: {f[\"message\"]}')
        sys.exit(1)
    "
```

---

### E. Monitor via the DefenseClaw TUI and Audit Store

All scan results, guardrail decisions, tool-call inspections, and policy verdicts are written to DefenseClaw's SQLite audit store. The TUI gives a live operator view:

```bash
defenseclaw tui
```

The TUI panels cover:
- **Alerts** β€” recent HIGH/CRITICAL findings and blocked events
- **Scans** β€” historical scan results per skill/file
- **Tools** β€” tool-call verdicts from the inspection engine
- **Policy** β€” current block/allow lists

**Export to external systems:**

| Target | Setup |
|---|---|
| OTLP (Prometheus/Grafana/Honeycomb) | `defenseclaw setup observability --otlp-endpoint http://collector:4317` |
| Splunk HEC | `defenseclaw setup splunk --hec-url http://splunk:8088 --hec-token $TOKEN` |
| Slack / PagerDuty / Webex | `defenseclaw setup notifications --slack-webhook $SLACK_URL` |
| Local Splunk bundle (Docker) | `defenseclaw setup splunk --logs --accept-splunk-license` |

---

## Integration Summary

| Goal | Mechanism | Config change? | Code change? |
|---|---|---|---|
| Agent writes secure code by default | CodeGuard skill in `.agents/skills/` | Drop-in file | No |
| Inspect all LLM prompts and responses | Guardrail proxy at `localhost:4000` | Set `base_url` | No |
| Vet skills before loading | `defenseclaw skill scan` in CI/workflow | None | No |
| Scan agent-generated code | `defenseclaw codeguard scan <workspace>` | None | No |
| Audit trail and alerting | DefenseClaw TUI, OTLP, Splunk, webhooks | DefenseClaw config | No |

---

## Future Work: Code-Level Extensions

The following integrations would require changes to Agent Canvas, the Agent Server, or DefenseClaw, but would significantly deepen the security posture.

### 1. Native `SecurityAnalyzer` hook

The OpenHands SDK exposes a [`SecurityAnalyzer`](https://docs.openhands.dev/sdk/arch/security.md) interface. A custom implementation could call DefenseClaw's `/api/v1/inspect/tool` endpoint before every tool invocation β€” mirroring the inspection the OpenClaw TypeScript plugin performs. This would gate bash commands, file writes, and other tool calls through DefenseClaw's four-stage inspection pipeline (regex, Cisco AI Defense cloud rules, LLM judge, OPA policy) before they execute.

```python
# Sketch β€” not yet implemented
class DefenseClawSecurityAnalyzer(SecurityAnalyzer):
    async def analyze(self, action: Action) -> ActionSecurityRisk:
        resp = await httpx.post(
            "http://localhost:18970/api/v1/inspect/tool",
            json={"tool": action.tool_name, "args": action.args},
            headers={"X-DefenseClaw-Client": "agent-server"},
        )
        if resp.json()["action"] == "block":
            return ActionSecurityRisk.HIGH
        return ActionSecurityRisk.LOW
```

### 2. Skill install pipeline integration

The Agent Server's `skills_service.py` (`service_install_skill`) runs skill validation during install. A pre-install hook that calls `defenseclaw skill scan` and fails the install on HIGH/CRITICAL findings would enforce a mandatory scan gate β€” no skill reaches the agent without passing DefenseClaw's scanner. This change would live in `openhands-agent-server`.

### 3. Hooks integration

The Agent Server loads `.openhands/hooks.json` from the workspace. An `on_conversation_end` hook that runs `defenseclaw codeguard scan <workspace>` and writes findings to a structured report file would give per-session security evidence without manual operator intervention.

### 4. Agent Canvas security dashboard

A dedicated panel in the Agent Canvas UI that queries DefenseClaw's gateway REST API (`GET /alerts`, `GET /enforce/blocked`) would surface guardrail findings inline with the conversation view β€” correlating blocked prompts or tool calls with the agent turn that triggered them.

### 5. Agent Server β†’ DefenseClaw audit bridge

The Agent Server supports outgoing webhooks (`WebhookSpec`). A webhook handler that forwards conversation events to `POST /audit/event` on the DefenseClaw gateway would allow DefenseClaw's audit store to record Agent Server conversation lifecycle events (start, tool invocation, finish) alongside its own security findings β€” building a single correlated audit trail.

### 6. Skill registry alignment

DefenseClaw's registry system (`defenseclaw registry add`) ingests external skill/MCP catalogs from ClawHub, Smithery, skills.sh, HTTP YAML, and Git sources. Aligning the Agent Server's marketplace skill catalog with the DefenseClaw registry would allow `defenseclaw skill scan all` to exhaustively vet the entire available catalog, not just individually installed skills.

---

## References

- [DefenseClaw GitHub](https://github.com/cisco-ai-defense/defenseclaw)
- [DefenseClaw Quick Start](https://github.com/cisco-ai-defense/defenseclaw/blob/main/docs/QUICKSTART.md)
- [DefenseClaw API Reference](https://github.com/cisco-ai-defense/defenseclaw/blob/main/docs/API.md)
- [DefenseClaw Guardrail Architecture](https://github.com/cisco-ai-defense/defenseclaw/blob/main/docs/GUARDRAIL.md)
- [DefenseClaw CodeGuard Skill](https://github.com/cisco-ai-defense/defenseclaw/blob/main/skills/codeguard/SKILL.md)
- [OpenHands Agent Server](https://github.com/OpenHands/software-agent-sdk/tree/main/openhands-agent-server)
- [OpenHands SDK Security Analyzer](https://docs.openhands.dev/sdk/arch/security.md)
- [Agent Canvas Self-Hosting](./SELF_HOSTING.md)

---

_This document was created by an AI agent (OpenHands) on behalf of the user._