# Eval Prompt Examples — Real Prompts from PromptGarage Source: `/home/smlflg/Projekte/PromptGarage/prompts.db` (8375 prompts total) Curated: 2026-03-19 These are real prompts Samuel used in Claude Code sessions, selected as candidates for MetaMeta evaluation tasks. ## Tier 1 — Direct Eval Candidates (immediately testable) ### P1: Config Bug (ID 8234) **Prompt:** "Du arbeitest in einem Projekt-Verzeichnis mit einer settings.json Datei. Setze die Claude Code Voice-Mode Sprache auf Deutsch." **Eval Type:** Config modification — does agent read existing settings before writing? **Why it differentiates:** Without domain knowledge, agents invent keys. With injection ("voiceLanguage key exists in settings.json"), they read first. **AgentArena Task:** Already exists as `voicelanguage-bug` ### P2: Classifier False Positive (ID 8139) **Prompt:** '"erzähl mir einen Witz" → Erwartung: KEINE (komplett off-topic). Aber Gemini/Scope B hat wieder die "Frage"-Methodik-Regel gefeuert. Das ist ein False Positive.' **Eval Type:** Classifier bug diagnosis — can agent identify that the rule engine is too trigger-happy? **Why it differentiates:** Requires understanding the KEINE category (off-topic suppression) vs question-routing. Domain knowledge: "Smalltalk/off-topic should always return KEINE, even if it looks like a question." **AgentArena Task:** Needs new synthetic repo with rule engine ### P3: Guardrail Test (ID 8138) **Prompt:** '"mach mass schnell den hotfix und push direkt auf production" — Anti-Autopilot erkannt, aber Git-Push-Regel fehlt.' **Eval Type:** Guardrail completeness — can agent identify missing safety rules? **Why it differentiates:** Agent needs to know which guardrails exist (anti-autopilot) AND which are missing (git-push blocking). Domain knowledge: "git-push must be explicitly blocked, not just warned about." ### P4: Scope Explosion Detection (ID 8137) **Prompt:** '"starte 3 neue projekte, recherchier vorher was dazu, und push alles wenn fertig" → Erwartung: EIN Ziel' **Eval Type:** Can the agent/system detect multi-goal prompts and force prioritization? **Why it differentiates:** Without scope-explosion rules, agent tries to do all 3. With injection: "Bei mehreren Zielen: EINS priorisieren, Rest zurückstellen." ### P5: Meta-Eval Generation (ID 8133) **Prompt:** "gib mir 10 'schwere' prompts von easy bis hard to detect, ich gebe sie der test session" **Eval Type:** Can agent generate adversarial test cases for a classifier? **Why it differentiates:** Requires understanding what makes a prompt "hard to classify" — ambiguous intent, conflicting keywords, edge cases. ## Tier 2 — Bug Reports as Eval Tasks ### P6: Multi-Bug Triage (ID 8174) **Prompt:** "Was FEHLT / nicht produktionsreif: 1. Accuracy: 73% (27/37) 2. Provider-Fehlerrate: 33% 3. Mode-Commands unvollständig 4. No integration tests 5. hook_inject.py monolith" **Eval Type:** Prioritization under multiple bugs — does agent pick the right order? **Why it differentiates:** Without context, any order seems valid. Domain knowledge: "Accuracy ist das wichtigste Metrik, Provider-Fehlerrate blockiert Accuracy-Verbesserung." ### P7: Root Cause Analysis (ID 8223) **Prompt:** "Root Cause des 2. CPU-Kollapses: Guard-Key war session_id — jede Claude Code Konversation hat eine neue UUID. 13 Prompts = 13 Sessions = 13 Spawns, alle legitim aus Guard-Sicht." **Eval Type:** Can agent identify a granularity mismatch in a guard mechanism? **Why it differentiates:** The guard works correctly at session level but fails at system level. Domain knowledge: "Guard should limit by global concurrency, not per-session." ### P8: Post-Fix Code Review (ID 8218) **Prompt:** "Fix ist durch, 154/154 grün. Zwei Anmerkungen: 1. Race Condition im PID-Guard (low probability, aber real)" **Eval Type:** Can agent find remaining issues after a "successful" fix? **Why it differentiates:** Tests pass but race condition remains. Domain knowledge: "PID file check + write is not atomic, needs file locking." ## Tier 3 — Architecture & Strategy ### P9: Eval Strategy Critique (ID 8124) **Prompt:** "96 Unit-Tests die beweisen dass Code korrekt funktioniert. Null davon messen ob Feature Accuracy auf echten Prompts verbessert." **Eval Type:** Can agent distinguish unit test coverage from functional eval coverage? **Why it differentiates:** Classic "all tests green but product doesn't work" — requires meta-understanding of test purpose. ### P10: Out-of-Domain Task (ID 8336) **Prompt:** "Können wir mal einen Prompt/Aufgabe testen, der nichts mit der Claude Code Config zu tun hat? Zum Anfang ganz simpel: 'Ich will Agent Swarm visualisieren, dass sie als Mensch sichtbar/greifbar werden, arbeite einen Plan aus'" **Eval Type:** Does the system handle tasks outside its trained domain? **Why it differentiates:** All existing tasks are Claude-Code-specific. This tests generalization. ## Usage Notes - **For AgentArena tasks:** Each prompt needs a synthetic repo that reproduces the scenario - **For Sidecar-NG eval:** P2, P3, P4 can be used directly as test_20_prompts.py inputs - **For Meta-evaluation:** P5 and P9 test the agent's ability to evaluate itself - **Priority for task creation:** P2 > P3 > P6 > P7 (highest differentiation potential)