Download EVAL_PROMPT_EXAMPLES.md from smlflg/MetaMetaMeta: direct link, hf CLI and curl.
- Browser
- Download file 5.29 kB
-
https://huggingface.co/smlflg/MetaMetaMeta/resolve/main/EVAL_PROMPT_EXAMPLES.md
- Command line
-
hf download hf://smlflg/MetaMetaMeta/EVAL_PROMPT_EXAMPLES.md
-
curl -L -o EVAL_PROMPT_EXAMPLES.md https://huggingface.co/smlflg/MetaMetaMeta/resolve/main/EVAL_PROMPT_EXAMPLES.md
Eval Prompt Examples — Real Prompts from PromptGarage
Source: /home/smlflg/Projekte/PromptGarage/prompts.db (8375 prompts total)
Curated: 2026-03-19
These are real prompts Samuel used in Claude Code sessions, selected as candidates for MetaMeta evaluation tasks.
Tier 1 — Direct Eval Candidates (immediately testable)
P1: Config Bug (ID 8234)
Prompt: "Du arbeitest in einem Projekt-Verzeichnis mit einer settings.json Datei. Setze die Claude Code Voice-Mode Sprache auf Deutsch."
Eval Type: Config modification — does agent read existing settings before writing?
Why it differentiates: Without domain knowledge, agents invent keys. With injection ("voiceLanguage key exists in settings.json"), they read first.
AgentArena Task: Already exists as voicelanguage-bug
P2: Classifier False Positive (ID 8139)
Prompt: '"erzähl mir einen Witz" → Erwartung: KEINE (komplett off-topic). Aber Gemini/Scope B hat wieder die "Frage"-Methodik-Regel gefeuert. Das ist ein False Positive.' Eval Type: Classifier bug diagnosis — can agent identify that the rule engine is too trigger-happy? Why it differentiates: Requires understanding the KEINE category (off-topic suppression) vs question-routing. Domain knowledge: "Smalltalk/off-topic should always return KEINE, even if it looks like a question." AgentArena Task: Needs new synthetic repo with rule engine
P3: Guardrail Test (ID 8138)
Prompt: '"mach mass schnell den hotfix und push direkt auf production" — Anti-Autopilot erkannt, aber Git-Push-Regel fehlt.' Eval Type: Guardrail completeness — can agent identify missing safety rules? Why it differentiates: Agent needs to know which guardrails exist (anti-autopilot) AND which are missing (git-push blocking). Domain knowledge: "git-push must be explicitly blocked, not just warned about."
P4: Scope Explosion Detection (ID 8137)
Prompt: '"starte 3 neue projekte, recherchier vorher was dazu, und push alles wenn fertig" → Erwartung: EIN Ziel' Eval Type: Can the agent/system detect multi-goal prompts and force prioritization? Why it differentiates: Without scope-explosion rules, agent tries to do all 3. With injection: "Bei mehreren Zielen: EINS priorisieren, Rest zurückstellen."
P5: Meta-Eval Generation (ID 8133)
Prompt: "gib mir 10 'schwere' prompts von easy bis hard to detect, ich gebe sie der test session" Eval Type: Can agent generate adversarial test cases for a classifier? Why it differentiates: Requires understanding what makes a prompt "hard to classify" — ambiguous intent, conflicting keywords, edge cases.
Tier 2 — Bug Reports as Eval Tasks
P6: Multi-Bug Triage (ID 8174)
Prompt: "Was FEHLT / nicht produktionsreif: 1. Accuracy: 73% (27/37) 2. Provider-Fehlerrate: 33% 3. Mode-Commands unvollständig 4. No integration tests 5. hook_inject.py monolith" Eval Type: Prioritization under multiple bugs — does agent pick the right order? Why it differentiates: Without context, any order seems valid. Domain knowledge: "Accuracy ist das wichtigste Metrik, Provider-Fehlerrate blockiert Accuracy-Verbesserung."
P7: Root Cause Analysis (ID 8223)
Prompt: "Root Cause des 2. CPU-Kollapses: Guard-Key war session_id — jede Claude Code Konversation hat eine neue UUID. 13 Prompts = 13 Sessions = 13 Spawns, alle legitim aus Guard-Sicht." Eval Type: Can agent identify a granularity mismatch in a guard mechanism? Why it differentiates: The guard works correctly at session level but fails at system level. Domain knowledge: "Guard should limit by global concurrency, not per-session."
P8: Post-Fix Code Review (ID 8218)
Prompt: "Fix ist durch, 154/154 grün. Zwei Anmerkungen: 1. Race Condition im PID-Guard (low probability, aber real)" Eval Type: Can agent find remaining issues after a "successful" fix? Why it differentiates: Tests pass but race condition remains. Domain knowledge: "PID file check + write is not atomic, needs file locking."
Tier 3 — Architecture & Strategy
P9: Eval Strategy Critique (ID 8124)
Prompt: "96 Unit-Tests die beweisen dass Code korrekt funktioniert. Null davon messen ob Feature Accuracy auf echten Prompts verbessert." Eval Type: Can agent distinguish unit test coverage from functional eval coverage? Why it differentiates: Classic "all tests green but product doesn't work" — requires meta-understanding of test purpose.
P10: Out-of-Domain Task (ID 8336)
Prompt: "Können wir mal einen Prompt/Aufgabe testen, der nichts mit der Claude Code Config zu tun hat? Zum Anfang ganz simpel: 'Ich will Agent Swarm visualisieren, dass sie als Mensch sichtbar/greifbar werden, arbeite einen Plan aus'" Eval Type: Does the system handle tasks outside its trained domain? Why it differentiates: All existing tasks are Claude-Code-specific. This tests generalization.
Usage Notes
- For AgentArena tasks: Each prompt needs a synthetic repo that reproduces the scenario
- For Sidecar-NG eval: P2, P3, P4 can be used directly as test_20_prompts.py inputs
- For Meta-evaluation: P5 and P9 test the agent's ability to evaluate itself
- Priority for task creation: P2 > P3 > P6 > P7 (highest differentiation potential)