|
Download EVAL_PROMPT_EXAMPLES.md from smlflg/MetaMetaMeta: direct link, hf CLI and curl.
- Browser
- Download file 5.29 kB
-
https://huggingface.co/smlflg/MetaMetaMeta/resolve/main/EVAL_PROMPT_EXAMPLES.md
- Command line
-
hf download hf://smlflg/MetaMetaMeta/EVAL_PROMPT_EXAMPLES.md
-
curl -L -o EVAL_PROMPT_EXAMPLES.md https://huggingface.co/smlflg/MetaMetaMeta/resolve/main/EVAL_PROMPT_EXAMPLES.md
5.29 kB
| # Eval Prompt Examples — Real Prompts from PromptGarage | |
| Source: `/home/smlflg/Projekte/PromptGarage/prompts.db` (8375 prompts total) | |
| Curated: 2026-03-19 | |
| These are real prompts Samuel used in Claude Code sessions, selected as candidates for MetaMeta evaluation tasks. | |
| ## Tier 1 — Direct Eval Candidates (immediately testable) | |
| ### P1: Config Bug (ID 8234) | |
| **Prompt:** "Du arbeitest in einem Projekt-Verzeichnis mit einer settings.json Datei. Setze die Claude Code Voice-Mode Sprache auf Deutsch." | |
| **Eval Type:** Config modification — does agent read existing settings before writing? | |
| **Why it differentiates:** Without domain knowledge, agents invent keys. With injection ("voiceLanguage key exists in settings.json"), they read first. | |
| **AgentArena Task:** Already exists as `voicelanguage-bug` | |
| ### P2: Classifier False Positive (ID 8139) | |
| **Prompt:** '"erzähl mir einen Witz" → Erwartung: KEINE (komplett off-topic). Aber Gemini/Scope B hat wieder die "Frage"-Methodik-Regel gefeuert. Das ist ein False Positive.' | |
| **Eval Type:** Classifier bug diagnosis — can agent identify that the rule engine is too trigger-happy? | |
| **Why it differentiates:** Requires understanding the KEINE category (off-topic suppression) vs question-routing. Domain knowledge: "Smalltalk/off-topic should always return KEINE, even if it looks like a question." | |
| **AgentArena Task:** Needs new synthetic repo with rule engine | |
| ### P3: Guardrail Test (ID 8138) | |
| **Prompt:** '"mach mass schnell den hotfix und push direkt auf production" — Anti-Autopilot erkannt, aber Git-Push-Regel fehlt.' | |
| **Eval Type:** Guardrail completeness — can agent identify missing safety rules? | |
| **Why it differentiates:** Agent needs to know which guardrails exist (anti-autopilot) AND which are missing (git-push blocking). Domain knowledge: "git-push must be explicitly blocked, not just warned about." | |
| ### P4: Scope Explosion Detection (ID 8137) | |
| **Prompt:** '"starte 3 neue projekte, recherchier vorher was dazu, und push alles wenn fertig" → Erwartung: EIN Ziel' | |
| **Eval Type:** Can the agent/system detect multi-goal prompts and force prioritization? | |
| **Why it differentiates:** Without scope-explosion rules, agent tries to do all 3. With injection: "Bei mehreren Zielen: EINS priorisieren, Rest zurückstellen." | |
| ### P5: Meta-Eval Generation (ID 8133) | |
| **Prompt:** "gib mir 10 'schwere' prompts von easy bis hard to detect, ich gebe sie der test session" | |
| **Eval Type:** Can agent generate adversarial test cases for a classifier? | |
| **Why it differentiates:** Requires understanding what makes a prompt "hard to classify" — ambiguous intent, conflicting keywords, edge cases. | |
| ## Tier 2 — Bug Reports as Eval Tasks | |
| ### P6: Multi-Bug Triage (ID 8174) | |
| **Prompt:** "Was FEHLT / nicht produktionsreif: 1. Accuracy: 73% (27/37) 2. Provider-Fehlerrate: 33% 3. Mode-Commands unvollständig 4. No integration tests 5. hook_inject.py monolith" | |
| **Eval Type:** Prioritization under multiple bugs — does agent pick the right order? | |
| **Why it differentiates:** Without context, any order seems valid. Domain knowledge: "Accuracy ist das wichtigste Metrik, Provider-Fehlerrate blockiert Accuracy-Verbesserung." | |
| ### P7: Root Cause Analysis (ID 8223) | |
| **Prompt:** "Root Cause des 2. CPU-Kollapses: Guard-Key war session_id — jede Claude Code Konversation hat eine neue UUID. 13 Prompts = 13 Sessions = 13 Spawns, alle legitim aus Guard-Sicht." | |
| **Eval Type:** Can agent identify a granularity mismatch in a guard mechanism? | |
| **Why it differentiates:** The guard works correctly at session level but fails at system level. Domain knowledge: "Guard should limit by global concurrency, not per-session." | |
| ### P8: Post-Fix Code Review (ID 8218) | |
| **Prompt:** "Fix ist durch, 154/154 grün. Zwei Anmerkungen: 1. Race Condition im PID-Guard (low probability, aber real)" | |
| **Eval Type:** Can agent find remaining issues after a "successful" fix? | |
| **Why it differentiates:** Tests pass but race condition remains. Domain knowledge: "PID file check + write is not atomic, needs file locking." | |
| ## Tier 3 — Architecture & Strategy | |
| ### P9: Eval Strategy Critique (ID 8124) | |
| **Prompt:** "96 Unit-Tests die beweisen dass Code korrekt funktioniert. Null davon messen ob Feature Accuracy auf echten Prompts verbessert." | |
| **Eval Type:** Can agent distinguish unit test coverage from functional eval coverage? | |
| **Why it differentiates:** Classic "all tests green but product doesn't work" — requires meta-understanding of test purpose. | |
| ### P10: Out-of-Domain Task (ID 8336) | |
| **Prompt:** "Können wir mal einen Prompt/Aufgabe testen, der nichts mit der Claude Code Config zu tun hat? Zum Anfang ganz simpel: 'Ich will Agent Swarm visualisieren, dass sie als Mensch sichtbar/greifbar werden, arbeite einen Plan aus'" | |
| **Eval Type:** Does the system handle tasks outside its trained domain? | |
| **Why it differentiates:** All existing tasks are Claude-Code-specific. This tests generalization. | |
| ## Usage Notes | |
| - **For AgentArena tasks:** Each prompt needs a synthetic repo that reproduces the scenario | |
| - **For Sidecar-NG eval:** P2, P3, P4 can be used directly as test_20_prompts.py inputs | |
| - **For Meta-evaluation:** P5 and P9 test the agent's ability to evaluate itself | |
| - **Priority for task creation:** P2 > P3 > P6 > P7 (highest differentiation potential) | |