BonusLockSMith commited on
Commit
0a2f66d
Β·
verified Β·
1 Parent(s): a5c7816

add demo gif to README

Browse files
Files changed (1) hide show
  1. README.md +64 -60
README.md CHANGED
@@ -1,60 +1,64 @@
1
- ---
2
- title: DIY Creator
3
- emoji: πŸ› οΈ
4
- colorFrom: gray
5
- colorTo: red
6
- sdk: docker
7
- app_port: 7860
8
- pinned: false
9
- ---
10
-
11
- # DIY Creator πŸ› οΈ β€” a multi-agent team you can watch work
12
-
13
- Name something you want to build. A **team of three agents** hands off to each other to produce a
14
- vetted, step-by-step build guide β€” and you watch them do it in real time.
15
-
16
- > Project **#17** of the *"30 AI Projects in 15 Days"* challenge. Focus: **multi-agent orchestration.**
17
-
18
- ## The team
19
-
20
- | Agent | Role |
21
- |-------|------|
22
- | πŸ” **Researcher** | Gathers a practical brief: materials (with quantities), tools, techniques, safety hazards, common mistakes, and honest difficulty/time/cost. |
23
- | ✍️ **Writer** | Turns the brief into a clear, start-to-finish guide a beginner could actually follow. |
24
- | 🧐 **Critic** | An **independent reviewer** β€” scores the draft, flags issues by severity, and **sends it back**. The writer revises; the critic re-checks. The loop runs until the guide passes (or hits the round cap). |
25
-
26
- ## The lesson: roles, handoffs, and an honest gate
27
-
28
- A multi-agent "team" is three things done well:
29
- 1. **Role-specialized agents** β€” each one has a single job and a prompt tuned for it.
30
- 2. **Explicit handoffs** β€” the researcher's brief feeds the writer; the writer's draft feeds the critic.
31
- 3. **A critique β†’ revise loop** β€” nobody ships their own first draft.
32
-
33
- Two design choices worth calling out:
34
-
35
- - **The critic is independent (the "diamond").** It reviews the writer's work, ideally as a *different*
36
- model, so the review isn't the author grading itself. Set `DIY_CRITIC_MODEL` to a second local model
37
- for a sharper, more independent review.
38
- - **The gate is honest-by-construction.** The critic *finds* issues (what LLMs are good at), but the
39
- pass/revise decision is made in **code** from issue severity β€” a guide passes only when no high- or
40
- medium-severity issues remain. The model can't rubber-stamp its own work.
41
-
42
- ## Run it
43
-
44
- ```bash
45
- pip install -r requirements.txt
46
- python app.py # http://127.0.0.1:7860
47
- # Local Ollama by default (free/private). For the cloud model instead:
48
- # DIY_LLM_BACKEND=claude ANTHROPIC_API_KEY=sk-... python app.py
49
- ```
50
-
51
- **Config (env):** `DIY_LLM_BACKEND` (`ollama`|`claude`) Β· `DIY_OLLAMA_URL` Β· `DIY_OLLAMA_MODEL`
52
- (default `qwen2.5:7b`) Β· `DIY_CRITIC_MODEL` (optional independent critic) Β· `ANTHROPIC_API_KEY` /
53
- `ANTHROPIC_MODEL` for the Claude backend.
54
-
55
- > Note on models: a small local model self-critiquing is genuinely noisy β€” it will sometimes ship
56
- > "after review, minor notes remain" rather than a clean pass. That's honest. A stronger model (or a
57
- > `DIY_CRITIC_MODEL` diamond) converges to a clean pass more often.
58
-
59
- ---
60
- Built by **Robert Lucyk** Β· [GritAI Solutions](https://gritai.solutions) Β· part of the 30-in-15 challenge.
 
 
 
 
 
1
+ ---
2
+ title: DIY Creator
3
+ emoji: πŸ› οΈ
4
+ colorFrom: gray
5
+ colorTo: red
6
+ sdk: docker
7
+ app_port: 7860
8
+ pinned: false
9
+ ---
10
+
11
+ # DIY Creator πŸ› οΈ β€” a multi-agent team you can watch work
12
+
13
+ Name something you want to build. A **team of three agents** hands off to each other to produce a
14
+ vetted, step-by-step build guide β€” and you watch them do it in real time.
15
+
16
+ > Project **#17** of the *"30 AI Projects in 15 Days"* challenge. Focus: **multi-agent orchestration.**
17
+
18
+ ![Demo: the Researcher gathers a brief, the Writer drafts a guide, the Critic scores it and sends it back to revise](demo.gif)
19
+
20
+ *Above (local model): πŸ” Researcher gathers the brief β†’ ✍️ Writer drafts a guide β†’ 🧐 Critic scores it **5/10 and sends it back** with concrete fixes β†’ Writer revises β†’ the guide ships (honestly noting any minor items that remain). On a stronger backend it converges to a clean pass.*
21
+
22
+ ## The team
23
+
24
+ | Agent | Role |
25
+ |-------|------|
26
+ | πŸ” **Researcher** | Gathers a practical brief: materials (with quantities), tools, techniques, safety hazards, common mistakes, and honest difficulty/time/cost. |
27
+ | ✍️ **Writer** | Turns the brief into a clear, start-to-finish guide a beginner could actually follow. |
28
+ | 🧐 **Critic** | An **independent reviewer** β€” scores the draft, flags issues by severity, and **sends it back**. The writer revises; the critic re-checks. The loop runs until the guide passes (or hits the round cap). |
29
+
30
+ ## The lesson: roles, handoffs, and an honest gate
31
+
32
+ A multi-agent "team" is three things done well:
33
+ 1. **Role-specialized agents** β€” each one has a single job and a prompt tuned for it.
34
+ 2. **Explicit handoffs** β€” the researcher's brief feeds the writer; the writer's draft feeds the critic.
35
+ 3. **A critique β†’ revise loop** β€” nobody ships their own first draft.
36
+
37
+ Two design choices worth calling out:
38
+
39
+ - **The critic is independent (the "diamond").** It reviews the writer's work, ideally as a *different*
40
+ model, so the review isn't the author grading itself. Set `DIY_CRITIC_MODEL` to a second local model
41
+ for a sharper, more independent review.
42
+ - **The gate is honest-by-construction.** The critic *finds* issues (what LLMs are good at), but the
43
+ pass/revise decision is made in **code** from issue severity β€” a guide passes only when no high- or
44
+ medium-severity issues remain. The model can't rubber-stamp its own work.
45
+
46
+ ## Run it
47
+
48
+ ```bash
49
+ pip install -r requirements.txt
50
+ python app.py # http://127.0.0.1:7860
51
+ # Local Ollama by default (free/private). For the cloud model instead:
52
+ # DIY_LLM_BACKEND=claude ANTHROPIC_API_KEY=sk-... python app.py
53
+ ```
54
+
55
+ **Config (env):** `DIY_LLM_BACKEND` (`ollama`|`claude`) Β· `DIY_OLLAMA_URL` Β· `DIY_OLLAMA_MODEL`
56
+ (default `qwen2.5:7b`) Β· `DIY_CRITIC_MODEL` (optional independent critic) Β· `ANTHROPIC_API_KEY` /
57
+ `ANTHROPIC_MODEL` for the Claude backend.
58
+
59
+ > Note on models: a small local model self-critiquing is genuinely noisy β€” it will sometimes ship
60
+ > "after review, minor notes remain" rather than a clean pass. That's honest. A stronger model (or a
61
+ > `DIY_CRITIC_MODEL` diamond) converges to a clean pass more often.
62
+
63
+ ---
64
+ Built by **Robert Lucyk** Β· [GritAI Solutions](https://gritai.solutions) Β· part of the 30-in-15 challenge.