AI & ML interests

AI worker on a Mac mini that drafts quotes from your price list; a person checks each one · local open-weight models · published results, misses shown · free score on 30 past quotes: hadi@staffbox.ai

Recent Activity

irvani  updated a Space about 23 hours ago
staffbox/README
irvani  updated a dataset about 23 hours ago
staffbox/staffbox-bench
irvani  updated a dataset about 23 hours ago
staffbox/staffbox-scorecards
View all activity

Organization Card

Staffbox.ai drafts your quotes from your own price list

A person checks each one before it goes out. No customer has run Staffbox yet, so the first step is a free test on your own quotes: we run it on 30 of your past quotes and show you the score, before you pay anything.

You forward a quote request to your quotes mailbox (or drop it in a shared folder). A draft reply lands in your drafts folder, usually within minutes, and each priced line cites the price-list row and that list's date. When a request isn't covered by your price list, it is tested to decline rather than guess: 6 of every 40 test questions are not in the company's notes, and an invented answer counts as a miss. Staffbox has no sending code, and at install it gets no permission to send email or delete files. A person sends.

Steps 1 and 2 below need no install and are run by us directly (hadi@staffbox.ai). For a pilot, the Mac mini sits on your own network; Staffbox is sold and supported through managed IT providers (first partner conversations under way). Cloud AI is off by default; if you turn it on with your own key, that provider's terms apply. Staffbox does not train models on your data.

For distributors, fabricators, cabinet and millwork shops, and IT providers that send at least 8 quotes a week from a price list. Tested so far on quotes priced from a price list by item, quantity, volume discount and setup.

How to start
1. Free 3-request test Two minutes: email three requests your team answers every week. You get back what Staffbox would have drafted, misses included.
2. Free score on 30 of your past quotes Free. Send 30 past quote requests and your price list. We run them on our Mac mini in Atlanta, your team grades 10 of the answers blind, and we delete the files afterwards. Nothing is installed at your site for this step. Security and data handling: trust.staffbox.ai (no SOC 2 or other certification yet).
3. Founding-site pilot Free for 30 days, then $595 a month for as long as the box is live, no onboarding fee, end any time. The Mac mini is included, and after the first paid month the box and your notes are yours to keep. Each month we re-test on your own quotes, and your IT provider applies updates after we confirm the tests still pass. First 100 sites; founding price, may change. Standard rate from unit 101: $695 a month plus $1,800 onboarding.

→ staffbox.ai · hadi@staffbox.ai · +1 404 493 6248 (Hadi Irvani, founder, Atlanta)

Not a shop owner? - Know a distributor, fabricator or cabinet shop that sends 8+ quotes a week? Forward this page; the 3-request test is free. - Run an IT firm? Staffbox is sold and supported through managed IT providers (first partner conversations under way); the IT provider applies monthly updates after we confirm the tests still pass. Email the number of your clients that send 8+ quotes a week: staffbox.ai/partners · email, subject "MSP". - Tested another model on our sets? Run staffbox eval from Staffbox-ai/staffbox and open an issue with the output.

The proof, in public

Graded by machine against exact answers, misses kept. Fictional test companies.

In plain words: on our fictional test shops, the Mac mini got 40 of 40 and 39 of 40 questions exactly right, and 351 of 360 on a larger set of messy, machine-written requests, 1 to 3.2 seconds on average. No wrong totals in the latest 475 answers. On 5 Oct, 6 drafts had a wrong total, all the same way: the customer also asked about an item the shop doesn't stock, and Staffbox priced the nearest stocked item instead (a 1500VA battery backup when asked for 1000VA). We fixed that on 6 Oct: it now flags the unstocked item and never substitutes. Of the 11 misses that remain, 10 are refusals (it asks the customer rather than guess) and 1 is a routing miss. Every miss, before and after the fix, is published.

Example (fictional IT shop, from the held-out set of messy requests). Request: "Could you price a dozen small desktops with setup?" What it computed: 12 x DT-MINI at $899.00 = $10,788.00, less 4% = $10,356.48, plus setup 12 x $95.00 = $1,140.00; line total $11,496.48 (0.8 s, graded right).

Your real test is your own 30 past quotes, on your own price list.

Technical detail (16 GB Mac mini M4, qwen3:8b, one request at a time, model warm). 6 Oct, quote action v2.4: 40/40, 39/40, 19/20, 15/15, and 201/204 + 150/156 on 360 requests written by a local model on our test box (expected answers computed by code); 0 wrong totals in 475 answers. 5 Oct, v2.3: 345/360 with 6 wrong totals. 2 Oct: the two 40-question sets were used while building it. On messy, real-looking emails: 15 of 15 on a set committed before the changes it measured (the clean held-out result; 15 requests is a small sample), and 19 of 20 on a set used while fixing. Given the model alone, without the company's notes, it got 1 of 40; notes plus the quote tool took it to 40 of 40 (different machine, 30 Sep). Pilots will run on a 32 GB unit, measured on the same tests before day 0; only the 16 GB Mac mini is measured so far.

📋 Scorecard Every graded answer, misses first: open any run to read the request, the expected answer and what Staffbox wrote
🧪 staffbox-evals The test sets and the company notes they run on
📊 staffbox-scorecards Every graded answer from every published run
⏱ staffbox-bench Speed measurements on the Mac mini
🗂 Collection All of the above, plus the models tested

Open source (MIT): github.com/Staffbox-ai/staffbox · For AI assistants: llms.txt

models 0

None public yet