julia-1-zorch
A compact SDLC control-plane classifier that reads a JIRA ticket and decides who owns it, what happens next, which review gate applies, who executes, how risky it is, and whether a human must approve.
v2: robust to missing fields, casing, truncation and typos.
Result. 100.00% decision accuracy on the held-out test split (base
SupersonicLabs/Julia-1: 25.06%) and 100.00% on 5 unseen, de-duplicated suites (26,400 decisions). Under 13 input perturbations the mean accuracy is 99.0% (v1: 84.8%), and the worst case, only one content field kept, is 87.1% (v1: 57.6%).
What changed in v2
v1 scored 99.7 % on clean tickets but collapsed when the Gherkin criteria were removed (about 79 %), when text was upper-cased (about 70 %) or truncated (about 81 %). v2 is trained on the same six decisions with a larger, de-duplicated dataset and label-preserving input augmentation, so the model no longer depends on any single field or on surface formatting.
| v1 | v2 | |
|---|---|---|
| Training records | 3,720 clean | 28,560 (14,280 clean + 14,280 perturbed) |
| Augmentation | none | field dropout, casing, truncation, typos, whitespace |
| Clean held-out accuracy | 99.74% | 100.00% |
| Mean accuracy under perturbation | 84.8% | 99.0% |
| Worst-case perturbation | 57.6% | 87.1% |
What it decides
| Decision | Question | Label space |
|---|---|---|
role_owner |
Which SDLC role should own the next unit of work? | product_manager, business_analyst, solution_architect, ux_designer, frontend_engineer, backend_engineer, mobile_engineer, qa_engineer, devops_engineer, security_engineer |
next_action |
What should happen next in the SDLC? | clarify_requirements, architecture_review, design_ui, implement, write_or_update_tests, security_review, deploy_or_configure, monitor_or_observe, investigate_incident, release_review |
review_gate |
Which review gate is required before this work can progress? | none, peer_review, architecture_review, qa_review, security_review, product_acceptance, change_advisory |
executor_profile |
Which executor profile should handle the next action? | product_agent, analysis_agent, architecture_agent, coding_agent, qa_agent, devops_agent, security_agent, human_only |
risk_level |
What is the delivery risk level of this change? | low, medium, high, critical |
human_required |
Is explicit human approval required before autonomous progression? | no, yes |
Results
Held-out test split
774 decisions from 129 scenarios not used for training or model selection.
| Metric | Fine-tuned | Base | Change |
|---|---|---|---|
| Accuracy | 100.00% | 25.06% | +74.94 pts |
| Negative log-likelihood | 0.0000 | 3.6520 | -3.6520 |
| Expected calibration error | 0.0000 | 0.4740 | -0.4740 |
| Decision | Fine-tuned | Base |
|---|---|---|
| Role owner | 100.00% | 12.40% |
| Next action | 100.00% | 17.83% |
| Review gate | 100.00% | 17.05% |
| Executor profile | 100.00% | 15.50% |
| Risk level | 100.00% | 35.66% |
| Human approval | 100.00% | 51.94% |
Unseen multi-seed suites
5 suites generated with seeds 101, 202, 303, 404, 505, none used in training. Each record was hashed and checked against the training, validation and test data first (0 exact duplicates found).
| Suite | Decisions | Accuracy | NLL | ECE |
|---|---|---|---|---|
seed_101 |
5,280 | 100.00% | 0.0000 | 0.0000 |
seed_202 |
5,280 | 100.00% | 0.0000 | 0.0000 |
seed_303 |
5,280 | 100.00% | 0.0000 | 0.0000 |
seed_404 |
5,280 | 100.00% | 0.0000 | 0.0000 |
seed_505 |
5,280 | 100.00% | 0.0000 | 0.0000 |
| Pooled | 26,400 | 100.00% |
| Decision | Pooled accuracy |
|---|---|
| Role owner | 100.00% |
| Next action | 100.00% |
| Review gate | 100.00% |
| Executor profile | 100.00% |
| Risk level | 100.00% |
| Human approval | 100.00% |
Robustness to degraded input
One unseen suite (seed_505, 1,320 decisions) is re-scored after perturbing every ticket. Trained families are the augmentations used in training; held-out families were never seen, so they show how well the robustness transfers.
| Perturbation | Family | v1 | v2 | Change |
|---|---|---|---|---|
| clean (reference) | 99.92% | 100.00% | +0.08 pts | |
| UPPERCASED | trained | 54.39% | 100.00% | +45.61 pts |
| only one content field kept | trained | 57.58% | 87.12% | +29.55 pts |
| no Gherkin / acceptance criteria | trained | 69.24% | 100.00% | +30.76 pts |
| Title Case | trained | 80.83% | 100.00% | +19.17 pts |
| text truncated (30-150 chars) | trained | 84.85% | 100.00% | +15.15 pts |
| one content field removed | trained | 86.14% | 100.00% | +13.86 pts |
| 3% character typos | trained | 90.15% | 99.85% | +9.70 pts |
| lowercased | trained | 98.11% | 100.00% | +1.89 pts |
| 1-3 metadata fields removed | trained | 99.47% | 100.00% | +0.53 pts |
| whitespace collapsed | trained | 99.62% | 100.00% | +0.38 pts |
| word dropout (20%) | held-out | 83.71% | 99.70% | +15.98 pts |
| shuffled lines | held-out | 98.26% | 100.00% | +1.74 pts |
| [URGENT] prefix | held-out | 99.47% | 100.00% | +0.53 pts |
Confidence and selective prediction
| Threshold | Coverage | Accuracy on covered |
|---|---|---|
| 0.50 | 100.00% | 100.00% |
| 0.70 | 100.00% | 100.00% |
| 0.80 | 100.00% | 100.00% |
| 0.90 | 100.00% | 100.00% |
| 0.95 | 100.00% | 100.00% |
Suggested policy: auto-accept at confidence >= 0.90 and escalate the rest to a stronger judge or a human.
Limitations
- Synthetic data. Training and every evaluation set come from one scenario generator. The numbers are in-distribution; they are not a claim about real JIRA projects. Validate on reviewed tickets from your own backlog before relying on the model.
- Weakest case. only one content field kept is the hardest perturbation at 87.1%; tickets that lose that much information should be escalated.
- Two ticket formats. Training covers Gherkin-style tickets (
background,feature,gherkin) and classic JIRA tickets (summary,description,acceptance_criteria). Other layouts are untested. - Scope. English only, six fixed decisions, 1,024-token window. It is a classifier: it does not plan or explain.
Inference
The weights use a custom JuliaDecisionModel head on top of the base model. The inference code is part of the companion zorch project that produced this release; this repository holds the weights, configuration and evaluation artefacts. Input is a ticket state object plus the decision to ask; the output is a probability distribution over that decision's labels.
{
"state": {
"key": "CORE-1000",
"product": "Mobile",
"component": "web-shell",
"feature": "Define success criteria for Mobile usage limits",
"labels": ["web-shell", "github-actions", "medium"],
"priority": "Medium",
"technology_context": "GitHub Actions",
"workflow_status": "Backlog",
"background": "Component web-shell uses GitHub Actions; priority is Medium.",
"gherkin": "Given multiple customer plans have different limits\nWhen the team prepares the next release\nThen the exact limit policy and success metric must be agreed before implementation"
}
}
Training
| Method | Full fine-tune, all weights unfrozen |
| Trainable parameters | 144,192,769 of 144,292,867 (99.93%) |
| Data | 14,280 clean records (original train + two extra seeds) plus one perturbed copy of each; no overlap with any evaluation set |
| Augmentation | drop Gherkin / acceptance criteria, drop a content field, keep a single field, drop 1-3 metadata fields, lower / UPPER / Title case, truncate to 30-150 characters, 3 % character typos, whitespace collapse (one or two per copy, labels unchanged) |
| Validation | original validation split, clean plus one perturbed copy; checkpoint chosen on NLL |
| Epochs | 2 |
| Effective batch size | 16 (8 x 2) |
| Learning rate | 2e-05, warm-up 0.05, weight decay 0.01 |
| Precision | bf16 |
| Wall-clock | 9.8 min on 1 x NVIDIA RTX 4090 (24 GB) |
| Best validation NLL | 0.0113 |
| Seed | 17 |
Files
| File | Purpose |
|---|---|
model.safetensors, julia_config.json, encoder/, tokenizer/ |
Weights and configuration |
eval/suite_metrics.json, eval/stress_compare.json |
Multi-seed and robustness results (v2 vs v1) |
test_metrics.json, test_predictions.json |
Held-out metrics and per-record predictions |
training_config.yaml, resolved_config.json, history.json, training_summary.json |
Training provenance |
dataset_generation_report.json, dataset_lint_report.json, TAXONOMY.md |
Data provenance |
Built on SupersonicLabs/Julia-1. Apache-2.0; the base model's terms also apply.
Model tree for niko0xdev/julia-1-zorch
Evaluation results
- Accuracy on ZOrch held-out test split (synthetic JIRA tickets)self-reported1.000
- Pooled accuracy on ZOrch unseen multi-seed suites (5 suites)self-reported1.000
- Mean accuracy under perturbation on ZOrch stress test, mean over 13 perturbationsself-reported0.990