julia-1-zorch

A compact SDLC control-plane classifier that reads a JIRA ticket and decides who owns it, what happens next, which review gate applies, who executes, how risky it is, and whether a human must approve.

v2: robust to missing fields, casing, truncation and typos.

Held-out Unseen Perturbed Params License

Result. 100.00% decision accuracy on the held-out test split (base SupersonicLabs/Julia-1: 25.06%) and 100.00% on 5 unseen, de-duplicated suites (26,400 decisions). Under 13 input perturbations the mean accuracy is 99.0% (v1: 84.8%), and the worst case, only one content field kept, is 87.1% (v1: 57.6%).

What changed in v2

v1 scored 99.7 % on clean tickets but collapsed when the Gherkin criteria were removed (about 79 %), when text was upper-cased (about 70 %) or truncated (about 81 %). v2 is trained on the same six decisions with a larger, de-duplicated dataset and label-preserving input augmentation, so the model no longer depends on any single field or on surface formatting.

v1 v2
Training records 3,720 clean 28,560 (14,280 clean + 14,280 perturbed)
Augmentation none field dropout, casing, truncation, typos, whitespace
Clean held-out accuracy 99.74% 100.00%
Mean accuracy under perturbation 84.8% 99.0%
Worst-case perturbation 57.6% 87.1%

What it decides

Decision Question Label space
role_owner Which SDLC role should own the next unit of work? product_manager, business_analyst, solution_architect, ux_designer, frontend_engineer, backend_engineer, mobile_engineer, qa_engineer, devops_engineer, security_engineer
next_action What should happen next in the SDLC? clarify_requirements, architecture_review, design_ui, implement, write_or_update_tests, security_review, deploy_or_configure, monitor_or_observe, investigate_incident, release_review
review_gate Which review gate is required before this work can progress? none, peer_review, architecture_review, qa_review, security_review, product_acceptance, change_advisory
executor_profile Which executor profile should handle the next action? product_agent, analysis_agent, architecture_agent, coding_agent, qa_agent, devops_agent, security_agent, human_only
risk_level What is the delivery risk level of this change? low, medium, high, critical
human_required Is explicit human approval required before autonomous progression? no, yes

Results

Held-out test split

774 decisions from 129 scenarios not used for training or model selection.

Metric Fine-tuned Base Change
Accuracy 100.00% 25.06% +74.94 pts
Negative log-likelihood 0.0000 3.6520 -3.6520
Expected calibration error 0.0000 0.4740 -0.4740
Decision Fine-tuned Base
Role owner 100.00% 12.40%
Next action 100.00% 17.83%
Review gate 100.00% 17.05%
Executor profile 100.00% 15.50%
Risk level 100.00% 35.66%
Human approval 100.00% 51.94%

Unseen multi-seed suites

5 suites generated with seeds 101, 202, 303, 404, 505, none used in training. Each record was hashed and checked against the training, validation and test data first (0 exact duplicates found).

Suite Decisions Accuracy NLL ECE
seed_101 5,280 100.00% 0.0000 0.0000
seed_202 5,280 100.00% 0.0000 0.0000
seed_303 5,280 100.00% 0.0000 0.0000
seed_404 5,280 100.00% 0.0000 0.0000
seed_505 5,280 100.00% 0.0000 0.0000
Pooled 26,400 100.00%
Decision Pooled accuracy
Role owner 100.00%
Next action 100.00%
Review gate 100.00%
Executor profile 100.00%
Risk level 100.00%
Human approval 100.00%

Robustness to degraded input

One unseen suite (seed_505, 1,320 decisions) is re-scored after perturbing every ticket. Trained families are the augmentations used in training; held-out families were never seen, so they show how well the robustness transfers.

Perturbation Family v1 v2 Change
clean (reference) 99.92% 100.00% +0.08 pts
UPPERCASED trained 54.39% 100.00% +45.61 pts
only one content field kept trained 57.58% 87.12% +29.55 pts
no Gherkin / acceptance criteria trained 69.24% 100.00% +30.76 pts
Title Case trained 80.83% 100.00% +19.17 pts
text truncated (30-150 chars) trained 84.85% 100.00% +15.15 pts
one content field removed trained 86.14% 100.00% +13.86 pts
3% character typos trained 90.15% 99.85% +9.70 pts
lowercased trained 98.11% 100.00% +1.89 pts
1-3 metadata fields removed trained 99.47% 100.00% +0.53 pts
whitespace collapsed trained 99.62% 100.00% +0.38 pts
word dropout (20%) held-out 83.71% 99.70% +15.98 pts
shuffled lines held-out 98.26% 100.00% +1.74 pts
[URGENT] prefix held-out 99.47% 100.00% +0.53 pts

Confidence and selective prediction

Threshold Coverage Accuracy on covered
0.50 100.00% 100.00%
0.70 100.00% 100.00%
0.80 100.00% 100.00%
0.90 100.00% 100.00%
0.95 100.00% 100.00%

Suggested policy: auto-accept at confidence >= 0.90 and escalate the rest to a stronger judge or a human.

Limitations

  • Synthetic data. Training and every evaluation set come from one scenario generator. The numbers are in-distribution; they are not a claim about real JIRA projects. Validate on reviewed tickets from your own backlog before relying on the model.
  • Weakest case. only one content field kept is the hardest perturbation at 87.1%; tickets that lose that much information should be escalated.
  • Two ticket formats. Training covers Gherkin-style tickets (background, feature, gherkin) and classic JIRA tickets (summary, description, acceptance_criteria). Other layouts are untested.
  • Scope. English only, six fixed decisions, 1,024-token window. It is a classifier: it does not plan or explain.

Inference

The weights use a custom JuliaDecisionModel head on top of the base model. The inference code is part of the companion zorch project that produced this release; this repository holds the weights, configuration and evaluation artefacts. Input is a ticket state object plus the decision to ask; the output is a probability distribution over that decision's labels.

{
  "state": {
    "key": "CORE-1000",
    "product": "Mobile",
    "component": "web-shell",
    "feature": "Define success criteria for Mobile usage limits",
    "labels": ["web-shell", "github-actions", "medium"],
    "priority": "Medium",
    "technology_context": "GitHub Actions",
    "workflow_status": "Backlog",
    "background": "Component web-shell uses GitHub Actions; priority is Medium.",
    "gherkin": "Given multiple customer plans have different limits\nWhen the team prepares the next release\nThen the exact limit policy and success metric must be agreed before implementation"
  }
}

Training

Method Full fine-tune, all weights unfrozen
Trainable parameters 144,192,769 of 144,292,867 (99.93%)
Data 14,280 clean records (original train + two extra seeds) plus one perturbed copy of each; no overlap with any evaluation set
Augmentation drop Gherkin / acceptance criteria, drop a content field, keep a single field, drop 1-3 metadata fields, lower / UPPER / Title case, truncate to 30-150 characters, 3 % character typos, whitespace collapse (one or two per copy, labels unchanged)
Validation original validation split, clean plus one perturbed copy; checkpoint chosen on NLL
Epochs 2
Effective batch size 16 (8 x 2)
Learning rate 2e-05, warm-up 0.05, weight decay 0.01
Precision bf16
Wall-clock 9.8 min on 1 x NVIDIA RTX 4090 (24 GB)
Best validation NLL 0.0113
Seed 17

Files

File Purpose
model.safetensors, julia_config.json, encoder/, tokenizer/ Weights and configuration
eval/suite_metrics.json, eval/stress_compare.json Multi-seed and robustness results (v2 vs v1)
test_metrics.json, test_predictions.json Held-out metrics and per-record predictions
training_config.yaml, resolved_config.json, history.json, training_summary.json Training provenance
dataset_generation_report.json, dataset_lint_report.json, TAXONOMY.md Data provenance

Built on SupersonicLabs/Julia-1. Apache-2.0; the base model's terms also apply.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for niko0xdev/julia-1-zorch

Finetuned
(8)
this model

Evaluation results

  • Accuracy on ZOrch held-out test split (synthetic JIRA tickets)
    self-reported
    1.000
  • Pooled accuracy on ZOrch unseen multi-seed suites (5 suites)
    self-reported
    1.000
  • Mean accuracy under perturbation on ZOrch stress test, mean over 13 perturbations
    self-reported
    0.990