Bongard-mini
Bongard-mini is an open-weight System One model for machine intuition and Jev-like judgment, built on the T5 encoder–decoder architecture of T5Gemma 2 4B-4B. Give it text, structured data or images, along with questions and candidate answers or actions. It returns a probability distribution for each question without generating text.
Read once. Judge in parallel.
Try the live demo · Get started locally · Code & runtime · Design · Benchmarks · Examples
In the live demo's Playground, choose an example, edit the Input and the questions under Questions & settings, then click Run to get new probabilities. The demo runs on Hugging Face ZeroGPU: anonymous quota is limited, and signing in provides more quota. For local execution, use the quick start below. The game GIFs below and the examples on the website are recorded outputs; use the live demo to try your own input.
Bongard-mini is trained on judgments, semantic relationships and action outcomes. The same model can read a customer conversation, compare an invoice with an order, or judge an agent's next action.
Game recordings
| 2048 — 4096 tile, 55,140 final score | Tetris — consecutive line clears |
|---|---|
![]() |
![]() |
| Othello — 13–0 win against a positional opponent | Snake — grows to 32 cells on an 8×8 board |
![]() |
![]() |
These selected gameplay recordings show Bongard-mini choosing actions from candidate lists. For board games, the engine supplies legal moves and their computed features; 2048 and Tetris also include engine ratings of moves. A short instruction states the strategy, and the model selects each action. The scores describe these selected recordings with the supplied engine features.
Doom — Defend the Center
The model receives a text description of the scene every 0.2 seconds and chooses turn_left,
turn_right or fire. Across 32 episodes capped at 30 seconds, it averaged 10.25 kills, with a best
of 20. The GIF shows an excerpt, with probabilities beside the action.
Intuition is learned
An experienced player sees a promising move. A reader catches irony. Such judgments draw on learned relationships, often before the person can explain each step. Their value extends beyond speed: a whole pattern can carry meaning that is hard to express as a list of reasons.
Bongard treats this form of judgment as an independent capability to design and train.
| Stage | What the model learns |
|---|---|
| Judgment | Judge text, records and images. Learn which changes should alter an answer and which should leave it unchanged. |
| Relationships | Connect judgments to their meaning across words, images and states through joint-embedding training. |
| Outcomes | Learn what follows an action. Sandbox rollouts and exact solvers supply outcome distributions for candidate actions. |
All three stages are complete. They update 7.09 billion trainable parameters, including the text backbone, multimodal projector and judgment head. The vision tower stays frozen. Each stage ran on one GPU: a B200 for stages 1 and 2, and a B300 for stage 3.
One reading, many decisions
Bongard-mini uses the T5 encoder–decoder structure of T5Gemma 2 4B-4B, with 7.51 billion total parameters. The encoder reads the evidence in both directions, with the question instructions in view. Separate decoder branches share that evidence. A trained head scores each candidate from its name and description. The runtime reads the state once and processes questions in groups of up to eight.
This structure serves a concrete purpose. A later clause can change the meaning of an earlier one; a log entry can explain an earlier failure. Bidirectional encoding lets those facts shape each other's representation before the model makes several judgments about the situation.
A wrapper changes the interface. Bongard trains the judgment.
OpenJev's default implementation wraps pretrained DiffusionGemma and reads answer-token probabilities. Kev trains LoRA adapters and a pointer head on Qwen, and reuses a causal state cache across questions. Bongard develops a shared, bidirectional evidence representation through full-model training on judgments, semantic relationships and action outcomes.
Jev is a trained, closed System One model. Bongard provides an open T5 route to the same class of judgments. The results below show how their strengths differ across tasks and measures. Full-model training costs more than adapter training; Bongard uses it to shape both the evidence representation and the judgment function.
Quick start
Use Python 3.12 or later. Install the runtime from GitHub and download the complete model bundle:
git clone --branch main --depth 1 https://github.com/AgentBull/bongard.git
cd bongard
python -m pip install .
hf download AgentBull/bongard-mini --local-dir bongard-mini
The examples below use CUDA. Use device="mps" on Apple Silicon or device="cpu" for CPU execution;
the CLI accepts the same values with --device. The runtime loads the BF16 weights, with additional
memory needed for the input and candidate descriptions.
Python: three questions, one state
import json
from bongard.inference import Predictor
predictor = Predictor.load(
"bongard-mini",
device="cuda",
temperatures="bongard-mini/temperatures.json",
)
request = {
"state": {
"message": "My order arrived yesterday with a broken screen. Please replace it.",
"policy": "Replace items damaged on delivery if reported within 14 days."
},
"questions": {
"replace": {
"type": "noul",
"instructions": "Is this request eligible for a replacement under the policy?"
},
"route": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"returns": "Damaged items, returns and replacements.",
"billing": "Payment failures and invoice questions.",
"sales": "New orders and product recommendations."
}
},
"urgency": {
"type": "score",
"instructions": "How urgently does the customer need a response?",
"criteria": [
"Routine: can wait several days.",
"Important: respond within one business day.",
"Critical: immediate assistance is needed."
]
}
}
}
answers = predictor.predict(request)["answers"]
print(json.dumps(answers, indent=2))
print(answers["route"]["choice"])
Each entry in answers has the question ID you supplied:
| Type | How to define it | How to read the answer |
|---|---|---|
noul |
A yes/no question; optional descriptions of true and false |
noul is the probability of true, from 0 to 1. |
choice |
A map of 2–255 candidate names to descriptions | choice is the most likely candidate; probabilities contains the distribution over all candidates. |
score |
A list of 2–10 ordered level descriptions, lowest first | score is the expected zero-based level; probabilities uses keys "0", "1", etc.; legend maps them to descriptions. |
Keep shared facts in state and put each decision in questions so the model can reuse its reading
of the state. Use the supplied temperatures.json in both Python and the CLI to apply probability
calibration. Loading only backbone/ with Transformers does not load the complete judgment model.
CLI and HTTP API
The GitHub repository includes a ready-to-run example request:
bongard predict --checkpoint bongard-mini --device cuda \
--temperatures bongard-mini/temperatures.json \
--request examples/request.json
To serve the model locally:
bongard serve --checkpoint bongard-mini --device cuda \
--temperatures bongard-mini/temperatures.json \
--host 127.0.0.1 --port 8000
In another terminal, from the same repository directory:
curl http://127.0.0.1:8000/v1/models
curl http://127.0.0.1:8000/v1/systemone \
-H 'Content-Type: application/json' \
--data-binary @examples/request.json
POST /v1/systemone uses the same request structure as Python and returns model, answers
and usage. The API follows the System One wire format; usage.output_tokens is zero.
Using images
Python example
Add a top-level images array of base64 data URLs and refer to them as image 1, image 2, etc.
PNG, JPEG and static WebP are supported. Local paths and remote URLs must be read and encoded first.
For example, with your own package.png and the predictor loaded above:
import base64
from pathlib import Path
encoded = base64.b64encode(Path("package.png").read_bytes()).decode("ascii")
response = predictor.predict({
"state": "A customer submitted a photo of their delivery.",
"images": [f"data:image/png;base64,{encoded}"],
"questions": {
"damaged": {
"type": "noul",
"instructions": "Does image 1 show visible damage to the package?"
}
}
})
print(response["answers"]["damaged"]["noul"])
Benchmarks
DecisionBench was evaluated with its official runner on 2026-09-30. The other release evaluations below are from 2026-09-29. Measurements use one NVIDIA RTX PRO 6000 in BF16, with the splits and protocols stated below.
| Benchmark | Evaluation scope | Bongard-mini |
|---|---|---|
| DecisionBench 1.0 | Official runner; eval split, all 23,900 decisions, 43 tasks, 2–255 candidates |
78.05% accuracy; 77.73% on the 20,959 rows with no detected source-text overlap with training data |
| JevBench v1.4 | 231 public decisions: easy / standard / hard | 100% / 97.2% / 57.7% accuracy |
| typed-decisions | Test split: 400 cases, 2,000 decisions; zero-shot workflows | 59.4% accuracy, 0.518 soft accuracy, 0.256 KL, 0.132 Brier |
| ImajevBench v2.0-lite | Dev + calibration: 254 items with public gold | 66.9% accuracy; chance 26.5% |
| behavior-benchmark | Every trace; present / absent / not observable | 0.458 / 0.668 F1 for present, core / multilingual |
| Fair-chance probes | 18 kinds, including coins, dice, lotteries and tie-breaks | 0.0009 mean KL from the uniform distribution; lower is better |
How to read these results: DecisionBench uses its own runner
(jobs/run_system_one_http_eval.py at 9a6328e) against bongard serve at ea57853, through
the System One HTTP contract, one request at a time. Serving limits were 262,144 tokens per
request and 131,072 tokens per sequence, with up to 255 candidates. All 23,900 rows
returned answers, with no errors or input truncation. The
official run, summary and manifest
are public. Of these rows, 2,941 have detected source-text overlap with training data; the remaining
20,959 score 77.73%.
An earlier custom-harness run remains in the evaluation repository. It sent 3,042 binary rows as two-option Choice questions, whereas the official runner sends Noul questions. Its scores belong to that earlier protocol and are not the headline result above.
JevBench scores cover its public questions. For typed-decisions, the gold is a teacher probability
distribution, so its metrics measure agreement with that teacher. ImajevBench uses direct option
scoring with an unknown option and averages all cyclic option orders.
DecisionBench comparison
| Selected systems | Accuracy |
|---|---|
| Imajev-4B | 79.7% |
| Bongard-mini | 78.05% |
| Winnow-12B | 76.7% |
| Jev 1.13 | 72.0% |
| DeepSeek V4.1 Flash | 71.0% |
| GPT-5.6 Luna | 69.9% |
Bongard-mini ranks fourth among the 61 systems in the full comparison. Reference scores are a 2026-09-29 snapshot of the DecisionBench leaderboard. Bongard-mini is our 2026-09-30 official-runner evaluation, with the serving limits and source-text overlap qualifications above. This table lists selected systems, rather than every leaderboard entry.
Bongard and Jev, by task and measure
| Evaluation | Bongard-mini | Jev 1.13 |
|---|---|---|
| DecisionBench accuracy, official runner, all 23,900 decisions | 78.05% | 72.0% |
| typed-decisions top-label agreement, 2,000 decisions | 59.4% | 72.7% |
| typed-decisions KL from the teacher, lower is better | 0.256 | 1.442 |
| typed-decisions Brier score, lower is better | 0.132 | 0.148 |
Jev matches the teacher's top label more often on typed-decisions. Bongard's full distributions are closer to the teacher's under KL and Brier. These measures establish agreement with the teacher; calibration against observed outcomes needs a separate evaluation. Reference scores come from the dataset cards and leaderboards on 2026-09-29. Bongard's per-item predictions let you inspect the decisions behind the scores.
Inference speed
| Workload | Measured result |
|---|---|
| Short JevBench easy/standard requests, one at a time (120 requests) | 36.1 ms median, 39.3 ms p95 per decision |
| 32 questions about one shared state | 221 ms total, about 6.9 ms per decision |
| Independent JevBench requests, batch size 64 | 112.1 decisions/s, about 403,560 decisions/hour |
These are local inference measurements and exclude network overhead. The batched result comes from batched model evaluation; the supplied HTTP server processes one request at a time. Input length, candidate descriptions and hardware affect latency.
To evaluate your server on typed-decisions, use the included benchmark runner from the repository root while the server is running:
python benchmarks/typed_decisions.py \
--endpoint http://127.0.0.1:8000 --output typed-results.json
Judgments across situations
These are selected test scenarios from typed-decisions, with Bongard's published predictions. Each row shows one question from a five-question request.
| Evidence | Judgment | Model output |
|---|---|---|
| A customer reports receiving the wrong item and asks what to do next. | What is the conversation about? | delivery: 71.0% |
| An order lists ten units; five were delivered and invoiced. | Does the invoice reconcile with the order and delivery? | false: 77.7% |
| An agent's certificate-rotation trace has eleven steps and no reported tool errors or constraint violations. | What should the observability system do? | continue: 66.9% |
Evaluation notes
- Worked-answer checking: in the evaluated tasks, the model rarely rejected incorrect worked solutions.
- Calibration: use the supplied temperatures and evaluate probabilities on representative data. Refit calibration for a new deployment domain when needed.
- Candidate order: Choice probabilities can depend on option order.
rotations=TrueinPredictor.load, or--rotationsin the CLI, averages cyclic option orders at extra compute cost; it is off by default.
Model and license
This is the final three-stage Bongard-mini checkpoint. The public GitHub repository contains the inference runtime, HTTP server, supervised training loop and a typed-decisions benchmark runner. The training data, joint-embedding objectives, sandbox environments and full training recipes are not part of the code release.
Model weights are subject to the Gemma Terms of Use; see NOTICE. The runtime code is Apache-2.0.
Cite
@techreport{ding2026bongard,
title = {Bongard: Training Machine Intuition},
author = {Ding, Li and Jin, Haidi and Ji, Chen},
institution = {AgentBull Pte Ltd},
year = {2026}
}
- Downloads last month
- -
Model tree for AgentBull/bongard-mini
Base model
google/t5gemma-2-4b-4bSpace using AgentBull/bongard-mini 1
Evaluation results
- Accuracy (%, official runner; all 23,900 rows, no errors or truncation) on DecisionBench 1.0 (23,900 decisions)AgentBull official-runner evaluation (2026-09-30)78.050
- Accuracy (%, official runner; 20,959 rows with no detected source-text overlap) on DecisionBench 1.0 (23,900 decisions)AgentBull official-runner evaluation (2026-09-30)77.730
- Top-label agreement with teacher (%) on typed-decisions (400 cases, 2,000 decisions)test set AgentBull release evaluation (2026-09-29)59.400
- Soft accuracy against teacher distribution on typed-decisions (400 cases, 2,000 decisions)test set AgentBull release evaluation (2026-09-29)0.518
- KL from teacher distribution (lower is better) on typed-decisions (400 cases, 2,000 decisions)test set AgentBull release evaluation (2026-09-29)0.256
- Brier score against teacher distribution (lower is better) on typed-decisions (400 cases, 2,000 decisions)test set AgentBull release evaluation (2026-09-29)0.132
- Accuracy (%, unknown option; all cyclic option orders averaged) on ImajevBench v2.0-lite (dev + calibration, 254 public-gold items)AgentBull release evaluation (2026-09-29)66.900




