Kodama Vision
Experimental image-and-text decision model combining a ModernBERT-base decision model, frozen SigLIP, and a trained image adapter. Formerly jevlike-mm-proto. The frozen encoder and its tokenizer are included under vision/.
All models return structured Choice, Score, and Noul decisions. They score answer options directly; they do not generate conversational responses. These are custom jevlike checkpoints, loaded with the bundled runtime below. Generic Transformers pipeline() / AutoModel loading is not the supported entry point.
Which Kodama model should I use?
| Model | Inputs | Typical reason to choose it |
|---|---|---|
| Kodama Core | Text; JSON serialized as text | Fast routing, classification, and small decisions inside an application |
| Kodama Vision | One image plus optional text | Experimental object, vehicle, or document-type classification |
| Kodama Sense | Text, images, PDFs, charts, audio, video, and mixed media | Decisions that require visual or audio evidence |
These models choose from the answers you supply. They are not conversational chatbots, open-ended report writers, transcription services, or general document extraction systems. The use cases below are applications to evaluate on your own examples; the benchmark results do not establish reliable performance on every such application.
Practical use cases
| Application | Example question | Important scope |
|---|---|---|
| Image-library categorization | “Which object category is shown?” | Classify among the labels you provide |
| Vehicle-photo experiments | “Which vehicle category best fits this image?” | Strong local car results do not establish accuracy on every vehicle collection |
| Document-type triage | “Is this an invoice, letter, form, or report?” | Coarse page appearance; small printed text may be unreadable |
| Visual checks in a workflow | “Does the picture show a red object?” | One image and a structured question |
Vision is an experimental specialist. Use Sense for native PDFs, chart interpretation, audio, video, or combined modalities. Vision resizes each image to 224×224 pixels; it is not an OCR engine or a document-field extraction model. In the small controlled chart/PDF comparison, Vision scored 0/8 in each group, so chart arithmetic is not a supported performance claim for this checkpoint.
Install and download
Tested environment: Python 3.12, PyTorch 2.10, Transformers 5.2.0, Linux, and an NVIDIA RTX 4050 Laptop GPU with 6 GB VRAM. Python 3.10+ is supported by the package. Install a CUDA-enabled PyTorch build appropriate for your system when using the GPU examples. Approximate model download size: 1.48 GB, excluding Python/CUDA dependencies and caches.
python -m pip install huggingface_hub
# For a private repository, authenticate once with: hf auth login
hf download Cem13/kodama-vision --local-dir kodama-vision
python -m pip install "./kodama-vision/runtime[model,ingest]" "transformers==5.2.0"
Run the examples from the directory containing kodama-vision. After the download and
dependency installation, prediction runs locally; no API key or paid inference service
is needed for prediction. Private downloads require access to the repository. An existing
HF_TOKEN can authenticate the Hub client without adding a token to source code.
Python: classify an image
from PIL import Image
from jevlike import Choice, Noul
from jevlike.mm import MultimodalSystemOne
model = MultimodalSystemOne.load(
"kodama-vision", device="cuda", vision="kodama-vision/vision"
)
questions = {
"color": Choice("Which color dominates the picture?", ["red", "green", "blue"]),
"red": Noul("The picture is predominantly red."),
}
image = Image.new("RGB", (224, 224), "red") # Self-contained demonstration.
result = model.predict(None, questions, image=image)
print(result["answers"])
# For your own input: model.predict("Optional context", questions, image="photo.jpg")
Pass the bundled vision/ path to use the included frozen SigLIP weights and tokenizer.
The image argument accepts a local image path, PIL image, or an RGB NumPy array.
None is a valid state when an image provides the evidence. A runnable example is included:
python kodama-vision/examples/quickstart.py
Multiple images and text-only calls
results = model.predict_batch([
(None, questions, "first.jpg"),
(None, questions, "second.jpg"),
])
Each tuple supplies one state, its questions, and one image. A call with no image uses
the underlying text model. For image calls, long text is truncated to the available
512-token sequence budget; image slots, instructions, and options also consume that budget.
Use concise context. CPU execution is available with device="cpu", but has not been
speed-benchmarked. This release does not expose a dedicated Vision image-upload HTTP server.
Understanding the answers
Every response contains model, latency_ms, and answers, keyed by your question names.
| Question type | How to define it | Main output |
|---|---|---|
Choice |
2–255 distinct labels, optionally with descriptions | choice and a label-to-probability dictionary |
Score |
2–10 ordered descriptions, lowest first | Zero-based level, its label, a probability list, and expected score scaled to 0–10 |
Noul |
A statement to test; no answer options | noul, the probability that the statement is true |
confidence is the largest answer probability. For Noul, a confident false answer
can have noul=0.02 and confidence=0.98; use noul when deciding whether the statement
is true. escalate and escalate_prob help flag uncertain answers for another step in
your application. Neither confidence nor escalation is a guarantee of correctness.
Use specific instructions and descriptive, distinct options. Probabilities are conditional
on the supplied options; adding an other option can be useful when your label set does
not cover every input, but the model's use of that option still needs validation.
Core and Vision apply saved temperature calibration by default. Their escalate_prob
comes from a separately trained escalation head with saved calibration; default escalation
is above 0.5. Set calibrated=False on prediction calls to inspect uncalibrated outputs.
Their answer dictionaries do not include Sense's per-answer calibrated flag.
Local evaluation
Measured on an NVIDIA RTX 4050 Laptop GPU (6 GB). Latency includes tokenization and excludes model loading and HTTP.
| Evaluation split | Decisions | Accuracy |
|---|---|---|
| Caltech101 objects | 1,371 | 92.85% |
| Cars | 1,153 | 97.66% |
| RVL, eight unseen document types | 1,675 | 58.99% |
| In-domain validation | 3,407 | 88.02% |
See evaluation/vision_full.json for exact metrics and split details. In-domain
validation is not an unseen test. The smaller shared comparison used 48 object,
48 car, and 64 document examples; those sample results should not be confused
with the full evaluation above.
Timing details and interpretation
Vision does not have a dedicated repeated-call latency benchmark comparable to the Core/Sense timing suites. In the earlier small shared evaluation, per-decision medians were 87.5 ms for objects (48 decisions), 72.4 ms for cars (48), and 46.5 ms for document types (64). These are evaluation-run timings, not controlled throughput or p95 measurements. No batch-throughput claim is made for Vision.
A separate packaged-model smoke check used 1.28 GiB peak PyTorch GPU allocation while loading and scoring a simple image. That is one smoke workload, not a worst-case memory bound. The frozen SigLIP branch uses BF16 on CUDA; the decision model uses float32 weights with BF16 autocast. Image decoding, size, option counts, and batch lengths affect latency.
The full accuracy table mixes held-out object/car/document splits with an explicitly
labelled in-domain validation split. It is not directly comparable to Sense's much smaller
shared image subset. On that shared car subset Vision scored 93.75% and Sense 83.33%, each
on 48 examples; on document types Sense scored 78.13% and Vision 57.81%, each on 64.
See evaluation/vision_full.json and
evaluation/shared_image_comparison.json.
In the shared comparison, existing means Vision for image groups and Core for text groups.
The memory smoke measurement is saved in evaluation/package_smoke.json.
Training and limitations
The base encoders were pretrained by their upstream authors. Kodama adds local training/adaptation; these foundation models were not pretrained from scratch. Included training metadata describes the local run. Evaluation data are not redistributed in this repository. Public pretraining overlap is unknown, and small benchmark samples should not be treated as broad capability guarantees. Accuracy depends on task and option wording. Confidence is not a guarantee of correctness; use suitable validation before relying on decisions.
Attribution and licensing
The pretrained base components identify their license as Apache-2.0. Upstream cards, source revisions, and the Apache license text are preserved in third_party/ and NOTICE. No separate public license has been selected for the original Kodama additions and runtime in this release; the upstream license statements should not be read as a new license grant for those additions.
Base sources:
Troubleshooting and reproducibility
- 401/403 while downloading: authenticate with an account allowed to read the repository.
- Model configuration / AutoModel error: use the bundled
jevlikeruntime and the custom loader shown above. - Import or tokenizer mismatch: use the documented Transformers version and install the runtime extras for this model.
- Different answers near a tie: numerical precision and batch shapes can affect probabilities; measure your own task before changing inference settings.
release_manifest.json records file hashes and pinned base revisions. The runtime is
bundled under runtime/; examples/quickstart.py provides a complete local example.
The pretrained foundations were not trained from scratch by Kodama. Earlier evaluation
artifacts use jevlike for Core, omni for Sense, and prototype for Vision.
- Downloads last month
- -
Model tree for Cem13/kodama-vision
Base model
answerdotai/ModernBERT-base