Kodama Vision

Experimental image-and-text decision model combining a ModernBERT-base decision model, frozen SigLIP, and a trained image adapter. Formerly jevlike-mm-proto. The frozen encoder and its tokenizer are included under vision/.

All models return structured Choice, Score, and Noul decisions. They score answer options directly; they do not generate conversational responses. These are custom jevlike checkpoints, loaded with the bundled runtime below. Generic Transformers pipeline() / AutoModel loading is not the supported entry point.

Which Kodama model should I use?

Model Inputs Typical reason to choose it
Kodama Core Text; JSON serialized as text Fast routing, classification, and small decisions inside an application
Kodama Vision One image plus optional text Experimental object, vehicle, or document-type classification
Kodama Sense Text, images, PDFs, charts, audio, video, and mixed media Decisions that require visual or audio evidence

These models choose from the answers you supply. They are not conversational chatbots, open-ended report writers, transcription services, or general document extraction systems. The use cases below are applications to evaluate on your own examples; the benchmark results do not establish reliable performance on every such application.

Practical use cases

Application Example question Important scope
Image-library categorization “Which object category is shown?” Classify among the labels you provide
Vehicle-photo experiments “Which vehicle category best fits this image?” Strong local car results do not establish accuracy on every vehicle collection
Document-type triage “Is this an invoice, letter, form, or report?” Coarse page appearance; small printed text may be unreadable
Visual checks in a workflow “Does the picture show a red object?” One image and a structured question

Vision is an experimental specialist. Use Sense for native PDFs, chart interpretation, audio, video, or combined modalities. Vision resizes each image to 224×224 pixels; it is not an OCR engine or a document-field extraction model. In the small controlled chart/PDF comparison, Vision scored 0/8 in each group, so chart arithmetic is not a supported performance claim for this checkpoint.

Install and download

Tested environment: Python 3.12, PyTorch 2.10, Transformers 5.2.0, Linux, and an NVIDIA RTX 4050 Laptop GPU with 6 GB VRAM. Python 3.10+ is supported by the package. Install a CUDA-enabled PyTorch build appropriate for your system when using the GPU examples. Approximate model download size: 1.48 GB, excluding Python/CUDA dependencies and caches.

python -m pip install huggingface_hub
# For a private repository, authenticate once with: hf auth login
hf download Cem13/kodama-vision --local-dir kodama-vision
python -m pip install "./kodama-vision/runtime[model,ingest]" "transformers==5.2.0"

Run the examples from the directory containing kodama-vision. After the download and dependency installation, prediction runs locally; no API key or paid inference service is needed for prediction. Private downloads require access to the repository. An existing HF_TOKEN can authenticate the Hub client without adding a token to source code.

Python: classify an image

from PIL import Image
from jevlike import Choice, Noul
from jevlike.mm import MultimodalSystemOne

model = MultimodalSystemOne.load(
    "kodama-vision", device="cuda", vision="kodama-vision/vision"
)
questions = {
    "color": Choice("Which color dominates the picture?", ["red", "green", "blue"]),
    "red": Noul("The picture is predominantly red."),
}
image = Image.new("RGB", (224, 224), "red")  # Self-contained demonstration.
result = model.predict(None, questions, image=image)
print(result["answers"])
# For your own input: model.predict("Optional context", questions, image="photo.jpg")

Pass the bundled vision/ path to use the included frozen SigLIP weights and tokenizer. The image argument accepts a local image path, PIL image, or an RGB NumPy array. None is a valid state when an image provides the evidence. A runnable example is included:

python kodama-vision/examples/quickstart.py

Multiple images and text-only calls

results = model.predict_batch([
    (None, questions, "first.jpg"),
    (None, questions, "second.jpg"),
])

Each tuple supplies one state, its questions, and one image. A call with no image uses the underlying text model. For image calls, long text is truncated to the available 512-token sequence budget; image slots, instructions, and options also consume that budget. Use concise context. CPU execution is available with device="cpu", but has not been speed-benchmarked. This release does not expose a dedicated Vision image-upload HTTP server.

Understanding the answers

Every response contains model, latency_ms, and answers, keyed by your question names.

Question type How to define it Main output
Choice 2–255 distinct labels, optionally with descriptions choice and a label-to-probability dictionary
Score 2–10 ordered descriptions, lowest first Zero-based level, its label, a probability list, and expected score scaled to 0–10
Noul A statement to test; no answer options noul, the probability that the statement is true

confidence is the largest answer probability. For Noul, a confident false answer can have noul=0.02 and confidence=0.98; use noul when deciding whether the statement is true. escalate and escalate_prob help flag uncertain answers for another step in your application. Neither confidence nor escalation is a guarantee of correctness.

Use specific instructions and descriptive, distinct options. Probabilities are conditional on the supplied options; adding an other option can be useful when your label set does not cover every input, but the model's use of that option still needs validation.

Core and Vision apply saved temperature calibration by default. Their escalate_prob comes from a separately trained escalation head with saved calibration; default escalation is above 0.5. Set calibrated=False on prediction calls to inspect uncalibrated outputs. Their answer dictionaries do not include Sense's per-answer calibrated flag.

Local evaluation

Measured on an NVIDIA RTX 4050 Laptop GPU (6 GB). Latency includes tokenization and excludes model loading and HTTP.

Evaluation split Decisions Accuracy
Caltech101 objects 1,371 92.85%
Cars 1,153 97.66%
RVL, eight unseen document types 1,675 58.99%
In-domain validation 3,407 88.02%

See evaluation/vision_full.json for exact metrics and split details. In-domain validation is not an unseen test. The smaller shared comparison used 48 object, 48 car, and 64 document examples; those sample results should not be confused with the full evaluation above.

Timing details and interpretation

Vision does not have a dedicated repeated-call latency benchmark comparable to the Core/Sense timing suites. In the earlier small shared evaluation, per-decision medians were 87.5 ms for objects (48 decisions), 72.4 ms for cars (48), and 46.5 ms for document types (64). These are evaluation-run timings, not controlled throughput or p95 measurements. No batch-throughput claim is made for Vision.

A separate packaged-model smoke check used 1.28 GiB peak PyTorch GPU allocation while loading and scoring a simple image. That is one smoke workload, not a worst-case memory bound. The frozen SigLIP branch uses BF16 on CUDA; the decision model uses float32 weights with BF16 autocast. Image decoding, size, option counts, and batch lengths affect latency.

The full accuracy table mixes held-out object/car/document splits with an explicitly labelled in-domain validation split. It is not directly comparable to Sense's much smaller shared image subset. On that shared car subset Vision scored 93.75% and Sense 83.33%, each on 48 examples; on document types Sense scored 78.13% and Vision 57.81%, each on 64. See evaluation/vision_full.json and evaluation/shared_image_comparison.json. In the shared comparison, existing means Vision for image groups and Core for text groups. The memory smoke measurement is saved in evaluation/package_smoke.json.

Training and limitations

The base encoders were pretrained by their upstream authors. Kodama adds local training/adaptation; these foundation models were not pretrained from scratch. Included training metadata describes the local run. Evaluation data are not redistributed in this repository. Public pretraining overlap is unknown, and small benchmark samples should not be treated as broad capability guarantees. Accuracy depends on task and option wording. Confidence is not a guarantee of correctness; use suitable validation before relying on decisions.

Attribution and licensing

The pretrained base components identify their license as Apache-2.0. Upstream cards, source revisions, and the Apache license text are preserved in third_party/ and NOTICE. No separate public license has been selected for the original Kodama additions and runtime in this release; the upstream license statements should not be read as a new license grant for those additions.

Base sources:

Troubleshooting and reproducibility

  • 401/403 while downloading: authenticate with an account allowed to read the repository.
  • Model configuration / AutoModel error: use the bundled jevlike runtime and the custom loader shown above.
  • Import or tokenizer mismatch: use the documented Transformers version and install the runtime extras for this model.
  • Different answers near a tie: numerical precision and batch shapes can affect probabilities; measure your own task before changing inference settings.

release_manifest.json records file hashes and pinned base revisions. The runtime is bundled under runtime/; examples/quickstart.py provides a complete local example. The pretrained foundations were not trained from scratch by Kodama. Earlier evaluation artifacts use jevlike for Core, omni for Sense, and prototype for Vision.

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Cem13/kodama-vision

Finetuned
(1507)
this model