Kodama Sense

Larger text, image, PDF, chart, audio, and video understanding model built from the Qwen2.5-Omni-3B Thinker with a locally trained low-rank adapter. Approximately 4.39 billion runtime parameters after narrowing the decision output head. Formerly jevlike-omni-3b. Built with Qwen.

All models return structured Choice, Score, and Noul decisions. They score answer options directly; they do not generate conversational responses. These are custom jevlike checkpoints, loaded with the bundled runtime below. Generic Transformers pipeline() / AutoModel loading is not the supported entry point.

Which Kodama model should I use?

Model Inputs Typical reason to choose it
Kodama Core Text; JSON serialized as text Fast routing, classification, and small decisions inside an application
Kodama Vision One image plus optional text Experimental object, vehicle, or document-type classification
Kodama Sense Text, images, PDFs, charts, audio, video, and mixed media Decisions that require visual or audio evidence

These models choose from the answers you supply. They are not conversational chatbots, open-ended report writers, transcription services, or general document extraction systems. The use cases below are applications to evaluate on your own examples; the benchmark results do not establish reliable performance on every such application.

Practical use cases

Application Example decision Input/API
Document and chart review Select a document category or the company with the highest plotted value predict_file("page.pdf", questions) or predict(..., images=...)
Short audio classification Select a topic or test whether a statement was spoken predict_file("recording.wav", questions)
Sampled video review Classify a clip using sampled frames and its audio track predict_file("clip.mp4", questions)
Mixed evidence checks Answer a question using an image and an accompanying audio recording predict(..., images=..., audio=...)
Local multimodal research Compare typed decisions across different input modalities The same Choice, Score, and Noul interface

Sense is useful when a decision needs evidence that Core cannot read directly. It is slower for text, and its size does not guarantee better accuracy. Temporal-order video accuracy was only 5/8 in a small controlled check. Long-document relationships, precise event timing, arbitrary arithmetic, and unrestricted audio/video reasoning remain limitations.

Install and download

Tested environment: Python 3.12, PyTorch 2.10, Transformers 5.2.0, Linux, and an NVIDIA RTX 4050 Laptop GPU with 6 GB VRAM. Python 3.10+ is supported by the package. Install a CUDA-enabled PyTorch build appropriate for your system when using the GPU examples. Approximate model download size: 10.01 GB, excluding Python/CUDA dependencies and caches.

python -m pip install huggingface_hub
# For a private repository, authenticate once with: hf auth login
hf download Cem13/kodama-sense --local-dir kodama-sense
python -m pip install "./kodama-sense/runtime[omni]" "transformers==5.2.0"

Run the examples from the directory containing kodama-sense. After the download and dependency installation, prediction runs locally; no API key or paid inference service is needed for prediction. Private downloads require access to the repository. An existing HF_TOKEN can authenticate the Hub client without adding a token to source code.

Install ffmpeg and ffprobe on PATH for audio/video decoding.

Python: text, JSON, images, and files

from jevlike import OmniSystemOne, Choice, Score, Noul

model = OmniSystemOne.load("kodama-sense", device="cuda")
questions = {
    "team": Choice("Which team should handle this request?", {
        "billing": "payments, invoices, refunds",
        "technical": "bugs, errors, broken software",
        "other": "another kind of request",
    }),
    "urgency": Score("How urgent is this request?", ["low", "normal", "urgent"]),
    "refund": Noul("The customer explicitly requests a refund."),
}
result = model.predict(
    {"message": "I was charged twice. Please refund the duplicate payment."}, questions
)
print(result["answers"])

Sense accepts either a string or a JSON-serializable state. A runnable example is included:

python kodama-sense/examples/quickstart.py

For media, replace these example paths with your files and choose questions that fit them:

visual_questions = {
    "kind": Choice("What kind of document is shown?", ["invoice", "letter", "report", "other"]),
}
pdf_result = model.predict_file("document.pdf", visual_questions)
image_result = model.predict(None, visual_questions, images="scan.png")

audio_questions = {
    "topic": Choice("What is the main topic of the recording?", ["billing", "technical support", "other"]),
}
audio_result = model.predict_file("recording.wav", audio_questions)
video_result = model.predict_file("clip.mp4", audio_questions)

combined = model.predict(
    "Use both the picture and the recording.",
    {"match": Noul("The spoken description matches the visible chart.")},
    images="chart.png", audio="commentary.wav",
)

These snippets show supported calls, not measured accuracy on these particular tasks. There is no transcript or generated explanation in the response. Video uses both its sampled frames and audio track when available.

Input limits and long files

Input Direct predict predict_file
Text Default maximum 4,096 processor tokens, including prompt/options Split into text windows, normally up to 1,800 text tokens each
Images Up to eight images; default 448ร—448 pixel budget per image Images or multi-page TIFF; each page is processed
PDF Use predict_file Rendered page plus available text layer, page by page
Audio Up to 30 seconds All consecutive 30-second windows
Video Up to eight seconds All consecutive eight-second windows; at most eight sampled frames each

Each call accepts 1โ€“128 questions. The default limit is 128 inference windows. Inputs exceeding configured limits raise errors rather than being silently clipped. predict_file reports coverage in result["media"]; pass return_windows=True to inspect window results. More media or longer answer options can reach the token limit sooner.

Default pooling uses a geometric mean for Choice/Score and maximum true probability for Noul. Across a long recording, that Noul rule means true in any window. aggregation="mean" changes the pooling rule, but does not create holistic reasoning across the entire recording. Sampled video can miss brief events between frames.

Throughput and hardware

Put questions about the same evidence into one predict call to use prefix caching and bounded batches. Defaults: text_batch_size=8, text_batch_tokens=4096, and text_prefix_cache=True. text_batch_size=1 selects the sequential reference path. predict_batch([(state, questions), ...]) returns one result per record; separate records still execute sequentially. Image/audio/video questions also execute sequentially.

The tested configuration uses CUDA, 4-bit NF4 weights, and BF16 computation. CPU loading is supported with device="cpu", quantize=False, requires substantially more host RAM, and has not been performance-benchmarked. The upstream/ shards include all required Thinker weights; unused speech-generation tensors may cause harmless unexpected-key warnings. No speech generation component is used by this decision API.

Understanding the answers

Every response contains model, latency_ms, and answers, keyed by your question names.

Question type How to define it Main output
Choice 2โ€“255 distinct labels, optionally with descriptions choice and a label-to-probability dictionary
Score 2โ€“10 ordered descriptions, lowest first Zero-based level, its label, a probability list, and expected score scaled to 0โ€“10
Noul A statement to test; no answer options noul, the probability that the statement is true

confidence is the largest answer probability. For Noul, a confident false answer can have noul=0.02 and confidence=0.98; use noul when deciding whether the statement is true. escalate and escalate_prob help flag uncertain answers for another step in your application. Neither confidence nor escalation is a guarantee of correctness.

Use specific instructions and descriptive, distinct options. Probabilities are conditional on the supplied options; adding an other option can be useful when your label set does not cover every input, but the model's use of that option still needs validation.

Sense reports calibrated per answer. It applies only existing calibration for the current adapter, modality, and question type. Multi-window pooled answers are marked uncalibrated. escalate_prob is 1 - confidence, with default escalation above 0.35; it is a heuristic rather than a trained error predictor.

Serve a local HTTP API

python -m jevlike.omni_server --checkpoint kodama-sense --port 8078

From another terminal:

curl http://127.0.0.1:8078/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{"state":{"message":"Please refund the duplicate charge."},"questions":{"team":{"type":"choice","instructions":"Which team?","criteria":["billing","technical","other"]}}}'

The server binds to localhost by default. GET /health reports readiness. Model loading happens once at startup. This repository does not supply an active hosted inference endpoint.

For a media file, the multipart endpoint accepts the same question schema:

curl http://127.0.0.1:8078/v1/systemone/file \
  -F 'file=@recording.wav' \
  -F 'questions={"topic":{"type":"choice","instructions":"What is the main topic?","criteria":["billing","technical support","other"]}}'

Use /v1/systemone/media with repeated files fields for combined image/audio/video inputs. The server caps combined uploads at 128 MiB; local Python file calls use the window and token limits described above.

Local evaluation

Measured on an NVIDIA RTX 4050 Laptop GPU (6 GB). Latency includes tokenization and excludes model loading and HTTP.

Dataset Decisions Optimized accuracy
Short news/finance 140 66.43%
Long news/finance 56 55.36%
Other held-out tasks 300 63.33%
Typed decisions 2,000 47.55%

Median latency after optimization: 127 ms for one short-text question, 465 ms for ten short-text questions, and 1.19 s for ten questions over long text (10 paired timed calls per workload). Peak PyTorch allocation, including loading: 3.19 GiB. Batching changed 51/2,496 top labels relative to sequential inference due to numerical differences; typed accuracy changed 47.65% โ†’ 47.55%.

Media checks included images, documents, charts, PDF, speech, environmental audio, and audiovisual clips. Synthetic groups have only eight examples each; these are integration checks, not evidence of general media accuracy. Video temporal-order accuracy was only 5/8. See evaluation/media_comparison.json for all groups.

Timing details and interpretation

Text workload Optimized median Optimized p95 Sequential โ†’ optimized speedup
short 1q 127.0 ms 132.9 ms 1.00ร—
short 5q 274.3 ms 277.4 ms 2.31ร—
short 10q 464.9 ms 620.4 ms 2.76ร—
short 50q 2151.7 ms 2440.6 ms 3.03ร—
long 1q 609.4 ms 615.3 ms 1.00ร—
long 5q 919.5 ms 945.1 ms 3.31ร—
long 10q 1193.6 ms 1226.2 ms 5.11ร—

Ten timed calls per workload, three warmups, four CPU threads, and CUDA synchronization. Before/after paths alternated on identical changing input strings in the same process. Long calls used the same complete document-window path. The 16-state ร— three-question batch measured 11.87 decisions/s, versus 7.66 for the sequential path (three repetitions). Single-question inference has no algorithmic speedup from this optimization.

The earlier Laya run measured 30.8 ms for one short-text question and 118.5 ms for ten; optimized Sense measured 127.0/464.9 ms. These are separate runs on the same GPU, and Laya's optional fast path was not tested. Sense's main benefit is broader input support. Its 55.36% long-news/finance accuracy was below Core's 71.43% on 56 decisions. The small short-text accuracy lead does not establish a reliable general advantage.

Media accuracy checks

Group Decisions Accuracy Scope
Objects 48 89.58% Small shared image subset
Cars 48 83.33% Small shared image subset
Document types 64 78.13% Small shared document subset
Charts 8 100% Controlled synthetic fixtures
Scanned PDF charts 8 100% Controlled synthetic fixtures
Speech questions 8 100% Controlled synthetic speech
Video order 8 62.50% Controlled synthetic clips
Video/audio questions 8 100% Controlled synthetic clips
Environmental audio 20 100% ESC-50 fold-5 subset, ten classes

The small 100% results do not show general audio, video, PDF, or chart reliability. The public audio dataset is not redistributed here. These are the earlier media accuracy checks, not a newly enlarged test suite.

Media timing smoke checks

Input Median latency
One chart image 291 ms
One scanned PDF page 307 ms
3.7-second audio clip 336 ms
Four-second video with audio 568 ms
Image plus audio 536 ms

Three timed calls after three warmups per case; decode and processing are included. These small fixture timings are not sustained streaming measurements. Peak allocation for the optimization process, including loading, was 3.19 GiB; peak reservation 3.52 GiB. Desktop GPU memory is additional. p95 estimates from small samples are exploratory, not latency guarantees. Changing input lengths, question counts, and hardware changes results.

Recorded evidence: accuracy, timing samples, media accuracy, media timings, and run metadata. In the media comparison JSON, existing denotes Core for text groups and Vision for image groups.

Training and limitations

The base encoders were pretrained by their upstream authors. Kodama adds local training/adaptation; these foundation models were not pretrained from scratch. Included training metadata describes the local run. Evaluation data are not redistributed in this repository. Public pretraining overlap is unknown, and small benchmark samples should not be treated as broad capability guarantees. Accuracy depends on task and option wording. Confidence is not a guarantee of correctness; use suitable validation before relying on decisions.

The active adapter trains 307,200 parameters; 3,072 mixed examples were used and the selected checkpoint was step 384 of a 768-step run, using validation NLL. Calibration is scoped to modality and question type and tied to the adapter hash. Pooled multi-window answers are marked uncalibrated. escalate_prob is the heuristic 1 - confidence, not a separately learned error model. Audio/video calibration has not been established. Video is sampled to at most eight frames per eight-second window, so brief events and cross-window relationships may be missed.

Attribution and licensing

The pinned Qwen2.5-Omni-3B base is governed by the Qwen Research License, included in LICENSE. It permits research/evaluation use and requires a separate upstream license for commercial use. See NOTICE for attribution and changes. Earlier local metadata incorrectly claimed Apache-2.0; this release corrects that error.

Base sources:

Troubleshooting and reproducibility

  • 401/403 while downloading: authenticate with an account allowed to read the repository.
  • Model configuration / AutoModel error: use the bundled jevlike runtime and the custom loader shown above.
  • Import or tokenizer mismatch: use the documented Transformers version and install the runtime extras for this model.
  • Different answers near a tie: numerical precision and batch shapes can affect probabilities; measure your own task before changing inference settings.

release_manifest.json records file hashes and pinned base revisions. The runtime is bundled under runtime/; examples/quickstart.py provides a complete local example. The pretrained foundations were not trained from scratch by Kodama. Earlier evaluation artifacts use jevlike for Core, omni for Sense, and prototype for Vision.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Cem13/kodama-sense

Finetuned
(30)
this model