- Kodama Sense
Kodama Sense
Larger text, image, PDF, chart, audio, and video understanding model built from the Qwen2.5-Omni-3B Thinker with a locally trained low-rank adapter. Approximately 4.39 billion runtime parameters after narrowing the decision output head. Formerly jevlike-omni-3b. Built with Qwen.
All models return structured Choice, Score, and Noul decisions. They score answer options directly; they do not generate conversational responses. These are custom jevlike checkpoints, loaded with the bundled runtime below. Generic Transformers pipeline() / AutoModel loading is not the supported entry point.
Which Kodama model should I use?
| Model | Inputs | Typical reason to choose it |
|---|---|---|
| Kodama Core | Text; JSON serialized as text | Fast routing, classification, and small decisions inside an application |
| Kodama Vision | One image plus optional text | Experimental object, vehicle, or document-type classification |
| Kodama Sense | Text, images, PDFs, charts, audio, video, and mixed media | Decisions that require visual or audio evidence |
These models choose from the answers you supply. They are not conversational chatbots, open-ended report writers, transcription services, or general document extraction systems. The use cases below are applications to evaluate on your own examples; the benchmark results do not establish reliable performance on every such application.
Practical use cases
| Application | Example decision | Input/API |
|---|---|---|
| Document and chart review | Select a document category or the company with the highest plotted value | predict_file("page.pdf", questions) or predict(..., images=...) |
| Short audio classification | Select a topic or test whether a statement was spoken | predict_file("recording.wav", questions) |
| Sampled video review | Classify a clip using sampled frames and its audio track | predict_file("clip.mp4", questions) |
| Mixed evidence checks | Answer a question using an image and an accompanying audio recording | predict(..., images=..., audio=...) |
| Local multimodal research | Compare typed decisions across different input modalities | The same Choice, Score, and Noul interface |
Sense is useful when a decision needs evidence that Core cannot read directly. It is slower for text, and its size does not guarantee better accuracy. Temporal-order video accuracy was only 5/8 in a small controlled check. Long-document relationships, precise event timing, arbitrary arithmetic, and unrestricted audio/video reasoning remain limitations.
Install and download
Tested environment: Python 3.12, PyTorch 2.10, Transformers 5.2.0, Linux, and an NVIDIA RTX 4050 Laptop GPU with 6 GB VRAM. Python 3.10+ is supported by the package. Install a CUDA-enabled PyTorch build appropriate for your system when using the GPU examples. Approximate model download size: 10.01 GB, excluding Python/CUDA dependencies and caches.
python -m pip install huggingface_hub
# For a private repository, authenticate once with: hf auth login
hf download Cem13/kodama-sense --local-dir kodama-sense
python -m pip install "./kodama-sense/runtime[omni]" "transformers==5.2.0"
Run the examples from the directory containing kodama-sense. After the download and
dependency installation, prediction runs locally; no API key or paid inference service
is needed for prediction. Private downloads require access to the repository. An existing
HF_TOKEN can authenticate the Hub client without adding a token to source code.
Install ffmpeg and ffprobe on PATH for audio/video decoding.
Python: text, JSON, images, and files
from jevlike import OmniSystemOne, Choice, Score, Noul
model = OmniSystemOne.load("kodama-sense", device="cuda")
questions = {
"team": Choice("Which team should handle this request?", {
"billing": "payments, invoices, refunds",
"technical": "bugs, errors, broken software",
"other": "another kind of request",
}),
"urgency": Score("How urgent is this request?", ["low", "normal", "urgent"]),
"refund": Noul("The customer explicitly requests a refund."),
}
result = model.predict(
{"message": "I was charged twice. Please refund the duplicate payment."}, questions
)
print(result["answers"])
Sense accepts either a string or a JSON-serializable state. A runnable example is included:
python kodama-sense/examples/quickstart.py
For media, replace these example paths with your files and choose questions that fit them:
visual_questions = {
"kind": Choice("What kind of document is shown?", ["invoice", "letter", "report", "other"]),
}
pdf_result = model.predict_file("document.pdf", visual_questions)
image_result = model.predict(None, visual_questions, images="scan.png")
audio_questions = {
"topic": Choice("What is the main topic of the recording?", ["billing", "technical support", "other"]),
}
audio_result = model.predict_file("recording.wav", audio_questions)
video_result = model.predict_file("clip.mp4", audio_questions)
combined = model.predict(
"Use both the picture and the recording.",
{"match": Noul("The spoken description matches the visible chart.")},
images="chart.png", audio="commentary.wav",
)
These snippets show supported calls, not measured accuracy on these particular tasks. There is no transcript or generated explanation in the response. Video uses both its sampled frames and audio track when available.
Input limits and long files
| Input | Direct predict |
predict_file |
|---|---|---|
| Text | Default maximum 4,096 processor tokens, including prompt/options | Split into text windows, normally up to 1,800 text tokens each |
| Images | Up to eight images; default 448ร448 pixel budget per image | Images or multi-page TIFF; each page is processed |
Use predict_file |
Rendered page plus available text layer, page by page | |
| Audio | Up to 30 seconds | All consecutive 30-second windows |
| Video | Up to eight seconds | All consecutive eight-second windows; at most eight sampled frames each |
Each call accepts 1โ128 questions. The default limit is 128 inference windows. Inputs exceeding configured limits raise
errors rather than being silently clipped. predict_file reports coverage in
result["media"]; pass return_windows=True to inspect window results.
More media or longer answer options can reach the token limit sooner.
Default pooling uses a geometric mean for Choice/Score and maximum true probability
for Noul. Across a long recording, that Noul rule means true in any window.
aggregation="mean" changes the pooling rule, but does not create holistic reasoning
across the entire recording. Sampled video can miss brief events between frames.
Throughput and hardware
Put questions about the same evidence into one predict call to use prefix caching
and bounded batches. Defaults: text_batch_size=8, text_batch_tokens=4096, and
text_prefix_cache=True. text_batch_size=1 selects the sequential reference path.
predict_batch([(state, questions), ...]) returns one result per record; separate records
still execute sequentially. Image/audio/video questions also execute sequentially.
The tested configuration uses CUDA, 4-bit NF4 weights, and BF16 computation. CPU loading
is supported with device="cpu", quantize=False, requires substantially more host RAM,
and has not been performance-benchmarked. The upstream/ shards include all required
Thinker weights; unused speech-generation tensors may cause harmless unexpected-key
warnings. No speech generation component is used by this decision API.
Understanding the answers
Every response contains model, latency_ms, and answers, keyed by your question names.
| Question type | How to define it | Main output |
|---|---|---|
Choice |
2โ255 distinct labels, optionally with descriptions | choice and a label-to-probability dictionary |
Score |
2โ10 ordered descriptions, lowest first | Zero-based level, its label, a probability list, and expected score scaled to 0โ10 |
Noul |
A statement to test; no answer options | noul, the probability that the statement is true |
confidence is the largest answer probability. For Noul, a confident false answer
can have noul=0.02 and confidence=0.98; use noul when deciding whether the statement
is true. escalate and escalate_prob help flag uncertain answers for another step in
your application. Neither confidence nor escalation is a guarantee of correctness.
Use specific instructions and descriptive, distinct options. Probabilities are conditional
on the supplied options; adding an other option can be useful when your label set does
not cover every input, but the model's use of that option still needs validation.
Sense reports calibrated per answer. It applies only existing calibration for the
current adapter, modality, and question type. Multi-window pooled answers are marked
uncalibrated. escalate_prob is 1 - confidence, with default escalation above 0.35;
it is a heuristic rather than a trained error predictor.
Serve a local HTTP API
python -m jevlike.omni_server --checkpoint kodama-sense --port 8078
From another terminal:
curl http://127.0.0.1:8078/v1/systemone \
-H 'Content-Type: application/json' \
-d '{"state":{"message":"Please refund the duplicate charge."},"questions":{"team":{"type":"choice","instructions":"Which team?","criteria":["billing","technical","other"]}}}'
The server binds to localhost by default. GET /health reports readiness. Model loading
happens once at startup. This repository does not supply an active hosted inference endpoint.
For a media file, the multipart endpoint accepts the same question schema:
curl http://127.0.0.1:8078/v1/systemone/file \
-F 'file=@recording.wav' \
-F 'questions={"topic":{"type":"choice","instructions":"What is the main topic?","criteria":["billing","technical support","other"]}}'
Use /v1/systemone/media with repeated files fields for combined image/audio/video
inputs. The server caps combined uploads at 128 MiB; local Python file calls use the
window and token limits described above.
Local evaluation
Measured on an NVIDIA RTX 4050 Laptop GPU (6 GB). Latency includes tokenization and excludes model loading and HTTP.
| Dataset | Decisions | Optimized accuracy |
|---|---|---|
| Short news/finance | 140 | 66.43% |
| Long news/finance | 56 | 55.36% |
| Other held-out tasks | 300 | 63.33% |
| Typed decisions | 2,000 | 47.55% |
Median latency after optimization: 127 ms for one short-text question, 465 ms for ten short-text questions, and 1.19 s for ten questions over long text (10 paired timed calls per workload). Peak PyTorch allocation, including loading: 3.19 GiB. Batching changed 51/2,496 top labels relative to sequential inference due to numerical differences; typed accuracy changed 47.65% โ 47.55%.
Media checks included images, documents, charts, PDF, speech, environmental audio,
and audiovisual clips. Synthetic groups have only eight examples each; these are
integration checks, not evidence of general media accuracy. Video temporal-order
accuracy was only 5/8. See evaluation/media_comparison.json for all groups.
Timing details and interpretation
| Text workload | Optimized median | Optimized p95 | Sequential โ optimized speedup |
|---|---|---|---|
| short 1q | 127.0 ms | 132.9 ms | 1.00ร |
| short 5q | 274.3 ms | 277.4 ms | 2.31ร |
| short 10q | 464.9 ms | 620.4 ms | 2.76ร |
| short 50q | 2151.7 ms | 2440.6 ms | 3.03ร |
| long 1q | 609.4 ms | 615.3 ms | 1.00ร |
| long 5q | 919.5 ms | 945.1 ms | 3.31ร |
| long 10q | 1193.6 ms | 1226.2 ms | 5.11ร |
Ten timed calls per workload, three warmups, four CPU threads, and CUDA synchronization. Before/after paths alternated on identical changing input strings in the same process. Long calls used the same complete document-window path. The 16-state ร three-question batch measured 11.87 decisions/s, versus 7.66 for the sequential path (three repetitions). Single-question inference has no algorithmic speedup from this optimization.
The earlier Laya run measured 30.8 ms for one short-text question and 118.5 ms for ten; optimized Sense measured 127.0/464.9 ms. These are separate runs on the same GPU, and Laya's optional fast path was not tested. Sense's main benefit is broader input support. Its 55.36% long-news/finance accuracy was below Core's 71.43% on 56 decisions. The small short-text accuracy lead does not establish a reliable general advantage.
Media accuracy checks
| Group | Decisions | Accuracy | Scope |
|---|---|---|---|
| Objects | 48 | 89.58% | Small shared image subset |
| Cars | 48 | 83.33% | Small shared image subset |
| Document types | 64 | 78.13% | Small shared document subset |
| Charts | 8 | 100% | Controlled synthetic fixtures |
| Scanned PDF charts | 8 | 100% | Controlled synthetic fixtures |
| Speech questions | 8 | 100% | Controlled synthetic speech |
| Video order | 8 | 62.50% | Controlled synthetic clips |
| Video/audio questions | 8 | 100% | Controlled synthetic clips |
| Environmental audio | 20 | 100% | ESC-50 fold-5 subset, ten classes |
The small 100% results do not show general audio, video, PDF, or chart reliability. The public audio dataset is not redistributed here. These are the earlier media accuracy checks, not a newly enlarged test suite.
Media timing smoke checks
| Input | Median latency |
|---|---|
| One chart image | 291 ms |
| One scanned PDF page | 307 ms |
| 3.7-second audio clip | 336 ms |
| Four-second video with audio | 568 ms |
| Image plus audio | 536 ms |
Three timed calls after three warmups per case; decode and processing are included. These small fixture timings are not sustained streaming measurements. Peak allocation for the optimization process, including loading, was 3.19 GiB; peak reservation 3.52 GiB. Desktop GPU memory is additional. p95 estimates from small samples are exploratory, not latency guarantees. Changing input lengths, question counts, and hardware changes results.
Recorded evidence: accuracy,
timing samples, media accuracy,
media timings, and run metadata.
In the media comparison JSON, existing denotes Core for text groups and Vision for image groups.
Training and limitations
The base encoders were pretrained by their upstream authors. Kodama adds local training/adaptation; these foundation models were not pretrained from scratch. Included training metadata describes the local run. Evaluation data are not redistributed in this repository. Public pretraining overlap is unknown, and small benchmark samples should not be treated as broad capability guarantees. Accuracy depends on task and option wording. Confidence is not a guarantee of correctness; use suitable validation before relying on decisions.
The active adapter trains 307,200 parameters; 3,072 mixed examples were used and
the selected checkpoint was step 384 of a 768-step run, using validation NLL.
Calibration is scoped to modality and question type and tied to the adapter hash.
Pooled multi-window answers are marked uncalibrated. escalate_prob is the heuristic
1 - confidence, not a separately learned error model. Audio/video calibration
has not been established. Video is sampled to at most eight frames per eight-second
window, so brief events and cross-window relationships may be missed.
Attribution and licensing
The pinned Qwen2.5-Omni-3B base is governed by the Qwen Research License, included in LICENSE. It permits research/evaluation use and requires a separate upstream license for commercial use. See NOTICE for attribution and changes. Earlier local metadata incorrectly claimed Apache-2.0; this release corrects that error.
Base sources:
Troubleshooting and reproducibility
- 401/403 while downloading: authenticate with an account allowed to read the repository.
- Model configuration / AutoModel error: use the bundled
jevlikeruntime and the custom loader shown above. - Import or tokenizer mismatch: use the documented Transformers version and install the runtime extras for this model.
- Different answers near a tie: numerical precision and batch shapes can affect probabilities; measure your own task before changing inference settings.
release_manifest.json records file hashes and pinned base revisions. The runtime is
bundled under runtime/; examples/quickstart.py provides a complete local example.
The pretrained foundations were not trained from scratch by Kodama. Earlier evaluation
artifacts use jevlike for Core, omni for Sense, and prototype for Vision.
Model tree for Cem13/kodama-sense
Base model
Qwen/Qwen2.5-Omni-3B