- Kodama Core
- Which Kodama model should I use?
- Practical use cases
- Install and download
- Python: route a ticket and score urgency
- Understanding the answers
- Serve a local HTTP API
- Local evaluation
- Timing details and interpretation
- Training and limitations
- Attribution and licensing
- Troubleshooting and reproducibility
- Which Kodama model should I use?
Kodama Core
Fast text/JSON decision model with a ModernBERT-large backbone and a trained typed-decision head. Approximately 412.9 million parameters. Formerly jevlike-large-nf.
All models return structured Choice, Score, and Noul decisions. They score answer options directly; they do not generate conversational responses. These are custom jevlike checkpoints, loaded with the bundled runtime below. Generic Transformers pipeline() / AutoModel loading is not the supported entry point.
Which Kodama model should I use?
| Model | Inputs | Typical reason to choose it |
|---|---|---|
| Kodama Core | Text; JSON serialized as text | Fast routing, classification, and small decisions inside an application |
| Kodama Vision | One image plus optional text | Experimental object, vehicle, or document-type classification |
| Kodama Sense | Text, images, PDFs, charts, audio, video, and mixed media | Decisions that require visual or audio evidence |
These models choose from the answers you supply. They are not conversational chatbots, open-ended report writers, transcription services, or general document extraction systems. The use cases below are applications to evaluate on your own examples; the benchmark results do not establish reliable performance on every such application.
Practical use cases
| Application | Example question | Decision returned |
|---|---|---|
| Support-ticket routing | “Which team should handle this request?” | Billing, technical support, sales, or other |
| Queue prioritization | “How urgent is this ticket?” | An ordered urgency level and its probabilities |
| Agent workflow checks | “Does the supplied log show that the task completed?” | A probability for the statement |
| News or document triage | “Which topic best describes this text?” | One label from a supplied taxonomy |
| Structured business records | “Which follow-up action fits this order record?” | A typed classification after JSON serialization |
Core is the first model to try when the evidence is already text and latency matters. Its 45.95% result on the broad typed-decisions test means task-specific validation is still necessary before automatically acting on these decisions. It does not read raw images, PDF pages, audio, or video; provide extracted text or choose Sense for those inputs.
Install and download
Tested environment: Python 3.12, PyTorch 2.10, Transformers 5.2.0, Linux, and an NVIDIA RTX 4050 Laptop GPU with 6 GB VRAM. Python 3.10+ is supported by the package. Install a CUDA-enabled PyTorch build appropriate for your system when using the GPU examples. Approximate model download size: 1.66 GB, excluding Python/CUDA dependencies and caches.
python -m pip install huggingface_hub
# For a private repository, authenticate once with: hf auth login
hf download Cem13/kodama-core --local-dir kodama-core
python -m pip install "./kodama-core/runtime[model,ingest]" "transformers==5.2.0"
Run the examples from the directory containing kodama-core. After the download and
dependency installation, prediction runs locally; no API key or paid inference service
is needed for prediction. Private downloads require access to the repository. An existing
HF_TOKEN can authenticate the Hub client without adding a token to source code.
Python: route a ticket and score urgency
import json
from jevlike import SystemOne, Choice, Score, Noul
model = SystemOne.load("kodama-core", device="cuda", long_state="chunk")
questions = {
"team": Choice("Which team should handle this request?", {
"billing": "payments, invoices, refunds",
"technical": "bugs, errors, broken software",
"other": "another kind of request",
}),
"urgency": Score("How urgent is this request?", ["low", "normal", "urgent"]),
"refund": Noul("The customer explicitly requests a refund."),
}
state = json.dumps({"message": "I was charged twice. Please refund the duplicate payment."})
result = model.predict(state, questions)
print(result["answers"])
The direct Core Python API takes a string state. Serialize a dictionary with
json.dumps, as above. The HTTP API below accepts JSON objects directly.
A runnable version is included at examples/quickstart.py:
python kodama-core/examples/quickstart.py
Multiple records and long text
results = model.predict_batch([
("Please refund this duplicate charge.", questions),
("The application crashes every time I log in.", questions),
])
Each result matches the corresponding input item. Questions across records are batched
within a token budget. Batch-result latency_ms covers the whole batch call, not an
independent per-record measurement. The checkpoint uses a 2,048-token context, including
the question and option text. long_state="chunk" reads overlapping windows; the default
"truncate" mode truncates long states. Chunk pooling does not reproduce a single
unlimited-context read, and escalation/calibration should be revalidated for long inputs.
For CPU execution, load with device="cpu"; CPU speed has not been benchmarked here.
The GPU benchmark used float32 model weights with BF16 autocast.
Understanding the answers
Every response contains model, latency_ms, and answers, keyed by your question names.
| Question type | How to define it | Main output |
|---|---|---|
Choice |
2–255 distinct labels, optionally with descriptions | choice and a label-to-probability dictionary |
Score |
2–10 ordered descriptions, lowest first | Zero-based level, its label, a probability list, and expected score scaled to 0–10 |
Noul |
A statement to test; no answer options | noul, the probability that the statement is true |
confidence is the largest answer probability. For Noul, a confident false answer
can have noul=0.02 and confidence=0.98; use noul when deciding whether the statement
is true. escalate and escalate_prob help flag uncertain answers for another step in
your application. Neither confidence nor escalation is a guarantee of correctness.
Use specific instructions and descriptive, distinct options. Probabilities are conditional
on the supplied options; adding an other option can be useful when your label set does
not cover every input, but the model's use of that option still needs validation.
Core and Vision apply saved temperature calibration by default. Their escalate_prob
comes from a separately trained escalation head with saved calibration; default escalation
is above 0.5. Set calibrated=False on prediction calls to inspect uncalibrated outputs.
Their answer dictionaries do not include Sense's per-answer calibrated flag.
Serve a local HTTP API
python -m pip install "./kodama-core/runtime[server]"
python -m jevlike.server --checkpoint kodama-core --port 8077
From another terminal:
curl http://127.0.0.1:8077/v1/systemone \
-H 'Content-Type: application/json' \
-d '{"state":{"message":"Please refund the duplicate charge."},"questions":{"team":{"type":"choice","instructions":"Which team?","criteria":["billing","technical","other"]}}}'
The server binds to localhost by default. GET /health reports readiness. Model loading
happens once at startup. This repository does not supply an active hosted inference endpoint.
Local evaluation
Measured on an NVIDIA RTX 4050 Laptop GPU (6 GB). Latency includes tokenization and excludes model loading and HTTP.
| Dataset | Decisions | Accuracy |
|---|---|---|
| Short news/finance | 140 | 63.57% |
| Long news/finance | 56 | 71.43% |
| Other held-out tasks | 300 | 58.00% |
| Typed decisions | 2,000 | 45.95% |
Median short-text latency: 26.5 ms for one question and 104.8 ms for ten questions (20 timed calls per workload). Batch throughput: 94.6 decisions/s. These are local exploratory measurements, not service guarantees.
Timing details and interpretation
| Text workload | Median | p95 |
|---|---|---|
| short 1q | 26.5 ms | 29.0 ms |
| short 5q | 60.1 ms | 65.5 ms |
| short 10q | 104.8 ms | 119.2 ms |
| short 50q | 682.2 ms | 754.4 ms |
| long 1q | 123.9 ms | 131.8 ms |
| long 5q | 685.8 ms | 728.8 ms |
| long 10q | 1321.1 ms | 1337.0 ms |
Twenty timed calls per workload, three warmups, four CPU threads, and CUDA synchronization. Short evidence strings contain approximately 80–120 ModernBERT tokens. Long timing strings are the same shared strings across backends, derived from 1,500 ModernBERT tokens. Batch throughput used 16 states × three questions and three repetitions. Peak PyTorch allocation in the comparison run was 2.80 GiB (allocator reservation 2.90 GiB); this is not a universal minimum VRAM requirement. The desktop shared the GPU, and clocks were not locked.
On the same earlier benchmark, standard Laya measured 30.8 ms for one short-text question and 118.5 ms for ten, compared with Core's 26.5/104.8 ms. Accuracy was task-dependent: Core scored 58.0% versus Laya's 66.0% on other held-out tasks. Laya's specialized checkpoint scored 74.25% on typed decisions versus Core's 45.95%; that checkpoint was trained for the task family. The results do not establish that Core universally outperforms Laya.
The 2,496 accuracy decisions use identical normalized states/questions across models. No calibration was refitted on that test set. Small news/finance groups have substantial sampling uncertainty. These local normalized tasks are not a reproduction of publisher headline benchmark scores. See evaluation/comparison.json for the recorded protocol, model accuracies, and Core's timing samples. Run metadata records the hardware, precision, and memory measurements.
Training and limitations
The base encoders were pretrained by their upstream authors. Kodama adds local training/adaptation; these foundation models were not pretrained from scratch. Included training metadata describes the local run. Evaluation data are not redistributed in this repository. Public pretraining overlap is unknown, and small benchmark samples should not be treated as broad capability guarantees. Accuracy depends on task and option wording. Confidence is not a guarantee of correctness; use suitable validation before relying on decisions.
Attribution and licensing
The pretrained base components identify their license as Apache-2.0. Upstream cards, source revisions, and the Apache license text are preserved in third_party/ and NOTICE. No separate public license has been selected for the original Kodama additions and runtime in this release; the upstream license statements should not be read as a new license grant for those additions.
Base sources:
Troubleshooting and reproducibility
- 401/403 while downloading: authenticate with an account allowed to read the repository.
- Model configuration / AutoModel error: use the bundled
jevlikeruntime and the custom loader shown above. - Import or tokenizer mismatch: use the documented Transformers version and install the runtime extras for this model.
- Different answers near a tie: numerical precision and batch shapes can affect probabilities; measure your own task before changing inference settings.
release_manifest.json records file hashes and pinned base revisions. The runtime is
bundled under runtime/; examples/quickstart.py provides a complete local example.
The pretrained foundations were not trained from scratch by Kodama. Earlier evaluation
artifacts use jevlike for Core, omni for Sense, and prototype for Vision.
- Downloads last month
- 1
Model tree for Cem13/kodama-core
Base model
answerdotai/ModernBERT-large