--- language: - en base_model: - answerdotai/ModernBERT-large tags: - kodama - typed-decisions - custom-code - safetensors - text-classification inference: false --- # Kodama Core Fast text/JSON decision model with a ModernBERT-large backbone and a trained typed-decision head. Approximately 412.9 million parameters. Formerly `jevlike-large-nf`. All models return structured `Choice`, `Score`, and `Noul` decisions. They score answer options directly; they do not generate conversational responses. These are custom `jevlike` checkpoints, loaded with the bundled runtime below. Generic Transformers `pipeline()` / `AutoModel` loading is not the supported entry point. ## Which Kodama model should I use? | Model | Inputs | Typical reason to choose it | |---|---|---| | [Kodama Core](https://huggingface.co/Cem13/kodama-core) | Text; JSON serialized as text | Fast routing, classification, and small decisions inside an application | | [Kodama Vision](https://huggingface.co/Cem13/kodama-vision) | One image plus optional text | Experimental object, vehicle, or document-type classification | | [Kodama Sense](https://huggingface.co/Cem13/kodama-sense) | Text, images, PDFs, charts, audio, video, and mixed media | Decisions that require visual or audio evidence | These models choose from the answers you supply. They are not conversational chatbots, open-ended report writers, transcription services, or general document extraction systems. The use cases below are applications to evaluate on your own examples; the benchmark results do not establish reliable performance on every such application. ## Practical use cases | Application | Example question | Decision returned | |---|---|---| | Support-ticket routing | “Which team should handle this request?” | Billing, technical support, sales, or other | | Queue prioritization | “How urgent is this ticket?” | An ordered urgency level and its probabilities | | Agent workflow checks | “Does the supplied log show that the task completed?” | A probability for the statement | | News or document triage | “Which topic best describes this text?” | One label from a supplied taxonomy | | Structured business records | “Which follow-up action fits this order record?” | A typed classification after JSON serialization | Core is the first model to try when the evidence is already text and latency matters. Its 45.95% result on the broad typed-decisions test means task-specific validation is still necessary before automatically acting on these decisions. It does not read raw images, PDF pages, audio, or video; provide extracted text or choose Sense for those inputs. ## Install and download Tested environment: Python 3.12, PyTorch 2.10, Transformers 5.2.0, Linux, and an NVIDIA RTX 4050 Laptop GPU with 6 GB VRAM. Python 3.10+ is supported by the package. Install a CUDA-enabled PyTorch build appropriate for your system when using the GPU examples. Approximate model download size: **1.66 GB**, excluding Python/CUDA dependencies and caches. ```bash python -m pip install huggingface_hub # For a private repository, authenticate once with: hf auth login hf download Cem13/kodama-core --local-dir kodama-core python -m pip install "./kodama-core/runtime[model,ingest]" "transformers==5.2.0" ``` Run the examples from the directory containing `kodama-core`. After the download and dependency installation, prediction runs locally; no API key or paid inference service is needed for prediction. Private downloads require access to the repository. An existing `HF_TOKEN` can authenticate the Hub client without adding a token to source code. ## Python: route a ticket and score urgency ```python import json from jevlike import SystemOne, Choice, Score, Noul model = SystemOne.load("kodama-core", device="cuda", long_state="chunk") questions = { "team": Choice("Which team should handle this request?", { "billing": "payments, invoices, refunds", "technical": "bugs, errors, broken software", "other": "another kind of request", }), "urgency": Score("How urgent is this request?", ["low", "normal", "urgent"]), "refund": Noul("The customer explicitly requests a refund."), } state = json.dumps({"message": "I was charged twice. Please refund the duplicate payment."}) result = model.predict(state, questions) print(result["answers"]) ``` The direct Core Python API takes a **string** state. Serialize a dictionary with `json.dumps`, as above. The HTTP API below accepts JSON objects directly. A runnable version is included at `examples/quickstart.py`: ```bash python kodama-core/examples/quickstart.py ``` ### Multiple records and long text ```python results = model.predict_batch([ ("Please refund this duplicate charge.", questions), ("The application crashes every time I log in.", questions), ]) ``` Each result matches the corresponding input item. Questions across records are batched within a token budget. Batch-result `latency_ms` covers the whole batch call, not an independent per-record measurement. The checkpoint uses a 2,048-token context, including the question and option text. `long_state="chunk"` reads overlapping windows; the default `"truncate"` mode truncates long states. Chunk pooling does not reproduce a single unlimited-context read, and escalation/calibration should be revalidated for long inputs. For CPU execution, load with `device="cpu"`; CPU speed has not been benchmarked here. The GPU benchmark used float32 model weights with BF16 autocast. ## Understanding the answers Every response contains `model`, `latency_ms`, and `answers`, keyed by your question names. | Question type | How to define it | Main output | |---|---|---| | `Choice` | 2–255 distinct labels, optionally with descriptions | `choice` and a label-to-probability dictionary | | `Score` | 2–10 ordered descriptions, lowest first | Zero-based `level`, its `label`, a probability list, and expected `score` scaled to 0–10 | | `Noul` | A statement to test; no answer options | `noul`, the probability that the statement is true | `confidence` is the largest answer probability. For `Noul`, a confident **false** answer can have `noul=0.02` and `confidence=0.98`; use `noul` when deciding whether the statement is true. `escalate` and `escalate_prob` help flag uncertain answers for another step in your application. Neither confidence nor escalation is a guarantee of correctness. Use specific instructions and descriptive, distinct options. Probabilities are conditional on the supplied options; adding an `other` option can be useful when your label set does not cover every input, but the model's use of that option still needs validation. Core and Vision apply saved temperature calibration by default. Their `escalate_prob` comes from a separately trained escalation head with saved calibration; default escalation is above 0.5. Set `calibrated=False` on prediction calls to inspect uncalibrated outputs. Their answer dictionaries do not include Sense's per-answer `calibrated` flag. ## Serve a local HTTP API ```bash python -m pip install "./kodama-core/runtime[server]" ``` ```bash python -m jevlike.server --checkpoint kodama-core --port 8077 ``` From another terminal: ```bash curl http://127.0.0.1:8077/v1/systemone \ -H 'Content-Type: application/json' \ -d '{"state":{"message":"Please refund the duplicate charge."},"questions":{"team":{"type":"choice","instructions":"Which team?","criteria":["billing","technical","other"]}}}' ``` The server binds to localhost by default. `GET /health` reports readiness. Model loading happens once at startup. This repository does not supply an active hosted inference endpoint. ## Local evaluation Measured on an NVIDIA RTX 4050 Laptop GPU (6 GB). Latency includes tokenization and excludes model loading and HTTP. | Dataset | Decisions | Accuracy | |---|---:|---:| | Short news/finance | 140 | 63.57% | | Long news/finance | 56 | 71.43% | | Other held-out tasks | 300 | 58.00% | | Typed decisions | 2,000 | 45.95% | Median short-text latency: **26.5 ms** for one question and **104.8 ms** for ten questions (20 timed calls per workload). Batch throughput: **94.6 decisions/s**. These are local exploratory measurements, not service guarantees. ## Timing details and interpretation | Text workload | Median | p95 | |---|---:|---:| | short 1q | 26.5 ms | 29.0 ms | | short 5q | 60.1 ms | 65.5 ms | | short 10q | 104.8 ms | 119.2 ms | | short 50q | 682.2 ms | 754.4 ms | | long 1q | 123.9 ms | 131.8 ms | | long 5q | 685.8 ms | 728.8 ms | | long 10q | 1321.1 ms | 1337.0 ms | Twenty timed calls per workload, three warmups, four CPU threads, and CUDA synchronization. Short evidence strings contain approximately 80–120 ModernBERT tokens. Long timing strings are the same shared strings across backends, derived from 1,500 ModernBERT tokens. Batch throughput used 16 states × three questions and three repetitions. Peak PyTorch allocation in the comparison run was 2.80 GiB (allocator reservation 2.90 GiB); this is not a universal minimum VRAM requirement. The desktop shared the GPU, and clocks were not locked. On the same earlier benchmark, standard Laya measured 30.8 ms for one short-text question and 118.5 ms for ten, compared with Core's 26.5/104.8 ms. Accuracy was task-dependent: Core scored 58.0% versus Laya's 66.0% on other held-out tasks. Laya's specialized checkpoint scored 74.25% on typed decisions versus Core's 45.95%; that checkpoint was trained for the task family. The results do not establish that Core universally outperforms Laya. The 2,496 accuracy decisions use identical normalized states/questions across models. No calibration was refitted on that test set. Small news/finance groups have substantial sampling uncertainty. These local normalized tasks are not a reproduction of publisher headline benchmark scores. See [evaluation/comparison.json](evaluation/comparison.json) for the recorded protocol, model accuracies, and Core's timing samples. [Run metadata](evaluation/run_metadata.json) records the hardware, precision, and memory measurements. ## Training and limitations The base encoders were pretrained by their upstream authors. Kodama adds local training/adaptation; these foundation models were not pretrained from scratch. Included training metadata describes the local run. Evaluation data are not redistributed in this repository. Public pretraining overlap is unknown, and small benchmark samples should not be treated as broad capability guarantees. Accuracy depends on task and option wording. Confidence is not a guarantee of correctness; use suitable validation before relying on decisions. ## Attribution and licensing The pretrained base components identify their license as Apache-2.0. Upstream cards, source revisions, and the Apache license text are preserved in `third_party/` and `NOTICE`. No separate public license has been selected for the original Kodama additions and runtime in this release; the upstream license statements should not be read as a new license grant for those additions. Base sources: - [answerdotai/ModernBERT-large](https://huggingface.co/answerdotai/ModernBERT-large/tree/45bb4654a4d5aaff24dd11d4781fa46d39bf8c13) ## Troubleshooting and reproducibility - **401/403 while downloading:** authenticate with an account allowed to read the repository. - **Model configuration / AutoModel error:** use the bundled `jevlike` runtime and the custom loader shown above. - **Import or tokenizer mismatch:** use the documented Transformers version and install the runtime extras for this model. - **Different answers near a tie:** numerical precision and batch shapes can affect probabilities; measure your own task before changing inference settings. `release_manifest.json` records file hashes and pinned base revisions. The runtime is bundled under `runtime/`; `examples/quickstart.py` provides a complete local example. The pretrained foundations were not trained from scratch by Kodama. Earlier evaluation artifacts use `jevlike` for Core, `omni` for Sense, and `prototype` for Vision.