LFM2.5 Jev-style schema model

This research model is part of the Audio Jev project by odunola. It uses LiquidAI/LFM2.5-350M, LoRA adapters, and a joint schema head. The model reads a state and a set of questions. It gives a probability for each permitted answer in one forward pass.

This checkpoint accepts text and JSON only. It cannot process audio. The project takes ideas from Jev and Clef. This is an independent implementation. TypeSafe, Cloudflare, and Liquid AI did not release this model.

Project purpose

We want a small model that can evaluate a state against specified questions. The questions define the permitted answers. Applications can use the answer probabilities to select their next action.

The long-term objective is to process audio with this type of model. We started with text to test the schema head, training loss, and data conversion. We have not selected or trained an audio input system for this checkpoint.

Work completed

  1. We implemented the LFM model with code for training and inference.
  2. We added a loader for the original Liquid AI weights.
  3. We added a joint schema head for Noul, Choice, and Score questions.
  4. We added a converter for Open-Jev states, questions, and answer distributions.
  5. We trained LoRA adapters and the schema head for 10,000 optimizer updates.
  6. We tested the checkpoint on a local CPU and Apple Silicon MPS.
  7. We built a local browser interface for states, questions, and results.

The browser interface shows the probability of each answer. It also shows the input token count and inference time. The model package below contains the code needed for command-line inference. The project repository contains the browser interface.

Checkpoint

step-10000.pt contains the trained LoRA parameters, schema head, and configuration data. It contains 10,632,452 trained parameters. The file size is approximately 41 MiB. It does not contain the frozen base weights or optimizer state. It is not a standard PEFT adapter.

The loader gets the base weights separately. It checks the adapter and head parameter names before it loads the checkpoint.

This is the final checkpoint from the 10,000-update run. We tested it in the local browser interface. We did not compare all saved checkpoints to identify the best one. The training configuration did not specify a fixed base model revision. Thus, we cannot claim exact reproduction from the recorded configuration alone.

Load and run

Use Python 3.14 to match the development environment. For a private repository, first use hf auth login to sign in. Run these commands to download and load the model package:

hf download odunola/lfm2.5_jev_style --local-dir lfm2.5_jev_style
cd lfm2.5_jev_style
python3.14 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python inference.py example.json --device cpu

On Apple Silicon, use --device mps. For a CUDA installation, use --device cuda. The first run downloads the base model if it is not in the local cache. Use the supplied tokenizer with this checkpoint. The loader uses torch.load(..., weights_only=True).

Use the supplied Python code to load the schema model. A standard AutoModelForCausalLM call cannot load this package directly. The output from inference.py maps question IDs to answer probabilities.

Input schema

{
  "state": "The customer explicitly asks for a refund.",
  "questions": {
    "refund_requested": {
      "type": "noul",
      "instructions": "Does the customer explicitly request a refund?",
      "criteria": {"true": "yes", "false": "no"}
    }
  }
}
  • Noul: true/false outcomes, with optional descriptions.
  • Choice: an object mapping option IDs to descriptions.
  • Score: an ordered array of level descriptions, indexed from zero.

The encoder accepts multiple questions. It rejects prompts with more than 2,048 tokens. It includes question IDs and option IDs in the prompt. Do not change the encoder format without tests for this checkpoint.

The request uses Jev-style question types. The response is a custom probability map. It does not have the official Jev API response format. For Score, calculate the weighted level with sum(index * probability). The supplied loader returns the full answer distribution.

Training

Setting Value
Dataset Open-Jev, release-v2-redistributable
Dataset revision c67699e13d0ae25e35b77165a4b6b079bedc8aba
Optimizer updates 10,000
Micro-batch / accumulation 8 / 1
Learning rate 0.0001
LoRA rank / alpha 8 / 16
Head width / attention heads 256 / 4
Routing layers / head layers 2 / 4
Feedforward width / dropout 1,024 / 0
Loss Cross-entropy with 0.1 label smoothing + Brier, weight 1
Precision BF16 autocast during GPU training

Training uses supervised targets, including probability distributions. It does not use RLCR, GRPO, or another reinforcement learning method. training_metadata.json contains the recorded training settings. Its epoch setting is a configuration value, not a verified count of completed epochs.

Evaluation and limitations

We tested inference on a local CPU and MPS. One small test used eight Open-Jev validation records. The model selected the correct answer for all 46 questions with one correct label. Two other questions had probability distributions as targets. This small test does not establish general accuracy.

We also changed the wording of two states without changing their intended meaning. Correct answers decreased from 10 of 10 to 8 of 10. This result shows that wording can affect the model's answers.

We tried hospital intake questions in the browser interface. These informal tests do not establish medical accuracy. The model is not validated for diagnosis, clinical triage, or treatment.

A high answer probability does not guarantee a correct answer. The tests do not establish that the probabilities are calibrated. Missing information does not mean that a condition is absent.

Next experiment

We plan to test calibration and responses to missing evidence. Calibration measures agreement between predicted probabilities and observed outcomes. The experiment will compare four loss settings:

Setting Label smoothing Brier weight
A 0 0
B 0 1
C 0.1 0
D 0.1 1

All settings will use cross-entropy. Each run will start from the same checkpoint. We plan three seeds and 1,000 optimizer updates per run.

The test states will contain complete, missing, negated, or conflicting evidence. We will keep related states and wording templates in the same data split. We will measure accuracy, Brier score, and errors with high predicted probability. We will also check for loss of performance on the original task.

This plan takes its calibration objective from the RLCR paper. It uses supervised training, not the paper's reinforcement learning procedure. This experiment has not started.

We also plan separate work on Jev API compatibility. Our encoder sends question IDs to the model. The official Jev interface does not send those IDs to the model. We will test any encoder change separately from the loss experiment.

Files and provenance

  • step-10000.pt: trained adapters and schema head.
  • tokenizer/: tokenizer used for local inference.
  • audio_jev/, inference.py: the matching minimal inference implementation.
  • example.json: an Open-Jev validation prompt in the training format, without targets.
  • training_metadata.json, SHA256SUMS: recorded settings and checkpoint checksum.
  • Project: https://github.com/odunola499/audio-jev

This package does not redistribute the frozen base weights. The base model and its derivatives are subject to Liquid AI's upstream license. Refer to the Open-Jev dataset card for the data source and license information. This model card does not grant additional rights to upstream materials.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for odunola/lfm2.5_jev_style

Adapter
(42)
this model

Dataset used to train odunola/lfm2.5_jev_style

Paper for odunola/lfm2.5_jev_style