jev-9b-coreml / README.md
alexwengg's picture
Add files using upload-large-folder tool
a718625 verified
|
Raw History Blame Contribute Delete
4.38 kB
metadata
license: apache-2.0
base_model: autotrust/JEV-9B
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: coreml
tags:
  - coreml
  - apple
  - decision-model
  - jev
  - system-one
  - typed-decisions
  - calibrated-probabilities
  - vision
  - qwen3.5
  - fluidinference

JEV-9B with vision (Core ML)

AutoTrust's JEV-9B System 1, with vision, converted to Core ML for Apple silicon. It makes typed decisions over text and images: yes/no (noul), a 0–5 score, or a choice among 2–256 options. Each decision returns a calibrated probability for every option from a single forward pass; nothing is generated.

JEV-9B with vision is the stock Qwen3.5-9B, including its unchanged vision tower, plus JEV-9B's System 1 LoRA and decision head (vl/adapter_vllm). In this repository:

  • The LoRA is merged into the decoder weights.
  • The decision head is reduced to the adapted lm_head rows it reads.
  • The prompt is exactly JEV's serve_decide.py System 1 template.

Runtime: JevManager in FluidUse (Swift, macOS 15+), with a SwiftUI demo where the model plays Super Mario Bros. 1-1 (JevMarioDemo).

Files

File Role Size Precision
part00 … part07.mlpackage decoder with the System 1 LoRA merged, 4 layers per part; functions L256 and L512 (sequence buckets), weights shared 6.5 GB total 8-bit weights (per channel), fp16 compute
VisionTower_P256.mlpackage, VisionTower_P576.mlpackage Qwen3.5-9B vision tower, one image per call, up to 256 / 576 patches (16 px) 1.7 GB each fp32 (fp16 is not accurate enough for this ViT)
pos_embed.f32 the vision tower's learned position grid; the host resamples it per image 10.6 MB fp32
embeddings.f16 input token embeddings (gathered on the host) 2.0 GB fp16
head_rows.f32, head.json adapted lm_head rows at the 24 verbalizer ids (false/true, 0–5, A–P) plus the other single-token option labels; per-slot bias and per-kind temperature 4.3 MB fp32
tokenizer.json, config.json tokenizer; the manifest with shapes, buckets and token ids for the host

Decision:

  1. Run the vision tower once per image.
  2. Splice its tokens into the prompt at the <|image_pad|> positions.
  3. Chain the 8 decoder parts. Their I/O is hidden [1, L, 4096] and cos/sin [L, 64], all fp16, with interleaved M-RoPE tables built on the host.
  4. Take the last prompt row times the head rows, add the bias, divide by the temperature, and softmax over the question's options.

Quality

The fp32 reference is Hugging Face's Qwen3_5VisionModel, plus the same decoder in fp32 with the LoRA merged (streamed 4 layers at a time), plus the head rows. It was run on 12 Super Mario Bros. frames with a choice, a yes/no and a score question on each:

8-bit Core ML vs fp32
answers that differ 0 / 36
max probability difference (choice / noul / score) 0.013 / 0.023 / 0.006
vision tokens, max relative error 1.4e-4

Speed (M5 Pro, 24 GB, GPU; Swift JevManager)

Median
vision tower, one image (256 / 576-patch bucket) ~60 / ~120 ms
decoder, one question (L256 bucket) ~270 ms
one image + 3 questions 0.95 s
load (compiled): 8 decoder parts + 2 vision towers ~40 s

The decoder needs ~7 GB of memory to itself. The Qwen3.5 decoder (Gated DeltaNet layers) does not run on the Neural Engine, so this is a GPU model.

Usage (Swift)

import FluidUse

let jev = try await JevManager.load(from: bundleDirectory, bucket: 256)
let result = try await jev.decide(
    state: [.text("Close-up of the area just ahead of Mario in a Super Mario Bros. level:"), .image(view)],
    questions: [JevQuestion(.noul, "Is there an enemy (a brown Goomba or a Koopa) in this picture?")])
print(result.decisions[0].yes ?? 0, result.totalMilliseconds)

License and credits

Apache-2.0, the same as autotrust/JEV-9B and Qwen/Qwen3.5-9B (LICENSE is Qwen's). All weights belong to AutoTrust and Qwen; this repository only changes their format and precision. Conversion by Fluid Inference.