OneJev 0.8B, Core ML W16A16, Neural Engine, 2K context

OneJev answers typed questions about what it is shown: which of several named options applies (Choice), whether a statement is true (Noul), or where something sits on a scale (Score), over text, screenshots and still frames, with a probability for every option. This is OneJev-0.8B converted to Core ML so that the whole model runs on the Mac's Neural Engine: 33 functions, 37,809 operations, none planned for the CPU or GPU.

Source: OmniJev/OneJev-0.8B, revision c3939d8bf4cad34549a2b13bbb6aee9bcb6afee8. Weights, activations and recurrent state are all FP16 (W16A16). The token-embedding lookup (an FP16 table) and the answer head (FP32) run on the host. This is not an official OmniJev release.

Releases

The same functions are published at two precisions:

Repository Precision Largest probability error
1of2/onejev-0.8b-coreml-w8 W8A16 0.0427
1of2/onejev-0.8b-coreml-w16 (this repository) W16A16 0.0210

Requirements

  • An Apple-silicon Mac with macOS 15 or later. Validation used an Apple M4 Max on macOS 27.0.1; other combinations are not qualified.
  • A host that implements the execution contract below. This repository holds the model, not a runtime.
  • Disk for the compiled programs, about 1.2 GB. The first process to load a function compiles it for the Neural Engine; later processes reuse it as long as the compiled program stays at the same path. With the converter project's runtime, the first process on an empty cache answered its first request after 328 s, a later process after 3.3 s. macOS may purge compiled functions from its cache when the disk runs low; a process then compiles those again, about 12 s each.

Download

hf download 1of2/onejev-0.8b-coreml-w16 --local-dir onejev-0.8b

The download is 1.7 GB. model/ is the sealed artifact: model/manifest.json describes every function and lists the SHA-256 of every file; verify them before loading, and keep the directory structure intact. Pin a Hub revision in production rather than following main.

Functions

Programs Functions
models/decoder_00_04.mlpackage to decoder_20_24: six shards of four layers b1_t32, b1_t128, b4_t32, b4_t128: 1 or 4 rows, 32 or 128 tokens a call
models/vision_00_04.mlpackage to vision_08_12: three shards of four blocks, the patch merger in the last p256, p1024, p4096: patches per image

manifest.json gives each function's inputs and outputs (names, shapes, all FP16), the layer schedule (every fourth layer is full attention, the others Gated DeltaNet), the answer slots and the upstream commits.

Execution contract

  • Prompt. Render each request with the upstream processor in model/processor/ and qev's prompt renderer at the commit the manifest records (source.qev_commit), prompt version qev-labels-v2, thinking off. The request's state and media are one shared prefix; each question is a suffix after it. At most 2,048 tokens, prefix, media and question together; refuse longer requests rather than truncate them.
  • Embedding and positions. Look tokens up in host/embedding.f16 (little-endian FP16, [248320, 1024]). Build rotary tables from the upstream multimodal positions in FP32 and narrow them to FP16: cos and sin [B,T,64].
  • Decoder. Run each chunk of 32 or 128 tokens through the six shards in order: the prefix at batch 1, then the questions at batch 4, each row forked from the prefix's state. A call takes hidden [B,T,1024], cos, sin, three controls and each of its layers' state, and returns hidden_out and the next state:
    • live [B,T]: 1 for a real token, 0 for padding.
    • mask [B,T,2048+T]: 0 where a query may attend (its row's cached prefix, then causally within the chunk), -30000 elsewhere.
    • select [B,3,T+32]: one-hot rows that pick the next convolution history, the three rows ending at the row's last real token. Columns below T index the chunk; columns from T index the previous history.
    • Gated DeltaNet layers carry conv [B,3,6144] (channels last) and ssm [B,16,128,128], the recurrent state at 64 times the upstream scale. Full-attention layers take key and value [B,2,2048,256] and return only the chunk's new keys and values, which the host writes at each row's next positions. Commit a chunk's state only after all six shards succeed; a padded row's state comes back unchanged.
  • Vision. Per image, or per temporal group of a clip: the processor's patches in merge-major order, padded to 256, 1,024 or 4,096, with kmask [1,N] (0 for a patch, -30000 for padding). The first shard also takes position, the learned table in host/vision_positions.f32 interpolated bilinearly (align corners) to the patch grid; cos and sin are spatial rotary tables from integer patch coordinates. The last shard returns one row per four patches, which replace the image tokens' embeddings.
  • Answer. At a question's last token, apply the final RMS norm (eps 1e-6, scale 1 + host/norm.f32), round to FP16, and project with host/head.f32 [256, 1024], the rows of the answer-slot tokens (slot_ids). A softmax over the question's slots gives its probabilities; Choice, Noul and Score answers follow qev's formulas. Temperatures are not fitted.

Selecting .cpuAndNeuralEngine does not by itself guarantee Neural Engine execution. The functions were checked with Core ML compute plans on the validation host; check the plan on your own device if you need the guarantee.

Quality

Measured through the converter project's Python runtime, against the upstream model run whole by Transformers in float32 on the GPU, over 28 requests and 165 decisions: synthetic coverage fixtures (every answer type, Unicode, ten questions in one request, an image, a four-frame clip, a prefix of three quarters of the context) and four screen and meeting requests.

Requests Decisions Top-1 agreement Largest probability error Largest total variation
Coverage fixtures 139 100% 0.0210 0.0210
Document stretch 7 100% 0.0155 0.0155
Meeting piece 5 100% 0.0113 0.0113
Screenshot, 256 patches 7 100% 0.0202 0.0202
Screenshot, 720p 7 100% 0.0198 0.0198

Over every decision: top-1 agreement 100%, largest probability error 0.0210, largest total variation 0.0210, largest expected-score error 0.0373. This release's limits are 0.025 for probability error, 0.03 for total variation and 0.05 for expected-score error, with at least 99% top-1 agreement. A request repeated through its cached prefix gives identical probabilities. This release's error comes from FP16 arithmetic alone, which the Neural Engine uses throughout; the INT8-weight release reaches at most 0.0427 on the same decisions. The project's native Swift runtime agrees with its Python runtime to 0.0075 over the same 165 decisions.

Temperatures are not fitted (identity calibration). The coverage fixtures check the conversion, not task accuracy.

Speed

On an Apple M4 Max on macOS 27.0.1, with the converter project's Python runtime in a warm process (medians of 20):

Request Tokens Questions New stretch Same stretch, new questions
Document stretch 808 7 0.61 s 0.49 s
Meeting piece 522 5 0.43 s 0.31 s
Screenshot, 256 patches 881 7 0.69 s 0.49 s
Screenshot, 720p 1,697 7 1.79 s 0.49 s

New stretch reads the shared state and answers every question; same stretch, new questions reuses the state already read.

Files

  • model/: the sealed artifact. models/ holds the nine multifunction Core ML packages, host/ the embedding table, answer head, final norm and vision positions, processor/ the upstream tokenizer, chat template and processor configuration.
  • qualification.json: every number on this card, the compute plan's placement, the software versions and the host.
  • LICENSE (Apache 2.0), NOTICE.

License

Apache 2.0, as the source model, OneJev-0.8B by the OmniJev team, a fine-tune of Qwen3.5-0.8B. See NOTICE.

@misc{onejev2026,
  title        = {{OneJev}: A Multimodal System One Decision Model},
  author       = {{OmniJev Team}},
  year         = {2026},
  howpublished = {\url{https://github.com/OmniJev/OneJev}}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 1of2/onejev-0.8b-coreml-w16

Quantized
(5)
this model