qwen3-0.6b-int8
A Core AI bundle of Qwen/Qwen3-0.6B for the Aether SDK (iOS and macOS 27+).
- Source:
Qwen/Qwen3-0.6Bat revisionc1899de289a04d12100db370d81485cdf75e47ca, licence Apache-2.0. The licence is included asLICENSE. - Changes from the source: converted from PyTorch to Core AI (
.aimodel) by Aether forge (recipeqwen3-0.6b-int8@1). Weights are int8-linear-perchannel (8-bit weights). The tokenizer files are the source's own.
Variants
| Variant | Platform | Arch | Compute | Compiled | Assets | Download |
|---|---|---|---|---|---|---|
macos-any-gpu |
macos | any | gpu | no (specialized on first load) | qwen3_0_6b_int8.aimodel 597.5 MB |
613.4 MB |
ios-any-gpu |
ios | any | gpu | no (specialized on first load) | qwen3_0_6b_int8.aimodel 597.5 MB |
613.4 MB |
ios-h18p-gpu |
ios | h18p | gpu | yes | qwen3_0_6b_int8.h18p.aimodelc 597.5 MB |
613.4 MB |
Verification
Every row is a record in verification/ about exactly these bytes (matched by bundle digest). Reference rows
are strict T2 passes of the unquantized export on the same fixture, in verification/reference/.
| Variant | Tier | Result | Detail | Device | OS build | Compute | Record |
|---|---|---|---|---|---|---|---|
ios-any-gpu |
T0 | pass | iPhone18,2 | 24A446 | target | 6f934189 |
|
ios-any-gpu |
T2 | pass | 19/19 strict; profile quantized-8bit; fixture cc6fd71af4ed3af5 |
iPhone18,2 | 24A446 | target | c8d2ae39 |
ios-any-gpu |
T3 | pass | copy-fidelity-v1; 100.0% vs reference 100.0%; 50 items |
iPhone18,2 | 24A446 | target | 3ec55862 |
ios-h18p-gpu |
T0 | pass | iPhone18,2 | 24A446 | target | 1858b220 |
|
ios-h18p-gpu |
T2 | pass | 19/19 strict; profile quantized-8bit; fixture cc6fd71af4ed3af5 |
iPhone18,2 | 24A446 | target | 42f46a81 |
ios-h18p-gpu |
T3 | pass | copy-fidelity-v1; 100.0% vs reference 100.0%; 50 items |
iPhone18,2 | 24A446 | target | b313af26 |
macos-any-gpu |
T0 | pass | Mac17,6 | 26A434 | target | a268e4d2 |
|
macos-any-gpu |
T1 | pass | Mac17,6 | 26A434 | target | fb8d5aab |
|
macos-any-gpu |
T2 | pass | 18/19 strict; 19/19 strict; profile quantized-8bit; budgeted: think-multiply; fixture cc6fd71af4ed3af5 |
Mac17,6 | 26A434 | target | cec042a2 |
macos-any-gpu |
T3 | pass | copy-fidelity-v1; 100.0% vs reference 100.0%; 50 items |
Mac17,6 | 26A434 | target | c965a95c |
| unquantized reference (not published) | T2 | pass | 19/19 strict; 19/19 strict; profile strict; fixture cc6fd71af4ed3af5 |
Mac17,6 | 26A434 | target | e73cf364 |
The quantized profile also requires: T2 strict on the unquantized reference export (met by the reference row).
Use
aether run qwen3-0.6b-int8 --prompt "Hello"
import Aether
let aether = try Aether()
let chat = try await aether.chat("qwen3-0.6b-int8")
let reply = try await chat.respond(to: "Hello")
print(reply.text)