File size: 2,504 Bytes
05b2581
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
---
license: mit
base_model: microsoft/Phi-4-mini-instruct
pipeline_tag: text-generation
tags:
- aether
- core-ai
- apple
- macos
- ios
---

# phi-4-mini-instruct

A Core AI bundle of [microsoft/Phi-4-mini-instruct](https://huggingface.co/microsoft/Phi-4-mini-instruct) for the Aether SDK (iOS and macOS 27+).

- **Source:** `microsoft/Phi-4-mini-instruct` at revision `cfbefacb99257ffa30c83adab238a50856ac3083`, licence MIT.
  The licence is included as `LICENSE`.
- **Changes from the source:** converted from PyTorch to Core AI (`.aimodel`) by Aether forge (recipe
  `phi-4-mini-instruct@2`). Weights are int8-linear-perblock32 (8-bit weights). The tokenizer files are the source's own.

## Variants

| Variant | Platform | Arch | Compute | Compiled | Assets | Download |
|---|---|---|---|---|---|---|
| `macos-any-gpu` | macos | any | gpu | no (specialized on first load) | `phi_4_mini_instruct.aimodel` 4.08 GB | 4.1 GB |
| `ios-h18p-gpu` | ios | h18p | gpu | yes | `phi_4_mini_instruct.h18p.aimodelc` 4.08 GB | 4.1 GB |

`ios-h18p-gpu` needs the `com.apple.developer.kernel.increased-memory-limit` entitlement.

## Verification

Every row is a record in `verification/` about exactly these bytes (matched by bundle digest). Reference rows
are strict T2 passes of the unquantized export on the same fixture, in `verification/reference/`.

| Variant | Tier | Result | Detail | Device | OS build | Compute | Record |
|---|---|---|---|---|---|---|---|
| `ios-h18p-gpu` | T0 | pass |  | iPhone18,2 | 24A446 | target | `799a7472` |
| `ios-h18p-gpu` | T2 | pass | 20/20 strict; profile quantized-8bit; fixture `e6d8aa51e5bbd5c4` | iPhone18,2 | 24A446 | target | `56352841` |
| `macos-any-gpu` | T0 | pass |  | Mac17,6 | 26A434 | target | `f8ec071b` |
| `macos-any-gpu` | T1 | pass |  | Mac17,6 | 26A434 | target | `06f8b021` |
| `macos-any-gpu` | T2 | pass | 20/20 strict; 20/20 strict; profile quantized-8bit; fixture `e6d8aa51e5bbd5c4` | Mac17,6 | 26A434 | target | `d86a849c` |
| unquantized reference (not published) | T2 | pass | 20/20 strict; 20/20 strict; profile strict; fixture `e6d8aa51e5bbd5c4` | Mac17,6 | 26A434 | target | `998a7aa8` |

The quantized profile also requires: T2 strict on the unquantized reference export (met by the reference row).

## Use

```sh
aether run phi-4-mini-instruct --prompt "Hello"
```

```swift
import Aether

let aether = try Aether()
let chat = try await aether.chat("phi-4-mini-instruct")
let reply = try await chat.respond(to: "Hello")
print(reply.text)
```