Laya for mobile: Core ML and ONNX conversions

This repository holds Core ML and ONNX files of Laya for iOS and Android. The files run the English Laya checkpoint at revision 55cf4c4e without PyTorch.

This is an unofficial conversion. It is not affiliated with or endorsed by Convai Innovations. The original model is convaiinnovations/laya, licensed under the Apache License 2.0 as its model card states. The runtime and the training code are at NandhaKishorM/laya. This repository uses the same license. See LICENSE and NOTICE.

The conversion changes only the file format and the numeric precision. It does not retrain or fine-tune the model. All measurements in this card were made on 2026-09-30 and 2026-10-01, unless the text names another source.

What Laya is

Laya is a non-generative decision model by Convai Innovations. You give it a state (text or JSON) and typed questions. It returns a probability for every option of every question in one forward pass, and it never generates text. The English checkpoint is a ModernBERT-large encoder (395M parameters) with a decision head: two transformer layers, an option scorer, and an act head. The total is 421M parameters. Laya scores each option at its own [MASK] token, so the options are defined at request time and need no retraining. The upstream card describes the RLCD training and the benchmarks.

These files were made for an on-device agent router app. The examples in this card use a customer-support question with five teams. The ONNX files answer any Laya question. The Core ML file answers any choice question with up to 5 options.

Files

Path What it is Size Use it for
coreml/laya-fp16-512.mlpackage.zip Core ML MLProgram, fp16, one fixed input length of 512 tokens, iOS 16 or later. One choice question with up to 5 options. 843,287,672 bytes iOS and macOS apps. Download this one file and unzip it.
coreml/laya-fp16-512.mlpackage/ The same package, not zipped 843,287,102 bytes Browse it in Xcode, or add it to an app target.
onnx/laya-q8.onnx ONNX, opset 17. Every MatMul weight is 8-bit (MatMulNBits, asymmetric, block 32, fp32 compute). 653,610,034 bytes Android, and any ONNX Runtime CPU target where memory matters
onnx/laya-fp32.onnx ONNX, opset 17, fp32. The full graph of the upstream DecisionModel, all question types, dynamic batch, length, and option count. 1,686,064,349 bytes Desktop and server, or a reference for your own conversion
tokenizer.json The upstream tokenizer (tokenizer/tokenizer.json), unchanged bytes 3,583,228 bytes Token ids for every file
rl_agent_config.json The upstream config with the calibrated temperatures, unchanged bytes 745 bytes The decode step
examples/ laya_sequence.py (the sequence builder and the example question), route_onnx.py, route_coreml.py, RouteCoreML.swift Runnable examples, all tested
tests/vectors.json 46 test cases for the example question: the state, the token ids, the marker positions, and the probabilities and the choice of the laya package 59,365 bytes Parity tests for your tokenizer and your runtime
tests/check.py Runs the cases of tests/vectors.json through the ONNX files and the Core ML package 7,298 bytes Reproduce the accuracy table
tests/make_vectors.py Writes tests/vectors.json with the laya package 7,605 bytes Change or extend the test cases
scripts/ The export and quantization scripts Reproduce the files

Which file to pick

Platform File Runtime and settings Why
iOS coreml/laya-fp16-512.mlpackage.zip Core ML, MLModelConfiguration.computeUnits = .cpuAndNeuralEngine The Neural Engine is the fastest place for Laya on a phone, and it uses the least memory. It needs one fixed input length.
iOS, if the Neural Engine path fails on a device the same file .cpuAndGPU It has no argmax or 0.40 side changes on the Mac (see Accuracy).
Android onnx/laya-q8.onnx ONNX Runtime CPU execution provider Peak memory is about half of the fp32 file.
Android, latency more important than memory onnx/laya-fp32.onnx ONNX Runtime CPU Faster on CPU, but twice the memory
Desktop, server, Python the upstream laya package PyTorch It is the reference implementation.

Evidence from a real phone

A test of Laya Core ML packages on an iPhone 13 Pro Max is in OpenJevSwift pull request 85. It measured one fp16 package per fixed length on the Neural Engine:

Tokens 128 256 512 1,024
Neural Engine, per question 27.9 ms 57.1 ms 137 ms 513 ms

The memory footprint was 106 MB or less, load included. The GPU was 2.6 to 2.9 times slower (82 to 1,346 ms) and used 855 MB. The same test found that a package with several lengths (a multifunction package) does not load on the Neural Engine. That test used the laya-typed-decisions checkpoint, which has the same ModernBERT-large architecture and 421M parameters. Its max probability difference to PyTorch fp32 was 0.0181, with 199 of 200 top answers unchanged.

The Core ML file in this repository uses the graph rewrites of that pull request, so the Neural Engine accepts it. It is not measured on a phone yet. On the private set of 67 routing inputs (see Accuracy), the median length is 134 tokens, 90% of inputs are 234 tokens or less, and the maximum is 512. One 512-token package covers every input. The phone test above measured 137 ms at 512 tokens.

No published measurement of ModernBERT-large or Laya on an Android phone was found.

Measurements on a Mac, a simulator, and an emulator (not phones)

These numbers were measured on 2026-09-30 and 2026-10-01 on an Apple M4 Mac (10 CPU cores, 16 GB RAM, macOS 26.6.2). They are not phone numbers. Other processes ran on the Mac during the runs, and the 1-minute load average was 2.5 to 3.3. Latency is the median and p90 of 30 predictions after one warm-up.

Core ML, laya-fp16-512.mlpackage, coremltools 9.0 from Python:

Compute units Median p90 First load Second load
CPU_AND_NE 71.3 ms 71.7 ms 11.2 s 0.14 s
CPU_AND_GPU 163.3 ms 165.6 ms 3.2 s 0.36 s
CPU_ONLY 153.2 ms 158.3 ms 2.6 s 0.22 s

The input was a message of 131 tokens, padded to 512. A 512-token input gave the same times, because the package always computes 512 positions. scripts/export_laya_coreml.py --skip-export times the 131-token case conv_lamp of tests/vectors.json. On 2026-10-01 it gave a CPU_AND_NE median of 71.4 ms in two runs. The CPU_AND_GPU median was 164 ms in one run and 206 ms in the other, so the GPU time changes with the load of the Mac. The first load on the Neural Engine includes the Neural Engine compile. Core ML caches the result, so later loads are fast. The Core ML compute plan (MLComputePlan) puts 1,178 of 1,190 operations on the Neural Engine for CPU_AND_NE. The 12 operations on the CPU build the masks and read the embeddings. The same file from Swift (examples/RouteCoreML.swift) gave the same plan.

ONNX Runtime and the Android emulator, measured on 2026-09-30:

File Runtime Median Peak memory
laya-fp32.onnx ORT 1.23.0 CPU, M4 Mac 448 ms not measured
laya-q8.onnx ORT 1.23.0 CPU, M4 Mac 601 ms not measured
laya-fp32.onnx ORT 1.23.0 CPU, Pixel 8 emulator (4 vCPU) 512 ms 2,136 MB RSS
laya-q8.onnx ORT 1.23.0 CPU, Pixel 8 emulator (4 vCPU) 917 ms 1,155 MB RSS

The emulator guest reports SME2 from the M4 host. A real Pixel 8 has no SME, so measure on your own phones.

Accuracy

The reference is the laya package 0.3.22 at the pinned revision, rounded to 4 places. The probability is softmax(logits / 1.76015) over the 5 options.

An app that routes with these files should check three things against the reference. Each case must keep the same top option (argmax). The max probability error must stay small. An app that adds a second route when the top probability is below 0.40 should also check that no case changes side of 0.40.

Public test vectors

These results are for the 46 cases of tests/vectors.json. The package made these vectors on the CPU in fp32. python tests/check.py reproduces the results. It checks every file whose runtime is installed.

File Runtime (macOS 26.6.2, M4) Max probability error Argmax changes 0.40 side changes
laya-fp16-512.mlpackage Core ML CPU_AND_NE 0.0069 0 0
laya-fp16-512.mlpackage Core ML CPU_AND_GPU 0.0071 0 0
laya-fp16-512.mlpackage Core ML CPU_ONLY 0.0062 0 0
laya-q8.onnx ONNX Runtime 1.23.0 CPU 0.0167 0 0
laya-fp32.onnx ONNX Runtime 1.23.0 CPU 0.00005 0 0

Core ML ran through coremltools 9.0 with numpy 2.3.5. For laya-q8.onnx, the largest error is on list_long, a conversation that fills all 512 tokens (0.5396 against 0.5563). The next largest error of that file is 0.0108. The vectors round to 4 places, so an error near 0.00005 is the floor. 15 of the 46 cases have a top probability below 0.40. Three cases (near_ship_cost, near_student_discount, near_order_page_error) have a top probability within 0.011 of 0.40, so the side check can fail.

Private routing set

The earlier gate of these files used a private set of 67 routing inputs and 55 labeled routing cases. That set is not in this repository, so you cannot reproduce these numbers. The routing column counts the labeled cases whose top option is a correct route. The upstream PyTorch model on the CPU in fp32 also gives 49 of 55.

File Runtime (macOS 26.6.2, M4) Max probability error Argmax changes 0.40 side changes Routing (private)
laya-fp16-512.mlpackage Core ML CPU_AND_NE 0.0061 0 0 49/55
laya-fp16-512.mlpackage Core ML CPU_AND_GPU 0.0018 0 0 49/55
laya-fp16-512.mlpackage Core ML CPU_ONLY 0.0123 0 1 49/55
laya-q8.onnx ONNX Runtime 1.23.0 CPU 0.0086 0 0 49/55
laya-fp32.onnx ONNX Runtime 1.23.0 CPU 0.00005 0 0 49/55

On CPU_ONLY, one input moves from 0.3952 to 0.4027 and crosses 0.40. The other 66 inputs stay on their side. The reference of this set ran on the Apple GPU (MPS) in fp32.

Interface

Core ML: laya-fp16-512.mlpackage

Name Type Shape Content
input_ids (input) int32 [1, 512] The token ids, right-padded with the pad token id 50283
attention_mask (input) int32 [1, 512] 1 for a real token, 0 for padding
marker_pos (input) int32 [1, 5] The position of the [MASK] marker of each option
logits (output) float32 [1, 5] One raw score per option, before temperature
  • MLProgram, fp16 compute, minimum deployment target iOS 16 (specification version 7).
  • The graph answers one choice question. The question type embedding is fixed to choice.
  • Each logit depends only on its own marker position. For fewer than 5 options, put 0 in the unused marker_pos slots and ignore their logits. For more than 5 options, use an ONNX file.
  • The attention masks come from attention_mask. A query and key pair is masked when either position is padding, in the global layers, the sliding-window layers (64 tokens on each side), and the two head layers. The padding does not change the logits of the real tokens.

ONNX: laya-q8.onnx and laya-fp32.onnx

Name Type Shape Content
input_ids (input) int64 [rows, tokens] The token ids. Pad rows of a batch with 50283.
attention_mask (input) int64 [rows, tokens] 1 for a real token, 0 for padding
marker_pos (input) int64 [rows, options] The marker position of each option
marker_mask (input) bool [rows, options] True for a real option
qtype (input) int64 [rows] 0 for choice, 1 for score, 2 for noul
logits (output) float32 [rows, options] One raw score per option, before temperature. Masked options get -10,000.
act_logits (output) float32 [rows, 2] The act head. The upstream card says that it carries no usable signal yet.

The inputs and outputs match laya.onnx_agent.ONNXAgent of the laya package.

The token sequence

Laya reads one sequence per question. examples/laya_sequence.py builds it with the tokenizers package only. It gives the same ids and markers as the laya package on all 46 cases of tests/vectors.json.

[CLS] "choice question: <instructions>" [SEP]
[MASK] " <label 1>: <description 1>"
[MASK] " <label 2>: <description 2>"
...
[SEP] <state> [SEP]
Token Id
[CLS] 50281
[SEP] 50282
[PAD] 50283
[MASK] 50284

The rules come from laya.common.build_sequence (laya 0.3.22):

  1. Tokenize every piece with add_special_tokens=False. Replace the text [MASK] in the input with a space first.
  2. Give each option at most 48 tokens after its [MASK] token.
  3. Give the question part (instructions and options) at most 192 tokens. If the options leave fewer than 16 tokens for the instructions, cut every option to max(4, (192 - 16) // options) tokens.
  4. Cut the instructions to max(8, 192 - option tokens) tokens.
  5. Give the state the rest of the 512 tokens, minus one for the last [SEP]. Cut a text or JSON state at the end. Cut a list state (a conversation, newest last) at the start.
  6. A JSON state is json.dumps(state, ensure_ascii=False).

The marker position of an option is the index of its [MASK] token.

Decode

t = rl_agent_config.json["temperature_by_options"]["choice:<bucket>"]
p = softmax(logits[:k] / t)

The bucket depends on the option count k: 2, 3-5, 6-10, or 11+. The laya package clamps every temperature to the range 0.5 to 5.0. Only choice:11+ (0.1006) is outside the range, so it becomes 0.5. When a bucket is missing, use temperature[qtype].

Bucket Temperature after clamping
choice:2 1.90636
choice:3-5 1.76015
choice:6-10 1.00002
choice:11+ 0.5 (shipped as 0.10058)

Usage

Python with ONNX Runtime

pip install onnxruntime==1.23.0 tokenizers numpy
python examples/route_onnx.py "I paid for express delivery but you sent it to my old address."

The core of examples/route_onnx.py:

import numpy as np
import onnxruntime as ort
from laya_sequence import INSTRUCTIONS, OPTIONS, build, probabilities, temperature
from tokenizers import Tokenizer

tok = Tokenizer.from_file("tokenizer.json")
session = ort.InferenceSession("onnx/laya-q8.onnx", providers=["CPUExecutionProvider"])

ids, markers = build(tok, INSTRUCTIONS, OPTIONS, message)
feed = {
    "input_ids": np.asarray([ids], np.int64),
    "attention_mask": np.ones((1, len(ids)), np.int64),
    "marker_pos": np.asarray([markers], np.int64),
    "marker_mask": np.ones((1, len(markers)), bool),
    "qtype": np.zeros((1,), np.int64),  # 0 = choice
}
logits = session.run(["logits"], feed)[0][0, :len(OPTIONS)]
p = probabilities(logits, temperature("rl_agent_config.json", len(OPTIONS)))

Output on the M4 Mac, ONNX Runtime 1.23.0. The message is the case pay_wrong_address of tests/vectors.json, so the example also prints what the laya package returns.

112 tokens, markers [12, 29, 47, 64, 81]
shipping   0.6160  (laya 0.6077)
billing    0.1246  (laya 0.1284)
other      0.1014  (laya 0.1029)
technical  0.0946  (laya 0.0965)
account    0.0634  (laya 0.0646)

Python with Core ML (macOS)

pip install coremltools==9.0 tokenizers numpy==2.3.5
python examples/route_coreml.py "I paid for express delivery but you sent it to my old address."

Output on the M4 Mac, CPU_AND_NE:

112 tokens, markers [12, 29, 47, 64, 81]
shipping   0.6071  (laya 0.6077)
billing    0.1282  (laya 0.1284)
other      0.1032  (laya 0.1029)
technical  0.0965  (laya 0.0965)
account    0.0650  (laya 0.0646)

Swift with Core ML

This is the core of examples/RouteCoreML.swift. Compile the package once after the download and keep the .mlmodelc directory.

import CoreML

func multiArray(_ values: [Int32], count: Int, fill: Int32) throws -> MLMultiArray {
    let array = try MLMultiArray(shape: [1, NSNumber(value: count)], dataType: .int32)
    let pointer = array.dataPointer.bindMemory(to: Int32.self, capacity: count)
    for i in 0..<count { pointer[i] = i < values.count ? values[i] : fill }
    return array
}

let compiled = try MLModel.compileModel(at: packageURL)  // laya-fp16-512.mlpackage
let configuration = MLModelConfiguration()
configuration.computeUnits = .cpuAndNeuralEngine
let model = try MLModel(contentsOf: compiled, configuration: configuration)

// ids and markers come from a port of examples/laya_sequence.py.
let features = try MLDictionaryFeatureProvider(dictionary: [
    "input_ids": try multiArray(ids, count: 512, fill: 50283),
    "attention_mask": try multiArray([Int32](repeating: 1, count: ids.count), count: 512, fill: 0),
    "marker_pos": try multiArray(markers, count: 5, fill: 0),
])
let logits = try model.prediction(from: features).featureValue(for: "logits")!.multiArrayValue!

let temperature = 1.7601518630981445  // choice:3-5
let z = (0..<5).map { logits[$0].doubleValue / temperature }
let e = z.map { exp($0 - z.max()!) }
let probabilities = e.map { $0 / e.reduce(0, +) }

Output of swiftc -O examples/RouteCoreML.swift -o route && ./route . pay_wrong_address on the M4 Mac. The token ids come from the case pay_wrong_address of tests/vectors.json. Without a case id, the example uses the first case.

pay_wrong_address: 112 tokens
billing    0.1282  (laya 0.1284)
technical  0.0965  (laya 0.0965)
account    0.0650  (laya 0.0646)
shipping   0.6071  (laya 0.6077)
other      0.1032  (laya 0.1029)
Swift .cpuAndNeuralEngine: median 71.5 ms, p90 72.1 ms
MLComputePlan preferred devices: ["CPU": 12, "NeuralEngine": 1178]

That run had a 1-minute load average of 2.7 to 3.3. Use the table above for the Mac latency.

Android with ONNX Runtime

Use com.microsoft.onnxruntime:onnxruntime-android (1.23.0 is the version checked here) and the CPU execution provider. Create one OrtSession from onnx/laya-q8.onnx. Feed the five inputs of the ONNX table as OnnxTensor values: LongBuffer for the int64 inputs and a boolean array for marker_mask. Read logits and decode it as in the Python example. The tokenizer must be a port of the ModernBERT byte-level BPE in tokenizer.json. Check the port against tests/vectors.json. It holds the expected token ids for 46 inputs, with Unicode, emoji, JSON states, conversations, and inputs longer than 512 tokens. This Android path was tested on the emulator, not on a phone.

Known limits

  • Not tested on a phone yet. The Core ML accuracy and latency in this card come from a Mac. The phone evidence comes from another checkpoint with the same architecture.
  • Core ML on the CPU only can be less exact. On the private routing set, CPU_ONLY had the largest error of the three compute units (0.0123) and moved one input across 0.40. On tests/vectors.json its error is 0.0062, close to the other units. Use .cpuAndNeuralEngine or .cpuAndGPU.
  • Fixed length. The Core ML package always computes 512 positions, so a short input costs the same as a long one. A package with several lengths, or with a flexible length, does not run on the Neural Engine. The OpenJevSwift test found this for multifunction packages. A flexible fp32 package of this conversion gave a prediction error (-7) on .cpuAndNeuralEngine on macOS.
  • NFC normalization. tokenizer.json declares NFC normalization. The Python tokenizers package applies it. A port without NFC, for example in Dart, gives the same ids only for text that is already NFC. Text with combining accents, such as e followed by U+0301, can get different ids. Normalize to NFC before you tokenize, if your platform can.
  • Dynamic int8 quantization is wrong on ONNX Runtime 1.23. On the private routing set, a quantize_dynamic int8 file (per channel, 443 MB) had a max error of 0.527 and 20 argmax changes on ORT 1.23.0, and 0.080 on ORT 1.30.0. Use the MatMulNBits file. 4-bit MatMulNBits files had errors of 0.12 to 0.31 and are not included.
  • fp16 ONNX does not load. An fp16 ONNX file made with onnxconverter-common fails to load on ORT 1.23 (type mismatch in the rotary and attention nodes).
  • The q8 file trades latency for memory. On ORT CPU it is slower than fp32 (601 ms against 448 ms on the Mac, 917 ms against 512 ms on the emulator).
  • English only. This is the English checkpoint, with a 512-token context. For other languages, use the upstream laya-multilingual checkpoint. It is not converted here.
  • Upstream limits apply. The upstream card lists them. For example, the checkpoint ships over-confident and works best with temperatures fitted on your own data.

Reproduce the files

The scripts run in the layout of this repository. They write to onnx/ and coreml/, and their checks read tests/vectors.json. The check-only commands below do not overwrite any file.

Upstream: convaiinnovations/laya at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851.

  1. Export the fp32 ONNX file. Tools: Python 3.12, torch 2.14.0, transformers 5.17.0, laya 0.3.22, onnx 1.23.0. The TorchScript exporter (dynamo=False) writes opset 17.

    uv run --no-project --python 3.12 --with laya==0.3.22 --with torch==2.14.0 \
      --with transformers==5.17.0 --with onnx==1.23.0 --with onnxruntime==1.30.0 \
      python scripts/export_laya_onnx.py --revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851
    

    The export traces one example question. --question FILE reads another question from a JSON file. The graph has dynamic rows, tokens, and options, so the trace question does not change what the file computes. The file in this repository was traced with another 5-option question, so a new export can have other bytes. To check the existing file only, run python scripts/export_laya_onnx.py --skip-export with onnxruntime and numpy.

  2. Quantize to 8-bit weights. Then check the file on ORT 1.23.0.

    uv run --no-project --python 3.12 --with onnxruntime==1.30.0 --with onnx==1.23.1 \
      --with onnx-ir==1.0.0 python scripts/quantize_laya_nbits.py
    uv run --no-project --python 3.12 --with onnxruntime==1.23.0 --with numpy \
      python scripts/quantize_laya_nbits.py --check-only
    

    Expected: 653,610,034 bytes, and on tests/vectors.json a max error of 0.0167 with no argmax or 0.40 side changes.

  3. Convert to Core ML and check the package on all three compute units. The conversion takes about 10 seconds. The checks take a few minutes.

    uv run --no-project --python 3.12 --with coremltools==9.0 --with torch==2.14.1 \
      --with transformers==5.17.0 --with laya==0.3.22 --with numpy==2.3.5 \
      python scripts/export_laya_coreml.py --precision fp16 --fixed-length 512
    

    To check the existing package only, add --skip-export. That mode needs only coremltools 9.0 and numpy 2.3.5.

    Pin numpy to 2.3. coremltools 9.0 calls int() on a one-element array, and numpy 2.4 and later reject that. Expected: weights directory hash d373ef10ee6cc5fce17129f88a31e492c690cd7876a30e5de9967ce4feb23b4a (the sha256_tree of scripts/export_laya_coreml.py). Two conversions gave the same weights. Manifest.json holds random identifiers, so the zip file hash changes with each conversion.

  4. To change or extend the test cases, edit tests/make_vectors.py and run it with the laya package.

    uv run --no-project --python 3.12 --with laya==0.3.22 python tests/make_vectors.py
    

The Core ML wrapper changes three things in the graph, after OpenJevSwift pull request 85. The head layers use rank-4 tensors and scaled_dot_product_attention. The question type embedding is a constant row. The option rows are picked with a one-hot matrix product instead of a gather. The rotary tables and the window band are constants. The masks use -10,000, which is finite in fp16. The PyTorch fp32 wrapper matches tests/vectors.json to 0.00005 before the conversion.

sha256

File sha256
coreml/laya-fp16-512.mlpackage.zip fa81e1e2d843c32419870db8721cbaf1be7c843d0a033ac1f229fb000520bf2c
coreml/laya-fp16-512.mlpackage/Manifest.json 1484bc40932dda7e10dbc3eff8a050f765952a94acbd9a5381846e63ac8584c7
coreml/laya-fp16-512.mlpackage/Data/com.apple.CoreML/model.mlmodel a3b4948bc744fcf47810b33d8623a0fc5fe3349ab7614416c92c32481cab4179
coreml/laya-fp16-512.mlpackage/Data/com.apple.CoreML/weights/weight.bin 93584529b5a43501e1ea47c8ff4fbb7a71cd8c468f66e70f0e82ec4b8059dd50
onnx/laya-q8.onnx 8ab449ce38f11dd98ef0d053c68f37d0a6f1daa2b0018923f9c8a2bfe1a2fdb6
onnx/laya-fp32.onnx f068ada76b84034b9a731c296cf97a1c78182e5fdda9d553bf9a788078c7b8b0
tokenizer.json 6c8aaa9a542084f2457eab775d4eeb51f92a70c0fd9de28d5edb0ddec3c08d30
rl_agent_config.json ae287b56bbcf5f8c4f4541ae9dfd00c914c4c48b940b8398c3058af37ba92bbd
examples/laya_sequence.py 89720535288053de184e2000e057bca9cdca79d1ef76ce881905d09ae7a762ba
examples/route_onnx.py 5b8a68ca5041870ed4fa52f1a07d609e08603549530cf563af4da65986b34fb8
examples/route_coreml.py 211fb02f59a252a912e946fc30eda16b1dc4cf37c164e810158ab568463ce606
examples/RouteCoreML.swift 73cdacac4904d29bae4aa4b267d0778b99b8dad7268920a4e2234fc87fb97256
tests/vectors.json 1567466378a6416f76a5d6f7bc730466763854dfefda95f7653bf1832a80a016
tests/check.py 75f190de5610b6bdbdcf011ae750fbb8d91725d338e35a1f4a862981ec92704d
tests/make_vectors.py da3785c3499df07e0cff3d198f26fa490782e1bd187b6aefb87dd6de3d90a245
scripts/export_laya_coreml.py c5cabc4daa74f95ac6927b9b1fc81b8adadf433ed29fe6af2eedcc01ab5c2a04
scripts/export_laya_onnx.py aa1d591542233d27d4d1f59528b60f58023fc30d2c0dc3f41e4907df01de8a9b
scripts/quantize_laya_nbits.py 8cf953b6a1489e1239717c41c5db7b4e3550f104be2536e48bbb29d31f1ffe16
LICENSE cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30
NOTICE 1e40db608c7c84d6e56f4b00affd714e30f7a8ff4e063901348bd11aefb50d73

Credit

Laya, its weights, its tokenizer files, and its calibrated temperatures are the work of Convai Innovations and the authors of NandhaKishorM/laya. The upstream card publishes no citation entry. Please credit the upstream model when you use these files, for example:

@misc{convai_laya,
  title        = {Laya},
  author       = {{Convai Innovations}},
  howpublished = {\url{https://huggingface.co/convaiinnovations/laya}},
  note         = {Revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851}
}

The Laya encoder is ModernBERT-large by Answer.AI and LightOn. The Neural Engine graph rewrites follow OpenJevSwift pull request 85 by Algorythm Canada.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for charioteer/laya-mobile

Quantized
(56)
this model