Laya MoE for the browser

ONNX Runtime Web files for layaMOE: the shared encoder of convaiinnovations/laya-typed-decisions plus one small decision head per expert (general is the original head; the others were fine-tuned on synthetic domain data).

  • encoder.onnx + parts: ModernBERT-large encoder, int8 weights (~400 MB, downloaded once)
  • head_<name>.onnx + parts: 2-layer decision head per expert, int8 weights
  • manifest.json: file list with SHA-256, router question, per-head calibration temperatures

Modified derivative of convaiinnovations/laya-typed-decisions (Apache-2.0, Copyright ConvAI Innovations); unofficial and not affiliated with ConvAI Innovations. Changes: split into encoder and head graphs, heads fine-tuned, weights quantized (weight-only int8), files split into parts. See LICENSE and NOTICE.md.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VishalMysore/laya-moe-web

Quantized
(12)
this model