Nitzotz (BrainboxAI/nitzotz) Copyright 2026 BrainboxAI This model is released under the Apache License, Version 2.0 (see LICENSE). It is built from the following third-party work. Their licences and attribution requirements are kept here. Encoder - HalleluBERT-large (HalleluBERT/HalleluBERT_large), MIT License. The encoder weights of this model were initialised from it and fine-tuned. Paper: arXiv 2510.21372. Architecture, training method and runtime - laya (github.com/NandhaKishorM/laya, NandhaKishorM / Convai Innovations), Apache License 2.0. The DecisionModel head architecture, the option-marker scoring and the laya runtime. The head of this model was trained from scratch (soft cross-entropy); no weights come from any laya checkpoint. - ggmlc (github.com/monatis/ggmlc), used to compile the GGUF files. ggmlc/nitzotz_trunk.py and ggmlc/compile_nitzotz.py are adapted from its examples/laya files (laya_trunk.py, compile_laya.py). Training data - HeQ, Hebrew Question Answering Dataset v1.1 (NNLP-IL, github.com/NNLP-IL/Hebrew-Question-Answering-Dataset, commit f46d2ff), CC BY 4.0. Created by Webiks for MAFAT and the National NLP Program of Israel. Includes passages from Hebrew Wikipedia and from Geektime, shared by the dataset authors under the dataset licence. Train split only. Used twice: as extractive question answering to teach the encoder to read (passages overlapping the test removed), and reformatted into yes/no and 4-option questions for the decision head. - MASSIVE (AmazonScience/massive, he-IL), CC BY 4.0, Amazon. FitzGerald et al., 2022. Train split only. Intent labels were translated into short Hebrew labels. - Synthetic Hebrew messages written by DeepSeek V4.1 Flash (deepseek-ai, MIT License) and labelled by DeepSeek V4.1 Flash and Gemma 4 31B (google/gemma-4-31b-it, Apache License 2.0), both served by DeepInfra through OpenRouter, with no data retention. These synthetic data are not redistributed here. - Hebrew reading questions and topic labels on passages from FineWeb-2 (HuggingFaceFW/fineweb-2, Hebrew subset), ODC-By 1.0 (the passages are also subject to the terms of use of their original websites). Questions written by DeepSeek V4.1 Flash and checked by Gemma 4 31B; topic labels by the same two models. - Part of the training data uses text written with OpenAI GPT (through a ChatGPT subscription, OpenAI terms of use, not an open licence): alternative wordings of the questions and answer options, and short scenario outlines from which DeepSeek V4.1 Flash wrote messages. GPT wrote no message text and no label. These data are not redistributed here. No personal data and no private customer data were used.