nitzotz / NOTICE
BrainboxAI's picture
Update NOTICE
0b2905b verified
Raw History Blame Contribute Delete
2.7 kB
Nitzotz (BrainboxAI/nitzotz)
Copyright 2026 BrainboxAI
This model is released under the Apache License, Version 2.0 (see LICENSE).
It is built from the following third-party work. Their licences and attribution requirements are kept here.
Encoder
- HalleluBERT-large (HalleluBERT/HalleluBERT_large), MIT License. The encoder weights of this model were initialised
from it and fine-tuned. Paper: arXiv 2510.21372.
Architecture, training method and runtime
- laya (github.com/NandhaKishorM/laya, NandhaKishorM / Convai Innovations), Apache License 2.0. The DecisionModel head
architecture, the option-marker scoring and the laya runtime. The head of this model was trained from scratch
(soft cross-entropy); no weights come from any laya checkpoint.
- ggmlc (github.com/monatis/ggmlc), used to compile the GGUF files. ggmlc/nitzotz_trunk.py and
ggmlc/compile_nitzotz.py are adapted from its examples/laya files (laya_trunk.py, compile_laya.py).
Training data
- HeQ, Hebrew Question Answering Dataset v1.1 (NNLP-IL, github.com/NNLP-IL/Hebrew-Question-Answering-Dataset, commit
f46d2ff), CC BY 4.0. Created by Webiks for MAFAT and the National NLP Program of Israel. Includes passages from
Hebrew Wikipedia and from Geektime, shared by the dataset authors under the dataset licence. Train split only.
Used twice: as extractive question answering to teach the encoder to read (passages overlapping the test removed),
and reformatted into yes/no and 4-option questions for the decision head.
- MASSIVE (AmazonScience/massive, he-IL), CC BY 4.0, Amazon. FitzGerald et al., 2022. Train split only. Intent
labels were translated into short Hebrew labels.
- Synthetic Hebrew messages written by DeepSeek V4.1 Flash (deepseek-ai, MIT License) and labelled by DeepSeek V4.1
Flash and Gemma 4 31B (google/gemma-4-31b-it, Apache License 2.0), both served by DeepInfra through OpenRouter,
with no data retention. These synthetic data are not redistributed here.
- Hebrew reading questions and topic labels on passages from FineWeb-2 (HuggingFaceFW/fineweb-2,
Hebrew subset), ODC-By 1.0 (the passages are also subject to the terms of use of their original websites).
Questions written by DeepSeek V4.1 Flash and checked by Gemma 4 31B; topic labels by the same two models.
- Part of the training data uses text written with OpenAI GPT (through a ChatGPT subscription, OpenAI terms of use, not an
open licence): alternative wordings of the questions and answer options, and short scenario outlines from which
DeepSeek V4.1 Flash wrote messages. GPT wrote no message text and no label. These data are not redistributed here.
No personal data and no private customer data were used.