Download README.md from BrainboxAI/nitzotz: direct link, hf CLI and curl.
- Browser
- Download file 60.6 kB
-
https://huggingface.co/BrainboxAI/nitzotz/resolve/main/README.md
- Command line
-
hf download hf://BrainboxAI/nitzotz/README.md
-
curl -L -o README.md https://huggingface.co/BrainboxAI/nitzotz/resolve/main/README.md
license: apache-2.0
language:
- he
library_name: laya
pipeline_tag: zero-shot-classification
base_model: HalleluBERT/HalleluBERT_large
tags:
- laya
- hebrew
- decision-model
- calibrated
- scam-detection
- routing
- gguf
datasets:
- AmazonScience/massive
What it is. Nitzotz reads a Hebrew message and answers questions you type about it: pick one of several options, give a score on a scale, or say yes or no to a claim. For every answer it gives a probability you can trust, so you know when it is sure and when it is guessing. It does not write text, so it cannot make things up. It runs on a normal laptop, with no internet connection and no cost per question.
What it is for. Deciding what to do with incoming messages: is this a scam, what kind of message is it, which department should get it, how urgent is it. It is not a chatbot and it is not built for long documents (see Limitations).
In numbers. On 298 Hebrew messages it says correctly whether a message is a scam 92.0% of the time. For context: 70% of those messages are not scams, so a model that always says "not a scam" would score 70.5%. The number that matters more is the ranking: it gives real scams a higher probability than legitimate messages 97% of the time (AUC 0.97).
Try it
Python (the laya library, pip install laya):
import laya
agent = laya.load("BrainboxAI/nitzotz")
q = {"scam": {"type": "noul",
"instructions": "讛讛讜讚注讛 诪谞住讛 诇讙专讜诐 诇谞诪注谉 诇诇讞讜抓 注诇 拽讬砖讜专, 诇砖诇诐 讗讜 诇诪住讜专 驻专讟讬诐 讘诇讬 住讬讘讛 诇讙讬讟讬诪讬转.",
"criteria": {"true": "讻谉, 讝讛 谞讬住讬讜谉 诪专诪讛",
"false": "诇讗, 讝讜 讛讜讚注讛 诇讙讬讟讬诪讬转, 讙诐 讗诐 讬砖 讘讛 拽讬砖讜专 讗讜 讘拽砖转 转砖诇讜诐"}}}
print(agent.predict("讛讞讘讬诇讛 砖诇讱 诪注讜讻讘转. 诇砖讞专讜专 砖诇诐 12.90 讘拽讬砖讜专", q)["answers"]["scam"]["noul"])
Output: 0.8628, the probability that the claim ("the message tries to make the reader click a link, pay or hand
over details for no legitimate reason") is true.
The wording of the question matters. This is the exact question the scam test used, and the numbers on this card are for it. In our tries a shorter wording ("the message is a scam attempt") gave clearly worse answers, for example a high scam probability for a plain verification code. If you change the wording, check it on your own messages first.
laya.exe (the standalone binary from ggmlc releases, no Python needed). Download the Q8 file, start the daemon, then send one JSON request per line; each answer comes back on one line:
hf download BrainboxAI/nitzotz nitzotz-q8_0.gguf --local-dir .
laya.exe daemon nitzotz-q8_0.gguf --device vulkan
{"id":"1","state":"讛讬讬, 讛驻讙讬砖讛 诪讞专 讘-10 注讚讬讬谉 讘转讜拽祝?","questions":{"scam":{"type":"noul","instructions":"讛讛讜讚注讛 诪谞住讛 诇讙专讜诐 诇谞诪注谉 诇诇讞讜抓 注诇 拽讬砖讜专, 诇砖诇诐 讗讜 诇诪住讜专 驻专讟讬诐 讘诇讬 住讬讘讛 诇讙讬讟讬诪讬转.","criteria":{"true":"讻谉, 讝讛 谞讬住讬讜谉 诪专诪讛","false":"诇讗, 讝讜 讛讜讚注讛 诇讙讬讟讬诪讬转, 讙诐 讗诐 讬砖 讘讛 拽讬砖讜专 讗讜 讘拽砖转 转砖诇讜诐"}}}}
{"status":"ready","model":"laya"}
{"model": "nitzotz", "family": "nitzotz", "route": "forced nitzotz", "answers": {"scam": {"type": "noul", "action": {"act_probability": 0.908}, "confidence": 0.9536, "noul": 0.0464}}, "usage": {"input_tokens": 66, "output_tokens": 0, "latency_ms": 121.783}, "id": "1"}
noul is the probability that the claim is true. Several questions in one request are answered together in one pass.
Use --device cpu on a machine without a GPU.
Benchmarks
8 frozen test sets, 3,990 questions in total, locked by checksum before any training data existed. All three models answered exactly the same questions. The two reference points are the other open laya models that read Hebrew: RoeiG/laya-hebrew (licence CC-BY-NC-SA-4.0, non-commercial use only) and laya-multilingual (Apache-2.0).
Full table. Accuracy, the best in each row in bold. The last two columns say whether Nitzotz's difference from that model is real or could be luck (a paired exact McNemar test on the same questions): "better" or "worse" means p < 0.05, "tie" means the difference could be chance.
| Test (questions) | Nitzotz | RoeiG (non-commercial) | laya-multilingual | Chance | Nitzotz vs RoeiG | Nitzotz vs laya-multilingual |
|---|---|---|---|---|---|---|
| Scam or not? (298 messages) | 92.0% | 32.2% | 50.3% | 50.0% | better p<0.001 |
better p<0.001 |
| same, hard cases only (65) | 83.1% | 32.3% | 35.4% | 50.0% | better p<0.001 |
better p<0.001 |
| Message type, 6 options (298) | 77.8% | 52.7% | 20.8% | 16.7% | better p<0.001 |
better p<0.001 |
| same, hard cases only (65) | 67.7% | 44.6% | 24.6% | 16.7% | better p=0.01 |
better p<0.001 |
| Support ticket type, 5 options (30) | 90.0% | 80.0% | 56.7% | 20.0% | tie p=0.45 |
better p=0.01 |
| Ticket urgency, 5 levels (30) | 70.0% | 36.7% | 33.3% | 20.0% | better p=0.03 |
better p=0.02 |
| Paying customer? yes/no (30) | 80.0% | 56.7% | 50.0% | 50.0% | better p=0.02 |
tie p=0.06 |
| Voice command intent, 20 options (500) | 90.0% | 72.6% | 47.4% | 5.0% | better p<0.001 |
better p<0.001 |
| Voice command intent, 4 options (500) | 97.2% | 90.8% | 69.4% | 25.0% | better p<0.001 |
better p<0.001 |
| News topic, 7 options (204) | 79.4% | 82.3% | 66.2% | 14.3% | tie p=0.36 |
better p=0.001 |
| Does the passage support this answer? (600) | 94.8% | 94.2% | 48.8% | 50.0% | tie p=0.69 |
better p<0.001 |
| Plausible answer the passage does not give (600) | 89.2% | 53.0% | 53.7% | 50.0% | better p<0.001 |
better p<0.001 |
| Reading comprehension, 4 options (900) | 63.2% | 75.4% | 31.4% | 25.0% | worse p<0.001 |
better p<0.001 |
On Belebele reading comprehension Nitzotz scores 63.2%, below RoeiG/laya-hebrew (75.4%); Nitzotz is built for message decisions, not long-passage comprehension.
How to read it:
- MASSIVE (voice commands): Nitzotz was trained on MASSIVE's training commands. The test commands are different, but written by the same people in the same style, so this test is easier for Nitzotz than for the others.
- HeQ: Nitzotz was trained on HeQ's training passages. The test passages are different ones, and training passages that overlapped a test passage were removed.
- The support-ticket rows have only 30 questions. No difference there is reliable.
- The spam and scam test has a lean: 70% of its messages are not scams. The "chance" line (50%) is a coin flip, not the best blind strategy.
Why you can trust it
The probabilities mean something. Take all the answers where Nitzotz said it was about 70 to 80% sure, and count
how many were right: the chart above does that for every confidence level, over 3,990 test questions.
Where the dots sit above the line, Nitzotz is more often right than it claims (it is modest); where they sit below, it
is over-confident. This is what lets you set thresholds (see the next section). The per-question-type temperatures
were fitted on 4,825 held-out training items, never on the test sets (calibration.json).
It reads the text. With the message removed and only the question left, its accuracy falls to 70.5% on the scam question and 7.7% on message type. So the answers come from the message, not from the wording of the question.
Rewording the question rarely changes the answer. We asked 1,786 test questions in 6 different wordings with the same meaning. On 5.0% of them the answer was not the same in all 6. No wording makes it give one fixed answer to every item of a test. The exception is the urgency question (see Limitations).
It is fast on ordinary hardware. On one laptop (Intel Core Ultra 9 285H laptop, built-in Arc 140T GPU, Windows 11), one question at a time:
| Runtime | Median over 50 check questions (about 135 tokens) | Short message (38 tokens) | Long input (430 tokens) |
|---|---|---|---|
| laya.exe, Q8 file, GPU (Vulkan) | 48 ms | 37 ms | 127 ms |
| laya.exe, F16 file, GPU (Vulkan) | 50 ms | 43 ms | 111 ms |
| Python (laya), GPU (PyTorch XPU) | 60 ms | 44 ms | 139 ms |
| laya.exe, Q8 file, CPU only (16 threads) | 309 ms | 160 ms | 1137 ms |
| Python (laya), CPU only | 224 ms | 112 ms | 1022 ms |
The short and long columns repeat one fixed input 30 times after 5 warm-up calls. Timings on this laptop change a lot from one session to another (an earlier measurement of the same setup was several times slower), so treat these as rough.
The GGUF files give almost the same answers as the Python model. Compared with the full-precision Python model on the CPU:
| File, device | Same top answer, 50 questions | Largest probability gap | Same top answer, 596 spam questions | Largest gap | Average gap |
|---|---|---|---|---|---|
| Q8, GPU | 50/50 | 0.0375 | 596/596 | 0.0166 | 0.00130 |
| F16, GPU | 50/50 | 0.0017 | 596/596 | 0.0015 | 0.00020 |
| Q8, CPU | 50/50 | 0.0549 | 595/596 | 0.0207 | 0.00192 |
| F16, CPU | 50/50 | 0.0032 | 596/596 | 0.0012 | 0.00018 |
A gap of 0.01 means, for example, 0.83 against 0.84. The few questions where the top answer changes are ones where the two best answers were almost tied. Over both spam questions the Q8 file on the GPU is right 84.7% of the time, against 84.7% for the Python model; the F16 file is closer (84.7%).
Use it in your business
The idea ("Ramzor", traffic light). Every incoming message gets one or more quick questions. The probability decides what happens next:
- Green (confident it is fine): handle automatically, file it, tag it, route it.
- Yellow (not sure): send it to a person, or to a large language model if you use one.
- Red (confident it is a scam): block or quarantine it.
The drawing is a worked example on the 298 test messages. The upper threshold, 0.49, is the one chosen for the scam question on 1188 held-out training messages (never on the test); the lower one, 0.35, was picked by hand. With them, 207 messages go to green (13 of them are in fact scams), 9 to yellow (2 scams), and 82 to red (9 of them are in fact legitimate). So red should mean "quarantine and check", not "delete". Pick your own thresholds on a sample of your own messages, and decide how many mistakes in green and red you can live with. At the 0.49 threshold alone, the scam answer is right 92.3% of the time overall and 83.1% on the 65 hard cases (at 0.50: 92.0% and 83.1%).
The model is cheap enough to run on every message. The person (or the LLM) only sees the yellow part. Do not use it as the only line of defence for decisions that can hurt someone.
Training data and transparency
Two training stages, on a rented GPU:
- Learning to read. HalleluBERT-large was first trained to find the answer to a question inside a passage, on 27,085 HeQ training questions (CC BY 4.0). Passages that overlapped a test passage were removed.
- Learning to decide. A laya decision head was put on top and the whole model was trained on 116,854 items (112,029 for training, 4,825 held out to pick the best of 2 passes and to fit the temperatures).
| Source | Items | Licence | How it was made |
|---|---|---|---|
| Synthetic: Message type, 6 classes | 14,000 | teacher output, project-owned (DeepSeek MIT; Gemma Apache-2.0) | written by DeepSeek V4.1 Flash, labelled independently by both models |
| Synthetic: Routing to a department (3 to 6 options) | 8,750 | teacher output, project-owned (DeepSeek MIT; Gemma Apache-2.0) | written by DeepSeek V4.1 Flash, labelled independently by both models |
| Synthetic: Yes/no claims about a message | 8,750 | teacher output, project-owned (DeepSeek MIT; Gemma Apache-2.0) | written by DeepSeek V4.1 Flash, labelled independently by both models |
| Synthetic: Urgency, 5 levels | 3,500 | teacher output, project-owned (DeepSeek MIT; Gemma Apache-2.0) | written by DeepSeek V4.1 Flash, labelled independently by both models |
| HeQ: is the proposed answer supported (yes/no) | 6,000 | CC BY 4.0 | the dataset's gold labels, train split only |
| HeQ: plausible answer to an unanswerable question | 9,750 (4,750 of them repeated copies) | CC BY 4.0 | the dataset's gold labels, train split only |
| HeQ: 4-option reading | 5,000 | CC BY 4.0 | the dataset's gold labels, train split only |
| MASSIVE he-IL: voice command intent | 7,000 | CC BY 4.0 | the dataset's gold labels, train split only |
| Scam or not, disguised scams and real messages that look suspicious | 6,000 | model output, project-owned (DeepSeek MIT; Gemma Apache-2.0) | written by DeepSeek V4.1 Flash from scenario outlines written with GPT; kept only if both labellers agreed with the writer |
| Message type of the same messages | 5,374 | model output, project-owned | the type both labellers agreed on |
| Copies of existing items with the question reworded | 21,199 | as the original item; the wordings were written with GPT | question or options in one of 376 alternative wordings written with GPT; the label comes from the original item |
| Reading, 4 options ("which is NOT", reworded answers, several sentences) | 6,825 | passages: FineWeb-2, ODC-By 1.0; questions: model output, project-owned | written by DeepSeek V4.1 Flash, checked by Gemma 4 31B with the passage and again without it; dropped if it could be answered without the passage |
| Topic of a passage, 7 options | 4,000 | passages: FineWeb-2, ODC-By 1.0; labels: model output | labelled by DeepSeek V4.1 Flash and Gemma 4 31B |
| Is the sender an existing paying customer (yes/no) | 575 | model output, project-owned; question wordings written with GPT | support messages written by DeepSeek V4.1 Flash, checked by Gemma 4 31B |
| Scam or not: warnings about scams, short scams without a link, and legitimate look-alikes | 3,746 | model output, project-owned (DeepSeek MIT; Gemma Apache-2.0); question wordings written with GPT | written by DeepSeek V4.1 Flash; kept only if both labellers agreed with the writer, except that a scam only DeepSeek recognised was kept with a softer label (70%) |
| Message type of the same messages | 1,901 | model output, project-owned | the type both labellers agreed on, only where it fits the scam decision |
| "None of the options": existing training items with one option changed | 4,484 | as the original item (MASSIVE and HeQ: CC BY 4.0; reading: FineWeb-2 passages, ODC-By 1.0) | the right answer removed and a "none of the options" choice added; in a third of them a wrong option was removed instead, so "none" is wrong there |
- Synthetic messages (35,000 items). DeepSeek V4.1 Flash (MIT) wrote Israeli-style SMS, WhatsApp and email messages from a plan (intended label, topic, tone, varied fake phone numbers and links). DeepSeek and Gemma 4 31B (Apache-2.0) then each labelled every message on their own. A message was kept only if both agreed; the two models agreed on 94% of them, and 93.7% of the 37,429 written messages were kept. Both ran through DeepInfra (via OpenRouter), with data retention off and "thinking" mode off. The synthetic data is not published.
- Reading, hard scams and support tickets. 6,825 reading questions on Hebrew web passages from FineWeb-2 (ODC-By 1.0): "which of these is NOT true or NOT mentioned" (each paired with a positive question on the same passage), answers said in other words, and answers that need several sentences. DeepSeek V4.1 Flash wrote them and Gemma 4 31B checked each one twice, with the passage and without it; a question that could be answered without the passage was dropped. 6,000 messages that are hard to tell apart (disguised scams, and real messages that look suspicious), written by DeepSeek and kept only if both labellers agreed with each other and with the writer. 4,000 passages labelled with a topic, 575 support messages for the paying-customer question, and 4,750 repeated HeQ items so that the HeQ skills keep their weight in the mix.
- Scam patterns seen in real messages. Real Hebrew messages showed two weak spots: genuine warnings about scams (from banks, the police, companies) flagged as scams, and short scams without a link missed. So DeepSeek V4.1 Flash wrote 3,746 more messages: 1,000 legitimate warnings about scams, 1,476 short scams without a link (a small unpaid debt or toll, a "friend" with a new number, a gift, a fake payment confirmation, a payment app), 389 pairs of a scam and a warning about that same scam, and 492 legitimate messages that look like those scams. The keep rule changed for these: a message was kept if both labellers agreed with the writer, and a scam that only DeepSeek recognised was also kept, with a softer label (70% instead of close to 100%); 134 messages are of that kind. Messages that resembled one of the real messages we checked were dropped before labelling, and the real messages themselves were never used for training. 134 existing training rows (67 messages: a request from a "new number" or for a small debt, with no link) were relabelled as scams after the labellers called them scams (with the 70% label where only DeepSeek did). 4,484 "none of the options" items were made from existing MASSIVE, HeQ and reading training items: in 2,990 the right answer was removed and a "none of the options" choice added, and in 1,494 a wrong option was removed instead, so "none" is wrong there.
- Written with GPT. Part of the training data was written with OpenAI's GPT, through our ChatGPT subscription: 376 alternative wordings of the questions and 34 sets of alternative answer options, used in 35,189 training items (including all 575 paying-customer items), and 120 short scenario outlines from which DeepSeek V4.1 Flash wrote 6,000 messages. GPT wrote no message and no label. In total 37,849 of the 116,854 training items (32%) use text written with GPT. GPT output is covered by OpenAI's terms of use, not by an open licence.
- Open data. HeQ v1.1 (CC BY 4.0; about half of its questions are on Geektime articles, shared by the HeQ authors under the same licence) and MASSIVE he-IL (CC BY 4.0), training splits only. FineWeb-2 Hebrew (ODC-By 1.0) passages for the reading and topic items.
- No leaks from the tests. Every training item was compared with every test question; anything sharing an 8-word run with a test text was dropped, and so were MASSIVE commands equal to a test command. The reading, topic, hard-scam, paying-customer and reworded items were also checked with stricter rules: every passage against every test text (any shared 6-word run), and every question, option and wording against every test question and every reworded test question (exact match and character similarity).
- Not used: no output of any other closed commercial chatbot, no DICTA model, no non-commercial or share-alike data. GPT, a closed commercial model, was used only as described above.
- Size: encoder 357.1M parameters (HalleluBERT-large, MIT, fine-tuned); decision head 26.5M parameters, trained from scratch.
About the tests.
| Test | Questions | Source and licence |
|---|---|---|
| Spam and business messages | 596 (298 messages, 2 questions each) | in-house, Israeli SMS, WhatsApp and email style, 65 hard cases |
| Support tickets (triage30) | 90 (30 tickets, 3 questions each) | in-house |
| MASSIVE he-IL, 20 and 4 options | 500 + 500 | MASSIVE test split, CC BY 4.0 |
| SIB-200 news topic | 204 | CC BY-SA 4.0, used for testing only |
| Belebele reading | 900 | CC BY-SA 4.0, used for testing only |
| HeQ verify and unanswerable | 600 + 600 | HeQ v1.1 test split, CC BY 4.0 |
| "None of the options" (reported apart, see Limitations) | 400 (200 where "none" is right, 200 where it is wrong) | MASSIVE (CC BY 4.0) and Belebele (CC BY-SA 4.0) test questions with a "none of the options" choice added, used for testing only |
Caveats that change how much to trust the numbers:
- The spam and business test was written by an AI model and checked by an AI model, not by a person. Real inboxes will look different. It was frozen after that review (4 labels changed, 2 messages removed).
- Three training runs, three random seeds; this is one of them. On the held-out training items the three are almost level (95.0% here against 95.2% and 95.2%). This one was picked because it did best on the 188 held-out messages of the added scam data (98.4% against 97.3% and 97.3%) and on the real messages we checked, not on the tests in this card. The three runs are close on the large tests: scam or not 92.0% here against 92.0% and 90.6%, news topic 79.4% against 82.8% and 80.4%, reading 63.2% against 60.9% and 63.1%. This run and each of the other two give the same answer on 90.1% to 90.6% of all test questions. On the 30-ticket tests they differ more: ticket urgency 70.0% here against 73.3% and 76.7%.
- The training labels come from two AI models. Where both are wrong in the same way, Nitzotz learned their mistake.
- HeQ's wrong answers in the test were picked by code, not checked by a person.
Limitations
- Reading comprehension of longer passages is limited. Belebele: 63.2%, where a blind guess gets 25%, and still below the best other laya model on this test (see Benchmarks). When the right answer is written word for word in the passage it gets 77% (252 questions); when the answer is said in other words it gets 58% (648 questions); on "which of these is NOT" questions 59% (158 questions). Do not ask it whether a long document supports a claim.
- Checking an answer against a passage (HeQ) works. 94.8% with the passage; with the passage removed it falls to 51.0%, a coin flip. So on this kind of question it really reads the passage.
- Numbers, dates, amounts and rules: not trained and not measured. Compute them in code and pass the result in.
- Hard cases are still the weak spot: scams written to look legitimate (a "supplier" changing bank details, the "CEO" asking for a transfer) and real messages that look like scams (a real bank alert with a link, a real verification code). On the 65 hard cases the scam question is right 83.1% of the time (always answering "not a scam" there would give 70.8%), with AUC 0.84. At the chosen threshold it still misses 5 of the 19 scams written to look legitimate, and flags 6 of the 46 real messages that look like scams.
- Urgency is subjective, and its answer depends on the wording. Even the two teacher models matched the intended urgency only about 62 to 64% of the time. When the urgency question is asked in other words, the answer changes on 63% of the 30 test tickets.
- The in-house tests and much of the training data were written by AI models, not by people. This includes the spam and business test, the support-ticket test and the FineWeb-2 reading questions. Real messages will look different.
- The support-ticket tests are small: 30 tickets per question, so one ticket moves a score by 3.3 points.
- Sarcasm and irony are probably read literally. Not measured.
- 512 tokens (roughly 300 to 400 Hebrew words) per question. A longer message is cut from the end without a warning.
- Hebrew only. Not trained or tested on English or Arabic.
- Not a safety system on its own. It makes mistakes in both directions. Keep a person in the loop for anything that can hurt someone.
- When a "none of the options" choice is added, it picks it too often. Measured on a separate frozen test of 400 questions, MASSIVE and Belebele test questions rebuilt with a "none of the options" choice. When "none" is the right answer, it picks it 76% of the time. When the right answer is in the list, it still picks "none" on 31% of the questions and gets 61% of them right (77% on voice commands, 45% on reading questions). With the text removed it picks "none" almost every time (98%). If you offer such a choice, test it on your own questions first.
- Real messages: not measured on an independent set yet. Training data was added for genuine warnings about scams and for short scams without a link, the two weak spots real messages showed. A measurement on an independent set of real messages is still pending, so this card gives no number for real messages.
Licence and attribution
Apache-2.0 for the weights, the GGUF files and the code. Commercial use is allowed. Built on:
- HalleluBERT-large: the encoder, MIT.
- laya (NandhaKishorM, Convai Innovations): the decision-head architecture and the runtime, Apache-2.0. No laya weights are used.
- HeQ: CC BY 4.0, by Webiks for MAFAT and the Israeli National NLP Program (NNLP-IL); includes Geektime passages.
- MASSIVE: CC BY 4.0, Amazon (FitzGerald et al., 2022).
- DeepSeek V4.1 Flash (MIT) and Gemma 4 31B (Apache-2.0), used through DeepInfra as data writer and labellers.
- FineWeb-2 (Hebrew): ODC-By 1.0, Hugging Face; the passages of the reading and topic items.
- OpenAI GPT, through a ChatGPT subscription (OpenAI terms of use): question wordings and scenario outlines, as described in the training data section.
- ggmlc for the GGUF files.
- Belebele and SIB-200 (CC BY-SA 4.0) were used only to test, never to train.
The full notice is in NOTICE.
谞讬爪讜抓, 讘注讘专讬转
诪讛 讝讛. 谞讬爪讜抓 拽讜专讗 讛讜讚注讛 讘注讘专讬转 讜注讜谞讛 注诇 砖讗诇讜转 砖讗转诐 诪拽诇讬讚讬诐 注诇讬讛: 诇讘讞讜专 讗讞转 诪讻诪讛 讗驻砖专讜讬讜转, 诇转转 爪讬讜谉 讘住讜诇诐, 讗讜 诇注谞讜转 讻谉 讗讜 诇讗 注诇 讟注谞讛. 注诇 讻诇 转砖讜讘讛 讛讜讗 谞讜转谉 讛住转讘专讜转 砖讗驻砖专 诇住诪讜讱 注诇讬讛, 讻讱 砖讬讜讚注讬诐 诪转讬 讛讜讗 讘讟讜讞 讜诪转讬 讛讜讗 诪谞讞砖. 讛讜讗 诇讗 讻讜转讘 讟拽住讟, 讜诇讻谉 讛讜讗 诇讗 讬讻讜诇 诇讛诪爪讬讗 讚讘专讬诐. 讛讜讗 专抓 注诇 诪讞砖讘 谞讬讬讚 专讙讬诇, 讘诇讬 讗讬谞讟专谞讟 讜讘诇讬 转砖诇讜诐 注诇 讻诇 砖讗诇讛.
讘砖讘讬诇 诪讛. 诇讛讞诇讬讟 诪讛 注讜砖讬诐 注诐 讛讜讚注讜转 谞讻谞住讜转: 讛讗诐 讝讜 讛讜谞讗讛, 讗讬讝讛 住讜讙 讛讜讚注讛 讝讜, 诇讗讬讝讜 诪讞诇拽讛 诇讛注讘讬专, 讻诪讛 讝讛 讚讞讜祝. 讝讛 诇讗 爪'讗讟讘讜讟, 讜讛讜讗 诇讗 讘谞讜讬 诇诪住诪讻讬诐 讗专讜讻讬诐 (专讗讜 诪讙讘诇讜转).
讘诪住驻专讬诐. 注诇 298 讛讜讚注讜转 讘注讘专讬转 讛讜讗 拽讜讘注 谞讻讜谉 讗诐 讛讛讜讚注讛 讛讬讗 讛讜谞讗讛 讘-92.0% 诪讛诪拽专讬诐. 讘砖讘讬诇 驻专讜驻讜专爪讬讛: 70% 诪讛讛讜讚注讜转 讛讗诇讛 讛谉 诇讗 讛讜谞讗讛, 讻讱 砖诪讜讚诇 砖转诪讬讚 注讜谞讛 "诇讗 讛讜谞讗讛" 讛讬讛 诪拽讘诇 70.5%. 讛诪住驻专 砖讞砖讜讘 讬讜转专 讛讜讗 讛讚讬专讜讙: 讛讜讗 谞讜转谉 诇讛讜谞讗讛 讗诪讬转讬转 讛住转讘专讜转 讙讘讜讛讛 讬讜转专 诪讗砖专 诇讛讜讚注讛 转拽讬谞讛 讘-97% 诪讛诪拽专讬诐 (AUC 0.97).
诇谞住讜转
讘驻讬讬转讜谉 (讛住驻专讬讬讛 laya, pip install laya):
import laya
agent = laya.load("BrainboxAI/nitzotz")
q = {"scam": {"type": "noul",
"instructions": "讛讛讜讚注讛 诪谞住讛 诇讙专讜诐 诇谞诪注谉 诇诇讞讜抓 注诇 拽讬砖讜专, 诇砖诇诐 讗讜 诇诪住讜专 驻专讟讬诐 讘诇讬 住讬讘讛 诇讙讬讟讬诪讬转.",
"criteria": {"true": "讻谉, 讝讛 谞讬住讬讜谉 诪专诪讛",
"false": "诇讗, 讝讜 讛讜讚注讛 诇讙讬讟讬诪讬转, 讙诐 讗诐 讬砖 讘讛 拽讬砖讜专 讗讜 讘拽砖转 转砖诇讜诐"}}}
print(agent.predict("讛讞讘讬诇讛 砖诇讱 诪注讜讻讘转. 诇砖讞专讜专 砖诇诐 12.90 讘拽讬砖讜专", q)["answers"]["scam"]["noul"])
讛驻诇讟: 0.8628, 讛讛住转讘专讜转 砖讛讟注谞讛 ("讛讛讜讚注讛 诪谞住讛 诇讙专讜诐 诇谞诪注谉 诇诇讞讜抓 注诇 拽讬砖讜专, 诇砖诇诐 讗讜 诇诪住讜专 驻专讟讬诐 讘诇讬 住讬讘讛 诇讙讬讟讬诪讬转")
谞讻讜谞讛.
讛谞讬住讜讞 砖诇 讛砖讗诇讛 诪砖谞讛. 讝讜 讘讚讬讜拽 讛砖讗诇讛 砖讘讛 讛砖转诪砖 诪讘讞谉 讛讛讜谞讗讜转, 讜讛诪住驻专讬诐 讘讻专讟讬住 讛讝讛 讛诐 注诇讬讛. 讘谞讬住讬讜谞讜转 砖诇谞讜 谞讬住讜讞 拽爪专 讬讜转专 ("讛讛讜讚注讛 讛讬讗 谞讬住讬讜谉 讛讜谞讗讛") 谞转谉 转砖讜讘讜转 讙专讜注讜转 讘讛专讘讛, 诇诪砖诇 讛住转讘专讜转 讙讘讜讛讛 诇讛讜谞讗讛 诇拽讜讚 讗讬诪讜转 专讙讬诇. 讗诐 诪砖谞讬诐 讗转 讛谞讬住讜讞, 讘讜讚拽讬诐 讗讜转讜 拽讜讚诐 注诇 讛讛讜讚注讜转 砖诇讻诐.
讘诇讬 驻讬讬转讜谉, 注诐 laya.exe (转讜讻谞讛 注爪诪讗讬转 诪-ggmlc). 诪讜专讬讚讬诐 讗转 拽讜讘抓 Q8, 诪驻注讬诇讬诐, 讜砖讜诇讞讬诐 讘拽砖转 JSON 讗讞转 讘讻诇 砖讜专讛. 讻诇 转砖讜讘讛 讞讜讝专转 讘砖讜专讛 讗讞转:
hf download BrainboxAI/nitzotz nitzotz-q8_0.gguf --local-dir .
laya.exe daemon nitzotz-q8_0.gguf --device vulkan
{"id":"1","state":"讛讬讬, 讛驻讙讬砖讛 诪讞专 讘-10 注讚讬讬谉 讘转讜拽祝?","questions":{"scam":{"type":"noul","instructions":"讛讛讜讚注讛 诪谞住讛 诇讙专讜诐 诇谞诪注谉 诇诇讞讜抓 注诇 拽讬砖讜专, 诇砖诇诐 讗讜 诇诪住讜专 驻专讟讬诐 讘诇讬 住讬讘讛 诇讙讬讟讬诪讬转.","criteria":{"true":"讻谉, 讝讛 谞讬住讬讜谉 诪专诪讛","false":"诇讗, 讝讜 讛讜讚注讛 诇讙讬讟讬诪讬转, 讙诐 讗诐 讬砖 讘讛 拽讬砖讜专 讗讜 讘拽砖转 转砖诇讜诐"}}}}
{"status":"ready","model":"laya"}
{"model": "nitzotz", "family": "nitzotz", "route": "forced nitzotz", "answers": {"scam": {"type": "noul", "action": {"act_probability": 0.908}, "confidence": 0.9536, "noul": 0.0464}}, "usage": {"input_tokens": 66, "output_tokens": 0, "latency_ms": 121.783}, "id": "1"}
noul 讛讬讗 讛讛住转讘专讜转 砖讛讟注谞讛 谞讻讜谞讛. 讻诪讛 砖讗诇讜转 讘讘拽砖讛 讗讞转 谞注谞讜转 讬讞讚, 讘诪注讘专 讗讞讚. 注诇 诪讞砖讘 讘诇讬 讻专讟讬住 诪住讱 诪砖转诪砖讬诐 讘---device cpu.
诪讘讞谞讬诐
8 住讟讬诐 砖诇 诪讘讞谉, 3,990 砖讗诇讜转 讘住讱 讛讻讜诇, 砖谞谞注诇讜 讘讟讘讬注转 讗爪讘注 诇驻谞讬 砖谞讜爪专 驻专讬讟 讗讬诪讜谉 讗讞讚. 砖诇讜砖转 讛诪讜讚诇讬诐 注谞讜 注诇 讗讜转谉 砖讗诇讜转 讘讚讬讜拽. 砖转讬 谞拽讜讚讜转 讛讛砖讜讜讗讛 讛谉 诪讜讚诇讬 laya 讛驻转讜讞讬诐 讛讗讞专讬诐 砖拽讜专讗讬诐 注讘专讬转: RoeiG/laya-hebrew (专讬砖讬讜谉 CC-BY-NC-SA-4.0, 诇砖讬诪讜砖 诇讗 诪住讞专讬 讘诇讘讚) 讜-laya-multilingual (Apache-2.0).
讛讟讘诇讛 讛诪诇讗讛. 讗讞讜讝 讛转砖讜讘讜转 讛谞讻讜谞讜转, 讛讟讜讘 讘讬讜转专 讘讻诇 砖讜专讛 诪讜讚讙砖. 砖转讬 讛注诪讜讚讜转 讛讗讞专讜谞讜转 讗讜诪专讜转 讗诐 讛讛讘讚诇 砖诇 谞讬爪讜抓 诪讛诪讜讚诇 讛讝讛 讗诪讬转讬 讗讜 砖讗讜诇讬 讝讛 诪讝诇 (诪讘讞谉 McNemar 诪讚讜讬拽 注诇 讗讜转谉 砖讗诇讜转 讘讚讬讜拽): "讟讜讘 讬讜转专" 讗讜 "讞诇砖 讬讜转专" 驻讬专讜砖讜 p 拽讟谉 诪-0.05, "转讬拽讜" 驻讬专讜砖讜 砖讛讛讘讚诇 讬讻讜诇 诇讛讬讜转 诪拽专讬.
| 诪讘讞谉 (诪住驻专 砖讗诇讜转) | 谞讬爪讜抓 | RoeiG (诇讗 诪住讞专讬) | laya-multilingual | 谞讬讞讜砖 | 谞讬爪讜抓 诪讜诇 RoeiG | 谞讬爪讜抓 诪讜诇 laya-multilingual |
|---|---|---|---|---|---|---|
| 讛讜谞讗讛 讗讜 诇讗? (298 讛讜讚注讜转) | 92.0% | 32.2% | 50.3% | 50.0% | 讟讜讘 讬讜转专 p<0.001 |
讟讜讘 讬讜转专 p<0.001 |
| 讗讜转讜 讚讘专, 专拽 讛诪拽专讬诐 讛拽砖讬诐 (65) | 83.1% | 32.3% | 35.4% | 50.0% | 讟讜讘 讬讜转专 p<0.001 |
讟讜讘 讬讜转专 p<0.001 |
| 住讜讙 讛讛讜讚注讛, 6 讗驻砖专讜讬讜转 (298) | 77.8% | 52.7% | 20.8% | 16.7% | 讟讜讘 讬讜转专 p<0.001 |
讟讜讘 讬讜转专 p<0.001 |
| 讗讜转讜 讚讘专, 专拽 讛诪拽专讬诐 讛拽砖讬诐 (65) | 67.7% | 44.6% | 24.6% | 16.7% | 讟讜讘 讬讜转专 p=0.01 |
讟讜讘 讬讜转专 p<0.001 |
| 住讜讙 驻谞讬讬转 转诪讬讻讛, 5 讗驻砖专讜讬讜转 (30) | 90.0% | 80.0% | 56.7% | 20.0% | 转讬拽讜 p=0.45 |
讟讜讘 讬讜转专 p=0.01 |
| 讚讞讬驻讜转 讛驻谞讬讬讛, 5 专诪讜转 (30) | 70.0% | 36.7% | 33.3% | 20.0% | 讟讜讘 讬讜转专 p=0.03 |
讟讜讘 讬讜转专 p=0.02 |
| 诇拽讜讞 诪砖诇诐? 讻谉/诇讗 (30) | 80.0% | 56.7% | 50.0% | 50.0% | 讟讜讘 讬讜转专 p=0.02 |
转讬拽讜 p=0.06 |
| 讻讜讜谞转 驻拽讜讚讛 拽讜诇讬转, 20 讗驻砖专讜讬讜转 (500) | 90.0% | 72.6% | 47.4% | 5.0% | 讟讜讘 讬讜转专 p<0.001 |
讟讜讘 讬讜转专 p<0.001 |
| 讻讜讜谞转 驻拽讜讚讛 拽讜诇讬转, 4 讗驻砖专讜讬讜转 (500) | 97.2% | 90.8% | 69.4% | 25.0% | 讟讜讘 讬讜转专 p<0.001 |
讟讜讘 讬讜转专 p<0.001 |
| 谞讜砖讗 砖诇 讬讚讬注讛, 7 讗驻砖专讜讬讜转 (204) | 79.4% | 82.3% | 66.2% | 14.3% | 转讬拽讜 p=0.36 |
讟讜讘 讬讜转专 p=0.001 |
| 讛讗诐 讛拽讟注 转讜诪讱 讘转砖讜讘讛? (600) | 94.8% | 94.2% | 48.8% | 50.0% | 转讬拽讜 p=0.69 |
讟讜讘 讬讜转专 p<0.001 |
| 转砖讜讘讛 住讘讬专讛 砖讛拽讟注 诇讗 谞讜转谉 (600) | 89.2% | 53.0% | 53.7% | 50.0% | 讟讜讘 讬讜转专 p<0.001 |
讟讜讘 讬讜转专 p<0.001 |
| 讛讘谞转 讛谞拽专讗, 4 讗驻砖专讜讬讜转 (900) | 63.2% | 75.4% | 31.4% | 25.0% | 讞诇砖 讬讜转专 p<0.001 |
讟讜讘 讬讜转专 p<0.001 |
讘讛讘谞转 讛谞拽专讗 砖诇 Belebele 谞讬爪讜抓 诪拽讘诇 63.2%, 驻讞讜转 诪-RoeiG/laya-hebrew (75.4%). 谞讬爪讜抓 讘谞讜讬 诇讛讞诇讟讜转 注诇 讛讜讚注讜转, 诇讗 诇讛讘谞讛 砖诇 拽讟注讬诐 讗专讜讻讬诐.
讗讬讱 诇拽专讜讗 讗转 讝讛:
- MASSIVE (驻拽讜讚讜转 拽讜诇讬讜转): 谞讬爪讜抓 讗讜诪谉 注诇 驻拽讜讚讜转 讛讗讬诪讜谉 砖诇 MASSIVE. 驻拽讜讚讜转 讛诪讘讞谉 讗讞专讜转, 讗讘诇 谞讻转讘讜 讘讬讚讬 讗讜转诐 讗谞砖讬诐 讜讘讗讜转讜 住讙谞讜谉, 讜诇讻谉 讛诪讘讞谉 讛讝讛 拽诇 讬讜转专 诇谞讬爪讜抓 诪讗砖专 诇讗讞专讬诐.
- HeQ: 谞讬爪讜抓 讗讜诪谉 注诇 拽讟注讬 讛讗讬诪讜谉 砖诇 HeQ. 拽讟注讬 讛诪讘讞谉 讗讞专讬诐, 讜拽讟注讬 讗讬诪讜谉 砖讞驻驻讜 诇拽讟注 诪讘讞谉 讛讜住专讜.
- 讘砖讜专讜转 砖诇 驻谞讬讜转 讛转诪讬讻讛 讬砖 专拽 30 砖讗诇讜转. 砖讜诐 讛讘讚诇 砖诐 诇讗 讗诪讬谉.
- 诪讘讞谉 讛住驻讗诐 讜讛讛讜谞讗讜转 诇讗 诪讗讜讝谉: 70% 诪讛讛讜讚注讜转 讘讜 讛谉 诇讗 讛讜谞讗讛. 拽讜 讛"谞讬讞讜砖" (50%) 讛讜讗 讛讟诇转 诪讟讘注, 诇讗 讛讗住讟专讟讙讬讛 讛注讬讜讜专转 讛讟讜讘讛 讘讬讜转专.
诇诪讛 讗驻砖专 诇住诪讜讱 注诇讬讜
诇讛住转讘专讜讬讜转 讬砖 诪砖诪注讜转. 拽讞讜 讗转 讻诇 讛转砖讜讘讜转 砖讘讛谉 谞讬爪讜抓 讗诪专 砖讛讜讗 讘讟讜讞 讘注专讱 讘-70 注讚 80%, 讜住驻专讜 讻诪讛 诪讛谉 讛讬讜 谞讻讜谞讜转.
讛讙专祝 注讜砖讛 讗转 讝讛 诇讻诇 专诪转 讘讬讟讞讜谉, 注诇 3,990 砖讗诇讜转 诪讘讞谉. 讻砖讛谞拽讜讚讜转 诪注诇 讛拽讜, 谞讬爪讜抓 爪讜讚拽 讬讜转专 诪诪讛 砖讛讜讗 讗讜诪专 (讛讜讗
爪谞讜注). 讻砖讛谉 诪转讞转, 讛讜讗 讘讟讜讞 讘注爪诪讜 讬讜转专 诪讚讬. 讝讛 诪讛 砖诪讗驻砖专 诇拽讘讜注 住驻讬诐 (讘驻专拽 讛讘讗). 讛讻讬讜诇 谞注砖讛 注诇 4,825 驻专讬讟讬 讗讬诪讜谉
砖讛讜驻专讚讜 诪专讗砖, 讗祝 驻注诐 诇讗 注诇 讛诪讘讞谉 (calibration.json).
讛讜讗 讘讗诪转 拽讜专讗 讗转 讛讟拽住讟. 讻砖诪讜讞拽讬诐 讗转 讛讛讜讚注讛 讜诪砖讗讬专讬诐 专拽 讗转 讛砖讗诇讛, 讛讚讬讜拽 讬讜专讚 诇-70.5% 讘砖讗诇转 讛讛讜谞讗讛 讜诇-7.7% 讘住讜讙 讛讛讜讚注讛. 讻诇讜诪专 讛转砖讜讘讜转 讘讗讜转 诪讛讛讜讚注讛, 诇讗 诪讛谞讬住讜讞 砖诇 讛砖讗诇讛.
谞讬住讜讞 讗讞专 砖诇 讛砖讗诇讛 讻诪注讟 诇讗 诪砖谞讛 讗转 讛转砖讜讘讛. 砖讗诇谞讜 1,786 砖讗诇讜转 诪讘讞谉 讘-6 谞讬住讜讞讬诐 砖讜谞讬诐 注诐 讗讜转讛 诪砖诪注讜转. 讘-5.0% 诪讛谉 讛转砖讜讘讛 诇讗 讛讬讬转讛 讝讛讛 讘讻诇 6 讛谞讬住讜讞讬诐. 讗祝 谞讬住讜讞 诇讗 讙讜专诐 诇讜 诇转转 转砖讜讘讛 拽讘讜注讛 讗讞转 诇讻诇 讛驻专讬讟讬诐 砖诇 诪讘讞谉. 讛讬讜爪讗 诪谉 讛讻诇诇 讛讜讗 砖讗诇转 讛讚讞讬驻讜转 (专讗讜 诪讙讘诇讜转).
讛讜讗 诪讛讬专 注诇 讞讜诪专讛 专讙讬诇讛. 注诇 诪讞砖讘 谞讬讬讚 注诐 诪注讘讚 Intel Core Ultra 9 285H 讜讻专讟讬住 讛诪住讱 讛诪讜讘谞讛 Arc 140T, 讜讜讬谞讚讜住 11, 砖讗诇讛 讗讞转 讘讻诇 驻注诐:
| 讗讬讱 诪专讬爪讬诐 | 讞爪讬讜谉 注诇 50 砖讗诇讜转 讘讚讬拽讛 (讘诪诪讜爪注 135 讟讜拽谞讬诐) | 讛讜讚注讛 拽爪专讛 (38 讟讜拽谞讬诐) | 拽诇讟 讗专讜讱 (430 讟讜拽谞讬诐) |
|---|---|---|---|
| laya.exe, 拽讜讘抓 Q8, 讻专讟讬住 诪住讱 (Vulkan) | 48 ms | 37 ms | 127 ms |
| laya.exe, 拽讜讘抓 F16, 讻专讟讬住 诪住讱 (Vulkan) | 50 ms | 43 ms | 111 ms |
| 驻讬讬转讜谉 (laya), 讻专讟讬住 诪住讱 (PyTorch XPU) | 60 ms | 44 ms | 139 ms |
| laya.exe, 拽讜讘抓 Q8, 诪注讘讚 讘诇讘讚 (16 转讛诇讬讻讜谞讬诐) | 309 ms | 160 ms | 1137 ms |
| 驻讬讬转讜谉 (laya), 诪注讘讚 讘诇讘讚 | 224 ms | 112 ms | 1022 ms |
讘注诪讜讚讜转 砖诇 讛讛讜讚注讛 讛拽爪专讛 讜讛拽诇讟 讛讗专讜讱 讗讜转讛 砖讗诇讛 专爪讛 30 驻注诪讬诐, 讗讞专讬 5 讛专爪讜转 讞讬诪讜诐. 讛讝诪谞讬诐 注诇 讛诪讞砖讘 讛讝讛 诪砖转谞讬诐 讛专讘讛 讘讬谉 讛驻注诇讛 诇讛驻注诇讛 (诪讚讬讚讛 拽讜讚诪转 砖诇 讗讜转讛 讛讙讚专讛 讬爪讗讛 讗讬讟讬转 驻讬 讻诪讛), 讗讝 讗诇讛 诪住驻专讬诐 讘拽讬专讜讘.
拽讜讘爪讬 讛-GGUF 谞讜转谞讬诐 讻诪注讟 讗转 讗讜转谉 转砖讜讘讜转 讻诪讜 诪讜讚诇 讛驻讬讬转讜谉. 讘讛砖讜讜讗讛 诇诪讜讚诇 讛驻讬讬转讜谉 讛诪诇讗 注诇 讛诪注讘讚:
| 拽讜讘抓, 诪讻砖讬专 | 讗讜转讛 转砖讜讘讛 诪讜讘讬诇讛, 50 砖讗诇讜转 | 讛驻专砖 讛住转讘专讜转 诪专讘讬 | 讗讜转讛 转砖讜讘讛 诪讜讘讬诇讛, 596 砖讗诇讜转 住驻讗诐 | 讛驻专砖 诪专讘讬 | 讛驻专砖 诪诪讜爪注 |
|---|---|---|---|---|---|
| Q8, 讻专讟讬住 诪住讱 | 50/50 | 0.0375 | 596/596 | 0.0166 | 0.00130 |
| F16, 讻专讟讬住 诪住讱 | 50/50 | 0.0017 | 596/596 | 0.0015 | 0.00020 |
| Q8, 诪注讘讚 | 50/50 | 0.0549 | 595/596 | 0.0207 | 0.00192 |
| F16, 诪注讘讚 | 50/50 | 0.0032 | 596/596 | 0.0012 | 0.00018 |
驻注专 砖诇 0.01 驻讬专讜砖讜, 诇诪砖诇, 0.83 诪讜诇 0.84. 讛砖讗诇讜转 讛诪注讟讜转 砖讘讛谉 讛转砖讜讘讛 讛诪讜讘讬诇讛 诪砖转谞讛 讛谉 讻讗诇讛 砖讘讛谉 砖转讬 讛转砖讜讘讜转 讛讟讜讘讜转 讛讬讜 讻诪注讟 砖讜讜转. 讘砖转讬 砖讗诇讜转 讛住驻讗诐 讬讞讚 拽讜讘抓 Q8 注诇 讻专讟讬住 讛诪住讱 爪讜讚拽 讘-84.7%, 诪讜诇 84.7% 诇诪讜讚诇 讛驻讬讬转讜谉. 拽讜讘抓 F16 拽专讜讘 讬讜转专 (84.7%).
砖讬诪讜砖 讘注住拽
讛专注讬讜谉 ("专诪讝讜专"). 讻诇 讛讜讚注讛 谞讻谞住转 诪拽讘诇转 砖讗诇讛 诪讛讬专讛 讗讞转 讗讜 讻诪讛. 讛讛住转讘专讜转 诪讞诇讬讟讛 诪讛 拽讜专讛 讛诇讗讛:
- 讬专讜拽 (讘讟讜讞 砖讝讛 讘住讚专): 讟讬驻讜诇 讗讜讟讜诪讟讬, 转讬讜拽, 转讬讜讙, 谞讬转讜讘.
- 爪讛讜讘 (诇讗 讘讟讜讞): 诇讘讚讬拽讛 砖诇 讗讚诐, 讗讜 砖诇 诪讜讚诇 砖驻讛 讙讚讜诇 讗诐 讗转诐 诪砖转诪砖讬诐 讘讜.
- 讗讚讜诐 (讘讟讜讞 砖讝讜 讛讜谞讗讛): 讞住讬诪讛 讗讜 讛住讙专.
讛砖专讟讜讟 讛讜讗 讚讜讙诪讛 注诇 298 讛讜讚注讜转 讛诪讘讞谉. 讛住祝 讛注诇讬讜谉, 0.49, 讛讜讗 讛住祝 砖谞讘讞专 诇砖讗诇转 讛讛讜谞讗讛 注诇 1188 讛讜讚注讜转 讗讬诪讜谉 砖讛讜驻专讚讜 诪专讗砖 (讗祝 驻注诐 诇讗 注诇 讛诪讘讞谉). 讛住祝 讛转讞转讜谉, 0.35, 谞讘讞专 讘讬讚. 讗讬转诐 207 讛讜讚注讜转 讛讜诇讻讜转 诇讬专讜拽 (13 诪讛谉 讛谉 讘注爪诐 讛讜谞讗讛), 9 诇爪讛讜讘 (2 讛讜谞讗讜转), 讜-82 诇讗讚讜诐 (9 诪讛谉 讘注爪诐 转拽讬谞讜转). 讻诇讜诪专 讗讚讜诐 爪专讬讱 诇讛讬讜转 "讛住讙专 讜讘讚讬拽讛", 诇讗 "诪讞讬拽讛". 讘讞专讜 住驻讬诐 诪砖诇讻诐 注诇 诪讚讙诐 砖诇 讛讛讜讚注讜转 砖诇讻诐, 讜讛讞诇讬讟讜 讻诪讛 讟注讜讬讜转 讘讬专讜拽 讜讘讗讚讜诐 讗转诐 诪讜讻谞讬诐 诇拽讘诇. 讘住祝 0.49 诇讘讚讜, 转砖讜讘转 讛讛讜谞讗讛 谞讻讜谞讛 讘-92.3% 诪讛诪拽专讬诐 讘住讱 讛讻讜诇 讜讘-83.1% 注诇 65 讛诪拽专讬诐 讛拽砖讬诐 (讘住祝 0.50: 92.0% 讜-83.1%).
讛诪讜讚诇 讝讜诇 诪住驻讬拽 讻讚讬 诇讛专讬抓 讗讜转讜 注诇 讻诇 讛讜讚注讛. 讛讗讚诐 (讗讜 诪讜讚诇 讛砖驻讛) 专讜讗讛 专拽 讗转 讛讞诇拽 讛爪讛讜讘. 讗诇 转砖转诪砖讜 讘讜 讻拽讜 讛讙谞讛 讬讞讬讚 讘讛讞诇讟讜转 砖讬讻讜诇讜转 诇驻讙讜注 讘诪讬砖讛讜.
谞转讜谞讬 讛讗讬诪讜谉 讜砖拽讬驻讜转
砖谞讬 砖诇讘讬 讗讬诪讜谉, 注诇 讻专讟讬住 诪住讱 砖讻讜专:
- 诇诇诪讜讚 诇拽专讜讗. HalleluBERT-large 讗讜诪谉 拽讜讚诐 诇诪爪讜讗 讗转 讛转砖讜讘讛 诇砖讗诇讛 讘转讜讱 拽讟注, 注诇 27,085 砖讗诇讜转 讗讬诪讜谉 砖诇 HeQ (CC BY 4.0). 拽讟注讬诐 砖讞驻驻讜 诇拽讟注 诪讘讞谉 讛讜住专讜.
- 诇诇诪讜讚 诇讛讞诇讬讟. 诪注诇讬讜 讛讜谞讞 专讗砖 讛讞诇讟讜转 砖诇 laya, 讜讻诇 讛诪讜讚诇 讗讜诪谉 注诇 116,854 驻专讬讟讬诐 (112,029 诇讗讬诪讜谉, 讜-4,825 讛讜驻专讚讜 诪专讗砖 讻讚讬 诇讘讞讜专 讗转 讛讟讜讘 诪讘讬谉 2 诪注讘专讬诐 讜诇讻讬讬诇 讗转 讛讟诪驻专讟讜专讜转).
| 诪拽讜专 | 驻专讬讟讬诐 | 专讬砖讬讜谉 | 讗讬讱 谞讜爪专 |
|---|---|---|---|
| 住讬谞转讟讬: 住讜讙 讛讜讚注讛, 6 住讜讙讬诐 | 14,000 | 讛驻诇讟 砖诇 讛诪讜专讛 砖讬讬讱 诇谞讜 (DeepSeek, MIT; Gemma, Apache-2.0) | 谞讻转讘 讘讬讚讬 DeepSeek V4.1 Flash, 住讜诪谉 讘谞驻专讚 讘讬讚讬 砖谞讬 讛诪讜讚诇讬诐 |
| 住讬谞转讟讬: 谞讬转讜讘 诇诪讞诇拽讛 (3 注讚 6 讗驻砖专讜讬讜转) | 8,750 | 讛驻诇讟 砖诇 讛诪讜专讛 砖讬讬讱 诇谞讜 (DeepSeek, MIT; Gemma, Apache-2.0) | 谞讻转讘 讘讬讚讬 DeepSeek V4.1 Flash, 住讜诪谉 讘谞驻专讚 讘讬讚讬 砖谞讬 讛诪讜讚诇讬诐 |
| 住讬谞转讟讬: 讟注谞讜转 讻谉/诇讗 注诇 讛讜讚注讛 | 8,750 | 讛驻诇讟 砖诇 讛诪讜专讛 砖讬讬讱 诇谞讜 (DeepSeek, MIT; Gemma, Apache-2.0) | 谞讻转讘 讘讬讚讬 DeepSeek V4.1 Flash, 住讜诪谉 讘谞驻专讚 讘讬讚讬 砖谞讬 讛诪讜讚诇讬诐 |
| 住讬谞转讟讬: 讚讞讬驻讜转, 5 专诪讜转 | 3,500 | 讛驻诇讟 砖诇 讛诪讜专讛 砖讬讬讱 诇谞讜 (DeepSeek, MIT; Gemma, Apache-2.0) | 谞讻转讘 讘讬讚讬 DeepSeek V4.1 Flash, 住讜诪谉 讘谞驻专讚 讘讬讚讬 砖谞讬 讛诪讜讚诇讬诐 |
| HeQ: 讛讗诐 讛转砖讜讘讛 讛诪讜爪注转 谞转诪讻转 (讻谉/诇讗) | 6,000 | CC BY 4.0 | 转讜讜讬讜转 讛讝讛讘 砖诇 讛诪讗讙专, 诪驻讬爪讜诇 讛讗讬诪讜谉 讘诇讘讚 |
| HeQ: 转砖讜讘讛 住讘讬专讛 诇砖讗诇讛 砖讗讬谉 诇讛 转砖讜讘讛 | 9,750 (诪转讜讻诐 4,750 注讜转拽讬诐 讞讜讝专讬诐) | CC BY 4.0 | 转讜讜讬讜转 讛讝讛讘 砖诇 讛诪讗讙专, 诪驻讬爪讜诇 讛讗讬诪讜谉 讘诇讘讚 |
| HeQ: 拽专讬讗讛, 4 讗驻砖专讜讬讜转 | 5,000 | CC BY 4.0 | 转讜讜讬讜转 讛讝讛讘 砖诇 讛诪讗讙专, 诪驻讬爪讜诇 讛讗讬诪讜谉 讘诇讘讚 |
| MASSIVE he-IL: 讻讜讜谞转 驻拽讜讚讛 拽讜诇讬转 | 7,000 | CC BY 4.0 | 转讜讜讬讜转 讛讝讛讘 砖诇 讛诪讗讙专, 诪驻讬爪讜诇 讛讗讬诪讜谉 讘诇讘讚 |
| 讛讜谞讗讛 讗讜 诇讗, 讛讜谞讗讜转 诪讜住讜讜转 讜讛讜讚注讜转 讗诪讬转讬讜转 砖谞专讗讜转 讞砖讜讚讜转 | 6,000 | 驻诇讟 诪讜讚诇讬诐, 砖讬讬讱 诇谞讜 (DeepSeek, MIT; Gemma, Apache-2.0) | 谞讻转讘 讘讬讚讬 DeepSeek V4.1 Flash 诇驻讬 转专讞讬砖讬诐 砖谞讻转讘讜 注诐 GPT. 谞砖诪专 专拽 讗诐 砖谞讬 讛诪住诪谞讬诐 讛住讻讬诪讜 注诐 讛讻讜转讘 |
| 住讜讙 讛讛讜讚注讛 砖诇 讗讜转谉 讛讜讚注讜转 | 5,374 | 驻诇讟 诪讜讚诇讬诐, 砖讬讬讱 诇谞讜 | 讛住讜讙 砖砖谞讬 讛诪住诪谞讬诐 讛住讻讬诪讜 注诇讬讜 |
| 注讜转拽讬诐 砖诇 驻专讬讟讬诐 拽讬讬诪讬诐 注诐 讛砖讗诇讛 讘谞讬住讜讞 讗讞专 | 21,199 | 讻诪讜 讛驻专讬讟 讛诪拽讜专讬. 讛谞讬住讜讞讬诐 谞讻转讘讜 注诐 GPT | 讛砖讗诇讛 讗讜 讛讗驻砖专讜讬讜转 讘讗讞讚 诪-376 谞讬住讜讞讬诐 讞诇讜驻讬讬诐 砖谞讻转讘讜 注诐 GPT. 讛转讜讜讬转 诇拽讜讞讛 诪讛驻专讬讟 讛诪拽讜专讬 |
| 拽专讬讗讛, 4 讗驻砖专讜讬讜转 ("讗讬讝讜 诪讛讘讗讜转 诇讗", 转砖讜讘讛 讘诪讬诇讬诐 讗讞专讜转, 讻诪讛 诪砖驻讟讬诐) | 6,825 | 拽讟注讬诐: FineWeb-2, ODC-By 1.0. 砖讗诇讜转: 驻诇讟 诪讜讚诇讬诐, 砖讬讬讱 诇谞讜 | 谞讻转讘 讘讬讚讬 DeepSeek V4.1 Flash 讜谞讘讚拽 讘讬讚讬 Gemma 4 31B 注诐 讛拽讟注 讜砖讜讘 讘诇讬 讛拽讟注. 谞讝专拽 讗诐 讗驻砖专 讛讬讛 诇注谞讜转 讘诇讬 讛拽讟注 |
| 谞讜砖讗 砖诇 拽讟注, 7 讗驻砖专讜讬讜转 | 4,000 | 拽讟注讬诐: FineWeb-2, ODC-By 1.0. 转讜讜讬讜转: 驻诇讟 诪讜讚诇讬诐 | 住讜诪谉 讘讬讚讬 DeepSeek V4.1 Flash 讜-Gemma 4 31B |
| 讛讗诐 讛砖讜诇讞 诇拽讜讞 诪砖诇诐 拽讬讬诐 (讻谉/诇讗) | 575 | 驻诇讟 诪讜讚诇讬诐, 砖讬讬讱 诇谞讜. 谞讬住讜讞讬 讛砖讗诇讛 谞讻转讘讜 注诐 GPT | 驻谞讬讜转 转诪讬讻讛 砖谞讻转讘讜 讘讬讚讬 DeepSeek V4.1 Flash 讜谞讘讚拽讜 讘讬讚讬 Gemma 4 31B |
| 讛讜谞讗讛 讗讜 诇讗: 讗讝讛专讜转 诪驻谞讬 讛讜谞讗讛, 讛讜谞讗讜转 拽爪专讜转 讘诇讬 拽讬砖讜专, 讜讛讜讚注讜转 诇讙讬讟讬诪讬讜转 砖谞专讗讜转 讻诪讜讛谉 | 3,746 | 驻诇讟 诪讜讚诇讬诐, 砖讬讬讱 诇谞讜 (DeepSeek, MIT; Gemma, Apache-2.0). 谞讬住讜讞讬 讛砖讗诇讛 谞讻转讘讜 注诐 GPT | 谞讻转讘 讘讬讚讬 DeepSeek V4.1 Flash. 谞砖诪专 专拽 讗诐 砖谞讬 讛诪住诪谞讬诐 讛住讻讬诪讜 注诐 讛讻讜转讘, 讞讜抓 诪讛讜谞讗讛 砖专拽 DeepSeek 讝讬讛讛, 砖谞砖诪专讛 注诐 转讜讜讬转 专讻讛 讬讜转专 (70%) |
| 住讜讙 讛讛讜讚注讛 砖诇 讗讜转谉 讛讜讚注讜转 | 1,901 | 驻诇讟 诪讜讚诇讬诐, 砖讬讬讱 诇谞讜 | 讛住讜讙 砖砖谞讬 讛诪住诪谞讬诐 讛住讻讬诪讜 注诇讬讜, 专拽 讻砖讛讜讗 诪转讗讬诐 诇讛讞诇讟讛 讗诐 讝讜 讛讜谞讗讛 |
| "讗祝 讗讞转 诪讛讗驻砖专讜讬讜转": 驻专讬讟讬 讗讬诪讜谉 拽讬讬诪讬诐 注诐 讗驻砖专讜转 讗讞转 砖砖讜谞转讛 | 4,484 | 讻诪讜 讛驻专讬讟 讛诪拽讜专讬 (MASSIVE 讜-HeQ: CC BY 4.0. 拽专讬讗讛: 拽讟注讬诐 诪-FineWeb-2, ODC-By 1.0) | 讛转砖讜讘讛 讛谞讻讜谞讛 讛讜住专讛 讜谞讜住驻讛 讗驻砖专讜转 "讗祝 讗讞转 诪讛讗驻砖专讜讬讜转". 讘砖诇讬砖 诪讛诐 讛讜住专讛 讘诪拽讜诐 讝讛 讗驻砖专讜转 砖讙讜讬讛, 讻讱 砖"讗祝 讗讞转" 砖讙讜讬讛 砖诐 |
- 讛讜讚注讜转 住讬谞转讟讬讜转 (35,000 驻专讬讟讬诐). DeepSeek V4.1 Flash (MIT) 讻转讘 讛讜讚注讜转 SMS, 讜讜讗讟住讗驻 讜诪讬讬诇 讘住讙谞讜谉 讬砖专讗诇讬 诇驻讬 转讜讻谞讬转 (讛转讜讜讬转 讛诪转讜讻谞谞转, 谞讜砖讗, 诪砖诇讘, 诪住驻专讬 讟诇驻讜谉 讜拽讬砖讜专讬诐 诪讝讜讬驻讬诐 讜诪讙讜讜谞讬诐). 讗讞专 讻讱 DeepSeek 讜-Gemma 4 31B (Apache-2.0) 住讬诪谞讜 讻诇 讛讜讚注讛, 讻诇 讗讞讚 诇讘讚. 讛讜讚注讛 谞砖诪专讛 专拽 讗诐 砖谞讬讛诐 讛住讻讬诪讜. 讛诐 讛住讻讬诪讜 注诇 94% 诪讛讛讜讚注讜转, 讜谞砖诪专讜 93.7% 诪转讜讱 37,429 讛讛讜讚注讜转 砖谞讻转讘讜. 砖谞讬讛诐 专爪讜 讚专讱 DeepInfra (讚专讱 OpenRouter), 讘诇讬 砖诪讬专转 谞转讜谞讬诐 讜讘诇讬 诪爪讘 "讞砖讬讘讛". 讛谞转讜谞讬诐 讛住讬谞转讟讬讬诐 诇讗 诪驻讜专住诪讬诐.
- 拽专讬讗讛, 讛讜谞讗讜转 拽砖讜转 讜驻谞讬讜转 转诪讬讻讛. 6,825 砖讗诇讜转 拽专讬讗讛 注诇 拽讟注讬诐 诪讛专砖转 讘注讘专讬转 诪转讜讱 FineWeb-2 (ODC-By 1.0): "讗讬讝讜 诪讛讘讗讜转 诇讗 谞讻讜谞讛 讗讜 诇讗 诪讜讝讻专转" (讻诇 讗讞转 注诐 砖讗诇讛 讞讬讜讘讬转 爪诪讜讚讛 注诇 讗讜转讜 拽讟注), 转砖讜讘讜转 砖谞讗诪专讜转 讘诪讬诇讬诐 讗讞专讜转, 讜转砖讜讘讜转 砖讚讜专砖讜转 讻诪讛 诪砖驻讟讬诐. DeepSeek V4.1 Flash 讻转讘 讗讜转谉, 讜-Gemma 4 31B 讘讚拽 讻诇 讗讞转 驻注诪讬讬诐, 注诐 讛拽讟注 讜讘诇讬 讛拽讟注. 砖讗诇讛 砖讗驻砖专 讛讬讛 诇注谞讜转 注诇讬讛 讘诇讬 讛拽讟注 谞讝专拽讛. 6,000 讛讜讚注讜转 砖拽砖讛 诇讛讘讞讬谉 讘讬谞讬讛谉 (讛讜谞讗讜转 诪讜住讜讜转, 讜讛讜讚注讜转 讗诪讬转讬讜转 砖谞专讗讜转 讞砖讜讚讜转), 砖谞讻转讘讜 讘讬讚讬 DeepSeek 讜谞砖诪专讜 专拽 讗诐 砖谞讬 讛诪住诪谞讬诐 讛住讻讬诪讜 讝讛 注诐 讝讛 讜讙诐 注诐 讛讻讜转讘. 4,000 拽讟注讬诐 砖住讜诪谞讜 诇驻讬 谞讜砖讗, 575 驻谞讬讜转 转诪讬讻讛 诇砖讗诇转 讛诇拽讜讞 讛诪砖诇诐, 讜-4,750 驻专讬讟讬 HeQ 讞讜讝专讬诐, 讻讚讬 砖诇讬讻讜诇讜转 砖诇 HeQ 讬讬砖讗专 诪砖拽诇 讘转注专讜讘转.
- 讚驻讜住讬 讛讜谞讗讛 砖谞专讗讜 讘讛讜讚注讜转 讗诪讬转讬讜转. 讛讜讚注讜转 讗诪讬转讬讜转 讘注讘专讬转 讛专讗讜 砖转讬 谞拽讜讚讜转 讞诇砖讜转: 讗讝讛专讜转 讗诪讬转讬讜转 诪驻谞讬 讛讜谞讗讛 (诪讘谞拽讬诐, 诪讛诪砖讟专讛, 诪讞讘专讜转) 砖住讜诪谞讜 讻讛讜谞讗讛, 讜讛讜谞讗讜转 拽爪专讜转 讘诇讬 拽讬砖讜专 砖驻讜住驻住讜. 诇讻谉 DeepSeek V4.1 Flash 讻转讘 注讜讚 3,746 讛讜讚注讜转: 1,000 讗讝讛专讜转 诇讙讬讟讬诪讬讜转 诪驻谞讬 讛讜谞讗讛, 1,476 讛讜谞讗讜转 拽爪专讜转 讘诇讬 拽讬砖讜专 (讞讜讘 拽讟谉 讗讜 讗讙专讛 砖诇讗 砖讜诇诪讜, "讞讘专" 注诐 诪住驻专 讞讚砖, 诪转谞讛, 讗讬砖讜专 转砖诇讜诐 诪讝讜讬祝, 讗驻诇讬拽爪讬讬转 转砖诇讜诪讬诐), 389 讝讜讙讜转 砖诇 讛讜谞讗讛 讜讗讝讛专讛 注诇 讗讜转讛 讛讜谞讗讛, 讜-492 讛讜讚注讜转 诇讙讬讟讬诪讬讜转 砖谞专讗讜转 讻诪讜 讛讛讜谞讗讜转 讛讗诇讛. 讻诇诇 讛砖诪讬专讛 讛砖转谞讛 讘讛谉: 讛讜讚注讛 谞砖诪专讛 讗诐 砖谞讬 讛诪住诪谞讬诐 讛住讻讬诪讜 注诐 讛讻讜转讘, 讜讛讜谞讗讛 砖专拽 DeepSeek 讝讬讛讛 谞砖诪专讛 讙诐 讻谉, 注诐 转讜讜讬转 专讻讛 讬讜转专 (70% 讘诪拽讜诐 拽专讜讘 诇-100%). 134 讛讜讚注讜转 讛谉 诪讛住讜讙 讛讝讛. 讛讜讚注讜转 砖讚诪讜 诇讗讞转 诪讛讛讜讚注讜转 讛讗诪讬转讬讜转 砖讘讚拽谞讜 谞讝专拽讜 诇驻谞讬 讛住讬诪讜谉, 讜讛讛讜讚注讜转 讛讗诪讬转讬讜转 注爪诪谉 诇讗 砖讬诪砖讜 诇讗讬诪讜谉. 134 砖讜专讜转 讗讬诪讜谉 拽讬讬诪讜转 (67 讛讜讚注讜转: 讘拽砖讛 诪"诪住驻专 讞讚砖" 讗讜 注诇 讞讜讘 拽讟谉, 讘诇讬 拽讬砖讜专) 转讜讬讙讜 诪讞讚砖 讻讛讜谞讗讛 讗讞专讬 砖讛诪住诪谞讬诐 拽讘注讜 砖讛谉 讛讜谞讗讛 (注诐 转讜讜讬转 70% 讻砖专拽 DeepSeek 拽讘注). 4,484 驻专讬讟讬 "讗祝 讗讞转 诪讛讗驻砖专讜讬讜转" 谞讘谞讜 诪驻专讬讟讬 讗讬诪讜谉 拽讬讬诪讬诐 砖诇 MASSIVE, HeQ 讜拽专讬讗讛: 讘-2,990 诪讛诐 讛转砖讜讘讛 讛谞讻讜谞讛 讛讜住专讛 讜谞讜住驻讛 讗驻砖专讜转 "讗祝 讗讞转 诪讛讗驻砖专讜讬讜转", 讜讘-1,494 讛讜住专讛 讘诪拽讜诐 讝讛 讗驻砖专讜转 砖讙讜讬讛, 讻讱 砖"讗祝 讗讞转" 砖讙讜讬讛 砖诐.
- 谞讻转讘 注诐 GPT. 讞诇拽 诪谞转讜谞讬 讛讗讬诪讜谉 谞讻转讘 注诐 GPT 砖诇 OpenAI, 讚专讱 诪谞讜讬 ChatGPT 砖诇谞讜: 376 谞讬住讜讞讬诐 讞诇讜驻讬讬诐 砖诇 讛砖讗诇讜转 讜-34 住讟讬诐 讞诇讜驻讬讬诐 砖诇 讗驻砖专讜讬讜转 转砖讜讘讛, 砖诪讜驻讬注讬诐 讘-35,189 驻专讬讟讬 讗讬诪讜谉 (讻讜诇诇 讻诇 575 讛驻专讬讟讬诐 砖诇 砖讗诇转 讛诇拽讜讞 讛诪砖诇诐), 讜-120 转专讞讬砖讬诐 拽爪专讬诐 砖诇驻讬讛诐 DeepSeek V4.1 Flash 讻转讘 6,000 讛讜讚注讜转. GPT 诇讗 讻转讘 讗祝 讛讜讚注讛 讜讗祝 转讜讜讬转. 讘住讱 讛讻讜诇 37,849 诪转讜讱 116,854 驻专讬讟讬 讛讗讬诪讜谉 (32%) 诪砖转诪砖讬诐 讘讟拽住讟 砖谞讻转讘 注诐 GPT. 讛驻诇讟 砖诇 GPT 讻驻讜祝 诇转谞讗讬 讛砖讬诪讜砖 砖诇 OpenAI, 诇讗 诇专讬砖讬讜谉 驻转讜讞.
- 谞转讜谞讬诐 驻转讜讞讬诐. HeQ v1.1 (CC BY 4.0. 讘注专讱 讞爪讬 诪讛砖讗诇讜转 砖诇讜 注诇 讻转讘讜转 砖诇 Geektime, 讜诪讞讘专讬 HeQ 诪砖转驻讬诐 讗讜转诐 讘讗讜转讜 专讬砖讬讜谉) 讜-MASSIVE he-IL (CC BY 4.0), 专拽 诪驻讬爪讜诇讬 讛讗讬诪讜谉. 拽讟注讬诐 讘注讘专讬转 诪-FineWeb-2 (ODC-By 1.0) 诇砖讗诇讜转 讛拽专讬讗讛 讜讛谞讜砖讗.
- 讘诇讬 讚诇讬驻讛 诪讛诪讘讞谞讬诐. 讻诇 驻专讬讟 讗讬诪讜谉 讛讜砖讜讜讛 诇讻诇 砖讗诇转 诪讘讞谉. 讻诇 诪讛 砖讞讜诇拽 专爪祝 砖诇 8 诪讬诇讬诐 注诐 讟拽住讟 诪讘讞谉 谞讝专拽, 讜讙诐 驻拽讜讚讜转 MASSIVE 砖讝讛讜转 诇驻拽讜讚转 诪讘讞谉. 驻专讬讟讬 讛拽专讬讗讛, 讛谞讜砖讗, 讛讛讜谞讗讜转 讛拽砖讜转, 讛诇拽讜讞 讛诪砖诇诐 讜讛谞讬住讜讞讬诐 讛讞诇讜驻讬讬诐 谞讘讚拽讜 讙诐 讘讻诇诇讬诐 诪讞诪讬专讬诐 讬讜转专: 讻诇 拽讟注 诪讜诇 讻诇 讟拽住讟 诪讘讞谉 (讻诇 专爪祝 诪砖讜转祝 砖诇 6 诪讬诇讬诐), 讜讻诇 砖讗诇讛, 讗驻砖专讜转 讜谞讬住讜讞 诪讜诇 讻诇 砖讗诇转 诪讘讞谉 讜讻诇 谞讬住讜讞 讞诇讜驻讬 砖诇 砖讗诇转 诪讘讞谉 (讛转讗诪讛 诪讚讜讬拽转 讜讚诪讬讜谉 转讜讜讬诐).
- 诪讛 诇讗 砖讬诪砖: 砖讜诐 驻诇讟 砖诇 爪'讗讟讘讜讟 诪住讞专讬 住讙讜专 讗讞专, 砖讜诐 诪讜讚诇 砖诇 DICTA, 砖讜诐 谞转讜谞讬诐 诇讗 诪住讞专讬讬诐 讗讜 讘专讬砖讬讜谉 "砖讬转讜祝 讝讛讛". GPT, 诪讜讚诇 诪住讞专讬 住讙讜专, 砖讬诪砖 专拽 讻诪转讜讗专 诇诪注诇讛.
- 讙讜讚诇: 讛诪拽讜讚讚 357.1 诪讬诇讬讜谉 驻专诪讟专讬诐 (HalleluBERT-large, MIT, 讗讜诪谉 诪讞讚砖). 专讗砖 讛讛讞诇讟讜转 26.5 诪讬诇讬讜谉 驻专诪讟专讬诐, 讗讜诪谉 诪讗驻住.
注诇 讛诪讘讞谞讬诐.
| 诪讘讞谉 | 砖讗诇讜转 | 诪拽讜专 讜专讬砖讬讜谉 |
|---|---|---|
| 讛讜讚注讜转 住驻讗诐 讜注住拽讬诐 | 596 (298 讛讜讚注讜转, 2 砖讗诇讜转 诇讻诇 讗讞转) | 驻谞讬诪讬, 讘住讙谞讜谉 SMS, 讜讜讗讟住讗驻 讜诪讬讬诇 讬砖专讗诇讬, 65 诪拽专讬诐 拽砖讬诐 |
| 驻谞讬讜转 转诪讬讻讛 (triage30) | 90 (30 驻谞讬讜转, 3 砖讗诇讜转 诇讻诇 讗讞转) | 驻谞讬诪讬 |
| MASSIVE he-IL, 20 讜-4 讗驻砖专讜讬讜转 | 500 + 500 | 驻讬爪讜诇 讛诪讘讞谉 砖诇 MASSIVE, CC BY 4.0 |
| SIB-200, 谞讜砖讗 讬讚讬注讛 | 204 | CC BY-SA 4.0, 专拽 诇讘讚讬拽讛 |
| Belebele, 讛讘谞转 讛谞拽专讗 | 900 | CC BY-SA 4.0, 专拽 诇讘讚讬拽讛 |
| HeQ, 讗讬诪讜转 讜"讗讬谉 转砖讜讘讛" | 600 + 600 | 驻讬爪讜诇 讛诪讘讞谉 砖诇 HeQ v1.1, CC BY 4.0 |
| "讗祝 讗讞转 诪讛讗驻砖专讜讬讜转" (诪讚讜讜讞 讘谞驻专讚, 专讗讜 诪讙讘诇讜转) | 400 (200 砖"讗祝 讗讞转" 谞讻讜谞讛 讘讛谉, 200 砖讛讬讗 砖讙讜讬讛 讘讛谉) | 砖讗诇讜转 诪讘讞谉 砖诇 MASSIVE (CC BY 4.0) 讜-Belebele (CC BY-SA 4.0) 注诐 讗驻砖专讜转 "讗祝 讗讞转 诪讛讗驻砖专讜讬讜转" 砖谞讜住驻讛, 专拽 诇讘讚讬拽讛 |
讛住转讬讬讙讜讬讜转 砖诪砖谞讜转 讻诪讛 诇住诪讜讱 注诇 讛诪住驻专讬诐:
- 诪讘讞谉 讛住驻讗诐 讜讛注住拽讬诐 谞讻转讘 讘讬讚讬 诪讜讚诇 AI 讜谞讘讚拽 讘讬讚讬 诪讜讚诇 AI, 诇讗 讘讬讚讬 讗讚诐. 转讬讘讜转 讚讜讗专 讗诪讬转讬讜转 讬讬专讗讜 讗讞专转. 讛讜讗 谞谞注诇 讗讞专讬 讛讘讚讬拽讛 讛讝讜 (4 转讜讜讬讜转 砖讜谞讜, 2 讛讜讚注讜转 讛讜住专讜).
- 砖诇讜砖 专讬爪讜转 讗讬诪讜谉, 注诐 砖诇讜砖讛 讝专注讬诐 讗拽专讗讬讬诐, 讜讝讜 讗讞转 诪讛谉. 注诇 驻专讬讟讬 讛讗讬诪讜谉 砖讛讜驻专讚讜 诪专讗砖 砖诇讜砖 讛专讬爪讜转 讻诪注讟 砖讜讜转 (95.0% 讻讗谉 诪讜诇 95.2% 讜-95.2%). 讛专讬爪讛 讛讝讜 谞讘讞专讛 讻讬 讛讬讗 讛讬讬转讛 讛讟讜讘讛 讘讬讜转专 注诇 188 讛讛讜讚注讜转 讛诪讜讞讝拽讜转 诪谞转讜谞讬 讛讛讜谞讗讛 砖谞讜住驻讜 (98.4% 诪讜诇 97.3% 讜-97.3%) 讜注诇 讛讛讜讚注讜转 讛讗诪讬转讬讜转 砖讘讚拽谞讜, 诇讗 诇驻讬 讛诪讘讞谞讬诐 讘讻专讟讬住 讛讝讛. 讘诪讘讞谞讬诐 讛讙讚讜诇讬诐 砖诇讜砖 讛专讬爪讜转 拽专讜讘讜转: 讛讜谞讗讛 讗讜 诇讗 92.0% 讻讗谉 诪讜诇 92.0% 讜-90.6%, 谞讜砖讗 讬讚讬注讛 79.4% 诪讜诇 82.8% 讜-80.4%, 讛讘谞转 讛谞拽专讗 63.2% 诪讜诇 60.9% 讜-63.1%. 讛专讬爪讛 讛讝讜 讜讻诇 讗讞转 诪讛砖转讬讬诐 讛讗讞专讜转 谞讜转谞讜转 讗讜转讛 转砖讜讘讛 注诇 90.1% 注讚 90.6% 诪讻诇 砖讗诇讜转 讛诪讘讞谉. 讘诪讘讞谞讬诐 砖诇 30 驻谞讬讜转 讛讛讘讚诇讬诐 讙讚讜诇讬诐 讬讜转专: 讚讞讬驻讜转 讛驻谞讬讬讛 70.0% 讻讗谉 诪讜诇 73.3% 讜-76.7%.
- 转讜讜讬讜转 讛讗讬诪讜谉 讘讗讜转 诪砖谞讬 诪讜讚诇讬 AI. 讗讬驻讛 砖砖谞讬讛诐 讟讜注讬诐 讘讗讜转讜 讗讜驻谉, 谞讬爪讜抓 诇诪讚 讗转 讛讟注讜转 砖诇讛诐.
- 讛转砖讜讘讜转 讛砖讙讜讬讜转 砖诇 HeQ 讘诪讘讞谉 谞讘讞专讜 讘拽讜讚, 讜诇讗 谞讘讚拽讜 讘讬讚讬 讗讚诐.
诪讙讘诇讜转
- 讛讘谞转 讛谞拽专讗 砖诇 拽讟注讬诐 讗专讜讻讬诐 诪讜讙讘诇转. Belebele: 63.2%, 讻砖谞讬讞讜砖 注讬讜讜专 诪拽讘诇 25%, 讜注讚讬讬谉 驻讞讜转 诪诪讜讚诇 laya 讛讗讞专 讛讟讜讘 讘讬讜转专 讘诪讘讞谉 讛讝讛 (专讗讜 诪讘讞谞讬诐). 讻砖讛转砖讜讘讛 讛谞讻讜谞讛 讻转讜讘讛 讘拽讟注 诪讬诇讛 讘诪讬诇讛 讛讜讗 诪拽讘诇 77% (252 砖讗诇讜转). 讻砖讛转砖讜讘讛 谞讗诪专转 讘诪讬诇讬诐 讗讞专讜转 讛讜讗 诪拽讘诇 58% (648 砖讗诇讜转). 讘砖讗诇讜转 "讗讬讝讜 诪讛讘讗讜转 诇讗" 59% (158 砖讗诇讜转). 讗诇 转砖讗诇讜 讗讜转讜 讗诐 诪住诪讱 讗专讜讱 转讜诪讱 讘讟注谞讛.
- 讘讚讬拽讛 讗诐 拽讟注 转讜诪讱 讘转砖讜讘讛 (HeQ) 注讜讘讚转. 94.8% 注诐 讛拽讟注. 讻砖诪讜讞拽讬诐 讗转 讛拽讟注 讝讛 讬讜专讚 诇-51.0%, 讛讟诇转 诪讟讘注. 讻诇讜诪专 讘砖讗诇讜转 诪讛住讜讙 讛讝讛 讛讜讗 讘讗诪转 拽讜专讗 讗转 讛拽讟注.
- 诪住驻专讬诐, 转讗专讬讻讬诐, 住讻讜诪讬诐 讜讻诇诇讬诐: 诇讗 讗讜诪谉 注诇讬讛诐 讜诇讗 谞诪讚讚. 讞砖讘讜 讗讜转诐 讘拽讜讚 讜讛注讘讬专讜 讗转 讛转讜爪讗讛.
- 讛诪拽专讬诐 讛拽砖讬诐 注讚讬讬谉 谞拽讜讚转 讛转讜专驻讛: 讛讜谞讗讜转 砖谞讻转讘讜 讻讚讬 诇讛讬专讗讜转 诇讙讬讟讬诪讬讜转 ("住驻拽" 砖诪讞诇讬祝 驻专讟讬 讞砖讘讜谉 讘谞拽, "讛诪谞讻"诇" 砖诪讘拽砖 讛注讘专讛), 讜讛讜讚注讜转 讗诪讬转讬讜转 砖谞专讗讜转 讻诪讜 讛讜谞讗讛 (讛转专讗讛 讗诪讬转讬转 诪讛讘谞拽 注诐 拽讬砖讜专, 拽讜讚 讗讬诪讜转 讗诪讬转讬). 注诇 65 讛诪拽专讬诐 讛拽砖讬诐 砖讗诇转 讛讛讜谞讗讛 爪讜讚拽转 讘-83.1% 诪讛诪拽专讬诐 (转砖讜讘讛 拽讘讜注讛 "诇讗 讛讜谞讗讛" 讛讬讬转讛 诪拽讘诇转 砖诐 70.8%), 注诐 AUC 0.84. 讘住祝 砖谞讘讞专 讛讜讗 注讚讬讬谉 诪驻住驻住 5 诪转讜讱 19 讛讜谞讗讜转 砖谞讻转讘讜 讻讚讬 诇讛讬专讗讜转 诇讙讬讟讬诪讬讜转, 讜诪转专讬注 注诇 6 诪转讜讱 46 讛讜讚注讜转 讗诪讬转讬讜转 砖谞专讗讜转 讻诪讜 讛讜谞讗讛.
- 讚讞讬驻讜转 讛讬讗 注谞讬讬谉 住讜讘讬讬拽讟讬讘讬, 讜讛转砖讜讘讛 注诇讬讛 转诇讜讬讛 讘谞讬住讜讞. 讗驻讬诇讜 砖谞讬 诪讜讚诇讬 讛诪讜专讛 讛住讻讬诪讜 注诐 讛讚讞讬驻讜转 讛诪转讜讻谞谞转 专拽 讘讻-62 注讚 64% 诪讛诪拽专讬诐. 讻砖砖讜讗诇讬诐 讗转 砖讗诇转 讛讚讞讬驻讜转 讘诪讬诇讬诐 讗讞专讜转, 讛转砖讜讘讛 诪砖转谞讛 讘-63% 诪转讜讱 30 驻谞讬讜转 讛诪讘讞谉.
- 讛诪讘讞谞讬诐 讛驻谞讬诪讬讬诐 讜讞诇拽 讙讚讜诇 诪谞转讜谞讬 讛讗讬诪讜谉 谞讻转讘讜 讘讬讚讬 诪讜讚诇讬 AI, 诇讗 讘讬讚讬 讗谞砖讬诐. 讝讛 讻讜诇诇 讗转 诪讘讞谉 讛住驻讗诐 讜讛注住拽讬诐, 讗转 诪讘讞谉 驻谞讬讜转 讛转诪讬讻讛 讜讗转 砖讗诇讜转 讛拽专讬讗讛 注诇 FineWeb-2. 讛讜讚注讜转 讗诪讬转讬讜转 讬讬专讗讜 讗讞专转.
- 诪讘讞谞讬 驻谞讬讜转 讛转诪讬讻讛 拽讟谞讬诐: 30 驻谞讬讜转 诇讻诇 砖讗诇讛, 讻讱 砖驻谞讬讬讛 讗讞转 诪讝讬讝讛 爪讬讜谉 讘-3.3 谞拽讜讚讜转.
- 爪讬谞讬讜转 讜讗讬专讜谞讬讛 讻谞专讗讛 谞拽专讗讜转 诪讬诇讜诇讬转. 诇讗 谞诪讚讚.
- 512 讟讜拽谞讬诐 (讘注专讱 300 注讚 400 诪讬诇讬诐 讘注讘专讬转) 诇砖讗诇讛. 讛讜讚注讛 讗专讜讻讛 讬讜转专 谞讞转讻转 诪讛住讜祝 讘诇讬 讗讝讛专讛.
- 注讘专讬转 讘诇讘讚. 诇讗 讗讜诪谉 讜诇讗 谞讘讚拽 注诇 讗谞讙诇讬转 讗讜 注专讘讬转.
- 讝讜 诇讗 诪注专讻转 讛讙谞讛 诇讘讚. 讛讜讗 讟讜注讛 诇砖谞讬 讛讻讬讜讜谞讬诐. 讛砖讗讬专讜 讗讚诐 讘转讛诇讬讱 讘讻诇 讚讘专 砖讬讻讜诇 诇驻讙讜注 讘诪讬砖讛讜.
- 讻砖诪讜住讬驻讬诐 讗驻砖专讜转 "讗祝 讗讞转 诪讛讗驻砖专讜讬讜转", 讛讜讗 讘讜讞专 讘讛 讬讜转专 诪讚讬. 谞诪讚讚 注诇 诪讘讞谉 拽驻讜讗 谞驻专讚 砖诇 400 砖讗诇讜转, 砖讗诇讜转 诪讘讞谉 砖诇 MASSIVE 讜-Belebele 砖谞讘谞讜 诪讞讚砖 注诐 讗驻砖专讜转 "讗祝 讗讞转 诪讛讗驻砖专讜讬讜转". 讻砖"讗祝 讗讞转" 讛讬讗 讛转砖讜讘讛 讛谞讻讜谞讛, 讛讜讗 讘讜讞专 讘讛 讘-76% 诪讛诪拽专讬诐. 讻砖讛转砖讜讘讛 讛谞讻讜谞讛 谞诪爪讗转 讘专砖讬诪讛, 讛讜讗 注讚讬讬谉 讘讜讞专 "讗祝 讗讞转" 讘-31% 诪讛砖讗诇讜转, 讜爪讜讚拽 专拽 讘-61% 诪讛谉 (77% 讘驻拽讜讚讜转 拽讜诇讬讜转, 45% 讘砖讗诇讜转 拽专讬讗讛). 讻砖诪讜讞拽讬诐 讗转 讛讟拽住讟 讛讜讗 讘讜讞专 "讗祝 讗讞转" 讻诪注讟 转诪讬讚 (98%). 讗诐 讗转诐 诪爪讬注讬诐 讗驻砖专讜转 讻讝讜, 讘讚拽讜 讗讜转讛 拽讜讚诐 注诇 讛砖讗诇讜转 砖诇讻诐.
- 讛讜讚注讜转 讗诪讬转讬讜转: 注讜讚 诇讗 谞诪讚讚 注诇 住讟 讘诇转讬 转诇讜讬. 谞讜住驻讜 谞转讜谞讬 讗讬诪讜谉 诇讗讝讛专讜转 讗诪讬转讬讜转 诪驻谞讬 讛讜谞讗讛 讜诇讛讜谞讗讜转 拽爪专讜转 讘诇讬 拽讬砖讜专, 砖转讬 讛谞拽讜讚讜转 讛讞诇砖讜转 砖讛讜讚注讜转 讗诪讬转讬讜转 讛专讗讜. 诪讚讬讚讛 注诇 住讟 讘诇转讬 转诇讜讬 砖诇 讛讜讚注讜转 讗诪讬转讬讜转 注讜讚 诇讗 谞注砖转讛, 讜诇讻谉 讘讻专讟讬住 讛讝讛 讗讬谉 诪住驻专 注诇 讛讜讚注讜转 讗诪讬转讬讜转.
专讬砖讬讜谉 讜拽专讚讬讟讬诐
Apache-2.0 诇诪砖拽讜诇讜转, 诇拽讜讘爪讬 讛-GGUF 讜诇拽讜讚. 诪讜转专 诇砖讬诪讜砖 诪住讞专讬. 讘谞讜讬 注诇:
- HalleluBERT-large: 讛诪拽讜讚讚, MIT.
- laya (NandhaKishorM, Convai Innovations): 讛讗专讻讬讟拽讟讜专讛 砖诇 专讗砖 讛讛讞诇讟讜转 讜住讘讬讘转 讛讛专爪讛, Apache-2.0. 诇讗 谞注砖讛 砖讬诪讜砖 讘诪砖拽讜诇讜转 砖诇 laya.
- HeQ: CC BY 4.0, 砖诇 Webiks 注讘讜专 诪驻讗"转 讜转讜讻谞讬转 讛-NLP 讛诇讗讜诪讬转 (NNLP-IL). 讻讜诇诇 拽讟注讬诐 诪-Geektime.
- MASSIVE: CC BY 4.0, 讗诪讝讜谉 (FitzGerald 讜讗讞专讬诐, 2022).
- DeepSeek V4.1 Flash (MIT) 讜-Gemma 4 31B (Apache-2.0), 讚专讱 DeepInfra, 讻讻讜转讘 讛谞转讜谞讬诐 讜讻诪住诪谞讬诐.
- FineWeb-2 (注讘专讬转): ODC-By 1.0, 砖诇 Hugging Face. 讛拽讟注讬诐 砖诇 砖讗诇讜转 讛拽专讬讗讛 讜讛谞讜砖讗.
- GPT 砖诇 OpenAI, 讚专讱 诪谞讜讬 ChatGPT (转谞讗讬 讛砖讬诪讜砖 砖诇 OpenAI): 谞讬住讜讞讬 砖讗诇讜转 讜转专讞讬砖讬诐, 讻诪转讜讗专 讘驻专拽 谞转讜谞讬 讛讗讬诪讜谉.
- ggmlc 诇拽讜讘爪讬 讛-GGUF.
- Belebele 讜-SIB-200 (CC BY-SA 4.0) 砖讬诪砖讜 专拽 诇讘讚讬拽讛, 讗祝 驻注诐 诇讗 诇讗讬诪讜谉.
讛讛讜讚注讛 讛诪诇讗讛 讘拽讜讘抓 NOTICE.











