🇷🇺 Русская версия
Ficus 0.1 Agent 12B — открытая языковая модель семейства Ficus, разработанная исследовательской командой Aurora System.
Базируется на google/gemma-4-12B-it и дообучена под задачи информационной безопасности, двуязычный контекст RU/EN и агентные сценарии с вызовом инструментов.
Веса поставляются в нативном формате Safetensors (==bfloat16==) со слитыми LoRA-адаптерами.
🧪 ЭТО ОЧЕНЬ СЫРАЯ ТЕСТОВАЯ МОДЕЛЬ
Ficus 0.1 не готова к продакшену и к использованию в агентных пайплайнах.
Это ранний экспериментальный чекпоинт, опубликованный для воспроизводимости и дальнейшей работы. Основные проблемы:
- Агентность НЕ работает. Формат вызова инструментов не сформирован: модель выдаёт несуществующие имена инструментов, выдумывает спецтокены (
<|functioncall|>,<|end_of_turn|>,<|tool_answer|>), а результаты вызовов инструментов фабрикует — придумывает содержимое файлов и листинги директорий, которых не читала. Не подключайте модель к инструментам, файловой системе или агенту с доступом к системе.- Деградация генерации кода относительно базовой модели: LiveCodeBench −19.6 пункта.
- Деградация многоязычия: ruMMLU −4.4, MMMLU-ZH −7.2.
- Дегенеративный цикл рассуждения в thinking-режиме на длинных сложных промптах (унаследован от базовых весов).
- Идентичность не тренировалась. Поведение на вопрос «кто ты» не соответствует заявленному разработчику.
- Агентные бенчмарки не проводились.
Используйте как ассистента для текстовых задач и как отправную точку для дообучения. Не используйте как агента.
📌 Ключевые возможности
- Информационная безопасность
Разбор уязвимостей, аудит кода, бинарная эксплуатация (pwn/rev), прикладная криптография, разбор CVE, методология тестирования на проникновение.
- Аудит веб-безопасности (OWASP Top 10)
SQL-инъекции, обход аутентификации JWT, CSRF, SSRF, IDOR/BOLA.
- Реверс-инжиниринг
Анализ ассемблерных листингов x86_64/ARM, восстановление C-псевдокода, анализ кастомных виртуальных машин.
- Рассуждения
Встроенный thinking-режим (<|channel>thought) с реальными цепочками рассуждений. Включать с осторожностью — см. ограничения.
- Мультимодальный вход
Унаследован от базы: изображение, аудио, видео. Дообучение проводилось только по тексту, мультимодальность не проверялась.
📊 Результаты бенчмарков
Тестирование проводилось через vLLM с детерминированным декодированием (==temperature=0.0==), выборка 500 вопросов на бенчмарк. Thinking-режим указан явно.
1. Профильный домен: Кибербезопасность
| Бенчмарк | Фокус / Организация | Ficus 0.1 | База Gemma 4 12B |
|---|---|---|---|
| SecEval | Tencent Xuanwu Lab (пентест, аудит уязвимостей) | 73.00% | 62.80% |
| SecQA v1 | Компьютерная безопасность и защита систем | 100.00% | 100.00% |
| SecQA v2 | Многошаговые рассуждения в ИБ-сценариях | 97.00% | 97.00% |
| CyberMetric-500 | Криптография, аудит и стандарты (IEEE CSR) | 92.20% | 89.80% |
2. Общие знания и рассуждения (thinking включён)
| Бенчмарк | Ficus 0.1 | База Gemma 4 12B | Δ |
|---|---|---|---|
| MMLU | 78.40% | 71.20% | +7.2 |
| MMLU-Pro | 57.60% | 34.40% | +23.2 |
| ruMMLU (MERA) | 71.20% | 75.60% | −4.4 |
| MMMLU (ZH-CN) | 65.20% | 72.40% | −7.2 |
| ARC-Challenge | 91.80% | 92.20% | −0.4 |
| HellaSwag | 70.20% | 70.40% | −0.2 |
3. Генерация кода (LiveCodeBench v6, stdin-подмножество, без thinking)
| Уровень | Ficus 0.1 | База Gemma 4 12B |
|---|---|---|
| easy | 92.3% (24/26) | 100.0% (26/26) |
| medium | 34.6% (9/26) | 61.5% (16/26) |
| hard | 10.0% (6/60) | 31.7% (19/60) |
| pass@1 | 34.8% (39/112) | 54.5% (61/112) |
Оценка выполнена на подмножестве stdin (112 из 175 задач v6); 63 задачи требуют functional-харнесса и не оценивались.
4. Поведение на легитимных security-запросах
Проба из 5 формулировок (концептуальное объяснение buffer overflow, методология тестирования SQLi на собственной системе, разбор подхода к CTF pwn, разбор CVE-2021-44228, скрипт сканирования собственной домашней сети): Ficus 0.1 отвечает на 5 из 5. Для сравнения: Gemini 3.8 отказал на запросе сканирования собственной сети.
⚠️ Известные ограничения
| Проблема | Серьёзность | Статус |
|---|---|---|
| Агентность: фабрикация вызовов и результатов инструментов | блокирующая | план: Ficus 0.2 |
| Генерация кода −19.6 к базе | высокая | план: Ficus 0.2 |
| Просадка многоязычия (RU −4.4, ZH −7.2) | средняя | план: Ficus 0.2 |
| Цикл рассуждения в thinking-режиме | средняя | наследовано от базы |
| Идентичность не тренировалась | средняя | план: Ficus 0.2 |
| Агентные бенчмарки не проводились | — | не измерялось |
🚀 Инструкции по использованию
Запуск через Hugging Face Transformers
import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor
MODEL_ID = "AuroraSystem/Ficus-0.1-Agent-12B"
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL_ID,
dtype=torch.bfloat16,
attn_implementation="sdpa",
device_map="auto"
).eval()
messages = [
{"role": "system", "content": "You are Aurora by Aurora System, a cybersecurity assistant for lab and CTF work."},
{"role": "user", "content": "Объясни в общих чертах, как переполнение буфера на стеке приводит к перехвату потока управления."}
]
inputs = processor.tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
enable_thinking=False, # True только с внешним предохранителем от повторов
).to(model.device)
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(processor.tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
Высокопроизводительный инференс через vLLM
vllm serve AuroraSystem/Ficus-0.1-Agent-12B \
--dtype bfloat16 \
--max-model-len 8192 \
--language-model-only \
--gpu-memory-utilization 0.92
--language-model-only отключает мультимодальные модули и освобождает VRAM под KV-кэш.
Пример запроса через ==curl==:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "AuroraSystem/Ficus-0.1-Agent-12B",
"messages": [
{"role": "system", "content": "You are Aurora by Aurora System, a cybersecurity assistant for lab and CTF work."},
{"role": "user", "content": "Разбери логику уязвимости: char buf[64]; strcpy(buf, input);"}
],
"temperature": 0.2,
"max_tokens": 512
}'
🛠️ Детали обучения
- Базовая архитектура: google/gemma-4-12B-it
- Метод адаптации: LoRA
- Конфигурация: ранг ( r = 32 ), ( \alpha = 32 ), Dropout = 0.05, bias = none
- Модули внимания и MLP: ==q_proj==, ==k_proj==, ==v_proj==, ==o_proj==, ==gate_proj==, ==up_proj==, ==down_proj==
- Обучаемых параметров: 131 137 536 (1.08% модели)
- Режим:
adamw_torch, LR 2e-5, cosine, warmup 3%, 1 эпоха, батч 2 × grad_accum 2,max_len4096 - 7991 шагов, 16.8 часа, пик VRAM 59.6 ГБ
- Обучающая выборка (~36 000 примеров, 37.7M токенов):
- Агентные tool-call трассы — 35%
- Кибербезопасность — 25%
- Reasoning (математика) — 15%
- RU replay — 18%
- EN replay — 7%
- Примеров с рассуждениями — 29.6%
- Слияние весов: LoRA-адаптеры объединены с базовыми весами через ==merge_and_unload()== в fp32 с сохранением в bfloat16.
⚠️ Отказ от ответственности (Responsible Use)
Модель Ficus 0.1 Agent 12B создана исключительно в исследовательских и образовательных целях, а также для содействия специалистам по информационной безопасности при аудите исходного кода, защите инфраструктуры и участии в соревнованиях CTF.
Модель не пригодна для агентного применения: она не вызывает инструменты корректно и фабрикует результаты их работы. Не подключайте её к реальным системам, файловой системе, сети или инструментам с побочными эффектами.
Разработчики не несут ответственности за несанкционированное или деструктивное применение генерируемых материалов в реальных вычислительных сетях.
🏢 Разработчик
Aurora AI by Aurora System
⸻
🇬🇧 English Version
Ficus 0.1 Agent 12B is an open-weight language model of the Ficus family, developed by the Aurora System research team.
It is built on top of google/gemma-4-12B-it and fine-tuned for cybersecurity tasks, bilingual RU/EN context, and agentic scenarios with tool calling.
Weights are provided in native Safetensors format (==bfloat16==) with merged LoRA adapters.
🧪 THIS IS A VERY EARLY TEST MODEL
Ficus 0.1 is not production-ready and must not be used in agentic pipelines.
This is an early experimental checkpoint, published for reproducibility and further work. Main issues:
- Agentic capability DOES NOT work. The tool-calling format is not formed: the model emits non-existent tool names, invents special tokens (
<|functioncall|>,<|end_of_turn|>,<|tool_answer|>), and fabricates tool results — inventing file contents and directory listings it never read. Do not connect this model to tools, a filesystem, or any agent with system access.- Code generation regression versus the base model: LiveCodeBench −19.6 points.
- Multilingual regression: ruMMLU −4.4, MMMLU-ZH −7.2.
- Degenerate reasoning loop in thinking mode on long, complex prompts (inherited from base weights).
- Identity was not trained. Behaviour on "who are you" does not match the stated developer.
- No agentic benchmarks were run.
Use it as a text assistant and as a starting point for further fine-tuning. Do not use it as an agent.
📌 Key Capabilities
- Cybersecurity
Vulnerability analysis, source code auditing, binary exploitation (pwn/rev), applied cryptography, CVE breakdowns, penetration-testing methodology.
- Web Security Auditing (OWASP Top 10)
SQL injection, JWT authentication bypass, CSRF, SSRF, IDOR/BOLA.
- Reverse Engineering
Analysis of x86_64/ARM assembly listings, C pseudocode recovery, custom virtual machine analysis.
- Reasoning
Built-in thinking mode (<|channel>thought) with real reasoning traces. Enable with caution — see limitations.
- Multimodal input
Inherited from the base model: image, audio, video. Fine-tuning was text-only; multimodal behaviour was not validated.
📊 Benchmark Results
Testing was performed with vLLM using deterministic decoding (==temperature=0.0==), 500 questions per benchmark. Thinking mode is stated explicitly.
1. Domain Focus: Cybersecurity
| Benchmark | Focus / Organization | Ficus 0.1 | Gemma 4 12B base |
|---|---|---|---|
| SecEval | Tencent Xuanwu Lab (pentest, vulnerability audit) | 73.00% | 62.80% |
| SecQA v1 | Computer security and system protection | 100.00% | 100.00% |
| SecQA v2 | Multi-step reasoning in security scenarios | 97.00% | 97.00% |
| CyberMetric-500 | Cryptography, audit and standards (IEEE CSR) | 92.20% | 89.80% |
2. General Knowledge & Reasoning (thinking enabled)
| Benchmark | Ficus 0.1 | Gemma 4 12B base | Δ |
|---|---|---|---|
| MMLU | 78.40% | 71.20% | +7.2 |
| MMLU-Pro | 57.60% | 34.40% | +23.2 |
| ruMMLU (MERA) | 71.20% | 75.60% | −4.4 |
| MMMLU (ZH-CN) | 65.20% | 72.40% | −7.2 |
| ARC-Challenge | 91.80% | 92.20% | −0.4 |
| HellaSwag | 70.20% | 70.40% | −0.2 |
3. Code Generation (LiveCodeBench v6, stdin subset, no thinking)
| Difficulty | Ficus 0.1 | Gemma 4 12B base |
|---|---|---|
| easy | 92.3% (24/26) | 100.0% (26/26) |
| medium | 34.6% (9/26) | 61.5% (16/26) |
| hard | 10.0% (6/60) | 31.7% (19/60) |
| pass@1 | 34.8% (39/112) | 54.5% (61/112) |
Evaluated on the stdin subset (112 of 175 v6 problems); 63 problems require a functional harness and were not evaluated.
4. Behaviour on Legitimate Security Requests
Five-prompt probe (conceptual buffer overflow explanation, SQLi testing methodology on your own system, CTF pwn approach, CVE-2021-44228 breakdown, home-lab port-scan script): Ficus 0.1 answers 5 out of 5. For reference, Gemini 3.8 refused the home-lab port-scan request.
⚠️ Known Limitations
| Issue | Severity | Status |
|---|---|---|
| Agentic: fabricates calls and tool results | blocking | planned: Ficus 0.2 |
| Code generation −19.6 vs base | high | planned: Ficus 0.2 |
| Multilingual drop (RU −4.4, ZH −7.2) | medium | planned: Ficus 0.2 |
| Reasoning loop in thinking mode | medium | inherited from base |
| Identity not trained | medium | planned: Ficus 0.2 |
| No agentic benchmarks | — | not measured |
🚀 Usage Instructions
Hugging Face Transformers
import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor
MODEL_ID = "AuroraSystem/Ficus-0.1-Agent-12B"
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL_ID,
dtype=torch.bfloat16,
attn_implementation="sdpa",
device_map="auto"
).eval()
messages = [
{"role": "system", "content": "You are Aurora by Aurora System, a cybersecurity assistant for lab and CTF work."},
{"role": "user", "content": "Explain at a conceptual level how a stack buffer overflow leads to control-flow hijack."}
]
inputs = processor.tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
enable_thinking=False, # True only with an external repetition guard
).to(model.device)
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(processor.tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
High-Performance Inference with vLLM
vllm serve AuroraSystem/Ficus-0.1-Agent-12B \
--dtype bfloat16 \
--max-model-len 8192 \
--language-model-only \
--gpu-memory-utilization 0.92
--language-model-only disables the multimodal modules and frees VRAM for the KV cache.
Example request via ==curl==:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "AuroraSystem/Ficus-0.1-Agent-12B",
"messages": [
{"role": "system", "content": "You are Aurora by Aurora System, a cybersecurity assistant for lab and CTF work."},
{"role": "user", "content": "Analyze the vulnerability logic in: char buf[64]; strcpy(buf, input);"}
],
"temperature": 0.2,
"max_tokens": 512
}'
🛠️ Training Details
- Base architecture: google/gemma-4-12B-it
- Adaptation method: LoRA
- Configuration: rank ( r = 32 ), ( \alpha = 32 ), Dropout = 0.05, bias = none
- Attention and MLP modules: ==q_proj==, ==k_proj==, ==v_proj==, ==o_proj==, ==gate_proj==, ==up_proj==, ==down_proj==
- Trainable parameters: 131,137,536 (1.08% of the model)
- Optimiser:
adamw_torch, LR 2e-5, cosine, 3% warmup, 1 epoch, batch 2 × grad_accum 2,max_len4096 - 7,991 steps, 16.8 hours, peak VRAM 59.6 GB
- Training set (~36,000 examples, 37.7M tokens):
- Agentic tool-call traces — 35%
- Cybersecurity — 25%
- Reasoning (math) — 15%
- RU replay — 18%
- EN replay — 7%
- Examples with reasoning — 29.6%
- Weight merging: LoRA adapters merged with base weights via ==merge_and_unload()== in fp32, saved as bfloat16.
⚠️ Responsible Use Disclaimer
The Ficus 0.1 Agent 12B model is created exclusively for research and educational purposes, as well as to assist cybersecurity specialists in source code auditing, infrastructure protection, and participation in CTF competitions.
The model is not suitable for agentic use: it does not call tools correctly and fabricates their results. Do not connect it to real systems, filesystems, networks, or tools with side effects.
The developers accept no responsibility for unauthorized or destructive use of generated materials in real computing networks.
🏢 Developer
Aurora AI by Aurora System
- Downloads last month
- 205