Instructions to use RavinduSen/JaneGPT-v2-Janus with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use RavinduSen/JaneGPT-v2-Janus with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="RavinduSen/JaneGPT-v2-Janus")# pip install -U transformers accelerate # Load model directly from transformers import JaneGPTJanusNLU model = JaneGPTJanusNLU.from_pretrained("RavinduSen/JaneGPT-v2-Janus", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- JaneGPT v2 Janus - Intent Classification Model
JaneGPT v2 Janus - Intent Classification Model
Hierarchical command understanding with state-aware runtime behavior for practical assistant workflows.
Update (Oct 2026): improved release. Higher accuracy and several bug fixes. The tokenizer now normalises casing, so capitalised speech-to-text output is handled like lowercase text; the model was retrained on cleaner, broader data; and the runtime handles wake phrases, cancelled questions and "close it"-style follow-ups. Use the new
weights/tokenizer.jsontogether with the new weights.
🏛️ The Temple of Janus (Web Experience)
We have deployed a dedicated interactive environment to showcase the essence of JaneGPT-v2 Janus.
Note: This is a visual and technical walkthrough; it does not feature a live chat interface.
- 🔗 Enter the Experience
- Best Viewed On: Desktop (Chrome/Edge) for full hardware-accelerated 3D effects.
Quickstart (2 minutes)
Install + first prediction
pip install -r requirements.txt
from janegpt_v2_janus import JaneGPTJanusNLU
nlu = JaneGPTJanusNLU(
model_path="weights/janegpt_v2_janus.pt",
tokenizer_path="weights/tokenizer.json",
)
state = {}
result = nlu.predict("set volume", state=state)
print(result)
if result.get("type") == "command":
state = nlu.update_state(result, state)
Runtime wrapper (recommended for assistant flows)
from runtime.jane_nlu_runtime import JaneNLURuntime
rt = JaneNLURuntime(base_dir=".")
state = {}
out, state = rt.handle_turn("set volume", state)
print(out) # expected: clarify prompt for missing VALUE
out, state = rt.handle_turn("55", state)
print(out) # expected: resolved local command
Run bundled demos
python examples/demo_inference.py
python examples/demo_runtime.py
python examples/demo_runtime_suite.py
What You Get
- Single-pass multitask prediction: domain + action + BIO slots.
- Runtime-safe clarification loops for missing required slots.
- Stateful follow-ups (for example, "that is not enough" after a volume change).
- Local command routing with controlled chat fallback.
- Compact deployment footprint: ~30.61 MB checkpoint.
Model Architecture
Interactive Architecture Visualization
Arch: Linear(256) → GELU → Dropout → Linear(10)
Output: 10 classes
Arch: Linear(256) → GELU → Dropout → Linear(33)
Output: 33 classes
Arch: Linear(256) → Linear(15 BIO)
Output: 15 labels/token
Training Objective
Architecture Specifications
| Component | Configuration | Details |
|---|---|---|
| Backbone Type | Transformer (GPT-style) | Bidirectional, non-causal attention |
| Vocabulary Size | 8,192 | BPE tokenization (NFKC + lowercase normalizer) |
| Embedding Dim | 256 | Token + Rotary Position embeddings |
| Attention Heads | 8 Query, 4 KV | Grouped Query Attention (GQA) for efficiency |
| Head Dimension | 32 | per head_dim = embed_dim / num_heads |
| Transformer Blocks | 8 Layers | Each with Attn + FFN + Residuals |
| Feed-Forward Hidden | 672 | SwiGLU gate activation |
| Position Encoding | RoPE | Rotary Position Embeddings (theta=10000) |
| Normalization | RMSNorm | Pre-layer normalization |
| Max Sequence Length | 96 tokens | Approximately 60-80 words |
| Dropout Rate | 0.1 | Applied during training |
| Total Parameters | 7,949,626 | All trainable |
| Parameter Breakdown | Backbone: 7.80M, Task Heads: 146K | Efficient multitask design |
Task Configuration
| Task | Type | Classes | Architecture |
|---|---|---|---|
| Domain Classification | Sequence-level | 10 domains | Pooled → Linear(256) → GELU → Linear(10) |
| Action Classification | Sequence-level | 33 actions | Pooled → Linear(256) → GELU → Linear(33) |
| Slot Tagging | Token-level | 15 BIO labels | Per-token → Linear(256) → Linear(15) |
Benchmark Results
Comprehensive Benchmark Summary
| Metric | Detail | Jane v2 | Janus |
|---|---|---|---|
| Speed (mean latency) | CUDA, batch=1 | 31.60 ms | 15.37 ms |
| Throughput | CUDA, single GPU | 32 pred/sec | 65 pred/sec, 0 errors across 82 turns |
| OOD F1 | BANKING77 | 94.31% | 97.30% |
| OOD F1 | CLINC OOS | 89.16% | 94.18% |
| OOD Precision | BANKING77 | 99.35% | 100.00% |
| OOD Precision | CLINC OOS | 99.14% | 100.00% |
| OOD Recall | BANKING77 | 89.75% | 94.75% |
| OOD Recall | CLINC OOS | 81.00% | 89.00% |
| Validation Accuracy | Domain (best epoch) | — | 98.87% |
| Validation Accuracy | Action (best epoch) | — | 98.47% |
| Validation Accuracy | Domain+Action pair (best epoch) | — | 98.41% |
| Slot Extraction F1 | All 15 slot types | — | 0.993 (99.33%) |
| Training Loss | Epoch 1 → 6 | — | 0.570 → 0.054 → 0.019 → 0.009 → 0.007 → 0.006 |
| Validation Loss | Epoch 1 → 6 | — | 0.1992 → 0.1086 → 0.1283 → 0.1877 → 0.1434 → 0.1301 (selected epoch) |
| Runtime Reliability | 82-turn conversation test | — | 0 errors, 0 crashes |
| Domain Confusion | 10 domains | — | 94.7%+ per-domain, minimal cross-confusion |
| Action Confusion | 33 actions | — | Strong diagonal: 25 of 33 actions at 99%+, lowest minimize at 90.6% |
Live Output Shapes (click to expand)
Command output
{
"type": "command",
"domain": "apps",
"action": "launch",
"slots": {
"APP_NAME": {
"text": "chrome",
"start": 5,
"end": 11,
"confidence": 0.999
}
},
"confidence": 0.97,
"route": "local"
}
Clarification output
{
"type": "clarify",
"question": "What value should I set it to?",
"debug": {
"domain": "volume",
"action": "set",
"reason": "missing_VALUE"
}
}
Label schema
- Domains (10): volume, brightness, media, apps, browser, productivity, screen, window, system, conversation
- Actions (33): up, down, set, mute, unmute, play, pause, next, previous, launch, close, switch, search, set_reminder, screenshot, read, explain, undo, quit, chat, minimize, maximize, restore, focus, copy, paste, cut, lock, sleep, wifi_on, wifi_off, bluetooth_on, bluetooth_off
- Slot labels (BIO, 15): VALUE, APP_NAME, QUERY, DURATION, TIME, WINDOW_NAME, TEXT
Visual Benchmark Evidence
Confusion Matrix — Interactive Breakdown
View original confusion matrix images
Additional diagnostics
Upload-Ready Layout
.
|- README.md
|- .gitattributes
|- LICENSE
|- requirements.txt
|- assets/
| |- jane-janus-glitch.webp
|- janegpt_v2_janus/
| |- __init__.py
| |- architecture.py
| |- dataset.py
| |- inference.py
| |- labels.py
| |- multitask.py
|- runtime/
| |- jane_nlu_runtime.py
|- examples/
| |- demo_inference.py
| |- demo_runtime.py
| |- demo_runtime_suite.py
|- weights/
| |- janegpt_v2_janus.pt
| |- tokenizer.json
|- reports/
| |- fair_benchmarks.json
| |- fair_benchmarks.md
| |- janus_model_report.json
| |- janus_model_report.md
| |- public_benchmarks.json
| |- *.png benchmark visuals
Limitations
- English-focused command language.
- Command NLU model, not an open-domain generative chatbot.
- MASSIVE and SNIPS mapped-intent accuracy is excluded from headline claims because mapping coverage is partial.
Use Cases
- Virtual assistant command routing
- Smart home intent classification
- Voice command understanding
- Chatbot intent detection
- Edge device deployment (small enough for embedded systems)
Part of the JANE Project
JANE — a fully offline, privacy-first AI voice assistant.
🔗 JANE AI Assistant on GitHub 🔗 JaneGPT-v2 on GitHub
Created By
Ravindu Senanayake
Built from scratch — architecture, tokenizer, and training pipeline designed and implemented by the author.
License
Apache-2.0 (see LICENSE).
- Downloads last month
- 40
Evaluation results
- OOD Precision on BANKING77self-reported1.000
- OOD F1 on BANKING77self-reported0.973
- OOD Recall on BANKING77self-reported0.948
- OOD Precision on CLINC OOSself-reported1.000
- OOD F1 on CLINC OOSself-reported0.942
- OOD Recall on CLINC OOSself-reported0.890
- Validation Domain Accuracyself-reported0.989
- Validation Action Accuracyself-reported0.985