You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Model Card for Falco Block-Sparse Multi-Query Router

The Falco Multi-Query Router is a specialized, fine-tuned sequence classification model exported to ONNX format. It is explicitly designed to act as the weights backend for the Rust-based Falco Neural Decision Routing Engine.

Model Details

  • Base Architecture: distilbert-base-uncased
  • Export Format: ONNX (Opset 18)
  • Execution Provider: CPU Optimized (ort)
  • Calibration: Post-training temperature scaled via LBFGS (Optimal $T = 1.82$)

Model Architecture & Custom Input Contract

Unlike standard Hugging Face sequence classifiers, this model uses a specialized multi-query architecture. It does not accept standard 2D attention masks. It requires:

  1. input_ids [Rank 2, int64]: Packed sequence [CLS] State [SEP] Query 1 [SEP] Query 2 [SEP].
  2. attention_mask [Rank 4, float32]: Additive block-sparse mask to prevent cross-contamination between parallel queries in the same batch.
  3. query_indices [Rank 2, int64]: Pointers to the exact token indices where query representations are gathered for the classification head.
  4. position_ids [Rank 2, int64]: Custom positional embeddings aligned with the block-sparse map.

Intended Use

This model is intended for deployment within the Falco Rust Engine. It is not designed to be loaded directly via the standard Python transformers pipeline due to its custom 4D attention mask and index-gathering classification head.

Primary Use Cases:

  • Semantic Tool Selection: Powering local MCP (Model Context Protocol) servers to select tools efficiently.
  • LLM Gateway Routing: Deciding which foundation model (e.g., fast/cheap vs. slow/expensive) should handle an incoming user prompt.
  • Automated SRE: Analyzing system state prefixes (log streams) to recommend immediate operational queries (e.g., "Scale Database Nodes").

Training and Calibration

The model was trained using Supervised Fine-Tuning (SFT) on paired state-query datasets. Following SFT, the model underwent Expected Calibration Error (ECE) optimization. The uncalibrated logits exhibited overconfidence. By applying temperature scaling ($T \approx 1.82$), the model's ECE was reduced to under the 0.05 production threshold, ensuring that output probability values represent accurate real-world confidences suitable for automated thresholding in production.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support