sentence-transformers
English
agentweave_semantic_router
agentweave
agentic-ai
tool-routing
semantic-routing
function-calling
cpu
minilm
pre-inference-routing
Instructions to use sauravsingla08/AgentWeave-Router-MiniLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use sauravsingla08/AgentWeave-Router-MiniLM with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("sauravsingla08/AgentWeave-Router-MiniLM") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
Publish AgentWeave Router MiniLM from 3791f8e
Browse files- README.md +161 -0
- config.json +11 -0
- requirements.txt +2 -0
- route_prototypes.json +42 -0
- router.py +84 -0
README.md
ADDED
|
@@ -0,0 +1,161 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
library_name: sentence-transformers
|
| 4 |
+
pipeline_tag: feature-extraction
|
| 5 |
+
base_model: sentence-transformers/all-MiniLM-L6-v2
|
| 6 |
+
tags:
|
| 7 |
+
- agentweave
|
| 8 |
+
- agentic-ai
|
| 9 |
+
- tool-routing
|
| 10 |
+
- semantic-routing
|
| 11 |
+
- function-calling
|
| 12 |
+
- cpu
|
| 13 |
+
- minilm
|
| 14 |
+
- sentence-transformers
|
| 15 |
+
- pre-inference-routing
|
| 16 |
+
language:
|
| 17 |
+
- en
|
| 18 |
+
---
|
| 19 |
+
|
| 20 |
+
# AgentWeave Router MiniLM 🧭
|
| 21 |
+
|
| 22 |
+
> **Route before you reason.**
|
| 23 |
+
|
| 24 |
+
A lightweight, CPU-first semantic capability router for **AgentWeave**. It uses `sentence-transformers/all-MiniLM-L6-v2` as a frozen embedding encoder and ranks route prototypes with cosine similarity before downstream model inference.
|
| 25 |
+
|
| 26 |
+
This repository is intentionally small: it publishes the AgentWeave routing configuration, route prototypes, and executable router code while reusing the upstream MiniLM encoder at runtime instead of copying its weights.
|
| 27 |
+
|
| 28 |
+
## Why this exists
|
| 29 |
+
|
| 30 |
+
Tool-rich agents can expose large action spaces to a language model. AgentWeave explores a complementary systems strategy: reduce the candidate action space *before* model reasoning. This model repository provides an experimental semantic routing companion to AgentWeave's default deterministic routing path.
|
| 31 |
+
|
| 32 |
+
### Route families
|
| 33 |
+
|
| 34 |
+
- 🔎 `research`
|
| 35 |
+
- 📚 `retrieval`
|
| 36 |
+
- 🧠 `analysis`
|
| 37 |
+
- 💻 `coding`
|
| 38 |
+
- 🗺️ `planning`
|
| 39 |
+
- ✅ `verification`
|
| 40 |
+
- 📝 `summarization`
|
| 41 |
+
- 📊 `data_analysis`
|
| 42 |
+
|
| 43 |
+
## Architecture
|
| 44 |
+
|
| 45 |
+
```text
|
| 46 |
+
Task / user request
|
| 47 |
+
│
|
| 48 |
+
▼
|
| 49 |
+
all-MiniLM-L6-v2
|
| 50 |
+
384-d embedding
|
| 51 |
+
│
|
| 52 |
+
├──────────────┐
|
| 53 |
+
▼ ▼
|
| 54 |
+
query vector route prototypes
|
| 55 |
+
│ │
|
| 56 |
+
└──── cosine ──┘
|
| 57 |
+
│
|
| 58 |
+
▼
|
| 59 |
+
ranked route set
|
| 60 |
+
│
|
| 61 |
+
▼
|
| 62 |
+
downstream AgentWeave
|
| 63 |
+
```
|
| 64 |
+
|
| 65 |
+
**No fine-tuning is claimed.** This is a prototype-based semantic router built on a frozen MiniLM encoder. Similarity scores are ranking signals, **not calibrated probabilities**.
|
| 66 |
+
|
| 67 |
+
## Quick start
|
| 68 |
+
|
| 69 |
+
```bash
|
| 70 |
+
pip install -r requirements.txt
|
| 71 |
+
python router.py "research the latest protocol changes, verify the sources, and summarize the findings"
|
| 72 |
+
```
|
| 73 |
+
|
| 74 |
+
Example output shape:
|
| 75 |
+
|
| 76 |
+
```json
|
| 77 |
+
[
|
| 78 |
+
{"route": "research", "score": 0.0},
|
| 79 |
+
{"route": "verification", "score": 0.0},
|
| 80 |
+
{"route": "summarization", "score": 0.0}
|
| 81 |
+
]
|
| 82 |
+
```
|
| 83 |
+
|
| 84 |
+
The numeric values above are placeholders showing the response schema; actual scores are computed locally from MiniLM embeddings.
|
| 85 |
+
|
| 86 |
+
## Python usage
|
| 87 |
+
|
| 88 |
+
```python
|
| 89 |
+
from router import AgentWeaveSemanticRouter
|
| 90 |
+
|
| 91 |
+
router = AgentWeaveSemanticRouter()
|
| 92 |
+
routes = router.route(
|
| 93 |
+
"inspect this code, identify correctness risks, and propose a fix",
|
| 94 |
+
top_k=3,
|
| 95 |
+
)
|
| 96 |
+
print(routes)
|
| 97 |
+
```
|
| 98 |
+
|
| 99 |
+
## CPU-first design
|
| 100 |
+
|
| 101 |
+
The router is designed for lightweight local execution:
|
| 102 |
+
|
| 103 |
+
- frozen MiniLM encoder
|
| 104 |
+
- no text generation
|
| 105 |
+
- no external inference API required
|
| 106 |
+
- normalized embeddings + cosine ranking
|
| 107 |
+
- small route-prototype file
|
| 108 |
+
|
| 109 |
+
The first run downloads the upstream MiniLM encoder. Subsequent runs can use the local Hugging Face cache.
|
| 110 |
+
|
| 111 |
+
## Relationship to AgentWeave
|
| 112 |
+
|
| 113 |
+
AgentWeave's documented default BYOM routing path is deterministic and provider-neutral. This MiniLM router is an **experimental semantic companion**, not a replacement for the default router and not the source of AgentWeave's published deterministic-router benchmark claims.
|
| 114 |
+
|
| 115 |
+
Relevance routing also does **not** grant permission to execute a tool. Policy filtering, scope controls, and authorization remain separate boundaries in AgentWeave.
|
| 116 |
+
|
| 117 |
+
## Files
|
| 118 |
+
|
| 119 |
+
| File | Purpose |
|
| 120 |
+
|---|---|
|
| 121 |
+
| `router.py` | CPU semantic router implementation |
|
| 122 |
+
| `route_prototypes.json` | Human-readable capability prototypes |
|
| 123 |
+
| `config.json` | Base model and routing configuration |
|
| 124 |
+
| `requirements.txt` | Minimal runtime dependencies |
|
| 125 |
+
|
| 126 |
+
## Intended use
|
| 127 |
+
|
| 128 |
+
Good fits:
|
| 129 |
+
|
| 130 |
+
- pre-inference capability routing
|
| 131 |
+
- agent/tool candidate reduction experiments
|
| 132 |
+
- CPU routing demos
|
| 133 |
+
- semantic route exploration
|
| 134 |
+
- research comparisons with deterministic routing
|
| 135 |
+
|
| 136 |
+
Not intended as:
|
| 137 |
+
|
| 138 |
+
- a calibrated confidence model
|
| 139 |
+
- an authorization engine
|
| 140 |
+
- a safety classifier
|
| 141 |
+
- a replacement for downstream function-calling evaluation
|
| 142 |
+
|
| 143 |
+
## Limitations
|
| 144 |
+
|
| 145 |
+
- English-focused route prototypes
|
| 146 |
+
- prototype wording influences ranking
|
| 147 |
+
- route scores are cosine similarities, not probabilities
|
| 148 |
+
- the route taxonomy is intentionally compact
|
| 149 |
+
- domain-specific tools may need custom prototypes
|
| 150 |
+
|
| 151 |
+
## Source
|
| 152 |
+
|
| 153 |
+
AgentWeave source code and research artifacts are maintained at:
|
| 154 |
+
|
| 155 |
+
`https://github.com/sauravsingla/AgentWeave`
|
| 156 |
+
|
| 157 |
+
The Hugging Face Space provides an interactive companion experience under the same project name.
|
| 158 |
+
|
| 159 |
+
## License
|
| 160 |
+
|
| 161 |
+
Apache-2.0. The upstream `sentence-transformers/all-MiniLM-L6-v2` model is loaded separately at runtime and remains subject to its own model card and license terms.
|
config.json
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model_type": "agentweave_semantic_router",
|
| 3 |
+
"architecture": "prototype_cosine_router",
|
| 4 |
+
"base_model": "sentence-transformers/all-MiniLM-L6-v2",
|
| 5 |
+
"embedding_dimension": 384,
|
| 6 |
+
"similarity": "cosine",
|
| 7 |
+
"normalize_embeddings": true,
|
| 8 |
+
"default_top_k": 3,
|
| 9 |
+
"training": "none; prototype-based semantic routing",
|
| 10 |
+
"version": "0.1.0"
|
| 11 |
+
}
|
requirements.txt
ADDED
|
@@ -0,0 +1,2 @@
|
|
|
|
|
|
|
|
|
|
| 1 |
+
sentence-transformers>=3.0.0
|
| 2 |
+
numpy>=1.26.0
|
route_prototypes.json
ADDED
|
@@ -0,0 +1,42 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"research": [
|
| 3 |
+
"research a technical topic using authoritative sources",
|
| 4 |
+
"compare alternatives and gather supporting evidence",
|
| 5 |
+
"investigate prior work, documentation, or literature"
|
| 6 |
+
],
|
| 7 |
+
"retrieval": [
|
| 8 |
+
"retrieve facts, records, documents, or external information",
|
| 9 |
+
"search for relevant evidence before answering",
|
| 10 |
+
"look up documentation, references, or source material"
|
| 11 |
+
],
|
| 12 |
+
"analysis": [
|
| 13 |
+
"analyze evidence, systems, tradeoffs, risks, or behavior",
|
| 14 |
+
"reason over multiple inputs and identify important patterns",
|
| 15 |
+
"evaluate a technical claim or architecture"
|
| 16 |
+
],
|
| 17 |
+
"coding": [
|
| 18 |
+
"write, inspect, debug, or modify source code",
|
| 19 |
+
"build a prototype or implementation",
|
| 20 |
+
"test code and identify correctness problems"
|
| 21 |
+
],
|
| 22 |
+
"planning": [
|
| 23 |
+
"create an implementation plan, workflow, or sequence of actions",
|
| 24 |
+
"design a migration, experiment, or engineering approach",
|
| 25 |
+
"organize tasks and optimize the execution path"
|
| 26 |
+
],
|
| 27 |
+
"verification": [
|
| 28 |
+
"verify a claim, result, constraint, or implementation",
|
| 29 |
+
"independently validate evidence or conclusions",
|
| 30 |
+
"check consistency, correctness, compliance, or reliability"
|
| 31 |
+
],
|
| 32 |
+
"summarization": [
|
| 33 |
+
"summarize findings into a concise structured response",
|
| 34 |
+
"compress multiple analyses while preserving key conclusions",
|
| 35 |
+
"produce a clear synthesis of technical information"
|
| 36 |
+
],
|
| 37 |
+
"data_analysis": [
|
| 38 |
+
"analyze a dataset and identify trends or anomalies",
|
| 39 |
+
"inspect metrics, measurements, tables, or experimental results",
|
| 40 |
+
"perform quantitative analysis and summarize findings"
|
| 41 |
+
]
|
| 42 |
+
}
|
router.py
ADDED
|
@@ -0,0 +1,84 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
from __future__ import annotations
|
| 2 |
+
|
| 3 |
+
import argparse
|
| 4 |
+
import json
|
| 5 |
+
from pathlib import Path
|
| 6 |
+
from typing import Dict, List
|
| 7 |
+
|
| 8 |
+
import numpy as np
|
| 9 |
+
from sentence_transformers import SentenceTransformer
|
| 10 |
+
|
| 11 |
+
|
| 12 |
+
ROOT = Path(__file__).resolve().parent
|
| 13 |
+
|
| 14 |
+
|
| 15 |
+
class AgentWeaveSemanticRouter:
|
| 16 |
+
"""Prototype-based semantic capability router built on MiniLM embeddings.
|
| 17 |
+
|
| 18 |
+
This is an experimental semantic companion to AgentWeave's default
|
| 19 |
+
deterministic routing path. It does not replace AgentWeave policy,
|
| 20 |
+
authorization, or execution controls.
|
| 21 |
+
"""
|
| 22 |
+
|
| 23 |
+
def __init__(
|
| 24 |
+
self,
|
| 25 |
+
config_path: str | Path = ROOT / "config.json",
|
| 26 |
+
prototypes_path: str | Path = ROOT / "route_prototypes.json",
|
| 27 |
+
) -> None:
|
| 28 |
+
self.config = json.loads(Path(config_path).read_text(encoding="utf-8"))
|
| 29 |
+
self.prototypes: Dict[str, List[str]] = json.loads(
|
| 30 |
+
Path(prototypes_path).read_text(encoding="utf-8")
|
| 31 |
+
)
|
| 32 |
+
self.model = SentenceTransformer(self.config["base_model"], device="cpu")
|
| 33 |
+
|
| 34 |
+
texts: List[str] = []
|
| 35 |
+
labels: List[str] = []
|
| 36 |
+
for label, examples in self.prototypes.items():
|
| 37 |
+
for example in examples:
|
| 38 |
+
labels.append(label)
|
| 39 |
+
texts.append(example)
|
| 40 |
+
|
| 41 |
+
self._prototype_labels = labels
|
| 42 |
+
self._prototype_embeddings = self.model.encode(
|
| 43 |
+
texts,
|
| 44 |
+
normalize_embeddings=bool(self.config.get("normalize_embeddings", True)),
|
| 45 |
+
convert_to_numpy=True,
|
| 46 |
+
show_progress_bar=False,
|
| 47 |
+
)
|
| 48 |
+
|
| 49 |
+
def route(self, query: str, top_k: int | None = None) -> List[dict]:
|
| 50 |
+
if not query or not query.strip():
|
| 51 |
+
raise ValueError("query must be a non-empty string")
|
| 52 |
+
|
| 53 |
+
top_k = int(top_k or self.config.get("default_top_k", 3))
|
| 54 |
+
query_embedding = self.model.encode(
|
| 55 |
+
[query],
|
| 56 |
+
normalize_embeddings=bool(self.config.get("normalize_embeddings", True)),
|
| 57 |
+
convert_to_numpy=True,
|
| 58 |
+
show_progress_bar=False,
|
| 59 |
+
)[0]
|
| 60 |
+
|
| 61 |
+
similarities = self._prototype_embeddings @ query_embedding
|
| 62 |
+
best_by_label: Dict[str, float] = {}
|
| 63 |
+
for label, score in zip(self._prototype_labels, similarities):
|
| 64 |
+
best_by_label[label] = max(best_by_label.get(label, -1.0), float(score))
|
| 65 |
+
|
| 66 |
+
ranked = sorted(best_by_label.items(), key=lambda item: item[1], reverse=True)
|
| 67 |
+
return [
|
| 68 |
+
{"route": label, "score": round(score, 6)}
|
| 69 |
+
for label, score in ranked[: max(1, min(top_k, len(ranked)))]
|
| 70 |
+
]
|
| 71 |
+
|
| 72 |
+
|
| 73 |
+
def main() -> None:
|
| 74 |
+
parser = argparse.ArgumentParser(description="AgentWeave MiniLM semantic router")
|
| 75 |
+
parser.add_argument("query", help="Task or request to route")
|
| 76 |
+
parser.add_argument("--top-k", type=int, default=None, help="Number of routes to return")
|
| 77 |
+
args = parser.parse_args()
|
| 78 |
+
|
| 79 |
+
router = AgentWeaveSemanticRouter()
|
| 80 |
+
print(json.dumps(router.route(args.query, args.top_k), indent=2))
|
| 81 |
+
|
| 82 |
+
|
| 83 |
+
if __name__ == "__main__":
|
| 84 |
+
main()
|