nomos-v1-nano-g1

nomos-v1-nano-g1 is a small local agentic tool-routing co-processor. Given a user objective, current agent state, and legal candidate registry, it ranks the tools most likely to be useful for the next step.

It is not an answer generator, tool executor, or fixed-vocabulary classifier. It sits beside an agent as a fast candidate-reduction layer, so the primary LLM can reason over a short ranked list instead of every tool description in a large registry.

Nomos v1 ranks candidates through semantic metadata rather than treating tool names as output classes. The same checkpoint can therefore rank registries supplied by different agents, including tools with previously unseen names.

Native V1 Output

A surrounding tool router can return an ordered candidate list:

Field Value Intended use
tool_id Registry-local tool identifier The legal candidate proposed to the agent.
tool_family Semantic tool family Routing analysis and family-level holdouts.
semantic_fingerprint Stable metadata fingerprint Candidate identity across aliases and registry presentations.
score Normalized cosine similarity Relative ranking within the supplied candidate set.

Output Contract

The raw Hugging Face model returns normalized 384-dimensional embeddings. It does not emit a tool name from a closed label set. A router embeds the current decision state and every legal candidate, computes cosine similarity, and returns the highest-scoring candidates:

[
  {
    "tool_id": "inspect_record",
    "tool_family": "code_inspection",
    "semantic_fingerprint": "a1b2c3d4...",
    "score": 0.81
  },
  {
    "tool_id": "search_catalog",
    "tool_family": "search",
    "semantic_fingerprint": "e5f6a7b8...",
    "score": 0.63
  }
]

Scores are relative to the legal candidate set supplied for that decision. Candidate filtering, argument validation, execution, provenance checks, side- effect policy, and recovery paging remain responsibilities of the surrounding agent runtime.

Intended Use

Use this model when an agent runtime needs fast local signals for:

  • reducing a large legal tool registry to a short top-k candidate list,
  • ranking tools from their capabilities and schemas rather than memorized names,
  • routing across registries owned by different agents,
  • proposing a fresh candidate page when an agent rejects the first page,
  • improving tool selection for language models with weaker agentic behavior,
  • keeping routing local on CPU-only systems.

This model is not intended to choose arguments, execute tools, verify facts, authorize side effects, or replace deterministic runtime validation.

Input Format

The query side should describe the objective and relevant decision state:

Objective: Find the implementation corresponding to this SDK symbol.
Current need: Inspect code structure before opening a specific definition.
Agent state: active execution; source inventory known; schema unknown.
History: searched documentation by exact symbol; no implementation inspected.
Governance: read-only operations are allowed.

Each candidate should carry semantic metadata rather than relying on its name:

Inspect symbols and code structure in a source repository.
Capabilities: inspect code structure.
Accepts: code, text. Returns: symbols.
Evidence role: observation. Prerequisites: repository available.
Constraints: read only. Side effects: none.
Arguments: query (string required), scope (string optional).

Useful candidate fields include the description, capabilities, accepted and returned modalities, evidence role, prerequisites, constraints, side effects, and argument schema. Production integrations should preserve one consistent serialization for both training and inference.

Quick Start

from sentence_transformers import SentenceTransformer

MODEL_ID = "yafitzdev/nomos-v1-nano-g1"
model = SentenceTransformer(MODEL_ID)

state = """Objective: Find the implementation corresponding to this SDK symbol.
Current need: Inspect code structure before opening a specific definition.
Agent state: active execution. Governance: read-only operations are allowed."""

candidates = [
    "Search source text for an exact pattern. Capabilities: exact pattern search.",
    "Inspect symbols and code structure. Capabilities: inspect code structure.",
    "Search public web pages. Capabilities: web search.",
]

query_embedding = model.encode([state], normalize_embeddings=True)
candidate_embeddings = model.encode(candidates, normalize_embeddings=True)
scores = (query_embedding @ candidate_embeddings.T)[0]

for index in scores.argsort()[::-1][:3]:
    print(float(scores[index]), candidates[index])

The Nomos V1 training pairs show the query and candidate serialization used for the final specialist training branch. An integration should keep that serialization consistent between training and inference.

CPU Runtime

This repository contains the native SentenceTransformers checkpoint. No ONNX export is included in this release. Candidate embeddings can be cached when a registry is loaded, leaving only the current decision state to encode during warm routing.

Unoptimized PyTorch CPU measurements on the release workstation:

Legal candidate pool Warm p50 Warm p95 Cold registry + query
10 tools 187 ms 200 ms 0.71 s
30 tools 292 ms 295 ms 1.07 s
100 tools 190 ms 203 ms 2.66 s

The representative state text differs between pool sizes, so warm latency is not expected to rise monotonically after candidate embeddings are cached. Results are workstation-specific and should be remeasured in each deployment.

Evaluation

All figures below measure raw encoder ranking. No downstream LLM choice, repair controller, deterministic override, or tool execution is included.

Frozen evaluation Recall@1 Recall@3
sealed unseen-style states 0.9625 0.9750
independent ToolRet sample 0.6333 0.8000
final multiview suite 0.9257 0.9865
promotion multiview suite 0.9737 1.0000

Additional frozen generic-registry Recall@3 is 0.9440. These suites exercise opaque tool names, changed registry presentation, legal-candidate filtering, hard negatives, and held-out workflows. ToolRet contains 60 independently sampled queries, so its result has substantial sampling uncertainty. None of these benchmarks establishes universal tool-routing accuracy.

Training Data

Field Value
Backbone BAAI/bge-small-en-v1.5
Parameters approximately 33M
Embedding dimension 384
Maximum sequence length 512 tokens
Final specialist input 40,181 rows; 35,812 distinct query/candidate pairs
Main objective in-batch multiple-negative ranking loss
Released weights 90% established base router + 10% specialist interpolation
Similarity normalized cosine similarity

The Nomos V1 dataset is the exact query/positive text snapshot for the final specialist branch recorded by its training manifest. It includes 20,000 generic portability, 4,185 agentic, 4,096 ToolRet-derived, 3,400 transition, 3,400 contrast, and 5,100 balanced hard-subset rows. Repeated pairs are retained because the trainer consumed those rows. The dataset card records source hashes, selection rules, row ranges, and the final file hash.

The base router was trained in earlier stages. The linked dataset therefore documents the final specialist input rather than every stage needed to reproduce the released checkpoint. The uploaded nomos_training_manifest.json is inherited from the base lineage; nomos_interpolation_manifest.json records the final 90/10 weight blend. No evaluation rows are included in the dataset.

Artifacts

This repository contains:

  • model.safetensors: standalone SentenceTransformers checkpoint,
  • tokenizer and BERT configuration files,
  • pooling and normalization modules,
  • nomos_training_manifest.json: inherited training-lineage metadata,
  • nomos_interpolation_manifest.json: final 90/10 interpolation record,
  • nomos_calibration.json: optional abstention-calibration metadata.

The SHA-256 digest of model.safetensors in this release is 429fb204e4a38c1514c5a232376600fb4d3c6d2e82e9e27951cdfdeada27fca8.

Limitations

  1. Top-k retrieval, not execution. The checkpoint ranks candidates but does not select arguments, execute tools, or establish provenance.
  2. English-centric synthetic training. Multilingual performance has not been established.
  3. Independent evaluation is still small. The ToolRet result is useful but does not cover the full diversity of real agent tool registries.
  4. Metadata quality matters. Missing or misleading tool descriptions and state fields can produce incorrect rankings.
  5. Abstention is separate. The included calibration is an optional policy signal and should be validated again for each integration distribution.
  6. Warm latency assumes caching. New or frequently changing registries pay the candidate-embedding cost before ranking.

License

Mixed-source research preview. The BGE backbone is MIT-licensed, while the complete release also reflects synthetic and benchmark-derived training and evaluation sources with their own obligations. Review the dataset card and upstream source terms before redistribution or commercial use.

Downloads last month
47
Safetensors
Model size
33.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yafitzdev/nomos-v1-nano-g1

Finetuned
(512)
this model

Dataset used to train yafitzdev/nomos-v1-nano-g1

Collection including yafitzdev/nomos-v1-nano-g1