graph-reasoner-1

A 99M-parameter graph-native reasoning model. Given a typed graph and a structured query, it returns typed answers together with verifier-checked evidence paths. It does not consume or emit natural language.

Architecture: relation-aware GraphSAGE encoder β†’ query-conditioned reasoning transformer β†’ Path-Aware Transformer (PAT) over typed path tokens β†’ learned path selection β†’ deterministic verifier.

98,994,604 parameters. Trained for 2,500 steps on synthetic typed graphs.

What it is actually good at

One thing: inferring which relation chain answers an intent, from graph structure, when the query does not name the path.

measurement value
exact match, in-distribution (150 stage-balanced examples) 0.473
answer F1 0.546
selection precision 0.455
endpoint precision 0.701
vs. traversal heuristics (exact match 0.093) 5.1Γ—

Ablations, same slice β€” each component is load-bearing:

ablation exact match
full model 0.473
without PAT 0.153
without query conditioning 0.187
without learned selection 0.120
greedy instead of beam 0.373

It knows when to say nothing

On examples where the supporting evidence has been deleted, abstaining is the only correct answer:

abstains on unanswerable false abstentions balanced accuracy
previous checkpoint 0.017 0.034 0.491 (chance)
this checkpoint 0.650 0.017 0.783

Real knowledge graphs, composition withheld

Both models receive only the question text, the start entities and a hop budget:

workload G-Reasoner-34M MRR this model MRR
FB15k-237 2-hop 0.084 0.175
WN18RR 2-hop 0.252 0.596

Limitations β€” read before using

These are measured, not hypothetical.

  1. The abstention gate does not transfer. Its evidence signals come from PAT query-compatibility, which saturates low on relation vocabulary the model has never seen. On real typed KGs it fires on 100% of queries and drives every metric to 0.000 β€” even when the gold relation path is supplied. Outside the training distribution you must pass abstain=False or lower abstention_threshold. Nothing warns you that it has fired on everything. This is the most important rough edge.
  2. It does not generalise across ontologies. On held-out graphs whose relation labels are renamed, exact match is 0.158 and PAT discrimination is 0.577 against a 0.5 floor β€” near chance. In practice this means you must train on your own ontology. It is a train-per-domain tool, not a drop-in model.
  3. It has never been run on a real product graph. Its entire experience is one synthetic ontology (10 entity types, 18 relations) plus FB15k-237, WN18RR and LLM-free HotpotQA entity graphs.
  4. If you already know the relation path, do not use this. A breadth-first traversal scores a perfect 1.000 where this model reaches 0.58–0.92, and runs at ~7,100 queries per second against this model's ~20.
  5. Throughput is low: ~49 ms / 20 qps at 100 nodes; ~323 ms / 3 qps at 1,000 nodes and 5 hops (A100).
  6. Scale is unverified above ~5,000 nodes.
  7. Calibration is mediocre: isotonic ECE 0.102. Use the calibrated confidence, not the raw head output.
  8. On text-derived graphs it merely ties a heuristic. HotpotQA bridge recall 0.518 against the heuristic's 0.518 β€” matching, not beating.
  9. 2,500 steps is deliberate. The same configuration trained to 6,000 steps was worse (abstention balanced accuracy 0.817 β†’ 0.684), with seeds diverging in opposite directions. This configuration needs early stopping and is probably not at the architecture's ceiling.

Not a language model

This is not comparable to an LLM and has not been benchmarked against one. It takes a typed graph plus a structured query; it does not read or write prose. For graphs small enough to fit in a context window, a frontier LLM would very likely be more accurate. The only argument for this model is graphs too large for a context window, and that argument is unsupported β€” nothing above ~5,000 nodes has been tested.

Usage

from graph.export import load_inference_artifact
from graph.graph import GraphStore
from graph.types import GraphQuery

model = load_inference_artifact("path/to/graph-reasoner-1")
graph = GraphStore.from_dict(your_graph)
query = GraphQuery(text="which supplier is behind this project?", start_entity_ids=("n1",), max_hops=3)

result = model.reason(graph, query)              # in-distribution
result = model.reason(graph, query, abstain=False)  # out-of-distribution: see limitation 1

for answer in result.answers:
    print(answer.entity_id, answer.score)
for path in result.supporting_paths:             # every edge verified to exist
    print(path.relation_sequence, path.endpoint_id)

Code: graph-reasoning-model (Apache-2.0). Training data: frontal-labs/graph-reasoner1-trainning-data.

Provenance

run_manifest.json ships alongside the weights and records the config hash, dataset hash, seed, parameter buckets, platform and torch version. It carries the version strings grm-architecture-v2 / grm-synthetic-2 because those are what the training run actually emitted β€” the project was renamed from "GRM" to "Graph" after this checkpoint was trained, and the manifest is left as produced rather than retconned.

Derived in part from GFM-RAG / G-Reasoner (Apache-2.0); the retrieval subpackage is vendored. The weights here are trained from scratch and share no parameters with upstream.

Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support