graph-reasoner-1
A 99M-parameter graph-native reasoning model. Given a typed graph and a structured query, it returns typed answers together with verifier-checked evidence paths. It does not consume or emit natural language.
Architecture: relation-aware GraphSAGE encoder β query-conditioned reasoning transformer β Path-Aware Transformer (PAT) over typed path tokens β learned path selection β deterministic verifier.
98,994,604 parameters. Trained for 2,500 steps on synthetic typed graphs.
What it is actually good at
One thing: inferring which relation chain answers an intent, from graph structure, when the query does not name the path.
| measurement | value |
|---|---|
| exact match, in-distribution (150 stage-balanced examples) | 0.473 |
| answer F1 | 0.546 |
| selection precision | 0.455 |
| endpoint precision | 0.701 |
| vs. traversal heuristics (exact match 0.093) | 5.1Γ |
Ablations, same slice β each component is load-bearing:
| ablation | exact match |
|---|---|
| full model | 0.473 |
| without PAT | 0.153 |
| without query conditioning | 0.187 |
| without learned selection | 0.120 |
| greedy instead of beam | 0.373 |
It knows when to say nothing
On examples where the supporting evidence has been deleted, abstaining is the only correct answer:
| abstains on unanswerable | false abstentions | balanced accuracy | |
|---|---|---|---|
| previous checkpoint | 0.017 | 0.034 | 0.491 (chance) |
| this checkpoint | 0.650 | 0.017 | 0.783 |
Real knowledge graphs, composition withheld
Both models receive only the question text, the start entities and a hop budget:
| workload | G-Reasoner-34M MRR | this model MRR |
|---|---|---|
| FB15k-237 2-hop | 0.084 | 0.175 |
| WN18RR 2-hop | 0.252 | 0.596 |
Limitations β read before using
These are measured, not hypothetical.
- The abstention gate does not transfer. Its evidence signals come from PAT
query-compatibility, which saturates low on relation vocabulary the model has never
seen. On real typed KGs it fires on 100% of queries and drives every metric to
0.000 β even when the gold relation path is supplied. Outside the training
distribution you must pass
abstain=Falseor lowerabstention_threshold. Nothing warns you that it has fired on everything. This is the most important rough edge. - It does not generalise across ontologies. On held-out graphs whose relation labels are renamed, exact match is 0.158 and PAT discrimination is 0.577 against a 0.5 floor β near chance. In practice this means you must train on your own ontology. It is a train-per-domain tool, not a drop-in model.
- It has never been run on a real product graph. Its entire experience is one synthetic ontology (10 entity types, 18 relations) plus FB15k-237, WN18RR and LLM-free HotpotQA entity graphs.
- If you already know the relation path, do not use this. A breadth-first traversal scores a perfect 1.000 where this model reaches 0.58β0.92, and runs at ~7,100 queries per second against this model's ~20.
- Throughput is low: ~49 ms / 20 qps at 100 nodes; ~323 ms / 3 qps at 1,000 nodes and 5 hops (A100).
- Scale is unverified above ~5,000 nodes.
- Calibration is mediocre: isotonic ECE 0.102. Use the calibrated confidence, not the raw head output.
- On text-derived graphs it merely ties a heuristic. HotpotQA bridge recall 0.518 against the heuristic's 0.518 β matching, not beating.
- 2,500 steps is deliberate. The same configuration trained to 6,000 steps was worse (abstention balanced accuracy 0.817 β 0.684), with seeds diverging in opposite directions. This configuration needs early stopping and is probably not at the architecture's ceiling.
Not a language model
This is not comparable to an LLM and has not been benchmarked against one. It takes a typed graph plus a structured query; it does not read or write prose. For graphs small enough to fit in a context window, a frontier LLM would very likely be more accurate. The only argument for this model is graphs too large for a context window, and that argument is unsupported β nothing above ~5,000 nodes has been tested.
Usage
from graph.export import load_inference_artifact
from graph.graph import GraphStore
from graph.types import GraphQuery
model = load_inference_artifact("path/to/graph-reasoner-1")
graph = GraphStore.from_dict(your_graph)
query = GraphQuery(text="which supplier is behind this project?", start_entity_ids=("n1",), max_hops=3)
result = model.reason(graph, query) # in-distribution
result = model.reason(graph, query, abstain=False) # out-of-distribution: see limitation 1
for answer in result.answers:
print(answer.entity_id, answer.score)
for path in result.supporting_paths: # every edge verified to exist
print(path.relation_sequence, path.endpoint_id)
Code: graph-reasoning-model (Apache-2.0). Training data: frontal-labs/graph-reasoner1-trainning-data.
Provenance
run_manifest.json ships alongside the weights and records the config hash, dataset
hash, seed, parameter buckets, platform and torch version. It carries the version strings
grm-architecture-v2 / grm-synthetic-2 because those are what the training run
actually emitted β the project was renamed from "GRM" to "Graph" after this checkpoint
was trained, and the manifest is left as produced rather than retconned.
Derived in part from GFM-RAG / G-Reasoner (Apache-2.0); the retrieval subpackage is vendored. The weights here are trained from scratch and share no parameters with upstream.
- Downloads last month
- 16