How to use from the
Use from the
Transformers library
# Gated model: Login with a HF token with gated access permission
hf auth login
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("feature-extraction", model="Elda-AI/memory-resoner", trust_remote_code=True)
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True, device_map="auto")
Quick Links

Access to memory-resoner

Released for research, evaluation, internal validation and education. Tell us who you are and we will grant access.

By requesting access you agree to the Elda Community License 1.0: no commercial use, no redistribution of the weights or derivatives, and attribution as "Built with Elda". Commercial licensing is available on request.

Log in or Sign Up to review the conditions and access this model content.

memory-resoner β€” the reference step of conversational memory

A conversation can only be stored if you know what γ€Œκ·Έ 동넀」 meant. memory-resoner reads the turns of a conversation and says, for each mention, which earlier mention it refers to β€” so the layer above can write a memory entry that points at the right words.

"합정동에 κ°€κ²Œλ₯Ό μ—΄μ–΄μš” . κ·Έ λ™λ„€λŠ” μœ λ™μΈκ΅¬κ°€ λ§Žλ‚˜μš” ?"

  mentions   합정동에   (s0, w0–0)
             κ·Έ λ™λ„€λŠ”  (s1, w0–1)
  chains     [합정동에 Β· κ·Έ λ™λ„€λŠ”]

version 0.2.6 Β· 308M parameters Β· fp32 Β· 33 ms per sentence on CPU when the mentions are supplied (124 ms when it finds them itself).



Where it sits

memory-resoner is one stage of a conversational stack, and it is multilingual β€” trained on Korean and English, measured on Japanese, and it transfers to languages it never saw. Each stage answers one question and hands the next stage a span, never prose.

utterance
   β”‚
   β”œβ”€ Elda-AI/intenter              what was said, and what kind of thing it is  ── spans come from here
   β”œβ”€ slot extraction               which of those the user said about themselves
   β”‚
   β–Ό
β˜… memory-resoner                    β˜… which earlier mention does this one refer to
   β”‚
   β–Ό
memory write                        a pointer into the user's own words β€” checkable, not generated

Where the spans come from. In the Elda stack, mention spans are produced upstream by Elda-AI/intenter and this model consumes them; its card is the reference for how a span is drawn, which types exist, and what the channels mean. Used standalone, memory-resoner will find its own mentions instead.

⚠ Those two paths are not equally good, and the gap is large. On KLUE-WoS dialogue reference, this version scores 0.5640 when the mentions are supplied and 0.2319 when it has to find them itself. Detection and linking are two abilities and the end-to-end number is their product β€” so if you already have spans upstream, pass them in.

⚠ The two sides count spans in different units, and the conversion is yours to make. Upstream spans are character offsets with the particle left outside (γ€Œν•©μ •λ™γ€). This model works in μ–΄μ ˆ (word) indices, and a μ–΄μ ˆ contains its particle, so the same mention is γ€Œν•©μ •λ™μ—γ€ here. Neither is wrong; they are different units. Align on the μ–΄μ ˆ that contains the upstream span.

What it returns. Span indices and chains. Not a rewritten sentence, not a summary β€” an index into the words you sent, which the layer above can verify before storing anything.

What it does not decide. Whether a fact is true, whether it is worth keeping, how long it lives. Those belong to the system around it. This stage answers one question and stops.



Why a small encoder

this model
Parameters 308M
Latency per sentence, CPU β€” linking, mentions given p50 33 ms
Latency per sentence, CPU β€” finding mentions and linking p50 124 ms
Output span index β€” checkable byte-for-byte
Determinism same input, same chains

Memory is written on every turn. A stage that runs that often has to be cheap, and its answer has to be something the next stage can check rather than trust. A span index is both.



Output contract

input        sentences, in order β€” a string per sentence, or a list of μ–΄μ ˆ
output       mentions Β· antecedents Β· chains (links closed transitively)
spans        inclusive word (μ–΄μ ˆ) indices β€” Korean particles stay attached (see the note above)
window       6 previous sentences of context Β· 40 previous mentions as candidates

A mention with no antecedent opens a chain of its own β€” that is how the next stage learns it has seen a new entity, so singleton chains are kept rather than dropped.



Usage

from transformers import AutoModel, AutoTokenizer

m   = AutoModel.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True).eval()
tok = AutoTokenizer.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True)

out = m.coref(["합정동에 κ°€κ²Œλ₯Ό μ—΄μ–΄μš” .",
               "κ·Έ λ™λ„€λŠ” μœ λ™μΈκ΅¬κ°€ λ§Žλ‚˜μš” ?"], tok)

out["mentions"]     # [{'sent': 0, 'words': [0, 0], 'text': '합정동에'}, ...]
out["antecedents"]  # [None, 0]   ← per mention: what it points back at
out["chains"]       # [[0, 1]]
call does
m.coref(sentences, tok) finds the mentions and links them
m.mentions(sentences, tok) finds mentions only β€” a span provider for another linker
m.resolve(units, mentions, tok) links only β€” you supply the mentions
m.coref_last(sentences, tok) β˜… serving shape β€” only the newest turn's mentions are returned, with the earlier turns kept as context
m.resolve_last(sentences, mentions, tok) β˜… serving shape, mentions supplied β€” the same, but you hand in the spans

The two _last calls are what a live conversation wants: one turn arrives, you want answers for that turn, and the turns before it are context rather than output. They take the whole conversation and cut the window themselves, so the caller never has to know how long the window is.

β›” On coref_last, do not count with antecedents

The two calls differ here, and it matters. resolve_last returns all the mentions you gave it, so its antecedents[i] indexes that same list and reads normally. coref_last returns only the newest turn's mentions β€” so when the antecedent sits in an earlier turn, which on a live conversation is almost always, antecedents[i] is None. There, None wears one face for two different facts: "it pointed at something outside this list" and "it pointed at nothing", and a consumer that counts non-None gets a structural zero even when the model is working.

field what it is use it for
linked[i] β˜… bool β€” did this mention resolve at all counting, on either call
n_linked sum of the above counting
antecedent_mentions[i] the antecedent itself ({sent, words, text}), or None reading the answer
antecedents[i] index into the returned list, or None safe on resolve_last; on coref_last only for within-turn links

resolve_last also reports what it had to drop or fix, and none of it is thrown away silently: mentions_outside_window Β· mentions_unencodable Β· window_truncated Β· mentions_out_of_order. ⚠ When mentions_out_of_order > 0 the layer re-sorted your mentions into reading order, and then antecedents[i] > i can occur β€” the index is into the list you gave, not into reading order. If that count is zero, "an antecedent index is always smaller than its own" holds.

Notes that matter in practice:

  • fp32. The config pins it. In bf16 near-ties flip and chains change.
  • Pass sentences, not a paragraph. The sentence boundary is what the window is counted in.
  • Send the resolved span downstream, not the reference. The output is an index into the words you sent, so whatever consumes it can verify what it was handed.
  • Strip the particle at display time, not at span time. Keeping it inside the μ–΄μ ˆ is what makes the span a plain index into your own input; trimming belongs to whatever renders the value.


Measurements

Public multilingual coreference, held out from training; document overlap with the training split is 0 and the benchmark script asserts it. Both scripts ship in this repository.

⚠ Korean is one of the languages here, not the only one. If your pipeline is Korean-only, measure before you swap: reaching other languages costs some Korean headroom, and the trade is visible below. The two scripts here let you check that against your own data rather than take our word for it.

pip install torch transformers datasets
python bench_corefud_ko.py --model Elda-AI/memory-resoner

Mention detection β€” boundaries must match exactly

precision recall F1
0.7917 0.8087 0.8001

Linking, mentions given

scored mentions 10263
accuracy 0.8803
non-NULL accuracy 0.7045 (n=2721)
deictic, non-NULL 0.7159 (n=271)
reference β€” always NULL 0.7349
reference β€” always most-recent 0.0448

Read the non-NULL rows. Most mentions open a chain, so answering NULL every time already scores 0.7349; the rows that require actually linking are the ones in bold.

End-to-end β€” (mention, antecedent) pairs, the model finding its own mentions

precision recall F1
0.6188 0.6661 0.6415

Every run first pushes the gold answers back through the scorer and asserts they survive intact (1.0000). A scorer that cannot return its own gold is not grading anything.

On dialogue

The numbers above are prose. On dialogue β€” references mined from a public multi-turn Korean corpus (KLUE-WoS), with the antecedent taken from that corpus's own annotation rather than chosen by us β€” this version resolves γ€Œκ·Έ 식당」-style references at 0.5640 when the mentions are supplied, and 0.2319 when it has to find them itself.

python bench_wos_references.py --model Elda-AI/memory-resoner

β˜… That difference is the whole point of the two paths. End-to-end is detection times linking, so a pipeline that already has spans should pass them in rather than ask for them again. ⚠ These are measured with candidates mined by the gold rule, which makes them a perfect-upstream ceiling β€” a real intenter will score lower. These numbers are for conversation; on plain prose the same model reads reference less reliably.



What comes next

Conversational memory here is built in three layers. They are not future versions of this model β€” they are different kinds of component, and saying so is the point.

what it is for how it is built
Layer 1 Β· conversation memory what was said in this conversation, and what refers to what β˜… this model + deterministic assembly
Layer 2 blackbox β€” how a memory is addressed and kept apart from every other conversation not described here
Layer 3 Β· authority and persona whether something holds and on what grounds, and what an agent is like across conversations a deterministic VM, and a decoder used off the request path

Layer 1 is the only one this model sits in. Layer 2 is deliberately not described: it is where memories are separated from one another, and it is the part of the system we keep closed. Layer 3 is where facts acquire grounds and where a persona is formed, and it is built from two very different machines β€” one that must be exact, and one that must be fluent.

The VM β€” the exact one. Memory is written in a small closed language and executed by a deterministic machine, so a write is either accepted, a duplicate, a replacement, or refused, and the same inputs always produce the same store. Reads are queries against that store, not recollection. This is what makes memory auditable rather than merely plausible, and it is why the per-turn path carries no generation at all.

The decoder β€” the fluent one. Never on the request path. Writing and reading memory happen on every turn and stay on the encoder-and-machine side. Generation is for prose a person will read, and for periodic passes that look across many conversations at once to form a persona β€” a pass that can fail without the conversation stopping.


Where this is heading β€” one meaning, several machines

Nothing in this section runs in this release. The model you download is the encoder and the Python in this repository; it does not call a VM and contains no circuit. This section says where the work is going, because it explains choices you can already see in the output contract.

The goal is for the reference step itself to have more than one exact implementation. A decision such as "this mention opens a new chain" or "these two mentions are the same entity" is defined once, as a deterministic map on bounded integers, and then carried out by different machines that must agree with that definition:

machine what it brings
the learned encoder (this model) reads free text; proposes, per input, which mention points where
a deterministic program on a small VM holds what must never be false β€” transitivity, one mention in one chain, capacity limits β€” and is auditable
a fixed-point integer circuit compiled from the same program the same decisions as a neural-style circuit, byte-identical across machines

The division of labour follows one rule that we measured on this model's own task: invariants go to the exact machine, input-dependent judgements stay with the learned part. Enforcing chain invariants on top of the link scores raised the benchmark score, while the same number of random merges lowered it; turning a per-input judgement into a hand-set constant only moved errors between types once it was trained in.

On the circuit side, small closed controllers already exist in both forms β€” four lowered by hand and one family lowered automatically by a compiler. For each one checked so far, the integer circuit produces the same output as the program on every reachable state with zero mismatches, and execution can be switched between the program and the circuit at a tick boundary without changing behaviour. That work is on game-agent controllers, not yet on reference resolution.

Why it is not in serving yet. Each machine added to the request path is another thing that has to agree with the others under load. The order is therefore staged:

  1. stabilise the learned model and the output contract (this release);
  2. move the chain invariants into a deterministic layer shipped with the model, with no VM;
  3. run the same invariants as a program on the VM, checked against step 2 for identical output;
  4. compile that program to an integer circuit and switch between the two at tick boundaries.

A step is taken only after the previous one is measured on the same data, and each step will be reported here when it ships.


Compared with another system

We ran the same public development sets through CorPipe 26 (ufal/corpipe26-onestage-corefud1.4-base-260702, CC BY-NC-SA 4.0 β€” used only as a point of comparison, never as a training signal) and scored both sides on linking: for each gold mention, is the chosen antecedent the previous mention of the same entity, or "none" when it opens a chain? The data is CorefUD 1.4.

The two sides do not see the same input, so there are two rows per language.

mentions this model CorPipe 26 gap (pt)
Korean (ko_ecmt) Β· all gold mentions β€” this model is given the mentions, CorPipe finds its own 10,263 87.8% 68.5% +19.4
Korean (ko_ecmt) Β· β˜… only mentions CorPipe also found β€” linking ability alone 8,176 87.9% 85.9% +1.9
English (en_gum) Β· all gold mentions β€” this model is given the mentions, CorPipe finds its own 4,641 84.6% 66.0% +18.5
English (en_gum) Β· β˜… only mentions CorPipe also found β€” linking ability alone 3,805 84.5% 80.5% +4.0
  • First row β€” the system as deployed. Mentions reach this model from an upstream component, so it is scored on every gold mention. CorPipe has to find the mentions itself; when it misses one, that row counts it wrong. Most of the gap in this row comes from that difference, not from linking.
  • Second row β€” like for like. Only the mentions CorPipe also found, so both sides are judged on choosing the antecedent. This is the comparison of linking ability.
  • It is still not perfectly equal: CorPipe is marked wrong when it found the mention but missed the antecedent mention. Running CorPipe with gold mentions as input would close that, and we have not done it.

What you can check yourself

0.2.6 reproducible here?
Japanese reference resolution † 1.0000 βœ—
CorefUD, 19 languages, macro end-to-end 0.2240 βœ“ public data
input positions this bundle uses 768 βœ“ config.max_len

† JMultiWOZ 1.0 (CC BY-SA 4.0), references mined from that corpus's own slot annotation, words segmented with janome, 113 in-window cases, mentions supplied. Read this row narrowly: every case in this set has an antecedent and only one candidate, so a model that always links also scores 1.0000 here. It measures whether the model links at all on Japanese input, not whether it chooses well. Rows marked βœ“ you can check yourself with the two scripts in this repository.


New knobs (all optional, all defaulting to the old behaviour)

field default what it does
detect_layers 1 BIO layers in the detection head. 1 emits outermost mentions only, as before.
word_joiner " " What goes between words when the surface is rebuilt. Pass "" for languages written without spaces (Japanese, Chinese, Thai) β€” otherwise the encoder is handed a surface it has never seen.
null_bias 1.0 How strongly to discourage the "no antecedent" choice. Larger means the model answers more often. Raise it only on a path where every mention really has an antecedent β€” see the warning below.
null_bias_by_kind {"PRON": 1.0, "DEMNP": 1.0, "NOM": -1.5} Per-mention override of null_bias by mention kind (PRON Β· DEMNP Β· NOM, from mention_kind.py). Remove it to get one global value. Passing null_bias= to a call overrides both.
Japanese reference resolution score
word_joiner="" (as written) 1.0000
word_joiner=" " (words spaced out) 0.4779

Same weights, same inputs, same knob positions otherwise β€” far apart. A bigger encoder would not have saved you from this.

Each can also be passed per call: model.resolve_last(sentences, mentions, tok, joiner="", null_bias=2.0).

β›” Do not set fix_mistral_regex=True

Recent transformers prints a warning when loading this tokenizer, suggesting the flag. We measured it on 5 sentences (ko, ja, en): with the flag, every one of them tokenizes to byte garbage β€” 강남역 becomes Γͺ Β° Δ· Γ« Δ€ Β¨ instead of ▁강 남 μ—­, and sequence length goes from 19 to 71. The model still runs and still returns spans; they are simply wrong. The default path is the correct one, and it is the one every number on this card was measured with.

⚠ null_bias: the number that looks free on the wrong test set

On a benchmark that only ever asks "which earlier mention does this pronoun refer to?", every abstention is wrong by construction, so pushing the model to always answer looks strictly better β€” and keeps looking better the harder you push. On text where "no antecedent" is sometimes the correct answer, the same push is destructive. We measured both:

null_bias dialogue (KLUE-WoS), mentions given β€” correct / silent CorefUD ko, linking with gold mentions
3 0.606 / 6 0.745
2 0.599 / 11 0.803
1 (global default) 0.564 / 28 0.844
0.5 0.544 / 48 0.858
0 0.478 / 93 0.871
-0.5 0.379 / 154 0.878
-1 0.269 / 225 0.880
-1.5 0.200 / 286 0.881
by kind β€” pronoun/demonstrative 1 Β· other noun phrases -1.5 (shipped) 0.564 / 28 0.879

This version ships null_bias by mention kind: pronouns and demonstratives 1, every other mention -1.5 (the kind is read from the mention's own words, see mention_kind.py). That keeps dialogue at the global-default numbers while full noun phrases stop linking where they should open a new chain. A lower value wins on CorefUD, where many mentions correctly open a new chain and silence is cheap β€” but on dialogue, where every pointing word has an antecedent, the same setting turns most answers into silence. Accuracy among answered cases barely moves across the range; what moves is how often the model answers at all. If your pipeline guarantees that the mentions you pass in are anaphoric, a higher value is safe; otherwise leave it alone.


License & access

Released under the Elda Community License 1.0 (see LICENSE). Provenance of the training data is recorded in LINEAGE.md; no user conversations were used.

  • βœ… Research, evaluation, internal validation, education β€” free of charge
  • ✳ Attribution: "Built with Elda"
  • β›” Commercial use and redistribution require a separate agreement

Access is gated: tell us who you are and access is granted automatically.

Citation

@software{memoryresoner2026,
  title  = {memory-resoner: multilingual reference resolution for conversational memory},
  author = {Elda AI},
  year   = {2026},
  url    = {https://huggingface.co/Elda-AI/memory-resoner}
}
Downloads last month
4
Safetensors
Model size
0.3B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support