memory-resoner / README.md
Jinou6502's picture
anaphora + judgement encoder β€” v0.3.0
f1ebcfe verified
|
Raw History Blame Contribute Delete
28.1 kB
---
license: other
license_name: elda-community-license-1.0
license_link: LICENSE
language:
- ko
- en
- ja
library_name: transformers
tags:
- multilingual
- korean
- english
- conversational-memory
- coreference-resolution
- anaphora-resolution
- mention-detection
- entity-linking
- conversational-ai
- real-time
- encoder
extra_gated_heading: Access to memory-resoner
extra_gated_description: >-
Released for research, evaluation, internal validation and education.
Tell us who you are and we will grant access.
extra_gated_prompt: >-
By requesting access you agree to the Elda Community License 1.0:
no commercial use, no redistribution of the weights or derivatives,
and attribution as "Built with Elda". Commercial licensing is available
on request.
extra_gated_fields:
Name: text
Email: text
Affiliation: text
Intended use: text
I agree to the Elda Community License: checkbox
---
# memory-resoner β€” the reference step of conversational memory
![memory-resoner β€” what happens on one turn](https://huggingface.co/datasets/Elda-AI/assets/resolve/main/memory-resoner/overview.png)
![memory-resoner β€” inside one call](https://huggingface.co/datasets/Elda-AI/assets/resolve/main/memory-resoner/inside.png)
A conversation can only be stored if you know what γ€Œκ·Έ 동넀」 meant. **memory-resoner** reads the
turns of a conversation and says, for each mention, **which earlier mention it refers to** β€” so the
layer above can write a memory entry that points at the right words.
```
"합정동에 κ°€κ²Œλ₯Ό μ—΄μ–΄μš” . κ·Έ λ™λ„€λŠ” μœ λ™μΈκ΅¬κ°€ λ§Žλ‚˜μš” ?"
mentions 합정동에 (s0, w0–0)
κ·Έ λ™λ„€λŠ” (s1, w0–1)
chains [합정동에 Β· κ·Έ λ™λ„€λŠ”]
```
`version 0.3.0` Β· **308M parameters** Β· fp32 Β· **32 ms** per sentence on CPU
when the mentions are supplied (122 ms when it finds them itself).
---
---
## Where it sits
memory-resoner is one stage of a conversational stack, and it is **multilingual** β€” trained on
Korean and English, measured on Japanese, and it transfers to languages it never saw. Each stage
answers one question and hands the next stage a **span**, never prose.
```
utterance
β”‚
β”œβ”€ Elda-AI/intenter what was said, and what kind of thing it is ── named spans come from here
β”œβ”€ slot extraction which of those the user said about themselves
β”‚
β–Ό
β˜… memory-resoner β˜… which earlier mention does this one refer to
pronouns, demonstratives and dropped subjects are found here, not upstream
β”‚
β–Ό
memory write a pointer into the user's own words β€” checkable, not generated
```
**Where the spans come from.** In the Elda stack, spans for named things are produced upstream by
[**Elda-AI/intenter**](https://huggingface.co/Elda-AI/intenter) and this model consumes them; its
card is the reference for how a span is drawn, which types exist, and what the channels mean. Used
standalone, memory-resoner will find its own mentions instead.
**Pointing words are the exception, by design.** γ€Œκ±”γ€, γ€Œκ±°κΈ°γ€, γ€Œκ·Έκ±°γ€, *she*, *there*, あそこ are found by
this model itself, and so is an argument a Korean or Japanese speaker simply leaves out β€” subject
or object, which are deliberately not told apart. Linking them
is this stage's job, so finding them here keeps one component responsible when either half goes
wrong. Send the raw turns; do not pre-select pointing words with a word list upstream.
⚠ **Those two paths are not equally good, and the gap is large.** On KLUE-WoS dialogue reference,
this version scores **0.5640** when the mentions are supplied and **0.2319** when it has
to find them itself. Detection and linking are two abilities and the end-to-end number is their product β€”
so if you already have spans upstream, pass them in.
⚠ **The two sides count spans in different units, and the conversion is yours to make.** Upstream
spans are **character offsets with the particle left outside** (γ€Œν•©μ •λ™γ€). This model works in
**μ–΄μ ˆ (word) indices**, and a μ–΄μ ˆ contains its particle, so the same mention is γ€Œν•©μ •λ™μ—γ€ here.
Neither is wrong; they are different units. Align on the μ–΄μ ˆ that contains the upstream span.
**What it returns.** Span indices and chains. Not a rewritten sentence, not a summary β€” an index
into the words you sent, which the layer above can verify before storing anything.
**What it does not decide.** Whether a fact is true, whether it is worth keeping, how long it lives.
Those belong to the system around it. This stage answers one question and stops.
---
---
## Why a small encoder
| | this model |
|---|---|
| Parameters | **308M** |
| Latency per sentence, CPU β€” linking, mentions given | **p50 32 ms** |
| Latency per sentence, CPU β€” finding mentions **and** linking | **p50 122 ms** |
| Output | span index β€” checkable byte-for-byte |
| Determinism | same input, same chains |
Memory is written on **every turn**. A stage that runs that often has to be cheap, and its answer
has to be something the next stage can check rather than trust. A span index is both.
---
---
## Output contract
```
input sentences, in order β€” a string per sentence, or a list of μ–΄μ ˆ
output mentions Β· antecedents Β· chains (links closed transitively)
spans inclusive word (μ–΄μ ˆ) indices β€” Korean particles stay attached (see the note above)
window 6 previous sentences of context Β· 40 previous mentions as candidates
```
A mention with no antecedent opens a chain of its own β€” that is how the next stage learns it has
seen a **new** entity, so singleton chains are kept rather than dropped.
---
---
## Usage
```python
from transformers import AutoModel, AutoTokenizer
m = AutoModel.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True).eval()
tok = AutoTokenizer.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True)
out = m.coref(["합정동에 κ°€κ²Œλ₯Ό μ—΄μ–΄μš” .",
"κ·Έ λ™λ„€λŠ” μœ λ™μΈκ΅¬κ°€ λ§Žλ‚˜μš” ?"], tok)
out["mentions"] # [{'sent': 0, 'words': [0, 0], 'text': '합정동에'}, ...]
out["antecedents"] # [None, 0] ← per mention: what it points back at
out["chains"] # [[0, 1]]
```
| call | does |
|---|---|
| `m.coref(sentences, tok)` | finds the mentions **and** links them |
| `m.mentions(sentences, tok)` | finds mentions only β€” a span provider for another linker |
| `m.resolve(units, mentions, tok)` | links only β€” you supply the mentions |
| `m.coref_last(sentences, tok)` | β˜… **serving shape** β€” only the newest turn's mentions are returned, with the earlier turns kept as context |
| `m.resolve_last(sentences, mentions, tok)` | β˜… **serving shape, mentions supplied** β€” the same, but you hand in the spans |
The two `_last` calls are what a live conversation wants: one turn arrives, you want answers for
**that** turn, and the turns before it are context rather than output. They take the whole
conversation and cut the window themselves, so the caller never has to know how long the window is.
> ### β›” On `coref_last`, do not count with `antecedents`
>
> **The two calls differ here, and it matters.** `resolve_last` returns *all* the mentions you gave
> it, so its `antecedents[i]` indexes that same list and reads normally. `coref_last` returns only
> the **newest turn's** mentions β€” so when the antecedent sits in an earlier turn, which on a live
> conversation is *almost always*, `antecedents[i]` is `None`. There, `None` wears one face for two
> different facts: **"it pointed at something outside this list"** and **"it pointed at nothing"**,
> and a consumer that counts non-`None` gets a structural zero even when the model is working.
>
> | field | what it is | use it for |
> |---|---|---|
> | `linked[i]` | β˜… bool β€” did this mention resolve at all | **counting**, on either call |
> | `n_linked` | sum of the above | **counting** |
> | `antecedent_mentions[i]` | the antecedent itself (`{sent, words, text}`), or `None` | **reading the answer** |
> | `antecedents[i]` | index into the returned list, or `None` | safe on `resolve_last`; on `coref_last` only for within-turn links |
>
> `resolve_last` also reports what it had to drop or fix, and none of it is thrown away silently:
> `mentions_outside_window` Β· `mentions_unencodable` Β· `window_truncated` Β· `mentions_out_of_order`.
> ⚠ When `mentions_out_of_order > 0` the layer re-sorted your mentions into reading order, and then
> **`antecedents[i] > i` can occur** β€” the index is into the list you gave, not into reading order.
> If that count is zero, "an antecedent index is always smaller than its own" holds.
Notes that matter in practice:
* **fp32.** The config pins it. In bf16 near-ties flip and chains change.
* **Pass sentences, not a paragraph.** The sentence boundary is what the window is counted in.
* **Send the resolved span downstream, not the reference.** The output is an index into the words
you sent, so whatever consumes it can verify what it was handed.
* **Strip the particle at display time, not at span time.** Keeping it inside the μ–΄μ ˆ is what makes
the span a plain index into your own input; trimming belongs to whatever renders the value.
---
---
## In a live conversation
A document has one author. A conversation has at least two, and that changes what a pointing word
can mean. This section describes how the reference step is used once it is wired into a chat, and
which parts of that run in this release.
**One table per turn.** Every newest turn produces a small table of links, and everything downstream
is read off that table rather than produced by another model:
| row kind | example | resolved by |
|---|---|---|
| pointing word β†’ earlier mention | γ€Œκ±” μ˜€λŠ˜λ„ 와?」 β†’ γ€Œλ―Όμˆ˜κ°€γ€ | β˜… this model (runs in this release) |
| first / second person β†’ **speaker** or **addressee** | γ€Œλ„ˆ λ‚΄ 이름 μ•Œμ•„?」 β†’ *the assistant*, *the user* | β˜… the speaker layer β€” the turn's speaker, applied over the model's scores (runs in this release) |
| left-out argument β†’ speaker, addressee or an earlier mention | γ€Œ(γ€€) 내일 가도 λΌμš”?」 | β˜… the dropped-argument head on this model's encoder, with the same speaker layer (runs in this release Β· needs predicate spans, see below) |
| nothing to point at | γ€Œκ·Έκ±° 뭐야?」 with no earlier candidate | β˜… `none` β€” an answer, not silence (runs) |
| could not see far enough | the antecedent may lie before the window | β˜… `unknown` β€” the conversation was cut, so downstream should look elsewhere (runs) |
**Why speaker is not learned from text.** γ€Œλ‚˜γ€ and γ€Œλ„ˆγ€ are fixed by who is talking, not by the words
around them, and training on documents never sees a second author. So the turn's speaker is passed
in as a field and used as a constraint on top of the link scores: it rules out links that cannot be
true (one speaker's γ€Œλ‚˜γ€ and the other speaker's γ€Œλ‚˜γ€), it does not invent links on its own. A generic
*you* that addresses nobody in particular stays with the model.
**A self-contained turn is a rendering, not a generation.** A turn such as γ€Œκ±°κΈ° κ°„νŒ μ˜ˆλ»μš”?」 is made
readable on its own by annotating the original words with what they point at β€”
`κ±°κΈ°[=합정동에] κ°„νŒ μ˜ˆλ»μš”?` β€” never by rewriting the sentence. The original stays byte-for-byte, the
annotation is a derived view stored next to it, and a turn that needs no help is left untouched.
Nothing is generated, so nothing can be invented.
**Signals on the same pass.** Reactions that depend on who a turn is aimed at (for example, a remark
addressed to the assistant itself) are read off the same link table. Content signals such as abuse
are separate heads on the same frozen encoder, trained from a larger judge's labels and compared
against it in shadow before they replace anything. The aim is for them to reuse the encoder pass already made for linking rather than run another one.
**What is in this release.** The speaker layer and the dropped-argument head ship with the weights
and code you download, behind one call:
```python
out = m.anaphora_turn({
"turns": [{"turn_index": 0, "speaker": "user", "text": "...", "intent": ...},
{"turn_index": 1, "speaker": "assistant", "text": "...", "intent": ...}],
"now": {"turn_index": 2, "speaker": "user", "text": "...", "intent": ...},
}, tok)
```
* `speaker` is `"user"` or `"assistant"` and is what fixes γ€Œλ‚˜γ€ and γ€Œλ„ˆγ€.
* `intent` is the output of [**Elda-AI/intenter**](https://huggingface.co/Elda-AI/intenter) for that
sentence, passed as it came. The dropped-argument head asks, for each predicate the intent model
found, whether something left unsaid belongs to it and what that is. Subject and object are not
told apart: a predicate may have none, one or several. With `intent: null` there are no
predicates to ask about, so that turn gets pointing-word links and the speaker layer only β€” the
missing rows are absent, not guessed.
* Every offset in the answer is a character offset into that turn's own `text`, the same unit as the
intent model's spans, and an antecedent that matches one of the intent model's entities carries
its index.
* The head sits on the frozen encoder; `coref`, `coref_last` and the other calls above return the
same fields as before.
**What is not in this release.** The turn signals (reactions that depend on who a turn is aimed at,
content heads) are not part of what you download. This card will say so when they ship.
**Conversation scope only.** Links never cross from one conversation into another. Memory here is
what was said in *this* conversation.
---
## Measurements
Public multilingual coreference, held out from training; document overlap with the training split is 0
and the benchmark script asserts it. Both scripts ship in this repository.
⚠ **Korean is one of the languages here, not the only one.** If your pipeline is Korean-only,
measure before you swap: reaching other languages costs some Korean headroom, and the trade is
visible below. The two scripts here let you check that against your own data rather than take our
word for it.
```bash
pip install torch transformers datasets
python bench_corefud_ko.py --model Elda-AI/memory-resoner
```
**Mention detection** β€” boundaries must match exactly
| precision | recall | F1 |
|---:|---:|---:|
| 0.7917 | 0.8087 | **0.8001** |
**Linking, mentions given**
| | |
|---|---:|
| scored mentions | 10263 |
| accuracy | 0.8803 |
| **non-NULL accuracy** | **0.7045** (n=2721) |
| **deictic, non-NULL** | **0.7159** (n=271) |
| reference β€” always NULL | 0.7349 |
| reference β€” always most-recent | 0.0448 |
Read the non-NULL rows. Most mentions open a chain, so answering NULL every time already scores
0.7349; the rows that require actually *linking* are the ones in bold.
**End-to-end** β€” (mention, antecedent) pairs, the model finding its own mentions
| precision | recall | F1 |
|---:|---:|---:|
| 0.6188 | 0.6661 | **0.6415** |
Every run first pushes the gold answers back through the scorer and asserts they survive intact
(1.0000). A scorer that cannot return its own gold is not grading anything.
### On dialogue
The numbers above are prose. On **dialogue** β€” references mined from a public multi-turn Korean
corpus (KLUE-WoS), with the antecedent taken from that corpus's own annotation rather than chosen
by us β€” this version resolves γ€Œκ·Έ 식당」-style references at **0.5640** when the mentions are
supplied, and **0.2319** when it has to find them itself.
```bash
python bench_wos_references.py --model Elda-AI/memory-resoner
```
β˜… **That difference is the whole point of the two paths.** End-to-end is detection **times**
linking, so a pipeline that already has spans should pass them in rather than ask for them again.
⚠ These are measured with candidates mined by the gold rule, which makes them a **perfect-upstream
ceiling** β€” a real intenter will score lower.
These numbers are for conversation; on plain prose the same model reads reference less reliably.
---
---
## What comes next
Conversational memory here is built in three layers. They are **not** future versions of this
model β€” they are different kinds of component, and saying so is the point.
| | what it is for | how it is built |
|---|---|---|
| **Layer 1 Β· conversation memory** | what was said in this conversation, and what refers to what | β˜… **this model** + deterministic assembly |
| **Layer 2** | **blackbox** β€” how a memory is addressed and kept apart from every other conversation | not described here |
| **Layer 3 Β· authority and persona** | whether something holds and on what grounds, and what an agent is like across conversations | a deterministic **VM**, and a **decoder** used off the request path |
Layer 1 is the only one this model sits in. Layer 2 is deliberately not described: it is where
memories are separated from one another, and it is the part of the system we keep closed. Layer 3
is where facts acquire grounds and where a persona is formed, and it is built from two very
different machines β€” one that must be exact, and one that must be fluent.
**The VM β€” the exact one.** Memory is written in a small closed language and executed by a
deterministic machine, so a write is either accepted, a duplicate, a replacement, or refused, and
the same inputs always produce the same store. Reads are queries against that store, not
recollection. This is what makes memory auditable rather than merely plausible, and it is why the
per-turn path carries no generation at all.
**The decoder β€” the fluent one.** Never on the request path. Writing and reading memory happen on
every turn and stay on the encoder-and-machine side. Generation is for prose a person will read,
and for periodic passes that look across many conversations at once to form a persona β€” a pass
that can fail without the conversation stopping.
---
## Where this is heading β€” one meaning, several machines
**Nothing in this section runs in this release.** The model you download is the encoder and the
Python in this repository; it does not call a VM and contains no circuit. This section says where
the work is going, because it explains choices you can already see in the output contract.
The goal is for the reference step itself to have more than one exact implementation. A decision
such as *"this mention opens a new chain"* or *"these two mentions are the same entity"* is defined
once, as a deterministic map on bounded integers, and then carried out by different machines that
must agree with that definition:
| machine | what it brings |
|---|---|
| the learned encoder (this model) | reads free text; proposes, per input, which mention points where |
| a deterministic program on a small VM | holds what must never be false β€” transitivity, one mention in one chain, capacity limits β€” and is auditable |
| a fixed-point integer circuit compiled from the same program | the same decisions as a neural-style circuit, byte-identical across machines |
The division of labour follows one rule that we measured on this model's own task: **invariants go to
the exact machine, input-dependent judgements stay with the learned part.** Enforcing chain
invariants on top of the link scores raised the benchmark score, while the same number of random
merges lowered it; turning a per-input judgement into a hand-set constant only moved errors between
types once it was trained in.
On the circuit side, small closed controllers already exist in both forms β€” four lowered by hand and
one family lowered automatically by a compiler. For each one checked so far, the integer circuit
produces the same output as the program on every reachable state with zero mismatches, and execution can be switched between the program and the
circuit at a tick boundary without changing behaviour. That work is on game-agent controllers, not
yet on reference resolution.
**Why it is not in serving yet.** Each machine added to the request path is another thing that has to
agree with the others under load. The order is therefore staged:
1. stabilise the learned model and the output contract (this release);
2. move the chain invariants into a deterministic layer shipped with the model, with no VM;
3. run the same invariants as a program on the VM, checked against step 2 for identical output;
4. compile that program to an integer circuit and switch between the two at tick boundaries.
A step is taken only after the previous one is measured on the same data, and each step will be
reported here when it ships.
---
## Compared with another system
CorPipe and this model are built for different ends β€” CorPipe is a general-purpose system that reads
whole documents and finds every mention itself, while this model is one stage of a live conversation
that is handed most of its spans and must answer within a single turn. The comparison below is
therefore not a ranking. It is a baseline check: on shared public data, does the reference step hold
up against an established system?
We ran the same public development sets through **CorPipe 26**
(`ufal/corpipe26-onestage-corefud1.4-base-260702`, CC BY-NC-SA 4.0 β€” used only as a point of
comparison, never as a training signal) and scored both sides on **linking**: for each gold mention,
is the chosen antecedent the previous mention of the same entity, or "none" when it opens a chain?
The data is CorefUD 1.4.
**The two sides do not see the same input, so there are two rows per language.**
| | mentions | this model | CorPipe 26 | gap (pt) |
|---|---:|---:|---:|---:|
| Korean (ko_ecmt) Β· all gold mentions β€” *this model is given the mentions, CorPipe finds its own* | 10,263 | **87.8%** | 68.5% | +19.4 |
| Korean (ko_ecmt) Β· β˜… only mentions CorPipe also found β€” *linking ability alone* | 8,176 | **87.9%** | 85.9% | +1.9 |
| English (en_gum) Β· all gold mentions β€” *this model is given the mentions, CorPipe finds its own* | 4,641 | **84.6%** | 66.0% | +18.5 |
| English (en_gum) Β· β˜… only mentions CorPipe also found β€” *linking ability alone* | 3,805 | **84.5%** | 80.5% | +4.0 |
- **First row β€” the system as deployed.** Mentions reach this model from an upstream component, so
it is scored on every gold mention. CorPipe has to find the mentions itself; when it misses one,
that row counts it wrong. Most of the gap in this row comes from that difference, not from linking.
- **Second row β€” like for like.** Only the mentions CorPipe also found, so both sides are judged on
choosing the antecedent. This is the comparison of linking ability.
- It is still not perfectly equal: CorPipe is marked wrong when it found the mention but missed the
*antecedent* mention. Running CorPipe with gold mentions as input would close that, and we have not
done it.
---
## What you can check yourself
| | **0.3.0** | reproducible here? |
|---|---:|---|
| Japanese reference resolution † | **1.0000** | βœ— |
| CorefUD, 19 languages, macro end-to-end | **0.2245** | βœ“ public data |
| input positions this bundle uses | **768** | βœ“ `config.max_len` |
† JMultiWOZ 1.0 (CC BY-SA 4.0), references mined from that corpus's own slot annotation, words
segmented with janome, **113 in-window cases**, mentions supplied. **Read this row narrowly:** every
case in this set has an antecedent and only one candidate, so a model that *always* links also scores
1.0000 here. It measures whether the model links at all on Japanese input, not whether it chooses well.
Rows marked βœ“ you can check yourself with the two scripts in this repository.
---
## New knobs (all optional, all defaulting to the old behaviour)
| field | default | what it does |
|---|---|---|
| `detect_layers` | 1 | BIO layers in the detection head. 1 emits outermost mentions only, as before. |
| `word_joiner` | `" "` | What goes **between words** when the surface is rebuilt. Pass `""` for languages written without spaces (Japanese, Chinese, Thai) β€” otherwise the encoder is handed a surface it has never seen. |
| `null_bias` | 1.0 | How strongly to discourage the "no antecedent" choice. Larger means the model answers more often. **Raise it only on a path where every mention really has an antecedent** β€” see the warning below. |
| `null_bias_by_kind` | `{"PRON": 1.0, "DEMNP": 1.0, "NOM": -1.5}` | Per-mention override of `null_bias` by mention kind (`PRON` Β· `DEMNP` Β· `NOM`, from `mention_kind.py`). Remove it to get one global value. Passing `null_bias=` to a call overrides both. |
| Japanese reference resolution | score |
|---|---:|
| `word_joiner=""` (as written) | **1.0000** |
| `word_joiner=" "` (words spaced out) | **0.7080** |
Same weights, same inputs, same knob positions otherwise β€” **far apart**. A bigger encoder
would not have saved you from this.
Each can also be passed per call: `model.resolve_last(sentences, mentions, tok, joiner="", null_bias=2.0)`.
> ### β›” Do **not** set `fix_mistral_regex=True`
>
> Recent `transformers` prints a warning when loading this tokenizer, suggesting the flag. We
> measured it on 5 sentences (ko, ja, en): **with the flag, every one of them tokenizes to byte
> garbage** β€” `강남역` becomes `Γͺ Β° Δ· Γ« Δ€ Β¨` instead of `▁강 남 μ—­`, and sequence length goes from 19
> to 71. The model still runs and still returns spans; they are simply wrong. The default path is
> the correct one, and it is the one every number on this card was measured with.
> ### ⚠ `null_bias`: the number that looks free on the wrong test set
>
> On a benchmark that only ever asks *"which earlier mention does this pronoun refer to?"*, **every
> abstention is wrong by construction**, so pushing the model to always answer looks strictly better
> β€” and keeps looking better the harder you push. On text where "no antecedent" is sometimes the
> **correct** answer, the same push is destructive. We measured both:
>
> | `null_bias` | dialogue (KLUE-WoS), mentions given β€” correct / silent | CorefUD ko, linking with gold mentions |
> |---:|---:|---:|
> | 3 | 0.606 / 6 | 0.745 |
> | 2 | 0.599 / 11 | 0.803 |
> | 1 (global default) | 0.564 / 28 | 0.844 |
> | 0.5 | 0.544 / 48 | 0.858 |
> | 0 | 0.478 / 93 | 0.871 |
> | -0.5 | 0.379 / 154 | 0.878 |
> | -1 | 0.269 / 225 | 0.880 |
> | -1.5 | 0.200 / 286 | 0.881 |
> | **by kind** β€” pronoun/demonstrative 1 Β· other noun phrases -1.5 (shipped) | 0.564 / 28 | 0.879 |
>
> This version ships `null_bias` **by mention kind**: pronouns and demonstratives **1**, every other mention **-1.5** (the kind is read from the mention's own words, see `mention_kind.py`). That keeps dialogue at the global-default numbers while full noun phrases stop linking where they should open a new chain. A lower value wins on CorefUD, where many mentions correctly open a new
> chain and silence is cheap β€” but on dialogue, where every pointing word has an antecedent, the same
> setting turns most answers into silence. Accuracy **among answered cases** barely moves across the
> range; what moves is how often the model answers at all. If your pipeline guarantees that the mentions
> you pass in are anaphoric, a higher value is safe; otherwise leave it alone.
---
## License & access
Released under the **Elda Community License 1.0** (see `LICENSE`). Provenance of the training data
is recorded in `LINEAGE.md`; no user conversations were used.
* βœ… Research, evaluation, internal validation, education β€” free of charge
* ✳ Attribution: **"Built with Elda"**
* β›” Commercial use and redistribution require a separate agreement
Access is gated: tell us who you are and access is granted automatically.
## Citation
```bibtex
@software{memoryresoner2026,
title = {memory-resoner: multilingual reference resolution for conversational memory},
author = {Elda AI},
year = {2026},
url = {https://huggingface.co/Elda-AI/memory-resoner}
}
```