Feature Extraction
Transformers
Safetensors
Korean
English
Japanese
ko_coref
multilingual
korean
english
conversational-memory
coreference-resolution
anaphora-resolution
mention-detection
entity-linking
conversational-ai
real-time
encoder
custom_code
Instructions to use Elda-AI/memory-resoner with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Elda-AI/memory-resoner with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="Elda-AI/memory-resoner", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from Elda-AI/memory-resoner: direct link, hf CLI and curl.
- Browser
- Download file 28.1 kB
-
https://huggingface.co/Elda-AI/memory-resoner/resolve/main/README.md
- Command line
-
hf download hf://Elda-AI/memory-resoner/README.md
-
curl -L -H "Authorization: Bearer $HF_TOKEN" -o README.md https://huggingface.co/Elda-AI/memory-resoner/resolve/main/README.md
28.1 kB
| license: other | |
| license_name: elda-community-license-1.0 | |
| license_link: LICENSE | |
| language: | |
| - ko | |
| - en | |
| - ja | |
| library_name: transformers | |
| tags: | |
| - multilingual | |
| - korean | |
| - english | |
| - conversational-memory | |
| - coreference-resolution | |
| - anaphora-resolution | |
| - mention-detection | |
| - entity-linking | |
| - conversational-ai | |
| - real-time | |
| - encoder | |
| extra_gated_heading: Access to memory-resoner | |
| extra_gated_description: >- | |
| Released for research, evaluation, internal validation and education. | |
| Tell us who you are and we will grant access. | |
| extra_gated_prompt: >- | |
| By requesting access you agree to the Elda Community License 1.0: | |
| no commercial use, no redistribution of the weights or derivatives, | |
| and attribution as "Built with Elda". Commercial licensing is available | |
| on request. | |
| extra_gated_fields: | |
| Name: text | |
| Email: text | |
| Affiliation: text | |
| Intended use: text | |
| I agree to the Elda Community License: checkbox | |
| # memory-resoner β the reference step of conversational memory | |
|  | |
|  | |
| A conversation can only be stored if you know what γκ·Έ λλ€γ meant. **memory-resoner** reads the | |
| turns of a conversation and says, for each mention, **which earlier mention it refers to** β so the | |
| layer above can write a memory entry that points at the right words. | |
| ``` | |
| "ν©μ λμ κ°κ²λ₯Ό μ΄μ΄μ . κ·Έ λλ€λ μ λμΈκ΅¬κ° λ§λμ ?" | |
| mentions ν©μ λμ (s0, w0β0) | |
| κ·Έ λλ€λ (s1, w0β1) | |
| chains [ν©μ λμ Β· κ·Έ λλ€λ] | |
| ``` | |
| `version 0.3.0` Β· **308M parameters** Β· fp32 Β· **32 ms** per sentence on CPU | |
| when the mentions are supplied (122 ms when it finds them itself). | |
| --- | |
| --- | |
| ## Where it sits | |
| memory-resoner is one stage of a conversational stack, and it is **multilingual** β trained on | |
| Korean and English, measured on Japanese, and it transfers to languages it never saw. Each stage | |
| answers one question and hands the next stage a **span**, never prose. | |
| ``` | |
| utterance | |
| β | |
| ββ Elda-AI/intenter what was said, and what kind of thing it is ββ named spans come from here | |
| ββ slot extraction which of those the user said about themselves | |
| β | |
| βΌ | |
| β memory-resoner β which earlier mention does this one refer to | |
| pronouns, demonstratives and dropped subjects are found here, not upstream | |
| β | |
| βΌ | |
| memory write a pointer into the user's own words β checkable, not generated | |
| ``` | |
| **Where the spans come from.** In the Elda stack, spans for named things are produced upstream by | |
| [**Elda-AI/intenter**](https://huggingface.co/Elda-AI/intenter) and this model consumes them; its | |
| card is the reference for how a span is drawn, which types exist, and what the channels mean. Used | |
| standalone, memory-resoner will find its own mentions instead. | |
| **Pointing words are the exception, by design.** γκ±γ, γκ±°κΈ°γ, γκ·Έκ±°γ, *she*, *there*, γγγ are found by | |
| this model itself, and so is an argument a Korean or Japanese speaker simply leaves out β subject | |
| or object, which are deliberately not told apart. Linking them | |
| is this stage's job, so finding them here keeps one component responsible when either half goes | |
| wrong. Send the raw turns; do not pre-select pointing words with a word list upstream. | |
| β **Those two paths are not equally good, and the gap is large.** On KLUE-WoS dialogue reference, | |
| this version scores **0.5640** when the mentions are supplied and **0.2319** when it has | |
| to find them itself. Detection and linking are two abilities and the end-to-end number is their product β | |
| so if you already have spans upstream, pass them in. | |
| β **The two sides count spans in different units, and the conversion is yours to make.** Upstream | |
| spans are **character offsets with the particle left outside** (γν©μ λγ). This model works in | |
| **μ΄μ (word) indices**, and a μ΄μ contains its particle, so the same mention is γν©μ λμγ here. | |
| Neither is wrong; they are different units. Align on the μ΄μ that contains the upstream span. | |
| **What it returns.** Span indices and chains. Not a rewritten sentence, not a summary β an index | |
| into the words you sent, which the layer above can verify before storing anything. | |
| **What it does not decide.** Whether a fact is true, whether it is worth keeping, how long it lives. | |
| Those belong to the system around it. This stage answers one question and stops. | |
| --- | |
| --- | |
| ## Why a small encoder | |
| | | this model | | |
| |---|---| | |
| | Parameters | **308M** | | |
| | Latency per sentence, CPU β linking, mentions given | **p50 32 ms** | | |
| | Latency per sentence, CPU β finding mentions **and** linking | **p50 122 ms** | | |
| | Output | span index β checkable byte-for-byte | | |
| | Determinism | same input, same chains | | |
| Memory is written on **every turn**. A stage that runs that often has to be cheap, and its answer | |
| has to be something the next stage can check rather than trust. A span index is both. | |
| --- | |
| --- | |
| ## Output contract | |
| ``` | |
| input sentences, in order β a string per sentence, or a list of μ΄μ | |
| output mentions Β· antecedents Β· chains (links closed transitively) | |
| spans inclusive word (μ΄μ ) indices β Korean particles stay attached (see the note above) | |
| window 6 previous sentences of context Β· 40 previous mentions as candidates | |
| ``` | |
| A mention with no antecedent opens a chain of its own β that is how the next stage learns it has | |
| seen a **new** entity, so singleton chains are kept rather than dropped. | |
| --- | |
| --- | |
| ## Usage | |
| ```python | |
| from transformers import AutoModel, AutoTokenizer | |
| m = AutoModel.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True).eval() | |
| tok = AutoTokenizer.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True) | |
| out = m.coref(["ν©μ λμ κ°κ²λ₯Ό μ΄μ΄μ .", | |
| "κ·Έ λλ€λ μ λμΈκ΅¬κ° λ§λμ ?"], tok) | |
| out["mentions"] # [{'sent': 0, 'words': [0, 0], 'text': 'ν©μ λμ'}, ...] | |
| out["antecedents"] # [None, 0] β per mention: what it points back at | |
| out["chains"] # [[0, 1]] | |
| ``` | |
| | call | does | | |
| |---|---| | |
| | `m.coref(sentences, tok)` | finds the mentions **and** links them | | |
| | `m.mentions(sentences, tok)` | finds mentions only β a span provider for another linker | | |
| | `m.resolve(units, mentions, tok)` | links only β you supply the mentions | | |
| | `m.coref_last(sentences, tok)` | β **serving shape** β only the newest turn's mentions are returned, with the earlier turns kept as context | | |
| | `m.resolve_last(sentences, mentions, tok)` | β **serving shape, mentions supplied** β the same, but you hand in the spans | | |
| The two `_last` calls are what a live conversation wants: one turn arrives, you want answers for | |
| **that** turn, and the turns before it are context rather than output. They take the whole | |
| conversation and cut the window themselves, so the caller never has to know how long the window is. | |
| > ### β On `coref_last`, do not count with `antecedents` | |
| > | |
| > **The two calls differ here, and it matters.** `resolve_last` returns *all* the mentions you gave | |
| > it, so its `antecedents[i]` indexes that same list and reads normally. `coref_last` returns only | |
| > the **newest turn's** mentions β so when the antecedent sits in an earlier turn, which on a live | |
| > conversation is *almost always*, `antecedents[i]` is `None`. There, `None` wears one face for two | |
| > different facts: **"it pointed at something outside this list"** and **"it pointed at nothing"**, | |
| > and a consumer that counts non-`None` gets a structural zero even when the model is working. | |
| > | |
| > | field | what it is | use it for | | |
| > |---|---|---| | |
| > | `linked[i]` | β bool β did this mention resolve at all | **counting**, on either call | | |
| > | `n_linked` | sum of the above | **counting** | | |
| > | `antecedent_mentions[i]` | the antecedent itself (`{sent, words, text}`), or `None` | **reading the answer** | | |
| > | `antecedents[i]` | index into the returned list, or `None` | safe on `resolve_last`; on `coref_last` only for within-turn links | | |
| > | |
| > `resolve_last` also reports what it had to drop or fix, and none of it is thrown away silently: | |
| > `mentions_outside_window` Β· `mentions_unencodable` Β· `window_truncated` Β· `mentions_out_of_order`. | |
| > β When `mentions_out_of_order > 0` the layer re-sorted your mentions into reading order, and then | |
| > **`antecedents[i] > i` can occur** β the index is into the list you gave, not into reading order. | |
| > If that count is zero, "an antecedent index is always smaller than its own" holds. | |
| Notes that matter in practice: | |
| * **fp32.** The config pins it. In bf16 near-ties flip and chains change. | |
| * **Pass sentences, not a paragraph.** The sentence boundary is what the window is counted in. | |
| * **Send the resolved span downstream, not the reference.** The output is an index into the words | |
| you sent, so whatever consumes it can verify what it was handed. | |
| * **Strip the particle at display time, not at span time.** Keeping it inside the μ΄μ is what makes | |
| the span a plain index into your own input; trimming belongs to whatever renders the value. | |
| --- | |
| --- | |
| ## In a live conversation | |
| A document has one author. A conversation has at least two, and that changes what a pointing word | |
| can mean. This section describes how the reference step is used once it is wired into a chat, and | |
| which parts of that run in this release. | |
| **One table per turn.** Every newest turn produces a small table of links, and everything downstream | |
| is read off that table rather than produced by another model: | |
| | row kind | example | resolved by | | |
| |---|---|---| | |
| | pointing word β earlier mention | γκ± μ€λλ μ?γ β γλ―Όμκ°γ | β this model (runs in this release) | | |
| | first / second person β **speaker** or **addressee** | γλ λ΄ μ΄λ¦ μμ?γ β *the assistant*, *the user* | β the speaker layer β the turn's speaker, applied over the model's scores (runs in this release) | | |
| | left-out argument β speaker, addressee or an earlier mention | γ(γ) λ΄μΌ κ°λ λΌμ?γ | β the dropped-argument head on this model's encoder, with the same speaker layer (runs in this release Β· needs predicate spans, see below) | | |
| | nothing to point at | γκ·Έκ±° λμΌ?γ with no earlier candidate | β `none` β an answer, not silence (runs) | | |
| | could not see far enough | the antecedent may lie before the window | β `unknown` β the conversation was cut, so downstream should look elsewhere (runs) | | |
| **Why speaker is not learned from text.** γλγ and γλγ are fixed by who is talking, not by the words | |
| around them, and training on documents never sees a second author. So the turn's speaker is passed | |
| in as a field and used as a constraint on top of the link scores: it rules out links that cannot be | |
| true (one speaker's γλγ and the other speaker's γλγ), it does not invent links on its own. A generic | |
| *you* that addresses nobody in particular stays with the model. | |
| **A self-contained turn is a rendering, not a generation.** A turn such as γκ±°κΈ° κ°ν μλ»μ?γ is made | |
| readable on its own by annotating the original words with what they point at β | |
| `κ±°κΈ°[=ν©μ λμ] κ°ν μλ»μ?` β never by rewriting the sentence. The original stays byte-for-byte, the | |
| annotation is a derived view stored next to it, and a turn that needs no help is left untouched. | |
| Nothing is generated, so nothing can be invented. | |
| **Signals on the same pass.** Reactions that depend on who a turn is aimed at (for example, a remark | |
| addressed to the assistant itself) are read off the same link table. Content signals such as abuse | |
| are separate heads on the same frozen encoder, trained from a larger judge's labels and compared | |
| against it in shadow before they replace anything. The aim is for them to reuse the encoder pass already made for linking rather than run another one. | |
| **What is in this release.** The speaker layer and the dropped-argument head ship with the weights | |
| and code you download, behind one call: | |
| ```python | |
| out = m.anaphora_turn({ | |
| "turns": [{"turn_index": 0, "speaker": "user", "text": "...", "intent": ...}, | |
| {"turn_index": 1, "speaker": "assistant", "text": "...", "intent": ...}], | |
| "now": {"turn_index": 2, "speaker": "user", "text": "...", "intent": ...}, | |
| }, tok) | |
| ``` | |
| * `speaker` is `"user"` or `"assistant"` and is what fixes γλγ and γλγ. | |
| * `intent` is the output of [**Elda-AI/intenter**](https://huggingface.co/Elda-AI/intenter) for that | |
| sentence, passed as it came. The dropped-argument head asks, for each predicate the intent model | |
| found, whether something left unsaid belongs to it and what that is. Subject and object are not | |
| told apart: a predicate may have none, one or several. With `intent: null` there are no | |
| predicates to ask about, so that turn gets pointing-word links and the speaker layer only β the | |
| missing rows are absent, not guessed. | |
| * Every offset in the answer is a character offset into that turn's own `text`, the same unit as the | |
| intent model's spans, and an antecedent that matches one of the intent model's entities carries | |
| its index. | |
| * The head sits on the frozen encoder; `coref`, `coref_last` and the other calls above return the | |
| same fields as before. | |
| **What is not in this release.** The turn signals (reactions that depend on who a turn is aimed at, | |
| content heads) are not part of what you download. This card will say so when they ship. | |
| **Conversation scope only.** Links never cross from one conversation into another. Memory here is | |
| what was said in *this* conversation. | |
| --- | |
| ## Measurements | |
| Public multilingual coreference, held out from training; document overlap with the training split is 0 | |
| and the benchmark script asserts it. Both scripts ship in this repository. | |
| β **Korean is one of the languages here, not the only one.** If your pipeline is Korean-only, | |
| measure before you swap: reaching other languages costs some Korean headroom, and the trade is | |
| visible below. The two scripts here let you check that against your own data rather than take our | |
| word for it. | |
| ```bash | |
| pip install torch transformers datasets | |
| python bench_corefud_ko.py --model Elda-AI/memory-resoner | |
| ``` | |
| **Mention detection** β boundaries must match exactly | |
| | precision | recall | F1 | | |
| |---:|---:|---:| | |
| | 0.7917 | 0.8087 | **0.8001** | | |
| **Linking, mentions given** | |
| | | | | |
| |---|---:| | |
| | scored mentions | 10263 | | |
| | accuracy | 0.8803 | | |
| | **non-NULL accuracy** | **0.7045** (n=2721) | | |
| | **deictic, non-NULL** | **0.7159** (n=271) | | |
| | reference β always NULL | 0.7349 | | |
| | reference β always most-recent | 0.0448 | | |
| Read the non-NULL rows. Most mentions open a chain, so answering NULL every time already scores | |
| 0.7349; the rows that require actually *linking* are the ones in bold. | |
| **End-to-end** β (mention, antecedent) pairs, the model finding its own mentions | |
| | precision | recall | F1 | | |
| |---:|---:|---:| | |
| | 0.6188 | 0.6661 | **0.6415** | | |
| Every run first pushes the gold answers back through the scorer and asserts they survive intact | |
| (1.0000). A scorer that cannot return its own gold is not grading anything. | |
| ### On dialogue | |
| The numbers above are prose. On **dialogue** β references mined from a public multi-turn Korean | |
| corpus (KLUE-WoS), with the antecedent taken from that corpus's own annotation rather than chosen | |
| by us β this version resolves γκ·Έ μλΉγ-style references at **0.5640** when the mentions are | |
| supplied, and **0.2319** when it has to find them itself. | |
| ```bash | |
| python bench_wos_references.py --model Elda-AI/memory-resoner | |
| ``` | |
| β **That difference is the whole point of the two paths.** End-to-end is detection **times** | |
| linking, so a pipeline that already has spans should pass them in rather than ask for them again. | |
| β These are measured with candidates mined by the gold rule, which makes them a **perfect-upstream | |
| ceiling** β a real intenter will score lower. | |
| These numbers are for conversation; on plain prose the same model reads reference less reliably. | |
| --- | |
| --- | |
| ## What comes next | |
| Conversational memory here is built in three layers. They are **not** future versions of this | |
| model β they are different kinds of component, and saying so is the point. | |
| | | what it is for | how it is built | | |
| |---|---|---| | |
| | **Layer 1 Β· conversation memory** | what was said in this conversation, and what refers to what | β **this model** + deterministic assembly | | |
| | **Layer 2** | **blackbox** β how a memory is addressed and kept apart from every other conversation | not described here | | |
| | **Layer 3 Β· authority and persona** | whether something holds and on what grounds, and what an agent is like across conversations | a deterministic **VM**, and a **decoder** used off the request path | | |
| Layer 1 is the only one this model sits in. Layer 2 is deliberately not described: it is where | |
| memories are separated from one another, and it is the part of the system we keep closed. Layer 3 | |
| is where facts acquire grounds and where a persona is formed, and it is built from two very | |
| different machines β one that must be exact, and one that must be fluent. | |
| **The VM β the exact one.** Memory is written in a small closed language and executed by a | |
| deterministic machine, so a write is either accepted, a duplicate, a replacement, or refused, and | |
| the same inputs always produce the same store. Reads are queries against that store, not | |
| recollection. This is what makes memory auditable rather than merely plausible, and it is why the | |
| per-turn path carries no generation at all. | |
| **The decoder β the fluent one.** Never on the request path. Writing and reading memory happen on | |
| every turn and stay on the encoder-and-machine side. Generation is for prose a person will read, | |
| and for periodic passes that look across many conversations at once to form a persona β a pass | |
| that can fail without the conversation stopping. | |
| --- | |
| ## Where this is heading β one meaning, several machines | |
| **Nothing in this section runs in this release.** The model you download is the encoder and the | |
| Python in this repository; it does not call a VM and contains no circuit. This section says where | |
| the work is going, because it explains choices you can already see in the output contract. | |
| The goal is for the reference step itself to have more than one exact implementation. A decision | |
| such as *"this mention opens a new chain"* or *"these two mentions are the same entity"* is defined | |
| once, as a deterministic map on bounded integers, and then carried out by different machines that | |
| must agree with that definition: | |
| | machine | what it brings | | |
| |---|---| | |
| | the learned encoder (this model) | reads free text; proposes, per input, which mention points where | | |
| | a deterministic program on a small VM | holds what must never be false β transitivity, one mention in one chain, capacity limits β and is auditable | | |
| | a fixed-point integer circuit compiled from the same program | the same decisions as a neural-style circuit, byte-identical across machines | | |
| The division of labour follows one rule that we measured on this model's own task: **invariants go to | |
| the exact machine, input-dependent judgements stay with the learned part.** Enforcing chain | |
| invariants on top of the link scores raised the benchmark score, while the same number of random | |
| merges lowered it; turning a per-input judgement into a hand-set constant only moved errors between | |
| types once it was trained in. | |
| On the circuit side, small closed controllers already exist in both forms β four lowered by hand and | |
| one family lowered automatically by a compiler. For each one checked so far, the integer circuit | |
| produces the same output as the program on every reachable state with zero mismatches, and execution can be switched between the program and the | |
| circuit at a tick boundary without changing behaviour. That work is on game-agent controllers, not | |
| yet on reference resolution. | |
| **Why it is not in serving yet.** Each machine added to the request path is another thing that has to | |
| agree with the others under load. The order is therefore staged: | |
| 1. stabilise the learned model and the output contract (this release); | |
| 2. move the chain invariants into a deterministic layer shipped with the model, with no VM; | |
| 3. run the same invariants as a program on the VM, checked against step 2 for identical output; | |
| 4. compile that program to an integer circuit and switch between the two at tick boundaries. | |
| A step is taken only after the previous one is measured on the same data, and each step will be | |
| reported here when it ships. | |
| --- | |
| ## Compared with another system | |
| CorPipe and this model are built for different ends β CorPipe is a general-purpose system that reads | |
| whole documents and finds every mention itself, while this model is one stage of a live conversation | |
| that is handed most of its spans and must answer within a single turn. The comparison below is | |
| therefore not a ranking. It is a baseline check: on shared public data, does the reference step hold | |
| up against an established system? | |
| We ran the same public development sets through **CorPipe 26** | |
| (`ufal/corpipe26-onestage-corefud1.4-base-260702`, CC BY-NC-SA 4.0 β used only as a point of | |
| comparison, never as a training signal) and scored both sides on **linking**: for each gold mention, | |
| is the chosen antecedent the previous mention of the same entity, or "none" when it opens a chain? | |
| The data is CorefUD 1.4. | |
| **The two sides do not see the same input, so there are two rows per language.** | |
| | | mentions | this model | CorPipe 26 | gap (pt) | | |
| |---|---:|---:|---:|---:| | |
| | Korean (ko_ecmt) Β· all gold mentions β *this model is given the mentions, CorPipe finds its own* | 10,263 | **87.8%** | 68.5% | +19.4 | | |
| | Korean (ko_ecmt) Β· β only mentions CorPipe also found β *linking ability alone* | 8,176 | **87.9%** | 85.9% | +1.9 | | |
| | English (en_gum) Β· all gold mentions β *this model is given the mentions, CorPipe finds its own* | 4,641 | **84.6%** | 66.0% | +18.5 | | |
| | English (en_gum) Β· β only mentions CorPipe also found β *linking ability alone* | 3,805 | **84.5%** | 80.5% | +4.0 | | |
| - **First row β the system as deployed.** Mentions reach this model from an upstream component, so | |
| it is scored on every gold mention. CorPipe has to find the mentions itself; when it misses one, | |
| that row counts it wrong. Most of the gap in this row comes from that difference, not from linking. | |
| - **Second row β like for like.** Only the mentions CorPipe also found, so both sides are judged on | |
| choosing the antecedent. This is the comparison of linking ability. | |
| - It is still not perfectly equal: CorPipe is marked wrong when it found the mention but missed the | |
| *antecedent* mention. Running CorPipe with gold mentions as input would close that, and we have not | |
| done it. | |
| --- | |
| ## What you can check yourself | |
| | | **0.3.0** | reproducible here? | | |
| |---|---:|---| | |
| | Japanese reference resolution β | **1.0000** | β | | |
| | CorefUD, 19 languages, macro end-to-end | **0.2245** | β public data | | |
| | input positions this bundle uses | **768** | β `config.max_len` | | |
| β JMultiWOZ 1.0 (CC BY-SA 4.0), references mined from that corpus's own slot annotation, words | |
| segmented with janome, **113 in-window cases**, mentions supplied. **Read this row narrowly:** every | |
| case in this set has an antecedent and only one candidate, so a model that *always* links also scores | |
| 1.0000 here. It measures whether the model links at all on Japanese input, not whether it chooses well. | |
| Rows marked β you can check yourself with the two scripts in this repository. | |
| --- | |
| ## New knobs (all optional, all defaulting to the old behaviour) | |
| | field | default | what it does | | |
| |---|---|---| | |
| | `detect_layers` | 1 | BIO layers in the detection head. 1 emits outermost mentions only, as before. | | |
| | `word_joiner` | `" "` | What goes **between words** when the surface is rebuilt. Pass `""` for languages written without spaces (Japanese, Chinese, Thai) β otherwise the encoder is handed a surface it has never seen. | | |
| | `null_bias` | 1.0 | How strongly to discourage the "no antecedent" choice. Larger means the model answers more often. **Raise it only on a path where every mention really has an antecedent** β see the warning below. | | |
| | `null_bias_by_kind` | `{"PRON": 1.0, "DEMNP": 1.0, "NOM": -1.5}` | Per-mention override of `null_bias` by mention kind (`PRON` Β· `DEMNP` Β· `NOM`, from `mention_kind.py`). Remove it to get one global value. Passing `null_bias=` to a call overrides both. | | |
| | Japanese reference resolution | score | | |
| |---|---:| | |
| | `word_joiner=""` (as written) | **1.0000** | | |
| | `word_joiner=" "` (words spaced out) | **0.7080** | | |
| Same weights, same inputs, same knob positions otherwise β **far apart**. A bigger encoder | |
| would not have saved you from this. | |
| Each can also be passed per call: `model.resolve_last(sentences, mentions, tok, joiner="", null_bias=2.0)`. | |
| > ### β Do **not** set `fix_mistral_regex=True` | |
| > | |
| > Recent `transformers` prints a warning when loading this tokenizer, suggesting the flag. We | |
| > measured it on 5 sentences (ko, ja, en): **with the flag, every one of them tokenizes to byte | |
| > garbage** β `κ°λ¨μ` becomes `Γͺ Β° Δ· Γ« Δ€ Β¨` instead of `βκ° λ¨ μ`, and sequence length goes from 19 | |
| > to 71. The model still runs and still returns spans; they are simply wrong. The default path is | |
| > the correct one, and it is the one every number on this card was measured with. | |
| > ### β `null_bias`: the number that looks free on the wrong test set | |
| > | |
| > On a benchmark that only ever asks *"which earlier mention does this pronoun refer to?"*, **every | |
| > abstention is wrong by construction**, so pushing the model to always answer looks strictly better | |
| > β and keeps looking better the harder you push. On text where "no antecedent" is sometimes the | |
| > **correct** answer, the same push is destructive. We measured both: | |
| > | |
| > | `null_bias` | dialogue (KLUE-WoS), mentions given β correct / silent | CorefUD ko, linking with gold mentions | | |
| > |---:|---:|---:| | |
| > | 3 | 0.606 / 6 | 0.745 | | |
| > | 2 | 0.599 / 11 | 0.803 | | |
| > | 1 (global default) | 0.564 / 28 | 0.844 | | |
| > | 0.5 | 0.544 / 48 | 0.858 | | |
| > | 0 | 0.478 / 93 | 0.871 | | |
| > | -0.5 | 0.379 / 154 | 0.878 | | |
| > | -1 | 0.269 / 225 | 0.880 | | |
| > | -1.5 | 0.200 / 286 | 0.881 | | |
| > | **by kind** β pronoun/demonstrative 1 Β· other noun phrases -1.5 (shipped) | 0.564 / 28 | 0.879 | | |
| > | |
| > This version ships `null_bias` **by mention kind**: pronouns and demonstratives **1**, every other mention **-1.5** (the kind is read from the mention's own words, see `mention_kind.py`). That keeps dialogue at the global-default numbers while full noun phrases stop linking where they should open a new chain. A lower value wins on CorefUD, where many mentions correctly open a new | |
| > chain and silence is cheap β but on dialogue, where every pointing word has an antecedent, the same | |
| > setting turns most answers into silence. Accuracy **among answered cases** barely moves across the | |
| > range; what moves is how often the model answers at all. If your pipeline guarantees that the mentions | |
| > you pass in are anaphoric, a higher value is safe; otherwise leave it alone. | |
| --- | |
| ## License & access | |
| Released under the **Elda Community License 1.0** (see `LICENSE`). Provenance of the training data | |
| is recorded in `LINEAGE.md`; no user conversations were used. | |
| * β Research, evaluation, internal validation, education β free of charge | |
| * β³ Attribution: **"Built with Elda"** | |
| * β Commercial use and redistribution require a separate agreement | |
| Access is gated: tell us who you are and access is granted automatically. | |
| ## Citation | |
| ```bibtex | |
| @software{memoryresoner2026, | |
| title = {memory-resoner: multilingual reference resolution for conversational memory}, | |
| author = {Elda AI}, | |
| year = {2026}, | |
| url = {https://huggingface.co/Elda-AI/memory-resoner} | |
| } | |
| ``` | |