YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
EvoRM: Evolvable Neuro-Symbolic Reasoning Framework
EvoRM is an evolvable neuro-symbolic reasoning framework for entity matching and schema matching. It combines the reasoning capabilities of LLMs with the efficiency of symbolic rules, achieving high accuracy with reduced LLM calls.
Architecture
EvoRM consists of four core components:
- RuleEncoding โ Distills LLM decisions into structured FOL (First-Order Logic) rules
- HypergraphStorage โ Weighted hypergraph for rules + entity context indexing
- TwoStageInferenceController โ Stage 1: symbolic filtering (fast), Stage 2: LLM judgment (accurate)
- RuleMaintenance โ Dynamic confidence tracking + periodic rule optimization
Plus an optional MLP Gate (gฯ) for neural routing between stages.
Key Features
- Plug-and-play: Integrates with any LLM-based entity matching baseline
- Two-stage inference: Rules route high-confidence decisions, saving LLM calls
- Continual learning: Rules evolve with each inference, building a knowledge base
- Confidence boost: Rules gain confidence as they are triggered correctly
Installation
pip install -r requirements.txt
Quick Start
from evorm_core import EvoRMPlugin
# Initialize EvoRM plugin
plugin = EvoRMPlugin(
client=openai_client,
persistence_dir="./evorm_state",
)
# Use with your entity matching pipeline
decision, triggered, candidates = plugin.stage1(eid_a, eid_b, es_ctx, et_ctx)
if decision is not None:
# Stage 1 routed (high confidence)
pass
else:
# Stage 2: LLM judgment
result = plugin.stage2_simple(eid_a, eid_b, es_ctx, et_ctx, triggered, prompt)
decision = result['decision']
Supported Baselines
EvoRM provides wrappers for 6 baseline methods:
| Baseline | Task | Wrapper |
|---|---|---|
| MatchGPT | Entity Resolution | baselines/evorm_wrappers/matchgpt.py |
| Anymatch | Entity Resolution | baselines/evorm_wrappers/anymatch.py |
| LELA | Entity Linking | baselines/evorm_wrappers/lela_el.py |
| ChatEA | Entity Alignment | baselines/evorm_wrappers/chat_ea.py |
| Matchmaker | Schema Matching | baselines/evorm_wrappers/cohard.py |
| AdaCoAgentEA | Entity Alignment | baselines/evorm_wrappers/zero_cot.py |
Datasets
The framework supports 14 datasets across 4 tasks:
- Entity Resolution (ER): Amazon-Google, Beer, DBLP-ACM, DBLP-GoogleScholar, Fodors-Zagats, iTunes-Amazon, Walmart-Amazon, Abt-Buy
- Entity Linking (EL): ZESHEL (FR, Lego, ST, YG)
- Schema Matching (SM): MIMIC, SYNTHEA
- Entity Alignment (EA): DBP-WIKI-V1, DBP-WIKI-V2
Results (V5)
Entity Resolution (MatchGPT)
| Dataset | Baseline F1 | +EvoRM F1 | ฮ | S1 Rate |
|---|---|---|---|---|
| DBLP-ACM | 1.0000 | 1.0000 | 0.00 | 16.7% |
| DBLP-GoogleScholar | 0.9474 | 0.9474 | 0.00 | 13.3% |
| Walmart-Amazon | 0.9714 | 0.9714 | 0.00 | 13.3% |
| Amazon-Google | 0.9756 | 0.9756 | 0.00 | 0% |
Schema Matching
| Dataset | Baseline | Baseline F1 | +EvoRM F1 | S1 Rate |
|---|---|---|---|---|
| MIMIC | matchmaker | 0.9697 | 1.0000 | 40% |
| SYNTHEA | matchmaker | 0.9697 | 1.0000 | 80% |
| SYNTHEA | rematch | 0.9697 | 1.0000 | 66.7% |
Configuration
from evorm_core import EvoRMConfig
config = EvoRMConfig(
theta_hi=0.65, # High-confidence match threshold
theta_prune=0.75, # Conservative non-match threshold
route_nonmatch=True, # Allow S1 to route non-match decisions
alpha=0.6, # Rule encoding weight
beta=0.4, # Context similarity weight
)
Citation
@article{evorm2024,
title={EvoRM: Evolvable Neuro-Symbolic Reasoning for Entity Matching},
journal={IEEE TKDE},
year={2024}
}
License
MIT License
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support