YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

EvoRM: Evolvable Neuro-Symbolic Reasoning Framework

License

EvoRM is an evolvable neuro-symbolic reasoning framework for entity matching and schema matching. It combines the reasoning capabilities of LLMs with the efficiency of symbolic rules, achieving high accuracy with reduced LLM calls.

Architecture

EvoRM consists of four core components:

  1. RuleEncoding โ€” Distills LLM decisions into structured FOL (First-Order Logic) rules
  2. HypergraphStorage โ€” Weighted hypergraph for rules + entity context indexing
  3. TwoStageInferenceController โ€” Stage 1: symbolic filtering (fast), Stage 2: LLM judgment (accurate)
  4. RuleMaintenance โ€” Dynamic confidence tracking + periodic rule optimization

Plus an optional MLP Gate (gฯ•) for neural routing between stages.

Key Features

  • Plug-and-play: Integrates with any LLM-based entity matching baseline
  • Two-stage inference: Rules route high-confidence decisions, saving LLM calls
  • Continual learning: Rules evolve with each inference, building a knowledge base
  • Confidence boost: Rules gain confidence as they are triggered correctly

Installation

pip install -r requirements.txt

Quick Start

from evorm_core import EvoRMPlugin

# Initialize EvoRM plugin
plugin = EvoRMPlugin(
    client=openai_client,
    persistence_dir="./evorm_state",
)

# Use with your entity matching pipeline
decision, triggered, candidates = plugin.stage1(eid_a, eid_b, es_ctx, et_ctx)
if decision is not None:
    # Stage 1 routed (high confidence)
    pass
else:
    # Stage 2: LLM judgment
    result = plugin.stage2_simple(eid_a, eid_b, es_ctx, et_ctx, triggered, prompt)
    decision = result['decision']

Supported Baselines

EvoRM provides wrappers for 6 baseline methods:

Baseline Task Wrapper
MatchGPT Entity Resolution baselines/evorm_wrappers/matchgpt.py
Anymatch Entity Resolution baselines/evorm_wrappers/anymatch.py
LELA Entity Linking baselines/evorm_wrappers/lela_el.py
ChatEA Entity Alignment baselines/evorm_wrappers/chat_ea.py
Matchmaker Schema Matching baselines/evorm_wrappers/cohard.py
AdaCoAgentEA Entity Alignment baselines/evorm_wrappers/zero_cot.py

Datasets

The framework supports 14 datasets across 4 tasks:

  • Entity Resolution (ER): Amazon-Google, Beer, DBLP-ACM, DBLP-GoogleScholar, Fodors-Zagats, iTunes-Amazon, Walmart-Amazon, Abt-Buy
  • Entity Linking (EL): ZESHEL (FR, Lego, ST, YG)
  • Schema Matching (SM): MIMIC, SYNTHEA
  • Entity Alignment (EA): DBP-WIKI-V1, DBP-WIKI-V2

Results (V5)

Entity Resolution (MatchGPT)

Dataset Baseline F1 +EvoRM F1 ฮ” S1 Rate
DBLP-ACM 1.0000 1.0000 0.00 16.7%
DBLP-GoogleScholar 0.9474 0.9474 0.00 13.3%
Walmart-Amazon 0.9714 0.9714 0.00 13.3%
Amazon-Google 0.9756 0.9756 0.00 0%

Schema Matching

Dataset Baseline Baseline F1 +EvoRM F1 S1 Rate
MIMIC matchmaker 0.9697 1.0000 40%
SYNTHEA matchmaker 0.9697 1.0000 80%
SYNTHEA rematch 0.9697 1.0000 66.7%

Configuration

from evorm_core import EvoRMConfig

config = EvoRMConfig(
    theta_hi=0.65,        # High-confidence match threshold
    theta_prune=0.75,     # Conservative non-match threshold
    route_nonmatch=True,  # Allow S1 to route non-match decisions
    alpha=0.6,            # Rule encoding weight
    beta=0.4,             # Context similarity weight
)

Citation

@article{evorm2024,
  title={EvoRM: Evolvable Neuro-Symbolic Reasoning for Entity Matching},
  journal={IEEE TKDE},
  year={2024}
}

License

MIT License

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support