You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

PV Duplication Embedder β€” v2

Hybrid embedding model that detects duplicate pharmacovigilance adverse-event reports by jointly encoding free-text narratives and structured metadata (dates, MedDRA/VeDDRA terms, reporter info, product).

Architecture

Component Details
Text encoder joe32140/ModernBERT-large-msmarco (frozen, 1024-dim, referenced not bundled)
Metadata encoder Fitted sklearn pipeline β†’ 768-dim dense vector, bottleneck autoencoder
Fusion MLPEncoder β€” 1792-dim concatenation projected to 1024-dim output

What changed since v1

v1 v2
Fusion MultiGatedFusionEncoder, per-dimension gate MLPEncoder
Metadata 512-dim, TruncatedSVD 768-dim, bottleneck autoencoder
Pickle by reference β€” needed a tree of stub modules beside it by value β€” self-contained
Aggregator named in the class body, so a new one meant a new wrapper read from the sidecar the fit wrote

The gate was replaced because it lost. On the closed-set test corpus the MLP reaches 0.9650 hit@20 against the gate's 0.8505, and under corrupted metadata it stays 21 points above narrative-only retrieval where the gate falls below it.

Artifacts in this repo

File Description
config.json Architecture config, text model reference, and component provenance
metadata_pipeline.pkl Fitted sklearn pipeline (cloudpickle, by value β€” no stub modules needed)
aggregator.pt Fusion encoder weights (inference-only, no training state)
pv_duplication_embedder.py Model entry point
model/aggregator/encoders.py Pure nn.Module architectures, reusable across projects
projects/pv_duplication/specs.py The metadata fields this pipeline consumes

model/ holds code that is not specific to this project; projects/ holds what is. A second project reusing these architectures adds a directory rather than a fork.

Usage

Each input record must be a dict with exactly two keys:

Key Type Description
text str Free-text adverse-event narrative
metadata dict Raw metadata fields consumed by the sklearn pipeline. The exact list is in projects/pv_duplication/specs.py; any field may be omitted and encodes as unknown
from huggingface_hub import snapshot_download
import sys

# `revision` pins the release. Without it you track `main`, which moves.
snapshot_dir = snapshot_download(
    "Ennov/pv_ae_document_duplication_embed", revision="v2"
)
sys.path.insert(0, snapshot_dir)

from pv_duplication_embedder import PVDuplicateEmbedder

model = PVDuplicateEmbedder.from_pretrained(snapshot_dir)

records = [
    {
        "text": "A 3-year-old Labrador received product X. Vomiting observed after 2 hours.",
        "metadata": {
            "recd_date": "2023-06-01",
            "pt_name": ["Vomiting"],
            "reporter_role": ["Veterinarian"],
            "prod_code": ["PROD123"],
            # ... see projects/pv_duplication/specs.py for the full field list
        },
    }
]

# Returns np.ndarray of shape (N, 1024), float32
embeddings = model.encode(records, batch_size=32)

from_pretrained also accepts a repo id directly, downloading on demand:

model = PVDuplicateEmbedder.from_pretrained(
    "Ennov/pv_ae_document_duplication_embed", revision="v2"
)

Versions

Releases are pinned with git tags, so a revision keeps working after the next push.

Tag Contents
v1 Gated fusion, 512-dim metadata, by-reference pickle. Different entry point β€” see that revision's own card
v2 This release

Authentication

Set your HuggingFace token before pushing:

# Option A β€” environment variable (CI/CD friendly)
export HF_TOKEN="hf_..."

# Option B β€” interactive login
huggingface-cli login

Then assemble, verify and publish β€” three commands, because publishing is the one step that is expensive to take back:

train save --wrapper v2
train load --wrapper v2          # refuses a bundle it cannot load
train push --wrapper v2 --tag v2

Training

The metadata vectorizer is fitted on the train split of a 70/30 mixture of clean and metadata-damaged reports. The fusion encoder is trained on cached component vectors with a triplet objective over annotated duplicate clusters, never on documents directly.

Component-level provenance β€” which registry version or run each part came from β€” is recorded in config.json.

Expected metadata structure

Every key is optional β€” a field that is absent, null or unrecognised encodes as unknown rather than raising, which is what makes the vectorizer resilient. Supplying a field the pipeline does not read is ignored.

Group Shape Fields
VeDDRA codes list of int llt_code, pt_code, hlt_code, soc_code
VeDDRA terms list of str llt_name, pt_name, hlt_name, soc_name
Dates date, or ISO-8601 str; sign_start_date a list of them recd_date, orig_recd_date, first_recd_date, reaction_start_date, sign_start_date, patient_dead_date
Categoricals str, except serious which is bool user_case_type, patient_sex, patient_species, veddra_class, serious
Numerics float patient_age, patient_weight, patient_exposed_number
Reporter list of str reporter_firstname, reporter_surname, reporter_country, reporter_city, reporter_institute, reporter_role, reporter_zip

The same list is importable from the bundle as projects.pv_duplication.specs.FIELD_GROUPS.

These are accepted but not encoded: patient_breed, patient_reacted_number, patient_dead_number, prod_code, prod_report, prod_country. They sit in the fitted frame yet reach no branch of the vectorizer β€” blanking every one of them leaves the output bit-identical β€” so there is nothing to gain by supplying them.

Model structure

The fitted metadata pipeline, as the object actually is β€” every branch, transformer and the parameters that decide what it does.

Fitted metadata pipeline

Identical parallel branches are drawn once with a multiplier; sequential steps are always shown in full.

metadata_pipeline.html is scikit-learn's own interactive version β€” download it and open it in a browser.

Downloads last month
50
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support