PV Duplication Embedder β v2
Hybrid embedding model that detects duplicate pharmacovigilance adverse-event reports by jointly encoding free-text narratives and structured metadata (dates, MedDRA/VeDDRA terms, reporter info, product).
Architecture
| Component | Details |
|---|---|
| Text encoder | joe32140/ModernBERT-large-msmarco (frozen, 1024-dim, referenced not bundled) |
| Metadata encoder | Fitted sklearn pipeline β 768-dim dense vector, bottleneck autoencoder |
| Fusion | MLPEncoder β 1792-dim concatenation projected to 1024-dim output |
What changed since v1
| v1 | v2 | |
|---|---|---|
| Fusion | MultiGatedFusionEncoder, per-dimension gate |
MLPEncoder |
| Metadata | 512-dim, TruncatedSVD |
768-dim, bottleneck autoencoder |
| Pickle | by reference β needed a tree of stub modules beside it | by value β self-contained |
| Aggregator | named in the class body, so a new one meant a new wrapper | read from the sidecar the fit wrote |
The gate was replaced because it lost. On the closed-set test corpus the MLP reaches 0.9650 hit@20 against the gate's 0.8505, and under corrupted metadata it stays 21 points above narrative-only retrieval where the gate falls below it.
Artifacts in this repo
| File | Description |
|---|---|
config.json |
Architecture config, text model reference, and component provenance |
metadata_pipeline.pkl |
Fitted sklearn pipeline (cloudpickle, by value β no stub modules needed) |
aggregator.pt |
Fusion encoder weights (inference-only, no training state) |
pv_duplication_embedder.py |
Model entry point |
model/aggregator/encoders.py |
Pure nn.Module architectures, reusable across projects |
projects/pv_duplication/specs.py |
The metadata fields this pipeline consumes |
model/ holds code that is not specific to this project; projects/ holds what is.
A second project reusing these architectures adds a directory rather than a fork.
Usage
Each input record must be a dict with exactly two keys:
| Key | Type | Description |
|---|---|---|
text |
str |
Free-text adverse-event narrative |
metadata |
dict |
Raw metadata fields consumed by the sklearn pipeline. The exact list is in projects/pv_duplication/specs.py; any field may be omitted and encodes as unknown |
from huggingface_hub import snapshot_download
import sys
# `revision` pins the release. Without it you track `main`, which moves.
snapshot_dir = snapshot_download(
"Ennov/pv_ae_document_duplication_embed", revision="v2"
)
sys.path.insert(0, snapshot_dir)
from pv_duplication_embedder import PVDuplicateEmbedder
model = PVDuplicateEmbedder.from_pretrained(snapshot_dir)
records = [
{
"text": "A 3-year-old Labrador received product X. Vomiting observed after 2 hours.",
"metadata": {
"recd_date": "2023-06-01",
"pt_name": ["Vomiting"],
"reporter_role": ["Veterinarian"],
"prod_code": ["PROD123"],
# ... see projects/pv_duplication/specs.py for the full field list
},
}
]
# Returns np.ndarray of shape (N, 1024), float32
embeddings = model.encode(records, batch_size=32)
from_pretrained also accepts a repo id directly, downloading on demand:
model = PVDuplicateEmbedder.from_pretrained(
"Ennov/pv_ae_document_duplication_embed", revision="v2"
)
Versions
Releases are pinned with git tags, so a revision keeps working after the next push.
| Tag | Contents |
|---|---|
v1 |
Gated fusion, 512-dim metadata, by-reference pickle. Different entry point β see that revision's own card |
v2 |
This release |
Authentication
Set your HuggingFace token before pushing:
# Option A β environment variable (CI/CD friendly)
export HF_TOKEN="hf_..."
# Option B β interactive login
huggingface-cli login
Then assemble, verify and publish β three commands, because publishing is the one step that is expensive to take back:
train save --wrapper v2
train load --wrapper v2 # refuses a bundle it cannot load
train push --wrapper v2 --tag v2
Training
The metadata vectorizer is fitted on the train split of a 70/30 mixture of clean and metadata-damaged reports. The fusion encoder is trained on cached component vectors with a triplet objective over annotated duplicate clusters, never on documents directly.
Component-level provenance β which registry version or run each part came from β is
recorded in config.json.
Expected metadata structure
Every key is optional β a field that is absent, null or unrecognised encodes as unknown rather than raising, which is what makes the vectorizer resilient. Supplying a field the pipeline does not read is ignored.
| Group | Shape | Fields |
|---|---|---|
| VeDDRA codes | list of int | llt_code, pt_code, hlt_code, soc_code |
| VeDDRA terms | list of str | llt_name, pt_name, hlt_name, soc_name |
| Dates | date, or ISO-8601 str; sign_start_date a list of them |
recd_date, orig_recd_date, first_recd_date, reaction_start_date, sign_start_date, patient_dead_date |
| Categoricals | str, except serious which is bool |
user_case_type, patient_sex, patient_species, veddra_class, serious |
| Numerics | float | patient_age, patient_weight, patient_exposed_number |
| Reporter | list of str | reporter_firstname, reporter_surname, reporter_country, reporter_city, reporter_institute, reporter_role, reporter_zip |
The same list is importable from the bundle as projects.pv_duplication.specs.FIELD_GROUPS.
These are accepted but not encoded: patient_breed, patient_reacted_number, patient_dead_number, prod_code, prod_report, prod_country. They sit in the fitted frame yet reach no branch of the vectorizer β blanking every one of them leaves the output bit-identical β so there is nothing to gain by supplying them.
Model structure
The fitted metadata pipeline, as the object actually is β every branch, transformer and the parameters that decide what it does.
Identical parallel branches are drawn once with a multiplier; sequential steps are always shown in full.
metadata_pipeline.html is scikit-learn's own interactive version β download it and open it in a browser.
- Downloads last month
- 50
