OSSoMeX: software mentions in research text

Experimental five-stage SciBERT pipeline for software names and versions, local version association, author intent, sentiment and aliases. This repository contains 006 and 012, with shared attribute heads. It is a custom pipeline, not a drop-in transformers.pipeline() model.

Code: krushiraj/OSSoMeX. Release tag: v0.0.12; source commit: 0019e3efbb7eccf0847fe6799caa272782d0b5ee. Python package metadata remains 0.0.1.

Models

Checkpoint Role Training and limitation
006 Earlier selected name-extraction POC Detector uses 161 annotated passages from 21 papers; version associations can be incorrect
012 Experimental Softcite adaptation Detector uses 900 source records; not established as better than 006 across domains

012 uses 011 detector weights with constrained wordpiece decoding, not continued training from 006. The four attribute/linking stages are shared. Both keep optional boundary repair disabled. registry.json and bundle-manifest.json record exact identities.

Installation and use

Install the code and pinned dependencies using the GitHub README. Download the complete Hub snapshot; keep the checkpoints/ sibling layout intact. This preview repository is public and can be downloaded without authentication.

from huggingface_hub import snapshot_download
model_root = snapshot_download("krushiraj/OSSoMeX", revision="v0.0.12")
print(model_root)

Use the printed path in place of /path/to/download below. Then pass the downloaded pipeline directory to the existing CLI:

.venv-scibert/bin/python -m research full-label predict \
  --model /path/to/download/checkpoints/scibert-full-label-006 \
  --device cpu --text 'We used NumPy 1.24.3.' --format pretty

Use scibert-full-label-012 to select the adapted detector. CPU and Apple MPS are supported CLI devices; CUDA is not exposed. Inference uses local safetensors/tokenizers and verifies hashes. No remote custom-code execution is required. Loading only a detector will not produce the five-stage pipeline's outputs.

public_rows contains names, versions, context, intent and sentiment. Inspect the full envelope's status, field predictions and review flags. Scores are uncalibrated; missing/unresolved fields and failed requests must not be treated as negative facts.

Evaluation

Exact software-name F1 (%), same frozen 480-token windows, 64-token overlap:

Slice 012 Softcite Wapiti Softcite SciBERT
OpenAlex fresh30 76.4 37.5 72.7
OpenAlex exposed50 56.6 44.3 42.9
OpenAlex reserve27 58.7 59.6 57.4
Seven regression excerpts 69.8 28.6 53.3
Recovered Softcite gold, 461 paragraphs 75.1 93.2 92.1

The OpenAlex/regression labels are provisional and development-exposed. Softcite name projection includes language fields and excludes implicit mentions. The fresh30 difference against SciBERT has a 95% interval spanning zero; no general winner is established. Gold covers 233 source articles rather than the entire published holdout. Do not average overlapping slices.

Selected SoMeSci10 recall is 71.7% versus 36.6%/51.0%; negative coverage is unaudited, so this is not a precision/F1 comparison. Regression7 version-link F1 is 79.2% versus 18.2%/27.8%, on only 28 reference links. Details are in the source repository's current benchmark document.

Limitations and intended use

  • Research and reviewed extraction, not production-validated author/software assertions.
  • 012 can miss MATLAB/Java in contexts where 006 succeeds; both recognize them in simpler text.
  • Training included contradictory positive-text/empty-label copies; the effect requires a controlled ablation.
  • Sentiment is weak (fresh30 observed macro-F1 20.5%); the three selected positive alias pairs were missed.
  • Dense package lists can produce expensive pairwise inference. Local speed comparisons used unmatched runtimes/hardware execution paths.
  • These models do not resolve mentions to repositories, registries or unique software identities.
  • No new training accompanied this release. A clean virtual environment on the development Mac installed the hash-pinned dependencies and built wheel; CPU inference succeeded for both pipelines outside the source checkout. This is not another-machine or production validation.

Training and attribution

The 006 pool contains 164 software occurrences and 83 normalized name strings; 150 of 161 passage annotations are agent provisional. The 012 detector pool contains 2,869 names, 906 versions and 1,045 normalized strings (1,117 exact spellings). These are strings, not resolved software projects. 012's attribute heads retain the earlier training lineage.

Base: SciBERT, Beltagy, Lo and Cohan (2019), paper. Adaptation data: Softcite dataset, documented under CC-BY-4.0; recovered input came from the Softcite software-mentions repository. Training/evaluation datasets and article text are not redistributed here.

The accompanying LICENSE applies to the project code; checkpoint reuse and redistribution terms have not been separately assigned. Third-party sources retain their own terms. Original manifest bytes are retained for hash verification, including historical local provenance paths.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for krushiraj/OSSoMeX

Finetuned
(18)
this model