Instructions to use pidakwo/rtc-ner-extended with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- spaCy
How to use pidakwo/rtc-ner-extended with spaCy:
!pip install https://huggingface.co/pidakwo/rtc-ner-extended/resolve/main/rtc-ner-extended-any-py3-none-any.whl # Using spacy.load(). import spacy nlp = spacy.load("rtc-ner-extended") # Importing as module. import rtc-ner-extended nlp = rtc-ner-extended.load() - Notebooks
- Google Colab
- Kaggle
- rtc-ner-extended
- Model Description
- Entity Labels
- Intended Use
- Data Anonymization Application
- Model Architecture
- Model Files
- Performance
- Model Usage Test
- Limitations
- Research Context
- Related Model
- Citation
- Disclaimer
- rtc-ner-extended is provided for research purposes. Model predictions are automated outputs and should be independently validated before being used in operational, medical, emergency-response, privacy-critical, or other high-stakes applications.
- Model Description
rtc-ner-extended
rtc-ner-extended is a domain-specific Named Entity Recognition (NER) model developed for extracting structured information from
unstructured road traffic crash (RTC) narratives. It works to identify entities not covered by the pidakwo/rtc-ner model
The model identifies RTC attributes such as rtc_time, weekday, no_vehicle, vehicle_type, persons_involved, and rtc_factors (causal
factors, collision types, collision effects, physical site attributes, environmental condition, actors – humans and animals, goods types,
and general vehicle category). Additionally, the model was also used as part of a data anonymization workflow in which it was used to identify
entities - person, plate_no, address, and victim_organization which indicate personal details of road users involved in the RTC incident.
In order to hide the personal details in the content feature, NER entities identified were replaced with their label names.
Model Description
rtc-ner-extended is an English-language spaCy NER model consisting of a tok2vec component and an NER component.
The model uses spaCy's:
MultiHashEmbedrepresentation for token featuresMaxoutWindowEncoderfor contextual token representations- Transition-based NER architecture for entity recognition
The model was configured and trained using spaCy 3.8.x.
Entity Labels
rtc-ner-extended recognizes the following ten domain-specific entity types:
| Entity | Description |
|---|---|
ADDRESS |
Address of RTC incident victims appearing in an RTC narrative |
FACTORS |
Reported factors or circumstances associated with the occurrence of the crash |
NO_VEHICLES |
Number of vehicles involved in the crash |
PERSON |
Name of a person mentioned in the RTC narrative |
PERSONS_INVOLVED |
Total number of persons involved in the crash |
PLATE_NO |
Vehicle registration or license plate number |
TIME |
Time information associated with the crash or reported event |
VEHICLE_TYPE |
Type or category of vehicle involved in the crash |
VICTIM_ORGANIZATION |
Organization associated with a victim or incident |
WEEKDAY |
Day of the week associated with the crash |
Intended Use
rtc-ner-extended is intended primarily for research and information-extraction applications involving road traffic crash narratives.
Potential applications include:
- Extraction of structured information from RTC reports
- Construction and enrichment of RTC datasets
- Identification of crash-related factors
- Extraction of personally identifiable information (PII)
- Extraction of temporal information
- Preparation of RTC narratives for downstream machine learning
- Natural language processing of road safety reports
- Data curation and information structuring
- Support for road traffic crash analysis and research
Data Anonymization Application
rtc-ner-extended was also used as part of a privacy-preserving data processing workflow.
Four entity categories were specifically identified as potentially containing personally identifiable information (PII):
PERSONPLATE_NOADDRESSVICTIM_ORGANIZATION
These entities can be identified in an RTC narrative and subsequently replaced with their entity labels or other designated placeholders during anonymization.
Important: NER-based anonymization should not be regarded as a guarantee that all personally identifiable or sensitive information has been removed. Human review and/or additional privacy-preserving processing should be considered before publicly releasing processed text.
Model Architecture
The model uses the following spaCy pipeline:
tok2vec → ner
Model Files
The repository contains the complete trained spaCy model and its associated resources, including:
config.cfg
meta.json
tokenizer
vocab/
ner/
tok2vec/
These components should be retained together when loading the model.
Performance
The performance figures are recorded in the model's meta.json metadata.
Model Usage Test
The RTC_NER_Extended_model_test.ipynb file contains the code for testing usage of the model.
The final output contains the entities identified by the model together with their corresponding entity labels.
Limitations
rtc-ner-extended is a domain-specific research model and should not be assumed to identify every relevant entity in every RTC narrative.
Performance may be affected by:
- Differences in narrative writing styles
- Spelling and typographical errors
- Ambiguous entity boundaries
- Unusual abbreviations
- Previously unseen terminology
- Domain shifts between training and application data
- Differences in reporting conventions
- Context-dependent interpretations of entities
In particular, automated anonymization should not be considered sufficient on its own to guarantee that an RTC narrative contains no personally identifiable information.
For applications involving public release of textual data, model predictions should be complemented by appropriate validation and privacy review.
Research Context
rtc-ner-extended was developed as part of research investigating the transformation of unstructured road traffic crash narratives into structured, machine-readable information.
The model extends the information-extraction capability of the domain-specific pidakwo/rtc-ner model pipeline by identifying entities associated with crash
circumstances, persons, vehicles, temporal information, and potentially privacy-sensitive information.
The extracted information can subsequently support data curation, analysis, machine learning, and road safety research.
Related Model
A separate model, pidakwo/rtc-ner, was developed for the extraction of geographic and selected incident-related entities from RTC narratives.
RTC-NER and rtc-ner-extended are separate but complementary models, and should not be treated as interchangeable.
The code for model training and evaluation as well as data extraction can be found at: https://github.com/PatUnoka/Geospatial-and-Contextual-Information-Extraction-from-Road-Traffic-Crash-Narratives.git
Citation
If you use rtc-ner in academic research, please cite the associated research publication and dataset from which the model was developed.
[1] P. O. Idakwo, O. Adekanmbi, A. Soronnadi, and A. David, “Geo-parsing and analysis of road traffic crash incidents for data-driven emergency response planning,” Heliyon, vol. 11, no. 4, p. e41067, 2025, doi: 10.1016/j.heliyon.2024.e41067.
[2] P. O. Idakwo, O. Adekanmbi, and A. David, “Nigerian Multi-modal Road Traffic Crash Data,” 2026, doi: 10.5281/ZENODO.15862127.
Disclaimer
rtc-ner-extended is provided for research purposes. Model predictions are automated outputs and should be independently validated before being used in operational, medical, emergency-response, privacy-critical, or other high-stakes applications.
- Downloads last month
- -