Text Classification
setfit
Safetensors
sentence-transformers
mpnet
generated_from_setfit_trainer
text-embeddings-inference
Instructions to use peter2000/setfit-vulnerability-groups with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- setfit
How to use peter2000/setfit-vulnerability-groups with setfit:
from setfit import SetFitModel model = SetFitModel.from_pretrained("peter2000/setfit-vulnerability-groups") preds = model.predict(["i loved the spiderman movie!", "pineapple on pizza is the worst"]) print(preds) - sentence-transformers
How to use peter2000/setfit-vulnerability-groups with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("peter2000/setfit-vulnerability-groups") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
|
Download README.md from peter2000/setfit-vulnerability-groups: direct link, hf CLI and curl.
- Browser
- Download file 8.92 kB
-
https://huggingface.co/peter2000/setfit-vulnerability-groups/resolve/main/README.md
- Command line
-
hf download hf://peter2000/setfit-vulnerability-groups/README.md
-
curl -L -o README.md https://huggingface.co/peter2000/setfit-vulnerability-groups/resolve/main/README.md
8.92 kB
metadata
tags:
- setfit
- sentence-transformers
- text-classification
- generated_from_setfit_trainer
widget:
- text: >-
The infrastructure requirement for collection based on the targets and
projections made is presented in Table 12.22. A total of about 149,000 km
length and 8,660km length of sewers are required for urban and rural
communities, respectively by 2047. In addition, a little over 4 million
facilities in urban areas and about 853,000 facilities in rural areas will
be required to meet on-site sanitation needs by 2033 nationwide.
- text: >-
The population of the Republic of Congo is among the most vulnerable,
insofar as it has limited room for adaptation, mainly due to poverty.
Maintaining the services provided by natural ecosystems (forests,
savannas, hydrological basins, etc.) is essential to ensure future
development relays, limit the impacts of climate change and offer
possibilities for adaptation to the most vulnerable groups, including are
part of women and young people of all socio-cultural categories of urban
and rural centers.
- text: >-
Finally, and from the generation of spaces for the exchange of experiences
and good practices, specialists in the subject,. In terms of raising
awareness, a discussion was held to integrate the gender perspective into
the climate change agenda, which included views from government management
and science. The discussion was aimed at the general public and had the
participation of youth organizations that work to raise awareness and
sensitize in the fight against climate change.
- text: >-
Over the past period, even the Party and the Government of Lao PDR were
aware of the importance of and have paid attention to gender role.
However, the status of women in Lao PDR in many fields is not equal to
that of men, and women were still taken advantage of in many forms. Hence,
in order to ensure that peoples of all gender and ages and all social
strata are able to participate in the process and receive the benefits
from the development in a comprehensive, inclusive and fair manner, the
National Green Growth Strategy of the Lao PDR has identified gender
role/protection and promotion of the advancement of women activities to be
an important focus of the green growth and will particularly focus on:
- text: >-
Construction of pipelines and connection to existing ones to transmit
water to demand centres. Reduce water loss during transmission by
investing on telemetric monitoring systems. Enhance conjunctive
groundwater-surface water use. Agriculture. Improve genetic
characteristics of the livestock breed such as Musi breed. Improve
livestock diet through supplementary feeding. A switch to crops with the
following traits:. Drought resistant. Tolerant to high temperatures. Short
maturity. Health. Public education and malaria campaigns. Malaria
Strategy. Control of Diarrhoeal Diseases
metrics:
- accuracy
pipeline_tag: text-classification
library_name: setfit
inference: false
base_model: sentence-transformers/paraphrase-mpnet-base-v2
SetFit with sentence-transformers/paraphrase-mpnet-base-v2
This is a SetFit model that can be used for Text Classification. This SetFit model uses sentence-transformers/paraphrase-mpnet-base-v2 as the Sentence Transformer embedding model. A OneVsRestClassifier instance is used for classification.
The model has been trained using an efficient few-shot learning technique that involves:
- Fine-tuning a Sentence Transformer with contrastive learning.
- Training a classification head with features from the fine-tuned Sentence Transformer.
Model Details
Model Description
- Model Type: SetFit
- Sentence Transformer body: sentence-transformers/paraphrase-mpnet-base-v2
- Classification head: a OneVsRestClassifier instance
- Maximum Sequence Length: 256 tokens
- Number of Classes: 17 classes
Model Sources
- Repository: SetFit on GitHub
- Paper: Efficient Few-Shot Learning Without Prompts
- Blogpost: SetFit: Efficient Few-Shot Learning Without Prompts
Uses
Direct Use for Inference
First install the SetFit library:
pip install setfit
Then you can load this model and run inference.
from setfit import SetFitModel
# Download from the 🤗 Hub
model = SetFitModel.from_pretrained("peter2000/setfit-vulnerability-groups")
# Run inference
preds = model("The infrastructure requirement for collection based on the targets and projections made is presented in Table 12.22. A total of about 149,000 km length and 8,660km length of sewers are required for urban and rural communities, respectively by 2047. In addition, a little over 4 million facilities in urban areas and about 853,000 facilities in rural areas will be required to meet on-site sanitation needs by 2033 nationwide.")
Training Details
Training Set Metrics
| Training set | Min | Median | Max |
|---|---|---|---|
| Word count | 15 | 71.2316 | 164 |
Training Hyperparameters
- batch_size: (16, 16)
- num_epochs: (1, 1)
- max_steps: -1
- sampling_strategy: oversampling
- num_iterations: 20
- body_learning_rate: (2e-05, 1e-05)
- head_learning_rate: 0.01
- loss: CosineSimilarityLoss
- distance_metric: cosine_distance
- margin: 0.25
- end_to_end: False
- use_amp: False
- warmup_proportion: 0.1
- l2_weight: 0.01
- seed: 42
- eval_max_steps: -1
- load_best_model_at_end: False
Training Results
| Epoch | Step | Training Loss | Validation Loss |
|---|---|---|---|
| 0.0011 | 1 | 0.3006 | - |
| 0.0526 | 50 | 0.2232 | - |
| 0.1053 | 100 | 0.1670 | - |
| 0.1579 | 150 | 0.1202 | - |
| 0.2105 | 200 | 0.0935 | - |
| 0.2632 | 250 | 0.0862 | - |
| 0.3158 | 300 | 0.0626 | - |
| 0.3684 | 350 | 0.0664 | - |
| 0.4211 | 400 | 0.0555 | - |
| 0.4737 | 450 | 0.0528 | - |
| 0.5263 | 500 | 0.0543 | - |
| 0.5789 | 550 | 0.0501 | - |
| 0.6316 | 600 | 0.0535 | - |
| 0.6842 | 650 | 0.0465 | - |
| 0.7368 | 700 | 0.0468 | - |
| 0.7895 | 750 | 0.0470 | - |
| 0.8421 | 800 | 0.0421 | - |
| 0.8947 | 850 | 0.0379 | - |
| 0.9474 | 900 | 0.0475 | - |
| 1.0 | 950 | 0.0449 | - |
Framework Versions
- Python: 3.12.12
- SetFit: 1.2.0
- Sentence Transformers: 6.1.0
- Transformers: 5.17.0
- PyTorch: 2.14.0+cu130
- Datasets: 5.0.1
- Tokenizers: 0.23.2
Citation
BibTeX
@article{https://doi.org/10.48550/arxiv.2209.11055,
doi = {10.48550/ARXIV.2209.11055},
url = {https://arxiv.org/abs/2209.11055},
author = {Tunstall, Lewis and Reimers, Nils and Jo, Unso Eun Seo and Bates, Luke and Korat, Daniel and Wasserblat, Moshe and Pereg, Oren},
keywords = {Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences},
title = {Efficient Few-Shot Learning Without Prompts},
publisher = {arXiv},
year = {2022},
copyright = {Creative Commons Attribution 4.0 International}
}