Text Classification
setfit
Safetensors
sentence-transformers
mpnet
generated_from_setfit_trainer
text-embeddings-inference
Instructions to use peter2000/setfit-vulnerability-groups with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- setfit
How to use peter2000/setfit-vulnerability-groups with setfit:
from setfit import SetFitModel model = SetFitModel.from_pretrained("peter2000/setfit-vulnerability-groups") preds = model.predict(["i loved the spiderman movie!", "pineapple on pizza is the worst"]) print(preds) - sentence-transformers
How to use peter2000/setfit-vulnerability-groups with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("peter2000/setfit-vulnerability-groups") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
|
Download README.md from peter2000/setfit-vulnerability-groups: direct link, hf CLI and curl.
- Browser
- Download file 8.92 kB
-
https://huggingface.co/peter2000/setfit-vulnerability-groups/resolve/main/README.md
- Command line
-
hf download hf://peter2000/setfit-vulnerability-groups/README.md
-
curl -L -o README.md https://huggingface.co/peter2000/setfit-vulnerability-groups/resolve/main/README.md
8.92 kB
| tags: | |
| - setfit | |
| - sentence-transformers | |
| - text-classification | |
| - generated_from_setfit_trainer | |
| widget: | |
| - text: The infrastructure requirement for collection based on the targets and projections | |
| made is presented in Table 12.22. A total of about 149,000 km length and 8,660km | |
| length of sewers are required for urban and rural communities, respectively by | |
| 2047. In addition, a little over 4 million facilities in urban areas and about | |
| 853,000 facilities in rural areas will be required to meet on-site sanitation | |
| needs by 2033 nationwide. | |
| - text: The population of the Republic of Congo is among the most vulnerable, insofar | |
| as it has limited room for adaptation, mainly due to poverty. Maintaining the | |
| services provided by natural ecosystems (forests, savannas, hydrological basins, | |
| etc.) is essential to ensure future development relays, limit the impacts of climate | |
| change and offer possibilities for adaptation to the most vulnerable groups, including | |
| are part of women and young people of all socio-cultural categories of urban and | |
| rural centers. | |
| - text: Finally, and from the generation of spaces for the exchange of experiences | |
| and good practices, specialists in the subject,. In terms of raising awareness, | |
| a discussion was held to integrate the gender perspective into the climate change | |
| agenda, which included views from government management and science. The discussion | |
| was aimed at the general public and had the participation of youth organizations | |
| that work to raise awareness and sensitize in the fight against climate change. | |
| - text: 'Over the past period, even the Party and the Government of Lao PDR were aware | |
| of the importance of and have paid attention to gender role. However, the status | |
| of women in Lao PDR in many fields is not equal to that of men, and women were | |
| still taken advantage of in many forms. Hence, in order to ensure that peoples | |
| of all gender and ages and all social strata are able to participate in the process | |
| and receive the benefits from the development in a comprehensive, inclusive and | |
| fair manner, the National Green Growth Strategy of the Lao PDR has identified gender | |
| role/protection and promotion of the advancement of women activities to be an | |
| important focus of the green growth and will particularly focus on: ' | |
| - text: Construction of pipelines and connection to existing ones to transmit water | |
| to demand centres. Reduce water loss during transmission by investing on telemetric | |
| monitoring systems. Enhance conjunctive groundwater-surface water use. Agriculture. | |
| Improve genetic characteristics of the livestock breed such as Musi breed. Improve | |
| livestock diet through supplementary feeding. A switch to crops with the following | |
| traits:. Drought resistant. Tolerant to high temperatures. Short maturity. Health. | |
| Public education and malaria campaigns. Malaria Strategy. Control of Diarrhoeal | |
| Diseases | |
| metrics: | |
| - accuracy | |
| pipeline_tag: text-classification | |
| library_name: setfit | |
| inference: false | |
| base_model: sentence-transformers/paraphrase-mpnet-base-v2 | |
| # SetFit with sentence-transformers/paraphrase-mpnet-base-v2 | |
| This is a [SetFit](https://github.com/huggingface/setfit) model that can be used for Text Classification. This SetFit model uses [sentence-transformers/paraphrase-mpnet-base-v2](https://huggingface.co/sentence-transformers/paraphrase-mpnet-base-v2) as the Sentence Transformer embedding model. A OneVsRestClassifier instance is used for classification. | |
| The model has been trained using an efficient few-shot learning technique that involves: | |
| 1. Fine-tuning a [Sentence Transformer](https://www.sbert.net) with contrastive learning. | |
| 2. Training a classification head with features from the fine-tuned Sentence Transformer. | |
| ## Model Details | |
| ### Model Description | |
| - **Model Type:** SetFit | |
| - **Sentence Transformer body:** [sentence-transformers/paraphrase-mpnet-base-v2](https://huggingface.co/sentence-transformers/paraphrase-mpnet-base-v2) | |
| - **Classification head:** a OneVsRestClassifier instance | |
| - **Maximum Sequence Length:** 256 tokens | |
| - **Number of Classes:** 17 classes | |
| <!-- - **Training Dataset:** [Unknown](https://huggingface.co/datasets/unknown) --> | |
| <!-- - **Language:** Unknown --> | |
| <!-- - **License:** Unknown --> | |
| ### Model Sources | |
| - **Repository:** [SetFit on GitHub](https://github.com/huggingface/setfit) | |
| - **Paper:** [Efficient Few-Shot Learning Without Prompts](https://arxiv.org/abs/2209.11055) | |
| - **Blogpost:** [SetFit: Efficient Few-Shot Learning Without Prompts](https://huggingface.co/blog/setfit) | |
| ## Uses | |
| ### Direct Use for Inference | |
| First install the SetFit library: | |
| ```bash | |
| pip install setfit | |
| ``` | |
| Then you can load this model and run inference. | |
| ```python | |
| from setfit import SetFitModel | |
| # Download from the 🤗 Hub | |
| model = SetFitModel.from_pretrained("peter2000/setfit-vulnerability-groups") | |
| # Run inference | |
| preds = model("The infrastructure requirement for collection based on the targets and projections made is presented in Table 12.22. A total of about 149,000 km length and 8,660km length of sewers are required for urban and rural communities, respectively by 2047. In addition, a little over 4 million facilities in urban areas and about 853,000 facilities in rural areas will be required to meet on-site sanitation needs by 2033 nationwide.") | |
| ``` | |
| <!-- | |
| ### Downstream Use | |
| *List how someone could finetune this model on their own dataset.* | |
| --> | |
| <!-- | |
| ### Out-of-Scope Use | |
| *List how the model may foreseeably be misused and address what users ought not to do with the model.* | |
| --> | |
| <!-- | |
| ## Bias, Risks and Limitations | |
| *What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.* | |
| --> | |
| <!-- | |
| ### Recommendations | |
| *What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.* | |
| --> | |
| ## Training Details | |
| ### Training Set Metrics | |
| | Training set | Min | Median | Max | | |
| |:-------------|:----|:--------|:----| | |
| | Word count | 15 | 71.2316 | 164 | | |
| ### Training Hyperparameters | |
| - batch_size: (16, 16) | |
| - num_epochs: (1, 1) | |
| - max_steps: -1 | |
| - sampling_strategy: oversampling | |
| - num_iterations: 20 | |
| - body_learning_rate: (2e-05, 1e-05) | |
| - head_learning_rate: 0.01 | |
| - loss: CosineSimilarityLoss | |
| - distance_metric: cosine_distance | |
| - margin: 0.25 | |
| - end_to_end: False | |
| - use_amp: False | |
| - warmup_proportion: 0.1 | |
| - l2_weight: 0.01 | |
| - seed: 42 | |
| - eval_max_steps: -1 | |
| - load_best_model_at_end: False | |
| ### Training Results | |
| | Epoch | Step | Training Loss | Validation Loss | | |
| |:------:|:----:|:-------------:|:---------------:| | |
| | 0.0011 | 1 | 0.3006 | - | | |
| | 0.0526 | 50 | 0.2232 | - | | |
| | 0.1053 | 100 | 0.1670 | - | | |
| | 0.1579 | 150 | 0.1202 | - | | |
| | 0.2105 | 200 | 0.0935 | - | | |
| | 0.2632 | 250 | 0.0862 | - | | |
| | 0.3158 | 300 | 0.0626 | - | | |
| | 0.3684 | 350 | 0.0664 | - | | |
| | 0.4211 | 400 | 0.0555 | - | | |
| | 0.4737 | 450 | 0.0528 | - | | |
| | 0.5263 | 500 | 0.0543 | - | | |
| | 0.5789 | 550 | 0.0501 | - | | |
| | 0.6316 | 600 | 0.0535 | - | | |
| | 0.6842 | 650 | 0.0465 | - | | |
| | 0.7368 | 700 | 0.0468 | - | | |
| | 0.7895 | 750 | 0.0470 | - | | |
| | 0.8421 | 800 | 0.0421 | - | | |
| | 0.8947 | 850 | 0.0379 | - | | |
| | 0.9474 | 900 | 0.0475 | - | | |
| | 1.0 | 950 | 0.0449 | - | | |
| ### Framework Versions | |
| - Python: 3.12.12 | |
| - SetFit: 1.2.0 | |
| - Sentence Transformers: 6.1.0 | |
| - Transformers: 5.17.0 | |
| - PyTorch: 2.14.0+cu130 | |
| - Datasets: 5.0.1 | |
| - Tokenizers: 0.23.2 | |
| ## Citation | |
| ### BibTeX | |
| ```bibtex | |
| @article{https://doi.org/10.48550/arxiv.2209.11055, | |
| doi = {10.48550/ARXIV.2209.11055}, | |
| url = {https://arxiv.org/abs/2209.11055}, | |
| author = {Tunstall, Lewis and Reimers, Nils and Jo, Unso Eun Seo and Bates, Luke and Korat, Daniel and Wasserblat, Moshe and Pereg, Oren}, | |
| keywords = {Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences}, | |
| title = {Efficient Few-Shot Learning Without Prompts}, | |
| publisher = {arXiv}, | |
| year = {2022}, | |
| copyright = {Creative Commons Attribution 4.0 International} | |
| } | |
| ``` | |
| <!-- | |
| ## Glossary | |
| *Clearly define terms in order to be accessible across audiences.* | |
| --> | |
| <!-- | |
| ## Model Card Authors | |
| *Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.* | |
| --> | |
| <!-- | |
| ## Model Card Contact | |
| *Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.* | |
| --> |