Instructions to use farid678/code-search-net-tokenizer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use farid678/code-search-net-tokenizer with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("farid678/code-search-net-tokenizer", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 10,166 Bytes
cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb cf1b2b2 5493aeb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 | ---
library_name: transformers
tags:
- gpt2
- tokenizer
- code
- python
- code-search-net
datasets:
- code_search_net
base_model: gpt2
---
# Model Card for farid678/gpt2-python-tokenizer
<!-- Provide a quick summary of what the model is/does. -->
A new byte-level BPE tokenizer trained from scratch for GPT-2, specialized for Python source code, using the `code_search_net` (Python subset) dataset.
## Model Details
### Model Description
<!-- Provide a longer summary of what this model is. -->
This repository contains a custom tokenizer trained from the base GPT-2 tokenizer architecture, re-trained on the Python portion of the `code_search_net` dataset. The goal of this tokenizer is to better capture Python-specific syntax, keywords, identifiers, and code patterns (e.g. indentation, operators, common function/variable naming conventions) compared to the original GPT-2 tokenizer, which was trained primarily on natural language web text.
This tokenizer can be paired with a GPT-2 model (either the original pretrained weights with an extended/adapted embedding layer, or a model trained from scratch) for downstream tasks involving Python code, such as code completion, code summarization, or code generation.
- **Developed by:** [farid678](https://huggingface.co/farid678)
- **Funded by [optional]:** [More Information Needed]
- **Shared by [optional]:** farid678
- **Model type:** Byte-level BPE tokenizer (GPT-2 architecture)
- **Language(s) (NLP):** Python (programming language); tokenizer vocabulary derived from source code rather than natural language
- **License:** [More Information Needed]
- **Finetuned from model:** `gpt2` (tokenizer re-trained from scratch on new data, using GPT-2's tokenizer architecture as the base)
### Model Sources [optional]
<!-- Provide the basic links for the model. -->
- **Repository:** https://huggingface.co/farid678
- **Paper [optional]:** [More Information Needed]
- **Demo [optional]:** [More Information Needed]
## Uses
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
### Direct Use
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
This tokenizer can be used directly to tokenize Python source code for input into a GPT-2-style language model. It is intended for use in code-related NLP pipelines such as tokenizing datasets before training/fine-tuning a language model on Python code.
### Downstream Use [optional]
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
Intended to be paired with a GPT-2 (or GPT-2-style) causal language model for tasks such as:
- Python code completion
- Python code generation
- Code summarization / docstring generation
- Code-to-text or text-to-code tasks
### Out-of-Scope Use
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
This tokenizer is optimized for Python code and is not expected to perform well on natural language text or other programming languages (e.g. Java, C++, JavaScript) since its vocabulary was derived specifically from Python source code in `code_search_net`. It should not be used as a general-purpose natural language tokenizer.
## Bias, Risks, and Limitations
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
- The tokenizer's vocabulary reflects patterns present in the `code_search_net` Python subset, which is sourced from public open-source GitHub repositories. As such, it may inherit biases present in that codebase (e.g. naming conventions, coding styles, or underrepresentation of certain coding domains).
- Performance on code written in significantly different styles, older Python versions, or non-English identifiers/comments may be degraded.
- This tokenizer alone does not generate code; it must be paired with a trained language model to be useful for downstream tasks.
### Recommendations
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the tokenizer. It is recommended to evaluate tokenization quality (e.g. compression rate, out-of-vocabulary handling) on your own target dataset before relying on it for production use.
## How to Get Started with the Model
Use the code below to get started with the tokenizer.
```python
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("farid678/gpt2-python-tokenizer")
code_sample = "def hello_world():\n print('Hello, world!')"
tokens = tokenizer.tokenize(code_sample)
print(tokens)
encoded = tokenizer(code_sample)
print(encoded["input_ids"])
```
## Training Details
### Training Data
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
The tokenizer was trained on the **Python subset** of the [`code_search_net`](https://huggingface.co/datasets/code_search_net) dataset:
```python
from datasets import load_dataset
raw_dataset = load_dataset("code_search_net", "python")
```
`code_search_net` contains functions and methods collected from open-source GitHub repositories, along with their associated docstrings/comments. The Python configuration used here consists of Python source code specifically.
### Training Procedure
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
A new byte-level BPE tokenizer was trained from scratch using the GPT-2 tokenizer architecture as a template (i.e. `tokenizer.train_new_from_iterator` from the 🤗 Tokenizers/Transformers library), using the raw code text from `code_search_net` (Python) as the training corpus.
#### Preprocessing [optional]
Python code and associated documentation strings from `code_search_net` were used as raw text input for tokenizer training. [More Information Needed] (exact preprocessing steps, e.g. whether docstrings/comments were included or code-only)
#### Training Hyperparameters
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
- **Vocabulary size:** [More Information Needed]
- **Base tokenizer:** `gpt2` (byte-level BPE)
#### Speeds, Sizes, Times [optional]
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
[More Information Needed]
## Evaluation
<!-- This section describes the evaluation protocols and provides the results. -->
### Testing Data, Factors & Metrics
#### Testing Data
<!-- This should link to a Dataset Card if possible. -->
[More Information Needed]
#### Factors
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
[More Information Needed]
#### Metrics
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
[More Information Needed] (e.g. average tokens per line of code, compression ratio vs. original GPT-2 tokenizer, out-of-vocabulary rate)
### Results
[More Information Needed]
#### Summary
This tokenizer is expected to produce more efficient, code-aware tokenization for Python source code compared to the original GPT-2 tokenizer, though formal benchmark results have not yet been recorded.
## Model Examination [optional]
<!-- Relevant interpretability work for the model goes here -->
[More Information Needed]
## Environmental Impact
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
- **Hardware Type:** [More Information Needed]
- **Hours used:** [More Information Needed]
- **Cloud Provider:** [More Information Needed]
- **Compute Region:** [More Information Needed]
- **Carbon Emitted:** [More Information Needed]
## Technical Specifications [optional]
### Model Architecture and Objective
Byte-level BPE tokenizer following the GPT-2 tokenizer architecture, retrained on a new corpus (Python code from `code_search_net`) rather than the original GPT-2 training data.
### Compute Infrastructure
[More Information Needed]
#### Hardware
[More Information Needed]
#### Software
- 🤗 `transformers`
- 🤗 `datasets`
- 🤗 `tokenizers`
## Citation [optional]
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
**BibTeX:**
```bibtex
@misc{husain2019codesearchnet,
title={CodeSearchNet Challenge: Evaluating the State of Semantic Code Search},
author={Husain, Hamel and Wu, Ho-Hsiang and Gazit, Tiferet and Allamanis, Miltiadis and Brockschmidt, Marc},
year={2019},
eprint={1909.09436},
archivePrefix={arXiv}
}
```
**APA:**
Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., & Brockschmidt, M. (2019). CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. arXiv preprint arXiv:1909.09436.
## Glossary [optional]
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
- **BPE (Byte-Pair Encoding):** A subword tokenization algorithm that iteratively merges the most frequent pairs of bytes/characters to build a vocabulary.
- **`train_new_from_iterator`:** A 🤗 Transformers method that allows retraining an existing tokenizer's vocabulary on a new corpus while keeping the same tokenization algorithm/architecture.
## More Information [optional]
[More Information Needed]
## Model Card Authors [optional]
farid678
## Model Card Contact
https://huggingface.co/farid678 |