Transformers
gpt2
tokenizer
code
python
code-search-net
File size: 10,166 Bytes
cf1b2b2
 
5493aeb
 
 
 
 
 
 
 
 
cf1b2b2
 
5493aeb
cf1b2b2
 
 
5493aeb
cf1b2b2
 
 
 
 
 
 
5493aeb
 
 
cf1b2b2
5493aeb
cf1b2b2
5493aeb
 
 
cf1b2b2
5493aeb
cf1b2b2
 
 
 
 
5493aeb
cf1b2b2
 
 
 
 
 
 
 
 
 
 
5493aeb
cf1b2b2
 
 
 
 
5493aeb
 
 
 
 
cf1b2b2
 
 
 
 
5493aeb
cf1b2b2
 
 
 
 
5493aeb
 
 
cf1b2b2
 
 
 
 
5493aeb
cf1b2b2
 
 
5493aeb
cf1b2b2
5493aeb
 
 
 
 
 
 
 
 
 
 
 
cf1b2b2
 
 
 
 
 
 
5493aeb
 
 
 
 
 
 
 
 
cf1b2b2
 
 
 
 
5493aeb
cf1b2b2
5493aeb
cf1b2b2
5493aeb
cf1b2b2
 
 
 
5493aeb
 
cf1b2b2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5493aeb
cf1b2b2
 
 
 
 
 
 
5493aeb
cf1b2b2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5493aeb
cf1b2b2
 
 
 
 
 
 
 
 
 
 
5493aeb
 
 
cf1b2b2
 
 
 
 
 
 
5493aeb
 
 
 
 
 
 
 
 
cf1b2b2
 
 
5493aeb
cf1b2b2
 
 
 
 
5493aeb
 
cf1b2b2
 
 
 
 
 
 
5493aeb
cf1b2b2
 
 
5493aeb
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
---
library_name: transformers
tags:
- gpt2
- tokenizer
- code
- python
- code-search-net
datasets:
- code_search_net
base_model: gpt2
---

# Model Card for farid678/gpt2-python-tokenizer

<!-- Provide a quick summary of what the model is/does. -->

A new byte-level BPE tokenizer trained from scratch for GPT-2, specialized for Python source code, using the `code_search_net` (Python subset) dataset.

## Model Details

### Model Description

<!-- Provide a longer summary of what this model is. -->

This repository contains a custom tokenizer trained from the base GPT-2 tokenizer architecture, re-trained on the Python portion of the `code_search_net` dataset. The goal of this tokenizer is to better capture Python-specific syntax, keywords, identifiers, and code patterns (e.g. indentation, operators, common function/variable naming conventions) compared to the original GPT-2 tokenizer, which was trained primarily on natural language web text.

This tokenizer can be paired with a GPT-2 model (either the original pretrained weights with an extended/adapted embedding layer, or a model trained from scratch) for downstream tasks involving Python code, such as code completion, code summarization, or code generation.

- **Developed by:** [farid678](https://huggingface.co/farid678)
- **Funded by [optional]:** [More Information Needed]
- **Shared by [optional]:** farid678
- **Model type:** Byte-level BPE tokenizer (GPT-2 architecture)
- **Language(s) (NLP):** Python (programming language); tokenizer vocabulary derived from source code rather than natural language
- **License:** [More Information Needed]
- **Finetuned from model:** `gpt2` (tokenizer re-trained from scratch on new data, using GPT-2's tokenizer architecture as the base)

### Model Sources [optional]

<!-- Provide the basic links for the model. -->

- **Repository:** https://huggingface.co/farid678
- **Paper [optional]:** [More Information Needed]
- **Demo [optional]:** [More Information Needed]

## Uses

<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->

### Direct Use

<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->

This tokenizer can be used directly to tokenize Python source code for input into a GPT-2-style language model. It is intended for use in code-related NLP pipelines such as tokenizing datasets before training/fine-tuning a language model on Python code.

### Downstream Use [optional]

<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->

Intended to be paired with a GPT-2 (or GPT-2-style) causal language model for tasks such as:
- Python code completion
- Python code generation
- Code summarization / docstring generation
- Code-to-text or text-to-code tasks

### Out-of-Scope Use

<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->

This tokenizer is optimized for Python code and is not expected to perform well on natural language text or other programming languages (e.g. Java, C++, JavaScript) since its vocabulary was derived specifically from Python source code in `code_search_net`. It should not be used as a general-purpose natural language tokenizer.

## Bias, Risks, and Limitations

<!-- This section is meant to convey both technical and sociotechnical limitations. -->

- The tokenizer's vocabulary reflects patterns present in the `code_search_net` Python subset, which is sourced from public open-source GitHub repositories. As such, it may inherit biases present in that codebase (e.g. naming conventions, coding styles, or underrepresentation of certain coding domains).
- Performance on code written in significantly different styles, older Python versions, or non-English identifiers/comments may be degraded.
- This tokenizer alone does not generate code; it must be paired with a trained language model to be useful for downstream tasks.

### Recommendations

<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->

Users (both direct and downstream) should be made aware of the risks, biases and limitations of the tokenizer. It is recommended to evaluate tokenization quality (e.g. compression rate, out-of-vocabulary handling) on your own target dataset before relying on it for production use.

## How to Get Started with the Model

Use the code below to get started with the tokenizer.

```python
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("farid678/gpt2-python-tokenizer")

code_sample = "def hello_world():\n    print('Hello, world!')"
tokens = tokenizer.tokenize(code_sample)
print(tokens)

encoded = tokenizer(code_sample)
print(encoded["input_ids"])
```

## Training Details

### Training Data

<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->

The tokenizer was trained on the **Python subset** of the [`code_search_net`](https://huggingface.co/datasets/code_search_net) dataset:

```python
from datasets import load_dataset

raw_dataset = load_dataset("code_search_net", "python")
```

`code_search_net` contains functions and methods collected from open-source GitHub repositories, along with their associated docstrings/comments. The Python configuration used here consists of Python source code specifically.

### Training Procedure

<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->

A new byte-level BPE tokenizer was trained from scratch using the GPT-2 tokenizer architecture as a template (i.e. `tokenizer.train_new_from_iterator` from the 🤗 Tokenizers/Transformers library), using the raw code text from `code_search_net` (Python) as the training corpus.

#### Preprocessing [optional]

Python code and associated documentation strings from `code_search_net` were used as raw text input for tokenizer training. [More Information Needed] (exact preprocessing steps, e.g. whether docstrings/comments were included or code-only)

#### Training Hyperparameters

- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
- **Vocabulary size:** [More Information Needed]
- **Base tokenizer:** `gpt2` (byte-level BPE)

#### Speeds, Sizes, Times [optional]

<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->

[More Information Needed]

## Evaluation

<!-- This section describes the evaluation protocols and provides the results. -->

### Testing Data, Factors & Metrics

#### Testing Data

<!-- This should link to a Dataset Card if possible. -->

[More Information Needed]

#### Factors

<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->

[More Information Needed]

#### Metrics

<!-- These are the evaluation metrics being used, ideally with a description of why. -->

[More Information Needed] (e.g. average tokens per line of code, compression ratio vs. original GPT-2 tokenizer, out-of-vocabulary rate)

### Results

[More Information Needed]

#### Summary

This tokenizer is expected to produce more efficient, code-aware tokenization for Python source code compared to the original GPT-2 tokenizer, though formal benchmark results have not yet been recorded.

## Model Examination [optional]

<!-- Relevant interpretability work for the model goes here -->

[More Information Needed]

## Environmental Impact

<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->

Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).

- **Hardware Type:** [More Information Needed]
- **Hours used:** [More Information Needed]
- **Cloud Provider:** [More Information Needed]
- **Compute Region:** [More Information Needed]
- **Carbon Emitted:** [More Information Needed]

## Technical Specifications [optional]

### Model Architecture and Objective

Byte-level BPE tokenizer following the GPT-2 tokenizer architecture, retrained on a new corpus (Python code from `code_search_net`) rather than the original GPT-2 training data.

### Compute Infrastructure

[More Information Needed]

#### Hardware

[More Information Needed]

#### Software

- 🤗 `transformers`
- 🤗 `datasets`
- 🤗 `tokenizers`

## Citation [optional]

<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->

**BibTeX:**

```bibtex
@misc{husain2019codesearchnet,
  title={CodeSearchNet Challenge: Evaluating the State of Semantic Code Search},
  author={Husain, Hamel and Wu, Ho-Hsiang and Gazit, Tiferet and Allamanis, Miltiadis and Brockschmidt, Marc},
  year={2019},
  eprint={1909.09436},
  archivePrefix={arXiv}
}
```

**APA:**

Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., & Brockschmidt, M. (2019). CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. arXiv preprint arXiv:1909.09436.

## Glossary [optional]

<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->

- **BPE (Byte-Pair Encoding):** A subword tokenization algorithm that iteratively merges the most frequent pairs of bytes/characters to build a vocabulary.
- **`train_new_from_iterator`:** A 🤗 Transformers method that allows retraining an existing tokenizer's vocabulary on a new corpus while keeping the same tokenization algorithm/architecture.

## More Information [optional]

[More Information Needed]

## Model Card Authors [optional]

farid678

## Model Card Contact

https://huggingface.co/farid678