Text Generation
TensorRT
Safetensors
PyTorch
English
llama
code
developer-agent
blackwell
rtx-5070
slm
conversational
Instructions to use Vivid86/MiniTransformer-91M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- TensorRT
How to use Vivid86/MiniTransformer-91M with TensorRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
File size: 1,088 Bytes
a8e9558 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 | from tokenizers import Tokenizer, models, trainers, pre_tokenizers, decoders
import os
def train_tokenizer(data_folder="dataRaw", vocab_size=8192, save_path="dataProcessed/tokenizer.json"):
# Collect all text files
files = []
for root, _, filenames in os.walk(data_folder):
for f in filenames:
if f.endswith(".txt"):
files.append(os.path.join(root, f))
if not files:
raise ValueError("No .txt files found in dataRaw/. Add training text first.")
# Initialize tokenizer with byte-level pre-tokenization and decoding
tokenizer = Tokenizer(models.BPE())
tokenizer.pre_tokenizer = pre_tokenizers.ByteLevel()
tokenizer.decoder = decoders.ByteLevel()
trainer = trainers.BpeTrainer(
vocab_size=vocab_size,
min_frequency=2,
special_tokens=["<pad>", "<unk>", "<bos>", "<eos>"]
)
tokenizer.train(files, trainer)
tokenizer.save(save_path)
print(f"Tokenizer trained and saved to {save_path}")
if __name__ == "__main__":
train_tokenizer()
|