Text Generation
Transformers
Safetensors
cohere2
conversational

Set tokenizer_class to PreTrainedTokenizerFast so that transformers v5 uses tokenizer.json

#3

Problem:

When the config says "tokenizer_class": "CohereTokenizerFast", transformers v5 loads CohereTokenizer. This class builds its own pre-tokenization rules in __init__ instead of using the rules in tokenizer.json. So under transformers v5, for the case print(type(tok).__name__, tok.tokenize("1234567")), every number is split differently from how the model was trained. We can apply exactly the same change as commit e9feb287 in the tiny-aya-l2-thinker repo. After this change, in our own tests, transformers 4.57.6 and 5.17.0 give identical results. The same change applies to tiny-aya-fire, -water and -global. I have also opened a PR on tiny-aya-base.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment