Text Generation
Transformers
Safetensors
English
bolmo
custom_code

Bwen 8B

Qwen3 8B Base retrofitted to operate over bytes instead of tokens via byteification through a short additional training procedure.

See our technical report for details: https://allenai.org/papers/bolmo.

Name Model Starting Point
Bolmo 1B Bolmo-1B Bolmo-1B-Stage1
Bolmo 7B Bolmo-7B Bolmo-7B-Stage1
Bwen 8B (you are here) Bwen-8B Bwen-8B-Stage1
Llama-B 8B Llama-B-8B Llama-B-8B-Stage1
Bolmo 1B (Stage 1) Bolmo-1B-Stage1 OLMo-2-1B
Bolmo 7B (Stage 1) Bolmo-7B-Stage1 Olmo-3-7B
Bwen 8B (Stage 1) Bwen-8B-Stage1 Qwen3-8B-Base
Llama-B 8B (Stage 1) Llama-B-8B-Stage1 Meta-Llama-3-8B

Installation

This model was tested with transformers 4.57.3 and Python 3.11:

pip install transformers>=4.57.3

It additionally requires the xlstm package (which needs Python>=3.11):

pip install xlstm==2.0.4

Inference

You can use this model with the standard HuggingFace transformers library:

from transformers import AutoModelForCausalLM, AutoTokenizer

device = "cuda"
model = AutoModelForCausalLM.from_pretrained("allenai/Bwen-8B", trust_remote_code=True).to(device)
tokenizer = AutoTokenizer.from_pretrained("allenai/Bwen-8B", trust_remote_code=True)

message = ["Language modeling is "]
input_ids = tokenizer(message, return_tensors="pt")["input_ids"].to(device)

# `max_new_tokens` is the amount of bytes to generate
response = model.generate(input_ids, max_new_tokens=256, do_sample=True, temperature=0.1)
print(tokenizer.decode(response[0], skip_special_tokens=True))

Model Description

  • Retrofitted from model: allenai/Bwen-8B-Stage1
  • Developed by: Allen Institute for AI (Ai2)
  • Model type: a byte-level autoregressive language model.
  • Language(s) (NLP): English
  • License: This model is licensed under Apache 2.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
  • Contact: Press: press@allenai.org

Model Sources

Bias, Risks, and Limitations

Like any base language model or fine-tuned model without safety filtering, these models can easily be prompted by users to generate harmful and sensitive content. Such content may also be produced unintentionally, especially in cases involving bias, so we recommend that users consider the risks when applying this technology. Additionally, many statements from Bolmo or any LLM are often inaccurate, so facts should be verified.

Citation

@misc{bolmo,
      title={Bolmo: Byteifying the Next Generation of Language Models}, 
      author={Benjamin Minixhofer and Tyler Murray and Tomasz Limisiewicz and Anna Korhonen and Luke Zettlemoyer and Noah A. Smith and Edoardo M. Ponti and Luca Soldaini and Valentin Hofmann},
      year={2025},
      eprint={2512.15586},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2512.15586}, 
}
Downloads last month
-
Safetensors
Model size
8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for allenai/Bwen-8B

Finetuned
(1)
this model

Dataset used to train allenai/Bwen-8B

Collection including allenai/Bwen-8B

Paper for allenai/Bwen-8B