Instructions to use lennartcb/pflm1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lennartcb/pflm1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="lennartcb/pflm1")# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("lennartcb/pflm1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use lennartcb/pflm1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "lennartcb/pflm1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lennartcb/pflm1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/lennartcb/pflm1
- SGLang
How to use lennartcb/pflm1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "lennartcb/pflm1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lennartcb/pflm1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "lennartcb/pflm1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lennartcb/pflm1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use lennartcb/pflm1 with Docker Model Runner:
docker model run hf.co/lennartcb/pflm1
PFLM: A Prior-Fitted Language Model
PFLM is a 300M-parameter byte-level model pretrained only on samples from a synthetic non-linguistic prior. Given the start of a byte sequence, such as a text in a language it has never seen, it infers the source in context and predicts what comes next.
This repository holds the weights and their configuration. The model code and a byte-level API for scoring streams and measuring in-context learning are at github.com/cbl/prior-fitted-language-model.
Quickstart
pip install "pflm1[hf]"
Watch it learn. The model has never seen prime numbers, yet it gets better at predicting them the more it reads.
import pflm1
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("lennartcb/pflm1", dtype="bfloat16").cuda().eval()
def primes(n): # 1 where the integer is prime, else 0
flags = bytearray([1]) * n
flags[:2] = b"\0\0"
for i in range(2, int(n ** 0.5) + 1):
if flags[i]:
flags[i * i::i] = bytes(len(flags[i * i::i]))
return bytes(48 + f for f in flags)
text = primes(10_000) # "0011010100..." ten thousand digits
bits = model.bits_per_byte(text) # bits per byte, one entry each
print(bits[:500].mean(), bits[-500:].mean()) # the first 500 digits vs. the last 500
Citation
@article{carstensbehrens2026learning,
title = {Learning to Learn a Language},
author = {Carstens-Behrens, Lennart and Fr{\"o}hlich, Holger},
journal = {arXiv preprint arXiv:2610.05879},
year = {2026},
}
- Downloads last month
- 327
Paper for lennartcb/pflm1
Paper • 2610.05879 • Published • 4