Instructions to use lowerkey/bard-small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lowerkey/bard-small with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="lowerkey/bard-small", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("lowerkey/bard-small", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use lowerkey/bard-small with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "lowerkey/bard-small" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lowerkey/bard-small", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/lowerkey/bard-small
- SGLang
How to use lowerkey/bard-small with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "lowerkey/bard-small" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lowerkey/bard-small", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "lowerkey/bard-small" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lowerkey/bard-small", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use lowerkey/bard-small with Docker Model Runner:
docker model run hf.co/lowerkey/bard-small
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("lowerkey/bard-small", trust_remote_code=True, device_map="auto")Play in the browser · Space · Use it from Python
Bard Small is a 137M parameter model that runs text adventures. You type a first line, it writes what happens next, tells the game what changed (you lost some money, picked up a key) and gives you a few choices. Pick one or type something else and it keeps going.
I trained it from scratch for a game I made, Stochastic Lands of Prophecy. The game runs the whole model in your browser, so there's no server and nothing you type leaves your machine. You can try it there or in the Space. This repo has the weights plus the code that goes around them (tokenizer, number handling, the game rules) so you can also use it from Python.
A sample
The opening was You arrive at the gates of the old castle. and I picked the first choice. I ran it a few times and
took one of the better ones.
You arrive at the gates of the old castle. “You’ve come a long way, little Prince. The guards say that you are only
here for a few days. If you wish for it, I can take you back home and live somewhere else. Just tell me where your
home is and the guards will let you pass.”
This request really brings a smile to your face. The king asks you to accompany him to his manor and there you will
learn more.
The gate opens and you follow him inside. He takes a deep breath before leading you into a room. You see a man sitting
on a chair.
“I have a question for you son,” he says with a smile.
1. “What do you want to ask me?”
2. “What will you be doing now?”
3. “What is your name?”
> “What do you want to ask me?”
“Well, Prince, I have been summoned to King Arthur’s castle and I have come to ask for your help. Do you know where
your home is located?”
You think about this for a moment. The king is waiting for you to speak but he already knows.
It loses the thread pretty quickly. More on that at the bottom.
Using it
# pip install torch transformers tokenmonster
from transformers import AutoModelForCausalLM
bard = AutoModelForCausalLM.from_pretrained("lowerkey/bard-small", trust_remote_code=True)
tale = bard.begin("You wake up in a dungeon.")
print(tale.text) # what the model wrote
print(tale.choices) # what you can do next
tale = bard.act(tale, tale.choices[0]) # take a choice
tale = bard.act(tale, "I ask the guard about the key") # or type your own
bard.play() runs the same loop in the terminal if you only want to poke at it.
begin takes a few extra arguments. length is "short", "medium" or "long" and controls how much it writes per turn.
state is what the player carries and defaults to {"money": 100}. seed makes a run repeatable. on_text is a
function that receives the text in pieces while it's being written, which is handy for streaming.
A Tale is a small dataclass and the whole story is in tale.history, a plain string. To save one use
tale.to_dict(), and to load it again Tale.from_dict(d).
On a GPU load it in bf16 with from_pretrained(..., dtype=torch.bfloat16).to("cuda") (the argument is called
torch_dtype before transformers 4.56). A turn takes about a second on a GPU and a few seconds on a CPU. I tried it
with transformers 4.46 and 5.19.
You need trust_remote_code=True because the architecture isn't part of transformers. The code it runs is the .py
files in this repo.
Why there's no generate()
Numbers don't go through the tokenizer. Every run of digits in a text turns into one <|num|> token, and the actual
value travels in a second stream next to the token ids. A small network turns that value into a vector that gets added
to the token's embedding, and when the model wants to write a number a second small network predicts its digits. The
model's input is therefore two tensors, input_ids and number_values, and the usual generate() only knows about
the first. bard.begin and bard.act take care of this. If you want raw logits, bard(input_ids, number_values)
gives you those.
The format
The model reads and writes tales in a small text format. A tale is a list of blocks, and every block is
<|start|>ROLE<|content|>text<|end|>
where ROLE is state, story or player. The role headers and all the other markers are single tokens.
A state block says what the player has. Each variable is its name, a newline and its value, and variables are
separated by <|sep|>. The game writes this block, not the model, and puts a new one in after every turn.
A player block is the player's move: the text of the choice they picked or whatever they typed. Also written by the
game.
A story block is the only one the model writes. First comes the narration. After that, if something changed,
<|vars|> and then the changes separated by <|sep|>, each one a variable name, then <|add|> or <|set|>, then
a number. Last comes <|choices|> and the choices, again separated by <|sep|>. When the model writes a story block
with no choices, the tale is over.
Put together it looks like this (the passage is shortened):
<|start|>state<|content|>money
100<|end|>
<|start|>story<|content|>The guard looks at your coin purse and smiles.<|vars|>money<|add|>-20<|choices|>Pay him<|sep|>Run<|end|>
<|start|>state<|content|>money
80<|end|>
<|start|>player<|content|>Pay him<|end|>
<|start|>story<|content|>He steps aside and the door swings open. ...
What the model puts in <|vars|> is a suggestion. A model this small will cheerfully write money -400 when you
have 100, or take your money after a fight where nobody had any. bard_rules.py looks at each change before it goes
through. A change to money has to be backed up by the passage actually talking about money, an amount can't drop below
zero or grow tenfold in one go, and so on. bard_steering.py keeps passages from running on forever and makes sure
there are at least two different choices.
The model
It has 36 layers, a width of 512 and 8 attention heads, for 136.7M parameters. Attention and the feed forward layer run in parallel in each block. Most layers look at a sliding window of 256 tokens with RoPE, and every fourth layer looks at everything with no position encoding. There's QK-norm and a gate on the attention output, and the logits are softcapped. The context is 4096 tokens, and older text falls off the back as a tale gets long.
The vocabulary has 16,128 ids: tokenmonster's fiction-16000-strict-v1 plus the special tokens above. The weights are
in bf16 (model.safetensors, 273 MB).
Training
Three stages, all from scratch. First plain English, about 11B tokens of FineWeb-Edu. Then fiction: Royal Road stories, Project Gutenberg fiction, WritingPrompts and BookCorpus, with some FineWeb mixed back in so it doesn't forget how to write normal sentences. Last, choose-your-own-adventure stories converted into the block format above, which is where it learns to write narration, changes and choices in that shape.
What it's bad at
Plenty. It's small. Names and details change between turns, the plot wanders off, and now and then a turn makes no sense at all. It reads more like a dream than a story. If it gets stuck, pick a different choice or type your own action, that usually shakes it loose.
It was trained on web text and fiction, some of it for adults, and I did no safety tuning. It can write violent or sexual things, and it can be biased. Don't use it for anything important.
License and data
The weights and the code here are Apache 2.0. The training data is publicly available text. Some of it, Royal Road and
BookCorpus in particular, belongs to its authors and has its own terms, so if you want to build something on these
weights, check that it's fine for what you have in mind. The tokenizer vocabulary comes from tokenmonster's
fiction-16000-strict-v1 (MIT).
- Downloads last month
- 23

# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="lowerkey/bard-small", trust_remote_code=True)