Bard Small

Play in the browser · Space · Use it from Python

Bard Small is a 137M parameter model that runs text adventures. You type a first line, it writes what happens next, tells the game what changed (you lost some money, picked up a key) and gives you a few choices. Pick one or type something else and it keeps going.

I trained it from scratch for a game I made, Stochastic Lands of Prophecy. The game runs the whole model in your browser, so there's no server and nothing you type leaves your machine. You can try it there or in the Space. This repo has the weights plus the code that goes around them (tokenizer, number handling, the game rules) so you can also use it from Python.

A sample

The opening was You arrive at the gates of the old castle. and I picked the first choice. I ran it a few times and took one of the better ones.

You arrive at the gates of the old castle. “You’ve come a long way, little Prince. The guards say that you are only
here for a few days. If you wish for it, I can take you back home and live somewhere else. Just tell me where your
home is and the guards will let you pass.”

This request really brings a smile to your face. The king asks you to accompany him to his manor and there you will
learn more.

The gate opens and you follow him inside. He takes a deep breath before leading you into a room. You see a man sitting
on a chair.

“I have a question for you son,” he says with a smile.

  1. “What do you want to ask me?”
  2. “What will you be doing now?”
  3. “What is your name?”

> “What do you want to ask me?”

“Well, Prince, I have been summoned to King Arthur’s castle and I have come to ask for your help. Do you know where
your home is located?”

You think about this for a moment. The king is waiting for you to speak but he already knows.

It loses the thread pretty quickly. More on that at the bottom.

Using it

# pip install torch transformers tokenmonster
from transformers import AutoModelForCausalLM

bard = AutoModelForCausalLM.from_pretrained("lowerkey/bard-small", trust_remote_code=True)

tale = bard.begin("You wake up in a dungeon.")
print(tale.text)       # what the model wrote
print(tale.choices)    # what you can do next

tale = bard.act(tale, tale.choices[0])                  # take a choice
tale = bard.act(tale, "I ask the guard about the key")  # or type your own

bard.play() runs the same loop in the terminal if you only want to poke at it.

begin takes a few extra arguments. length is "short", "medium" or "long" and controls how much it writes per turn. state is what the player carries and defaults to {"money": 100}. seed makes a run repeatable. on_text is a function that receives the text in pieces while it's being written, which is handy for streaming.

A Tale is a small dataclass and the whole story is in tale.history, a plain string. To save one use tale.to_dict(), and to load it again Tale.from_dict(d).

On a GPU load it in bf16 with from_pretrained(..., dtype=torch.bfloat16).to("cuda") (the argument is called torch_dtype before transformers 4.56). A turn takes about a second on a GPU and a few seconds on a CPU. I tried it with transformers 4.46 and 5.19.

You need trust_remote_code=True because the architecture isn't part of transformers. The code it runs is the .py files in this repo.

Why there's no generate()

Numbers don't go through the tokenizer. Every run of digits in a text turns into one <|num|> token, and the actual value travels in a second stream next to the token ids. A small network turns that value into a vector that gets added to the token's embedding, and when the model wants to write a number a second small network predicts its digits. The model's input is therefore two tensors, input_ids and number_values, and the usual generate() only knows about the first. bard.begin and bard.act take care of this. If you want raw logits, bard(input_ids, number_values) gives you those.

The format

The model reads and writes tales in a small text format. A tale is a list of blocks, and every block is

<|start|>ROLE<|content|>text<|end|>

where ROLE is state, story or player. The role headers and all the other markers are single tokens.

A state block says what the player has. Each variable is its name, a newline and its value, and variables are separated by <|sep|>. The game writes this block, not the model, and puts a new one in after every turn.

A player block is the player's move: the text of the choice they picked or whatever they typed. Also written by the game.

A story block is the only one the model writes. First comes the narration. After that, if something changed, <|vars|> and then the changes separated by <|sep|>, each one a variable name, then <|add|> or <|set|>, then a number. Last comes <|choices|> and the choices, again separated by <|sep|>. When the model writes a story block with no choices, the tale is over.

Put together it looks like this (the passage is shortened):

<|start|>state<|content|>money
100<|end|>
<|start|>story<|content|>The guard looks at your coin purse and smiles.<|vars|>money<|add|>-20<|choices|>Pay him<|sep|>Run<|end|>
<|start|>state<|content|>money
80<|end|>
<|start|>player<|content|>Pay him<|end|>
<|start|>story<|content|>He steps aside and the door swings open. ...

What the model puts in <|vars|> is a suggestion. A model this small will cheerfully write money -400 when you have 100, or take your money after a fight where nobody had any. bard_rules.py looks at each change before it goes through. A change to money has to be backed up by the passage actually talking about money, an amount can't drop below zero or grow tenfold in one go, and so on. bard_steering.py keeps passages from running on forever and makes sure there are at least two different choices.

The model

It has 36 layers, a width of 512 and 8 attention heads, for 136.7M parameters. Attention and the feed forward layer run in parallel in each block. Most layers look at a sliding window of 256 tokens with RoPE, and every fourth layer looks at everything with no position encoding. There's QK-norm and a gate on the attention output, and the logits are softcapped. The context is 4096 tokens, and older text falls off the back as a tale gets long.

The vocabulary has 16,128 ids: tokenmonster's fiction-16000-strict-v1 plus the special tokens above. The weights are in bf16 (model.safetensors, 273 MB).

Training

Three stages, all from scratch. First plain English, about 11B tokens of FineWeb-Edu. Then fiction: Royal Road stories, Project Gutenberg fiction, WritingPrompts and BookCorpus, with some FineWeb mixed back in so it doesn't forget how to write normal sentences. Last, choose-your-own-adventure stories converted into the block format above, which is where it learns to write narration, changes and choices in that shape.

What it's bad at

Plenty. It's small. Names and details change between turns, the plot wanders off, and now and then a turn makes no sense at all. It reads more like a dream than a story. If it gets stuck, pick a different choice or type your own action, that usually shakes it loose.

It was trained on web text and fiction, some of it for adults, and I did no safety tuning. It can write violent or sexual things, and it can be biased. Don't use it for anything important.

License and data

The weights and the code here are Apache 2.0. The training data is publicly available text. Some of it, Royal Road and BookCorpus in particular, belongs to its authors and has its own terms, so if you want to build something on these weights, check that it's fine for what you have in mind. The tokenizer vocabulary comes from tokenmonster's fiction-16000-strict-v1 (MIT).

Downloads last month
23
Safetensors
Model size
0.1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using lowerkey/bard-small 1