RivaQuant

A small (162M-param), from-scratch decoder-only transformer with BitNet b1.58 ternary weights ({-1, 0, 1}) in every attention/MLP projection, trained on TinyStories. Architecture and training code: entropy-om/rivaquant.

RivaQuant is the project/repo name, not a self-chosen identity โ€” see "Does it know its own name?" below for why.

What this actually is

  • Architecture: BitNet b1.58 ternary linear layers (weight quantization via absmean + sign, per-token int8 activation quantization, straight-through estimator for gradients), RoPE, standard decoder-only transformer block structure otherwise. Adapted from real reference implementations (kyegomez/BitNet's math, Microsoft's BitNet b1.58 paper), not reinvented from memory.
  • Params: 162,213,888
  • Trained: 17,500 steps (best checkpoint by validation loss, out of 20,000 total steps run) on TinyStories, batch size 8 with 4x gradient accumulation (effective batch 32), block size 256.
  • Validation perplexity: 5.455

What this is not

Not a general-purpose assistant, not instruction-tuned, not evaluated on anything beyond TinyStories-style short story completion. Read the actual architecture and eval code before drawing conclusions from perplexity alone.

Does it know its own name?

No. Prompted with "My name is" / "I am called" / "You can call me" (15 samples, unfiltered), it produces a different plausible children's-story character name every time โ€” a direct artifact of training purely on TinyStories, which is full of characters introducing themselves that way. There is no consistent self-identity to report, so none is claimed:

  • `My name is Max. What is your name?"

Lily said, "That is Max, the owl`

  • `My name is Daisy. I like your hat."

Mia smiled and said, "That is very nice`

  • `My name is Lily. I live in a small village. Do you want to play with me?"

L`

  • `My name is Lily," said Lily.

The boy said, "I'm happy to meet you. I`

  • `My name is Ben. What is your name?"

Ben says, "I am Ben, and this is`

  • `I am called it. I love you. You are my best friend."

She hugs Ben and says,`

  • I am called Tim. I like to play with you." They played and had fun together. Lily learned that sharing
  • `I am called a horse. I am looking for a home. Do you want to come with me?"

`

  • I am called Tom. I want to play with you. But you have to wait for me. I am taking
  • `I am called an elephant. I like to eat honey. Do you want a banana?"

Lily and`

  • `You can call me if you want. I will give you a hug."

Lily smiled and said, "`

  • `You can call me. I have a phone. It is called my phone. Do you like it?"

Ben`

  • `You can call me and get up, but please stay down," said the dog.

Mama smiled and said`

  • `You can call me and ask me if you want to come. There's no need to do you."

Anna`

  • `You can call me, or I will call your dad."

Lily and Ben went back to the fence.`

Training bugs hit and fixed along the way

  1. torch.nn.RMSNorm requires torch>=2.4; the training environment shipped 2.1.0. Custom RMSNorm module, same math, no version dependency.
  2. CUDA OOM at batch=32/block=512 on a 24GB card โ€” BitNet's straight-through estimator keeps extra activation copies per layer, more memory-hungry than plain nn.Linear at the same nominal size. Fixed with a smaller micro-batch + gradient accumulation to keep the same effective batch.

Both caught mid-training by an automated cost-safety watcher that stops the GPU on any crash signature, understood, fixed, verified with a smoke test, and relaunched โ€” not discovered after the fact.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support