roneneldan/TinyStories
Viewer โข Updated โข 2.14M โข 106k โข 1.18k
How to use thekosmix/kids-scroll-models with llama.cpp:
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf thekosmix/kids-scroll-models # Run inference directly in the terminal: llama cli -hf thekosmix/kids-scroll-models
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf thekosmix/kids-scroll-models # Run inference directly in the terminal: llama cli -hf thekosmix/kids-scroll-models
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf thekosmix/kids-scroll-models # Run inference directly in the terminal: ./llama-cli -hf thekosmix/kids-scroll-models
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf thekosmix/kids-scroll-models # Run inference directly in the terminal: ./build/bin/llama-cli -hf thekosmix/kids-scroll-models
docker model run hf.co/thekosmix/kids-scroll-models
How to use thekosmix/kids-scroll-models with vLLM:
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "thekosmix/kids-scroll-models"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "thekosmix/kids-scroll-models",
"prompt": "Once upon a time,",
"max_tokens": 512,
"temperature": 0.5
}'docker model run hf.co/thekosmix/kids-scroll-models
How to use thekosmix/kids-scroll-models with Ollama:
ollama run hf.co/thekosmix/kids-scroll-models
How to use thekosmix/kids-scroll-models with Docker Model Runner:
docker model run hf.co/thekosmix/kids-scroll-models
How to use thekosmix/kids-scroll-models with Lemonade:
# Download Lemonade from https://lemonade-server.ai/ lemonade pull thekosmix/kids-scroll-models
lemonade run user.kids-scroll-models-{{QUANT_TAG}}lemonade list
This repository hosts the on-device story generation model (story-model.gguf) and WebAssembly runtime (wllama.wasm) used by Kids Scroll โ a lightweight, offline-first Progressive Web App (PWA) designed for toddlers and young children.
story-model.gguf (~26 MB)wllama.wasm (~7.4 MB)roneneldan/TinyStories consisting of ~2.1 million synthetic short stories generated by GPT-3.5 / GPT-4.wllama (WebAssembly)
import { Wllama } from '@wllama/wllama';
const wllama = new Wllama({
'wllama.wasm': 'https://huggingface.co/thekosmix/kids-scroll-models/resolve/main/wllama.wasm'
});
// Load the model
const response = await fetch('https://huggingface.co/thekosmix/kids-scroll-models/resolve/main/story-model.gguf');
const blob = await response.blob();
await wllama.loadModel([blob], { n_ctx: 192 });
// Generate a story segment
const output = await wllama.createCompletion({
prompt: 'Once upon a time, there was a little rabbit.',
max_tokens: 100,
temperature: 0.8,
top_p: 0.9,
stop: ['\n\n', 'The end.']
});
console.log(output.choices[0].text);
llama-cli -m story-model.gguf -p "Once upon a time, there was a friendly lion." -n 100 --temp 0.8
If you use this model or dataset, please cite the original TinyStories research:
@article{eldan2023tinystories,
title={TinyStories: How Small Can Language Models Be and Still Speak Coherent English?},
author={Eldan, Ronen and Li, Yuanzhi},
journal={arXiv preprint arXiv:2305.07759},
year={2023}
}
We're not able to determine the quantization variants.
Base model
roneneldan/TinyStories-33M