Instructions to use dreamidotme/fable2notbyclaude with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use dreamidotme/fable2notbyclaude with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf dreamidotme/fable2notbyclaude:F16 # Run inference directly in the terminal: llama cli -hf dreamidotme/fable2notbyclaude:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf dreamidotme/fable2notbyclaude:F16 # Run inference directly in the terminal: llama cli -hf dreamidotme/fable2notbyclaude:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf dreamidotme/fable2notbyclaude:F16 # Run inference directly in the terminal: ./llama-cli -hf dreamidotme/fable2notbyclaude:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf dreamidotme/fable2notbyclaude:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf dreamidotme/fable2notbyclaude:F16
Use Docker
docker model run hf.co/dreamidotme/fable2notbyclaude:F16
- LM Studio
- Jan
- vLLM
How to use dreamidotme/fable2notbyclaude with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dreamidotme/fable2notbyclaude" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dreamidotme/fable2notbyclaude", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dreamidotme/fable2notbyclaude:F16
- Ollama
How to use dreamidotme/fable2notbyclaude with Ollama:
ollama run hf.co/dreamidotme/fable2notbyclaude:F16
- Unsloth Desktop
- Docker Model Runner
How to use dreamidotme/fable2notbyclaude with Docker Model Runner:
docker model run hf.co/dreamidotme/fable2notbyclaude:F16
- Lemonade
How to use dreamidotme/fable2notbyclaude with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull dreamidotme/fable2notbyclaude:F16
Run and chat with the model
lemonade run user.fable2notbyclaude-F16
List all available models
lemonade list
- Atomic Chat
Why
This model is a joke, a parody making fun of how Anthropic made a model that won't reply by default in fables, so this joke AI remedied that.
Fable 2
A 89m-parameter language model trained from scratch on public-domain fables and folklore. Not a fine-tune, not a distillation, not a LoRA on somebody else's base β the weights start from random init and the whole corpus is out of copyright. Compliant with EU article 50. Runs on a laptop CPU. No GPU, no API, no account.
What it does well, and what it does not
It writes fable prose with the right cadence, and it produces genuinely well-formed morals β "Presence of mind and quick thinking can save you from treachery" is real output.
It mixes fables up. It will hand the cheese to a Deer, or put the Crow at the Fox's dinner table. it has fables vocabulary and rhythm without reliable bindings between characters and their stories. That is the honest ceiling of a model this size, not a bug to report.
Use it for a laugh. Do not use it as a reference for what any particular fable actually says.
Running it
llama.cpp β the prompt template is baked into the GGUF, so conversation mode needs no configuration:
llama-cli -m fable-2-f16.gguf -cnv
LM Studio / Ollama / any GGUF runner β load the file and go.
Raw prompting, if you are driving it programmatically. Match this exactly; off-template the model reverts to continuing a story instead of answering:
### Instruction:
What is the moral of The Fox and the Grapes?
### Response:
Generation stops at EOS (<|endoftext|>, id 50256). Replies are short by
design β the training targets have a median of 26 words.
Specification
| Parameters | 89 M |
| Architecture | GPT-2 style β learned positional embeddings, LayerNorm, GELU MLP, fused QKV, weight-tied embeddings |
| Layers / heads / width | 12 / 8 / 512 |
| Context | 512 tokens |
| Tokenizer | GPT-2 BPE, vocab 50257 |
| Biases | none (trained with bias off; the GGUF carries explicit zeros, which llama.cpp's gpt2 graph requires) |
| GGUF arch tag | gpt2 |
| Precision | f16 |
Training data
Public-domain texts from Project Gutenberg β multiple Aesop editions plus other out-of-copyright folklore and period fiction β with an instruction-formatted fable dataset folded in.
Everything the model saw is in the public domain. That is the point of the project, not an afterthought.
Licence
CC0 1.0 β public domain dedication. Do whatever you want with the weights.
Transparency
Output from this model is machine-generated.
The GGUF applies an invisible watermark to replies.
Limitations and risks
- Confidently misattributes fables (see above). Do not cite it.
- 512-token context, hard limit.
- English only.
- Trained on 19th and early-20th century public-domain text, and carries the assumptions and language of that period.
- No safety tuning of any kind.
- Downloads last month
- 17
16-bit