Instructions to use paradigma-inc/limite-1b-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use paradigma-inc/limite-1b-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="paradigma-inc/limite-1b-base", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("paradigma-inc/limite-1b-base", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use paradigma-inc/limite-1b-base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "paradigma-inc/limite-1b-base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "paradigma-inc/limite-1b-base", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/paradigma-inc/limite-1b-base
- SGLang
How to use paradigma-inc/limite-1b-base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "paradigma-inc/limite-1b-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "paradigma-inc/limite-1b-base", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "paradigma-inc/limite-1b-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "paradigma-inc/limite-1b-base", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use paradigma-inc/limite-1b-base with Docker Model Runner:
docker model run hf.co/paradigma-inc/limite-1b-base
Limite 1B - Base
Limite 1B - Base is the post-pretraining checkpoint of the Limite 1B model family.
This model card focuses on loading and running the checkpoint with Hugging Face Transformers. For the latest checkpoint and its full model card, see Limite 1B - Violetto.
Run with Transformers
Limite 1B - Base supports inference through the standard Hugging Face Transformers APIs. The custom architecture code is downloaded from this repository, so loading the model requires trust_remote_code=True.
Validated on an NVIDIA H100 with Python 3.12 · PyTorch 2.11.0 (CUDA 13.0) · Transformers 5.6.2. Other version and hardware combinations have not yet been formally qualified. SDPA is the portable default.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "paradigma-inc/limite-1b-base"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype="auto",
device_map="auto",
attn_implementation="sdpa",
)
messages = [{"role": "user", "content": "Solve: If x + 3 = 8, what is x?"}]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
with torch.inference_mode():
output_ids = model.generate(**inputs, max_new_tokens=256)
answer = tokenizer.decode(
output_ids[0, inputs.input_ids.shape[1]:],
skip_special_tokens=True,
)
print(answer)
Supported inference paths
| Attention backend | Cache | Status |
|---|---|---|
| SDPA | DynamicCache |
Supported |
| SDPA | mixed global/sliding-window StaticCache |
Supported |
| SDPA | StaticCache + torch.compile |
Supported |
| FlashAttention 2 | DynamicCache |
Supported |
| FlashAttention 2 | StaticCache |
Unsupported; rejected with an explicit error |
FlashAttention 2 with StaticCache is intentionally rejected because that combination does not produce numerically correct logits for Limite's hybrid local/global attention layout. Use SDPA with StaticCache, or FlashAttention 2 with DynamicCache.
FlashAttention 2 was validated through Transformers' kernels-community/flash-attn2 integration with kernels==0.12.3. The separately installed native flash_attn package has not been independently qualified.
Official support currently covers inference. The Transformers implementation is differentiable and exposes the standard causal-language-model loss, but full training, gradient checkpointing, PEFT/LoRA, and distributed-training workflows have not yet been formally validated and are not part of the supported interface.
License
The model weights are released under Apache-2.0.
- Downloads last month
- 650