Instructions to use TheFinAI/FinLLaMA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TheFinAI/FinLLaMA with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="TheFinAI/FinLLaMA")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("TheFinAI/FinLLaMA") model = AutoModelForCausalLM.from_pretrained("TheFinAI/FinLLaMA", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use TheFinAI/FinLLaMA with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "TheFinAI/FinLLaMA" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TheFinAI/FinLLaMA", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/TheFinAI/FinLLaMA
- SGLang
How to use TheFinAI/FinLLaMA with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "TheFinAI/FinLLaMA" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TheFinAI/FinLLaMA", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "TheFinAI/FinLLaMA" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TheFinAI/FinLLaMA", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use TheFinAI/FinLLaMA with Docker Model Runner:
docker model run hf.co/TheFinAI/FinLLaMA
Install from pip and serve model
# Install vLLM from pip:
pip install vllm# Start the vLLM server:
vllm serve "TheFinAI/FinLLaMA"# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "TheFinAI/FinLLaMA",
"prompt": "Once upon a time,",
"max_tokens": 512,
"temperature": 0.5
}'Use Docker
docker model run hf.co/TheFinAI/FinLLaMAFinLLaMA
📄 Paper · 🤗 Collection · 🌐 The Fin AI
Part of Open-FinLLMs — Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications (arXiv:2408.11878).
FinLLaMA is a financial foundation model obtained by continual pre-training of LLaMA3-8B on a 52-billion-token financial corpus covering text, tables and time-series data from seven financial domains. It is a base (non-instruction-tuned) model: use it for further fine-tuning, few-shot prompting or research on domain adaptation.
Model Details
| Base model | LLaMA3-8B |
| Architecture | LlamaForCausalLM, 32 layers, hidden size 4096, vocab 128,256 |
| Context length | 8,192 tokens |
| Weights | float32 (torch_dtype in config.json) |
| Training data | 52B-token financial corpus (text, tables, time series) |
| Training | 1 epoch with DeepSpeed on 64 × A100 80GB (≈250 h); LR 1e-5, cosine schedule, weight decay 1e-5, warm-up ratio 0.05, batch size 2 per device, max sequence length 8,192 |
| License | Llama 3 Community License (inherited from the base model) |
Intermediate checkpoints
23 intermediate pre-training checkpoints are published as git tags, from step204 to step4692, so you can study how financial knowledge emerges during continual pre-training. Load one with revision:
from transformers import AutoModelForCausalLM, AutoTokenizer
ckpt = "step2448" # any of: step204, step408, step612, step816, step1020, step1224, step1428, step1632, step1836, step2040, step2244, step2448, step2652, step2856, step3060, step3264, step3468, step3672, step3876, step4079, step4284, step4488, step4692
tokenizer = AutoTokenizer.from_pretrained("TheFinAI/FinLLaMA", revision=ckpt)
model = AutoModelForCausalLM.from_pretrained("TheFinAI/FinLLaMA", revision=ckpt, torch_dtype="auto", device_map="auto")
The main branch holds the final model.
Quick Start
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("TheFinAI/FinLLaMA")
model = AutoModelForCausalLM.from_pretrained("TheFinAI/FinLLaMA", torch_dtype="auto", device_map="auto")
prompt = "The company's operating margin improved because"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
print(tokenizer.decode(model.generate(**inputs, max_new_tokens=64)[0], skip_special_tokens=True))
Related models
- FinLLaMA-instruct — FinLLaMA instruction-tuned on 573K financial instructions.
- FinLLaVA — multimodal model trained with 1.43M image–text instructions.
Evaluation
The paper evaluates FinLLaMA on 19 datasets (9 tasks) zero-shot and 4 datasets (3 tasks) few-shot, comparing with LLaMA3-8B, LLaMA3.1-8B and BloombergGPT, and in trading simulations (Sharpe ratio). See the paper for results.
Intended Use & Limitations
- Research on financial language modeling and as a starting point for fine-tuning.
- Not instruction-tuned: it continues text rather than following instructions.
- Not investment advice; may produce incorrect or outdated financial information.
Citation
@misc{huang2025openfinllmsopenmultimodallarge,
title={Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications},
author={Jimin Huang and Mengxi Xiao and Dong Li and Zihao Jiang and Yuzhe Yang and Yifei Zhang and Lingfei Qian and Yan Wang and Xueqing Peng and Yang Ren and Ruoyu Xiang and Zhengyu Chen and Xiao Zhang and Yueru He and Weiguang Han and Shunian Chen and Lihang Shen and Daniel Kim and Yangyang Yu and Yupeng Cao and Zhiyang Deng and Haohang Li and Duanyu Feng and Yongfu Dai and VijayaSai Somasundaram and Peng Lu and Guojun Xiong and Zhiwei Liu and Zheheng Luo and Zhiyuan Yao and Ruey-Ling Weng and Meikang Qiu and Kaleb E Smith and Honghai Yu and Yanzhao Lai and Min Peng and Jian-Yun Nie and Jordan W. Suchow and Xiao-Yang Liu and Benyou Wang and Alejandro Lopez-Lira and Qianqian Xie and Sophia Ananiadou and Junichi Tsujii},
year={2025},
eprint={2408.11878},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2408.11878},
}
- Downloads last month
- -
# Gated model: Login with a HF token with gated access permission hf auth login