Instructions to use CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE", filename="gguf_releases/DeepSeek-R1-Distill-Llama-8B.F16.gguf", )
llm.create_chat_completion( messages = "No input example has been defined for this model task." )
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE:F16 # Run inference directly in the terminal: llama cli -hf CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE:F16 # Run inference directly in the terminal: llama cli -hf CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE:F16 # Run inference directly in the terminal: ./llama-cli -hf CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE:F16
Use Docker
docker model run hf.co/CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE:F16
- LM Studio
- Jan
- Ollama
How to use CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE with Ollama:
ollama run hf.co/CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE:F16
- Unsloth Studio
How to use CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE to start chatting
- Atomic Chat new
- Docker Model Runner
How to use CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE with Docker Model Runner:
docker model run hf.co/CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE:F16
- Lemonade
How to use CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE:F16
Run and chat with the model
lemonade run user.DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE-F16
List all available models
lemonade list
- DeepSeek-R1-Llama-8B Indian Finance Quant (v2.1 RELEASE)
DeepSeek-R1-Llama-8B Indian Finance Quant (v2.1 RELEASE)
Codename: Celeste v2.1 (Dual-Brain Architecture)
Celeste v2.1 is a high-fidelity reasoning model fine-tuned specifically for the Indian Equity Market. Built on the DeepSeek-R1-Distilled-Llama-8B architecture, it combines advanced technical analysis with deep regulatory and legal awareness (SEBI/RBI/Income Tax).
🧠 Why Celeste v2.1?
While base models provide general financial advice, Celeste v2.1 is trained to think like a Senior Indian Quant Analyst. It utilizes a native <think> chain to weigh technical indicators against macro-catalysts and institutional risks.
📊 Performance Benchmark: The "Ethics" Stress Test
In a head-to-head comparison on April 2, 2026, regarding HDFC Bank's leadership crisis:
- Base Model (DeepSeek-R1): Suggested a "Value Buy" based solely on 52-week lows and RSI.
- Celeste v2.1: Corrected identified the Chairman's resignation over 'ethics differences' as a primary risk, issuing a Cautionary/Avoid verdict despite the "cheap" price.
🛠 Model Capabilities
- Technical Synthesis: RSI, EMA Crosses, Volume Profile, and Mean Reversion analysis.
- Legal & Regulatory Awareness: Trained on SEBI circulars and April 2026 Indian tax updates (STT & Buyback changes).
- Dual-Brain Logic: Separates internal "Reasoning" from the "Final Verdict" for maximum transparency.
📁 Available Formats
- Merged 16-bit (Safetensors): The full-fidelity weights for professional inference.
- F16 GGUF: Maximum precision for local power users (LM Studio/Ollama).
- Q8_0 GGUF: Balanced efficiency for 12GB-16GB VRAM hardware.
🚀 Usage (Transformers)
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
torch_dtype=torch.bfloat16
)
prompt = "Analyze PC Jeweller (PCJEWELLER) given it is trading below 200-day EMA with high retail volume."
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
💻 For C# / .NET Users (LLamaSharp Implementation)
As a Senior .NET developer, I have validated this model for local C# integration using LLamaSharp. This is ideal for building local Indian Finance desktop tools or integrating with existing .NET portfolio trackers.
using LLama.Common;
using LLama;
// Initialize the model with local GGUF path
var parameters = new ModelParams("Celeste_v2.1_Q8_0.gguf")
{
ContextSize = 4096,
GpuLayerCount = 33 // Fully offload to your 3090/A4000
};
using var weights = LLamaWeights.LoadFromFile(parameters);
using var context = weights.CreateContext(parameters);
var executor = new InteractiveExecutor(context);
var session = new ChatSession(executor);
Console.WriteLine("🧠 Celeste is analyzing (Local C# Inference)...");
var prompt = "Explain the impact of the April 2026 STT hike on intraday trading margins.";
await foreach (var text in session.ChatAsync(
new ChatHistory.Message(AuthorRole.User, prompt),
new InferenceParams { Temperature = 0.7f }))
{
Console.Write(text);
}
📜 v2.1 Platinum Release: Change Log
The v2.1 Platinum edition represents a significant architectural shift from the initial v2.0 release, moving from general fine-tuning to specialized financial reasoning.
🛠️ Key Technical Upgrades:
- Logic Engine: Transitioned from Standard Prompting to Chain of Thought (CoT) Reasoning, allowing the model to "think through" banking ratios before providing a final verdict.
- Dataset Curation: Replaced raw transcripts with Expert-Labelled Financial Pairs, focusing on Nifty 50 quarterly guidance and SEBI 2025/2026 regulatory shifts.
- Optimization: Fully integrated with the Unsloth Engine, resulting in 2.3x faster inference and 70% less VRAM usage compared to v1.0.
- Banking Vertical: Specialized training on Net Interest Margin (NIM), CASA ratios, and Loan-to-Deposit (LDR) dynamics unique to the Indian private and PSU banking sectors.
- Context Stability: Rock-solid performance up to 8k tokens, enabling the analysis of full "Management Discussion & Analysis" sections.
🎯 Sample Prompting Guide (Gold Standard Tests)
To see the v2.1 Platinum difference, try these complex "Institutional-Grade" queries. These are designed to test the model's ability to reason, not just retrieve facts.
Test Case 1: Banking NIM Pressure
Prompt: "A private sector bank reports a 10% increase in credit growth but a 20% spike in bulk deposit reliance. How will this impact their NIM in the next two quarters?"
Expected Reasoning: The model should identify that a reliance on high-cost bulk deposits, despite credit growth, will lead to a "cost of funds" spike that outpaces "yield on advances," resulting in compressed Net Interest Margins.
Test Case 2: Regulatory Impact (SEBI/RBI)
Prompt: "Analyze the impact of the latest RBI circular on 'Unsecured Lending' risk weights for a mid-sized Indian NBFC."
Expected Reasoning: The model should reason that higher risk weights lead to a lower Capital Adequacy Ratio (CAR), forcing the NBFC to either raise fresh equity capital or slow down its high-yield personal loan book.
Test Case 3: Corporate Actions & Sentiment
Prompt: "A Nifty 50 company announces a 1:10 stock split alongside a surprise 50% dividend hike. What is the likely short-term impact on retail liquidity and institutional sentiment?"
Expected Reasoning: It should differentiate between the "psychological liquidity" boost for retail investors (due to the split) and the "strong cash-flow signal" for institutions (due to the dividend), likely leading to an accumulation phase.
🏗️ Technical Forge & Infrastructure
- Model Type: LoRA Adapter (PEFT / Unsloth)
- Architecture: DeepSeek-R1-Llama-8B (v2.1 Platinum Release)
- Training Workstation: Dual-GPU (NVIDIA RTX 3090 24GB + RTX A4000 16GB)
- Memory: 64GB DDR4
- Engine: Unsloth (Optimized for zero-latency financial reasoning)
📜 License & Disclaimer
License: This project is licensed under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0).
What this means:
- Attribution: You must give appropriate credit to Celeste Imperia (Abhishek Jaiswal).
- Non-Commercial: You may not use this model or its outputs for commercial purposes.
- ShareAlike: If you remix or build upon this work, you must distribute your contributions under the same license.
Disclaimer: This model is provided as an educational and research tool only. It does not constitute financial advice. Financial markets involve significant risk. Always consult a SEBI-registered professional before making any investment decisions. Celeste Imperia and its architects are not liable for any financial losses incurred through the use of this AI.
☕ Support the Forge
Maintaining a dual-GPU AI workstation and hosting high-bandwidth models requires significant resources. If our open-source tools power your projects, consider supporting our development:
| Platform | Support Link |
|---|---|
| Global & India | Support via Razorpay |
Scan to support via UPI (India Only):
Connect with the architect: Abhishek Jaiswal on LinkedIn
- Downloads last month
- 7
8-bit
16-bit
Model tree for CelesteImperia/DeepSeek-R1-Llama-8B-Indian-Finance-Quant-v2.1-RELEASE
Base model
deepseek-ai/DeepSeek-R1-Distill-Llama-8B