YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Bharosa

Bharosa is a ~140M-parameter decoder-only language model trained entirely from scratch. It is designed as a compact general-purpose model emphasizing language understanding, commonsense knowledge, educational content, mathematics, reasoning, and code-related knowledge. Bharosa was pretrained for approximately 50 billion tokens using a capacity-aware curriculum and a diverse dataset mixture.

Bharosa is a base language model, not an instruction-tuned or chat model. It is intended primarily for research, text continuation, evaluation, local inference, and downstream fine-tuning.

Model Overview

Property Value
Model Bharosa
Model type Decoder-only causal language model
Parameters ~140M (138.97M actual)
Training tokens ~50B
Layers 24
Hidden size 640
Intermediate size 1,920
Attention heads 8 (Query) / 4 (KV)
Head dimension 80
Vocabulary size 32,768
Context length 3,072 tokens
Attention Grouped-Query Attention (GQA)
Attention normalization QK Normalization
Position encoding RoPE (\theta = 100,000)
MLP SwiGLU
Normalization RMSNorm (\epsilon = 1\text{e-}6)
Embeddings Tied input/output embeddings
Precision bfloat16
Weight format Safetensors
Framework PyTorch + Transformers
Benchmark Accuracy
:--- :---
ARC Easy 59.01%
ARC Challenge 25.51%
PIQA 67.63%
HellaSwag 34.94%
Winogrande 51.54%
OpenBookQA 20.80%

The benchmarks demonstrate strong physical-world commonsense (PIQA) and general knowledge relative to parameter size, alongside expected performance bottlenecks in multi-step reasoning and complex QA. Usage & Code Examples Install dependencies:

pip install -U torch transformers safetensors

Standard GPU Inference

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "pihu21057w/bharosa"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
dtype = torch.bfloat16 if torch.cuda.is_available() and torch.cuda.is_bf16_supported() else torch.float32
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=dtype
).to("cuda" if torch.cuda.is_available() else "cpu")
model.eval()
# Prompt format should be text completion style
prompt = "The capital of France is"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
    output = model.generate(
        **inputs,
        max_new_tokens=80,
        do_sample=True,
        temperature=0.7,
        top_p=0.9,
        repetition_penalty=1.1,
        use_cache=True,
    )
print(tokenizer.decode(output[0], skip_special_tokens=True))

Deterministic Generation

with torch.no_grad():
    output = model.generate(
        **inputs,
        max_new_tokens=80,
        do_sample=False,
        use_cache=True,
    )
print(tokenizer.decode(output[0], skip_special_tokens=True))

Limitations & Safety

  • Base Model Constraints: As an unaligned pretrained base model, Bharosa predicts text continuations and will not consistently format outputs as a conversational assistant.
  • Capacity Bounds: Complex reasoning, multi-step math, coreference resolution, and strict factual reliability are limited by scale (~140M parameters).
  • Safety: The model has not undergone RLHF, preference alignment, or refusal training. Deployments require appropriate input/output filtering mechanisms. Citation
@misc{bharosa2026,
  title        = {Bharosa: A 140M Parameter Language Model Trained on 50B Tokens},
  author       = {BananaMind},
  year         = {2026},
  publisher    = {Hugging Face},
  url          = {https://huggingface.co/pihu21057w/bharosa}
}

License Released under the Apache License 2.0. Copyright © 2026 Bharosa contributors.

Downloads last month
40
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support