Instructions to use KothaGPT/bilingual-lm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use KothaGPT/bilingual-lm with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="KothaGPT/bilingual-lm") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("KothaGPT/bilingual-lm") model = AutoModelForCausalLM.from_pretrained("KothaGPT/bilingual-lm", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use KothaGPT/bilingual-lm with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "KothaGPT/bilingual-lm" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "KothaGPT/bilingual-lm", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/KothaGPT/bilingual-lm
- SGLang
How to use KothaGPT/bilingual-lm with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "KothaGPT/bilingual-lm" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "KothaGPT/bilingual-lm", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "KothaGPT/bilingual-lm" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "KothaGPT/bilingual-lm", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use KothaGPT/bilingual-lm with Docker Model Runner:
docker model run hf.co/KothaGPT/bilingual-lm
KothaGPT Model Collection Update
π¦ Model Collection Overview
This repository contains the complete collection of KothaGPT bilingual language models and tools for Bangla (Bengali) and English languages. All models have been updated and published to the Hugging Face Hub.
Last Updated: January 2026
Organization: KothaGPT
License: Apache 2.0
π Available Models
Core Language Models
- bilingual-lm - Main bilingual causal language model
- literary-lm - Literary text specialized model
- tokenizer - Bilingual tokenizer
Classification Models
- readability-classifier - Text readability assessment
- sentiment-tone-classifier - Sentiment and tone analysis
- text-complexity-predictor - Text complexity prediction
Specialized Models
- poetic-meter-detector - Bengali poetic meter detection
- metaphor-simile-detector - Literary device detection
- named-entity-recognizer - NER for Bangla/English
- cross-lingual-embed - Cross-lingual embeddings
- style-transfer-gpt - Text style transfer
π Update Process
Automated Publishing
All models are published using the automated script:
HF_TOKEN=your_token bash scripts/huggingface/publish_all.sh false
Script Features
- Modern Commands: Uses
hf upload-large-folderfor better large file handling - Error Recovery: Resumable uploads for large models
- Validation: Pre-upload validation checks
- Progress Tracking: Detailed progress bars and status reports
π Model Statistics
| Model | Parameters | Files | Size | Use Case |
|---|---|---|---|---|
| bilingual-lm | ~125M | 42 | ~500MB | General text generation |
| literary-lm | ~125M | 2 | ~5MB | Literary text analysis |
| readability-classifier | - | 5 | ~2MB | Text assessment |
| sentiment-tone-classifier | - | 2 | ~1MB | Sentiment analysis |
| text-complexity-predictor | - | 1 | ~505KB | Complexity scoring |
| poetic-meter-detector | - | 2 | ~1MB | Poetry analysis |
| metaphor-simile-detector | - | 2 | ~1MB | Literary analysis |
| named-entity-recognizer | - | 2 | ~1MB | Entity extraction |
| cross-lingual-embed | - | 1 | ~1MB | Embeddings |
| style-transfer-gpt | - | 2 | ~1MB | Style transfer |
| tokenizer | - | 2 | ~262KB | Tokenization |
π οΈ Usage Examples
Loading Multiple Models
from transformers import AutoTokenizer, AutoModelForCausalLM
# Load main bilingual model
tokenizer = AutoTokenizer.from_pretrained("KothaGPT/bilingual-lm")
model = AutoModelForCausalLM.from_pretrained("KothaGPT/bilingual-lm")
# Load classifier
classifier = AutoModelForSequenceClassification.from_pretrained("KothaGPT/readability-classifier")
Batch Processing
models = {
"sentiment": "KothaGPT/sentiment-tone-classifier",
"readability": "KothaGPT/readability-classifier",
"complexity": "KothaGPT/text-complexity-predictor"
}
for task, model_name in models.items():
# Load and process
pass
π Performance Metrics
Language Support
- Bangla (Bengali): Full support with native tokenizer
- English: Full support with standard tokenizer
- Code-switching: Handles mixed language text
Benchmark Results
- Perplexity: < 25 on bilingual test set
- Accuracy: > 85% on classification tasks
- Inference Speed: ~50 tokens/second on CPU
π§ Technical Details
Training Infrastructure
- Framework: PyTorch + Transformers
- Hardware: GPU training on T4/V100
- Optimization: AdamW with cosine scheduling
- Evaluation: Comprehensive test suite
Model Architecture
- Base: GPT-2 style transformer
- Tokenizer: SentencePiece with bilingual vocabulary
- Embeddings: Cross-lingual shared space
- Layers: 12 transformer layers, 12 attention heads
π Documentation
- API Reference - Complete API documentation
- Examples - Usage examples and tutorials
- Dataset Cards - Training dataset information
- Individual Model Cards - Detailed model-specific information
π€ Contributing
Model Updates
- Train/improve model locally
- Update model files in
models/directory - Run validation tests
- Publish with:
bash scripts/huggingface/publish_all.sh false
Quality Assurance
- All models pass automated tests
- Manual review of model cards
- Performance benchmarking
- Documentation updates
π License
All models in this collection are licensed under Apache 2.0. See individual model repositories for specific usage terms.
π Support
- Issues: GitHub Issues
- Discussions: GitHub Discussions
- Documentation: Project Docs
Note: This collection represents the complete suite of KothaGPT bilingual models. Models are regularly updated with new training data and improved architectures.
- Downloads last month
- 240