Text Generation
Transformers
English
phi3
finance
entity-extraction
ner
phi-3
production
indian-banking
custom_code
4-bit precision
Instructions to use Ranjit0034/finance-entity-extractor with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Ranjit0034/finance-entity-extractor with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Ranjit0034/finance-entity-extractor", trust_remote_code=True)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Ranjit0034/finance-entity-extractor", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("Ranjit0034/finance-entity-extractor", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Ranjit0034/finance-entity-extractor with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Ranjit0034/finance-entity-extractor" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ranjit0034/finance-entity-extractor", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Ranjit0034/finance-entity-extractor
- SGLang
How to use Ranjit0034/finance-entity-extractor with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Ranjit0034/finance-entity-extractor" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ranjit0034/finance-entity-extractor", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Ranjit0034/finance-entity-extractor" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ranjit0034/finance-entity-extractor", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Ranjit0034/finance-entity-extractor with Docker Model Runner:
docker model run hf.co/Ranjit0034/finance-entity-extractor
| """ | |
| Tests for FinEE Cache (Tier 0). | |
| """ | |
| import time | |
| from finee.cache import LRUCache, ExtractionResult | |
| def test_cache_hashing(): | |
| cache = LRUCache() | |
| text = " Rs.500 spent " | |
| # Should normalize whitespace and case | |
| key1 = cache.hash_text("Rs.500 spent") | |
| key2 = cache.hash_text("rs.500 SPENT") | |
| assert key1 == key2 | |
| def test_cache_operations(): | |
| cache = LRUCache(max_size=2) | |
| # Add item | |
| res1 = ExtractionResult(amount=100.0) | |
| cache.set("tx1", res1) | |
| # Get item | |
| cached = cache.get("tx1") | |
| assert cached.amount == 100.0 | |
| assert cached.from_cache is True | |
| # Check stats | |
| stats = cache.get_stats() | |
| assert stats.hits == 1 | |
| assert stats.size == 1 | |
| def test_lru_eviction(): | |
| cache = LRUCache(max_size=2) | |
| # Fill cache | |
| cache.set("tx1", ExtractionResult(amount=1)) | |
| cache.set("tx2", ExtractionResult(amount=2)) | |
| # Access tx1 to make it recent | |
| cache.get("tx1") | |
| # Add 3rd item (should evict tx2, because tx1 was just used) | |
| cache.set("tx3", ExtractionResult(amount=3)) | |
| assert cache.contains("tx1") # Kept | |
| assert cache.contains("tx3") # New | |
| assert not cache.contains("tx2") # Evicted | |
| def test_cache_threading_safety(): | |
| # Basic check ensuring no crash on rapid updates | |
| cache = LRUCache(max_size=100) | |
| for i in range(200): | |
| cache.set(f"tx{i}", ExtractionResult(amount=i)) | |
| assert len(cache) == 100 | |
| assert cache.get("tx199").amount == 199.0 | |