Instructions to use SurjoLabs/Surjo-100m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SurjoLabs/Surjo-100m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="SurjoLabs/Surjo-100m", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("SurjoLabs/Surjo-100m", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SurjoLabs/Surjo-100m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SurjoLabs/Surjo-100m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SurjoLabs/Surjo-100m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/SurjoLabs/Surjo-100m
- SGLang
How to use SurjoLabs/Surjo-100m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SurjoLabs/Surjo-100m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SurjoLabs/Surjo-100m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SurjoLabs/Surjo-100m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SurjoLabs/Surjo-100m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use SurjoLabs/Surjo-100m with Docker Model Runner:
docker model run hf.co/SurjoLabs/Surjo-100m
Architecture questions
A couple questions, as I'm currently working on my own model(s).
- Have you performed any ablation studies i.e. what are the advantages of your architecture vs. a naïve approach of just using a basic transformer?
- Why GDN vs. once again, the naïve approach of full attention? You're not dealing with a very large context window so the quadratic blowup at these scales likely isn't extreme and you may be sacrificing representational ability at the scales you're working with while adding computational complexity, reducing throughput and thus able to train on fewer tokens (and overtraining small models seems to be the go-to approach with memory limitations).
I ask because I think we're training on similar hardware and at similar scales and I'm frankly a bit bored by the idea of "attention is all you need" but paradoxically, the approach of "full transformer" may be the smartest choice at this scale as much as I don't like that answer. But I could be completely mistaken here; these are honest questions.
- We did not do extensive ablation studies for this generation. However we did studies to ensure recurrent layers which is the backbone of our architecture is better when you are memory limited or parameter limited.
- Our current models are meant to be experiments for a larger model which will natively support large contexts. Using it here gave us some experience with GDN-2. Also GDN-2 does help during training even on lower context windows.
We did make some mistakes when designing the architecture for this model. We aim to fix them in Surjo2
What have you noticed that it does with smaller contexts? My understanding is that attention strictly has more representational capacity at the expensive of tracking a lot more things at once (hence the memory blowup)
The cost of hybrid attentions are O(N) while not using it would be O(N²)
If we were to change the seq len of a training run from 1024 to 2048 normally we would have to square root the micro batch size to fit in memory. However in our runs we just have to halve it.