Text Generation
Transformers
Safetensors
English
German
multilingual
k2_horizon
ocean-horizon
3.7b
dense
open-weights
oceanlabs
long-context
reasoning
agentic
custom_code
Instructions to use OceanLabs/Ocean-3.7B-coding with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OceanLabs/Ocean-3.7B-coding with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="OceanLabs/Ocean-3.7B-coding", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("OceanLabs/Ocean-3.7B-coding", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use OceanLabs/Ocean-3.7B-coding with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "OceanLabs/Ocean-3.7B-coding" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OceanLabs/Ocean-3.7B-coding", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/OceanLabs/Ocean-3.7B-coding
- SGLang
How to use OceanLabs/Ocean-3.7B-coding with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "OceanLabs/Ocean-3.7B-coding" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OceanLabs/Ocean-3.7B-coding", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "OceanLabs/Ocean-3.7B-coding" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OceanLabs/Ocean-3.7B-coding", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use OceanLabs/Ocean-3.7B-coding with Docker Model Runner:
docker model run hf.co/OceanLabs/Ocean-3.7B-coding
|
Download README.md from OceanLabs/Ocean-3.7B-coding: direct link, hf CLI and curl.
- Browser
- Download file 5.21 kB
-
https://huggingface.co/OceanLabs/Ocean-3.7B-coding/resolve/main/README.md
- Command line
-
hf download hf://OceanLabs/Ocean-3.7B-coding/README.md
-
curl -L -o README.md https://huggingface.co/OceanLabs/Ocean-3.7B-coding/resolve/main/README.md
5.21 kB
| pipeline_tag: text-generation | |
| library_name: transformers | |
| model_name: Ocean-Horizon-3.7B | |
| language: | |
| - en | |
| - de | |
| - multilingual | |
| license: apache-2.0 | |
| tags: | |
| - ocean-horizon | |
| - 3.7b | |
| - dense | |
| - open-weights | |
| - oceanlabs | |
| - long-context | |
| - reasoning | |
| - agentic | |
| # Ocean-Horizon-3.7B | |
| **OceanLabs / Ocean-Horizon-3.7B** is a high-performance 3.7B-parameter dense decoder-only language model optimized for advanced reasoning, coding, agentic workflows, and ultra-long context understanding. | |
| Built upon a rigorously engineered dense architecture with a native **524,288-token (512K)** context window, Ocean-Horizon-3.7B delivers frontier-level capabilities in a compact, efficient package. | |
| ## Key Capabilities | |
| - **Exceptional small-model performance.** Strong results across agentic, coding, mathematical, and scientific reasoning benchmarks. | |
| - **Native 512K context.** Full 524,288-token context window available from mid-training stages onward, enabling deep document analysis, long-horizon planning, and complex multi-turn interactions. | |
| - **Optimized reasoning depth.** Supports configurable reasoning effort (high / medium / low) with dedicated thinking tokens for transparent chain-of-thought. | |
| - **Production-ready tool use.** Native support for structured tool calling with multiple presentation formats (JSON, XML, Markdown). | |
| - **Fully open.** Architecture, configuration, and evaluation resources are openly available for research and commercial use under Apache 2.0. | |
| ## Architecture Summary | |
| | Parameter | Value | | |
| |-----------------------------|----------------| | |
| | Parameters | 3.7B | | |
| | Architecture | Dense | | |
| | Hidden size | 2560 | | |
| | Intermediate size | 10240 | | |
| | Layers | 36 | | |
| | Attention heads | 32 | | |
| | Key-value heads | 8 (GQA) | | |
| | Head dimension | 128 | | |
| | Vocabulary size | 250,624 | | |
| | Context length | 524,288 | | |
| | Activation | SiLU | | |
| | Precision | bfloat16 | | |
| | RoPE θ | 10,000,000 | | |
| ## Recommended Inference Settings | |
| - **Reasoning effort:** always prefer `"high"` for maximum quality. | |
| - **Sampling:** `temperature=1.0`, `top_p=0.95`. | |
| - **Maximum output tokens:** at least 32,768 to avoid truncating reasoning traces. | |
| - **Serving backends:** Validated with vLLM and SGLang (BF16, FlashAttention-3 recommended). | |
| ### Example with Transformers | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "OceanLabs/Ocean-Horizon-3.7B" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, | |
| device_map="auto", | |
| dtype="bfloat16", | |
| low_cpu_mem_usage=True, | |
| trust_remote_code=True | |
| ) | |
| inputs = tokenizer("Explain the advantages of a 512K context window for agentic systems.", return_tensors="pt").to(model.device) | |
| inputs.pop("token_type_ids", None) | |
| outputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True) | |
| print(tokenizer.decode(outputs[0], skip_special_tokens=True)) | |
| ``` | |
| ### Example with OpenAI-compatible API (vLLM / SGLang) | |
| ```python | |
| from openai import OpenAI | |
| client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY") | |
| response = client.chat.completions.create( | |
| model="OceanLabs/Ocean-Horizon-3.7B", | |
| messages=[{"role": "user", "content": "Solve this step by step."}], | |
| temperature=1.0, | |
| top_p=0.95, | |
| max_tokens=32768, | |
| extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}}, | |
| ) | |
| message = response.choices[0].message | |
| print("Reasoning:", getattr(message, "reasoning_content", None)) | |
| print("Answer:", message.content) | |
| ``` | |
| ## Training Pipeline Overview | |
| Ocean-Horizon-3.7B follows a multi-stage curriculum: | |
| 1. **Pretraining** – large-scale general knowledge acquisition (22.9T tokens). | |
| 2. **Midtraining** – progressive context extension (32K → 128K → 512K) plus agentic and reasoning data. | |
| 3. **Reinforcement Learning** – specialized experts for mathematics, code, and STEM-code, followed by model merging. | |
| 4. **Supervised Fine-Tuning** – high-quality multi-domain instruction tuning with learning-rate decay. | |
| This staged approach yields strong generalization while preserving long-context fidelity. | |
| ## Best Practices | |
| 1. Always request high reasoning effort for evaluation and complex tasks. | |
| 2. Allocate sufficient output length (≥ 32k tokens) so that reasoning is never truncated. | |
| 3. Use the native `k2_horizon` reasoning and tool-call parsers when serving with vLLM or SGLang. | |
| 4. Pin a specific revision or commit hash for reproducible deployments. | |
| ## License | |
| Apache License 2.0 | |
| ## Citation | |
| ```bibtex | |
| @misc{oceanhorizon2026, | |
| title = {Ocean-Horizon-3.7B: Compact Dense Model with Frontier Reasoning and 512K Context}, | |
| author = {{OceanLabs}}, | |
| year = {2026}, | |
| url = {https://huggingface.co/OceanLabs/Ocean-Horizon-3.7B}, | |
| } | |
| ``` | |
| --- | |
| **OceanLabs** – Advancing open, high-capability language models. | |