OceanLabs
/

Ocean-3.7B-coding / README.md
roskosmos19's picture
Upload 8 files
2b2ff9a verified
|
Raw History Blame Contribute Delete
5.21 kB
---
pipeline_tag: text-generation
library_name: transformers
model_name: Ocean-Horizon-3.7B
language:
- en
- de
- multilingual
license: apache-2.0
tags:
- ocean-horizon
- 3.7b
- dense
- open-weights
- oceanlabs
- long-context
- reasoning
- agentic
---
# Ocean-Horizon-3.7B
**OceanLabs / Ocean-Horizon-3.7B** is a high-performance 3.7B-parameter dense decoder-only language model optimized for advanced reasoning, coding, agentic workflows, and ultra-long context understanding.
Built upon a rigorously engineered dense architecture with a native **524,288-token (512K)** context window, Ocean-Horizon-3.7B delivers frontier-level capabilities in a compact, efficient package.
## Key Capabilities
- **Exceptional small-model performance.** Strong results across agentic, coding, mathematical, and scientific reasoning benchmarks.
- **Native 512K context.** Full 524,288-token context window available from mid-training stages onward, enabling deep document analysis, long-horizon planning, and complex multi-turn interactions.
- **Optimized reasoning depth.** Supports configurable reasoning effort (high / medium / low) with dedicated thinking tokens for transparent chain-of-thought.
- **Production-ready tool use.** Native support for structured tool calling with multiple presentation formats (JSON, XML, Markdown).
- **Fully open.** Architecture, configuration, and evaluation resources are openly available for research and commercial use under Apache 2.0.
## Architecture Summary
| Parameter | Value |
|-----------------------------|----------------|
| Parameters | 3.7B |
| Architecture | Dense |
| Hidden size | 2560 |
| Intermediate size | 10240 |
| Layers | 36 |
| Attention heads | 32 |
| Key-value heads | 8 (GQA) |
| Head dimension | 128 |
| Vocabulary size | 250,624 |
| Context length | 524,288 |
| Activation | SiLU |
| Precision | bfloat16 |
| RoPE θ | 10,000,000 |
## Recommended Inference Settings
- **Reasoning effort:** always prefer `"high"` for maximum quality.
- **Sampling:** `temperature=1.0`, `top_p=0.95`.
- **Maximum output tokens:** at least 32,768 to avoid truncating reasoning traces.
- **Serving backends:** Validated with vLLM and SGLang (BF16, FlashAttention-3 recommended).
### Example with Transformers
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "OceanLabs/Ocean-Horizon-3.7B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
dtype="bfloat16",
low_cpu_mem_usage=True,
trust_remote_code=True
)
inputs = tokenizer("Explain the advantages of a 512K context window for agentic systems.", return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
### Example with OpenAI-compatible API (vLLM / SGLang)
```python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="OceanLabs/Ocean-Horizon-3.7B",
messages=[{"role": "user", "content": "Solve this step by step."}],
temperature=1.0,
top_p=0.95,
max_tokens=32768,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
message = response.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print("Answer:", message.content)
```
## Training Pipeline Overview
Ocean-Horizon-3.7B follows a multi-stage curriculum:
1. **Pretraining** – large-scale general knowledge acquisition (22.9T tokens).
2. **Midtraining** – progressive context extension (32K → 128K → 512K) plus agentic and reasoning data.
3. **Reinforcement Learning** – specialized experts for mathematics, code, and STEM-code, followed by model merging.
4. **Supervised Fine-Tuning** – high-quality multi-domain instruction tuning with learning-rate decay.
This staged approach yields strong generalization while preserving long-context fidelity.
## Best Practices
1. Always request high reasoning effort for evaluation and complex tasks.
2. Allocate sufficient output length (≥ 32k tokens) so that reasoning is never truncated.
3. Use the native `k2_horizon` reasoning and tool-call parsers when serving with vLLM or SGLang.
4. Pin a specific revision or commit hash for reproducible deployments.
## License
Apache License 2.0
## Citation
```bibtex
@misc{oceanhorizon2026,
title = {Ocean-Horizon-3.7B: Compact Dense Model with Frontier Reasoning and 512K Context},
author = {{OceanLabs}},
year = {2026},
url = {https://huggingface.co/OceanLabs/Ocean-Horizon-3.7B},
}
```
---
**OceanLabs** – Advancing open, high-capability language models.