Spaces:
Running
Running
File size: 2,243 Bytes
dbf7d82 f8dc6ed ba91dca f8dc6ed ba91dca 4b2d495 ba91dca f8dc6ed ba91dca 1352dc3 ba91dca cc3bd29 ba91dca f8dc6ed ba91dca f8dc6ed ba91dca f8dc6ed ba91dca f8dc6ed ba91dca | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 | ---
title: README
emoji: π
colorFrom: blue
colorTo: green
sdk: static
pinned: false
---
# IntelCS-AI
**Low-cost inference for open-weight models.** We serve open models on
efficient GPU capacity at some of the lowest per-token prices on the
platform β with no compromise on features.
## Why IntelCS-AI
- **Price-first**: open-weight models at floor prices β see the table below.
- **Full feature parity**: tool calling (function calling) and structured
output (`response_format: json_schema`) on every conversational model.
- **Long context**: up to 1M tokens of context on supported models.
- **Low latency**: time-to-first-token well under the 5 s provider budget
(measured ~0.9 s non-streaming).
- **Autoscaling fleet**: capacity scales out automatically with demand;
routing, metering, and billing run on edge infrastructure.
## Pricing
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Context |
|---|---|---|---|
| `google/gemma-3-4b-it` | $0.05 | $0.10 | 131K |
Our lineup rotates as we add capacity β check back for new models.
## Usage
OpenAI-compatible, through the standard Hugging Face clients:
```python
from huggingface_hub import InferenceClient
client = InferenceClient(
model="google/gemma-3-4b-it",
provider="intelcs-ai-iaas",
)
# Chat with tool calling
response = client.chat.completions.create(
messages=[{"role": "user", "content": "What's the weather in Berlin?"}],
tools=[{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {"type": "object", "properties": {"location": {"type": "string"}}},
},
}],
)
print(response.choices[0].message.tool_calls)
# Structured output
response = client.chat.completions.create(
messages=[{"role": "user", "content": "Extract the city: 'Flight to Berlin delayed'"}],
response_format={"type": "json_schema", "json_schema": {"name": "city", "schema": {
"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"],
}}},
)
print(response.choices[0].message.content) # {"city": "Berlin"}
```
Streaming (`stream=True`) is fully supported and metered per token.
## Resources
- **Website**: [intelcs.ai](https://intelcs.ai)
|