File size: 2,243 Bytes
dbf7d82
 
 
 
 
 
 
 
 
f8dc6ed
 
ba91dca
 
 
f8dc6ed
ba91dca
 
 
 
 
4b2d495
ba91dca
f8dc6ed
 
 
 
ba91dca
 
 
 
 
 
1352dc3
ba91dca
 
 
 
 
 
 
 
 
 
cc3bd29
ba91dca
 
 
 
 
 
 
 
 
 
 
 
 
 
f8dc6ed
ba91dca
 
 
 
 
 
 
 
 
f8dc6ed
ba91dca
f8dc6ed
ba91dca
f8dc6ed
ba91dca
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
---
title: README
emoji: πŸ“‰
colorFrom: blue
colorTo: green
sdk: static
pinned: false
---

# IntelCS-AI

**Low-cost inference for open-weight models.** We serve open models on
efficient GPU capacity at some of the lowest per-token prices on the
platform β€” with no compromise on features.

## Why IntelCS-AI

- **Price-first**: open-weight models at floor prices β€” see the table below.
- **Full feature parity**: tool calling (function calling) and structured
  output (`response_format: json_schema`) on every conversational model.
- **Long context**: up to 1M tokens of context on supported models.
- **Low latency**: time-to-first-token well under the 5 s provider budget
  (measured ~0.9 s non-streaming).
- **Autoscaling fleet**: capacity scales out automatically with demand;
  routing, metering, and billing run on edge infrastructure.

## Pricing

| Model | Input (per 1M tokens) | Output (per 1M tokens) | Context |
|---|---|---|---|
| `google/gemma-3-4b-it` | $0.05 | $0.10 | 131K |

Our lineup rotates as we add capacity β€” check back for new models.

## Usage

OpenAI-compatible, through the standard Hugging Face clients:

```python
from huggingface_hub import InferenceClient

client = InferenceClient(
    model="google/gemma-3-4b-it",
    provider="intelcs-ai-iaas",
)

# Chat with tool calling
response = client.chat.completions.create(
    messages=[{"role": "user", "content": "What's the weather in Berlin?"}],
    tools=[{
        "type": "function",
        "function": {
            "name": "get_weather",
            "parameters": {"type": "object", "properties": {"location": {"type": "string"}}},
        },
    }],
)
print(response.choices[0].message.tool_calls)

# Structured output
response = client.chat.completions.create(
    messages=[{"role": "user", "content": "Extract the city: 'Flight to Berlin delayed'"}],
    response_format={"type": "json_schema", "json_schema": {"name": "city", "schema": {
        "type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"],
    }}},
)
print(response.choices[0].message.content)  # {"city": "Berlin"}
```

Streaming (`stream=True`) is fully supported and metered per token.

## Resources

- **Website**: [intelcs.ai](https://intelcs.ai)