Spaces:
Running
Running
Long context: up to 1M tokens on supported models
Browse files
README.md
CHANGED
|
@@ -18,7 +18,7 @@ platform — with no compromise on features.
|
|
| 18 |
- **Price-first**: open-weight models at floor prices — see the table below.
|
| 19 |
- **Full feature parity**: tool calling (function calling) and structured
|
| 20 |
output (`response_format: json_schema`) on every conversational model.
|
| 21 |
-
- **Long context**: up to
|
| 22 |
- **Low latency**: time-to-first-token well under the 5 s provider budget
|
| 23 |
(measured ~0.9 s non-streaming).
|
| 24 |
- **Autoscaling fleet**: capacity scales out automatically with demand;
|
|
|
|
| 18 |
- **Price-first**: open-weight models at floor prices — see the table below.
|
| 19 |
- **Full feature parity**: tool calling (function calling) and structured
|
| 20 |
output (`response_format: json_schema`) on every conversational model.
|
| 21 |
+
- **Long context**: up to 1M tokens of context on supported models.
|
| 22 |
- **Low latency**: time-to-first-token well under the 5 s provider budget
|
| 23 |
(measured ~0.9 s non-streaming).
|
| 24 |
- **Autoscaling fleet**: capacity scales out automatically with demand;
|