Respite-API / API_REFERENCE.md
cazyundee's picture
Update API reference with news/images backends and extraction support
c638c2d verified
|
Raw
History Blame Contribute Delete
9.93 kB
# Respite API Reference
**Base URL:** `https://cazyundee-respite-api.hf.space`
**Version:** 1.0.0
**Authentication:** None required
**Content Types:** `application/json`, `audio/wav`
---
## Overview
Respite is a general-purpose AI API running on Hugging Face ZeroGPU infrastructure. It currently exposes two capabilities: **web search** (DuckDuckGo, no API key) and **audio generation** (Stable Audio 3 Small). The API is self-describing β€” call the discovery endpoints to enumerate capabilities and server specs at runtime.
---
## Endpoints
### `GET /respite/specs`
Returns server metadata, hardware specs, and runtime limits.
**Response:**
```json
{
"name": "Respite API",
"version": "1.0.0",
"description": "General-purpose AI API server with audio generation capabilities.",
"base_path": "/respite",
"authentication": "none",
"content_types": ["application/json", "audio/wav"],
"resources_endpoint": "/respite/resources",
"specs_endpoint": "/respite/specs",
"limits": {
"max_concurrent_requests": 1,
"max_queue_size": 4,
"audio_max_duration_seconds": 120
},
"runtime": {
"platform": "Linux-6.12.94-123.192.amzn2023.x86_64-x86_64-with-glibc2.36",
"python_version": "3.10.13",
"cpu_cores": 16,
"host_cpu_cores": 192,
"ram_bytes": 104000000000,
"storage_total_bytes": 8589854879744,
"storage_used_bytes": 6852819365888,
"storage_free_bytes": 1737035513856,
"hardware": "zero-a10g",
"gpu": {
"type": "NVIDIA RTX Pro 6000 Blackwell",
"vram_bytes": 51539607552,
"shared_zero_gpu": true
}
}
}
```
**Fields:**
| Field | Description |
|---|---|
| `limits.max_concurrent_requests` | Only 1 request can be processed at a time (ZeroGPU queue) |
| `limits.max_queue_size` | Max 4 requests queued; beyond that, 503 |
| `limits.audio_max_duration_seconds` | Max audio length per generation |
| `runtime.gpu.shared_zero_gpu` | GPU time is shared β€” requests are queued and batched |
| `runtime.ram_bytes` | Container memory limit (not host total) |
| `runtime.cpu_cores` | Container CPU limit (not host total) |
---
### `GET /respite/resources`
Lists all callable capabilities.
**Response:**
```json
{
"resources": {
"audio_generation": {
"name": "Audio generation",
"description": "Generate music or sound effects from a text prompt.",
"endpoint": "/respite/audio/generate",
"method": "POST",
"input": {
"prompt": "string",
"duration": "number (1-120 seconds)",
"steps": "integer (1-50)",
"cfg_scale": "number (0-10)",
"seed": "integer (-1 for random)",
"model": "small-music | small-sfx"
},
"output": "WAV audio file"
},
"web_search": {
"name": "Web search",
"description": "Search the web via DuckDuckGo. Supports general, news, and image search. No API key required.",
"endpoint": "/respite/search",
"method": "GET",
"input": {
"q": "string (search query)",
"backend": "string: 'text' (default), 'news', or 'images'",
"max_results": "integer (1-50, default 10)",
"extract_top": "integer (0-5, default 0) - fetch full page content for top N results",
"region": "string (default 'wt-wt') - e.g. 'us-en', 'gb-en', 'de-de'"
},
"output": "JSON with title, url, content per result. News results include date and source."
}
}
}
```
---
### `GET /respite/search`
Search the web via DuckDuckGo. Supports general text, news, and image search. No API key required.
**Query Parameters:**
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| `q` | string | Yes | β€” | Search query |
| `backend` | string | No | `"text"` | `"text"` (general), `"news"` (news articles), or `"images"` |
| `max_results` | integer | No | 10 | Number of results (1-50) |
| `extract_top` | integer | No | 0 | Fetch full page content for top N results (0-5). Adds latency but gives richer content. |
| `region` | string | No | `"wt-wt"` | Region code: `"us-en"`, `"gb-en"`, `"de-de"`, `"wt-wt"` (worldwide) |
**Request (text search):**
```
GET /respite/search?q=what+is+python&backend=text&max_results=3
```
**Request (news search):**
```
GET /respite/search?q=latest+AI+news&backend=news&max_results=5
```
**Request (news + extraction):**
```
GET /respite/search?q=nvidia+earnings&backend=news&max_results=3&extract_top=2
```
**Response (text):**
```json
{
"query": "what is python",
"backend": "text",
"results": [
{
"title": "Python (programming language) - Wikipedia",
"url": "https://en.wikipedia.org/wiki/Python_(programming_language)",
"content": "Python is a high-level, general-purpose programming language that emphasizes code readability..."
}
],
"total_results": 1
}
```
**Response (news):**
```json
{
"query": "latest AI news",
"backend": "news",
"results": [
{
"title": "Strong AI chip demand fuels Nvidia's Q2 results",
"url": "https://apnews.com/article/...",
"content": "Nvidia's latest quarterly results once again blew past Wall Street's expectations...",
"source": "Associated Press News",
"date": "2026-08-27T00:00:00+00:00",
"image": "https://..."
}
],
"total_results": 1
}
```
**Response (with extraction):**
```json
{
"query": "nvidia earnings",
"backend": "news",
"results": [
{
"title": "Nvidia earnings report",
"url": "https://...",
"content": "Short snippet...",
"source": "Reuters",
"date": "2026-08-27T00:00:00+00:00",
"extracted_content": "Full article text, up to 8000 characters..."
}
],
"total_results": 1
}
```
**Error responses:**
| Status | Body | Cause |
|---|---|---|
| 400 | `{"error": "Missing query parameter 'q'"}` | No query provided |
| 400 | `{"error": "Invalid backend 'x'. Use 'text', 'news', or 'images'."}` | Bad backend value |
| 502 | `{"error": "..."}` | Search backend unreachable |
---
### `POST /respite/audio/generate`
Generate music or sound effects from a text prompt using Stable Audio 3.
**Request body (JSON):**
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| `prompt` | string | Yes | β€” | Text description of desired audio |
| `duration` | number | No | 30 | Duration in seconds (1-120) |
| `steps` | integer | No | 8 | Inference steps (1-50). Higher = better quality, slower |
| `cfg_scale` | number | No | 1.0 | Classifier-free guidance scale (0-10). Higher = more prompt adherence |
| `seed` | integer | No | -1 | Random seed. -1 = random each time |
| `model` | string | No | "small-music" | Either `"small-music"` or `"small-sfx"` |
**Request:**
```bash
curl -X POST "https://cazyundee-respite-api.hf.space/respite/audio/generate" \
-H "Content-Type: application/json" \
-d '{"prompt": "upbeat synthwave melody", "duration": 30, "model": "small-music"}'
```
**Response:**
- **Content-Type:** `audio/wav`
- **Body:** Raw WAV audio binary (44100 Hz sample rate)
- The response is a downloadable audio file, not JSON
**Error responses:**
| Status | Body | Cause |
|---|---|---|
| 503 | `{"detail":"Queue is full. Please try again later."}` | All 4 queue slots occupied |
| 500 | Server error | Generation failed |
**Important notes:**
- ZeroGPU shares one GPU across all requests. The server processes requests serially.
- First request triggers model loading (~30-60s). Subsequent requests use the cached model.
- The `model` field selects between `"small-music"` (general music) and `"small-sfx"` (sound effects).
- The returned audio is always WAV format, 44100 Hz sample rate, 16-bit PCM.
---
## Rate Limits
- **Max concurrent requests:** 1 (GPU locked)
- **Max queue size:** 4 (503 if full)
- **Audio generation:** ~5-15s per request after model is loaded (varies by `steps` and `duration`)
- **Search:** No limit (DuckDuckGo backend, no API key)
---
## Error Handling
All errors return JSON with an `"error"` or `"detail"` field:
```json
{"error": "Missing query parameter 'q'"}
{"detail": "Queue is full. Please try again later."}
```
Check the HTTP status code:
| Code | Meaning |
|---|---|
| 200 | Success |
| 400 | Bad request (missing/invalid parameters) |
| 502 | Upstream service unreachable |
| 503 | Queue full, try again |
---
## Integration Notes for Chat Agents
1. **Discover capabilities at runtime** β€” Call `GET /respite/resources` to enumerate what the API can do. Do not hardcode resource names.
2. **Search strategies by query type:**
| Query type | Backend | extract_top | Region |
|---|---|---|---|
| Current events, news | `news` | 2-3 | user's region |
| Factual questions | `text` | 1 | `wt-wt` |
| How-to, tutorials | `text` | 0 | `wt-wt` |
| Recent developments | `news` | 0 | `wt-wt` |
| Images, visuals | `images` | 0 | user's region |
3. **Use extraction for depth** β€” When a user wants detailed information (e.g., "summarize this article"), set `extract_top=1` to get the full page content. Don't overuse it β€” it adds latency.
4. **Audio generation is queued** β€” Only 1 generation runs at a time. If the queue is full (503), retry with exponential backoff. Do not retry rapidly.
5. **Model selection** β€” Use `"small-music"` for music, melodies, ambient sound. Use `"small-sfx"` for sound effects, foley, short percussive sounds.
6. **Response handling** β€” `/respite/audio/generate` returns binary WAV, not JSON. Save to a `.wav` file and play/serve it. Do not try to JSON-parse the response body.
7. **Specs for capability planning** β€” `GET /respite/specs` tells you GPU type, VRAM, and queue limits. Use this to decide what to offer users (e.g., don't promise 4-minute tracks β€” max is 120s).
8. **Graceful degradation** β€” If the API is down (503, 502), tell the user the service is temporarily unavailable rather than hanging. ZeroGPU Spaces can go cold when idle.