Spaces:
Running on Zero
Running on Zero
| # Respite API Reference | |
| **Base URL:** `https://cazyundee-respite-api.hf.space` | |
| **Version:** 1.0.0 | |
| **Authentication:** None required | |
| **Content Types:** `application/json`, `audio/wav` | |
| --- | |
| ## Overview | |
| Respite is a general-purpose AI API running on Hugging Face ZeroGPU infrastructure. It currently exposes two capabilities: **web search** (DuckDuckGo, no API key) and **audio generation** (Stable Audio 3 Small). The API is self-describing β call the discovery endpoints to enumerate capabilities and server specs at runtime. | |
| --- | |
| ## Endpoints | |
| ### `GET /respite/specs` | |
| Returns server metadata, hardware specs, and runtime limits. | |
| **Response:** | |
| ```json | |
| { | |
| "name": "Respite API", | |
| "version": "1.0.0", | |
| "description": "General-purpose AI API server with audio generation capabilities.", | |
| "base_path": "/respite", | |
| "authentication": "none", | |
| "content_types": ["application/json", "audio/wav"], | |
| "resources_endpoint": "/respite/resources", | |
| "specs_endpoint": "/respite/specs", | |
| "limits": { | |
| "max_concurrent_requests": 1, | |
| "max_queue_size": 4, | |
| "audio_max_duration_seconds": 120 | |
| }, | |
| "runtime": { | |
| "platform": "Linux-6.12.94-123.192.amzn2023.x86_64-x86_64-with-glibc2.36", | |
| "python_version": "3.10.13", | |
| "cpu_cores": 16, | |
| "host_cpu_cores": 192, | |
| "ram_bytes": 104000000000, | |
| "storage_total_bytes": 8589854879744, | |
| "storage_used_bytes": 6852819365888, | |
| "storage_free_bytes": 1737035513856, | |
| "hardware": "zero-a10g", | |
| "gpu": { | |
| "type": "NVIDIA RTX Pro 6000 Blackwell", | |
| "vram_bytes": 51539607552, | |
| "shared_zero_gpu": true | |
| } | |
| } | |
| } | |
| ``` | |
| **Fields:** | |
| | Field | Description | | |
| |---|---| | |
| | `limits.max_concurrent_requests` | Only 1 request can be processed at a time (ZeroGPU queue) | | |
| | `limits.max_queue_size` | Max 4 requests queued; beyond that, 503 | | |
| | `limits.audio_max_duration_seconds` | Max audio length per generation | | |
| | `runtime.gpu.shared_zero_gpu` | GPU time is shared β requests are queued and batched | | |
| | `runtime.ram_bytes` | Container memory limit (not host total) | | |
| | `runtime.cpu_cores` | Container CPU limit (not host total) | | |
| --- | |
| ### `GET /respite/resources` | |
| Lists all callable capabilities. | |
| **Response:** | |
| ```json | |
| { | |
| "resources": { | |
| "audio_generation": { | |
| "name": "Audio generation", | |
| "description": "Generate music or sound effects from a text prompt.", | |
| "endpoint": "/respite/audio/generate", | |
| "method": "POST", | |
| "input": { | |
| "prompt": "string", | |
| "duration": "number (1-120 seconds)", | |
| "steps": "integer (1-50)", | |
| "cfg_scale": "number (0-10)", | |
| "seed": "integer (-1 for random)", | |
| "model": "small-music | small-sfx" | |
| }, | |
| "output": "WAV audio file" | |
| }, | |
| "web_search": { | |
| "name": "Web search", | |
| "description": "Search the web via DuckDuckGo. Supports general, news, and image search. No API key required.", | |
| "endpoint": "/respite/search", | |
| "method": "GET", | |
| "input": { | |
| "q": "string (search query)", | |
| "backend": "string: 'text' (default), 'news', or 'images'", | |
| "max_results": "integer (1-50, default 10)", | |
| "extract_top": "integer (0-5, default 0) - fetch full page content for top N results", | |
| "region": "string (default 'wt-wt') - e.g. 'us-en', 'gb-en', 'de-de'" | |
| }, | |
| "output": "JSON with title, url, content per result. News results include date and source." | |
| } | |
| } | |
| } | |
| ``` | |
| --- | |
| ### `GET /respite/search` | |
| Search the web via DuckDuckGo. Supports general text, news, and image search. No API key required. | |
| **Query Parameters:** | |
| | Parameter | Type | Required | Default | Description | | |
| |---|---|---|---|---| | |
| | `q` | string | Yes | β | Search query | | |
| | `backend` | string | No | `"text"` | `"text"` (general), `"news"` (news articles), or `"images"` | | |
| | `max_results` | integer | No | 10 | Number of results (1-50) | | |
| | `extract_top` | integer | No | 0 | Fetch full page content for top N results (0-5). Adds latency but gives richer content. | | |
| | `region` | string | No | `"wt-wt"` | Region code: `"us-en"`, `"gb-en"`, `"de-de"`, `"wt-wt"` (worldwide) | | |
| **Request (text search):** | |
| ``` | |
| GET /respite/search?q=what+is+python&backend=text&max_results=3 | |
| ``` | |
| **Request (news search):** | |
| ``` | |
| GET /respite/search?q=latest+AI+news&backend=news&max_results=5 | |
| ``` | |
| **Request (news + extraction):** | |
| ``` | |
| GET /respite/search?q=nvidia+earnings&backend=news&max_results=3&extract_top=2 | |
| ``` | |
| **Response (text):** | |
| ```json | |
| { | |
| "query": "what is python", | |
| "backend": "text", | |
| "results": [ | |
| { | |
| "title": "Python (programming language) - Wikipedia", | |
| "url": "https://en.wikipedia.org/wiki/Python_(programming_language)", | |
| "content": "Python is a high-level, general-purpose programming language that emphasizes code readability..." | |
| } | |
| ], | |
| "total_results": 1 | |
| } | |
| ``` | |
| **Response (news):** | |
| ```json | |
| { | |
| "query": "latest AI news", | |
| "backend": "news", | |
| "results": [ | |
| { | |
| "title": "Strong AI chip demand fuels Nvidia's Q2 results", | |
| "url": "https://apnews.com/article/...", | |
| "content": "Nvidia's latest quarterly results once again blew past Wall Street's expectations...", | |
| "source": "Associated Press News", | |
| "date": "2026-08-27T00:00:00+00:00", | |
| "image": "https://..." | |
| } | |
| ], | |
| "total_results": 1 | |
| } | |
| ``` | |
| **Response (with extraction):** | |
| ```json | |
| { | |
| "query": "nvidia earnings", | |
| "backend": "news", | |
| "results": [ | |
| { | |
| "title": "Nvidia earnings report", | |
| "url": "https://...", | |
| "content": "Short snippet...", | |
| "source": "Reuters", | |
| "date": "2026-08-27T00:00:00+00:00", | |
| "extracted_content": "Full article text, up to 8000 characters..." | |
| } | |
| ], | |
| "total_results": 1 | |
| } | |
| ``` | |
| **Error responses:** | |
| | Status | Body | Cause | | |
| |---|---|---| | |
| | 400 | `{"error": "Missing query parameter 'q'"}` | No query provided | | |
| | 400 | `{"error": "Invalid backend 'x'. Use 'text', 'news', or 'images'."}` | Bad backend value | | |
| | 502 | `{"error": "..."}` | Search backend unreachable | | |
| --- | |
| ### `POST /respite/audio/generate` | |
| Generate music or sound effects from a text prompt using Stable Audio 3. | |
| **Request body (JSON):** | |
| | Field | Type | Required | Default | Description | | |
| |---|---|---|---|---| | |
| | `prompt` | string | Yes | β | Text description of desired audio | | |
| | `duration` | number | No | 30 | Duration in seconds (1-120) | | |
| | `steps` | integer | No | 8 | Inference steps (1-50). Higher = better quality, slower | | |
| | `cfg_scale` | number | No | 1.0 | Classifier-free guidance scale (0-10). Higher = more prompt adherence | | |
| | `seed` | integer | No | -1 | Random seed. -1 = random each time | | |
| | `model` | string | No | "small-music" | Either `"small-music"` or `"small-sfx"` | | |
| **Request:** | |
| ```bash | |
| curl -X POST "https://cazyundee-respite-api.hf.space/respite/audio/generate" \ | |
| -H "Content-Type: application/json" \ | |
| -d '{"prompt": "upbeat synthwave melody", "duration": 30, "model": "small-music"}' | |
| ``` | |
| **Response:** | |
| - **Content-Type:** `audio/wav` | |
| - **Body:** Raw WAV audio binary (44100 Hz sample rate) | |
| - The response is a downloadable audio file, not JSON | |
| **Error responses:** | |
| | Status | Body | Cause | | |
| |---|---|---| | |
| | 503 | `{"detail":"Queue is full. Please try again later."}` | All 4 queue slots occupied | | |
| | 500 | Server error | Generation failed | | |
| **Important notes:** | |
| - ZeroGPU shares one GPU across all requests. The server processes requests serially. | |
| - First request triggers model loading (~30-60s). Subsequent requests use the cached model. | |
| - The `model` field selects between `"small-music"` (general music) and `"small-sfx"` (sound effects). | |
| - The returned audio is always WAV format, 44100 Hz sample rate, 16-bit PCM. | |
| --- | |
| ## Rate Limits | |
| - **Max concurrent requests:** 1 (GPU locked) | |
| - **Max queue size:** 4 (503 if full) | |
| - **Audio generation:** ~5-15s per request after model is loaded (varies by `steps` and `duration`) | |
| - **Search:** No limit (DuckDuckGo backend, no API key) | |
| --- | |
| ## Error Handling | |
| All errors return JSON with an `"error"` or `"detail"` field: | |
| ```json | |
| {"error": "Missing query parameter 'q'"} | |
| {"detail": "Queue is full. Please try again later."} | |
| ``` | |
| Check the HTTP status code: | |
| | Code | Meaning | | |
| |---|---| | |
| | 200 | Success | | |
| | 400 | Bad request (missing/invalid parameters) | | |
| | 502 | Upstream service unreachable | | |
| | 503 | Queue full, try again | | |
| --- | |
| ## Integration Notes for Chat Agents | |
| 1. **Discover capabilities at runtime** β Call `GET /respite/resources` to enumerate what the API can do. Do not hardcode resource names. | |
| 2. **Search strategies by query type:** | |
| | Query type | Backend | extract_top | Region | | |
| |---|---|---|---| | |
| | Current events, news | `news` | 2-3 | user's region | | |
| | Factual questions | `text` | 1 | `wt-wt` | | |
| | How-to, tutorials | `text` | 0 | `wt-wt` | | |
| | Recent developments | `news` | 0 | `wt-wt` | | |
| | Images, visuals | `images` | 0 | user's region | | |
| 3. **Use extraction for depth** β When a user wants detailed information (e.g., "summarize this article"), set `extract_top=1` to get the full page content. Don't overuse it β it adds latency. | |
| 4. **Audio generation is queued** β Only 1 generation runs at a time. If the queue is full (503), retry with exponential backoff. Do not retry rapidly. | |
| 5. **Model selection** β Use `"small-music"` for music, melodies, ambient sound. Use `"small-sfx"` for sound effects, foley, short percussive sounds. | |
| 6. **Response handling** β `/respite/audio/generate` returns binary WAV, not JSON. Save to a `.wav` file and play/serve it. Do not try to JSON-parse the response body. | |
| 7. **Specs for capability planning** β `GET /respite/specs` tells you GPU type, VRAM, and queue limits. Use this to decide what to offer users (e.g., don't promise 4-minute tracks β max is 120s). | |
| 8. **Graceful degradation** β If the API is down (503, 502), tell the user the service is temporarily unavailable rather than hanging. ZeroGPU Spaces can go cold when idle. | |