# Respite API Reference **Base URL:** `https://cazyundee-respite-api.hf.space` **Version:** 1.0.0 **Authentication:** None required **Content Types:** `application/json`, `audio/wav` --- ## Overview Respite is a general-purpose AI API running on Hugging Face ZeroGPU infrastructure. It currently exposes two capabilities: **web search** (DuckDuckGo, no API key) and **audio generation** (Stable Audio 3 Small). The API is self-describing — call the discovery endpoints to enumerate capabilities and server specs at runtime. --- ## Endpoints ### `GET /respite/specs` Returns server metadata, hardware specs, and runtime limits. **Response:** ```json { "name": "Respite API", "version": "1.0.0", "description": "General-purpose AI API server with audio generation capabilities.", "base_path": "/respite", "authentication": "none", "content_types": ["application/json", "audio/wav"], "resources_endpoint": "/respite/resources", "specs_endpoint": "/respite/specs", "limits": { "max_concurrent_requests": 1, "max_queue_size": 4, "audio_max_duration_seconds": 120 }, "runtime": { "platform": "Linux-6.12.94-123.192.amzn2023.x86_64-x86_64-with-glibc2.36", "python_version": "3.10.13", "cpu_cores": 16, "host_cpu_cores": 192, "ram_bytes": 104000000000, "storage_total_bytes": 8589854879744, "storage_used_bytes": 6852819365888, "storage_free_bytes": 1737035513856, "hardware": "zero-a10g", "gpu": { "type": "NVIDIA RTX Pro 6000 Blackwell", "vram_bytes": 51539607552, "shared_zero_gpu": true } } } ``` **Fields:** | Field | Description | |---|---| | `limits.max_concurrent_requests` | Only 1 request can be processed at a time (ZeroGPU queue) | | `limits.max_queue_size` | Max 4 requests queued; beyond that, 503 | | `limits.audio_max_duration_seconds` | Max audio length per generation | | `runtime.gpu.shared_zero_gpu` | GPU time is shared — requests are queued and batched | | `runtime.ram_bytes` | Container memory limit (not host total) | | `runtime.cpu_cores` | Container CPU limit (not host total) | --- ### `GET /respite/resources` Lists all callable capabilities. **Response:** ```json { "resources": { "audio_generation": { "name": "Audio generation", "description": "Generate music or sound effects from a text prompt.", "endpoint": "/respite/audio/generate", "method": "POST", "input": { "prompt": "string", "duration": "number (1-120 seconds)", "steps": "integer (1-50)", "cfg_scale": "number (0-10)", "seed": "integer (-1 for random)", "model": "small-music | small-sfx" }, "output": "WAV audio file" }, "web_search": { "name": "Web search", "description": "Search the web via DuckDuckGo. Supports general, news, and image search. No API key required.", "endpoint": "/respite/search", "method": "GET", "input": { "q": "string (search query)", "backend": "string: 'text' (default), 'news', or 'images'", "max_results": "integer (1-50, default 10)", "extract_top": "integer (0-5, default 0) - fetch full page content for top N results", "region": "string (default 'wt-wt') - e.g. 'us-en', 'gb-en', 'de-de'" }, "output": "JSON with title, url, content per result. News results include date and source." } } } ``` --- ### `GET /respite/search` Search the web via DuckDuckGo. Supports general text, news, and image search. No API key required. **Query Parameters:** | Parameter | Type | Required | Default | Description | |---|---|---|---|---| | `q` | string | Yes | — | Search query | | `backend` | string | No | `"text"` | `"text"` (general), `"news"` (news articles), or `"images"` | | `max_results` | integer | No | 10 | Number of results (1-50) | | `extract_top` | integer | No | 0 | Fetch full page content for top N results (0-5). Adds latency but gives richer content. | | `region` | string | No | `"wt-wt"` | Region code: `"us-en"`, `"gb-en"`, `"de-de"`, `"wt-wt"` (worldwide) | **Request (text search):** ``` GET /respite/search?q=what+is+python&backend=text&max_results=3 ``` **Request (news search):** ``` GET /respite/search?q=latest+AI+news&backend=news&max_results=5 ``` **Request (news + extraction):** ``` GET /respite/search?q=nvidia+earnings&backend=news&max_results=3&extract_top=2 ``` **Response (text):** ```json { "query": "what is python", "backend": "text", "results": [ { "title": "Python (programming language) - Wikipedia", "url": "https://en.wikipedia.org/wiki/Python_(programming_language)", "content": "Python is a high-level, general-purpose programming language that emphasizes code readability..." } ], "total_results": 1 } ``` **Response (news):** ```json { "query": "latest AI news", "backend": "news", "results": [ { "title": "Strong AI chip demand fuels Nvidia's Q2 results", "url": "https://apnews.com/article/...", "content": "Nvidia's latest quarterly results once again blew past Wall Street's expectations...", "source": "Associated Press News", "date": "2026-08-27T00:00:00+00:00", "image": "https://..." } ], "total_results": 1 } ``` **Response (with extraction):** ```json { "query": "nvidia earnings", "backend": "news", "results": [ { "title": "Nvidia earnings report", "url": "https://...", "content": "Short snippet...", "source": "Reuters", "date": "2026-08-27T00:00:00+00:00", "extracted_content": "Full article text, up to 8000 characters..." } ], "total_results": 1 } ``` **Error responses:** | Status | Body | Cause | |---|---|---| | 400 | `{"error": "Missing query parameter 'q'"}` | No query provided | | 400 | `{"error": "Invalid backend 'x'. Use 'text', 'news', or 'images'."}` | Bad backend value | | 502 | `{"error": "..."}` | Search backend unreachable | --- ### `POST /respite/audio/generate` Generate music or sound effects from a text prompt using Stable Audio 3. **Request body (JSON):** | Field | Type | Required | Default | Description | |---|---|---|---|---| | `prompt` | string | Yes | — | Text description of desired audio | | `duration` | number | No | 30 | Duration in seconds (1-120) | | `steps` | integer | No | 8 | Inference steps (1-50). Higher = better quality, slower | | `cfg_scale` | number | No | 1.0 | Classifier-free guidance scale (0-10). Higher = more prompt adherence | | `seed` | integer | No | -1 | Random seed. -1 = random each time | | `model` | string | No | "small-music" | Either `"small-music"` or `"small-sfx"` | **Request:** ```bash curl -X POST "https://cazyundee-respite-api.hf.space/respite/audio/generate" \ -H "Content-Type: application/json" \ -d '{"prompt": "upbeat synthwave melody", "duration": 30, "model": "small-music"}' ``` **Response:** - **Content-Type:** `audio/wav` - **Body:** Raw WAV audio binary (44100 Hz sample rate) - The response is a downloadable audio file, not JSON **Error responses:** | Status | Body | Cause | |---|---|---| | 503 | `{"detail":"Queue is full. Please try again later."}` | All 4 queue slots occupied | | 500 | Server error | Generation failed | **Important notes:** - ZeroGPU shares one GPU across all requests. The server processes requests serially. - First request triggers model loading (~30-60s). Subsequent requests use the cached model. - The `model` field selects between `"small-music"` (general music) and `"small-sfx"` (sound effects). - The returned audio is always WAV format, 44100 Hz sample rate, 16-bit PCM. --- ## Rate Limits - **Max concurrent requests:** 1 (GPU locked) - **Max queue size:** 4 (503 if full) - **Audio generation:** ~5-15s per request after model is loaded (varies by `steps` and `duration`) - **Search:** No limit (DuckDuckGo backend, no API key) --- ## Error Handling All errors return JSON with an `"error"` or `"detail"` field: ```json {"error": "Missing query parameter 'q'"} {"detail": "Queue is full. Please try again later."} ``` Check the HTTP status code: | Code | Meaning | |---|---| | 200 | Success | | 400 | Bad request (missing/invalid parameters) | | 502 | Upstream service unreachable | | 503 | Queue full, try again | --- ## Integration Notes for Chat Agents 1. **Discover capabilities at runtime** — Call `GET /respite/resources` to enumerate what the API can do. Do not hardcode resource names. 2. **Search strategies by query type:** | Query type | Backend | extract_top | Region | |---|---|---|---| | Current events, news | `news` | 2-3 | user's region | | Factual questions | `text` | 1 | `wt-wt` | | How-to, tutorials | `text` | 0 | `wt-wt` | | Recent developments | `news` | 0 | `wt-wt` | | Images, visuals | `images` | 0 | user's region | 3. **Use extraction for depth** — When a user wants detailed information (e.g., "summarize this article"), set `extract_top=1` to get the full page content. Don't overuse it — it adds latency. 4. **Audio generation is queued** — Only 1 generation runs at a time. If the queue is full (503), retry with exponential backoff. Do not retry rapidly. 5. **Model selection** — Use `"small-music"` for music, melodies, ambient sound. Use `"small-sfx"` for sound effects, foley, short percussive sounds. 6. **Response handling** — `/respite/audio/generate` returns binary WAV, not JSON. Save to a `.wav` file and play/serve it. Do not try to JSON-parse the response body. 7. **Specs for capability planning** — `GET /respite/specs` tells you GPU type, VRAM, and queue limits. Use this to decide what to offer users (e.g., don't promise 4-minute tracks — max is 120s). 8. **Graceful degradation** — If the API is down (503, 502), tell the user the service is temporarily unavailable rather than hanging. ZeroGPU Spaces can go cold when idle.