Spaces:
Running on Zero
A newer version of the Gradio SDK is available: 6.28.0
Respite API Reference
Base URL: https://cazyundee-respite-api.hf.space
Version: 1.0.0
Authentication: None required
Content Types: application/json, audio/wav
Overview
Respite is a general-purpose AI API running on Hugging Face ZeroGPU infrastructure. It currently exposes two capabilities: web search (DuckDuckGo, no API key) and audio generation (Stable Audio 3 Small). The API is self-describing β call the discovery endpoints to enumerate capabilities and server specs at runtime.
Endpoints
GET /respite/specs
Returns server metadata, hardware specs, and runtime limits.
Response:
{
"name": "Respite API",
"version": "1.0.0",
"description": "General-purpose AI API server with audio generation capabilities.",
"base_path": "/respite",
"authentication": "none",
"content_types": ["application/json", "audio/wav"],
"resources_endpoint": "/respite/resources",
"specs_endpoint": "/respite/specs",
"limits": {
"max_concurrent_requests": 1,
"max_queue_size": 4,
"audio_max_duration_seconds": 120
},
"runtime": {
"platform": "Linux-6.12.94-123.192.amzn2023.x86_64-x86_64-with-glibc2.36",
"python_version": "3.10.13",
"cpu_cores": 16,
"host_cpu_cores": 192,
"ram_bytes": 104000000000,
"storage_total_bytes": 8589854879744,
"storage_used_bytes": 6852819365888,
"storage_free_bytes": 1737035513856,
"hardware": "zero-a10g",
"gpu": {
"type": "NVIDIA RTX Pro 6000 Blackwell",
"vram_bytes": 51539607552,
"shared_zero_gpu": true
}
}
}
Fields:
| Field | Description |
|---|---|
limits.max_concurrent_requests |
Only 1 request can be processed at a time (ZeroGPU queue) |
limits.max_queue_size |
Max 4 requests queued; beyond that, 503 |
limits.audio_max_duration_seconds |
Max audio length per generation |
runtime.gpu.shared_zero_gpu |
GPU time is shared β requests are queued and batched |
runtime.ram_bytes |
Container memory limit (not host total) |
runtime.cpu_cores |
Container CPU limit (not host total) |
GET /respite/resources
Lists all callable capabilities.
Response:
{
"resources": {
"audio_generation": {
"name": "Audio generation",
"description": "Generate music or sound effects from a text prompt.",
"endpoint": "/respite/audio/generate",
"method": "POST",
"input": {
"prompt": "string",
"duration": "number (1-120 seconds)",
"steps": "integer (1-50)",
"cfg_scale": "number (0-10)",
"seed": "integer (-1 for random)",
"model": "small-music | small-sfx"
},
"output": "WAV audio file"
},
"web_search": {
"name": "Web search",
"description": "Search the web via DuckDuckGo. Supports general, news, and image search. No API key required.",
"endpoint": "/respite/search",
"method": "GET",
"input": {
"q": "string (search query)",
"backend": "string: 'text' (default), 'news', or 'images'",
"max_results": "integer (1-50, default 10)",
"extract_top": "integer (0-5, default 0) - fetch full page content for top N results",
"region": "string (default 'wt-wt') - e.g. 'us-en', 'gb-en', 'de-de'"
},
"output": "JSON with title, url, content per result. News results include date and source."
}
}
}
GET /respite/search
Search the web via DuckDuckGo. Supports general text, news, and image search. No API key required.
Query Parameters:
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
q |
string | Yes | β | Search query |
backend |
string | No | "text" |
"text" (general), "news" (news articles), or "images" |
max_results |
integer | No | 10 | Number of results (1-50) |
extract_top |
integer | No | 0 | Fetch full page content for top N results (0-5). Adds latency but gives richer content. |
region |
string | No | "wt-wt" |
Region code: "us-en", "gb-en", "de-de", "wt-wt" (worldwide) |
Request (text search):
GET /respite/search?q=what+is+python&backend=text&max_results=3
Request (news search):
GET /respite/search?q=latest+AI+news&backend=news&max_results=5
Request (news + extraction):
GET /respite/search?q=nvidia+earnings&backend=news&max_results=3&extract_top=2
Response (text):
{
"query": "what is python",
"backend": "text",
"results": [
{
"title": "Python (programming language) - Wikipedia",
"url": "https://en.wikipedia.org/wiki/Python_(programming_language)",
"content": "Python is a high-level, general-purpose programming language that emphasizes code readability..."
}
],
"total_results": 1
}
Response (news):
{
"query": "latest AI news",
"backend": "news",
"results": [
{
"title": "Strong AI chip demand fuels Nvidia's Q2 results",
"url": "https://apnews.com/article/...",
"content": "Nvidia's latest quarterly results once again blew past Wall Street's expectations...",
"source": "Associated Press News",
"date": "2026-08-27T00:00:00+00:00",
"image": "https://..."
}
],
"total_results": 1
}
Response (with extraction):
{
"query": "nvidia earnings",
"backend": "news",
"results": [
{
"title": "Nvidia earnings report",
"url": "https://...",
"content": "Short snippet...",
"source": "Reuters",
"date": "2026-08-27T00:00:00+00:00",
"extracted_content": "Full article text, up to 8000 characters..."
}
],
"total_results": 1
}
Error responses:
| Status | Body | Cause |
|---|---|---|
| 400 | {"error": "Missing query parameter 'q'"} |
No query provided |
| 400 | {"error": "Invalid backend 'x'. Use 'text', 'news', or 'images'."} |
Bad backend value |
| 502 | {"error": "..."} |
Search backend unreachable |
POST /respite/audio/generate
Generate music or sound effects from a text prompt using Stable Audio 3.
Request body (JSON):
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
prompt |
string | Yes | β | Text description of desired audio |
duration |
number | No | 30 | Duration in seconds (1-120) |
steps |
integer | No | 8 | Inference steps (1-50). Higher = better quality, slower |
cfg_scale |
number | No | 1.0 | Classifier-free guidance scale (0-10). Higher = more prompt adherence |
seed |
integer | No | -1 | Random seed. -1 = random each time |
model |
string | No | "small-music" | Either "small-music" or "small-sfx" |
Request:
curl -X POST "https://cazyundee-respite-api.hf.space/respite/audio/generate" \
-H "Content-Type: application/json" \
-d '{"prompt": "upbeat synthwave melody", "duration": 30, "model": "small-music"}'
Response:
- Content-Type:
audio/wav - Body: Raw WAV audio binary (44100 Hz sample rate)
- The response is a downloadable audio file, not JSON
Error responses:
| Status | Body | Cause |
|---|---|---|
| 503 | {"detail":"Queue is full. Please try again later."} |
All 4 queue slots occupied |
| 500 | Server error | Generation failed |
Important notes:
- ZeroGPU shares one GPU across all requests. The server processes requests serially.
- First request triggers model loading (~30-60s). Subsequent requests use the cached model.
- The
modelfield selects between"small-music"(general music) and"small-sfx"(sound effects). - The returned audio is always WAV format, 44100 Hz sample rate, 16-bit PCM.
Rate Limits
- Max concurrent requests: 1 (GPU locked)
- Max queue size: 4 (503 if full)
- Audio generation: ~5-15s per request after model is loaded (varies by
stepsandduration) - Search: No limit (DuckDuckGo backend, no API key)
Error Handling
All errors return JSON with an "error" or "detail" field:
{"error": "Missing query parameter 'q'"}
{"detail": "Queue is full. Please try again later."}
Check the HTTP status code:
| Code | Meaning |
|---|---|
| 200 | Success |
| 400 | Bad request (missing/invalid parameters) |
| 502 | Upstream service unreachable |
| 503 | Queue full, try again |
Integration Notes for Chat Agents
Discover capabilities at runtime β Call
GET /respite/resourcesto enumerate what the API can do. Do not hardcode resource names.Search strategies by query type:
Query type Backend extract_top Region Current events, news news2-3 user's region Factual questions text1 wt-wtHow-to, tutorials text0 wt-wtRecent developments news0 wt-wtImages, visuals images0 user's region Use extraction for depth β When a user wants detailed information (e.g., "summarize this article"), set
extract_top=1to get the full page content. Don't overuse it β it adds latency.Audio generation is queued β Only 1 generation runs at a time. If the queue is full (503), retry with exponential backoff. Do not retry rapidly.
Model selection β Use
"small-music"for music, melodies, ambient sound. Use"small-sfx"for sound effects, foley, short percussive sounds.Response handling β
/respite/audio/generatereturns binary WAV, not JSON. Save to a.wavfile and play/serve it. Do not try to JSON-parse the response body.Specs for capability planning β
GET /respite/specstells you GPU type, VRAM, and queue limits. Use this to decide what to offer users (e.g., don't promise 4-minute tracks β max is 120s).Graceful degradation β If the API is down (503, 502), tell the user the service is temporarily unavailable rather than hanging. ZeroGPU Spaces can go cold when idle.