Respite-API / API_REFERENCE.md
cazyundee's picture
Update API reference with news/images backends and extraction support
c638c2d verified
|
Raw
History Blame Contribute Delete
9.93 kB

A newer version of the Gradio SDK is available: 6.28.0

Upgrade

Respite API Reference

Base URL: https://cazyundee-respite-api.hf.space

Version: 1.0.0
Authentication: None required
Content Types: application/json, audio/wav


Overview

Respite is a general-purpose AI API running on Hugging Face ZeroGPU infrastructure. It currently exposes two capabilities: web search (DuckDuckGo, no API key) and audio generation (Stable Audio 3 Small). The API is self-describing β€” call the discovery endpoints to enumerate capabilities and server specs at runtime.


Endpoints

GET /respite/specs

Returns server metadata, hardware specs, and runtime limits.

Response:

{
  "name": "Respite API",
  "version": "1.0.0",
  "description": "General-purpose AI API server with audio generation capabilities.",
  "base_path": "/respite",
  "authentication": "none",
  "content_types": ["application/json", "audio/wav"],
  "resources_endpoint": "/respite/resources",
  "specs_endpoint": "/respite/specs",
  "limits": {
    "max_concurrent_requests": 1,
    "max_queue_size": 4,
    "audio_max_duration_seconds": 120
  },
  "runtime": {
    "platform": "Linux-6.12.94-123.192.amzn2023.x86_64-x86_64-with-glibc2.36",
    "python_version": "3.10.13",
    "cpu_cores": 16,
    "host_cpu_cores": 192,
    "ram_bytes": 104000000000,
    "storage_total_bytes": 8589854879744,
    "storage_used_bytes": 6852819365888,
    "storage_free_bytes": 1737035513856,
    "hardware": "zero-a10g",
    "gpu": {
      "type": "NVIDIA RTX Pro 6000 Blackwell",
      "vram_bytes": 51539607552,
      "shared_zero_gpu": true
    }
  }
}

Fields:

Field Description
limits.max_concurrent_requests Only 1 request can be processed at a time (ZeroGPU queue)
limits.max_queue_size Max 4 requests queued; beyond that, 503
limits.audio_max_duration_seconds Max audio length per generation
runtime.gpu.shared_zero_gpu GPU time is shared β€” requests are queued and batched
runtime.ram_bytes Container memory limit (not host total)
runtime.cpu_cores Container CPU limit (not host total)

GET /respite/resources

Lists all callable capabilities.

Response:

{
  "resources": {
    "audio_generation": {
      "name": "Audio generation",
      "description": "Generate music or sound effects from a text prompt.",
      "endpoint": "/respite/audio/generate",
      "method": "POST",
      "input": {
        "prompt": "string",
        "duration": "number (1-120 seconds)",
        "steps": "integer (1-50)",
        "cfg_scale": "number (0-10)",
        "seed": "integer (-1 for random)",
        "model": "small-music | small-sfx"
      },
      "output": "WAV audio file"
    },
    "web_search": {
      "name": "Web search",
      "description": "Search the web via DuckDuckGo. Supports general, news, and image search. No API key required.",
      "endpoint": "/respite/search",
      "method": "GET",
      "input": {
        "q": "string (search query)",
        "backend": "string: 'text' (default), 'news', or 'images'",
        "max_results": "integer (1-50, default 10)",
        "extract_top": "integer (0-5, default 0) - fetch full page content for top N results",
        "region": "string (default 'wt-wt') - e.g. 'us-en', 'gb-en', 'de-de'"
      },
      "output": "JSON with title, url, content per result. News results include date and source."
    }
  }
}

GET /respite/search

Search the web via DuckDuckGo. Supports general text, news, and image search. No API key required.

Query Parameters:

Parameter Type Required Default Description
q string Yes β€” Search query
backend string No "text" "text" (general), "news" (news articles), or "images"
max_results integer No 10 Number of results (1-50)
extract_top integer No 0 Fetch full page content for top N results (0-5). Adds latency but gives richer content.
region string No "wt-wt" Region code: "us-en", "gb-en", "de-de", "wt-wt" (worldwide)

Request (text search):

GET /respite/search?q=what+is+python&backend=text&max_results=3

Request (news search):

GET /respite/search?q=latest+AI+news&backend=news&max_results=5

Request (news + extraction):

GET /respite/search?q=nvidia+earnings&backend=news&max_results=3&extract_top=2

Response (text):

{
  "query": "what is python",
  "backend": "text",
  "results": [
    {
      "title": "Python (programming language) - Wikipedia",
      "url": "https://en.wikipedia.org/wiki/Python_(programming_language)",
      "content": "Python is a high-level, general-purpose programming language that emphasizes code readability..."
    }
  ],
  "total_results": 1
}

Response (news):

{
  "query": "latest AI news",
  "backend": "news",
  "results": [
    {
      "title": "Strong AI chip demand fuels Nvidia's Q2 results",
      "url": "https://apnews.com/article/...",
      "content": "Nvidia's latest quarterly results once again blew past Wall Street's expectations...",
      "source": "Associated Press News",
      "date": "2026-08-27T00:00:00+00:00",
      "image": "https://..."
    }
  ],
  "total_results": 1
}

Response (with extraction):

{
  "query": "nvidia earnings",
  "backend": "news",
  "results": [
    {
      "title": "Nvidia earnings report",
      "url": "https://...",
      "content": "Short snippet...",
      "source": "Reuters",
      "date": "2026-08-27T00:00:00+00:00",
      "extracted_content": "Full article text, up to 8000 characters..."
    }
  ],
  "total_results": 1
}

Error responses:

Status Body Cause
400 {"error": "Missing query parameter 'q'"} No query provided
400 {"error": "Invalid backend 'x'. Use 'text', 'news', or 'images'."} Bad backend value
502 {"error": "..."} Search backend unreachable

POST /respite/audio/generate

Generate music or sound effects from a text prompt using Stable Audio 3.

Request body (JSON):

Field Type Required Default Description
prompt string Yes β€” Text description of desired audio
duration number No 30 Duration in seconds (1-120)
steps integer No 8 Inference steps (1-50). Higher = better quality, slower
cfg_scale number No 1.0 Classifier-free guidance scale (0-10). Higher = more prompt adherence
seed integer No -1 Random seed. -1 = random each time
model string No "small-music" Either "small-music" or "small-sfx"

Request:

curl -X POST "https://cazyundee-respite-api.hf.space/respite/audio/generate" \
  -H "Content-Type: application/json" \
  -d '{"prompt": "upbeat synthwave melody", "duration": 30, "model": "small-music"}'

Response:

  • Content-Type: audio/wav
  • Body: Raw WAV audio binary (44100 Hz sample rate)
  • The response is a downloadable audio file, not JSON

Error responses:

Status Body Cause
503 {"detail":"Queue is full. Please try again later."} All 4 queue slots occupied
500 Server error Generation failed

Important notes:

  • ZeroGPU shares one GPU across all requests. The server processes requests serially.
  • First request triggers model loading (~30-60s). Subsequent requests use the cached model.
  • The model field selects between "small-music" (general music) and "small-sfx" (sound effects).
  • The returned audio is always WAV format, 44100 Hz sample rate, 16-bit PCM.

Rate Limits

  • Max concurrent requests: 1 (GPU locked)
  • Max queue size: 4 (503 if full)
  • Audio generation: ~5-15s per request after model is loaded (varies by steps and duration)
  • Search: No limit (DuckDuckGo backend, no API key)

Error Handling

All errors return JSON with an "error" or "detail" field:

{"error": "Missing query parameter 'q'"}
{"detail": "Queue is full. Please try again later."}

Check the HTTP status code:

Code Meaning
200 Success
400 Bad request (missing/invalid parameters)
502 Upstream service unreachable
503 Queue full, try again

Integration Notes for Chat Agents

  1. Discover capabilities at runtime β€” Call GET /respite/resources to enumerate what the API can do. Do not hardcode resource names.

  2. Search strategies by query type:

    Query type Backend extract_top Region
    Current events, news news 2-3 user's region
    Factual questions text 1 wt-wt
    How-to, tutorials text 0 wt-wt
    Recent developments news 0 wt-wt
    Images, visuals images 0 user's region
  3. Use extraction for depth β€” When a user wants detailed information (e.g., "summarize this article"), set extract_top=1 to get the full page content. Don't overuse it β€” it adds latency.

  4. Audio generation is queued β€” Only 1 generation runs at a time. If the queue is full (503), retry with exponential backoff. Do not retry rapidly.

  5. Model selection β€” Use "small-music" for music, melodies, ambient sound. Use "small-sfx" for sound effects, foley, short percussive sounds.

  6. Response handling β€” /respite/audio/generate returns binary WAV, not JSON. Save to a .wav file and play/serve it. Do not try to JSON-parse the response body.

  7. Specs for capability planning β€” GET /respite/specs tells you GPU type, VRAM, and queue limits. Use this to decide what to offer users (e.g., don't promise 4-minute tracks β€” max is 120s).

  8. Graceful degradation β€” If the API is down (503, 502), tell the user the service is temporarily unavailable rather than hanging. ZeroGPU Spaces can go cold when idle.