Spaces:
Sleeping
Sleeping
| # API Documentation | |
| The Unified Embedding API provides an OpenAI-compatible interface for dense embeddings, sparse embeddings, and document reranking. | |
| **Base URL** | |
| ```text | |
| http://localhost:7860/api/v1 | |
| ``` | |
| For Hugging Face Spaces: | |
| ```text | |
| https://YOUR_SPACE.hf.space/api/v1 | |
| ``` | |
| --- | |
| # Authentication | |
| Authentication is not required by default. | |
| When using the official OpenAI SDK, an API key is still required by the client but its value is ignored. | |
| ```python | |
| client = OpenAI( | |
| base_url="https://YOUR_SPACE.hf.space/api/v1", | |
| api_key="not-required" | |
| ) | |
| ``` | |
| --- | |
| # Endpoints | |
| | Method | Endpoint | Description | | |
| | ------ | --------------- | ------------------------------- | | |
| | POST | `/embeddings` | Generate dense embeddings | | |
| | POST | `/embed_sparse` | Generate sparse embeddings | | |
| | POST | `/rerank` | Rank documents based on a query | | |
| | GET | `/models` | List available models | | |
| | GET | `/health` | Health check | | |
| --- | |
| # Dense Embeddings | |
| Generate dense vector embeddings. | |
| ## Request | |
| **POST** | |
| ```text | |
| /api/v1/embeddings | |
| ``` | |
| ### Body | |
| ```json | |
| { | |
| "input": [ | |
| "Machine learning", | |
| "Artificial Intelligence" | |
| ], | |
| "model": "qwen3-0.6b" | |
| } | |
| ``` | |
| | Field | Type | Required | Description | | |
| | --------------- | ----------------- | -------- | ---------------------------------- | | |
| | input | string | string[] | Yes | Text or list of texts | | |
| | model | string | Yes | Configured embedding model | | |
| | encoding_format | string | No | Currently returns float embeddings | | |
| --- | |
| ## Response | |
| ```json | |
| { | |
| "object": "list", | |
| "data": [ | |
| { | |
| "object": "embedding", | |
| "index": 0, | |
| "embedding": [ | |
| 0.124, | |
| -0.554, | |
| 0.873 | |
| ] | |
| } | |
| ], | |
| "model": "qwen3-0.6b", | |
| "usage": { | |
| "prompt_tokens": 2, | |
| "total_tokens": 2 | |
| } | |
| } | |
| ``` | |
| --- | |
| ## Python (OpenAI SDK) | |
| ```python | |
| from openai import OpenAI | |
| client = OpenAI( | |
| base_url="https://YOUR_SPACE.hf.space/api/v1", | |
| api_key="not-required" | |
| ) | |
| response = client.embeddings.create( | |
| input="Large Language Models", | |
| model="qwen3-0.6b" | |
| ) | |
| embedding = response.data[0].embedding | |
| ``` | |
| --- | |
| ## cURL | |
| ```bash | |
| curl -X POST \ | |
| http://localhost:7860/api/v1/embeddings \ | |
| -H "Content-Type: application/json" \ | |
| -d '{ | |
| "input":"Large Language Models", | |
| "model":"qwen3-0.6b" | |
| }' | |
| ``` | |
| --- | |
| # Sparse Embeddings | |
| Generate sparse embeddings (SPLADE). | |
| ## Request | |
| ```text | |
| POST /api/v1/embed_sparse | |
| ``` | |
| ### Body | |
| ```json | |
| { | |
| "input": [ | |
| "Hybrid Search" | |
| ], | |
| "model": "splade" | |
| } | |
| ``` | |
| --- | |
| ## Response | |
| ```json | |
| { | |
| "object":"list", | |
| "data":[ | |
| { | |
| "index":0, | |
| "indices":[ | |
| 24, | |
| 91, | |
| 118 | |
| ], | |
| "values":[ | |
| 0.87, | |
| 0.41, | |
| 0.29 | |
| ] | |
| } | |
| ] | |
| } | |
| ``` | |
| --- | |
| # Document Reranking | |
| Ranks retrieved documents according to their relevance to a query. | |
| ## Request | |
| ```text | |
| POST /api/v1/rerank | |
| ``` | |
| ### Body | |
| ```json | |
| { | |
| "query":"What is LangGraph?", | |
| "documents":[ | |
| "LangGraph is an orchestration framework.", | |
| "React is a frontend library.", | |
| "FastAPI is a Python framework." | |
| ], | |
| "model":"bge-reranker-base", | |
| "top_k":2 | |
| } | |
| ``` | |
| | Field | Type | Required | | |
| | --------- | -------- | -------- | | |
| | query | string | Yes | | |
| | documents | string[] | Yes | | |
| | model | string | Yes | | |
| | top_k | integer | No | | |
| --- | |
| ## Response | |
| ```json | |
| { | |
| "results":[ | |
| { | |
| "index":0, | |
| "score":0.992, | |
| "text":"LangGraph is an orchestration framework." | |
| }, | |
| { | |
| "index":2, | |
| "score":0.843, | |
| "text":"FastAPI is a Python framework." | |
| } | |
| ] | |
| } | |
| ``` | |
| --- | |
| # Available Models | |
| Returns every configured model. | |
| ## Request | |
| ```text | |
| GET /api/v1/models | |
| ``` | |
| ## Response | |
| ```json | |
| { | |
| "models":[ | |
| { | |
| "id":"qwen3-0.6b", | |
| "type":"embeddings" | |
| }, | |
| { | |
| "id":"splade", | |
| "type":"sparse-embeddings" | |
| }, | |
| { | |
| "id":"bge-reranker-base", | |
| "type":"rerank" | |
| } | |
| ] | |
| } | |
| ``` | |
| --- | |
| # Health Check | |
| Check server status. | |
| ## Request | |
| ```text | |
| GET /health | |
| ``` | |
| ## Response | |
| ```json | |
| { | |
| "status":"healthy" | |
| } | |
| ``` | |
| --- | |
| # Error Responses | |
| The API follows standard HTTP status codes. | |
| | Status | Description | | |
| | ------ | --------------------- | | |
| | 200 | Success | | |
| | 400 | Invalid request | | |
| | 404 | Resource not found | | |
| | 422 | Validation error | | |
| | 500 | Internal server error | | |
| Example: | |
| ```json | |
| { | |
| "detail":"Model 'unknown-model' is not configured." | |
| } | |
| ``` | |
| --- | |
| # Model Configuration | |
| Available models are defined in: | |
| ```text | |
| config/models.yaml | |
| ``` | |
| Example: | |
| ```yaml | |
| models: | |
| qwen3: | |
| name: Qwen/Qwen3-Embedding-0.6B | |
| type: embeddings | |
| splade: | |
| name: prithivida/Splade_PP_en_v1 | |
| type: sparse-embeddings | |
| reranker: | |
| name: BAAI/bge-reranker-base | |
| type: rerank | |
| ``` | |
| Adding a new model only requires updating this configuration file. No application code changes are needed. | |
| --- | |
| # OpenAI Compatibility | |
| The `/embeddings` endpoint implements the OpenAI Embeddings API specification. | |
| Migrating from OpenAI usually requires changing only the client configuration. | |
| ```python | |
| client = OpenAI( | |
| base_url="https://YOUR_SPACE.hf.space/api/v1", | |
| api_key="not-required" | |
| ) | |
| ``` | |
| Your existing embedding requests can remain unchanged. | |