MiniSearch / docs /api.md
system's picture
system HF Staff
Sync from felladrin/MiniSearch@ff7e32b
5748ede verified
|
Raw
History Blame Contribute Delete
20.8 kB

HTTP API Reference

Every server-side route is a Vite middleware hook declared in vite.config.ts (see the server hook system section of docs/overview.md), so the endpoints below behave identically in dev (vite) and in the production preview (vite preview) that the Docker image runs.

These endpoints exist for the browser app rather than as a public API: they are unversioned, and their shapes are the ones the client happens to need. They are documented here because the owner of an instance can reach them, and because reading middleware should not be the only way to learn what a request has to carry.

Method Path Token Rate limit Purpose
GET /search/text required shared Text results from SearXNG, reranked
GET /search/images required shared Image results from SearXNG, reranked
GET /page-content required shared Readable text from result pages, for answer grounding
GET /thumbnail required own budget One image-result thumbnail, fetched server-side
POST /inference required shared Streaming chat completion from the internal API
GET /status none none Uptime, counters, and service health
GET /api/config none none Runtime client configuration, including the search token
POST /api/validate-access-key none shared Access key check
GET /dictation-models/<version>/<file> none none Speech-to-text model files, served from the instance

A path that matches no hook falls through to Vite's static handler, which is also what a method mismatch does on /api/config and /api/validate-access-key: those two answer only their documented method and call next() otherwise, so a GET on the validation endpoint returns the app shell, not a 405.

Those two compare the whole request target, query string included, so /api/config?v=2 falls through as well. The GET endpoints in the table above do not check the method at all: a POST to /search/text is served like a GET. Only /inference answers 405.

Authentication

Search token

The token is CSRF protection, not authorization: it proves a request came from a page this server served. server/searchToken.ts generates it on first use in each process and holds it for that process's life, and /api/config hands it out. It is not a build artifact: a token written during an image build would be shared by every container of that build.

A client never sends the raw token. It hashes it with argon2id using the parameters in shared/argon2Parameters.ts (m=512, t=16, p=1, 32-byte output, encoded form) and sends the resulting hash as the token query parameter:

TOKEN=$(curl -s http://localhost:7860/api/config | jq -r .searchToken)
# hash $TOKEN with argon2id using the parameters above; see
# client/modules/searchTokenHash.ts
curl "http://localhost:7860/search/text?q=hello&token=$ARGON2_HASH"

The hash is salted per client, so it differs between browsers and every one of them verifies. The server checks the encoded parameter block before running argon2Verify, so a hash carrying different parameters is refused for free. The whole lifecycle is in the search token lifecycle section of docs/security.md.

Verification runs before parameters are parsed, so a malformed request from an unauthenticated caller still costs a rate-limit point. The exception is /inference, which answers 405 and 415 before it looks at the token, so those two rejections need no token and spend no rate-limit point. The failures below are the same on every token-gated endpoint:

Status Body When
429 {"error":"Too many requests."} The bucket for this client is empty
400 {"error":"Missing token."} No token parameter
401 {"error":"Invalid token."} The hash does not verify against this server's token

A 401 that persists across reloads usually means two processes hold different tokens: the server logs that once, on the way out.

The 400 and 401 rejections have a second shape for a person rather than a program. The token rides in the query string, so a search URL can be bookmarked, shared or set as a browser's search engine, and every such URL dies when the instance rotates its token. When the request's Accept header names text/html, which is what a browser sends when it navigates to a URL, the same status comes back with Content-Type: text/html; charset=utf-8 and a short static page that says the token was rotated and links to /, where the app takes a current token through /api/config. The page carries no token, sets no cookie and redirects nowhere, and nothing from the request is echoed into it. Any other Accept, including */* and none at all, gets the JSON body above, so API clients see no change; 429 is JSON either way. The rejection is counted the same whichever shape it is answered with.

Access keys

ACCESS_KEYS gates the app's UI, not these endpoints. The client asks /api/validate-access-key before rendering and stores the accepted hash locally; nothing about that key travels on a later search or inference request. An instance that must not answer unauthenticated callers needs something in front of it.

Rate limits

server/verifyTokenAndRateLimit.ts keeps two in-memory buckets, both keyed by client IP:

Bucket Budget Consumed by
shared 10 requests / 10 s /search/text, /search/images, /page-content, /inference, /api/validate-access-key
thumbnail 60 requests / 10 s /thumbnail

/thumbnail has its own budget because one image search fans out into up to 30 tile loads, which would otherwise consume a user's whole search budget. The point is consumed before token verification and before any argon2 work, so a flood of bogus tokens is bounded too.

The key is the TCP peer address. X-Forwarded-For and X-Real-IP are honored only when TRUST_PROXY is true or 1; on a directly-exposed instance they are client-controlled, and trusting them would hand every request a fresh identity.

Endpoints

GET /search/text

Query parameters:

Name Required Description
q yes Search query, trimmed, 1 to 2000 characters
token yes Search token hash
limit no How many of SearXNG's results to consider, default and maximum 30; a value that is not a positive integer falls back to the default. SearXNG is not sent this number: it caps a single SearXNG response after duplicate URLs are dropped, and the score filter below can return fewer

Responds with a JSON array of tuples, in the order the client renders them:

[
  ["Result title", "Snippet text", "https://example.com/page", 3.7]
]

Between SearXNG and the response the server drops duplicate URLs, keeps the first limit of what remains, and discards results with no title or no snippet. Index 0 is the first result that survives that.

The order is not a score sort. That first surviving result is pinned at index 0 and is the one result the score filter never drops; the next nine are the first nine survivors in SearXNG's order, sorted by score among themselves; everything after them is sorted by score too. So a result SearXNG ranked eleventh can outscore index 1 and still arrive at index 10 (server/rankSearchResults.ts, and the preserve top results section of docs/reranking.md).

Reranking drops results rather than only reordering them. Every score is first shifted by the absolute value of the lowest score in the batch, which puts the lowest at zero only when it is negative; everything below mean - 0.3 * standardDeviation on that shifted scale is filtered out. If that leaves fewer than 40% of the batch, the threshold becomes 40% of the highest shifted score instead, so the fallback keeps far more of an all-positive batch than of one whose lowest score is negative. The response carries only the survivors, so neither threshold can be recomputed from it. On /search/text the filter sees results 2..N only, since index 0 is exempt; on /search/images it sees all of them. Fewer results than limit is therefore normal on both, not a sign of an outage.

The fourth element is the reranker's raw relevance logit, deliberately not passed through a sigmoid (see docs/reranking.md). It is absent when the reranker is unhealthy or a rerank call fails, in which case SearXNG's own order is returned unranked and unfiltered, so a client must treat the score as optional.

Failures:

Status Body When
400 {"error":"Missing query parameter"} q missing or blank
400 {"error":"Query parameter must not exceed 2000 characters"} q too long
502 {"error":"Search service unavailable"} SearXNG unreachable, so an outage is distinguishable from a search with no matches
500 {"error":"Internal server error"} Anything else

GET /search/images

Same parameters and failures as /search/text. The hook claims the whole /search/ prefix and treats every path that does not start with /search/text as an image search, so /search/anything is an image search.

The ordering above does not carry over: image results are a plain score sort, nothing is pinned at index 0, and no result is exempt from the score filter (preserveTopResults is only passed for text, see docs/reranking.md).

Responds with a JSON array of four-element tuples:

[
  [
    "Image title",
    "https://example.com/page-that-embeds-it",
    "https://example.com/thumb.jpg",
    "https://example.com/full.jpg"
  ]
]

In order: the title, cut to 100 characters because that is the length the reranker sees, the page the image was found on, the thumbnail URL exactly as SearXNG returned it, and the full image. The last is the embeddable player URL for a video result, and an empty string when there is nothing to link to.

The response never waits on a thumbnail host; the client loads each tile through /thumbnail. A result SearXNG returned without a thumbnail URL is dropped before the response, since the grid would have nothing to show for it.

GET /page-content

Reads the pages behind a handful of results and returns the passages that match the query. Unlike /search/, this endpoint fetches URLs the caller chose, so it is SSRF-guarded, byte-capped and timed out per page (see docs/page-content.md).

The path is matched exactly, so /page-content/anything falls through.

Name Required Description
q yes Query the passages are ranked against, 1 to 2000 characters
token yes Search token hash
url yes Page to read, http or https, up to 2048 characters. Repeat for more; at most 6 per request, counted before duplicates are removed, so 7 URLs are refused even when two of them are the same

Responds with an object keyed by URL:

{
  "https://example.com/page": "The passage that best covers the query.\nThe next best passage from the same page."
}

A page that could not be read, was refused by the SSRF guard, or yielded too little text is simply absent from the object; {} means nothing was read. That is not an error, and the client degrades to snippet-only answers for those results.

Failures: 400 with the first validation message (Missing query parameter, Query parameter must not exceed 2000 characters, Missing url parameter, Invalid URL parameter, No more than 6 URLs can be read per request), or 500 {"error":"Internal server error"}. A url over 2048 characters is answered Invalid URL parameter, the same as any other unusable URL.

GET /thumbnail

Fetches one image-result thumbnail server-side, so the browser never requests a search-result URL directly.

Name Required Description
u yes Thumbnail URL, up to 2048 characters
token yes Search token hash

A hit responds with the image bytes and the upstream's content type, restricted to image/avif, image/bmp, image/gif, image/jpeg, image/png, image/webp, image/x-icon and image/vnd.microsoft.icon. SVG is refused because it would be a scriptable document on this origin. Responses carry X-Content-Type-Options: nosniff, Content-Security-Policy: default-src 'none'; sandbox, and Cache-Control: private, max-age=3600.

The fetch is bounded by a 3 s deadline covering DNS and every hop, at most 3 redirects, each of them re-validated, and 500 KB of body. Successful fetches are held in an in-process LRU (100 entries, 50 MB); failures never are.

Status Body When
400 {"error":"Missing thumbnail URL"} u missing
400 {"error":"Thumbnail URL too long"} u over 2048 characters
403 {"error":"Refusing to fetch a thumbnail from a non-public or unresolvable address"} The host is in private space, does not resolve, or u is not a parseable http/https URL. /page-content answers 400 for that last case; this endpoint does not distinguish it
502 {"error":"Thumbnail could not be fetched"} Upstream failed, timed out, redirected more than 3 times, answered with a type outside the list, or sent an empty body
500 {"error":"Internal server error"} Anything the hook itself threw, caught so Vite's connect stack does not see an unhandled rejection

Error responses carry Cache-Control: no-store, since neither a refusal nor an upstream failure is a stable property of the URL.

POST /inference

Streams a chat completion from the API configured through INTERNAL_OPENAI_COMPATIBLE_API_* (see docs/configuration.md), so an instance can offer a model without publishing its key.

Requires Content-Type: application/json and the token as the token query parameter. The body is at most 1 MiB:

{
  "messages": [{ "role": "user", "content": "Hello" }],
  "temperature": 0.7,
  "top_p": 0.9,
  "max_tokens": 512
}

messages needs at least one entry, each with a role of system, user or assistant and a string content. temperature is clamped to 0-2, top_p to 0-1, and max_tokens to the server's defaultMaxTokens (server/config/modelConfig.ts).

The response is an OpenAI-compatible SSE stream of chat.completion.chunk objects, ending with a chunk whose finish_reason is "stop" and then data: [DONE]:

data: {"id":"chatcmpl-1730000000000","object":"chat.completion.chunk","created":1730000000,"model":"some-model","choices":[{"index":0,"delta":{"content":"Hel"},"finish_reason":null}]}

data: {"id":"chatcmpl-1730000000000","object":"chat.completion.chunk","created":1730000000,"model":"some-model","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

When a model stalls or fails before the first token, the server retries with another model from the provider's list, up to 5 attempts. That ladder needs the list: with INTERNAL_OPENAI_COMPATIBLE_API_MODEL pinned there is no second model to try, so the first failure is the last attempt. Once bytes have been written the status line is already sent, so a later failure arrives as a data frame carrying an error field, followed by [DONE], rather than as a status code.

The stream carries Content-Type: text/event-stream, Cache-Control: no-cache and Connection: keep-alive.

Failures before the stream starts:

Status Body When
405 {"error":"Method Not Allowed"} Not a POST; the response carries Allow: POST. Checked before the token, so it needs no token and spends no rate-limit point
415 {"error":"Unsupported Media Type"} Content-Type is not JSON; also checked before the token
400 {"error":"Invalid request body"} or {"error":"Invalid request body: <field> <message>"} Unparseable or schema-invalid body
413 {"error":"Request body too large"} Body over 1 MiB
500 {"error":"OpenAI API configuration is missing"} INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL or _API_KEY unset
500 {"error":"Failed to fetch available models"} No model configured and the provider's listing failed
500 {"error":"No model available"} The listing succeeded but was empty
503 {"error":"Service unavailable - all models failed","lastError":"..."} Every attempt failed before the first token
500 {"error":"Internal server error","message":"..."} Anything else the hook threw

GET /status

The hook claims the /status prefix, so /status/anything answers the same way.

Unauthenticated and not rate-limited: uptime, the counters accumulated since the last restart, and the health of the reranker, the bi-encoder and SearXNG. The field reference is the /status section of docs/overview.md.

Nothing in the response is per-user: queries, URLs and client addresses are never recorded, only aggregate outcomes (see the privacy section of docs/security.md).

GET /api/config

Unauthenticated by design: the client needs it before it can prove anything, and the access key page depends on it. Served with Cache-Control: no-store, so a restart with different environment variables takes effect on the next reload.

{
  "accessKeysEnabled": false,
  "accessKeyTimeoutHours": 0,
  "wllamaDefaultModelId": "...",
  "internalApiEnabled": false,
  "internalApiName": "...",
  "defaultInferenceType": "...",
  "searchToken": "..."
}

The shape is ServerConfig in shared/serverConfig.ts. It reports whether a feature is on plus its display defaults, and never ACCESS_KEYS, INTERNAL_OPENAI_COMPATIBLE_API_KEY, or any other secret: a field added to that interface is published to anyone who can reach the instance (see the /api/config exposure section of docs/security.md).

POST /api/validate-access-key

Checks a client-hashed access key against ACCESS_KEYS. Consumes a point from the shared bucket before the argon2 loop, since a wrong hash costs one full verification per configured key.

{ "accessKeyHash": "$argon2id$v=19$m=512,t=16,p=1$..." }

Responds {"valid":true} or {"valid":false}. A hash whose parameter block differs from shared/argon2Parameters.ts is answered {"valid":false} without any verification running.

The body is capped at 4 KiB, which one encoded hash is nowhere near. Past that the answer is a 413 sent while the caller is still uploading, so it carries Connection: close and the socket is dropped once it has flushed.

Status Body When
429 {"error":"Too many requests."} Rate limited, kept distinct from a wrong key so the UI can say "try again"
400 {"valid":false,"error":"Invalid request"} Body is not JSON
413 {"error":"Request body too large"} Body over 4 KiB

GET /dictation-models/<version>/<file>

Serves the pinned speech-to-text model, so the page never contacts download.moonshine.ai. The version segment is DICTATION_MODEL_VERSION from shared/dictationModel.ts, currently quantized_26_07_30.

Only seven filenames resolve: frontend.ort, encoder.ort, adapter.ort, cross_kv.ort, decoder_kv.ort, streaming_config.json and tokenizer.bin. Anything else is a 404, so the route cannot be used as a proxy against the upstream host. No token and no rate-limit budget: the whitelist is what bounds it, and every file is on disk after the first request.

The first request for a file fetches it from the upstream, checks it against a pinned SHA-256, and writes it under DICTATION_MODELS_DIR in a subdirectory named after the version. Concurrent first requests share one transfer. Responses are streamed and carry Cache-Control: public, max-age=31536000, immutable, which the version segment makes safe.

Status Body When
404 Unknown dictation model file The filename is not one of the seven
502 Could not serve <file>: <reason> The upstream failed, sent no body, exceeded the 64 MiB cap, or the bytes did not match the pinned digest

A failure once the body has started cannot be reported, because the headers are already gone; the response is ended rather than having an error appended to a truncated file.

Related Topics

  • Overview: docs/overview.md - Server hook system and the /status field reference
  • Security: docs/security.md - Token lifecycle, access control, and privacy model
  • Configuration: docs/configuration.md - Environment variables the endpoints read
  • Page Content: docs/page-content.md - What /page-content does with the pages it reads
  • Reranking: docs/reranking.md - Where the search score comes from