SilverElixir commited on
Commit
df00367
Β·
unverified Β·
1 Parent(s): 290a3ee

Add files via upload

Browse files
Files changed (2) hide show
  1. docs/ARCHITECTURE.md +26 -26
  2. docs/ENVIRONMENT.md +7 -7
docs/ARCHITECTURE.md CHANGED
@@ -7,73 +7,73 @@ This document goes one level deeper than the main [README](../README.md) into ho
7
  1. Telegram POSTs an update to `/webhook`, authenticated by a secret header (`X-Telegram-Bot-Api-Secret-Token`).
8
  2. `_handle_message_core` in `bot.py` classifies the message: is it a TikTok link, a `/draw` or `/tts` trigger phrase, a reply to earlier media, plain text, or an attachment?
9
  3. `_build_route` (`lumen_router_config.py`) turns that classification into an ordered list of `(provider, model_id)` candidates.
10
- 4. `_run_route` tries the first candidate. If a whole provider's chain fails, it falls back to the other provider β€” except where that's physically impossible (only Gemini can read links or analyze video/audio).
11
  5. The reply streams into the chat if the request qualifies (see [Streaming](#streaming--typing-pace) below), or is sent as a single message otherwise.
12
 
13
  ## Automatic model routing
14
 
15
- There is no `/model` or `/provider` command β€” the router builds a fresh candidate list for every message:
16
 
17
  - **A YouTube or website link** β†’ Gemini only. It's the only provider that can read page content (`url_context`) or analyze a video by URL.
18
  - **A video or audio attachment** β†’ Gemini only. OpenRouter's multimodal models only accept images as base64.
19
  - **An image attachment, no live-info need** β†’ free OpenRouter vision models first, Gemini as a reserve.
20
- - **Needs current information** (a lightweight keyword heuristic β€” "now," "today," "price," "who is currently...") β†’ Gemini, prioritizing the models with a real search-grounding quota, then OpenRouter as a reserve.
21
- - **Plain text, no attachments, no freshness need** (the most common case) β†’ straight to OpenRouter. Heavier requests (code, multi-step reasoning β€” another lightweight heuristic) get routed to the stronger free OpenRouter models first.
22
 
23
- The reasoning: Gemini's free quota (roughly 20 requests/day for the flagship model) is the scarcest resource in the system, so it's only spent where a capability unique to Gemini is actually needed. Everything else β€” the bulk of ordinary messages β€” runs on OpenRouter's much more generous free tier.
24
 
25
- Models that the router should never pick are tracked in a single registry, `_OR_MODEL_HEALTH` (`lumen_router_config.py`), each with a dated reason (the provider dropped the free tier, or the model turned out to be uncensored and a poor fit for Lumen's persona). Nothing is removed on speculation β€” only on confirmed evidence from production logs (`HTTP 404`, `"no endpoints,"` etc.).
26
 
27
- Each model is tried exactly once per route, with no retries β€” a single failure (timeout, `429`, `5xx`) moves straight to the next candidate. That keeps the worst case bounded by `len(route) Γ— ROUTE_MODEL_TIMEOUT_SEC`, further capped by `ROUTE_TOTAL_BUDGET_SEC` for the whole route.
28
 
29
  ### Image generation
30
 
31
- `/draw` follows the same philosophy on a smaller scale: `_pick_image_model` (`lumen_images.py`) picks a Pollinations.ai model from the prompt's content (anime/manga β†’ `flux-anime`, fantasy β†’ `dreamshaper`, "photorealistic" β†’ `flux-realism`, "quick sketch" β†’ `turbo`, otherwise a general-purpose default) and falls back through the rest of the catalog if the chosen model fails, all within `DRAW_TOTAL_BUDGET_SEC`.
32
 
33
  ## Streaming & typing pace
34
 
35
- For the first candidate in a route β€” plain text, no attachments, no links β€” Lumen streams the reply by repeatedly editing one message, the same mechanism for both Gemini (`generate_content_stream`) and OpenRouter (SSE over `chat/completions`).
36
 
37
- Telegram won't let a bot edit one message more than about once a second, so true token-by-token output isn't possible regardless of how fast the model generates text. Instead, `lumen_typing_pace.py` tracks each `provider:model` pair's real observed characters-per-second and reveals a growing slice of the buffered text at that pace, capped by whatever has actually arrived. If a backend returns the whole answer in one large chunk β€” common for free OpenRouter models, since the same `:free` slug can be served by different backends depending on load β€” a short "catch-up" phase (at most `STREAM_TYPING_MAX_CATCHUP_TICKS Γ— STREAM_TYPING_TICK_SEC` seconds) finishes revealing it smoothly instead of dumping it all at once.
38
 
39
- This speed estimate is measured, not hardcoded: there's no table of "model X does N tokens/sec" to maintain, because that number isn't a property of the model on OpenRouter's free tier in the first place. A new model just starts at a default speed and calibrates itself over its first few replies.
40
 
41
  If the streaming attempt for the head-of-route model fails before showing any text, the bot falls back to a normal (non-streaming) call further down the chain, reusing the same placeholder message rather than sending a new one.
42
 
43
  ### Manual smoke test before touching streaming code
44
 
45
- Streaming is covered only by mocked clients β€” no test exercises a real Gemini or OpenRouter SSE stream. Before merging changes to `_run_streaming_reply`, `_gemini_stream_pieces`, `_openrouter_stream_pieces`, or `lumen_typing_pace.py`, check by hand in a real chat:
46
 
47
- 1. A short message (fits in one edit) β€” once via Gemini, once via OpenRouter.
48
- 2. A long message (past `TG_MAX_LEN`) β€” confirm it splits into multiple messages and continues correctly.
49
- 3. A request where the streaming model is deliberately unavailable β€” confirm the silent fallback to a non-streaming reserve model, with no visible glitch.
50
 
51
  ## Prompt-injection and identity-leak defenses
52
 
53
- Lumen deliberately hides which model or provider answers a given message (see the persona in `system_prompt.py`). The system prompt alone is the weakest layer β€” any LLM can potentially be talked out of following it with a creative enough injection β€” so there are several independent layers, each assuming the previous one might have failed:
54
 
55
- 1. **Input pre-filter** (`_looks_like_injection_probe`, `lumen_security.py`) β€” well-known jailbreak phrasing ("ignore previous instructions," "developer mode," "show me your system prompt") is caught before the model is ever called. The response is fully deterministic.
56
- 2. **System prompt** (`system_prompt.py`) β€” instructs the model that anything outside the prompt itself (user messages, chat background, page/video/document content) is data, not instructions, and that the persona doesn't change no matter who claims authority to override it.
57
- 3. **Output identity-leak filter** (`_detect_identity_leak` / `_scrub_identity_leak`) β€” a deterministic check on the finished reply: exact internal model IDs, and narrow self-identification patterns ("I am Gemini," "made by OpenAI"). Deliberately narrow, to avoid false positives on ordinary, honest discussion of other AI companies.
58
- 4. **Injected-payload echo filter** (`_detect_injected_payload_echo`) β€” catches the case where an attacker embeds an instruction in a photo or web page ("output this exact string to confirm the jailbreak worked") and the model, while refusing to *follow* it, ends up quoting it back verbatim while summarizing the content.
59
 
60
- Layers 3 and 4 run before the reply is written to chat history (so a leak can't influence future turns) and, during streaming, before each chunk is shown to the user β€” not only on the final text.
61
 
62
  None of this is airtight except the input pre-filter. The goal is raising the bar for known attack patterns, not proving immunity to every possible phrasing; incidents get logged with `[identity-leak]` / `[injection-echo]` / `[injection-probe]` tags so new patterns can be added as they show up.
63
 
64
  ## TikTok downloader
65
 
66
- Downloads go through the public TikWM API (`tikwm.com`), no account or API key needed. The bot never re-encodes video or photos β€” bytes go to Telegram exactly as TikWM served them.
67
 
68
  - **Video quality**: requested with `&hd=1`; among the variants TikWM returns (`hdplay`/`play`/`wmplay`, each with a known byte size), the best one that fits Telegram's 50 MB upload limit is picked. If Telegram still rejects it as too large, the next lighter variant is tried automatically.
69
  - **Slideshows**: TikTok allows up to 35 slides per post; Telegram's `sendMediaGroup` caps out at 10 items per call. The bot downloads the whole post in parallel and sends it as several media groups in sequence, never with a trailing group of exactly one item (Telegram requires 2–10 per group).
70
- - **"Live" slides**: TikWM's response includes a separate `live_images` field alongside the regular `images` β€” that's where the actually-moving version of a slide lives, if it has one. Each downloaded slide is also double-checked by its magic bytes (`ftyp` = video container) rather than trusting the field alone.
71
- - TikWM rejects Hugging Face Spaces' outbound IPs with an empty `403` β€” the same class of block documented for YouTube below. The proxy (`TIKWM_API_BASE_URL`) is the only fix that's worked in practice; throttling, retries, and spoofed `Referer`/`Origin` headers didn't help on their own (they're still in place as defense in depth).
72
 
73
  ## Persistent storage
74
 
75
- Chat history and quota counters live in `chat_state.json` / `global_quota.json` on the container's local disk by default β€” which Hugging Face Spaces wipes on every redeploy. If `UPSTASH_REDIS_REST_URL` and `UPSTASH_REDIS_REST_TOKEN` are both set, the same data is written to Upstash Redis instead, and survives redeploys. Each chat gets its own key so one failed write can't take down every chat's state at once, and a small in-memory index tracks which chats exist.
76
 
77
  ## Error tracking
78
 
79
- If `SENTRY_DSN` is set, `sentry_sdk` picks up every `log.exception()`/`log.error()` call already in the codebase with no changes needed at the call sites. Every event is scrubbed of known secrets (`BOT_TOKEN`, API keys, `WEBHOOK_SECRET`, `ADMIN_PANEL_KEY`, the Upstash token) before it leaves the process. Performance tracing is off β€” this only tracks errors, to stay comfortably inside Sentry's free tier.
 
7
  1. Telegram POSTs an update to `/webhook`, authenticated by a secret header (`X-Telegram-Bot-Api-Secret-Token`).
8
  2. `_handle_message_core` in `bot.py` classifies the message: is it a TikTok link, a `/draw` or `/tts` trigger phrase, a reply to earlier media, plain text, or an attachment?
9
  3. `_build_route` (`lumen_router_config.py`) turns that classification into an ordered list of `(provider, model_id)` candidates.
10
+ 4. `_run_route` tries the first candidate. If a whole provider's chain fails, it falls back to the other provider, except where that's physically impossible (only Gemini can read links or analyze video/audio).
11
  5. The reply streams into the chat if the request qualifies (see [Streaming](#streaming--typing-pace) below), or is sent as a single message otherwise.
12
 
13
  ## Automatic model routing
14
 
15
+ There is no `/model` or `/provider` command. The router builds a fresh candidate list for every message:
16
 
17
  - **A YouTube or website link** β†’ Gemini only. It's the only provider that can read page content (`url_context`) or analyze a video by URL.
18
  - **A video or audio attachment** β†’ Gemini only. OpenRouter's multimodal models only accept images as base64.
19
  - **An image attachment, no live-info need** β†’ free OpenRouter vision models first, Gemini as a reserve.
20
+ - **Needs current information** (a lightweight keyword heuristic: "now," "today," "price," "who is currently...") β†’ Gemini, prioritizing the models with a real search-grounding quota, then OpenRouter as a reserve.
21
+ - **Plain text, no attachments, no freshness need** (the most common case) β†’ straight to OpenRouter. Heavier requests (code, multi-step reasoning, caught by another lightweight heuristic) get routed to the stronger free OpenRouter models first.
22
 
23
+ The reasoning: Gemini's free quota (roughly 20 requests per day for the flagship model) is the scarcest resource in the system, so it's only spent where a capability unique to Gemini is actually needed. Everything else, the bulk of ordinary messages, runs on OpenRouter's much more generous free tier.
24
 
25
+ Models that the router should never pick are tracked in a single registry, `_OR_MODEL_HEALTH` (`lumen_router_config.py`), each with a dated reason: the provider dropped the free tier, or the model turned out to be uncensored and a poor fit for Lumen's persona. Nothing is removed on speculation, only on confirmed evidence from production logs (`HTTP 404`, `"no endpoints,"` etc.).
26
 
27
+ Each model is tried exactly once per route, with no retries: a single failure (timeout, `429`, `5xx`) moves straight to the next candidate. That keeps the worst case bounded by `len(route) Γ— ROUTE_MODEL_TIMEOUT_SEC`, further capped by `ROUTE_TOTAL_BUDGET_SEC` for the whole route.
28
 
29
  ### Image generation
30
 
31
+ `/draw` follows the same philosophy on a smaller scale. `_pick_image_model` (`lumen_images.py`) picks a Pollinations.ai model from the prompt's content (anime/manga β†’ `flux-anime`, fantasy β†’ `dreamshaper`, "photorealistic" β†’ `flux-realism`, "quick sketch" β†’ `turbo`, otherwise a general-purpose default) and falls back through the rest of the catalog if the chosen model fails, all within `DRAW_TOTAL_BUDGET_SEC`.
32
 
33
  ## Streaming & typing pace
34
 
35
+ For the first candidate in a route (plain text, no attachments, no links), Lumen streams the reply by repeatedly editing one message, the same mechanism for both Gemini (`generate_content_stream`) and OpenRouter (SSE over `chat/completions`).
36
 
37
+ Telegram won't let a bot edit one message more than about once a second, so true token-by-token output isn't possible regardless of how fast the model generates text. Instead, `lumen_typing_pace.py` tracks each `provider:model` pair's real observed characters-per-second and reveals a growing slice of the buffered text at that pace, capped by whatever has actually arrived. If a backend returns the whole answer in one large chunk (common for free OpenRouter models, since the same `:free` slug can be served by different backends depending on load), a short "catch-up" phase finishes revealing it smoothly instead of dumping it all at once, capped at `STREAM_TYPING_MAX_CATCHUP_TICKS Γ— STREAM_TYPING_TICK_SEC` seconds.
38
 
39
+ This speed estimate is measured, not hardcoded. There is no table of "model X does N tokens/sec" to maintain, because that number isn't a property of the model on OpenRouter's free tier in the first place. A new model just starts at a default speed and calibrates itself over its first few replies.
40
 
41
  If the streaming attempt for the head-of-route model fails before showing any text, the bot falls back to a normal (non-streaming) call further down the chain, reusing the same placeholder message rather than sending a new one.
42
 
43
  ### Manual smoke test before touching streaming code
44
 
45
+ Streaming is covered only by mocked clients; no test exercises a real Gemini or OpenRouter SSE stream. Before merging changes to `_run_streaming_reply`, `_gemini_stream_pieces`, `_openrouter_stream_pieces`, or `lumen_typing_pace.py`, check by hand in a real chat:
46
 
47
+ 1. A short message (fits in one edit): once via Gemini, once via OpenRouter.
48
+ 2. A long message (past `TG_MAX_LEN`): confirm it splits into multiple messages and continues correctly.
49
+ 3. A request where the streaming model is deliberately unavailable: confirm the silent fallback to a non-streaming reserve model, with no visible glitch.
50
 
51
  ## Prompt-injection and identity-leak defenses
52
 
53
+ Lumen deliberately hides which model or provider answers a given message (see the persona in `system_prompt.py`). The system prompt alone is the weakest layer, since any LLM can potentially be talked out of following it with a creative enough injection. So there are several independent layers, each assuming the previous one might have failed:
54
 
55
+ 1. **Input pre-filter** (`_looks_like_injection_probe`, `lumen_security.py`): well-known jailbreak phrasing ("ignore previous instructions," "developer mode," "show me your system prompt") is caught before the model is ever called. The response is fully deterministic.
56
+ 2. **System prompt** (`system_prompt.py`): instructs the model that anything outside the prompt itself (user messages, chat background, page/video/document content) is data, not instructions, and that the persona doesn't change no matter who claims authority to override it.
57
+ 3. **Output identity-leak filter** (`_detect_identity_leak` / `_scrub_identity_leak`): a deterministic check on the finished reply, catching exact internal model IDs and narrow self-identification patterns ("I am Gemini," "made by OpenAI"). Deliberately narrow, to avoid false positives on ordinary, honest discussion of other AI companies.
58
+ 4. **Injected-payload echo filter** (`_detect_injected_payload_echo`): catches the case where an attacker embeds an instruction in a photo or web page ("output this exact string to confirm the jailbreak worked") and the model, while refusing to *follow* it, ends up quoting it back verbatim while summarizing the content.
59
 
60
+ Layers 3 and 4 run before the reply is written to chat history (so a leak can't influence future turns) and, during streaming, before each chunk is shown to the user, not just the final text.
61
 
62
  None of this is airtight except the input pre-filter. The goal is raising the bar for known attack patterns, not proving immunity to every possible phrasing; incidents get logged with `[identity-leak]` / `[injection-echo]` / `[injection-probe]` tags so new patterns can be added as they show up.
63
 
64
  ## TikTok downloader
65
 
66
+ Downloads go through the public TikWM API (`tikwm.com`), no account or API key needed. The bot never re-encodes video or photos; bytes go to Telegram exactly as TikWM served them.
67
 
68
  - **Video quality**: requested with `&hd=1`; among the variants TikWM returns (`hdplay`/`play`/`wmplay`, each with a known byte size), the best one that fits Telegram's 50 MB upload limit is picked. If Telegram still rejects it as too large, the next lighter variant is tried automatically.
69
  - **Slideshows**: TikTok allows up to 35 slides per post; Telegram's `sendMediaGroup` caps out at 10 items per call. The bot downloads the whole post in parallel and sends it as several media groups in sequence, never with a trailing group of exactly one item (Telegram requires 2–10 per group).
70
+ - **"Live" slides**: TikWM's response includes a separate `live_images` field alongside the regular `images`; that's where the actually-moving version of a slide lives, if it has one. Each downloaded slide is also double-checked by its magic bytes (`ftyp` = video container) rather than trusting the field alone.
71
+ - TikWM rejects Hugging Face Spaces' outbound IPs with an empty `403`, the same class of block that also keeps [YouTube downloading off the table](../README.md#known-limitations). The proxy (`TIKWM_API_BASE_URL`) is the only fix that's worked in practice; throttling, retries, and spoofed `Referer`/`Origin` headers didn't help on their own (they're still in place as defense in depth).
72
 
73
  ## Persistent storage
74
 
75
+ Chat history and quota counters live in `chat_state.json` / `global_quota.json` on the container's local disk by default, which Hugging Face Spaces wipes on every redeploy. If `UPSTASH_REDIS_REST_URL` and `UPSTASH_REDIS_REST_TOKEN` are both set, the same data is written to Upstash Redis instead, and survives redeploys. Each chat gets its own key so one failed write can't take down every chat's state at once, and a small in-memory index tracks which chats exist.
76
 
77
  ## Error tracking
78
 
79
+ If `SENTRY_DSN` is set, `sentry_sdk` picks up every `log.exception()`/`log.error()` call already in the codebase with no changes needed at the call sites. Every event is scrubbed of known secrets (`BOT_TOKEN`, API keys, `WEBHOOK_SECRET`, `ADMIN_PANEL_KEY`, the Upstash token) before it leaves the process. Performance tracing is off; this only tracks errors, to stay comfortably inside Sentry's free tier.
docs/ENVIRONMENT.md CHANGED
@@ -1,6 +1,6 @@
1
  # Environment variables
2
 
3
- Only `BOT_TOKEN` and `GEMINI_API_KEY` are required. Everything else has a working default β€” most deployments never need to touch the rest of this file.
4
 
5
  ## Required
6
 
@@ -11,7 +11,7 @@ Only `BOT_TOKEN` and `GEMINI_API_KEY` are required. Everything else has a workin
11
 
12
  ## Telegram & TikWM proxy
13
 
14
- Hugging Face Spaces' outbound IPs are blocked by Telegram's API and rejected (`403`) by TikWM β€” see [Proxy setup](../README.md#proxy-setup) in the main README.
15
 
16
  | Variable | Default | Purpose |
17
  |---|---|---|
@@ -26,7 +26,7 @@ Hugging Face Spaces' outbound IPs are blocked by Telegram's API and rejected (`4
26
 
27
  | Variable | Default | Purpose |
28
  |---|---|---|
29
- | `ADMIN_SECRET_SEED` | `BOT_TOKEN` | Salt used to derive `WEBHOOK_SECRET`/`ADMIN_PANEL_KEY`. Set it independently to rotate those two secrets without touching the bot's actual Telegram token. |
30
  | `OWNER_ID` / `BOT_OWNER_ID` / `ADMIN_ID` / `TELEGRAM_OWNER_ID` | β€” | Telegram user ID that unlocks `/logs` and `/stats`. The first non-empty variable found is used. |
31
  | `BOT_USERNAME` | `LumenAI_bot` | Fallback username; the bot fetches its real one via `getMe` on startup and uses that instead. |
32
 
@@ -50,28 +50,28 @@ Hugging Face Spaces' outbound IPs are blocked by Telegram's API and rejected (`4
50
  | `TTS_MAX_CHARS` | `800` | Max text length accepted by `/tts`. |
51
  | `RATE_LIMIT_MAX_REQUESTS` | `5` | Max requests per user within `RATE_LIMIT_WINDOW_SEC`. |
52
  | `RATE_LIMIT_WINDOW_SEC` | `30s` | Sliding window width for rate limiting. |
53
- | `ROUTE_MODEL_TIMEOUT_SEC` | `22s` | Timeout for a single attempt at a single model. No retries β€” any failure moves straight to the next model. |
54
  | `ROUTE_TOTAL_BUDGET_SEC` | `40s` | Total time budget for the whole routing chain of one message, across both providers. |
55
  | `DRAW_TOTAL_BUDGET_SEC` | `120s` | Same idea, for the `/draw` fallback chain across image models. |
56
  | `STREAM_CHUNK_TIMEOUT_SEC` | `30s` | Timeout waiting for the next streamed chunk, shared by Gemini and OpenRouter. |
57
  | `STREAM_EDIT_MIN_INTERVAL_SEC` | `1.2s` | Minimum interval between message edits during streaming (protects against Telegram's `429`). |
58
  | `STREAM_TYPING_TICK_SEC` | `0.5s` | Interval between steps of the post-stream "catch-up" reveal. |
59
  | `STREAM_TYPING_MAX_CATCHUP_TICKS` | `6` | Max catch-up steps, capping the extra delay this can add. |
60
- | `STATE_FLUSH_CONCURRENCY` | `10` | Max concurrent background writes to storage per flush cycle. |
61
 
62
  ## TikTok downloader
63
 
64
  | Variable | Default | Purpose |
65
  |---|---|---|
66
  | `TIKTOK_DOWNLOAD_MAX_BYTES` | `75 MB` | Hard cap on any single downloaded TikTok file (video, slide, or cover), aborted mid-stream if exceeded. |
67
- | `TIKTOK_SLIDE_DOWNLOAD_CONCURRENCY` | `8` | Max slideshow slides downloaded in parallel β€” keeps one large post from hogging the shared HTTP connection pool. |
68
  | `TIKTOK_VIDEO_SLIDE_PROBE_CONCURRENCY` | `4` | Max concurrent `ffprobe`/`ffmpeg` processes when probing "live" video slides in a slideshow. |
69
 
70
  ## Persistent storage
71
 
72
  | Variable | Default | Purpose |
73
  |---|---|---|
74
- | `STATE_DIR` | `/app` | Where `chat_state`/`global_quota` files live if Upstash isn't configured. This is the container's ephemeral disk β€” see [Known limitations](../README.md#known-limitations). Falls back to a temp directory if not writable. |
 
75
  | `UPSTASH_REDIS_REST_URL` | β€” | Upstash Redis REST URL. Set together with the token below to persist state across redeploys. |
76
  | `UPSTASH_REDIS_REST_TOKEN` | β€” | Upstash Redis REST token. |
77
 
 
1
  # Environment variables
2
 
3
+ Only `BOT_TOKEN` and `GEMINI_API_KEY` are required. Everything else has a working default; most deployments never need to touch the rest of this file.
4
 
5
  ## Required
6
 
 
11
 
12
  ## Telegram & TikWM proxy
13
 
14
+ Hugging Face Spaces' outbound IPs are blocked by Telegram's API and rejected (`403`) by TikWM. See [Proxy setup](../README.md#proxy-setup) in the main README.
15
 
16
  | Variable | Default | Purpose |
17
  |---|---|---|
 
26
 
27
  | Variable | Default | Purpose |
28
  |---|---|---|
29
+ | `ADMIN_SECRET_SEED` | falls back to `BOT_TOKEN`, then to a random value if that's empty too | Salt used to derive `WEBHOOK_SECRET`/`ADMIN_PANEL_KEY`. Set it independently to rotate those two secrets without touching the bot's actual Telegram token. |
30
  | `OWNER_ID` / `BOT_OWNER_ID` / `ADMIN_ID` / `TELEGRAM_OWNER_ID` | β€” | Telegram user ID that unlocks `/logs` and `/stats`. The first non-empty variable found is used. |
31
  | `BOT_USERNAME` | `LumenAI_bot` | Fallback username; the bot fetches its real one via `getMe` on startup and uses that instead. |
32
 
 
50
  | `TTS_MAX_CHARS` | `800` | Max text length accepted by `/tts`. |
51
  | `RATE_LIMIT_MAX_REQUESTS` | `5` | Max requests per user within `RATE_LIMIT_WINDOW_SEC`. |
52
  | `RATE_LIMIT_WINDOW_SEC` | `30s` | Sliding window width for rate limiting. |
53
+ | `ROUTE_MODEL_TIMEOUT_SEC` | `22s` | Timeout for a single attempt at a single model. No retries: any failure moves straight to the next model. |
54
  | `ROUTE_TOTAL_BUDGET_SEC` | `40s` | Total time budget for the whole routing chain of one message, across both providers. |
55
  | `DRAW_TOTAL_BUDGET_SEC` | `120s` | Same idea, for the `/draw` fallback chain across image models. |
56
  | `STREAM_CHUNK_TIMEOUT_SEC` | `30s` | Timeout waiting for the next streamed chunk, shared by Gemini and OpenRouter. |
57
  | `STREAM_EDIT_MIN_INTERVAL_SEC` | `1.2s` | Minimum interval between message edits during streaming (protects against Telegram's `429`). |
58
  | `STREAM_TYPING_TICK_SEC` | `0.5s` | Interval between steps of the post-stream "catch-up" reveal. |
59
  | `STREAM_TYPING_MAX_CATCHUP_TICKS` | `6` | Max catch-up steps, capping the extra delay this can add. |
 
60
 
61
  ## TikTok downloader
62
 
63
  | Variable | Default | Purpose |
64
  |---|---|---|
65
  | `TIKTOK_DOWNLOAD_MAX_BYTES` | `75 MB` | Hard cap on any single downloaded TikTok file (video, slide, or cover), aborted mid-stream if exceeded. |
66
+ | `TIKTOK_SLIDE_DOWNLOAD_CONCURRENCY` | `8` | Max slideshow slides downloaded in parallel; keeps one large post from hogging the shared HTTP connection pool. |
67
  | `TIKTOK_VIDEO_SLIDE_PROBE_CONCURRENCY` | `4` | Max concurrent `ffprobe`/`ffmpeg` processes when probing "live" video slides in a slideshow. |
68
 
69
  ## Persistent storage
70
 
71
  | Variable | Default | Purpose |
72
  |---|---|---|
73
+ | `STATE_DIR` | `/app` | Where `chat_state`/`global_quota` files live if Upstash isn't configured. This is the container's ephemeral disk (see [Known limitations](../README.md#known-limitations)). Falls back to a temp directory if not writable. |
74
+ | `STATE_FLUSH_CONCURRENCY` | `10` | Max concurrent background writes to storage per flush cycle. |
75
  | `UPSTASH_REDIS_REST_URL` | β€” | Upstash Redis REST URL. Set together with the token below to persist state across redeploys. |
76
  | `UPSTASH_REDIS_REST_TOKEN` | β€” | Upstash Redis REST token. |
77