Spaces:
Running on Zero
Download README.md from spacedout-bits/Oracle: direct link, hf CLI and curl.
- Browser
- Download file 12.5 kB
-
https://huggingface.co/spaces/spacedout-bits/Oracle/resolve/main/README.md
- Command line
-
hf download hf://spaces/spacedout-bits/Oracle/README.md
-
curl -L -o README.md https://huggingface.co/spaces/spacedout-bits/Oracle/resolve/main/README.md
A newer version of the Gradio SDK is available: 6.30.0
title: Oracle
emoji: πΈ
colorFrom: green
colorTo: indigo
sdk: gradio
sdk_version: 5.50.0
app_file: app.py
pinned: false
short_description: Expense tracking and budget coaching over Telegram
coco-finbot
Deployed as the Oracle Space on
sdk: gradio.The bot is FastAPI; Gradio is only there to satisfy the platform. Since mid-2026, Docker Spaces require PRO, and the one free hardware tier for personal accounts β ZeroGPU β refuses any SDK but Gradio. So
app.pybuilds the FastAPI app, mounts a one-page Gradio UI at/viagr.mount_gradio_app, and serves both. No GPU is used: inference goes out over HTTP to HF Inference Providers.The
Dockerfileis still current and is what you'd use to self-host.
A personal finance bot that lives in Telegram and runs as a Hugging Face Docker Space. Tell it what you spent, send it receipts and bank statements, and ask it what your money is doing.
you βΈ coffee 250
bot βΈ β
Logged βΉ250.00 β coffee
Food & Dining Β· 2026-08-15
you βΈ [photo of a restaurant bill]
bot βΈ β
Logged βΉ1,840.00 β Toit Brewpub
Food & Dining Β· 2026-08-14
you βΈ [statement.pdf]
bot βΈ Found 47 transactions in that PDF Β· βΉ68,204.00
3 already logged (skipped)
These were read by a model and not saved yet.
/confirm to save Β· /cancel to discard
Read this before you deploy
Space disk is ephemeral. Hugging Face wipes it on every restart, sleep and
rebuild. If you do not set HF_DATASET_REPO, your entire ledger disappears the
first time the Space sleeps β silently, because nothing errors. The bot mirrors
SQLite to a private dataset repo on a timer and restores it on boot. /status
tells you which mode you are in.
Make the Space private. This is your spending history. A public Space exposes
the container, its logs and its /health endpoint to anyone.
Decide who can use it. ALLOWED_TELEGRAM_USER_IDS starts empty, which
authorises nobody β a forgotten allowlist stays closed rather than silently
opening to the world.
Set OPEN_ACCESS=true to let anyone message the bot. Ledgers are keyed by
Telegram user id, so open access makes the bot multi-tenant: each person
gets their own private ledger and nobody can see yours. What open access does
expose is your wallet β every stranger's receipt and statement is read using
your HF_TOKEN. UPLOADS_PER_USER_PER_DAY (default 20) caps that; typed
expenses are parsed locally and never metered. Note also that running an open
bot makes you the custodian of other people's financial data in your dataset
repo.
Deploy
1. Create the bot. Message @BotFather β /newbot β
keep the token.
2. Push the Space and dataset. deploy.sh creates both (private, Docker SDK)
and uploads the source. It takes the owner from your logged-in account:
./.venv/bin/hf auth login && ./deploy.sh
./deploy.sh --dry-run # show the target repos, create nothing
./deploy.sh Uniphy # deploy under an org instead of your account
Use this rather than
git push. This directory lives inside the larger UniPhy repo, so agit pushfrom here would send that entire tree β unrelated repos and any secrets they hold β to a public-by-default host.deploy.shuploads only this folder, minus.venv,.env, caches and any local database.
3. Set secrets under Space Settings β Secrets (use the UI, not the CLI, so they stay out of your shell history):
| Secret | Required | What it does |
|---|---|---|
TELEGRAM_BOT_TOKEN |
β | From BotFather |
TELEGRAM_WEBHOOK_SECRET |
β | Any random string; rejects forged webhook calls |
HF_DATASET_REPO |
β | <you>/finbot-data β without this your data is lost on restart |
HF_TOKEN |
β | Write-scoped token. Needed to mirror the ledger; also used for remote inference if you pick that backend |
ALLOWED_TELEGRAM_USERNAMES |
β | Your Telegram @handle (comma-separated for more). Not needed if OPEN_ACCESS=true |
ALLOWED_TELEGRAM_USER_IDS |
Numeric IDs β the only option for users with no @handle set | |
PUBLIC_BASE_URL |
Origin for dashboard links. Derived from SPACE_HOST on a Space |
|
DASHBOARD_LINK_MINUTES |
Sign-in link lifetime. Default 30 |
|
OPEN_ACCESS |
true lets anyone use the bot, each with their own ledger. Default false |
|
UPLOADS_PER_USER_PER_DAY |
Caps model-backed uploads per user. Default 20, 0 disables |
|
LLM_BACKEND |
auto (default), local, remote, off. See below |
|
LOCAL_LLM_MODEL / LOCAL_VISION_MODEL |
Override the on-GPU models | |
DEFAULT_CURRENCY |
ISO code for bare amounts. Default INR |
|
TIMEZONE |
Decides what "today" means. Default Asia/Kolkata |
A read-only token is not enough: CommitScheduler writes the ledger with it,
and a read token means your data silently never persists.
Generate a webhook secret with:
python3 -c "import secrets; print(secrets.token_urlsafe(32))"
4. Restart the Space so it picks up the secrets. On boot it reads
SPACE_HOST and registers its own Telegram webhook β nothing else to wire up.
5. Message the bot. Don't know your Telegram ID? Send it anything: it
replies with the ID to paste into ALLOWED_TELEGRAM_USER_IDS. Then send
/status and confirm storage reads mirroring, not local only.
Using it
Anything with a number in it is logged. No command needed.
coffee 250
βΉ1,240 groceries at bigbasket
45.50 lunch yesterday
320 uber #transport β #tag overrides the category guess
β¬40 dinner 12 aug
Send a photo of a receipt and a vision model reads the total. Send a
PDF/CSV/XLSX statement and it extracts the transactions, drops any already
in your ledger, and shows you the total before saving β nothing model-extracted
is written until you /confirm.
| Command | |
|---|---|
/today /week /month |
Spending summaries with a category breakdown |
/report <period> |
today, week, month, last_month, 30d, year, all |
/list [n] |
Recent transactions |
/recurring |
Subscriptions and repeat charges, with the annual cost |
/advice [question] |
Observations grounded in your own numbers |
/budget <category> <amount> |
Set a monthly cap |
/budgets |
How you're tracking |
/categories |
Category names for tags and budgets |
/undo |
Remove the last entry, or the last import wholesale |
/export |
Everything as CSV |
/status |
Config and storage health |
How it works
Telegram ββwebhookβββΆ FastAPI (:7860)
β
ββ registry βββΆ features: expenses β budgets β coaching β core
β β
β ββ parser deterministic, free, first
β ββ extract CSV Β· XLSX Β· PDF
β ββ llm βββΆ router.huggingface.co/v1
β
ββ store: SQLite ββCommitSchedulerβββΆ private HF dataset
Deterministic first, model second. The regex parser handles coffee 250
instantly and identically every time. Inference is spent only where judgement
is genuinely needed: reading a receipt, or working out which column of an
unfamiliar bank export holds the amount. If inference is unavailable, text
logging and every summary still work.
Two interchangeable inference backends. local runs a small instruct model
(Qwen2.5-3B, plus a 3B VL model for receipts) on the Space's own borrowed A10G;
remote calls HF Inference Providers. Same interface, chosen by LLM_BACKEND,
and auto picks local whenever a CUDA device is present. Local is the default
on a Space because it costs nothing per call β which is what makes
OPEN_ACCESS affordable, since otherwise every stranger's receipt would spend
your inference credits.
Money is never a float. Amounts are integer minor units (paise, cents, fils) plus an ISO 4217 code. Currencies are never summed together β a multi-currency month reports one total per currency rather than a fictional combined figure.
Statements are staged, not trusted. A model that misreads a column would
otherwise quietly poison months of history, so imports become a pending batch
you confirm, tagged with a batch_id so /undo can reverse the whole thing.
Inference is provider-agnostic. LLM_BASE_URL is an OpenAI-compatible
endpoint. It defaults to HF's router β one token, many providers β but points
just as happily at Ollama, llama.cpp or OpenAI.
Adding a feature
The registry is the extension point. A feature declares commands and optionally
claims plain messages or uploaded files; the first one to return a Reply wins.
Nothing else in the codebase changes.
from finbot.features.base import BotContext, CommandSpec, Reply, SimpleFeature
class RemindersFeature(SimpleFeature):
name = "reminders"
description = "β° Reminders"
def commands(self):
return (CommandSpec("remind", "Set a reminder", "<when> <what>"),)
def cmd_remind(self, ctx: BotContext) -> Reply:
return Reply(f"Noted: {ctx.args}")
Add it to default_features() in finbot/features/__init__.py. /help and
Telegram's command menu build themselves from what is registered, and duplicate
commands raise at startup rather than shadowing silently.
Development
python3.11 -m venv .venv && ./.venv/bin/pip install -r requirements.txt pytest
cp .env.example .env # fill in what you need
./.venv/bin/python -m pytest tests/ -q
./.venv/bin/uvicorn app:app --reload --port 7860
To test against real Telegram locally, expose port 7860 with a tunnel and register the webhook:
curl -X POST localhost:7860/admin/setup -H "X-Admin-Secret: $TELEGRAM_WEBHOOK_SECRET" -H 'Content-Type: application/json' -d '{"url":"https://<your-tunnel>/telegram/webhook"}'
268 tests cover the pure logic β money arithmetic, parsing, categorisation, analytics, the registry contract β plus an end-to-end pass over the webhook with a stubbed Telegram transport. None of them touch the network.
Known limits
- Credits are not stored. Statement imports keep debits only; incoming money is counted and reported but not written, because mixing it into the same table would corrupt every "total spent" figure. Income tracking is a real feature, not a small one.
- Scanned PDFs have no text layer. The bot detects this and tells you to send it as a photo instead.
- Legacy
.xlsis not supported β re-save as.xlsxor CSV. - 20 MB is Telegram's own cap on bot file downloads.
- Pending imports live in memory, so a Space restart between the preview and
/confirmdrops them. Re-send the file. - Recurring detection is lexical. It unifies
NETFLIX.COM 447192withNetflix, but notAMZN MktpwithAmazon. - Advice is budgeting only. It comments on your logged spending and is explicitly prevented from recommending investments or giving tax advice.
- Free Spaces sleep after 48h idle, and the first request after that pays a cold start plus a model load. Telegram retries, so the message lands β it is just slow.
- A 3B model is not a frontier model. It is good enough for "which column is
the amount" and short coaching; expect worse statement extraction than the
remote backend. Switch with
LLM_BACKEND=remoteif you would rather spend credits for quality. - ZeroGPU has a daily GPU-second quota on free accounts. Heavy statement days can exhaust it; typed expenses keep working regardless.
ssr_mode=Falseis load-bearing. Spaces setsGRADIO_SSR_MODE, which switches the frontend to a Node-rendered SvelteKit build served from/_app/immutable/*. Because this app callsdemo.launch()itself rather than going through Gradio's standard Space entrypoint, that SSR server never starts, so every CSS/JS asset 404s. The symptom is nasty: the page still returns HTTP 200 with a server-rendered shell, so it looks "up" while no JavaScript ever runs β meaningdemo.loadnever fires and the dashboard silently ignores its sign-in token. Do not remove it.