--- title: AIM Autonomous Ingestion Agent emoji: ๐Ÿค– colorFrom: blue colorTo: green sdk: docker app_port: 7860 pinned: true short_description: Autonomous DOI-deduped literature harvesting for AIM DB --- # AIM Autonomous Ingestion Agent An autonomous agent that continuously grows the **AIM Composites Materials Database** โ€” the same Postgres the [MaterialsDatabase Space](https://huggingface.co/spaces/aim4composites/MaterialsDatabase) serves โ€” by harvesting open-access literature **with DOI provenance**: ``` plan โ”€โ–ถ discover โ”€โ–ถ dedupe โ”€โ–ถ download โ”€โ–ถ ingest โ”€โ–ถ report (OpenAlex, (by DOI, (validated (Gemini โ†’ (metrics, arXiv, URL, and %PDF, grounding, run log, Unpaywall) sha256) size caps) unit-norm, query status) rotation) ``` * **discover** โ€” query rotation over an editable frontier (8 builtin intents + a ~450-query materials ร— reinforcement ร— property/process grid seeded at boot); per query: OpenAlex full-text search (OA gold/hybrid/green), arXiv (ANDed terms) and Semantic Scholar **bulk** search (up to 1,000 OA-PDF papers per call); per cycle: one page of the **OpenAlex topic walk** (1-credit filter calls over the fibre-reinforced-composites topics, cursor persisted) and, when due, a **manufacturer datasheet** seed crawl. * **dedupe** โ€” a durable `agent_doi_seen` registry: a paper already harvested (by DOI, then URL, then content sha256) is never downloaded twice. * **download** โ€” polite (rate-limited, robots-aware), magic-byte + PyMuPDF validated, size-bounded; URL chain per work: repository copies first, then the publisher PDF, then Unpaywall by DOI, then the landing page's `citation_pdf_url`. * **ingest** โ€” `extraction.py` (grounded, unit-aware Gemini extraction, prompt v2.1: every value also carries the processing route of its specimen) via `batch_ingest.process_pdf` with the `pg_mirror` Postgres backend: every value is grounded in the PDF text, unit-normalized (`value_si`, `unit_canonical`), classified, deduped at measurement grain, and inserted with a `status` โ€” flagged rows are quarantined for review, never dropped. * **report** โ€” per-cycle metrics in `agent_runs` (dashboards + paper ยง8), event stream in `agent_events`, optional Gemini query expansion when yields dry up. The graph runs on LangGraph; the agent loop runs in-process on an APScheduler interval, so the Space works autonomously while it is up โ€” no viewers needed. ## Pages | Page | What it shows | |---|---| | **Dashboard** | agent status pills, KPI tiles, rows/PDFs per cycle charts, latest activity feed, latest DOIs, recent cycles | | **Run Control** | autonomy on/off, interval, per-cycle budget, source toggles, figure-mining toggles, query frontier editor, manual "run one cycle" | | **Recently Added** | newest ingested documents with DOI links + verified-row counts | | **Review Queue** | all `status != 'ok'` rows with origin filter, the figure crops behind figure-estimate rows, CSV export, origin-scoped human promote-to-ok | | **Runs & Logs** | cycle history (status chips, durations, live-refreshing while a cycle runs) with per-PDF results, event stream grouped per run | | **Database** | material-first browser (one expander per material with the process type and conditions of each property; ๐Ÿ“ˆ marks figure-derived estimates) + raw-table browser (incl. the `figures` provenance table) with filters, pagination, CSV export | | **Knowledge Graph** | every row of the three material tables as one graph, the review queue included and marked as such: whole-database map, matrix and fiber pairs, single-node view with the measured values; each value opens to its page, quoted sentence, flag reason, figure picture, processing route and date added; rebuilt when the tables change and swapped into the open page; JSON and Neo4j exports | UI: single light theme + design system in `agent/ui_theme.py` (tokens, CSS, altair theme, KPI/pill/chip/empty-state components); pages compose those helpers and contain no styling of their own. Since 28 Sep the Dashboard carries a "Candidates queued" tile, Run Control a "Candidate queue" card (toggle, pass interval, low-water mark, per-pass budgets, queue status by source, "Run a discovery pass now") and Runs & Logs a `Kind` column (`cycle` / `discovery`) plus a `Queued` counter for passes. ## Deploy (โ‰ˆ5 minutes) 1. **Create the Space**: huggingface.co โ†’ New Space โ†’ owner `aim4composites`, name e.g. `AutonomousAgent`, SDK **Docker**, visibility your call. 2. **Push this folder** to it: ```bash git clone https://huggingface.co/spaces/aim4composites/AutonomousAgent cp -r /* /.streamlit /.gitignore AutonomousAgent/ cd AutonomousAgent && git add -A && git commit -m "Autonomous ingestion agent" && git push ``` (or upload the files through the web UI โ€” Files โ†’ Add file.) 3. **Set Space secrets** (Settings โ†’ Variables and secrets): | Secret | Value | |---|---| | `DB_HOST` / `DB_PORT` / `DB_NAME` / `DB_USER` / `DB_PASSWORD` | same values the MaterialsDatabase Space uses | | `GEMINI_API_KEY` | a **fresh** Gemini key (rotate the leaked ones!) | | `AGENT_ADMIN_PASSWORD` | password that unlocks Run Control / promote | | `CONTACT_EMAIL` *(variable, optional)* | polite-crawling contact for OpenAlex/Unpaywall | | `OPENALEX_API_KEY` | **required since 2026** โ€” OpenAlex bills per call: anonymous = $0.10/day *per IP* (shared with every other Space on the egress IP), a free account key = $1/day โ‰ˆ 1,000 searches or 10,000 filter calls. Get it at openalex.org โ†’ Settings โ†’ API | | `S2_API_KEY` *(optional)* | dedicated Semantic Scholar quota (the bulk lane works without a key on the shared bucket; a free key is a form away) | 4. **One-time schema check**: the shared Postgres must already have the hardening columns + `sources` table (`python pg_migrate.py` dry-run, then `--apply`, from any machine with the DB env set). If the July pipeline runs already did this, there is nothing to do โ€” the agent refuses to write until the schema is migrated, so it fails safe either way. 5. Open the Space โ†’ **Run Control** โ†’ unlock admin โ†’ *Run one cycle now* to smoke-test โ†’ toggle **Autonomous crawling** ON. ## Process type and conditions per property (prompt 2.1, Oct 2026) A composite's properties depend on how the part was made, so every property row carries the processing route of the specimen it was measured on: | Column | Content | |---|---| | `process_type` | route, normalized to `extraction.PROCESS_TYPE_ENUM` (injection molding, compression molding / hot press, autoclave, out-of-autoclave, automated fiber placement / tape laying, thermoforming / stamp forming, filament winding, pultrusion, extrusion / compounding, additive manufacturing, liquid molding, welding / joining, fiber spinning, film / solution casting, electrospinning, heat treatment / annealing, other) | | `process_name` | the route as the document names it | | `process_conditions` | processing parameters as printed, `; `-joined (temperature, pressure, hold time, cooling rate, nozzle/bed temperature, print speed, ...) | | `process_quote`, `process_page` | the verbatim sentence that names the route, and its page | | `process_status` | `grounded` / `grounded_off_page` (sentence found in the PDF text and every number of the conditions found on digit boundaries), `conditions_unverified` (sentence found, a number is not), `ungrounded` (sentence not found), `unchecked` (no text layer) | * The model returns the routes once per document (`processes`, ids `P1`, `P2`, ...) and each property points at one through `process_id`, so the output grows by one short field per property, not by a repeated paragraph. The same route at two settings is two routes. * All six columns are NULL when the document does not say how the material was made (most vendor datasheets, as-received constituents, values quoted from other literature). The prompt forbids guessing. * The route is context, not a value: its check never changes a row's `status`. An unverified route stays attached and marked (โš  in the Database page). * `test_condition` is still the *test* condition (23 ยฐC, 50% RH); the `Processing` section still holds process settings a document reports as values of their own. * Schema: the columns are appended to `migrate.EXTRA_COLUMNS`, so the boot migration adds them to the three material tables (`ADD COLUMN IF NOT EXISTS`); nothing is run by hand. The dedup grain is unchanged. * Rows stored before this change (prompt 2.0) keep an empty route. The source PDFs are deleted after ingest, so filling them needs a re-download pass; that is not part of this change. * Pages: Database (Process and Process conditions columns per material, a Process filter, a "With processing route" tile, the columns in the raw browser and its CSV), Review Queue (three process columns, also in the CSV), Dashboard ("Processing routes" card: rows per route, share grounded). * Tests: `test_process_e2e.py` (parse, grounding, rows, boot migration of an old database, a stubbed cycle into Postgres, the three pages). ## Knowledge graph (live, Oct 2026) The **Knowledge Graph** page shows the whole database as one graph and keeps it current without anyone running anything. - **What is in it.** `agent/kg.py` reads every row of `Polymers`, `Fibers` and `Composites_materials` through a server-side cursor (no row limit, the `image` bytes left out) and every row of `figures`, and builds the graph: materials, properties, property categories, polymer and fiber families, manufacturers, source documents (DOI, year and lane from `agent_doi_seen`), figures and process types. Every measured value keeps its status, page, SI value, route, the day it was added and the figure it was read off or cites. - **The review queue is in it.** Rows with `status != 'ok'` are not left out: each is marked "review queue" with its status and `flag_reason`, and the switch above the graph shows all values, the verified ones, or only the queue (counts, map and lists follow the switch). Promoting a row in the Review Queue page changes the fingerprint, so the graph follows. - **Where a value was found.** Opening a value (the arrow at the start of its row) shows the page, the quoted sentence (`source_quote`), the model's comment, the figure with its picture, the route's own name, quote, page and grounding status, and the date, model and prompt version of the record. These fields live in separate detail files of 2,000 values each (`static/kg/d//.json.gz`), fetched when a value is opened, so the main file stays small as the database grows. - **Figures.** Every harvested figure is a node, linked to its document and to the values read off it or citing it. The pictures of figures with values are copied from `figures.image_bytes` to `static/kg/fig/.png` in the background (20 per query, at most 4 minutes per run of the timer job, only what is not on disk yet; after a restart this takes a few minutes). `AGENT_KG_FIGURES` = `linked` (default) | `all` (every stored crop) | `off`. - **When it rebuilds.** A fingerprint (per table: rows, verified rows, newest `extracted_at`, rows with a process type; and the number of figures) is compared with the one the current graph was built from. It is read right after every cycle, every `AGENT_KG_REFRESH_MINUTES` (default 10; 0 switches the timer off) by the scheduler job `kg-refresh`, and when someone opens the page (at most every 20 s). The graph is rebuilt only when the fingerprint changed, so an idle database costs four small aggregate queries per check. - **How the open page follows.** The graph is written to `static/kg/` (`kg-data.json`, `kg-data.json.gz`, `kg-meta.json`) and served by Streamlit at `/app/static/kg/...` (`enableStaticServing` in `.streamlit/config.toml`; the Dockerfile makes the folder writable; `AGENT_KG_STATIC_DIR` moves it). The explorer (`agent/kg_explorer.html`, canvas + d3 from cdnjs) asks for `kg-meta.json` every `AGENT_KG_POLL_SECONDS` (default 60) and swaps the new graph in, keeping the open node and the zoom. Each build has its own detail folder (the one before is kept for pages still showing it). If the folder cannot be written, the graph and the value details are embedded in the page instead, a reload picks up changes, and figure pictures are not shown. - **What is inferred.** Polymer and fiber families are assigned by pattern rules, and property names are merged for case, spelling and a short synonym list. Everything else is the database as stored. The map is computed from names and counts, not simulated: the same database always gives the same picture, and it shifts only where families grow. - **Exports.** The published JSON can be read from outside the Space (`https:///app/static/kg/kg-data.json`; `detail.dir` in it names the folder of the detail files). The page also offers a Neo4j zip (nodes.csv, relationships.csv, load script; one Measurement node per value with its quote, flag reason, day added and review-queue mark; Figure nodes). `python -m agent.kg --out DIR` writes the same files, the figure pictures and a stand-alone explorer.html from the configured database. - **Cost.** On the 1 Oct snapshot (34,768 rows, 1,093 figures in the test copy): build about 3 s, 5.6 MB of JSON, 0.9 MB compressed. At ten times the rows: build about 20 s, 46 MB / 7 MB, about 0.7 GB of memory at the peak. No model calls. ## Figure & graph mining (in-cycle) Every ingested PDF also passes through `figures.py` (the same module the local pipeline uses): plots and table images are harvested from the PDF's layout (embedded rasters + vector clusters, junk filters, caption pairing), classified with **one** Gemini vision call per PDF, and mineable kinds (`property_plot`, `table_image`) are read for axis ranges and salient values. **A number read off a graph is an estimate, not a grounded fact.** Figure rows are inserted as `origin='figure'`, `status='figure_estimate'` โ€” they never publish (`status='ok'`) at insert, whatever the caller does (`batch_ingest._row_values()` downgrades them), and only a human promote in the Review Queue can bless one. The crop PNG itself is stored in the `figures` table (`image_bytes`, capped at ~1.5 MB) so the Review Queue can show the reviewer the exact plot behind every estimate โ€” the Space's disk is ephemeral, the database is not. Knobs (Run Control, or env for first-boot defaults): `AGENT_FIGURES` (on by default), `AGENT_FIGURE_MINING` (off = harvest+classify only), `AGENT_MAX_FIGURES` (per-PDF cap, default 12, hard cap 20). Per-run counters (`figures_found`, `figures_mined`, `figure_rows`, `vision_calls`) land in `agent_runs` for the Runs & Logs table and the paper's cycle accounting. ## Operations notes * **Budgets**: per-cycle caps (default 5 PDFs) bound Gemini spend; a hard ceiling of 25 is compiled in. Cycle lock is durable and self-expiring, so overlapping schedules or container restarts can't double-run. * **Keeps running on its own**: the container entry point (`run_space.py`) boots the agent (schema check, stale-run sweep, scheduler) at process start, so cycles run from the moment the container is up, with or without a visitor. A keep-alive thread GETs the Space's own URL (`https://$SPACE_HOST/_stcore/health`) every `AGENT_KEEPALIVE_MINUTES` (default 10; 0 disables; `AGENT_KEEPALIVE_URL` overrides) so a free Space never reaches its 48-h idle sleep. Keep an external hourly pinger as a backup: if the container ever does sleep, any request wakes it and the agent boots itself. * **Watchdog**: every cycle runs under a wall-clock budget (`AGENT_CYCLE_TIMEOUT_MINUTES`, default 60). A cycle still running past it is marked `timeout` in Runs & Logs, its lock is released and the next scheduled cycle proceeds. Without this, one cycle stuck in a network wait blocked the scheduler for good (the single-instance job never returned), which is what the long runs of `stale` cycles in August/September were. The crawler's `Retry-After` sleeps are capped at 300 s (the cap had been lost in the Aug 25 crawler sync), so that wait should no longer happen. * **Stall detector (22 Sep)**: besides the hard budget, a cycle whose last agent event is older than `AGENT_CYCLE_STALL_MINUTES` (default 15) is abandoned with the stall written into its verdict; a slow but reporting cycle is bound only by the hard budget. `heartbeat_at` on `agent_runs` is the liveness signal (every event touches it). * **Budgeted sources (22 Sep)**: a 429 whose `Retry-After` exceeds the 300 s cap, or that reports a zero remaining budget (OpenAlex's per-request billing), closes that host for its window with no retry and no sleep; the lane is skipped with one warn event per cycle and the reason lands in `agent_runs.sources`. Set the Space secret `OPENALEX_API_KEY` (a free account key is ten times the shared anonymous allowance). `AGENT_USE_NTRS` (default on) adds the NASA Technical Reports Server lane: free, no key, direct PDFs. * **Cost column (22 Sep)**: Gemini `usageMetadata` is summed per cycle into `tokens_in` / `tokens_out` (text + vision calls); Runs & Logs derives dollars at read time from `AGENT_PRICE_IN_PER_M` / `AGENT_PRICE_OUT_PER_M` (defaults 0.30 / 2.50 USD per million, Gemini 2.5 Flash list price). * **Figure citation linking (22 Sep)**: `AGENT_LINK_FIGURES` (default on, Run Control toggle) links each text row to the figure its evidence cites and embeds the PNG in the row's `image` column (no S3 needed). Adds three columns (`figure_ref`, `figure_link_score`, `figure_link_signals`) that the boot migration creates; `rows_linked` is counted per cycle. * **Frontier rotation on dry cycles (28 Sep)**: every query a cycle searched is marked used in `node_report`, whether or not the cycle downloaded anything. Until then the bookkeeping sat behind `node_ingest`'s early return, so a dry cycle left `last_used_at` untouched and `next_queries` (LRU, never-used first) handed the same 20 queries back every half hour: from 22 to 28 Sep every cycle rediscovered the identical 109 exhausted candidates (42 DOIs already registered, 66 URLs already final) while the Gemini-expanded queries piled up unreached. Symptom to watch for in Runs & Logs: byte-identical Candidates / Skipped / URL seen columns across consecutive cycles. The harvest loop also no longer fetches (and marks) a frontier slice past its last round. * **Downloads that never got ingested (28 Sep)**: a cycle that finds the material schema unmigrated ends as `error` (not `crashed`) and keeps its PDFs and their `downloaded` registry rows; the next cycle adopts and ingests whatever is still on disk (`adopted N PDF(s)` event). At boot, `downloaded` rows older than 3 h whose file is gone are deleted and their URLs re-opened in the crawler seen-set (`boot repair: re-queued ...`), so discovery finds and fetches those papers again โ€” this is what frees the five PDFs runs #135/#139 lost on 15โ€“16 Sep. The Inspect-run panel now shows why a run ended (crash traceback, watchdog note or cycle error). * **More papers (28 Sep)**: discovery lanes on by default โ€” OpenAlex phrase search *and* topic walk (`AGENT_OPENALEX_TOPICS`, 1-credit filter calls with persisted cursors, round-robin), Semantic Scholar bulk (โ‰ค1,000 OA-PDF papers per call; `S2_API_KEY` optional), arXiv with ANDed terms, NASA NTRS, and a datasheet round over the manufacturer seed pages (each seed โ‰ค once a day); the query frontier is seeded once with the 455-intent grid (`agent/query_grid.py`, origin `grid`). Downloads try repository copies before publisher URLs and fall back to the landing page's `citation_pdf_url`; records without an abstract get relevance-gate relief. Note that publisher hosts (MDPI, Elsevier, Wiley, โ€ฆ) 403 any cloud egress address, so the yield comes from repositories, NTRS, arXiv and datasheets. * **Parallel ingestion (28 Sep)**: a cycle extracts its PDFs on a thread pool (`ingest_workers` in Run Control, default 3, hard max `AGENT_HARD_MAX_INGEST_WORKERS` = 6), one Postgres connection per worker; counters, events and the registry are written by the cycle thread as each PDF completes. Run Control also exposes the cycle's target (`Target new PDFs per cycle`, default 10: below it the cycle takes another frontier slice, up to `Max rounds per cycle`) โ€” together with the 25-PDF cap these are the throughput knobs. A cycle's OpenAlex cost is 10 credits per query per round plus 1 per topic page, so 4 queries ร— 5 rounds ร— 48 cycles is โ‰ˆ 9,700 credits/day: right at the free key's allowance; prepaid credits or a longer interval buy headroom. * **Candidate queue (28 Sep)**: discovery and download are decoupled. A *discovery pass* (`agent/discovery.py`, an `agent_runs` row of kind `discovery`) searches the next `discovery_queries` frontier queries on every lane โ€” Semantic Scholar bulk, arXiv, NTRS, UDSpace (UD's DSpace theses, direct bitstreams), CORE (needs the free `CORE_API_KEY` secret; inactive without it) and, for the first `discovery_openalex_queries`, the 10-credit OpenAlex phrase search โ€” walks `discovery_topic_pages` pages per OpenAlex topic (1 credit each) and crawls every datasheet seed that is due, then queues every relevant candidate in `agent_candidates` (a DOI already in the registry or a URL the crawler has given up on lands as `seen`). It runs every `discovery_interval_hours` (24), right after any cycle that finds the queue below `queue_low_water` (150) or drains it to that, at boot when the queue is empty, and on demand (Run Control โ†’ "Run a discovery pass now"). Cycles pop the queue best-score-first (repository-hosted copies rank ahead of publisher URLs) and download with no search calls at all; they search inline only when the queue is empty. Every popped row is settled: `ingested`, `seen` (registry / final URL), back to `queued` after a transient failure (never twice in one cycle; `failed` after 3 attempts). Turn it off with the Run Control toggle to get the pre-28 Sep inline behaviour. OpenAlex spend drops from up to ~9,700 credits/day to a few hundred per pass. * **Secrets never reach the database (28 Sep)**: `agdb.redact()` strips credential query parameters (`?key=`, `&api_key=`, `token=`) and bearer tokens from every event, run report, registry status and candidate error before it is stored, and `extraction` raises Gemini errors that name the endpoint and the API's own detail instead of the keyed URL (requests' `HTTPError` / urllib3's connection errors quote it). At boot, `agdb.redact_stored_secrets` scrubs rows written before this existed and logs `boot repair: redacted credential parameters โ€ฆ` โ€” if that event ever appears, rotate the key it refers to: it was readable on Runs & Logs. * **Gemini budget refusals (28 Sep)**: a 429 whose body says `RESOURCE_EXHAUSTED` / "monthly spending cap" is not retried (it cost 14 s per PDF and cleared nothing). The PDF stays on disk with its registry row `downloaded`, the cycle's remaining PDFs are not attempted, one warn event says `Gemini refused for budget reasons (โ€ฆ); N PDF(s) kept on disk for the next cycle`, and the run ends `error` only if nothing was ingested at all. The next cycle adopts the kept PDFs (or the boot repair re-queues them after a restart). Raise the cap at ai.studio/spend. While the refusal lasts the agent is on a **budget hold** (`agent_state gemini_budget_hold`): cycles download nothing and probe with one kept PDF (or one fresh download if the disk was wiped); the first successful extraction lifts the hold and downloads resume next cycle. Adoption of kept PDFs is capped per cycle at the download cap, so a backlog is worked off over several cycles. * **Cost (28 Sep)**: Gemini's thinking tokens are billed as output at 6ร— the input price, and with no `thinkingConfig` the model thinks with an unbounded dynamic budget: a 25-paper cycle cost ~$7.50, $6.69 of it output (~30k tokens per PDF for a ~4k-token extraction). `gemini_thinking` (Run Control โ†’ Gemini cost; env `GEMINI_THINKING`; default `low`) is applied to every Gemini call โ€” `dynamic` | `low` | `minimal` | `high` | `off` | an integer budget; a model that rejects the option gets the call repeated without it and the option is dropped for the process. Figure *mining* (one thinking vision call per mineable figure, into quarantined `figure_estimate` rows) is off by default; harvest + classify (one call per PDF) and figure linking stay on. Both defaults are applied once to a stored config at boot (`cost defaults applied: โ€ฆ` event) and can be changed back in Run Control. Expected: roughly a third of the previous bill per paper โ€” watch the Cost column, and rows per paper / flagged share for recall. * **Cost (30 Sep) โ€” real prices, models, near-duplicates**: the cost column used to price every token at Gemini 2.5 Flash rates ($0.30 / $2.50 per M) while the agent runs `gemini-3.5-flash` ($1.50 / $9.00): it showed about a quarter of the bill (runs 544โ€“593: $12.98 shown, ~$48.50 real, ~$0.41 per PDF on dynamic thinking with mining on). Runs now store `cost_usd`, priced per model (`agent/config.py` `MODEL_PRICES`), text and vision calls each at their own model; older runs are re-priced at gemini-3.5-flash in Runs & Logs, which also shows thinking tokens, the model, $/PDF and a spend line. Run Control โ†’ Gemini cost adds the **text model** (every row is stamped with it: change it only after `cost_ab.py` on the gold set), a **vision model** for the figure calls (default: same), and the **near-duplicate gate** (`agent/textdedupe.py`, default 0.9): a PDF whose text is that similar to an ingested one โ€” regional datasheet variants, a preprint and its published version โ€” is skipped before any Gemini call (`duplicate_text` in the registry and the Near-dup column). * **Ephemeral disk**: PDFs live under `/tmp` only until ingested; provenance (DOI, sha256, title, URL, query) is durable in `agent_doi_seen`. * **Safety**: agent bookkeeping is in namespaced `agent_*` tables; controls are password-gated; read-only without the password. At boot the Space additively migrates the **connected** database (hardening columns incl. `origin`/`figure_id`, the origin-aware dedup index, the `figures` table) โ€” `aim_agent` is the agent's own DB, so this is on by default. Before pointing `DB_NAME` at the team's shared database, either agree on the additive migration or set `AUTO_MIGRATE_MATERIAL_SCHEMA=0` (the cycle then refuses to write until `pg_migrate.py --apply` is run by hand). ## Repo map Vendored pipeline (single source of truth, synced from the project repo โ€” the only Space-side deltas are the PG keepalive/statement-timeout hardening and the Retry-After cap): `extraction.py`, `batch_ingest.py`, `figures.py`, `pdf_crawler.py`, `pg_mirror.py`, `pg_migrate.py`, `migrate.py`, `generate_queries.py`. `FIGURES.md` documents the figure stage. Agent layer (new): `agent/` (orchestrator, scheduler with watchdog, autostart, keepalive, DB, config, `query_grid`), `run_space.py` (container entry point), `app.py` + `page_files/` (Streamlit UI), `agent/kg.py` + `agent/kg_explorer.html` (knowledge graph builder and explorer; `test_kg.py`). Tests: `test_agent_state.py`, `test_harvest_loop.py`, `test_sources.py` (no DB, no network), `test_e2e.py`, `test_figures_e2e.py`, `test_keep_running.py`, `test_counters_e2e.py`, `test_frontier_rotation_e2e.py`, `test_more_papers_e2e.py`, `test_parallel_ingest_e2e.py`, `test_candidate_queue_e2e.py` (scratch Postgres via `DB_*`; the network sources are switched off by env), `test_gemini_errors.py` (no DB, no network).