AutonomousAgent / README.md
Claude
Claude Opus 5.5
Knowledge Graph: figures with pictures, the review queue marked, where each value was found
3d34953 unverified
|
Raw History Blame Contribute Delete
28.2 kB
metadata
title: AIM Autonomous Ingestion Agent
emoji: πŸ€–
colorFrom: blue
colorTo: green
sdk: docker
app_port: 7860
pinned: true
short_description: Autonomous DOI-deduped literature harvesting for AIM DB

AIM Autonomous Ingestion Agent

An autonomous agent that continuously grows the AIM Composites Materials Database β€” the same Postgres the MaterialsDatabase Space serves β€” by harvesting open-access literature with DOI provenance:

        plan ─▢ discover ─▢ dedupe ─▢ download ─▢ ingest ─▢ report
                (OpenAlex,   (by DOI,   (validated  (Gemini β†’   (metrics,
                 arXiv,       URL, and    %PDF,      grounding,   run log,
                 Unpaywall)   sha256)     size caps) unit-norm,   query
                                                     status)      rotation)
  • discover β€” query rotation over an editable frontier (8 builtin intents + a ~450-query materials Γ— reinforcement Γ— property/process grid seeded at boot); per query: OpenAlex full-text search (OA gold/hybrid/green), arXiv (ANDed terms) and Semantic Scholar bulk search (up to 1,000 OA-PDF papers per call); per cycle: one page of the OpenAlex topic walk (1-credit filter calls over the fibre-reinforced-composites topics, cursor persisted) and, when due, a manufacturer datasheet seed crawl.
  • dedupe β€” a durable agent_doi_seen registry: a paper already harvested (by DOI, then URL, then content sha256) is never downloaded twice.
  • download β€” polite (rate-limited, robots-aware), magic-byte + PyMuPDF validated, size-bounded; URL chain per work: repository copies first, then the publisher PDF, then Unpaywall by DOI, then the landing page's citation_pdf_url.
  • ingest β€” extraction.py (grounded, unit-aware Gemini extraction, prompt v2.1: every value also carries the processing route of its specimen) via batch_ingest.process_pdf with the pg_mirror Postgres backend: every value is grounded in the PDF text, unit-normalized (value_si, unit_canonical), classified, deduped at measurement grain, and inserted with a status β€” flagged rows are quarantined for review, never dropped.
  • report β€” per-cycle metrics in agent_runs (dashboards + paper Β§8), event stream in agent_events, optional Gemini query expansion when yields dry up.

The graph runs on LangGraph; the agent loop runs in-process on an APScheduler interval, so the Space works autonomously while it is up β€” no viewers needed.

Pages

Page What it shows
Dashboard agent status pills, KPI tiles, rows/PDFs per cycle charts, latest activity feed, latest DOIs, recent cycles
Run Control autonomy on/off, interval, per-cycle budget, source toggles, figure-mining toggles, query frontier editor, manual "run one cycle"
Recently Added newest ingested documents with DOI links + verified-row counts
Review Queue all status != 'ok' rows with origin filter, the figure crops behind figure-estimate rows, CSV export, origin-scoped human promote-to-ok
Runs & Logs cycle history (status chips, durations, live-refreshing while a cycle runs) with per-PDF results, event stream grouped per run
Database material-first browser (one expander per material with the process type and conditions of each property; πŸ“ˆ marks figure-derived estimates) + raw-table browser (incl. the figures provenance table) with filters, pagination, CSV export
Knowledge Graph every row of the three material tables as one graph, the review queue included and marked as such: whole-database map, matrix and fiber pairs, single-node view with the measured values; each value opens to its page, quoted sentence, flag reason, figure picture, processing route and date added; rebuilt when the tables change and swapped into the open page; JSON and Neo4j exports

UI: single light theme + design system in agent/ui_theme.py (tokens, CSS, altair theme, KPI/pill/chip/empty-state components); pages compose those helpers and contain no styling of their own.

Since 28 Sep the Dashboard carries a "Candidates queued" tile, Run Control a "Candidate queue" card (toggle, pass interval, low-water mark, per-pass budgets, queue status by source, "Run a discovery pass now") and Runs & Logs a Kind column (cycle / discovery) plus a Queued counter for passes.

Deploy (β‰ˆ5 minutes)

  1. Create the Space: huggingface.co β†’ New Space β†’ owner aim4composites, name e.g. AutonomousAgent, SDK Docker, visibility your call.

  2. Push this folder to it:

    git clone https://huggingface.co/spaces/aim4composites/AutonomousAgent
    cp -r <this folder>/* <this folder>/.streamlit <this folder>/.gitignore AutonomousAgent/
    cd AutonomousAgent && git add -A && git commit -m "Autonomous ingestion agent" && git push
    

    (or upload the files through the web UI β€” Files β†’ Add file.)

  3. Set Space secrets (Settings β†’ Variables and secrets):

    Secret Value
    DB_HOST / DB_PORT / DB_NAME / DB_USER / DB_PASSWORD same values the MaterialsDatabase Space uses
    GEMINI_API_KEY a fresh Gemini key (rotate the leaked ones!)
    AGENT_ADMIN_PASSWORD password that unlocks Run Control / promote
    CONTACT_EMAIL (variable, optional) polite-crawling contact for OpenAlex/Unpaywall
    OPENALEX_API_KEY required since 2026 β€” OpenAlex bills per call: anonymous = $0.10/day per IP (shared with every other Space on the egress IP), a free account key = $1/day β‰ˆ 1,000 searches or 10,000 filter calls. Get it at openalex.org β†’ Settings β†’ API
    S2_API_KEY (optional) dedicated Semantic Scholar quota (the bulk lane works without a key on the shared bucket; a free key is a form away)
  4. One-time schema check: the shared Postgres must already have the hardening columns + sources table (python pg_migrate.py dry-run, then --apply, from any machine with the DB env set). If the July pipeline runs already did this, there is nothing to do β€” the agent refuses to write until the schema is migrated, so it fails safe either way.

  5. Open the Space β†’ Run Control β†’ unlock admin β†’ Run one cycle now to smoke-test β†’ toggle Autonomous crawling ON.

Process type and conditions per property (prompt 2.1, Oct 2026)

A composite's properties depend on how the part was made, so every property row carries the processing route of the specimen it was measured on:

Column Content
process_type route, normalized to extraction.PROCESS_TYPE_ENUM (injection molding, compression molding / hot press, autoclave, out-of-autoclave, automated fiber placement / tape laying, thermoforming / stamp forming, filament winding, pultrusion, extrusion / compounding, additive manufacturing, liquid molding, welding / joining, fiber spinning, film / solution casting, electrospinning, heat treatment / annealing, other)
process_name the route as the document names it
process_conditions processing parameters as printed, ; -joined (temperature, pressure, hold time, cooling rate, nozzle/bed temperature, print speed, ...)
process_quote, process_page the verbatim sentence that names the route, and its page
process_status grounded / grounded_off_page (sentence found in the PDF text and every number of the conditions found on digit boundaries), conditions_unverified (sentence found, a number is not), ungrounded (sentence not found), unchecked (no text layer)
  • The model returns the routes once per document (processes, ids P1, P2, ...) and each property points at one through process_id, so the output grows by one short field per property, not by a repeated paragraph. The same route at two settings is two routes.
  • All six columns are NULL when the document does not say how the material was made (most vendor datasheets, as-received constituents, values quoted from other literature). The prompt forbids guessing.
  • The route is context, not a value: its check never changes a row's status. An unverified route stays attached and marked (⚠ in the Database page).
  • test_condition is still the test condition (23 Β°C, 50% RH); the Processing section still holds process settings a document reports as values of their own.
  • Schema: the columns are appended to migrate.EXTRA_COLUMNS, so the boot migration adds them to the three material tables (ADD COLUMN IF NOT EXISTS); nothing is run by hand. The dedup grain is unchanged.
  • Rows stored before this change (prompt 2.0) keep an empty route. The source PDFs are deleted after ingest, so filling them needs a re-download pass; that is not part of this change.
  • Pages: Database (Process and Process conditions columns per material, a Process filter, a "With processing route" tile, the columns in the raw browser and its CSV), Review Queue (three process columns, also in the CSV), Dashboard ("Processing routes" card: rows per route, share grounded).
  • Tests: test_process_e2e.py (parse, grounding, rows, boot migration of an old database, a stubbed cycle into Postgres, the three pages).

Knowledge graph (live, Oct 2026)

The Knowledge Graph page shows the whole database as one graph and keeps it current without anyone running anything.

  • What is in it. agent/kg.py reads every row of Polymers, Fibers and Composites_materials through a server-side cursor (no row limit, the image bytes left out) and every row of figures, and builds the graph: materials, properties, property categories, polymer and fiber families, manufacturers, source documents (DOI, year and lane from agent_doi_seen), figures and process types. Every measured value keeps its status, page, SI value, route, the day it was added and the figure it was read off or cites.
  • The review queue is in it. Rows with status != 'ok' are not left out: each is marked "review queue" with its status and flag_reason, and the switch above the graph shows all values, the verified ones, or only the queue (counts, map and lists follow the switch). Promoting a row in the Review Queue page changes the fingerprint, so the graph follows.
  • Where a value was found. Opening a value (the arrow at the start of its row) shows the page, the quoted sentence (source_quote), the model's comment, the figure with its picture, the route's own name, quote, page and grounding status, and the date, model and prompt version of the record. These fields live in separate detail files of 2,000 values each (static/kg/d/<build>/<n>.json.gz), fetched when a value is opened, so the main file stays small as the database grows.
  • Figures. Every harvested figure is a node, linked to its document and to the values read off it or citing it. The pictures of figures with values are copied from figures.image_bytes to static/kg/fig/<figure_id>.png in the background (20 per query, at most 4 minutes per run of the timer job, only what is not on disk yet; after a restart this takes a few minutes). AGENT_KG_FIGURES = linked (default) | all (every stored crop) | off.
  • When it rebuilds. A fingerprint (per table: rows, verified rows, newest extracted_at, rows with a process type; and the number of figures) is compared with the one the current graph was built from. It is read right after every cycle, every AGENT_KG_REFRESH_MINUTES (default 10; 0 switches the timer off) by the scheduler job kg-refresh, and when someone opens the page (at most every 20 s). The graph is rebuilt only when the fingerprint changed, so an idle database costs four small aggregate queries per check.
  • How the open page follows. The graph is written to static/kg/ (kg-data.json, kg-data.json.gz, kg-meta.json) and served by Streamlit at /app/static/kg/... (enableStaticServing in .streamlit/config.toml; the Dockerfile makes the folder writable; AGENT_KG_STATIC_DIR moves it). The explorer (agent/kg_explorer.html, canvas + d3 from cdnjs) asks for kg-meta.json every AGENT_KG_POLL_SECONDS (default 60) and swaps the new graph in, keeping the open node and the zoom. Each build has its own detail folder (the one before is kept for pages still showing it). If the folder cannot be written, the graph and the value details are embedded in the page instead, a reload picks up changes, and figure pictures are not shown.
  • What is inferred. Polymer and fiber families are assigned by pattern rules, and property names are merged for case, spelling and a short synonym list. Everything else is the database as stored. The map is computed from names and counts, not simulated: the same database always gives the same picture, and it shifts only where families grow.
  • Exports. The published JSON can be read from outside the Space (https://<space host>/app/static/kg/kg-data.json; detail.dir in it names the folder of the detail files). The page also offers a Neo4j zip (nodes.csv, relationships.csv, load script; one Measurement node per value with its quote, flag reason, day added and review-queue mark; Figure nodes). python -m agent.kg --out DIR writes the same files, the figure pictures and a stand-alone explorer.html from the configured database.
  • Cost. On the 1 Oct snapshot (34,768 rows, 1,093 figures in the test copy): build about 3 s, 5.6 MB of JSON, 0.9 MB compressed. At ten times the rows: build about 20 s, 46 MB / 7 MB, about 0.7 GB of memory at the peak. No model calls.

Figure & graph mining (in-cycle)

Every ingested PDF also passes through figures.py (the same module the local pipeline uses): plots and table images are harvested from the PDF's layout (embedded rasters + vector clusters, junk filters, caption pairing), classified with one Gemini vision call per PDF, and mineable kinds (property_plot, table_image) are read for axis ranges and salient values.

A number read off a graph is an estimate, not a grounded fact. Figure rows are inserted as origin='figure', status='figure_estimate' β€” they never publish (status='ok') at insert, whatever the caller does (batch_ingest._row_values() downgrades them), and only a human promote in the Review Queue can bless one. The crop PNG itself is stored in the figures table (image_bytes, capped at ~1.5 MB) so the Review Queue can show the reviewer the exact plot behind every estimate β€” the Space's disk is ephemeral, the database is not.

Knobs (Run Control, or env for first-boot defaults): AGENT_FIGURES (on by default), AGENT_FIGURE_MINING (off = harvest+classify only), AGENT_MAX_FIGURES (per-PDF cap, default 12, hard cap 20). Per-run counters (figures_found, figures_mined, figure_rows, vision_calls) land in agent_runs for the Runs & Logs table and the paper's cycle accounting.

Operations notes

  • Budgets: per-cycle caps (default 5 PDFs) bound Gemini spend; a hard ceiling of 25 is compiled in. Cycle lock is durable and self-expiring, so overlapping schedules or container restarts can't double-run.
  • Keeps running on its own: the container entry point (run_space.py) boots the agent (schema check, stale-run sweep, scheduler) at process start, so cycles run from the moment the container is up, with or without a visitor. A keep-alive thread GETs the Space's own URL (https://$SPACE_HOST/_stcore/health) every AGENT_KEEPALIVE_MINUTES (default 10; 0 disables; AGENT_KEEPALIVE_URL overrides) so a free Space never reaches its 48-h idle sleep. Keep an external hourly pinger as a backup: if the container ever does sleep, any request wakes it and the agent boots itself.
  • Watchdog: every cycle runs under a wall-clock budget (AGENT_CYCLE_TIMEOUT_MINUTES, default 60). A cycle still running past it is marked timeout in Runs & Logs, its lock is released and the next scheduled cycle proceeds. Without this, one cycle stuck in a network wait blocked the scheduler for good (the single-instance job never returned), which is what the long runs of stale cycles in August/September were. The crawler's Retry-After sleeps are capped at 300 s (the cap had been lost in the Aug 25 crawler sync), so that wait should no longer happen.
  • Stall detector (22 Sep): besides the hard budget, a cycle whose last agent event is older than AGENT_CYCLE_STALL_MINUTES (default 15) is abandoned with the stall written into its verdict; a slow but reporting cycle is bound only by the hard budget. heartbeat_at on agent_runs is the liveness signal (every event touches it).
  • Budgeted sources (22 Sep): a 429 whose Retry-After exceeds the 300 s cap, or that reports a zero remaining budget (OpenAlex's per-request billing), closes that host for its window with no retry and no sleep; the lane is skipped with one warn event per cycle and the reason lands in agent_runs.sources. Set the Space secret OPENALEX_API_KEY (a free account key is ten times the shared anonymous allowance). AGENT_USE_NTRS (default on) adds the NASA Technical Reports Server lane: free, no key, direct PDFs.
  • Cost column (22 Sep): Gemini usageMetadata is summed per cycle into tokens_in / tokens_out (text + vision calls); Runs & Logs derives dollars at read time from AGENT_PRICE_IN_PER_M / AGENT_PRICE_OUT_PER_M (defaults 0.30 / 2.50 USD per million, Gemini 2.5 Flash list price).
  • Figure citation linking (22 Sep): AGENT_LINK_FIGURES (default on, Run Control toggle) links each text row to the figure its evidence cites and embeds the PNG in the row's image column (no S3 needed). Adds three columns (figure_ref, figure_link_score, figure_link_signals) that the boot migration creates; rows_linked is counted per cycle.
  • Frontier rotation on dry cycles (28 Sep): every query a cycle searched is marked used in node_report, whether or not the cycle downloaded anything. Until then the bookkeeping sat behind node_ingest's early return, so a dry cycle left last_used_at untouched and next_queries (LRU, never-used first) handed the same 20 queries back every half hour: from 22 to 28 Sep every cycle rediscovered the identical 109 exhausted candidates (42 DOIs already registered, 66 URLs already final) while the Gemini-expanded queries piled up unreached. Symptom to watch for in Runs & Logs: byte-identical Candidates / Skipped / URL seen columns across consecutive cycles. The harvest loop also no longer fetches (and marks) a frontier slice past its last round.
  • Downloads that never got ingested (28 Sep): a cycle that finds the material schema unmigrated ends as error (not crashed) and keeps its PDFs and their downloaded registry rows; the next cycle adopts and ingests whatever is still on disk (adopted N PDF(s) event). At boot, downloaded rows older than 3 h whose file is gone are deleted and their URLs re-opened in the crawler seen-set (boot repair: re-queued ...), so discovery finds and fetches those papers again β€” this is what frees the five PDFs runs #135/#139 lost on 15–16 Sep. The Inspect-run panel now shows why a run ended (crash traceback, watchdog note or cycle error).
  • More papers (28 Sep): discovery lanes on by default β€” OpenAlex phrase search and topic walk (AGENT_OPENALEX_TOPICS, 1-credit filter calls with persisted cursors, round-robin), Semantic Scholar bulk (≀1,000 OA-PDF papers per call; S2_API_KEY optional), arXiv with ANDed terms, NASA NTRS, and a datasheet round over the manufacturer seed pages (each seed ≀ once a day); the query frontier is seeded once with the 455-intent grid (agent/query_grid.py, origin grid). Downloads try repository copies before publisher URLs and fall back to the landing page's citation_pdf_url; records without an abstract get relevance-gate relief. Note that publisher hosts (MDPI, Elsevier, Wiley, …) 403 any cloud egress address, so the yield comes from repositories, NTRS, arXiv and datasheets.
  • Parallel ingestion (28 Sep): a cycle extracts its PDFs on a thread pool (ingest_workers in Run Control, default 3, hard max AGENT_HARD_MAX_INGEST_WORKERS = 6), one Postgres connection per worker; counters, events and the registry are written by the cycle thread as each PDF completes. Run Control also exposes the cycle's target (Target new PDFs per cycle, default 10: below it the cycle takes another frontier slice, up to Max rounds per cycle) β€” together with the 25-PDF cap these are the throughput knobs. A cycle's OpenAlex cost is 10 credits per query per round plus 1 per topic page, so 4 queries Γ— 5 rounds Γ— 48 cycles is β‰ˆ 9,700 credits/day: right at the free key's allowance; prepaid credits or a longer interval buy headroom.
  • Candidate queue (28 Sep): discovery and download are decoupled. A discovery pass (agent/discovery.py, an agent_runs row of kind discovery) searches the next discovery_queries frontier queries on every lane β€” Semantic Scholar bulk, arXiv, NTRS, UDSpace (UD's DSpace theses, direct bitstreams), CORE (needs the free CORE_API_KEY secret; inactive without it) and, for the first discovery_openalex_queries, the 10-credit OpenAlex phrase search β€” walks discovery_topic_pages pages per OpenAlex topic (1 credit each) and crawls every datasheet seed that is due, then queues every relevant candidate in agent_candidates (a DOI already in the registry or a URL the crawler has given up on lands as seen). It runs every discovery_interval_hours (24), right after any cycle that finds the queue below queue_low_water (150) or drains it to that, at boot when the queue is empty, and on demand (Run Control β†’ "Run a discovery pass now"). Cycles pop the queue best-score-first (repository-hosted copies rank ahead of publisher URLs) and download with no search calls at all; they search inline only when the queue is empty. Every popped row is settled: ingested, seen (registry / final URL), back to queued after a transient failure (never twice in one cycle; failed after 3 attempts). Turn it off with the Run Control toggle to get the pre-28 Sep inline behaviour. OpenAlex spend drops from up to ~9,700 credits/day to a few hundred per pass.
  • Secrets never reach the database (28 Sep): agdb.redact() strips credential query parameters (?key=, &api_key=, token=) and bearer tokens from every event, run report, registry status and candidate error before it is stored, and extraction raises Gemini errors that name the endpoint and the API's own detail instead of the keyed URL (requests' HTTPError / urllib3's connection errors quote it). At boot, agdb.redact_stored_secrets scrubs rows written before this existed and logs boot repair: redacted credential parameters … β€” if that event ever appears, rotate the key it refers to: it was readable on Runs & Logs.
  • Gemini budget refusals (28 Sep): a 429 whose body says RESOURCE_EXHAUSTED / "monthly spending cap" is not retried (it cost 14 s per PDF and cleared nothing). The PDF stays on disk with its registry row downloaded, the cycle's remaining PDFs are not attempted, one warn event says Gemini refused for budget reasons (…); N PDF(s) kept on disk for the next cycle, and the run ends error only if nothing was ingested at all. The next cycle adopts the kept PDFs (or the boot repair re-queues them after a restart). Raise the cap at ai.studio/spend. While the refusal lasts the agent is on a budget hold (agent_state gemini_budget_hold): cycles download nothing and probe with one kept PDF (or one fresh download if the disk was wiped); the first successful extraction lifts the hold and downloads resume next cycle. Adoption of kept PDFs is capped per cycle at the download cap, so a backlog is worked off over several cycles.
  • Cost (28 Sep): Gemini's thinking tokens are billed as output at 6Γ— the input price, and with no thinkingConfig the model thinks with an unbounded dynamic budget: a 25-paper cycle cost $7.50, $6.69 of it output (30k tokens per PDF for a ~4k-token extraction). gemini_thinking (Run Control β†’ Gemini cost; env GEMINI_THINKING; default low) is applied to every Gemini call β€” dynamic | low | minimal | high | off | an integer budget; a model that rejects the option gets the call repeated without it and the option is dropped for the process. Figure mining (one thinking vision call per mineable figure, into quarantined figure_estimate rows) is off by default; harvest + classify (one call per PDF) and figure linking stay on. Both defaults are applied once to a stored config at boot (cost defaults applied: … event) and can be changed back in Run Control. Expected: roughly a third of the previous bill per paper β€” watch the Cost column, and rows per paper / flagged share for recall.
  • Cost (30 Sep) β€” real prices, models, near-duplicates: the cost column used to price every token at Gemini 2.5 Flash rates ($0.30 / $2.50 per M) while the agent runs gemini-3.5-flash ($1.50 / $9.00): it showed about a quarter of the bill (runs 544–593: $12.98 shown, ~$48.50 real, ~$0.41 per PDF on dynamic thinking with mining on). Runs now store cost_usd, priced per model (agent/config.py MODEL_PRICES), text and vision calls each at their own model; older runs are re-priced at gemini-3.5-flash in Runs & Logs, which also shows thinking tokens, the model, $/PDF and a spend line. Run Control β†’ Gemini cost adds the text model (every row is stamped with it: change it only after cost_ab.py on the gold set), a vision model for the figure calls (default: same), and the near-duplicate gate (agent/textdedupe.py, default 0.9): a PDF whose text is that similar to an ingested one β€” regional datasheet variants, a preprint and its published version β€” is skipped before any Gemini call (duplicate_text in the registry and the Near-dup column).
  • Ephemeral disk: PDFs live under /tmp only until ingested; provenance (DOI, sha256, title, URL, query) is durable in agent_doi_seen.
  • Safety: agent bookkeeping is in namespaced agent_* tables; controls are password-gated; read-only without the password. At boot the Space additively migrates the connected database (hardening columns incl. origin/figure_id, the origin-aware dedup index, the figures table) β€” aim_agent is the agent's own DB, so this is on by default. Before pointing DB_NAME at the team's shared database, either agree on the additive migration or set AUTO_MIGRATE_MATERIAL_SCHEMA=0 (the cycle then refuses to write until pg_migrate.py --apply is run by hand).

Repo map

Vendored pipeline (single source of truth, synced from the project repo β€” the only Space-side deltas are the PG keepalive/statement-timeout hardening and the Retry-After cap): extraction.py, batch_ingest.py, figures.py, pdf_crawler.py, pg_mirror.py, pg_migrate.py, migrate.py, generate_queries.py. FIGURES.md documents the figure stage.

Agent layer (new): agent/ (orchestrator, scheduler with watchdog, autostart, keepalive, DB, config, query_grid), run_space.py (container entry point), app.py + page_files/ (Streamlit UI), agent/kg.py + agent/kg_explorer.html (knowledge graph builder and explorer; test_kg.py). Tests: test_agent_state.py, test_harvest_loop.py, test_sources.py (no DB, no network), test_e2e.py, test_figures_e2e.py, test_keep_running.py, test_counters_e2e.py, test_frontier_rotation_e2e.py, test_more_papers_e2e.py, test_parallel_ingest_e2e.py, test_candidate_queue_e2e.py (scratch Postgres via DB_*; the network sources are switched off by env), test_gemini_errors.py (no DB, no network).