AI & ML interests
molecular foundation models, Bayesian optimization, knowledge graphs, low-data regimes, preclinical drug development
Recent Activity
Optagon Labs
We build digital twins for preclinical drug development. A twin reasons across hit-to-lead, lead optimization, candidate selection and early CMC at once, so a program can co-optimize properties that are usually traded off against each other stage by stage. It stays with the program after the engagement, as a system the team can keep tuning.
The work sits on three pieces: a molecular foundation model, a state representation that carries a program's history, and Bayesian optimization adapted to the very low data regimes real preclinical work operates in. Every recommendation carries its uncertainty and where that uncertainty came from.
What is published here
This organization holds the datasets behind that work: structured knowledge graphs extracted from the literature and from partner data, together with the manifests that record every source a graph was built from.
Each corpus is one directory:
tables/— the graph itself, one JSONL file per tablemanifest.csv— every source document, with its DOI where one existsREADME.md— row counts, the schema version, and the exact revision of the extraction code that produced it
The extraction pipelines, the schema and the loaders live on GitHub at Optagon-Labs. Each consumer pins the dataset revision it was built against, so a result can always be traced back to the graph and the code that produced it.
Access
Corpora are private. Graphs built from partner data stay with that partner, and source publications are referenced rather than redistributed.
For access or collaboration: optagonlabs.com