Abstract
A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM. Together, these engines span phase identification, multiphase decomposition and physics-constrained whole-pattern modelling. Gan Jiang converts analytical experience into executable skills by diagnosing failures, revising skill instructions and code, and validating revisions before reuse, without retraining the language model or changing the underlying physical models. Skills selected using development data and frozen before held-out evaluation achieve higher refinement scores than the original expert-designed skills across FullProf, GSAS-II and PyWPEM. The agent resolves strongly overlapping reflections, quantifies a five-phase ancient Egyptian cosmetic, tracks lattice evolution in an operating battery and compares atomic configurations in a disordered oxide catalyst. On DeltaXRDbench, it leads the evaluated methods in single- and multiphase identification across simulated and experimental data. Without supplied composition, single-phase top-1 accuracies reach 96.30\%, 81.78\% and 40.83\% on MP500, RRUFF and opXRD, respectively, compared with 58.00\%, 58.47\% and 26.45\% for the strongest comparator. These results demonstrate how an integrated scientific tool ecosystem can support agents that extract structural knowledge from measurements while accumulating validated analytical expertise that transfers to new samples.
Community
π Gan Jiang connects phase identification, decomposition and refinement, keeping every scientific judgement grounded in evidence. Feel free to give it a try at: https://ganjiang.asia/
I've run skill libraries in agents, and the failure mode is never the obviously broken skill. It's the one that passes on the case that generated it, then quietly misfires on a slightly different input weeks later while still returning something plausible. So the validation gate before a revised skill gets reused is the part I'd read closely β the self-learning loop is the easy half. What I want reported is skill churn: how many revised skills get reverted or superseded, and how long one survives before a new failure diagnosis overwrites it. High churn means the bank is a cache with extra steps, and retrieval cost matters too, because at some point picking a skill costs more than re-deriving the analysis. XRD is a good testbed for exactly this β ground truth is physical, so a bad skill shows up as a wrong phase, not a wrong vibe. https://huggingface.co/papers/2610.07862
Thank you
this is a really important point. The validation gate, skill churn, and retrieval cost are exactly why we focus on XRD: ground truth is physical, so a bad skill surfaces as a wrong phase rather than a plausible-sounding one. Weβre taking this seriously in the project. Please keep an eye on it
weβre launching soon.
Get this paper in your agent:
hf papers read 2610.07862 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper