Autonomous agentsCrystal chemistry

Two million hypotheses refuted.
Eight crystal laws left.

Agents proposed, implemented and actively refuted their own candidate laws. The eight survivors do not return pass or fail: they name the physicochemical constraint that breaks.

Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening
Zhilong Song1,2  ·  Lixue Cheng1,2,3,✉
1Department of Chemistry, Hong Kong University of Science and Technology  ·  2IAS Center for AI for Scientific Discoveries, HKUST  ·  3Department of Chemical and Biological Engineering, HKUST
2,037,606
candidate law evaluations
572
numbered investigations · 11 self-refutations
8
surviving one-line laws  ·  PRIS
Every rejection points to one of five mechanisms
Short-range repulsion
Ionic contact & packing
Electrostatic balance
Bond-valence conservation
Site complexity
87.9%
of chemically damaged structures detected by the strictest law set
the 0.5- and 0.7-Å distance cutoffs used in generative pipelines detect 1.6–3.2%
82–99%
of experimental structures satisfy the law sets
Pauling's rules 2–5 applied jointly are satisfied by only 6.5%
83.7%
of hard-to-synthesize structures screened by the PRIS-derived synthesis score
while retaining 80.7% of experimental structures, and no synthesis label entered the discovery
−67.3%
DFT validation queue in a property-conditioned inverse-design run
keeping 99.2% of the candidates whose DFT-validated bulk moduli reach the target
Abstract

From a pass-or-fail verdict to a chemical diagnosis

Crystal generators and tool-using agents propose structures faster than density functional theory (DFT) energy and phonon calculations or experiments can assess them. Deciding which candidates merit expensive assessment is therefore the bottleneck, yet most screens test little beyond atomic overlap and give no chemical reason for failure. Here, our agents generate, test and actively refute two million candidate laws, leaving eight Plausibility Rules for Inorganic Structures (PRIS). These laws encode five mechanisms: short-range repulsion, ionic contact and packing, electrostatic balance, bond-valence conservation and crystallographic site complexity. Experimental structures satisfy our law sets at 82–99%, but satisfy Pauling's rules 2–5 together at only 6.5%. The strictest set detects 87.9% of damaged crystal structures, whereas distance cutoffs detect only 1.6–3.2%.

PRIS plausibility is linearly correlated with synthesizability, so the PRIS-derived synthesis score (PSS) explainably screens 83.7% of hard-to-synthesize structures while retaining 80.7% of experimental structures. In a property-conditioned inverse-design run, PRIS and PSS can reduce the DFT validation queue by up to 67.3% and keep 99.2% of the candidates whose DFT-validated bulk moduli reach the design target. Beyond screening, PRIS explains why GNoME remains enriched in rare low-symmetry structures and reveals how wrong-element assignments in falsified crystal reports hide behind plausible coordinates. PRIS moves screening from a pass-or-fail verdict to a chemical reason for failure, showing that autonomous agents can discover, by active refutation, physicochemical laws that guide calculations and experiments.

The gap

Distance cutoffs are element-blind. Pauling's rules are too idealised.

A good plausibility law must keep experimental structures, detect damaged ones at a high rate, and say why a structure is implausible. Measured on both axes at once, the two classical answers err in opposite directions.

Minimum-distance cutoff

0.5 Å or 0.7 Å, as deployed in generative pipelines
1.6–3.2%damage detected

Inexpensive and almost always satisfied, but it examines neither chemical ordering nor elemental identity. Avoiding gross overlap does not make coordination, electrostatics or bond valence plausible.

Pauling's rules 2–5, jointly

1929 · radius table and formal charges
6.5%experimental structures satisfied

Chemically reasoned but too strict when applied together: an earlier audit found 13% of about 5,000 oxides satisfied rules 2–5, and among charge-balanced ionic experimental structures only 6.5% did.

PRIS, Set 4

eight one-line laws · five mechanisms
82–99%satisfied · up to 91% detected

Keeps experimental structures, detects damage, and every failure names the violated law and the mechanism to review. Runs on the structure as given, without relaxation or a phase-hull reference.

Experimental structure
✓ All eight laws satisfied
Alternating cations and anions on a regular lattice.
Wrong element
✗ Fails Law 6 · charge topology
Same coordinates, one species exchanged. Formal charges, Madelung energies and bond valences all change.
All sites split
✗ Fails Law 7 · site complexity
Every site made crystallographically distinct, the pattern behind artificial ordering.
All three pass the 0.7 Å distance cutoff commonly used in generative pipelines. Geometrically clean does not mean chemically sound.
The laws

Eight laws, each a single computable and refutable statement

Every law is a one-line predicate over quantities computable from a structure file and a radius table: no relaxation, no training, no synthesis label. A conditional law is satisfied by any structure whose trigger is not met, so the clauses of Law 2, Law 3 and Law 6 confine each law to its domain.

L1
ρ ≥ τ  (τ = 0.735 permissive, 0.804 strict)
The shortest cation–anion contact, relative to the sum of the two Shannon radii, may not fall below a floor.
short-range repulsionSet 1 · 1′ · 2 · 3 · 4
L2
fi > 0.50  ⇒  ρ ≤ 1.05
An upper bound on contact applied only where the ionic model holds: strongly ionic compounds may not be too loose.
ionic contactSet 1′
L3
mean anion CN ≤ 3.333  ⇒  mean d/(rcat+ran) ≤ 1.081
Low-coordination structures may not also be loosely packed on average.
packingSet 2 · 3 · 4
L4
rangei of EM(i)/|zi| ≤ 31.45 eV
No site may be electrostatically far out of line with the others, judged by the site Madelung energy per unit charge.
electrostatic balanceSet 2 · 3 · 4
L5
maxi EM(i) ≤ 15.17 eV
No single site may have an implausible Madelung energy.
electrostatic balanceSet 2 · 3 · 4
L6
fi > 0.55  ⇒  no like-charge bonds
Charge topology: ionic compounds do not bond like to like. This is the law a wrong-element assignment breaks without moving an atom.
electrostatic balanceSet 3 · 4
L7
inequivalent sites / sites ≤ 2/3
A distinct-site bound consistent with a preference for simpler structures. Needs no charge assignment, so it can audit a whole catalogue.
site complexitySet 4
L8
mean |BV sum − vi| / vi ≤ 0.7143
A permissive tail bound on the bond-valence sum, the quantity of Pauling's rule 2.
bond-valence conservationSet 4

ρ is the reduced contact ratio (shortest cation–anion distance over the sum of the Shannon radii), fi is Pauling's composition-based estimate of ionic character, EM(i) is the site Madelung energy from an Ewald sum over formal charges, zi the formal charge at site i, and site complexity is inequivalent sites over sites at spglib symprec 0.01. Charges are formal oxidation states inferred from composition, never from bond lengths, so a bond-length law is never tested on its own conclusion.

Five nested law sets: experimental-structure satisfaction vs damage detection
Held-out benchmark: 5,297 real and 3,612 damaged structures never seen by any fitting step. Pauling's rules 2–5 jointly are satisfied by 6.5% of experimental structures, and distance cutoffs detect almost no damage.
0.000.250.500.751.00 0.800.850.900.951.00 0.5 / 0.7 Å distance cutoffs Set 1 Set 1′ Set 2 Set 3 Set 4 ↖ up and to the left: more complete law sets, higher detection Satisfaction of experimental structures (held-out) Damage detection
The sets are nested, and each is named for the crystal model it enforces
set · crystal model · lawssatisfieddetected
Set 1 hard-sphere floor
Law 1 (τ = 0.735)
0.99190.2890
Set 1′ two-sided window
Law 1 (τ = 0.735) + Law 2
0.98940.3837
Set 2 rigid-ion lattice
Law 1 (τ = 0.804), Law 3–Law 5
0.95790.6121
Set 3 ionic network
Set 2 + Law 6
0.91710.7004
Set 4 crystal chemistry
Set 3 + Law 7, Law 8
0.81800.9111
Satisfaction and detection are always quoted from one population, and the population is always named. Set 4's worst damage class is still detected at 0.7338. Set 1′ stands beside the chain rather than inside it, which is why the catalogue holds eight laws while Set 4 applies seven.
How they were found

Proposed, tested and actively refuted by the agents themselves

A pre-registered protocol fixed the criteria, the data split and the vocabulary before any evaluation. The agents then ran 572 numbered investigations over 99,162 experimental crystal structures. Refuted claims re-enter the search, so the cycles form a sequence of falsifiable experiments.

PROPOSE
41 literature queries and 14 agent workflows turn crystal-chemical intuition into candidate one-line laws.
TABULATE
84 descriptor tables computed over 99,162 structures: contacts, Madelung energies, bond valences, site counts.
SEARCH
Systematic search over interpretable law combinations and fitted thresholds, benchmarked against provably optimal decision trees.
TEST
Held-out data and negative controls. Thresholds are fitted only on the discovery split and never touched again.
REFUTE
11 claims written down as results were overturned by the agent that produced them, and returned to the search.
11
refuted conclusions are published with the surviving results, each paired with the cheap diagnostic that exposed it (Supplementary Figs. S1–S2).
72,583
result files archived and 615 analysis scripts written along the way, all kept as the audit trail of what was actually executed.
0
synthesis labels entered the discovery of the laws, so the correlation between PRIS plausibility and synthesizability is a finding rather than a fitting target.
Beyond the benchmark

Plausibility tracks synthesizability, and shortens the DFT queue

Set 1 to Set 4 are conservative discrete screens that name a violated mechanism. PSS, the PRIS-derived synthesis score, refits the same quantities to the experimental record of what has been made and gives continuously tunable control over how strongly the combined evidence shortens a queue.

Hard-to-synthesize structures screened  ·  no synthesis labels used in discovery
83.7%while retaining 80.7% of experimental structures
PSS synthesis score83.7%
Ehull threshold (needs relaxation)72.0%
At matched satisfaction PSS screens 31.8 percentage points more than a MatterSim-computed hull-energy threshold, which additionally requires a relaxation and a phase-hull reference. On the most confident fifth of same-composition pairs PSS reaches 0.944 accuracy against 0.844 for DFT energy above the hull.
DFT validation queue reduced  ·  MatterGen inverse design, target bulk modulus ≥ 400 GPa
−67.3%of the validation queue
Queue reduction67.3%
On-target candidates retained99.2%
In a property-conditioned inverse-design run, PRIS and PSS reduce the DFT validation queue by up to 67.3% and keep 99.2% of the candidates whose DFT-validated bulk moduli reach the design target.
Ordering artefact · identity failure

The coordinates pass. The chemistry doesn't.

Artificial element ordering in generated catalogues and swapped element identities in falsified crystal reports both slip past geometric checks and energy relaxation. Neither slips past the eight laws.

CASE 1Why GNoME has so many low-symmetry structures5,000 uniformly drawn entries · pre-fixed audit protocol
whole sample 40.6% (n = 5,000) P11.000n = 727 Pm0.966n = 757 Cm0.884n = 345 Amm20.166n = 380 Others0.077n = 2,653 C2/m0.014n = 138
Fraction not satisfying Law 7 (distinct-site fraction ≤ 2/3), by source space group
78%Merge chemically similar elements into one label and the failing entries' distinct-site fraction drops back within the Law 7 threshold, not one atom moved
79%Space-group symmetry jumps after merging. Of 27 equally mergeable experimental entries, not one does this
0.0001
vs 0.036
Energy cost of scrambling the ordering (eV per atom), GNoME vs experimental, from DFT over every symmetry-distinct ordering. 18 of 23 GNoME entries are interchangeable below 300 K. 0 of 10 experimental structures are
These structures are not "crushed": unmodified GNoME parents release a median of under 0.001 eV per atom on relaxation, while compression and displacement damage release 0.22 and 6.05 eV per atom. A small relaxation energy rules out gross distortion, not electrostatic, bond-valence or site-complexity failures.
CASE 2Right coordinates, wrong element: the falsifier's trickLattice and coordinates fixed; only the species exchanged
Cation ↔ cationn = 69 Coordinate checks27.5% PRIS89.9% Cation ↔ anionn = 83 Coordinate checks48.2% PRIS98.8%
Damage detection for parent-matched element exchanges
≥70falsified crystal structures on record: genuine diffraction data kept, element identities rewritten (Harrison et al., 2010)
HirshfeldCaught only by rigid-bond alerts and odd metal–ligand distances, then pinned down by structure factors: after the fact, not at the gate
L1 + L8Four retracted entries on one M(C2H2N3Cl) framework, M labelled Cu, Ni, Mn or Fe, all give ρc = 0.53 and bond-valence deviations of 0.77–0.78, violating both laws at once
Swap the element and the formal charges, Madelung energies and bond valences change at once, without a single atom moving. Coordinate checks cannot see this. Electrostatic and bond-valence laws can.
CPU time
per structure
0.05–10.7 ms
contact law alone, 8 → 216 atoms
58 ms – 2.4 s
all eight laws, reference code
~102 CPU-hours
one DFT relaxation, literature estimate

Law 1 evaluates 10,000 cells of 6–20 atoms in under one second, and all eight laws on the same queue take approximately ten minutes. The same queue would cost approximately one million CPU-hours of DFT relaxation, so a full PRIS evaluation costs under one millionth as much. The eight laws do not replace DFT or experiment. They decide which structures earn them.

Figures

The five main figures

Click any figure to enlarge. Every figure is drawn from aggregate data committed in the repository, and figures/manifest.json maps each one to the script that draws it.

Figure 1: autonomous discovery by proposal, testing and refutation
FIG. 1Autonomous discovery by proposal, testing and refutation. a, Pre-specified workflow: candidate-law proposal, descriptor tabulation, systematic search, held-out testing and attempted refutation. Refuted claims return to the next cycle. Below, the eight surviving laws judge a candidate structure, and an implausible verdict names the unsatisfied law and the mechanism to review. b, t-SNE projection of the archived one-line candidate-law statements, with red circles locating the statements nearest Law 1–Law 8. c, Counts of candidate evaluations, archived result files, analysis scripts, investigations, refuted claims and surviving laws. d, Law 1–Law 8, their predicates and their membership in the five nested sets, with held-out satisfaction and damage detection for each set. e, Running best held-out performance versus investigation index.
Figure 2: satisfaction, damage detection and threshold transfer
FIG. 2Experimental-structure satisfaction, damage detection and threshold transfer. a, Damage detection versus satisfaction on the discovery split: the depth-limited decision-tree frontier, the systematic search over interpretable law combinations and the five law sets. b, Held-out damage detection by composition-preserving perturbation class and pooled. c, Satisfaction of Pauling rules 2–5, individually and jointly, and of the five law sets on the same held-out structures. d, Law thresholds re-derived at their defining percentiles on held-out data, shown as the ratio to the frozen values, with the percentage of held-out verdicts that change.
Figure 3: anatomy of the laws
FIG. 3Physical basis of the PRIS plausibility laws. a, Experimental MgAl2O4 and its five chemically damaged variants (uniaxial compression, cation–cation exchange, random displacement, isotropic expansion and cation–anion exchange), with the relevant descriptors and law verdicts beneath each structure (green satisfied, red unsatisfied); purple rings mark exchanged atoms. b, Held-out damage detection by perturbation class as laws are added, from Law 1 alone to Set 4; red outlines mark the combinations referenced in the text. c, DFT energy above each compound's own minimum for twenty experimental compounds rigidly scaled along the reduced-contact coordinate ρc (median and interquartile band, logarithmic left axis), together with the experimental distribution (right axis). The two Law 1 floors and the conditional Law 2 ceiling are marked with their domains; the dotted line gives the pre-registered 0.1 eV per atom cost.
Figure 4: physicochemical screening across validation tasks
FIG. 4Physicochemical screening across validation tasks. a, Overall damage detection for fixed-distance, composition-only and successive PRIS criteria. b, Damage detection by perturbation class for baselines, contact laws and law sets. c, Satisfaction versus the fraction of hard-to-synthesize structures screened, for Set 1–Set 4, the PSS threshold sweep, distance cutoffs and a MatterSim-computed hull-energy threshold. d, Set 4 violation and mean PSS across within-model CLscore deciles for CGCNN-PU and MatterSim-1M-MLP-PU. e, Held-out same-composition pair accuracy for PSS and DFT Ehull against the fraction of most-confident pairs retained. f, DFT-validation queue reduction versus retention of candidates predicted to have bulk modulus ≥ 400 GPa, across PSS thresholds.
Figure 5: physicochemical screening of generated and database structures
FIG. 5Physicochemical screening of generated and database structures. a, Minimum-distance-cutoff satisfaction versus Set 4 satisfaction among charge-assignable outputs from seven generators and from MP-20. b, Fraction of the GNoME sample failing Law 7 by source space group, and DFT order–disorder temperatures for 23 label-merged GNoME entries and 10 experimental structures. c, Energy released on MatterSim relaxation for unmodified GNoME parents and their damaged counterparts, with a DFT curve for twenty parents. d, Law-set verdicts against relaxation energy. e, Damage detection for parent-matched element exchanges: coordinate checks versus PRIS. f, Measured CPU time per structure for Law 1 and for all eight laws, against a literature estimate for one DFT relaxation.
Try it

Give it a structure file. It reports every law, the verdict and the mechanism.

Thresholds and PSS weights are read from the frozen artefacts in agent_loop/frozen/, so a verdict produced today is the verdict the manuscript reports.

Install and run

# Python ≥ 3.10
git clone https://github.com/AI4QC/PRIS.git
cd PRIS
pip install -r requirements.txt

python src/pris_analyze.py mystructure.cif
python src/pris_analyze.py --quiet *.cif   # one verdict line per file
python src/pris_analyze.py --json POSCAR   # machine-readable

When a structure fails, the point is the last column

$ python src/pris_analyze.py --quiet damaged/*.cif
IMPLAUSIBLE  compressed.cif   bond-valence conservation, short-range repulsion
IMPLAUSIBLE  expanded.cif     bond-valence conservation

Roughly 19% of structures cannot be judged (multiple anions, complex molecular groups, no integer or fractional charge assignment). "Skipped" does not mean "passed."

Example: MgAl2O4

MgAl2O4, 14 sites, charges from integer charge balancing, f_i = 0.759

law    quantity                        measured  thresh  verdict  mechanism
Law 1  reduced contact rho               0.9865  0.8040   ok   short-range repulsion
Law 2  reduced contact rho               0.9865  1.0500   ok   ionic contact
Law 3  mean reduced cation-anion contact 4.0000  1.0810   --   packing
Law 4  range of site Madelung / valence  3.6122 31.4500   ok   electrostatic balance
Law 5  largest site Madelung energy    -20.2144 15.1700   ok   electrostatic balance
Law 6  fraction of like-charge bonds     0.0000  0.0001   ok   electrostatic balance
Law 7  inequivalent sites / sites        0.2143  0.6667   ok   site complexity
Law 8  mean |BV sum - v_i| / v_i         0.0384  0.7143   ok   bond-valence conservation

Set 4  crystal chemistry           plausible
PSS  +3.915        VERDICT  PLAUSIBLE

Law 3's trigger (mean anion CN ≤ 3.333) is not met here, so the law is satisfied by its clause and reported as --. Formal charges come from composition. BVAnalyzer is never called, because it infers valence from bond lengths and would make the conclusion the premise.

Share

Posters and cover

One-page summaries of the paper, sized for social media. Every number on them is quoted from the manuscript.

Poster: Two million hypotheses refuted. Eight crystal laws left.
Two million hypotheses refuted. Eight crystal laws left.
The discovery funnel, the five nested law sets, the synthesis and inverse-design results, and the eight laws on one page.
Poster: Beyond coordinates, there is chemistry to ask
The coordinates pass. The chemistry doesn't.
Why GNoME has so many low-symmetry structures, the falsifier's right-coordinates-wrong-element trick, and the cost of a PRIS check next to a DFT relaxation.
PRIS cover image
Cover
The PRIS wordmark and the refutation funnel, in square and 3:4 formats.
Citation

Cite the preprint

@article{song2026pris,
  title         = {Autonomous discovery of new structure-plausibility laws for explainable
                   and rapid crystal diagnosis and screening},
  author        = {Song, Zhilong and Cheng, Lixue},
  year          = {2026},
  eprint        = {2609.01209},
  archivePrefix = {arXiv},
  primaryClass  = {cond-mat.mtrl-sci},
  url           = {https://arxiv.org/abs/2609.01209},
}

To cite the software and the frozen law definitions specifically, use the @software entry in the repository README, or GitHub's "Cite this repository" button, which reads CITATION.cff.