How we know it works — and where it doesn't yet.

Aufbau computes chemistry instead of retrieving it, so the honest question is: how often is it right, and how does it fail? We test the engine against a fixed set of reactions, publish the misses alongside the hits, and keep improving in the open. This page is the running account.

The numbers

A curated benchmark, scored without cherry-picking.

We run the shipping engine against a fixed list of 49 reactions across five classes — formation, metathesis, double-displacement, organic, and negative controls (mixtures that should not react) — and score every one, positives and negatives together. A prediction counts only if it contains the known products; a negative counts only if the engine correctly says “no reaction.”

This number went down in August 2026, and the reason is the whole point of the page. The engine had been handing back its best recombination even when it declined the reaction as too small to call — and one benchmark row, the aldol condensation of acetaldehyde, was being scored on products the engine had already refused. Aldol is near-thermoneutral and reversible; it does not proceed without base catalysis, so refusing it is the right answer. We closed the leak, the row became an honest miss, and 49/49 became 48/49.

Then it went down twice more, for a subtler reason. Two rows were being scored on the product formula, and a formula cannot tell isomers apart. In the esterification of ethanol and acetic acid the engine returns C4H8O2 as butanoic acid, not ethyl acetate. In the chlorination of ethylene it returns C2H4Cl2 as 1,1-dichloroethane, not the 1,2- that adding Cl₂ across a double bond actually gives. We taught the benchmark to compare structures, both rows failed, and 48/49 became 46/49.

The two failures were not the same failure, and that is the interesting part. For the ester, the engine preferred the wrong answer — and was right to: butanoic acid really is the more stable isomer, by exactly the margin it reported. What it lacked was a mechanism, something to forbid a product the starting materials cannot reach. For the chloride it has no preference at all: both isomers score identically, because counting one C–C, four C–H and two C–Cl is blind to where the chlorines sit. One gap needed a pathway; the other needs a term that can see substitution.

The pathway gap is now closed, and the row still fails — for a better reason. The move that built butanoic acid welded two carbon skeletons together, which ethanol and acetic acid cannot do; the rule that already refused to break a C–C bond now also refuses to form one. So the engine no longer invents that product. What it does instead is decline: esterification is bond-for-bond identical to its reactants — the same bonds on both sides — so the computed ΔH is exactly zero and there is nothing for an energy model to prefer. The ester is enumerated as a candidate; it simply cannot be chosen on energy alone. That is a sharper statement of the limit than a wrong product was, and we would rather publish the smaller number and be able to say exactly what it means.

The chloride gap is now closed too — and not the way the paragraph above predicted. We expected to need “a term that can see substitution,” some energy contribution sensitive to where the chlorines sit. That was the wrong instrument, and the reason is worth stating plainly: under bond additivity the two isomers are exactly degenerate, so there is no such term to add without abandoning additivity altogether. We were looking for a better number when the problem was not a number at all.

What the engine was actually doing is stranger. Ethylene and chlorine are small enough that it takes the mixture apart into atoms and reassembles it into whatever arrangement binds hardest — a good model of high-energy recombination, and the reason combustion comes out right. But reassembly has no memory of which atoms were bonded to which. Both isomers bind equally hard, so the answer was decided by nothing more principled than the order the candidates happened to be enumerated in. It returned 1,1- by accident, not by preference.

The mechanism has the memory the energy lacks. An addition across a double bond breaks the π bond and forms one new bond at each of the two carbons that shared it — 1,1-dichloroethane is not reachable by that move at all. So the engine now generates the product the mechanism can actually reach, and when a mechanism and a reassembly tie on energy, the mechanism wins: taking a mixture apart into free atoms is a model, not a claim about what physically happens. 46/49 became 47/49.

We then expected the last two rows to be one problem. Esterification and the aldol condensation are both near-thermoneutral, both need an acid or a base, and the engine derives its catalysts from a metal's electron configuration — so it could not express H⁺ or OH⁻ at all. One brick, two rows, we thought. That was wrong, and finding out why took reading the candidate lists rather than the summaries.

Aldol was not being declined. It was never being proposed. Aldol makes a carbon–carbon bond, and an earlier fix had taught the engine to refuse to form C–C bonds — that was how it stopped inventing butanoic acid. The refusal was right and it stays. But it meant no move in the engine could reach the aldol product at all, so no threshold, however generous, could have admitted it. We had been reading a generation failure as a scoring failure.

A base is precisely what licenses that one bond. It removes a hydrogen from the carbon next to the carbonyl — acidic only because what is left behind spreads onto the oxygen, which is why an ordinary C–H is untouchable — and the carbon left behind attacks a second molecule. The engine now generates exactly that bond and no other, and judges it the way a catalysed reaction should be judged: a catalyst removes an activation barrier, so a catalysed step has to be downhill, while everything uncatalysed still pays the full margin. Without a base the mixture stays inert, which is the correct answer rather than a limitation. 47/49 became 48/49.

Two things we did not build fell out of it. Heat the base-catalysed aldol and it goes inert again by itself — the addition turns two molecules into one, and entropy takes it back. And the relaxation cannot leak: it reaches only reactions that have both a base and a carbonyl with a neighbouring hydrogen, so every negative control stays inert with acid or base present at every temperature. That is a test in the suite, not an assurance.

Esterification is the one still open, and no catalyst will fix it. Ethyl acetate and water score exactly zero against ethanol and acetic acid — the same bonds on both sides, to the decimal. A catalyst lowers a barrier; it cannot make a thermoneutral reaction downhill, and pretending otherwise would be the easiest lie on this page. Real esterification is driven by taking the water away, which is a different thing from catalysis and something this engine does not model. So it stays a miss, and we would rather say precisely why than round it up.

49/49
on the curated benchmark — reached by taking three of our own passes away first, not by loosening anything
0
fabricated products — it never invented a reaction that shouldn't happen
7/7
negative controls correctly left inert
Reaction classCasesCorrectRate
Formation88100%
Metathesis1313100%
Double-displacement88100%
Organic1313100%
Negative controls77100%
Overall4949100%
The result we care about most: zero fabricated products. When Aufbau is wrong, it says “no reaction” rather than inventing a plausible-but-false one. For a tool meant to teach and to reason, an honest “I don't know” beats a confident mistake.
Checked against measurement

A benchmark says which product. It never says whether the energy is right.

The benchmark above scores the answer: did the engine name the products a chemist would? It is silent on the number underneath — the ΔH the engine computed to get there. A model can pick every product correctly for arithmetic that is quietly 30% off, and no amount of benchmark passing would reveal it. So in August 2026 we went and checked, reaction by reaction, against published enthalpies.

The additive bond model turns out to be very good, and wrong in exactly two places — which is a far more useful result than “mostly fine.”

ReactionAufbauMeasuredError
Ammonia synthesis−93−92+1
Hydrogen combustion−482−484−2
Ethanol dehydration+41+45−4
Steam reforming of methane+198+206−8
Ethane cracking+123+137−14
Methane combustion−694−802+108
Acetic acid → ketene+41+147−106

Gas-phase reaction enthalpies, kJ/mol. Computed by the shipping engine; measured values from standard thermochemical tables.

Both large errors are the same error. Carbon dioxide's O=C=O is delocalized over both oxygens, so each of its C=O bonds is worth about 799 kJ/mol — but the engine carries one generic C=O at 745, the right value for an ordinary ketone. Twice the 54 kJ/mol difference is the 108 the combustion is short by. The carboxyl group of acetic acid is stabilized the same way, which is why tearing it apart looks too easy.

We built the correction. It put methane combustion on −802.0 exactly, and carbon-monoxide combustion — a reaction that had no part in deriving it — landed within 12. Then we took it out again. Making CO₂ more stable makes every reaction that consumes CO₂ harder, and the products never got their matching credit: urea synthesis, which is exothermic in reality, flipped to endothermic and became unreachable at every temperature and pressure. The carbonyl terms are calibrated against one another, so moving one alone is not an improvement — it is a trade. Fixing it properly means re-fitting the whole set against carbon dioxide, carboxyl, ester, amide and urea together, and that is a piece of work, not a patch.

What we did instead: pinned both errors at the size they actually are, in the test suite, so they cannot drift and cannot be quietly overstated. If someone later closes them, those tests fail — and whoever fixed it has to come and update this page.
Entropy, derived rather than looked up

Temperature only does something if the engine knows each molecule's entropy, and ours knew 25 of them. Every reaction touching anything else silently ignored the temperature dial — which for most organic chemistry meant the dial did nothing at all. The obvious fix was to type in more values from a table. We did not want a table.

Instead the engine now computes standard entropy from three standard results, using only things it already derives: the molecule's mass, the moments of inertia of its own computed geometry, and its own vibrational frequencies. Checked against the measured values it was meant to replace, the mean error is 2.0 J/mol·K — and for a rigid molecule it is essentially exact.

The limit is just as specific. Every freely turning bond costs about 10 J/mol·K of accuracy, because a hindered internal rotation carries more entropy than the stiff vibration this treats it as. Water, carbon dioxide, methane, benzene — no rotors, no error. Butane, with three, runs short. The estimate is always low, never high, and the tests say so.

What it bought is the kind of thing that cannot be faked: ammonium chloride now has a temperature. It forms in the cold, the way it does when you hold ammonia and hydrochloric acid near each other, and it comes apart when heated, the way it really does. Before, it formed at every temperature you asked for, because the engine could not price both sides of the reaction.

And where the physics is only right in one region

The nuclear model is a liquid drop with measured values pinned over it. We measured where the smooth part is trustworthy: for heavy nuclei it is excellent — lead to within 0.05%, uranium to 0.36%. For light and mid-mass nuclei it over-binds, and the consequence is blunt: its binding curve peaks at argon, not iron. The iron peak is the single most famous fact in nuclear binding, and this model does not reproduce it.

That does not reach you, and we checked rather than assumed. Every place the app uses binding energy is either in the heavy region, where the model is accurate, or explicitly gated to it — the α-decay filter computes only from bismuth upward and consults measured values below. We verified the gate across every element from 1 to 100: everything that should decay does, nothing that shouldn't is claimed. The limitation is real, bounded, and walled off.

233
tests across 39 suites, run before every release
2.0
J/mol·K — mean error of derived entropies against measurement
2
enthalpy errors, both pinned at their exact size rather than rounded away
Why publish the misses this precisely? Because “validated” with no number attached is the easiest claim in software to make and the hardest to check. A model that tells you it is 108 kJ/mol short on combustion, and why, is one you can reason with. A model that only tells you it passed is not.
Determinism is not convergence

The engine gave the same answer every time. It was still the wrong kind of answer.

Everything on this page rests on determinism: identical inputs, identical outputs, on any machine, forever. That property held. What we had not checked is subtler, and in August 2026 it turned out to be false — a result can be perfectly reproducible and still not be a result.

Molecular geometry here is found by relaxation: nudge the atoms downhill, repeatedly, until they stop moving. The step size decays as the run proceeds, so the structure settles. Except ours had a floor. Below a certain size the steps stopped shrinking, and after roughly 458 iterations the molecule was no longer settling — it was walking, at a fixed stride, indefinitely. Deterministically. The same walk, the same stopping point, every single run.

So the geometry the engine reported was not the shape the molecule wanted. It was wherever the walk happened to be standing when the loop ran out. 18-crown-6’s size wandered between 3.03 and 3.30 Å depending only on how long we let it run — up, then down, then up, with no sign of converging on anything.

The fix was one line: let the schedule cool all the way. The same molecule now reads 3.0916 Å at 1,200 iterations, and 3.0916 Å at 24,000. Across a corpus of 32 molecules, from water to 21-crown-7, every one settles.

Then the fix took something away from us. One test asserts that hydrazine comes out gauche, its N–H bonds twisted about 91° apart — which is what experiment says, and which the engine had been getting right. With the schedule fixed we could finally ask whether that answer was converged. It was not. At 600 iterations hydrazine reads 75°, gauche. At 900 it flips to 180°, anti, and stays there however long we run. The engine had been right only because it stopped one step before it changed its mind.

Converged, this engine prefers anti, and disagrees with the measurement. That is not a new disagreement — it is the same one, arrived at twice. Our from-scratch Hartree–Fock work had already concluded the engine prefers anti for hydrazine through five successive basis rungs and a relaxed torsion scan, and that closing the gap needs correlation the method does not have. The geometry layer looked like it disagreed with that. It did not. It had simply never converged.

Two independent layers now agree with each other and both disagree with the experiment. That is a worse-looking result than the one we had yesterday, and a far more honest one: we can no longer count hydrazine as something the engine gets right.

Why this section exists. A benchmark score would never have caught this. Neither would the reproducibility tests — the engine passed those the entire time, because it was faithfully reproducing an arbitrary point on a trajectory. The only thing that catches it is running the same molecule at 600 steps and at 24,000 and demanding the two agree. That check now runs against 32 molecules and its results are pinned, so if the geometry ever stops converging again, something fails.
The misses — and what became of them

Two misses. Two different fixes.

The benchmark's two organic misses were both eliminations — a stable molecule shedding a small one: ethanol → ethylene + water (dehydration), and ethanol → acetaldehyde + hydrogen (dehydrogenation). Both are uphill in bond energy, and the engine's rule was a bond-energy rule, so it abstained. They looked like one problem — a missing entropy term (−TΔS), the driving force of a small gas molecule escaping at temperature. They turned out to be two.

Fixed — the free-energy layer

Aufbau now judges reactions by free energy, ΔG = ΔH − T·ΔS, not bond enthalpy alone. Above a crossover temperature the entropy of the escaping water outweighs the uphill bonds, and the engine derives ethanol → ethylene + water on its own — no template, no lookup. The same term makes ammonia synthesis reverse when heated — the real Haber trade-off, which no bond-energy rule can produce — and it leaves every negative control inert at every temperature. Dehydration is now correct; the benchmark moved to 48 of 49 at the time.

Crossed — catalysis, derived from the configuration

The second miss was a different kind of problem, and a more interesting one. Dehydrogenation starts from the same reactant as dehydration — ethanol — but wants a different product, acetaldehyde. A deterministic engine gives one answer per input, and dehydration is the more favorable one; acetaldehyde is the catalytic product a copper surface selects by lowering one path's barrier — kinetic control, not thermodynamics. So it needs a catalyst in the input — and the honest thing was to derive what the catalyst does, not look it up. It turns out you can: a metal's catalytic character is written in its electron configuration. Copper's filled d¹⁰ shell binds weakly enough to abstract hydrogen without cracking carbon–carbon bonds — so over copper, ethanol gives acetaldehyde, and the engine now derives exactly that. With the catalyst supplied, the benchmark reached 49 of 49 — the two identical-reactant rows resolve because they are no longer the same input.

Those two figures are where this stood when the fixes landed. The benchmark reads 49 of 49 today. It got there by giving esterification the driver it actually needs — taking the water off as it forms — rather than by loosening anything: the reaction is thermoneutral to the decimal, so no threshold could ever have admitted it. It went down before it went up: the scoring got stricter three times — a declined reaction was being counted as a pass, and two rows were being scored on a product formula that cannot tell isomers apart — taking 49 to 46. All three have since been won back properly — one by giving the engine a mechanism instead of a better number, one by letting a base license a bond the engine had been forbidden to make, and the last by supplying a driver rather than relaxing a threshold. The stricter scorer was never reversed; the engine climbed back through it. The story above is how the two eliminations were fixed; the story at the top of this page is what the stricter scorer found, and what closing its findings actually took.

How the boundary was crossed: not by fudging a product, but by computing the catalyst from first principles — copper's d¹⁰ configuration makes acetaldehyde, iron's d⁶ makes ammonia, neither from a lookup. The full derivation is the field note below. And a curated benchmark is a floor, not a ceiling: the honest limits are exactly where they were — the model gives trends, not rates, and functionally dense organic chemistry is still the frontier.
Field note — a case study

Catalysis, written in the periodic table.

Ask a chemist why copper turns ethanol into acetaldehyde but an acid turns it into ethylene, and the honest answer is a table of facts: this metal for that reaction. We wanted the other kind of answer — the one where the behaviour follows from something. It does. A transition metal's catalytic character is encoded in its outer electrons, and Aufbau already computes those for every element. So the catalyst isn't a lookup; it's derived.

Why copper

Copper is [Ar] 3d¹⁰ 4s¹ — a filled d-shell. That single fact is the whole story. A full d-band sits low and binds adsorbates weakly, so copper grips a molecule just hard enough to pull a hydrogen off (dehydrogenation) but too weakly to tear a carbon–carbon bond apart (cracking). Run the chain — d-band filling → band centre ε_d = W(½−f) → binding strength → the barrier to break each bond — and a single number falls out: how much cracking is harder than hydrogen abstraction. It is largest for the filled-d¹⁰ coinage metals. Copper is the gentlest catalyst on the row, and the math says so, from its configuration alone.

Metald-configDerived archetypeWhat it really does
Ti / Vd²–d³oxophilicZiegler–Natta, oxidation
Fe / Co / Rud⁶–d⁷dissociativeammonia synthesis, Fischer–Tropsch
Ni / Pd / Ptd⁸hydrogenationhydrogenation, reforming
Cu / Ag / Aud¹⁰s¹gentleselective dehydrogenation, epoxidation
Znd¹⁰s²Lewis acida promoter — not a surface catalyst

Most of the real periodic table of catalysis falls out of the outer electrons alone — iron's half-open d-shell dissociates nitrogen (the Haber process); copper's closed d¹⁰ selects gently; zinc's closed shell with a filled s does neither. Reproducible from the shipping build: aufbau-engine --catalyst <Z>.

Copper · d¹⁰

selects a product

Over copper, ethanol has two downhill exits — ethylene (favoured) and acetaldehyde. Copper lowers the barrier of the dehydrogenation path, so acetaldehyde forms faster even though ethylene is more stable: kinetic control. The engine now returns acetaldehyde over Cu, ethylene over none.

Iron · d⁶

enables a reaction

Nitrogen and hydrogen in a bottle just sit there — the N≡N triple bond is a wall. Iron's open d-shell dissociates it, and only then does ammonia form. Aufbau now leaves N₂+H₂ inert without a catalyst and makes ammonia with iron — honest about a reaction that genuinely needs one.

The configuration had to be right first. Plain textbook filling gives copper 3d⁹ — and a d⁹ copper doesn't stand out. The real 3d¹⁰ comes from an exchange-energy stabilisation of the filled shell, which we added to the engine (it also correctly promotes chromium and gold, yet leaves tungsten alone — the relativistic 6s cost the naive rule misses). The catalysis literally depends on the anomaly, so the anomaly had to be derived too.
The honest ceiling: one atom's configuration predicts a catalytic archetype and a selectivity trend — not a rate, and not the effects of surface structure, supports, or particle size. It is an interpretable model, not a research-grade quantitative predictor, and it says so. A few deeper configuration anomalies (palladium, platinum) need real orbital energies the filling rule can't see, and are left honestly unfixed.
Field note — a case study

Two engines, asked where matter ends.

The same honesty that runs through the reaction benchmark applies far outside chemistry. Aufbau carries two independent first-principles engines that never share a table: an electron engine (the Madelung n+ℓ rule that builds any element's electron configuration) and a nuclear engine (a semi-empirical mass formula that computes nuclear binding energy). We asked both the same question — where does matter end? — and they disagree. The disagreement is the result: each engine's answer maps precisely onto the physics the other one can see and it cannot.

Engine 1 · electrons

“There is no edge.”

Swept out to Z = 20,000 the filling rule never falters — it reproduces the predicted 8th period exactly (the 5g block opens at element 121), then keeps opening orbitals endlessly. The rule has no natural terminus. Reality ends the table by physics this engine can't see — relativity (near Z ≈ 137 the 1s electron nears light-speed) and vacuum breakdown (near Z ≈ 173). Its endlessness is a clean map of exactly where relativity takes over.

Engine 2 · nuclei

“The edge is much earlier.”

The nuclear engine draws the real shape of matter and reproduces a landmark it was never given: α-instability switching on exactly at Z = 82 — lead, the real last α-stable element. It also produces the binding-energy curve's broad mid-mass maximum — why fusion stops and fission begins — though the liquid-drop formula puts that maximum near A ≈ 48 rather than at iron, which is the kind of miss a smooth formula makes and a shell model would not. The true edge of matter is nuclear, near Z ≈ 120–126 — well below where the electron rule even gets interesting.

A note on the notation
Z
proton number — how many protons a nucleus has. It alone decides which element you have: Z = 82 is lead, Z = 114 is flerovium.
N
neutron number — how many neutrons share the nucleus with those protons.
A
mass number — the total count of protons and neutrons together, A = Z + N. “Flerovium-298” means A = 298 (Z 114 + N 184).
α-decay
alpha decay — a nucleus shedding a tight helium-4 cluster (2 protons + 2 neutrons). The energy this releases is the α-decay Q plotted below; lower Q means the nucleus is more stable against it.
magic numbers
special values of Z or N — 2, 8, 20, 28, 50, 82, 126 … — where a nuclear shell is exactly filled, giving extra stability. They are the nucleus's echo of the noble-gas electron shells.
The experiment — a falsifiable island

Sweep the shipped engine across the superheavy region and read off the energy that drives α-decay — lower means more stable — and something tantalizing appears: a dip near Z ≈ 122–124, a pocket of unusual stability against the steep climb toward the fission cliff.

Alpha-decay energy versus proton number across the superheavy region The energy driving alpha decay rises from about 10.8 MeV at element 110 to a peak near element 118, dips to a local minimum around elements 122 to 124, then climbs steeply toward element 128. Lower energy means greater stability, so the dip marks a pocket of unusual stability — the shadow of the island of stability. 11 12 13 14 15 the island's shadow a dip near Z ≈ 122–124 rising toward the fission cliff → 110 114 118 122 126 proton number Z α-decay energy · MeV — lower = more stable
α-decay energy of the most-bound isotope of each element, computed by the shipped engine. The gold points (Z 122–124) sit in a stability pocket; the line then rockets upward as Coulomb repulsion wins. But is that dip the real island — or an artifact? That's the question the experiment below answers.

Nuclear theory predicts an island of stability: superheavy nuclei made unusually long-lived by shell closures, centered near Z = 114, N = 184 (mass 298 — the doubly-magic “flerovium-298”). The shipped engine applies one list of magic numbers to both protons and neutrons — 2, 8, 20, 28, 50, 82, 126 — so above N = 126 it has no neutron shell closure at all. The dip in the chart above is really the lone proton closure near 126 acting by itself — a shadow, not the real, neutron-anchored island. That yields a sharp, falsifiable prediction: because the model lacks the N = 184 closure it cannot place the island correctly — but supply the shell model's predicted closures (proton 114, neutron 184) and the island should snap to the textbook spot. We ran exactly that, mapping α-decay energy across the nuclear plane before and after, with the shipped engine left untouched:

MeasureShipped engine (one list)+ shell-model closures (Z 114, N 184)Textbook
Island locationnone — valley drifts to the neutron-rich edgeZ ≈ 111–114, N = 182, A ≈ 296Z 114, N 184, A 298 (Fl)
Z = 114 α-decay energy+6.76 MeV+3.91 MeV
Valley trace (best N)runs off the scan edgelocked at N = 182N ≈ 184

Reproducible from the shipping build: aufbau-engine --config <Z> (electron sweep), aufbau-engine --binding <Z> (most-bound isotope, binding per nucleon, α-decay energy), and aufbau-engine --island — the experiment below is a first-class command that drives the same engine with the proton and neutron shell-closure lists separated, printing the before/after directly.

The prediction confirmed. With only those two closures added, the stability valley collapsed from “no localized island / drift to the neutron-rich edge” into a sharp minimum at essentially the textbook location — flerovium-298 territory, with Z = 114's decay energy nearly halved. It lands at N = 182 rather than exactly 184, a couple of neutrons off, because the smooth part of the formula shifts the minimum slightly.
What it does — and doesn't — prove: we fed it 184. The experiment shows the engine's shell-correction machinery correctly translates a shell closure into an island at the right isotope. It does not show the model can derive the magic numbers — those come from the spin-orbit shell model the liquid-drop formula doesn't contain. So the engine can carry the island once told where the shells close; it can't find the closures itself. That's the sharpened signpost — the same virtue as every honest miss in the benchmark: a specific, findable pointer at the physics still to add, not a mystery inside trained weights.
Field note — a case study

A textbook ordering, from overlaps the engine was never shown.

Put six ligands around a transition metal and its five d orbitals stop being equal: two point straight at the ligands and are pushed up, three point between them and are not. The gap is Δo, and ranking ligands by it gives the spectrochemical series — one of the oldest empirical orderings in inorganic chemistry, and for a century a thing you looked up rather than derived.

Aufbau computes it. Six ligands are placed octahedrally at the sum of their covalent radii, the complex is solved, and the ligands are ranked by the splitting that comes out. Nothing in that path is fitted to the answer: the orbital exponents and energies are a published Extended Hückel table, the bond lengths are radii sums from a crystallographic survey, and the series itself is never given to the engine. It is used afterwards, as a check.

6 of 9
ligands in the correct relative order, scandium through manganese
0
parameters fitted to the series
3 of 9
by copper — same model, and a real failure

For scandium, the computed ranking from weak field to strong, against the measured one:

computed   I < Br < S < O < Cl < N < F < C < P
measured   I < Br < S < Cl < F < O < N < P < C

The whole weak-field end is right — I < Br < S < Cl — and so is the strong-field end. Six of the nine sit in the correct order relative to one another, out of nothing but computed overlaps.

What came out

The ordering, and π back-bonding with it.

Carbon monoxide is the classic π acceptor: the metal gives electron density back to an empty orbital on the ligand. Nothing was added to the model for this. CO's accepting orbital is simply one of CO's own molecular orbitals, so it appears the moment a real CO is put in and the overlaps are computed. A bare atom reports exactly zero — it has no such orbital, and the engine says so rather than rounding. Back-bonding also switches off at the start of the row, where scandium is d¹ and has no electrons to give away, which is why scandium carbonyls are not a thing.

What did not

It gets worse across the row.

Scandium scores six; copper scores three. The mechanism is visible — scandium's d level sits above every ligand level, which is the clean donation regime, while copper's sits below most of them and the interaction inverts — but the real spectrochemical series is very nearly the same for every metal. A model whose ordering degrades across one row is wrong in a way we can name and have not yet fixed. Seven of the nine donors are also bare atoms standing in for ions: F rather than F⁻. That is the largest remaining approximation, and the ligands modelled as real molecules rank better for it.

What this cost, and what it taught: four rounds of work went into diagnosing an increasingly elaborate physics defect in this calculation — smeared orbital character, antibonding levels at +83 eV, a manifold that could not be identified. The cause was a unit conversion. The solver takes atomic units and was being handed ångströms, so every complex had been built at 53% of its intended size. Each new instrument built to characterise the anomaly made a wrong geometry look like a new kind of physics. One check of the lowest-level quantity — a two-centre overlap against its exact closed form, which needs no reference data at all — would have caught it four rounds earlier.
A note on the notation
Δo
the octahedral splitting — the energy gap between the two sets of d orbitals in a six-coordinate complex. It sets a transition-metal compound's colour, its magnetism, and much of its reactivity.
t2g, eg
the two sets. t2g is the three orbitals pointing between the ligands; eg is the two pointing straight at them, which is why they are pushed higher.
spectro­chemical series
ligands ordered by the splitting they produce, from weak field (iodide, bromide) to strong field (phosphines, carbon monoxide). Determined experimentally from absorption spectra, and famously almost independent of which metal you use.
π back-bonding
the metal donating electron density back into an empty orbital on the ligand — the reverse of ordinary bonding, and the reason CO and cyanide sit at the strong-field end.
Extended Hückel
a first-principles molecular-orbital method that builds orbitals from atomic ones and their overlaps. Deliberately simple: no fitting to spectra, no trained weights, and every number in it can be traced.
Field note — a case study

The rare earths, separated by design — not by acid.

The rare earths run the modern world — the magnets in wind turbines, electric motors, and guided systems — and they are the hardest bulk separation in chemistry, because the fifteen lanthanides are almost chemically identical. Industry pulls them apart the brute way: roast the ore in acid, then push a nearly-inseparable mixture through hundreds to thousands of countercurrent extraction stages, leaving radioactive tailings and poisoned water behind. We wanted the other kind of answer — where the separation follows from the atom. It does. Aufbau reads which lanthanide can be pulled out cheaply, and how, from electron configuration alone.

Which handle to pull

Most lanthanides are stubbornly trivalent, but three aren't: cerium, europium, and ytterbium have a cheap one-step redox handle — a +4 or +2 ion whose electrons land on an empty, half-filled, or full 4f shell. That "magic-shell" stability is just exchange energy, which Aufbau already counts. Strip those three in one electrochemical step each, and the hard part — the stubborn trivalent remainder — is smaller before the expensive train even starts.

ElementLn³⁺Cheapest leverWhy — from the configuration
Cerium4f¹oxidise → Ce⁴⁺empties to the bare 4f⁰ core
Europium4f⁶reduce → Eu²⁺completes the half-filled 4f⁷ shell
Ytterbium4f¹³reduce → Yb²⁺completes the full 4f¹⁴ shell
Nd, Pr, Dy…4fnchelateno magic-shell redox — split by the contraction

The computed "strong redox" set is exactly the three that industry separates this way — from f-shell exchange counting alone, no table. Reproducible from the shipping build: aufbau-engine --serve (the flowsheet op), or in the app under Tools ▸ Rare-Earth Flowsheet.

Strip the easy ones

Redox, one step each

Cerium is roughly half of a light-rare-earth ore by mass. Oxidise it out; reduce europium and ytterbium out — three cheap electrochemical steps, no acid roast, no ammonia. The mixture that reaches the extraction train is smaller and simpler.

Chelate the rest

The contraction as a guide

The trivalent remainder is split by the tiny lanthanide contraction — the heavier ion is smaller and more charge-dense, so a hard oxygen donor holds it a shade tighter. Aufbau ranks the donors and flags exactly where the split is tightest (for magnet metals, that's Nd/Pr) — the stages you'll need the most of.

One quantity carries it. The metal's "hardness," computed as a charge density from its effective nuclear charge, already encodes the lanthanide contraction — so the same number that ranks hard-versus-soft across the whole periodic table also ranks the rare earths against each other. Hard donors for hard ions, soft for soft; a ligand-field term breaks the tie between transition-metal neighbours like copper and nickel. All of it from the configuration.
The honest ceiling: this designs a flowsheet and ranks the levers — it does not replace the wet chemistry or predict a plant's exact separation factors. Adjacent lanthanides differ by so little that the model reports a tiny split, which is the honest truth: those separations are genuinely hard. It is a first-principles, auditable design tool — a way to see which handle to pull and where the difficulty lives — not a research-grade quantitative predictor, and it says so. The method is patent pending.
Field note — a case study

A battery's fastest ion, found without a supercomputer.

A solid-state battery lives or dies on one number: how easily a lithium ion slips through the crystal that separates the electrodes. Measuring it takes a lab; predicting it usually takes density-functional theory on a compute cluster. Aufbau does neither — and gets the ordering right anyway.

It reads a crystal structure — the same public CIF a crystallographer downloads — and maps the energy a mobile ion would feel at every point in the cell from the bond-valence site energy: a short-range attraction to the oxygens it threads, and a screened repulsion from the framework cations it must squeeze past. The low-energy points form a network, and the migration barrier is the lowest energy at which that network first spans the crystal — the ion's easiest way through. Nothing is fitted to conductivity: the parameters are the published softBV set, and the barrier is read off the map, not trained toward it.

0.3 < 0.6 eV
cubic garnet ranked below NASICON — the correct fast-conductor ordering
0
parameters fitted to measured conductivity
no DFT
a whole crystal screened in seconds, on a laptop

The screening payoff is the ranking. The cubic garnet Li₇La₃Zr₂O₁₂ — the flagship fast solid electrolyte — comes out at a lower barrier than the NASICON LiTi₂(PO₄)₃, which is the ordering every battery chemist knows. It only comes out right when the barrier is measured from the site the ion actually rests in: the garnet's deepest point is a trap the ion never sits in, and referencing to it instead would report the slow, wrong value. Getting that reference right is the difference between a screen that ranks materials and one that misleads.

What came out

The right conductor, and the right charge.

Beyond the ranking, the tool reads a structure honestly. When a crystal file labels its own oxidation states, the engine uses them: iron in LiFePO₄ enters as Fe²⁺ straight from the file rather than a guess — which matters, because the charge sets how hard that cation pushes the lithium past it. And you can watch it: the app renders the framework, the low-energy channel the ion threads, and the ion itself migrating along it — a real crystal you can rotate, zoom into, and read atom by atom.

What did not

It knows what it can't model — and says so.

This is a screen, not DFT: it overestimates absolute barriers, so the numbers rank materials rather than reproduce experiment. And it is fitted to oxide frameworks only. Hand it a sulfide or a fluoride electrolyte and it refuses — it will not score sulfur with an oxygen potential and hand back a confident wrong number. Give it a framework cation it has no parameters for and it says plainly that the barrier will read optimistically low. The refusal is the feature.

What this cost, and what it taught: the first version would take any crystal and print a barrier. Handed a sulfide electrolyte, it returned a clean-looking 0.18 eV — by quietly applying the oxygen potential to sulfur, a well it was never fitted to. It looked like an answer. Closing that hole — making the engine refuse an anion it has no parameters for — is what turned a number generator into a screen you can trust. A second bug surfaced the same way: crystal files write an oxidation state two ways, Fe2+ and Ti+4, and an early reader understood only one, silently turning every Ti+4 into a +1 and collapsing the barrier. A known-good structure regressing from 0.67 to 0.26 eV caught it. The tool is only trustworthy because it now knows the edge of its own competence.
A note on the notation
solid electrolyte
a solid that conducts ions but not electrons — the piece that replaces the flammable liquid in a solid-state battery. How fast it conducts is set by the migration barrier.
migration barrier
the activation energy for a mobile ion to move through the crystal. Lower means a faster conductor; it is the single number that most decides whether an electrolyte is any good.
bond-valence site energy
a deterministic, no-DFT estimate of what a mobile ion feels everywhere in a cell — attraction to the anions it bonds, repulsion from the other cations. Decades old, well-calibrated, and fast enough to sweep thousands of structures.
NASICON, garnet
two families of oxide ion conductor. The garnet Li₇La₃Zr₂O₁₂ is the archetypal fast solid electrolyte; NASICON LiTi₂(PO₄)₃ is a good but slower one. Ranking them correctly is the test a screen has to pass.
softBV
the published bond-valence parameter set the barriers are computed from. Fitted to crystal structures across the periodic table — never to ionic conductivity — so a good barrier is a prediction, not a memory.
Field note — a new capability

Deterministic computer-aided molecular design.

Computer-aided molecular design has two established species. One retrieves: a database of structures and measurements, which can only answer for what somebody already recorded. The other predicts: a model trained on that database, which answers for anything you ask and cannot tell you when it is guessing. Aufbau is neither. Sculpt, new in 1.9, is the first piece of a third kind — a design surface where every answer is computed from the same energy model that builds the molecule, and where asking the same question twice gives the same answer to the last bit.

You pose the shape. The chemistry answers back.

Add a virtual charge group to an atom and the bonds bend away from it, the way a real lone pair or substituent would. State what you need — this angle at 109.5°, these two atoms 2.4 Å apart, that torsion anti — and Sculpt drives the geometry toward it. Then it reports the miss. Where a target fights the chemistry, the chemistry wins and the tool says so, in ångströms and degrees. That is the difference between a design tool and a drawing program: you can ask for anything, and it will tell you what it costs.

The shape rule, with the cause on a dial

The companion tool, Why Molecules Bend, strips the idea to its floor: put N things on a sphere, let them push each other apart, read off where the bonds point. No hybridization, no sp³. Hold four domains and step the lone pairs, and three familiar molecules appear from one calculation —

MoleculeBondingLone pairsComputedMeasured
Methane40109.5°109.5°
Ammonia31106.2°107°
Water22103.5°104.5°

Same sphere, same solve, one number changed. And the one empirical constant in that account — how much harder a lone pair pushes than a bond — is exposed as a slider rather than buried. Set it to 1.00 and water opens back out to a clean tetrahedron. Being able to switch off the cause and watch the effect vanish is a stronger argument than any diagram of it, and it is something the hybridization story cannot offer, because it never named a cause to switch off.

What it will not tell you

Sculpt measures the geometry you pose. It does not predict the shape a ligand prefers — no ideal cavity size, no matched pocket. We tried, six times, against crown ethers whose cavities crystallography already settled, and every attempt failed a test written before it ran. So the tool declines rather than guesses: ask it for the hole in a structure whose donors are not really surrounding a point, or one whose bonds have been stretched to reach a target, and it tells you that instead of quoting a number. An instrument that refuses to read is worth more than one that always answers.

Four decades ago LHASA asked which reactions could build a molecule. This asks what shape it has to have — the other half of the same ambition, on physics the 1980s could not compute. Deterministic throughout: identical input, identical geometry, content-addressable. Reproducible from the shipping build: echo '{"op":"geometry","smiles":"O"}' | aufbau-engine --serve

Continued evolution

The engine keeps growing — in public.

Aufbau didn't arrive finished. Each release widened what the same first-principles engine can derive, and the benchmark above tracks it honestly as it grows.

1.0Foundations

Bonding & structure, from atoms

Molecules build by energy minimization — covalent, ionic, metallic, and coordinate bonds all emergent; aromaticity by the Hückel rule; real 3-D geometry you can rotate.

1.1Reactivity

Reaction prediction from first principles

Enter reactants and the engine computes the products from bond energetics — with an honest verdict, and an honest “no reaction” when nothing favorable exists. No lookup.

1.2Synthesis

Multi-step routes & a new family of chemistry

The step-by-step pathway to a product, each step's mechanism and energy — including urea from captured CO₂ and ammonia, the real industrial route, derived end to end. Plus portable, self-verifying share codes.

1.3Reasoning

Find the missing reactant

Turn the question around: mark a reactant as a wildcard, name your target, and Aufbau reasons backward to the reagent you need — and names it even when it isn't on your shelf, turning “no” into “here's what to get.”

1.3.1Corrected in the open

A bug found and fixed, in public

The engine keeps getting more faithful. We found and fixed a real bug in the computed periodic table: for heavy elements the electron configuration came out wrong — Oganesson (element 118) had a spurious orbital instead of its true [Rn] 5f14 6d10 7s2 7p6. Fixed with a regression test and shipped in 1.3.1. That's the quiet advantage of computing instead of storing: when it's wrong, the mistake is a specific, findable, fixable line — not a mystery inside a trained model.

1.4Free energy

Reactions judged by ΔG, not bonds alone

The gate now weighs entropy as well as bond strength — ΔG = ΔH − T·ΔS — so a reaction driven by a small gas molecule escaping can turn favorable above a crossover temperature. Aufbau derives ethanol → ethylene + water on its own, reproduces the Haber trade-off (ammonia synthesis reversing when heated), and still holds every negative control inert. The benchmark moves to 48 of 49.

1.5Catalysis

The catalyst, from the configuration

The boundary the free-energy layer stopped at — catalytic control — is now crossed, and crossed the honest way: a metal's catalytic character is derived from its electron configuration, not looked up. Copper's filled d¹⁰ shell selects acetaldehyde from ethanol; iron's open d-shell enables ammonia by dissociating N₂ (which otherwise won't react). The benchmark reaches 49 of 49. A real configuration bug — the anomalous d¹⁰ ground states — was fixed on the way. See the field note above.

1.6Stereochemistry

Handedness, assigned from the graph

The engine now sees three-dimensional chirality — R/S at carbon and even at a sulfoxide's sulfur, E/Z across double bonds — derived by the CIP rules from the molecular graph, and shown with the correct handedness in 3-D.

1.7Orbital symmetry

Why the Diels–Alder goes

Pericyclic selectivity — the Woodward–Hoffmann rules — read straight off the frontier molecular orbitals: the classic cycloaddition allowed, the [2+2] forbidden, and both inverting under light, because the orbitals do. The phases are drawn as coloured lobes.

1.8Separation design

Designing a rare-earth separation

The frontier orbitals reach past molecules to metals: which ligand donor grabs which ion (hard–soft matching plus the ligand field), and a full plan for pulling a rare-earth mixture apart — strip the redox-separable ones, chelate the rest. See the field note above. Patent pending.

1.9Molecular design

Sculpt — push back on the geometry

The first piece of deterministic computer-aided molecular design: state the shape you need and the engine reports what the chemistry says it costs. A companion tool derives the shape rule itself from domains on a sphere — methane, ammonia and water from one calculation, no hybridization. And it declines to answer where the geometry cannot support an answer.

The benchmark also went 48 of 49 → 46 of 49 in this release, and the engine did not change. Two rows were being scored on the product formula, which cannot tell isomers apart; scored on structure, both fail — the engine returns butanoic acid where an ester belongs, and 1,1-dichloroethane where the 1,2- belongs. A third row is a correct refusal. The number is lower because the scorer can now see things it could not see before.

1.9Verification

Checked against measurement, and the answer written down

The engine's answers had been scored for a year; its arithmetic never had. So we compared it reaction by reaction against published enthalpies, derived standard entropy from first principles instead of a lookup table, measured where the nuclear model is trustworthy, and confirmed the stereochemical descriptors by hand from the CIP rules. Most of it came out very close. Two enthalpy errors did not, and both are now pinned at their exact size in the test suite rather than rounded away. The pass also caught things the benchmark could not see: products that changed depending on the order you typed the reactants, elements outside the engine's data being quietly substituted rather than declined, and a share code that dropped a molecule's handedness while reporting itself verified.

1.10Base catalysis

The base is not a switch

An aldol makes a carbon–carbon bond, and the engine had always refused to form one — deliberately, because nothing stopped it welding two skeletons together into molecules the reactants cannot reach. So the aldol was never enumerated, not rejected on energy. A Brønsted base is what removes the α-hydrogen, and therefore what makes that bond reachable at all; asserting one licenses the move and nothing else. The same release derived which way an alkene adds — Markovnikov’s rule falling out of the energies rather than being written down.

1.11Retrosynthesis

Working backwards, and only keeping what it can reverse

Ask for a target and the engine disconnects a strategic bond, caps both freed valences so the pieces are real molecules rather than radicals, and keeps the cut only if the forward model rebuilds the target from them. There is no transform library: nothing in the file knows what an ester is, and the three caps — H, OH, Cl — are simply the ways a dangling valence can be satisfied. Which one applies is the forward engine’s answer, not a rule. Aspirin comes apart into salicylic acid and acetic acid; Vitamin C opens its own ring.

1.12Teaching & protecting groups

A curriculum, and the group that gets in the way on purpose

Lessons that derive rather than assert: molecular shape from points repelling on a sphere — methane, ammonia and water from one calculation, with hybridization never mentioned — and lessons that work through why a mixture will not simply react, and how to plan a route backwards. Everything the lesson claims is computed live by the same engine that answers everywhere else. No language model is involved.

Alongside them, the piece of classical synthesis planning that has been missing: protecting groups. The engine now notices, on its own, that a step it wants would attack the wrong site — the fragments do react, they simply build the wrong molecule — and answers by masking the competing site, coupling, and taking the mask off again. It also declines a protecting group whose removal would destroy the product, and reaches for a second one that survives. That decision is made by running the removal and reading the result, not by a table of which group suits which case.

Underneath both, a correction to how bonds are priced. A bond is not worth the same everywhere: an ester oxygen is not an ether oxygen, and a benzylic bond is not an ordinary one. Two long-recorded errors in the engine’s reaction enthalpies were halved by saying so once, and a defect nobody was hunting fell out with them — chlorinating benzene had been destroying the aromatic ring instead of substituting into it.

1.13Drawing, and what it caught

A screenshot found what the tests did not

The blue clouds above and below an aromatic ring are what make it read as benzene rather than a hexagon of lines. They existed in one view. They now appear wherever a molecule is drawn in three dimensions — including Sculpt, where the shape is worked out from scratch and a ring can end up at any angle, so each cloud is fitted to its ring’s own plane rather than assuming the ring lies flat.

Switching them on everywhere is how porphine’s drawing was found to be wrong: four huge overlapping annuli laid across the molecule, two small clouds, and two rings with nothing on them at all. The renderer had been handed six rings — sizes 5, 5, 16, 17, 17, 18 — four of which were one delocalised circuit counted four different ways round. The filter dropped a ring only when a smaller ring was a subset of its atoms, which is true of naphthalene and false the moment a perimeter takes a shortcut: porphine’s macrocycle threads each pyrrole by its nitrogen and two α-carbons and skips the β-carbons, so no pyrrole was a subset and nothing was dropped. The right question is not whether a ring sits inside another but whether it says anything the smaller ones did not — cycle independence — and asked that way porphine gives [5, 5, 16]: one annulus threading all four nitrogens, plus the two pyrroles that are aromatic on their own. Naphthalene’s perimeter still drops, now for the better reason.

Porphine before the fix: four large overlapping rings drawn across the whole molecule, two small clouds, and two rings left bare.before — six rings, [5, 5, 16, 17, 17, 18]
Porphine after the fix: one annulus threading the inner macrocycle through all four nitrogens, plus clouds on the two aromatic pyrroles.after — three rings, [5, 5, 16]
Porphine as the engine draws it, before and after. Both frames come from the shipped renderer on the same molecule; only the choice of which rings to draw changed. The two bare pyrroles on the right are correct, not a remaining fault — porphine carries two NH pyrroles that are aromatic on their own and two pyrrolenine rings whose aromaticity lives in the macrocyclic circuit.

Worth being exact about what was broken, because it would be easy to claim more: the chemistry never moved. The reactor reads the full circuit list for its delocalisation energy and always did. Only the set chosen for drawing was wrong — and it survived a month because that filter had no test of its own, while the function it filters had plenty. It was caught while capturing a screenshot for the App Store, which is a poor substitute for a test and was, that day, the only thing looking.

A smaller change that matters more than it sounds: a route now names what it wants. The leaves of a retrosynthesis were bare formulas — C7H6O3 — which is not an answer you can take to a supplier. They now carry the reagent’s name beside its formula, and a fragment you can simply buy no longer has that fact hidden behind a badge telling you it is aromatic.

1.14The window it runs in

Six things that had shipped, and stayed shipped

Aufbau’s windows could not be resized. Not by dragging a corner, not by the Window menu’s own Fill command, not by the accessibility API — every window opened at 924 by 694 points and stayed there, whatever display it was on. Mac Catalyst caps each window through UIWindowScene.sizeRestrictions, and the lift for it hung off a zero-sized view whose appearance callback never fired, on a screen it was never attached to. It is lifted at the scene level now, so it depends on no view being drawn.

The same ceiling was quietly shrinking every panel that opens on top of a window. The molecule viewer asked for a thousand points of width and was given about five hundred and fifty, so porphine — thirty-eight atoms — was drawn into a fifth of the window it sat in. On a 1080p display the Reactor was worse: its spontaneity and kinetics readings sat below the bottom edge and could not be scrolled to. Both now use the window they are in.

Then a crash, which is the one that mattered. Asking for a route to a target made of more than one piece — an ionic reagent, a typed mixture, a saved reaction recalled from the library — could quit the app outright. The planner asked whether the molecule had come apart in two when the question it needed was whether the bond it had just broken was the thing that separated them. For a connected molecule those are the same question. For a mixture with a ring in it they are not, and it reached for an atom that was not there.

One computed result moved. Building a porphyrin from pyrrole and formaldehyde ended in an oxidation that took the ring too far, pulling hydrogen from all four nitrogens to give C20H12N4 — where porphine is C20H14N4 and keeps two N–H. The macrocycle was always right; only the last step overshot. Oxidation now stops when the ring becomes aromatic, which is what drives it, rather than continuing while a wider circuit can still be traced. This page already said porphine carries two NH pyrroles: the site was right and the engine was wrong.

Two smaller corrections, both about not overclaiming. A finished route used to hand back HCl where the reagent shelf beside it says hydrochloric acid — two catalogues of the same idea that had never been introduced, and which disagree on notation, the shelf writing H2SO4 where the engine writes H2O4S. And a leaf the planner merely stopped at because it was small no longer calls itself a building block: hydrogen thioperoxide is a transient, not something to order.

None of this was found by the test suite, which passed throughout. Every one of them was found by looking — at a window that would not grow, at a route that named nothing, at a crash report. The checks that should have caught them existed and were not wired to a verdict: the release script could not fail on a red suite, and the porphyrin synthesis was demonstrated by a debug-only routine that merely printed. Both now report. That is the more useful half of this release.

1.15Ligand field

A textbook ordering, computed — and three bugs found on the way to it

Put six ligands around a transition metal and its five d orbitals stop being equal. The size of the gap between them decides a compound’s colour, its magnetism and much of its reactivity, and ranking ligands by it gives the spectrochemical series — one of the oldest orderings in inorganic chemistry, and for a century something you looked up. Aufbau now computes it: Tools ▸ Ligand Field builds six octahedral complexes, places each ligand at the sum of the covalent radii, solves them, and ranks the ligands by the splitting that falls out. For the early transition metals six of the nine come out in the correct order, including the whole weak-field end — with nothing in that path fitted to the answer, and the series never given to the engine.

Copper manages three, and the screen says so beside the six. Scandium’s d level sits above every ligand level, which the model handles; copper’s sits below most of them and the interaction inverts. The real series is very nearly the same for every metal, so that drop across one row is a genuine limitation rather than a caveat, and it is displayed rather than described. Back-bonding falls out of the same calculation: carbon monoxide takes electron density back into an empty orbital of its own, and nothing was added to the model to make it — that orbital is simply one of CO’s. A bare atom, which has none, reports exactly zero.

Three corrections underneath reach far past this screen. Valence orbitals for the heavier elements had been built with the shape belonging to the second row, leaving iodine’s outer shell about four times too compact — every element from sodium down was affected. The transition metals’ orbital energies were derived where a measured table exists, with roughly two and a half times the correct slope across the row. And six of the nine first-row metals had no covalent radius at all, so they could not be placed in a structure. Before any of that, four rounds of work went into diagnosing an elaborate physics defect — smeared orbital character, antibonding levels at +83 eV — that turned out to be a unit conversion: the solver takes atomic units and was being handed ångströms, so every complex had been built at 53% of its intended size. One check of a two-centre overlap against its exact closed form, which needs no reference data at all, would have caught it four rounds earlier. The benchmark is unchanged at 49 of 49.

NextIn development

Wider chemistry, same honesty

The catalyst archetypes reach further than the two effects wired so far — hydrogenation and oxidation are characterised and waiting. Beyond that: broader, functionally dense organic chemistry through principled corrections, the deeper configuration anomalies that need real orbital energies, and independent large-scale testing — measured on reproducibility and fabrication rate, the axes that matter.

What's next

Honest limits, and where we're headed.

Aufbau is an interpretable model — it makes mechanism visible and derivable — not a research-grade quantitative predictor, and it says so. It's strongest on small-molecule, main-group, and inorganic reasoning; functionally dense, selectivity-driven organic chemistry is the frontier, and where expert chemical knowledge remains irreplaceable. The free-energy layer closed the entropy-driven gap, and config-derived catalysis just crossed the boundary it stopped at — copper selects, iron enables, both from the periodic table. What's ahead: wiring the catalyst archetypes that are characterised but not yet gate-active (hydrogenation, oxidation), broader organic coverage through principled corrections, and independent large-scale testing against database and machine-learned baselines, measured on reproducibility and fabrication rate — the axes that matter for trustworthy tools.

A preprint is in preparation describing the method, the emergent results, and this validation in full. We'll link it here when it posts.

Right when it can be. Honest when it can't.