In reaction or material optimisation, the hard part is not proposing many possibilities but deciding which expensive option to test next. A language model alone can suggest plausible-sounding yet invalid conditions. GOLLuM therefore uses a language model to represent an experiment's description and a Gaussian process to estimate both outcome and uncertainty.

The probabilistic objective also reshapes the language-model representations so experiments with similar outcomes move closer together. Selection is therefore not based on textual resemblance alone and can balance exploiting promising regions with exploring uncertain ones.

The authors ran the same configuration across 23 benchmark tasks in organic synthesis, materials, process chemistry and molecular design, starting each from ten below-median observations. GOLLuM ranked first on average and matched standard Bayesian optimisation with over 40% fewer trials. Across five Buchwald–Hartwig tasks it found 43% high-performing conditions, compared with 24–25% for the approaches used as comparators.

One possible use is controlled selection of the next measurement in self-driving laboratories, with separate safety constraints and human oversight. The evidence is still benchmark-based; Gaussian processes also scale poorly to thousands of trials, and text alone is insufficient for complex 3D structures. Prospective laboratory pilots could appear within 2–4 years, while routine use will require independent validation.