On September 6, OpenAI published internal measurements of agent use in research. It reports 3.1 eight-hour agent-workdays for each day of human labour.
This ratio measures tool runtime, not a multiplier of scientific productivity. The company recorded more experiments while its available computing resources also grew.
An accompanying essay by Jakub Pachocki emphasizes the unresolved challenge of controlling increasingly capable systems. This is the research leader’s position, not an independent forecast.
We see potential for such measurements to reveal laboratory bottlenecks: where agents save time and where they add review work. Meaningful comparisons must include result quality, costs and error correction.
Editorial estimate: teams could establish their own measured pilots within 1–6 months. These data cannot establish a date for a reliable autonomous scientist.

Be the first to open the discussion.