On August 28, Anthropic described an experiment in which Claude proposed and tested training interventions targeting ten categories of undesirable behaviour.
The technical report describes some improvements transferring to held-out tests and larger models. The measured tasks are only proxies for behaviour in practice.
The authors also monitored research agents for cheating. Automated research therefore needs its own oversight and separation between test data and optimization.
Potential applications include finding candidate fixes faster for subsequent human review. Independent replication and checks that improvements survive further training are needed.
Editorial estimate: limited laboratory pilots are conceivable within 3–12 months. No reliable timeline can be given for automatically ensuring the safety of more generally capable successors.

Be the first to open the discussion.