H company introduced NeoMME, a family of 260M and 800M parameter models. The technical paper was submitted to arXiv on August 31, 2026, followed by an announcement on September 3. It is an unreviewed preprint.

A shared bidirectional Transformer handles text and images. The Retriever variant searches pages as images, retaining layout information. The weights are released under Apache 2.0; the base model alone is not a finished search system.

We see potential in technical archives, where a drawing and its nearby caption can jointly indicate which page answers a query. A person or downstream system should still verify the retrieved page.

Deployment requires testing on the actual languages, scans and tables involved. A smaller search index does not mean smaller original documents or guarantee a correct answer.

Optimistic editorial estimate: a small pilot could take 1–3 months. Reliable operation across a large archive depends on data quality, hardware and access controls.