Google released EmbeddingGemma 2 on October 6 with open weights under Apache 2.0. It maps different types of content into a shared numerical space for similarity search. The full version has 740 million parameters; audio and vision components can be omitted when unnecessary. It outputs representations, not finished answers.

Imagine a sound engineer who remembers a noise's character but not its filename. Or a family archive where a brief moment is hidden among hundreds of recordings. If retrieval worked on our material, a question could yield a few concrete matches. The answer would be playable evidence, not a newly invented scene.

Local processing is important, but it does not by itself guarantee the privacy of an entire application. Synchronization, logs, backups and surrounding services matter too. A similar file is not necessarily the correct one. Users of professional or personal archives should see the source and decide whether the match really fits.

Before deployment, we would test our own Czech queries, quiet speech and easily confused scenes. The model card warns that quality differs across languages and that aggressive representation truncation sacrifices some quality. Testing must cover missed items and false matches as well as speed, power consumption and initial indexing time.

Our optimistic editorial estimate is 2–4 weeks for a small local prototype with a limited archive, suitable hardware and an experienced integration team. This is not Google's promise. A reliable application across phones and long recordings needs further validation; a model does not replace a complete media archive system.