Set the boundary for multimodal search
EmbeddingGemma 2 brings text, images, video and audio into one local search space. The architecture decision is which content and derived vectors may stay on a device, and which may enter a shared index.
Trace the interfaces between applications and records before changing a system.
On 6 October, Google released EmbeddingGemma 2, an open-weight model that maps text, code, images, video and audio into one 768-dimensional space. A text query can therefore be compared directly with an image or video representation, without first turning every asset into a caption or transcript. Google’s model card describes a 270-million-parameter text-and-code base with optional vision and audio encoders, for 740 million parameters in total. (Google DeepMind launch and model card)
That changes what a search system can do at the device boundary. It does not decide where the boundary belongs. A team can now test local search across mixed media, but still has to choose which source files, queries, vectors and result metadata stay on a device and which enter shared infrastructure.
A common space removes a conversion step
Earlier multimedia search pipelines often converted images into captions and audio into transcripts, then embedded that text. Google’s edge implementation demonstrates text-to-image search and finding moments in video without first transcribing audio or generating intermediate captions. The model instead gives the query and media items comparable vectors, while specialised encoders process the different modalities. This is a reported product demonstration, not independent evidence that every media-search task can drop its existing conversion or indexing steps. (Google AI Edge implementation)
The footprint makes a local trial plausible on some devices. Google reports about 191 MB of active RAM for text-only weights and 567 MB for the full multimodal model on a Pixel 11 Pro. Those are provider measurements for one named device, not a specification for other phones, laptops or embedded hardware. The model card reports benchmark scores on the full-precision checkpoint, while Google’s on-device figures refer to quantised weights. Neither tells an organisation how quickly its own media library can be indexed, how much battery that will use, or whether the top results are useful. (Google AI Edge and model card evaluation tables)
The shared 768-dimensional output is also not a reason to discard retrieval configuration. Google’s guide specifies different text prefixes for query and document sides of asymmetric search, while images, video and audio are passed without those text prefixes. The model card also says truncated vectors must be re-normalised before cosine similarity and that query and corpus vectors must use the same dimension. These settings belong with the index version, alongside the model and encoder configuration. (Google model card, task instructions and truncation guidance)
Decide which search path crosses the device boundary
Consider a hypothetical maintenance team that wants to find a faulty pump by typing a symptom, selecting a reference image or searching video recorded during a repair. If technicians need results while disconnected and the relevant material is already authorised on their managed tablets, local embeddings and a local vector store could keep the search path on those devices. If the aim is to search every depot’s recordings from a central service, a shared index may be a better fit, but the media or derived representations must cross into infrastructure with the right access controls. A unified embedding space does not choose between those designs.
| Search boundary |
|---|
| Local. Query and vectors stay on the managed device during search. Decide which assets are available offline and how access, device loss, replacement and user departure affect the local index. |
| Shared. Media ingestion and query processing use a service-side index. Preserve source permissions in retrieval, define tenant boundaries and retention, and record which model and index version returned a match. |
| Split. The device searches its local and an authorised shared collection. Decide which metadata or vectors synchronise, how revocation reaches both stores, and how the separate result sets are ranked. |
This comparison is an architectural design aid, not a claim that EmbeddingGemma 2 supplies these controls. In particular, local inference does not make a whole application local if it uploads media for backup, synchronises vectors, sends telemetry or asks a hosted model to explain retrieved items. Map each outbound path, not only the encoder call. Treat vectors and their identifying metadata according to the sensitivity and permissions of the source material they represent.
Measure the workload before choosing the footprint
Start with one collection and a set of representative questions that require crossing modalities: text to image, text to a video moment, image to similar images, and audio to a related document. Include near misses and the languages, labels and terminology people actually use. Compare the proposed local route with the current search method using a reviewed relevance set. Measure useful results at the chosen rank, index-build time, query latency on the target device, peak memory, battery use and behaviour without a network. Do not infer production quality from the showcase.
Vector truncation offers another trade-off to test. Google’s model card reports that reducing from 768 to 256 dimensions cuts vector size to one third while its reported MMEB v2 overall score falls from 59.01 to 56.24. At 128 dimensions the score is 45.65. These are Google’s benchmark results, not an enterprise corpus test, and the overall multimodal score does not show which media type or query fails. Select a dimension against the actual quality and storage requirements, then re-normalise as the model card instructs. (EmbeddingGemma 2 model card, truncation results)
This is distinct from an embedding-model migration, which also needs compatible query and corpus vectors, a rebuild plan and a controlled index cutover. The site’s earlier migration analysis covers that work. The new choice comes first: which data boundary should this search service respect, and can the device actually meet its retrieval, access and operating requirements there?
The product owner and data owner should choose one use case and one authorised collection for a bounded local trial. Record what stays on the device, what is synchronised, and the retrieval and device measures that would justify expansion. The service owner should not widen that collection until the data owner has approved every path by which its media, vectors and results leave the device.