THE SIGNAL IN ONE SENTENCE

An embedding model is not a chatbot. It does not write the answer. It builds the shelves. Give an embedding model a document, photo, sound clip or video frame and it turns that item into a list of numbers. Items with similar meaning end up near one another in that numerical space. A search system can then take a question such as "Where is the clip with the squeaky bicycle wheel?" and retrieve the matching video or audio without requiring the query and result to share the same words. Google DeepMind's new EmbeddingGemma 2 puts text, code, images, video and audio into one shared 768-dimensional space. The whole model has 740 million parameters and is designed to run on phones and laptops. That combination matters. Multimodal search has usually asked developers to stitch together separate encoders, send private files to a server or accept a model too large for ordinary consumer hardware. A smaller shared model makes a different architecture possible: index the files where they already live, search them locally and send only the retrieved material into any later reasoning step. In plain English, your phone may be able to search a voice memo for the video it describes, find a function in a codebase from a natural-language question, or locate a scanned receipt from a spoken description without uploading the entire collection. The useful trick is that 740 million is a ceiling, not a fixed cover charge. The text backbone and embedder total 270 million parameters. Developers can add the 170-million-parameter vision encoder for a 440-million-parameter text-and-image system, or the 300-million-parameter audio encoder for a 570-million-parameter text-and-audio system. Loading all three reaches 740 million. That modularity lets an app carry only the senses it needs. A local code-search tool does not need an audio encoder loitering in memory. A field recorder that searches spoken interviews can skip vision. A personal media archive can load the full set and compare a typed query directly with pictures, video frames and sounds. Google also trained the output vectors so developers can shorten them. The native vector has 768 dimensions. It can be truncated to 512, 256 or 128 dimensions, then normalized again. Shorter vectors take less storage and can make similarity search faster. Google reports up to a sixfold reduction in vector storage at 128 dimensions. Its own benchmark table also shows the tradeoff: 128 dimensions loses considerably more multimodal quality than 256 dimensions. Smaller is not free. It is an engineering dial. The same is true of memory. On a Pixel 11 Pro with quantization, Google reports about 191 megabytes of active RAM for text-only weights and about 567 megabytes for the full multimodal model. Those figures are promising for a demo device. They are not a guarantee for every phone, runtime or production app. Index storage, media decoding, the user interface and any generative model all need memory too. The model has an 8,192-token context window shared across every modality. Google says that can hold roughly 5.5 minutes of audio, 29 images or 58 video frames at default settings when used for only one media type. Mix text, images, sound and video, and they draw from the same budget. Again, the box has knobs. Video is sampled at one frame per second by default. Image and video token budgets can be raised for finer detail or lowered for speed. Audio should arrive as mono at 16 kilohertz. A team searching short interviews has a different optimization problem from one matching product photos or finding moments in a two-hour lecture. The benchmark story is encouraging but belongs in the correct frame. Google reports nearly unchanged multilingual text performance compared with the first EmbeddingGemma and a jump on the MTEB code benchmark from 68.76 to 78.68. It also reports results across image, visual-document, video and audio benchmarks, where the first model had no equivalent multimodal scores. Those are vendor-run evaluations on public benchmark collections. They show the model is worth testing. They do not show that it will understand a hospital's scanned forms, an Indigenous-language oral archive, a mechanic's noisy recordings or your family's gloriously chaotic photo library. The model card says it understands more than 100 languages. It also says quality may not be equal across them. That second sentence should travel with the first. There are several small implementation traps with unusually large consequences. Text inputs work best with task prefixes that tell the model whether it is doing search, question answering, fact checking, code retrieval, classification or similarity. Omitting the prefix still produces a vector, but Google says precision drops. Truncated vectors must be normalized again. Queries and indexed items must use the same dimension. Google warns that skipping normalization can silently degrade rankings while still producing plausible scores. The model card also says not to run inference in float16 because values can become invalid or quietly deteriorate. Use bfloat16 on supported hardware or float32 elsewhere. This is the deeply unglamorous part of AI that decides whether a product works. A retrieval system can return something that looks reasonable while missing the file the user actually needed. There may be no crash, red light or theatrical robot apology. The wrong neighbor simply arrives first. That is why every deployment needs its own test set. Build queries from real users, label the right results, measure misses by language and media type, and preserve a basic keyword or metadata search as a fallback. Check not only the first result, but whether the correct item appears anywhere in the useful shortlist. Privacy needs the same discipline. Running the encoder locally can keep raw files off a remote inference service. It can also reduce latency and allow search without a network connection. Those are substantial benefits. But local is not a magical privacy adjective. The app may still upload analytics, synchronize its vector index, expose filenames in logs, hand retrieved files to a cloud chatbot or retain queries in crash reports. Embeddings themselves can encode sensitive semantic relationships. A shared family computer can reveal that two documents are similar even if it does not display their full contents. Private by topology, not by adjective. To claim private search, a developer should publish the whole data path: where raw media is decoded, where vectors are stored, whether queries leave the device, what is logged, what is synchronized, which later models receive retrieved material and how everything can be deleted. The training story deserves attention too. Google says the model learned from web documents, code, images, video, audio and cross-modal pairs, with a January 2025 cutoff. It reports filtering for child sexual abuse material and certain personal information. The model card does not provide a full itemized training corpus. Embedding models do not produce prose that can be moderated after generation. Their risks surface in rankings, classifications and clusters. A biased neighborhood in vector space can make some people, accents, languages or cultural references harder to retrieve. It can also group harmful material in ways a downstream app did not anticipate. Google did not apply post-training alignment, safety tuning or output moderation to this model. It assigns developers responsibility for retrieval filtering, fairness testing and other application-level safeguards. That is reasonable as a division of labor. It is also a transfer of work. The weights are downloadable and the model card lists an Apache 2.0 license. Deployments must also follow the Gemma Prohibited Use Policy. Teams should review both before treating "open" as a complete legal summary. What should people build? Local search across notes, screenshots, recordings and video is the obvious start. Field teams could search offline manuals with a photo or spoken question. Journalists could find a moment in an interview archive without uploading confidential recordings. Developers could retrieve related code, diagrams and bug screenshots from one query. Museums and libraries could connect descriptions, catalog text, images and audio guides while keeping sensitive collections on site. Accessibility tools may be especially interesting. A person who cannot remember a filename could describe the sound or scene. A spoken query could reach a visual collection. A photo could retrieve a related explanation. None of those applications needs the embedding model to know the truth. It needs the model to surface useful candidates. A person or a separate system still has to inspect the result, understand the context and decide what follows. That modest job description is a strength. EmbeddingGemma 2 does not promise a pocket oracle. It offers a compact filing clerk with several senses, a large multilingual vocabulary and a somewhat fussy instruction manual. For local search, that may be exactly the machine worth having.

01

WHAT ACTUALLY CHANGED

Google DeepMind released EmbeddingGemma 2 on October 6 with downloadable weights and an Apache 2.0 license

The model maps text, code, images, video and audio into one shared 768-dimensional vector space

Developers can load a 270-million-parameter text system, add vision or audio selectively, or use the full 740-million-parameter model

Output vectors can be reduced to 512, 256 or 128 dimensions to trade retrieval quality for lower storage and faster search

Google reports on-device memory and benchmark gains, while its model card places downstream safety testing and retrieval filtering on deployers

02

WHY THIS MATTERS

One compact model can support cross-modal search without sending an entire private media collection to a server

Selective encoders make local retrieval possible for products that need only text, images or audio rather than every modality

A shared space can connect a spoken or typed query to code, pictures, recordings and video moments without exact keyword matches

Vendor benchmarks do not predict quality for a particular language, archive, device or operational setting

Local inference improves privacy only when indexing, logging, synchronization and later reasoning steps also stay under control

FIG. 335How local multimodal search works
1Choose the text, image, audio and video encoders the device actually needs→
2Turn each local file into a numerical vector and store it in a local index→
3Convert the user query into a vector with the correct task instruction and dimension→
4Compare the query with the index and return the nearest candidate files→
5Let the user inspect the results, record misses and improve the test set
The model builds a searchable map of meaning. Privacy and usefulness depend on where the map is stored, what leaves the device and how retrieval errors are measured.

03

WHERE IT COULD HELP

  • Search personal notes, screenshots, voice memos and video from one local query
  • Build offline field tools that retrieve manuals from a photo, spoken question or short clip
  • Index code, diagrams, bug screenshots and developer recordings in one project search layer
  • Search confidential interview, legal or research archives without uploading the raw collection for embedding
  • Create accessibility tools that let one modality, such as speech or an image, retrieve useful material in another
  • Measure retrieval quality on real user queries by language, modality and failure cost before production use

KEEP A HAND ON THE WHEEL

Watch for independent tests on older phones and laptops, end-to-end memory and battery measurements, retrieval quality in lower-resource languages, audit results for bias and sensitive-content clustering, community comparisons with larger multimodal embedders, application-level privacy disclosures, fine-tuned domain models, and evidence that shorter vectors preserve quality on real archives rather than only public benchmarks.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on October 7, 2026.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US