Google released EmbeddingGemma 2 on 6 October 2026 under Apache 2.0. It is a 740M-parameter embedder built on Gemma 4 that puts text, images, audio and video into one vector space. It is modular: the text-only core is 270M parameters, and the vision (170M) and audio (300M) encoders are optional. Output is 768 dimensions, and Matryoshka training lets you cut vectors to 512, 256 or 128 dimensions. Google says this means "up to 6x storage reduction". The context window is 8K tokens, four times the first EmbeddingGemma.
The question this post answers is practical: if you store 128-dim vectors instead of 768, how much of your retrieval do you lose on your own data? Google's announcement does not publish a recall-versus-dimension table, and the model card only says quality is close to lossless down to 256 dimensions. So below is a setup that encodes once, stores three sizes side by side in Qdrant, and prints recall@10 for each. It takes about 15 minutes plus encoding time.
What it costs to run
Google gives on-device memory figures measured on a Pixel 11 Pro with quantized weights: about 191 MB of active RAM for the text-only weights and about 567 MB for the full multimodal model. The gap is the price of the vision and audio encoders. If your corpus is text only, load the text-only variant.

Storage scales linearly with dimensions. With float32 vectors (4 bytes per dimension), the four Matryoshka sizes work out as below. That is where the "6x" comes from: 3,072 bytes against 512 per vector, before any quantization in the vector database.
Google lists support in transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio, with Qdrant for storage. In Ollama the library page shows embeddinggemma-2 tags for 270m (text), 440m (text and image), 570m (text) and 740m (text and image). Ollama is fine for text, but this post uses sentence-transformers so that images go through the same model card code.
1. Start Qdrant and install the libraries
docker run -d --name qdrant -p 6333:6333 -v "$(pwd)/qdrant_storage:/qdrant/storage" qdrant/qdrant
python3 -m venv .venv
. .venv/bin/activate
pip install -U sentence-transformers transformers qdrant-client numpy pillow
2. Describe your corpus
One JSON object per line in docs.jsonl. The image field is optional. Use it for a screenshot that belongs to the doc, such as an error dialog or a settings page.
{"id": "csv-import", "title": "Importing a CSV", "text": "Columns must match the template. Error 42 means a missing header row.", "image": "shots/error-42.png"}
{"id": "reset-password", "title": "Resetting your password", "text": "Open Settings, then Security, then Reset password."}
3. Encode once, store three sizes
Matryoshka truncation means the first k values of the 768-dim vector are themselves a usable embedding once renormalized. So you only need one model pass per document. The script slices the vector three times and writes each slice to a named vector in the same Qdrant point. Documents use the card's title: … | text: … format, and images use the card's interleaved <|image|> placeholder.

# index.py - usage: python index.py docs.jsonl
import json
import sys
import numpy as np
from qdrant_client import QdrantClient, models
from sentence_transformers import SentenceTransformer
DIMS = [768, 256, 128]
model = SentenceTransformer("google/embeddinggemma-2")
client = QdrantClient(url="http://localhost:6333")
def shrink(vec, dim):
v = np.asarray(vec[:dim], dtype=np.float32)
return (v / np.linalg.norm(v)).tolist()
def embed_doc(doc):
text = f"title: {doc['title']} | text: {doc['text']}"
if doc.get("image"):
return model.encode({"text": text + " <|image|>", "image": [doc["image"]]})
return model.encode(text)
docs = [json.loads(line) for line in open(sys.argv[1]) if line.strip()]
if not client.collection_exists("docs"):
client.create_collection(
"docs",
vectors_config={
f"d{d}": models.VectorParams(size=d, distance=models.Distance.COSINE)
for d in DIMS
},
)
points = []
for i, doc in enumerate(docs):
full = embed_doc(doc)
points.append(
models.PointStruct(
id=i,
vector={f"d{d}": shrink(full, d) for d in DIMS},
payload={"doc_id": doc["id"], "title": doc["title"]},
)
)
client.upsert("docs", points=points)
print(f"indexed {len(points)} docs at {DIMS} dims")
Storing all three sizes on every point triples the cost of the experiment, not of production. Once you have picked a size, recreate the collection with that one vector.



