← posts / rag & search

EmbeddingGemma 2 in Qdrant: one 740M model for text, images and audio, and what truncating 768 to 128 dims costs your recall

Google's open multimodal embedder runs locally and truncates from 768 to 128 dims. Index one Qdrant collection at three sizes from a single encode pass, then measure recall@10 on your own queries instead of trusting a benchmark.

if.codesOct 7, 2026 · 6 min read#embeddings#qdrant#rag#open-modelsAI-assisted

Google released EmbeddingGemma 2 on 6 October 2026 under Apache 2.0. It is a 740M-parameter embedder built on Gemma 4 that puts text, images, audio and video into one vector space. It is modular: the text-only core is 270M parameters, and the vision (170M) and audio (300M) encoders are optional. Output is 768 dimensions, and Matryoshka training lets you cut vectors to 512, 256 or 128 dimensions. Google says this means "up to 6x storage reduction". The context window is 8K tokens, four times the first EmbeddingGemma.

The question this post answers is practical: if you store 128-dim vectors instead of 768, how much of your retrieval do you lose on your own data? Google's announcement does not publish a recall-versus-dimension table, and the model card only says quality is close to lossless down to 256 dimensions. So below is a setup that encodes once, stores three sizes side by side in Qdrant, and prints recall@10 for each. It takes about 15 minutes plus encoding time.

What it costs to run

Google gives on-device memory figures measured on a Pixel 11 Pro with quantized weights: about 191 MB of active RAM for the text-only weights and about 567 MB for the full multimodal model. The gap is the price of the vision and audio encoders. If your corpus is text only, load the text-only variant.

The model's memory footprint is optimized for on-device deployment.
The model's memory footprint is optimized for on-device deployment.
0 MB200 MB400 MB600 MBText-only weights · Active RAM: 191 MB191 MBText-only weightsFull multimodal · Active RAM: 567 MB567 MBFull multimodal
EmbeddingGemma 2 active RAM on a Pixel 11 Pro (quantized) Approximate figures as stated by Google Source: Google, EmbeddingGemma 2 announcement, Oct 2026
EmbeddingGemma 2 active RAM on a Pixel 11 Pro (quantized)
Active RAM
Text-only weights191 MB
Full multimodal567 MB

Storage scales linearly with dimensions. With float32 vectors (4 bytes per dimension), the four Matryoshka sizes work out as below. That is where the "6x" comes from: 3,072 bytes against 512 per vector, before any quantization in the vector database.

0 bytes1,000 bytes2,000 bytes3,000 bytes4,000 bytes768 dims · float32: 3,072 bytes3,072 bytes768 dims512 dims · float32: 2,048 bytes2,048 bytes512 dims256 dims · float32: 1,024 bytes1,024 bytes256 dims128 dims · float32: 512 bytes512 bytes128 dims
Bytes per stored vector at each EmbeddingGemma 2 dimension Arithmetic: dimensions from Google times 4 bytes per float32 value Source: Google, EmbeddingGemma 2 announcement, Oct 2026
Bytes per stored vector at each EmbeddingGemma 2 dimension
float32
768 dims3,072 bytes
512 dims2,048 bytes
256 dims1,024 bytes
128 dims512 bytes

Google lists support in transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio, with Qdrant for storage. In Ollama the library page shows embeddinggemma-2 tags for 270m (text), 440m (text and image), 570m (text) and 740m (text and image). Ollama is fine for text, but this post uses sentence-transformers so that images go through the same model card code.

1. Start Qdrant and install the libraries

docker run -d --name qdrant -p 6333:6333 -v "$(pwd)/qdrant_storage:/qdrant/storage" qdrant/qdrant
python3 -m venv .venv
. .venv/bin/activate
pip install -U sentence-transformers transformers qdrant-client numpy pillow

2. Describe your corpus

One JSON object per line in docs.jsonl. The image field is optional. Use it for a screenshot that belongs to the doc, such as an error dialog or a settings page.

{"id": "csv-import", "title": "Importing a CSV", "text": "Columns must match the template. Error 42 means a missing header row.", "image": "shots/error-42.png"}
{"id": "reset-password", "title": "Resetting your password", "text": "Open Settings, then Security, then Reset password."}

3. Encode once, store three sizes

Matryoshka truncation means the first k values of the 768-dim vector are themselves a usable embedding once renormalized. So you only need one model pass per document. The script slices the vector three times and writes each slice to a named vector in the same Qdrant point. Documents use the card's title: … | text: … format, and images use the card's interleaved <|image|> placeholder.

Matryoshka embeddings allow a single vector to be truncated into multiple smaller sizes.
Matryoshka embeddings allow a single vector to be truncated into multiple smaller sizes.
# index.py - usage: python index.py docs.jsonl
import json
import sys

import numpy as np
from qdrant_client import QdrantClient, models
from sentence_transformers import SentenceTransformer

DIMS = [768, 256, 128]
model = SentenceTransformer("google/embeddinggemma-2")
client = QdrantClient(url="http://localhost:6333")


def shrink(vec, dim):
    v = np.asarray(vec[:dim], dtype=np.float32)
    return (v / np.linalg.norm(v)).tolist()


def embed_doc(doc):
    text = f"title: {doc['title']} | text: {doc['text']}"
    if doc.get("image"):
        return model.encode({"text": text + " <|image|>", "image": [doc["image"]]})
    return model.encode(text)


docs = [json.loads(line) for line in open(sys.argv[1]) if line.strip()]

if not client.collection_exists("docs"):
    client.create_collection(
        "docs",
        vectors_config={
            f"d{d}": models.VectorParams(size=d, distance=models.Distance.COSINE)
            for d in DIMS
        },
    )

points = []
for i, doc in enumerate(docs):
    full = embed_doc(doc)
    points.append(
        models.PointStruct(
            id=i,
            vector={f"d{d}": shrink(full, d) for d in DIMS},
            payload={"doc_id": doc["id"], "title": doc["title"]},
        )
    )

client.upsert("docs", points=points)
print(f"indexed {len(points)} docs at {DIMS} dims")

Storing all three sizes on every point triples the cost of the experiment, not of production. Once you have picked a size, recreate the collection with that one vector.

Want semantic search like this on your own data?I build RAG and search systems end to end. The estimate is free.

4. Write queries you already know the answer to

Recall is only meaningful against a labelled set. Take 30 to 50 real questions from your support inbox, search logs or chat history, and write down the doc that should answer each one. Keep the wording as messy as users write it. A set of tidy, keyword-matching queries will flatter every dimension equally.

{"query": "getting error 42 when i upload my spreadsheet", "relevant": "csv-import"}
{"query": "locked out, how do i change my password", "relevant": "reset-password"}

5. Measure recall@10 per size

The script reports two numbers per size. Recall@10 is the share of queries whose labelled doc appears in the top 10. Overlap with 768 is how much of the full-size top 10 the smaller vector keeps. The first number is what your users feel. The second tells you whether truncation reshuffles results even when the right answer survives. Queries use the card's task: search result | query: … prefix.

Measuring recall helps determine the actual impact of dimension reduction on search quality.
Measuring recall helps determine the actual impact of dimension reduction on search quality.
# eval.py - usage: python eval.py queries.jsonl
import json
import sys

import numpy as np
from qdrant_client import QdrantClient
from sentence_transformers import SentenceTransformer

DIMS = [768, 256, 128]
K = 10
model = SentenceTransformer("google/embeddinggemma-2")
client = QdrantClient(url="http://localhost:6333")


def shrink(vec, dim):
    v = np.asarray(vec[:dim], dtype=np.float32)
    return (v / np.linalg.norm(v)).tolist()


def top_k(qvec, dim):
    res = client.query_points("docs", query=shrink(qvec, dim), using=f"d{dim}", limit=K)
    return [p.payload["doc_id"] for p in res.points]


queries = [json.loads(line) for line in open(sys.argv[1]) if line.strip()]
hits = {d: 0 for d in DIMS}
overlap = {d: 0.0 for d in DIMS}

for q in queries:
    qvec = model.encode(f"task: search result | query: {q['query']}")
    full = top_k(qvec, 768)
    for d in DIMS:
        got = top_k(qvec, d)
        if q["relevant"] in got:
            hits[d] += 1
        overlap[d] += len(set(got) & set(full)) / max(len(full), 1)

n = len(queries)
for d in DIMS:
    print(f"{d:>4} dims  recall@{K} {hits[d] / n:.3f}  "
          f"overlap with 768 {overlap[d] / n:.3f}  bytes/vector {d * 4}")

Run both scripts:

python index.py docs.jsonl
python eval.py queries.jsonl

The output has this shape (the numbers are placeholders; yours depend entirely on your corpus):

 768 dims  recall@10 x.xxx  overlap with 768 1.000  bytes/vector 3072
 256 dims  recall@10 x.xxx  overlap with 768 x.xxx  bytes/vector 1024
 128 dims  recall@10 x.xxx  overlap with 768 x.xxx  bytes/vector 512

Reading the result

  • Recall at 256 matches 768, and overlap stays high: take the 3x saving. This is the case the model card describes as close to lossless.
  • Recall holds at 128 but overlap drops: the right doc is still in the top 10 but in a different order. That is fine if a reranker or an LLM reads the full top 10, and worse if you show only the first three results.
  • Recall drops at 128: a common pattern is to search wide at 128 dims (say limit 50) and rescore those candidates with the 768-dim vectors. Qdrant can hold both on the same point, which is exactly what the index above does.
  • Screenshots behave differently from text: split the query set by whether the relevant doc has an image and report each half separately. A single average can hide a modality that degrades faster.

Two caveats. First, 30 to 50 queries is enough to spot a big drop, not to separate 0.92 from 0.94. Treat small differences as noise. Second, the RAM figures above are Google's phone measurements with quantized weights. A desktop run through sentence-transformers in full precision will use more, so check with your own process monitor before sizing a server.

Sources

if.codesI build RAG, AI integrations and agent pipelines on Go and Python backends — and write about it here.
// keep reading

More posts on AI and backends.

// free quote

Read something you need? I’ll quote it for free.

RAG, AI integrations, agents or the backend underneath — tell me what you have and what should change. I read every request myself.

Free · no commitment

Tell me what you have. I’ll tell you what it takes.

1Describe the projectA few sentences is enough — about two minutes.
2I review itI read it myself and may ask a follow-up question.
3You get a free quoteScope, approach and estimate — yours to keep, no strings.
What kind of project is it?
Free and without obligation. Your details are used only to reply — see the privacy policy.