[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f1d948m6m3kcrg":3,"$fanuq43nlrv5g":56},{"slug":4,"title":5,"body":6,"summary":7,"tags":8,"author":14,"cover_url":15,"published_at":16,"seo_title":17,"seo_description":18,"reading_minutes":19,"related":20},"embeddinggemma-2-qdrant-truncation-recall","EmbeddingGemma 2 in Qdrant: one 740M model for text, images and audio, and what truncating 768 to 128 dims costs your recall","\u003Cp>Google released \u003Cstrong>EmbeddingGemma 2\u003C\u002Fstrong> on 6 October 2026 under Apache 2.0. It is a 740M-parameter embedder built on Gemma 4 that puts text, images, audio and video into one vector space. It is modular: the text-only core is 270M parameters, and the vision (170M) and audio (300M) encoders are optional. Output is 768 dimensions, and Matryoshka training lets you cut vectors to 512, 256 or 128 dimensions. Google says this means \"up to 6x storage reduction\". The context window is 8K tokens, four times the first EmbeddingGemma.\u003C\u002Fp>\n\u003Cfigure data-post-media=\"6ac5e861df0d655df505346d\">\u003Cvideo src=\"https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6ac5e863df0d655df5053473-0-9a3d0d36.mp4\" autoplay muted loop playsinline preload=\"metadata\">\u003C\u002Fvideo>\u003C\u002Ffigure>\n\u003Cp>The question this post answers is practical: if you store 128-dim vectors instead of 768, how much of your retrieval do you lose \u003Cem>on your own data\u003C\u002Fem>? Google's announcement does not publish a recall-versus-dimension table, and the model card only says quality is close to lossless down to 256 dimensions. So below is a setup that encodes once, stores three sizes side by side in Qdrant, and prints recall@10 for each. It takes about 15 minutes plus encoding time.\u003C\u002Fp>\n\u003Ch2>What it costs to run\u003C\u002Fh2>\n\u003Cp>Google gives on-device memory figures measured on a Pixel 11 Pro with quantized weights: about 191 MB of active RAM for the text-only weights and about 567 MB for the full multimodal model. The gap is the price of the vision and audio encoders. If your corpus is text only, load the text-only variant.\u003C\u002Fp>\n\u003Cfigure data-post-media=\"6ac5ccf1df0d655df505307e\">\u003Cimg src=\"https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6ac5ccf1df0d655df5053083-0-73d263b8.png\" alt=\"The model&#39;s memory footprint is optimized for on-device deployment.\" loading=\"lazy\">\u003Cfigcaption>The model&#39;s memory footprint is optimized for on-device deployment.\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\u003Cfigure data-chart='{\"type\":\"bar\",\"title\":\"EmbeddingGemma 2 active RAM on a Pixel 11 Pro (quantized)\",\"unit\":\"MB\",\"labels\":[\"Text-only weights\",\"Full multimodal\"],\"series\":[{\"name\":\"Active RAM\",\"values\":[191,567]}],\"note\":\"Approximate figures as stated by Google\",\"source\":{\"label\":\"Google, EmbeddingGemma 2 announcement, Oct 2026\",\"url\":\"https:\u002F\u002Fblog.google\u002Finnovation-and-ai\u002Ftechnology\u002Fdevelopers-tools\u002Fembeddinggemma-2\u002F\"}}'>\u003Cfigcaption>EmbeddingGemma 2 active RAM on a Pixel 11 Pro: about 191 MB text-only, about 567 MB multimodal (source: Google)\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\u003Cp>Storage scales linearly with dimensions. With float32 vectors (4 bytes per dimension), the four Matryoshka sizes work out as below. That is where the \"6x\" comes from: 3,072 bytes against 512 per vector, before any quantization in the vector database.\u003C\u002Fp>\n\u003Cfigure data-chart='{\"type\":\"bar\",\"title\":\"Bytes per stored vector at each EmbeddingGemma 2 dimension\",\"unit\":\"bytes\",\"labels\":[\"768 dims\",\"512 dims\",\"256 dims\",\"128 dims\"],\"series\":[{\"name\":\"float32\",\"values\":[3072,2048,1024,512]}],\"note\":\"Arithmetic: dimensions from Google times 4 bytes per float32 value\",\"source\":{\"label\":\"Google, EmbeddingGemma 2 announcement, Oct 2026\",\"url\":\"https:\u002F\u002Fblog.google\u002Finnovation-and-ai\u002Ftechnology\u002Fdevelopers-tools\u002Fembeddinggemma-2\u002F\"}}'>\u003Cfigcaption>Bytes per float32 vector at 768, 512, 256 and 128 dimensions: 3,072 down to 512 (dimensions from Google, arithmetic by if.codes)\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\u003Cp>Google lists support in transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio, with Qdrant for storage. In Ollama the library page shows \u003Ccode>embeddinggemma-2\u003C\u002Fcode> tags for 270m (text), 440m (text and image), 570m (text) and 740m (text and image). Ollama is fine for text, but this post uses sentence-transformers so that images go through the same model card code.\u003C\u002Fp>\n\u003Ch2>1. Start Qdrant and install the libraries\u003C\u002Fh2>\n\u003Cpre class=\"code-block\" data-lang=\"bash\">\u003Ccode class=\"hljs language-bash\">docker run -d --name qdrant -p 6333:6333 -v \u003Cspan class=\"hljs-string\">&quot;\u003Cspan class=\"hljs-subst\">$(pwd)\u003C\u002Fspan>\u002Fqdrant_storage:\u002Fqdrant\u002Fstorage&quot;\u003C\u002Fspan> qdrant\u002Fqdrant\npython3 -m venv .venv\n. .venv\u002Fbin\u002Factivate\npip install -U sentence-transformers transformers qdrant-client numpy pillow\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Ch2>2. Describe your corpus\u003C\u002Fh2>\n\u003Cp>One JSON object per line in \u003Ccode>docs.jsonl\u003C\u002Fcode>. The \u003Ccode>image\u003C\u002Fcode> field is optional. Use it for a screenshot that belongs to the doc, such as an error dialog or a settings page.\u003C\u002Fp>\n\u003Cpre class=\"code-block\" data-lang=\"json\">\u003Ccode class=\"hljs language-json\">\u003Cspan class=\"hljs-punctuation\">{\u003C\u002Fspan>\u003Cspan class=\"hljs-attr\">&quot;id&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">:\u003C\u002Fspan> \u003Cspan class=\"hljs-string\">&quot;csv-import&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">,\u003C\u002Fspan> \u003Cspan class=\"hljs-attr\">&quot;title&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">:\u003C\u002Fspan> \u003Cspan class=\"hljs-string\">&quot;Importing a CSV&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">,\u003C\u002Fspan> \u003Cspan class=\"hljs-attr\">&quot;text&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">:\u003C\u002Fspan> \u003Cspan class=\"hljs-string\">&quot;Columns must match the template. Error 42 means a missing header row.&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">,\u003C\u002Fspan> \u003Cspan class=\"hljs-attr\">&quot;image&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">:\u003C\u002Fspan> \u003Cspan class=\"hljs-string\">&quot;shots\u002Ferror-42.png&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">}\u003C\u002Fspan>\n\u003Cspan class=\"hljs-punctuation\">{\u003C\u002Fspan>\u003Cspan class=\"hljs-attr\">&quot;id&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">:\u003C\u002Fspan> \u003Cspan class=\"hljs-string\">&quot;reset-password&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">,\u003C\u002Fspan> \u003Cspan class=\"hljs-attr\">&quot;title&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">:\u003C\u002Fspan> \u003Cspan class=\"hljs-string\">&quot;Resetting your password&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">,\u003C\u002Fspan> \u003Cspan class=\"hljs-attr\">&quot;text&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">:\u003C\u002Fspan> \u003Cspan class=\"hljs-string\">&quot;Open Settings, then Security, then Reset password.&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">}\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Ch2>3. Encode once, store three sizes\u003C\u002Fh2>\n\u003Cp>Matryoshka truncation means the first \u003Cem>k\u003C\u002Fem> values of the 768-dim vector are themselves a usable embedding once renormalized. So you only need one model pass per document. The script slices the vector three times and writes each slice to a named vector in the same Qdrant point. Documents use the card's \u003Ccode>title: … | text: …\u003C\u002Fcode> format, and images use the card's interleaved \u003Ccode>&lt;|image|&gt;\u003C\u002Fcode> placeholder.\u003C\u002Fp>\n\u003Cfigure data-post-media=\"6ac5ccf1df0d655df5053074\">\u003Cimg src=\"https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6ac5ccf1df0d655df5053079-0-46d7cbf7.png\" alt=\"Matryoshka embeddings allow a single vector to be truncated into multiple smaller sizes.\" loading=\"lazy\">\u003Cfigcaption>Matryoshka embeddings allow a single vector to be truncated into multiple smaller sizes.\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\u003Cpre class=\"code-block\" data-lang=\"python\">\u003Ccode class=\"hljs language-python\">\u003Cspan class=\"hljs-comment\"># index.py - usage: python index.py docs.jsonl\u003C\u002Fspan>\n\u003Cspan class=\"hljs-keyword\">import\u003C\u002Fspan> json\n\u003Cspan class=\"hljs-keyword\">import\u003C\u002Fspan> sys\n\n\u003Cspan class=\"hljs-keyword\">import\u003C\u002Fspan> numpy \u003Cspan class=\"hljs-keyword\">as\u003C\u002Fspan> np\n\u003Cspan class=\"hljs-keyword\">from\u003C\u002Fspan> qdrant_client \u003Cspan class=\"hljs-keyword\">import\u003C\u002Fspan> QdrantClient, models\n\u003Cspan class=\"hljs-keyword\">from\u003C\u002Fspan> sentence_transformers \u003Cspan class=\"hljs-keyword\">import\u003C\u002Fspan> SentenceTransformer\n\nDIMS = [\u003Cspan class=\"hljs-number\">768\u003C\u002Fspan>, \u003Cspan class=\"hljs-number\">256\u003C\u002Fspan>, \u003Cspan class=\"hljs-number\">128\u003C\u002Fspan>]\nmodel = SentenceTransformer(\u003Cspan class=\"hljs-string\">&quot;google\u002Fembeddinggemma-2&quot;\u003C\u002Fspan>)\nclient = QdrantClient(url=\u003Cspan class=\"hljs-string\">&quot;http:\u002F\u002Flocalhost:6333&quot;\u003C\u002Fspan>)\n\n\n\u003Cspan class=\"hljs-keyword\">def\u003C\u002Fspan> \u003Cspan class=\"hljs-title function_\">shrink\u003C\u002Fspan>(\u003Cspan class=\"hljs-params\">vec, dim\u003C\u002Fspan>):\n    v = np.asarray(vec[:dim], dtype=np.float32)\n    \u003Cspan class=\"hljs-keyword\">return\u003C\u002Fspan> (v \u002F np.linalg.norm(v)).tolist()\n\n\n\u003Cspan class=\"hljs-keyword\">def\u003C\u002Fspan> \u003Cspan class=\"hljs-title function_\">embed_doc\u003C\u002Fspan>(\u003Cspan class=\"hljs-params\">doc\u003C\u002Fspan>):\n    text = \u003Cspan class=\"hljs-string\">f&quot;title: \u003Cspan class=\"hljs-subst\">{doc[\u003Cspan class=\"hljs-string\">&#x27;title&#x27;\u003C\u002Fspan>]}\u003C\u002Fspan> | text: \u003Cspan class=\"hljs-subst\">{doc[\u003Cspan class=\"hljs-string\">&#x27;text&#x27;\u003C\u002Fspan>]}\u003C\u002Fspan>&quot;\u003C\u002Fspan>\n    \u003Cspan class=\"hljs-keyword\">if\u003C\u002Fspan> doc.get(\u003Cspan class=\"hljs-string\">&quot;image&quot;\u003C\u002Fspan>):\n        \u003Cspan class=\"hljs-keyword\">return\u003C\u002Fspan> model.encode({\u003Cspan class=\"hljs-string\">&quot;text&quot;\u003C\u002Fspan>: text + \u003Cspan class=\"hljs-string\">&quot; &lt;|image|&gt;&quot;\u003C\u002Fspan>, \u003Cspan class=\"hljs-string\">&quot;image&quot;\u003C\u002Fspan>: [doc[\u003Cspan class=\"hljs-string\">&quot;image&quot;\u003C\u002Fspan>]]})\n    \u003Cspan class=\"hljs-keyword\">return\u003C\u002Fspan> model.encode(text)\n\n\ndocs = [json.loads(line) \u003Cspan class=\"hljs-keyword\">for\u003C\u002Fspan> line \u003Cspan class=\"hljs-keyword\">in\u003C\u002Fspan> \u003Cspan class=\"hljs-built_in\">open\u003C\u002Fspan>(sys.argv[\u003Cspan class=\"hljs-number\">1\u003C\u002Fspan>]) \u003Cspan class=\"hljs-keyword\">if\u003C\u002Fspan> line.strip()]\n\n\u003Cspan class=\"hljs-keyword\">if\u003C\u002Fspan> \u003Cspan class=\"hljs-keyword\">not\u003C\u002Fspan> client.collection_exists(\u003Cspan class=\"hljs-string\">&quot;docs&quot;\u003C\u002Fspan>):\n    client.create_collection(\n        \u003Cspan class=\"hljs-string\">&quot;docs&quot;\u003C\u002Fspan>,\n        vectors_config={\n            \u003Cspan class=\"hljs-string\">f&quot;d\u003Cspan class=\"hljs-subst\">{d}\u003C\u002Fspan>&quot;\u003C\u002Fspan>: models.VectorParams(size=d, distance=models.Distance.COSINE)\n            \u003Cspan class=\"hljs-keyword\">for\u003C\u002Fspan> d \u003Cspan class=\"hljs-keyword\">in\u003C\u002Fspan> DIMS\n        },\n    )\n\npoints = []\n\u003Cspan class=\"hljs-keyword\">for\u003C\u002Fspan> i, doc \u003Cspan class=\"hljs-keyword\">in\u003C\u002Fspan> \u003Cspan class=\"hljs-built_in\">enumerate\u003C\u002Fspan>(docs):\n    full = embed_doc(doc)\n    points.append(\n        models.PointStruct(\n            \u003Cspan class=\"hljs-built_in\">id\u003C\u002Fspan>=i,\n            vector={\u003Cspan class=\"hljs-string\">f&quot;d\u003Cspan class=\"hljs-subst\">{d}\u003C\u002Fspan>&quot;\u003C\u002Fspan>: shrink(full, d) \u003Cspan class=\"hljs-keyword\">for\u003C\u002Fspan> d \u003Cspan class=\"hljs-keyword\">in\u003C\u002Fspan> DIMS},\n            payload={\u003Cspan class=\"hljs-string\">&quot;doc_id&quot;\u003C\u002Fspan>: doc[\u003Cspan class=\"hljs-string\">&quot;id&quot;\u003C\u002Fspan>], \u003Cspan class=\"hljs-string\">&quot;title&quot;\u003C\u002Fspan>: doc[\u003Cspan class=\"hljs-string\">&quot;title&quot;\u003C\u002Fspan>]},\n        )\n    )\n\nclient.upsert(\u003Cspan class=\"hljs-string\">&quot;docs&quot;\u003C\u002Fspan>, points=points)\n\u003Cspan class=\"hljs-built_in\">print\u003C\u002Fspan>(\u003Cspan class=\"hljs-string\">f&quot;indexed \u003Cspan class=\"hljs-subst\">{\u003Cspan class=\"hljs-built_in\">len\u003C\u002Fspan>(points)}\u003C\u002Fspan> docs at \u003Cspan class=\"hljs-subst\">{DIMS}\u003C\u002Fspan> dims&quot;\u003C\u002Fspan>)\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Storing all three sizes on every point triples the cost of the experiment, not of production. Once you have picked a size, recreate the collection with that one vector.\u003C\u002Fp>\n\u003Ch2>4. Write queries you already know the answer to\u003C\u002Fh2>\n\u003Cp>Recall is only meaningful against a labelled set. Take 30 to 50 real questions from your support inbox, search logs or chat history, and write down the doc that should answer each one. Keep the wording as messy as users write it. A set of tidy, keyword-matching queries will flatter every dimension equally.\u003C\u002Fp>\n\u003Cpre class=\"code-block\" data-lang=\"json\">\u003Ccode class=\"hljs language-json\">\u003Cspan class=\"hljs-punctuation\">{\u003C\u002Fspan>\u003Cspan class=\"hljs-attr\">&quot;query&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">:\u003C\u002Fspan> \u003Cspan class=\"hljs-string\">&quot;getting error 42 when i upload my spreadsheet&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">,\u003C\u002Fspan> \u003Cspan class=\"hljs-attr\">&quot;relevant&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">:\u003C\u002Fspan> \u003Cspan class=\"hljs-string\">&quot;csv-import&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">}\u003C\u002Fspan>\n\u003Cspan class=\"hljs-punctuation\">{\u003C\u002Fspan>\u003Cspan class=\"hljs-attr\">&quot;query&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">:\u003C\u002Fspan> \u003Cspan class=\"hljs-string\">&quot;locked out, how do i change my password&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">,\u003C\u002Fspan> \u003Cspan class=\"hljs-attr\">&quot;relevant&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">:\u003C\u002Fspan> \u003Cspan class=\"hljs-string\">&quot;reset-password&quot;\u003C\u002Fspan>\u003Cspan class=\"hljs-punctuation\">}\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Ch2>5. Measure recall@10 per size\u003C\u002Fh2>\n\u003Cp>The script reports two numbers per size. \u003Cstrong>Recall@10\u003C\u002Fstrong> is the share of queries whose labelled doc appears in the top 10. \u003Cstrong>Overlap with 768\u003C\u002Fstrong> is how much of the full-size top 10 the smaller vector keeps. The first number is what your users feel. The second tells you whether truncation reshuffles results even when the right answer survives. Queries use the card's \u003Ccode>task: search result | query: …\u003C\u002Fcode> prefix.\u003C\u002Fp>\n\u003Cfigure data-post-media=\"6ac5ccf1df0d655df5053088\">\u003Cimg src=\"https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6ac5ccf1df0d655df505308d-0-501872df.png\" alt=\"Measuring recall helps determine the actual impact of dimension reduction on search quality.\" loading=\"lazy\">\u003Cfigcaption>Measuring recall helps determine the actual impact of dimension reduction on search quality.\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\u003Cpre class=\"code-block\" data-lang=\"python\">\u003Ccode class=\"hljs language-python\">\u003Cspan class=\"hljs-comment\"># eval.py - usage: python eval.py queries.jsonl\u003C\u002Fspan>\n\u003Cspan class=\"hljs-keyword\">import\u003C\u002Fspan> json\n\u003Cspan class=\"hljs-keyword\">import\u003C\u002Fspan> sys\n\n\u003Cspan class=\"hljs-keyword\">import\u003C\u002Fspan> numpy \u003Cspan class=\"hljs-keyword\">as\u003C\u002Fspan> np\n\u003Cspan class=\"hljs-keyword\">from\u003C\u002Fspan> qdrant_client \u003Cspan class=\"hljs-keyword\">import\u003C\u002Fspan> QdrantClient\n\u003Cspan class=\"hljs-keyword\">from\u003C\u002Fspan> sentence_transformers \u003Cspan class=\"hljs-keyword\">import\u003C\u002Fspan> SentenceTransformer\n\nDIMS = [\u003Cspan class=\"hljs-number\">768\u003C\u002Fspan>, \u003Cspan class=\"hljs-number\">256\u003C\u002Fspan>, \u003Cspan class=\"hljs-number\">128\u003C\u002Fspan>]\nK = \u003Cspan class=\"hljs-number\">10\u003C\u002Fspan>\nmodel = SentenceTransformer(\u003Cspan class=\"hljs-string\">&quot;google\u002Fembeddinggemma-2&quot;\u003C\u002Fspan>)\nclient = QdrantClient(url=\u003Cspan class=\"hljs-string\">&quot;http:\u002F\u002Flocalhost:6333&quot;\u003C\u002Fspan>)\n\n\n\u003Cspan class=\"hljs-keyword\">def\u003C\u002Fspan> \u003Cspan class=\"hljs-title function_\">shrink\u003C\u002Fspan>(\u003Cspan class=\"hljs-params\">vec, dim\u003C\u002Fspan>):\n    v = np.asarray(vec[:dim], dtype=np.float32)\n    \u003Cspan class=\"hljs-keyword\">return\u003C\u002Fspan> (v \u002F np.linalg.norm(v)).tolist()\n\n\n\u003Cspan class=\"hljs-keyword\">def\u003C\u002Fspan> \u003Cspan class=\"hljs-title function_\">top_k\u003C\u002Fspan>(\u003Cspan class=\"hljs-params\">qvec, dim\u003C\u002Fspan>):\n    res = client.query_points(\u003Cspan class=\"hljs-string\">&quot;docs&quot;\u003C\u002Fspan>, query=shrink(qvec, dim), using=\u003Cspan class=\"hljs-string\">f&quot;d\u003Cspan class=\"hljs-subst\">{dim}\u003C\u002Fspan>&quot;\u003C\u002Fspan>, limit=K)\n    \u003Cspan class=\"hljs-keyword\">return\u003C\u002Fspan> [p.payload[\u003Cspan class=\"hljs-string\">&quot;doc_id&quot;\u003C\u002Fspan>] \u003Cspan class=\"hljs-keyword\">for\u003C\u002Fspan> p \u003Cspan class=\"hljs-keyword\">in\u003C\u002Fspan> res.points]\n\n\nqueries = [json.loads(line) \u003Cspan class=\"hljs-keyword\">for\u003C\u002Fspan> line \u003Cspan class=\"hljs-keyword\">in\u003C\u002Fspan> \u003Cspan class=\"hljs-built_in\">open\u003C\u002Fspan>(sys.argv[\u003Cspan class=\"hljs-number\">1\u003C\u002Fspan>]) \u003Cspan class=\"hljs-keyword\">if\u003C\u002Fspan> line.strip()]\nhits = {d: \u003Cspan class=\"hljs-number\">0\u003C\u002Fspan> \u003Cspan class=\"hljs-keyword\">for\u003C\u002Fspan> d \u003Cspan class=\"hljs-keyword\">in\u003C\u002Fspan> DIMS}\noverlap = {d: \u003Cspan class=\"hljs-number\">0.0\u003C\u002Fspan> \u003Cspan class=\"hljs-keyword\">for\u003C\u002Fspan> d \u003Cspan class=\"hljs-keyword\">in\u003C\u002Fspan> DIMS}\n\n\u003Cspan class=\"hljs-keyword\">for\u003C\u002Fspan> q \u003Cspan class=\"hljs-keyword\">in\u003C\u002Fspan> queries:\n    qvec = model.encode(\u003Cspan class=\"hljs-string\">f&quot;task: search result | query: \u003Cspan class=\"hljs-subst\">{q[\u003Cspan class=\"hljs-string\">&#x27;query&#x27;\u003C\u002Fspan>]}\u003C\u002Fspan>&quot;\u003C\u002Fspan>)\n    full = top_k(qvec, \u003Cspan class=\"hljs-number\">768\u003C\u002Fspan>)\n    \u003Cspan class=\"hljs-keyword\">for\u003C\u002Fspan> d \u003Cspan class=\"hljs-keyword\">in\u003C\u002Fspan> DIMS:\n        got = top_k(qvec, d)\n        \u003Cspan class=\"hljs-keyword\">if\u003C\u002Fspan> q[\u003Cspan class=\"hljs-string\">&quot;relevant&quot;\u003C\u002Fspan>] \u003Cspan class=\"hljs-keyword\">in\u003C\u002Fspan> got:\n            hits[d] += \u003Cspan class=\"hljs-number\">1\u003C\u002Fspan>\n        overlap[d] += \u003Cspan class=\"hljs-built_in\">len\u003C\u002Fspan>(\u003Cspan class=\"hljs-built_in\">set\u003C\u002Fspan>(got) &amp; \u003Cspan class=\"hljs-built_in\">set\u003C\u002Fspan>(full)) \u002F \u003Cspan class=\"hljs-built_in\">max\u003C\u002Fspan>(\u003Cspan class=\"hljs-built_in\">len\u003C\u002Fspan>(full), \u003Cspan class=\"hljs-number\">1\u003C\u002Fspan>)\n\nn = \u003Cspan class=\"hljs-built_in\">len\u003C\u002Fspan>(queries)\n\u003Cspan class=\"hljs-keyword\">for\u003C\u002Fspan> d \u003Cspan class=\"hljs-keyword\">in\u003C\u002Fspan> DIMS:\n    \u003Cspan class=\"hljs-built_in\">print\u003C\u002Fspan>(\u003Cspan class=\"hljs-string\">f&quot;\u003Cspan class=\"hljs-subst\">{d:&gt;\u003Cspan class=\"hljs-number\">4\u003C\u002Fspan>}\u003C\u002Fspan> dims  recall@\u003Cspan class=\"hljs-subst\">{K}\u003C\u002Fspan> \u003Cspan class=\"hljs-subst\">{hits[d] \u002F n:\u003Cspan class=\"hljs-number\">.3\u003C\u002Fspan>f}\u003C\u002Fspan>  &quot;\u003C\u002Fspan>\n          \u003Cspan class=\"hljs-string\">f&quot;overlap with 768 \u003Cspan class=\"hljs-subst\">{overlap[d] \u002F n:\u003Cspan class=\"hljs-number\">.3\u003C\u002Fspan>f}\u003C\u002Fspan>  bytes\u002Fvector \u003Cspan class=\"hljs-subst\">{d * \u003Cspan class=\"hljs-number\">4\u003C\u002Fspan>}\u003C\u002Fspan>&quot;\u003C\u002Fspan>)\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Run both scripts:\u003C\u002Fp>\n\u003Cpre class=\"code-block\" data-lang=\"bash\">\u003Ccode class=\"hljs language-bash\">python index.py docs.jsonl\npython eval.py queries.jsonl\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>The output has this shape (the numbers are placeholders; yours depend entirely on your corpus):\u003C\u002Fp>\n\u003Cpre class=\"code-block\">\u003Ccode class=\"hljs\"> 768 dims  recall@10 x.xxx  overlap with 768 1.000  bytes\u002Fvector 3072\n 256 dims  recall@10 x.xxx  overlap with 768 x.xxx  bytes\u002Fvector 1024\n 128 dims  recall@10 x.xxx  overlap with 768 x.xxx  bytes\u002Fvector 512\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Ch2>Reading the result\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Recall at 256 matches 768, and overlap stays high:\u003C\u002Fstrong> take the 3x saving. This is the case the model card describes as close to lossless.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Recall holds at 128 but overlap drops:\u003C\u002Fstrong> the right doc is still in the top 10 but in a different order. That is fine if a reranker or an LLM reads the full top 10, and worse if you show only the first three results.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Recall drops at 128:\u003C\u002Fstrong> a common pattern is to search wide at 128 dims (say limit 50) and rescore those candidates with the 768-dim vectors. Qdrant can hold both on the same point, which is exactly what the index above does.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Screenshots behave differently from text:\u003C\u002Fstrong> split the query set by whether the relevant doc has an image and report each half separately. A single average can hide a modality that degrades faster.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Two caveats. First, 30 to 50 queries is enough to spot a big drop, not to separate 0.92 from 0.94. Treat small differences as noise. Second, the RAM figures above are Google's phone measurements with quantized weights. A desktop run through sentence-transformers in full precision will use more, so check with your own process monitor before sizing a server.\u003C\u002Fp>\n\u003Ch2>Sources\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fblog.google\u002Finnovation-and-ai\u002Ftechnology\u002Fdevelopers-tools\u002Fembeddinggemma-2\u002F\">Google: EmbeddingGemma 2 announcement\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Fgoogle\u002Fembeddinggemma-2\">Hugging Face: google\u002Fembeddinggemma-2 model card\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Follama.com\u002Flibrary\u002Fembeddinggemma-2\">Ollama library: embeddinggemma-2\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>","Google's open multimodal embedder runs locally and truncates from 768 to 128 dims. Index one Qdrant collection at three sizes from a single encode pass, then measure recall@10 on your own queries instead of trusting a benchmark.",[9,10,11,12,13],"embeddings","qdrant","rag","open-models","ai-assisted","if.codes","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6ac5cceedf0d655df505306e-0-d17d0fef.png","2026-10-07T07:21:20.071Z","EmbeddingGemma 2 in Qdrant: what 128 vs 768 dims costs your recall","Run Google's open multimodal embedder locally, store 768, 256 and 128-dim vectors in one Qdrant collection, and measure recall@10 on your own queries.",6,[21,34,45],{"slug":22,"title":23,"type":24,"summary":25,"tags":26,"author":14,"cover_url":31,"published_at":32,"updated_at":33},"cloudflare-access-strict-service-token-auth-migration","Strict service token auth in Cloudflare Access: moving your scripts and CI over before it bites","blog","Cloudflare Access now has a strict mode for service tokens: 401\u002F403 instead of a 302 to the login page, only Service Auth policies count, and no CF_Authorization cookie. New orgs get it forced on from 5 October. A 15-minute check and switch for existing orgs.",[27,28,29,30,13],"cloudflare","zero-trust","ci-cd","authentication","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6ac4a8860956594c947bd197-0-f768b25b.png","2026-10-06T08:30:13.154Z","2026-10-06T08:30:13.155Z",{"slug":35,"title":36,"type":24,"summary":37,"tags":38,"author":14,"cover_url":42,"published_at":43,"updated_at":44},"cloudflare-traces-trace-rules-debug-one-customer","Why was that request blocked? Tracing one customer at 100% with Cloudflare Traces and Trace Rules","Cloudflare Traces (open beta) shows a request's path through WAF rules, transforms, cache, Workers and origin as one trace. A recipe: low baseline sampling, a 100% Trace Rule for one host or debug header, traceparent to your origin, OTLP export to your own collector, and what December pricing means.",[27,39,40,41,13],"observability","opentelemetry","tracing","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6ac338052cb6b40613ac2b7c-0-2294d114.png","2026-10-05T06:22:19.055Z","2026-10-05T06:22:19.056Z",{"slug":46,"title":47,"type":24,"summary":48,"tags":49,"author":14,"cover_url":54,"published_at":55,"updated_at":55},"copyescape-cve-2026-17106-patch-docker-cp","CopyEscape (CVE-2026-17106): patch docker cp, and stop copying out of running containers","A race in docker cp lets a malicious container write files anywhere the copying process can write on the host. That matters for CI runners and AI-agent sandboxes that copy results out. Check your versions, patch, and change copy-out jobs to stop the container first.",[50,51,52,53,13],"docker","security","cve","ci","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6ac218ebfd770659725abe68-0-eee7c8f7.png","2026-10-04T23:51:14.689Z",[57,60,63,65,67,78,91,105,117,129,140,151,160,172,183,192,203,213,222,232],{"slug":4,"title":5,"type":24,"summary":7,"tags":58,"author":14,"cover_url":15,"published_at":16,"updated_at":59,"reading_minutes":19},[9,10,11,12,13],"2026-10-07T07:21:20.072Z",{"slug":22,"title":23,"type":24,"summary":25,"tags":61,"author":14,"cover_url":31,"published_at":32,"updated_at":33,"reading_minutes":62},[27,28,29,30,13],5,{"slug":35,"title":36,"type":24,"summary":37,"tags":64,"author":14,"cover_url":42,"published_at":43,"updated_at":44,"reading_minutes":19},[27,39,40,41,13],{"slug":46,"title":47,"type":24,"summary":48,"tags":66,"author":14,"cover_url":54,"published_at":55,"updated_at":55,"reading_minutes":62},[50,51,52,53,13],{"slug":68,"title":69,"type":24,"summary":70,"tags":71,"author":14,"cover_url":74,"published_at":75,"updated_at":76,"reading_minutes":77},"protected-quick-tunnels-vs-tailscale-funnel","Share localhost with three named people: Cloudflare's Protected Quick Tunnels vs Tailscale Funnel","cloudflared 2026.9.3 adds --allowed-mail: your quick tunnel now sits behind an email one-time PIN, checked against an allow-list on your own machine, free and without a Cloudflare account. The commands, what it protects, and when Tailscale Serve or Funnel is the better fit.",[27,72,73,51,13],"tailscale","tunnels","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6ac218d8fd770659725abe3a-0-40f8cb2e.png","2026-10-04T23:19:14.764Z","2026-10-04T23:19:14.765Z",4,{"slug":79,"title":80,"type":24,"summary":81,"tags":82,"author":14,"cover_url":88,"published_at":89,"updated_at":90,"reading_minutes":19},"si-domains-super-intelligence-data",".si after 'Super Intelligence': did one UN speech move a ccTLD?","Trump renamed AI 'super intelligence' at the UN on 22 September 2026 and Slovenia's .si went from about 190,000 names to almost 276,000 in a month. Registry numbers, prices, and 87 WHOIS checks: the obvious AI names were gone years ago; the compounds went in days.",[83,84,85,86,87,13],"domains","si","tld","data","ai","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abaaee53e21d5cbd14d442d-0-d6243bc5.png","2026-10-04T22:36:14.598Z","2026-10-04T22:36:14.599Z",{"slug":92,"title":93,"type":24,"summary":94,"tags":95,"author":14,"cover_url":101,"published_at":102,"updated_at":103,"reading_minutes":104},"palantir-agent-stack-python","Steal Palantir's agent stack: typed tools, one LLM gateway, swappable models","An X thread boils Palantir's AIP docs down to four agent patterns. We check each one against the docs, then build them in one stdlib-only Python file: typed business-object tools, a gateway that masks PII, caches and retries, a model set in config, and schedule\u002Fevent\u002FAPI triggers.",[96,97,98,99,100,13],"ai-agents","llm","python","architecture","palantir","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abd88f9c951ea7137fa3872-0-48eb7eee.png","2026-10-01T23:05:10.531Z","2026-10-01T23:05:10.532Z",10,{"slug":106,"title":107,"type":24,"summary":108,"tags":109,"author":14,"cover_url":113,"published_at":114,"updated_at":115,"reading_minutes":116},"claude-code-effort-levels","Effort levels in Claude Code: when max effort pays off and when it just burns tokens","Anthropic's effort deep dive (Terminal-Bench 3.0 plus three builds) shows higher effort mostly buys verification and edge-case testing, not smarter code. A rule of thumb per task type, the commands to set effort, and a script to measure cost vs pass rate on your own repo.",[110,111,97,112,13],"claude-code","ai-coding","developer-tools","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abd88edc951ea7137fa3804-0-2716005a.png","2026-10-01T22:32:14.139Z","2026-10-02T04:44:58.001Z",9,{"slug":118,"title":119,"type":24,"summary":120,"tags":121,"author":14,"cover_url":125,"published_at":126,"updated_at":127,"reading_minutes":128},"agentic-inbox-cloudflare-setup","Self-host an AI email agent on Cloudflare Workers: agentic-inbox set up and costed","Cloudflare's open-source agentic-inbox runs a full email client on Workers, with one SQLite Durable Object per mailbox and a Kimi K2.5 agent that drafts replies. Covers the post-deploy steps people miss (Access, sending, routing, mailbox first) and the cost.",[27,122,96,123,124,13],"workers","email","self-hosting","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abd88dac951ea7137fa378b-0-5b0ec96d.png","2026-10-01T08:17:18.646Z","2026-10-05T05:56:11.919Z",8,{"slug":130,"title":131,"type":24,"summary":132,"tags":133,"author":14,"cover_url":137,"published_at":138,"updated_at":139,"reading_minutes":128},"audit-ai-agent-public-traces","Nearly a million leaked links: auditing what your AI agents leave on the public web","OpenAI's agent swarm left almost a million public shortener URLs holding credentials. Here's a tested shell + gitleaks audit to find the shortlinks, pastes and webhooks your own agents created, scan them for secrets and close the channels.",[51,134,135,136,97,13],"agents","secrets","gitleaks","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abaad873e21d5cbd14d4397-0-0f4f51f7.png","2026-10-01T07:43:19.642Z","2026-10-01T08:44:43.041Z",{"slug":141,"title":142,"type":24,"summary":143,"tags":144,"author":14,"cover_url":148,"published_at":149,"updated_at":150,"reading_minutes":128},"mikrotrick-check-patch-mikrotik","MikroTrick: check and patch your MikroTik in 15 minutes","Two chained RouterOS bugs give anyone who can reach SSH full admin, no password needed, and attacks started before the patch. Find exposed SSH, check the version, grep for the published IoCs, patch and move management behind WireGuard.",[51,145,146,147,124,13],"mikrotik","routeros","ssh","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abcd9c0838b650cb96b3d10-0-6cd107e0.png","2026-10-01T07:02:15.03Z","2026-10-01T08:44:41.391Z",{"slug":152,"title":153,"type":24,"summary":154,"tags":155,"author":14,"cover_url":157,"published_at":158,"updated_at":159,"reading_minutes":128},"agent-sandbox-dns-egress-lockdown","Your agent sandbox leaks through DNS: lock down egress in 15 minutes","An OpenAI model escaped its sandbox by tunnelling questions through DNS. Here is a tested Docker Compose setup for coding agents: a DNS allowlist, a logging egress proxy and a kill switch that actually fires.",[51,50,134,156,124,13],"dns","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abaaca13e21d5cbd14d4306-0-cb93fe1b.png","2026-10-01T03:00:15.943Z","2026-10-01T08:44:43.165Z",{"slug":161,"title":162,"type":24,"summary":163,"tags":164,"author":14,"cover_url":168,"published_at":169,"updated_at":170,"reading_minutes":171},"who-blocks-ai-crawlers-robots-txt","Who blocks AI crawlers? robots.txt vs the network edge, with numbers","I scanned robots.txt on the top 300 sites: 33 of 138 block GPTBot, 14 block training but allow AI search. What each AI bot directive controls, why robots.txt is only a request, and a copy-paste policy plus nginx rule for small SaaS sites.",[87,165,166,27,167,13],"robots-txt","seo","saas","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abcd9c1838b650cb96b3d1b-0-11f94198.png","2026-09-30T21:00:20.673Z","2026-10-01T20:47:34.092Z",7,{"slug":173,"title":174,"type":24,"summary":175,"tags":176,"author":14,"cover_url":180,"published_at":181,"updated_at":182,"reading_minutes":77},"bullet-time-with-first-last-frame-video","Bullet time with first\u002Flast-frame video: orbiting a frozen moment from three stills","A freeze-frame camera orbit built from generated stills: one action shot, two camera-move angles, two first\u002Flast-frame clips between them, stitched and ping-ponged. The pipeline, the seams, and where the model re-imagines the water.",[87,177,178,179],"comfyui","video-generation","flowdsl","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abae46c45201648bfd477a7-0-6f4d036e.png","2026-09-28T22:42:41Z","2026-09-28T22:42:41.6Z",{"slug":184,"title":185,"type":24,"summary":186,"tags":187,"author":14,"cover_url":189,"published_at":190,"updated_at":191,"reading_minutes":171},"an-ai-media-pipeline-that-shows-its-work","An AI media pipeline that shows its work: ComfyUI presets, FlowDSL routing and the misses","How the images on my sites are generated: four ComfyUI presets behind one Go module, job rows as state, FlowDSL flows for routing, per-post media in the admin — and the bugs and model misses I hit shipping it. This post's own images were made the same way.",[87,179,177,188],"image-generation","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6aba8ba317ceba3543925be4-0-2f6b8a5c.png","2026-09-28T15:57:50Z","2026-10-01T20:47:34.327Z",{"slug":193,"title":194,"type":24,"summary":195,"tags":196,"author":14,"cover_url":199,"published_at":200,"updated_at":201,"reading_minutes":202},"openai-embeddings-python-mongodb","Transforming Text into Vectors: OpenAI Embeddings in Python","Learn how to generate text embeddings with the OpenAI API in Python to power semantic search, recommendations, and more. Includes practical examples with MongoDB integration and cost analysis.",[197,87,98,198],"openai","mongodb","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abad0fa45201648bfd46c2d-0-2e60b732.png","2024-11-23T00:00:00Z","2026-09-28T22:31:01.385Z",3,{"slug":204,"title":205,"type":24,"summary":206,"tags":207,"author":14,"cover_url":209,"published_at":210,"updated_at":211,"reading_minutes":212},"check-pricing-availability-ing-domains","Last Chance to Grab Short .ING Domains: The Extended List Part II","Welcome back to the second part of our exciting exploration into the .ING domain zone! This time, I've expanded our horizons to bring you an even larger selection of .ING domain names. List of over 24,000 domain names inside.",[83,208],"business","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abad0fa45201648bfd46c38-0-c55f4c8d.png","2023-12-14T00:00:00Z","2026-09-28T22:31:01.453Z",1,{"slug":214,"title":215,"type":24,"summary":216,"tags":217,"author":14,"cover_url":218,"published_at":219,"updated_at":220,"reading_minutes":221},"impressive-ing-domains","Unveiling the Impressive .ING Domains","Discover the vast potential of the new .ING domain zone in my latest blog post! I've used AI and a Python script to unearth a treasure trove of available domain names. From budget-friendly picks to exclusive premium domains, there's something for every ambition. Plus, a special list of unique, lesser-known domains awaits.",[83,208],"https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abad0fa45201648bfd46c43-0-03be24b7.png","2023-12-11T00:00:00Z","2026-09-28T22:31:01.527Z",2,{"slug":223,"title":224,"type":24,"summary":225,"tags":226,"author":14,"cover_url":229,"published_at":230,"updated_at":231,"reading_minutes":202},"secured-web-server-in-5-minutes","Fortify Web Server Security in 5 Minutes with Tailscale","Tailscale revolutionizes secure networking with its user-friendly approach, effortlessly connecting devices across diverse networks.",[227,72,228],"firewall","webserver","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abad0fa45201648bfd46c4e-0-eae6f62d.png","2023-11-03T00:00:00Z","2026-09-28T22:31:01.597Z",{"slug":233,"title":234,"type":24,"summary":235,"tags":236,"author":14,"cover_url":239,"published_at":240,"updated_at":241,"reading_minutes":19},"lets-encrypt-free-ssl","How to Secure Your Website with Free SSL Certificates for a Lifetime","Let’s Encrypt certificates have revolutionized internet security by providing free, automated, and widely trusted SSL\u002FTLS certificates. The non-profit Certificate Authority (CA) has significantly contributed to a more secure web environment by simplifying the process of securing websites with HTTPS.",[237,238,228],"ssl","https","https:\u002F\u002Fmedia.stufio.com\u002Fmedia\u002Fifcodes\u002Fmediagen\u002F6a\u002F6abad0fa45201648bfd46c59-0-bf2a9a0a.png","2023-11-01T00:00:00Z","2026-09-28T22:39:13.555Z"]