EmbeddingGemma 2 Local Retrieval Test: Mac 270M, PDF/DOCX, 256d vs 768d

A reproducible, bounded Mac CPU LiteRT-LM pilot: convert 16 published bilingual technical articles into 8 PDFs and 8 DOCX files, extract 278 chunks and evaluate EmbeddingGemma 2 270M against SQLite FTS5 and RRF on 32 authored questions. Not a natural-document, iPhone, or production RAG benchmark.

Published · 2026-10-108 min readXBSTACK
Evidence 7 cited sources
On this page

Direct answer: EmbeddingGemma 2’s quantized text-only 270M variant ran locally on Apple Silicon Mac with LiteRT-LM CPU, and it retrieved the right source documents from extracted PDF/DOCX text in a small, reproducible pilot. Sixteen published bilingual technical articles were converted into eight PDFs and eight DOCX files, extracted back into 278 text chunks, and searched with 32 English/Chinese queries. On this dataset, the 256d and 768d embeddings each achieved document-level Recall@5 of 1.0000 and MRR of 0.9688. Adding equal-weight FTS5 and RRF did not improve MRR.

The limitation belongs in the opening, not a footnote. Those PDFs and DOCX files were generated from real published articles rather than collected as natural business documents. Queries and relevance labels were written by the author. We did not evaluate scanned documents, OCR, evidence-passage precision, a reranker, full RAG answers, the production RecalAI application or a physical iPhone 13. This is a bounded engineering pilot, not a model leaderboard or production benchmark.

1. The decision: does a smaller local embedding index still find the right documents?

A practical knowledge base cannot optimize vector accuracy in isolation. Readers import files, expect immediate keyword search, then ask questions that may use different words or languages from the original document. Developers must choose the model footprint, index dimensions, extraction quality and search strategy without locking older devices into slow foreground indexing.

Google released EmbeddingGemma 2 on October 6, 2026. The full family covers text, image, audio and video embedding, but the checkpoint in this article is the LiteRT-LM 270M text-only quantized package, not the larger 740M multimodal variant. It produces 768-dimensional vectors that can be truncated and re-normalized at 512d, 256d and 128d through Matryoshka Representation Learning (MRL). Google’s own benchmark tables and hardware results describe Google’s evaluation settings; none of them are XBSTACK measurements. Google’s LiteRT-LM runtime documentation covers the separate deployment toolchain.

We focused on two questions: What did 256d give up relative to 768d on a small bilingual document collection? Did combining FTS5 with semantic ranking actually help?

2. The three tasks, environment and dataset

ComponentActual run
MachineApple Silicon Mac, Darwin arm64
RuntimePython 3.12.13, LiteRT-LM 0.18.0, CPU with four threads
Modelembeddinggemma-2-text-270m.litertlm, quantized text-only
Model SHA2562d079ee2f6f066b1f368e8d7c819f55214eaef1d0513b312321901f30ab286fb
Original contentEight XBSTACK technical topics, 16 published English/Chinese Markdown articles
Generated formatsEight PDFs and eight DOCX files
Parserpypdf and python-docx; not OCR
Chunks850 characters, 120-character overlap, 278 total
Questions16 basic + 16 harder rephrasings authored before the final run
BaselinesSQLite FTS5 unicode61, FTS5 trigram, pure vector, equal-weight RRF
OutcomesParent-document Hit@1, Hit@5, Recall@5 and MRR
ExcludedNatural-document OCR, tables, passage labels, reranker, production RAG and iPhone tests

We used source articles covering large MCP tool responses, n8n error workflows, LangGraph human approval and checkpoint storage, Google ADK memory deletion, Agent context cost, Responses API stream aborts and tool-authorization boundaries. Source paths, article hashes, parser output lengths and model checksum were written to a machine-readable evidence file.

Task A asked whether file creation and parsing actually worked, not whether a filename extension changed. Task B evaluated document retrieval from the parsed chunks. Task C compared 768/512/256/128-dimensional indexes and storage on the same original embeddings. Keeping these tasks distinct prevents a high semantic-search score from being misreported as PDF vision or a correct final AI answer.

Document conversion and local retrieval from authentic published bilingual articles: temporary PDF and DOCX files, actual parsing, FTS5 and EmbeddingGemma 2

3. Task A: parse converted PDF and DOCX files into searchable chunks

The input was not a synthetic five-sentence toy dataset. The original articles are published technical Markdown with real domain details and varying lengths. Our script stripped frontmatter, fenced code and image markup, and used ReportLab to construct real PDF files and python-docx to construct real Word files. It then separately opened those files with pypdf and python-docx and extracted their textual content.

All 16 conversions cleared an extraction-length sanity check. Source texts ranged from roughly 4,154 to 29,135 characters. Some extracted strings were slightly longer because conversion and text extraction changed whitespace; that is not evidence of extra content being recovered. The resulting normalized text was split into overlapping 850-character windows with a 120-character overlap, producing 278 chunks.

That test verifies an actual text-extraction boundary. It does not verify scanned PDF OCR, diagrams, native image understanding, handwriting, complex table reading order, natural financial documents or code-block exact matching. In fact, removing fenced code means code retrieval is outside the evaluation scope. Engineers should treat all of those as independent ingestion acceptance tests.

4. Task B: lexical search, pure embedding and RRF compared fairly

We used exactly the same 32 authored queries and the same 278-chunk corpus for every search strategy. Each query had two relevant parent documents—the English and Chinese articles on the same topic. We ranked chunks, deduplicated by parent document, and then measured parent-document quality.

This matters when interpreting the scores: Hit@5 asks whether the first five documents include at least one relevant parent. Recall@5 asks what fraction of the two related documents were retrieved. MRR uses the position of the first relevant parent. None is equivalent to retrieving the exact answer-bearing passage.

MethodHit@1Hit@5Recall@5MRR
FTS5 unicode610.62500.87500.78120.7396
FTS5 trigram0.65621.00000.78120.7755
Embedding 768d0.93751.00001.00000.9688
Embedding 512d0.93751.00001.00000.9688
Embedding 256d0.93751.00001.00000.9688
Embedding 128d0.87501.00000.98440.9323
unicode61 + 768d RRF0.93751.00000.98440.9688
trigram + 768d RRF0.87501.00000.90620.9375

Document-level MRR from 32 bilingual queries: FTS5 unicode61/trigram, embedding 768d/256d/128d and fused RRF

The failure case is important: RRF was not an automatic win. Unicode61 plus 768d tied the pure 768d model’s MRR, but its Recall@5 fell slightly. Trigram plus 768d also had lower MRR than embedding-only. These are results from one default equal-weight RRF setting, not a tuned production fusion pipeline. Lexical tokenizer behavior, rank cutoffs, weights and reranking may change the outcome; no model should be declared generally superior on these scores alone.

The other clear weak point was lexical Chinese search. FTS5 unicode61 tokenization was not customized for Chinese morphological segmentation. The trigram baseline caught more queries in the first five, but still did not recover the same fraction of the bilingual documents. This is a baseline design choice, not proof that lexical retrieval is dispensable: it often remains essential for exact identifiers, filenames, codes and immediately available first-import search.

5. Task C: what 256d actually saves over 768d

On these 32 queries, the embedding-only 768d, 512d and 256d document rankings each produced Hit@1=0.9375, Recall@5=1.0000 and MRR=0.9688. At 128d, Hit@1 dropped to 0.8750 and MRR to 0.9323. These observations justify further investigation of 256d, not a blanket claim that 256d and 768d are equivalent on real production knowledge bases.

Our implementation calculates the original 768-dimensional vectors once, slices each vector to the requested dimension and L2-normalizes both documents and queries again. That is the important part of applying MRL truncation correctly:

import numpy as np

# vectors: original 768-dimensional document vectors
# qvec: original 768-dimensional query vector
dim = 256
index = vectors[:, :dim]
index = index / np.maximum(
    np.linalg.norm(index, axis=1, keepdims=True), 1e-12
)
query = qvec[:dim]
query = query / max(float(np.linalg.norm(query)), 1e-12)
scores = index @ query

For this 278-vector corpus, raw float32 coordinate storage only was:

DimensionVector bytesRelative to 768d
768d854,016100%
512d569,34466.7%
256d284,67233.3%
128d142,33616.7%

Those numbers exclude SQLite or ANN index overhead, IDs, original files, metadata, model weights and loaded runtime buffers. Going from 768d to 256d saves two-thirds of the raw float32 vector array; it does not reduce the 270M model file to one-third or prove a threefold improvement in inference speed. Whenever the model checkpoint, quantization, normalization or dimension changes, keep indexes versioned and rebuild or isolate incompatible embeddings.

6. Actual Mac CPU encoding times—not mobile performance

The instrumented Mac CPU execution produced:

StepObserved
Model initialization286.37 ms, one launch
278 document-chunk embeddings53.99 seconds total
Mean embedding time per chunk194.22 ms
Mean query encoding time across 32 queries79.08 ms
Script-sampled query encoding p9582.47 ms

These are not a benchmark of complete UI search latency. Search-service scheduling, database IO, FTS query time, reranking, local generation, energy consumption and cold-start variance are excluded. A single launch and one measured query set do not establish sustained performance or an SLA.

A physical iPhone 13 was not connected and tested. Even though the model family advertises on-device use and Google publishes device-specific measurements, this article does not inherit any of those numbers as XBSTACK results. For a local-first app, the safer product workflow is to make normalized text immediately available through FTS while asynchronous embedding catches up in the background. That architectural decision still requires application-level validation on each target device.

7. Bilingual search: does it recover the opposite-language document?

Our corpus includes English and Chinese versions of each topic. In a separate metric, we treated only the matching document in the other language as relevant. Under this stricter setup, 768d obtained cross-language Hit@5=1.0000 and MRR=0.5000; 256d had Hit@5=1.0000 and MRR=0.5130. FTS5 unicode61 recorded Hit@5=0.6875 and MRR=0.3087.

That metric needs careful interpretation. The same-language version often ranks above its translation. A cross-language article ranking second does not mean the model failed semantic matching; equally, 16 bilingual source documents and a small authored question set are not independent proof of reliable multilingual retrieval across domains.

Real-world acceptance should include typos, terminology, entity IDs, long factual sections, queries with no answer and distractor documents about adjacent topics. The missing key is human judgment of answer-bearing evidence chunks, not merely a parent-document label.

8. Reproduction assets, tests and unsupported claims

We executed actual quantized LiteRT-LM inference, recorded the exact model SHA, hashed each published source, and retained full per-query parent rankings and resource measurements in local project files:

# From the XBSTACK blog repository
uv pip install --python research/embeddinggemma2/.venv-litert/bin/python \
  -r research/embeddinggemma2/requirements-document-pilot.txt

research/embeddinggemma2/.venv-litert/bin/python \
  research/embeddinggemma2/document_format_chunk_pilot.py

research/embeddinggemma2/.venv-litert/bin/python -m unittest \
  discover -s research/embeddinggemma2/tests -v

The local research test suite passed 20 out of 20 tests. The raw evidence file is research/embeddinggemma2/evidence/mac-litert-converted-pdf-docx-chunks-2026-10-10.json. The runnable script, dedicated dependencies and complete engineering report are under research/embeddinggemma2/.

The public XBSTACK EmbeddingGemma 2 reproduction repository now contains the actual PDF/DOCX conversion script, complete 32-query evidence, original 16 public source articles with verified SHA256 hashes and research tests. See the verified October 10 upload commit. Model weights are not redistributed; obtain them separately from the official model repository.

Other limitations include no independently sourced natural PDF/DOCX set, scanned OCR, image content, complex layouts, manually annotated relevant passages, reranker, generated-answer attribution, production RecalAI pipeline or on-device iPhone power/thermal tests. The question set was authored rather than collected from blind external users, and rankings were evaluated on one small selected topic corpus. There are no confidence intervals or statistically supported general win rates.

9. Adoption decision: promising for a prototype, unproven for production

A reasonable prototype candidate: primarily text-based offline knowledge search, limited vector storage, developer control over document parsing, and willingness to version embeddings. The 270M text-only quantized model with 256d vectors is worth testing further because this pilot preserved the 768d document-level scores while storing one-third as many float32 coordinate bytes.

A poor reason to migrate immediately: a production library dominated by scanned PDFs, tables, OCR errors, source-code snippets, precise legal/financial evidence, or stringent citation correctness. The same applies to iPhone hardware whose CPU, memory, background behavior and thermals have not been measured. A real migration decision needs representative held-out documents, passage-level human judgments, comparison with the existing embedding model and a fair FTS/reranker baseline.

The useful outcome is not a universal model winner. It is a reproducible boundary: on 32 authored queries over 16 converted documents, 256d matched 768d document retrieval, 128d lost some ranking quality, and equal-weight RRF did not improve embedding-only MRR. Production design begins with verifying where those statements stop being true.

Next step: continue the Agent system

Full decision path →

Choose the next architecture, tool, memory, evaluation, security, or production decision instead of starting over.

Continue with state and memoryFunes Agent Memory Tested: Codex Recall, Stale Memory, and Local Privacy BoundariesLocal Funes 1.3.0+dev test on real Codex traces: 5/5 Hit@1 and Hit@5, 3.14s mean recall, plus unrelated-query, stale-memory, scrub, and privacy limits.Continue with tool and protocol boundaries2026 Full-Stack Productivity Build Guide: Assembling a Sovereign Individual's Rigorous Tool Matrix2026 Full-Stack Productivity Build Guide: 2026 Full-Stack Productivity System Selection Report. This article breaks down the essential build list for beginners, revealing how to esContinue with tool and protocol boundariesGemini 3 and MCP Protocol in Practice: A Hands-On Guide to Building a Local AI Financial Audit SystemGemini 3 and MCP Protocol in Practice: A practical demonstration of refactoring a financial audit system using the MCP protocol, handling 2M+ token contexts, and enabling second-le
Comments & evidence

DISCUSSION

Questions, verification and corrections

Sign in to comment. Every new comment is reviewed before publication; while pending, it is visible only to you and the administrator.

Sign-in required Reviewed before public
Loading the discussion…