Embedding Gemma 2: Multimodal Search & RAG (Free Colab)
Share

Post Content

 

 [[{“value”:”EmbeddingGemma 2 is a new open embedding model from Google DeepMind that puts text, code, images, video and audio into one shared vector space. It has 740M parameters, an Apache 2.0 license, and it’s small enough to run on a phone. In this video I explain why multimodal retrieval matters for RAG and search, how people solved it before (caption-and-transcribe pipelines, CLIP, ImageBind, ColPali), and how EmbeddingGemma 2 works under the hood: modality encoders, one shared Gemma 4 backbone, mean pooling, contrastive training, Matryoshka embeddings, and modular loading. Then we walk through a Colab notebook you can run on a free T4 GPU: photo search, multilingual search, voice search with no transcription, sound search, finding moments in a video, searching PDF pages without OCR, and how small you can make the vectors.

Let me know in the comments what you’d build with this, and whether you want a follow-up on fine-tuning EmbeddingGemma 2.

Colab notebook: https://colab.research.google.com/drive/1mXzvOCo-_y4r1yQ0AqAJJtU8BgU3LlCP
Model card: https://huggingface.co/google/embeddinggemma-2
Launch blog: https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/
Developer guide: https://developers.googleblog.com/embeddinggemma-2-the-developer-guide/
On-device guide (LiteRT, MediaPipe): https://developers.googleblog.com/google-ai-edge-with-embeddinggemma-2/
Gemini Embedding 2 paper: https://arxiv.org/abs/2605.27295
CLIP paper: https://arxiv.org/abs/2103.00020
ImageBind paper: https://arxiv.org/abs/2305.05665
ColPali paper: https://arxiv.org/abs/2407.01449

My ColPali videos:
https://www.youtube.com/watch?v=rhJJynv47Pw
https://www.youtube.com/watch?v=DI9Q60T_054

LocalGPT: https://github.com/PromtEngineer/localGPT

My voice to text App: whryte.com
Website: https://engineerprompt.ai/
RAG Beyond Basics Course:
https://prompt-s-site.thinkific.com/courses/rag
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0

💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off).

00:00 – EmbeddingGemma 2: Multimodal Embeddings
00:40 – Demo: Search Photos With Your Voice
01:20 – Why Multimodal Retrieval Matters for RAG
02:25 – Before: Convert Everything to Text
03:21 – Before: CLIP, ImageBind and ColPali
04:14 – One Backbone for Every Modality
04:33 – How It Works: Embeddings, Encoders, Tokens
06:07 – Mean Pooling and the 768-Number Vector
06:29 – Contrastive Training
07:15 – Matryoshka Embeddings: Smaller Vectors
08:05 – Modular Loading: 270M to 740M
09:00 – Google Colab Notebook
19:11 – Verdict

Credits: Big Buck Bunny © Blender Foundation (CC BY 3.0). ESC-50 by K. Piczak (CC BY-NC 3.0). Flickr30k test images. Benchmark numbers in the explainer are Google’s; results in the notebook section are from my own Colab T4 run.”}]] Read More Prompt Engineering 

#Promptengineering #AI

By ali

Leave a Reply