Google 开源 EmbeddingGemma 2:多模态嵌入模型可在手机本地运行
Google 开源了个 7.4 亿参数的嵌入模型,一句话就能在手机上搜照片视频和录音,还不用联网,免费商用,做本地搜索的可以去试试。
Google 发布 EmbeddingGemma 2,参数量 7.4 亿,运行内存占用 191MB 到 567MB,手机和电脑都能本地跑。这是 Google 首个原生多模态开源嵌入模型,把文字、代码、图片、视频、音频放进同一套向量空间,支持 8K 上下文,相当于 5.5 分钟音频或 29 张图片。上一代 EmbeddingGemma 只支持文字,这一代可以跨格式检索,Google 称其性能超过一些体量大两倍以上的专业模型。模型基于 Gemma 4 架构,可与 Gemma 4 搭配做设备端 RAG,文件不出设备。权重以 Apache 2.0 协议开源,可在 Hugging Face 和 Kaggle 下载,允许免费商用。
Google 发布 EmbeddingGemma 2:在手机上用一句话搜照片、视频和录音 Google 开源了 EmbeddingGemma 2,这是一个能在手机和电脑本地运行的多模态嵌入模型。文字、图片、视频、音频和代码都能放进同一套“坐标系”,所以你可以用一句话找到一段视频,整个过程不用联网。 嵌入模型(embedding model)是搜索背后的一层技术:它把内容转成一串数字,意思相近的内容,数字也相近。上一代 EmbeddingGemma 只处理文字,这一代把几种格式统一了起来,可以跨格式找东西。比如对着手机说一句话,找出相册里对应的视频片段;或者输入一行字,在几个小时的录音里找到某段对话。 模型有 7.4 亿参数,运行时占用 191MB 到 567MB 内存,手机上跑得动。Google 说它在同体量模型里表现最好,还超过了一些比它大一倍多的模型。它一次能处理的内容是上一代的 4 倍(8K 上下文),相当于 5.5 分钟音频、29 张图片或 58 帧视频。 它可以和 Google 的开源大模型 Gemma 4 搭配,在设备上做 RAG(检索增强生成,先查资料再回答)。EmbeddingGemma 2 从你的本地文件里找出相关内容,Gemma 4 读完后给出答案。文件不离开设备,合同、病历这类不想上传到云端的资料也能这样处理。 它基于 Gemma 4 架构,用 Apache 2.0 协议开源,可以免费商用。现在能在 Hugging Face 和 Kaggle 下载,Google 企业 AI 平台的模型库稍后上线。 Sundar Pichai @sundarpichai Introducing EmbeddingGemma 2, a new open multimodal model that sets the standard for on-device efficiency. - our first open, natively multimodal embedding model - handles text, code, image, video, and audio tasks within a lightweight, modular 740M parameter form factor - ideal for offline, privacy-first RAG when paired with Gemma 4 - outperforms some specialist models more than twice its size Weights available now on Hugging Face. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 6 🔄 3 ❤️ 17 👀 2323 📊 8 ⚡