Gemini实现跨视频内容动态推理

Instead of scanning an entire file, Gemini reasons across the video’s transcript, audio, and frames,...

精选理由

GoogleDeepMind让Gemini能同时分析视频文本、音频和画面,只提取关键帧,长视频处理效率大增。

AI 摘要

Gemini不再扫描整个文件,而是基于视频转录文本、音频和帧进行推理,动态调整帧率提取所需时刻。这一方法在10分钟指南至多小时的长时间内容中效率提升显著。Agentic视频理解功能已通过API在Google AI Studio中面向3.7 Flash、3.6 Flash和3.5 Flash-Lite模型推出,并即将在Gemini App中上线。

原文 · Google DeepMind

Instead of scanning an entire file, Gemini reasons across the video’s transcript, audio, and frames,...

Instead of scanning an entire file, Gemini reasons across the video’s transcript, audio, and frames, dynamically adjusting the frame rate to pull the exact moments needed. The efficiency gains are most significant for long-form content, from 10-minute guides to multi-hour recordings. Agentic video understanding is rolling out to 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite via API in @GoogleAIStudio and coming soon in the @GeminiApp . Find out more → goo.gle/4gDKuGo 💬 0 🔄 1 ❤️ 21 👀 1414 📊 3 ⚡