omar sar教你怎么用语音+屏幕操作提示Agent,比纯文字提示聪明多了,能省下大量调试时间。
研究员omar sar分享了多模态提示工作流,通过录制语音、屏幕注释、鼠标点击等输入,预处理后传递给Agent,显著提升任务完成效率。该方法已为他节省数小时工作时间,减少与Agent的挫败交互。他将这些录制的任务作为可复用数据集,不断改进并打包成工作流/模式/技能。该技巧应用于Web开发、设计、原型制作、研究等多个场景。
Multimodal prompting is clearly the future. I love experimenting with new ways to interact with ag...
Multimodal prompting is clearly the future. I love experimenting with new ways to interact with agents. As a researcher and engineer, I've found that the richer the inputs to the agent and the richer the outputs I consume, the better the overall results of the collaboration. In this little walkthrough, I go over what I mean by a multimodal prompt and when you might find it useful. It's more than simple text prompting, so I call it a "task" for lack of a better word. It helps me record my voice, annotate the screen, click/mouse actions, and more. Then all of that is preprocessed and passed to the agent to complete the task more efficiently. The agent has the high-level prompt, but it also has the raw transcriptions if needed. So naturally, I am also using this to build out multimodal skills that I reuse in workflows where agents tend to struggle. This has saved me hours of work. And even the older models are pretty great at understanding the tasks more clearly. Some noise is introduced in the process, but it doesn't seem to hurt the performance. I've also found that this new way of prompting has reduced the number of frustrating interactions I have with agents. This is something I have been thinking about for some time now because we are going to move more into multimodal AI models. And so the interactions are going to evolve with models being able to handle a variety of modalities natively. Currently, I process all the recorded tasks with another model in the background, but it's not crazy to think that all of it will just be naturally consumed by omnimodel in the future. All of these recorded tasks (which you can also think of as rich annotated datasets) are things I mine and recursively improve over time and, in some cases, package as reusable workflows/patterns/skills. This process has really elevated how I use coding agents for all kinds of work. I use multimodal prompting in things like web development, designing, artifact creation, prototyping, researching, reading, simulations, AI-assisted writing, and much more. So it's not just about prompting. It's going deeper into understanding and exploring the right level of detail that the agents need to make the right decision and to push/maximize their capabilities. Your browser does not support the video tag. 🔗 View on Twitter 💬 0 🔄 2 ❤️ 11 👀 1432 📊 3 ⚡