Clairy 让 AI 决策过程透明化,解决了模型输出错误时只能靠猜的痛点,做 AI 安全、审计或模型调试的团队值得关注,可以直接用它来诊断和修正模型偏见。
Guide Labs 推出了首个可解释 AI 平台 Clairy,旨在解决 AI 的“黑箱”问题。该模型以文本块形式生成内容,用户可点击某一块查看模型生成时使用的概念(如“海洋生物”、“计算机科学”等)。Clairy 还提供训练数据归因功能,将生成的文本块与相似训练样本关联,便于诊断错误。此外,用户可通过概念引导直接增强或抑制特定概念,无需重写提示或重新训练模型。
This is brilliant. The first inherently interpret…
This is brilliant.
The first inherently interpretable AI platform just launched, "Clairy" by Guide Labs.
Attacks the "Black box" problem of AI.
The model generates text in chunks. You can click a chunk and see what concepts the model used to generate it.
With normal LLMs: if the model gives a wrong or biased answer, you mostly have to guess which words to change in the prompt.
Clarity changes that by trying to show the concepts the model is using while generating the answer, such as “marine life,” “African wildlife,” “computer science,” or “male role descriptions.”
i.e. you are not only seeing the final answer, you are seeing some of the hidden ingredients that pushed the model toward that answer.
Clarity also adds training data attribution, which connects generated chunks to similar training chunks so mistakes can be diagnosed instead of treated as mystery failures.
The new control layer is concept steering, where users amplify or suppress a concept directly, so, e.g. “marine life” can be raised without rewriting the question and unwanted concept families can be reduced without retraining.