后训练防护栏使LLM文本可被检测

LLMs could write like humans but post-training guardrails make their text detectable

精选理由

Pangram CTO说,不是LLM写不好,是安全防护栏限制了它们。没加防护栏的基础模型,写作风格反而更多样。

AI 摘要

Pangram CTO Bradley Emi认为,LLM无法形成可识别的写作风格并非能力不足。后训练和安全防护栏大幅缩小了其表达范围。未受约束的基础模型已展现出更丰富的写作多样性。

原文 · Decoder

LLMs could write like humans but post-training guardrails make their text detectable

LLMs don't write in a recognizable style because they can't do better. Post-training and safety guardrails sharply narrow their expressive range, argues Pangram CTO Bradley Emi. Base models without these constraints already write with far more variety. The article LLMs could write like humans but post-training guardrails make their text detectable appeared first on The Decoder .