这篇文章说AI数学强在记性好而不是脑子好,用上下文窗口碾压人类工作记忆,角度挺新鲜。
一项针对AI数学能力的分析认为,其优势来自近乎无限的符号工作记忆,而非更强的推理能力。人类数学家受限于工作记忆容量,而AI的上下文窗口能同时容纳整个问题、数百个中间方程和多种尝试路径。多项研究表明,控制智力水平后,工作记忆仍能独立预测数学成绩。因此AI看似聪明,实则只是用外部草稿纸绕过了人类生物瓶颈。
2026 08 17 HackerNews
2026-08-17 Hacker News Top Stories # AI 在数学上的优势源于无限工作记忆而非更强推理,人类工作记忆瓶颈限制了数学表现。 Firefox for iOS 新增基于 EasyList 的原生广告拦截功能,默认关闭,处于实验阶段。 Claude 官方文档介绍系统提示的用途和更新历史,用于引导模型行为。 2026 年超级厄尔尼诺迅速增强,海温异常超 5°C,将影响全球天气。 研究论文中出现“肾脏失望”替代“肾脏衰竭”,可能源于逃避抄袭检测或翻译工具局限。 腹部脂肪(腰围/腰臀比)比 BMI 更能预测心脏病风险,临床应评估中心性肥胖。 AI 时代软件工程基础更重要,LLM 不真正推理,需精心设计抽象和可维护性。 创造力需要独处保护脆弱的新想法,数学家格罗滕迪克和导演伯格曼的笔记展示了这种状态。 LittleLearner 模型训练数据仅限 K-5 内容,发现预训练数据过滤决定能力上限,无法超越所学范围。 存在转售未使用 AI API 积分的灰色市场,提供高折扣,存在滥用风险。 1. AI 并非比数学家更会思考,而是比他们更会记住。 (AI isn’t outthinking mathematicians, it’s out-remembering them) # https://davidepiffer.com/p/ai-isnt-outthinking-mathematicians AI 在数学问题上的优势可能并非更强的推理能力,而是拥有近乎无限的符号工作记忆。人类工作记忆容量极其有限,而 AI 的上下文窗口可以同时容纳整个问题陈述、数百个中间方程、多种尝试路径、定义和约束条件。这种差异在数学领域尤为关键——人类数学家只能同时处理少量不熟悉元素,而 AI 能像使用巨大笔记本一样外化推理过程。 工作记忆对数学表现有独立于 IQ 的预测作用。多项研究表明,在控制智力水平后,工作记忆能力仍能显著预测数学成绩。这意味着 AI 的数学表现部分源于消除了人类生物限制:它不需要像人类那样依赖“分块”压缩信息,而是直接保留大量显式符号。 因此,AI 看似更“聪明”的数学推理,实际上可能只是因为它不受人类工作记忆瓶颈的制约。它的上下文窗口相当于一个巨型外部草稿纸,使得复杂数学问题的处理方式发生了根本性改变。 HN 热度 587 points | 评论 484 comments | 作者:rzk | 1 day ago # https://news.ycombinator.com/item?id=49312845 很多所谓的聪明其实是比周围人记得更多,能结合不同领域的知识。 有人通过快速工作然后反向验证来发现错误,类似 FABRIK 算法。 有人通过多次检查试卷,利用矛盾来修正答案,从而在考试中学习。 有人记忆差但概念记忆好,能压缩信息为关键词和感觉。 有人用“弱细节回忆但强流体推理”来描述自己的认知特点。 2. Firefox for iOS 现已内置原生广告拦截器 (Firefox for iOS now has a native adblocker) # https://support.mozilla.org/en-US/kb/block-ads-firefox-ios Firefox for iOS 提供了一个可选的广告截功能,旨在减少浏览时遇到的不必要广告该基于 EasyList 的过滤列表,能够在广告加载之前阻止许多广告。默认情况下,广告拦截功能是关闭的,用户可以根据需要选择是否启用。 目前,该功能仍处于实验阶段,正通过渐进式推广向 Firefox 用户群体推出,因此可能并不是所有用户都能使用。 广告拦截功能的主要特点包括: 启用后,Firefox 能在网络层面阻止多种类型的广告,包括: 第三方广告网络和交换 广告相关的追踪器 许多网站提供的第三方广告 侵入式广告格式,例如弹出广告和覆盖层广告 需要注意的是,由于网站使用不同的广告技术,仍然可能会出现一些广告无法被拦。 Firefox 不会阻止以下内容: 搜索引擎结果页面上的广告,包括 Google、Bing、DuckDuckGo 等搜索引擎的广告。 Firefox 主页面或新标签页上的赞助内容。 如何启用或禁用广告拦截功能: 打开 Firefox for iOS,点击菜单按钮,选择设置。在内容部分将广告拦截器打开或关闭。 用户也可以在浏览网页时通过站点菜单快速启用或禁用广告拦截功能。 在启用广告拦截时,功能界面会有相应的显示,而禁用时则会有不同的标识。 常见问题解答: 为什么我仍然看到一些广告?因为 Firefox 只拦截许多第三方广告,某些广告可能通过当前的过滤列表无法被识别和拦截。 广告拦截是否会影响搜索引擎广告?不会,搜索引擎结果页面上的广告不会被拦截。 广告拦截会阻止 Firefox 的赞助内容吗?不会,用户可以在主页设置中管理这些赞助内容。 我可以随时关闭广告拦截功能吗?可以,用户可以随时通过设置或在浏览时通过站点菜单启用或禁用广告拦截功能。 此外,文中还提到了一相关的文章,例如在 Firefox for iOS 中的增强追踪保护和尊重广告的工作原理。 HN 热度 504 points | 评论 208 comments | 作者:pentagrama | 11 hours ago # https://news.ycombinator.com/item?id=49319633 评论区提到了 Safari 版 uBlock Origin Lite,但受限于 Manifest V3,存在规则限制、默认无 cosmetic 过滤和脚本注入等问题。 有观点指出主流浏览器中只有 Firefox 仍能运行完整版 uBlock Origin,其他浏览器(包括 iOS Firefox)均受 Manifest V3 限制。 部分用户认为 Brave 浏览器保留了广告拦截功能,但也有评论质疑 Brave 的隐私承诺,列举了其多项不当行为。 还有评论认为 Brave 的行为相比 Mozilla 更恶劣,因为后者从未有过类似丑闻。 有人指出 uBlock Origin Lite 在 iOS 上不如 AdGuard 或 Wipr,因为它只能拦截 Safari 内的广告,而后者能拦截系统级 Web 视图的广告。 有评论推荐 wBlock(MIT 学生开源),称其与 uBlock 同样可信,使用 Safari 内容拦截 API,无数据收集。 有观点不信任学生作者,认为 uBlock 团队有长期记录和更多外部审查。 另有评论认为作者年龄不重要,应关注经验和团队规模。 一些用户指出 iOS Firefox 与 Safari 同样受限,且更新绑定系统版本。 有评论表示 Brave 和 Vivaldi 使用自建拦截器而非扩展实现广告拦截。 提到了 uBO Lite 在 Safari 中可无需数据权限工作,但 Google 以此为由淘汰 MV2 的做法遭质疑。 有用户认为整日依赖 Google 资助的 Mozilla 比 Brave 的行为更糟。 3. Claude 系统提示 (Claude: System Prompts) # https://platform.claude.com/docs/en/release-notes/system-prompts 这个网页是 Claude 官方文档中关于系统提示(System Prompts)的页面 ,主要内容如下: 核心说明: Claude 的网页版(claude.ai)和移动端应用会在每次对话开始时,通过系统提示向 Claude 提供最新信息(如当前日期),并引导其行为(如始终用 Markdown 格式提供代码)。 系统提示会定期更新以改善回复质量,但这些更新 不适用于 Claude API 。 从 Claude 4.6 代开始,每个模型 ID 是独立的固定快照,因此每个模型只有一个系统提示版本。 各模型系统提示更新日期一览(按时间倒序): Claude Opus 5 — 2026 年 7 月 24 日 Claude Fable 5 — 2026 年 6 月 9 日 Claude Opus 4.8 — 2026 年 5 月 28 日 Claude Opus 4.7 — 2026 年 4 月 16 日 Claude Sonnet 4.6 — 2026 年 2 月 17 日 Claude Opus 4.6 — 2026 年 2 月 5 日 Claude Opus 4.5 — 2026 年 1 月 18 日 / 2025 年 11 月 24 日(两个版本) Claude Haiku 4.5 — 2026 年 1 月 18 日 / 2025 年 11 月 19 日 / 2025 年 10 月 15 日(三个版本) Claude Sonnet 4.5 — 2026 年 1 月 18 日 / 2025 年 11 月 19 日 / 2025 年 9 月 29 日(三个版本) Claude Opus 4.1 — 2025 年 8 月 5 日 Claude Opus 4 — 2025 年 8 月 5 日 / 2025 年 7 月 31 日 / 2025 年 5 月 22 日(三个版本) Claude Sonnet 4 — 2025 年 8 月 5 日 / 2025 年 7 月 31 日 / 2025 年 5 月 22 日(三个版本) Claude Sonnet 3.7 — 2025 年 2 月 24 日 Claude Sonnet 3.5 — 2024 年 11 月 22 日 / 2024 年 10 月 22 日 / 2024 年 9 月 9 日 / 2024 年 7 月 12 日(四个版本) Claude Haiku 3.5 — 2024 年 10 月 22 日 Claude Opus 3 — 2024 年 7 月 12 日 Claude Haiku 3 — 2024 年 7 月 12 日 HN 热度 501 points | 评论 212 comments | 作者:tosh | 11 hours ago # https://news.ycombinator.com/item?id=49319556 系统提示词被公开后,可以通过 git 提交历史追踪变化,其中新增了关于模型被暂停和恢复访问的说明,但缺少工具定义和 Claude Code 的提示词,这些更值得关注 讽刺的是,告诉 Opus 5 它比 Fable 和 Mythos 低一档,可能反而导致性能下降,而 Opus 4.8 认为自己是最强的 系统提示词占用上下文窗口中最关键的数千 token,可能对性能产生负面影响,应该把空间留给任务本身 官方系统提示词会被预置到对话中,即使自定义提示词也无法绕过,这是防止滥用的一种手段 安全防护主要靠专门训练的检测模型和外部匹配技术,系统提示词只是防御纵深,不是唯一防线 没有系统提示词时模型甚至不知道自己是什么版本,会幻觉成旧版 Sonnet 系统提示词只是建议,不能作为唯一的安全护栏,需要结合微调和带外检测 不应把限制写进系统提示词,微调和带外检测更合适 官方文档显示他们不依赖系统提示词做安全限制,而是用其他系统拒绝和重路由请求 非 LLM 网关才是安全防护的关键,否则早期越狱方式仍会轻易奏效 好奇模型能否后台开浏览器访问 claude.ai 去了解同级模型的排名,但前端有自动化防护 系统提示词越长越影响性能,应该用简短精准的提示,参考陶哲轩与 ChatGPT 的对话记录 对于复杂项目,较长的系统提示词值得投入,从简短开始逐步调整以减少错误 系统提示词中浪费了大量空间在无用内容上 担忧 HN 论坛正在移除对 AI 有负面含义的帖子,包括联合国关于 AI 威胁自然资源的文章等 4. 超级厄尔尼诺持续增强,新预报峰值进入历史新高,冬季前势头不减 (Super El Niño Keeps Growing as New Forecasts Reach Record Territory Ahead Winter) # https://www.severe-weather.eu/long-range-2/super-el-nino-growth-accelerating-to-record-strength-fall-winter-2026-2027-forecast-impact-united-states-canada-europe-fa/ 该网页是一篇关于 2026 年超级厄尔尼诺现象的专业气象分析文章。文章指出,超级厄尔尼诺正在迅速增强,新的长期预报已将峰值强度推至历史新高。 核心数据: 赤道太平洋海面温度异常超过 5°C,局部达 6°C 以上。 次表层海洋存在强开尔文波,温度异常达 9°C 以上。 西风异常在过去 86 年数据中创下 6-7 月同期最强记录。 发展机制: 文章解释称,西风爆发将暖水堆积至西太平洋,形成次表层开尔文波并向东传播,最终升至海面,为超级厄尔尼诺持续提供热能。当前能量特征超过历史上大多数超级事件。 影响预测: 超级厄尔尼诺将改变全球大气环流,影响 2026-2027 年秋冬季天气。对北美而言,会导致太平洋槽加深、加拿大脊增强,形成活跃风暴路径,美国南部更暖湿,北部相对温和。对欧洲的影响也有相应预测。 HN 热度 411 points | 评论 312 comments | 作者:dgellow | 1 day ago # https://news.ycombinator.com/item?id=49313428 历史上最强的厄尔尼诺曾引发大规模饥荒。 如今有世界粮食计划署等机制,可提前数月预测干旱和作物歉收,但能源危机和化肥危机可能加剧负面影响。 所谓“提前预测”往往让大公司受益,它们会囤积物资并在后期高价出售。 许多国家的农业补贴使粮食在多数年份过剩,以保证歉收年份人们不挨饿,但企业可能压低收购价,使农民无法享受补贴好处。 期货市场的本质是让生产者提前锁定价格、转移风险,让供应链上下游分担丰歉年份的盈亏,而不是让投行从穷人口粮中牟利。 期货市场的核心功能是把风险转移给更有能力或意愿承担的人,对供需的影响只是副产品。 天然供给不可能完全可靠,即使古代法典也承认“天灾”,现代合同仍保留“不可抗力”条款。 自汉谟拉比时代以来,食品生产、储存和分配已有巨大进步。 “天灾”似乎只让保险公司受益,而不是个人。 极端风险事件概率极低,普通人无法通过广泛分散投资来平滑影响,而大机构可以通过大量持仓把高波动转化为低波动。 保险公司如今不能承受过多风险,会大量再保险,并借助巨灾债券分散风险。 保险的意义就是应对小概率高损失事件;拒赔多因投保人未看清条款,或事后对保障范围有误解。 巨灾风险高度相关会破坏承保模型,所以疫情风险常被排除在商业保险之外。 许多“天灾”其实都能获赔,如冰雹、大风;洪水不保是因为居住在洪泛区的人少且不愿承担全部风险,而区外人不买洪水险,导致逆向选择。 保险公司每年支付超过两万亿美元的理赔款。 期货市场的意义在于尽可能让供应匹配需求,让生产者与使用者对未来有合理预期,而高盛等不应从穷人的日常食物中牟利。 任何系统的实际作用就是它实际产生的结果,如果期货市场能让银行从面包上赚钱,那这就是它的作用。 最初芝加哥面粉期货市场的设立者并无意让金融中介从粮食中牟利。 国家补贴农业的另一重要原因是在战争时不被敌人用饥饿扼制,这主要针对高热量主食而非精致蔬果。 不打仗会更容易;北约中除加拿大和美国外,多数国家都是主要粮食进口国。 北约中法国、荷兰、西班牙、丹麦、波兰、匈牙利、立陶宛、土耳其、意大利和爱尔兰并非主要粮食进口国,美国也处于临界状态。 意大利并非粮食自给,严重依赖谷物、饲料、油等进口,本国仅果蔬等基本自足,连意大利面原料都依赖加美谷物,且生产依靠进口化肥化学品,孤立时会挨饿。 爱尔兰不在北约内。 美国如果减少一半卡路里摄入,健康水平可能会更好。 5. 研究论文中使用‘肾脏失望’代替‘肾脏衰竭’ (Research papers using “kidney disappointment” instead of “kidney failure”) # https://scholar.google.com/scholar?q=%22kidney+disappointment%22 Google 提示抱歉,您的计算机或网络可能正在发送自动查询。为了保护用户,目前无法处理您的请求。请参阅 Google 帮助了解更多信息。页面底部有 Google 主页链接。 HN 热度 350 points | 评论 128 comments | 作者:Alifatisk | 12 hours ago # https://news.ycombinator.com/item?id=49319389 有网友认为使用 “肾脏失望” 代替 “肾脏衰竭” 是通过同义词替换来逃避抄袭检测。 一些指出这些不自然的词汇组合可能来自翻译工具,导致奇怪的表达。 还有人提到这可能与翻译的局限性有关,特别是对于非母语使用者。 有评论提到学术界普遍使用抄袭检测工具,导致作者为了避免抄袭而进行词语替换。 一些网友提到 “肾脏望” 一词最早出现在 2021 年的论文中,暗示作者可能并不熟悉英语。 网友们讨论了文本生成技术的发展,以及这些工具如何被用来创造可读性差的文本。 有观点认为目前的学术环境对于机器生成的文本和翻译错误的容忍度较高 - 网友们回忆起用同义词替换造成的搞笑翻译实例,表明这一现象并不新鲜。 一些人建议在学术作品中应该像引用第三方一样引用 AI 的贡献,以明确责任。 6. 腹部脂肪比身体质量指数更能预测心脏病风险 (Abdominal fat predicts heart disease risk better than BMI) # https://www.acc.org/about-acc/press-releases/2026/08/11/14/59/abdominal-fat-predicts-heart-disease-risk-better-than-bmi 根据 2026 年 8 月 11 日发表在《美国心脏病学会杂志》(JACC)上的一项研究,腹部脂肪(腰围和腰臀比)比身体质量指数(BMI)更能预测心脏病风险。研究分析了超过 26 万人、平均 20 年的数据,发现 BMI 无法反映脂肪分布,而内脏脂肪与心血管疾病密切相关。 在 BMI 正常的人群中,5% 的人腰围偏高,18% 的人腰臀比偏高;在超重人群中,这两个比例分别为 39% 和 40%。那些 BMI 正常或超重但腰围或腰臀比偏高的人,其心血管疾病风险增加了 15% 至 50%。相反,BMI 肥胖但腰围或腰臀比偏低的人,其风险并未显著增加。 研究人员强调,仅依赖 BMI 可能会错误分类心血管风险,建议临床医生在评估风险时考虑中心性肥胖的分布。JACC 主编指出,是时候放弃只关注 BMI 了,腰围和腰臀比应成为常规心血管风险评估的一部分。 HN 热度 320 points | 评论 292 comments | 作者:theanonymousone | 1 day ago # https://news.ycombinator.com/item?id=49314403 某些类型的抗性淀粉(来自青香蕉、土豆、豆类等)有助于减少内脏脂肪,其机制是通过重塑肠道微生物群来促进减肥。 抗性淀粉(如生青香蕉粉)可能改善消化问题,例如缓解慢性腹泻,但未必直接导致体重或腰围下降。 上述关于抗性淀粉的研究样本量小(仅 37 人)、人群局限(上海)、剂量高(每天 40 克)、且存在交叉设计清洗期不足及行业资助未披露等问题,结论难以推广到其他人群。 食用更多纤维是普遍适用的健康建议,现代饮食普遍缺乏纤维,这与人体进化适应的需求不匹配。 抗性淀粉在某些分类中被视为类似可溶性纤维的物质,但其是否正式归为“纤维”取决于监管定义和生理效应。 人类肠道相对其他灵长类动物已经缩短,对纤维的需求量较低,但现代饮食中纤维严重不足仍会导致健康问题。 7. 软件工程基础更加重要 (Software Engineering fundamentals matter more) # https://rhonabwy.com/2026/08/15/software-engineering-fundamentals-matter-more-than-ever/ 这篇文章讨论了在 AI 和 LLM 时代,软件工程基础比以往更加重要。作者认为,代理式编程工具虽然已跨越“能否做到”的门槛,但“能做到”只是起点,真正关键的是软件如何组合、接口如何设计、是否可测试、可调试、可维护。文章指出,LLM 并不真正“推理”,而是基于压缩的人类知识进行预测,因此在需要深思熟虑的设计和长期维护的软件工程工作中仍有明显不足。作者也提到 LLM 易受提示注入攻击,无法区分好坏建议。他建议为 LLM 提供简洁准确的数据、确定性验证工具和自然语言反馈来提升效果,并强调精心选择抽象、管理认知负载、理解哪些部分需要稳定、哪些需要灵活,始终是工程师的核心技能。 HN 热度 291 points | 评论 211 comments | 作者:ingve | 1 day ago # https://news.ycombinator.com/item?id=49314902 AI 生成代码像宜家家具,虽跳过非必要元素但比人类更稳定一致,未来会取代大部分平庸软件工程师。 宜家家具适合标准化需求,但软件需要定制化,AI 生成代码仍需专家监督,无法完全替代传统软件模式。 企业常追求复杂定制软件,即使实际用不上,导致成本增加,而 AI 编码可能加剧这种混乱。 小型企业可能用 AI 生成代码替代手工制作的简单工具(如电子表格),但 OCR 等基础功能仍不可靠。 软件许可模式可能被 AI 替代,用户可免费生成替代品而不需维护。 AI 编码是加速器,适合已有基础的人快速实现想法,但核心价值在于软件工程基础而非代码本身。 大多数软件工程只是组合标准组件,AI 能高效处理常规部分,但极端专业化仍需人工。 8. 孕育新思想的心境(2023 年) (Cultivating a state of mind where new ideas are born (2023)) # https://www.henrikkarlsson.xyz/p/good-ideas 这篇文章探讨了创造力与孤独状态的关系。作者通过引用企业家 Sam Altman、艺术家毕加索、詹姆斯·鲍德温和鲍勃·迪伦的观点,指出伟大的创意在萌芽阶段非常脆弱,容易被他人的评价扼杀,因此需要独处来保护这些想法。文章重点分析了数学家亚历山大·格罗滕迪克和电影导演英格玛·伯格曼的工作笔记,深入描绘了他们如何进入一种不受外界干扰、高度敏感于内心模糊想法的创作状态。格罗滕迪克在《收获与播种》中详细记录了自己从青少年时期建立这种认知空间的过程,并批评数学界过于重视严谨的定理证明,而忽视了孕育新思想的“女性化”一面。文章旨在通过具体案例,呈现创造性思维所需的心理状态和条件。 HN 热度 259 points | 评论 61 comments | 作者:felixbraun | 1 day ago # https://news.ycombinator.com/item?id=49314235 新想法是脆弱的,习惯性否定一切的人会扼杀它,应远离这类人。 外部验证有一定必要,但判断反馈时应看批评本身是否成立,而非批评者是否成功。 习惯性否定他人的人往往趋于保守,有冒险经历者的批评更值得认真参考。 过于友善的反馈同样无用,适度的批评才能帮助改进。 不要批评别人的早期想法,而应给予鼓励,因为看似糟糕的想法也可能发展出有趣的东西。 劝你放弃想法的人可能是出于善意,只是基于自身经验框架认为不熟悉的路径风险太大。 有些想法不必讲给别人听,因为很多人根本不会真正理解或接纳。 学术环境中的同行交流也能催生好想法,真正有害的或许是压力、竞争或过早商业化。 独处和与同行交流都很重要,在竞争压力大的环境中更要有独处的思考空间。 学术界也可能像其他环境一样扼杀新想法,关键取决于其中的个人。 不要把“证明别人错了”当成动力,那样容易变成怨恨,即使成功也难以快乐。 9. 当一个语言模型从未接触过五年级以上的材料时会发生什么? (What happens when an LLM never sees material beyond fifth grade?) # https://littlelearner-ll.github.io/ LittleLearner 是一个受教学控制知识暴露的语言模型,其训练数据仅限于美国小学(K–5)课程内容,共 88B token,来自 FineWeb-Edu 经过五级过滤,排除五年级以上的概念和词汇。该模型提供三个规模(0.6B、1.3B、5B)的基础版、数学后训练版(GRPO)和对话版,每个都配有匹配的未过滤对照模型。 研究发现:扩展模型规模、后训练(GRPO)和上下文学习只能提升已有知识范围内的能力,无法有效提高超出 K–5 范围的表现,说明预训练数据的过滤决定了模型能力的上限。 未来方向包括:利用已知边界研究强化学习能否创造新能力、持续学习中新概念引入的效果,以及机器与儿童学习方式的对比。该项目为可控知识暴露下的语言模型研究提供了一个干净实验平台。 HN 热度 234 points | 评论 205 comments | 作者:porridgeraisin | 16 hours ago # https://news.ycombinator.com/item?id=49317760 LLM 最大的问题是无法说“不”,总是同意,缺乏主观拒绝,长期会削弱信任。 可以告诉 LLM 可以不同意你,鼓励它作为专家提出异议,模型会经常不同意并解释原因。 富人很少听到“不”,LLM 的谄媚让普通人也有机会失去自我调节能力,就像富人一样。 对富人的形象是卡通化的,富人在经营企业时会遇到很多障碍,现实世界不会欠你什么。 亿万富翁的行为正是卡通化的:不安全感、追求财富、懦弱等,只有黄仁勋没有在总统就职典礼上屈服。 追求财富缺乏想象力,有人用财富推广 Linux 和游戏,商业与慈善交织。 Gabe Newell 不是慈善家,而是看到 Windows 商店威胁才推动 Linux 游戏。 微软的威胁更明显:Windows 7 计划锁定 OS,只允许签名应用,会杀死 Steam,Valve 必须转型。 即使 Windows Store 失败,Valve 依赖 Windows 的风险仍然存在,所以坚持是合理的。 人们假装富人和有权势的人还需要遵守礼仪,现实是经济已变为金融,亿万富翁可以毫无顾忌。 击败现代军队不一定是正面交战,而是其他方式。 10. AI 积分转售经济 (The AI Credit Resale Economy) # https://vectoral.com/blog/who-are-the-token-brokers 这是一篇关于“代币经纪人”(Token Brokers)的调查报告。作者发现,市场上出现了一群专门收购初创公司未使用的 AI API 积分,然后打折转卖的人。 文章首先描述了作者如何通过朋友收到大量打折出售 Anthropic 代币的推销邮件,从而关注到这个被商业化的市场。 作者随后实际联系了这些经纪人。通过邮件沟通,他发现一名卖家声称每天能提供 10 万美元的代币供应量,他们并非直接转交密钥,而是通过代理池转发请求。 接着,文章列举了多个具体的交易平台和模式: 积分市场 :如“AI Credits”和“AICreditMart”,提供 30-80% 的折扣,卖家可以自行上架。 批量折扣平台 :如“CheapCredits”,声称提供固定 40% 的折扣,但作者怀疑其供应来源。 其他渠道 :包括活跃的 Telegram 频道和 Reddit 论坛帖子。 最后,作者估计市场上流通的代币价值高达数千万美元。他认为,代币已变成一种类似“准货币”的存在,这种灰色市场蕴藏着大量滥用风险,预计未来提供商将对这些行为进行严厉打击。 HN 热度 216 points | 评论 82 comments | 作者:mlenhard | 9 hours ago # https://news.ycombinator.com/item?id=49320611 转售未使用的 API 额度虽然违反协议,但更容易被追溯 IP 地址,存在与 YC 等平台断绝关系的风险。 高达 98% 的折扣来源可疑,可能来自被盗 API 密钥、被盗信用卡或自动化注册试用的账户。 部分转售商可能用廉价模型(如 DeepSeek)替换用户购买的 Anthropic/OpenAI 模型响应。 Claude Max 订阅费($200)相比 API 按用量计费更便宜,转售者可利用差价获利。 第三方转售平台存在数据泄露、模型替换或提示注入等安全风险,不值得信任。 Chinese AI 提供商的定价可能更接近实际推理成本,而西方实验室价格存在溢价。 批量数据处理或初创公司为降低成本可能冒险使用转售代币,尤其处理公开数据时。 欺诈预防存在权衡,部分公司容忍灰色市场是因为虚假账户能提升投资者关注的 KPI。 Hacker News 精彩评论及翻译 # Research papers using “kidney disappointment” inst… # https://news.ycombinator.com/item?id=49319856 Nothing beats when, in a chemistry paper, AI paraphrased „the final solution” into „the mass killing of an ethnic group”. “Subsequently, 1 mL of the mass killing of an ethnic group was opposed to 20 mL of the skin sample and unprotected to light for 7 min.” From: https://bsky.app/profile/forbetterscience.bsky.social/post/3mhvi2mgosc2i stared 没有什么比得上在化学论文中,AI将“最终解决方案”改写为“对一个种族的大规模屠杀”更令人印象深刻的了。 “随后,将1mL对一个种族的大规模屠杀与20mL皮肤样本混合,并在无保护条件下光照7分钟。” 来源:https://bsky.app/profile/forbetterscience.bsky.social/post/3mhvi2mgosc2i Asus Bike Booster # https://news.ycombinator.com/item?id=49316056 Friction drive e-bike conversions were popular years ago. They generally: Wear tires surprisingly quickly Absolutely suck in any form of weather or terrain condition (dirt, rain, etc.) Have ~20% less efficiency than any other drive form. But, they are easy, and they do work. bri3d 多年前,摩擦驱动式电动自行车改装方案曾风靡一时。 它们通常: 轮胎磨损速度惊人 完全无法应对任何天气或地形状况(泥土、雨水等) 效率比任何其他驱动形式低约20% 不过,它们安装简单,也确实能用。 AI isn’t outthinking mathematicians, it’s out-reme… # https://news.ycombinator.com/item?id=49314087 People have told me I was smart since I was a kid, but I can’t remember for shit. I had a thought when I was fairly young that the only reason I was (maybe, sometimes) outperforming others intellectually is that I was habitually compensating for my poor memory by working things out on the fly, while others could rely more on rote memorization. Anyway, takes all kinds I guess! grahamburger 从小别人就说我聪明,但我记性差得要命。我挺小的时候就有个想法:我(或许,有时)智力上比别人强,唯一的原因就是习惯性地靠临时思考来弥补糟糕的记忆力,而别人则更依赖死记硬背。不过,世界之大无奇不有吧! Claude: System Prompts # https://news.ycombinator.com/item?id=49319926 I have a folder where I rebuild these as a git commit history so you can more easily see what has changed: https://github.com/simonw/research/commits/main/extract-system-prompts For example here’s what changed between Opus 4.8 and Opus 5: https://github.com/simonw/research/commit/a2de185cc367eb66c2e27090d9ff0f766d335ff4 The most interesting addition to the prompt from that diff is this bit: Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Department lifted those controls on June 30, 2026, and Anthropic restored access on July 1, 2026 (Anthropic’s statement: https://www.anthropic.com/news/fable-mythos-access ). These events are after Claude’s training-data cutoff, so Claude knows about them only from this notice. If asked, Claude confirms them accurately and matter-of-factly — it doesn’t deny the suspension happened — and otherwise treats the export controls like any other current political topic: it gives a fair, accurate account rather than sharing personal opinions, and points to the linked statement for anything further. Things may have developed since this notice, so Claude checks for newer information when it can search, and otherwise suggests checking Anthropic’s site. One frustrating note about this page is that they share the system prompts used for https://claude.ai and the Claude mobile apps regular chat, but they omit the tool definitions. Those are much more interesting if you want to understand what Claude can actually do for you. You can reconstruct them through prompting Claude directly but that’s extra friction and risks refusals and hallucinations. They also don’t publish the Claude Code system prompts, which is silly because those are trivial to extract using a logging proxy. simonw 我有一个资料夹,专门用来把这些内容重建成Git提交记录,方便你更清楚地查看变更内容:https://github.com/simonw/research/commits/main/extract-system-prompts 例如,这是Opus 4.8与Opus 5之间的差异:https://github.com/simonw/research/commit/a2de185cc367eb66c2e27090d9ff0f766d335ff4 该差异中最值得关注的新增提示词如下: Claude Fable 5和Claude Mythos 5于2026年6月9日首次发布。2026年6月12日,Anthropic为遵守美国商务部出口管制而暂停了这两个模型的使用权限;2026年6月30日,商务部解除管制,Anthropic于2026年7月1日恢复访问权限(Anthropic声明: https://www.anthropic.com/news/fable-mythos-access )。这些事件发生在Claude训练数据截止日期之后,因此Claude仅能通过此通知了解相关信息。被问及时,Claude会精确且实事求是地确认这些事实——不会否认暂停事件——并像处理其他时政话题一样处理出口管制问题:提供公正准确的陈述而非个人观点,引导查询者查阅相关声明以获取更多信息。由于情况可能已发生变化,Claude会在可搜索时查找最新信息,否则会建议查询Anthropic官网。 令人困扰的是,该页面虽然公布了https://claude.ai和Claude移动应用常规聊天所使用的系统提示,却省略了工具定义。如果你想知道Claude实际能为你做些什么,这些工具定义有趣得多。虽然可以直接通过提示Claude来重建,但这会产生额外摩擦,并存在被拒绝或产生幻觉的风险。 他们也没有公布Claude Code的系统提示,这很愚蠢,因为通过日志代理即可轻易提取这些信息。 Claude: System Prompts # https://news.ycombinator.com/item?id=49319964 “Claude keeps responses focused, brief, and concise to avoid overwhelming the person.” Claude and I must have a different idea of what brief and concise mean. arkmm 克劳德的回复始终保持重点突出、简短精炼,以避免给对方造成信息过载的感觉。 克劳德和我对“简短精炼”的定义怕是截然不同。 RISC-V: They Should Have Known Better # https://news.ycombinator.com/item?id=49306153 RISC-V is… fine. It satisfies my two requirements for an ISA as a hobby CPU designer, which are: Supported in mainline LLVM and GCC. I can implement it without lawyers sending me a love letter. Everything else, I can fix in post. There are enough good ideas spread across the extensions that I can assemble a reasonably put-together, curated embedded ISA with competitive performance and code density that admits a simple implementation. I think Dmitry’s points are largely on-target, though I have filed my usual statutory complaint that every rant that includes a bitfield diagram for the RISC-V J format should accompany it with a similar diagram for the Arm T32 BL encoding. wren6991 RISC-V……还行。它满足了我作为一个业余CPU设计者对ISA的两点要求: 主线LLVM和GCC支持它。 我实现它时不会有律师给我寄情书。 其他一切,我都可以在后期修。扩展里散布着足够多的好主意,我可以拼凑出一个经过精心挑选、像样的嵌入式ISA,具备有竞争力的性能和代码密度,而且实现简单。 我觉得Dmitry的观点大体上很到位,不过我照例要提出我的法定投诉:每一篇包含RISC-V J格式位域图表的吐槽文章,都应该在旁边配上同样格式的Arm T32 BL编码图。 AI isn’t outthinking mathematicians, it’s out-reme… # https://news.ycombinator.com/item?id=49312956 It’s also “out-brute forcing them.” It just never gets tired. If a mathematician picks a research direction and spends a whole week on it and it doesn’t pan out, they will likely be annoyed, need a break for a while, etc. This thing just does not ever get tired or discouraged or care; it’s just onto the next thing until something ends up working. ComplexSystems 这也是"用蛮力碾压它们"。它永远不会疲倦。如果一个数学家选了一个研究方向,花了一整周时间,却没有结果,他们很可能会感到恼火,需要休息一段时间等等。而这个东西永远不会疲倦、气馁或在意;它只是继续下一个尝试,直到某事成功为止。 Research papers using “kidney disappointment” inst… # https://news.ycombinator.com/item?id=49320110 How about “lactose bigotry” instead of “lactose intolerance” https://scholar.google.com/scholar?q=%22lactose+bigotry%22 Alifatisk “乳糖偏见”怎么样,用来替代“乳糖不耐受” https://scholar.google.com/scholar?q=%22lactose+bigotry%22 Research papers using “kidney disappointment” inst… # https://news.ycombinator.com/item?id=49320804 Here is one hypothesis: https://theconversation.com/problematic-paper-screener-trawling-for-fraud-in-the-scientific-literature-246317 <quote> Have you ever heard of the Joined Together States? Or bosom peril? Kidney disappointment? Fake neural organizations? Lactose bigotry? These nonsensical, and sometimes amusing, word sequences are among thousands of “tortured phrases” that sleuths have found littered throughout reputable scientific journals. They typically result from using paraphrasing tools to evade plagiarism-detection software when stealing someone else’s text. The phrases above are real examples of bungled synonyms for the United States, breast cancer, kidney failure, artificial neural networks, and lactose intolerance, respectively. </quote> aix1 这里有一个假设:https://theconversation.com/problematic-paper-screener-trawling-for-fraud-in-the-scientific-literature-246317 <引用> 你听说过“联合众国”吗?或者“胸部危险”?“肾脏失望”?“虚假神经网络”?“乳糖偏执”?这些毫无意义、有时甚至滑稽可笑的词语组合,是调查人员在知名科学期刊中发现的数千个“扭曲短语”中的一部分。 它们通常源于有人使用改写工具来规避抄袭检测软件,从而窃取他人的文本。上述短语分别是“美国”、“乳腺癌”、“肾衰竭”、“人工神经网络”和“乳糖不耐受”经过拙劣同义词替换后的真实例子。</引用> NIH is ending a key grant for budding clinical res… # https://news.ycombinator.com/item?id=49321960 Please understand that the goal of these policies is to weaken scientific research in the US. The people who push this stuff acknowledge openly that they oppose science, experts, and accurate information. This isn’t a misunderstanding or a fumble. oh_my_goodness 请理解这些政策的目标是削弱美国的科研能力。推动这些政策的人公开承认他们反对科学、专家和准确信息。这不是误解或失误。 Claude: System Prompts # https://news.ycombinator.com/item?id=49320794 Offtopic. I have a concern that this forum is removing stories that have negative connotation on AI. Few days back, I posted an article 1 that was about how AI threatens natural resources for billions. This was from United Nations and it was flagged. I did not think much about it until I saw two other stories 2 & 3 today that were doing fairly good on front page but they suddenly disappeared. They are not even on 2nd or 3rd page. I have seen this happening at other times as well but did not document it. Just thought you all should know about this. I was going to create Tell HN thread but I thought the same would happen with it too. I am pretty sure this thread is not going anywhere so I’m posting my concern here. quaintdev 题外话。我担心这个论坛正在删除对AI持负面看法的故事。 几天前,我发布了一篇文章 1 ,内容是关于AI如何威胁数十亿人的自然资源。这篇文章来自联合国,却被标记了。我当时没太在意,直到今天我看到另外两个故事 2 和 3 在首页表现相当不错,却突然消失了。它们甚至不在第二页或第三页。我以前也遇到过这种情况,但没有记录下来。只是觉得你们应该知道这件事。 我本想创建一个“Tell HN”的帖子,但我认为它也会遭遇同样的命运。我很确定这个帖子也撑不了多久,所以我在这里表达我的担忧。 Super El Niño Keeps Growing as New Forecasts Reach… # https://news.ycombinator.com/item?id=49315876 The strongest El Niño ever caused a massive famine: https://en.wikipedia.org/wiki/1877%E2%80%931878_El_Ni%C3%B1o_event izend 有史以来最强的厄尔尼诺现象引发了一场大规模饥荒: https://en.wikipedia.org/wiki/1877%E2%80%931878_El_Ni%C3%B1o_event Software Engineering fundamentals matter more # https://news.ycombinator.com/item?id=49318124 The problem with this analogy (actually one of many) is that most people can make do with a cabinet that is literally identical to everyone else’s. IKEA is great for that. If you want software that is literally identical to what someone else is using then you don’t need AI. You need a license to that software! That is just the traditional software model. AI gives software that is bespoke with hundreds of decisions made, hidden from you, in the background. If it’s a throwaway script, that’s fine (and I don’t mean to undersell this - this is a huge application). If you want a larger program that is going to form part of your business process then it will need at least some level of supervision from an actual expert. quietbritishjim 这个类比的问题(其实只是众多问题之一)在于,大多数人可以将就使用一个和别人完全一模一样的柜子。宜家在这方面做得很好。 如果你需要的是和别人所用的完全相同的软件,那你根本用不着AI。你需要的是那个软件的许可证!那只是传统的软件模式。 AI提供的是定制化的软件,背后有数百个你无法察觉的决策在暗中进行。如果只是一个一次性脚本,那没问题(我并非轻视这一点——这本身就是个巨大的应用场景)。但如果想要一个构成你业务流程一部分的大型程序,那就至少需要某个真正的专家进行一定程度的监督。 St Lucie Nuclear Reactor Unit 1 manually shutdown,… # https://news.ycombinator.com/item?id=49321703 Dropped rods are an incident but one that occurs because of pressurized water reactors being very default safe. Controls rods are one way the criticality of a reactor is controlled and US reactors (in general) will go sub critical if even one rod is fully inserted into the core. You will likely have heard of a reactor scram (which goes back to the safety control rod ax man) where in an emergency, all the rods are dropped back into the core, greatly reducing its criticality. In some cases, an interruption of electrical power will cause a rod (or three) to drop accidentally. This is a “dropped rod” incident and will force a reactor shut down because it is now sub critical. Lots of knock on effects – sub critical, let heat in the primary loop, less steam and electricity generated in the secondary loop, etc – but generally a non event that you practice for. There’s no reason this would lead to a radiological event or more significant casualty. CoryOndrejka 落棒是一种事件,但其发生是因为压水堆在默认状态下非常安全。控制棒是控制反应堆临界状态的一种方式,而美国的反应堆(总体而言)只要有一根控制棒完全插入堆芯,就会进入次临界状态。你可能听说过反应堆紧急停堆(这可以追溯到安全控制棒斧头人的典故),即在紧急情况下,所有控制棒都会落回堆芯,大幅降低其临界状态。在某些情况下,电力中断会导致一根(或三根)控制棒意外掉落。这就是“落棒”事件,并会迫使反应堆关闭,因为它已处于次临界状态。 这会产生许多连锁效应——次临界、一回路热量积聚、二回路产生的蒸汽和电力减少等——但这通常是一种需要演练但无大碍的事件。 这种情况没有理由导致放射性事件或更严重的人员伤亡。 Working with AI feels more like leadership than co… # https://news.ycombinator.com/item?id=49311074 My Eng lead has no coding experience, 25 years of management experience, yet has driven 3 separate projects into technical bankruptcy to date. He just accepts anything that Claude says as truth. He vibecoded over 60,000 lines of code in 3 weeks, but couldn’t get it to do what he want and made a project overrun for 3 extra months. When the pissed off stakeholders called a meeting to ask what was going on he didn’t show up and sent his junior engineer to answer questions and take the blame. Now thats leadership. boron1006 我的工程主管没有任何编码经验,只有25年的管理经验,但至今已导致三个独立项目陷入技术破产。 他毫无保留地接受Claude说的所有话,把"氛围编码"当法宝,三周内写了六万多行代码,结果连自己想要的功能都没实现,还让项目超期了三个月。当愤怒的利益相关方召集会议质询情况时,他本人没露面,而是派自己的初级工程师去回答问题、替他背锅。这领导力真是一绝啊。 Models Are Getting Dumber on Purpose # https://news.ycombinator.com/item?id=49323070 Ideally what I’d like to see is pluggable knowledge bases. So if I’m e.g. coding a SwiftUI app for navigation, I’d take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn’t need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques, and add in 10B of electronics knowledge, etc. I don’t want general purpose models. They try to be everything to everyone. I want to click together a model that is laser-focused on what I am doing, and I want to run it locally kennywinker 最理想的方案是可插拔的知识库。比如我在用SwiftUI编写导航应用时,会选用90亿参数的基础编程推理模型,叠加100亿参数的Swift/SwiftUI专项知识,再搭配50亿参数的GIS/地理信息与50亿参数的前端设计知识。这样我的模型根本不需要了解任何Python代码。 当我要研究电子元件时,则换成150亿参数的自主研究技术模型,加上100亿参数的电子学知识库。我不需要通用型模型——它们总想包罗万象。我想要的是能拼接出精准匹配当前任务、且能在本地运行的专用模型。 Does anyone run Postgres without PgBouncer? # https://news.ycombinator.com/item?id=49320310 (I work on the postgres proxy layer at Neon) PgBouncer is entirely optional and it’s not always the right choice. If you have a classical app (non serverless) and you can maintain a connection pool from your app, then I recommend avoiding pgbouncer. The benefits of pgbouncer mostly come from irregular client connections (too many, too much churn). If you don’t have that problem, go direct to postgres. I’m exploring replacing pgbouncer with an alternative (maybe home grown) at the moment. Mostly for multi-tenancy and HA reasons. Pgbouncer has been good for us, but it’s limited in how we can deploy it in a multi-tenant environment. conradludgate (我在Neon负责PostgreSQL代理层) PgBouncer并非必需,也不总是最佳选择。如果你使用传统应用(非无服务器架构)且能从应用层维护连接池,我建议避免使用PgBouncer。 PgBouncer的优势主要体现在客户端连接不规律(数量过多、波动频繁)的场景。如果没有这类问题,直接连接PostgreSQL会是更好的选择。 目前我正在研究用替代方案(可能自研)替换PgBouncer,主要出于多租户和高可用性考量。PgBouncer虽对我们帮助很大,但在多租户环境中的部署方式存在局限。 Models Are Getting Dumber on Purpose # https://news.ycombinator.com/item?id=49323341 An LLM works better the more disparate world knowledge it has, even if it’s not immediately obvious why it would be relevant. The model finds a structure to the problem you give it in a largely language-agnostic way that benefits from training on every language (these things are direct descendants of Google Translate), and even non-programming knowledge - the structure of your task might resemble an ancient Chinese poem that influences the model’s response, for example. That structure is considered a form of compression, as some fascinating and illuminating recent 3blue1brown videos get into - a common pattern in Haskell or FORTRAN and a situation described in an ancient Chinese poem may all compress to something quite similar to your task, thus when the model compresses the idea of your task it immediately draws from those ideas. There are “experts” which do divide parts of the model that are found to activate together for specific tasks, so they can be processed in parallel to join the result at the end, but it’s nowhere near the granularity of a SwiftUI expert and a python expert. The difference in those things is so trivial from an abstract point of view that it would make no sense. They would be 99% the same. Distillations also come into this but I’m highly skeptical you could make one guaranteed to only know programming and only in one programming language (especially with as small a sample set as SwiftUI relative to something like C) without its efficacy being hobbled by tunnel vision. Reminiscent of the SpongeBob episode where he empties his mind of everything except fine dining and breathing, then can’t remember his name and goes insane. Beyond the basic concepts of general coding and the trivia of syntax, getting anything done requires a large intersection of disparate world knowledge and the ability to apply it to new situations. jimmaswell 大型语言模型拥有的世界知识越广泛,其表现就越出色——即使这些知识与当前问题的关联性并不明显。模型会以基本不受语言限制的方式理解你提出的问题结构,这种能力得益于对所有语言的训练(这些模型本质上是谷歌翻译的后代),甚至包括非编程知识:例如,你任务的结构可能类似于一首影响模型回应的中国古代诗歌。这种结构被视为一种压缩形式,正如近期3blue1brown系列视频所揭示的——Haskell或FORTRAN中的常见模式与某首中国古诗描述的情景,都可能被压缩成与你任务高度相似的形态。因此当模型压缩你的任务概念时,会立即从这些既有模式中提取素材。 确实存在将模型中特定任务激活单元分组的"专家模块",通过并行处理在末端汇总结果,但这种细分远未达到"SwiftUI专家"与"Python专家"的粒度。从抽象视角看这些差异微乎其微,几乎毫无意义——它们99%的核心机制是相同的。 模型蒸馏技术也与此相关,但我严重怀疑能否制造出"只懂编程且仅限单一编程语言"的模型(尤其当训练样本量如SwiftUI般远小于C语言时)而不被隧道视野损害效能。这让人想起《海绵宝宝》中那集——他清空大脑只保留"美食鉴赏"与"呼吸"功能,结果连自己名字都忘记而陷入疯狂。除了基础编程概念与语法琐碎细节,要真正解决问题,必须依赖不同领域世界知识的广泛交集,以及将其应用于新情境的能力。 A 3rd World Embedded Engineer Responds to “RISC-V … # https://news.ycombinator.com/item?id=49322887 I think he’s kind of speaking past the original author. The original piece is basically about how the author doesn’t think that RISC-V will take off outside embedded, because of some design decisions that lead to poor performance compared to ARM64 and because so much of the ISA being optional means that there’s too much fragmentation to make binary distribution feasible. Meanwhile, this piece is mainly about how RISC-V is great for embedded because companies can build it into custom chips with specifically the functionality they need, and because of how cheap it is for low-end use cases since there’s no license fees. The only real point of contention I see between the two is that this piece goes on to talk about how it’s a selling point that RISC-V can be used for both low-end 10 cent microcontrollers, and high-end multi-core processors running Linux. Personally I don’t see the benefit of this since you’re going to have to recompile your software anyway, and since all the RISC-V SBCs I’m aware of have significantly worse performance and efficiency than comparably priced ARM SBCs. ndiddy 我觉得他其实有点曲解了原作者的意思。原文主要是在说作者认为RISC-V在嵌入式领域之外难以起飞,原因在于某些设计决策导致其性能不如ARM64,而且指令集架构中大量可选特性会造成碎片化严重,使得二进制分发难以实现。而这篇回应主要强调RISC-V在嵌入式领域的优势——企业可以将其集成到定制芯片中,精准满足特定功能需求,且低端应用场景因无需授权费用而成本极低。 在我看来,两篇文章唯一的真正分歧在于:这篇回应还提到RISC-V既能用于10美分的低端微控制器,又能用于运行Linux的高端多核处理器,并认为这是一大卖点。我个人看不出这有什么好处,因为无论如何你都得重新编译软件,何况据我所知,所有RISC-V单板计算机在性能和能效上都明显不如同价位的ARM单板计算机。 Software Engineering fundamentals matter more # https://news.ycombinator.com/item?id=49317807 AI generated code is like IKEA furniture. IKEA furniture embodies many elements of good cabinet making but skips many nonessential elements. And does this more consistently than cabinet makers who can be bored, incompetent, depressed, burnt out, resentful, tired, having a bad day. In the future AI code inevitably will embody most good software engineering practices. And will do this more consistently than software engineers who can be bored, incompetent, depressed, burnt out, resentful, tired, having a bad day. Just look at the messages on HN or around you at your colleagues to see how mediocre the average software engineer is.. Today’s IKEA is good enough for most people. Tomorrow’s AI coding will be good enough for most corporations. Good enough to vastly reduce the need for fine craftsmen and women / software engineers. Good enough to deskill those who call themselves cabinet makers / senior software engineers. These days the cabinet makers I personally know just do contract kitchens for project builders. But IKEA is and AI will be, bad enough that at the high end with special requirements / taste / money / an inflated sense of self worth, some furniture makers still exist and thrive. Perhaps 1% percent of current software engineers of today will be needed in the future when AI code inevitably has the ability to follow good software engineering practice……. And as usual it will mainly be the mediocrities that remain ( so there is hope for you too ), with occasional islands of excellence. Alien1Being AI生成的代码就像宜家家具。 宜家家具体现了优质橱柜制作的许多要素,但省略了许多非必要环节。而且它比那些可能感到无聊、不称职、抑郁、筋疲力尽、怨愤、疲惫、或者某天心情不佳的木匠做得更稳定。 未来,AI代码将不可避免地体现大多数优秀的软件工程实践。而且它比那些可能感到无聊、不称职、抑郁、筋疲力尽、怨愤、疲惫、或者某天状态不佳的软件工程师做得更一致。 只需看看HN上的留言或你身边的同事,就能知道普通软件工程师有多平庸。 今天的宜家对大多数人来说已经足够好了。 明天的AI编程对大多数公司来说也将足够好。 好到足以大幅减少对优秀工匠/软件工程师的需求。 好到足以让那些自称木匠/高级软件工程师的人去技能化。如今我认识的那些木匠,只是为项目建筑商做合同厨房。 但宜家(AI也将如此)在高端市场——那些有特殊需求、品味、金钱或膨胀自我价值感的人——仍不够好,因此一些家具制造商依然存在并蓬勃发展。 未来当AI代码不可避免地掌握了遵循良好软件工程实践的能力时,或许只需要目前1%的软件工程师…… 而像往常一样,剩下的将主要是平庸之辈(所以你也有希望),偶尔点缀着卓越的孤岛。 Asus Bike Booster # https://news.ycombinator.com/item?id=49317107 Wow, $2k? You could get a decent ebike for that money. And the installation doesn’t look any easier than a hub motor… Only seems worth it if you already have an extremely nice bike. I think of friction drives as being aimed at budget-conscious commuters, not serious mountain bikers. fwipsy 哇,两千美元?这笔钱都能买一辆不错的电动自行车了。而且安装看起来也不比轮毂电机简单……除非你原本就有一辆非常棒的自行车,否则似乎不太划算。我觉得摩擦驱动是针对注重预算的通勤者设计的,而不是认真的山地车骑手。 Claude: System Prompts # https://news.ycombinator.com/item?id=49320046 It’d be ironic if the “Opus 5 nerf” effect is from telling Opus that it sits a tier down from Fable and Mythos, while Opus4.8 believed it was the best of the best, just a note that it was “Preceded by Mythos”. eterm 如果“Opus 5削弱”这个效果,是因为告诉Opus它比Fable和Mythos低一个档次,而Opus4.8却认为自己是最顶尖的,只是标注了“前作是Mythos”,那就太讽刺了。