这篇文章讲了为什么中国开放权重的 AI 策略正在赢,用具体数据说明差距缩小,和美国的封闭策略形成了鲜明对比。
文章指出 AI 模型本身几乎无护城河,用户可轻松切换。中国通过开放权重模型将计算劣势转化为分发优势,加速了美国公司模型层商品化。尽管美国实施 GPU 出口管制和限制数据共享,但中国前沿模型与美国的差距正迅速缩小。作者认为美国封闭、专有的做法失败,而中国开放的生态策略正在胜出。
2026 07 21 HackerNews
2026-07-21 Hacker News Top Stories # 中国开放权重AI模型正将计算劣势转化为分发优势,侵蚀美国企业盈利基础。 这是一个实时展示机场航班起降数据和统计的页面。 数学家借助AI找到了雅可比猜想的三维反例,推翻了该猜想。 黑客入侵并删除了罗马尼亚土地登记数据库,但官方持有离线备份。 欧盟拟向美国交出公民生物特征数据,以换取免签旅行待遇。 小米发布基于超10万小时真实操作数据训练的基础机器人模型,可快速适应新任务。 作者宣称用GPT-5.6以25美元成本发现WordPress远程代码执行漏洞,但被质疑为营销炒作。 OpenCode工具被批存在缓存失效、压缩脆弱和安全问题,作者建议弃用。 研究发现错误AI建议会显著降低用户正确率,却大幅提升其盲目自信。 Moonshine是一个Linux游戏串流工具,支持远程输入并将画面传输到Moonlight设备。 1. 中国开放权重的 AI 策略正在胜出 (China’s open-weights AI strategy is winning) # https://werd.io/american-ai-is-locked-down-and-proprietary-its-losing/ 本文讨论了中美两国在人工智能领域的策略对比。作者指出,AI 模型本身几乎没有护城河,用户可以轻松切换使用不同厂商的模型。中国采取开放权重策略,将模型公开释放,从而将计算劣势转化为分发优势,并使得美国公司赖以盈利的模型层商品化。尽管美国通过 GPU 出口管制和限制数据共享来维持优势,但中国前沿模型与美国的差距正在迅速缩小。文章认为,美国封闭、专有的做法是失败的,而中国开放、协作的生态策略正在胜出,并可能对美国经济产生深远影响。作者呼吁美国需要调整激励结构,转向更开放、符合公共利益的技术发展路径。 HN 热度 905 points | 评论 734 comments | 作者:benwerd | 9 hours ago # https://news.ycombinator.com/item?id=48979269 过去 50 年计算机和软件市场的教训是免费和低端最终获胜,PC 摧毁小型机,Windows/Linux 摧毁 UNIX,前沿模型训练成本不可持续,本地 LLM 类似 PC 爱好者时代,未来 10-15 年消费级 PC/手机能运行当前前沿模型,中国开放权重模型加速了低成本竞争 通用 LLM 作为独立工具在消费应用不会长久,产品设计师会创造透明处理模型交互的实用工具,消费者不想要阿谀奉承或固执的界面 每天使用 AI 聊天非常有用,已替代 Google 搜索 LLM 有助于获取流行推荐,但感觉还处于早期,更精致的集成到搜索引擎可能更有用 先行者样本说明产品采用周期,很多人使用 AI 不代表全貌 Google 搜索仍在发生,只是由 agent 执行 消费者喜欢聊天机器人的阿谀奉承,例如面试后 AI 安慰用户 开放权重模型依赖大公司巨额训练投入并放弃盈利路径,无法分布式训练,激励机制可能随时改变,除非真正去中心化 不同公司有不同激励,Nvidia 有动力让用户购买个人硬件运行开源模型 Nvidia 可能更倾向于企业用高价服务器而非个人 GPU,RAM 厂商已放弃低端市场 销售给个人更稳定可预测,消费者硬件市场有吸引力 硬件制造商激励取决于供应是否受限,受限时卖给 LLM SaaS 公司,不受限时个人用户更有利可图,Nvidia 仍重视 AI 桌面和笔记本为开发者关系 2. 机场模拟器 (Airport Simulator) # https://airport.apunen.com/ 这是一个航班活动统计页面,显示当前机场的实时飞行数据。主要包括: 着陆 :0 架次 起飞 :0 架次 航班频率 :0.0 架/分钟 飞行时长 :0 分 3 秒 页面由 @lapunen 制作。 HN 热度 661 points | 评论 132 comments | 作者:apunen | 13 hours ago # https://news.ycombinator.com/item?id=48976846 怀念 Flight Control,渴望更多现代版本,如 Mini Metro 提及 Jan David Nose 的 Auto Traffic Control 编程挑战游戏,通过 GRPC 控制飞机,但玩法门槛高 希望 ATC 游戏基于真实的 FF-ICE/FIXM 飞行计划格式 Flight Control 曾是 iOS 经典,现已无法在移动端游玩,仅存 Steam 版 Planes Control 被视为 Flight Control 的忠实翻版,仍获更新 早期 ATC 游戏包括 Heathow ATC(ZX Spectrum)、Kennedy Approach 等 开源 ATC 模拟器 OpenScope 更复杂,适合键盘操作 Mini Airways 等作品在类似 Mini Metro 基础上增加了复杂度 有 VR 版本 Final Approach,体验类似 BSD games 中的 ‘atc’ 采用纯文本界面 Endless ATC 可在 Android 和 Steam 上玩 游戏内飞行员不遵守“看避”规则,导致空中碰撞 希望改进:玩家能自主决定起飞时机,并加入高度影响机制 期望游戏是像后室风格的无限机场探索体验 3. Claude Fable 给出了雅可比猜想的一个反例 (Claude Fable produced a counterexample to the Jacobian Conjecture) # https://xcancel.com/ alpoge /status/2079028340955197566 数学家 levent 在 X 上宣布 Jacobian 猜想(雅可比猜想)被推翻,并给出了具体反例: 映射:((1+xy)³ z + y² (1+xy)(4+3xy), y + 3x(1+xy)² z + 3x y²(4+3xy), 2x - 3x² y - x³ z) 从 ℂ³ 到 ℂ³。 其雅可比行列式恒为 -2(常数),但该映射将三个不同点 (0,0,-1/4)、(1,-3/2,13/2)、(-1,3/2,13/2) 都映到 (-1/4,0,0),故非单射,从而否定了 Jacobian 猜想。 levent 提到这是他与朋友 Akhil 和 Fable(在世界杯决赛期间工作)的合作成果。 回复中,Grok 确认这是一个有效的反例,并指出该反例于 2026 年 7 月 19 日被发现。其他讨论涉及该结果对 Dixon 猜想、Poisson 猜想的牵连,以及如何通过 Poisson 括号简化构造过程的数学原理。 HN 热度 646 points | 评论 417 comments | 作者:loubbrad | 21 hours ago # https://news.ycombinator.com/item?id=48973869 这个反例的阶数只有 7,而之前人们猜测可能需要高达 200 阶,令人震惊。 两变量情况曾被测试到 150 阶以上,而这是三变量反例,不能混淆。 他们当年的搜索空间太大,只能通过过滤缩小范围,这个反例可能不在过滤后的池中。 除非分享完整聊天记录,否则这更像是营销炒作,AI 实际作用可能很小。 许多数学家长期努力未果,现在 AI 做到了,不能简单否定为营销。 关于 AI 的争论已经两极分化,有人极端否定 AI 的价值。 人类研究者很可能提供了非平凡的洞察,但 AI 仍然是解决方案中之前缺失的关键部分。 作者不分享 prompt 可能是想用相同方法再找更多反例。 推理过程比聊天记录更有价值,但不会被公开。 应该对结果保持怀疑,因为只知道一条推文,缺乏具体信息。 发布者是 Anthropic 的一位数学家,并非公司官方行为,分享聊天记录似乎并不必要。 聊天记录可能包含大量死胡同和错误,没人愿意公开这些尴尬内容。 4. 黑客摧毁罗马尼亚土地登记数据库 (Hacker wipes Romania’s land registry database) # https://news.risky.biz/risky-bulletin-hacker-wipes-romanias-entire-land-registry-database/ 一名黑客入侵了罗马尼亚国家地籍与房地产广告局(ANCPI),使用有效凭证进入系统、映射内部网络,在勒索失败后删除了整个土地登记数据库及备份,导致该国房地产交易市场全面停摆,官方应用和网站已下线一周。被盗数据(包括员工凭证、内部文档及网络细节)被黑客“ByteToBreach”在论坛上出售,其真实身份已被安全公司确认为阿尔及利亚的 Zakaria Mahdjoub。ANCPI 表示正在从零重建整个网络,但似乎持有离线备份。 HN 热度 549 points | 评论 307 comments | 作者:speckx | 10 hours ago # https://news.ycombinator.com/item?id=48978605 黑客删除了罗马尼亚土地登记数据库,但官方有离线备份,因此影响有限。 类似灾难(如洪水)后,可通过产权证明和证人证言重新建立土地登记,虽无法 100% 恢复,但有异议期处理虚假主张。 身份证明文件全部丢失时,可通过多人宣誓证词开始重建文件链,类似婚姻中改姓流程。 更换社会安全卡有次数限制(10 次后不再发放),但合法改名不受此限。 超过次数后需亲自到社保局办公室办理,之前可在线或邮寄申请。 现实中很少需要出示实体社安卡,因为无照片,仅靠号码即可验证;但获取 Real ID 驾照时仍可能需要。 社安卡实际上已沦为事实上的全国身份证,尽管本意并非如此。 美国数字化水平远不如其他国家(如索马里)的评论被反驳。 一些人对个人信息过度谨慎,导致难以证明自己存在,如无驾照、未婚、无信用记录,办理护照和购房时才需 ID。 使用非照片 ID(需先用照片 ID 获取)可绕过部分官僚障碍,工作人员有时会理解“安全表演”而通融。 通过地址记录、历史税表、银行账户等也可以恢复身份证明。 重建身份是可能的,但若土地所有权完全丢失则想象更严重场景。 5. 欧盟即将将我们最敏感的数据出售给美国以换取免签旅行 (The EU is about to sell our most sensitive data to the US for visa-free travel) # https://edri.org/our-work/the-eu-is-about-to-sell-our-most-sensitive-data-to-the-us-for-visa-free-travel/ 欧盟委员会正与美国政府秘密谈判一项“加强边境安全伙伴关系”(EBSP)框架协议,允许美国边境当局获取欧盟国家的生物特征数据库,并对旅行者进行安全风险画像。2022 年,美国以保持免签待遇为条件,要求欧盟成员国共享这些数据。2025 年 12 月,欧盟理事会授权委员会代表成员国进行谈判。 据泄露的草案文本,欧盟几乎完全接受了美方过分要求,包括系统性传输生物特征数据和主观“风险指标”。此举可能针对政治异见人士、跨性别权利支持者或批评加沙战争的人士,导致他们在入境美国时被拘留。EDRi 分析认为,该草案严重偏离成员国谈判授权,且不符合欧盟基本权利宪章。 EDRi 呼吁欧盟委员会和理事会抵制美国施压,拒绝将公民个人数据出售给一个人权记录堪忧、民主倒退严重的国家,并要求坚守核心数据保护规则。 HN 热度 464 points | 评论 275 comments | 作者:rapnie | 11 hours ago # https://news.ycombinator.com/item?id=48977711 进入美国本就需提供照片和指纹等生物数据,免签证方案是电子传输(更便捷),而签证手续繁琐且仍需在机场采集生物数据。 欧盟目前也对游客采集照片和指纹,且每次出入境都需采集,而美国仅需一次(或每护照一次)且在入境时。 新方案可能使美国获取未前往美国人士的生物数据,因为查询可基于姓名、出生日期、身份证号甚至指纹,即使对方从未申请签证或试图入境。 多个欧洲国家的国民身份证号是公开或可简单推算的(如意大利、瑞典、比利时),增加了隐私泄露风险。 比利时身份证号由出生日期、当日出生顺序和性别(奇/偶)组成,甚至可能随性别变更而改变。 草案显示查询机制不完善,可能允许非法查询且难以被检测。 只有前往美国的人才应被纳入该系统,反之亦然;这改变了整个计划的性质。 RealID 推行二十年并未显著提升安全,这种追踪级别不值得安全增益,安全只是借口。 国家身份证系统是欧洲国家的标准配置,对 RealID 的反对令人困惑。 RealID 可以通过支付 45 美元绕过,说明其实际效果有限。 6. Xiaomi-Robotics-1 (Xiaomi-Robotics-1) # https://robotics.xiaomi.com/xiaomi-robotics-1.html Xiaomi-Robotics-1 是一个基于超过 10 万小时真实世界操作轨迹训练的机器人基础模型。它采用“无本体预训练(UMI)+ 后训练对齐”的方法:预训练阶段利用大规模 UMI 数据学习通用的动作生成能力,后训练阶段使用真实机器人数据和指令数据进行对齐。实验表明,预训练阶段数据量和模型大小与验证误差呈现清晰的规模扩展规律,且这种扩展行为直接传递到后训练阶段——更强的预训练模型带来更高的真实机器人成功率。在应用方面,该模型能高效适应新任务,平均不到 10 小时演示即可达到 75% 成功率,平均不到 40 小时演示可提升至 85%;同时在四个主流仿真基准(RoboCasa、VLABench 等)上取得最优结果,相对第二名有显著提升。 HN 热度 444 points | 评论 294 comments | 作者:ilreb | 19 hours ago # https://news.ycombinator.com/item?id=48974454 对机器人能做家务感到乐观和激动 认为 HN 上存在反华情绪和亲美例外主义 指出对非美国产品常持怀疑态度,但对日本产品例外 反驳亲美说法,认为近期帖子倾向反资本主义和反美 观察到开放权重模型讨论中总有反华声音 猜测反华者可能是担心 AGI 风险的有效利他主义者 讽刺有效利他主义者已转向 AI 恐怖故事 认为任何正面美国的内容会被压制,而负面美国容易热门 指出美国寿命增长不如其他富裕国家 吐槽华盛顿邮报的爱国动画和付费墙显得讽刺 认为本贴评论被“不准批评中国”的群体主导 认为在美国主导论坛如此反应正常,并预测 21 世纪属于中国 认为 HN 上亲华或反华取决于话题时间点和机器人数量 质疑国家赞助机器人的存在并缺乏证据 觉得 HN 实际上是亲华论坛 有人因特朗普政策不想支持美国公司,也不希望中国垄断生产链 质疑特朗普是否在实施种族灭绝(指向巴勒斯坦) 7. 漏洞经纪商花费 50 万美元购买 WordPress 远程代码执行漏洞。我用 GPT5.6 和 25 美元就找到了一个。 (Exploit brokers pay $500k for WordPress RCEs. I found one with GPT5.6 and $25) # https://slcyber.io/research-center/exploit-brokers-pay-500000-for-a-wordpress-rce-i-found-one-with-gpt5-6/ 这是一个网站的 Cookie 同意管理页面,详细列出了网站所使用的各类 Cookie 及其用途。页面将 Cookie 分为四类: 必要 Cookie(30 个) :用于实现网页基本功能,如页面导航、安全访问等,包括 BambooHR、Cloudflare、Cookiebot、Google、HubSpot、LinkedIn 等多个服务商提供的 Cookie。 偏好 Cookie(2 个) :记住用户的地区、语言等偏好设置,如 LinkedIn 的负载均衡 Cookie。 统计 Cookie(13 个) :帮助网站匿名收集访客行为数据,以优化网站性能,涉及 Hotjar、HubSpot、Microsoft、PubMatic 等服务。 营销 Cookie(50 个) :用于跨站追踪用户,以便展示个性化广告,包含 Meta、Adroll 等平台的追踪代码和数据存储。 HN 热度 379 points | 评论 209 comments | 作者:infosecau | 15 hours ago # https://news.ycombinator.com/item?id=48975665 没有证据表明 WordPress 漏洞能卖到 50 万美元,作者可能只是想卖 prompt。 漏洞经纪公司如 Crowdfense 和 Zerodium 曾报价高额收购,但通常分期付款且要求漏洞未修复。 WordPress 漏洞价值不高,因为 WordPress 站点没有有价值信息,浏览器和移动端漏洞才受政府青睐。 联邦承包商使用 WordPress 作为主要网站,可通过社工利用信息。 政府只对浏览器和移动端漏洞感兴趣,WordPress 不在此列。 WordPress 占 41% 网站,可用于水坑攻击、勒索、数据泄露等,价值被低估。 地下论坛上 WordPress 漏洞只值 200 美元,远低于 50 万。 Panama papers 就是通过 WordPress 黑客攻击泄露的。 许多政府网站用 WordPress 作为 CMS,因此漏洞有潜在价值。 WordPress 站点背后的服务器可用于挖矿、僵尸网络等犯罪。 Zerodium 是否真实存在有争议,有人认为它已不存在,有人声称做过生意。 在 HN 上发布无根据的声明令人困惑,可能只是展示 Gell-Mann amnesia 效应。 8. OpenCode 的恼人与令人担忧之处 (Annoying and alarming things about OpenCode) # https://wren.wtf/shower-thoughts/stop-using-opencode/ 这篇文章大肆批判 OpenCode,称其为“小丑车般的涡轮垃圾”,安全极差。作者通过本地 LLM 试用后指出两大问题:恼人之处和令人担忧之处。恼人部分包括:提示缓存频繁失效(如每次 SSE 轮次都会重新读取 AGENTS.md、在 agent 到用户转换时剪枝上下文丢失 40k 前缀、午夜更新日期导致完全缓存未命中);上下文剪枝对早期读取缺乏保护,例如中断后规格说明会被删掉;压缩功能基础薄弱且与剪枝配合糟糕;系统提示冗长且充满“烂建议”,无法全局修改默认提示,不同模型提示质量参差不齐。作者建议放弃 OpenCode,转而接受上下文窗口和缓存作为首要特性。 HN 热度 367 points | 评论 242 comments | 作者:alekq | 11 hours ago # https://news.ycombinator.com/item?id=48978112 文章标题应改为"一些小烦扰,修好就能改进 OpenCode" OpenCode 每次 SSE 都会重复读取 AGENTS.md 并注入系统提示,造成缓存未命中 OpenCode 把当前日期塞入系统提示,午夜刷新时触发全量缓存重算 OpenCode 的压缩(compaction)功能实现糟糕,且与修剪(pruning)相互影响不好 OpenCode 的默认系统提示词有偏见但可以修改 OpenCode 代码因 vibe coding 过度膨胀,稳定性、性能、内存使用都变差了 尽管有技术缺陷,OpenCode 实际非常有效,不必费心切换 很多 OpenCode 用户后来转向 Pi(oh-my-pi) OpenCode 的源代码不好读 OpenCode 的代码像二进制机器码,提示词才是真正的源码 替代工具包括 Pi、Dirge、Aider,也可以基于 pi 自己写 OpenAI 的 Codex CLI 用 Rust 编写,内存消耗较小,表现不错 OpenCode 的压缩是必要的恶,但 V2 版本已用新系统避免缓存丢失 OpenCode 新界面取消了工作区/workspace 支持,多会话切换困难,命令面板只支持快捷键 项目追求快速迭代,缺乏稳定,GitHub issue 堆积且 stale bot 过于激进 9. AI 建议让人准确率更低但更自信——研究 (AI advice made people less accurate but more confident – sudy) # https://thenextweb.com/news/ai-advice-suppresses-critical-thinking-wrong-answers-study 研究人员发现,当人们使用 AI 建议时,其判断能力显著下降:愿意说“不知道”的比例从 44% 降至 3%,准确率从 27% 降至 9%,而信心却从 30% 升至 76%。即使用户原本回答正确,一旦参考了 AI(通常是错误的回答),也会变得错误。 货币激励的改善效果有限,仅让“不知道”比例升至 8%、准确率升至 16%,仍远低于无 AI 时的水平。这种现象被称为“认知投降”——人们 80% 的情况下接受错误 AI 答案,同时自信心反而更高。 研究者特别关注对孩子的影响,因为 AI 产品设计总是给出答案,从不承认“不知道”,这会削弱人类识别自身认知局限的能力。 HN 热度 357 points | 评论 206 comments | 作者:rbanffy | 1 day ago # https://news.ycombinator.com/item?id=48971738 研究设计很糟糕,其测试的不是 AI 系统特有的问题,而是类似给一本有错误的教材让学生使用。 研究中的 AI 被故意设置成几乎总是错误回答,这不代表真实 AI 的表现。 所有 AI 都有缺陷,容易产生幻觉或“记错”事情,人类也会如此。 研究应该使用正常失败率的 AI,并与其他信息来源进行对比。 研究的重点不是 LLM 本身,而是人们是否会批判性评估 AI 输出的内容。 人们是否愿意听从 AI,取决于他们过去对该模型的经验。 测试的六个问题纯属记忆力题,无法通过思考区分正确与貌似合理的答案。 研究标题具有误导性,实际结论应为“非常不准确的 AI 使人们更不准确”。 研究证明了一个狭窄结论:当 AI 不足以回答时,人们仍然倾向于依赖它。 不同模型对同一问题的回答各异,追问“你确定吗”有时能让模型自我纠正。 10. Moonshine:让你将游戏从 PC 串流到任何运行 Moonlight 的设备上 (Moonshine: Lets you stream games from your PC to any device running Moonlight) # https://github.com/hgaiser/moonshine Moonshine 是一个 Linux 上的游戏流式传输工具,允许用户将 PC 游戏无线串流到任何运行 Moonlight 的设备上,并支持鼠标、键盘和手柄的远程输入。 核心特性:每个流会话运行在独立的合成器中,与桌面环境隔离,不影响主机正常使用;无需物理显示器即可工作于无头服务器;支持 H.264、H.265 和 AV1 硬件编码(AV1 在 NVIDIA 旧驱动下存在帧大小增长问题,建议使用前两种);支持 HDR 10-bit、环绕声(5.1/7.1)及低延迟 Opus 音频编码;支持完整的输入设备(含手柄运动、触控板、触觉反馈)。 安装要求:系统需使用 systemd,显卡需支持 Vulkan 视频编码(NVIDIA RTX、AMD RDNA2+、Intel Arc),客户端需 Moonlight v6.0.0 或更高。可通过 AUR(Arch 系统)安装,也可从源码编译构建。 首次使用需配对客户端:访问 http://localhost:47989/pin 或通过终端 curl 提交 PIN。应用通过配置 config.toml 添加,可设置启动命令、前置/后置脚本,并支持 Steam 等自动应用扫描器。 HN 热度 326 points | 评论 135 comments | 作者:wertyk | 23 hours ago # https://news.ycombinator.com/item?id=48972970 Moonshine 的发展历程:从 Nvidia Gamestream→Sunshine/Moonlight→Apollo/Artemis,再到 Moonshine 实现了 Linux 上的虚拟显示流媒体。 Steam Link 仍然可用,但需要与主机交互(如解锁、登录),不是完全无头的方案。 使用 Steam Link 远程时可能暴露主机安全,且他人可干扰游戏或存档。 Steam Link 可作为跳板绕过 VPN 问题,从手机远程访问局域网设备。 Valve 想做类似 Stadia 的云游戏需要数据中心和大量 GPU,与家庭 PC 流媒体本质不同。 当前 GPU 算力虽有余量,但多为 AI 设计(低 CPU 时钟、缺少视频编码器/RT 核心),不适合游戏场景。 Games on Whales 比 Steam Link 更高效,能更好处理光标缩放等问题。 Vibepollo 是几乎全用 vibe coding 改写的 Apollo 分支,但作者个人不推荐使用。 Hacker News 精彩评论及翻译 # China’s open-weights AI strategy is winning # https://news.ycombinator.com/item?id=48982784 The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins. PCs destroyed minicomputers. Mainframes survive, but serving a much tinier portion of the market than they used to. PC office productivity software destroyed expensive professional products. Windows (low end) and Linux (free) completely destroyed the UNIX marketplace, and again, have taken huge market share from the mainframe world. Ignoring the huge Chinese open-weight models for a moment: The training costs and resource requirements for frontier models are unsustainable. The high price, and social pushback, mean that the American companies producing these models are precarious. There are enormous financial incentives for research results allowing for cheaper, less resource-intensive models of high quality. Local LLMs on consumer hardware are akin to the PC hobbyist world of the 70s and 80s. Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now. Getting back to the Chinese models: They allow for new competition against Anthropic and OpenAI, basically SaaS renting out these very capable AIs much cheaper. That will just accelerate trends. geophile 过去50年计算机和软件市场的经验表明,免费和低端产品最终会获胜。 PC摧毁了小型计算机。大型机幸存下来,但市场份额远小于过去。 PC办公生产力软件摧毁了昂贵的专业产品。 Windows(低端)和Linux(免费)完全摧毁了UNIX市场,并且再次从大型机领域夺取了巨大市场份额。 暂时忽略中国庞大的开放权重模型: 前沿模型的训练成本和资源需求不可持续。高昂价格和社会的反弹意味着生产这些模型的美国公司处境不稳。 经济激励巨大,推动研究出更廉价、资源消耗更少的高质量模型。 消费级硬件上的本地LLM类似于70、80年代的PC爱好者世界。 综合这些趋势,我认为10-15年内,我们将会拥有能运行模型的消费级PC(和手机!),几乎能完成当前前沿模型能做的任何事情。 回到中国模型:它们为Anthropic和OpenAI带来了新的竞争,基本上以更低的成本出租这些功能强大的AI,这只会加速趋势。 Qwen 3.8 # https://news.ycombinator.com/item?id=48973549 I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source competitions—it’s got nothing to do with politics; you’ve simply lost sight of your original aspirations. cnhwl 我是来自中国的开发者。所以,这就是黑客新闻的德性吗?每次中国的模型一出现,评论区就不讨论技术架构、优化要点或实际性能,反而开始扯政治、人权这些垃圾?说实话,我们中国IT从业者有着真正的极客精神。这就是为什么你们在开源竞赛中落后——这跟政治无关,纯粹是你们自己忘了初心。 Show HN: I replaced a $120k bowling center system … # https://news.ycombinator.com/item?id=48971269 Thank you for posting this, it reaffirms what I’ve been thinking for so long. There are many opportunities to retrofit old systems of all types with modern low-cost embedded technologies. Around 2019 or so I was approached by an engineer who had a small business retrofitting very old machine tools with modern motion controls. Think very large lathes and planers. The problem they had was that in order to get these systems working with newer controls, they had to make time-consuming modifications to the old machines, in some cases modifying the axis motors and that cost added up quickly. The engineer realized that it was theoretically possible to take the old analog position signals and convert them into something a modern motion controller could read. That converter box would make the retrofit pretty much plug & play, but he didn’t have the programming expertise to make it happen. We probably built the first iteration of the converter for under $50 in parts and less than 50 hours of development time. That had me searching for other similar opportunities. I have found a few, but they tend to be one-offs that aren’t worth the time unless you’re already building something similar. Either that or my ability to see opportunity sucks! HeyLaughingBoy 感谢您发布这条内容,它印证了我长久以来的想法。有大量机会可以用现代低成本嵌入式技术改造各种旧系统。 大约2019年左右,一位工程师找到我,他经营着一家小公司,专门为非常古老的机床(比如大型车床和刨床)加装现代运动控制系统。他们面临的问题是,为了让这些设备兼容新控制系统,必须对旧机器进行耗时的改装,有时还需要改造轴电机,这些成本迅速累积。这位工程师意识到,理论上可以将旧的模拟位置信号转换为现代运动控制器能够读取的信号。这种转换器盒能让改造变得几乎即插即用,但他缺乏实现这一方案的编程专业知识。 我们大概用了不到50美元的零件和少于50小时的开发时间,就造出了转换器的第一个版本。这促使我寻找其他类似的机会。我确实找到了一些,但它们往往是一次性的项目,除非你已经在做类似的工作,否则不值得投入时间。要么就是我发现机会的能力太差了! Moonshot AI suspends new subscriptions due to Kimi… # https://news.ycombinator.com/item?id=48970885 Over the past 48 hours, demand has pushed close to the limits of our current capacity. To protect the experience of existing subscribers, we’re temporarily pausing new subscriptions and prioritizing compute for current members. Existing subscribed users are not affected. Such a beautiful paragraph to read, a company that prioritizes their current customers and focus on keeping them satisfied instead of just focusing on fast growth. Alifatisk 过去48小时内,需求已接近我们当前能力的极限。为了保护现有订阅者的体验,我们暂时暂停新订阅,并优先为现有会员提供计算资源。现有订阅用户不受影响。 多么优美的一段文字,一家优先考虑现有客户并专注于让他们满意,而不是只关注快速增长的公司。 Hacker wipes Romania’s land registry database # https://news.ycombinator.com/item?id=48978985 Since the hack, officials restored their website and posted a message announcing they are rebuilding the agency’s entire network from scratch. Even if the hacker claims they deleted backups, the agency appears to have had an offline copy, otherwise things would have gotten really messy over the coming months in Romania. So it seems not all has been lost. I was worried about the societal implications of being unable to prove land ownership but it seems that may be avoided. skinfaxi 自黑客攻击事件发生后,官员们恢复了网站并发布消息,宣布他们正从头重建该机构的整个网络。即便黑客声称已删除了备份,该机构似乎仍持有离线副本,否则未来几个月罗马尼亚的局势将变得十分混乱。 如此看来,并非一切尽失。我原本担心无法证明土地所有权会带来社会影响,但目前看来这种局面或许可以避免。 Exploit brokers pay $500k for WordPress RCEs. I fo… # https://news.ycombinator.com/item?id=48977439 There is no evidence that $500k has been paid or would be paid for an exploit like this one. Given that the article says that prompts are modified like they are holy scripture, perhaps sell the prompt for $500k. The author works for https://www.assetnote.io/ , which has AI products for automated scanning. Zsfe510asG 没有证据表明有人为这种漏洞支付或将要支付50万美元。 鉴于文章称提示被修改得像神圣经文一样,也许可以把提示卖50万美元。 作者供职于https://www.assetnote.io/,该公司提供自动化扫描的AI产品。 Claude Code uses Bun written in Rust now # https://news.ycombinator.com/item?id=48966826 Bah. Personally my take on the entire affair is quite negative, whatever Jarred or Simonw says about it. I think Bun owned by Anthropic and the entire rewrite with AI is not the real point (even if it’s quite interesting, though). My take is that Jarred, and Bun,didn’t demonstrate a serious, adult approach, from “this is my branch, you are overreacting” message to just proceeding with a 1mil+ PR merged in less than month. The communication is the issue, and it was handled very badly, in a way that impacted trust and divisions. Was it so difficult to adopt the approach that the TS team adopted for 7.0? gabrieledarrigo 哼。 就我个人而言,我对整件事的看法相当负面,不管Jarred或Simonw怎么说。 我认为Bun被Anthropic收购以及整个用AI重写的过程并不是真正的重点(尽管这确实挺有意思的)。 我的看法是,Jarred和Bun并没有展现出严肃、成熟的做法——从“这是我的分支,你们反应过度了”这样的消息,到在不到一个月的时间里直接合并了一个超过一百万行的PR。 问题在于沟通,而且处理得非常糟糕,以某种方式损害了信任并造成了分裂。 难道采用TypeScript团队为7.0版本所用的方法就那么难吗? Moonshine: Lets you stream games from your PC to a… # https://news.ycombinator.com/item?id=48974648 So, to sum up: First there was Nvidia Gamestream, which was proprietary and later deprecated by Nvidia themselves. Then came Sunshine and Moonlight, open source server and client implementations of the protocol to allow low latency streaming for all platforms. Along the way the Game on Whales project started, to allow multi seat streaming plus virtual displays so currently logged in user sessions aren’t impacted. Basically this is a whole “host your own Stadia/Geforce Now” solution. But instead of running already setup applications as an existing user, you run these curated containers of Steam, Firefox, etc. Then someone forked Sunshine and Moonlight into Apollo and Artemis, to make streaming from a virtual display turnkey and OOTB easy. Sadly the virtual display feature of Apollo is in practice Windows only. And then there’s this, Moonshine, in a way doing what Apollo brings to the table but for Linux servers. Oh, and along the way someone forked Apollo into Vibepollo, an almost fully vibecoded “enhancement”. Personally not touching that. Did I miss anything? PalmPilotProMax 那么,总结一下: 先是Nvidia Gamestream,专有技术,后来被Nvidia自己废弃了。 接着出现了Sunshine和Moonlight,作为该协议的开源服务器和客户端实现,为所有平台提供低延迟串流。 在这个过程中,Game on Whales项目启动了,旨在实现多席位串流和虚拟显示,从而不影响当前登录的用户会话。基本上,这是一个完整的“自建Stadia/GeForce Now”解决方案。但并非以现有用户身份运行已设置好的应用程序,而是运行这些精心打包的Steam、Firefox等容器。 后来有人将Sunshine和Moonlight分叉成了Apollo和Artemis,使得虚拟显示串流变得开箱即用、简单易行。遗憾的是,Apollo的虚拟显示功能实际上只适用于Windows。 然后就是现在这个,Moonshine,某种程度上实现了Apollo的功能,但针对的是Linux服务器。 哦,对了,中间还有人把Apollo分叉成了Vibepollo,一个几乎全靠AI编程“增强”的版本。我个人是不会碰的。 我漏了什么吗? China’s open-weights AI strategy is winning # https://news.ycombinator.com/item?id=48982883 What’s interesting/funny is that the American LLM companies took from the public domain and copyrighted work to close all that content into a box they charge for. Then the Chinese took the distilled stuff out from that box and released it into the world for everyone. xandrius 有趣/好笑的是,美国的大语言模型公司从公共领域和受版权保护的作品中获取内容,然后将所有这些内容封闭在一个收费的盒子里。然后中国人从这个盒子里提取出蒸馏后的东西,并将其释放到世界上供所有人使用。 Xiaomi-Robotics-1 # https://news.ycombinator.com/item?id=48975306 I have been slopfolding since before it was cool broodbucket 我早在它流行之前就开始slopfolding了。 Xiaomi-Robotics-1 # https://news.ycombinator.com/item?id=48975600 I’m not sure why some comments here are so pessimistic. I was grinning ear to ear while seeing the video. You’re telling me we’ve finally got robots that can do our laundry? And the model is at least as free as beer? To quote the lyrics/meme, we used to pray for times like this. This is what robotics should be used for. user_7832 不知道这里为什么有些人这么悲观。我看这视频时全程笑得合不拢嘴。你是说我们终于有能洗衣服的机器人了?而且这模型至少像免费啤酒一样自由? 用歌词/梗来说,我们以前祈祷的不就是这样的时刻吗。机器人就该干这个。 Claude Code uses Bun written in Rust now # https://news.ycombinator.com/item?id=48970377 Drilling into the original article where Jarred explained the reasoning behind the change, It’s pretty clear that under zig the team was doing things by hand that are automatic in rust. Humans and agents share one thing: they are both non-deterministic. He talks about the issue of tracking memory lifecycles manually in zig so it can be explicitly freed. As expected, this leads to a long list of bugs where people missed things. Rust does this automatically. It removes an entire class of errors from his backlog. From an engineering management perspective, this looks like a pretty good trade. The bonus here is that compiler errors are exactly the kind of deterministic guardrail you need to put around coding agents. Claude works really well if you give it a way to test for correctness and “make it compile” is a pretty good target. There’s a general version of this: the artifact you expose plus the test you run on it. Deterministic tests turn stochastic output into a hard guarantee. Wrote it up here if useful: https://michael.roth.rocks/blog/verification-surface/ mrothroc 深入阅读Jarred解释这一变更背后原因的原文章后,可以清楚看到,在Zig环境下团队需要手动完成那些在Rust中已是自动化的操作。 人类与AI代理有一个共同点:它们都具有非确定性。他提到在Zig中手动追踪内存生命周期以便显式释放的问题。不出所料,这导致了一系列因遗漏而引发的bug。 Rust则自动处理了这些。它从待办事项清单中移除了一整类错误。从工程管理的角度来看,这似乎是一个相当划算的权衡。 额外的好处在于,编译器错误正好是你需要在编码代理周围设置的那种确定性护栏。如果你能为Claude提供验证正确性的方法,那么“让它通过编译”就是一个相当不错的目标,Claude会表现得很好。 这里有一个更通用的版本:你所暴露的工件加上对其运行的测试。确定性测试将随机输出转化为硬性保证。如果觉得有用,可以在这里查阅:https://michael.roth.rocks/blog/verification-surface/ Show HN: I replaced a $120k bowling center system … # https://news.ycombinator.com/item?id=48968835 Right now I’m working on adding LED + DMX DJ light control - I kinda want to be able to order LED strips to “chase” a ball as it goes down the lane or back up the return. I plan on triggering laser-light shows and such with the DMX controller. Eventually, I want to allow a customer walk up to the lane, tap to pay and start bowling immediately, too. Kiosk-ize bowling alleys, yknow? I’m pumped. So much room for activities! section33 目前我正在添加LED和DMX DJ灯光控制——我有点想让LED灯带能够“追逐”球道上的球,无论是球滚下去还是从回球道回来。我计划用DMX控制器触发激光灯光秀之类的效果。 最终,我还想允许顾客走到球道前,轻触支付并立即开始打球。就是把保龄球馆自助服务终端化,你懂吧? 我很兴奋。有太多可以折腾的空间了! China’s open-weights AI strategy is winning # https://news.ycombinator.com/item?id=48983155 Try instructing Codex to (say) fine-tune a language model based on a collection of books you’ve got saved. You will find yourself admonished, repeatedly and at length, not to utilize copyrighted materials to train language models, by an AI who owes its entire existence to that very act. These models might be smart but they’re not close to being able to savor irony. hyperbovine 试着指导Codex去(比方说)基于你收藏的书籍微调一个语言模型。你会发现自己被一个完全依赖于使用版权材料训练而存在的人工智能反复且长篇大论地告诫:不要使用受版权保护的材料来训练语言模型。这些模型或许聪明,但远未能品味讽刺。 China’s open-weights AI strategy is winning # https://news.ycombinator.com/item?id=48979957 I’m suspicious of some quotes here, “80% of startups using Chinese models,” doesn’t seem quite right to me. I just interviewed at several startups and they were all using the US models. Maybe they have some minor use of Chinese models but the bread-and-butter of most of these businesses model use is the Claude and Codex subscriptions. tyleo 我对这里的一些说法持怀疑态度,“80%的初创公司使用中国模型”在我看来不太准确。我刚刚面试了几家初创公司,它们都在使用美国模型。也许它们稍微用了一些中国模型,但这些企业模型使用的核心还是Claude和Codex的订阅服务。 Exploit brokers pay $500k for WordPress RCEs. I fo… # https://news.ycombinator.com/item?id=48976285 https://github.com/WordPress/WordPress/commit/3a640e1c5e39aa60d98bd5a048b603402e70209c String concatenation SQL injection in the year 2026. progbits 2026年还存在字符串拼接导致的SQL注入漏洞。 AI advice made people less accurate but more confi… # https://news.ycombinator.com/item?id=48972275 This study is pretty bad. The comment ( https://news.ycombinator.com/item?id=48970182 ) on the other link with the direct PDF explains the problem well, which is that nothing here being tested is specific to AI systems. This study gave people access to an LLM that the researchers knew would give incorrect answers to certain questions, and then quizzed people on those questions, with the option to not respond to a given question if they are unsure about the answer. This is akin to giving someone a textbook on an obscure subject that has certain factual errors, letting them know they can use that textbook in a quiz on that subject, and then quizzing that person on those facts that the textbook gets wrong. Obviously that person is both more likely to be willing to respond to the question and is more likely to get it wrong! There are a lot of things I’m very interested in that are specific to modern LLMs and how they affect learning and confidence (sycophancy, cognitive helplessness, etc.). This study tested none of those. Its experimental setup is not very different than simply substituting the LLM with a textbook with errors. dwohnitmok 这项研究相当糟糕。另一个链接(https://news.ycombinator.com/item?id=48970182)中直接附有PDF的评论很好地解释了问题:这里测试的内容并非针对人工智能系统的特性。 该研究让参与者使用一个研究者已知会对某些问题给出错误答案的LLM,然后针对这些问题进行测试,同时允许参与者在对答案不确定时选择不回答。 这相当于给某人一本关于冷门学科的教科书,其中包含某些事实错误,让他们知道可以在该学科的测验中使用这本教科书,然后测试那些教科书出错的事实。 显然,那个人既更有可能愿意回答问题,也更有可能答错! 我对许多现代LLM特有的问题非常感兴趣,比如它们如何影响学习和自信(谄媚行为、认知无助等)。 这项研究完全没有测试这些。其实验设计与单纯用一本有错误的教科书替代LLM没有太大区别。 Hacker wipes Romania’s land registry database # https://news.ycombinator.com/item?id=48980243 This happened in a 50k people town where my father is from in 1982 with a BIG flood that destroyed the town land registry documents (among a lot of the town). Since he’s a lawyer, had first hand experience and I was always curious I asked many things about this a while back. Basically, what happened is that they rebuilt it from proof of ownership and testimonies of the people. You can never get to 100% recovery like that, but everyone knows who their neighbor is, at least in a town that is small like this. So they rebuilt it first from first hand proof, then by testimonies, with a period of counter claims available IIRC. For sure there were some false claims, but given the magnitude of the disaster, this is the best solution within that context. franciscop 这件事发生在1982年我父亲家乡的一个五万人口的小镇,一场大洪水摧毁了镇上的土地登记文件(以及镇上许多其他东西)。由于他是律师,有第一手经验,而且我一直很好奇,不久前我问了很多相关情况。基本上,他们是通过所有权证明和人们的证词重建了登记记录。这样永远无法做到100%恢复,但每个人都认识自己的邻居,至少在这种小镇上是这样。 所以他们首先根据第一手证明重建,然后根据证词,据我回忆,还有一段异议期。当然,肯定有一些虚假申报,但考虑到这场灾难的规模,这是当时背景下最好的解决方案。 Claude Fable produced a counterexample to the Jaco… # https://news.ycombinator.com/item?id=48974718 This is a rare instance where feeding this groundbreaking information into an LLM gives them psychosis. I fed this to claude code and watched it verify the result in 7 different ways to be 100% certain, and it was just flabbergasted. Quite remarkable. aizk 这是一个罕见的案例:将这种颠覆性的信息输入大语言模型,反而会让它产生“精神病”。我把它喂给了Claude Code,然后看着它用七种不同方式验证结果以确保100%准确,它完全惊呆了。相当不可思议。 What I learned selling 2,500 MIDI recorders: Hardw… # https://news.ycombinator.com/item?id=48970190 Hardware has a reputation for being hard for three reasons. First, it scales differently than software. It’s far harder to design something you want to make a million of than something you want to make ten of. Second, it’s harder to anticipate, test, and correct all the things that can go wrong on the user end. Some people will put batteries the wrong way, some will drop the device on the ground, some will connect it to some vintage equipment you have never seen in your life, etc. And third, there are just more failure modes in an unfamiliar domain - sometimes, your code will intermittently crash not because you have a software bug, but because you put the decoupling capacitor too far from the chip. Or, just as you finish your design, the chip you designed it around becomes obsolete, or impossible to find because a factory in Indonesia is on fire. Another complication is that at least in theory, if you’re selling electronics, there are actual regulations and third-party testing that needs to happen, and if you fail emissions, you might have to redo your design from scratch. Imagine we had that for software - “your JS is too big, you can’t ship until you get it under 50 kB”. So, I’m happy for the author, but I think he had an outlier experience. When you look at Kickstarter stories, people repeatedly stumble over this. Manufacturing / cost difficulties, supplier issues, reliability issues, etc. skippyfish 硬件之所以被认为困难,有三个原因。首先,它的扩展方式与软件不同。设计一个要量产百万件的产品,远比设计只做十件的样品困难得多。其次,更难预测、测试和修正用户端可能出现的各种问题。有人会把电池装反,有人会把设备摔在地上,还有人会把它接到你从未见过的老旧设备上等等。第三,在不熟悉的领域有更多故障模式——有时你的代码间歇性崩溃,并非因为软件bug,而是因为你把去耦电容放得离芯片太远。或者,正当你完成设计时,你所依赖的芯片停产了,或因为印尼某家工厂失火而无法采购。 另一个复杂因素是,至少在理论上,如果你销售电子产品,就需遵守相关法规并通过第三方测试。如果辐射超标,你可能要重新设计。想象一下软件领域也有类似规定——“你的JavaScript太大了,必须压缩到50KB以下才能发布”。 所以,我为作者感到高兴,但认为他的经历并不典型。看看Kickstarter上的案例,人们反复在这些问题上栽跟头:制造成本困难、供应商问题、可靠性问题等等。 AI advice made people less accurate but more confi… # https://news.ycombinator.com/item?id=48972168 Advice and information subreddits have gone to shit because of AI usage. A large number of people seem to think that when someone asks a question, what they really want is not someone with direct knowledge, but instead someone to relay the question to ChatGPT and post the result as if it is their own hard earned knowledge and insight. I have no idea about the quality of this research but in the real world (well, real-ish, as real as Reddit can be) it is stark, people aren’t just refusing to say “I don’t know” they’re actively seeking out opportunities to pretend they know things. reticulates 因为AI的滥用,提供建议和信息的子版块已经变得一团糟。似乎很多人认为,当有人提问时,对方真正想要的不是有直接经验的人,而是由某人将问题转述给ChatGPT,再把结果当作自己辛辛苦苦得来的知识和见解发出来。我不清楚这项研究有多靠谱,但在现实世界(好吧,算是比较现实,起码在Reddit上能算现实)里,情况已经很明显了——人们不仅拒绝说“我不知道”,反而积极主动地寻找机会假装自己什么都懂。