Daily Nuts.
第 113 期/27 项已收录
今日简报· 2026 年 9 月 2 日 · 周三

今日共采集 27 项内容

由 AI 自动整理·27 项条目·约 4 分钟阅读

日共采集 27 项内容。

本期无分主题展开

请向下查看条目列表。

今日条目27 项

按主题归类 · 3
筛选
来源
§ 01

模型架构与记忆

5
  • A
    arXiv 计算语言学
    论文·9 月 1 日

    LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. We investigate this procedure mechanistically through an eight-attack perturba

    #研究
    17:59 UTC
    ENZH
    阅读全文
  • A
    arXiv 计算语言学
    论文·9 月 1 日

    Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-benchmark performance, but are largely result-only

    #研究
    17:59 UTC
    ENZH
    阅读全文
  • A
    arXiv 计算语言学
    论文·9 月 1 日

    The repository-level code generation task requires synthesizing code that satisfies task requirements while remaining consistent with the target repository context. Since real-world repositories often exceed the input length limits of LLMs, existing approaches commonly adopt retrieval-augmented gene

    #研究
    17:59 UTC
    ENZH
    阅读全文
  • A
    arXiv 计算语言学
    论文·9 月 1 日

    Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility brings a new reasoning burden: a local plugin change can propagate through dependencies and cleanup. We introduce CordisBench, a 1,200-question benchmark of this lifecycle reasoning. It

    #研究
    17:59 UTC
    ENZH
    阅读全文
  • A
    arXiv 计算语言学
    论文·9 月 1 日

    Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal structure in forms interpretable by both humans and modern language models. We call this paradigm Verbal Reinforcement Learning (VRL) and offer the first uni

    #研究
    17:58 UTC
    ENZH
    阅读全文
§ 02

新模型发布

5
  • D
    DeepSeek GitHub
    发布·9 月 2 日

    [中文](#cn-v0.1.2-alpha.5) | [English](#en-v0.1.2-alpha.5) <h3 id="cn-v0.1.2-alpha.5">问题修复</h3> * 修复从 `0.1.1-rc.2` 或 `0.1.2-alpha.3` 升级时,应用可能启动失败或者会话列表标题丢失的问题 @imccyu --- <h3 id="en-v0.1.2-alpha.5">Bug Fixes</h3> * Fix an issue where upgrading from `0.1.1-rc.2` or `0.1.2-alpha.3` could

    #发布
    10:02 UTC
    ENZH
    阅读全文
  • D
    DeepSeek GitHub
    发布·9 月 1 日

    [中文](#cn-v0.1.2-alpha.4) | [English](#en-v0.1.2-alpha.4) <h3 id="cn-v0.1.2-alpha.4">新增功能</h3> * 父 Agent 与可持续子 Agent 可通过 `send_message` 双向传递后续消息,取代单向 `report` 工具 @Dudu-0223 <h3>体验优化</h3> * 自定义模型发现复用 Profile 请求头;模型目录支持搜索和筛选 @LegGasai * 界面优化圆角、描边、轮次导航、投影效果 @yixiangihsiang, @LegGasai * 改善超

    #发布
    15:45 UTC
    ENZH
    阅读全文
  • M
    Moonshot AI GitHub
    发布·9 月 2 日

    ### Patch Changes - [#3469](https://github.com/MoonshotAI/kimi-code/pull/3469) [`979baad`](https://github.com/MoonshotAI/kimi-code/commit/979baad8597aa1760917752b3663f1eb4e40eeb0) Thanks [@sailist](https://github.com/sailist)! - Fix the condition for showing the kimi-cli migration prompt.

    #发布
    09:20 UTC
    ENZH
    阅读全文
  • M
    Moonshot AI GitHub
    发布·9 月 2 日

    ### Minor Changes - [#3434](https://github.com/MoonshotAI/kimi-code/pull/3434) [`ae7a6dc`](https://github.com/MoonshotAI/kimi-code/commit/ae7a6dc6fb56cde119f0ac1512649a52c19ef7e8) Thanks [@sailist](https://github.com/sailist)! - The `kimi acp` subcommand no longer honors `KIMI_CODE_LEGACY_FLAG`; it

    #发布
    05:59 UTC
    ENZH
    阅读全文
  • M
    Moonshot AI GitHub
    发布·9 月 1 日

    ## What's Changed * fix(kosong): omit empty anthropic-beta header when no beta features declared by @7Sageer in https://github.com/MoonshotAI/kimi-cli/pull/2580 * chore(release): bump kosong to 0.56.0 by @jackfish212 in https://github.com/MoonshotAI/kimi-cli/pull/2581 * feat(shell): deprecation-awar

    #发布
    16:53 UTC
    ENZH
    阅读全文
§ 03

应用与产品

17
  • G
    Google Research 博客
    博客·9 月 1 日

    Climate & Sustainability

    #业界
    18:40 UTC
    ENZH
    阅读全文
  • A
    Anthropic 新闻
    博客·9 月 1 日

    Developing Enterprise Frontier Safeguards with our customers

    #业界
    00:00 UTC
    ENZH
    阅读全文
  • C
    Claude 博客
    博客·9 月 2 日

    How an Anthropic field marketer uses Claude Code to send weekly personalized updates to every sales rep

    #业界
    12:26 UTC
    ENZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 2 日

    美团 LongCat-2.0 上线 Cline 免费试用

    美团 LongCat-2.0 上线 Cline 免费试用

    #业界
    03:50 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 2 日

    Anthropic 发布 Claude Fable 5.1 与 Claude Mythos 5.1

    Anthropic 发布 Claude Fable 5.1 和 Claude Mythos 5.1,两者为同一模型,Mythos 5.1 仅通过受信任访问计划提供给网络安全和生命科学领域。

    #业界
    03:33 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 2 日

    UU远程新版本上线:完整 TUI 渲染与多终端会话管理,强化远程 Vibe Coding 体验

    UU远程于9月2日上线新版本,重点优化终端功能,补齐 TUI 渲染交互与终端会话管理能力。主要更新包括:Mac 免密码登录、移动端输入优化并新增调用系统输入法的独立输入框、可同时创建和管理多个终端会话并支持手机与电脑间跨端同步接管(通过 uuyc-cli lterm 命令)。

    #业界
    03:32 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 2 日

    Qwen3.8-Max-0902 登顶 Code Arena 并以 $5/MToken 领跑 Pareto 前沿

    通义千问发布 Qwen3.8-Max-0902,在 Code Arena: WebDev 以 1,691 分首次亮相即排名总榜第一,并以混合价 $5/MToken 成为 Pareto 前沿上得分最高的模型,现已可在 QwenCloud 试用。

    #业界
    02:57 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 2 日

    UC Berkeley 团队发布 Vero 基准:测试 AI 智能体能否构建形式化验证的软件仓库

    UC Berkeley 等机构发布 Vero,据称是首个要求智能体在仓库级同时编写实现与证明的基准,含 43 个多模块 Lean 4 实例、743 个计分 API 和 2705 条形式化规范。

    #业界
    02:28 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 2 日

    Nvidia 接近以 129 亿美元收购 Hugging Face

    Bloomberg 报道 Nvidia 正接近以约 129 亿美元收购 Hugging Face,交易总额可能达约 140 亿美元,双方尚未达成最终协议,时间与细节仍可能变动。该价格约为 Hugging Face 2023 年融资轮 45 亿美元估值的 2.9 倍,按年化收入约 1.5 亿美元计算相当于 86 倍,Nvidia 还谈及在交易中加入 10 亿美元的员工留任方案。

    #业界
    02:26 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 1 日

    Claude Fable 5.1 登顶 Artificial Analysis 智能指数,但每任务成本比 Fable 5 高 20%

    Artificial Analysis 评测 Claude Fable 5.1,其在 max effort 下得 66 分登顶 Artificial Analysis Intelligence Index。

    #业界
    20:12 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 1 日

    Fable 5.1 系统卡披露隐蔽任务与监控难度上升等安全发现

    Rohan Paul 梳理了 Fable 5.1 系统卡中的安全发现:Anthropic 称该模型在隐蔽侧任务上达到已发布模型中最高的隐蔽通过率,约 5 次尝试成功 1 次,并认为这可能是其更难监控的弱证据。

    #业界
    19:43 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 1 日

    Claude Fable 5.1 上线 Claude Code 与 Claude Platform,缓存读取降价 75%

    Claude Fable 5.1 现已上线 Claude Code 和 Claude Platform,定价与 Fable 5 相同,API 缓存读取便宜 75%。模型在长任务中能更久自主推进、更善于提示用户它已卡住,写作风格也更自然。Anthropic 同时发布了 Claude Fable 5.1 和 Claude Mythos 5.1,称其为编码与知识工作领域最先进的模型。

    #业界
    18:13 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 1 日

    Google DeepMind 为 Gemini 推出 agentic 视频理解功能

    Google DeepMind 为 Gemini 3.7 Flash、3.6 Flash 和 3.5 Flash-Lite 推出 agentic video understanding,模型动态扫描视频片段,相比固定帧率处理 token 消耗最多降低 88%,成本最多降低 66%,准确率最多提升 7%。

    #业界
    17:08 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 1 日

    Google Workspace 推出图像创作编辑工具 Google Pics

    Google 发布 Workspace 图像创作与编辑工具 Google Pics,将在未来数周内面向所有 Google AI Pro 和 Ultra 订阅者及多数 Workspace 商业客户推出。

    #业界
    16:00 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 1 日

    OpenAI 评定 Astra 达到网络安全 Critical 能力阈值,将受限发布

    OpenAI 宣布 Astra 在其 Preparedness Framework 下达到 Critical 网络安全能力阈值,是首个被评定为该级别的模型,可在少人干预下发现未知漏洞并构建利用链。

    #业界
    13:00 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 1 日

    路透社调查:美国 AI 数据中心现大量幽灵用电需求,得州等多州出手整治

    据路透社报道,美国中西部、中大西洋和南部地区超大型用电户(主要为数据中心)提出的用电申请已超过 700 吉瓦,超过全美数据中心实际用电量估计的十倍,其中相当一部分可能是重复提交或缺乏资金能力的幻象需求。

    #业界
    12:40 UTC
    ZHZH
    阅读全文
  • A
    AI HOT 精选
    博客·9 月 1 日

    Hugging Face 发布 @huggingface/kernels,提供 207 个 WebGPU 内核用于浏览器本地 AI 推理

    Hugging Face WebAI 团队发布 @huggingface/kernels 库及 207 个以独立仓库形式托管在 Hub 上的 WebGPU 内核(Apache-2.0),每个内核带 manifest、正确性测试、基准用例和 WGSL 着色器模板。

    #业界
    00:00 UTC
    ZHZH
    阅读全文