今日共采集 27 项内容
今日共采集 27 项内容。
本期无分主题展开
请向下查看条目列表。
今日条目27 项
模型架构与记忆
5 项- AarXiv 计算语言学论文·9 月 1 日
LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. We investigate this procedure mechanistically through an eight-attack perturba
#研究17:59 UTCEN → ZH阅读全文 → - AarXiv 计算语言学论文·9 月 1 日
Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-benchmark performance, but are largely result-only
#研究17:59 UTCEN → ZH阅读全文 → - AarXiv 计算语言学论文·9 月 1 日
The repository-level code generation task requires synthesizing code that satisfies task requirements while remaining consistent with the target repository context. Since real-world repositories often exceed the input length limits of LLMs, existing approaches commonly adopt retrieval-augmented gene
#研究17:59 UTCEN → ZH阅读全文 → - AarXiv 计算语言学论文·9 月 1 日
Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility brings a new reasoning burden: a local plugin change can propagate through dependencies and cleanup. We introduce CordisBench, a 1,200-question benchmark of this lifecycle reasoning. It
#研究17:59 UTCEN → ZH阅读全文 → - AarXiv 计算语言学论文·9 月 1 日
Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal structure in forms interpretable by both humans and modern language models. We call this paradigm Verbal Reinforcement Learning (VRL) and offer the first uni
#研究17:58 UTCEN → ZH阅读全文 →
新模型发布
5 项- DDeepSeek GitHub发布·9 月 2 日
[中文](#cn-v0.1.2-alpha.5) | [English](#en-v0.1.2-alpha.5) <h3 id="cn-v0.1.2-alpha.5">问题修复</h3> * 修复从 `0.1.1-rc.2` 或 `0.1.2-alpha.3` 升级时,应用可能启动失败或者会话列表标题丢失的问题 @imccyu --- <h3 id="en-v0.1.2-alpha.5">Bug Fixes</h3> * Fix an issue where upgrading from `0.1.1-rc.2` or `0.1.2-alpha.3` could
#发布10:02 UTCEN → ZH阅读全文 → - DDeepSeek GitHub发布·9 月 1 日
[中文](#cn-v0.1.2-alpha.4) | [English](#en-v0.1.2-alpha.4) <h3 id="cn-v0.1.2-alpha.4">新增功能</h3> * 父 Agent 与可持续子 Agent 可通过 `send_message` 双向传递后续消息,取代单向 `report` 工具 @Dudu-0223 <h3>体验优化</h3> * 自定义模型发现复用 Profile 请求头;模型目录支持搜索和筛选 @LegGasai * 界面优化圆角、描边、轮次导航、投影效果 @yixiangihsiang, @LegGasai * 改善超
#发布15:45 UTCEN → ZH阅读全文 → - MMoonshot AI GitHub发布·9 月 2 日
### Patch Changes - [#3469](https://github.com/MoonshotAI/kimi-code/pull/3469) [`979baad`](https://github.com/MoonshotAI/kimi-code/commit/979baad8597aa1760917752b3663f1eb4e40eeb0) Thanks [@sailist](https://github.com/sailist)! - Fix the condition for showing the kimi-cli migration prompt.
#发布09:20 UTCEN → ZH阅读全文 → - MMoonshot AI GitHub发布·9 月 2 日
### Minor Changes - [#3434](https://github.com/MoonshotAI/kimi-code/pull/3434) [`ae7a6dc`](https://github.com/MoonshotAI/kimi-code/commit/ae7a6dc6fb56cde119f0ac1512649a52c19ef7e8) Thanks [@sailist](https://github.com/sailist)! - The `kimi acp` subcommand no longer honors `KIMI_CODE_LEGACY_FLAG`; it
#发布05:59 UTCEN → ZH阅读全文 → - MMoonshot AI GitHub发布·9 月 1 日
## What's Changed * fix(kosong): omit empty anthropic-beta header when no beta features declared by @7Sageer in https://github.com/MoonshotAI/kimi-cli/pull/2580 * chore(release): bump kosong to 0.56.0 by @jackfish212 in https://github.com/MoonshotAI/kimi-cli/pull/2581 * feat(shell): deprecation-awar
#发布16:53 UTCEN → ZH阅读全文 →
应用与产品
17 项- GGoogle Research 博客博客·9 月 1 日
Climate & Sustainability
#业界18:40 UTCEN → ZH阅读全文 → - AAnthropic 新闻博客·9 月 1 日
Developing Enterprise Frontier Safeguards with our customers
#业界00:00 UTCEN → ZH阅读全文 → - CClaude 博客博客·9 月 2 日
How an Anthropic field marketer uses Claude Code to send weekly personalized updates to every sales rep
#业界12:26 UTCEN → ZH阅读全文 → - AAI HOT 精选博客·9 月 2 日
美团 LongCat-2.0 上线 Cline 免费试用
美团 LongCat-2.0 上线 Cline 免费试用
#业界03:50 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 2 日
Anthropic 发布 Claude Fable 5.1 与 Claude Mythos 5.1
Anthropic 发布 Claude Fable 5.1 和 Claude Mythos 5.1,两者为同一模型,Mythos 5.1 仅通过受信任访问计划提供给网络安全和生命科学领域。
#业界03:33 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 2 日
UU远程新版本上线:完整 TUI 渲染与多终端会话管理,强化远程 Vibe Coding 体验
UU远程于9月2日上线新版本,重点优化终端功能,补齐 TUI 渲染交互与终端会话管理能力。主要更新包括:Mac 免密码登录、移动端输入优化并新增调用系统输入法的独立输入框、可同时创建和管理多个终端会话并支持手机与电脑间跨端同步接管(通过 uuyc-cli lterm 命令)。
#业界03:32 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 2 日
Qwen3.8-Max-0902 登顶 Code Arena 并以 $5/MToken 领跑 Pareto 前沿
通义千问发布 Qwen3.8-Max-0902,在 Code Arena: WebDev 以 1,691 分首次亮相即排名总榜第一,并以混合价 $5/MToken 成为 Pareto 前沿上得分最高的模型,现已可在 QwenCloud 试用。
#业界02:57 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 2 日
UC Berkeley 团队发布 Vero 基准:测试 AI 智能体能否构建形式化验证的软件仓库
UC Berkeley 等机构发布 Vero,据称是首个要求智能体在仓库级同时编写实现与证明的基准,含 43 个多模块 Lean 4 实例、743 个计分 API 和 2705 条形式化规范。
#业界02:28 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 2 日
Nvidia 接近以 129 亿美元收购 Hugging Face
Bloomberg 报道 Nvidia 正接近以约 129 亿美元收购 Hugging Face,交易总额可能达约 140 亿美元,双方尚未达成最终协议,时间与细节仍可能变动。该价格约为 Hugging Face 2023 年融资轮 45 亿美元估值的 2.9 倍,按年化收入约 1.5 亿美元计算相当于 86 倍,Nvidia 还谈及在交易中加入 10 亿美元的员工留任方案。
#业界02:26 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 1 日
Claude Fable 5.1 登顶 Artificial Analysis 智能指数,但每任务成本比 Fable 5 高 20%
Artificial Analysis 评测 Claude Fable 5.1,其在 max effort 下得 66 分登顶 Artificial Analysis Intelligence Index。
#业界20:12 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 1 日
Fable 5.1 系统卡披露隐蔽任务与监控难度上升等安全发现
Rohan Paul 梳理了 Fable 5.1 系统卡中的安全发现:Anthropic 称该模型在隐蔽侧任务上达到已发布模型中最高的隐蔽通过率,约 5 次尝试成功 1 次,并认为这可能是其更难监控的弱证据。
#业界19:43 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 1 日
Claude Fable 5.1 上线 Claude Code 与 Claude Platform,缓存读取降价 75%
Claude Fable 5.1 现已上线 Claude Code 和 Claude Platform,定价与 Fable 5 相同,API 缓存读取便宜 75%。模型在长任务中能更久自主推进、更善于提示用户它已卡住,写作风格也更自然。Anthropic 同时发布了 Claude Fable 5.1 和 Claude Mythos 5.1,称其为编码与知识工作领域最先进的模型。
#业界18:13 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 1 日
Google DeepMind 为 Gemini 推出 agentic 视频理解功能
Google DeepMind 为 Gemini 3.7 Flash、3.6 Flash 和 3.5 Flash-Lite 推出 agentic video understanding,模型动态扫描视频片段,相比固定帧率处理 token 消耗最多降低 88%,成本最多降低 66%,准确率最多提升 7%。
#业界17:08 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 1 日
Google Workspace 推出图像创作编辑工具 Google Pics
Google 发布 Workspace 图像创作与编辑工具 Google Pics,将在未来数周内面向所有 Google AI Pro 和 Ultra 订阅者及多数 Workspace 商业客户推出。
#业界16:00 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 1 日
OpenAI 评定 Astra 达到网络安全 Critical 能力阈值,将受限发布
OpenAI 宣布 Astra 在其 Preparedness Framework 下达到 Critical 网络安全能力阈值,是首个被评定为该级别的模型,可在少人干预下发现未知漏洞并构建利用链。
#业界13:00 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 1 日
路透社调查:美国 AI 数据中心现大量幽灵用电需求,得州等多州出手整治
据路透社报道,美国中西部、中大西洋和南部地区超大型用电户(主要为数据中心)提出的用电申请已超过 700 吉瓦,超过全美数据中心实际用电量估计的十倍,其中相当一部分可能是重复提交或缺乏资金能力的幻象需求。
#业界12:40 UTCZH → ZH阅读全文 → - AAI HOT 精选博客·9 月 1 日
Hugging Face 发布 @huggingface/kernels,提供 207 个 WebGPU 内核用于浏览器本地 AI 推理
Hugging Face WebAI 团队发布 @huggingface/kernels 库及 207 个以独立仓库形式托管在 Hub 上的 WebGPU 内核(Apache-2.0),每个内核带 manifest、正确性测试、基准用例和 WGSL 着色器模板。
#业界00:00 UTCZH → ZH阅读全文 →