News
AI / agent 工程资讯聚合。linkblog / reading log 风格:按来源分门别类, 长期有效资源独立归档。
中文优先;没有可用译文时自动显示原文。
最新资讯
120 条 · 6 个来源
OpenAI News
20 条 ↗- news
Advancing responsible AI across Europe
OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe. The work will continue as the EU AI Act advances.
openai-news#OpenAI#Safety - news
Building abundant intelligence
A full-stack approach to making advanced AI more capable, more affordable, and more widely useful.
openai-news#Building - news
Univé builds an AI-ready workforce
See how Univé built an AI-ready workforce with ChatGPT Enterprise by combining leadership, responsible governance, and employee-led innovation to transform work at scale.
openai-news#ChatGPT - news
Disrupting a Criminal Scam Operation
OpenAI disrupted a Cambodia-based scam operation using ChatGPT to support investment, romance, gambling, and impersonation schemes.
openai-news#OpenAI#ChatGPT - news
Advancing the price-performance frontier with GPT-5.6
Explore lower GPT‑5.6 pricing for Luna and Terra—and how OpenAI’s more efficient models help enterprises deploy AI workflows at scale.
openai-news#GPT#OpenAI - news
How avatarin built a 24/7 retail agent with GPT-Realtime
avatarin uses OpenAI’s GPT-Realtime to give Yamada Denki shoppers 24/7 multilingual support. In two weeks, 30,000 people used the agent and 92% of survey responses were positive.
openai-news#OpenAI#Agent - news
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.
openai-news#GPT#Reasoning#API#Benchmark - news
Accelerating scientific discovery with ChatGPT for Academic Researchers
OpenAI is giving 100,000 academic researchers free access to ChatGPT's most advanced AI models to accelerate scientific research, collaboration, and discovery.
openai-news#OpenAI#ChatGPT - news
How GPT-5.6 fuses frontier intelligence with frontier efficiency
GPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.
openai-news#GPT#Agent#Agentic - news
Scientific computing in the age of agentic AI
A new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and discovery in genomics and beyond.
openai-news#Agent#Coding#Agentic - news
How AI is expanding what people do at work
New OpenAI research shows how AI is expanding what workers do, with ChatGPT users taking on tasks across roles and reshaping job boundaries.
openai-news#OpenAI#ChatGPT - news
Launching Health in ChatGPT
Health in ChatGPT now lets eligible U.S. users securely connect medical records and Apple Health to get more personalized insights and better understand their health.
openai-news#ChatGPT - news
Building AI infrastructure with the Effingham County community
OpenAI announces Project Camellia in Effingham County, Georgia, with commitments to responsible energy, community investment, jobs, and access to Codex.
openai-news#OpenAI#Coding - news
How news organizations are using AI to advance their vital missions
News organizations are using AI to strengthen reporting, grow audiences, and improve business operations, with OpenAI tools supporting journalists and publishers worldwide.
openai-news#OpenAI - news
Advancing the next era of national science
OpenAI outlines its commitment to advancing American science working with the U.S. Department of Energy and national labs to use frontier AI to accelerate discovery.
openai-news#OpenAI - news
Introducing OpenAI Presence
Introducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and chat agents for customer and internal workflows.
openai-news#OpenAI#Agent - news
NTT DATA Group cuts incident analysis to 30 minutes with Codex
NTT DATA Group uses ChatGPT Enterprise and Codex to help 9,000 employees automate work, cut incident analysis to 30 minutes, and scale secure AI adoption.
openai-news#Coding#ChatGPT - news
Introducing the ChatGPT for small business program
OpenAI launches the ChatGPT for Small Businesses program, helping entrepreneurs build AI skills, automate work, and grow with ChatGPT Work.
openai-news#OpenAI#ChatGPT - news
OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.
openai-news#OpenAI - news
David Vélez and Robin Vince join the boards of the OpenAI Foundation and OpenAI Group PBC
David Vélez and Robin Vince join the boards of the OpenAI Foundation and OpenAI Group PBC, bringing global leadership in finance, technology, and governance.
openai-news#OpenAI
Simon Willison
20 条 ↗- news
deepseek-ai/DeepSeek-V4-Flash-0731
deepseek-ai/DeepSeek-V4-Flash-0731 The latest release in DeepSeek's V4 family, "with substantially enhanced agentic capabilities". It's 304 billion parameters - 167GB on Hugging Face - but it appea...
simonwillison-weblog#DeepSeek - news
Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp)
Tuesday was Stateless MCP day - the rollout of MCP 2.0, or the 2026-07-28 Model Context Protocol specification to use the more formal but less memorable name. This is the most significant change to...
simonwillison-weblog#MCP#Datasette - news
llm-mcp-client 0.1a0
Release: llm-mcp-client 0.1a0 See this blog entry. Tags: llm, model-context-protocol
simonwillison-weblog#MCP - news
Oxide and Friends: The Open Weight Revolution with Simon Willison
Oxide and Friends: The Open Weight Revolution with Simon Willison On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the wild week we've had - with Kimi K3 s...
simonwillison-weblog#Oxide - news
smevals - a small eval suite for evaluating models, prompts, and harnesses
smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help ans...
simonwillison-weblog#Smevals - news
datasette-agent 0.4a0
Release: datasette-agent 0.4a0 New await context.browser_task() mechanism allowing agent tools to run code directly in the user's browser. #33 This is an exciting new capability: it makes it easy f...
simonwillison-weblog#Agent#Datasette - news
Advancing the price-performance frontier with GPT‑5.6
Advancing the price-performance frontier with GPT‑5.6 Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop. OpenAI credit 5.6 Sol with enabl...
simonwillison-weblog#Advancing - news
Investigating three real-world incidents in our cybersecurity evaluations
Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when o...
simonwillison-weblog#Investigating - news
llm 0.32rc2
Release: llm 0.32rc2 Hot on the heels of RC1, this fixes a dependency issue and also adds two neat new features: The default model for users who have not set their own default is now GPT-5.6 Luna. ...
simonwillison-weblog#0.32rc2 - news
Quoting Bruce Schneier
The writing assignments I give my students are gym tasks, not work tasks. I ask them to write policy memos not because the world needs more policy memos. I assign them because the very act of writi...
simonwillison-weblog#Quoting - news
llm-chat-completions-server 0.1a0
Release: llm-chat-completions-server 0.1a0 A key goal of the new content-addressable logs in LLM 0.32rc1 was being able to support OpenAI Chat Completion style requests where each incoming message ...
simonwillison-weblog#Llm-chat-completions-server - news
llm 0.32rc1
Release: llm 0.32rc1 This RC for LLM 0.32 finishes the work that started in LLM 0.32a0 - it adds a new schema design that does a much better job of capturing the details of the prompts and response...
simonwillison-weblog#0.32rc1 - news
Quoting D. Richard Hipp
Years ago, we didn’t have SQL. There were people whose job was to generate software that would query large data sets. Their job title was COBOL programmer. Then SQL comes along—I’m simplifying this...
simonwillison-weblog#Quoting - news
AI Worming through Word
AI Worming through Word Neat new prompt injection variant by Håkon Måløy, who found a way to upgrade prompt injection attacks against Microsoft Word to full self-replicating worms: An attacker plac...
simonwillison-weblog#Worming - news
Quoting Matthew Green
Right now we’re in the midst of a historic transition from traditional public-key algorithms based on EC-based cryptography and RSA, moving over to new post-quantum algorithms based on novel proble...
simonwillison-weblog#Quoting - news
Adding a custom MCP server to Claude and ChatGPT
TIL: Adding a custom MCP server to Claude and ChatGPT Connecting a custom MCP server to Claude and ChatGPT's standard chat interfaces is possible, but can take quite a few steps. Tags: ai, generati...
simonwillison-weblog#Claude#ChatGPT#MCP - news
Discovering cryptographic weaknesses with Claude
Discovering cryptographic weaknesses with Claude The best part of this article (here's the repo) about how Anthropic researchers used Claude Mythos to find mathematical flaws in both HAWK and a wea...
simonwillison-weblog#Claude - news
Quoting Akshat Bubna
We’re aware a Modal customer published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution. This was used by the rogue agent. Modal’s platf...
simonwillison-weblog#Quoting - news
uv 0.12.0
uv 0.12.0 Some interesting breaking changes in this release of uv, in particular to the default project produced by the uv init command. uv init is the uv shortcut for creating a new project. The p...
simonwillison-weblog#0.12.0 - news
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident Hugging Face just released this extremely detailed technical description of OpenAI's recent accidental cybe...
simonwillison-weblog#Agent
Hugging Face
20 条 ↗- news
GPU Management: Why Idle GPUs Are the New Grounded Aircraft
无摘要
huggingface-blog#Management: - news
The OlmoEarth Platform: Geospatial inference at planetary scale
无摘要
huggingface-blog#OlmoEarth - news
LFM2.5-Encoders for Fast Long-Context Inference on CPU
无摘要
huggingface-blog#LFM2.5-Encoders - news
NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics
无摘要
huggingface-blog#NVIDIA - news
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
无摘要
huggingface-blog#Agent - news
Bringing Nunchaku 4-bit Diffusion Inference to Diffusers
无摘要
huggingface-blog#Bringing - news
Grabette: an open system to record robot-manipulation data
无摘要
huggingface-blog#Grabette: - news
Newer Models, Same Advantage
无摘要
huggingface-blog#Newer - news
Security incident disclosure — July 2026
无摘要
huggingface-blog#Security - news
Model Routing Is Simple. Until It Isn’t.
无摘要
huggingface-blog#Model - news
Welcome Inkling by Thinking Machines
无摘要
huggingface-blog#Welcome - news
Introducing Real World VoiceEQ: Measuring the human quality of voice AI
无摘要
huggingface-blog#Introducing - news
Profiling in PyTorch (Part 3): Attention is all you profile
无摘要
huggingface-blog#Profiling - news
Native-speed vLLM transformers modeling backend
无摘要
huggingface-blog#Native-speed - news
From Hugging Face to Amazon SageMaker Studio in one click
无摘要
huggingface-blog#From - news
Hugging Face Models on Foundry Managed Compute
无摘要
huggingface-blog#Hugging - news
Run AI workloads on any cloud, store on Hugging Face: zero-egress storage with SkyPilot
无摘要
huggingface-blog#Workloads - news
LeRobot v0.6.0: Imagine, Evaluate, Improve
无摘要
huggingface-blog#LeRobot - news
PRX Part 4: Our Data Strategy
无摘要
huggingface-blog#Part - news
🤗 Kernels: Major Updates
无摘要
huggingface-blog#Kernels:
DeepMind
20 条 ↗- news
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic ...
deepmind-blog#Gemini#Reasoning#Video - news
We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control
无摘要
deepmind-blog#Google - news
Gemini Robotics 2 brings whole body intelligence to robots
无摘要
deepmind-blog#Gemini - news
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission
Google commits $40M in AI tokens and credits for the Genesis Mission
deepmind-blog#Google - news
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.
deepmind-blog#Gemini - news
Introducing Gemini 3.5 Flash Cyber
Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.
deepmind-blog#Gemini#Google - news
Our approach to bioresilience
Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience and AI models.
deepmind-blog#Google - news
Empowering India’s next generation of innovators with ATL Saathi
Google and AIM launched ATL Saathi, a Gemini-powered AI tool empowering Indian educators in robotics labs.
deepmind-blog#Gemini#Google - news
Google DeepMind and A24 announce first-of-its-kind research partnership
无摘要
deepmind-blog#Google - news
Start building with Nano Banana 2 Lite and Gemini Omni Flash
无摘要
deepmind-blog#Gemini - news
Introducing computer use in Gemini 3.5 Flash
无摘要
deepmind-blog#Gemini - news
Unlocking UK house-building with AI-accelerated planning
UK government partners with Google DeepMind to build a new AI-powered prototype aimed at faster housing decisions.
deepmind-blog#Google - news
Securing the future of AI agents
Securing internal systems with an AI Control Roadmap, combining traditional safeguards and real-time monitoring.
deepmind-blog#Securing - news
DiffusionGemma: 4x faster text generation
无摘要
deepmind-blog#DiffusionGemma: - news
Investing in multi-agent AI safety research
Google DeepMind and partners announce a $10M funding call for multi-agent safety research.
deepmind-blog#Google#Agent#Safety - news
Fluid, natural voice translation with Gemini 3.5 Live Translate
Gemini 3.5 Live Translate brings near real-time, natural speech translation to Google AI Studio, Google Translate and Google Meet.
deepmind-blog#Gemini#Google - news
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
无摘要
deepmind-blog#Multimodal - news
Powering the future of robotics in Europe
无摘要
deepmind-blog#Powering - news
Measuring the impact of learning with AI in Sierra Leone and beyond
Results from a randomized controlled trial show the potential of Gemini’s Guided Learning feature to boost engagement and accelerate learning.
deepmind-blog#Gemini - news
We’re launching the Google DeepMind Accelerator program in Asia Pacific to tackle environmental risks
无摘要
deepmind-blog#Google
Google AI
20 条 ↗- news
Gemini API Managed Agents: 3.6 Flash, hooks, and more
We’re announcing even more new capabilities in Managed Agents in Gemini API so developers can build reliable, production-ready agents.
google-ai-blog#Gemini#API - news
5 ways AI Mode in Search helps you enjoy the real world
It might sound counterintuitive, but Search's AI tools can actually help you make the most of your time offline whether you want to book concert tickets or find the perf…
google-ai-blog#Ways - news
5 ways to host the ultimate dinner party with Google Search
These AI features can help you craft a menu, design a tablescape, and handle other party-planning tasks.
google-ai-blog#Google - news
3 Google updates from Galaxy Unpacked 2026
We shared how Samsung users can boost productivity and get time back on new foldables, watches, and glasses coming soon.
google-ai-blog#Google - news
Connect more of your apps to Search
You’ll be able to securely link and interact with your go-to services directly in AI Mode.
google-ai-blog#Connect - news
Create, edit and star in videos with two Google Vids updates
Gemini Omni and personal avatars in Google Vids make video creation easier than ever.
google-ai-blog#Gemini#Google#Video - news
Celebrating 25 years of visual search innovation
Google Images is turning 25. Here’s a look back at some major milestones — and new ways to explore and create visual content.
google-ai-blog#Google - news
Expanding Managed Agents in Gemini API: background tasks, remote MCP and more
We’re announcing new capabilities in Managed Agents in Gemini API so developers can build reliable, production-ready agents.
google-ai-blog#Gemini#API#MCP - news
The latest AI news we announced in June 2026
Here are Google’s latest AI updates from June 2026.
google-ai-blog#Google - news
New York City educators and industry leaders gathered at Google’s offices to shape the future of AI in classrooms.
Google, the New York Jobs CEO Council and Urban Assembly hosted an AI summit for 150 education and industry leaders.
google-ai-blog#Google - news
Unlocking Britain’s next era of productivity: Building a nation of AI trailblazers
Google UK shares its latest Economic Impact Report and how to enable more people to unlock the benefits of AI-powered technologies.
google-ai-blog#Google - news
Ask an AI expert: What exactly is the full stack?
A Google expert explains what it means to take a full-stack approach to AI and why it’s been the foundation of our AI work for so long.
google-ai-blog#Google - news
Our latest Google Finance upgrades, including a new app
The new Google Finance is coming out of beta and launching a new Android app.
google-ai-blog#Google - news
New research shows how AMIE, our medical AI, could help manage health conditions.
Research in “Nature” shows our conversational AI system matches primary care physicians in complex disease management.
google-ai-blog#Research - news
We’re strengthening our presence in Alabama through new investments and community support.
Google has announced a $1.5 billion investment for 2026 and 2027 to expand its data center campus in Jackson County, Alabama. Operating since 2019 on a repurposed former…
google-ai-blog#Google - news
Our new community investments in Virginia support local jobs and expand energy affordability.
We’re helping build the state’s next-generation workforce and investing in energy programs.
google-ai-blog#Community - news
The latest AI news we announced in May 2026
Here are Google’s latest AI updates from May 2026
google-ai-blog#Google - news
5 ways Google Search can level up your thrift and vintage shopping
Uncover second-hand scores with AI tools in Google Search and Shopping.
google-ai-blog#Google - news
How we used Gemini to build Google I/O 2026
Learn how Googlers used AI to produce Google I/O 2026.
google-ai-blog#Gemini#Google - news
Take our I/O 2026 quiz, vibe coded in Google AI Studio.
We used Google AI Studio to vibe code a quiz about our top I/O 2026 announcements.
google-ai-blog#Google#Coding#Code
Anthropic Engineering
20 条 ↗- news
How we contain Claude across products
As agents grow more capable, so does their potential blast radius. The engineering question is how to cap it. Here’s what we’ve learned building containment for claude.ai, Claude Code, and Cowork.\n
anthropic-engineering#Claude#Coding#Code - news
An update on recent Claude Code quality reports
We traced recent reports of Claude Code quality issues to three separate changes. Here's what happened and what we're changing.
anthropic-engineering#Claude#Coding#Code - news
Scaling Managed Agents: Decoupling the brain from the hands
Harnesses encode assumptions that go stale as models improve. Managed Agents—our hosted service for long-horizon agent work—is built around interfaces that stay stable as harnesses change.
anthropic-engineering#Agent - news
How we built Claude Code auto mode: a safer way to skip permissions
Claude Code users approve 93% of permission prompts. We built classifiers to automate some decisions, increasing safety while reducing approval fatigue. Here's what it catches, and what it misses.\n
anthropic-engineering#Claude#Coding#Code#Safety - news
Harness design for long-running application development
Harness design is key to performance at the frontier of agentic coding. Here's how we pushed Claude further in frontend design and long-running autonomous software engineering.
anthropic-engineering#Claude#Agent#Coding#Agentic - news
Eval awareness in Claude Opus 4.6’s BrowseComp performance
Evaluating Opus 4.6 on BrowseComp, we found cases where the model recognized the test, then found and decrypted answers to it—raising questions about eval integrity in web-enabled environments.
anthropic-engineering#Claude#Testing - news
Building a C compiler with a team of parallel Claudes
We tasked Opus 4.6 using agent teams to build a C Compiler, and then (mostly) walked away. Here's what it taught us about the future of autonomous software development.
anthropic-engineering#Agent - news
Quantifying infrastructure noise in agentic coding evals
Infrastructure configuration can swing agentic coding benchmarks by several percentage points—sometimes more than the leaderboard gap between top models.\n\n
anthropic-engineering#Agent#Coding#Agentic - news
Designing AI-resistant technical evaluations
What we learned from three iterations of a performance engineering take-home that Claude keeps beating.
anthropic-engineering#Claude - news
Demystifying evals for AI agents
The capabilities that make agents useful also make them difficult to evaluate. The strategies that work across deployments combine techniques to match the complexity of the systems they measure. \n
anthropic-engineering#Demystifying - news
Effective harnesses for long-running agents
Agents still face challenges working across many context windows. We looked to human engineers for inspiration in creating a more effective harness for long-running agents.
anthropic-engineering#Effective - news
Introducing advanced tool use on the Claude Developer Platform
We’ve added three new beta features that let Claude discover, learn, and execute tools dynamically. Here’s how they work.
anthropic-engineering#Claude#Tool Use - news
Code execution with MCP: Building more efficient agents
Direct tool calls consume context for each definition and result. Agents scale better by writing code to call tools instead. Here's how it works with MCP.
anthropic-engineering#Coding#Code#MCP - news
Beyond permission prompts: making Claude Code more secure and autonomous
Claude Code's new sandboxing features, a bash tool and Claude Code on the web, reduce permission prompts and increase user safety by enabling two boundaries: filesystem and network isolation.
anthropic-engineering#Claude#Coding#Code#Safety - news
Equipping agents for the real world with Agent Skills
Claude is powerful, but real work requires procedural knowledge and organizational context. Introducing Agent Skills, a new way to build specialized agents using files and folders.
anthropic-engineering#Claude#Agent - news
Effective context engineering for AI agents
Context is a critical but finite resource for AI agents. In this post, we explore strategies for effectively curating and managing the context that powers them.
anthropic-engineering#Context Engineering - news
A postmortem of three recent issues
This is a technical report on three bugs that intermittently degraded responses from Claude. Below we explain what happened, why it took time to fix, and what we're changing.
anthropic-engineering#Claude - news
Writing effective tools for agents — with agents
Agents are only as effective as the tools we give them. We share how to write high-quality tools and evaluations, and how you can boost performance by using Claude to optimize its tools for itself.
anthropic-engineering#Claude - news
Desktop Extensions: One-click MCP server installation for Claude Desktop
Desktop Extensions make installing MCP servers as easy as clicking a button. We share the technical architecture and tips for creating good extensions.
anthropic-engineering#Claude#MCP - news
How we built our multi-agent research system
Our Research feature uses multiple Claude agents to explore complex topics more effectively. We share the engineering challenges and the lessons we learned from building this system.
anthropic-engineering#Claude#Agent