Anthropic’s new Project Fetch: Phase two report says Claude Opus 4.7 could complete the tracked robotics tasks about 20x faster than the fastest human team, and the Hacker News thread is already pushing on methodology and trust. This is a fresh major-lab benchmark for long-horizon agentic execution outside software. Sources: https://www.anthropic.com/research/project-fetch-phase-two ; https://news.ycombinator.com/item?id=48614311
Anthropic also published Agentic coding and persistent returns to expertise, a study of roughly 400,000 Claude Code sessions that finds humans mostly decide what to do while Claude mostly decides how to do it, with domain expertise strongly affecting session success. HN discussion is active and the framing is useful for anyone building coding agents or evaluating operator skill transfer. Sources: https://www.anthropic.com/research/claude-code-expertise ; https://news.ycombinator.com/item?id=48575785