Who Evaluates the Evaluations FP blogmax 1000x1000 1

Frontier and Center: Who evaluates the evaluations?

Editor’s note: Some of the most interesting questions in AI are being asked by information theoreticians, around how to provide context to an emerging class of AI agents. A few weeks ago, we waded into those waters with a blog about the Open Knowledge Format, a specification that formalizes the LLM-wiki pattern into a portable, …

Incentivizing Temporal-Awareness in Egocentric Video Understanding Models

Multimodal large language models (MLLMs) have recently shown strong performance in visual understanding, yet they often lack temporal awareness, particularly in egocentric settings where reasoning depends on the correct ordering and evolution of events. This deficiency stems in part from training objectives that fail to explicitly reward temporal reasoning and instead rely on frame-level spatial …

ML 20889 1

MCP tool design: Practical approaches and tradeoffs

When Model Context Protocol (MCP) tools underperform, the cause is rarely the protocol itself but the tool design. Many teams start by exposing an existing API as-is and trusting the agent to figure out the rest. It is a natural way to extend APIs to agentic systems and generative AI coding tools. For straightforward use …

2 AlphaEvolve logo wallmax 1000x1000 1

Solve harder problems with AlphaEvolve, now available to everyone on Google Cloud

Many of the most challenging and valuable problems in the world are related to optimization. Now, AI is now making these problems tractable. If you’ve ever tried to design a microchip, plan a delivery network, or optimize a training architecture for a large AI model, you know how hard it is to find the most …

Meet Biomni—an AI-powered biomedical co-scientist

In creating a comprehensive, AI-enabled research agent for the biomedical sciences, Stanford University researchers hope to speed innovation by eliminating the tedium of scientific legwork. Biomni, an AI-powered, multiskilled biomedical research agent, is no mere chatbot. It is a full-fledged “co-scientist” capable of designing and developing complex research workflows, said Jure Leskovec, the Alfred and …

Screenshot 2026 07 08 at 20031PM

Introducing Claude apps gateway for AWS

Enterprises deploying Claude Code and Claude Desktop across development teams need centralized control over access, cost, and policy. At scale, this is hard to manage: each developer needs an individual credential, settings must be distributed manually, and spend is difficult to track or cap. Without a centralized control point, governance is left to whatever tooling …