Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is particularly problematic for proactive video reasoning. We ask whether models can learn to think visually during …

ml 21715 2 create cluster

Introducing new Ray capabilities on SageMaker HyperPod

Today, we are announcing new Ray capabilities on Amazon SageMaker HyperPod that integrate Ray with the HyperPod purpose-built infrastructure for foundation model training and serving. Ray is an open-source framework that data scientists use to scale distributed Python workloads across clusters of GPUs, from distributed training with Ray Train to model serving with Ray Serve. …

From X-ray speckles to solar magnetic fields, AI shrinks data while keeping crucial details

Next-generation science experiments will collect more data than ever—so much so that they’ll surpass the capabilities of current data storage and analysis methods. To help, researchers at the Department of Energy’s SLAC National Accelerator Laboratory developed an AI method to compress large amounts of raw data without losing subtle details critical to scientific discovery. They …

AI agents can build consensus on a scale humans can’t

Everyone is familiar with the situation: A larger group of people plans to visit a restaurant together, but it can take time and sometimes a great deal of patience to agree on a time and place to meet. The better the participants know each other and their preferences, the faster they will reach an agreement. …