Categories: FAANG

FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations

This paper was accepted at the Workshop on Foundation Models in the Wild at ICLR 2025.
Visual understanding is inherently contextual – what we focus on in an image depends on the task at hand. For instance, given an image of a person holding a bouquet of flowers, we may focus on either the person such as their clothing, or the type of flowers, depending on the context of interest. Yet, most existing image encoding paradigms represent an image as a fixed, generic feature vector, overlooking the potential needs of prioritizing varying visual information for different downstream use cases. In…
AI Generated Robotic Content

Recent Posts

15 Best Office Chairs of 2026—We Tested 70 to Pick Them

Upgrade your WFH setup and work in style with these comfy, WIRED-tested seats.

1 day ago

AI reduces sensory hallucinations, even at night or in smoke

Multimodal large language models (MLLMs), which process multiple types of sensory information, such as text,…

1 day ago

Modeling Device Capabilities for Analytics

by Aarti Laddha, Richard Diaz-Cool, Rishika Idnani, Venkatesh SelverajNetflix supports a vast and evolving set…

2 days ago

Announcing the Agentic Catalog Experience in Amazon Quick

As organizations embrace AI-powered analytics, the value of a natural language (Text2SQL) answer is only…

2 days ago

What’s new in AI infrastructure and orchestration this month

At Google, AI is a soup-to-nuts endeavor. Obviously, we make leading AI models like Gemini…

2 days ago

SpaceX’s Falcon 9 Rocket Is About to Crash Into the Moon—and It Could Be Visible From Earth

The impact will kick up a plume of debris so high, it’ll likely be visible…

2 days ago