Categories: FAANG

FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations

This paper was accepted at the Workshop on Foundation Models in the Wild at ICLR 2025.
Visual understanding is inherently contextual – what we focus on in an image depends on the task at hand. For instance, given an image of a person holding a bouquet of flowers, we may focus on either the person such as their clothing, or the type of flowers, depending on the context of interest. Yet, most existing image encoding paradigms represent an image as a fixed, generic feature vector, overlooking the potential needs of prioritizing varying visual information for different downstream use cases. In…
AI Generated Robotic Content

Recent Posts

Minimax H3 + RefMod = consistent location trick

Hey, I found a pretty cool way to keep locations consistent across generations. I took…

5 hours ago

The Nvidia Shield TV Is 7 Years Old. It Just Got a $100 Price Hike

The price of anything with memory is skyrocketing thanks to AI. Aging streaming devices are…

6 hours ago

What image model was used here?

Anyone knows what could've been used here? Which model generates such photorealism? I've been using…

1 day ago

Language Discrimination Improves Linguistic Learning in Multilingual Speech Models

Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched…

1 day ago

Early Talent Hiring at Palantir

What Hiring Managers value — and how they’ve built their careers at PalantirEditor’s Note: Technical Recruiter Rachel Vogel…

1 day ago

Sweep thousands of leases for compliance using Amazon Quick and the Adjudicated Query pattern

Checking tens of thousands of apartment leases against constantly changing state landlord-tenant laws, and proving…

1 day ago