Categories: FAANG

Referring to Screen Texts with Voice Assistants

Voice assistants help users make phone calls, send messages, create events, navigate, and do a lot more. However, assistants have limited capacity to understand their users’ context. In this work, we aim to take a step in this direction. Our work dives into a new experience for users to refer to phone numbers, addresses, email addresses, URLs, and dates on their phone screens. Our focus lies in reference understanding, which becomes particularly interesting when multiple similar texts are present on screen, similar to visual grounding. We collect a dataset and propose a lightweight…
AI Generated Robotic Content

Recent Posts

OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree

At the Black Hat security conference, the AI giant revealed new details about how its…

5 mins ago

AI models nearly erase female characters when they write kids stories about animals

Last year, Melanie Walsh, a University of Washington assistant professor in the Information School, wrote…

5 mins ago

Measuring Performance of Transformer Inference

This chapter is divided into eight parts; they are: • Metrics for LLM Inference •…

23 hours ago

Static vs. Dynamic vs. Continuous Batching in LLM Inference

In this article, you will learn how static, dynamic, and continuous batching work in LLM…

23 hours ago

Introducing Web Search on Amazon Bedrock for foundation model grounding

When a foundation model needs to answer a question about last week’s earnings call, yesterday’s…

23 hours ago

How Deutsche Bank unlocked agility with an API-ready ecosystem

When people think about digital transformation in banking, they often focus on the visible results:…

23 hours ago