ai/ml

Optimizing cost and latency with Amazon Bedrock prompt caching

Prompt caching in Amazon Bedrock can reduce your input token costs by up to 90 percent when you repeatedly send…

5 days ago

Abnormal AI: Amazon Bedrock AgentCore for agentic email security at scale

AI agents now run in production at a scale of billions of operations a day, and a recurring architectural pattern…

6 days ago

DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

Sign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension.…

1 week ago

Monitoring production agent lifecycle with AWS DevOps Agent and AgentCore Evaluations

Multi-agent systems in production experience issues in ways that traditional monitoring misses. For example, the agent can’t invoke its foundation…

1 week ago

Reminder: Live Today — Building AI Agents, The Loop

Quick note — The Loop’s first session is today, 4:30 PM PDT, live on Zoom.Free, monthly, and genuinely hands-on: what building an AI…

1 week ago

Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference

When you build an application on top of a large language model (LLM), the prompt you send to the model…

1 week ago

Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

On August 12, 2026, Alibaba’s Qwen team released Qwen3.8-2.4T-A95B. This is the first time a Qwen-Max-class model has been made…

2 weeks ago

Google is a Leader in the 2026 Gartner® Magic Quadrant™ for Enterprise AI Assistants

We are excited to share that Gartner has named Google a Leader in its inaugural 2026 Magic Quadrant for Enterprise…

2 weeks ago

Take on your most ambitious work with GPT-6 Astra on Amazon Bedrock

GPT-6 Astra from OpenAI brings greater depth and judgment to your most demanding tasks and runs on the Amazon Bedrock…

2 weeks ago

How KDDI built Buffmee, a faster, reliable consumer RAG app

When building consumer-facing generative AI applications,  balancing high generation quality with fast response times across diverse media types, can be…

2 weeks ago