From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers

Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holistic scalar objective. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during …

ML 21711 1

Build agentic creative workflows with Amazon Quick and fal

Creative teams face growing demand for more assets, formats, and revisions, while their scripts, references, models, and outputs often remain fragmented across tools. Creators must repeatedly transfer context and assemble results manually. With 78% of creative leaders saying demand exceeds their teams’ capacity, faster generation alone does not solve the underlying workflow problem. To address …

Reimagining work: How Pythian’s internal AI playbook delivers customer ROI

When Pythian rolled out Google Cloud’s Gemini Enterprise across our 500-person company in 27 countries, the goal was simple: use our own company as a proving ground to discover how enterprise AI actually delivers ROI. What we found changed our strategy entirely. Since the rollout of Gemini Enterprise and our previous enterprise AI deployments, Pythian …

Enterprise AI’s real risk isn’t autonomous agents. It’s the complexity between them.

Presented by Gravitee Agent complexity is the insidious shadow lurking inside enterprises right now that needs a light shone on it. That’s because enterprises don’t deploy a single agent and watch it run, they deploy fleets, each one calling APIs, calling other agents, reaching into applications that were never built with a machine decision-maker in …

PROOF-Gen: From Optimized Data to Better Distillation

Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher’s passing trajectories, discard the rest) and each cycle …