Categories: FAANG

PROOF-Gen: From Optimized Data to Better Distillation

Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher’s passing trajectories, discard the rest) and each cycle leaves behind the same hard scenarios because failures supply no signal. On τ 2-bench, 57% of teacher trials fail, two-thirds of them near-misses (most tool calls correct, undone…
AI Generated Robotic Content

Recent Posts

GoT cast as Lebanese families

submitted by /u/Rokkit_man [link] [comments]

3 hours ago

Optimizing cost and latency with Amazon Bedrock prompt caching

Prompt caching in Amazon Bedrock can reduce your input token costs by up to 90…

3 hours ago

AI ‘Actor’ Tilly Norwood Told Me That ‘All Lives Matter’

The virtual character, which is promoting its upcoming movie Misaligned, tries to evade politics by…

4 hours ago

The shape behind the Einstein problem just revealed strange new physics

A mathematical shape famous for covering a surface without ever repeating has revealed an unexpected…

4 hours ago

AI can sound empathetic and human—but not at the same time

AI-generated texts are increasingly perceived as human, but people can still recognize human writing as…

4 hours ago

I trained the missing encoder for YuE2, so we can all bring our own music into it

YuE2 is an impressive open music model. Give it a style prompt and lyrics, and…

1 day ago