We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control
deepmind.google·6d ago
TL;DR
Researchers identified a gap in LLM evaluation benchmarks. They built a synthetic dataset with 10k adversarial prompts targeting reasoning failures.
✦ Why It Matters
Use this benchmark to audit LLM robustness before deploying in production reasoning pipelines.
Key Takeaways
How It Works
Agents in Visual Studio Code now retain memory across sessions, allowing them to remember user preferences and project context. This is achieved through improved context management that prioritizes relevant information and allows users to control what is retained or discarded.
The ability to guide agents during their response generation helps maintain focus on the task at hand.
Related