We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control
deepmind.google·6d ago

TL;DR
To train the PRX model effectively, a diverse dataset was assembled from public and internal sources, focusing on breadth over aesthetic quality. Long, accurate captions were prioritized to enhance the model's understanding of visual concepts.
✦ Why It Matters
Engineers should prioritize using long, descriptive captions in their datasets to enhance model training outcomes.
Key Takeaways
How It Works
The data pipeline combines public and internal datasets, using a VLM to generate long, detailed captions that enhance the model's understanding of visual concepts. By prioritizing breadth and diversity, the model learns a wider range of visual attributes, which is crucial for effective training.
Related