
TL;DR
Google released Gemini 3.5 Flash, a faster language model variant, directly to general availability at Google I/O without a preview phase. The model is being integrated across multiple Google products including the Gemini app and AI Mode, reaching billions of users globally.
✦ Why It Matters
Engineers can expect Gemini 3.5 Flash as a standard dependency in Google products and should plan integrations accordingly.
Key Takeaways
Full Summary
Gemini 3.5 Flash was unveiled at Google I/O and is now available globally across multiple platforms, including the Gemini app and Google Search. This model, identified as gemini-3.5-flash, supports 1,048,576 input tokens and 65,536 output tokens, with a knowledge cut-off in January 2025.
Notably, the pricing has increased significantly, costing three times more than the previous Flash Preview model and six times more than the Flash-Lite version. The new Interactions API, currently in beta, introduces server-side history management, enhancing user experience.
Despite the price hike, Google is deploying this model in many free consumer products, indicating a strategic move to test market tolerance for AI services. Benchmarking costs reveal that Gemini 3.5 Flash is more expensive to run than its predecessors, aligning with trends seen in other AI models from competitors.
Related