Gemini AI Models Signal a Shift to Agentic, Low-Latency Coding
Google’s latest Gemini AI models and Android Studio Quail 2 update together mark a shift from simple code completion toward agentic workflows, where AI systems run multiple tasks, follow complex instructions, and respond with low latency while maintaining high token efficiency in developer tools and consumer search products. This is not just another model refresh; it is a structural change in how coding and search will feel day-to-day. By combining faster Gemini AI models with deeper IDE and search integration, Google is betting that AI-assisted development will move from occasional helper to always-on agent. Developers and ordinary users alike will notice the difference less in benchmark charts and more in how quickly they can ask, iterate, and ship.

Android Studio Quail 2: AI-Assisted Development Goes Parallel
Android Studio Quail 2 takes AI-assisted development from a single-threaded novelty to a practical, parallel co-worker. Agent Mode has been redesigned to allow multiple AI conversations at once, so the familiar bottleneck of waiting for one long response before starting another task is gone. Developers can now kick off a UI refactor in one tab, fix a ProGuard rule in a second, and generate documentation in a third, all powered by Gemini AI models and other large language models. This is the first time Android Studio feels like a place where AI agents can run alongside you, not in front of you. LeakCanary integration moves heap analysis off resource-constrained devices and onto the developer machine, making leak tracing up to five times faster and jank-free. When a leak or crash appears, the agent can explain the root cause, propose a step-by-step fix, and let the developer apply or edit that plan. That is real, opinionated assistance, not autocomplete dressed up as intelligence.
Gemini 3.6 Flash and 3.5 Flash-Lite: Token Efficiency With Teeth
Google’s new Gemini 3.6 Flash and 3.5 Flash-Lite models are built for speed and cost control, and the numbers matter because they reshape where AI is economically viable. Gemini 3.6 Flash consumes 17% fewer output tokens overall than 3.5 Flash, with savings reaching up to 65% on coding benchmarks like DeepSWE. For intensive workloads, Gemini 3.5 Flash-Lite delivers up to 350 output tokens per second and is priced at USD 0.30 (approx. RM1.23) per million input tokens and USD 2.50 (approx. RM10.24) per million output tokens, compared with USD 1.50 (approx. RM6.14) input and USD 7.50 (approx. RM30.71) output for 3.6 Flash. That gap is not a minor optimization; it is a design choice that makes high-volume agentic workflows viable rather than a budget risk. Both models add computer use as a built-in tool through the Gemini API and enterprise offerings, and include enhanced safety safeguards against chemical, biological, radiological, nuclear, and cyber risks. In practice, these Gemini AI models encourage developers to wire AI deeper into their systems without flinching at latency or token efficiency trade-offs.
3.5 Flash-Lite in Search: Agentic Workflows Leave the Lab
The most telling sign of confidence in Gemini 3.5 Flash-Lite is that Google Search is already using it for agentic search experiences, with likely use in AI Overviews and AI Mode. According to Google, “3.5 Flash-Lite is also rolling out in Google Search” and is designed for low-latency, high-throughput scenarios such as agentic search and document processing. This model is described as Google’s fastest, most cost-effective 3.5-class option, again delivering 350 output tokens per second and significantly outperforming prior Flash-Lite versions in agentic workflows. It improves instruction following and understanding of user intent, so conversations in search and AI modes flow more smoothly for everyday users, not just developers. Importantly, 3.5 Flash-Lite enables efficient scaling for agentic systems, and on many coding and agentic evaluations it even beats some 3 Flash variants. Search information agents, first promised at I/O for AI Pro and Ultra subscribers this summer, now have credible infrastructure behind them. This is AI-assisted discovery, not static search results.

From Experimental Labs to Default Workflows
Taken together, Quail 2’s redesigned Agent Mode and the new Gemini AI models show Google’s intent: AI agents should live where work happens, not in separate playgrounds. In Android Studio, Studio Labs is now stable, letting developers try experimental AI features without upgrading the IDE, which effectively turns the editor into a rolling test bed for agentic workflows. In Search, 3.5 Flash-Lite lays a foundation for information agents that promise to handle multi-step tasks and multi-document reasoning at consumer scale. The strategy is clear and opinionated: normalize AI-assisted development and discovery by making it fast enough, cheap enough, and integrated enough that saying “no” feels like opting out of a productivity baseline. There are risks in that confidence, but the direction is unmistakable. AI-assisted development is moving from a bonus feature to the default expectation in Google’s ecosystem, and developers who ignore agentic workflows now may find themselves behind not on hype, but on everyday efficiency.






