Google Vids’ New Direction: From Slides to Full-Fledged AI Video Studio
Google Vids is an AI video creation tool inside Google Workspace that now combines personal AI video avatars, text to video generation, image to video generation, and conversational AI video editing in a single workflow for enterprise teams that want to produce professional videos without cameras, studios, or specialist editors. Google has rolled out two major Google Vids features: Gemini Omni, which enables multimodal text- and image-to-video generation and conversational editing, and personal avatars, which let employees create a digital likeness of themselves from a selfie and short voice recording. The real story is not the technology alone; it is that this capability arrives inside a bundle most large organizations already pay for, turning AI video avatars from a niche add-on into a default option for enterprise video creation.

Personal AI Avatars: Your Selfie Becomes the Presenter
The personal avatar feature is Google’s most visible shot across the bow of HeyGen and Synthesia. To create an avatar, a user submits a selfie and records about ten seconds of their voice; Vids then builds a digital version of that person that can deliver any typed script as video. That means the next training module, onboarding clip, or company announcement can star an employee without them ever stepping in front of a camera or memorizing lines. In practical terms, this kills a lot of friction: no studio time, no reshoots, no scheduling around executive calendars. Because avatars are tied to the user’s Google Account and restricted to their own likeness, Google is signaling a conservative stance on identity misuse. But it is also creating a future where being “on camera” is a configurable setting rather than a physical act.
Gemini Omni: Text-to-Video Meets Step-by-Step AI Video Editing
Gemini Omni is the less flashy but more transformative half of the update. It brings multimodal text to video generation and image to video generation into Google Vids, allowing users to describe scenes in plain language, attach reference images, and receive synthesized clips that match those inputs. More importantly, Omni fixes the editing loop that has hamstrung many avatar platforms: editing happens through a chat-based, step-by-step interface where users refine generated clips or phone footage by asking to swap backgrounds, adjust lighting, or add effects instead of starting again. As one review noted, Omni “handles the gap that avatar-generation tools have traditionally left open: the editing loop.” This is AI video editing designed for non-editors—product managers, HR leads, trainers—who can finally iterate like they do on documents and slides, not like film crews.
Bundling Changes the Economics: Direct Pressure on HeyGen and Synthesia
On technology alone, Vids now sits in the same category as independent avatar-video vendors. HeyGen, Synthesia, Captions, and D-ID have built subscription businesses around letting companies produce video at scale without production crews or camera-shy executives. The difference is that Google is putting comparable AI video avatars and text to video generation inside Workspace, which costs large organizations about USD 12–22 (approx. RM55–100) per user per month and is already deployed widely. Standalone avatar tools often start around USD 50 (approx. RM230) per month for individuals and can exceed USD 500 (approx. RM2,300) at enterprise tiers. When your IT department is already paying for Google Workspace, every extra subscription now faces a brutal question: why pay separately for enterprise video creation when Vids is built into the stack? History suggests bundling tends to crush adjacent categories that cannot differentiate fast enough.
Enterprise-Ready—but Not Risk-Free
From an enterprise perspective, Google Vids looks more “ready” than many point solutions. Both the avatar and Gemini Omni features are available to Google AI Pro and Ultra subscribers and Google Workspace business customers, and every AI-generated clip includes an invisible SynthID watermark so companies can verify that footage was AI-made. This is a direct nod to compliance teams worried about synthetic media. Yet the rollout lands amid regulatory pressure on Google’s bundled AI services and a class action over how Gemini was trained. There are also glaring gaps: Google has not disclosed what happens to the selfies and voice clips that power personal avatars, nor spelled out the data retention and access controls over this biometric information. For regulated industries, that silence is more than a footnote—it is a deployment blocker. The feature is limited to over-18 users in unspecified regions, but enforcement will depend on administrators, not individual employees.






