Study Guide2,996 words

Unit 1.7 study guide — Choose among Google model families

Generative AI Leader › Unit 1 › Topic 7

Choose among Google model families

Study guide for Generative AI Leader, Unit 1 · Topic 7. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.

What the exam guide asks, quoted. Gemini Gemma Imagen Veo

This hive's learning objectives for the topic:

  1. Recognize Gemini use cases and verify model-specific input and output capabilities.
  2. Recognize Gemma open-weight use cases and the resulting deployment and license responsibilities.
  3. Recognize Imagen image-generation and supported editing use cases.
  4. Recognize Veo video-generation use cases and verify version-specific capabilities.

Choosing among Google's model families

Gemini, Gemma, Imagen and Veo — and why you read the model page

Gemini — multimodal models for most tasks. Gemma — lightweight open models you can run and tune yourself. Imagen — specialized image generation, now with Gemini as Google's starting point. Veo — video from text and images. Capabilities differ by version: check the model's own page.

The guide names four of Google's model families and asks for their use cases and strengths. Gemini is the general family, built from the ground up for multimodality, so it can reason across text, images, video, audio and code. Gemma is a set of lightweight, generative artificial intelligence open models that you can run on your own hardware or hosted services and tune yourself. Imagen is Google's specialized image generation model — though, as we will see, Google now recommends Gemini as the starting point for generating images. And Veo generates video from text and images. One habit runs through the whole topic: capabilities differ by model and by version, so a leader confirms them on the model's own page instead of trusting a summary — including this deck, which names the version behind every number it gives.

Gemini: one family, different strengths

Pick the Gemini model by the job, then verify what that model accepts

  • Built for multimodality: text, images, video, audio and code
  • 3.8 Flash: long-horizon coding and autonomous agents
  • 3.5 Flash: near-Pro intelligence at Flash-tier cost and speed
  • 3.1 Flash-Lite: most cost-efficient, for high-volume, low-latency traffic
  • Gemini image models: conversational editing and multi-image fusion

Worked example (synthetic). A support team routes a million short questions a day and needs low cost per request. Flash-Lite's documented strength fits; a coding agent that works for hours fits 3.8 Flash's.

Gemini is a family, not one model. Google's documentation says Gemini models are built from the ground up for multimodality and can reason seamlessly across text, images, video, audio and code. Within the family, each model page states a strength. Gemini 3.8 Flash is described as Google's most intelligent workhorse model yet, built for long-horizon coding and autonomous agents. Gemini 3.5 Flash delivers near-Pro intelligence at Flash-tier cost and speed. Gemini 3.1 Flash-Lite is the most cost-efficient model, optimized for low-latency use cases with high-volume, cost-sensitive traffic. And the Gemini image models add conversational editing, multi-image fusion and character consistency for creative work. These descriptions are current as of this deck's retrieval date and will change — which is why the second half of this objective is verification: before relying on a capability, check the features that each model supports, on that model's own page.

What a model page tells you

Modalities, limits and capabilities are listed per model and version

Figure. A model page card for Gemini 3.5 Flash, retrieved 6 October 2026, with three panels: token limits showing a context window of 1,048,576 tokens and a maximum output of 65,536 tokens; modalities, meaning which inputs it accepts and which outputs it returns; and capabilities such as tuning, grounding and function calling, each marked supported or not.

Worked example (synthetic). A team plans to send 2,000-page document sets in one request. The model page's context window answers whether that fits before any test is run.

Here is what verification looks like in practice. Each Gemini model has its own page listing its modalities, its token limits and its capabilities. Gemini 3.5 Flash's page, for example, lists a context window of one million, forty-eight thousand, five hundred and seventy-six tokens and a maximum output of sixty-five thousand, five hundred and thirty-six tokens. Those numbers belong to that model and that version; another Gemini model can list different ones, and Google revises them over time. So the leader's habit is simple: when a use case depends on an input type, a request size or a capability such as tuning, confirm it on the model's own page — Google's guidance is to check the features that are supported by each model.

Gemma: open models you run and tune

Open weights give control — and hand you the work of running them

  • Lightweight generative AI open models, based on Gemini models
  • Run in your apps, on your hardware, mobile devices or hosted services
  • Open weight: tune with the framework of your choice
  • Smaller sizes need fewer resources and deploy more flexibly
  • Pretrained versions need tuning before real use

Worked example (synthetic). A device maker needs a model that runs on a phone with no network connection. A small Gemma variant fits that requirement in a way a hosted model cannot.

Gemma is Google's open model family. Google describes Gemma as a set of lightweight, generative artificial intelligence open models, available to run in your applications and on your hardware, mobile devices or hosted services. Because the models are open weight, you can tune any of them using the framework of your choice. Size is a deployment decision too: lower parameter sizes mean lower resource requirements and more deployment flexibility, and Gemma 4, for example, is listed as supporting text and image input on all variants and audio on its smallest variants. Open weights come with responsibilities. Pretrained versions are not ready for users as they are — Google does not recommend using them without performing some tuning. And whoever runs the model takes on its operation. Google's guidance is to use Agent Platform if you want end-to-end operations tooling and a serverless experience, and to self-manage on Google Kubernetes Engine when you have in-house expertise and need granular control.

Who runs an open model, and what that costs

Managed serving trades control for less operational work

Serving choiceWho operates itWhat you gain
Managed API (MaaS)Google provisions, scales and maintainsFocus on the application
Self-deployedYour team, in your project and networkControl of the data path; lower cost at scale
On GKEYour team, with in-house operations skillsGranular control of the workload

Worked example (synthetic). A bank whose regulator forbids multi-tenant services self-deploys Gemma in its own network and staffs the operations work that comes with it.

The deployment responsibility the guide mentions becomes concrete in Google's serving options for open models. With a managed API — model as a service — Google handles all provisioning, scaling and maintenance, which suits teams focused on the application rather than operations. Self-deploying lets you run the model within your own Google Cloud project and network, with complete control over the data path; it requires a greater upfront engineering investment but can lower total cost of ownership at scale. Running on Google Kubernetes Engine suits organizations with existing Kubernetes investments and in-house operations expertise that need granular control. The pattern to remember: the more control you take over an open model, the more of its running becomes your team's job.

Imagen: specialized image generation

Know its use cases — and that Google now points to Gemini first

  • Generate novel images from a text prompt
  • Edit or expand an image using a mask you define; upscale images
  • Strong for photorealism, branding, logos and typography
  • Google recommends Gemini as the starting point for images
  • Imagen endpoints are listed as deprecated, migrating to Gemini

Worked example (synthetic). A marketing team needs a product logo with exact lettering. That is a documented Imagen strength — but a new project starts on Gemini's image model, because that is where Google directs new image work.

Imagen is the image family the guide names, and Google's documentation still describes its use cases: generate novel images using only a text prompt, edit or expand an uploaded or generated image using a mask area you define, and upscale existing images. Google lists its strengths as image quality, photorealism and specific styles; infusing branding or generating logos and product designs; and advanced spelling or typography — with low latency, optimized for near-real-time performance. But the same documentation now steers new work elsewhere. It recommends Gemini as a starting point for generating images, and its table of deprecated image generation endpoints lists the Imagen endpoints — imagen-4.0-generate-001 among them — with gemini-2.5-flash-image as the replacement, recommending updates before June 30, 2026. These pages also carry the notice that Vertex AI's services are now part of Gemini Enterprise Agent Platform. So know Imagen's use cases for the exam, and know that Gemini's image models are where Google points new image projects.

Gemini image models against Imagen 4

Google's own comparison: flexibility against specialized quality

AspectGemini imageImagen 4
Google's stanceStarting point for imagesSpecialized; Ultra for best quality
EditingMulti-turn conversational editingMask-based editing, upscaling
StrengthsMulti-image fusion, consistencyPhotorealism, logos, typography
LatencyHigherNear-real-time

Worked example (synthetic). A design team iterates on a poster through several rounds of spoken changes — Gemini's conversational editing fits. A catalogue needs a thousand photorealistic product shots quickly — the Imagen strengths fit.

Google's documentation compares the two image options directly. Gemini's image models are Google's recommended starting point for generating images; they are uniquely capable of multi-turn conversational editing, and support multi-image fusion and character consistency, at higher latency. Imagen is Google's specialized image generation model: it edits with masks and upscales, it is strongest where image quality, photorealism, branding, logos or typography are the priority, and it is optimized for near-real-time performance — Google advises choosing Imagen 4 Ultra when you need the best image quality. Hold the earlier caveat alongside the table: Google calls Imagen 4 its latest line of image generation models, yet lists its endpoints as deprecated with a Gemini image model as the replacement. In a scenario, prefer the answer that starts from Gemini and reaches for Imagen only for its specialized strengths.

Two facts about Imagen to hold together

Its use cases are documented; its endpoints are being retired

Figure. Two cards. The first, what Imagen does, lists text-to-image generation, mask-based editing and expansion, upscaling, and strengths in photorealism, logos and typography. The second, where Google points now, lists Gemini as the starting point for images, the imagen-4.0-generate-001 endpoint migrating to gemini-2.5-flash-image, and updating endpoints before June 30, 2026.

Worked example (synthetic). An exam item asks what Imagen is used for: answer with its use cases. A project plan asks which image model to build on: start with Gemini, as Google recommends.

Two facts about Imagen are both true, and a good answer keeps them apart. On the left, what Imagen does: it generates novel images from text prompts, edits or expands images using a mask you define, upscales images, and is strongest at photorealism, branding and logos, and typography. On the right, where Google points today: Google recommends Gemini as a starting point for generating images, and its table of deprecated endpoints maps imagen-4.0-generate-001 to gemini-2.5-flash-image, with a recommendation to update endpoints before June 30, 2026. So a question about Imagen's use cases is answered from the left card, and a question about what to build a new image workflow on is answered from the right.

Veo: video from text and images

Know the use case, then check what this Veo version supports

  • Veo generates videos from text prompts and images
  • Veo 3.1 is Google's latest line of video generation models
  • Veo 3.1 Fast adds low latency to high quality
  • Version facts: 4, 6 or 8 second clips; 9:16 or 16:9
  • Gemini Omni Flash can also generate and edit video

Worked example (synthetic). A retailer wants short vertical product clips for phones. Veo 3.1's documented 9:16 aspect ratio and 8-second limit decide whether one clip or several stitched clips are needed.

Veo is Google's video family. Google's model list describes Veo 3.1 Generate as generating videos from text prompts and images with high quality, and Veo 3.1 Fast as doing the same with high quality and low latency. Veo 3.1 is described as Google's latest line of video generation models. On Agent Platform you can generate video with either Gemini Omni Flash or Veo, in Media Studio or through their APIs, and Gemini Omni Flash is listed as able to generate video from text or reference assets, or edit existing videos. Version-specific facts matter here more than anywhere. Veo 3.1's page lists video lengths of 4, 6 or 8 seconds — with reference-image-to-video supporting only 8 — aspect ratios of 9:16 and 16:9, output resolutions up to 4K, and English as the prompt language. A use case that needs a two-minute landscape film is not answered by any single Veo 3.1 clip, which is exactly the kind of check a leader makes on the model page first.

Veo 3.1, as its page documents it

Requirements to check before promising a video use case

RequirementVeo 3.1 documentsCheck before
Clip length4, 6 or 8 secondsPromising longer scenes
Aspect ratio9:16 or 16:9Square or other formats
Output resolution720p, 1080p, 4K (Generate)Print or cinema output
Prompt languageEnglishNon-English prompts

Worked example (synthetic). A travel brand plans 30-second landscape adverts. Veo 3.1's 8-second maximum means planning several clips per advert from the start.

This table turns the version facts into checks. Clip length: Veo 3.1 lists 4, 6 or 8 seconds, so longer scenes need planning. Aspect ratio: 9:16 for vertical and 16:9 for landscape, so any other format needs another plan. Output resolution: the Generate model lists 720p, 1080p and 4K. Prompt language: English. All of these come from Veo 3.1's own page on the retrieval date, and Google states that Veo and Gemini Omni Flash are designed with its AI Principles in mind — which still leaves testing and responsible deployment to you. The figures will change with the next version; the habit of checking them should not.

Which family fits the job?

Start from what must be produced, then where it must run

Loading Diagram...
Figure 1 — Mermaid diagram

Figure: A flowchart. A business job first asks: must it run on your own hardware or device? Yes leads to Gemma. No leads to a second question: what must it produce? Video leads to Veo; images lead to a Gemini image model, with Imagen for its specialized strengths; text, code or analysis leads to Gemini.

Worked example (synthetic). Three requests in one week: an offline field-inspection app (Gemma), a social video campaign (Veo), and a contract-analysis tool (Gemini).

Put the four families together as a first-pass choice. Start with where the model must run. If it must run on your own hardware or a device, the open Gemma models are the family designed for that — Google says Gemma runs on your hardware, mobile devices or hosted services. Otherwise, ask what must be produced. Video points to Veo. Images point to a Gemini image model first, because Google recommends Gemini as the starting point, with Imagen for its specialized strengths such as typography. Text, code and analysis point to Gemini, built for multimodality across all of them. This is a first pass, not a final answer: the earlier topic's method still applies — screen, evaluate on your own workload, then compare.

What this topic actually tests

Match the family to the job, then check the version

Gemini for most text, code and multimodal work. Gemma when you must run or tune the model yourself — and run it. Imagen for its specialized image strengths, with Gemini as Google's starting point. Veo for video, within each version's documented limits.

Close with the four families and one habit. Gemini is the general, multimodal family; pick the model within it by the job. Gemma is for when you must run or tune the model yourself — open weights give control, and hand you the work of running the model. Imagen has documented strengths in photorealism, logos and typography, but Google now recommends Gemini as the starting point for images and lists Imagen endpoints as deprecated. Veo produces video from text and images, within limits that each version documents. And the habit: every capability belongs to a model and a version, so confirm it on that model's page before a business plan depends on it. Model Garden is where Google's proprietary and select open models are discovered, tested and deployed, and it is the natural place to start that check.

Official sources for this topic

Ready to study Generative AI Leader (GCP-GAIL)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free