Match five requests to a model family — decision exercise
Exercise: match five requests to a model family
Original fictional scenario. No cloud account, API calls or paid services are required. Difficulty: intermediate · Estimated duration: 15–20 minutes
A consumer-electronics company receives five requests in one week. All names and details below are invented.
| # | Request | Hard requirement |
|---|---|---|
| 1 | Field technicians need an assistant on rugged tablets in basements with no signal | Must run on the device, offline |
| 2 | Marketing wants a new logo with the brand's exact lettering | Legible typography |
| 3 | Social media wants a vertical product video for phones | 30 seconds long |
| 4 | Support wants a system that reads warranty photos and writes a claim summary | Image in, text out |
| 5 | Design wants to refine a poster over several rounds of spoken feedback | Conversational editing |
Your decision
- Choose a model family for each request: Gemini, Gemma, Imagen or Veo.
- For each choice, name the one fact you would confirm on the model's own page before committing.
- Request 3 cannot be met by one clip as stated. Explain why and propose a plan.
- Request 2 points to Imagen's strengths. State what Google's current image documentation says that affects a new project, and decide.
Rubric (10 house points)
- 4 points: a defensible family for each of requests 1, 3, 4 and 5.
- 2 points: a specific page fact to confirm for each.
- 2 points: request 3 answered from Veo 3.1's documented clip lengths.
- 2 points: request 2 decided with both Imagen's strengths and Google's current recommendation.
Reference solution
1 — Gemma. It runs on your hardware and mobile devices; confirm a size small enough for the tablet, since lower parameter sizes need fewer resources. 4 — Gemini. It is built for multimodality across images and text; confirm the chosen model's page lists image input. 5 — a Gemini image model, which is uniquely capable of multi-turn conversational editing. 3 — Veo, which generates video from text prompts and images.
Request 3: Veo 3.1's page lists clip lengths of 4, 6 or 8 seconds and 9:16 for vertical video, so a single 30-second clip is not available as stated. Plan four 8-second vertical clips edited together, or shorten the brief. Rejected alternative: "Veo 3.1 is the latest, so it will do 30 seconds".
Request 2: typography and logos are documented Imagen strengths. But Google recommends Gemini as the starting point for generating images and lists Imagen endpoints such as imagen-4.0-generate-001 as deprecated, with gemini-2.5-flash-image as the replacement. Start the new project on a Gemini image model and test the lettering; treat Imagen as a documented option, not a new dependency.
Sources and scope
All names, requests and rubric points above are house-authored. The model capabilities applied are grounded in:
- https://docs.cloud.google.com/docs/ai-ml/generative-ai/develop-generative-ai-application
- https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/open-models/use-gemma
- https://docs.cloud.google.com/vertex-ai/generative-ai/docs/image/overview
- https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/veo/3-1-generate
- https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/google-models