Unit 1.1 study guide — Choose a foundation model for a business use case
Generative AI Leader › Unit 1 › Topic 1
Choose a foundation model for a business use case
Study guide for Generative AI Leader, Unit 1 · Topic 1. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks, quoted. Identifying how to choose the appropriate foundation model for a business use case (e.g., modality, context window, security, availability and reliability, cost, performance, fine-tuning, and customization).
This hive's learning objectives for the topic:
- Screen model candidates against required inputs, context and deployment constraints.
- Choose workload-specific evaluation evidence before selecting a model.
- Compare eligible candidates using measured quality, latency and workload cost.
Choosing a foundation model
Screen what must fit, measure what must work, compare what is left
Screen — modality, context, features, data controls and availability rule a candidate in or out. Evaluate — evidence from your own workload, not a leaderboard. Compare — the most affordable model that still meets your quality and latency bar.
This topic is one decision a generative AI leader is asked to make again and again: which foundation model should this business use case run on? Google's guidance starts from three properties — when you choose a model, consider the model's modality, size and cost — and ends with a rule worth memorising: choose the most affordable model that still meets your response quality and latency requirements. Between those two sentences sits a method, and this deck teaches it in three steps. First, screen. Some requirements are pass or fail: a model that cannot take the inputs your users send, or that your data policy cannot accept, is out however impressive it is. Second, evaluate. For the candidates that survive, gather evidence from your own workload, because a public leaderboard does not measure your task. Third, compare. Among the models that meet the quality and latency bar, the cheapest one for your actual workload wins. Keep those three verbs in order and most exam scenarios about model choice become easy to read.
Screen candidates before you score them
A model that cannot take your inputs is out, however good it is
- Modality: the data categories a model is trained for, like text, images and video
- Context window: how much a model can take in at one time
- Features: not every model supports tuning or distillation
- Data controls: what the service keeps, and what you must configure
- Availability: endpoints and regions that keep the model reachable
Worked example (synthetic). An insurer wants claim photos in and a written damage summary out. A text-only model fails the first gate before anyone looks at its benchmark scores; it is not a weaker candidate, it is not a candidate.
The first objective is screening, and the point of screening is that some requirements are not trade-offs at all. Start with modality. Google defines the modality of a model as the high-level data categories it is trained for, like text, images and video, and adds that your use case and the model's modality are closely associated. If photos must go in, a model that cannot accept images is excluded. Needing several modalities is possible too, but Google warns that cost and latency might be higher. Next, the context window, which Google compares to short-term memory: how much the model can take in at one time. Historically, large language models, or LLMs, were significantly limited by the amount of text that could be passed in at once, so a task that needs a long contract in a single request screens on this. Then features. Google says not all models support features like tuning and distillation, and if those capabilities matter to you, you should check what each model supports. Then data controls. Google states that it won't use your data to train or fine-tune any AI or machine learning model without your prior permission or instruction — but retention is a configuration question, not a slogan, which the next slide shows. Finally availability: for global availability and resilience, Google recommends deploying models to endpoints in multiple regions or using the global endpoint. A candidate passes only when every gate the business actually requires is green.
Screening is a sequence of gates
Each gate removes candidates; only survivors are worth evaluating
Figure: A left-to-right flowchart. A business requirement passes through four yes-or-no gates in turn: does the model take and return the modalities needed, does it fit the context the task needs, does it support required features such as tuning, and do its data controls and availability meet policy. A no at any gate leads to Out. Four yeses lead to Eligible: evaluate on your workload.
Worked example (synthetic). A legal team needs a whole contract in one request and tuning on its clause library. Two of five candidates fail the context gate and one fails the features gate, so only two go on to evaluation.
Drawn as a flow, screening is a series of gates, and the order matters less than the rule that any single no removes a candidate. The first gate is modality — your use case and the model's modality are closely associated, so the model must take in and return the kinds of data the use case involves. The second is context: the context window works like short-term memory, and a task that must hold a long document at once needs a model whose window fits it. The third is features: not all models support features like tuning and distillation, so if the plan depends on tuning, check what each model supports before you invest in evaluating it. The fourth gathers the policy checks — data controls and availability. Only the candidates that pass every gate are eligible, and that is where the next objective begins. Notice what is absent from the gates: no quality score and no price. Those are comparisons, and comparing a model that fails a hard requirement wastes the evaluation budget.
Data controls are configured, not assumed
Training use is restricted; retention depends on what you turn on
| Question to ask | What Google documents |
|---|---|
| Will our data train Google's models? | Not without your prior permission or instruction |
| Can we reach zero data retention? | Only by taking specific actions; may not be possible with some Advanced AI features |
| Does grounding on Google Search keep logs? | Yes — no way to disable that storage |
| Will the model stay reachable? | Use endpoints in multiple regions or the global endpoint |
Worked example (synthetic). A bank requires zero data retention for a customer chatbot. The team plans Grounding with Google Search, then learns that service stores logs with no off switch, so it plans a different grounding approach before going further.
Security and availability appear in the guide's list as considerations, and the trap is to treat them as properties of a model rather than of how a service is configured. Google states the training restriction plainly: it won't use your data to train or fine-tune any AI or machine learning model without your prior permission or instruction. Retention is different. Google's zero data retention page lists specific actions customers must take to reach zero data retention, and warns that it may not be possible when using some Advanced AI features. Some features retain data by design — for Grounding with Google Search, Google says there is no way to disable the storage of that information. So a requirement like zero data retention screens not only the model but the features you plan to use with it. Availability is a deployment choice as well: Google's reliability guidance says that for global availability and resilience you deploy models to endpoints across multiple regions or use the global endpoint. In an exam scenario, the candidate that silently assumes a data or availability guarantee is usually the wrong answer.
Gather evidence from your own workload
A public benchmark is not evidence about your task
- Evaluate on your specific tasks and criteria, not leaderboards
- Every method needs an evaluation dataset of prompts and ideal responses
- Include diverse examples that match the task you are evaluating
- Use a wide range of metrics in combination with human evaluation
- Once live, sample production logs to see real-world use
Worked example (synthetic). A support team shortlists two models that both lead public charts. On 200 of its own past tickets, one answers refund questions far more accurately — the only result that predicts how the bot will do for these customers.
The second objective is about evidence. Once a candidate has passed the screen, the question becomes how well it does the actual job, and Google is direct about where that answer comes from: its evaluation service lets you see how a model performs on your specific tasks and against your unique criteria, providing insights which cannot be derived from public leaderboards and general benchmarks. Model evaluation, in Google's words, helps you assess how your prompts and customizations affect a model's performance. Every evaluation method needs the same raw material — an evaluation dataset, made of prompt and ground-truth pairs that you create, where ground truth means the ideal response. Build it with a diverse set of examples that align with the task you are evaluating, or the results will not mean much. Then measure broadly. Automated metrics are quick and scalable, but each method has weaknesses, and Google's remedy is to use a wide range of metrics in combination with human evaluations. Finally, evidence does not stop at launch: Google suggests sampling directly from production logs to evaluate real-world usage. The exam favours the answer that tests on representative business data over the one that trusts a general benchmark.
Four ways to score the responses
Pick the method from what you know about a good answer
| Method | What it does | Reach for it when |
|---|---|---|
| Adaptive rubrics | Pass or fail tests generated for each prompt | Google's recommended default |
| Static rubrics | One fixed set of criteria for every prompt | The same criteria apply everywhere |
| Computation-based metrics | Deterministic scores such as ROUGE or BLEU | You have a ground truth answer |
| Custom functions | Your own scoring logic in Python | Requirements nothing else covers |
Worked example (synthetic). A team scoring product-description rewrites has no single correct text, so a computation-based metric against one reference would mislead; rubrics that test each rewrite for the required facts fit better.
Google's evaluation service supports four common methods, and the choice follows from what you know about a good answer. Adaptive rubrics, which Google marks as recommended, generate a unique set of pass or fail rubrics for each individual prompt in your dataset — rubrics work like unit tests in software. Static rubrics apply a fixed set of scoring criteria across all prompts, which suits a workload where the same standard applies to every request. Computation-based metrics use deterministic algorithms when a ground truth is available, so they fit tasks with one right answer. And custom functions let you define your own evaluation logic in Python for specialized requirements. None of these is a leaderboard: each one scores responses to your prompts. The same tooling helps later too — Google lists model migrations as a use, comparing model versions to understand behavioural differences before you switch.
The evaluation workflow, end to end
Same dataset, same metrics, every candidate
Figure: A left-to-right flowchart of five steps: an evaluation dataset of prompts and ideal responses, define metrics, generate responses from each candidate model, run the evaluation, and interpret the scores and individual responses. An arrow loops from interpret back to generate, labelled adjust prompts, then re-run.
Worked example (synthetic). Two eligible models answer the same 150 billing questions; both are scored by the same rubric, and reviewers read the twenty lowest-scoring answers from each before anyone looks at an average.
The workflow Google describes makes the comparison fair by holding everything constant except the model. You start from the evaluation dataset. You define evaluation metrics — choosing the metrics you want to use to measure model performance. You generate model responses, selecting one or more models to answer the same dataset. You run the evaluation job, which assesses each model's responses against the selected metrics. And you interpret the results, reviewing the aggregated scores and also the individual responses. That last instruction matters: an average can hide the handful of failures a business cannot accept, so read the worst answers as well as the summary. The loop back to generation is where prompt changes are tested, which is also why Google recommends starting with prompting to find the optimal prompt before reaching for heavier customization.
Compare only the candidates that passed
The most affordable model that still clears quality and latency wins
- Larger models can answer better, at higher latency and cost
- Metering differs: input and output tokens for some, node hours for others
- Cost comes from your workload's volumes, not a rate card alone
- Tuning can lower latency and cost by shortening prompts
Worked example (synthetic). Two models both clear a 90% acceptance bar and a two-second limit. The smaller one costs a third as much for the forecast volume, so it is chosen — even though the larger one scores higher.
The third objective is the comparison itself, and Google's rule sets its shape: choose the most affordable model that still meets your response quality and latency requirements. Notice that quality and latency are thresholds in that sentence, not scores to maximise. Size is the usual tension. In general, Google says, a larger model can learn more complex patterns, which can result in higher quality responses — but larger models in the same family can have higher latency and costs, so you may need to experiment and evaluate to find which size works best. Cost needs care because models can be metered and charged differently: some are charged by the number of input and output tokens, others by the node hours used while the model is deployed. So the honest comparison multiplies each candidate's prices by your own workload — how many requests, how long the prompts, how long the answers. Customization changes the arithmetic too: Google lists lower inference latency and cost, due to shorter prompts, among the benefits of tuning. And at deployment, if lower cost is a priority, you might be able to tolerate higher latency with lower-cost machines.
Thresholds first, then price
A higher score does not win if a cheaper model also clears the bar
| Candidate | Accepted / 100 | p95 latency | Monthly cost | Verdict |
|---|---|---|---|---|
| Alder | 93 | 1.6 s | $1,200 | Chosen — clears both bars, cheapest |
| Birch | 97 | 1.9 s | $3,400 | Clears both bars, costs more |
| Cedar | 95 | 2.6 s | $900 | Out — misses the 2 s latency limit |
| Dogwood | 84 | 0.9 s | $500 | Out — misses the 90 accepted bar |
Worked example (synthetic). Invented figures for a retail assistant that must reach at least 90 accepted answers per 100 and a 95th-percentile latency of 2 seconds or less.
All the numbers on this slide are invented, so the reasoning is the point. The business set two bars before any model was tested: at least ninety accepted answers out of a hundred reviewed, and a ninety-fifth percentile response time of two seconds or less. Dogwood is the cheapest and fastest, but it misses the quality bar, so it is out. Cedar has a strong score and a low price, but it misses the latency limit, so it is out too. Two candidates remain. Birch has the highest score of all, and Alder costs about a third as much. Google's rule decides it — choose the most affordable model that still meets your response quality and latency requirements — so Alder is chosen. The extra four points Birch scores are real, but the business did not ask for them, and paying for them is a choice to justify separately. If the bars were wrong, change the bars; do not let the leaderboard instinct change them silently.
Price the workload, not the rate card
A cheaper input rate can still cost more for an output-heavy job
Figure. Two cards with invented prices. Model E charges 1 dollar per million input tokens and 8 per million output tokens; Model F charges 3 and 4. For 20 million input and 10 million output tokens each totals 100 dollars. A band beneath shows that doubling output to 20 million tokens makes E cost 180 dollars and F 140.
Worked example (synthetic). Invented rates only. A summariser that writes long reports is output-heavy; a classifier that reads long documents and returns one label is input-heavy. The same two rate cards rank differently for each.
Here is why cost has to be computed from the workload. Google notes that some models are charged based on the number of input and output tokens, and the input and output rates can differ. The two invented rate cards on this slide come to exactly one hundred dollars each for a job with twenty million input tokens and ten million output tokens. Now double the output — the same reports, but twice as long. Model E, with its cheap input and expensive output, rises to one hundred and eighty dollars; Model F only to one hundred and forty. Neither card is cheaper in general; each is cheaper for some shape of workload. And a model billed by node hours while it is deployed, which Google also describes, cannot be compared on tokens at all — you estimate how long it must run to serve your traffic. When you plan deployment, Google advises considering your anticipated traffic, latency requirements and budget together, which is exactly this comparison.
What this topic actually tests
Three questions, in this order
Can it do the job at all? modality, context, features, data controls, availability. Does it do our job well? evidence from our own prompts, scored broadly and read by people. Is it the cheapest that clears the bar? priced on our workload, with quality and latency as thresholds.
Close on three questions, asked in order. First, can it do the job at all? That is screening: the modality must match the use case, the context window must hold the task, required features such as tuning must be supported, and data controls and availability must meet policy — any no removes the candidate. Second, does it do our job well? That is evidence: an evaluation dataset of your own prompts and ideal responses, a wide range of metrics combined with human evaluation, never a public leaderboard standing in for your task. Third, is it the most affordable model that still clears the quality and latency bar? That is the comparison, priced on your workload's real volumes, because models are metered differently. The next topic in this unit turns to Google's own model families, which is where these three questions get concrete names.
Official sources for this topic
- Develop a generative AI application — Google Cloud
- Gen AI evaluation service overview — Gemini Enterprise Agent Platform
- Long context — Gemini Enterprise Agent Platform
- Gemini Enterprise Agent Platform and zero data retention
- AI and ML perspective: Reliability — Cloud Architecture Center
- Introduction to tuning — Gemini Enterprise Agent Platform