Unit 3.4 study guide — Ground outputs and control generation
Generative AI Leader › Unit 3 › Topic 4
Ground outputs and control generation
Study guide for Generative AI Leader, Unit 3 · Topic 4. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks, quoted. Describing the concept of grounding in LLMs and differentiating between grounding with first-party enterprise data, third-party data, and world data. Describing how retrieval-augmented generation (RAG) can affect the generated output from your gen AI models. Google Cloud grounding offerings: Identifying how sampling parameters and settings are used to control the behavior of gen AI models (e.g., token count, temperature, top-p [nucleus sampling], safety settings, and output length).
This hive's learning objectives for the topic:
- Distinguish grounding with first-party enterprise data, third-party data and world data.
- Explain how retrieval-augmented generation changes a model's output.
- Choose among prebuilt RAG with Agent Search, RAG APIs and Grounding with Google Search.
- Predict the effect of temperature, top-p and token or output-length limits on generated output.
- Explain what safety settings control and when to adjust them.
Grounding outputs and controlling generation
Connect the model to facts, then set how freely and safely it writes
Ground — connect output to verifiable sources: your data, a partner's data, or the world's. Retrieve — RAG adds those facts to the prompt. Choose — Agent Search, RAG APIs or Google Search. Control — sampling, length and safety settings.
The previous topic was about what you say to a model. This one is about two things around it: what the model is connected to, and how it is allowed to generate. Google defines grounding as the ability to connect model output to verifiable sources of information, and says grounding tethers output to those sources and reduces the chances of inventing content. The deck takes five steps. First, the kinds of data you can ground on — your own enterprise data, data a third party supplies, and world knowledge from the web. Second, how retrieval-augmented generation, or RAG, changes what the model writes. Third, how to choose among Google Cloud's grounding offerings: prebuilt RAG with Agent Search, the RAG APIs, and Grounding with Google Search. Fourth, the sampling and length parameters that set how freely the model writes. Fifth, the safety settings that decide what it is allowed to return.
Grounding, and the three kinds of data
Ground on your data, a partner's data, or the world's — by need
- Grounding connects model output to verifiable sources
- First-party: your own documents and data, often not public
- Third-party: data another provider supplies, such as partner web data
- World: up-to-date public knowledge through Google Search
- Grounded answers carry links to their sources
Worked example (synthetic). An insurer's assistant answers policy questions from its own policy documents (first-party), checks weather events through a partner's web data (third-party), and finds today's public health guidance through Google Search (world).
Grounding, in Google's words, is the ability to connect model output to verifiable sources of information. It reduces model hallucinations — content that isn't factual — anchors responses to your data sources, and provides auditability through grounding support, which are links to sources. The first objective asks you to tell apart three kinds of data you can ground on. First-party enterprise data is your organization's own: documents, websites and records. Google's grounding options for it include Agent Search, which uses retrieval-augmented generation, or RAG, to connect the model to your website data or document sets, and RAG Engine, a configurable managed RAG service; connecting your own search API is, as Google puts it, particularly useful for enterprise-specific data that isn't publicly available. Third-party data comes from another provider. Google's grounding pages describe partner offerings such as Exa, which connects Gemini models to public web data provided by Exa's search API. World data is the public web: Grounding with Google Search connects the model to world knowledge, a wide range of topics, or up-to-date information on the internet. The business question is simply where the facts the answer needs actually live.
Where the facts live decides how you ground
Three kinds of data, three families of grounding source
| Kind of data | What it is | Google Cloud grounding option |
|---|---|---|
| First-party enterprise | Your documents, sites and records | Agent Search, RAG Engine, your search API |
| Third-party | Data a partner provider supplies | Partner offerings such as Exa web data |
| World | Public, up-to-date web knowledge | Grounding with Google Search |
| World, regulated | A compliance-ready subset of the web | Web Grounding for Enterprise |
Worked example (synthetic). A hospital group wants web facts but cannot have customer data logged. Web Grounding for Enterprise, which doesn't log customer data, fits better than Grounding with Google Search.
Read the table by asking where the facts the answer needs actually live. Every row is a source that retrieval-augmented generation, or RAG, can draw on. If they are in your own documents, websites or databases, you ground on first-party enterprise data — with Agent Search, with RAG Engine, or with your own search API for enterprise-specific data that isn't publicly available. If a partner provides them, you ground on third-party data — Google's documentation describes, for example, an Exa offering that connects Gemini models to public web data provided by Exa's search API. If they are general, public and current, you ground on world data with Grounding with Google Search. The last row is a refinement of world data for regulated industries: Web Grounding for Enterprise indexes a subset of what's available on Google Search, suited to finance, healthcare and the public sector, and the service doesn't log customer data. One answer can draw on more than one row — that is the subject of objective three.
How retrieval-augmented generation changes the output
Retrieved facts become context, so the answer reflects them
- A model alone is limited to its pre-trained data
- RAG retrieves information relevant to the query first
- The retrieved facts are added to the prompt as context
- Answers become more accurate, current and specific
- Irrelevant retrieval gives grounded but off-topic answers
Worked example (synthetic). Asked about this year's travel policy, a plain model guesses from general knowledge. With RAG, the policy document is retrieved and added to the prompt, and the answer quotes this year's limits.
The second objective is what retrieval-augmented generation, or RAG, actually does to the output. Start from the limitation. Google notes that large language models, or LLMs, are limited to their pre-trained data, and that a common problem is that they don't understand private knowledge — your organization's data. RAG adds a step before generation. When a user asks a question, a retrieval component searches a knowledge base to find information relevant to the query; the retrieved information then becomes the context added to the original query, as a guide for the model to generate factually grounded and relevant responses. Google's grounding guidance names this the recommended best practice: to implement grounding, you must retrieve relevant source data, and RAG is the technique to do it. The effect on output is what the exam asks about: by combining your data and world knowledge with the model's language skills, grounded generation is more accurate, up to date and relevant to your needs, and RAG helps the model reduce hallucinations. But note the honest caveat Google gives too — if your retrieved information is irrelevant, the generation could be grounded but off-topic or incorrect. Retrieval quality limits answer quality.
Grounding at request time
The question fans out to sources; the facts come back as context
Figure. A diagram inside Gemini Enterprise Agent Platform. An employee question goes as a prompt to the Gemini model, which sends a query to a retrieve-facts step. Retrieval fans out, as search, to two grounding sources: an Agent Search data store and Google Search. Both feed a grounded answer with links. A callout says retrieved facts become context and the answer links to its sources.
Worked example (synthetic). An employee asks which expense limits changed this year. The policy store supplies the internal limits, Google Search supplies a public tax-rate change, and the answer links to both.
This diagram puts the steps of retrieval-augmented generation, or RAG, in order. An employee's question reaches the Gemini model as a prompt. Before the model answers, a retrieval step queries the grounding sources. Here there are two, because Google documents that grounding to your data supports up to ten Agent Search data sources and can be combined with Grounding with Google Search — so the internal policy store and the public web are both searched. The results come back as context: in Google's words, the retrieved information becomes the context added to the original user query, as a guide for the model to generate factually grounded and relevant responses. And the answer is not just text; grounding support provides links to the sources, which gives the auditability Google describes. Notice what is not drawn. RAG Engine's retrieval tool is not shown beside Google Search: Google lists it among the non-search tools that cannot share a request with search tools, and says multiple tools are supported only when they are all search tools.
Choosing a Google Cloud grounding offering
Prebuilt when you can, RAG APIs when you must, Search for the world
- Prebuilt RAG with Agent Search: Google Search for your data, managed
- RAG APIs: assemble your own pipeline from components
- RAG Engine: a configurable managed RAG service
- Grounding with Google Search: public, up-to-date world knowledge
- Agent Search can be combined with Google Search grounding
Worked example (synthetic). A retailer wants product-manual answers fast with little engineering: prebuilt RAG with Agent Search. A bank needing its own chunking and a grounding check before every reply builds with the RAG APIs.
The guide names three Google Cloud grounding offerings, two of them built on retrieval-augmented generation, or RAG, and the third objective is choosing among them. Prebuilt RAG with Agent Search is the managed path. Google describes Agent Search as Google Search for your data, a fully managed, out-of-the-box search and RAG builder; if you want to do RAG over your website data or your sets of documents, you use Grounding with Agent Search. The RAG APIs are the build-it-yourself path. Google calls this mix-and-match building: you can implement a RAG solution from separate services and APIs — among them a check grounding API that compares RAG output with the retrieved facts and helps ensure all statements are grounded before the response goes to the user. RAG Engine sits in between, a configurable managed RAG service. Grounding with Google Search is for a different need: connecting the model to world knowledge, a wide range of topics, or up-to-date information on the internet. The decision rule a leader needs is about control versus effort: take the prebuilt option when it meets the need, assemble from the RAG APIs when you need control over each stage, and add Google Search when the facts are public and current — Google documents that Agent Search grounding can be combined with it.
Which grounding offering, for which need
Start from the need, not the product list
| The need | Choose | Because |
|---|---|---|
| Answers over our documents, fast | Prebuilt RAG with Agent Search | Managed, out-of-the-box search and RAG |
| Control over every RAG stage | RAG APIs | Mix-and-match components, incl. a grounding check |
| Managed RAG we can configure | RAG Engine | A configurable managed RAG service |
| Current public facts | Grounding with Google Search | World knowledge and up-to-date web results |
Worked example (synthetic). A university wants course-catalogue answers this term with one engineer: prebuilt RAG with Agent Search. Next year it adds a grounding check before replies, which the RAG APIs provide.
Use the table from the need outward; every row but the last is a way to run retrieval-augmented generation, or RAG, on your own data. When the requirement is answers over your own documents with the least engineering, prebuilt RAG with Agent Search fits — Google calls it a fully managed, out-of-the-box search and RAG builder. When the requirement is control over each stage, such as how documents are parsed or a check that every statement is grounded before a reply goes out, the RAG APIs fit — Google's mix-and-match building lets you implement a RAG solution from separate services, and its check grounding API compares output with the retrieved facts. When you want a managed service you can still configure, RAG Engine is described as exactly that. And when the facts are public and current, Grounding with Google Search connects the model to world knowledge and up-to-date information. An exam distractor will usually offer the most powerful option; the right answer is usually the least effort that meets the stated need.
Sampling and length parameters
Temperature and top-P set how freely the model chooses its words
- Sampling parameters shape how the next token is selected
- Temperature: lower is more predictable, higher more diverse
- Top-P: sample only the most probable tokens up to a threshold
- Max output tokens caps length; stop sequences end it early
- Gemini 3.6 Flash and later ignore custom sampling values
Worked example (synthetic). A contract summariser needs consistent wording, so it uses a low temperature and a token cap; a slogan generator wants variety, so it uses a higher temperature — on a model that honours those settings.
The fourth objective is about the dials that control generation. Google explains that sampling parameters influence how the model selects the next token from its vocabulary, and by adjusting them you can control the randomness and diversity of the generated text. Temperature controls the degree of randomness in token selection: lower temperatures are good for prompts that need a less open-ended or creative response, while higher temperatures can lead to more diverse or creative results; at a temperature of zero, the highest-probability tokens are always selected. Top-P works on the candidate list: tokens are selected from most probable to least probable until their probabilities add up to the top-P value. For both, Google's rule is the same — a lower value for less random responses, a higher value for more random ones. Length has its own controls. Set maxOutputTokens to limit the number of tokens generated — a token is roughly four characters — and define stop sequences to tell the model to stop if a given string appears. One scope note matters on the exam and in practice: the parameters available for each model may differ, and for Gemini 3.6 Flash and later models, custom values for token sampling parameters aren't supported and are ignored if set.
Top-P, worked through
Add probabilities from the top until you reach the threshold
Figure. Three cards for one top-P example. Token A, probability 0.3, running total 0.3, kept. Token B, probability 0.2, running total 0.5, kept because the threshold is reached. Token C, probability 0.1, beyond 0.5, excluded. A band beneath says that with top-P 0.5 the model chooses between A and B using temperature and C is never a candidate; lower top-P keeps fewer candidates, higher keeps more.
Worked example (synthetic). Google's own numbers. A customer-service reply that must stay on script benefits from the narrower candidate list a lower top-P gives.
This is Google's own example of top-P, worked through. Suppose the next token could be A, B or C, with probabilities of 0.3, 0.2 and 0.1, and the top-P value is 0.5. The model adds probabilities from the most probable token down: A brings the total to 0.3, B brings it to 0.5, and the threshold is reached. So, in Google's words, the model will select either A or B as the next token by using temperature, and excludes C as a candidate. That is the whole mechanism: top-P decides which tokens are allowed into the draw, and temperature decides how randomly the draw is made among them. Lower the top-P and fewer tokens qualify, so output is less random; raise it and more qualify, so output is more varied — the same direction Google gives for every sampling parameter.
Four controls, four different effects
Randomness, candidate pool, length and stopping point
| Control | What it changes | Turn it down when… |
|---|---|---|
| Temperature | Randomness of the token choice | Answers must be predictable |
| Top-P | Which tokens may be chosen at all | Output wanders off the likely path |
| Max output tokens | The maximum length of the response | Responses run long or cost too much |
| Stop sequences | A string that ends generation | Output must stop at a marker |
Worked example (synthetic). A support bot keeps writing five paragraphs when two would do. Lowering max output tokens caps it; raising temperature would only make it wordier in more varied ways.
The four controls are often confused on the exam because they all change the output. Separate them by what they act on. Temperature changes the randomness of each token choice. Top-P changes which tokens may be chosen at all. Max output tokens changes the maximum length — Google says to set it to limit the number of tokens generated, and to set a low value to limit the length of the response; when the limit is hit, the response reports that generation stopped because the model reached the maximum number of tokens specified. Stop sequences end generation at a chosen string. One more piece of Google's guidance runs the other way: if the model returns a response that's too generic, too short, or a fallback response, try increasing the temperature. And remember the version scope from the previous slide — newer Gemini models ignore custom sampling values, so of these four controls, only the length and stopping rows still apply on those models.
What safety settings control
Content filters block harmful output at thresholds you choose
- Models can still produce harmful responses when prompted
- Filters score harm categories by probability and severity
- Thresholds are set per category for your use case
- From gemini-3.5-flash, the configurable filter defaults to OFF
- Filters block output; they do not change model behaviour
Worked example (synthetic). A children's tutoring app sets strict thresholds on every harm category; an internal security-research tool relaxes the dangerous-content threshold for its analysts, after review.
The fifth objective is safety settings. Google is direct that its models can still generate harmful responses, especially when they're explicitly prompted, so you can configure content filters to block potentially harmful responses. The filters assess content against harm categories — among them hate speech, harassment, sexually explicit content and dangerous content, which Google defines as content that promotes or enables access to harmful goods, services and activities. For each category, a filter assigns two scores: one for the probability that content is harmful, another for its severity. You then choose a blocking threshold for each category based on what is appropriate for your use case and business — the strictest blocks at low scores and above, the most permissive only at high scores. Two facts change how a leader should act. First, the default: for gemini-3.5-flash and subsequent models, the configurable filter's default is OFF, so a consumer-facing deployment must set thresholds deliberately. Second, the limits: filters act as a barrier that prevents harmful output, but they don't directly influence the model's behaviour. Steering behaviour is the job of system instructions, which is why Google recommends a multi-layered approach to safety.
Blocking thresholds, strict to permissive
Stricter thresholds block more — including some benign content
| Threshold | Blocks when the probability or severity is | Fits |
|---|---|---|
| BLOCK_LOW_AND_ABOVE | Low, medium or high | Audiences that need the most protection |
| BLOCK_MEDIUM_AND_ABOVE | Medium or high | General-purpose assistants |
| BLOCK_ONLY_HIGH | High only | Expert users who need wider latitude |
| OFF / BLOCK_NONE | Never — no automated blocking | Only with your own review layer |
Worked example (synthetic). House judgements in the right-hand column. A retailer's public chatbot chooses medium and above, then reviews blocked replies weekly for false positives.
Google's thresholds run from strict to permissive, and each one blocks when either the probability score or the severity score reaches its level. Block low and above stops content scored low, medium or high; block medium and above stops medium or high; block only high stops content only when a score is high. At the far end, off means no automated blocking, and block none removes automated blocking while still returning the scores so you can apply your own guidelines. Remember why there are two scores: content can have a low probability and a high severity, or the reverse. The right-hand column is a house judgement, not Google's wording, but the trade-off behind it is Google's: configurable filters may occasionally block benign content — false positives — or miss some harmful content — false negatives. Stricter is not automatically better; it is a business decision about the audience.
Safety is layered, not a single switch
Each layer catches what the one before it can miss
Figure: A left-to-right flowchart of safety layers. A user prompt passes Model Armor, which protects against prompt injection and jailbreaks, then system instructions, which guide preferred behaviour, then the Gemini model. The model's output passes content filters with thresholds per harm category before becoming the response.
Worked example (synthetic). A bank's assistant uses Model Armor against prompt injection, system instructions to keep to banking topics, and medium-and-above filters on every harm category.
Google's safety overview is explicit that a multi-layered approach to safety is critical, and the layers do different jobs. Model Armor provides enterprise-grade protection against prompt injection and jailbreaks, content harms, sensitive data protection, and malware detection. System instructions provide direct guidance to the model on preferred behaviour and topics to avoid — they steer. Content filters let you set specific thresholds for common harm types — they block, and because they are an independent layer from the model they are robust against jailbreaks. The order on the slide is a simplification of a real deployment, but the lesson holds: Google warns that its default safety might not meet your organization's needs, so the leader's job is to decide which layers a use case needs and at what settings, not to trust one switch.
What this topic actually tests
Where are the facts, and how free and how safe should the output be?
Which data? first-party, third-party or world. Why RAG? retrieved facts become context — good retrieval in, grounded answer out. Which offering? prebuilt Agent Search, RAG APIs, or Google Search. Which settings? temperature and top-P for freedom, tokens for length, thresholds for safety.
Close with the four questions this topic asks. Which data does the answer need? Your own enterprise data, data a partner provides, or world knowledge from the web — the answer picks the grounding source. Why does retrieval-augmented generation, or RAG, help? Because retrieved facts are added to the prompt as context, the answer becomes more accurate and current — as long as retrieval finds the right facts, since irrelevant retrieval gives grounded but off-topic answers. Which offering? Prebuilt RAG with Agent Search when a managed path meets the need, the RAG APIs when you need control over each stage, Grounding with Google Search for public and current facts. Which settings? Temperature and top-P set how freely the model writes, on models that honour them; max output tokens and stop sequences set length; content filter thresholds, layered with system instructions and Model Armor, decide what may be returned. The next unit turns from techniques to the business strategy around them.
Official sources for this topic
- Grounding overview — Gemini Enterprise Agent Platform
- Ground responses using RAG — Gemini Enterprise Agent Platform
- Content generation parameters — Gemini Enterprise Agent Platform
- Safety and content filters — Gemini Enterprise Agent Platform
- Grounding with your search API — Gemini Enterprise Agent Platform
- Grounding with Exa — Gemini Enterprise Agent Platform
- Grounding with Google Search — Gemini Enterprise Agent Platform
- What is Retrieval-Augmented Generation (RAG)? — Google Cloud
- RAG Engine on Gemini Enterprise Agent Platform overview
- Grounding with Agent Search — Gemini Enterprise Agent Platform
- Safety in Gemini Enterprise Agent Platform
- Web Grounding for Enterprise — Gemini Enterprise Agent Platform