Study Guide3,221 words

Unit 2.5 study guide — Choose building blocks for a custom AI solution

Generative AI Leader › Unit 2 › Topic 5

Choose building blocks for a custom AI solution

Study guide for Generative AI Leader, Unit 2 · Topic 5. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.

What the exam guide asks, quoted. Recognizing the functionality, use cases, and business value of Agent Platform (e.g., Model Garden, Agent Search, Agent Platform AutoML). Recognizing the functionality, use cases, and business value of Google Cloud’s RAG offerings (e.g., prebuilt RAG with Agent Search, RAG APIs). Recognizing the functionality, use cases, and business value of using Agent Platform to build custom agents.

This hive's learning objectives for the topic:

  1. Match Model Garden, Agent Search and AutoML to their distinct developer tasks.
  2. Choose prebuilt search-based RAG or RAG APIs according to retrieval needs.
  3. Recognize when a custom agent is appropriate and identify its tools and operating controls.

Building blocks for a custom AI solution

Pick the model, wire in your data, and decide whether you need an agent

Models — Model Garden to find and deploy them, AutoML to train your own with minimal technical effort. Your data — Agent Search as a ready-made RAG system, or RAG APIs and RAG Engine when you need control. Agents — build a custom agent only when the task needs tools and decisions, and run it with the controls the platform provides.

When a business builds its own generative AI solution on Google Cloud, it assembles it from building blocks on Gemini Enterprise Agent Platform, which Google describes as a unified platform to build, deploy, govern and optimize enterprise-grade AI agents and model-based solutions, and as an evolution of Vertex AI. This topic asks three questions about those blocks. First, which block fits which developer task — finding and deploying a model in Model Garden, adding Google-quality search with Agent Search, or training a model with AutoML? Second, when a solution must answer from your own data, should you use the ready-made retrieval-augmented generation, or RAG, system that Agent Search provides, or the component APIs and RAG Engine that give you more control? Third, when is a custom agent the right answer, and which tools and operating controls come with it? We take them in that order.

Three building blocks, three different jobs

Find a model, search your content, or train a model without code

  • Model Garden: discover, test, customize and deploy models
  • Agent Search: Google-quality search and an out-of-the-box RAG system
  • AutoML: train a model with minimal technical effort, to prototype
  • Custom training: full control, for a team that writes code

Worked example (synthetic). A media company wants three things: an open model for captioning, search across its archive, and a quick model that predicts which articles subscribers will read. That is Model Garden, Agent Search and AutoML — one block each.

The first objective is matching three building blocks to the developer task each one serves. Model Garden is, in Google's words, an AI and machine learning model library that helps you discover, test, customize and deploy models and assets from Google and Google partners, giving access to Google's frontier models such as Gemini, third-party models and open-source models. Reach for it when the job is choosing and deploying a model. Agent Search is for search over your own content: Google describes it as functioning as an out-of-the-box retrieval-augmented generation, or RAG, system for information retrieval. Reach for it when the job is helping people or models find the right information in your data. AutoML is for training your own model with little effort. Google says that with AutoML you create and train a model with minimal technical effort, and that you can use it to quickly prototype models and explore new datasets before investing in development. Its contrast is custom training, where a team writes a training application optimized for its targeted outcome. So when a scenario mentions a team without machine learning programmers that wants a model from its own data, AutoML is the block; when it mentions browsing and deploying existing models, it is Model Garden.

Match the developer task to the block

The verb in the requirement usually names the block

Developer taskBuilding blockWhy
Compare and deploy an open modelModel GardenDiscover, test, customize and deploy models
Let customers search help articlesAgent SearchGoogle-quality search over your content
Predict churn from a table, no codersAutoMLMinimal-effort training to prototype quickly
A bespoke loss function and algorithmCustom trainingFull control of the training application

Worked example (synthetic). Four requests from a fictional retailer's data team, each answered by a different block.

Put the tasks in a table and the verbs give the answer away. Comparing and deploying an existing model is Model Garden's job, because it exists to help you discover, test, customize and deploy models. Letting customers search help articles is Agent Search, the out-of-the-box retrieval system. Predicting churn from a table when nobody on the team writes training code is AutoML, which Google describes as creating and training a model with minimal technical effort, suited to quickly prototyping models. And a requirement for a bespoke algorithm or loss function points past AutoML to custom training, where you create a training application optimized for your targeted outcome. A frequent distractor offers Model Garden for a training-from-scratch task, or AutoML for choosing among existing foundation models; neither matches the verb.

Choosing from Model Garden is governed

An administrator can decide which models are allowed at all

Figure. Two cards. Organization policy: set at organization, folder or project level to allow vetted models and deny the rest. Security scanning: models scanned as unsafe are blocked from deployment in Model Garden. A band beneath says a team picks the model and the organization decides the menu it picks from.

Worked example (synthetic). A fictional bank's platform team allows three vetted models in its production projects, so a developer's request to deploy an unvetted open model is refused by policy, not by a meeting.

Model Garden is not a free-for-all, which matters to a leader weighing risk. Google says you can set a Model Garden organization policy at the organization, folder or project level to control access to specific models — for example, allowing the models you have vetted and denying the rest. There is a safety net beneath that too: for third-party models from the Hugging Face hub, models deemed unsafe by scanning are flagged and blocked from deployment in Model Garden. So the honest picture is that development teams choose a model, but the organization decides the menu they choose from. In an exam scenario about controlling which models developers may deploy, the answer is a Model Garden organization policy, not a written guideline.

Prebuilt RAG or RAG you assemble

Take the managed pipeline unless you need control of its parts

  • RAG is Google's recommended way to ground answers in your data
  • Agent Search: an out-of-the-box RAG system, set up in a few clicks
  • RAG APIs: the same components exposed for granular control
  • RAG Engine: a managed framework that ingests, indexes and retrieves

Worked example (synthetic). An insurer wants a policy chatbot by next quarter with no specialists. Prebuilt RAG with Agent Search fits. A research firm that must control chunking and ranking for its own corpus reaches for the RAG APIs instead.

The second objective is choosing between prebuilt retrieval-augmented generation and building it yourself. Google's starting point is that the recommended best practice for grounding is to use the RAG technique. The prebuilt option is Agent Search, which functions as an out-of-the-box RAG system; Google says it has simplified the end-to-end process — the extraction, chunking, embedding, indexing, storage and retrieval — to just a few clicks. Building your own is possible, but Google is candid that developing a well-functioning RAG system for do-it-yourself grounding can be complex. For teams that need it anyway, the RAG application programming interfaces expose the underlying components of Agent Search's system, so developers can address custom use cases or serve customers who want granular control, combining them in what Google calls mix-and-match building. And RAG Engine, a component of Agent Platform, is a managed framework that lets you enrich a large language model's context with private information so the model can reduce hallucinations and answer more accurately. The decision rule: take the managed pipeline unless the requirement names a part of it you must control.

Retrieval-augmented generation, drawn

Retrieve from your sources first, then let the model answer

Figure. A customer sends a question to a support app inside a Google Cloud project. The app sends a query to a retrieval engine, which retrieves from three of your data sources — a help-center site, policy PDFs and order records. The retrieved results flow on to a Gemini answer. The callout says retrieved passages become context for the model's answer.

Worked example (synthetic). A fictional airline's support app answers "can I change my flight?" by retrieving the fare rules page and the booking record before the model writes a word.

Here is the shape every retrieval-augmented generation solution shares. A customer's question reaches the application. The application sends a query to a retrieval engine — Agent Search in the prebuilt case, RAG Engine or your own components in the assembled case. The retrieval engine searches the sources you connected; in RAG Engine's words, the retrieval component searches through its knowledge base to find information relevant to the query. Then generation happens: the retrieved information becomes the context added to the original user query, as a guide for the generative model to produce factually grounded and relevant responses. The point to carry into the exam is the order. Retrieval comes first and the model answers second, which is why RAG can answer from private data the model was never trained on.

Three ways to get RAG

More control costs more building

OptionYou manageChoose it when
Prebuilt RAG with Agent SearchData sources and settingsSpeed matters and defaults are good enough
RAG APIs (mix-and-match)Which components, and how they combineYou need granular control of the pipeline
RAG EngineA corpus you ingest and queryYou want a managed framework for private context

Worked example (synthetic). Three fictional teams: a call center needing answers in weeks, a legal-tech firm tuning its own ranking, and a developer team building an internal assistant over Drive documents.

Compare the three ways to get retrieval-augmented generation by what you manage. With prebuilt RAG on Agent Search you manage little beyond your data sources, because Google has simplified the pipeline to a few clicks — choose it when speed matters and the defaults are good enough. With the RAG application programming interfaces you choose and combine the components yourself, which Google offers for developers addressing custom use cases or customers who want granular control. Among those components is a check grounding service that compares RAG output with the retrieved facts and helps ensure every statement is grounded before the response reaches the user. And with RAG Engine you work with a managed framework: Google says it creates an index called a corpus over your ingested data, and enriches the model's context with that private information. The exam's usual contrast is between the first two — pick the prebuilt option unless the scenario names a control requirement.

What a RAG pipeline actually does

Five steps happen before the model writes its answer

Loading Diagram...
Figure 1 — Mermaid diagram

Figure: A left-to-right flowchart of six steps: ingest from files, Cloud Storage and Drive; transform by splitting into chunks; embed, turning meaning into numbers; index into a corpus; retrieve passages for the query; and generate a grounded answer.

Worked example (synthetic). A fictional HR team loads 400 policy PDFs once; every later question only runs the last two steps.

The documentation for RAG Engine, Google's managed retrieval-augmented generation framework, lists the work in the order it happens, and seeing it explains why the prebuilt option saves so much effort. Data is ingested from sources such as local files, Cloud Storage and Google Drive. It is transformed for indexing — for example, split into chunks. Each chunk is embedded, turned into numbers that capture its meaning. RAG Engine then creates an index called a corpus. Only then, when a user asks something, does retrieval search the corpus for relevant information, and generation add what was retrieved to the user's query as context for a grounded answer. The first four steps are preparation that happens once and again on updates; the last two happen on every question. Agent Search runs all of this for you, which is exactly what Google means by a few clicks.

When a custom agent is the right block

Build an agent when the task needs tools, steps and decisions

  • Agent Development Kit: open-source framework for building agents
  • Tools connect the agent to third-party apps and your own code
  • Multi-agent teams can collaborate and delegate tasks
  • Agent Runtime: managed runtime that removes infrastructure work

Worked example (synthetic). A logistics firm wants an assistant that checks a shipment, re-books a carrier and emails the customer. Search alone cannot act, so this is a custom agent with three tools.

The third objective is recognizing when a custom agent is warranted. Search and retrieval answer questions; an agent is for tasks that need tools, several steps and decisions along the way. Google's framework for building them is the Agent Development Kit, or ADK, an open-source agent development framework that lets you build, debug and deploy reliable AI agents at enterprise scale. Its tool ecosystem integrates third-party applications and your own custom code, and it natively supports multi-agent architectures, where specialized agents collaborate and delegate tasks. Once built, an agent needs somewhere to run: ADK agents can run locally or scale globally on Agent Runtime, Cloud Run or Google Kubernetes Engine, and Google says Agent Runtime abstracts away the underlying infrastructure so you can focus on agent logic instead of operations. Not every team needs to start from code, either: Agent Garden offers a library of prebuilt agents and templates, and Agent Studio is a low-code canvas for designing agent workflows. The test is the task: if it only needs an answer, use search or retrieval-augmented generation; if it must act, build an agent.

The controls an agent runs under

Identity, a gateway, a registry and runtime security

ControlWhat it does for the business
Agent IdentityA unique identity per agent, for access control and auditing
Agent GatewayOne policy point that governs every tool call
Agent RegistryA catalog of all agents and tools in the organization
Agent Runtime securityAuthentication, IAM and perimeter compliance settings

Worked example (synthetic). A fictional bank lets its claims agent read policy records but not move money: the gateway enforces the rule and the agent's identity proves which agent asked.

An agent acts on the business's behalf, so the platform pairs it with operating controls, and a leader should know their names and purposes. Agent Identity provides a fully managed, unique identity for each agent, enabling secure access control and auditing — so you can tell which agent did what. Agent Gateway is a central policy enforcement point to govern all agent tool calls, manage authentication and apply security policies — the place where rules like read but do not write are enforced. Agent Registry is a centralized catalog for discovering, tracking and managing all agents and tools across the organization, which stops agents from multiplying unseen. And Agent Runtime itself brings security features including configuration of authentication and identity and access management. In a scenario that asks how to limit what an agent may do, the answer is a control from this table rather than a sentence in the agent's instructions.

Search, prebuilt agent, or custom agent?

Start with the least you need, and build only what is missing

Figure. Three cards left to right. Answer only, questions over your content: use Agent Search or RAG. Common agent task, a pattern others share: start from Agent Garden templates. Your own actions, with tools, steps and your systems: build with ADK and run on Agent Runtime.

Worked example (synthetic). Three fictional requests in one week: an intranet Q&A, a meeting-notes summarizer and a refund agent that touches the payments system.

Put the three options on one line and choose the least that does the job. If the requirement is only to answer questions over your content, an agent is unnecessary — Agent Search, Google's out-of-the-box retrieval-augmented generation system, does it. If the task is a common agent pattern, start from Agent Garden, the library of prebuilt agents and templates that Google offers to accelerate development. And if the agent must take actions in your own systems with tools and several steps, build it with the Agent Development Kit, whose tool ecosystem integrates third-party applications and your own code, and run it on Agent Runtime. Building more than the task needs adds cost and risk without adding value, which is the reasoning the exam rewards.

What this topic actually tests

Name the task, then pick the smallest block that does it

Model? Model Garden to find and deploy; AutoML to train with minimal effort. Your data? Agent Search's prebuilt RAG unless you must control the parts — then RAG APIs or RAG Engine. Action? a custom agent with ADK on Agent Runtime, under identity, gateway and registry controls.

Close on one habit: name the task, then pick the smallest block that does it. If the task is about a model, Model Garden helps you discover, test, customize and deploy one, while AutoML trains your own with minimal technical effort. If the task is answering from your own data, retrieval-augmented generation is the recommended technique, and Agent Search provides it out of the box; reach for the RAG application programming interfaces or RAG Engine only when the requirement names a part of the pipeline you must control. And if the task needs actions through tools, build a custom agent with the Agent Development Kit, run it on Agent Runtime, and govern it with Agent Identity, Agent Gateway and Agent Registry. The next topic looks closely at those tools.

Official sources for this topic

Ready to study Generative AI Leader (GCP-GAIL)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free