Unit 1.6 study guide — Locate decisions in the generative AI landscape
Generative AI Leader › Unit 1 › Topic 6
Locate decisions in the generative AI landscape
Study guide for Generative AI Leader, Unit 1 · Topic 6. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks, quoted. Infrastructure Models Platforms Agents Applications
This hive's learning objectives for the topic:
- Explain the infrastructure layer and its compute, storage and networking responsibilities.
- Distinguish reusable models from the applications that use them.
- Explain how platforms support development, evaluation, deployment and operations.
- Explain how agents combine model decisions with tools and bounded workflows.
- Identify the user-facing application layer and its business workflow.
Locating decisions in the generative AI landscape
Five layers: infrastructure, models, platforms, agents, applications
Infrastructure — the compute, storage and networking everything runs on. Models — reusable engines such as Gemini. Platforms — where teams build, deploy and govern. Agents — models given goals and tools. Applications — what people actually use at work.
A generative AI leader is constantly asked to decide things — which model, which tool, who may approve what — and the guide asks you to place each decision on one of five layers. From the bottom: infrastructure, the compute, storage and networking that everything runs on; models, the reusable engines such as Gemini; platforms, where teams build, deploy and govern solutions; agents, which give a model goals and tools; and applications, the products employees and customers actually use. The layers matter because a decision made on one layer depends on the ones beneath it. An application cannot be more reliable than the platform it runs on, and an agent can only do what its tools and permissions allow. Google's own phrase for the relationship between the middle two is that foundation models form the core upon which numerous generative AI applications are built.
Infrastructure: what everything runs on
Accelerators, networking and storage — usually chosen for you
- Accelerators: NVIDIA GPUs and Google Cloud TPUs
- TPUs: Google-designed chips for large matrix operations
- Networking and storage move and hold the data
- AI Hypercomputer: Google Cloud's unified AI infrastructure portfolio
Worked example (synthetic). A bank calls a hosted model through an API and never sees a chip. A research lab training its own model must choose accelerators and storage — the same layer, decided by very different people.
The first layer is infrastructure, and its job is compute, storage and networking. Google describes AI Hypercomputer as an integrated supercomputing system optimized for artificial intelligence and machine learning workloads, and as the unified portfolio for Google Cloud's AI infrastructure — performance-optimized hardware, open software and flexible consumption models. Its hardware layer provides accelerators, both NVIDIA GPUs — graphics processing units — and Google Cloud TPUs, along with networking and storage resources. TPUs, or Tensor Processing Units, are Google's custom-developed chips used to accelerate machine learning workloads; they train models efficiently with hardware designed for the large matrix operations machine learning relies on. For most businesses this layer is a decision made for them: a team calling a hosted model never chooses a chip. It becomes a leader's decision when the organization trains or serves its own models and must weigh hardware, scale and how capacity is bought.
The five layers in order
Infrastructure is the base everything else depends on
Figure. A band titled generative AI landscape holds five boxes in order: 1 Infrastructure, with TPUs, GPUs, storage and networking; 2 Models, such as Gemini, Gemma and Veo; 3 Platforms, to build, deploy, govern and optimize; 4 Agents, models plus tools and memory; 5 Applications, what employees and customers use. A note reads that a business decision usually lands on one layer and depends on the layers below it.
Worked example (synthetic). A team wants its customer-service app to answer faster. The fix might be in the application, a different model, or more serving capacity — locating the decision tells you whose job it is.
Here are the five layers the guide names, in its order. Infrastructure sits first because everything else runs on it. Models come next — reusable engines such as Gemini, Gemma and Veo. Platforms are where those models are turned into solutions: Google describes Gemini Enterprise Agent Platform as a unified platform to build, deploy, govern and optimize AI agents and model-based solutions. Agents add goals, tools and memory to a model. And applications are what people actually use. The note on the figure is the skill this topic tests: a business decision usually lands on one layer and depends on the layers beneath it. Asking which layer a decision belongs to also tells you who usually makes it — a procurement and engineering team for infrastructure, a builder for platforms and agents, a business owner for applications.
Models and the applications built on them
One model can serve many applications; an app serves one purpose
- A foundation model is the core many applications are built on
- A simple prompt can steer one model to many different tasks
- Model Garden offers Google's, third-party and open-source models
- The application adds the workflow, data and users
Worked example (synthetic). One company runs a contract summariser and a sales-email drafter on the same model. Changing the model affects both apps; changing the drafter's screens affects only the drafter.
The second objective is the line between a model and an application. A foundation model is a reusable engine, and Google's description is that foundation models form the core upon which numerous generative AI applications are built. Their flexibility comes from what Google calls emergent abilities: with a simple text prompt, a foundation model can learn to perform a variety of tasks — translating languages, answering questions, writing code — without explicit training for each one. Google Cloud's Model Garden gives access to Google's frontier models such as Gemini, plus third-party and open-source models. An application is different in kind. It wraps a model in a particular workflow, connects it to particular data and puts it in front of particular users. That is why a model decision ripples into every application built on it, while an application decision stays local to that application.
Model or application?
Ask whether the thing is reused across purposes, or built for one
| Question | Model | Application |
|---|---|---|
| What is it for? | Many tasks, steered by prompts | One business purpose |
| Who uses it? | Builders, through a platform | Employees or customers |
| What does it add? | General capability | Workflow, data and interface |
| Who feels a change? | Every app built on it | That app's users |
Worked example (synthetic). Upgrading the model under five internal tools needs testing in all five. Redesigning one tool's screen needs testing in one.
Four questions separate the two. What is it for? A model serves many tasks, steered by prompts; an application serves one business purpose. Who uses it? Builders reach models through a platform such as Model Garden, while employees or customers use applications. What does it add? A model brings general capability; the application brings the workflow, the data connections and the interface. And who feels a change? Because foundation models form the core of numerous applications, changing a model is felt by every application built on it, while changing one application is felt by that application's users. In an exam scenario, a decision that affects several products at once is usually a model or platform decision; one that affects a single workflow is usually an application decision.
Platforms: where solutions are built and run
Build, scale, govern and optimize, in one place
- Agent Platform: build, deploy, govern and optimize agents and solutions
- Covers the lifecycle from 200+ foundation models to running agents
- Low-code: Agent Studio designs agents without code
- Code: Agent Development Kit builds complex agents
- Integrated security and governance for enterprise needs
Worked example (synthetic). A retailer's data team builds a product-question agent in Agent Studio, connects its catalogue through RAG Engine, and governs who can deploy it — three platform decisions before any shopper sees it.
The third layer is the platform, where models become solutions. Google describes Gemini Enterprise Agent Platform as a unified platform to build, deploy, govern and optimize enterprise-grade AI agents and model-based solutions. As an evolution of Vertex AI, it supports the complete AI lifecycle, from accessing over two hundred foundation models to deploying and managing agents. It meets builders at different skill levels: Agent Studio lets you design agents and interact with models without code, while the Agent Development Kit is a modular, model-agnostic framework for building and deploying complex agents. It connects models to company knowledge too — RAG Engine securely connects private enterprise data to large language models to improve answer accuracy and reduce hallucinations. And it answers enterprise requirements with integrated security and governance. A platform decision is therefore a decision about how solutions are built, evaluated and controlled, not about any single application.
The platform's four pillars
Each pillar is a set of decisions a team makes on the platform
| Pillar | Decisions it holds | Example capability |
|---|---|---|
| Build | Design, prototype and develop agents | Agent Studio, Model Garden |
| Scale | Deploy and manage agents in production | Agent Runtime |
| Govern | Secure, catalog and govern agents | Agent Identity permissions |
| Optimize | Assess, monitor and refine quality | Agent evaluation |
Worked example (synthetic). A team debating who may deploy a new agent is in Govern; a team debating Agent Studio against code is in Build.
Google organizes Agent Platform around four pillars: build, scale, govern and optimize. Reading them as four sets of decisions makes the platform layer concrete. Build is where you design, prototype and develop agents — Agent Studio for low-code design, Model Garden for the choice of model. Scale is deploying and managing them, with Agent Runtime as the scalable runtime environment. Govern is securing, cataloguing and governing agents with enterprise-grade controls and policies — Agent Identity, for example, lets you grant granular permissions to agents. Optimize is assessing, monitoring and refining agents to keep quality and reliability high. When an exam scenario describes a team arguing about permissions, deployment or evaluation, the argument is happening on the platform layer, whatever application it is for.
Agents: a model with goals and tools
Agents reason, plan and act — within the tools and permissions given
- Agents use AI to pursue goals and complete tasks for users
- They reason, plan and remember, with some autonomy
- Tools are how an agent acts on its environment
- Large language models are the foundation they reason with
- Permissions and supervision bound what they may do
Worked example (synthetic). An expense agent reads a receipt, checks the travel policy and files the claim. It can file claims because it was given that tool — and not approve payments, because it was not.
The fourth layer is agents. Google defines AI agents as software systems that use AI to pursue goals and complete tasks on behalf of users. They show reasoning, planning and memory, and have a level of autonomy to make decisions, learn and adapt. Two ingredients make that possible. Large language models, or LLMs, serve as the foundation, giving agents the ability to understand, reason and act. And tools are the functions or external resources an agent uses to interact with its environment — looking something up, filing a form, calling a system. That second ingredient is also the boundary. An agent can only do what its tools let it do, and on Agent Platform, Agent Identity lets you grant granular permissions to agents. Google's framing keeps people involved as well: agents can reason and take action on users' behalf with their supervision. A leader's agent decision is therefore mostly about scope — which goals, which tools, which permissions, and where a human checks the work.
How an agent works through a task
Reason, act with a tool, observe, repeat — inside its permissions
Figure: A flowchart. A goal from a user leads to reasoning and planning with the model, then acting through a permitted tool, then observing the result. If the task is not done, the flow loops back to reasoning; when it is done, the agent reports back for review.
Worked example (synthetic). Asked to rebook a cancelled flight, an agent searches flights, holds a seat, checks the fare rule, and returns the option to the traveller to confirm.
This loop is the shape of most agent work. A user gives a goal. The agent reasons and plans with its model — Google names reasoning and planning among an agent's defining traits. It acts through a tool, the function or external resource that lets it touch the world. It observes the result, and if the task is not done it reasons again. When it is done, it reports back, because agents act on users' behalf with their supervision. Two things in the figure are business decisions rather than technical ones: which tools the agent is permitted to use, and where the loop stops for a person to review. Those two choices decide how much autonomy an agent really has.
Agents, assistants and bots
The difference is how much they decide on their own
| System | Autonomy | Typical behaviour |
|---|---|---|
| AI agent | Most | Pursues a goal; plans and acts |
| AI assistant | Less | Needs user input and direction |
| Bot | Least | Follows pre-programmed rules |
Worked example (synthetic). A website menu that answers fixed questions is a bot. A chat helper that drafts a reply when asked is an assistant. A system that resolves a refund end to end is an agent.
Google separates three things that are often all called agents. AI agents pursue goals and complete tasks with a level of autonomy. AI assistants, Google says, are less autonomous, requiring user input and direction. Bots are the least autonomous, typically following pre-programmed rules. The scale matters when you place a decision. Moving from a bot to an assistant changes how the model is used; moving from an assistant to an agent changes what the system is allowed to do without asking — which is why that step brings governance questions such as tool permissions and review points with it.
Applications: where people meet AI
The layer users see, built around a business workflow
- Gemini Enterprise: intranet search, AI assistant and agentic platform
- Permissions-aware search across enterprise information
- Gemini in Workspace: help in the flow of work in Gmail, Docs and Meet
- An application is judged by the workflow it improves
Worked example (synthetic). A sales team uses Gemini in Gmail to draft follow-ups and Gemini Enterprise to search past deals. Neither team member chooses a model or a chip; they choose how their work changes.
The fifth layer is applications — the products people actually use, organized around a business workflow. Google describes Gemini Enterprise as an intranet search, AI assistant and agentic platform that empowers knowledge workers with generative AI and agentic workflows by drawing on data sources from across the organization. It gives employees a single, multimodal search interface with permissions-aware access to enterprise information, so people find only what they are allowed to see. Google Workspace with Gemini puts help directly in the flow of work: help me write in Gmail and Docs to compose emails and documents, and take notes for me in Meet so people can focus on the conversation. Application decisions are a business owner's decisions — which workflow to change, for which people, and how to measure whether it got better — and they rest on every layer underneath.
Two Google applications, two workflows
Each one is defined by the work it changes, not the model under it
| Application | Who uses it | What it changes |
|---|---|---|
| Gemini Enterprise | Knowledge workers | Search, assistance and agents across company data |
| Workspace with Gemini | People in Gmail, Docs, Meet | Help in the flow of work: drafting, notes |
Worked example (synthetic). A legal team needs to find past clauses across many systems — a Gemini Enterprise workflow. The same team wants first drafts inside Docs — a Workspace with Gemini workflow.
Two of Google's applications show how this layer is defined by work rather than technology. Gemini Enterprise is an intranet search, AI assistant and agentic platform for knowledge workers, drawing on data sources from across the organization, with permissions-aware access so people see only what they may. Google Workspace with Gemini puts AI help directly in the flow of work — help me write in Gmail and Docs, take notes for me in Meet. Both may run on the same family of models; what separates them is the workflow each changes and the people who use it. That is the question an application decision answers.
Which layer does this decision belong to?
Name the layer, and you have named who decides
Figure. A two-column table of decisions and layers. Train on TPUs or GPUs: infrastructure. Which model powers our tools: models. Who may deploy agents: platforms. May it issue refunds alone: agents. Which team's workflow changes: applications.
Worked example (synthetic). Five questions from one project, each placed on a different layer of the guide's landscape.
Here the skill is applied to five questions from a single project. Whether to train on TPUs or GPUs — Tensor Processing Units or graphics processing units — is an infrastructure decision. Which model powers the company's tools is a model decision, felt by every application built on that model. Who may deploy agents is a platform decision — governance, in Agent Platform's terms. Whether an agent may issue refunds on its own is an agent decision about tools and permissions. And which team's workflow changes is an application decision. Naming the layer also names who usually decides, which is what the exam means by locating a decision in the landscape.
What this topic actually tests
Place the decision, then ask what it depends on
Infrastructure runs everything. Models are reused across applications. Platforms build, deploy and govern. Agents act through tools within permissions. Applications change a workflow for real people.
Close by placing decisions. Infrastructure — accelerators such as Tensor Processing Units, TPUs, and graphics processing units, GPUs, with networking and storage — runs everything. Models are reusable engines; foundation models are the core numerous applications are built on, so a model change ripples outward. Platforms such as Agent Platform are where teams build, deploy, govern and optimize. Agents give a model goals and tools, and act only within the permissions they are granted. Applications such as Gemini Enterprise and Gemini in Workspace change a workflow for real people. When a scenario describes a decision, name its layer first, then ask what it depends on underneath. The next topic opens the model layer and looks at Google's model families one by one.
Official sources for this topic
- AI Hypercomputer overview — Google Cloud
- Develop a generative AI application — Google Cloud
- Agent Platform overview — Gemini Enterprise Agent Platform
- What are AI agents? — Google Cloud
- What is Gemini Enterprise? — Google Cloud
- Introduction to Cloud TPU — Google Cloud
- Google Workspace with Gemini — Google Workspace Help