Unit 2.2 study guide — Connect infrastructure and data controls to delivery
Generative AI Leader › Unit 2 › Topic 2
Connect infrastructure and data controls to delivery
Study guide for Generative AI Leader, Unit 2 · Topic 2. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks, quoted. Identifying the essential components of Google Cloud’s AI-optimized infrastructure and its benefits (e.g., hypercomputer, Google’s custom-designed TPUs, GPUs, data centers, cloud computing). Explaining how Google Cloud's AI platform provides users with control over their data (e.g., security, privacy, governance, open and leading first party models, pre-built and customizable solutions, agents). Describing how Google Cloud's AI platform democratizes AI development (e.g., low-code and no-code tools, pre-trained models, APIs).
This hive's learning objectives for the topic:
- Explain AI Hypercomputer, TPU, GPU, data-center and cloud roles at a business level.
- Identify data-control choices across models, governance, prebuilt solutions and agents.
- Choose low-code, no-code, pretrained-model or API paths for an organization capability.
Infrastructure, data control and who can build
What runs the models, who controls the data, and who can build with them
Infrastructure — AI Hypercomputer, TPUs, GPUs and the regions they run in, in business terms. Data control — what each layer does with your data, and what you configure. Paths to build — no-code, low-code, pre-trained APIs and code, matched to the people doing the work.
This topic continues Google Cloud's strengths at the level a business leader needs. The guide groups three ideas here. First, Google Cloud's AI-optimized infrastructure, where AI means artificial intelligence: the hypercomputer, Google's custom-designed Tensor Processing Units, or TPUs, graphics processing units, or GPUs, data centers and cloud computing. You are not asked to size a cluster; you are asked what each piece is for. Second, how the platform gives users control over their data — across models, governance, prebuilt solutions and agents. Third, how the platform democratizes development: low-code and no-code tools, pre-trained models and application programming interfaces, or APIs, so that more of an organization can build. Each objective ends in a decision a leader actually makes: which layer fits a workload, which control must be switched on, and which path fits the people who will do the work.
AI infrastructure, in business terms
One integrated system: accelerators, open software, flexible capacity
- AI Hypercomputer: Google Cloud's integrated system for AI workloads
- Accelerators: GPUs and Google's TPUs, with networking and storage
- Open software: familiar machine learning frameworks, optimized
- Consumption options: provisioning that fits each workload
- Regions and zones: resilience against failures at each level
Worked example (synthetic). A retailer plans a recommendation model. Its leaders need to know the work runs on accelerators sized and bought to fit the job — not which chip generation to order.
Start with the infrastructure, described the way a leader needs it. Google calls AI Hypercomputer an integrated supercomputing system that is optimized to support artificial intelligence, or AI, and machine learning workloads, and the unified portfolio for Google Cloud's AI infrastructure — one entry point to performance-optimized hardware, open software and flexible consumption models. Its hardware layer provides accelerators — NVIDIA graphics processing units, or GPUs, and Google Cloud Tensor Processing Units, or TPUs — together with networking and storage. TPUs are Google's custom-developed, application-specific chips used to accelerate machine learning workloads, designed for the large matrix operations machine learning relies on. The software layer offers optimized versions of familiar frameworks such as PyTorch, JAX and TensorFlow, and the consumption layer offers flexible provisioning based on your workload's needs. Underneath it all is the cloud itself: high availability helps ensure resilience against failures at the component, zone or region level. Google's word for why the parts are integrated is goodput — the measure of actual machine learning productivity.
Three integrated parts
Hardware, software and capacity, designed to work as one system
Figure. A framework figure titled AI Hypercomputer: three integrated parts. One band labelled AI Hypercomputer holds three boxes: performance-optimized infrastructure, with GPUs, TPUs, networking and storage; open software, with PyTorch, JAX and TensorFlow; and consumption options, provisioning that fits the workload. A callout reads: integrated to maximize goodput, productive machine learning work.
Worked example (synthetic). A media company's training job spends much of its time waiting on storage. The integrated view tells its leaders that faster chips alone will not fix it; the whole system decides productivity.
The figure draws the three parts Google's overview describes. Performance-optimized infrastructure provides the accelerators — graphics processing units, or GPUs, and Tensor Processing Units, or TPUs — along with networking and storage. Open software offers optimized versions of machine learning frameworks such as PyTorch, JAX and TensorFlow, plus orchestration platforms. Consumption options offer flexible provisioning based on workload needs. The point of drawing them as one band is Google's own: AI Hypercomputer, Google's artificial intelligence system, vertically integrates these three components to maximize goodput, the measure of actual machine learning productivity. For a leader the lesson is that speed comes from the system, not from any single chip — a slow storage path or the wrong way of buying capacity can waste the best accelerator.
CPU, GPU or TPU: the fit in plain terms
TPUs suit specific workloads; Google names when the others fit better
| Processor | Google's examples of a good fit | A leader's reading |
|---|---|---|
| CPU | Quick prototyping that needs maximum flexibility | Early experiments |
| GPU | Medium-to-large models with larger batch sizes | Broad, flexible acceleration |
| TPU | Models dominated by matrix computations | Large, steady training and serving |
| Not TPU | Workloads needing high-precision arithmetic | Choose another processor |
Worked example (synthetic). A bank's quants run a risk model that needs high-precision arithmetic. Google's TPU guidance lists exactly that kind of workload as unsuited to TPUs, so the team looks at other processors.
Google's own guidance on when to use each processor keeps leaders honest about the custom chip. Cloud TPUs, it says, are optimized for specific workloads, and in some situations you might want to use GPUs or CPUs instead. Its examples translate well into business terms. Central processing units, or CPUs, fit quick prototyping that requires maximum flexibility. Graphics processing units, or GPUs, fit medium-to-large models with larger effective batch sizes. Tensor Processing Units, or TPUs, fit models dominated by matrix computations. And Google lists workloads TPUs do not suit, including workloads that require high-precision arithmetic. Wherever the work lands, Google notes you can use TPUs through Compute Engine, Google Kubernetes Engine and Gemini Enterprise Agent Platform — so the choice of processor does not dictate the choice of tools.
Control over your data, layer by layer
Each layer has its own data promise, and its own switch
- Models: no training on your data without permission
- Governance: agent identity, registry and protection policies
- Prebuilt solutions: Workspace and Cloud assistance keep data in bounds
- Agents and search: answers respect who is allowed to see what
- Retention: zero data retention needs actions you take
Worked example (synthetic). A law firm adopts both Workspace with Gemini and a custom agent. Its data-protection officer checks two different documents: the Workspace privacy commitments, and the agent's identity and retention settings.
The second objective is the control Google Cloud's artificial intelligence, or AI, platform gives users over their data, and the useful way to hold it is layer by layer. For models, Google states it won't use your data to train or fine-tune any AI or machine learning model without your prior permission or instruction — yet to achieve zero data retention, customers must take specific actions. For governance, Agent Platform gives each agent a fully managed, unique identity for access control and auditing, a centralized registry of agents and tools, and governance policies such as content protection to mitigate risks like data leakage. Model Armor is a service designed to enhance the security and safety of AI applications. For prebuilt solutions, Google's Workspace privacy hub says your interactions with Gemini stay within your organization, and that content is not human reviewed or used for model training outside your domain without permission; Gemini in Google Cloud does not use your prompts or responses to train its models. For agents and search, Gemini Enterprise enforces access-controlled search results and generative answers, and Gemini Notebook Enterprise keeps your data within your Google Cloud project.
Where each data control lives
Ask about the layer you are actually using
| Layer | Google's commitment or control | What you decide |
|---|---|---|
| Models | No training on your data without permission | Retention settings |
| Governance | Agent identity, registry, protection policies | Permissions and policies |
| Prebuilt: Workspace | Interactions stay within your organization | Which users get it |
| Agents and search | Access-controlled results and answers | Which data sources connect |
| Notebooks | Data stays within your project | What sources are added |
Worked example (synthetic). A hospital's CIO is asked, 'Does Gemini keep our data private?' She answers per layer: Workspace commitments for staff email, project-scoped notebooks for research, retention settings for the patient-facing app.
The table is the same idea arranged for a meeting. Ask about the layer in use, because each has its own commitment and its own decision. Models: Google won't use your data to train models without permission, and you decide the retention settings. Governance: agents get a fully managed identity enabling secure access control and auditing, and you set the permissions and policies. Prebuilt Workspace assistance: your interactions with Gemini stay within your organization, and you decide which users get it. Agents and search in Gemini Enterprise: results and answers are access-controlled, and you choose which data sources to connect. Notebooks in Gemini Notebook Enterprise: your data is always within your Google Cloud project. A question like 'does the artificial intelligence, the AI, keep our data private?' has no single answer — it has one answer per layer.
The controls a single request passes
Identity, policy and screening sit between the user and the model
Figure: A left-to-right flowchart of one request. An employee asks; search and answers are limited to what that employee may see; an agent acts under its own identity; governance policies and Model Armor screen the interaction; the model receives the request and does not train on the data.
Worked example (synthetic). An analyst asks an agent for last quarter's contract terms. The agent can only reach documents its identity is granted, and protection policies block it from pasting a client's personal data into the answer.
Follow one request through the controls. An employee asks a question; in Gemini Enterprise, search results and generative answers are access-controlled, so they are limited to what that employee may see. If an agent acts on the request, it does so under its own fully managed identity, which enables access control and auditing of what it touches. Governance policies such as content protection mitigate risks like data leakage, and Model Armor is designed to enhance the security and safety of the artificial intelligence, or AI, application. Finally the model handles the request — and Google won't use that data to train or fine-tune models without permission. Each box is a place where a leader can ask who owns the setting.
Four paths to a capability
Match the path to the task and to the people who will build it
- Pre-trained APIs: translation, text recognition, document processing
- No-code: Workspace Studio and Workflow Builder for business users
- Low-code: Agent Studio's visual canvas for designing agents
- Code: notebooks and the Agent Development Kit for developers
Worked example (synthetic). An insurer needs to pull fields from scanned claim forms. A pre-trained document service already does that; building a custom model would spend months on a solved problem.
The third objective is how the platform democratizes artificial intelligence, or AI, development — low-code and no-code tools, pre-trained models and application programming interfaces, or APIs — and the decision it supports is which path fits a capability. When a pre-trained API already does the job, use it: Cloud Translation's Basic API gives quick, plug-and-play access to a standard translation model; the Vision API performs optical character recognition, turning text in images into machine-coded text; and Document AI targets document workflows that are usually time-intensive and manual — getting raw text from documents and extracting the fields that matter. When business users should build, no-code tools fit: Workspace Studio automates everyday work with no coding required, and Gemini Enterprise's Workflow Builder is an interactive no-code, low-code platform for creating chat agents and workflows. When the task needs a designed agent, Agent Studio lets you design agents and interact with models without code, on a low-code visual canvas. And when developers need full control, code-based development in notebooks and the Agent Development Kit, a model-agnostic framework for complex agents, are there.
Choosing the path
Start from what already exists, and from who will build
Figure: A left-to-right decision flowchart. From a capability needed, if a pre-trained API already does it, use the API. If not, and business users will build it, use no-code tools. If not, and it does not need custom agent logic in code, use the low-code Agent Studio. If it does, use code: notebooks and the Agent Development Kit.
Worked example (synthetic). An HR team wants an onboarding helper that answers policy questions and books training. No API does that, HR staff will build it, so the team starts in a no-code tool and escalates only if it hits a limit.
The flowchart turns the objective into an order of questions. First, does a pre-trained API already do it? Translation, text recognition and document field extraction are solved problems, and Cloud Translation's Basic API, for example, gives plug-and-play access to a standard model. If not, will business users build it? Then a no-code tool fits — Workspace Studio needs no coding, and Workflow Builder is no-code and low-code. If not, does the capability need custom agent logic in code? If it does not, Agent Studio's low-code visual canvas lets a team design agents without code. Only when the answer is yes do you reach for code: notebooks for code-based development and the Agent Development Kit, abbreviated ADK, for complex agents. Each step down the chart costs more skill and time, which is why the exam usually rewards the highest step that still works.
Four paths, side by side
Who builds, and what Google offers on each path
| Path | Google example | Who typically builds |
|---|---|---|
| Pre-trained API | Cloud Translation, Vision API, Document AI | Developers wire it in |
| No-code | Workspace Studio, Workflow Builder | Business users |
| Low-code | Agent Studio visual canvas | Analysts and builders |
| Code | Notebooks, Agent Development Kit | Developers |
Worked example (synthetic). A support director lists five automation ideas and assigns each to a row: two to APIs, two to no-code staff projects, one to the developers. The developers' queue shrinks to the one idea that needs them.
Side by side, the four paths differ mostly in who builds. Pre-trained application programming interfaces, or APIs, for artificial intelligence, or AI — Cloud Translation, the Vision API with its text recognition, Document AI for document workflows — are wired in by developers but need no model building. No-code tools — Workspace Studio, which needs no coding, and Gemini Enterprise's Workflow Builder — put building in business users' hands. Low-code means Agent Studio's visual canvas for designing and managing agent workflows. And code means notebooks and the model-agnostic Agent Development Kit. Google also lets developers customize, tune and deploy Gemini models at scale on Agent Platform to build differentiated applications — the code path at its fullest. Democratizing development means more of these rows are open to more of the organization.
What this topic actually tests
Three decisions a leader makes
Which layer fits? an integrated system; TPUs for specific workloads, GPUs and CPUs for others. Which control applies? ask per layer — models, governance, prebuilt solutions, agents. Which path? a pre-trained API first, then no-code, low-code, code.
Close on three decisions. Which infrastructure fits? AI Hypercomputer, Google's artificial intelligence system, integrates accelerators, open software and flexible consumption to maximize productive machine learning work, and Google's own guidance places Tensor Processing Units, or TPUs, on specific workloads, with graphics processing units, or GPUs, and central processing units, or CPUs, fitting others. Which data control applies? Ask per layer: models will not be trained on your data without permission, agents carry their own identities, Workspace interactions stay within your organization, and zero data retention requires actions you take. Which path builds the capability? Use a pre-trained API when one already does the job, then the no-code, low-code and code paths in that order. The next topic moves from the platform to the prebuilt tools employees use every day.
Official sources for this topic
- AI Hypercomputer overview
- Introduction to Cloud TPU
- Gemini Enterprise Agent Platform and zero data retention
- Agent Platform overview — Gemini Enterprise Agent Platform
- AI and ML perspective: Reliability — Cloud Architecture Center
- Model Armor overview
- Generative AI in Google Workspace Privacy Hub
- How Gemini products in Google Cloud use your data
- What is Gemini Enterprise? — Google Cloud
- What is Gemini Notebook Enterprise? — Google Cloud
- Cloud Translation overview
- Vision API features list
- Document AI overview
- Google Workspace with Gemini — Google Workspace Help
- Workflow Builder overview — Gemini Enterprise
- Google AI Studio vs. Gemini Enterprise Agent Platform vs. Gemini Enterprise app — Google Cloud