This map is built from Google's Cloud Digital Leader exam guide (retrieved 2026-09-22) and from this hive's own contents. Google publishes no exam code, no guide version and no passing score for this certification; "GCP-CDL" is this hive's internal label, not Google's.
The exam, as Google publishes it
Format: 50-60 multiple choice and multiple select questions
Time: 90 minutes
Validity: 3 years
Languages: English, Japanese, Spanish, Portuguese, French
Passing score: not published by Google
Sections and weights
Google attaches an approximate percentage to each section and to nothing below it. The counts are this hive's.
Section
Title
Weight (Google, approx.)
Topics
Objectives
Bank questions
Flashcards
Seats per mock
Unit 1
Digital Transformation with Google Cloud
~17%
3
17
82
85
9
Unit 2
Exploring Data Transformation with Google Cloud
~16%
3
15
78
75
9
Unit 3
Innovating with Google Cloud Artificial Intelligence
~16%
3
13
78
65
9
Unit 4
Modernize Infrastructure and Applications with Google Cloud
~17%
6
18
82
91
9
Unit 5
Trust and Security with Google Cloud
~17%
3
14
82
70
9
Unit 6
Scaling with Google Cloud Operations
~17%
3
13
82
65
9
How this hive is organised
A lecture deck for every topic — 21 decks, one per topic in the exam guide, each teaching every objective under it. Every factual claim on a slide is pinned to an exact excerpt from an official Google Cloud page.
484 practice questions, each citing the Google Cloud documentation page its answer comes from, with the supporting quote recorded.
8 blueprint-weighted mock exams of 54 questions, dealt in proportion to the section weights (9/9/9/9/9/9 seats per section). No question appears in more than one mock.
451 flashcards, one collection per objective, each card traceable to Google's documentation.
A study guide per topic (below): the topic's lecture in reading form.
Names in the exam guide that Google has since changed
The exam guide still uses these names, so expect them on the exam. Google's current documentation uses the second column, and the decks teach both.
The exam guide says
Google's documentation now
Taught in
Vertex AI
Gemini Enterprise Agent Platform, which Google describes as an evolution of Vertex AI
Unit 3 Topics 2–3
Anthos (as a single control panel)
GKE fleet management — GKE Enterprise features are now part of standard GKE
Unit 4 Topics 4 and 6
Cloud Functions
Cloud Run functions
Unit 4 Topic 3
preemptible VMs
Spot VMs, which Google calls the latest version of preemptible VMs and recommends instead
Unit 4 Topic 2
Cloud Spanner, Cloud Bigtable
Spanner, Bigtable
Unit 2 Topic 2
SecOps
Google Security Operations (Google SecOps)
Unit 5 Topic 2
Unit Roadmap604 words
Unit 1 roadmap — Digital Transformation with Google Cloud
Unit 1 roadmap — Digital Transformation with Google Cloud
Unit 1 at a glance
Exam weight (Google, approx.)
~17%
Topics
3
Objectives
17
Bank questions
82
Flashcards
85
Seats per mock
9
Every objective below is quoted from Google's Cloud Digital Leader exam guide, retrieved 2026-09-22. Under each topic is that topic's own summary of what the exam actually tests, taken from its lecture.
Topic 1 — Why Cloud Technology is Transforming Business
CDL-U1.T1 · 8 objectives · lecture deck of 16 slides
Define the terms: cloud, cloud technology, data, digital transformation, cloud-native, open source, open standard.
Describe the differences between cloud technology and traditional or on-premises technology.
Explain the benefits of cloud technology to a business’ digital transformation: this technology is scalable, flexible, agile, secure, cost-effective and offers strategic value.
Describe the primary benefits of on-premises infrastructure, public cloud, private cloud, hybrid cloud, and multicloud and differentiate between them.
Describe the main business transformation benefits of Google Cloud: intelligence, freedom, collaboration, trust, and sustainability.
Describe the implications and risks for organizations that do not adopt new technology.
Describe the drivers and challenges that lead organizations to undergo a digital transformation.
Describe the transformation cloud and how it accelerates an organization’s digital transformation through app and infrastructure modernization, data democratization, people connections, and trusted transactions.
What this topic actually tests.Which subject is the definition about? cloud, cloud native, transformation. Which axis is the estate on? deployment model, and separately vendor count. Which mechanism is under the benefit? a claim with nothing under it is the distractor.
Topic 2 — Fundamental Cloud Concepts
CDL-U1.T2 · 5 objectives · lecture deck of 13 slides
Describe how transitioning to a cloud infrastructure affects flexibility, scalability, reliability, elasticity, agility, and total cost of ownership (TCO). Apply these concepts to various business use cases.
Explain how an organization’s transition from an on-premises environment to the cloud shifts their capital expenditures (CapEx) to operational expenditures (OpEx), and how that affects their total cost of ownership (TCO).
Identify when private, hybrid, or multicloud infrastructures best apply to different business use cases.
Define basic network infrastructure terminology, including: IP address; internet service provider (ISP); domain name server (DNS), regions, and zones; fiber optics; subsea cables; network edge data centers, latency; and bandwidth.
Discuss how Google Cloud supports digital transformation with global infrastructure and data centers connected by a fast, reliable network.
What this topic actually tests.Scalability or elasticity? only one comes back down. Cheaper up front or cheaper in total? TCO counts managing. Which constraint is in the scenario? it picks private, hybrid or multicloud. How long, or how much? latency is time, bandwidth is capacity — and zones sit inside a region while the edge does not.
Topic 3 — Cloud Computing Models and Shared Responsibility
CDL-U1.T3 · 4 objectives · lecture deck of 11 slides
Define IaaS, PaaS, and SaaS.
Compare and contrast the benefits and tradeoffs of IaaS, PaaS, and SaaS including total cost of ownership (TCO), flexibility, shared responsibilities, management level, and necessary staffing and technical expertise.
Determine which computing model (IaaS, PaaS, SaaS) applies to various business scenarios and use cases.
Describe the cloud shared responsibility model. Compare which responsibilities are the cloud provider’s, and which responsibilities are the customer’s for on-premises and cloud computing models (IaaS, PaaS, SaaS).
What this topic actually tests.What do you still manage? operating system (IaaS), code (PaaS), data (SaaS). What does the scenario want to stop doing? that picks the model — and mixing models is normal. Who secures what? the line moves with the model; your data and access policies never leave you, and the infrastructure never leaves Google.
Study Guide4,268 words
Unit 1.1 study guide — Why Cloud Technology is Transforming Business
Study guide for Cloud Digital Leader, Unit 1 · Topic 1. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks. Explain why and how the cloud is revolutionizing businesses.
Objectives, quoted from the exam guide:
Define the terms: cloud, cloud technology, data, digital transformation, cloud-native, open source, open standard.
Describe the differences between cloud technology and traditional or on-premises technology.
Explain the benefits of cloud technology to a business’ digital transformation: this technology is scalable, flexible, agile, secure, cost-effective and offers strategic value.
Describe the primary benefits of on-premises infrastructure, public cloud, private cloud, hybrid cloud, and multicloud and differentiate between them.
Describe the main business transformation benefits of Google Cloud: intelligence, freedom, collaboration, trust, and sustainability.
Describe the implications and risks for organizations that do not adopt new technology.
Describe the drivers and challenges that lead organizations to undergo a digital transformation.
Describe the transformation cloud and how it accelerates an organization’s digital transformation through app and infrastructure modernization, data democratization, people connections, and trusted transactions.
Why cloud technology is transforming business
Three questions, asked eight ways
This topic asks what the words mean, what changes when you leave the data center, and what an organization gains — or gives up — by moving or standing still. Everything on the exam in this topic is one of those three.
This is the first topic of the Cloud Digital Leader exam, and it sets the vocabulary every later unit assumes. Its eight objectives look like eight separate things, but they are three questions asked eight ways. The first is definitional: what does Google mean by cloud, by cloud native, by digital transformation. The second is comparative: what actually changes when an application leaves a company's own data center. The third is consequential: what a business gains by moving, what the different deployment models each buy it, and what it risks by not moving at all. Hold those three questions and the objectives stop being a list to memorize.
The words, as Google defines them
Cloud, cloud native and digital transformation
Cloud computing: on-demand computing resources as services over the internet
It removes self-managed physical resources; you pay for what you use
Cloud native: an approach to building and running scalable applications
Digital transformation: redefining relationships using new technologies
Worked example (synthetic). A retailer's platform team is asked to "go cloud native". They rent virtual machines and keep the same single deployable. That is cloud hosting, not cloud native — nothing about how the application is built or run has changed.
Start with the three definitions the exam actually tests. Cloud computing, in Google's words, is the on-demand availability of computing resources such as storage and infrastructure, delivered as services over the internet. The second half of that definition matters as much as the first: it eliminates the need to self-manage physical resources, and you pay only for what you use. Cloud native is a different idea entirely. It is an approach to building and running scalable applications so they take full advantage of cloud-based services and delivery models — a statement about how software is designed, not about where it runs. Digital transformation is different again: it is when an organization takes advantage of new technologies to redesign and redefine its relationships with customers, employees and partners. A company can be on the cloud without being cloud native, and cloud native without being transformed. The exam will ask you to tell them apart.
Four terms, four different subjects
What each definition is actually about
Term
What it is a statement about
Cloud computing
Where resources come from — on demand, over the internet, paid by use
Cloud native
How an application is built and run — scalable, decomposed, designed for the cloud
Monolithic application
How an application is released — built, tested and deployed as a single unit
Digital transformation
What an organization changes — its relationships with customers, employees and partners
Worked example (synthetic). An exam item describes a bank that moved its core system unchanged onto rented servers. The correct reading is cloud computing without cloud native: the subject of each definition is different.
Put the four terms side by side and the distinction stops being subtle, because each one is a statement about a different subject. Cloud computing is about where resources come from. Cloud native is about how an application is built and run. A monolithic application is a statement about release: Google's contrast is explicit — unlike monolithic applications, which must be built, tested and deployed as a single unit, cloud-native architectures decompose components into loosely coupled services. And digital transformation is about what the organization changes, not what the technology is. When an exam item mixes two of these, work out which subject the question is really asking about.
Cloud against traditional on-premises technology
What actually changes when you leave the data center
Question
Traditional on-premises
Cloud
Who owns the infrastructure?
The organization, in its own data centers
A third-party provider, reached over the internet
What does a new application wait for?
Infrastructure to be procured and provisioned
Nothing — production without worrying about the underlying infrastructure
How is capacity reached?
Bought ahead of demand
Scaled up or down as needed, from anywhere with a connection
What is paid for?
What was bought
Only the computing resources used
Worked example (synthetic). A team is told a new service must launch in three weeks. On-premises the clock starts with a purchase order; on cloud it starts with a deployment. The difference the exam tests is what the team waits for.
The second objective asks for the difference between cloud technology and traditional on-premises technology, and the honest answer is not a feature list — it is a change in what a team waits for. Google states the operational half plainly: enterprises can develop new applications and rapidly get them into production without worrying about the underlying infrastructure. The access half is just as plain: because of the architecture of cloud computing, enterprises and their users can reach cloud services from anywhere with an internet connection, scaling services up or down as needed. And the commercial half: whatever service model is used, enterprises pay only for the computing resources they use. Cloud native adds a fourth difference, about design rather than operations — Google describes it as adapting to the many new possibilities, but a very different set of architectural constraints, offered by the cloud compared with traditional on-premises infrastructure.
Scale and reach: services up or down, from anywhere with a connection
Security: stronger than enterprise data centers, per Google, by depth and breadth
Cost: pay only for resources used; no overbuilt capacity for spikes
Strategic value: providers carry the latest innovations, so you do not buy obsolescence
Worked example (synthetic). A media company's traffic triples for one week a year. On-premises it owns that peak all year. On cloud it rents the peak for the week and redeploys the IT staff who used to plan for it.
The third objective is a benefits list, and the exam will test whether you can attach each benefit to the right mechanism. Scalability and flexibility come from the architecture: services are reachable from anywhere with an internet connection and scale up or down as needed. Security is a claim Google makes directly — cloud computing security is generally recognized as stronger than that in enterprise data centers, because of the depth and breadth of the security mechanisms cloud providers put into place. Cost-effectiveness has two parts: enterprises pay only for the resources they use, and they no longer overbuild data center capacity to handle unexpected spikes in demand or business growth — which also frees information technology (IT) staff to work on more strategic initiatives. Strategic value is the subtlest one. Because providers stay on top of the latest innovations and offer them as services, an enterprise gets more competitive advantage and a higher return on investment than it would investing in soon-to-be obsolete technologies. That is an argument about what you are not buying.
Each benefit, and the mechanism under it
A claim with nothing under it is not a benefit
Figure. Five benefit cards, each naming the mechanism Google states beneath it: scalable and flexible, agile, secure, cost-effective, and strategic value.
Worked example (synthetic). In an exam item a candidate is offered "the cloud is more secure" with no mechanism. The graded answer is the one naming depth and breadth of provider security mechanisms.
This figure exists to stop a benefit being memorized as a slogan. Each of the five claims on the exam has a stated mechanism underneath it, and the mechanism is what distinguishes a right answer from a plausible one. Scalable and flexible rests on services scaling as needed and being reachable from anywhere. Agile rests on production being reachable without worrying about infrastructure. Secure rests on the depth and breadth of the security mechanisms providers put in place. Cost-effective rests on paying for use and not overbuilding. And strategic value rests on providers carrying the latest innovations so the customer does not buy technology that is about to be obsolete. When an item offers you a benefit with no mechanism, that is usually the distractor.
The deployment models, and what each is for
Three models — and multicloud, which is not one of them
Model
Who runs it
What it is chosen for
On-premises / private cloud
A single organization, in its own data centers
Greater control, security and management of data, over a shared internal pool
Public cloud
Third-party providers
Compute, storage and network over the internet, shared and on demand
Hybrid cloud
Both, combined
Public cloud services while keeping private-cloud security and compliance
Multicloud
Two or more providers
Freedom to pick capabilities per vendor and minimize vendor lock-in
Worked example (synthetic). A regulated insurer keeps claims data in its own data center and runs its public quote engine on a provider. That is hybrid. If the quote engine also runs on a second provider, it is additionally multicloud.
The fourth objective is the one candidates most often get half right, because multicloud sits beside the deployment models without being one of them. Google is explicit: there are three different cloud computing deployment models — public cloud, private cloud, and hybrid cloud. Public clouds are run by third-party cloud service providers, offering compute, storage and network resources over the internet as shared on-demand resources. Private clouds are built, managed and owned by a single organization and privately hosted in their own data centers — what is commonly called on-premises — and they provide greater control, security and management of data while still giving internal users a shared pool. Hybrid clouds combine the two, letting a company use public cloud services and maintain the security and compliance capabilities commonly found in private cloud architectures. Multicloud is a separate axis: an organization using cloud computing services from at least two cloud providers to run their applications. A hybrid estate can also be multicloud, and a multicloud estate need not be hybrid.
Two axes, not one list
Hybrid is about model; multicloud is about vendor count
Loading Diagram...
Figure 1 — Mermaid diagram
Figure: A flow chart showing deployment model and vendor count as two independent axes that together describe an estate, with three example combinations: private only, hybrid with one provider, and hybrid with two providers which is both hybrid and multicloud.
Worked example (synthetic). An item says a company runs on-premises plus two public providers and asks for the single best description. Both axes are in play, so the answer names both.
This is the figure that keeps the fourth objective from collapsing into a list. Deployment model and vendor count are independent. Google's own wording supports reading them separately: the three deployment models are public, private and hybrid, while multicloud is defined by a count — cloud computing services from at least two cloud providers. So an estate can be private only, hybrid with one provider, or hybrid with two providers, which is hybrid and multicloud at the same time. The multicloud page adds the reason organizations choose the second axis deliberately: having the freedom to create a strategy that uses multiple vendors lets you pick and choose the capabilities that best suit your business needs and minimize vendor lock-in.
Intelligence: automate complex tasks, gain insight, personalize experiences
Freedom: pick capabilities across vendors and minimize lock-in
Collaboration: speed work across all teams and geographies
Sustainability: a resilient, optimized architecture compounds into lower cost and impact
Worked example (synthetic). A logistics firm adopts one provider's forecasting service and a second's mapping data. The freedom benefit is not that it uses two — it is that neither can hold the workload hostage.
The fifth objective names five business transformation benefits, and each one has something specific behind it. Intelligence is artificial intelligence and machine learning: Google says these technologies can empower organizations to automate complex tasks, gain insights from vast amounts of data, and deliver personalized customer experiences. Freedom is the multicloud argument — the freedom to create a strategy that uses multiple vendors lets you pick and choose the capabilities that best suit your specific business needs and minimize vendor lock-in — and it rests technically on open source: multicloud solutions built on open source technologies like Kubernetes provide the flexibility and portability to migrate, build and optimize applications across multiple clouds. Collaboration is stated as using modern digital tools to speed collaboration across all teams and geographies to deliver more value and faster results for customers. Sustainability is the one candidates skip, and it is the most interesting: a sustainable architecture is resilient and optimized, which creates a positive feedback loop of higher efficiency, lower cost, and lower environmental impact. Efficiency and cost move together with impact, which is why it is a business benefit and not only an ethical one.
Each benefit, and what Google actually publishes about it
Five names, five different kinds of claim
Benefit
What Google publishes
Intelligence
AI and ML automate complex tasks, give insight from vast data, and personalize customer experiences
Freedom
Multi-vendor strategy lets you pick capabilities and minimize vendor lock-in
Collaboration
Modern digital tools speed collaboration across all teams and geographies
Sustainability
A resilient, optimized architecture creates a positive feedback loop of efficiency, cost and impact
Worked example (synthetic). Asked which benefit answers "our two business units cannot work on the same data", the graded answer is collaboration, not intelligence — the obstacle is distance between teams, not analysis.
Reading the five benefits side by side shows they are not five versions of the same claim. Intelligence is about what the technology does to work — automating complex tasks, producing insight, personalizing experience. Freedom is about the business's position relative to its vendors, and the phrase that matters is minimize vendor lock-in. Collaboration is about distance: speeding work across all teams and geographies. Sustainability is about a loop rather than a one-time gain — resilient and optimized systems produce higher efficiency, lower cost and lower environmental impact together. Trust is the fifth name on the exam guide, and it is taught in Unit 5, where Google's security, compliance and transparency commitments are the evidence for it; nothing on this deck asserts it beyond naming it, because the pages that ground it belong to that unit.
The cost of not adopting
Standing still is a decision with consequences
Figure. Three cards naming what an organization forgoes by not adopting new technology: competitiveness, infrastructure efficiency, and speed of collaboration.
Worked example (synthetic). A manufacturer defers modernization for three years to protect margin. Nothing breaks — its competitors simply ship in weeks what it still ships in quarters.
This objective asks for the implications and risks of not adopting new technology, and it needs care, because Google's pages argue the upside rather than publishing a warning. The honest way to teach it is to read the stated upside and name what declining it forgoes. Google says organizations choose digital transformation frameworks as a way to reimagine themselves, staying competitive in their respective businesses and industries — so competitiveness is what is at stake. It says to use technology infrastructure more effectively, whether using containers, moving to serverless computers, or taking advantage of a chosen cloud platform's global network — so efficiency is forgone, not merely postponed. It says to use modern digital tools to speed collaboration across all teams and geographies to deliver more value and faster results for customers — so a competitor that adopts is faster to the same customer. And it is blunt about the direction of travel: digital transformation is what propels businesses and industries forward. An organization that declines is not holding position; the industry moves and it does not.
Drivers and challenges
What pushes a transformation, and what it demands back
Driver: meeting changing business and market dynamics
Driver: the ability to collect, process and analyze large volumes of data
Scope: foundational change to operations, internal resources and value delivered
Challenge: strong commitment from business and IT, and willingness to support the change
Worked example (synthetic). A bank's driver is a regulator's new reporting window. Its challenge is that the team that must change the process does not report to the team that wants the change.
The seventh objective pairs drivers with challenges, and the exam tests both halves. On drivers: Google defines digital transformation as using modern digital technologies — including all types of public, private and hybrid cloud platforms — to create or modify business processes, culture and customer experiences to meet changing business and market dynamics. That last phrase is the driver. A second driver is data itself: the ability to collect, process and analyze large volumes of data is crucial for digital transformation. On scope, the same page is explicit that this is not a tooling change — digital transformation drives foundational change in how an organization operates, how it optimizes internal resources, and how it delivers value to customers. And on challenges, one sentence carries the objective: it requires a strong commitment from both businesses and information technology (IT) teams, as well as a willingness to support the resulting changes. Notice that the challenge Google names is organizational, not technical. Nothing on the page says the hard part is the migration.
From driver to foundational change
The challenge sits between the two, not after them
Loading Diagram...
Figure 2 — Mermaid diagram
Figure: A flow chart with two drivers and one requirement feeding digital transformation, which in turn produces foundational change in how the organization operates, optimizes internal resources, and delivers value to customers.
Worked example (synthetic). An item offers four reasons a transformation stalled. The one Google actually names is the absence of commitment from both business and IT — not budget, not tooling.
Drawn out, the seventh objective has a shape. Two drivers push: changing business and market dynamics, and the volume of data an organization now has to collect, process and analyze. One requirement gates: a strong commitment from both businesses and information technology (IT) teams, and a willingness to support the resulting changes. And three outcomes follow, which are the foundational change Google names — how the organization operates, how it optimizes internal resources, and how it delivers value to customers. The reason to draw it rather than list it is the position of the requirement. It is not an outcome and not an obstacle encountered later; it is a condition on the arrow. A transformation with drivers and no commitment does not proceed slowly, it does not proceed.
The transformation cloud, as four moves
App and infrastructure, data, people, trust
Move
What Google says it does
App and infrastructure modernization
Modernize processes and applications to quickly pinpoint and solve problems in how a business operates and interacts with customers
Off physical data centers
Move to a modern cloud platform to speed deployment time, boost stability, and get new products to market faster
Data democratization
Sharpen the focus on smarter business analytics with the most advanced tools, for keener insight into data
Managing data at scale
Use new tools to better manage the huge amounts of data coming from different devices, sources and systems
Worked example (synthetic). A retailer's checkout defects take a week to locate across four systems. The modernization move is the one that changes that week, not the one that changes the hosting bill.
The eighth objective asks about the transformation cloud and how it accelerates a digital transformation. Google frames it as a set of moves rather than a product, and four of them are stated plainly enough to teach. Application and infrastructure modernization: implement new digital technologies and modernize processes and applications to quickly pinpoint and solve problems related to how a business operates and interacts with customers. Leaving the data center: move from physical data centers to a modern cloud platform to speed deployment time, boost stability, and get new products to market faster. Data democratization, in its analytics half: sharpen the focus on smarter business analytics with the most advanced tools to drive keener insight into data. And the data scale problem underneath it: use new tools and capabilities to better manage the huge amounts of data coming from different devices, sources and systems. People connections and trusted transactions are the guide's other two headings; the first is carried by the collaboration claim you have already seen, and the second belongs to Unit 5's trust and security material, which is where this hive grounds it.
Which move answers which complaint
Match the symptom to the move, not the product
Figure. A two-column table matching four business complaints to the transformation-cloud move that answers each: modernize applications, move off physical data centers, sharpen business analytics, and manage data at scale.
Worked example (synthetic). Given "our quarterly release is the risk", the graded answer is the data-center move — speed of deployment and stability — not the analytics one.
Exam items on this objective almost always arrive as a business complaint rather than as the name of a move, so it is worth rehearsing the mapping in that direction. A complaint about how long it takes to locate a problem points at modernizing processes and applications, which Google describes precisely as quickly pinpointing and solving problems related to how a business operates and interacts with customers. A complaint about release speed or stability points at moving off physical data centers, which is described as speeding deployment time and boosting stability. A complaint about who can answer a data question points at sharper business analytics with the most advanced tools. And a complaint about data arriving from everywhere points at using new tools and capabilities to better manage the huge amounts of data coming from different devices, sources and systems. The trap in these items is a product name in the answer; the objective is about the move.
What this topic actually tests
Three discriminations, eight objectives
Which subject is the definition about? cloud, cloud native, transformation. Which axis is the estate on? deployment model, and separately vendor count. Which mechanism is under the benefit? a claim with nothing under it is the distractor.
Close the topic on the three discriminations it really tests. First, when a definition is in play, ask what the definition is a statement about — where resources come from, how an application is built, or what an organization changes. Second, when an estate is described, remember there are two axes: the deployment model, which is public, private or hybrid, and the vendor count, which is what makes something multicloud. An estate can be on both at once. Third, when a benefit is offered, look for the mechanism underneath it — scaling as needed, paying only for what is used, the depth and breadth of provider security mechanisms, providers carrying the latest innovations so you do not buy obsolescence. Those three habits answer more of this unit than any list you could memorize, and they carry into Units 4 and 5 where the same words come back with technical weight.
Study guide for Cloud Digital Leader, Unit 1 · Topic 2. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks. Explain general cloud concepts.
Objectives, quoted from the exam guide:
Describe how transitioning to a cloud infrastructure affects flexibility, scalability, reliability, elasticity, agility, and total cost of ownership (TCO). Apply these concepts to various business use cases.
Explain how an organization’s transition from an on-premises environment to the cloud shifts their capital expenditures (CapEx) to operational expenditures (OpEx), and how that affects their total cost of ownership (TCO).
Identify when private, hybrid, or multicloud infrastructures best apply to different business use cases.
Define basic network infrastructure terminology, including: IP address; internet service provider (ISP); domain name server (DNS), regions, and zones; fiber optics; subsea cables; network edge data centers, latency; and bandwidth.
Discuss how Google Cloud supports digital transformation with global infrastructure and data centers connected by a fast, reliable network.
Fundamental cloud concepts
What moving to cloud changes: the system, the money, the place, the path
The system — six properties, each with its own mechanism. The money — capital expenditure becomes operational expenditure, and total cost still counts management. The place — private, hybrid or multicloud, read from the scenario. The path — how a user's request reaches Google's network, and the words for every step.
Topic one asked why cloud is transforming business. Topic two asks what, concretely, changes when a workload moves — and the guide's summary for it is just four words: explain general cloud concepts. Its five objectives fall into four questions. What happens to the system itself: six properties — flexibility, scalability, reliability, elasticity, agility and total cost — each with a different mechanism underneath. What happens to the money: spending moves from buying assets up front to paying as resources are consumed, and that changes what total cost of ownership means. Where the workload should live: private, hybrid or multicloud, which the exam asks as a scenario rather than a definition. And how a user actually reaches it: the network vocabulary — addresses, names, providers, regions, zones, cables, the edge, latency and bandwidth — and the global network Google runs underneath all of it. Keep those four questions in view and the objectives stop being a list.
Six properties, six different mechanisms
What the cloud changes about each one
Property
What changes in the cloud
Flexibility
Services reachable from anywhere with an internet connection; the business adapts quickly to changing markets
Scalability
Grow without buying and maintaining your own data centers and servers
Elasticity
Capacity follows load: autoscaling adds instances as load rises and deletes them as it falls
Reliability
Resources spread across zones, and across regions for even higher failure independence
Agility
New applications reach production without worrying about the underlying infrastructure
Total cost of ownership
Counts provisioning, use and managing the resources — not the price alone
Worked example (synthetic). A ticketing site sells out twice a year. Scalability says it can grow for the sale; elasticity says it shrinks again afterwards without anyone deciding to — and only the second changes the bill.
The first objective lists six properties, and the exam separates candidates who can attach each one to its mechanism from those who treat them as synonyms. Flexibility is about reach and adaptation: Google says enterprises and their users can access cloud services from anywhere with an internet connection, and calls flexibility and versatility one of the hallmarks of cloud, allowing enterprises to quickly adapt to changing markets or metrics. Scalability is about growth without ownership — scaling faster and more efficiently without the burden of buying and maintaining your own data centers and servers. Elasticity is the one most often confused with it. Google describes cloud-native applications as designed to take advantage of the elasticity of the cloud, and the concrete mechanism is autoscaling: managed instance groups automatically add or delete virtual machine instances based on increases or decreases in load. Scalability is being able to grow; elasticity is capacity that follows the load in both directions. Reliability is about failure: putting resources in different zones reduces the risk of one infrastructure outage affecting everything at once, and different regions give an even higher degree of failure independence. Agility is about time — new applications reach production without worrying about the underlying infrastructure. And total cost of ownership is about what you count: Google's framework says to consider not only the cost of provisioning and using resources, but also the cost of managing them.
Which property answers which business case
The exam gives you the situation, never the property's name
Figure. Six cards, each pairing a business complaint with the cloud property that answers it: elasticity for a one-week traffic peak, agility for an earlier launch, reliability for a facility failure, flexibility for staff in many countries, total cost of ownership for patching overhead, and scalability for long-term growth.
Worked example (synthetic). An item describes a firm whose peak lasts a week and asks which property cuts its idle spend. Scalability is the tempting answer; elasticity is the graded one, because the saving comes from scaling back in.
The objective does not stop at naming the properties — it says apply these concepts to various business use cases, and that is how the exam asks. You get a complaint, not a property name, so rehearse the mapping in that direction. Paying for idle servers after a short peak is elasticity: Google says autoscaling helps apps gracefully handle increases in traffic and reduce costs when the need for resources is lower. A launch date that moved earlier is agility. A facility failure that must not stop the business is reliability, answered by copies in different zones. Staff spread across many countries is flexibility — access from anywhere with an internet connection. The fifth card is the one candidates miss: virtual machines that looked cheap, and a team spending its weeks patching them. Google's framework says exactly this — virtual machines might be a cost-effective option, but when you consider the overhead to maintain, patch and scale them, the total cost of ownership can increase. And steady long-term growth is scalability. Two of these cards sound alike — the peak and the growth — and the difference is direction: growth only goes up, while elasticity comes back down.
From capital expenditure to operational expenditure
Buying assets up front becomes paying as you consume
On-premises: assets are acquired, then depreciated over their operating life
Cloud: most costs are OpEx, incurred when the resources are consumed
No capacity bought ahead of demand, and none overbuilt for spikes
TCO still counts managing what you run, not just upfront costs
Worked example (synthetic). A clinic network plans an imaging archive. On-premises, finance approves a hardware purchase and depreciates it for years. On cloud, the archive is an operating cost that tracks what is stored — cancelled in month four, it stops costing in month four.
The second objective is the money, and Google's cost optimization framework states the difference in one line: the cost models for on-premises and cloud workloads differ significantly. On-premises information technology (IT) costs include both capital expenditure, or CapEx, and operational expenditure, or OpEx. The capital part works like this: hardware and software assets are acquired, and the acquisition costs are depreciated over the operating life of the assets. The organization pays first and uses later, for years, whether or not the capacity is needed. In the cloud, the costs for most resources are treated as OpEx, incurred when the resources are consumed. That is why Google's cloud computing page says enterprises get scalable resources without needing to worry about capital expenditures or limited fixed infrastructure, only pay for the computing resources they use, and do not need to overbuild data center capacity to handle unexpected spikes. Two cautions keep this from becoming a slogan. First, the word is most: Google notes you might be able to classify the cost of some services, such as Compute Engine sole-tenant nodes, as capital expenditure. Second, the shift does not by itself lower total cost of ownership, or TCO — and TCO is the number Google says to manage: maximize the business value the cloud resources provide and minimize the total cost of ownership. The framework also tells teams to adopt a holistic attitude toward cloud spending, with an emphasis on TCO and not just upfront costs — so a lower upfront bill is the start of the calculation, not the answer.
Two cost timelines, one total
Where the money goes first, and what total cost still counts
Loading Diagram...
Figure 1 — Mermaid diagram
Figure: Two flows side by side. On-premises: acquire hardware and software, then depreciate over the operating life. Cloud: consume a resource, and the cost is incurred as it is consumed. Both flows end in the same box, total cost of ownership, which counts provisioning, use and managing.
Worked example (synthetic). Two proposals land on a finance director's desk: a hardware refresh and a cloud migration. The figure's point is that both must be priced to the same box — including the staff time to run them.
Drawn out, the second objective has a shape that a list hides. On the left, the on-premises timeline: assets are acquired, and the acquisition cost is depreciated over the operating life of the assets. On the right, the cloud timeline: a resource is consumed, and the cost is incurred as it is consumed — which is why Google can say enterprises only pay for the computing resources they use. The two timelines look like opposites, and the tempting conclusion is that the right-hand one is always cheaper. The figure refuses that conclusion by sending both arrows into the same box. Total cost of ownership is the thing actually being compared, and Google's framework defines what goes in it: consider not only the cost of provisioning and using the resources, but also the cost of managing them. A cloud design that trades a purchase order for months of manual maintenance has moved its cost, not removed it.
Private, hybrid or multicloud: which fits
Read the scenario for the constraint, then choose the model
Private: one organization's data centers, for control of data
Hybrid: regulated or resident data on-premises, public scale beside it
Multicloud: best capability per vendor, and no single cloud as a failure point
Worked example (synthetic). A bank must keep customer records in-country but wants a public provider's analytics. That is the hybrid signal: data residency on one side, public scale on the other.
Topic one defined the deployment models. The third objective here asks when each one best applies, and the exam asks it as a scenario. Private cloud is chosen for control: Google says private clouds provide greater control, security and management of data while still giving internal users a shared pool of resources. Hybrid has the most signals, so learn them. It is popular with companies in highly regulated industries that have strict data privacy requirements. It suits you if you want the scale and security of a public cloud while keeping your data on-premises to comply with data residency laws, or supporting computing needs closer to your customers. Migrations lead to it naturally, because organizations often have to transition applications and data slowly and systematically. And it is the answer for industries that demand edge computing for low latency — Google's examples are kiosks in retail and networks in telecom. Multicloud has two signals of its own. One is fit: it lets you match specific features and capabilities to your workloads based on speed, performance, reliability, geographical location, and security and compliance requirements. The other is failure: an outage in one cloud will not necessarily impact services in other clouds. And underneath both sits the freedom Google names directly — a strategy that uses multiple vendors lets you pick the capabilities that best suit your needs and minimize vendor lock-in. Every choice has a price, and Google names hybrid's plainly — you still have to invest in and maintain in-house hardware.
The signal in the scenario, and the model it points to
Each model has a sentence that gives it away
Signal in the scenario
Best fit
Why
Strict control of data in the organization's own data centers
Private
Greater control, security and management of data
Data must stay on-premises under data residency laws
Hybrid
Public scale while the data stays on-premises
Moving applications a few at a time; a mainframe that is hard to move
Hybrid
Migrations transition slowly and systematically
Retail kiosks or telecom sites that need low latency
Hybrid, at the edge
Select apps run at the edge
One provider's outage must not stop the business
Multicloud
An outage in one cloud need not reach the others
Worked example (synthetic). An item says a retailer migrates one application per quarter and asks which model it runs in the meantime. The graded answer is hybrid — the gradual migration is the signal, not the retailer's destination.
Here are the five signals as the exam writes them. Strict control of data in the organization's own data centers points to private cloud. Data that must stay on-premises for data residency laws points to hybrid, because the scale of a public cloud can sit beside it. A migration that moves applications a few at a time, or a mainframe system that is difficult to move to the cloud, points to hybrid as well — Google says regulated applications that need to remain on-premises and mainframe systems are both reasons for it. Retail kiosks or telecom sites that need low latency point to hybrid at the edge. And a requirement that one provider's outage must not stop the business points to multicloud, because an outage in one cloud will not necessarily impact services in the others. Notice that three of the five rows say hybrid. That is not an accident of this table; it reflects how many constraints hybrid is Google's answer to, which is why a scenario with any on-premises constraint in it deserves a second look before you answer public cloud.
Addresses, names and providers
How a request finds the thing it is looking for
Term
What it is
IP address
Lets resources communicate within Google Cloud, with on-premises networks, or on the public internet
External vs internal IP address
External addresses are reachable by any host on the internet; an internal address is not publicly routed
DNS
A hierarchical distributed database that stores IP addresses and looks them up by name
Internet service provider (ISP)
The user's network; ISPs hand traffic to one another, and each hop is part of the latency path
Worked example (synthetic). A customer types a shop's name into a browser. DNS turns the name into an address; the address is external, so any host on the internet can reach it; the customer's ISP carries the request to Google's network.
The fourth objective is vocabulary, and it is easiest to learn in the order a request uses it. An internet protocol address, or IP address, is what lets a resource communicate: Google says these IP addresses let Google Cloud resources communicate with other resources in Google Cloud, in on-premises networks, or on the public internet. There are two kinds that matter here. External IP addresses are publicly advertised, meaning they are reachable by any host on the internet. An internal IP address, by contrast, is not publicly routed. Nobody types an address, though, so the next term is DNS. The exam guide expands it as domain name server; Google's own documentation expands it as the domain name system, and defines it as a hierarchical distributed database that lets you store IP addresses and other data and look them up by name. Google's managed version, Cloud DNS, publishes your domain names to the global DNS. Last in this group is the internet service provider, or ISP — the user's own network. Google's region-selection guidance mentions it precisely because it is where the user's side of the latency path begins: it notes that many architects consider only the network latency, or distance, between the user's ISP and the virtual machine instance. And ISPs pass traffic along to each other: Google describes an ISP handing off traffic to a downstream ISP as quickly as it can, which is why a request can cross several providers before it reaches Google's network.
Regions, zones and the edge
Zones sit inside a region; the point of presence does not
Figure. Users reach a Google point of presence first, outside any region. Traffic is then carried to a region that contains three zones, each a single failure domain running a copy of the application. The callout notes that zones are inside the region and the point of presence is not.
Worked example (synthetic). A candidate proposes adding a zone to fix slow pages for users on another continent. The figure shows why that misses: a zone sits inside the same region, and distance to the user is an edge question.
This is the figure the vocabulary needs, because the terms are defined by what contains what. A region, in Google's words, is an independent geographic area that consists of zones, and a region consists of three or more zones housed in three or more physical data centers. A zone is a deployment area for resources within a region, and Google says zones should be considered a single failure domain. So a copy of the application in each zone survives one zone failing. Outside the region, on the left, is the edge. Google operates a global network of peering points of presence, which means customer traffic can travel within the Google network until it is close to its destination. The point of presence is where the user's traffic enters Google's network — near the user, not inside the region. Hold that containment and two common errors disappear: adding a zone does not bring an application closer to a distant user, and a point of presence is not somewhere you run a copy of the application for failover.
Latency, bandwidth, and the cables under both
Latency is how long; bandwidth is how much
Figure. Two cards and a band beneath them. Latency: round-trip time, which grows with distance at roughly one millisecond per hundred kilometres. Bandwidth: how much data a connection carries, a separate property from latency. Beneath both: fiber optics and subsea cables, the physical links, where every extra hop between internet service providers can add latency and limit bandwidth.
Worked example (synthetic). Users three thousand kilometres from the region complain pages feel slow, though the link is nowhere near full. That is latency, not bandwidth — more capacity would not shorten the trip.
The last pair of terms is the one the exam most likes to swap. Latency is about time. Google's region-selection guidance measures it as round-trip time — the time it takes to send an internet protocol packet and to receive the acknowledgment — and gives a rule of thumb: estimate one millisecond of round-trip latency for every hundred kilometres traveled. It also says why it matters: latency is often the key consideration for region selection, because high user latency can lead to an inferior user experience. Bandwidth is about quantity — how much data a connection can carry at once. Google treats the two as separate properties in a single sentence: zones have high-bandwidth, low-latency network connections to other zones in the same region. A link can have plenty of one and too little of the other, and a slow page far from its region is usually a latency problem that more bandwidth will not fix. Beneath both is the physical layer the objective names. Google describes engineering capacity into fiber optics that can be underground or laid at the bottom of the ocean, and notes that subsea cables carry ninety-nine percent of international network traffic. The same explainer ties the two properties back to the path a request takes: on the public internet, internet service providers hand traffic to one another as quickly as they can, and between the many hops you may face higher latency and limited bandwidth capacity.
Google's global network and data centers
Choose where to run; let Google's network carry the rest
Choose locations to meet latency, availability and durability requirements
Traffic enters at a nearby point of presence and stays on Google's network
Every cable cross section has multiple cables and no single point of failure
Plan for losing a whole region: that is what a second region is for
Worked example (synthetic). A game studio launches in Europe and Asia. It picks a region near each player base; players' traffic enters Google's network at a nearby point of presence instead of crossing the public internet to reach it.
The fifth objective asks how Google Cloud supports a digital transformation with global infrastructure connected by a fast, reliable network — and the answer has three layers. The first is choice of place. Google Cloud infrastructure services are available in locations across North America, South America, Europe, Asia, the Middle East and Australia, and Google says you can choose where to locate your applications to meet your latency, availability and durability requirements. The second is the path. Because Google operates peering points of presence around the world, a user's traffic enters Google's network close to the user; Google says its global backbone reduces end-user latency by having interconnects close to you, and its Premium Tier keeps traffic on that private network backbone, requiring fewer hops between internet service providers. The third is redundancy in the cables themselves: fiber optic cables are part of the network backbone of the internet that links data centers together, and Google says it designs the network with extra capacity — each cross section has multiple cables and no single point of failure. Redundancy still leaves one decision to the customer. Google's own advice is that to protect against the loss of an entire region due to natural disaster, you should have a disaster recovery plan and know how to bring up your application if your primary region is lost. The network makes a second region cheap to reach; it does not choose one for you.
Two paths from a user to a region
Where traffic enters Google's network is a choice
Loading Diagram...
Figure 2 — Mermaid diagram
Figure: A flow chart with two paths from an internet user through the user's internet service provider to the region running the application. On the Premium Tier path, traffic enters Google's network at a point of presence near the user and travels on Google's private backbone. On the Standard Tier path, traffic crosses regular ISP and transit networks and enters Google's network at a point of presence near the region.
Worked example (synthetic). A media company's viewers are worldwide and its application runs in one region. The Premium Tier path is the one that keeps their traffic on Google's backbone for most of the distance.
This figure makes the fifth objective concrete, because Google sells the network as two tiers and the difference is exactly where traffic joins it. Google's summary is one sentence: Premium Tier delivers traffic on Google's premium backbone, while Standard Tier uses regular internet service provider, or ISP, networks. On the Premium path, traffic from an internet user enters Google's network through peering or transit networks at a Google point of presence that is as close as possible to the user, and from there it rides Google's backbone to the region. On the Standard path, traffic crosses ordinary networks first and enters Google's network at a point of presence as close as possible to the region instead. Google's own explainer names the two routing styles: on the public internet an ISP hands off traffic to a downstream ISP as quickly as it can, which is hot potato routing, while Premium Tier keeps traffic on its private network backbone and offloads it to ISPs at the last possible moment, when the data is closest to the end user. Both arrive. The difference is how much of the journey happens on Google's network — and that is the practical meaning of a global infrastructure connected by a fast, reliable network. It is not only that the data centers exist; it is that a user far away can be on Google's network almost from the start.
What this topic actually tests
Four discriminations, five objectives
Scalability or elasticity? only one comes back down. Cheaper up front or cheaper in total? TCO counts managing. Which constraint is in the scenario? it picks private, hybrid or multicloud. How long, or how much? latency is time, bandwidth is capacity — and zones sit inside a region while the edge does not.
Close the topic on the four discriminations the exam leans on. First, scalability against elasticity: both involve growth, and only elasticity scales back in when the load falls, which is where autoscaling reduces costs. Second, cheaper up front against cheaper in total: the move from capital to operational expenditure changes when you pay, and total cost of ownership still counts the cost of managing what you run. Third, the constraint in the scenario: control of data points to private, residency or gradual migration or edge latency points to hybrid, and surviving one provider's outage points to multicloud. Fourth, the network: latency is how long a round trip takes and grows with distance, bandwidth is how much a connection carries, and zones sit inside a region while the point of presence where users join Google's network does not. Topic three takes the next step — who manages which layer under infrastructure, platform and software as a service — and it reuses every one of these words.
Study guide for Cloud Digital Leader, Unit 1 · Topic 3. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks. Discuss the benefits and tradeoffs of using infrastructure as a service (IaaS); platform as a service (PaaS); and software as a service (SaaS).
Objectives, quoted from the exam guide:
Define IaaS, PaaS, and SaaS.
Compare and contrast the benefits and tradeoffs of IaaS, PaaS, and SaaS including total cost of ownership (TCO), flexibility, shared responsibilities, management level, and necessary staffing and technical expertise.
Determine which computing model (IaaS, PaaS, SaaS) applies to various business scenarios and use cases.
Describe the cloud shared responsibility model. Compare which responsibilities are the cloud provider’s, and which responsibilities are the customer’s for on-premises and cloud computing models (IaaS, PaaS, SaaS).
Cloud computing models and shared responsibility
Three models, one question: how much do you still manage?
What each model is — infrastructure, platform or software as a service. What each costs you in control, customization, staff and total cost. Which one a scenario calls for.Who secures what — and the two things that never change hands.
This topic's summary in the guide is to discuss the benefits and trade-offs of infrastructure as a service, platform as a service and software as a service — IaaS, PaaS and SaaS. Google gives the one idea that holds all four objectives together: every one of these terms refers to how you use the cloud and the degree of management you are responsible for. So read the topic as a single dial. At one end you manage almost everything and control almost everything; at the other the provider manages almost everything and you control very little. The objectives ask you to define the three settings on that dial, weigh what each one costs you, choose one for a scenario, and — because security follows management — say who is responsible for what at each setting. One thing to hold from the start: the dial does not reach zero at either end. There are responsibilities that never leave the customer, and we will draw them.
Infrastructure, platform and software as a service
What the provider delivers, and what you still bring
Model
What the provider delivers
What you still manage
Google Cloud example
IaaS
On-demand compute, storage, networking and virtualization
Operating system, middleware, virtual machines, apps and data
Compute Engine, Cloud Storage
PaaS
All the hardware and software to develop applications
Your code, your data and your applications
Cloud Run, App Engine
SaaS
The entire application stack, ready to use
Your data — the rest is managed for you
Google Workspace
Worked example (synthetic). A team says it "runs on PaaS" but patches the operating system on its own virtual machines every month. Whatever it calls itself, it is on IaaS: patching the operating system is the IaaS customer's job.
Google names three main cloud service models, and each is defined by what the provider delivers. Infrastructure as a service, or IaaS, delivers on-demand infrastructure resources — compute, storage, networking and virtualization. Customers no longer manage their own data center infrastructure, but they are responsible for the operating system, middleware, virtual machines, and any apps or data. Compute Engine and Cloud Storage are Google's examples. Platform as a service, or PaaS, delivers and manages all the hardware and software resources to develop applications. Customers still write the code and manage their data and applications, but the environment to build and deploy them is managed by the provider; Cloud Run and App Engine are the examples. Software as a service, or SaaS, provides the entire application stack — a complete cloud-based application, completely managed by the provider, including all updates, bug fixes and maintenance. Google Workspace is the example. And the phrase as a service itself has a meaning Google spells out: the service model is offered by a third party in the cloud. Read the third column top to bottom and the pattern is the whole objective: what you manage shrinks from an operating system, to your code, to your data.
Build a house, hire a builder, rent it furnished, move in
Google's own housing analogy, one step per model
Figure. Four cards following Google's housing analogy: on-premises is building the house yourself; IaaS is hiring a contractor, so you rent the hardware but manage the operating system, runtime, scale and data; PaaS is renting a furnished house, so you bring only your code; SaaS is moving into a finished house, so you look after only your data. A note places containers as a service between IaaS and PaaS.
Worked example (synthetic). Asked which model leaves a team responsible only for its own data, the graded answer is SaaS — the finished house — even though the team still pays for upkeep.
Google teaches these models with a housing analogy, and it is worth borrowing because each step removes one kind of work. On-premises is building the house from scratch: you own everything from the hardware to your applications and scaling. Infrastructure as a service, or IaaS, is hiring a contractor: you rent the hardware to run your application on, but you are responsible for managing the operating system, runtime, scale and all the data. Platform as a service, or PaaS, is renting a furnished house: you bring your own code and deploy it, and the server management and scaling are left to the cloud provider. Software as a service, or SaaS, is moving into a finished house: you pay to use a complete application that is managed, maintained and secured by the provider, and you remain responsible for taking care of your own data. Google also slots in containers as a service, or CaaS, between the first two settings: with containers you bring a containerized application, so you do not worry about the underlying operating system but still control scale and runtime. The exam names three models; knowing where CaaS sits simply stops Google Kubernetes Engine looking like a fourth.
What each model costs you, and buys you
Control and customization trade against management and staffing
IaaS: the most control over infrastructure, and hands-on maintenance
PaaS: provider maintains and secures the infrastructure; less control
SaaS: provider manages everything; little to no customization
TCO: managed services cut operational overhead that VMs quietly add
Worked example (synthetic). Two teams host the same web app. One runs it on virtual machines and spends a day a week patching; the other runs it on a managed platform and spends that day on features. The second team's hourly price may be higher and its total cost lower.
The second objective asks you to compare the three models on five dimensions: total cost of ownership, flexibility, shared responsibility, management level, and the staffing and expertise each one needs. Google's framing is that the difference comes down to the level of control and responsibility, and that the provider manages different elements of the computing stack depending on which model you choose. On flexibility and control, infrastructure as a service, or IaaS, gives the highest level of control over infrastructure — and in exchange requires hands-on configuration and maintenance, which is a staffing statement: someone skilled has to do it. Platform as a service, or PaaS, moves maintenance and securing of infrastructure to the provider, at the cost of less control over operations and more limited customization. Software as a service, or SaaS, has the provider manage and maintain everything from hardware to software, and offers little to no customization. On total cost of ownership, or TCO, Google's cost framework is explicit that the cheapest-looking option is not the cheapest one: virtual machines might look cost-effective, but when you consider the overhead to maintain, patch and scale them, the TCO can increase. Its advice is to choose managed services and serverless products whenever possible to reduce operational overhead and maintenance costs — the lower overhead lets your team focus on core activities, and serverless services like Cloud Run can offer greater business value. One cost runs the other way: for PaaS and SaaS, Google lists vendor lock-in as a possible issue.
The five dimensions, side by side
Read across a row to see one trade-off move
Dimension
IaaS
PaaS
SaaS
Control and flexibility
Highest over infrastructure
Less over operations
None over infrastructure
Customization
Full
More limited
Little to none
Management level
Hands-on configuration and maintenance
Provider maintains the infrastructure
Provider manages everything
Staffing and expertise
Operating-system and infrastructure skills
Application developers
Users; no expertise to build the app
TCO lever
Upfront capital becomes operating expense
Lower operational overhead
Lowest management overhead
Worked example (synthetic). A hospital has no developers and needs scheduling software next quarter. Read the staffing row: SaaS is the only column that does not assume expertise the hospital lacks.
Here are the five dimensions as a grid, and the reason to draw it is that each row moves in the same direction. Control and flexibility fall from left to right: IaaS has the highest level of control over infrastructure, PaaS less control over operations, and SaaS no control over any of the infrastructure or security controls. Customization falls with it, to little or none for SaaS. Management level falls the other way for you and rises for the provider: IaaS requires hands-on configuration and maintenance, with PaaS the provider is responsible for maintenance and securing infrastructure, and with SaaS the provider manages and maintains everything. Staffing follows management, and so does security work: Google lists being responsible for your own data security and recovery as an IaaS disadvantage. Google's shared-responsibility page makes the SaaS end concrete: you use SaaS when your enterprise does not have the internal expertise or business requirement to build the application itself. And the total-cost row is where people go wrong. One of the primary reasons businesses choose IaaS is to turn capital expenditure into operational expense, but each step right also removes operational overhead, which Google's framework counts in total cost of ownership. The columns are infrastructure as a service, platform as a service and software as a service, and TCO is total cost of ownership.
Which model does the scenario call for
Look for what the organization wants to stop doing
Lift-and-shift, or specific VMs, databases and network configs: IaaS
Building an app and focusing on development, not infrastructure: PaaS
No internal expertise to build it, but the work must get done: SaaS
Most estates are a mix — the models are not mutually exclusive
Worked example (synthetic). A retailer moves its warehouse system to Compute Engine unchanged, builds a new storefront on Cloud Run, and runs email on Google Workspace. That is one company on all three models at once — and nothing about it is unusual.
The third objective turns the definitions into a choice, and Google's shared-responsibility page states the three triggers almost word for word. You use infrastructure as a service, or IaaS, if you plan on migrating an existing on-premises workload to the cloud using lift-and-shift, or if you want to run your application on particular virtual machines, using specific databases or network configurations. You use platform as a service, or PaaS, if you are building an application such as a website and want to focus on development, not on the underlying infrastructure. And you use software as a service, or SaaS, when your enterprise does not have the internal expertise or business requirement to build the application itself, but does need the work done. A useful habit is to read a scenario for what the organization wants to stop doing: stop running a data center but keep the configuration — IaaS; stop managing servers — PaaS; stop building software — SaaS. The last point on the slide is the one the exam uses as a trap. Google says the three are not mutually exclusive: you can combine one with another, or use a mix of all three along with more traditional information technology (IT) infrastructure. An answer that forces a single model on a varied estate is usually wrong.
Three questions that pick the model
Ask them in order; the first yes decides
Loading Diagram...
Figure 1 — Mermaid diagram
Figure: A decision flow of three questions asked in order. If the organization needs a finished application and has no expertise to build one, choose SaaS, such as Google Workspace. Otherwise, if it is writing the application and wants to focus on the code, choose PaaS, such as Cloud Run or App Engine. Otherwise, if it is moving an existing workload as it is or needs specific virtual machines and configurations, choose IaaS, such as Compute Engine.
Worked example (synthetic). An item describes a firm that wants to rehost a licensed application on the exact operating system it runs today. The first two answers are no, the third is yes: IaaS.
Drawn as a decision, the third objective becomes three questions asked in order, from the most managed model to the least. First: does the organization need a finished application and lack the expertise to build one? Then it is software as a service, or SaaS, and Google Workspace is the example. If not: is it writing the application itself and wanting to focus on the code rather than the infrastructure? Then platform as a service, or PaaS — Cloud Run or App Engine. If not: is it moving an existing workload unchanged, or does it need particular virtual machines, databases or network configurations? Then infrastructure as a service, or IaaS, and Compute Engine. The order matters because it matches Google's advice elsewhere to prefer managed services whenever possible: you only fall through to the model where you manage the most when the more managed ones cannot do the job.
The shared responsibility model
Two things never change hands, whichever model you pick
Figure. Two stacked zones. Above, always the customer's under every model: your data and access policies. Below, always Google Cloud's once the workload runs on it: the underlying infrastructure, the underlying network and physical security. A callout notes that everything between the two moves with the model chosen.
Worked example (synthetic). A company moves its email to a SaaS product and assumes security is now entirely the provider's. It still owns who has access and what data goes in — those two are the customer's under every model.
The fourth objective is the shared responsibility model, which Google defines as describing the tasks you have when it comes to security in the cloud, and how those tasks differ from the cloud provider's. Before comparing models, fix the two things that never move, because every question on this objective leans on them. Google's page says it directly: the cloud provider always remains responsible for the underlying network and infrastructure, and customers always remain responsible for their access policies and data. That is what the figure draws. Underneath, under every model, is Google Cloud's: the infrastructure, the network, and — Google's own words for the infrastructure-as-a-service case — physical security. On top, under every model including software as a service, is yours: your data and who can access it. The callout is the rest of the objective. Everything between those two lines moves as you choose a more or less managed model, and the next slide follows it.
Where the line sits, model by model
On-premises you hold everything; each step hands some over
Model
The provider's
The customer's
On-premises
Nothing
Everything, from the hardware to the applications and scaling
IaaS
Underlying infrastructure and physical security
The bulk of security: data, applications, network controls, operating system, user access
PaaS
More controls than in IaaS; application-level controls and IAM shared
Data security and client protection
SaaS
The bulk of security responsibilities
Access controls, and the data stored in the application
Worked example (synthetic). An item asks who patches the operating system of a Compute Engine virtual machine. The IaaS row answers it: the operating system is the customer's.
Now follow the line as it moves. On-premises there is no provider in the picture: you own everything from the hardware to your applications and scaling. In infrastructure as a service, or IaaS, Google says the bulk of the security responsibilities are yours, and its own are focused on the underlying infrastructure and physical security; its IaaS page lists what you secure as your data, applications, virtual network controls, operating system and user access. In platform as a service, or PaaS, Google is responsible for more controls than in IaaS, and you share responsibility for application-level controls and identity and access management (IAM); you remain responsible for your data security and client protection. In software as a service, or SaaS, Google owns the bulk of the security responsibilities, and you remain responsible for your access controls and the data you choose to store in the application. Functions — function as a service, or FaaS — sit with SaaS: Google says FaaS has a similar shared responsibility list. Read the customer column bottom to top and it never empties; it only shrinks to the two things on the previous slide.
Shared responsibility, and Google's shared fate
A division of tasks, and a partnership on top of it
Figure. Two cards. Shared responsibility divides security tasks between provider and customer, and is hard to apply because many breaches come from misconfiguration and most estates mix several service types. Shared fate, Google's extension, treats the relationship as an ongoing partnership to improve security, starting with a trusted cloud platform.
Worked example (synthetic). A bank on IaaS, PaaS and SaaS at once finds three different responsibility lists. Shared responsibility tells it which list applies where; shared fate is Google's commitment to help it configure each one well.
Google adds one idea on top of the model, and it is worth knowing by name. The shared responsibility model is a division of tasks, and Google is candid that applying it is hard. Many cloud security breaches, it says, are the direct result of misconfiguration, and though dividing items by cloud service is helpful, many enterprises have workloads that require multiple cloud service types — so a single company is reading several responsibility lists at once. Google's answer is what it calls shared fate. Shared fate builds on the shared responsibility model because it views the relationship between cloud provider and customer as an ongoing partnership to improve security, and it includes Google building and operating a trusted cloud platform for your workloads. Shared fate does not move the line on the previous two slides — your data and access policies are still yours. It changes how much help you get on your side of it.
What this topic actually tests
One dial, and the two ends it never reaches
What do you still manage? operating system (IaaS), code (PaaS), data (SaaS). What does the scenario want to stop doing? that picks the model — and mixing models is normal. Who secures what? the line moves with the model; your data and access policies never leave you, and the infrastructure never leaves Google.
Close on the three habits this topic rewards. First, when a model is named, ask what the customer still manages: the operating system and everything above it under infrastructure as a service, the code and data under platform as a service, and the data alone under software as a service. Second, when a scenario is described, ask what the organization wants to stop doing — running hardware, managing servers, or building software — and remember that most real estates use more than one model at once. Third, when security is in the question, remember the line moves with the model but has two fixed ends: the provider always remains responsible for the underlying network and infrastructure, and customers always remain responsible for their access policies and data. Unit 5 returns to that line in far more detail; this topic is where it is first drawn.
Unit 2 roadmap — Exploring Data Transformation with Google Cloud
Unit 2 at a glance
Exam weight (Google, approx.)
~16%
Topics
3
Objectives
15
Bank questions
78
Flashcards
75
Seats per mock
9
Every objective below is quoted from Google's Cloud Digital Leader exam guide, retrieved 2026-09-22. Under each topic is that topic's own summary of what the exam actually tests, taken from its lecture.
Topic 1 — The Value of Data
CDL-U2.T1 · 6 objectives · lecture deck of 13 slides
Explain how data generates business insights, drives decision making, and creates new value.
Differentiate between basic data management concepts, in particular: databases; data warehouses; data lakes.
Explain how organizations can create value by using their current data, collecting new data, and sourcing data externally.
Describe how the cloud unlocks business value from all types of data, including structured data and previously untapped unstructured data.
Discuss the main data value chain concepts and terms.
Explain how data governance is essential to a successful data journey.
What this topic actually tests.Which store? database for related records, warehouse for known reporting, lake for raw and unstructured data. Which source? integrate what you hold, stream what you lack, subscribe to what others hold. Which stage? value is realized at action, not ingestion. What does governance do? controls that make greater access safe.
Topic 2 — Google Cloud Data Management Solutions
CDL-U2.T2 · 5 objectives · lecture deck of 12 slides
Differentiate between Google Cloud data management options including data type and common business use case, including: Cloud Storage; Cloud Spanner; Cloud SQL; Cloud Bigtable; BigQuery; Firestore.
Define key data management concepts and terms, including: relational; non-relational; object storage; structured query language (SQL); NoSQL.
Describe the benefits of using BigQuery as a serverless, managed data warehouse and analytics engine that can be used in a multicloud environment.
Differentiate between storage classes in Cloud Storage regarding cost and frequency of access, including: Standard; Nearline; Coldline; Archive.
Describe the ways that an organization can migrate or modernize their current database in the cloud.
What this topic actually tests.Which shape is the data? objects, relational, documents, keyed series, or analysis. Relational at what scale? Cloud SQL, or Spanner for global consistency. Serve or analyze? Bigtable serves; BigQuery analyzes. How often is it read? monthly, quarterly, yearly — and colder costs more to read, never offline. What changes in the move? the engine, and separately the downtime.
Topic 3 — Making Data Useful and Accessible
CDL-U2.T3 · 4 objectives · lecture deck of 12 slides
Describe how Looker democratizes access to data by empowering individuals to self-serve business intelligence and create insights.
Discuss the value of analyzing and visualizing data from BigQuery in Looker to create real-time reports, dashboards, and integrating data into workflows.
Describe how streaming analytics in real time makes data more useful and generates business value.
Describe the main Google Cloud products that modernize data pipelines, including Pub/Sub and Dataflow.
What this topic actually tests.Who writes the SQL? the LookML model and the SQL generator, not the business user. Is a dashboard enough? alerts and deliveries put data into workflows. Batch or stream? ask how fast the data goes stale. Which product? Pub/Sub ingests, Dataflow transforms, BigQuery stores, Looker presents.
Study guide for Cloud Digital Leader, Unit 2 · Topic 1. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks. Describe the intrinsic role that data plays in an organizations’ digital transformation.
Objectives, quoted from the exam guide:
Explain how data generates business insights, drives decision making, and creates new value.
Differentiate between basic data management concepts, in particular: databases; data warehouses; data lakes.
Explain how organizations can create value by using their current data, collecting new data, and sourcing data externally.
Describe how the cloud unlocks business value from all types of data, including structured data and previously untapped unstructured data.
Discuss the main data value chain concepts and terms.
Explain how data governance is essential to a successful data journey.
The value of data
Where value comes from, where data lives, and what keeps it usable
Where value comes from — insight, decisions, new products; current, new and external data; structured and unstructured. Where data lives — database, warehouse, lake, and the stages between them. What keeps it usable — governance, without which none of the rest is worth trusting.
Unit two is about data, and this first topic asks one thing: what role data plays in an organization's digital transformation. Its six objectives answer in three parts. The first is where value comes from — how data turns into insight, how insight turns into decisions and new products, and which data an organization can draw on: what it already has, what it can start collecting, and what it can source from outside. The second is where data lives — the difference between a database, a data warehouse and a data lake, the difference between structured and unstructured data, and the stages data passes through on its way from being collected to being used. The third is what keeps any of it usable, which is governance. Google puts that third part bluntly, and it is worth hearing first: the value of data is only realized when it is trustworthy, discoverable and governed. Everything before governance in this topic is potential value; governance is what lets an organization actually collect it.
From data to insight to decision to new value
Data is worth nothing until it changes what someone does
Analytics now sits at the center of all core business activities
Business intelligence: collect and analyze data for strategic and daily decisions
More visibility gives better decisions and uncovers growth opportunities
Insight accelerates new products, features and updates
Only the right data, effectively analyzed, yields insight
Worked example (synthetic). A grocer's loyalty data shows that one store's customers buy baby products on Saturday mornings. The data was always there; the value appears when the Saturday delivery schedule changes because of it.
The first objective asks how data generates business insights, drives decision making and creates new value — three steps, and the exam tests each. Start with where data now sits. Google says today's data analytics activities have transformed to the center of all core business activities, including revenue generation, cost containment, improving operations and enhancing customer experiences. Insight is the first step. Business intelligence, or BI, is the process of using people and technologies to collect and analyze data for an organization's strategic and daily decision-making. Decision is the second. Google's statement of the whole idea is that the more visibility you have into anything, the more effectively you can gain insights to make better decisions, uncover growth opportunities and improve your business model. New value is the third, and it is more than doing the old things better: these insights can guide and accelerate the planning, production and launch of new products, features and updates. Insight also points inward — analytics helps you find where to reduce costs, save time and increase efficiency. And there is a condition on all of it: big data must contain the right data and then be effectively analyzed to yield insights that drive decisions. Volume alone produces nothing.
Four places the value lands
Same data, four different kinds of return
Figure. Four cards naming where the value of data lands: better decisions, new products, leaner operations, and better customer experience.
Worked example (synthetic). Asked what a bank gains by analyzing call-center transcripts beside account data, the graded answer is customer experience — the gain comes from combining the two, not from either alone.
This figure turns the first objective into four landing places, because the exam describes a business outcome and asks you to recognize data behind it. Better decisions come from visibility: Google's central claim for big data is that more visibility yields insight to make better decisions and uncover growth opportunities. New products come from insight that can guide and accelerate the planning, production and launch of new products, features and updates. Leaner operations come from generating insights that show where you can reduce costs, save time and increase efficiency. And better customer experience has a specific mechanism worth remembering: combining and analyzing structured data sources together with unstructured ones provides more useful insights for consumer understanding and personalization. That last card previews objective four — some value only exists when two kinds of data are read together.
Database, data warehouse, data lake
Three stores, three different jobs
Store
What it holds
What it is for
Relational database
Information structured in tables, rows and columns, with predefined relationships
Seeing how records relate to each other
Data warehouse
Structured and semi-structured data from multiple sources, current and historical
Analysis and reporting; a long-range view over time
Data lake
Structured, semistructured and unstructured data in its native format
Diverse, raw data for machine learning and advanced analytics
Worked example (synthetic). A retailer keeps each order in a database, loads a year of orders into a warehouse to report seasonal trends, and drops product photos and review text into a lake to train a recommendation model.
The second objective asks you to differentiate three data management concepts, and the fastest way is to ask what each one holds and what it is for. A relational database is a way of structuring information in tables, rows and columns; Google adds that it organizes data in predefined relationships, making it easy to see how different data structures relate to each other. A data warehouse is an enterprise system for the analysis and reporting of structured and semi-structured data from multiple sources, such as point-of-sale transactions, marketing automation and customer relationship management. It can store both current and historical data in one place and is designed to give a long-range view of data over time, which is why Google calls it a primary component of business intelligence. A data lake is a centralized, scalable and secure repository designed to store, process and analyze large amounts of structured, semistructured and unstructured data in its native format. The contrast Google draws between the last two is precise: a traditional data warehouse is optimized for repeatable business reporting and structured analysis, while a data lake excels at handling the diverse, raw data required for machine learning.
How data reaches each store
A schema before loading, or none at all
Loading Diagram...
Figure 1 — Mermaid diagram
Figure: A flow chart. Many source systems feed two paths. One passes through extract, transform and load into a data warehouse with a rigid schema, loaded in batches, which serves repeatable reporting and business intelligence. The other goes straight into a data lake in native format with no pre-defined schema, which serves machine learning on raw data.
Worked example (synthetic). A team asks why its warehouse cannot answer a question nobody planned for. The figure's answer: the warehouse's schema was decided before loading, and data outside it was never loaded.
Drawn out, the warehouse and the lake differ most in what happens before data arrives. On the warehouse path, data is shaped first. Google describes extract, transform and load, or ETL, as a traditionally accepted way to combine data from multiple systems into a single database, data store, data warehouse or data lake. A traditional warehouse is typically designed to capture a subset of data in batches and store it based on rigid schemas — which Google says makes it unsuitable for spontaneous queries or real-time analysis. That rigidity is also its strength: it serves repeatable reporting. On the lake path, nothing is decided first. A data lake lets enterprises ingest any data from any source, on-premises, cloud or edge, without the constraints of pre-defined schemas. So the useful question in an exam item is not which store is better, but whether the questions are known in advance. Known questions favor the warehouse; unknown ones, and machine learning on raw data, favor the lake.
Current data, new data, external data
Three places an organization finds value it does not yet use
Current data: combine what is already held into one unified view
New data: collect continuously, at any speed and volume
External data: discover third-party datasets and combine them with your own
Cloud warehouses collect from internal AND external sources
Worked example (synthetic). A logistics firm unifies its dispatch and billing data, starts streaming truck sensor readings, and subscribes to a third-party weather dataset — three sources, three different kinds of value.
The third objective asks how organizations create value from three sources: the data they already have, data they start collecting, and data they source externally. Current data first. Much of it sits in separate systems, and the value is in combining it: Google defines data integration as discovering, moving and combining this data into a unified view to drive insights and power artificial intelligence (AI) driven analytics. New data second. Big data lets you integrate automated, real-time data streaming with advanced analytics to continuously collect data, find new insights and discover new opportunities for growth and value — and a data lake lets enterprises ingest data at any speed and volume. External data third. Google's own mechanism for it is BigQuery sharing, formerly Analytics Hub, a data exchange platform to securely share, discover and access data across organizational boundaries without replicating it. With it you can discover curated third-party and Google datasets, and combine them with your internal data to augment analytics and machine learning. The cloud data warehouse joins all three: like a traditional warehouse, it collects, integrates and stores data from internal and external data sources.
Which source, which mechanism
Name the source in the scenario, then the move that unlocks it
Source
The move
What Google says it gives
Data already held, in separate systems
Integrate it
A unified view that drives insights
Data not yet collected
Stream it continuously
New insights and new opportunities for growth
Data held by someone else
Subscribe through a data exchange
Third-party datasets combined with internal data, without replicating it
Worked example (synthetic). An item says a retailer wants foot-traffic data it has never owned. The graded move is external sourcing through a data exchange — not a new collection program it would have to build.
The exam will describe an organization and ask where its next value comes from, so rehearse the three sources as moves. Data already held in separate systems calls for integration — discovering, moving and combining it into a unified view that drives insights. Data not yet collected calls for streaming, which Google describes as integrating automated, real-time data streaming with analytics to continuously collect data and discover new opportunities for growth and value. Data held by someone else calls for sourcing it, and Google's mechanism is a data exchange: BigQuery sharing lets you discover curated third-party and Google datasets and combine them with your internal data, and it does so across organizational boundaries without replicating the data. The trap in these items is to answer every one with collect more. Sometimes the value is already inside the organization, and sometimes it belongs to somebody else.
Structured and unstructured data
Most new data has no fixed schema — the cloud can keep it anyway
Figure. Three cards for structured, semi-structured and unstructured data with examples of each, above a band explaining that traditional warehouses discarded data to save storage while a cloud warehouse with a cloud data lake keeps unstructured data, and combining the two kinds gives new insight.
Worked example (synthetic). An insurer's claims table says what was paid; the adjusters' photos and notes say why. Only when the cloud keeps both can the insurer find which damage patterns predict the largest claims.
The fourth objective asks how the cloud unlocks value from all types of data, including previously untapped unstructured data — and the word untapped is the point. Google separates the types clearly. More traditional structured data, such as data in spreadsheets or relational databases, is now supplemented by unstructured text, images, audio and video files, or semi-structured formats like sensor data that cannot be organized in a fixed data schema. Why was it untapped? Partly scale: big data sets are so complex in volume, velocity and variety that traditional data management systems cannot store, process and analyze them. And partly economics: in a traditional warehouse, storage is typically limited compared to compute, so data is transformed quickly and then discarded to keep storage space free, and an on-premises warehouse means buying your own hardware and software, making it expensive to scale. The cloud changes both. A cloud data warehouse can be used with a cloud data lake to collect and store unstructured data. And the payoff is in the combination: analyzing structured and unstructured sources together provides more useful insights for consumer understanding and personalization.
The data value chain
The stages data passes through, and the word for each
Loading Diagram...
Figure 2 — Mermaid diagram
Figure: A left-to-right flow of the data life cycle: acquisition and ingestion, processing by extract transform and load, storage in a warehouse or lake, analytics and AI, action in real or near-real time, and secure disposal. A governance box connects to the first and last stages to show it spans the whole chain.
Worked example (synthetic). A team measures success by how much data it ingests. The chain shows why that is the wrong stage to measure: value is realized at action, three stages later.
The fifth objective asks for the main data value chain concepts and terms. The value chain is the guide's phrase rather than one Google's pages use, so this deck teaches it with the life cycle Google does define. Data governance, Google says, is a principled approach to managing data during its life cycle, from acquisition and ingestion to artificial intelligence (AI) data analytics and secure disposal. Read the stages in order. Acquisition and ingestion bring data in. Processing makes it usable — Google describes extract, transform and load, or ETL, as the end-to-end process by which a company takes its full breadth of data and gets it to a state where it is actually useful for business purposes. Storage holds it, in a warehouse or a lake. Analytics turns it into insight. Action is where the value is realized: Google says cloud data warehouses should support streaming use cases to activate on data in real or near-real time. And the chain ends deliberately, with secure disposal. Governance is drawn beside the chain, not inside it, because it applies to every stage.
The vocabulary: six Vs of big data
Three original Vs, three that decide whether data is worth it
Term
What it describes
Volume
How much data there is — the most common characteristic of big data
Velocity
The speed at which data is generated
Variety
Data from many sources: structured, unstructured or semi-structured
Veracity
Quality and accuracy — the higher the veracity, the more trustworthy
Variability
Meaning that changes over time, leading to inconsistency
Value
Whether the data you collect is worth anything to the business
Worked example (synthetic). An item describes sensor data that arrives every second but is often wrong. Two Vs are in play — velocity is high, veracity is low — and the graded answer names the one the question asks about.
The rest of the value chain vocabulary the exam uses comes from Google's big data page, and it is organized as Vs. Google says big data definitions may vary slightly, but it will always be described in terms of volume, velocity and variety. Volume is the amount — the most common characteristic associated with big data. Velocity refers to the speed at which data is generated. Variety means data can come from many sources and be structured, unstructured or semi-structured. Then three more, which Google says are often mentioned in relation to harnessing the power of big data: veracity, variability and value. Veracity is quality — the higher the veracity of the data, the more trustworthy it is. Variability is meaning that shifts: the meaning of collected data is constantly changing, which can lead to inconsistency over time. And value is the question the whole chain exists to answer: it is essential to determine the business value of the data you collect. The first three describe data; the last three decide whether it is worth keeping.
Why governance decides the journey
Governance is what makes more access safe, not what prevents it
Secure, private, accurate, available and usable — for people and for machine learning
Defines who can access sensitive information
Lets data be democratized without security or compliance breaches
Controls that allow GREATER access, because the access is controlled
Worked example (synthetic). A bank wants every analyst to query customer data. Without governance the answer is no; with column-level controls and defined access, the answer becomes yes for most columns.
The last objective asks why data governance is essential to a successful data journey, and Google's answer turns on a word that sounds like a contradiction: access. Data governance is everything you do to ensure data is secure, private, accurate, available and usable, for human analysis, machine learning and building agents. It means setting internal standards for how data is gathered and processed, and it involves defining who can access sensitive information and ensuring that the democratization of data does not lead to security risks or compliance breaches. Here is the part candidates get backwards. Governance is not mainly a brake. Google says data governance allows setting and enforcing controls that allow greater access to data, gaining the security and privacy from the controls on data. The controls are what make it safe to open the data up. That is why governance is essential rather than optional: effective governance lets organizations go from raw data to action faster while maintaining strict security and compliance standards, and without it, the value of data is simply not realized — Google says that value comes only when data is trustworthy, discoverable and governed.
What governance actually does
Three uses, each with its own word
Use
What it means
Data stewardship
Accountability for the data and the processes that ensure its proper use, given to data stewards
Data quality
Six dimensions: accuracy, completeness, consistency, timeliness, validity, uniqueness
Compliance
Data is safe, secure, private, usable, and compliant with internal and external policies
Worked example (synthetic). A dashboard shows two different revenue totals for the same month. That is a data quality failure on the consistency dimension — and deciding who fixes it is a stewardship question.
Google lists the common uses of governance, and each carries a term the exam uses. Stewardship is about people: data governance often means giving accountability and responsibility for both the data itself and the processes that ensure its proper use to data stewards. Quality is about fitness for use, and Google names six dimensions it is judged on: accuracy, completeness, consistency, timeliness, validity and uniqueness. It matters because, as Google's big data page puts it, data quality directly impacts the quality of decision-making, data analytics and planning strategies — so bad data does not just sit there, it produces bad decisions. And compliance is about obligations on both sides of the organization's boundary: governance is necessary to assure that data is safe, secure, private, usable, and in compliance with both internal and external data policies. When an exam item describes a data problem, ask which of the three it is — a person question, a quality question, or a policy question.
What this topic actually tests
Four discriminations, six objectives
Which store? database for related records, warehouse for known reporting, lake for raw and unstructured data. Which source? integrate what you hold, stream what you lack, subscribe to what others hold. Which stage? value is realized at action, not ingestion. What does governance do? controls that make greater access safe.
Close the topic on four discriminations. First, which store: a relational database structures related records in tables, a warehouse serves known, repeatable reporting over current and historical data, and a lake keeps raw data of every type in its native format, which is what machine learning needs. Second, which source of value: integrate the data you already hold, stream the data you are not yet collecting, and subscribe to data someone else holds through a data exchange. Third, which stage of the chain: data is acquired, processed, stored and analyzed, but its value is realized when someone acts on it — which is why Google wants warehouses that can activate on data in near-real time. Fourth, what governance is for: not to lock data away, but to set the controls that make greater access safe, because the value of data is only realized when it is trustworthy, discoverable and governed. Topic two turns these ideas into Google Cloud's products for managing data.
Study guide for Cloud Digital Leader, Unit 2 · Topic 2. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks. Determine which Google Cloud data management products are applicable to different business use cases.
Objectives, quoted from the exam guide:
Differentiate between Google Cloud data management options including data type and common business use case, including: Cloud Storage; Cloud Spanner; Cloud SQL; Cloud Bigtable; BigQuery; Firestore.
Define key data management concepts and terms, including: relational; non-relational; object storage; structured query language (SQL); NoSQL.
Describe the benefits of using BigQuery as a serverless, managed data warehouse and analytics engine that can be used in a multicloud environment.
Differentiate between storage classes in Cloud Storage regarding cost and frequency of access, including: Standard; Nearline; Coldline; Archive.
Describe the ways that an organization can migrate or modernize their current database in the cloud.
Google Cloud data management solutions
Six products, one vocabulary, four storage classes, one move
Which product fits the data? Cloud Storage, Cloud SQL, Spanner, Bigtable, BigQuery or Firestore. What do the words mean? relational, non-relational, object storage, SQL and NoSQL. Why BigQuery? serverless, managed, and able to reach data in other clouds. How cold is the data? Standard, Nearline, Coldline or Archive. How does an existing database get here? migrate as-is, or modernize on the way.
The guide's summary for this topic is a single instruction: determine which Google Cloud data management products are applicable to different business use cases. That is a matching exercise, and every objective supplies either the options or the vocabulary to match with. The first objective names six products — Cloud Storage, Cloud SQL, Spanner, Bigtable, BigQuery and Firestore — and asks what kind of data and which use case each is for. The guide still writes two of them with an older prefix, Cloud Spanner and Cloud Bigtable; Google's documentation now calls them simply Spanner and Bigtable. The second objective is the vocabulary that makes the matching possible: relational and non-relational, object storage, and the query languages. The third singles out BigQuery. The fourth is about Cloud Storage's four storage classes, which are a trade between how often data is read and what it costs to keep. And the fifth asks how an organization gets its existing database into the cloud. Five objectives, one question: what shape is the data, and what will be done with it?
Six products, six kinds of data
What each one stores, and the use case it is for
Product
Kind of data
Common business use
Cloud Storage
Objects in buckets — files of any kind
Images, video, backups, website content
Cloud SQL
Relational, fully managed MySQL, PostgreSQL and SQL Server
An existing application's database, without running the server
Spanner
Relational with graph, key-value and search, at global scale
Transactions that must stay consistent worldwide
Bigtable
Wide-column, single-keyed data at very large scale
Time series, Internet of Things readings, financial ticks
BigQuery
Analytical: a data warehouse with analytic tools
Reporting and analysis across the whole business
Firestore
Documents, in a serverless document database
Mobile and web apps, with optional Firebase integration
Worked example (synthetic). A streaming service stores its film files, its subscriber accounts, its viewing events and its quarterly analysis. That is four products — Cloud Storage, Cloud SQL, Bigtable and BigQuery — because it is four kinds of data.
Here are the six, read across the two columns the objective asks about: data type and common business use. Cloud Storage is object storage — Google calls it a scalable and managed storage service that lets you store data as objects in containers called buckets, and its own example is a photos bucket for the image files an app generates. Cloud SQL is a fully managed relational database service for MySQL, PostgreSQL and SQL Server, and it handles backups, high availability and failover, and maintenance and updates for you: the familiar database engines, without running the server. Spanner is also relational, but it brings together relational, graph, key-value and search, and offers transactional consistency at global scale. Bigtable is the wide-column one — ideal for storing large amounts of single-keyed data with low latency, and Google lists time-series data, Internet of Things data and financial data among its uses. BigQuery is the analytical one: a fully managed data platform that combines a cloud-based data warehouse and powerful analytic tools. And Firestore is a fully managed document database, which also supports, but does not require, integration with the Firebase mobile and web app development platform. Two pairs are worth separating now. Cloud SQL and Spanner are both relational; the difference is global scale. Bigtable and BigQuery both hold enormous datasets; the difference is that Bigtable serves an application low-latency reads and BigQuery answers analytical questions.
Choosing a product: two questions
What shape is the data, and what is done with it?
Loading Diagram...
Figure 1 — Mermaid diagram
Figure: A decision tree. The first question is the shape of the data: files and media lead to Cloud Storage; documents for an application lead to Firestore; huge single-keyed series lead to Bigtable; data held for analysis leads to BigQuery; tables and relations lead to a second question — must transactions stay consistent at global scale — where no leads to Cloud SQL and yes leads to Spanner.
Worked example (synthetic). An item describes a bank whose account ledger must stay correct across three continents at once. The relational branch, then the global-consistency question, gives Spanner; Cloud SQL is the distractor.
Exam items on this objective never ask you to recite a product's description — they describe a business and its data and ask which product fits. So rehearse it as a decision. The first question is the shape of the data. Files and media — images, video, backups — are objects, and objects go to Cloud Storage. Documents that an application reads and writes go to Firestore, a document database that auto scales to match your load. Very large volumes of single-keyed data — readings over time, events keyed by device — go to Bigtable, which supports high read and write throughput at low latency. Data kept for analysis, whatever its source, goes to BigQuery. Only the relational branch needs a second question: must transactions stay consistent at global scale? If not, Cloud SQL, the fully managed service for MySQL, PostgreSQL and SQL Server. If so, Spanner, which offers transactional consistency at global scale with automatic, synchronous replication for high availability. The tree is deliberately short. If you find yourself needing a third question, re-read the stem for the shape of the data.
The words underneath the choice
Relational, non-relational, object storage, SQL and NoSQL
Relational: tables of rows and columns, linked by primary and foreign keys
SQL: the query language relational databases share
NoSQL (not only SQL): non-relational, non-tabular, flexible schema
Object storage: data plus metadata plus a unique identifier, in a flat space
Worked example (synthetic). An online shop keeps customers and orders in two tables joined by customer ID — relational. Its product reviews arrive as documents of varying fields — a NoSQL fit. Its product photos are objects.
The second objective is five terms, and they sort into two families: how structured data is organized, and how everything else is stored. A relational database, in Google's words, organizes data in predefined relationships where data is stored in one or more tables of columns and rows. The relationships are made with keys: every table has a primary key, a unique identifier of a row, and a foreign key in one table refers to a primary key in another — which is how a customer is joined to their orders. Structured query language, or SQL, is the language these databases share, and Google notes it is easy to run complex queries using SQL, even for non-technical users. The opposite family is NoSQL, which Google expands as not only SQL: non-relational databases that store data in a non-tabular format, with a flexible schema model that supports documents, key-value, wide columns and graphs. There is a cost to that flexibility — NoSQL has no lingua franca like SQL, and each database may have its own query language. Google names Bigtable, Memorystore and Firestore as its NoSQL databases. The last term is object storage: an architecture for unstructured data that sections it into objects and stores them in a structurally flat data environment, each object carrying the data, metadata and a unique identifier.
Three ways to hold data
Tables, flexible records, and objects
Figure. Three cards. Relational: tables of rows and columns in predefined relationships, queried with SQL — Cloud SQL and Spanner. Non-relational or NoSQL: non-tabular with a flexible schema — documents, key-value, wide columns and graphs — Firestore and Bigtable. Object storage: data plus metadata plus a unique identifier in a flat space for unstructured data — Cloud Storage.
Worked example (synthetic). Asked where email attachments, audio files and web pages belong, the graded answer is object storage — Google lists exactly these as unstructured data that does not fit easily into traditional databases.
Side by side, the three families differ in one word each. Relational data sits in predefined relationships — the schema comes first — and Cloud SQL and Spanner are Google Cloud's relational products. Non-relational data has a flexible schema: Google describes NoSQL — not only SQL — databases as supporting documents, key-value, wide columns and graphs, and names Firestore and Bigtable among its NoSQL databases. Object storage is flat: each object carries its data, metadata and a unique identifier, and there is no table and no hierarchy to fit it into. That is why it suits unstructured data, which Google lists as email, media and audio files, web pages, sensor data and other digital content that does not fit easily into traditional databases. One caution on the middle card: BigQuery is absent on purpose. It is a warehouse for analysis rather than a place an application keeps its records, and it is the subject of the next objective.
BigQuery: serverless, managed, multicloud
A warehouse you query, not a server you run
Serverless: no resources to provision or manually scale
Managed: Google's engineering team handles updates and maintenance
Storage and compute are separate layers, each scaling on its own
Multicloud: BigQuery Omni analyzes data held in other public clouds
Worked example (synthetic). A retailer keeps sales data in Google Cloud and its web logs in another provider's object store. With BigQuery Omni it queries both from one interface instead of copying the logs across first.
The third objective asks for BigQuery's benefits as a serverless, managed data warehouse and analytics engine that can be used in a multicloud environment — and each of those words has a mechanism behind it. Serverless: Google says BigQuery's serverless architecture lets you answer your organization's biggest questions with zero infrastructure management, so you do not need to provision or manually scale resources and can focus on delivering value instead of traditional database management tasks. Managed: it is a fully managed serverless data warehouse in which the BigQuery engineering team handles updates and maintenance. Under both sits the architecture. BigQuery has a storage layer that ingests, stores and optimizes data, and a compute layer that provides analytics, and their separation lets each allocate resources without impacting the performance or availability of the other. Google contrasts this with legacy databases, which usually have to share resources between read and write operations and analytical operations, slowing queries. Multicloud: many organizations store data in multiple public clouds, and it ends up siloed. With BigQuery Omni, you can run BigQuery analytics on data stored in Amazon Simple Storage Service or Azure Blob Storage, using BigLake tables.
Why separation matters
Storage and compute scale apart — and the data need not move
Loading Diagram...
Figure 2 — Mermaid diagram
Figure: A flow from analysts writing SQL and Python to the BigQuery compute layer. The compute layer reads the BigQuery storage layer, and through BigQuery Omni and BigLake tables it also reaches data held in another public cloud.
Worked example (synthetic). A quarter-end reporting rush multiplies the number of queries but adds almost no data. Because compute and storage are separate layers, the query side absorbs the rush without the storage side changing.
Drawn out, the benefits stop being adjectives. On the left, analysts use languages like SQL and Python. Their queries run in the compute layer, which reads the storage layer — and because the two are separate, each can allocate resources without impacting the performance or availability of the other. That is the practical reason BigQuery can be serverless: nobody sizes a server that must hold both the data and the peak query load. The lower branch is the multicloud one. The compute layer can also reach data that never moved into Google Cloud, because BigQuery Omni runs BigQuery analytics on data stored in Amazon Simple Storage Service or Azure Blob Storage, using BigLake tables. So the answer to siloed data across clouds is not always to copy it; sometimes it is to query it where it sits.
The four Cloud Storage classes
Read less often, pay less to keep — and more to read
Class
Access pattern it is ideal for
Minimum storage duration
Typical use
Standard
Frequently accessed ("hot") data, or data kept briefly
None
Serving website content, active data
Nearline
Read or modified about once a month or less
30 days
Infrequently accessed data
Coldline
Read or modified at most once a quarter
90 days
Very infrequently accessed data
Archive
Accessed less than once a year
365 days
Archiving, online backup, disaster recovery
Worked example (synthetic). A law firm keeps closed case files it expects to open perhaps once in several years. Archive is the class built for that; Standard would charge it the hot-data rate for data nobody reads.
The fourth objective is Cloud Storage's storage classes, and the exam asks about them in terms of cost and frequency of access — so learn them as a ladder of access frequency. Standard storage is best for data that is frequently accessed — hot data — as well as data stored for only brief periods. Nearline storage is ideal for data you plan to read or modify on average once per month or less. Coldline storage is ideal for data you plan to read or modify at most once a quarter. Archive storage is the best choice for data you plan to access less than once a year, and Google describes it as the lowest-cost, highly durable storage service for data archiving, online backup and disaster recovery. Each step down the ladder carries a longer minimum storage duration: thirty days for Nearline, ninety for Coldline, three hundred and sixty-five for Archive. That minimum is part of the exam's cost question — a class chosen for data that will be deleted next week is a class chosen wrongly. Cloud Storage also now offers a Rapid storage class, a high-performance class optimized for input and output intensive workloads; it is not one of the four the guide lists.
What each step colder trades
Cheaper to keep, dearer to read, longer to commit
Figure. Four classes left to right from hot to cold: Standard, Nearline (monthly), Coldline (quarterly), Archive (yearly or less). Beneath them, three rows: at-rest storage cost falls with each step; data access costs and minimum storage duration rise with each step; every class keeps low latency with no offline retrieval, while colder classes have slightly lower availability.
Worked example (synthetic). An item claims Archive data must be restored offline before it can be read. That is false — every class has low latency with no offline data retrieval; what Archive changes is the cost of reading.
The table tells you where each class sits; this figure tells you what moving one step colder trades, in Google's own terms. Google describes Nearline as a better choice than Standard where slightly lower availability, a thirty-day minimum storage duration and costs for data access are acceptable trade-offs for lowered at-rest storage costs. It describes Coldline the same way against both warmer classes, with a ninety-day minimum and higher costs for data access. And Archive has higher costs for data access and operations, and a three-hundred-and-sixty-five-day minimum. So the pattern is one sentence: each step colder is cheaper to keep, more expensive to read, and a longer commitment. The row that candidates get wrong is the bottom one. Google lists low latency with no offline data retrieval as a feature of every storage class — Archive is not tape, and reading it is costlier, not slower. If an access pattern is unpredictable, Google offers Autoclass, which lets Cloud Storage manage storage class transitions automatically.
Migrating or modernizing a database
Move it as it is, or improve it on the way
Rehost (lift and shift): minor or no changes; quickest, not cloud-optimized
Replatform or refactor: change the workload to use cloud capabilities
Homogeneous: same engine both sides; heterogeneous: different engines
Database Migration Service: into Cloud SQL or AlloyDB for PostgreSQL
Worked example (synthetic). A retailer's MySQL database moves to Cloud SQL for MySQL unchanged — a homogeneous rehost. Its Oracle database moves to Cloud SQL for PostgreSQL — heterogeneous, and a modernization rather than a move.
The fifth objective asks how an organization can migrate or modernize its current database, and Google gives the vocabulary for both halves. For the approach, Google's migration guide defines the major types. In a rehost migration — lift and shift — you move workloads with minor or no modifications; it is the quickest, but afterwards the workloads aren't optimized for the cloud. In a replatform migration — lift and optimize — you lift the existing workloads and then optimize them for the new environment. In a refactor migration — move and improve — you modify the workloads to take advantage of cloud capabilities. For the database itself, Google's Database Migration Service helps you migrate data to Google Cloud and supports migrations into Cloud SQL and AlloyDB for PostgreSQL. It distinguishes two kinds of move. Homogeneous migrations take place between the same database technology. Heterogeneous migrations, such as Oracle to Cloud SQL for PostgreSQL, change the technology — which makes a heterogeneous migration a modernization as well as a move. And it offers two timings: a one-time migration, a single point-in-time snapshot, or a continuous migration, which keeps changes flowing after an initial full dump and load so that switching over while source and destination are in sync gives minimal downtime.
Two independent choices in every database move
How much to change, and how to cut over
Choice
Option
What it means
How much changes
Homogeneous (rehost)
Same database technology on both sides; the quickest move
How much changes
Heterogeneous (modernize)
Different technology, such as Oracle to Cloud SQL for PostgreSQL
How to cut over
One-time
A single point-in-time snapshot of the database
How to cut over
Continuous
Changes keep flowing; switch when in sync for minimal downtime
Worked example (synthetic). A payments company cannot take its database offline for a weekend. Whatever engine it lands on, the timing choice is continuous migration, switched over when source and destination are in sync.
A database move is two decisions, and the exam likes to test one while describing the other. The first decision is how much changes. A homogeneous migration keeps the same database technology on both sides, and pairs naturally with a rehost, which Google calls the quickest migration because the refactoring is kept to a minimum. A heterogeneous migration changes the technology — Google's example is Oracle to Cloud SQL for PostgreSQL — and is the modernization path. The second decision is how to cut over. A one-time migration is a single point-in-time snapshot of the database, which means the source stops changing while it is taken. A continuous migration is a continuous flow of changes from source to destination that follows an initial full dump and load, and doing the switch when the two are in sync gives minimal downtime. Any combination is possible: a heterogeneous, continuous migration is modernizing while staying online. When an item stresses downtime, it is asking the second question, whatever it says about engines.
What this topic actually tests
Five discriminations, five objectives
Which shape is the data? objects, relational, documents, keyed series, or analysis. Relational at what scale? Cloud SQL, or Spanner for global consistency. Serve or analyze? Bigtable serves; BigQuery analyzes. How often is it read? monthly, quarterly, yearly — and colder costs more to read, never offline. What changes in the move? the engine, and separately the downtime.
Close on the five discriminations this topic leans on. First, the shape of the data picks the product family: objects to Cloud Storage, relational tables to Cloud SQL or Spanner, documents to Firestore, huge single-keyed series to Bigtable, and analysis to BigQuery. Second, relational at what scale: Cloud SQL for the familiar engines, fully managed; Spanner when transactions must stay consistent at global scale. Third, serve or analyze: Bigtable gives applications low-latency reads, and BigQuery answers questions across the business, with its storage and compute separated and its reach extended to other clouds by BigQuery Omni. Fourth, how often the data is read: Nearline monthly, Coldline quarterly, Archive less than yearly — cheaper to keep, costlier to read, and never offline. Fifth, what changes in a migration: the engine, which is homogeneous or heterogeneous, and separately the cutover, which is one-time or continuous. The next topic asks how that data becomes useful to the people who need it.
Study guide for Cloud Digital Leader, Unit 2 · Topic 3. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks. Discuss how smart analytics, business intelligence tools, and streaming analytics can add value in different business use cases.
Objectives, quoted from the exam guide:
Describe how Looker democratizes access to data by empowering individuals to self-serve business intelligence and create insights.
Discuss the value of analyzing and visualizing data from BigQuery in Looker to create real-time reports, dashboards, and integrating data into workflows.
Describe how streaming analytics in real time makes data more useful and generates business value.
Describe the main Google Cloud products that modernize data pipelines, including Pub/Sub and Dataflow.
Making data useful and accessible
Who can ask the data a question, and how soon
Who can ask — Looker lets people who do not write SQL self-serve answers. What they get from BigQuery — reports, dashboards, real-time insight, and data pushed into the tools they already use. How soon — streaming analytics, and the two products that carry the stream: Pub/Sub and Dataflow.
The last topic of Unit 2 is about the moment data turns into a decision, and the guide's summary names three ways it happens: smart analytics, business intelligence tools, and streaming analytics. The four objectives split cleanly into two questions. The first is who can ask the data a question. Business intelligence, or BI, used to mean waiting for an analyst to write a report; Looker's promise is that people who do not write structured query language, or SQL, can get their own answers, and the second objective asks what that looks like on top of BigQuery. The second question is how soon. Streaming analytics processes data as it arrives rather than in nightly batches, and the fourth objective names the two Google Cloud products that make a pipeline streaming: Pub/Sub, which receives and distributes the events, and Dataflow, which transforms them. Hold both questions — who, and how soon — and every objective in this topic is one of them.
Looker: business intelligence people can serve themselves
Analysts define the data once; everyone else asks it questions
Analysts describe the data once, in LookML, as a semantic model
Business users build queries in the Explore interface, without SQL
Looker's SQL generator turns each question into SQL for the database
One definition of each metric, reused by every report built on it
Worked example (synthetic). A retailer's regional managers each want margin by store. Instead of filing five report requests, they open the same Explore, pick store and margin, and get answers built on one agreed definition of margin.
The first objective asks how Looker democratizes access to data by letting individuals self-serve business intelligence, or BI. Start with what Looker is: Google calls it a product that helps you explore, share and visualize your company's data so that you can make better business decisions. The self-service comes from a division of labour, and it is worth knowing exactly where the line is. On one side, data analysts use LookML — short for Looker Modeling Language — to create and maintain data models that define the data structure and business rules for the data being analyzed. On the other, business users use the Looker query builder, or the Explore interface, to create queries based on that model. Between them sits the piece that makes it work: the Looker SQL generator translates LookML into structured query language, or SQL, which lets business users query without writing any LookML or SQL. The result is that a business user can build complex queries while focusing only on the content they need, not the complexities of SQL structure. There is a second benefit that is easy to miss. Because analysts write each SQL expression once, in one place, every report that uses a metric uses the same definition of it. Google puts it as two promises: Looker offers a unified surface to access all of an organization's data, and it lets you create consistent data models on top of all your data. Democratized access is not everyone writing their own numbers; it is everyone asking questions of the same numbers.
Where the SQL comes from
The business user never writes it; the model does
Loading Diagram...
Figure 1 — Mermaid diagram
Figure: A flow chart. A data analyst writes a LookML model once, holding structure and business rules. A business user picks fields in the Explore interface. Both feed the Looker SQL generator, which sends SQL to a database such as BigQuery; the formatted results return to the Explore interface.
Worked example (synthetic). An exam item says a sales director builds her own report without knowing SQL and asks what makes that possible. The graded answer is the analyst-built LookML model plus the SQL generator — not the director learning to query.
Drawn out, self-service business intelligence, or BI, is two inputs meeting in one place. On the left, a data analyst writes the LookML model once — the structure of the data and the business rules for it. Google describes LookML as the language used in Looker to create semantic data models. On the right, a business user picks fields in the Explore interface. Both arrive at the Looker SQL generator, which is the only thing in the picture that writes structured query language, or SQL. Google describes the round trip exactly: when a user creates a query, it is sent to the SQL generator; the SQL query is executed against the database, and then Looker returns the formatted results to the user in the Explore interface. Notice what the figure does not contain: an arrow from the business user to the database. That missing arrow is the whole objective. The user has direct access to answers without direct access to SQL, and the model in the middle is what keeps those answers consistent.
Looker on BigQuery: what the pairing adds
Reports, speed, real time, and data that goes where work happens
Capability
What Google says it gives you
Reports and dashboards
LookML tells Looker how to query, so everyone can create easy-to-read reports and dashboards
Speed
BI Engine is an in-memory service that accelerates queries, including those behind dashboards
Real time
BigQuery streaming supports continuous ingestion and analysis; Looker gives real-time insight from it
Alerts
A condition in the data, when met or exceeded, notifies chosen recipients
Workflows
Scheduled deliveries, and actions that send content to third-party services
Worked example (synthetic). A logistics firm streams delivery scans into BigQuery. Its operations dashboard in Looker shows late parcels as they happen, an alert tells the depot lead when a route crosses a threshold, and the daily summary lands in the team's chat channel.
The second objective asks for the value of analyzing and visualizing BigQuery data in Looker, and names three results: real-time reports, dashboards, and integrating data into workflows. Start with the pair. BigQuery is Google's fully managed data platform with business intelligence among its built-in features; Looker is an enterprise platform for business intelligence, data applications and embedded analytics. Google is careful to say you don't need Looker to use BigQuery — the pairing is a choice, and this table is the case for it. Reports and dashboards come from the model: LookML tells Looker how to query data, so everyone in the organization can create easy-to-read reports and dashboards. Speed comes from BigQuery BI Engine, a fast, in-memory analysis service that accelerates many structured query language, or SQL, queries, including queries used for BI dashboards. Real time comes from the data arriving continuously: BigQuery streaming supports continuous data ingestion and analysis, and Google's own reference pipeline ends with Looker providing real-time BI insights from the data stored in BigQuery. The last two rows are the workflow half of the objective. Alerts let you specify conditions in your data that, when met or exceeded, trigger a notification to specific recipients. And Looker can schedule immediate or recurring deliveries of dashboards, including to third-party services integrated with Looker, such as Slack.
From looking at data to acting on it
A dashboard someone must remember to open is only half the value
Figure. Three cards in a row: See, a fast dashboard on BigQuery data; Notice, an alert that tells people when a condition is met; Act, scheduled deliveries and actions that push content into other services.
Worked example (synthetic). Asked why a retailer's stock-out dashboard 'is not being used', the graded fix is an alert or a scheduled delivery to the store team — not a faster dashboard.
This figure is here because the phrase integrating data into workflows is the part of the objective candidates skip. A dashboard delivers value only when someone opens it. Looker closes that gap in two moves. The first is noticing: alerts are set on dashboard tiles, and Looker checks at a set frequency whether the condition has been met or exceeded and, if so, notifies the recipients. The second is acting where the work happens: beyond Looker's built-in destinations, you can use actions — also called integrations — to deliver content to third-party services that are integrated with Looker through an action hub server. So the progression runs see, notice, act. The same idea exists outside Looker, too: Connected Sheets brings the scale of BigQuery to the familiar Google Sheets interface — but that is BigQuery meeting a spreadsheet, not Looker, and the exam will expect you to tell them apart.
Streaming analytics: acting while the data is fresh
Continuous, not batched — because some data goes stale
Streaming analytics processes records continuously, not in batches
Batches bring long latency; time-sensitive data goes stale waiting
Real-time actions need continuous processing and analysis
Value: act on a click, a transaction or a reading as it happens
Worked example (synthetic). A bank scores card transactions overnight and learns of a fraud pattern the next morning. Streamed, the same analysis flags the pattern while the card is still being used.
The third objective asks how streaming analytics in real time makes data more useful and generates business value, and Google's definition already contains the argument. Streaming analytics is the processing and analyzing of data records continuously rather than in batches. It suits sources that send data in small pieces in a continuous flow as the data is generated — telemetry from connected devices, log files from web applications, ecommerce transactions, or information from social networks and geospatial services. The contrast is batch processing, which often processes large volumes of data at the same time, with long periods of latency. Google is fair to batch: it can be an efficient way to handle large volumes of data. But it does not work with time-sensitive data, because that data can be stale by the time it is processed. That is the business case in one word — stale. A price that should have changed an hour ago, or an alert about a transaction that has already cleared, has lost most of its value. Google states the requirement directly: generating real-time actions requires continuous stream processing and analysis. And it states the payoff: stream analytics makes data more organized, useful and accessible from the instant it is generated.
Batch against streaming
Choose by how fast the data loses its value
Question
Batch
Streaming
When is data processed?
Large volumes at the same time
Record by record, continuously
How long until a result?
Long periods of latency
As the data is generated
What does it suit?
Large volumes where a delay costs little
Small pieces in a continuous flow; time-sensitive data
What is it used for?
Efficient bulk processing
Real-time aggregation and correlation, filtering, or sampling
Worked example (synthetic). Monthly payroll is a batch problem: nothing is lost by waiting for the month to close. A delivery van's location is a streaming problem: an hour-old position is useless to the customer tracking it.
The exam rarely asks you to define streaming; it asks you to choose it, and this table is the choosing rule. Batch processes large volumes of data at the same time, with long periods of latency, which Google calls an efficient way to handle large volumes. Streaming processes records continuously, as the data is generated, and suits sources that send small pieces in a continuous flow. The deciding question is the third row: how fast does this data lose its value? If waiting costs little, batch is fine and often cheaper to reason about. If the data is time-sensitive, batch fails outright, because the data can be stale by the time it is processed. The last row lists what streaming is typically used for — real-time aggregation and correlation, filtering, or sampling. Neither is better in general. The graded answer is always the one that matches how perishable the data in the scenario is.
Three businesses, three kinds of value
Revenue, risk and relevance — each from acting sooner
Figure. Three cards: ecommerce, where streamed clickstreams drive real-time pricing, promotions and inventory; financial services, where streamed account activity reveals anomalous behavior and raises a security alert; news media, where streamed clicks are enriched to serve relevant articles.
Worked example (synthetic). An item describes a bank that wants to stop suspicious transfers before they settle and asks which capability delivers it. Streaming analytics with anomaly detection is the answer; a nightly report is the distractor.
Google gives three worked use cases for streaming analytics, and they are worth holding as three different kinds of business value, because exam scenarios are usually one of them in disguise. In ecommerce the value is revenue: analyze user clickstreams to optimize the shopping experience with real-time pricing, promotions, and inventory management. In financial services the value is risk: analyze account activity to detect anomalous behavior in the data stream and generate a security alert for abnormal behavior. In news media the value is relevance: stream user click records from various platforms and enrich the data with demographic information to better serve articles that are relevant to the targeted audience. In all three, the same event analyzed tomorrow would be worth much less — the shopper has left, the money has moved, the story is old. That is the sense in which streaming makes data more useful: not more data, but data used while it still matters. And Google adds one accessibility point that ties back to the first objective: its stream analytics provisioning reduces complexity and makes stream analytics accessible to both data analysts and data engineers.
Pub/Sub and Dataflow: modernizing the pipeline
One carries the events; the other transforms them
Pub/Sub: asynchronous messaging that decouples producers from processors
Pub/Sub delivers each event to every service that reacts to it
Dataflow: unified stream and batch processing — read, transform, write
Both managed: Dataflow provisions and removes its own worker VMs
Worked example (synthetic). A game studio's servers publish match events to Pub/Sub without knowing who uses them. A Dataflow pipeline turns them into per-minute leaderboards; a separate service reading the same events flags cheating. Neither slows the servers down.
The fourth objective names the two products that modernize a data pipeline, and the exam separates candidates who can say what each one does from those who treat them as a pair. Pub/Sub is the messaging layer. Google defines it as an asynchronous and scalable messaging service that decouples services producing messages from services processing those messages. Publishers send events to Pub/Sub without regard to how or when they are to be processed, and Pub/Sub then delivers the events to all the services that react to them. The decoupling is the modernization: in systems communicating through remote procedure calls, publishers must wait for subscribers to receive the data, while the asynchronous integration in Pub/Sub increases the flexibility and robustness of the whole system. Latencies are typically on the order of a hundred milliseconds, and Google names streaming analytics and data integration pipelines as what it is used for. Dataflow is the processing layer: a Google Cloud service that provides unified stream and batch data processing at scale, used to create pipelines that read from one or more sources, transform the data, and write it to a destination. It is fully managed — Google manages all of the resources needed to run it — and it can autoscale by provisioning extra worker virtual machines, or VMs, or shutting some down when fewer are needed. So: Pub/Sub moves events and does not transform them; Dataflow transforms data and is not a message bus.
Google's reference pipeline, end to end
Ingest, transform, store, present — one product per stage
Loading Diagram...
Figure 2 — Mermaid diagram
Figure: A left-to-right flow chart of Google's reference pipeline: an external system sends events to Pub/Sub, which ingests them; Dataflow reads from Pub/Sub and transforms or aggregates the data; Dataflow writes to BigQuery, the data warehouse; Looker provides real-time business intelligence insights from BigQuery.
Worked example (synthetic). An item lists four products and asks which one ingests a stream of device events. Pub/Sub is the graded answer; Dataflow is the tempting one, but it reads from Pub/Sub rather than receiving the events itself.
This is the whole topic on one line, and it is Google's own figure, described in the Dataflow overview stage by stage. Pub/Sub ingests data from an external system. Dataflow reads the data from Pub/Sub and writes it to BigQuery, and during this stage Dataflow might transform or aggregate the data. BigQuery acts as a data warehouse, allowing data analysts to run ad hoc queries on the data. And Looker provides real-time business intelligence, or BI, insights from the data stored in BigQuery. Read it as four jobs, one product each: ingest, transform, store, present. Exam items on this objective usually hand you a job and ask for the product, and the two errors to avoid are giving Pub/Sub's job to Dataflow, or BigQuery's job to Looker. It is also the picture that ties the unit together — the first two objectives of this topic live at the right-hand end, and the streaming objectives live at the left.
Pub/Sub and Dataflow, side by side
What makes each one a modernization
Aspect
Pub/Sub
Dataflow
Job
Messaging: receives and distributes events
Processing: reads, transforms, writes data
What it replaces
Publishers waiting on subscribers over direct calls
Separate code for batch and streaming — one model covers both
Operations
Scalable, asynchronous service
Fully managed; autoscales its worker VMs
Openness and reach
Distributes each event to every reacting service
Built on open source Apache Beam; templates need no Beam knowledge
Worked example (synthetic). A team keeps two codebases, one for nightly batch jobs and one for a live feed of the same data. The Dataflow row is the fix: the same programming model serves both.
Side by side, each product modernizes something different. Pub/Sub's job is messaging, and what it replaces is waiting: publishers that had to wait for subscribers to receive the data now hand events to a service that delivers them asynchronously to every service that reacts to them. Dataflow's job is processing, and what it replaces is duplication: Dataflow uses the same programming model for both batch and stream analytics, and a solution built on it can grow with your needs as you move from batch to streaming. Operationally, Dataflow is fully managed and autoscales its worker virtual machines, or VMs. Two further facts are exam favourites. Dataflow is built on the open source Apache Beam project, which matters for portability. And Google provides templates for common scenarios that you can deploy without knowing any Apache Beam programming concepts — which is the self-service argument of objective one, arriving at the engineering end of the pipeline. One more detail is worth knowing because distractors misuse it: by default, Dataflow provides exactly-once processing of every record.
What this topic actually tests
Who can ask, how soon, and which product does which job
Who writes the SQL? the LookML model and the SQL generator, not the business user. Is a dashboard enough? alerts and deliveries put data into workflows. Batch or stream? ask how fast the data goes stale. Which product? Pub/Sub ingests, Dataflow transforms, BigQuery stores, Looker presents.
Close the topic on four discriminations. First, self-service: when a business user gets their own answers in Looker, the structured query language, or SQL, is written by the Looker SQL generator from a model analysts defined — that is what democratized access means. Second, value beyond the dashboard: Looker on BigQuery gives real-time reports and dashboards, but the objective also says integrating data into workflows, which is alerts, scheduled deliveries and actions. Third, batch against streaming: choose by how quickly the data goes stale, because streaming analytics processes records continuously and batch brings long periods of latency. Fourth, the pipeline: Pub/Sub ingests the events, Dataflow transforms them, BigQuery stores them for analysis, and Looker provides the insight. Unit 3 picks up where this line ends — what artificial intelligence and machine learning can do with data that is already clean, current and accessible.
Unit 3 roadmap — Innovating with Google Cloud Artificial Intelligence
Unit 3 at a glance
Exam weight (Google, approx.)
~16%
Topics
3
Objectives
13
Bank questions
78
Flashcards
65
Seats per mock
9
Every objective below is quoted from Google's Cloud Digital Leader exam guide, retrieved 2026-09-22. Under each topic is that topic's own summary of what the exam actually tests, taken from its lecture.
Topic 1 — AI and ML Fundamentals
CDL-U3.T1 · 6 objectives · lecture deck of 14 slides
Define artificial intelligence (AI) and machine learning (ML).
Differentiate the capabilities of AI and ML from data analytics and business intelligence.
Discuss the types of problems that ML can solve.
Explain the business value ML creates, including: ability to work with large datasets; scaling business decisions; and unlocking unstructured data.
Explain why high-quality, accurate data is essential for successful ML models.
Discuss the importance of explainable and responsible AI
What this topic actually tests.Learned or programmed? only learning from data is ML. Past or next? BI explains what happened; ML predicts what might. Labeled or not? it picks supervised or unsupervised. Good data in, and an explanation out? without both, accuracy is not trust.
Topic 2 — Google Cloud’s AI and ML solutions
CDL-U3.T2 · 2 objectives · lecture deck of 8 slides
Explain which decisions and tradeoffs organizations need to consider when selecting Google Cloud AI/ML solutions and products, including: speed; effort; differentiation; required expertise.
Discuss which Google Cloud AI and ML solutions and products might apply given different business use cases, including: pre-trained APIs; AutoML; build custom models.
What this topic actually tests.General task? pre-trained API. Your data, a standard objective, no coders? AutoML — or BigQuery ML if it is already in BigQuery. Your own objective or metric, and data scientists on hand? custom training. Speed and ease run one way; differentiation runs the other.
Topic 3 — Building and using Google Cloud AI and ML solutions
CDL-U3.T3 · 5 objectives · lecture deck of 12 slides
Discuss how BigQuery ML lets users create and execute machine learning models in BigQuery by using standard SQL queries.
Select which Google Cloud pre-trained API best applies to different business use cases, including: Natural Language API, Vision API, Cloud Translation API, Speech-to-Text API, and Text-to-Speech API.
Explain how an organization can create business value by using their own data to train custom ML models with AutoML.
Discuss how building custom models by using Google Cloud’s Vertex AI can create opportunities for business differentiation.
Recognize TensorFlow as an end-to-end open source set of tools for building and training machine learning models and that Cloud Tensor Processing Unit (TPU) is Google’s proprietary hardware optimized for TensorFlow and ML performance.
What this topic actually tests.Does the business need its own model? No → a pre-trained API, chosen by input. Is the data in BigQuery and the team fluent in SQL? → BigQuery ML. Own data, no coders? → AutoML. Is the model the differentiator? → custom training. And TensorFlow is software; the TPU is hardware.
Try 15 sample questions from a bank of 703. Answers and detailed explanations included.
Q1medium
A company will run the same steady database workload for at least three more years and wants a lower price for it. Which discount mechanism fits?
A.
A sustained use discount, which must be purchased before the term starts.
B.
A spend cap budget, which lowers the rate once spending reaches the cap.
C.
A committed use discount: commit to a minimum level of use for a term.
D.
An anomaly alert, which refunds costs that exceed historical patterns.
Show answer & explanation
Correct Answer: C
A sustained use discount, which must be purchased before the term starts.: Incorrect. Sustained use discounts are not purchased at all: Compute Engine applies them automatically, with no action required.
A spend cap budget, which lowers the rate once spending reaches the cap.: Incorrect. A spend cap budget pauses usage until the cap is lifted; it is a control, not a discount.
A committed use discount: commit to a minimum level of use for a term.: Correct. With a commitment you commit to a minimum level of resources or spend for a one- or three-year term, and Google says committed use discounts suit resources with predictable and steady usage.
An anomaly alert, which refunds costs that exceed historical patterns.: Incorrect. An anomaly is a spike or deviation from expected spend; detecting one refunds nothing.
Answer: C
Q2medium
A bank's system flags any transfer above a limit its engineers set by hand, and it never changes unless they edit the rule. Why is it not machine learning?
A.
It runs on a bank's own servers rather than in the cloud.
B.
It uses fewer than three layers, so it cannot count as learning.
C.
It makes a decision, and machine learning only produces reports.
D.
It follows explicit programming rather than learning from data.
Show answer & explanation
Correct Answer: D
It runs on a bank's own servers rather than in the cloud.: Incorrect. Nothing in Google's definition of machine learning concerns where a system runs; the test is whether it learns from data.
It uses fewer than three layers, so it cannot count as learning.: Incorrect. Three layers is the threshold Google gives for deep learning, not for machine learning in general.
It makes a decision, and machine learning only produces reports.: Incorrect. Google says machine learning analyzes data, learns from the insights and then makes informed decisions.
It follows explicit programming rather than learning from data.: Correct. Google defines machine learning by what replaces the programmer: instead of explicit programming, it uses algorithms to analyze large amounts of data and learn from the insights.
Answer: D
Q3medium
A payments company must migrate its database without a long outage. Which Database Migration Service approach fits?
A.
A one-time migration, because a single snapshot is fastest.
B.
A continuous migration, switching over when source and destination are in sync.
C.
A rehost, because lift and shift always avoids downtime.
D.
A move to Archive storage first, because every class has no offline retrieval.
Show answer & explanation
Correct Answer: B
A one-time migration, because a single snapshot is fastest.: Incorrect. A one-time migration is a single point-in-time snapshot, which does not keep later changes flowing during the move.
A continuous migration, switching over when source and destination are in sync.: Correct. Google describes continuous migration as a flow of changes following an initial full dump and load, and says switching when source and destination are in sync gives minimal downtime.
A rehost, because lift and shift always avoids downtime.: Incorrect. A rehost describes how much the workload changes, not how the cutover is timed.
A move to Archive storage first, because every class has no offline retrieval.: Incorrect. Storage classes govern object storage cost and access, not database cutover.
Answer: B
Q4easy
A charity with no IT staff needs shared documents and email for thirty volunteers by the end of the week. They have no application of their own to run. Which cloud service model should they adopt?
A.
Infrastructure as a service, so they control the operating system
B.
Platform as a service, so that they can deploy their code without managing servers
C.
Software as a service, so they pay to use a complete managed application
D.
Containers as a service, so their workloads stay portable
Show answer & explanation
Correct Answer: C
The charity has no code to deploy and no staff to run infrastructure, so the finished-application model is the fit. Google frames software as a service exactly this way: "you pay to use a complete application for a specific purpose that is managed, maintained, and secured by the cloud provider, but you are responsible for taking care of your own data." Google Workspace is the example the same page gives.
Infrastructure as a service would hand them operating systems to maintain, which is work they cannot staff. Platform as a service is for deploying an application they do not have. Containers as a service also assumes containerized applications of their own, and portability is not a need they expressed.
Besides certifications, what other information does Google say its compliance site provides?
A.
Customer contact details
B.
General information about certain region or sector-specific regulations
C.
Competitor compliance gaps
D.
The internal penetration test scripts Google itself runs
Show answer & explanation
Correct Answer: B
The site covers both: it contains information about certifications and standards "as well as general information about certain region or sector-specific regulations." That matters because responsibilities are shaped by industry and location.
Customer contact details, competitor compliance gaps and internal penetration test scripts are each either private or not something a provider would publish.
A P3 case becomes urgent when production goes down. What does Google recommend?
A.
Raise the priority to match the new impact.
B.
Escalate at once, since escalation is the fastest route.
C.
Open a second case for the same issue.
D.
Close it and wait for Customer Care to reopen it.
Show answer & explanation
Correct Answer: A
Raise the priority to match the new impact.: Correct. Google says you can change a case's priority based on urgency and business impact, and that escalation doesn't generally make a high-impact case proceed faster.
Escalate at once, since escalation is the fastest route.: Incorrect. Google says escalation doesn't generally make a high-impact case proceed faster; it flags a broken process.
Open a second case for the same issue.: Incorrect. Google asks for one support case per issue, and duplicates are closed.
Close it and wait for Customer Care to reopen it.: Incorrect. Closing ends the case; a closed case can be reopened only within 15 days, and closing it does nothing to speed the work.
Answer: A
Q7hard
An operations lead worries that Google's hardware maintenance will force regular reboots of the company's Compute Engine VMs. What does Google state?
A.
Maintenance reboots are avoided only by running on Spot VMs.
B.
Live migration lets maintenance happen without interrupting or rebooting the VM.
C.
Reboots are unavoidable; the answer is a load balancer in front of each VM.
D.
Maintenance requires the customer to rehost the VM onto new hardware first.
Show answer & explanation
Correct Answer: B
Maintenance reboots are avoided only by running on Spot VMs.: Incorrect. Google says Spot VMs can't live migrate or automatically restart on a host event — the opposite of avoiding interruption.
Live migration lets maintenance happen without interrupting or rebooting the VM.: Correct. Google says live migration lets it perform maintenance without interrupting a workload, rebooting an instance, or modifying the instance's properties.
Reboots are unavoidable; the answer is a load balancer in front of each VM.: Incorrect. A load balancer spreads traffic across instances; it does not change whether a single VM is rebooted.
Maintenance requires the customer to rehost the VM onto new hardware first.: Incorrect. Rehost is a migration path into the cloud, not a maintenance procedure.
Answer: B
Q8easy
How many zones does a Google Cloud region consist of, and where are they housed?
A.
Three or more zones, housed in three or more physical data centers
B.
Exactly two zones, housed in a single physical data center for low latency
C.
One zone per region, with the data center chosen by Google at creation time
D.
Ten zones, housed across three continents to guarantee global redundancy
Show answer & explanation
Correct Answer: A
The rule and its exceptions are both published: "A region consists of three or more zones housed in three or more physical data centers. The regions Stockholm, Mexico, Osaka, and Montreal have three zones housed in one or two physical data centers. These regions are in the process of expanding to at least three physical data centers."
Two zones in one data centre, one zone per region and ten zones across three continents are none of them what Google states.
An operations team keeps learning about incidents from customers. Which benefit of modernized operations addresses that?
A.
A guarantee that Google Cloud services never experience an outage.
B.
Seeing behavior early, so changes are handled quickly.
C.
Moving responsibility for incidents to Customer Care.
D.
Removing the need to review incidents afterwards.
Show answer & explanation
Correct Answer: B
A guarantee that Google Cloud services never experience an outage.: Incorrect. Google's reliability guidance is to plan for failure, which assumes failures happen; no service is promised to be outage-free.
Seeing behavior early, so changes are handled quickly.: Correct. Google says understanding how applications behave and how components connect helps you anticipate, identify and respond to unexpected changes quickly and effectively.
Moving responsibility for incidents to Customer Care.: Incorrect. Customer Care provides advice, troubleshooting and operational knowledge; it does not take over the customer's own monitoring.
Removing the need to review incidents afterwards.: Incorrect. Google's framework still requires a post-incident review to find root cause and lessons learned.
Answer: B
Q10easy
How does Google describe the way AI systems work?
A.
They follow a rule set that an engineer has written out in advance for each and every situation that the system is expected to encounter
B.
They learn from vast amounts of data, identifying patterns to make predictions or decisions without being explicitly programmed for every scenario
C.
They query a knowledge base assembled by subject matter experts before answering any question
D.
They run a fixed statistical model whose coefficients are set once and never revised afterwards
Show answer & explanation
Correct Answer: B
Learning from data rather than from rules is the point, and Google draws the contrast itself: "AI systems learn from vast amounts of data, identifying patterns to make predictions or decisions without being explicitly programmed for every scenario. Think of it as teaching a computer by showing it a million examples instead of writing a million rules."
A hand-written rule set is exactly what that sentence contrasts against. An expert knowledge base and a fixed statistical model likewise describe something other than learning from examples.
Which statement correctly separates cloud modernization from cloud migration?
A.
Modernization moves applications and databases to a cloud computing environment unchanged.
B.
Modernization transforms applications, data and infrastructure to use the cloud better.
C.
Modernization must be finished before any migration can start.
D.
Modernization means decommissioning applications and compute at source.
Show answer & explanation
Correct Answer: B
Modernization moves applications and databases to a cloud computing environment unchanged.: Incorrect. Moving applications, databases, storage and infrastructure to a cloud environment is Google's definition of cloud migration.
Modernization transforms applications, data and infrastructure to use the cloud better.: Correct. Google defines cloud modernization as a strategy for transforming existing applications, data and infrastructure to take better advantage of the benefits offered by cloud computing — as opposed to migration, which moves them.
Modernization must be finished before any migration can start.: Incorrect. Google lists optimization or modernization as a phase of the migration process itself, after assessment, planning and migration.
Modernization means decommissioning applications and compute at source.: Incorrect. Decommissioning the application and compute at source is the definition of retire.
Answer: B
Q12medium
Why does Google say traditional tools struggle where machine learning helps?
A.
Because the traditional tools are simply unable to connect to any form of cloud storage at all
B.
Because the sheer volume coupled with complexity often makes data difficult to analyze using traditional tools
C.
Because traditional tools cannot produce charts
D.
Because traditional tools require a data warehouse
Show answer & explanation
Correct Answer: B
Google names volume and complexity together: data helps businesses make better decisions, "But the sheer volume coupled with complexity often makes data difficult to analyze using traditional tools." It adds that building and iterating analytical models by hand "eats up employees' time in a way that scales poorly."
An inability to connect to cloud storage is a integration detail, not the stated limitation. Producing charts is something traditional business intelligence tools do well. And requiring a data warehouse is not a defect — warehouses are a primary component of business intelligence.
What is a data warehouse also called, and what is it used for?
A.
An enterprise data warehouse (EDW) — an enterprise data platform used for the analysis and reporting of structured and semi-structured data
B.
A data lakehouse — a store that keeps every record in its raw native format until it is queried
C.
An operational data store — a system tuned for high-volume single-row transactional lookups
D.
A content delivery network — a caching tier that holds the reporting output close to whichever people happen to be reading through it at the time
Show answer & explanation
Correct Answer: A
Both the alternative name and the purpose are given: "A data warehouse, also called an enterprise data warehouse (EDW), is an enterprise data platform used for the analysis and reporting of structured and semi-structured data from multiple data sources."
A lakehouse, an operational data store and a content delivery network are each real things, and none of them is what the page defines here.
An operations team must apply the same production-wide change to twenty clusters spread across ten projects and three environments. What business value does a fleet provide here?
A.
It removes the need for the clusters to run Kubernetes.
B.
It moves all twenty clusters onto Google Cloud first.
C.
It requires the change to be made on each cluster, but faster.
D.
The change is managed for the group from one fleet host project.
Show answer & explanation
Correct Answer: D
It removes the need for the clusters to run Kubernetes.: Incorrect. A fleet is a logical grouping of Kubernetes clusters; it does not replace Kubernetes.
It moves all twenty clusters onto Google Cloud first.: Incorrect. A fleet can include clusters outside Google Cloud; registering them does not move them.
It requires the change to be made on each cluster, but faster.: Incorrect. Making the change on individual clusters is what Google says happens WITHOUT fleets.
The change is managed for the group from one fleet host project.: Correct. Google says that without fleets a production-wide change must be made on individual clusters in multiple projects, while fleets let you group clusters and manage them from a single fleet host project.
Answer: D
Q15hard
Which capabilities does Google say Artificial General Intelligence is expected to have, in contrast with narrow intelligence?
A.
Performing one single specific narrow task extremely well indeed, and nothing else
B.
Running without any training data
C.
Operating faster than existing models
D.
Being adaptive, autonomous, and capable of learning from its actions
Show answer & explanation
Correct Answer: D
Google contrasts the two: "Unlike ANI, AGI is expected to be adaptive, autonomous, and capable of learning from its actions." It also stresses that "AGI does not yet exist."
Performing a single specific task extremely well is the definition of narrow intelligence, not what distinguishes general intelligence. Running without training data is not claimed for either. And operating faster than existing models is a performance property rather than a capability difference.
451 flashcards for spaced-repetition study. Showing 30 sample cards below.
CDL — AI and ML Against Analytics and BI(5 cards shown)
Question
How does an AI system learn, as against being programmed?
Answer
It learns from vast amounts of data, identifying patterns to predict or decide without explicit programming for every scenario — teaching a computer by showing it a million examples instead of writing a million rules.
Question
What makes enterprise data hard to analyze with traditional tools?
Answer
Sheer volume coupled with complexity. Businesses are inundated with data, and the two together defeat the traditional toolkit.
Question
Which part of the traditional analytics workflow scales poorly, and what does ML change?
Answer
Building, testing, iterating and deploying analytical models eats employee time in a way that scales poorly. Machine learning lets an organization derive insights quickly as data scales.
Question
What is a neural network?
Answer
A model using a system of artificial neurons — computational nodes used to classify and analyze data.
Question
Name the ML approaches Google lists alongside supervised and unsupervised learning.
Answer
Decision trees and linear regression — listed beside supervised and unsupervised learning as the techniques in play.
CDL — Analyzing BigQuery Data in Reports and Workflows(5 cards shown)
Question
What does BigQuery streaming support, and how fast is the analysis engine behind it?
Answer
Continuous data ingestion and analysis. The scalable, distributed engine queries terabytes in seconds and petabytes in minutes.
Question
Which two first-party interfaces does BigQuery offer?
Answer
The Google Cloud console interface and the BigQuery command-line tool.
Question
How do existing third-party tools and utilities reach BigQuery?
Answer
Through ODBC and JDBC drivers, which provide interaction with applications that already exist.
Question
Which programmatic routes into BigQuery do developers and data scientists have?
Answer
Client libraries for Python, Java, JavaScript and Go, plus BigQuery's REST API and RPC API, for transforming and managing data.
Question
What lets one BigQuery-backed dashboard sit over a mixed estate?
Answer
A uniform way to work with structured and unstructured data, plus support for open table formats — Apache Iceberg, Delta and Apache Hudi.
CDL — Authentication, Authorization and Auditing(5 cards shown)
Question
Which of the three A's is IAM, and what question does it answer?
Answer
Authorization — fine-grained authorization for Google Cloud. It controls who can do what on which resources.
Question
What does IAM do when someone tries an action they lack permission for?
Answer
It prevents them from performing it. Every action requires certain permissions, and IAM checks first.
Question
Name the three components of giving someone permissions in IAM.
Answer
Principal — the identity of the person or system. Role — the collection of permissions. Resource — the Google Cloud resource they may access.
Question
How is a role actually granted, and what does the resource hierarchy add?
Answer
By granting the principal a role on the resource, using an allow policy. Allow policies attach to resources organized hierarchically, so access can be granted to a single resource or to a container of resources.
Question
What does a principal represent, and what were they called before?
Answer
One or more identities that have authenticated to Google Cloud. They were formerly called members, and some APIs still use that term.
CDL — BigQuery as a Serverless Data Warehouse(5 cards shown)
Question
What is BigQuery, and which capabilities are built into it?
Answer
A fully managed, AI-ready data platform with built-in machine learning, search, geospatial analysis and business intelligence.
Question
What does BigQuery's serverless architecture let you skip, and which languages do you use instead?
Answer
Zero infrastructure management — no provisioning and no manual scaling. Questions are asked in languages like SQL and Python.
Question
Which two parts is BigQuery's architecture made of, and what does splitting them buy?
Answer
A storage layer that ingests, stores and optimizes data, and a compute layer for analytics. Each can allocate resources dynamically without affecting the other's performance or availability.
Question
Which legacy-database problem does separating compute and storage remove?
Answer
Resource conflict. Legacy databases share resources between read, write and analytical operations, so queries slow down while data is being written to or read from storage.
Question
What makes BigQuery workable across a multicloud estate rather than only its own storage?
Answer
It gives a uniform way to work with structured and unstructured data and supports open table formats — Apache Iceberg, Delta and Apache Hudi.
CDL — BigQuery ML and SQL-Native Machine Learning(5 cards shown)
Question
How does BigQuery ML let you create and run machine learning models?
Answer
With either GoogleSQL queries or the Google Cloud console.
Question
Where does a BigQuery ML model live?
Answer
In a BigQuery dataset, like a table or a view. BigQuery ML also reaches Agent Platform models and Cloud AI APIs from there.
Question
Who does BigQuery ML empower, and with which tools?
Answer
Data analysts — the primary data warehouse users — building and running models with existing business intelligence tools and spreadsheets, so predictive analytics can guide decisions across the organization.
Question
Which languages does BigQuery ML let a team avoid, and what does it use instead?
Answer
No need to program in Python or Java — models are trained and AI resources reached through SQL, a language data analysts already know.
Question
Which step does BigQuery ML remove to speed up model development?
Answer
Moving data out of the data warehouse. It brings ML to the data instead, which reduces complexity because fewer tools are required.
CDL — Business Benefits of Cloud Technology(5 cards shown)
Question
What two things does cloud architecture let users do from anywhere with an internet connection?
Answer
Access cloud services, and scale those services up or down as needed.
Question
Why is cloud security generally recognized as stronger than security in an enterprise data center?
Answer
Because of the depth and breadth of the security mechanisms providers put in place — and their security teams are known as top experts in the field.
Question
Whatever service model is used, how much computing resource does an enterprise pay for?
Answer
Only what it uses. That holds across every cloud computing service model.
Question
What does an enterprise stop overbuilding on cloud, and where do the freed IT staff go?
Answer
It stops overbuilding data center capacity for unexpected spikes in demand or business growth, and can deploy IT staff to more strategic initiatives.
Question
Where does Google locate the strategic value of cloud, set against buying technology outright?
Answer
Providers stay on top of the latest innovations and offer them as services, so an enterprise gets more competitive advantage and a higher return on investment than it would investing in soon-to-be obsolete technologies.
Flowchart, left to right. Deployment model<br/>public · private · hybrid connects to An estate is described<br/>by BOTH axes. Vendor count<br/>single cloud · multicloud connects to C. C connects to Private only. C connects to Hybrid, one provider. C connects to Hybrid, two providers<br/>= hybrid AND multicloud.
Loading Diagram...
Flowchart, top to bottom. Changing business and<br/>market dynamics connects to Digital transformation. Volumes of data to collect,<br/>process and analyze connects to T. Requires: strong commitment from<br/>business AND IT teams connects to T. T connects to How the organization operates. T connects to How it optimizes internal resources. T connects to How it delivers value to customers.
Loading Diagram...
Flowchart, left to right. Acquire hardware<br/>and software connects to Depreciate over<br/>operating life. Consume<br/>a resource connects to Cost incurred<br/>as it is consumed. B connects to Total cost of ownership:<br/>provisioning + use + managing. D connects to T.
Loading Diagram...
Flowchart, left to right. Internet user connects to User's internet<br/>service provider. I connects to Google point of presence<br/>near the USER (Premium Tier). P1 connects to Google's private<br/>backbone"] R["Region running<br/>the application. I connects to Regular ISP and<br/>transit networks (Standard Tier). T connects to P2["Google point of presence<br/>near the REGION"] R.
Loading Diagram...
Flowchart, top to bottom. Do you need a finished application,<br/>with no expertise to build one? connects to SaaS<br/>e.g. Google Workspace (yes). Do you need a finished application,<br/>with no expertise to build one?"} -->|yes| S["SaaS<br/>e.g. Google Workspace connects to Are you writing the application,<br/>and want to focus on the code? (no). Q2 connects to PaaS<br/>e.g. Cloud Run, App Engine (yes). Q2 connects to Moving an existing workload as it is,<br/>or need specific VMs and configs? (no). Q3 connects to IaaS<br/>e.g. Compute Engine (yes).
Loading Diagram...
Flowchart, left to right. Many source systems connects to Extract, transform, load. E connects to Data warehouse<br/>rigid schema, batches. Many source systems"] --> E["Extract, transform, load connects to Data lake<br/>native format, no pre-defined schema. W connects to Repeatable reporting<br/>and business intelligence. L connects to Machine learning<br/>on raw data.
Loading Diagram...
Flowchart, left to right. Acquisition<br/>and ingestion connects to Processing<br/>extract, transform, load. P connects to Storage<br/>warehouse or lake"] N["Analytics<br/>and AI. N connects to Action<br/>in real or near-real time. X connects to Secure disposal.
Loading Diagram...
Flowchart, top to bottom. What shape is the data? connects to Cloud Storage (Files and media). What shape is the data? connects to Must transactions stay<br/>consistent at global scale? (Tables and relations). Q2 connects to Cloud SQL (No). Q2 connects to Spanner (Yes). What shape is the data? connects to Firestore (Documents for an app). What shape is the data? connects to Bigtable (Huge single-keyed series). What shape is the data? connects to BigQuery (Anything, for analysis).
Loading Diagram...
Flowchart, left to right. Analysts: SQL<br/>and Python connects to BigQuery<br/>compute layer. C connects to BigQuery<br/>storage layer. C connects to BigQuery Omni<br/>via BigLake tables. O connects to Data held in<br/>another public cloud.
Loading Diagram...
Flowchart, left to right. Data analyst connects to LookML model<br/>structure + business rules (writes once). Business user connects to Explore interface (picks fields). M connects to Looker SQL generator. E connects to G. G connects to ("Database,<br/>such as BigQuery") (SQL). D connects to E (formatted results).
Loading Diagram...
Flowchart, left to right. External system<br/>devices, apps, logs connects to Pub/Sub<br/>ingests (events). P connects to Dataflow<br/>transforms or aggregates. F connects to ("BigQuery<br/>data warehouse"). Q connects to Looker<br/>real-time BI insights.