Unit 2.1 study guide — The Value of Data
Cloud Digital Leader › Unit 2 › Topic 1
The Value of Data
Study guide for Cloud Digital Leader, Unit 2 · Topic 1. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks. Describe the intrinsic role that data plays in an organizations’ digital transformation.
Objectives, quoted from the exam guide:
- Explain how data generates business insights, drives decision making, and creates new value.
- Differentiate between basic data management concepts, in particular: databases; data warehouses; data lakes.
- Explain how organizations can create value by using their current data, collecting new data, and sourcing data externally.
- Describe how the cloud unlocks business value from all types of data, including structured data and previously untapped unstructured data.
- Discuss the main data value chain concepts and terms.
- Explain how data governance is essential to a successful data journey.
The value of data
Where value comes from, where data lives, and what keeps it usable
Where value comes from — insight, decisions, new products; current, new and external data; structured and unstructured. Where data lives — database, warehouse, lake, and the stages between them. What keeps it usable — governance, without which none of the rest is worth trusting.
Unit two is about data, and this first topic asks one thing: what role data plays in an organization's digital transformation. Its six objectives answer in three parts. The first is where value comes from — how data turns into insight, how insight turns into decisions and new products, and which data an organization can draw on: what it already has, what it can start collecting, and what it can source from outside. The second is where data lives — the difference between a database, a data warehouse and a data lake, the difference between structured and unstructured data, and the stages data passes through on its way from being collected to being used. The third is what keeps any of it usable, which is governance. Google puts that third part bluntly, and it is worth hearing first: the value of data is only realized when it is trustworthy, discoverable and governed. Everything before governance in this topic is potential value; governance is what lets an organization actually collect it.
From data to insight to decision to new value
Data is worth nothing until it changes what someone does
- Analytics now sits at the center of all core business activities
- Business intelligence: collect and analyze data for strategic and daily decisions
- More visibility gives better decisions and uncovers growth opportunities
- Insight accelerates new products, features and updates
- Only the right data, effectively analyzed, yields insight
Worked example (synthetic). A grocer's loyalty data shows that one store's customers buy baby products on Saturday mornings. The data was always there; the value appears when the Saturday delivery schedule changes because of it.
The first objective asks how data generates business insights, drives decision making and creates new value — three steps, and the exam tests each. Start with where data now sits. Google says today's data analytics activities have transformed to the center of all core business activities, including revenue generation, cost containment, improving operations and enhancing customer experiences. Insight is the first step. Business intelligence, or BI, is the process of using people and technologies to collect and analyze data for an organization's strategic and daily decision-making. Decision is the second. Google's statement of the whole idea is that the more visibility you have into anything, the more effectively you can gain insights to make better decisions, uncover growth opportunities and improve your business model. New value is the third, and it is more than doing the old things better: these insights can guide and accelerate the planning, production and launch of new products, features and updates. Insight also points inward — analytics helps you find where to reduce costs, save time and increase efficiency. And there is a condition on all of it: big data must contain the right data and then be effectively analyzed to yield insights that drive decisions. Volume alone produces nothing.
Four places the value lands
Same data, four different kinds of return
Figure. Four cards naming where the value of data lands: better decisions, new products, leaner operations, and better customer experience.
Worked example (synthetic). Asked what a bank gains by analyzing call-center transcripts beside account data, the graded answer is customer experience — the gain comes from combining the two, not from either alone.
This figure turns the first objective into four landing places, because the exam describes a business outcome and asks you to recognize data behind it. Better decisions come from visibility: Google's central claim for big data is that more visibility yields insight to make better decisions and uncover growth opportunities. New products come from insight that can guide and accelerate the planning, production and launch of new products, features and updates. Leaner operations come from generating insights that show where you can reduce costs, save time and increase efficiency. And better customer experience has a specific mechanism worth remembering: combining and analyzing structured data sources together with unstructured ones provides more useful insights for consumer understanding and personalization. That last card previews objective four — some value only exists when two kinds of data are read together.
Database, data warehouse, data lake
Three stores, three different jobs
| Store | What it holds | What it is for |
|---|---|---|
| Relational database | Information structured in tables, rows and columns, with predefined relationships | Seeing how records relate to each other |
| Data warehouse | Structured and semi-structured data from multiple sources, current and historical | Analysis and reporting; a long-range view over time |
| Data lake | Structured, semistructured and unstructured data in its native format | Diverse, raw data for machine learning and advanced analytics |
Worked example (synthetic). A retailer keeps each order in a database, loads a year of orders into a warehouse to report seasonal trends, and drops product photos and review text into a lake to train a recommendation model.
The second objective asks you to differentiate three data management concepts, and the fastest way is to ask what each one holds and what it is for. A relational database is a way of structuring information in tables, rows and columns; Google adds that it organizes data in predefined relationships, making it easy to see how different data structures relate to each other. A data warehouse is an enterprise system for the analysis and reporting of structured and semi-structured data from multiple sources, such as point-of-sale transactions, marketing automation and customer relationship management. It can store both current and historical data in one place and is designed to give a long-range view of data over time, which is why Google calls it a primary component of business intelligence. A data lake is a centralized, scalable and secure repository designed to store, process and analyze large amounts of structured, semistructured and unstructured data in its native format. The contrast Google draws between the last two is precise: a traditional data warehouse is optimized for repeatable business reporting and structured analysis, while a data lake excels at handling the diverse, raw data required for machine learning.
How data reaches each store
A schema before loading, or none at all
Figure: A flow chart. Many source systems feed two paths. One passes through extract, transform and load into a data warehouse with a rigid schema, loaded in batches, which serves repeatable reporting and business intelligence. The other goes straight into a data lake in native format with no pre-defined schema, which serves machine learning on raw data.
Worked example (synthetic). A team asks why its warehouse cannot answer a question nobody planned for. The figure's answer: the warehouse's schema was decided before loading, and data outside it was never loaded.
Drawn out, the warehouse and the lake differ most in what happens before data arrives. On the warehouse path, data is shaped first. Google describes extract, transform and load, or ETL, as a traditionally accepted way to combine data from multiple systems into a single database, data store, data warehouse or data lake. A traditional warehouse is typically designed to capture a subset of data in batches and store it based on rigid schemas — which Google says makes it unsuitable for spontaneous queries or real-time analysis. That rigidity is also its strength: it serves repeatable reporting. On the lake path, nothing is decided first. A data lake lets enterprises ingest any data from any source, on-premises, cloud or edge, without the constraints of pre-defined schemas. So the useful question in an exam item is not which store is better, but whether the questions are known in advance. Known questions favor the warehouse; unknown ones, and machine learning on raw data, favor the lake.
Current data, new data, external data
Three places an organization finds value it does not yet use
- Current data: combine what is already held into one unified view
- New data: collect continuously, at any speed and volume
- External data: discover third-party datasets and combine them with your own
- Cloud warehouses collect from internal AND external sources
Worked example (synthetic). A logistics firm unifies its dispatch and billing data, starts streaming truck sensor readings, and subscribes to a third-party weather dataset — three sources, three different kinds of value.
The third objective asks how organizations create value from three sources: the data they already have, data they start collecting, and data they source externally. Current data first. Much of it sits in separate systems, and the value is in combining it: Google defines data integration as discovering, moving and combining this data into a unified view to drive insights and power artificial intelligence (AI) driven analytics. New data second. Big data lets you integrate automated, real-time data streaming with advanced analytics to continuously collect data, find new insights and discover new opportunities for growth and value — and a data lake lets enterprises ingest data at any speed and volume. External data third. Google's own mechanism for it is BigQuery sharing, formerly Analytics Hub, a data exchange platform to securely share, discover and access data across organizational boundaries without replicating it. With it you can discover curated third-party and Google datasets, and combine them with your internal data to augment analytics and machine learning. The cloud data warehouse joins all three: like a traditional warehouse, it collects, integrates and stores data from internal and external data sources.
Which source, which mechanism
Name the source in the scenario, then the move that unlocks it
| Source | The move | What Google says it gives |
|---|---|---|
| Data already held, in separate systems | Integrate it | A unified view that drives insights |
| Data not yet collected | Stream it continuously | New insights and new opportunities for growth |
| Data held by someone else | Subscribe through a data exchange | Third-party datasets combined with internal data, without replicating it |
Worked example (synthetic). An item says a retailer wants foot-traffic data it has never owned. The graded move is external sourcing through a data exchange — not a new collection program it would have to build.
The exam will describe an organization and ask where its next value comes from, so rehearse the three sources as moves. Data already held in separate systems calls for integration — discovering, moving and combining it into a unified view that drives insights. Data not yet collected calls for streaming, which Google describes as integrating automated, real-time data streaming with analytics to continuously collect data and discover new opportunities for growth and value. Data held by someone else calls for sourcing it, and Google's mechanism is a data exchange: BigQuery sharing lets you discover curated third-party and Google datasets and combine them with your internal data, and it does so across organizational boundaries without replicating the data. The trap in these items is to answer every one with collect more. Sometimes the value is already inside the organization, and sometimes it belongs to somebody else.
Structured and unstructured data
Most new data has no fixed schema — the cloud can keep it anyway
Figure. Three cards for structured, semi-structured and unstructured data with examples of each, above a band explaining that traditional warehouses discarded data to save storage while a cloud warehouse with a cloud data lake keeps unstructured data, and combining the two kinds gives new insight.
Worked example (synthetic). An insurer's claims table says what was paid; the adjusters' photos and notes say why. Only when the cloud keeps both can the insurer find which damage patterns predict the largest claims.
The fourth objective asks how the cloud unlocks value from all types of data, including previously untapped unstructured data — and the word untapped is the point. Google separates the types clearly. More traditional structured data, such as data in spreadsheets or relational databases, is now supplemented by unstructured text, images, audio and video files, or semi-structured formats like sensor data that cannot be organized in a fixed data schema. Why was it untapped? Partly scale: big data sets are so complex in volume, velocity and variety that traditional data management systems cannot store, process and analyze them. And partly economics: in a traditional warehouse, storage is typically limited compared to compute, so data is transformed quickly and then discarded to keep storage space free, and an on-premises warehouse means buying your own hardware and software, making it expensive to scale. The cloud changes both. A cloud data warehouse can be used with a cloud data lake to collect and store unstructured data. And the payoff is in the combination: analyzing structured and unstructured sources together provides more useful insights for consumer understanding and personalization.
The data value chain
The stages data passes through, and the word for each
Figure: A left-to-right flow of the data life cycle: acquisition and ingestion, processing by extract transform and load, storage in a warehouse or lake, analytics and AI, action in real or near-real time, and secure disposal. A governance box connects to the first and last stages to show it spans the whole chain.
Worked example (synthetic). A team measures success by how much data it ingests. The chain shows why that is the wrong stage to measure: value is realized at action, three stages later.
The fifth objective asks for the main data value chain concepts and terms. The value chain is the guide's phrase rather than one Google's pages use, so this deck teaches it with the life cycle Google does define. Data governance, Google says, is a principled approach to managing data during its life cycle, from acquisition and ingestion to artificial intelligence (AI) data analytics and secure disposal. Read the stages in order. Acquisition and ingestion bring data in. Processing makes it usable — Google describes extract, transform and load, or ETL, as the end-to-end process by which a company takes its full breadth of data and gets it to a state where it is actually useful for business purposes. Storage holds it, in a warehouse or a lake. Analytics turns it into insight. Action is where the value is realized: Google says cloud data warehouses should support streaming use cases to activate on data in real or near-real time. And the chain ends deliberately, with secure disposal. Governance is drawn beside the chain, not inside it, because it applies to every stage.
The vocabulary: six Vs of big data
Three original Vs, three that decide whether data is worth it
| Term | What it describes |
|---|---|
| Volume | How much data there is — the most common characteristic of big data |
| Velocity | The speed at which data is generated |
| Variety | Data from many sources: structured, unstructured or semi-structured |
| Veracity | Quality and accuracy — the higher the veracity, the more trustworthy |
| Variability | Meaning that changes over time, leading to inconsistency |
| Value | Whether the data you collect is worth anything to the business |
Worked example (synthetic). An item describes sensor data that arrives every second but is often wrong. Two Vs are in play — velocity is high, veracity is low — and the graded answer names the one the question asks about.
The rest of the value chain vocabulary the exam uses comes from Google's big data page, and it is organized as Vs. Google says big data definitions may vary slightly, but it will always be described in terms of volume, velocity and variety. Volume is the amount — the most common characteristic associated with big data. Velocity refers to the speed at which data is generated. Variety means data can come from many sources and be structured, unstructured or semi-structured. Then three more, which Google says are often mentioned in relation to harnessing the power of big data: veracity, variability and value. Veracity is quality — the higher the veracity of the data, the more trustworthy it is. Variability is meaning that shifts: the meaning of collected data is constantly changing, which can lead to inconsistency over time. And value is the question the whole chain exists to answer: it is essential to determine the business value of the data you collect. The first three describe data; the last three decide whether it is worth keeping.
Why governance decides the journey
Governance is what makes more access safe, not what prevents it
- Secure, private, accurate, available and usable — for people and for machine learning
- Defines who can access sensitive information
- Lets data be democratized without security or compliance breaches
- Controls that allow GREATER access, because the access is controlled
Worked example (synthetic). A bank wants every analyst to query customer data. Without governance the answer is no; with column-level controls and defined access, the answer becomes yes for most columns.
The last objective asks why data governance is essential to a successful data journey, and Google's answer turns on a word that sounds like a contradiction: access. Data governance is everything you do to ensure data is secure, private, accurate, available and usable, for human analysis, machine learning and building agents. It means setting internal standards for how data is gathered and processed, and it involves defining who can access sensitive information and ensuring that the democratization of data does not lead to security risks or compliance breaches. Here is the part candidates get backwards. Governance is not mainly a brake. Google says data governance allows setting and enforcing controls that allow greater access to data, gaining the security and privacy from the controls on data. The controls are what make it safe to open the data up. That is why governance is essential rather than optional: effective governance lets organizations go from raw data to action faster while maintaining strict security and compliance standards, and without it, the value of data is simply not realized — Google says that value comes only when data is trustworthy, discoverable and governed.
What governance actually does
Three uses, each with its own word
| Use | What it means |
|---|---|
| Data stewardship | Accountability for the data and the processes that ensure its proper use, given to data stewards |
| Data quality | Six dimensions: accuracy, completeness, consistency, timeliness, validity, uniqueness |
| Compliance | Data is safe, secure, private, usable, and compliant with internal and external policies |
Worked example (synthetic). A dashboard shows two different revenue totals for the same month. That is a data quality failure on the consistency dimension — and deciding who fixes it is a stewardship question.
Google lists the common uses of governance, and each carries a term the exam uses. Stewardship is about people: data governance often means giving accountability and responsibility for both the data itself and the processes that ensure its proper use to data stewards. Quality is about fitness for use, and Google names six dimensions it is judged on: accuracy, completeness, consistency, timeliness, validity and uniqueness. It matters because, as Google's big data page puts it, data quality directly impacts the quality of decision-making, data analytics and planning strategies — so bad data does not just sit there, it produces bad decisions. And compliance is about obligations on both sides of the organization's boundary: governance is necessary to assure that data is safe, secure, private, usable, and in compliance with both internal and external data policies. When an exam item describes a data problem, ask which of the three it is — a person question, a quality question, or a policy question.
What this topic actually tests
Four discriminations, six objectives
Which store? database for related records, warehouse for known reporting, lake for raw and unstructured data. Which source? integrate what you hold, stream what you lack, subscribe to what others hold. Which stage? value is realized at action, not ingestion. What does governance do? controls that make greater access safe.
Close the topic on four discriminations. First, which store: a relational database structures related records in tables, a warehouse serves known, repeatable reporting over current and historical data, and a lake keeps raw data of every type in its native format, which is what machine learning needs. Second, which source of value: integrate the data you already hold, stream the data you are not yet collecting, and subscribe to data someone else holds through a data exchange. Third, which stage of the chain: data is acquired, processed, stored and analyzed, but its value is realized when someone acts on it — which is why Google wants warehouses that can activate on data in near-real time. Fourth, what governance is for: not to lock data away, but to set the controls that make greater access safe, because the value of data is only realized when it is trustworthy, discoverable and governed. Topic two turns these ideas into Google Cloud's products for managing data.
Official sources for this topic
- Cloud Digital Leader exam guide — Section 2
- What is data governance? — Google Cloud
- What is a data warehouse? — Google Cloud
- What is business intelligence? — Google Cloud
- What is big data? — Google Cloud
- What is a relational database? — Google Cloud
- What is a data lake? — Google Cloud
- What is data integration? — Google Cloud
- Introduction to BigQuery sharing — BigQuery
- What is ETL? — Google Cloud