Unit 1.5 study guide — Judge whether data is fit for the task
Generative AI Leader › Unit 1 › Topic 5
Judge whether data is fit for the task
Study guide for Generative AI Leader, Unit 1 · Topic 5. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks, quoted. Explaining the characteristics and importance of data quality and data accessibility in AI (e.g., completeness, consistency, relevance, availability, cost, format). Identifying the differences between structured and unstructured data, and identifying real-world examples of each type. Identifying the differences between labeled and unlabeled data.
This hive's learning objectives for the topic:
- Evaluate completeness, consistency, relevance, availability, cost and format before using data.
- Distinguish structured and unstructured data using business examples.
- Distinguish labeled and unlabeled examples and their learning uses.
Judging whether data is fit for the task
Before a model learns from data, decide whether the data deserves it
Is it good enough? completeness, consistency, relevance, availability, cost and format. What kind is it? structured rows or unstructured text, images, audio and video. Does it carry answers? labeled examples teach; unlabeled examples are what the model later predicts on.
Every generative AI project rests on data, and this topic is about judging that data before anyone builds on it. Google's machine learning course puts the stakes in one line: regardless of the format, your ML model — your machine learning model — is only as good as the data it trains on. The guide asks three things of a generative AI leader here. First, is the data good enough? It names completeness, consistency, relevance, availability, cost and format, and we will map each onto what Google documents. Second, what kind of data is it — structured, like the rows of a spreadsheet, or unstructured, like text, images, audio and video? Third, does the data carry answers? Labeled examples contain the answer a model should learn; unlabeled examples do not. Each question changes what a project can realistically do, and how much it costs.
Six checks before data reaches a model
Quality is fitness for a purpose, not a property of data in general
- Completeness: does it hold everything the purpose requires?
- Consistency: do the same facts agree across tables and systems?
- Relevance: is it recent, and does it match what production will see?
- Availability: can the people building the model actually access it?
- Cost and format: what do labels cost, and does it meet format standards?
Worked example (synthetic). A retailer wants a returns assistant trained on three years of support tickets. Half the tickets lack the product code, and the ticket system and the order system disagree on refund amounts — two failed checks before any model is chosen.
The first objective is to evaluate data before using it, and Google's course defines quality pragmatically: a high-quality dataset helps your model accomplish its goal. That makes quality a question about purpose. Completeness, in Google's data catalog, assesses whether the data contains all of the information that's required for its intended purpose. Consistency refers to having the same values across multiple instances such as tables and columns — Google's example is a product whose revenue differs between a sales database and a usage database. The guide's word relevance maps onto two things Google documents: freshness, which helps you determine whether the data is recent enough to be useful, and Google's tuning advice that training data should reflect the prompt distribution, format and context the model will encounter in production. Availability is access: Google describes modern data governance as ensuring that data scientists and data engineers can access high-quality data to build accurate models and agents. Format is validity — whether data conforms to standards for format and acceptable ranges. And cost shows up most sharply in labels: Google notes you typically pay human raters, so human-generated data can be expensive. A data source that fails a check the use case depends on is not a weaker option; it is a blocked one until it is fixed.
The guide's words, mapped to what Google documents
Each check is a question you can answer before training starts
| Guide's word | Question to ask | Google's term |
|---|---|---|
| Completeness | Does it hold all that the purpose needs? | Completeness |
| Consistency | Do two systems agree on the same fact? | Consistency |
| Relevance | Is it recent, and like production? | Freshness; match production |
| Availability | Can the builders reach it? | Governed access |
| Format | Does it meet the standards set? | Validity |
| Cost | What will labels cost? | Paid human raters |
Worked example (synthetic). A bank's model must read loan applications from two systems. One stores dates as day-month-year, the other as month-day-year — a validity and consistency problem to fix before training, not after.
Here the six words the guide lists are set beside the terms Google actually uses, because an exam item may use either vocabulary. Completeness and consistency are Google's own dimension names: does the data hold everything the purpose needs, and do two systems agree on the same fact? Relevance is the guide's word; Google's nearest dimension is freshness — whether the data is recent enough to be useful — together with the tuning advice that training data should match what the model will see in production. Availability is a governance question: can the people building the model reach high-quality data at all? Format corresponds to validity, which evaluates whether data conforms to standards for format, acceptable ranges or other criteria. And cost is most visible in labelling, because human raters are paid. Google's own list of data-quality dimensions — accuracy, completeness, consistency, timeliness, validity and uniqueness — overlaps the guide's six without matching it word for word, which is why reading the question carefully matters more than memorising one list.
What to do with incomplete examples
Never train on them as they are: delete them or fill the gaps
Figure: A flowchart. An example with a missing value leads to a question: do enough complete examples remain? Yes leads to deleting the incomplete example; no leads to imputing a well-reasoned guess for the gap. Both paths lead to training on complete examples. A dotted arrow shows automation flagging unreliable data before the question.
Worked example (synthetic). A clinic has 20,000 complete patient-intake forms and 300 missing an age field. With so many complete forms, deleting the 300 costs little; with only 400 forms in total, imputing would be the better fix.
Completeness has a concrete rule in Google's course: don't train a model on incomplete examples. There are two ways to fix them. You can delete the incomplete examples, which is sensible when enough complete examples remain to train a useful model. Or you can impute missing values — in Google's words, convert the incomplete example to a complete example by providing well-reasoned guesses for the missing values — which is the fallback when deleting would leave too little data. Before either, find the problems at scale: Google recommends using automation to flag unreliable data, since reliability is the degree to which you can trust your data. For a generative AI leader the lesson is about planning. Data repair is real work with a real schedule, and Google's application guidance is blunt about skipping it: if you give the model suboptimal data, like inaccurate or incomplete data, you can't expect optimal performance.
Structured and unstructured data
Rows and columns, or text, images, audio and video
- Structured: data in spreadsheets or relational databases
- Unstructured: text, images, audio and video files
- Semi-structured: formats like sensor data without a fixed schema
- Combining both gives more useful insight than either alone
Worked example (synthetic). An airline holds structured booking records and unstructured complaint emails. A model that reads both can link a complaint to the flight and fare it concerns, which neither source does alone.
The second objective is telling structured data from unstructured data, with real examples. Google's description is a contrast: more traditional structured data, such as data in spreadsheets or relational databases, is now supplemented by unstructured text, images, audio and video files, or semi-structured formats like sensor data that can't be organized in a fixed data schema. So structured data has a fixed shape — named columns, typed values — while unstructured data is content whose meaning lives inside the file. Generative AI changes the business value of the second kind, because models can read documents, look at images and listen to audio. And the two are most useful together: Google says combining structured data sources with unstructured ones provides more useful insights for consumer understanding and personalization. On Google Cloud, unstructured files do not have to stay out of analysis either — BigQuery object tables exist primarily to make unstructured data accessible for analysis.
Business examples of each kind
Ask where the meaning lives: in a column, or inside the file
| Kind | Business examples | Where the meaning lives |
|---|---|---|
| Structured | Orders table, payroll spreadsheet | Named columns, fixed schema |
| Unstructured | Contracts, call recordings, product photos | Inside the file's content |
| Semi-structured | Sensor readings | Fields without a fixed schema |
Worked example (synthetic). An insurer asks which of its data a model could use to spot damage in claims. The claims table says what was claimed; the photos show what happened — the photos are the unstructured source the use case needs.
Business examples make the distinction stick. An orders table or a payroll spreadsheet is structured: its meaning sits in named columns with a fixed schema, which is exactly what Google means by data in spreadsheets or relational databases. Contracts, call-centre recordings and product photos are unstructured: the meaning is inside the content, in the text, the audio or the image. Sensor readings are Google's example of semi-structured data — formats that can't be organized in a fixed data schema. The useful question for an exam scenario is where the meaning lives. If the answer is a column, the data is structured; if you have to read, look or listen to get it, the data is unstructured.
One customer, two kinds of data
Each kind answers a question the other cannot
Figure. Two cards. The structured card shows an order record with two items, refunded, and notes that columns tell you what happened. The unstructured card shows a customer's complaint about a crushed box and a missing lid, and notes that text tells you why. A band beneath reads: together they reveal refunds caused by packaging damage, an insight neither source gives alone.
Worked example (synthetic). A synthetic order record and complaint. Counting refunds needs only the table; explaining them needs the text.
This invented example shows why Google recommends combining the two kinds. The structured record says what happened: an order of two items, refunded. The unstructured complaint says why: the box arrived crushed. A report built on the table alone counts refunds; a report built on the emails alone tells stories without totals. Put together, they show how many refunds packaging damage caused — the kind of more useful insight Google says comes from combining structured data sources with unstructured ones. For a business leader, the practical point is that unstructured data is often where the reasons are, and that generative AI is what makes those reasons readable at scale.
Labeled and unlabeled examples
A label is the answer; an unlabeled example has none
- Labeled example: features plus a label, the value to predict
- Unlabeled example: features, but no label
- A trained model predicts labels for new, unlabeled data
- Semi-supervised learning uses data where only some is labeled
- Labels can be costly: human raters are paid
Worked example (synthetic). A bank has 50,000 past transactions marked fraud or not fraud — labeled data. Tomorrow's transactions arrive unmarked, and the trained model's job is to supply that missing label.
The third objective separates labeled from unlabeled data and connects each to how it is used. Google's machine learning course defines the parts: features are the values that a supervised model uses to predict the label, and the label is the answer — the value the model should predict. Examples that contain both features and a label are called labeled examples. In contrast, unlabeled examples contain features, but no label. The two have different jobs. Labeled examples are what a supervised model learns from. Unlabeled data is what a trained model works on: once trained and evaluated, models can be used for inference, making predictions on new, unlabeled data in real-world applications. Between the two sits a mixed approach Google describes, semi-supervised learning, in which only some data is labeled. Labels also carry a cost, because you typically pay human raters, so human-generated data can be expensive — which is why the quality of labels matters more than their number: Google's tuning guidance says high-quality, well-labeled data is better than quantity.
Where labeled and unlabeled data go
Labels teach the model; unlabeled data is what it then predicts on
Figure: A flowchart. Labeled examples, made of features and a label, feed training and evaluation, which produces a trained model. New unlabeled data, features only, goes into the trained model, which outputs a predicted label.
Worked example (synthetic). A support team labels 5,000 past emails by department, trains a router, then lets the router label each new email as it arrives.
The flow puts each kind of data in its place. Labeled examples — features plus the answer — go into training and evaluation, and a trained model comes out. New data then arrives without labels, because in real use nobody knows the answer yet; the model's job is to predict it. That is Google's description of inference: once trained and evaluated, models make predictions on new, unlabeled data. So a scenario that asks what data is needed to build a classifier is asking about labeled examples, and one that asks what the model processes after launch is describing unlabeled data. Good training data has two more properties Google names: good datasets are both large and highly diverse.
What this topic actually tests
Three questions to ask of any dataset
Is it fit for this purpose? complete, consistent, recent, reachable, valid, affordable to label. Where does its meaning live? in columns, or inside files. Does it carry the answer? labeled to learn from, unlabeled to predict on.
Close on three questions to ask of any dataset a project proposes. First, is it fit for this purpose? Check completeness, consistency, relevance — recent and like production — availability, format and the cost of labelling, and remember that incomplete examples are fixed by deleting or imputing, never trained on as they are. Second, where does its meaning live? In columns, which makes it structured, or inside files of text, images, audio and video, which makes it unstructured — and combining the two usually tells you more. Third, does it carry the answer? Labeled examples teach a model; unlabeled data is what the trained model predicts on. A machine learning model is only as good as the data it trains on, so these questions come before any choice of model.
Official sources for this topic
- Datasets: Data characteristics — Machine Learning Crash Course
- What is big data? — Google Cloud
- Supervised learning — Introduction to Machine Learning
- Auto data quality overview — Knowledge Catalog
- Introduction to tuning — Gemini Enterprise Agent Platform
- What is data governance? — Google Cloud
- Datasets: Labels — Machine Learning Crash Course
- Introduction to object tables — BigQuery
- What is machine learning? — Google Cloud
- Develop a generative AI application — Google Cloud