Decide which data a support assistant may learn from — decision exercise
Exercise: decide which data a support assistant may learn from
Original fictional scenario. No cloud account, API calls or paid services are required. Difficulty: beginner · Estimated duration: 15–20 minutes
A furniture retailer wants to tune a support assistant that answers short chat messages about orders and returns. Its data team proposes four sources. All names and numbers below are invented.
| Source | Examples | Missing a required field | Agrees with the order system | Last updated | Labeled with the right answer | Looks like production chats |
|---|---|---|---|---|---|---|
| A. Chat transcripts | 40,000 | 2% | Yes | Last month | Yes | Yes |
| B. Formal email archive | 90,000 | 1% | Yes | Last month | Yes | No — long formal letters |
| C. Returns spreadsheet | 12,000 rows | 35% lack a product code | No — refund amounts differ | Last month | n/a | No |
| D. Product photos | 8,000 images | 0% | Yes | Three years ago | No | No |
Your decision
- For each source, name every data-fitness check it fails: completeness, consistency, relevance (recent and like production), format, or labels.
- Classify each source as structured or unstructured, with the reason.
- Choose the source(s) to tune on first, and explain why source B is not first despite being largest.
- Source C is valuable but 35% of rows lack a product code. Describe your two options for those rows and when each fits.
Rubric (10 house points)
- 3 points: correct failed checks for all four sources.
- 2 points: correct structured/unstructured classification with reasons.
- 3 points: chooses A first and explains B's mismatch with production.
- 2 points: names deleting or imputing incomplete rows, never training on them as they are.
Reference solution
A passes every check: complete, consistent, recent, labeled and like production. B is large and labeled, but it fails relevance: tuning data should reflect the format and context the model meets in production, and long formal letters do not look like short chats. C fails completeness (35% missing a code) and consistency (refund amounts disagree with the order system). D is out of date for current products and unlabeled, and it does not match a text chat task.
A and B are unstructured text; D is unstructured images; C is structured, rows and columns in a spreadsheet.
Tune on A first. Size does not rescue a mismatch: high-quality data matched to the task beats a larger set that is not. Rejected alternative: "use B because it is biggest".
For C's incomplete rows, either delete them, if enough complete rows remain to be useful, or impute the missing code with a well-reasoned guess when they do not. Never train on incomplete rows as they are. Fix the refund disagreement first, because consistent values across systems are part of quality.
Sources and scope
All company names, sources, numbers and rubric points above are house-authored. The data-quality principles applied are grounded in:
- https://developers.google.com/machine-learning/crash-course/overfitting/data-characteristics
- https://docs.cloud.google.com/knowledge-catalog/docs/auto-data-quality-overview
- https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning
- https://developers.google.com/machine-learning/crash-course/overfitting/labels
- https://cloud.google.com/learn/what-is-big-data
- https://developers.google.com/machine-learning/intro-to-ml/supervised