Hands-on Lab524 words

Decide which data a support assistant may learn from — decision exercise

Exercise: decide which data a support assistant may learn from

Original fictional scenario. No cloud account, API calls or paid services are required. Difficulty: beginner · Estimated duration: 15–20 minutes

A furniture retailer wants to tune a support assistant that answers short chat messages about orders and returns. Its data team proposes four sources. All names and numbers below are invented.

SourceExamplesMissing a required fieldAgrees with the order systemLast updatedLabeled with the right answerLooks like production chats
A. Chat transcripts40,0002%YesLast monthYesYes
B. Formal email archive90,0001%YesLast monthYesNo — long formal letters
C. Returns spreadsheet12,000 rows35% lack a product codeNo — refund amounts differLast monthn/aNo
D. Product photos8,000 images0%YesThree years agoNoNo

Your decision

  1. For each source, name every data-fitness check it fails: completeness, consistency, relevance (recent and like production), format, or labels.
  2. Classify each source as structured or unstructured, with the reason.
  3. Choose the source(s) to tune on first, and explain why source B is not first despite being largest.
  4. Source C is valuable but 35% of rows lack a product code. Describe your two options for those rows and when each fits.

Rubric (10 house points)

  • 3 points: correct failed checks for all four sources.
  • 2 points: correct structured/unstructured classification with reasons.
  • 3 points: chooses A first and explains B's mismatch with production.
  • 2 points: names deleting or imputing incomplete rows, never training on them as they are.

Reference solution

A passes every check: complete, consistent, recent, labeled and like production. B is large and labeled, but it fails relevance: tuning data should reflect the format and context the model meets in production, and long formal letters do not look like short chats. C fails completeness (35% missing a code) and consistency (refund amounts disagree with the order system). D is out of date for current products and unlabeled, and it does not match a text chat task.

A and B are unstructured text; D is unstructured images; C is structured, rows and columns in a spreadsheet.

Tune on A first. Size does not rescue a mismatch: high-quality data matched to the task beats a larger set that is not. Rejected alternative: "use B because it is biggest".

For C's incomplete rows, either delete them, if enough complete rows remain to be useful, or impute the missing code with a well-reasoned guess when they do not. Never train on incomplete rows as they are. Fix the refund disagreement first, because consistent values across systems are part of quality.

Sources and scope

All company names, sources, numbers and rubric points above are house-authored. The data-quality principles applied are grounded in:

Ready to study Generative AI Leader (GCP-GAIL)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free