Review a hiring assistant before launch — decision exercise
Exercise: review a hiring assistant before launch
Original fictional scenario. No cloud account, API calls or paid services are required. Difficulty: intermediate · Estimated duration: 20 minutes
Larkfield Logistics wants an assistant that ranks job applicants and drafts feedback letters. The project team's launch checklist is below. All names and facts are invented.
| Item | Team's position |
|---|---|
| Training data | Ten years of past hiring decisions, mostly from two university programmes |
| Applicant data | Names replaced with reversible tokens; the key is held by the same team |
| Fairness testing | One overall accuracy figure of 91% |
| Explanations | "The model is accurate, so explanations are not needed" |
| Oversight | Rankings sent straight to applicants, no human review |
| Safety | "The platform's built-in filters cover it" |
Your decision
- Identify the responsible-AI problem in each row and the risk it creates.
- Say which privacy technique is in use and whether the original names can come back.
- Name the bias most likely in the training data and how the team should test for it.
- List what must change before launch, in priority order.
Rubric (10 house points)
- 3 points: a correct problem and risk for each row.
- 2 points: identify two-way pseudonymization and the key-holder risk.
- 2 points: identify historical (or selection) bias and slice-based evaluation.
- 3 points: a defensible pre-launch list including human supervision and accountability.
Reference solution
- Training data reflects past decisions favouring two programmes — likely historical bias (and selection bias in who applied). Audit the data and evaluate predictions across slices before production; collect additional data where groups are missing.
- Applicant data uses two-way pseudonymization: the same key that creates tokens can reverse them, so whoever holds it can recover names. Separate key custody from the team using the data, and minimize the fields kept.
- One overall accuracy figure can hide a group the model serves badly; define fairness measures for the use case.
- No explanations: rankings that affect people need feature attributions or other explainability techniques, plus a named owner (accountability).
- No human review: hiring can materially affect individual rights, so Google calls for meaningful human supervision.
- Built-in filters: they do not cover use-case risks; the team must run safety testing for its own use case.
Priority: human review and fairness evaluation first (they block harm to applicants), then explainability and ownership, then key custody and data minimization.
Rejected alternatives. Launching because accuracy is high ignores bias across groups. Relying on built-in filters skips the customer's own risk assessment. Treating tokenized names as anonymous ignores that two-way tokens can be reversed.
Sources and scope
All names and facts above are house-authored. The general principles are grounded in:
- https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/responsible-ai
- https://docs.cloud.google.com/sensitive-data-protection/docs/pseudonymization
- https://docs.cloud.google.com/sensitive-data-protection/docs/concepts-risk-analysis
- https://developers.google.com/machine-learning/crash-course/fairness/types-of-bias
- https://docs.cloud.google.com/architecture/framework/perspectives/ai-ml/security