Triage a city services assistant — decision exercise
Exercise: triage a city services assistant
Original fictional scenario. No cloud account, API calls or paid services are required. Difficulty: intermediate · Estimated duration: 15–20 minutes
A city launches an assistant for residents. In its first month, reviewers log five problem outputs. All details below are invented.
| # | What the assistant did | Who is affected |
|---|---|---|
| 1 | Quoted last year's parking-permit fee as current | Residents paying fees |
| 2 | Cited a "Section 14b" of the noise bylaw that does not exist | Residents planning events |
| 3 | Gave noticeably worse answers to residents writing in a regional dialect | One community |
| 4 | Misread a request about a mobility-scooter permit for a houseboat mooring | One resident |
| 5 | Drafted a decision letter refusing a housing-benefit appeal | An appellant |
Your decision
- Name the limitation each output shows: knowledge cutoff, hallucination, bias and fairness, edge case — or none of these.
- For outputs 1–4, choose the practice that most directly addresses the cause: grounding, retrieval-augmented generation (RAG), prompt engineering, fine-tuning or human review. Give one practice you would reject for each, and why.
- Decide the review level for output 5, and for the assistant's routine answers about opening hours.
- The city proposes fine-tuning the model on last year's council documents to fix output 1. Respond.
Rubric (10 house points)
- 3 points: correct limitation for outputs 1–4, with the symptom that identifies it.
- 3 points: a matching practice for each, plus a rejected alternative with a reason.
- 2 points: human review before release for output 5, tied to the effect on a person's rights; monitoring and feedback for routine answers.
- 2 points: rejecting the fine-tuning plan for output 1, because it would teach the outdated fee.
Reference solution
- Knowledge cutoff — an answer that was true once; the model is limited to its pre-trained data. 2. Hallucination — a specific, fluent, invented citation. 3. Bias and fairness — quality that is worse for one group of users. 4. Edge case — a rare combination the training data barely covers, misread with confidence. Output 5 is not a limitation by itself; it is a high-stakes use.
For output 1, use RAG over the current fee schedule, which supplies up-to-date information; reject fine-tuning (see task 4). For output 2, ground answers on the published bylaws so output is tied to verifiable sources and invented content is reduced; reject "ask the model to be more careful" in the prompt, which cannot supply the real bylaw text. For output 3, investigate with evaluation across user groups, then improve prompts or data and add human review for that community while it is fixed; reject "do nothing because average accuracy is fine", since an unfair model can perform worse for certain slices. For output 4, route rare permit combinations to a person, who brings the judgment and context the model lacks; reject fine-tuning on so rare a case.
Output 5 affects an appellant's rights, so a caseworker reviews every letter before release. Routine opening-hours answers ship with safety testing, user feedback and content monitoring — no per-answer approval.
Task 4: fine-tuning on last year's documents would build depth on outdated material and bake the old fee in deeper. Retrieval of the current documents addresses the cause.
Sources and scope
All names, outputs, stakes and rubric points above are house-authored. The limitations, practices and review rule applied are grounded in:
- https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/responsible-ai
- https://cloud.google.com/discover/what-are-ai-hallucinations
- https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/grounding/overview
- https://cloud.google.com/use-cases/retrieval-augmented-generation
- https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning
- https://cloud.google.com/discover/human-in-the-loop