Hands-on Lab594 words

Write a production monitoring plan — decision exercise

Exercise: write a production monitoring plan

Original fictional scenario. No cloud account, API calls or paid services are required. Difficulty: intermediate · Estimated duration: 20 minutes

An insurer runs a gen AI assistant that summarises claim files for adjusters. The business goal is to settle straightforward claims faster. A colleague proposes the plan below. All numbers are invented.

Proposed itemDetail
Success measure"The model's grounding score"
MonitoringCheck outputs once a quarter, by hand
Model upgradesProduction calls the registry default version; the platform team moves the default when a newer version is approved
AlertsNone; adjusters will report problems
Feature dataA claims-risk model alongside the assistant reads features from ad-hoc spreadsheets

Your decision

  1. Replace the success measure with a KPI derived from the business goal, and place the grounding score where it belongs.
  2. Propose continuous measurements in three layers (output quality, service health, human judgment), naming at least two metrics in each automated layer.
  3. Explain the risk in the upgrade arrangement and design the release steps for the next model version.
  4. Propose alerts and the action each one should trigger.
  5. Recommend what to do about the claims-risk model's feature data.

Rubric (10 house points)

  • 2 points: a business KPI (for example, days to settle a straightforward claim) with grounding moved to the output-quality layer.
  • 2 points: continuous output metrics (quality, safety, grounding) and endpoint metrics (request rate, error rate, latency, utilization), plus human-in-the-loop review.
  • 3 points: naming that a moved default upgrades production silently; regression tests to at least the old quality, a subset release first, and a rollback plan.
  • 2 points: threshold alerts on output-quality shifts and drift, each mapped to an action.
  • 1 point: Feature Store for consistent, versioned feature data with feature monitoring for drift.

Reference solution

KPI. Business objectives are broken down into measurable KPIs: here, median days to settle a straightforward claim. The grounding score is an output-quality metric that helps explain the KPI; it is not the KPI.

Layers. Continuously monitor output with the Gen AI evaluation service — response quality, safety, instruction adherence and grounding. Track the endpoint's request rate, error rate, latency and resource utilization with anomaly alerts. Add human-in-the-loop review of a weekly sample, and check that the automated scores agree with the reviewers. Rejected: quarterly manual checks only — behaviour can change over time as data, environment and models change.

Upgrades. Because production calls the default version, moving the default changes what production serves with no code change — a silent upgrade. Before any move, run regression tests so the new version at least maintains the old quality; release to a subset of users first; keep a rollback plan to the most recent stable version. Rejected: "it is newer, so it is better" — changes are hard to predict without testing prompts on the new version.

Alerts. Alert on shifts in output quality or harmful content (action: roll back or fix the prompt) and on drift in the claims-risk model's inputs (action: retrain through a pipeline). Rejected: waiting for adjusters to report problems.

Feature data. Move the features into Agent Platform Feature Store for consistency and versioning, and schedule feature monitoring jobs to detect feature drift.

Sources and scope

All names, numbers, thresholds and rubric points above are house-authored. The practices applied are grounded in:

Ready to study Generative AI Leader (GCP-GAIL)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free