Study Guide3,325 words

Unit 3.2 study guide — Monitor and evaluate gen AI in production

Generative AI Leader › Unit 3 › Topic 2

Monitor and evaluate gen AI in production

Study guide for Generative AI Leader, Unit 3 · Topic 2. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.

What the exam guide asks, quoted. Recognizing Google-recommended practices for continuous monitoring and evaluation of gen AI models (e.g., automatic model upgrades, key performance indicators, security patches and updates, versioning, performance tracking, drift monitoring, Agent Platform Feature Store).

This hive's learning objectives for the topic:

  1. Choose key performance indicators and evaluation methods for a gen AI application.
  2. Plan for versioning, automatic model upgrades and security patches.
  3. Explain drift monitoring and performance tracking, including the role of Agent Platform Feature Store.

Gen AI in production

Measure it, change it safely, notice when the world moves

Measure — KPIs from business objectives, plus continuous checks on output quality, safety and service health. Change safely — versions, upgrades and patches, tested and released gradually with a way back. Notice — drift and degradation, caught by monitoring before users do.

The previous topic designed around a model's limitations. This one follows the model into production, where Google's operational guidance starts from an uncomfortable fact: the behavior of AI and machine learning systems can change over time due to changes in the data or environment and updates to the models. So a launched model is not finished; it is watched. The guide lists the practices — key performance indicators, automatic model upgrades, security patches and updates, versioning, performance tracking, drift monitoring and Agent Platform Feature Store — and they sort into three jobs. Measure: decide what success means for the business and keep measuring it after launch. Change safely: every new model version, prompt change or patch is tested, released gradually and reversible. Notice: detect when production data drifts away from what the model learned, or when output quality slips, and alert someone before customers notice. We take the three jobs in that order.

What to measure once the model is live

Start from the business objective, then watch quality and health

  • Break business objectives down into measurable KPIs
  • Continuously evaluate output quality, safety and grounding
  • Track request rate, error rate, latency and resource use
  • Add human-in-the-loop review for quality and relevance
  • Keep automated metrics aligned with human judgment

Worked example (synthetic). A support team's objective is faster resolution. Its KPI is minutes per resolved ticket; underneath sit an automatic grounding score on every reply, an endpoint latency alert, and a weekly review of fifty replies by senior agents.

The first objective is choosing what to measure in production, and Google's operational guidance is layered. The top layer is the business: start with an outline of the business objectives and break them down into measurable key performance indicators, or KPIs. Those keep the project honest about why the model exists. The second layer is the model's output. For generative AI, Google recommends using the Gen AI evaluation service to continuously monitor model output in production, and says you can enable automatic evaluation metrics for response quality, safety, instruction adherence, grounding, writing style and verbosity — chosen, it adds, to align with the intended use case and the potential risks. The third layer is service health: for endpoints, track request rate, error rate, latency and resource utilization, and set up alerts for anomalies. And the fourth is people: to assess outputs for quality, relevance, safety and adherence to guidelines, you can incorporate human-in-the-loop evaluation. One warning ties the layers together. Automated metrics are only useful if they agree with people, so make sure your automated evaluation aligns with human judgment.

Four layers of production measurement

A KPI says whether it matters; the other layers say why it moved

LayerWhat you measureHow Google suggests measuring it
BusinessKPIs broken down from business objectivesDefined before the project starts
Output qualityQuality, safety, instruction adherence, groundingGen AI evaluation service, continuously
Service healthRequest rate, error rate, latency, utilizationEndpoint metrics with anomaly alerts
Human judgmentQuality, relevance, adherence to guidelinesHuman-in-the-loop evaluation

Worked example (synthetic). A KPI for an HR assistant dips one week. The output-quality layer shows grounding scores fell after a policy library moved, while latency and error rates held steady — so the fix is the data source, not the model or the servers.

Laid out as layers, the measurements answer different questions. The business layer — key performance indicators, or KPIs, broken down from business objectives — tells you whether the application still matters, but not why a number moved. The output-quality layer, measured continuously with the Gen AI evaluation service, covers response quality, safety, instruction adherence and grounding. The service-health layer covers the endpoint's request rate, error rate, latency and resource utilization, with alerts for anomalies. And the human layer, human-in-the-loop evaluation, checks quality, relevance and adherence to guidelines in ways a metric may miss. The value of keeping all four is diagnosis. When a KPI falls, the lower layers say whether the cause is the content, the service or something people notice and metrics do not — and each points to a different fix.

From objective to alert

Each alert should trace back to something the business cares about

Figure. Four cards in a row: Objective — resolve customer questions faster; KPI — minutes per resolved ticket; Metrics — grounding score, safety and 95th-percentile latency; Alert — grounding below 0.8 for an hour. A band beneath says to read right to left: an alert that cannot be traced to a KPI and an objective is noise, and a KPI with no metrics beneath it cannot be diagnosed.

Worked example (synthetic). Invented thresholds for one support assistant. The numbers are a team's choice; the structure — objective, KPI, metrics, alert — is the guidance.

Here the layers become a chain for one invented support assistant. The objective is to resolve customer questions faster. Following Google's advice, it is broken down into a measurable key performance indicator, or KPI: minutes per resolved ticket. Beneath the KPI sit the metrics that explain it — a grounding score and a safety score from continuous evaluation, and the endpoint's latency — each picked, as Google advises, to align with the intended use case and its potential risks. And at the end of the chain is an alert, set at a threshold the team chose. Read the chain backwards to test a monitoring plan. An alert you cannot trace to a KPI and an objective is noise that teams learn to ignore. A KPI with nothing beneath it tells you that something went wrong but never what. The thresholds on this slide are invented; the structure is the part to remember.

Changing a running model safely

Every new version, prompt or patch is tested, staged and reversible

  • Version models, prompts and data so a change can be traced
  • Retirements come with at least 45 days' notice to migrate
  • A new model version must at least match the old one's quality
  • Release gradually, with a rollback plan to a stable version
  • Find and apply security patches to your code and containers

Worked example (synthetic). A team's prompt was tuned against one model version. When it moves to a newer version, a regression run on 300 saved prompts shows two task types got worse, so it adjusts the prompt before any customer sees the change.

The second objective is change. Google's operations guidance assumes models will change — developers constantly adjust foundation models through prompting techniques and by swapping the models out for newer versions — so the question is how to change them safely. Versioning comes first. Model versioning lets you create multiple versions of the same model, and Google notes that a prompt can work well with one model version and not as well with another, so experiments must record the prompt, component versions, model version, metrics and output data. Upgrades come with notice: when Google schedules a model for retirement, it posts a fixed date that gives you at least forty-five days to migrate, and the date will not be moved earlier. Before switching, test: it is hard to predict changes without first testing your prompts on the new version, and the regression bar is that the new version at least maintains the previous version's quality. Release gradually, to a subset of users first, and keep a robust rollback plan to the most recent stable version. And security is part of change too: find and apply all security patches for vulnerabilities in your code or system, including custom model containers.

Pinned versions and moving references

An upgrade with no code change is effortless — and silent

How production chooses the modelWhat happens when a new version appearsThe control you need
A pinned model versionNothing — it stays active until a retirement date is announcedPlan the migration inside the notice period
A moving reference (default or alias)Production follows the reference without a code changeRegression tests and monitoring on every move
A tuned modelCannot move to the new versionRun a new tuning job, then test

Worked example (synthetic). Two teams share a model in Model Registry. One calls a fixed version; the other calls the default. When someone sets a newer version as default, only the second team's application changes — overnight, and without a release note.

The guide's phrase automatic model upgrades is easiest to understand by asking how production picks a model. If your application pins a specific version, nothing changes on its own: Google says that even after a replacement launches, a model remains active until a retirement date is announced — and then you have the notice period to migrate. If production follows a moving reference, upgrades are automatic. In Model Registry, if you don't specify a model version for production, the default model is used, and an alias is a mutable, named reference that can be moved to another version, letting you deploy by reference without knowing the version's ID. Moving that reference upgrades production with no code change, which is convenient and also silent, so every move needs the regression tests and monitoring a migration would get. Tuned models are the exception that catches teams out: Google says you cannot move an existing tuned model to a newer Gemini version — you run a new tuning job for it.

A safe upgrade, step by step

Test offline, release to a few, watch, then decide

Loading Diagram...
Figure 1 — Mermaid diagram

Figure: A left-to-right flowchart. A new model version or prompt change goes through regression tests asking whether quality is at least as good. If not, fix prompts or stay on the current version. If so, release to a subset of users, then ask whether monitoring is healthy: yes leads to full rollout, no leads to rolling back to the most recent stable version.

Worked example (synthetic). A new version passes offline regression tests, goes to 5% of users, and within a day the safety alert fires on one language. The team rolls back to the stable version and investigates before trying again.

Put together, a safe upgrade has two gates. The first is offline: Google recommends thorough testing before fully migrating, and its regression bar is that the new version provides outputs that at least maintain the previous version's quality. A version that fails is fixed — often by adjusting prompts, since a prompt tuned to one version may not suit another — or not adopted yet. The second gate is live. For applications that use managed models like Gemini, Google recommends gradually releasing new application versions to a subset of users before the full deployment, and then watching. If monitoring stays healthy, roll out to everyone. If it does not, the robust rollback plan Google asks for comes into play: Model Registry's lineage lets you revert to the most recent stable version. Security patches follow the same path in miniature — tested, released, and reversible — because a patch is a change like any other.

Drift and performance tracking

The model stays the same while the world it serves moves

  • Drift: production data deviates from the data the model learned from
  • Watch input drift, output drift and changing feature attributions
  • Run monitoring on a schedule; alert when a threshold is passed
  • For gen AI, alert on shifts in output quality or harmful content
  • Feature Store keeps feature data consistent, versioned and monitored

Worked example (synthetic). A retailer's demand model was trained before a new loyalty scheme. Shoppers' basket sizes shift, the monitoring job's input-drift score crosses its threshold, and the alert reaches the team weeks before forecast errors would have shown up in stock reports.

The third objective is noticing change you did not make. Google's example is a model that predicts customer lifetime value: as customer habits change, the factors that predict spending change too, and the features the model was trained on may no longer be relevant. This deviation in the data is known as drift. After deployment, input data can deviate from the training data, or shift significantly over time, and predicted outcomes can shift too. Feature attributions add a third view, showing how much each feature contributed to each inference, so a change in what drives decisions is visible. Model Monitoring runs monitoring jobs as needed or on a regular schedule, and if you set alerts it informs you when metrics surpass a specified threshold. For generative AI, Google suggests tracking relevant metrics and alerting on shifts in output quality or the emergence of harmful content, using model evaluation with Cloud Logging and Cloud Monitoring. And Agent Platform Feature Store is the guide's named tool for the data side: Google recommends it to maintain the consistency and versioning of feature data, and it can run feature monitoring jobs that detect feature drift.

Watching a model in production

Serve and log; check on a schedule; alert at a threshold

Figure. Users send requests to a model endpoint inside a Google Cloud project. The endpoint's production data is logged. A group of scheduled monitoring jobs reads the logged data: one runs output quality checks, the other drift monitoring. Both send results to threshold alerts. A callout says an alert triggers retraining, a rollback or a prompt fix.

Worked example (synthetic). An insurer's claims assistant logs every request and response. A nightly job scores grounding and safety; another compares this week's inputs with the training data. One morning the drift job alerts, and the team retrains before accuracy visibly falls.

This diagram shows the shape monitoring takes. Users send requests to the model endpoint, and production data is logged as it is served. Monitoring does not sit in the request path: Model Monitoring lets you run monitoring jobs as needed or on a regular schedule, reading what was logged. Two kinds of check run side by side. Output quality checks, for generative AI, follow Google's suggestion to track relevant metrics with model evaluation along with Cloud Logging and Cloud Monitoring. Drift monitoring compares production data with what the model learned from — Google recommends Model Monitoring to detect data drift and performance degradation in traditional machine learning systems. Both feed alerts at the thresholds you set. And the callout is the point of the whole loop: an alert should trigger an action. Google describes feedback loops that automatically retrain models with Agent Platform Pipelines when Model Monitoring triggers an alert; for a generative AI application, the action may instead be a rollback or a prompt fix.

Reading a monitoring signal

Each signal says something different about what changed

SignalWhat it tells youWhere it comes from
Input driftProduction inputs no longer look like training dataDrift monitoring on input features
Output driftThe distribution of predictions has shiftedDrift monitoring on outputs
Attribution changeDifferent features now drive decisionsFeature attributions over time
Output quality shiftGen AI answers are getting worse or less safeModel evaluation with logging and monitoring
Feature driftThe feature data itself has changedFeature Store monitoring jobs

Worked example (synthetic). Input drift with steady output quality suggests the model is coping for now; output-quality decline with no input drift points to a changed model version or prompt instead. Reading the signals together narrows the cause.

Each monitoring signal answers a different question. Input drift says production inputs no longer look like the data the model was trained on — Google's description of data that can deviate from training data or shift over time. Output drift says the distribution of predictions has moved. A change in feature attributions, which show how much each feature contributed to each inference, says the model is now relying on different inputs to decide. For generative AI, an output-quality shift — measured with model evaluation alongside Cloud Logging and Cloud Monitoring — says answers are getting worse or less safe. And feature drift, detected by Feature Store's feature monitoring jobs, says the feature data itself has changed, which is why Google recommends Feature Store to keep feature data consistent and versioned in the first place. Read together, the signals narrow the cause, and Google's guidance is to configure them proactively, because degradation, data drift and unexpected behavior are cheaper to catch than to explain afterwards.

What this topic actually tests

Measure, change safely, notice

Measure: KPIs from business objectives, continuous output evaluation, endpoint health, human review. Change safely: versioned, regression-tested, released to a few first, with rollback — and patched. Notice: drift and quality shifts, caught by scheduled monitoring with threshold alerts; Feature Store keeps feature data consistent.

Close on the three jobs. Measure: break business objectives into measurable key performance indicators, or KPIs, monitor generative output continuously for quality, safety and grounding, track endpoint request rate, error rate and latency, and keep humans in the evaluation loop so metrics stay aligned with judgment. Change safely: version models, prompts and data; treat a retirement notice of at least forty-five days as a migration window; require a new version to at least match the old one's quality; release to a subset of users first with a rollback plan; remember that a moving default or alias upgrades production silently and a tuned model must be retuned; and apply security patches. Notice: drift is production data deviating from what the model learned, so run monitoring on a schedule with threshold alerts, watch output quality for generative models, and use Agent Platform Feature Store to keep feature data consistent, versioned and monitored. The next topic turns to the cheapest lever of all — the prompt.

Official sources for this topic

Ready to study Generative AI Leader (GCP-GAIL)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free