Unit 1.3 study guide — Follow the machine learning lifecycle
Generative AI Leader › Unit 1 › Topic 3
Follow the machine learning lifecycle
Study guide for Generative AI Leader, Unit 1 · Topic 3. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks, quoted. Identifying the stages of the machine learning lifecycle; data ingestion, data preparation, model training, model deployment, and model management; and the Google Cloud tools for each stage.
This hive's learning objectives for the topic:
- Sequence ingestion, preparation, training, deployment and ongoing management.
- Match lifecycle needs to relevant Google Cloud tools without confusing Model Garden and Model Registry.
The machine learning lifecycle
Five stages, each with its own question and its own Google Cloud tools
Ingest — bring the data in. Prepare — make it fit to learn from. Train — learn a model, then evaluate it. Deploy — serve a chosen version. Manage — monitor, version and retrain. The stages loop: what monitoring finds sends work back to the data.
This topic follows a model from raw data to production and back again. Google calls the discipline machine learning operations, or MLOps, and defines it as the process of managing the machine learning, or ML, life cycle, from development to deployment and monitoring. The guide names five stages — data ingestion, data preparation, model training, model deployment and model management — and asks for the Google Cloud tools that serve each one. The first objective is the order of the stages and what each one is for; the second is matching a need to the right tool, with one pair the guide singles out because candidates mix them up: Model Garden, where you find models, and Model Registry, where you manage your own. Keep one idea in view throughout: the lifecycle is a loop, not a line, because what you learn from monitoring a deployed model sends work back to the start.
Five stages, in order
Each stage answers a question the next one depends on
- Ingest: load and distribute the data the model will learn from
- Prepare: clean, transform and format the raw data for training
- Train and evaluate: learn a model, then check it on the task
- Deploy: serve a chosen version to the applications that call it
- Manage: monitor for drift, version models, retrain when needed
Worked example (synthetic). A parcel firm streams scanner events, removes duplicates and mislabeled scans, trains a delay model, deploys the approved version, then watches for drift when a new depot opens.
The first objective is sequence. Ingestion comes first because nothing can be learned from data the system has not received; Google's Pub/Sub, for example, is used for streaming analytics and data integration pipelines to load and distribute data. Preparation follows, and Google defines it as cleaning, transforming and formatting the raw data to make it suitable for model training — the step that decides whether training learns the task or the noise. Training learns a model from the prepared data, and evaluation checks it; Google's own machine learning operations, or MLOps, workflow lists model training, then model evaluation and iteration, as consecutive steps. Deployment makes a chosen model available: for online inference, Google says you must first deploy the model resource to an endpoint. Management is everything after: monitoring for drift, keeping versions organised, and deciding when to retrain. Each stage produces something the next one needs, which is why skipping one — training on unprepared data, deploying an unevaluated model — shows up later as a failure somewhere else.
The lifecycle is a loop
Monitoring sends work back to the data, not just to the model
Figure: A left-to-right flowchart of five stages: ingest the data, prepare the data, train and evaluate, deploy a chosen version, and monitor and manage. An arrow labelled drift or decay detected loops from monitor and manage back to ingest the data.
Worked example (synthetic). A demand model trained before a store opened weekend hours starts under-forecasting Saturdays. Monitoring flags the drift, and the fix starts with new data, not a new model.
Drawn as a flow, the five stages close into a loop. The arrow back from management is what Google's monitoring guidance is about: as with conventional monitoring in machine learning operations, or MLOps, you must deploy an alerting process to notify application owners when drift, skew or performance decay is detected. Drift has a precise meaning here. Google's Model Monitoring page explains that the factors that predicted an outcome when the model was trained may stop being relevant as behaviour changes, and that this deviation in the data is known as drift. When drift is detected, the right response usually begins at the left of the diagram — new or newly prepared data, then retraining and re-evaluation — rather than at the model alone. That is why the guide treats management as a stage of the lifecycle and not as an afterthought to deployment.
Each stage proves something different
A clean dataset is not a good model; a good score is not a deployment
| Stage | What it establishes | What it does not prove |
|---|---|---|
| Prepare | Data is suitable for training | That a model will perform well |
| Train and evaluate | A model meets the task on test data | That a version is serving users |
| Deploy | A version is reachable at an endpoint | That it stays accurate as data changes |
| Manage | Drift and decay are detected and acted on | Nothing on its own — it triggers the loop |
Worked example (synthetic). A handover note says 'model trained, 94% accuracy'. It proves the training stage, but not which version is deployed, where, or who watches it — so the handover is incomplete.
A useful way to hold the sequence is to ask what each stage proves and what it leaves open. Preparation makes data suitable for model training, but a clean dataset says nothing yet about how well a model will do. Training and evaluation show that a model meets the task on the evaluation data — Google's workflow has model evaluation and iteration as its own step — but a score is not a service. Deployment makes a version reachable: Google says that before sending a request for an online inference you must first deploy the model resource to an endpoint. And management is where you find out whether that version stays good, because Model Monitoring can track and alert you when deviations exceed a specified threshold. Exam scenarios often hand you evidence from one stage and ask about another — the trap answer treats a training score as proof of production readiness.
The same loop, for a generative AI application
Discovery replaces training from scratch; monitoring still closes the loop
| Phase | What happens |
|---|---|
| Discovery | Identify which foundation model suits the use case |
| Development and experimentation | Prompt engineering, then tuning if needed |
| Deployment | Manage prompt templates, chains, data stores and adapters |
| Continuous monitoring | Improve performance and maintain safety in production |
Worked example (synthetic). A team building a support assistant spends its 'training' stage choosing and prompting a foundation model, then deploys prompts and a retrieval store alongside it.
For a generative AI application the stages keep their order but change their content, and Google's architecture guide lists the phases. Discovery: developers and engineers identify which foundation model is most suitable for their use case. Development and experimentation: developers use prompt engineering to create and refine input prompts, with tuning methods available when needed. Deployment: developers must manage many artifacts, including prompt templates, chain definitions, retrieval data stores and fine-tuned model adapters. And continuous monitoring in production, where administrators improve performance and maintain safety standards. Map it onto the five stages and it fits: discovery and prompting take the place of training from scratch, the deployed thing is more than one model file, and monitoring still sends work back to the start.
Google Cloud tools for each stage
Match the need to the tool, not the tool to a guess
- Ingest and prepare: Pub/Sub, Dataflow, BigQuery, Cloud Storage
- Find a model to start from: Model Garden
- Train your own: AutoML code-free, or custom training code
- Manage versions and deploy: Model Registry and endpoints
- Watch and automate: Model Monitoring and Agent Platform Pipelines
Worked example (synthetic). A retailer streams till events with Pub/Sub, cleans them in Dataflow into BigQuery, trains with AutoML, registers each version in Model Registry, and alerts on drift with Model Monitoring.
The second objective attaches Google Cloud tools to the stages. For ingestion and preparation, Pub/Sub loads and distributes streaming data; Dataflow provides unified stream and batch data processing at scale, reading from sources, transforming the data and writing it to a destination; BigQuery supports continuous data ingestion and analysis; and Cloud Storage stores data as objects in containers called buckets. For a starting model, Model Garden is an AI and machine learning, or ML, model library that helps you discover, test, customize and deploy models from Google and its partners. To train your own, AutoML lets you build a code-free ML model from the training data you provide, and if AutoML doesn't address your needs, you can run your own training code on managed infrastructure. Model Registry is a central repository where you manage the lifecycle of your models, and from it you can assign a version to an endpoint. Finally, Model Monitoring tracks and alerts on drift, and Agent Platform Pipelines lets you automate, monitor and govern ML systems by orchestrating the workflow.
Ingestion and preparation, as a pipeline
Data streams in, is transformed, lands, and only then trains a model
Figure. A Google Cloud project. Apps and devices send events to Pub/Sub, which streams them to Dataflow. Dataflow transforms the data and writes it to a group labelled prepared data, holding BigQuery and Cloud Storage. Both feed Agent Platform training. A callout reads: ingest, prepare and land the data before any model is trained.
Worked example (synthetic). A fleet company's vehicles send telemetry events; Dataflow drops malformed readings and writes clean rows to BigQuery and raw files to Cloud Storage before a maintenance model is trained.
Here are the first stages as a pipeline. On the left, apps and devices produce events. Pub/Sub receives them — Google says it is used for streaming analytics and data integration pipelines to load and distribute data. Dataflow is the preparation step in the middle: it creates data pipelines that read from one or more sources, transform the data, and write the data to a destination, for batch and streaming alike. The prepared data lands in two places shown in the group: BigQuery, whose streaming supports continuous data ingestion and analysis, and Cloud Storage, which stores data as objects in buckets. Only then does Agent Platform training learn from it. The callout is the point of the picture — ingestion and preparation happen before training, and a model trained straight from raw events learns their duplicates and errors along with their patterns.
Model Garden against Model Registry
One is where you find models; the other is where you manage yours
| Question | Model Garden | Model Registry |
|---|---|---|
| What is it? | A model library from Google and partners | A central repository for your models |
| When do you use it? | Before you have a model: discover and test | After you have models: organize and track versions |
| Leads to | Customizing or deploying a chosen model | Assigning a version to an endpoint |
Worked example (synthetic). Team A has no model yet and wants to compare candidates — Model Garden. Team B has twelve trained versions and must know which one serves production — Model Registry.
The guide names one pair explicitly, because the two names sound alike and do different jobs. Model Garden is an AI and machine learning, or ML, model library that helps you discover, test, customize and deploy models and assets from Google and Google partners — it is where you look before you have chosen a model. Model Registry is a central repository where you can manage the lifecycle of your ML models: from it you have an overview of your models so you can better organize, track and train new versions, and when you have a version you would like to deploy, you can assign it to an endpoint directly from the registry. So the deciding question in a scenario is whether the team is looking for a model or keeping track of its own. A team asking which version is in production needs the registry; a team asking which model to start from needs the garden.
From a trained model to a watched service
Train, register, deploy, monitor — and automate the chain
| Need | Tool | In Google's words |
|---|---|---|
| Train without writing code | AutoML | A code-free model from your data |
| Train with your own code | Custom training | Your code on managed infrastructure |
| Serve online requests | Endpoint | Deploy the model before sending requests |
| Catch changing data | Model Monitoring | Alerts when deviations exceed a threshold |
| Repeat the whole chain | Agent Platform Pipelines | Orchestrate ML workflows |
Worked example (synthetic). A credit team with tabular data and no ML engineers starts with AutoML; a research team with its own model architecture uses custom training. Both deploy to an endpoint and monitor for drift.
The model side of the lifecycle has its own set of tools, each answering a different need. To train without code, AutoML lets you build a code-free machine learning, or ML, model based on the training data you provide; if AutoML doesn't address your needs, you can provide your own training code and run it on Agent Platform's managed infrastructure. To serve online requests, Google is precise about order: online inferences are synchronous requests to a model deployed to an endpoint, so you must first deploy the model resource to an endpoint. To catch a changing world, Model Monitoring can track and alert you when deviations exceed a specified threshold. And to run the whole chain again without doing it by hand, Agent Platform Pipelines lets you automate, monitor and govern ML systems by using pipelines to orchestrate your ML workflows. In a scenario, read the need in the first column and the tool follows.
What this topic actually tests
Order the stages, then match the need to the tool
Order: ingest, prepare, train and evaluate, deploy, manage — and loop back on drift. Proof: each stage proves only its own step. Tools: Pub/Sub and Dataflow move and shape data; AutoML or custom training learn; Model Garden finds models; Model Registry manages yours; Model Monitoring watches.
Close on what the exam asks. First, order: ingestion, preparation, training with evaluation, deployment, and management, with monitoring that sends work back to the data when drift is detected. Second, proof: a prepared dataset, a training score and a deployed endpoint each establish one stage, and evidence from one is not evidence for the next. Third, tools: Pub/Sub loads and distributes data, Dataflow transforms it, BigQuery and Cloud Storage hold it; AutoML trains without code and custom training runs your own; Model Garden is where you discover models, Model Registry is where you manage the lifecycle of your own versions; an endpoint serves them; and Model Monitoring alerts when deviations exceed a threshold. The next topic, choosing a foundation model, zooms into the discovery step this lifecycle starts from.
Official sources for this topic
- What is MLOps? — Google Cloud
- Overview of Model Garden — Gemini Enterprise Agent Platform
- Introduction to Model Registry — Gemini Enterprise Agent Platform
- What is Pub/Sub? — Google Cloud
- Deploy and operate generative AI applications — Cloud Architecture Center
- Overview of getting inferences — Gemini Enterprise Agent Platform
- Introduction to Model Monitoring — Gemini Enterprise Agent Platform
- Dataflow overview — Google Cloud
- BigQuery overview — Google Cloud
- Cloud Storage overview — Google Cloud
- Train and use your own models — Gemini Enterprise Agent Platform
- Introduction to Agent Platform Pipelines — Gemini Enterprise Agent Platform