Unit 1.2 study guide — Understand AI concepts and learning approaches
Generative AI Leader › Unit 1 › Topic 2
Understand AI concepts and learning approaches
Study guide for Generative AI Leader, Unit 1 · Topic 2. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.
What the exam guide asks, quoted. Defining core gen AI concepts (e.g., artificial intelligence, natural language processing, machine learning, generative AI, foundation models, multimodal foundation models, diffusion models, prompt tuning, prompt engineering, large language models). Describing the machine learning approaches (e.g., supervised, unsupervised, reinforcement).
This hive's learning objectives for the topic:
- Distinguish artificial intelligence, machine learning, natural language processing and generative AI.
- Distinguish foundation, large language, multimodal and diffusion models.
- Distinguish prompt engineering from learned prompt tuning and model tuning.
- Choose supervised, unsupervised or reinforcement learning for a stated learning problem.
AI concepts and learning approaches
Name the field, the model, the adaptation and the learning signal
What kind of system? artificial intelligence, machine learning, natural language processing, generative AI. What kind of model? foundation, large language, multimodal, diffusion. How is it adapted? prompt engineering, prompt tuning, fine-tuning. How did it learn? supervised, unsupervised, reinforcement.
This topic is the vocabulary every later topic in the hive assumes, and the guide tests it as distinctions rather than definitions. Four questions organise it. First, what kind of system is this? Artificial intelligence, or AI, is the broad field; machine learning, or ML, is the method of training a model from data; natural language processing works on human language; and generative AI creates new content. Second, what kind of model is it? Foundation, large language, multimodal and diffusion models are descriptions of a model's scale, material and method, and one model can carry several of them. Third, how is its behaviour adapted? By writing better prompts, by learning a soft prompt, or by training the model further. Fourth, how did it learn in the first place? From labeled examples, from patterns in unlabeled data, or from rewards. Each question has its own slide and a figure, and the last slide pulls the four together.
Four labels that answer four different questions
Field, method, material and output are not rivals
- AI: a non-human program or model that solves sophisticated tasks
- ML: a sub-field of AI that trains a model from input data
- NLP: processing what people say or type
- Generative AI: does not just analyze data, it creates new content
Worked example (synthetic). A returns desk classifies each customer email into a queue and drafts a reply. Classifying the email is natural language processing done with machine learning; drafting the reply is generative AI. One system, three correct labels.
The first objective asks you to tell apart four terms that are often used as if they competed. Google's glossary defines artificial intelligence, or AI, as a non-human program or model that can solve sophisticated tasks, and adds that, formally, machine learning is a sub-field of artificial intelligence. Machine learning, or ML, is defined by its method: a program or system that trains a model from input data. Natural language processing, or NLP, is defined by its material — Google's glossary calls it the field of teaching computers to process what a user said or typed — and Google's NLP page notes that it uses machine learning to reveal the structure and meaning of text. Generative AI is defined by its output: Google describes it as a type of artificial intelligence that doesn't just analyze data; it creates new content. Because these answer different questions — what field, what method, what material, what output — one system can carry several labels at once. The exam rewards reading which question a scenario is actually asking.
Nested, not parallel
Machine learning sits inside AI; NLP and generative AI describe what it works on
Figure. Nested boxes. The outer box is artificial intelligence, a program or model that solves sophisticated tasks. Inside it is machine learning, which trains a model from input data. Inside machine learning sit two dashed boxes side by side: natural language processing, which works on what people say or type and uses machine learning, and generative AI, which creates new content from learned patterns.
Worked example (synthetic). A chatbot that writes answers to typed questions is all four at once: AI, built with machine learning, doing natural language processing, generating new text.
Drawn as boxes, the terms nest rather than compete. The outer box is artificial intelligence, or AI. Inside it is machine learning, or ML — Google's glossary says that formally, machine learning is a sub-field of artificial intelligence. The two dashed boxes inside describe what a machine-learned system works on and what it produces. Natural language processing, which Google's AI page says enables computers to understand, interpret and generate human language, uses machine learning to reveal the structure and meaning of text. Generative AI, in Google's words, learns the patterns and structures within vast amounts of data — text, images, code and more — and then uses that knowledge to produce entirely new, original content based on prompts. The dashed boxes overlap in practice: a model that reads a question and writes an answer is doing language processing and generation in one step. So when a question offers two labels that both apply, look for the one the stem's verb points at — understand, or create.
Generative AI against traditional AI
Classify and predict, or create and summarize
| Approach | What it excels at | Typical output |
|---|---|---|
| Traditional AI | Learning from existing data to classify or predict | A category, a number, a forecast |
| Generative AI | Create summaries, uncover hidden correlations, generate new content | Text, images or videos in the style of its training data |
Worked example (synthetic). Routing a support email to the billing queue needs a category — traditional AI. Drafting the reply needs new text — generative AI. A single product can use both.
Google draws the line between the two kinds of artificial intelligence, or AI, by what they produce. Traditional AI models, its guidance says, excel at learning from existing data to classify information or predict future outcomes based on historical patterns — the output is a category, a value or a forecast. Generative AI models expand these capabilities to create summaries, uncover complex hidden correlations, or generate new content — like text, images or videos — that reflect the style and patterns within the training data. That gives you a fast test for an exam scenario: if the business needs a label or a number, a traditional model may be the better fit; if it needs a draft, a summary or an image, it needs generation. And the two are not exclusive — the worked example routes an email with one and answers it with the other.
Four ways to describe a model
Scale, language, modalities and method are separate descriptions
- Foundation model: very large, pre-trained on an enormous and diverse set
- Large language model: a language model with a very high number of parameters
- Multimodal model: inputs, outputs or both span more than one modality
- Diffusion model: learns to reverse added noise to generate outputs
Worked example (synthetic). A team reads that a model is a large multimodal foundation model. That tells them it is broadly pre-trained and handles more than one kind of data — not which input and output pairs it supports for their task.
The second objective covers four model descriptions, and the trap is to treat them as four kinds of model when they are four kinds of statement. A foundation model, in Google's glossary, is a very large pre-trained model trained on an enormous and diverse training set, and one thing it can do is serve as a base model for additional fine-tuning or other customization. A large language model, or LLM, is at a minimum a language model having a very high number of parameters; Google's LLM page adds that it is trained on a massive amount of data and can be used to generate and translate text and perform other natural language processing tasks. A multimodal model is a model whose inputs, outputs, or both include more than one modality — images, videos and text, for example. A diffusion model is described by its method: it works by adding noise to data and then learning how to reverse the process iteratively to generate new, realistic outputs. One model can match several of these descriptions at once, which is why a description never replaces checking what the model actually accepts and returns.
How a diffusion model learns
Add noise, learn to remove it, then generate from noise
Figure: A left-to-right flowchart: training data has noise added step by step to become noisy data; the model learns to reverse that process; applying the learned reverse process iteratively produces a new, realistic output.
Worked example (synthetic). A conceptual sequence only — it describes the mechanism, not any named Google image or video model.
The diffusion description is the one in the guide that names a method, so it helps to see the method. Google's definition, which appears in the glossary of its WeatherNext weather models, says a diffusion model works by adding noise to data and then learning how to reverse the process iteratively to generate new, realistic outputs. Read the flow left to right. During training, noise is added to real data step by step; the model learns how to undo that corruption. To generate, the learned reverse process is applied over and over until something new and realistic emerges. Treat this as a picture of the idea, not a claim about any particular product: the definition sits on a weather glossary, and this deck does not assert which Google image or video models use the technique.
What each description tells you, and what it does not
A label is a starting point; the task decides
| Description | Tells you | Still check |
|---|---|---|
| Foundation model | Broadly pre-trained; a base for customization | Whether it suits your task |
| Large language model | Language-focused; very many parameters | Which languages and tasks |
| Multimodal model | More than one modality in, out or both | The exact input and output pairs |
| Diffusion model | Generates by reversing added noise | What it produces for your use |
Worked example (synthetic). A support team needs photos in and written advice out. 'Multimodal' is promising, but only the model's own documentation can confirm photo-in, text-out.
Each description narrows the field without settling the choice. Foundation tells you a model is broadly pre-trained and can serve as a base for additional fine-tuning or other customization — not that it fits your task. Large language model tells you the material is language and the parameter count is very high — not which languages it handles well. Multimodal is the one most often over-read: the glossary's definition is that inputs, outputs or both include more than one modality, so a multimodal model may take images in and give only text out. Google's multimodal page gives an example of the direction mattering — Gemini can receive a photo of a plate of cookies and generate a written recipe, and vice versa — but a different model may support only one direction. Diffusion tells you how a model generates. In every row, the last column is the work: check the input and output pairs your use case actually needs.
Three ways to change what a model does
Edit the prompt, learn a prefix, or retrain the parameters
- Prompt engineering: craft prompts that elicit the desired responses
- Prompt tuning: learn a soft-prompt prefix while freezing the model
- Fine-tuning: a second, task-specific training pass on a pre-trained model
- Google recommends starting with prompting, then tuning only if required
Worked example (synthetic). A team rewrites its instructions and adds two examples to the prompt — prompt engineering. Nothing about the model has been trained, so nothing about it has changed.
The third objective is three ways of adapting a model, and the distinction is what actually changes. Prompt engineering, in Google's glossary, is the art of creating prompts that elicit the desired responses from a large language model; Google's longer page calls it the art and science of designing and optimizing prompts. The model itself is untouched — only the text sent to it changes. Prompt tuning is a parameter-efficient tuning mechanism that learns a prefix the system prepends to the actual prompt. That prefix is learned, not written: the glossary says the system learns the soft prompt by freezing all other model parameters and fine-tuning on a specific task. Fine-tuning is a second, task-specific training pass performed on a pre-trained model to refine its parameters for a specific use case. The order Google recommends follows the cost: start with prompting to find the optimal prompt, then move on to fine-tuning, if required, to further boost performance or fix recurrent errors.
What changes, and what stays fixed
Only one of the three leaves the model exactly as it was
| Method | What changes | What stays fixed |
|---|---|---|
| Prompt engineering | The prompt you write | Every model parameter |
| Prompt tuning | A learned soft prompt prepended to the input | All other model parameters (frozen) |
| Parameter-efficient tuning | A relatively small subset of parameters | The rest of the parameters |
| Full fine-tuning | All parameters of the model | Nothing |
Worked example (synthetic). A legal team's clause classifier keeps making the same mistake after weeks of prompt edits. Moving to tuning changes something the prompt never could: the learned values themselves.
The table sorts the methods by what they change. Prompt engineering changes only the text you send, so every model parameter stays exactly as it was. Prompt tuning changes a learned soft prompt that is prepended to the input, while the system freezes all other model parameters. Fine-tuning comes in two strengths on Google's tuning page: parameter-efficient tuning updates a relatively small subset of the model's parameters during the tuning process, and full fine-tuning updates all parameters of the model, which suits highly complex tasks with the potential of higher quality — at higher cost. Notice that prompt tuning appears in Google's glossary as a parameter-efficient mechanism, so it sits at the light end of the tuning family. The exam's usual distinction is the first row against the rest: writing a better prompt is not training, and a team that has only edited its prompt has not tuned anything.
Prompting first, tuning if required
Escalate only when the evidence says the prompt is not enough
Figure: A left-to-right flowchart. Design and refine the prompt, then ask whether it meets the bar on your evaluation. If yes, ship the prompt. If no, tune the model with parameter-efficient or full fine-tuning, then evaluate again.
Worked example (synthetic). A support bot's tone is fixed by adding a role and two examples to the prompt, so no tuning is needed; a second bot still mislabels product codes after prompt changes, so it moves to tuning.
Google's tuning guidance gives an order of operations, and the flowchart draws it. We recommend starting with prompting to find the optimal prompt, the page says; then move on to fine-tuning, if required, to further boost performance or fix recurrent errors. The decision in the middle is an evaluation, not a feeling: if the prompted model meets the bar on your own examples, the cheaper option has already succeeded. If it does not, tuning is the next step, and the lighter, parameter-efficient form updates only a relatively small subset of the model's parameters. Whatever was changed, it goes back through the same evaluation. In scenarios, the wrong answers usually jump straight to retraining before anyone has tried a better prompt, or claim a prompt edit has tuned the model.
Three learning signals
Labels, patterns or rewards — read what the data provides
- Supervised: features with their corresponding labels
- Unsupervised: patterns in a typically unlabeled dataset
- Reinforcement: an agent maximizes return through interaction
- The available data and feedback decide the approach
Worked example (synthetic). A bank has two years of transactions each marked fraud or not fraud. Those marks are labels, so predicting fraud on new transactions is a supervised problem.
The fourth objective asks you to choose a learning approach for a stated problem, and the reliable way is to ask what signal the data provides. Supervised machine learning, in Google's glossary, is training a model from features and their corresponding labels — Google's comparison page puts it simply: supervised learning requires labeled datasets. Unsupervised machine learning is training a model to find patterns in a dataset, typically an unlabeled one; it is more helpful for discovering new patterns and relationships in raw, unlabeled data. Reinforcement learning, or RL, is a family of algorithms that learn an optimal policy whose goal is to maximize return when interacting with an environment. Google's RL page describes an agent that, rather than relying on explicit programming or labeled datasets, learns by trial and error, receiving rewards or penalties for its actions. So: known answers mean supervised, no answers and a search for structure mean unsupervised, and actions with feedback mean reinforcement.
Signal, then typical tasks
Match the problem to what the training data looks like
| Approach | Learning signal | Typical tasks Google lists |
|---|---|---|
| Supervised | Labeled examples | Classification and regression: spam detection, sentiment, pricing changes |
| Unsupervised | Patterns in unlabeled data | Clustering, anomaly detection, customer segmentation |
| Reinforcement | Rewards or penalties from actions | Learning a policy by interacting with an environment |
Worked example (synthetic). A retailer wants customer groups nobody has defined yet — unsupervised clustering. A warehouse robot learning which picking routes finish fastest from trial runs — reinforcement learning.
Google's pages attach typical tasks to each approach, which turns the definitions into a lookup. Supervised machine learning is suited for classification and regression tasks — Google lists weather forecasting, pricing changes, sentiment analysis and spam detection — because each needs examples with known answers. Unsupervised learning is more commonly used for exploratory data analysis and clustering tasks, such as anomaly detection, big data visualization or customer segmentation, where the groups are what you want to discover. Reinforcement learning, or RL, fits problems where an agent acts, receives rewards or penalties, and improves its policy through interaction with an environment. When a scenario is ambiguous, go back to the signal column: if nobody has labeled the data, a supervised model has nothing to learn from, however well the task sounds like classification.
What this topic actually tests
Four distinctions, in Google's words
Which question does the label answer? field, method, material or output. What does the description prove? scale, language, modalities or method — not fit. What changed? the prompt, a learned prefix, or the parameters. What signal? labels, patterns or rewards.
Close on the four distinctions. First, the labels artificial intelligence, machine learning, natural language processing and generative AI answer different questions — field, method, material, output — so several can be true of one system, and generative AI is the one that creates new content. Second, model descriptions narrow the field without settling fit: a foundation model is a broadly pre-trained base, a large language model is a language model with very many parameters, a multimodal model has more than one modality in its inputs, outputs or both, and a diffusion model generates by learning to reverse added noise. Third, adapting a model means changing the prompt, learning a soft prompt with the model frozen, or retraining parameters — and Google says to start with prompting. Fourth, supervised learning needs labels, unsupervised learning finds patterns in unlabeled data, and reinforcement learning learns from rewards. The next topic follows a model through its lifecycle, using this vocabulary at every stage.
Official sources for this topic
- Machine Learning Glossary — Google for Developers
- What is artificial intelligence? — Google Cloud
- What is natural language processing? — Google Cloud
- What is a large language model? — Google Cloud
- Glossary — WeatherNext, Google for Developers
- What is prompt engineering? — Google Cloud
- Introduction to tuning — Gemini Enterprise Agent Platform
- Supervised vs. unsupervised learning — Google Cloud
- What is reinforcement learning? — Google Cloud
- When to use generative AI or traditional AI — Google Cloud
- Multimodal AI — Google Cloud