Study Guide3,481 words

Unit 2.6 study guide — Give agents the right tools

Generative AI Leader › Unit 2 › Topic 6

Give agents the right tools

Study guide for Generative AI Leader, Unit 2 · Topic 6. This is the topic's lecture in reading form — every slide's teaching, figures and worked examples, in order — followed by the official Google Cloud pages its claims rest on.

What the exam guide asks, quoted. Identifying how agents use tools to interact with the external environment and achieve tasks (e.g., extensions, functions, data stores, and plugins). Identifying relevant Google Cloud services and pre-built AI APIs for agent tooling (e.g., Cloud Storage, databases, Cloud Functions, Cloud Run, Agent Platform, Speech-to-Text API, Text-to-Speech API, Translation API, Document Translation API, Document AI API, Cloud Vision API, Cloud Video Intelligence API, Natural Language API, Google Cloud API Library). Determining when to use Agent Studio and Google AI Studio.

This hive's learning objectives for the topic:

  1. Explain how functions, extensions, data stores and plugins connect agents to external tasks.
  2. Match Google storage, database, execution and prebuilt AI services to tool inputs and outputs.
  3. Distinguish Agent Studio and Google AI Studio by intended workflow and service context.

Giving agents the right tools

A model decides; tools let it fetch, act and reach your systems

How tools connect — functions, extensions, data stores and API tools. Which services sit behind them — storage, databases, Cloud Run, and pre-built AI APIs for speech, language, vision and documents. Where you build — Google AI Studio to try Gemini fast, Agent Studio to build for production on Google Cloud.

A model on its own can only produce text from what it already knows. Tools are what let an agent reach beyond that: Google says tools connect an agent to external systems or inline code, so it can fetch, update, format or analyze information. This topic covers three things about them. First, how tools connect — functions, extensions, data stores and tools that call an external application programming interface, or API, from its schema. Second, which Google Cloud services and pre-built AI APIs a tool can call: storage, databases, execution on Cloud Run, and APIs for speech, translation, documents, vision and language. Third, where you build: Google AI Studio, the fast path to try Gemini models, or Agent Studio on Google Cloud, which turns ideas into production-ready solutions. Two of these products are changing, and the slides will say which.

How a tool reaches the outside world

The model chooses the tool; something else actually runs it

  • Function: the model names it, and your application runs it
  • Extension: the platform runs it — deprecated, shutting down in 2026
  • Data store tool: answers from website content and uploaded data
  • API tool: the agent calls an external API from its OpenAPI schema

Worked example (synthetic). A travel agent asks "what is the weather in Lisbon tomorrow?". The model outputs a call to a get_forecast function with the city; the travel app runs it and returns the forecast, and the model writes the reply.

The first objective is how tools connect an agent to the outside world, and the key idea is that the model chooses a tool but does not itself run it. With function calling, Google says the model determines whether a tool is needed and, if so, outputs structured data naming the tool and its parameters; your application then executes the tool and feeds the result back, so the model can finish its response with real-world information or the outcome of an action. That covers two uses: fetching data, such as current information from knowledge bases and application programming interfaces, or APIs, and taking action, such as submitting forms or handing off a conversation. An extension differs in who runs it: Google says extensions are executed automatically by the platform, whereas functions must be executed by the user or client. But Vertex AI Extensions was deprecated on May 26, 2026, and shuts down after November 26, 2026, so treat extensions as a concept the exam names rather than something to build on. A data store tool gives an agent AI-generated answers from website content and uploaded data. And an API tool — the kind of connection the guide calls a plugin — lets an agent connect to an external API by providing its OpenAPI schema, after which the agent calls the API on your behalf.

The function-calling loop

Decide, run, return, answer

Loading Diagram...
Figure 1 — Mermaid diagram

Figure: A left-to-right flowchart: a user request goes to the model, which decides a tool is needed and outputs a structured call naming the tool and its parameters. Your application runs the tool, and the result is fed back to the model, which completes the answer.

Worked example (synthetic). An expense assistant is asked to file a taxi receipt. The model outputs submit_expense with an amount and date; the finance app files it and returns a reference number, which the model includes in its reply.

Here is the loop every function call follows. The user's request reaches the model. The model determines that a tool is needed and outputs structured data specifying the tool to call and its parameters — it does not run anything itself. Your application then executes the tool and feeds the result back to the model, allowing it to complete its response with dynamic, real-world information or the outcome of an action. Google sums up the effect in one sentence: this effectively bridges the model with your systems and extends its capabilities. Keep the division of labour in mind for the exam: the model decides and the application acts, which is also why the application, not the model, is where permissions and checks belong.

Five ways to connect, compared

Who runs the tool, and what it is for

ConnectionWho runs itTypical use
FunctionYour applicationFetch data or take an action in your systems
Extension (deprecated)The platform, automaticallyShuts down after 26 Nov 2026 — migrate
Data store toolThe agent, over your data storesAnswers from website content and uploads
OpenAPI toolThe agent, calling the API for youAn external service with a published schema
Code executionThe model decides when to use itCalculations and data processing in code

Worked example (synthetic). A fictional retailer's agent uses all four live kinds in a day: a function to check stock, a data store for its returns policy, an OpenAPI tool for a courier's tracking service, and code execution to total an order.

Line the connection types up by who runs the tool. A function is run by your application, after the model names it. An extension was run automatically by the platform — but with the deprecation of Vertex AI Extensions, existing workflows must be migrated to Agent Platform, so expect it only as an exam term. A data store tool lets the agent find answers to end users' questions from your data stores. An OpenAPI tool connects the agent to an external application programming interface, or API, from its schema, and the agent calls that API on your behalf. And code execution is a tool too: once you add it, the model decides when to use it. Exam questions usually turn on the first distinction — whether the application or the platform executes the call — or on recognising which tool fits a data source.

Pick the service from the input and the output

What goes in and what must come out names the API

  • Audio in, text out: Speech-to-Text. Text in, audio out: Text-to-Speech
  • Text to another language: Translation; whole documents: Document Translation
  • Scanned forms to fields: Document AI. Images to text or labels: Vision
  • Text to sentiment: Natural Language. Run code: Cloud Run and its functions
  • Store files in Cloud Storage; rows in a database such as Cloud SQL

Worked example (synthetic). A claims agent receives a voicemail and a photographed form. Speech-to-Text turns the voicemail into text, Document AI pulls fields from the form, and a Cloud Run function writes the claim to Cloud SQL.

The second objective is matching Google Cloud services to what a tool must take in and give back, and the input and output usually name the service. Audio in and text out is Speech-to-Text: send audio and receive a text transcription. Text in and audio out is Text-to-Speech, which creates natural-sounding synthetic speech as playable audio. Translating text is Cloud Translation; translating a formatted file such as a PDF or a Word document is Document Translation, which preserves the original formatting and layout. Turning a scanned form into fields is Document AI, which takes unstructured data from documents and transforms it into structured data suitable for a database. Pulling text or labels out of an image is the Vision application programming interface, or API. Judging whether text is positive, negative or neutral is the Natural Language API's sentiment analysis. When a tool must run code, Cloud Run is a fully managed platform for running your code, function or container, and Cloud Run functions — the guide's Cloud Functions — can respond to events such as a change in a Cloud Storage bucket. And data lands in storage: objects in Cloud Storage buckets, or rows in a fully managed relational database like Cloud SQL.

One agent, three tools

The model picks the call; each tool is a Google Cloud service

Figure. Inside a Google Cloud project, an employee sends a request to an agent, which prompts Gemini to pick a tool. Gemini's function call goes to one of three tools — a Cloud Run function, a data store tool or the Translation API — and each tool's results flow back to the agent. The callout says the agent returns each result to the model to finish the answer.

Worked example (synthetic). A fictional HR agent answers "how many leave days do I have, in Spanish?": a Cloud Run function reads the balance, the data store tool finds the leave policy, and Translation renders the reply.

Drawn as an architecture, an agent with tools looks like this. An employee's request reaches the agent, which prompts the model. Following the function-calling pattern, the model determines which tool is needed and outputs the call, and the call goes to one of the tools the agent was given. Here there are three, each a Google Cloud service: a Cloud Run function, because Cloud Run runs your code or function on fully managed infrastructure; a data store tool, which provides answers based on website content and uploaded data; and the Translation application programming interface, or API, which translates text. Each tool's result returns to the agent, and the application feeds it back to the model so it can complete the response. The shape stays the same however many tools you add; what changes is which services sit behind them.

Input, output, API

Read the tool's job off its input and its output

Tool takesTool returnsService
AudioA text transcriptionSpeech-to-Text
TextPlayable speech audioText-to-Speech
A PDF or DOCX fileThe file, translated, layout keptDocument Translation
A scanned formStructured fields for a databaseDocument AI
An imageIts text, or labelsVision API

Worked example (synthetic). Five tools a fictional insurer's claims agent might need, read straight off their inputs and outputs.

This table is the quickest way to answer tooling questions. A tool that takes audio and returns text is Speech-to-Text. One that takes text and returns playable audio is Text-to-Speech. One that takes a formatted file such as a PDF or Word document and returns it translated, keeping the original formatting and layout, is Document Translation rather than plain text translation. One that takes a scanned form and returns specific fields suitable for a database is Document AI. And one that takes an image and returns the text in it, or general labels for it, is the Vision application programming interface, or API. Two more belong on the list with care. The Natural Language API returns sentiment — positive, negative or neutral — for text. The Video Intelligence API annotates video, but Google says it can be used until September 14, 2027, when it will be shut down, and recommends migrating to the Gemini family of models — so a new design would not choose it.

Where a tool runs and where its data lands

Cloud Run executes; Cloud Storage and databases keep the results

Figure. Three cards. Run: Cloud Run runs your code, function or container, and Cloud Run functions react to events. Store: Cloud Storage keeps objects in buckets, and Cloud SQL keeps relational rows. Enable: find APIs in the API Library and enable any not on by default.

Worked example (synthetic). A fictional team builds a receipt tool: a Cloud Run function triggered when a photo lands in a Cloud Storage bucket, which writes the parsed total to Cloud SQL.

Behind most custom tools sit three ordinary services. Execution: Cloud Run is a fully managed application platform for running your code, function or container, and Cloud Run functions can respond to asynchronous events such as a change in a Cloud Storage bucket — the guide still calls these Cloud Functions. Storage: Cloud Storage keeps data as objects in containers called buckets, and Cloud SQL is a fully managed relational database for MySQL, PostgreSQL and SQL Server. And enablement: Google says to use the Google Cloud console's API Library to browse available APIs and discover the ones that meet your business needs, and that an application programming interface, or API, not enabled by default must be enabled for your project before a tool can call it. A scenario in which a tool fails because a pre-built API was never switched on is testing exactly that last point.

Agent Studio or Google AI Studio?

Try Gemini fast in one; build for the enterprise in the other

  • Google AI Studio: the fast path to try Gemini, in the browser
  • Agent Studio: a Google Cloud workspace for models, instructions and prompts
  • Agent Studio targets production-ready solutions on Google Cloud
  • Regulated workloads belong on Agent Platform, not the Gemini API alone

Worked example (synthetic). A student prototypes a recipe bot in an afternoon in Google AI Studio. A bank building a customer-facing agent with audit and SLA needs works in Agent Studio on its Google Cloud project.

The third objective is telling two similarly named tools apart by workflow and service context. Google AI Studio is, in Google's words, the fast path for developers, students and researchers who want to try Gemini models and get started building with the Gemini Developer application programming interface, or API — a web-based tool that lets you prototype and run prompts right in your browser. Agent Studio lives inside Gemini Enterprise Agent Platform on Google Cloud: it provides a centralized, collaborative workspace for discovering models, refining system instructions and optimizing prompts, and by integrating with the broader Google Cloud ecosystem it helps you turn ideas into production-ready generative AI solutions. The service context is the difference that matters. Google's comparison of the Gemini API with Agent Platform notes that the Gemini API on its own has no enterprise-level support or service level agreements, while Agent Platform offers round-the-clock enterprise support and SLAs, authentication using identity and access management, or IAM, and — in Google's words — regulated customers should use Gemini Enterprise Agent Platform instead.

Two studios, two service contexts

The workflow is similar; the guarantees are not

AspectGoogle AI StudioAgent Studio
Built forTrying Gemini fast, in the browserProduction solutions on Google Cloud
Sits onThe Gemini Developer APIGemini Enterprise Agent Platform
SupportNo enterprise support or SLAs24/7 enterprise support and SLAs
Your dataFree tier: may improve Google productsNever used to improve Google products

Worked example (synthetic). A fictional startup prototypes in Google AI Studio, then moves to Agent Studio when its first enterprise customer asks for an SLA.

Compare the two along the dimensions a leader cares about. Google AI Studio is built for trying Gemini fast, in the browser, on the Gemini Developer application programming interface, or API. Agent Studio is built for production solutions on Google Cloud, inside Gemini Enterprise Agent Platform. Support differs: Google's comparison lists no enterprise-level support or service level agreements, or SLAs, for the Gemini API on its own, against round-the-clock enterprise support and SLAs on Agent Platform. And data use differs: on the Gemini API's free tier, Google says your prompts and responses may be used to improve Google products, while on Agent Platform your prompts, responses and data are never used to improve Google products. That last row is often the deciding fact in an exam scenario involving confidential data.

When a prototype grows up

Prompts move to Agent Studio; models trained elsewhere are retrained

Figure. Three cards. Why move: the app needs a more expansive, end-to-end platform. Prompts: saved in a Google Drive folder, then migrated to Agent Studio. Models: any created in Google AI Studio are retrained on Agent Platform.

Worked example (synthetic). A fictional recipe startup migrates thirty prompts from Drive into Agent Studio in a day, but schedules a week to retrain the one model it had tuned.

The two tools are often stages of one journey. Google's migration guide opens with the reason to move: as your Gemini application programming interface, or API, applications mature, you might find that you need a more expansive platform for building and deploying generative AI applications end to end. The move itself has two parts with different effort. Prompts are easy: Google AI Studio prompt data is saved in a Google Drive folder, and the guide shows how to migrate those prompts to Agent Studio. Models are not: any models you created in Google AI Studio need to be retrained in Gemini Enterprise Agent Platform. So a scenario that asks what carries over unchanged is testing that distinction — prompts migrate, trained models are retrained.

What this topic actually tests

Who runs the tool, what it takes and returns, and where you build

Who runs it? a function is run by your application; an extension by the platform — and extensions are retiring. What goes in and out? audio, text, documents, images and video each name their API; code runs on Cloud Run; data lands in Cloud Storage or a database. Where do you build? Google AI Studio to try Gemini fast; Agent Studio for production on Google Cloud.

Three questions carry this topic. Who runs the tool? With function calling the model outputs the call and your application executes it; extensions were executed by the platform, and Vertex AI Extensions is being shut down after November 26, 2026. What does the tool take in and return? Audio to text is Speech-to-Text, text to audio is Text-to-Speech, formatted documents to translated documents is Document Translation, scanned forms to fields is Document AI, images to text or labels is Vision — and the Video Intelligence application programming interface, or API, is scheduled to shut down, with Gemini recommended instead. Custom code runs on Cloud Run, and data lands in Cloud Storage or a database such as Cloud SQL. Where do you build? Google AI Studio is the fast path to try Gemini models; Agent Studio is the Google Cloud workspace for production-ready solutions with enterprise support.

Official sources for this topic

Ready to study Generative AI Leader (GCP-GAIL)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free