Introduction¶
This workshop helps you explore some of the agentic capabilities of agentgateway.
The workshop is designed to be self-contained with a minimum of external dependencies.
Prerequisites¶
To work through this workshop, please ensure that you have the following installed on your machine:
- python3 (version 3.11 or above)
- jq
- docker
- ollama (instructions for installation are below)
LLM Provider¶
This workshop assumes a local inference model using the Ollama project.
If you don't already have Ollama running, on a mac you can install it with homebrew:
For other platforms, consult the Ollama docs for the install instructions.
Pull the qwen3 model.
Make sure that qwen3:8b is now showing in the local models list:
By default, the ollama server listens on port 11434.
To make sure that the model is available and produces a response, send a test request to the local LLM:
If you'd rather not use a local model, here are some instructions for Google Gemini (no credit card is required for the free tier).
Free-tier quota
The free tier is often not enough to complete every lab in this workshop.
If you have your own Gemini API key, export that as GEMINI_API_KEY and skip minting a free key below.
Mint a free key:
- Open Google AI Studio and sign in with a personal Google account.
- Accept the Generative AI terms if prompted. Studio will create a default Cloud project for you.
- Click Create API key. Prefer Create key in a new project if you just want a sandbox.
Copy the key and configure the GEMINI_API_KEY environment variable:
Gemini exposes an OpenAI-compatible endpoint. Send it targeting the "Flash-Lite" model, a generous model for experiments:
The agentic scenario: TrendWatch¶
TrendWatch is an AI agent that tells you what is hot in agentic AI today. It leverages a set of MCP servers to read trending discussions, build and save a digest on the topics you care about, and can "publish" these digests.
Clone the GitHub repository for the project:
Navigate into the directory:
The code consists of an agent named "TrendWatch", and a set of example MCP servers, written in python.
Inspect the contents of the agent/ and mcp-servers/ subdirectories, to take account of the main project files:
Setup¶
Create a python virtual environment for the project:
Activate the environment:
Activating the virtual environment for different shells
The python3 virtual environment provides different activate scripts for different types of shells.
If you happen to be running the fish shell, substitute the above command with this instead:
Install the project's dependencies:
Run the distributed tracing project Jaeger in a docker container:
You will use Jaeger to inspect distributed traces illustrating the call flows between the agent, the LLM, and MCP servers.
Configure and run the agent¶
Agents generally are configured with a System prompt, an LLM, and tools to accomplish a specific job. The TrendWatch agent is configured to receive some of that information from environment variables, as follows:
- LLM_BASE_URL - the endpoint for making calls to the LLM.
- LLM_MODEL - the name of the model to target.
- MCP_URL - the URL for the MCP server whose tools the agent can call.
Let us walk through an example.
Configure the three environment variables as follows:
export LLM_BASE_URL="http://localhost:11434/v1"
export LLM_MODEL="qwen3:8b"
export MCP_URL="stdio:./mcp-servers/trends_server.py"
Above, we configure the agent to call ollama, to use the preconfigured qwen model, and to use the trend_server MCP server over the stdio transport (runs as a child process).
Try it out by running:
Configure the environment variables as follows:
export LLM_BASE_URL="https://generativelanguage.googleapis.com/v1beta/openai/"
export LLM_MODEL="gemini-3.5-flash-lite"
export LLM_API_KEY="$GEMINI_API_KEY"
export MCP_URL="stdio:./mcp-servers/trends_server.py"
Above, we configure the agent to call Gemini's OpenAI-compatible endpoint, to use the Flash-Lite model, and to use the trends_server MCP server over the stdio transport (runs as a child process). LLM_API_KEY is the Gemini key from the previous step.
Try it out by running:
The agent outputs some logging information such as:
- Its configuration.
- The tools made visible to the model.
- Information for each "turn", including token consumption, tools called.
Ultimately, the agent outputs the response to the user, in this case the list of trending conversations.
In the first turn, the agent should respond with a tool request for trending_discussions.
The tool fetches and filters trending discussions pertaining to AI from HackerNews, then responds with the top 5 trending discussions.
In the second turn, the agent takes that information and presents it to the user.
Summary¶
So far, we explored a local setup to run an agentic loop: an agent has access to a local model and MCP servers, and can answer questions. As the loop runs, the LLM is consulted, requests for specific tools to be called, and incorporates the responses to further reason about the user's query, and ultimately produces a response.
In the next sections, we explore the agentgateway project, and how it plays a crucial role as a proxy to both LLM and MCP traffic.