gpt-oss, locally.
Ollama bundles model weights, configuration, and data into a single package, defined by a Modelfile. It optimizes setup and configuration details, including GPU usage.
For a complete list of supported models and model variants, see the Ollama model library.
Overview
Integration details
Model features
Setup
First, follow these instructions to set up and run a local Ollama instance:- Download and install Ollama onto the available supported platforms (including Windows Subsystem for Linux aka WSL, macOS, and Linux)
- macOS users can install via Homebrew with
brew install ollamaand start withbrew services start ollama
- macOS users can install via Homebrew with
- Fetch available LLM model via
ollama pull <name-of-model>- View a list of available models via the model library
- e.g.,
ollama pull gpt-oss:20b
- This will download the default tagged version of the model. Typically, the default points to the latest, smallest sized-parameter model.
On Mac, the models will be download to~/.ollama/modelsOn Linux (or WSL), the models will be stored at/usr/share/ollama/.ollama/models
- Specify the exact version of the model of interest as such
ollama pull gpt-oss:20b(View the various tags for theVicunamodel in this instance) - To view all pulled models, use
ollama list - To chat directly with a model from the command line, use
ollama run <name-of-model> - View the Ollama documentation for more commands. You can run
ollama helpin the terminal to see available commands.
Installation
The LangChain Ollama integration lives in thelangchain-ollama package:
Instantiation
Now we can instantiate our model object and generate chat completions:Invocation
Tool calling
Ollama tool calling uses the OpenAI compatible web server specification, and you can use it with the defaultBaseChatModel.bind_tools() methods as described in the LangChain tools documentation.
Make sure to select an ollama model that supports tool calling.
We can use tool calling with an LLM that has been fine-tuned for tool use such as gpt-oss:
@tool decorator on a normal python function.
Multi-modal
Ollama has limited support for multi-modal LLMs, such as gemma3 Be sure to update Ollama so that you have the most recent version to support multi-modal.Log probabilities
ChatOllama supports token-level log probabilities via the logprobs and top_logprobs parameters. Log probabilities indicate how likely each token was at each generation step.
Basic usage
Top-K alternatives per token
Usetop_logprobs to return the most likely alternative tokens at each position:
Reasoning models and custom message roles
Some models, such as IBM’s Granite 3.2, support custom message roles to enable thinking processes. To access Granite 3.2’s thinking features, pass a message with a"control" role with content set to "thinking". Because "control" is a non-standard message role, we can use a ChatMessage object to implement it:
API reference
For detailed documentation of allChatOllama features and configurations head to the API reference.
Connect these docs to Claude, VSCode, and more via MCP for real-time answers.

