A Vercel AI SDK provider for Ollama, built on the official ollama package. Type-safe, cross-provider compatible, with access to native Ollama features.
Version compatibility
Each
ai-sdk-ollamamajor targets one AI SDK major. Every line stays published on npm, so pick the one that matches youraiversion:
ai-sdk-ollamaai(peer)Status 4.x(@latest)ai@^7Active. New features land here. 3.xai@^6Maintenance. Critical fixes only. 2.xai@^5Unmaintained. AI SDK v7 is stable, so v4 installs from
latest. The v3 line still works for AI SDK v6.
# AI SDK v7 (stable)
npm install ai-sdk-ollama ai
# AI SDK v6
npm install ai-sdk-ollama@^3 ai@^6import { ollama } from 'ai-sdk-ollama';
import { generateText } from 'ai';
// Works in both Node.js and browsers
const { text } = await generateText({
model: ollama('llama3.2'),
prompt: 'Write a haiku about coding',
temperature: 0.8,
});
console.log(text);- Tool calling reliability - Enhanced response synthesis for consistent tool execution
- Wrapper functions -
generateTextandstreamTextwith improved response handling - Built-in reliability - Default reliability features enabled automatically
- Automatic JSON repair - Cascade repair (jsonrepair + Ollama-specific fallback) for malformed object generation output
- Web search integration - Built-in web search and fetch tools powered by Ollama's web search API
- Reranking - Document relevance ranking using embedding-based similarity
- Decision models - Typed choice, boolean and score answers with probabilities via
experimental_evaluate - Middleware system - Wrap models with
defaultSettingsMiddlewareandextractReasoningMiddleware - ToolLoopAgent - Autonomous agents that run tool loops with configurable stop conditions
- Streaming utilities -
smoothStreamandparsePartialJsonfor stream handling - Type-safe - Full TypeScript support with strict typing
- Cross-environment - Runs in Node.js and browsers automatically
- Native Ollama features - Use advanced options like
mirostat,repeat_penalty,num_ctx - Production-ready - Handles the tool-call and JSON edge cases that otherwise need manual workarounds
The problem this solves: Standard Ollama providers can execute a tool and then return empty text. These wrapper functions synthesize a response so you get complete output.
import { generateText, streamText } from 'ai-sdk-ollama';
// Enhanced generateText with reliable response synthesis
const { text } = await generateText({
model: ollama('llama3.2'),
tools: {
/* your tools */
},
prompt: 'Use the tools and explain the results',
});
// Enhanced streaming with tool execution
const { textStream } = await streamText({
model: ollama('llama3.2'),
tools: {
/* your tools */
},
prompt: 'Stream with tools',
});Tools + structured output: The
enableToolsWithStructuredOutputoption lets you use tool calling and structured output in the same call.
import { generateText } from 'ai-sdk-ollama';
import { Output, tool } from 'ai';
import { z } from 'zod';
const weatherTool = tool({
description: 'Get current weather for a location',
inputSchema: z.object({
location: z.string().describe('City name'),
}),
execute: async ({ location }) => ({
location,
temperature: 22,
condition: 'sunny',
humidity: 60,
}),
});
// AI SDK v7: tools and structured output work together by default
const result = await generateText({
model: ollama('llama3.2'),
prompt: 'Get weather for San Francisco and provide a structured summary',
tools: { getWeather: weatherTool },
output: Output.object({
schema: z.object({
location: z.string(),
temperature: z.number(),
summary: z.string(),
}),
}),
toolChoice: 'required',
});
// Result: Tool is called AND structured output is generatedWeb search and fetch: Built-in tools powered by Ollama's web search API for current information.
import { generateText } from 'ai';
import { ollama } from 'ai-sdk-ollama';
// 🔍 Web search for current information
const { text } = await generateText({
model: ollama('qwen3-coder:480b-cloud'), // Cloud models recommended for web search
prompt: 'What are the latest developments in AI this week?',
tools: {
webSearch: ollama.tools.webSearch({ maxResults: 5 }),
},
});
// 📄 Fetch specific web content
const { text: summary } = await generateText({
model: ollama('gpt-oss:120b-cloud'),
prompt: 'Summarize this article: https://example.com/article',
tools: {
webFetch: ollama.tools.webFetch({ maxContentLength: 5000 }),
},
});
// 🔄 Combine search and fetch for comprehensive research
const { text: research } = await generateText({
model: ollama('gpt-oss:120b-cloud'),
prompt: 'Research recent TypeScript updates and provide a detailed analysis',
tools: {
webSearch: ollama.tools.webSearch({ maxResults: 3 }),
webFetch: ollama.tools.webFetch(),
},
});- Ollama API Key: Set
OLLAMA_API_KEYenvironment variable - Cloud Models: Use cloud models for optimal web search performance:
qwen3-coder:480b-cloud- Best for general web searchgpt-oss:120b-cloud- Best for complex reasoning with web data
# Set your API key
export OLLAMA_API_KEY="your_api_key_here"
# Get your API key from: https://ollama.com/account
# Run web search examples
npx tsx examples/node/src/web-search-ai-sdk-ollama.ts basic # Run basic example only
npx tsx examples/node/src/web-search-ai-sdk-ollama.ts combined # Run combined search and fetch
npx tsx examples/node/src/web-search-ai-sdk-ollama.ts streaming # Run streaming example
npx tsx examples/node/src/web-search-ai-sdk-ollama.ts error # Run error handling exampleimport { ollama } from 'ai-sdk-ollama';
import { generateText, streamText, Output, embed, tool } from 'ai';
import { z } from 'zod';
// Text generation - works exactly like OpenAI, Anthropic, etc.
const { text } = await generateText({
model: ollama('llama3.2'), // Just swap the model
prompt: 'Write a haiku about coding',
temperature: 0.8,
});
// Streaming text
const { textStream } = await streamText({
model: ollama('llama3.2'),
prompt: 'Tell me a story',
});
// Structured object generation
const { output } = await generateText({
model: ollama('llama3.2'),
output: Output.object({
schema: z.object({
name: z.string(),
age: z.number(),
interests: z.array(z.string()),
}),
}),
prompt: 'Generate a random person profile',
});
// Streaming structured objects
const { partialOutputStream } = await streamText({
model: ollama('llama3.2'),
output: Output.object({
schema: z.object({
step: z.string(),
result: z.string(),
}),
}),
prompt: 'Break down the process of making coffee',
});
// Embeddings
const { embedding } = await embed({
model: ollama.embedding('nomic-embed-text'),
value: 'Hello world',
});
console.log('Embedding dimensions:', embedding.length); // 768
// Tool calling (with enhanced reliability)
const { text, toolCalls } = await generateText({
model: ollama('llama3.2'),
prompt: 'What is the weather in San Francisco?',
tools: {
getWeather: tool({
description: 'Get current weather for a location',
inputSchema: z.object({
location: z.string().describe('City name'),
}),
execute: async ({ location }) => ({ temp: 18, condition: 'sunny' }),
}),
},
});
// Image analysis (vision models like llava, bakllava)
const { text } = await generateText({
model: ollama('llava'),
prompt: [
{
role: 'user',
content: [
{ type: 'text', text: 'Describe this image:' },
{
type: 'file',
data: new URL('https://example.com/image.jpg'),
mediaType: 'image/jpeg',
},
],
},
],
});
// Note: Image generation is not supported by Ollama
// Use other providers like OpenAI DALL-E for image generation// Access Ollama's advanced sampling while keeping portability
const { text } = await generateText({
model: ollama('llama3.2', {
options: {
mirostat: 2, // Advanced sampling algorithm
repeat_penalty: 1.1, // Fine-tune repetition
num_ctx: 8192, // Larger context window
},
}),
prompt: 'Write a detailed analysis',
temperature: 0.8, // Standard AI SDK parameters still work
});import { generateText, Output } from 'ai';
import { z } from 'zod';
// Auto-detection: structured outputs enabled automatically
const { output } = await generateText({
model: ollama('llama3.2'),
output: Output.object({
schema: z.object({
name: z.string(),
age: z.number(),
interests: z.array(z.string()),
}),
}),
prompt: 'Generate a random person profile',
});
console.log(output);
// { name: "Alice", age: 28, interests: ["reading", "hiking"] }MCP support: OAuth authentication, resources, prompts, and elicitation.
import { generateText } from 'ai';
import { ollama } from 'ai-sdk-ollama';
import { createMCPClient } from '@ai-sdk/mcp';
import { Experimental_StdioMCPTransport } from '@ai-sdk/mcp/mcp-stdio';
// Connect to MCP server (stdio transport)
const mcpClient = await createMCPClient({
transport: new Experimental_StdioMCPTransport({
command: 'path/to/mcp-server',
args: [], // Optional arguments
}),
// Or use HTTP transport for remote servers
// transport: new HTTPMCPTransport({
// url: 'https://mcp-server.example.com',
// headers: { 'Authorization': 'Bearer token' },
// }),
});
// Get tools from MCP server
const tools = await mcpClient.tools();
// Use MCP tools with Ollama
const { text } = await generateText({
model: ollama('llama3.2'),
prompt: 'Calculate 15 + 27 and get the current time',
tools, // MCP tools automatically work with Ollama
});
// Clean up
await mcpClient.close();Prerequisites: Install @ai-sdk/mcp package:
npm install @ai-sdk/mcpSee the MCP Tools documentation for OAuth, resources, prompts, and elicitation support.
MCP Apps: The provider also works with MCP Apps — tools that declare ui:// HTML resources rendered in a sandboxed iframe. See examples/node/src/mcp-apps-example.ts for the host flow and the browser example for an interactive demo with experimental_MCPAppRenderer.
// Automatic environment detection - same code works everywhere
import { ollama } from 'ai-sdk-ollama';
import { generateText } from 'ai';
const { text } = await generateText({
model: ollama('llama3.2'),
prompt: 'Hello from the browser!',
});import { embed } from 'ai';
// Single embedding
const { embedding } = await embed({
model: ollama.embedding('nomic-embed-text'),
value: 'Hello world',
});
console.log('Embedding length:', embedding.length); // 768 dimensions
// Multiple embeddings
const texts = ['Hello world', 'How are you?', 'AI is amazing'];
const results = await Promise.all(
texts.map((text) =>
embed({
model: ollama.embedding('nomic-embed-text'),
value: text,
}),
),
);Reranking: Rank documents to improve search results and RAG pipelines.
import { rerank } from 'ai';
import { ollama } from 'ai-sdk-ollama';
// Rerank documents by relevance to a query
const { ranking, rerankedDocuments } = await rerank({
model: ollama.embeddingReranking('nomic-embed-text'),
query: 'What is machine learning?',
documents: [
'Machine learning is a subset of AI that learns from data.',
'The weather is sunny today.',
'Deep learning uses neural networks for complex patterns.',
'I like pizza.',
],
topN: 2, // Return top 2 most relevant
});
// Results sorted by relevance score
ranking.forEach((item, i) => {
console.log(
`${i + 1}. Score: ${item.score.toFixed(3)} - ${rerankedDocuments[i]}`,
);
});Ollama 0.35 serves decision models such as nimble and tev1. You ask typed questions about some state and get each answer with its probabilities in one local call. Use them for triage, routing and moderation.
import { experimental_evaluate as evaluate } from 'ai';
import { ollama } from 'ai-sdk-ollama';
const { answers } = await evaluate({
model: ollama.evaluationModel('nimble'),
state: { ticket: 'I was charged twice. Please refund the extra payment.' },
questions: {
team: {
type: 'choice',
instructions: 'Which team should handle this ticket?',
criteria: {
billing: 'Payments and refunds',
technical: null,
other: null,
},
},
refund: {
type: 'boolean',
instructions: 'Does the customer ask for a refund?',
},
urgency: {
type: 'score',
instructions: 'How urgent is this ticket?',
criteria: ['Routine', 'Soon', 'Urgent'],
},
},
});
answers.team.choice; // 'billing'
answers.refund.probability; // P(true), e.g. 0.99
answers.urgency.score; // 0 to 2, e.g. 0.8import {
ollama,
wrapLanguageModel,
defaultSettingsMiddleware,
extractReasoningMiddleware,
} from 'ai-sdk-ollama';
// Apply default settings to all calls
const modelWithDefaults = wrapLanguageModel({
model: ollama('llama3.2'),
middleware: defaultSettingsMiddleware({
settings: { temperature: 0.7, maxOutputTokens: 500 },
}),
});
// Extract reasoning from <think> tags
const modelWithReasoning = wrapLanguageModel({
model: ollama('llama3.2'),
middleware: extractReasoningMiddleware({ tagName: 'think' }),
});
// Combine multiple middlewares
const enhancedModel = wrapLanguageModel({
model: ollama('llama3.2'),
middleware: [
defaultSettingsMiddleware({ settings: { temperature: 0.5 } }),
extractReasoningMiddleware({ tagName: 'thinking' }),
],
});import { ollama } from 'ai-sdk-ollama';
import { ToolLoopAgent, stepCountIs, hasToolCall } from 'ai';
import { tool } from 'ai';
import { z } from 'zod';
const agent = new ToolLoopAgent({
model: ollama('llama3.2'),
instructions: 'You are a helpful assistant.',
tools: {
weather: tool({
description: 'Get weather for a location',
inputSchema: z.object({ location: z.string() }),
execute: async ({ location }) => ({ temp: 72, condition: 'sunny' }),
}),
done: tool({
description: 'Signal task completion',
inputSchema: z.object({ summary: z.string() }),
execute: async ({ summary }) => ({ completed: true, summary }),
}),
},
stopWhen: [stepCountIs(10), hasToolCall('done')],
onStepFinish: (step, index) =>
console.log(`Step ${index + 1}:`, step.finishReason),
});
const result = await agent.generate({
prompt: 'Get the weather in Tokyo, then call done with a summary.',
});import { parsePartialJson, simulateReadableStream, smoothStream } from 'ai';
// Parse incomplete JSON from streams
const result = await parsePartialJson('{"name": "John", "age": 30');
if (result.state === 'repaired-parse' || result.state === 'successful-parse') {
console.log(result.value); // { name: "John", age: 30 }
}
// Testing utility with controlled timing
const testStream = simulateReadableStream({
chunks: ['Hello', ' ', 'World'],
chunkDelayInMs: 100,
});
// Smooth streaming with chunking
import { streamText } from 'ai';
const result = streamText({
model: ollama('llama3.2'),
prompt: 'Write a story',
experimental_transform: smoothStream({ chunking: 'word' }),
});Note: Most streaming utilities are now available directly from the 'ai' package. Import them from 'ai' instead of 'ai-sdk-ollama'.
- Ollama installed and running locally
- Node.js 22+ for development
- AI SDK v7 (
aipackage)
# Start Ollama
ollama serve
# Pull a model
ollama pull llama3.2Complete working examples in examples/node:
# Run any example directly
npx tsx examples/node/src/basic-chat.ts
npx tsx examples/node/src/dual-parameter-example.ts
npx tsx examples/node/src/simple-tool-test.ts
npx tsx examples/node/src/mcp-tools-example.ts # Model Context Protocol integration
npx tsx examples/node/src/mcp-apps-example.ts # MCP Apps host flow (ui:// resources)
npx tsx examples/node/src/embedding-example.ts # Vector embeddings
npx tsx examples/node/src/streaming-simple-test.ts
npx tsx examples/node/src/web-search-ai-sdk-ollama.ts # Web search and fetch tools
npx tsx examples/node/src/web-search-ai-sdk-ollama.ts basic # Run specific example to avoid rate limits
npx tsx examples/node/src/reasoning-example.ts # Chain-of-thought reasoning
npx tsx examples/node/src/image-handling-example.ts # Vision models
npx tsx examples/node/src/v6-reranking-example.ts # Document reranking
npx tsx examples/node/src/smooth-stream-example.ts # Smooth chunked streaming
npx tsx examples/node/src/middleware-example.ts # Middleware system
npx tsx examples/node/src/tool-loop-agent-example.ts # Autonomous tool agents
npx tsx examples/node/src/v6-tool-approval-example.ts # Tool execution approval (AI SDK v6)
npx tsx examples/node/src/v6-structured-output-example.ts # Structured output + tools (AI SDK v6)
npx tsx examples/node/src/v6-agent-example.ts # Advanced agent patterns (AI SDK v6)
npx tsx examples/node/src/json-repair-example.ts # JSON repair and objectGenerationOptions
npx tsx examples/node/src/test-cascade-repair.ts # Cascade repair (jsonrepair → enhancedRepairText)
npx tsx examples/node/src/test-cascade-repair.ts --llm # Same with LLM object generationTry the live browser example at examples/browser:
cd examples/browser
npm install
npm run devFeatures real-time text generation, model configuration UI, and proper CORS setup.
- Full Documentation - Complete API reference and advanced features
- Custom Ollama Instance - Connect to remote Ollama servers
- Tool Calling Guide - Function calling with Ollama models
- Reasoning Support - Chain-of-thought with DeepSeek-R1
- Browser Setup - CORS configuration and proxy setup
- Reranking - Document relevance ranking with embeddings
- Middleware System - Model wrapping and customization
- ToolLoopAgent - Autonomous agents with tool loops
- Streaming Utilities - Stream manipulation helpers
- Automatic JSON Repair - Cascade repair (jsonrepair + Ollama-specific) for object generation
Compatible with any model in your Ollama installation:
- Chat:
llama3.2,mistral,phi4-mini,qwen2.5,codellama - Vision:
llava,bakllava,llama3.2-vision,minicpm-v - Embeddings:
nomic-embed-text,all-minilm,mxbai-embed-large - Reasoning:
deepseek-r1:7b,deepseek-r1:1.5b,deepseek-r1:8b - Cloud Models (for web search):
qwen3-coder:480b-cloud,gpt-oss:120b-cloud
This project uses pnpm workspaces and Turborepo. Quick commands:
# Clone and setup
git clone https://github.com/jagreehal/ai-sdk-ollama.git
cd ai-sdk-ollama && pnpm install
# Build everything
pnpm build
# Run tests
pnpm test
# Run examples
npx tsx examples/node/src/basic-chat.tsContributing: Fork → feature branch → tests → PR. See CLAUDE.md for development guidelines.
MIT © Jag Reehal
See LICENSE for details.