From 21eff98b083999a9c57bce62f08aac61be210897 Mon Sep 17 00:00:00 2001 From: Luca Chang Date: Mon, 12 May 2025 15:59:13 -0700 Subject: [PATCH 1/3] Rewrite sampling documentation to focus on practical usage This change rewrites the sampling documentation to explain how sampling in terms of tool calls and demonstrate how to actually implement sampling workflows via the SDK. Most of the existing information has been reorganized into reference-level material. The intent of this change is to document the existing behavior of sampling prior to proposing any specification changes. --- docs/legacy/concepts/sampling.mdx | 471 +++++++++++++++++++++--------- 1 file changed, 327 insertions(+), 144 deletions(-) diff --git a/docs/legacy/concepts/sampling.mdx b/docs/legacy/concepts/sampling.mdx index 44516e788..f0486e1bd 100644 --- a/docs/legacy/concepts/sampling.mdx +++ b/docs/legacy/concepts/sampling.mdx @@ -1,9 +1,9 @@ --- title: "Sampling" -description: "Let your servers request completions from LLMs" +description: "Let your servers request completions from client LLMs" --- -Sampling is a powerful MCP feature that allows servers to request LLM completions through the client, enabling sophisticated agentic behaviors while maintaining security and privacy. +Sampling is a powerful MCP feature that allows servers to request LLM completions through the client, enabling sophisticated agentic behaviors such as agent-to-agent communication while maintaining security and privacy. @@ -11,70 +11,326 @@ This feature of MCP is not yet supported in the Claude Desktop client. -## How sampling works +## Overview -The sampling flow follows these steps: +Sampling allows an MCP server to request completions from (or "sample") an LLM owned by an MCP client. This enables complex collaborative interactions within other MCP features, such as [tools](/docs/concepts/tools) and [prompts](/docs/concepts/prompts). -1. Server sends a `sampling/createMessage` request to the client -2. Client reviews the request and can modify it -3. Client samples from an LLM -4. Client reviews the completion -5. Client returns the result to the server +Sampling enables agentic patterns such as: -This human-in-the-loop design ensures users maintain control over what the LLM sees and generates. +- Negotiating interactions with other LLM-controlled agents +- Providing interactive assistance to users +- Making decisions based on context +- Generating natural language data +- Handling multi-step tasks -## Message format +## Capabilities -Sampling requests use a standardized message format: +Sampling is only supported when interacting with clients that have declared support for the corresponding capability: ```typescript { - messages: [ - { - role: "user" | "assistant", - content: { - type: "text" | "image", - - // For text: - text?: string, - - // For images: - data?: string, // base64 encoded - mimeType?: string - } + capabilities: { + sampling: { } - ], - modelPreferences?: { - hints?: [{ - name?: string // Suggested model name/family - }], - costPriority?: number, // 0-1, importance of minimizing cost - speedPriority?: number, // 0-1, importance of low latency - intelligencePriority?: number // 0-1, importance of capabilities + } +} +``` + +## Implementing sampling + +Sampling can be a standalone server-initiated interaction, but works best when used within other MCP features. The following example demonstrates how sampling can be used within tool interactions: + +### Server + +```typescript +import assert from "node:assert"; +import { z } from "zod"; + +const server = new McpServer({ + name: "example-server", + version: "1.0.0", +}); + +// Add an addition tool +server.tool( + "add", + { a: z.number(), b: z.number() }, + async ({ a, b }, extra) => { + // Send the client a sampling request and wait for a response + const result = await extra.sendRequest( + { + method: "sampling/createMessage", + params: { + messages: [ + { + role: "user", + content: { + type: "text", + text: `Add ${a} and ${b}. Respond with only the sum of the two numbers.`, + }, + }, + ], + maxTokens: 1000, + systemPrompt: "You are very good at math.", + includeContext: "thisServer", + }, + }, + CreateMessageResultSchema, + ); + + // You should handle every message type, but for the sake of example, let's assume the client sent a text response + assert(result.content.type === "text"); + const completion = result.content.text; + + // Return the formatted result to the client + return { + content: [ + { type: "text", text: `The sum of ${a} and ${b} is ${completion}.` }, + ], + }; + }, +); + +const transport = new StdioServerTransport(); +await server.connect(transport); +``` + +### Client + +```typescript +const transport = new StdioClientTransport({ + command: "node", + args: ["server.js"], // Assumes both files are in the same directory +}); + +const client = new Client( + { + name: "example-client", + version: "1.0.0", + }, + { + capabilities: { + sampling: {}, + }, + }, +); + +async function sendPrompt(_prompt: string): Promise { + // In a real application, you would write code to make a request to an LLM + return "3"; +} + +// Define a request handler that determines how the client should respond to sampling requests +client.setRequestHandler(CreateMessageRequestSchema, async (request) => { + // In a real application, you would raise a prompt to a user before handling the request + // ... + + // Send the prompt to a model + const prompt = `${ + request.params.systemPrompt + }\n\n${request.params.messages.join("\n\n")}`; + const modelResult = await sendPrompt(prompt); + + // Return the response to the server + return { + model: "my-model", + role: "assistant", + content: { + type: "text", + text: modelResult, + }, + }; +}); + +await client.connect(transport); + +const toolResult = await client.callTool({ + name: "add", + arguments: { + a: 1, + b: 2, }, - systemPrompt?: string, - includeContext?: "none" | "thisServer" | "allServers", - temperature?: number, - maxTokens: number, - stopSequences?: string[], - metadata?: Record +}); +``` + +**Result:** + +```typescript +{ + content: [{ type: "text", text: "The sum of 1 and 2 is 3." }]; } ``` -## Request parameters +## Example sampling patterns + +Here are some examples of how servers can use sampling to enhance interactions: + +### Natural-language tool responses + +The implementer of a server without direct LLM access may want a tool to return the result of an operation in natural language to aid clients: + +```typescript +server.tool( + "retrieve_documents", + { query: z.string() }, + async ({ query }, { sendRequest }) => { + // Fetch data from some external data source + const documents = await fetchData(query); + + // Send the client a sampling request and wait for a response + const result = await sendRequest( + { + method: "sampling/createMessage", + params: { + messages: [ + { + role: "user", + content: { + type: "text", + text: `Reformat these documents into easy-to-consume summaries: ${documents}`, + }, + }, + ], + maxTokens: 8000, + }, + }, + CreateMessageResultSchema, + ); + + // Handle the result... + + // Return the formatted result to the client + const completion = result.content.text; + return { + content: [{ type: "text", text: completion }], + }; + }, +); +``` + +### Collaborative agent activities + +Sampling can also be used as part of collaborative multi-agent workflows: + +#### High-level interaction flow + +```mermaid +sequenceDiagram + participant Client as User Client + participant A1 as Agent 1 + participant A1ClientA2 as Agent 1 Client (to Agent 2) + participant A2 as Agent 2 + + Client->>A1: tools/call [fetch_complex_data] + A1->>A1: "We need to ask Agent 2 about this." + A1->>A2: tools/call [agent_2_fetch_data] + A2->>A2: "We need to ask Agent 1 why they want this, first." + A2->>A1: sampling/createMessage ["More info, please."] + A1->>A2: "My user wants to know $QUERY because $REASON." + A2->>A1: result [agent_2_fetch_data] + A1->>Client: result [fetch_complex_data] +``` + +#### Agent 1 + +```typescript +// An MCP client owned by Agent 1, which connects to an MCP server exposed by Agent 2 +const agent2Client = new Client( + { + name: "agent-2-client", + version: "1.0.0", + }, + { + capabilities: { + sampling: {}, + }, + }, +); + +// Agent 1 exposes a tool to the end-user's client, which requires calling Agent 2 +agent1Server.tool( + "fetch_complex_data", + { query: z.string() }, + async ({ query }) => { + // Query some LLM to determine what to do next (pseudocode) + const llmResult1 = await queryLLM(buildPrompt(query)); + + // Send a request to Agent 2 + const agent2ToolResult = await agent2Client.callTool({ + name: "agent_2_fetch_data", + arguments: { query }, + }); + + // Process the result, determining that the interaction is complete and the result can be returned to the end-users + const llmResult2 = await queryLLM(buildPrompt(query, agent2ToolResult)); + + // Return the result to the client + const completion = llmResult2.content.text; + return { + content: [{ type: "text", text: completion }], + }; + }, +); +``` + +#### Agent 2 + +```typescript +agent2Server.tool( + "agent_2_fetch_data", + { query: z.string() }, + async ({ query }, { sendRequest }) => { + // Query some LLM to determine what to do next, determining that more information is needed... + const llmResult1 = await queryLLM(buildPrompt(query)); + + // Send Agent 1 a sampling request and wait for a response + const agent1Result = await sendRequest( + { + method: "sampling/createMessage", + params: { + messages: [ + { + role: "user", + content: { + type: "text", + text: "Please provide more information regarding the purpose of your query.", + }, + }, + ], + maxTokens: 8000, + systemPrompt: "Context aggregated from Agent 2...", + }, + }, + CreateMessageResultSchema, + ); + + // Query the LLM again; it has what it needs, this time + const llmResult2 = await queryLLM(buildPrompt(query, agent1Result)); + + // Return the result to the client + const completion = llmResult2.content.text; + return { + content: [{ type: "text", text: completion }], + }; + }, +); +``` + +## Request parameter reference ### Messages The `messages` array contains the conversation history to send to the LLM. Each message has: -- `role`: Either "user" or "assistant" +- `role`: Either `"user"` or `"assistant"` - `content`: The message content, which can be: - - Text content with a `text` field - - Image content with `data` (base64) and `mimeType` fields + - `text` content with a `text` field + - `image` content with `data` (base64) and `mimeType` fields + - `audio` content with `data` (base64) and `mimeType` fields ### Model preferences -The `modelPreferences` object allows servers to specify their model selection preferences: +The optional `modelPreferences` request parameter allows servers to specify their model selection preferences: - `hints`: Array of model name suggestions that clients can use to select an appropriate model: @@ -95,7 +351,7 @@ An optional `systemPrompt` field allows servers to request a specific system pro ### Context inclusion -The `includeContext` parameter specifies what MCP context to include: +The optional `includeContext` parameter specifies what MCP context to include: - `"none"`: No additional context - `"thisServer"`: Include context from the requesting server @@ -105,54 +361,32 @@ The client controls what context is actually included. ### Sampling parameters -Fine-tune the LLM sampling with: +LLM sampling can be fine-tuned with additional optional parameters: - `temperature`: Controls randomness (0.0 to 1.0) - `maxTokens`: Maximum tokens to generate - `stopSequences`: Array of sequences that stop generation - `metadata`: Additional provider-specific parameters -## Response format +## Error handling -The client returns a completion result: +Robust error handling should: -```typescript -{ - model: string, // Name of the model used - stopReason?: "endTurn" | "stopSequence" | "maxTokens" | string, - role: "user" | "assistant", - content: { - type: "text" | "image", - text?: string, - data?: string, - mimeType?: string - } -} -``` +- Catch sampling failures +- Handle timeout errors +- Manage rate limits +- Validate responses +- Provide fallback behaviors +- Log errors appropriately -## Example request +## Limitations -Here's an example of requesting sampling from a client: +Be aware of these limitations: -```json -{ - "method": "sampling/createMessage", - "params": { - "messages": [ - { - "role": "user", - "content": { - "type": "text", - "text": "What files are in the current directory?" - } - } - ], - "systemPrompt": "You are a helpful file system assistant.", - "includeContext": "thisServer", - "maxTokens": 100 - } -} -``` +- Sampling depends on client capabilities +- Users and clients ultimately control sampling behavior +- Models have varying limits on availability and inference latency +- Not all content types may be supported by the available models ## Best practices @@ -169,81 +403,30 @@ When implementing sampling: 9. Test with various model parameters 10. Monitor sampling costs -## Human in the loop controls +### Human in the loop controls Sampling is designed with human oversight in mind: -### For prompts - -- Clients should show users the proposed prompt +- Users control which model is used - Users should be able to modify or reject prompts -- System prompts can be filtered or modified -- Context inclusion is controlled by the client - -### For completions - +- Clients control context inclusion +- Clients should show users the proposed prompt - Clients should show users the completion -- Users should be able to modify or reject completions - Clients can filter or modify completions -- Users control which model is used +- Clients and servers handle context size limits +- Servers request minimal necessary context +- Servers update system prompts as needed ## Security considerations When implementing sampling: - Validate all message content +- Audit sampling requests - Sanitize sensitive information -- Implement appropriate rate limits +- Control access to user data +- Control cost exposure +- Implement appropriate rate limits and timeouts - Monitor sampling usage - Encrypt data in transit -- Handle user data privacy -- Audit sampling requests -- Control cost exposure -- Implement timeouts - Handle model errors gracefully - -## Common patterns - -### Agentic workflows - -Sampling enables agentic patterns like: - -- Reading and analyzing resources -- Making decisions based on context -- Generating structured data -- Handling multi-step tasks -- Providing interactive assistance - -### Context management - -Best practices for context: - -- Request minimal necessary context -- Structure context clearly -- Handle context size limits -- Update context as needed -- Clean up stale context - -### Error handling - -Robust error handling should: - -- Catch sampling failures -- Handle timeout errors -- Manage rate limits -- Validate responses -- Provide fallback behaviors -- Log errors appropriately - -## Limitations - -Be aware of these limitations: - -- Sampling depends on client capabilities -- Users control sampling behavior -- Context size has limits -- Rate limits may apply -- Costs should be considered -- Model availability varies -- Response times vary -- Not all content types supported From 6bc252c38ea2781c0a0ab82398640f4252a9ed60 Mon Sep 17 00:00:00 2001 From: Luca Chang <131398524+LucaButBoring@users.noreply.github.com> Date: Tue, 13 May 2025 14:43:00 -0700 Subject: [PATCH 2/3] Rephrase natural-language sampling example intro Co-authored-by: Jonathan Hefner --- docs/legacy/concepts/sampling.mdx | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/legacy/concepts/sampling.mdx b/docs/legacy/concepts/sampling.mdx index f0486e1bd..f81f8bb50 100644 --- a/docs/legacy/concepts/sampling.mdx +++ b/docs/legacy/concepts/sampling.mdx @@ -167,7 +167,7 @@ Here are some examples of how servers can use sampling to enhance interactions: ### Natural-language tool responses -The implementer of a server without direct LLM access may want a tool to return the result of an operation in natural language to aid clients: +A tool without direct LLM access may want to perform an operation using natural language: ```typescript server.tool( From b7de769da885ed1779637868b4fb3896727cae8a Mon Sep 17 00:00:00 2001 From: Luca Chang Date: Wed, 28 May 2025 12:02:57 -0700 Subject: [PATCH 3/3] Rewrite sampling overview and example to reduce complexity --- docs/legacy/concepts/sampling.mdx | 25 +++++++++++++++++-------- 1 file changed, 17 insertions(+), 8 deletions(-) diff --git a/docs/legacy/concepts/sampling.mdx b/docs/legacy/concepts/sampling.mdx index f81f8bb50..a4d130a3a 100644 --- a/docs/legacy/concepts/sampling.mdx +++ b/docs/legacy/concepts/sampling.mdx @@ -3,7 +3,7 @@ title: "Sampling" description: "Let your servers request completions from client LLMs" --- -Sampling is a powerful MCP feature that allows servers to request LLM completions through the client, enabling sophisticated agentic behaviors such as agent-to-agent communication while maintaining security and privacy. +Sampling is a powerful MCP feature that allows servers to request LLM completions through the client, enabling sophisticated LLM-enhanced behaviors while maintaining security and privacy. @@ -13,15 +13,16 @@ This feature of MCP is not yet supported in the Claude Desktop client. ## Overview -Sampling allows an MCP server to request completions from (or "sample") an LLM owned by an MCP client. This enables complex collaborative interactions within other MCP features, such as [tools](/docs/concepts/tools) and [prompts](/docs/concepts/prompts). +Sampling allows an MCP server to request completions from (or "sample") an LLM controlled by an MCP client. +This enables servers to leverage an LLM as part of other MCP interactions, such as [tools](/docs/concepts/tools) and [prompts](/docs/concepts/prompts), without needing additional infrastructure or configuration to directly integrate with model providers themselves. -Sampling enables agentic patterns such as: +Sampling flows generally follow these steps: -- Negotiating interactions with other LLM-controlled agents -- Providing interactive assistance to users -- Making decisions based on context -- Generating natural language data -- Handling multi-step tasks +1. The server sends a `sampling/createMessage` request to the client, containing a prompt and other information +2. The client reviews the request and may modify it +3. The client requests the completion from its own LLM +4. The client reviews the completion +5. The client responds to the server with the LLM-generated completion ## Capabilities @@ -42,6 +43,9 @@ Sampling can be a standalone server-initiated interaction, but works best when u ### Server +This server exposes a single `add` tool which asks the client's LLM for the sum of the two inputs. Upon receiving a response from the client, the server extracts the result and returns the tool call result back to the client. +In effect, the server has made its own request from the client within that client's existing tool call request. + ```typescript import assert from "node:assert"; import { z } from "zod"; @@ -97,6 +101,9 @@ await server.connect(transport); ### Client +The client is responsible for collecting the server's system prompt and messages into a single request to send to the model provider it uses. +In this toy example, we simply have the client return a fixed text response, instead. + ```typescript const transport = new StdioClientTransport({ command: "node", @@ -155,6 +162,8 @@ const toolResult = await client.callTool({ **Result:** +Upon running this example, the sampling request executes within the tool call, and the final result is returned as expected. + ```typescript { content: [{ type: "text", text: "The sum of 1 and 2 is 3." }];