Sponsored by Deepsite.site

Together.ai Mcp

创建者
Leonfinn4 months ago
A Node.js Model Context Protocol (MCP) server that exposes Together AI's inference endpoints — chat completions, image generation, vision, and embeddings — as tools callable from Claude Desktop, Cursor, VS Code, and any other MCP-compatible client. I created this MCP due to an issue I was having accessing reasoning models through Together AI. Together AI's largest reasoning models (GLM-5, Qwen3.5-397B, MiniMax M2.5, Kimi K2.5) use a non-standard response format. During chain-of-thought generation, these models write their reasoning trace into choices[0].message.reasoning while leaving choices[0].message.content as an empty string. The final answer only appears in message.content once thinking is complete.
概览

together-ai-mcp

A Node.js Model Context Protocol (MCP) server that exposes Together AI's inference endpoints — chat completions, image generation, vision, and embeddings — as tools callable from Claude Desktop, Cursor, VS Code, and any other MCP-compatible client.

Why this exists

I created this MCP due to several issues I was having accessing models through Together AI.

1. Reasoning model silent empty responses

Together AI's largest reasoning models (GLM-5, Qwen3.5-397B, MiniMax M2.5, Kimi K2.5) write their chain-of-thought into non-standard response fields, and they exhaust the OpenAI SDK's default token budget before producing a final answer.

Two problems compound each other:

Token budget exhaustion. The OpenAI SDK sets a default max_tokens of 2048. For reasoning models, this budget is consumed entirely by the thinking phase — message.content is never populated. You get charged for tokens, no error is raised, and the response is silently empty.

Fragmented response fields. Different model families on Together AI write their output to different fields:

FieldUsed by
message.contentStandard models; Qwen (inline <think> tags)
message.reasoning_contentDeepSeek-style format
message.reasoningTogether AI format (GLM-5, MiniMax, Kimi)

Any code that only reads message.content — or even message.content \|\| message.reasoning — silently returns an empty string for some models.

// Broken — misses reasoning_content (DeepSeek format):
const text = message.content || message.reasoning || '';

// Fixed — covers all Together AI reasoning model formats:
const text = message.content || message.reasoning_content || message.reasoning || '';

The default max_tokens is raised to 8192 to give reasoning models enough budget to complete their chain of thought before producing a final answer.

2. Vision model failures

Using the OpenAI SDK's chat.completions.create() for vision requests fails silently against Together AI's vision API. Together AI requires stream: false to be set explicitly; the SDK may not send it. When it does fail, the SDK error contains no response body, making the root cause invisible.

// Broken — SDK may omit stream:false; errors are opaque:
const response = await openai.chat.completions.create({ model, messages });

// Fixed — raw fetch, explicit stream:false, full error body in exception:
const response = await fetch('https://api.together.xyz/v1/chat/completions', {
  method: 'POST',
  headers: { Authorization: `Bearer ${apiKey}`, 'Content-Type': 'application/json' },
  body: JSON.stringify({ model, messages, max_tokens, stream: false }),
});
if (!response.ok) {
  const body = await response.text();
  throw new Error(`Vision API error ${response.status}: ${body.slice(0, 200)}`);
}

Features

  • Chat completions — any Together AI text or reasoning model, with full prompt and multi-turn message support
  • Reasoning model support — correctly handles GLM-5, Qwen3.5-397B, MiniMax M2.5, Kimi K2.5 (see above)
  • Image generation — FLUX.1-dev, FLUX.1-schnell, Stable Diffusion XL; images saved to disk
  • Vision — analyse images via Llama 3.2 Vision or Qwen 2.5 VL
  • Embeddings — generate vectors for RAG/retrieval pipelines via BGE and Snowflake Arctic models

Installation

Prerequisites

Setup

git clone https://github.com/your-username/together-ai-mcp
cd together-ai-mcp
npm install
cp .env.example .env
# Edit .env and add your TOGETHER_API_KEY

Add to Claude Desktop

Edit ~/Library/Application Support/Claude/claude_desktop_config.json:

{
  "mcpServers": {
    "together-ai": {
      "command": "node",
      "args": ["/absolute/path/to/together-ai-mcp/index.js"],
      "env": {
        "TOGETHER_API_KEY": "your_api_key_here",
        "IMAGE_OUTPUT_DIR": "/path/to/save/images"
      }
    }
  }
}

See examples/claude-config.md for Cursor and VS Code configuration.


Tools

together_chat

Call any Together AI chat or reasoning model.

ParameterTypeDefaultDescription
modelstringmeta-llama/Llama-3.3-70B-Instruct-TurboModel ID
promptstringUser message (use this OR messages)
messagesarrayMulti-turn [{role, content}] array
systemstringSystem prompt (used with prompt only)
temperaturenumber0.70.0–2.0
max_tokensinteger8192Raised from SDK default to give reasoning models enough budget for chain-of-thought

together_generate_image

Generate images using FLUX or SDXL models.

ParameterTypeDefaultDescription
promptstringrequiredImage description
modelstringblack-forest-labs/FLUX.1-schnellModel ID
widthinteger1024Image width in pixels
heightinteger1024Image height in pixels
stepsinteger4Diffusion steps
ninteger1Number of images
negative_promptstringWhat to exclude

Images are saved as PNG files to IMAGE_OUTPUT_DIR.

Note: Image generation uses a direct fetch call rather than the OpenAI SDK's images.generate() because the SDK strips custom parameters like steps when calling Together AI's endpoint.

together_vision

Analyse an image using a vision model.

ParameterTypeDefaultDescription
promptstringrequiredQuestion or instruction
modelstringmeta-llama/Llama-3.2-11B-Vision-InstructModel ID
image_urlstringPublic image URL
image_pathstringLocal file path (converted to base64)
max_tokensinteger1024Max response length

together_embed

Generate text embeddings for RAG and retrieval pipelines.

ParameterTypeDefaultDescription
inputstring | string[]requiredText to embed
modelstringBAAI/bge-large-en-v1.5Embedding model ID

Models

The server works with any model available on Together AI's serverless API — just pass its model ID. No configuration changes are needed.

The tables below list the models I personally use. They are provided as a reference, not as a hard limit.

Finding model IDs

Browse all available models at api.together.ai/models. Each model's page shows its exact ID string. Pass that ID as the model parameter to any tool:

{
  "tool": "together_chat",
  "params": {
    "model": "any-model-id-from-together-ai",
    "prompt": "Hello"
  }
}

The only constraint is that image generation models must be called via together_generate_image, vision models via together_vision, and embedding models via together_embed — you cannot call an image model through together_chat.

Dedicated endpoints: Some models (e.g. meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8) require a dedicated endpoint rather than the serverless API. Calling these via this server will return a 400 error from Together AI.


Models I use

Chat / Reasoning

ModelIDNotes
Llama 3.3 70Bmeta-llama/Llama-3.3-70B-Instruct-TurboDefault — fast general-purpose
DeepSeek V3deepseek-ai/DeepSeek-V3Strong at code and reasoning
DeepSeek R1deepseek-ai/DeepSeek-R1Reasoning model
GLM-5 (744B)zai-org/GLM-5Reasoning model — requires fix above
Qwen3.5 397BQwen/Qwen3.5-397B-A17BReasoning model — requires fix above
MiniMax M2.5MiniMaxAI/MiniMax-M2.5Reasoning model — requires fix above
Kimi K2.5moonshotai/Kimi-K2.5Reasoning model — requires fix above
Qwen 2.5 7BQwen/Qwen2.5-7B-Instruct-TurboLightweight / low cost

Image generation

ModelID
FLUX.1-schnellblack-forest-labs/FLUX.1-schnell
FLUX.1-devblack-forest-labs/FLUX.1-dev
Stable Diffusion XLstabilityai/stable-diffusion-xl-base-1.0

Vision

ModelID
Llama 3.2 11B Visionmeta-llama/Llama-3.2-11B-Vision-Instruct
Qwen 2.5 VL 72BQwen/Qwen2.5-VL-72B-Instruct

Embeddings

ModelID
BGE LargeBAAI/bge-large-en-v1.5
M2-BERT 32Ktogethercomputer/m2-bert-80M-32k-retrieval
Snowflake ArcticSnowflake/snowflake-arctic-embed-m

Running tests

npm test

The test suite uses Node.js's built-in test runner and mocks all external dependencies — no API key required to run tests.


Project structure

together-ai-mcp/
├── index.js              # MCP server and handler logic
├── package.json
├── .env.example
├── test/
│   └── index.test.js     # Full test suite (node:test, no external framework)
└── examples/
    ├── chat.md           # Example prompts for each tool and model
    └── claude-config.md  # Configuration for Claude Desktop, Cursor, VS Code

Dependencies


License

MIT

服务器配置

{
  "mcpServers": {
    "together-ai": {
      "command": "node",
      "args": [
        "/absolute/path/to/together-ai-mcp/index.js"
      ],
      "env": {
        "TOGETHER_API_KEY": "your_api_key_here",
        "IMAGE_OUTPUT_DIR": "/path/to/save/images"
      }
    }
  }
}
推荐的 MCP Server
TraeBuild with Free GPT-4.1 & Claude 3.7. Fully MCP-Ready.
MiniMax MCPOfficial MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech, image generation and video generation APIs.
Y GuiA web-based graphical interface for AI chat interactions with support for multiple AI models and MCP (Model Context Protocol) servers.
Howtocook Mcp基于Anduin2017 / HowToCook (程序员在家做饭指南)的mcp server,帮你推荐菜谱、规划膳食,解决“今天吃什么“的世纪难题; Based on Anduin2017/HowToCook (Programmer's Guide to Cooking at Home), MCP Server helps you recommend recipes, plan meals, and solve the century old problem of "what to eat today"
RedisA Model Context Protocol server that provides access to Redis databases. This server enables LLMs to interact with Redis key-value stores through a set of standardized tools.
MCP AdvisorMCP Advisor & Installation - Use the right MCP server for your needs
Baidu Map百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Zhipu Web SearchZhipu Web Search MCP Server is a search engine specifically designed for large models. It integrates four search engines, allowing users to flexibly compare and switch between them. Building upon the web crawling and ranking capabilities of traditional search engines, it enhances intent recognition capabilities, returning results more suitable for large model processing (such as webpage titles, URLs, summaries, site names, site icons, etc.). This helps AI applications achieve "dynamic knowledge acquisition" and "precise scenario adaptation" capabilities.
BlenderBlenderMCP connects Blender to Claude AI through the Model Context Protocol (MCP), allowing Claude to directly interact with and control Blender. This integration enables prompt assisted 3D modeling, scene creation, and manipulation.
Serper MCP ServerA Serper MCP Server
AiimagemultistyleA Model Context Protocol (MCP) server for image generation and manipulation using fal.ai's Stable Diffusion model.
Amap Maps高德地图官方 MCP Server
DeepChatYour AI Partner on Desktop
EdgeOne Pages MCPAn MCP service designed for deploying HTML content to EdgeOne Pages and obtaining an accessible public URL.
WindsurfThe new purpose-built IDE to harness magic
Jina AI MCP ToolsA Model Context Protocol (MCP) server that integrates with Jina AI Search Foundation APIs.
ChatWiseThe second fastest AI chatbot™
Tavily Mcp
Playwright McpPlaywright MCP server
Visual Studio Code - Open Source ("Code - OSS")Visual Studio Code
CursorThe AI Code Editor