← back

mantisagent

2026

open-source agent SDK for any LLM

the mantis agent loop. press play, or step through it
Agent(model="qwen2.5:7b")ollama · path B, tool calls written as text
model call 0 of 54.8k / 32k
loading the scene…drag to look around
ready

Press play to watch one run

A local qwen2.5:7b through Ollama, on path B, asked to make an HTTP client respect Retry-After. It takes four tool calls. One comes back as broken JSON, and one needs your permission.

messages[] sent in full every step4.8k / 32k
  1. systempreamble + 3 tool schemas4.8k

a reimplementation of the Claude Agent SDK that works with any model, local or hosted. write the agent once and run it on Ollama, Groq, Together, Fireworks or vLLM by changing the model name. it handles tool calls for models that have no tool support of their own, and comes with sessions, budget limits, MCP and a terminal coding agent.

0:00 / 0:00

anthropic-style api, nine provider adapters, three ways to get a tool call out of a model

python · asyncio · anyio · httpx · msgspec · sse · sqlite · jsonl · mcp · json-schema · ollama · vllm · oauth

the problemEvery provider puts tool calls somewhere different

Anthropic's claude-agent-sdk is well designed, but it only works with Anthropic's models. Running the same agent against Qwen on my own GPU box meant writing a second agent loop.

Translating field names between APIs only gets you so far. Anthropic and OpenAI-compatible servers both have a dedicated field for tool calls, in different places. Most open-weight models have no such field. They write the call out as text, and often go on to make up the result.

mantis-agent-sdk
the PyPI package, currently 2.63.0
mantis_agent
the import name, and a drop-in replacement for claude_agent_sdk
mantis
the terminal coding agent, built on the same engine
where each kind of API puts a tool call
anthropicopenai-compatibleopen weights
a tool_use block in the messagea tool_calls array next to the contentno tool field
a tool_result block in the next user turna separate tool message after itthe call written out as text
a dedicated, typed fielda dedicated, typed fieldoften followed by a made-up result

the ideaDecide how to call tools once, per model and backend

Whether tool calling works depends on the model and the server together. Qwen2.5 returns proper tool_calls through vLLM, but not through a bare llama.cpp server. So both sides are described in tables, and one function picks the approach when the agent is created. The agent loop never has to check again.

The tables are kept by hand: 38 open-weight models, 24 hosted ones, and 17 family defaults for anything it doesn't recognise. Nobody publishes reliable data on which checkpoints produce well-formed tool calls, and testing at runtime would cost a full generation.

three ways to get a tool call out of a model
A · nativeB · promptedC · prompted + grammar
tools[] sent in the requestthe protocol goes in the system promptB with constrained sampling
the server parses the callcalls are parsed out of the textthe server enforces valid json
can't come back malformedmalformed calls are commonno text between calls
fastest, fewest tokensschemas take up prompt spacefalls back to B for now

how it worksWhat happens in one turn

You can pass model="qwen2.5:7b" and nothing else. The backend is worked out from the model name. Profile matching is anchored, so api.deepseek.com picks up the DeepSeek profile and a self-hosted machine called deepseek-box doesn't.

On path B, each tool's schema is pretty-printed into the system prompt. I tried minifying it once and the failure rate went up. Smaller models lose track of which properties belong to which tool when the schemas run together.

one turn, from the model name to a tool result
callersdk coretool callsbackendAgent(model=…)no backend givenroutingname to backendurlcapabilitycheckpath A or Btool executorhooks,permissionstools[] +tool_calls<tool_call> inprompttext streamparsermodel serverpath Apath Bsse deltas

the hard partKeeping small models to the protocol

The protocol is five lines of system prompt. Getting a 7B model to follow it for a whole coding session is much harder, and most of the path B code is error handling.

Truncated calls were the hardest case. A call cut off at max_tokens and a call with a syntax error both fail to parse, but they need opposite instructions. Tell a truncated call to resend valid JSON and it gets cut off again, every turn, at the cost of a full generation each time.

parser
two states, parses when the tag closes, buffers under a kilobyte
split tags
holds back up to 16 characters that could be the start of <tool_call>
escaping
tool output is escaped, so it can't pass itself off as a tool call
repairs, tried in order
  1. 1rawparse as sent
  2. 2trimfirst balanced {…}
  3. 3escapefix stray quotes
  4. 4escape(trim)
  5. 5trim(escape)
  6. 6give upreturns is_error
telling a cut-off call from a broken one
cut offmalformed
valid so far, brackets still openbrackets close, the json is invalid
ran out of max_tokensa syntax mistake by the model
tell it: send less, add the rest aftertell it: resend valid json
the wrong advice loops foreverthe wrong advice truncates again

the stackWhat it's built on

It doesn't use the openai, anthropic or ollama client libraries. Every adapter makes raw HTTP calls against the documented API, because vendor SDKs tend to check model names against their own list, and that rejects the self-hosted models this is for.

Retries happen in a custom httpx transport, so every provider gets them without extra code. It retries POST requests, which is normally a bad idea. Here it's safe, because a completion request has no side effects and a partly read stream is never replayed.

the layers, top to bottom
  1. surfacesmantis terminal · mantis-agent cli · query() / Agent · claude_agent_sdk shim
  2. agent coretool registry · permissions + hooks · budget tracker · sessions + jsonl
  3. normalisationcapability tables · path A/B/C resolver · <tool_call> parser · lenient json repair
  4. providersanthropic messages · openai chat completions · ollama · llama.cpp · tgi · bedrock / vertex
  5. runtimeanyio · httpx + retry transport · msgspec · sqlite

parityHow close it is to the Claude Agent SDK

PARITY.md records the differences, checked against the Claude Code docs and the deobfuscated CLI. Some behaviour is copied exactly, such as only honouring the ultracode keyword in text a person typed, since that boundary protects against prompt injection.

The specific claim is that every official Claude Agent SDK Python example runs unchanged once you swap the import and the options class name. There's a test that checks it.

what parity covers
coverednot covered
the same query() and @tool APIstatuslines, themes, output styles
the same permission and hook rulesthe plugin marketplace
byte-compatible jsonl transcriptsSlack and Chrome integrations
official examples run after an import swapremote and mobile

where it standsWhat works, and what doesn't yet

The three tool-call paths, sessions with fork and resume, the MCP client and server with OAuth, and the mantis terminal agent all work. Most of the 6,878 tests run against a mock provider, so the wire-format code can be tested without a GPU or an API key.

The first thing I'd change is capability detection. Capabilities are declared by hand today. They should still be declared, but then checked once per backend with a cheap handshake and cached on disk, so a new endpoint gets the best path without waiting for a release.

works, and known issues
worksknown issues
paths A and B across 21 profilespath C is chosen, then falls back to B
fork, resume, jsonl transcriptstoken counts estimated as len // 4
27 hook events, 7 can block a callOpenRouter prices are placeholder zeros
per-model cost tracking in USDnew checkpoints need a row added by hand