open-source agent SDK for any LLM
Agent(model="qwen2.5:7b")ollama · path B, tool calls written as textA local qwen2.5:7b through Ollama, on path B, asked to make an HTTP client respect Retry-After. It takes four tool calls. One comes back as broken JSON, and one needs your permission.
a reimplementation of the Claude Agent SDK that works with any model, local or hosted. write the agent once and run it on Ollama, Groq, Together, Fireworks or vLLM by changing the model name. it handles tool calls for models that have no tool support of their own, and comes with sessions, budget limits, MCP and a terminal coding agent.
mantisagent.cc ↗github ↗pypi ↗
anthropic-style api, nine provider adapters, three ways to get a tool call out of a model
python · asyncio · anyio · httpx · msgspec · sse · sqlite · jsonl · mcp · json-schema · ollama · vllm · oauth
Anthropic's claude-agent-sdk is well designed, but it only works with Anthropic's models. Running the same agent against Qwen on my own GPU box meant writing a second agent loop.
Translating field names between APIs only gets you so far. Anthropic and OpenAI-compatible servers both have a dedicated field for tool calls, in different places. Most open-weight models have no such field. They write the call out as text, and often go on to make up the result.
| anthropic | openai-compatible | open weights |
|---|---|---|
| a tool_use block in the message | a tool_calls array next to the content | no tool field |
| a tool_result block in the next user turn | a separate tool message after it | the call written out as text |
| a dedicated, typed field | a dedicated, typed field | often followed by a made-up result |
Whether tool calling works depends on the model and the server together. Qwen2.5 returns proper tool_calls through vLLM, but not through a bare llama.cpp server. So both sides are described in tables, and one function picks the approach when the agent is created. The agent loop never has to check again.
The tables are kept by hand: 38 open-weight models, 24 hosted ones, and 17 family defaults for anything it doesn't recognise. Nobody publishes reliable data on which checkpoints produce well-formed tool calls, and testing at runtime would cost a full generation.
| A · native | B · prompted | C · prompted + grammar |
|---|---|---|
| tools[] sent in the request | the protocol goes in the system prompt | B with constrained sampling |
| the server parses the call | calls are parsed out of the text | the server enforces valid json |
| can't come back malformed | malformed calls are common | no text between calls |
| fastest, fewest tokens | schemas take up prompt space | falls back to B for now |
You can pass model="qwen2.5:7b" and nothing else. The backend is worked out from the model name. Profile matching is anchored, so api.deepseek.com picks up the DeepSeek profile and a self-hosted machine called deepseek-box doesn't.
On path B, each tool's schema is pretty-printed into the system prompt. I tried minifying it once and the failure rate went up. Smaller models lose track of which properties belong to which tool when the schemas run together.
The protocol is five lines of system prompt. Getting a 7B model to follow it for a whole coding session is much harder, and most of the path B code is error handling.
Truncated calls were the hardest case. A call cut off at max_tokens and a call with a syntax error both fail to parse, but they need opposite instructions. Tell a truncated call to resend valid JSON and it gets cut off again, every turn, at the cost of a full generation each time.
| cut off | malformed |
|---|---|
| valid so far, brackets still open | brackets close, the json is invalid |
| ran out of max_tokens | a syntax mistake by the model |
| tell it: send less, add the rest after | tell it: resend valid json |
| the wrong advice loops forever | the wrong advice truncates again |
It doesn't use the openai, anthropic or ollama client libraries. Every adapter makes raw HTTP calls against the documented API, because vendor SDKs tend to check model names against their own list, and that rejects the self-hosted models this is for.
Retries happen in a custom httpx transport, so every provider gets them without extra code. It retries POST requests, which is normally a bad idea. Here it's safe, because a completion request has no side effects and a partly read stream is never replayed.
PARITY.md records the differences, checked against the Claude Code docs and the deobfuscated CLI. Some behaviour is copied exactly, such as only honouring the ultracode keyword in text a person typed, since that boundary protects against prompt injection.
The specific claim is that every official Claude Agent SDK Python example runs unchanged once you swap the import and the options class name. There's a test that checks it.
| covered | not covered |
|---|---|
| the same query() and @tool API | statuslines, themes, output styles |
| the same permission and hook rules | the plugin marketplace |
| byte-compatible jsonl transcripts | Slack and Chrome integrations |
| official examples run after an import swap | remote and mobile |
The three tool-call paths, sessions with fork and resume, the MCP client and server with OAuth, and the mantis terminal agent all work. Most of the 6,878 tests run against a mock provider, so the wire-format code can be tested without a GPU or an API key.
The first thing I'd change is capability detection. Capabilities are declared by hand today. They should still be declared, but then checked once per backend with a cheap handshake and cached on disk, so a new endpoint gets the best path without waiting for a release.
| works | known issues |
|---|---|
| paths A and B across 21 profiles | path C is chosen, then falls back to B |
| fork, resume, jsonl transcripts | token counts estimated as len // 4 |
| 27 hook events, 7 can block a call | OpenRouter prices are placeholder zeros |
| per-model cost tracking in USD | new checkpoints need a row added by hand |