
Agentfit
- 1 repo stars
- Updated July 18, 2026
- MukundaKatta/agentfit-mcp
Agentfit is an MCP server that token-aware truncates chat histories to fit a model context budget.
About
Agentfit is an MCP server that token-aware truncates chat message histories so they fit inside your model’s context budget. developers shipping Claude Code, Cursor, or Codex agents often accumulate long threads of user, assistant, and tool messages; without a consistent strategy, requests fail, costs spike, or the model silently drops early context. Agentfit exposes that fitting step as a reusable tool over stdio, distributed as @mukundakatta/agentfit-mcp on npm at version 0.1.0. It suits agent-first products and internal coding assistants where you control the orchestration layer and need predictable context usage before each completion. It is not a full memory or RAG system—it focuses on budgeted truncation. Install the package, register the server in your MCP client, and invoke it when history length threatens to exceed limits during build and iterate loops.
- Token-aware truncation of message histories against a configurable context budget
- stdio MCP server via npm package @mukundakatta/agentfit-mcp
- Designed for agent workflows that outgrow default context limits mid-session
- Keeps truncation policy in one MCP tool instead of scattered prompt hacks
Agentfit by the numbers
- Data as of Jul 19, 2026 (Skillselion catalog sync)
claude mcp add agentfit -- npx -y @mukundakatta/agentfit-mcpAdd your badge
Show developers this MCP server is listed on Skillselion. Paste this into your README.
| repo stars | ★ 1 |
|---|---|
| Package | @mukundakatta/agentfit-mcp |
| Transport | STDIO |
| Auth | None |
| Last updated | July 18, 2026 |
| Repository | MukundaKatta/agentfit-mcp ↗ |
What it does
Shrink long agent chat histories to a token budget before each model call without hand-rolling truncation logic.
Who is it for?
Best when you're orchestrating multi-turn agent chats and need deterministic context sizing before API calls.
Skip if: Skip if you only need a one-shot prompt with no history or want semantic memory instead of truncation.
What you get
You get a repeatable MCP step that fits message history to your token budget before the model runs.
- Registered stdio MCP server
- Truncated message lists within a declared token budget
By the numbers
- Package @mukundakatta/agentfit-mcp version 0.1.0
- Transport type stdio
- Repository github.com/MukundaKatta/agentfit-mcp
README.md
agentfit-mcp
MCP server for @mukundakatta/agentfit. Lets Claude Desktop, Cursor, Cline, Windsurf, Zed, or any other MCP client estimate token counts and fit a chat history into a model's context budget on demand.
npx -y @mukundakatta/agentfit-mcp
Three tools:
count_tokens— estimate tokens in a string or chat-message array, with per-model estimator families (openai, anthropic, google, llama, default).fit_messages— drop messages from a chat history until under amaxTokensbudget. Supports drop-oldest, drop-middle, and priority strategies; honorspreserveSystem,preserveFirstN,preserveLastN.list_estimators— list the built-in estimator families.
Add to your client
Claude Desktop
Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"agentfit": {
"command": "npx",
"args": ["-y", "@mukundakatta/agentfit-mcp"]
}
}
}
Cursor
~/.cursor/mcp.json:
{
"mcpServers": {
"agentfit": {
"command": "npx",
"args": ["-y", "@mukundakatta/agentfit-mcp"]
}
}
}
Cline / Windsurf / Zed
Same shape as above. The server speaks plain MCP over stdio, so any client that supports stdio MCP servers will work.
Tool examples
count_tokens:
{ "input": "hello world", "model": "claude-sonnet-4-6" }
Returns:
{ "tokens": 4, "model": "claude-sonnet-4-6" }
fit_messages:
{
"messages": [
{ "role": "system", "content": "You are precise." },
{ "role": "user", "content": "long context..." },
{ "role": "assistant", "content": "..." },
{ "role": "user", "content": "final question" }
],
"maxTokens": 8000,
"model": "claude-sonnet-4-6",
"preserveSystem": true,
"preserveLastN": 2,
"strategy": "drop-oldest"
}
Returns:
{
"messages": [...],
"dropped": [...],
"tokens": { "before": 12000, "after": 7800, "budget": 8000 },
"fit": true
}
fit_messages always returns a structured result and never throws across the wire: if the budget is unreachable even after dropping all non-protected messages, you get fit: false with the partial result so the caller can decide what to do.
Why a separate MCP server
@mukundakatta/agentfit is a zero-dependency JavaScript library. This package wraps it as an MCP server so it's accessible from inside any MCP-aware AI assistant: ask Claude "how many tokens is this transcript?" or "trim this chat to 8k tokens preserving the system prompt and last 2 turns" and the assistant calls these tools directly.
Sibling MCP servers
Part of the agent-stack series, all @mukundakatta/*-mcp:
@mukundakatta/agentfit-mcp— Fit it. (this)@mukundakatta/agentguard-mcp— Sandbox it.@mukundakatta/agentsnap-mcp— Test it.@mukundakatta/agentvet-mcp— Vet it.@mukundakatta/agentcast-mcp— Validate it.
License
MIT
Recommended MCP Servers
How it compares
MCP context-budget utility, not an agent skill or embedding store.
FAQ
Who is agentfit for?
Developers building agent workflows who need to cap token usage on rolling chat histories.
When should I use agentfit?
Use it during agent development and operation whenever conversation length regularly nears your model’s context window.
How do I add agentfit to my agent?
Install @mukundakatta/agentfit-mcp from npm, add a stdio MCP server entry in Claude Code or Cursor, and call the truncation tool before sending history to the model.