Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
MukundaKatta avatar

Agentfit

  • 1 repo stars
  • Updated July 18, 2026
  • MukundaKatta/agentfit-mcp

Agentfit is an MCP server that token-aware truncates chat histories to fit a model context budget.

About

Agentfit is an MCP server that token-aware truncates chat message histories so they fit inside your model’s context budget. developers shipping Claude Code, Cursor, or Codex agents often accumulate long threads of user, assistant, and tool messages; without a consistent strategy, requests fail, costs spike, or the model silently drops early context. Agentfit exposes that fitting step as a reusable tool over stdio, distributed as @mukundakatta/agentfit-mcp on npm at version 0.1.0. It suits agent-first products and internal coding assistants where you control the orchestration layer and need predictable context usage before each completion. It is not a full memory or RAG system—it focuses on budgeted truncation. Install the package, register the server in your MCP client, and invoke it when history length threatens to exceed limits during build and iterate loops.

  • Token-aware truncation of message histories against a configurable context budget
  • stdio MCP server via npm package @mukundakatta/agentfit-mcp
  • Designed for agent workflows that outgrow default context limits mid-session
  • Keeps truncation policy in one MCP tool instead of scattered prompt hacks

Agentfit by the numbers

  • Data as of Jul 19, 2026 (Skillselion catalog sync)
terminal
claude mcp add agentfit -- npx -y @mukundakatta/agentfit-mcp

Add your badge

Show developers this MCP server is listed on Skillselion. Paste this into your README.

Listed on Skillselion
repo stars1
Package@mukundakatta/agentfit-mcp
TransportSTDIO
AuthNone
Last updatedJuly 18, 2026
RepositoryMukundaKatta/agentfit-mcp

What it does

Shrink long agent chat histories to a token budget before each model call without hand-rolling truncation logic.

Who is it for?

Best when you're orchestrating multi-turn agent chats and need deterministic context sizing before API calls.

Skip if: Skip if you only need a one-shot prompt with no history or want semantic memory instead of truncation.

What you get

You get a repeatable MCP step that fits message history to your token budget before the model runs.

  • Registered stdio MCP server
  • Truncated message lists within a declared token budget

By the numbers

  • Package @mukundakatta/agentfit-mcp version 0.1.0
  • Transport type stdio
  • Repository github.com/MukundaKatta/agentfit-mcp
README.md

agentfit-mcp

MCP server for @mukundakatta/agentfit. Lets Claude Desktop, Cursor, Cline, Windsurf, Zed, or any other MCP client estimate token counts and fit a chat history into a model's context budget on demand.

npx -y @mukundakatta/agentfit-mcp

Three tools:

  • count_tokens — estimate tokens in a string or chat-message array, with per-model estimator families (openai, anthropic, google, llama, default).
  • fit_messages — drop messages from a chat history until under a maxTokens budget. Supports drop-oldest, drop-middle, and priority strategies; honors preserveSystem, preserveFirstN, preserveLastN.
  • list_estimators — list the built-in estimator families.

Add to your client

Claude Desktop

Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):

{
  "mcpServers": {
    "agentfit": {
      "command": "npx",
      "args": ["-y", "@mukundakatta/agentfit-mcp"]
    }
  }
}

Cursor

~/.cursor/mcp.json:

{
  "mcpServers": {
    "agentfit": {
      "command": "npx",
      "args": ["-y", "@mukundakatta/agentfit-mcp"]
    }
  }
}

Cline / Windsurf / Zed

Same shape as above. The server speaks plain MCP over stdio, so any client that supports stdio MCP servers will work.

Tool examples

count_tokens:

{ "input": "hello world", "model": "claude-sonnet-4-6" }

Returns:

{ "tokens": 4, "model": "claude-sonnet-4-6" }

fit_messages:

{
  "messages": [
    { "role": "system", "content": "You are precise." },
    { "role": "user", "content": "long context..." },
    { "role": "assistant", "content": "..." },
    { "role": "user", "content": "final question" }
  ],
  "maxTokens": 8000,
  "model": "claude-sonnet-4-6",
  "preserveSystem": true,
  "preserveLastN": 2,
  "strategy": "drop-oldest"
}

Returns:

{
  "messages": [...],
  "dropped": [...],
  "tokens": { "before": 12000, "after": 7800, "budget": 8000 },
  "fit": true
}

fit_messages always returns a structured result and never throws across the wire: if the budget is unreachable even after dropping all non-protected messages, you get fit: false with the partial result so the caller can decide what to do.

Why a separate MCP server

@mukundakatta/agentfit is a zero-dependency JavaScript library. This package wraps it as an MCP server so it's accessible from inside any MCP-aware AI assistant: ask Claude "how many tokens is this transcript?" or "trim this chat to 8k tokens preserving the system prompt and last 2 turns" and the assistant calls these tools directly.

Sibling MCP servers

Part of the agent-stack series, all @mukundakatta/*-mcp:

License

MIT

Recommended MCP Servers

How it compares

MCP context-budget utility, not an agent skill or embedding store.

FAQ

Who is agentfit for?

Developers building agent workflows who need to cap token usage on rolling chat histories.

When should I use agentfit?

Use it during agent development and operation whenever conversation length regularly nears your model’s context window.

How do I add agentfit to my agent?

Install @mukundakatta/agentfit-mcp from npm, add a stdio MCP server entry in Claude Code or Cursor, and call the truncation tool before sending history to the model.

AI & LLM Toolsagentsautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.