Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
sickn33 avatar

Context Window Management

  • 1.3k installs
  • 44k repo stars
  • Updated July 27, 2026
  • sickn33/antigravity-awesome-skills

context-window-management provides LLM context strategies for summarization, trimming, routing, and tiered token budgets.

About

The context-window-management skill teaches strategies for managing LLM context windows including summarization, trimming, routing, and avoiding context rot in multi-turn systems. Capabilities span context-engineering, summarization, trimming, routing, token-counting, and prioritization with prerequisites in LLM fundamentals and prompt engineering. Scope covers optimization strategies not RAG implementation, fine-tuning, or embedding model details. Tiered Context Strategy defines maxTokens thresholds choosing full, summarize, or rag strategies with model selection per tier from haiku through sonnet classes. prepareContext switches on strategy to return full messages, summarized old content plus recent tail, or retrieved relevant chunks plus recent messages. Serial Position Optimization places system prompts first, critical context immediately after, summarized history in the middle, and current query at the end leveraging primacy and recency effects. Ecosystem tools include tiktoken counting, LangChain utilities, and Claude API caching support. Patterns address when to summarize versus retrieve and how to prevent unbounded conversation growth. Does not cover RAG pipeline implement.

  • Tiered strategy selects full, summarize, or rag based on token count thresholds.
  • Serial position optimization places critical context after system prompt at start.
  • Capabilities include trimming, routing, token counting, and prioritization.
  • Explicit boundaries: no RAG implementation, fine-tuning, or embedding model coverage.
  • prepareContext merges summaries or retrieved chunks with recent message tail.

Context Window Management by the numbers

  • 1,277 all-time installs (skills.sh)
  • +37 installs in the week ending Jul 28, 2026 (Skillselion tracking)
  • Ranked #876 of 16,659 AI & Agent Building skills by installs in the Skillselion catalog
  • Security screen: LOW risk (skills.sh audit)
  • Data as of Jul 28, 2026 (Skillselion catalog sync)
At a glance

context-window-management capabilities & compatibility

Capabilities
tiered context selection · summarize and rag strategy switching · serial position prompt layout · token counting guidance · context rot avoidance
Use cases
orchestration · planning
From the docs

What context-window-management says it does

Strategies for managing LLM context windows including summarization, trimming, routing, and avoiding context rot
SKILL.md
Does_not_cover: RAG implementation details, Model fine-tuning, Embedding models
SKILL.md
npx skills add https://github.com/sickn33/antigravity-awesome-skills --skill context-window-management

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs1.3k
repo stars44k
Security audit3 / 3 scanners passed
Last updatedJuly 27, 2026
Repositorysickn33/antigravity-awesome-skills

How do I keep long agent conversations within context limits without losing important information?

Manage LLM context windows with summarization, trimming, routing, and token-aware tier strategies.

Who is it for?

Agent developers building multi-turn systems needing context budget discipline.

Skip if: Skip when implementing full RAG pipelines; this skill covers strategy not RAG details.

When should I use this skill?

User asks about context window limits, summarization, trimming, or context rot in agents.

What you get

Tier-selected context with summarization or retrieval plus serial-position-optimized prompts.

  • context budget plan
  • summarized history blocks
  • routed sub-session prompts

By the numbers

  • Documents six named capabilities including summarization, trimming, routing, and token counting
  • Catalog entry dated 2026-02-27 under Apache 2.0 license

Files

SKILL.mdMarkdownGitHub ↗

Context Window Management

Strategies for managing LLM context windows including summarization, trimming, routing, and avoiding context rot

Capabilities

  • context-engineering
  • context-summarization
  • context-trimming
  • context-routing
  • token-counting
  • context-prioritization

Prerequisites

  • Knowledge: LLM fundamentals, Tokenization basics, Prompt engineering
  • Skills_recommended: prompt-engineering

Scope

  • Does_not_cover: RAG implementation details, Model fine-tuning, Embedding models
  • Boundaries: Focus is context optimization, Covers strategies not specific implementations

Ecosystem

Primary_tools

  • tiktoken - OpenAI's tokenizer for counting tokens
  • LangChain - Framework with context management utilities
  • Claude API - 200K+ context with caching support

Patterns

Tiered Context Strategy

Different strategies based on context size

When to use: Building any multi-turn conversation system

interface ContextTier { maxTokens: number; strategy: 'full' | 'summarize' | 'rag'; model: string; }

const TIERS: ContextTier[] = [ { maxTokens: 8000, strategy: 'full', model: 'claude-3-haiku' }, { maxTokens: 32000, strategy: 'full', model: 'claude-3-5-sonnet' }, { maxTokens: 100000, strategy: 'summarize', model: 'claude-3-5-sonnet' }, { maxTokens: Infinity, strategy: 'rag', model: 'claude-3-5-sonnet' } ];

async function selectStrategy(messages: Message[]): ContextTier { const tokens = await countTokens(messages);

for (const tier of TIERS) { if (tokens <= tier.maxTokens) { return tier; } } return TIERS[TIERS.length - 1]; }

async function prepareContext(messages: Message[]): PreparedContext { const tier = await selectStrategy(messages);

switch (tier.strategy) { case 'full': return { messages, model: tier.model };

case 'summarize': const summary = await summarizeOldMessages(messages); return { messages: [summary, ...recentMessages(messages)], model: tier.model };

case 'rag': const relevant = await retrieveRelevant(messages); return { messages: [...relevant, ...recentMessages(messages)], model: tier.model }; } }

Serial Position Optimization

Place important content at start and end

When to use: Constructing prompts with significant context

// LLMs weight beginning and end more heavily // Structure prompts to leverage this

function buildOptimalPrompt(components: { systemPrompt: string; criticalContext: string; conversationHistory: Message[]; currentQuery: string; }): string { // START: System instructions (always first) const parts = [components.systemPrompt];

// CRITICAL CONTEXT: Right after system (high primacy) if (components.criticalContext) { parts.push(## Key Context\n${components.criticalContext}); }

// MIDDLE: Conversation history (lower weight) // Summarize if long, keep recent messages full const history = components.conversationHistory; if (history.length > 10) { const oldSummary = summarize(history.slice(0, -5)); const recent = history.slice(-5); parts.push(## Earlier Conversation (Summary)\n${oldSummary}); parts.push(## Recent Messages\n${formatMessages(recent)}); } else { parts.push(## Conversation\n${formatMessages(history)}); }

// END: Current query (high recency) // Restate critical requirements here parts.push(## Current Request\n${components.currentQuery});

// FINAL: Reminder of key constraints parts.push(Remember: ${extractKeyConstraints(components.systemPrompt)});

return parts.join('\n\n'); }

Intelligent Summarization

Summarize by importance, not just recency

When to use: Context exceeds optimal size

interface MessageWithMetadata extends Message { importance: number; // 0-1 score hasCriticalInfo: boolean; // User preferences, decisions referenced: boolean; // Was this referenced later? }

async function smartSummarize( messages: MessageWithMetadata[], targetTokens: number ): Message[] { // Sort by importance, preserve order for tied scores const sorted = [...messages].sort((a, b) => (b.importance + (b.hasCriticalInfo ? 0.5 : 0) + (b.referenced ? 0.3 : 0)) - (a.importance + (a.hasCriticalInfo ? 0.5 : 0) + (a.referenced ? 0.3 : 0)) );

const keep: Message[] = []; const summarizePool: Message[] = []; let currentTokens = 0;

for (const msg of sorted) { const msgTokens = await countTokens([msg]); if (currentTokens + msgTokens < targetTokens * 0.7) { keep.push(msg); currentTokens += msgTokens; } else { summarizePool.push(msg); } }

// Summarize the low-importance messages if (summarizePool.length > 0) { const summary = await llm.complete(` Summarize these messages, preserving:

  • Any user preferences or decisions
  • Key facts that might be referenced later
  • The overall flow of conversation

Messages: ${formatMessages(summarizePool)} `);

keep.unshift({ role: 'system', content: [Earlier context: ${summary}] }); }

// Restore original order return keep.sort((a, b) => a.timestamp - b.timestamp); }

Token Budget Allocation

Allocate token budget across context components

When to use: Need predictable context management

interface TokenBudget { system: number; // System prompt criticalContext: number; // User prefs, key info history: number; // Conversation history query: number; // Current query response: number; // Reserved for response }

function allocateBudget(totalTokens: number): TokenBudget { return { system: Math.floor(totalTokens 0.10), // 10% criticalContext: Math.floor(totalTokens 0.15), // 15% history: Math.floor(totalTokens 0.40), // 40% query: Math.floor(totalTokens 0.10), // 10% response: Math.floor(totalTokens * 0.25), // 25% }; }

async function buildWithBudget( components: ContextComponents, modelMaxTokens: number ): PreparedContext { const budget = allocateBudget(modelMaxTokens);

// Truncate/summarize each component to fit budget const prepared = { system: truncateToTokens(components.system, budget.system), criticalContext: truncateToTokens( components.criticalContext, budget.criticalContext ), history: await summarizeToTokens(components.history, budget.history), query: truncateToTokens(components.query, budget.query), };

// Reallocate unused budget const used = await countTokens(Object.values(prepared).join('\n')); const remaining = modelMaxTokens - used - budget.response;

if (remaining > 0) { // Give extra to history (most valuable for conversation) prepared.history = await summarizeToTokens( components.history, budget.history + remaining ); }

return prepared; }

Validation Checks

No Token Counting

Severity: WARNING

Message: Building context without token counting. May exceed model limits.

Fix action: Count tokens before sending, implement budget allocation

Naive Message Truncation

Severity: WARNING

Message: Truncating messages without summarization. Critical context may be lost.

Fix action: Summarize old messages instead of simply removing them

Hardcoded Token Limit

Severity: INFO

Message: Hardcoded token limit. Consider making configurable per model.

Fix action: Use model-specific limits from configuration

No Context Management Strategy

Severity: WARNING

Message: LLM calls without context management strategy.

Fix action: Implement context management: budgets, summarization, or RAG

Collaboration

Delegation Triggers

  • retrieval|rag|search -> rag-implementation (Need retrieval system)
  • memory|persistence|remember -> conversation-memory (Need memory storage)
  • cache|caching -> prompt-caching (Need caching optimization)

Complete Context System

Skills: context-window-management, rag-implementation, conversation-memory, prompt-caching

Workflow:

1. Design context strategy
2. Implement RAG for large corpuses
3. Set up memory persistence
4. Add caching for performance

Related Skills

Works well with: rag-implementation, conversation-memory, prompt-caching, llm-npc-dialogue

When to Use

  • User mentions or implies: context window
  • User mentions or implies: token limit
  • User mentions or implies: context management
  • User mentions or implies: context engineering
  • User mentions or implies: long context
  • User mentions or implies: context overflow

Limitations

  • Use this skill only when the task clearly matches the scope described above.
  • Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
  • Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.

Related skills

How it compares

Use context-window-management for live session hygiene; retrieval and embedding pipelines solve knowledge access, not overflowing chat history.

FAQ

What strategies exist per tier?

full under small limits, summarize at medium, rag at very large token counts.

Does it implement RAG?

No; RAG implementation details are explicitly out of scope.

Why serial position matters?

LLMs weight beginning and end more; structure prompts accordingly.

Is Context Window Management safe to install?

skills.sh reports 3 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.

AI & Agent Buildingagentsautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.