AI
Tech
September 29, 2026

Context Window Bloat: How RAG Quietly Balloons Your Token Spend by 10x

Girish Bhat
SVP, Revefi

A Retrieval-Augmented Generated (RAG) pipeline that works correctly and a RAG pipeline that works efficiently often look identical in a demo. Both retrieve relevant documents, both generate a correct answer, both pass every quality check anyone runs before shipping. The difference only shows up in the token bill, and it usually shows up as a number nobody can fully explain, because the bloat accumulated one small decision at a time rather than arriving as a single obvious mistake.

Image 01: Prompt Engineering vs. Context Engineering | Source: Anthropic

Here's where that bloat actually comes from, and why it tends to get worse over time rather than better.

What Is Context Window Bloat and Why Does It Happen in RAG Pipelines?

Context window bloat is the accumulation of unnecessary tokens in what gets sent to the model on every call: retrieved documents that don't get used, conversation history that keeps growing, tool outputs that stay in context long after they've served their purpose. None of it is a bug in the traditional sense. Every individual piece usually got added for a reason. The bloat is the sum of a dozen reasonable-sounding decisions, each one small, none of them ever revisited once the system was working.

Why Does “Retrieve More Just in Case” Multiply Token Costs?

The most common source is over-retrieval. A RAG pipeline set to retrieve the top ten most relevant chunks, when the query only needed three, pays for seven chunks of tokens the model never uses meaningfully. This setting almost always gets chosen conservatively during development, when the priority is making sure the model has enough context to answer correctly, and it almost never gets revisited once the system ships, because tightening it feels like a risk to answer quality for a saving that's hard to see on any single call. Multiply that by every query the system handles and it becomes one of the largest line items in the whole pipeline, invisible because it's distributed across millions of calls rather than concentrated in one place anyone would notice.

Overlapping chunks compound this. Many chunking strategies deliberately overlap chunk boundaries to avoid splitting a relevant sentence across two chunks, which is a reasonable choice for accuracy, but it means a chunk of content can get retrieved and paid for more than once if it sits near a boundary that gets crossed by multiple queries.

How Does Context Bloat Compound in Agentic Workflows?

Chat-style RAG has one retrieval step per user turn, which caps how bad this problem gets. Agentic workflows don't have that cap. An agent that retrieves context, calls a tool, retrieves more context to interpret the tool's output, calls another tool, and retries a step that failed, carries all of that forward in its context window unless something actively prunes it. Tool outputs in particular tend to be verbose, a full API response when the agent only needs one field, and they often sit in context for the rest of the task even after their useful information has been extracted and acted on.

Conversation history has the same problem in any multi-turn agent. Each new turn appends to what came before by default, and unless something is actively summarizing or trimming, a long-running agentic session can be carrying thousands of tokens of context that stopped being relevant several steps ago.

Image 02: LLM Context Bloat | Source: LangChain

Few-shot examples are a quieter version of the same pattern. Examples get added to a prompt to fix a specific failure case, they work, and they stay in the prompt indefinitely because nobody wants to risk removing them and reintroducing the failure. Over months, a prompt can accumulate a dozen examples where three would do the same job.

How Do You Measure and Fix Context Window Bloat?

The fix starts with visibility most teams don't have: tokens per call broken down by source, how many came from retrieved documents, how many from conversation history, how many from tool outputs, how many from the base prompt and examples. Without that breakdown, bloat is invisible, because the total token count alone doesn't tell you which part of the pipeline is responsible.

Once that visibility exists, the fixes are mostly straightforward. Tighten retrieval counts and add reranking so the model gets fewer, more relevant chunks instead of more, less relevant ones. Summarize or truncate conversation history past a certain point rather than carrying it forward indefinitely. Extract only the fields an agent actually needs from a tool response instead of passing the full payload into context. Periodically audit few-shot examples and remove the ones that are no longer pulling their weight.

None of these fixes require a different model or a different architecture. They require someone to actually look at what's in the context window on a representative sample of real calls, which is a smaller effort than it sounds like and is usually the single highest-leverage audit a team can run on an expensive pipeline.

The Takeaway

Context window bloat rarely arrives as one bad decision. It accumulates from a series of individually reasonable choices, generous retrieval settings, growing conversation history, verbose tool outputs, examples nobody wants to remove, none of which look expensive in isolation. The only way to catch it is to actually break down where the tokens in a call are coming from, because the total spend number alone will never point you to the fix.

Girish Bhat
SVP, Revefi
Girish Bhat is a seasoned technology expert with Engineering, Product and B2B marketing, product marketing and go-to-market (GTM) experience building and scaling high-impact teams at pioneering AI, data, observability, security, and cloud companies.
Blog FAQs
What is context window bloat in AI pipelines?
Context window bloat is the buildup of unnecessary tokens sent to a model on every call, including unused retrieved documents, growing conversation history, and verbose tool outputs left in context after they've served their purpose. It increases cost without improving answer quality.
Why do RAG pipelines waste so many tokens?
RAG pipelines often retrieve more documents than a query needs, using a conservative retrieval count set during development that's rarely revisited. Overlapping chunk boundaries can also cause the same content to be retrieved and paid for more than once.
How does context bloat get worse in agentic workflows?
Agentic workflows involve multiple retrieval and tool-call steps within a single task, and without active pruning, all of that context carries forward. Verbose tool outputs and accumulating conversation history are the two biggest contributors in multi-step agent sessions.
How do you find out how much context window bloat you have?
Break down token usage per call by source: retrieved documents, conversation history, tool outputs, and base prompt content. This reveals which part of the pipeline is driving cost, which a single total token count cannot show.
Does reducing retrieved context hurt answer quality?
Not necessarily. Adding reranking so the model receives fewer but more relevant chunks often maintains or improves answer quality while reducing token spend, since irrelevant retrieved content can dilute the model's attention rather than helping it.