Why RAG Fails on Error Codes, Config Keys and CLI Flags

Written by

Anton Malling

•

Updated

Short answer

Embeddings encode meaning. An identifier does not have meaning, it has identity.

ERR_CONN_REFUSED_42 and ERR_CONN_REFUSED_43 mean almost exactly the same thing to an embedding model, which is why a vector search for one will cheerfully return the other. --max-retries and --max-redirects are close neighbours in vector space and completely different flags at the command line. The model is not malfunctioning. It is doing precisely what it was trained to do, which is to map things that mean similar things to nearby points, and that is the wrong operation for a string whose entire job is to be distinguishable from its neighbours.

The fix is keyword search running alongside embedding search, because keyword scoring is inverted: rare terms count for more, not less. Everything else here is detail about where that goes wrong anyway.

kapa.ai is an LLM-powered agentic retrieval platform purpose-built for technical knowledge, and exact-identifier retrieval is a large part of why purpose-built matters. Specifics below about how one system does it are documentation, not a recommendation.

Why this matters more than it used to

For a long time the consumer of a retrieval system was a person reading an answer on a screen. A person who gets back the wrong error code has a reasonable chance of noticing, because they were already looking at the real one in their terminal.

That is no longer the main consumer. A retrieval layer today is queried by coding agents over MCP, by support automation, by internal assistants, and by the company's own product through an API. None of those read critically. An agent that asks your knowledge base for the right flag and receives --max-redirects instead of --max-retries does not pause. It writes the flag into a config file, runs it, and reports success.

That changes the cost of this particular failure. A near-miss on an identifier used to produce a confused reader. It now produces a wrong action taken confidently, several steps upstream of anyone who could catch it.

The failure, precisely

Ask a knowledge base "what does error E_QUOTA_0142 mean" and watch one of three things happen.

It returns E_QUOTA_0143, with a citation. The citation is real, the page is real, the answer is wrong, and nothing in the response signals the substitution.

It returns a general passage about quota errors that never mentions the specific code. This is the honest failure, and it still wastes the request.

It declines. If your system declines rather than guessing, this is the correct outcome and it is dramatically better than the first case, but it is still a retrieval failure and it should still show up in your analytics.

The same shape applies to configuration keys, CLI flags, API parameter names, environment variables, status codes, SDK method names, event names and metric names. These are the things technical questions are most precise about, and they are exactly the things dense retrieval handles worst.

Three separate mechanisms, not one

It is tempting to treat this as a single problem. It is three, and they need different fixes.

Tokenization shreds the identifier

An embedding model does not see max_retries_backoff_ms. It sees that string cut into subword pieces, several of which are meaningless fragments, and it builds a representation from the pieces. Two identifiers sharing most of their pieces end up sharing most of their representation. This is the mechanism behind near-miss retrieval, and it gets worse the more structured your naming convention is, which is to say it gets worse the better engineered your product is.

Rare terms get averaged away

A chunk's embedding is a single vector representing the whole chunk. A 400-word passage that mentions E_QUOTA_0142 once, in a table row, produces a vector dominated by the surrounding prose about quotas and limits. The identifier is in there somewhere, contributing a small fraction of a vector that is mostly about something else. Query for the identifier alone and the chunk you want may not be in the top hundred.

This is the mechanism that surprises people, because the content is unambiguously present in the index. Presence is not retrievability.

Nearest-neighbour is the wrong question

Underneath both of the above is a simpler point. You wanted exact match and you performed approximate nearest-neighbour search. Vector search has no concept of exact, only of close, and it will always return its best guess at close. There is no threshold that turns approximate search into exact search, because the distances are not calibrated to mean anything absolute.

Reranking does not save you

This is the part teams get wrong most often, and it is worth being blunt about.

A reranker reorders the candidates that retrieval already returned. If the chunk containing E_QUOTA_0142 was not in the candidate set, no cross-encoder will put it at the top, because it was never a candidate. Adding a reranker to a pipeline with an identifier recall problem improves your numbers on the queries that were already working and does nothing at all for the ones that were failing.

This is a recall failure wearing the costume of a ranking failure. The two look similar in an evaluation summary and have nothing in common in their fix.

There is a related point worth knowing even when recall is fine. Conventional rerankers score each chunk against the query independently, so they cannot tell that three of the chunks they ranked highly say the same thing. kapa's work on context pruning grades the retrieved set together rather than chunk by chunk, dropping about 68% of the context while keeping about 96% of recall. That helps precision, and precision is a real problem. It is a different problem from the one in this post.

What actually fixes it

Keyword search, scored so rare terms win

Keyword scoring works on the opposite principle to embeddings. A term appearing in few documents is treated as highly informative, and a term appearing everywhere is treated as nearly worthless. E_QUOTA_0142 appears on one page in your entire corpus, which makes it the single most discriminating token available, which is exactly the property you need.

kapa's retrieval documentation puts it directly: "Rare words count for much more than common ones, so an error code, a configuration key, or a command name leads straight to the passages that mention it."

Running keyword and embedding search together and fusing the results is usually described as hybrid retrieval. The fusion step matters, because the two methods produce scores on incompatible scales. Reciprocal rank fusion, which combines by position rather than by score, is the common answer and is covered in more detail in how to build a RAG pipeline from scratch.

Pulling the identifier out of the request

Requests do not arrive as E_QUOTA_0142. They arrive as "we started seeing E_QUOTA_0142 in prod this morning after bumping our plan, what changed?"

Embed that whole sentence and the identifier is one token in fifty, contributing almost nothing to the resulting vector. Hand the whole sentence to keyword search and the common words drag in noise. The identifier has to be extracted and searched for on its own, as a separate query, before any of the above helps.

This is why hybrid retrieval alone is not sufficient. Something has to decide that E_QUOTA_0142 is the load-bearing part of that request and issue a query for it specifically.

Context on the chunk

The second failure mode, where a config key appears on forty pages, is not a search problem at all.

If a chunk reads timeout: 30 with no surrounding indication of which service, which config file or which version it belongs to, then no retrieval method can pick the right one, because the chunk does not contain the information needed to distinguish it. You cannot rank your way out of missing information.

The fix is upstream: chunks need enough context attached at ingestion that they are individually identifiable. Contextual retrieval, where a short description of the chunk's place in its parent document is prepended before indexing, exists for this reason.

How kapa approaches it

Retrieval runs in one of two modes, selectable per request. The deep mode is an agent with six tools rather than a single search: a topology of the knowledge base built from the sources themselves, keyword search, embedding search, rerank, read document, and prune. It plans, searches, reads what it found, and searches again until it is confident it has gathered what the request needs. The default mode carries the same tools excluding pruning and reading full documents, because those add latency and require more reasoning.

kapa.ai

Three properties of that design matter for identifiers specifically.

Keyword search is a first-class tool rather than a fallback, and the agent can choose it directly when the request contains something that looks like an exact term.

Several queries can be issued for one request. A question containing one identifier and three concepts does not have to be answered from a single search over the whole sentence.

Read document exists because identifiers frequently live in tables. The relevant row is often retrievable while the header that gives it meaning sits outside the chunk. Being able to go back and read the rest of the table is the difference between returning a value and returning a value the caller can use.

The search is bounded three ways, by a maximum number of rounds, a time limit, and a ceiling on how many passages the context can hold. Most queries take about two rounds, and roughly one run in a hundred reaches a budget.

Two things outside retrieval matter as much. The system declines when the sources do not cover a request, which for this failure mode is the difference between a wrong error code and no error code. And coverage gap analytics cluster what could not be answered, which for identifier queries usually surfaces something more useful than a retrieval bug: a list of error codes nobody ever documented.

One structural point. Because the same retrieval layer serves the widget, the support surfaces, the internal assistant and whatever calls the API or the MCP server, fixing identifier recall once fixes it for all of them. The alternative, where each surface has its own search, means finding and fixing this failure separately in each place, and discovering it separately in each place too.

How to test whether you have this problem

Generic evaluation sets will not find this. They are built from natural questions, and natural questions under-represent exact identifiers relative to how often real callers ask about them.

Build a separate identifier test set. It takes an afternoon and it is the highest-yield evaluation work available for a technical knowledge base.

  1. Pull thirty identifiers straight out of your own content. Error codes, config keys, CLI flags, API parameters, environment variables. Take them from different parts of the corpus.

  2. Ask for each one in a full sentence, the way a real request arrives, not as a bare string. The bare string is the easy case and it is not the case you have.

  3. Score exact match only. The result is correct if it describes that identifier. A response about the general topic is a failure, not partial credit.

  4. Deliberately include five near-miss pairs, identifiers differing by one character or one word. This is where near-miss retrieval shows up and it will not show up anywhere else.

  5. Include three identifiers that appear on many pages, like a generic timeout or region key. These test chunk context rather than search, and they usually fail for a different reason than the rest.

  6. Include two identifiers that have been renamed or deprecated. You are checking whether the system returns the old one as though it were current.

Run the set before and after any retrieval change. Steps 4 and 5 catch the real damage, and they are the ones people skip.

Two notes on scoring. Count declines separately from wrong answers: a system that declines on eight of thirty is in much better shape than one that answers all thirty with six near-misses, and a single accuracy percentage ranks them the other way round. And run the set against the retrieval layer itself rather than through one surface, since that is what every consumer shares.

The short version

Embeddings are built to find things that mean the same. Identifiers are built to be distinguishable from things that mean the same. Those two facts are in direct conflict and no amount of reranking resolves it.

Run keyword search alongside embedding search, extract the identifier from the incoming request and query for it separately, give chunks enough context to be individually identifiable, and test with an identifier set rather than a natural-question set. Then make sure the system declines rather than returning the adjacent error code, because when the caller is an agent rather than a reader, a confident wrong answer becomes a wrong action.


Frequently Asked Questions

Frequently Asked Questions

FAQ

Why does my RAG system return the wrong error code?

Because vector search finds text that means something similar, and two error codes differing by one digit mean almost exactly the same thing to an embedding model. The retrieved page is genuinely about the right topic, which is why the answer looks correct. Adding keyword search alongside embedding search is what separates them, because keyword scoring treats rare terms as more informative rather than less.

Does hybrid search fix exact match in RAG?

It fixes most of it. Keyword scoring ranks rare terms highly, so an error code or config key leads straight to the passages containing it. Two things it does not fix are identifiers appearing on many pages, which is a chunking problem rather than a search problem, and identifiers buried inside long natural-language requests, which have to be extracted and queried separately.

Why does my coding agent get the wrong CLI flag from our knowledge base?

Flags that differ by one word sit close together in vector space, so retrieval returns a near neighbour and the agent has no way to detect the substitution. Unlike a human reader, an agent acts on the result immediately. Keyword search over the exact flag string is the fix, along with declining rather than returning the closest match.

Will a reranker solve my exact-match problem?

No. A reranker only reorders candidates that retrieval already returned, so if the passage containing your identifier was never retrieved, reranking cannot recover it. This is a recall failure rather than a ranking failure, and the two look similar in an evaluation summary while having completely different fixes.

How do I test whether my retrieval handles exact identifiers?

Build a separate test set of about thirty identifiers taken from your own content, ask for each inside a full sentence rather than as a bare string, and score exact match only. Include near-miss pairs differing by one character, and a few identifiers appearing on many pages. Natural-question evaluation sets under-represent exact identifiers and will not surface this.

Should retrieval guess when it cannot find an exact identifier?

No. Returning the adjacent error code is worse than returning nothing, because the caller has no way to detect the substitution and every signal in the response says the answer is correct. A system that declines turns a silent wrong answer into a visible coverage gap you can then fill.

TRUSTED BY 200+ INDUSTRY-LEADING ENTERPRISES WITH COMPLEX PRODUCTS
  • Silicon Labs
    Ask anything...
  • Logitech
    Ask anything...
  • n8n
    Ask anything...
  • monday.com
    Ask anything...

Turn technical documentation into customer-facing AI assistants