Fairs Place
SAM.gov activeEnterprise & government softwareWashington D.C. · Maryland · VirginiaSBA small business
Insights

Most GenAI Failures Are Context Failures, Not Model Failures

Fairs Place ·
A friendly doodle robot confidently answering while messy, crumpled papers feed into it

When a production GenAI system gives a wrong answer, the first instinct is to blame the model. Swap in a bigger one. Fine-tune. Add more parameters. In our experience building AI systems for federal programs and regulated enterprises, the model is rarely the problem. The problem is what the model was given to work with.

We call these context failures: the retrieval missed, the context was stale, the sources conflicted, or nobody can trace the answer back to the data that produced it. The model did exactly what it was supposed to do with bad input. Fixing the model when the context is broken is like tuning an engine when the fuel is contaminated.

A friendly doodle mechanic pouring muddy, contaminated fuel into a shiny, well-built engine

This is the problem our Context Engine is being built to solve. But whether or not you ever use our product, the failure modes below are worth knowing — because they’re what actually breaks GenAI systems in production.

Diagram of the five context failure modes — stale, miss, bloat, conflict, no trace — as hand-drawn doodle icons

The five context failure modes

1. Stale context

The knowledge base says the policy changed last quarter. The retrieval layer is still serving last year’s version. The model answers confidently from outdated information, and nobody notices until an operator spots it weeks later.

Why it happens: Ingestion pipelines run on schedules nobody monitors. Documents get updated in the source system but the update never propagates to the vector store. There is no freshness metadata on retrieved chunks, so the model can’t tell a 2024 memo from a 2026 one.

How to measure it: Track the age distribution of retrieved documents. If your top-k retrieval regularly surfaces documents older than the source system’s current version, you have a staleness problem.

How to fix it: Version every ingested document. Attach last_updated and source_version metadata at ingestion time, surface it in retrieval, and alert when retrieved content is older than a threshold you define with the program office.

2. Retrieval misses

The answer exists in the corpus. The system doesn’t find it. The user gets a confident hallucination instead of the document sitting three hops away in the index.

Why it happens: Chunking strategies that made sense for one document type destroy another. Embedding models trained on web text underperform on federal acronyms, program names, and domain jargon. And nobody evaluates retrieval quality against a labeled set of real questions — so the miss rate is unknown, not zero.

How to measure it: Build an evaluation set of real questions with known-good source documents. Measure recall@k and mean reciprocal rank on every index change. If you can’t quote your retrieval recall, you don’t have a retrieval system — you have a hope.

How to fix it: Evaluate retrieval independently from generation. Tune chunking per document type. Consider hybrid retrieval (dense + keyword) for acronym-heavy corpora, which describes nearly every federal document set we’ve seen.

3. Context bloat

The retrieval works — too well. Forty chunks get stuffed into the prompt, the model attends to the wrong ones, and the answer degrades. More context is not better context; past a point, it’s noise the model has to wade through.

Why it happens: Teams set k=20 “to be safe” and never measure whether the extra chunks help. Reranking is skipped because it adds latency. Nobody tests the answer quality curve as context grows.

How to measure it: Plot answer quality (against your evaluation set) as a function of k. You’ll typically find a knee: quality rises, plateaus, then falls. Operate at the knee, not at the maximum.

How to fix it: Rerank before stuffing. Compress or summarize low-relevance chunks. Set per-query context budgets and enforce them.

4. Conflicting sources

Two retrieved documents disagree. One is the current policy; the other is a superseded draft. The model picks one — or worse, blends them into an answer that matches no actual source.

Why it happens: Corpora accumulate drafts, superseded versions, and contradictory guidance. Without source hierarchy metadata, the model treats a draft memo and a signed policy as equals.

How to measure it: Sample retrieved sets and count how often top-k results contain mutually contradictory statements on the same question. In our experience, this is disturbingly common in program document sets.

How to fix it: Encode source authority in metadata: document type, signature status, effective dates, supersession links. Prefer authoritative sources in ranking, and when conflict is detected, surface it to the user instead of silently resolving it.

5. No traceability

The system answered. Nobody can say which documents it used, which version of each, or what the retrieval query was. When the program office asks “show me the evidence behind this answer,” there’s nothing to show.

Why it happens: Traceability is treated as a nice-to-have and deferred. Logging stores the final prompt but not the retrieval decisions that built it.

How to measure it: Ask your system to reproduce, for any past answer, the exact document set and versions it retrieved. If it can’t, you have no traceability.

How to fix it: Log the per-decision trace: query, retrieval parameters, returned document IDs and versions, ranking scores, and the final context assembly. This is an audit requirement in federal environments, not a feature — and it’s the core of what we’re building with the Context Engine.

What good looks like

Diagram of a healthy context pipeline — versioned docs, measured retrieval, rerank, right-sized context, per-decision trace

A production GenAI system with its context under control has:

  • Measured retrieval quality — recall@k and MRR tracked in CI, with regressions blocking deployment
  • Freshness guarantees — versioned documents, staleness alerts, defined SLAs with the data owners
  • Right-sized context — k chosen from measured quality curves, reranking in the pipeline
  • Source authority — metadata-driven ranking that prefers signed, current sources
  • Per-decision traces — every answer reproducible from its retrieval log

None of this is about the model. A mid-tier model with excellent context engineering will outperform a frontier model fed stale, bloated, untraceable context — every time, and at a fraction of the cost.

The infrastructure lesson

Machine learning learned this lesson a decade ago with feature stores: the model is only as good as the data pipeline feeding it. Generative AI is learning it now, in production, the expensive way. Context engineering — schemas, measured retrieval, versioned stores, per-decision traces — is the feature-store discipline applied to a new kind of system.

That’s what we build at Fairs Place: AI systems for federal programs and enterprise teams where the context layer is engineered, measured, and auditable — not hoped for. If your GenAI pilot works in the demo and fails in the program office, the context layer is the first place to look. Talk to us — you’ll reach the engineer, not an account manager.