Skip to content
Back to blog
AI appsWhen a company knowledge base needs retrieval

AI Knowledge Base Without RAG: Which Route Fits?

A small, stable corpus may be cheaper to test in the model's context before building retrieval. Compare full-context testing, existing search and permission-aware retrieval using your own questions, data and token budget.

The least expensive route is a measured test. The corpus size, answer quality, permissions and update rhythm decide whether full context or retrieval fits.

Author

Syntalith

Published Updated 8 min read

A small and stable document set may be worth testing in full context before a company builds retrieval. Anthropic's Contextual Retrieval article uses about 200,000 tokens, roughly 500 pages of text, as a starting heuristic. The number is not a promise: language, tables, OCR, user permissions and update frequency determine the result.

Four routes for one knowledge question

RouteFits whenMain cost
Existing searchDocuments are well named, users share access and questions are easy to locateTeam time and document maintenance
Full contextA small corpus fits the context, changes rarely and users share accessInput tokens on every question
RetrievalThe corpus outgrows context or answers need selected passagesIndexing, permissions, updates and retrieval quality
Rules plus searchThe question has stable fields and the answer must be exactProcess design and source maintenance

At Syntalith, a simple automation starts from €3,500 net and a dedicated knowledge application starts from €6,000 net. Those are offer starting points; first test whether an existing search or full-context route solves the actual questions.

Measure the corpus

Take a representative sample of documents and count tokens with the provider used for the test. Record:

  • language and format,
  • tables, scans and repeated page elements,
  • number of documents and versions,
  • users and their document access,
  • monthly question volume,
  • update rhythm and document owner.

The same 200,000-token corpus can have very different costs at ten questions per week and at thousands of questions per day. Full context sends the repeated corpus with each request unless the provider's caching feature changes that calculation.

Build a fair question set

Use real questions from the people who need the information. Include:

  1. direct lookups with an answer in one document,
  2. questions that require two documents,
  3. questions using employee language rather than procedure terms,
  4. outdated or conflicting versions,
  5. questions with no answer in the collection,
  6. a question from a user who lacks permission for the relevant file.

Score source accuracy, completeness, refusal when evidence is missing, response time, token cost and reviewer effort. A fluent answer without a source should fail the same check as a slow answer with a source.

Why context length is only a capacity

The Context Rot research reports uneven quality as inputs become longer. NoLiMa, published at ICML 2025, studies cases where the question and relevant information share little vocabulary. These sources support a test plan; they do not predict the result for a company's documents.

The practical rule is simple: measure a question set at the corpus size you intend to send. A context window says what can fit. It does not guarantee that the model will find and use every relevant passage equally well.

Permission is a separate decision

If every user sees the same documents, full context can be a reasonable pilot route. If access differs by role, retrieval must apply the user's permission before the model receives a passage. A small corpus with restricted HR or finance files still needs access-aware retrieval.

Use an allowlist of document identifiers and record:

  • the user or role requesting the answer,
  • the documents considered,
  • the passages supplied,
  • the answer and source links,
  • the version and date of each source.

The permission-aware retrieval guide explains the data-access design in more detail.

A decision ladder

  1. Clean names, owners and current versions.
  2. Test the existing search on real questions.
  3. Test the full corpus in context when it fits and access is shared.
  4. Compare retrieval on the same question set when answers, cost or permissions require it.
  5. Add hybrid retrieval, reranking or version controls only after the test identifies the problem they solve.

This order protects the budget and gives the team a baseline for every later change.

When a knowledge application earns its place

Choose a build when the corpus changes often, documents are spread across systems, users have different permissions, the full-context bill grows with every question or the question set shows missing evidence. A build should include document ownership, update handling, source links, access checks, evaluation and a person responsible for the result.

Keep the existing search when the collection is small, stable and easy to maintain. A better file name or one agreed location can solve more than a model in that situation.

Questions for the buyer and supplier

  1. How many tokens and documents enter one request?
  2. Which user permissions apply before retrieval?
  3. How are old versions removed or marked?
  4. Can every answer open its source passage?
  5. What happens when the collection has no answer?
  6. Which costs recur for storage, retrieval, models and maintenance?
  7. Who owns document updates and evaluation?

To compare a route on a real corpus, book a free process scan. Bring a small document sample and questions from the people who use it.

Frequently asked questions

When can a company skip RAG?
A small, stable corpus can be tested in full context when every user may see the same documents and the answers meet the agreed quality test. Measure the tokens and the result on real questions before committing to a build.
How many tokens fit in the full-context test?
Anthropic's Contextual Retrieval article uses about 200,000 tokens, roughly 500 pages of text, as a practical starting heuristic. Language, tables, scans and repeated headers change the count, so measure a representative sample.
Does long context remove the need for retrieval?
No general rule follows from a context-window size. Long-context quality can vary by task and position. Permission differences, frequent updates and poor answers make retrieval relevant at smaller corpus sizes.
What should a knowledge-base pilot measure?
Use real employee questions, including missing-answer cases. Measure source accuracy, answer quality, permission handling, response time, token cost and review effort. Compare the full-context route with the simplest search or retrieval route.

Free process scan

Start with a free process scan.

  • A 30-minute call with the engineer who would lead the work.
  • A review of the processes that cost you the most time and money.
  • A written summary of what to automate first and the likely cost range.

The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.

€0

30 minutes · written takeaway within 2 business days

Book a free process scan (30 min)

Times are shown in your own time zone. We work with clients across time zones.

Describe the process in the form