Skip to main content

RAG — Prior Art and the Package Ledger

This is a living document

It is written to be re-read and revised as the work lands, not archived. Every claim here is either measured, cited, or explicitly marked as unverified. When a stage ships and changes one of these answers, change it here in the same PR.

Last substantive review: 2026-08-13, during issue #403 Stage 4 design.

Three pages cover retrieval design; this is the fourth, and it answers the question the others assume: what already exists, and why did we build anything at all?

PageQuestion
RAG Chat (Internals)How does the pipeline work?
RAG Evaluation MethodologyHow is quality measured?
Design DecisionsWhy this shape, and what nearly fooled us?
This pageWhat is the prior art, and what do we import versus write?

Part 1 — How other systems summarize a large corpus

The canonical taxonomy

There are five patterns in general use. All of them are well-named, and using the wrong name for one of them is how teams end up building a sixth by accident.

PatternWhat it doesWhere it breaks
StuffPut every document in one promptThe context window. Hard stop.
Map-reduceSummarize each chunk (map), then summarize the summaries (reduce)Loses cross-chunk narrative; reduce prompt can itself overflow
RefineCarry a running summary forward, chunk by chunkStrictly sequential — no parallelism, latency scales with corpus
DocumentSummaryIndexPrecompute a per-document summary; use it as the retrieval handleSummaries go stale when the document changes
Hierarchical / RAPTORRecursively cluster and summarize into a tree; retrieve at any levelBuild cost; clustering quality is corpus-dependent

Map-reduce and tree_summarize are the same pattern. The LlamaIndex documentation states the equivalence outright — "In LlamaIndex this is referred to as tree_summarize, in LangChain this is referred to as map-reduce" — and they differ only in recursion depth: LangChain's classic map-reduce fans in once, tree_summarize repeats the fan-in until one root remains. If you find yourself designing a "novel" summarization strategy, check first that it is not one of these five wearing different vocabulary.

What Open WebUI actually does

Open WebUI is the most common open-source point of comparison, so it is worth being precise: it does not solve corpus-scale summarization, and its maintainers say so.

It offers two modes, neither of which is summarization:

  1. Default RAG. Chunk the document, embed, retrieve top-k, send those chunks. Users consistently report thin, partial summaries, because top-k retrieval answers "which passages match this query" — a fundamentally different question from "what does this document say."
  2. Full Context Mode. Send the entire document. Works until the context window, then truncates. Practically it requires an 80K–256K model plus the VRAM to serve it.

There is an open proposal to pick between the two automatically by token threshold, and an older issue observing that a real summarization pipeline "requires significantly more implementation effort." It has not been built. A widely-reported symptom is that the same model which summarizes a PDF well in ChatGPT returns a stub in Open WebUI — the model is not the variable, the retrieval strategy is.

This matters for our roadmap in one specific way: on the "summarize everything" query class there is no open-source prior art to import. We are not behind it.

Closed-source systems

Less can be stated with confidence here, so less is stated. What is publicly documented: the frontier assistants lean on very long contexts plus agentic chunked reading rather than a fixed map-reduce chain, and Anthropic has published contextual retrieval — prepending chunk-specific context to each chunk before embedding, which is cheap and improves retrieval independently of the summarization strategy. That last one is adjacent to our Stage 5 work and is not yet evaluated here; treat it as an open idea, not a decision.

Part 2 — The package ledger

The standing rule in this repository is do not hand-roll what a package already does well. The rule that qualifies it is do not import a framework to obtain a control-flow pattern. Most disagreements about "should we use library X" dissolve once you ask which of those two applies.

The test we apply, in order:

  1. Does the package do genuinely hard work? Format parsing, tokenization, ranking metrics, embeddings, OCR — yes. Fan out N calls and combine the results — no.
  2. Does it duplicate a layer we already run? Our embeddings execute inside OpenSearch via ML Commons, and fusion is OpenSearch-native. A framework that brings its own vector store and retriever adds a second implementation of a layer we already have.
  3. What does it cost the published image? We publish containers. Dependency weight and licence are release concerns, not preferences.
  4. Is the licence compatible with publishing? Not just with using. See trec_eval below.

Adopted — packages doing the hard work

JobPackage / serviceNote
Lexical ranking (BM25)OpenSearchNative; not reimplemented
Dense vectorsOpenSearch ML Commons + all-MiniLM-L6-v2 (384-dim)Embeddings run in the cluster, not in Python
Hybrid fusionOpenSearch score-ranker-processor (RRF, rank_constant 30)Native Reciprocal Rank Fusion
Sentence splittingnltk punktOne splitter shared by the transcript chunker, digests, and the document chunker — see the note below
Retrieval metricspytrec_eval_terrier (NIST trec_eval C code)nDCG@10 / recall@k / MRR
Document parsingDocling (tiered) + pypdfium2See Documents
Legacy OLE2 parsingApache Tika.doc / .ppt / .xls — Docling handles OOXML only
LLM servingvLLMGemma 4 E4B
Reranking seamreranker.get_reranker()Deliberately a seam; the model is Stage 5's bake-off
trec_eval is an eval-only dependency for a licence reason, not a size one

Its C sources carry a "research, non-commercial purposes" header, and we publish images. It lives in backend/requirements-eval.txt; every module that uses it imports lazily and every test importorskips it with that reason recorded. Never move it into requirements.txt.

Rejected — and the specific reason

These are recorded so they can be overruled with evidence rather than re-argued from scratch.

PackageRejected because
LangChain summarization chainsThree reasons, in increasing order of how much they cost us. (a) It brings a framework to obtain a control-flow pattern — fan out N calls, combine the results — which is test 1. (b) The classic chains (load_summarize_chain with stuff / map_reduce / refine) are deprecated in favour of LangGraph, so adopting them means adopting a migration we did not need. (c) The decisive one: every chain assumes the map step is an LLM call. Ours is not — the per-file map output is the extractive digest, computed deterministically at ingest, which is exactly what makes a summary over 1,000 recordings cost zero map-time work. A chain cannot express "the map already happened", so adopting one would have meant paying for the thing the design exists to avoid. We reuse the shape (bounded thread pool, pre-filled error slots, index-keyed results) and none of the code.
LlamaIndex response synthesizersDuplicates a layer we already run (test 2): its value is ingestion, vector stores, and retrievers, and ours live in OpenSearch. We adopt its vocabularytree_summarize, DocumentSummaryIndex — and name our components accordingly.
semantic-routerRoutes by embedding similarity, i.e. a client-side encoder call per turn on the critical path; and its base dependencies pull litellm + openai + aiohttp + tiktoken. Our routing requirement is explicitly rules-first with zero LLM calls before retrieval on turn 1.
RAPTOR (for now)Genuinely additive and not rejected on principle — see below. Deferred until it can be measured against a Stage 4 baseline that does not yet exist.

Written ourselves — and why that is the right call

Three components are ours, and in each case the reason is the same: they are the transcript-aware part. A generic package cannot know about speaker turns, and speaker turns are the product.

  • Speaker-turn chunking. Every general-purpose chunker splits on characters, tokens, or sentences. Ours splits on who is talking, which is why a retrieved passage is a coherent exchange rather than a window that begins mid-sentence in one speaker and ends mid-sentence in another. This is the differentiator, not an implementation detail.
  • Timestamp-anchored citations. A citation resolves to a real position in the media, so a claim in an answer is clickable back to the audio.
  • Rename propagation. Renaming a speaker updates the indexed text, because the indexed text contains the name.
Sharing beats reimplementing, even internally

When the document chunker needed to split over-long blocks, it imported chunking_service.split_into_sentences rather than writing a third splitter — the transcript chunker and the digest builder already shared it. Three splitters over the same words means three sets of boundaries and three sets of off-by-one bugs. The same test applies inside the codebase as outside it.

Part 3 — What we implement, in industry terms

So that the mapping is unambiguous:

Our componentIndustry name
Hybrid chunk retrievalBM25 + dense kNN fused by Reciprocal Rank Fusion
The digest plane (Stage 3)DocumentSummaryIndex / parent-document (small-to-big) retrieval
The query router (Stage 4)Query routing — "route, don't fuse"
Two-level summarization (Stage 4)Map-reduce = tree_summarize — see the map step is a read
Rerank stage (Stage 5)Two-stage retrieve-then-rerank

We built the standard patterns. For a while we simply were not calling them by their names — the digest plane was DocumentSummaryIndex before anyone wrote that down. Naming them is not cosmetic: it is how a reader knows which known failure modes apply.

The map step is a read, not a call

Our tree_summarize differs from every published implementation in one respect, and it is the respect that makes corpus scale tractable:

transcript chunks ──(TextRank, at ingest, NO LLM)──▶ file digest      ← the MAP
file digests ──(code, or N small bounded calls)─▶ collection view ← the REDUCE

Level 1 already ran when the file was ingested. A summary over 1,000 recordings therefore costs zero map-time work — the map is a database read. That is the whole reason the digest is deterministic and extractive rather than LLM-written: a map step that needed a model would leave the corpus-scale case exactly as impossible as it was before.

Two levels only, deliberately. Not recursive — see RAPTOR below.

Ranking picks the best passages; mapping covers every document

This distinction is load-bearing, it is not obvious, and it was learned by shipping the wrong one. Anyone reading mapreduce.scope_digest_hits will notice that a ranked digest retrieval already exists (retrieve_digests) and wonder why the map does not simply use it. It did, once.

Asked for 50 digest sections over a 25-file scope, the ranked leg returned 50 sections drawn from 8 files — sections cluster by relevance, which is precisely what a ranker is for. The composed overview was therefore headed recordings: 8, and the model faithfully answered "The recordings cover 8 vendor review board sessions" over a scope of twenty-five. Nothing looked broken. The number was simply wrong, and confidently so.

The two operations are not interchangeable:

question it answerscorrect behaviour
Ranking (retrieve_digests)"which passages best match this query?"return the top K by relevance, wherever they cluster
Mapping (scope_digest_hits)"what is in each document?"return one summary per document, ignoring relevance entirely

So for a bounded scope the map reads file_facts for every file in it and ignores ranking. The ranked leg survives only for the unbounded "all accessible" case, where mapping over everything is not possible — and there the header says how much it covered rather than reporting the covered count as the total.

Do not "simplify" the map back to the ranked leg

It will look like a redundant second retrieval path and it is not. Increasing size does not fix it: ranking gives you no coverage guarantee at any K. tests/unit/test_chat_mapreduce.py::test_the_scope_map_covers_every_file_not_the_best_ranked_ones fails if the map is replaced by a ranked retrieval.

RAPTOR — the one open idea worth measuring

RAPTOR builds a recursive summary tree by clustering nodes semantically before summarizing, and indexes every level so a query can retrieve at whatever abstraction it needs. It is the natural generalization of our digest plane.

The reason it is interesting here specifically: RAPTOR clusters because it has no better grouping signal. We have one — real speaker turns, meeting boundaries, and timestamps. Applying the product's actual moat to the summarization tier, rather than to retrieval alone, is a genuinely novel direction and a candidate for the whitepaper.

The reason it is not scheduled: the RAPTOR paper reports no significant gain from its clustering variant over the simpler sequential approach. It has to beat a Stage 4 baseline that does not exist yet. Measure first.

Sources