PageIndex is a document index for retrieval-augmented generation that drops vector databases entirely and instead has a language model reason its way through a table-of-contents tree. It is built for developers and machine-learning engineers who run RAG over long professional documents — reports, filings, manuals — and who have watched a similarity search return passages that look right and are not.
What it does
The premise behind the project is a complaint most RAG teams recognise: vector databases retrieve by similarity, not by relevance. Embeddings find text that resembles the question, which is not the same thing as text that answers it, and chunking a long document into fixed windows destroys exactly the structure a human reader uses to navigate it.
PageIndex takes the other road. It builds a hierarchical tree index of a document — essentially a table of contents — and then lets the model search that tree the way a domain expert reads a report: skim the headings, pick the section that matters, descend into it. There is no embedding step, no chunk size to tune, and no vector store to run alongside the application. Because retrieval happens as an explicit walk over named sections, the path the model took is visible, which makes the resulting answers traceable and explainable rather than a list of opaque nearest neighbours.
How it works
Indexing produces the tree first. PageIndex Flash, the fast index-generation method for text-based PDFs, is the default in the SDK's local mode; the project's own figures put indexing a thousand pages at roughly a dollar and under five minutes, which is cheap enough to re-index rather than patch. Retrieval is then a reasoning loop over that tree, with the model choosing which branch to open rather than being handed a top-k set of fragments.
Two pieces extend the idea past a single file. The PageIndex File System is a file-level tree indexing layer, so the same reasoning can range over an entire corpus instead of one document at a time. The PageIndex App applies the approach as a document analysis agent aimed at long professional documents. The channel's video points to sixty-two real questions where the answers came from reasoning rather than similarity — with, as advertised, no vector database anywhere in the stack.
Getting started
The SDK installs with pip install -U pageindex. From there you have two paths that share one client:
- Local mode — index, retrieve and chat entirely on your own machine, supplying your own LLM key.
- Cloud mode — point the same client at PageIndex Cloud with an API key instead.
That symmetry matters more than it sounds. You can prototype locally on a handful of documents, decide the retrieval quality is worth it, and move to the hosted service without rewriting the calling code. The project also publishes documentation and a developer site, and the repository is MIT-licensed Python, so vendoring or forking the indexing logic is on the table if you would rather not depend on anyone's service.
When to use it / when not
This fits when documents are long, structured and professional — the kind with a real table of contents and sections that mean something. It fits when an answer has to be defensible: finance, legal, compliance and technical documentation all care less about latency than about being able to point at the paragraph the answer came from. And it fits when the operational cost of a vector database, with its embeddings to refresh and its index to keep in sync, is out of proportion to the size of the corpus.
It fits less well on short, flat, unstructured text — chat logs, support tickets, scraped pages — where there is no hierarchy for the model to reason over. Retrieval here also means model calls at query time rather than a millisecond vector lookup, so workloads that need very high throughput or very low latency per query should measure before committing. Indexing is cheap, but it is not free, and it is aimed at text-based PDFs first.
Alternatives
The default alternative is the conventional stack: a chunker, an embedding model and a vector store, wired together by any of the mainstream RAG frameworks. That path is mature, fast at query time and well understood, and for large corpora of short documents it remains the sensible default. PageIndex is the deliberate opposite bet — spend compute on structure and reasoning up front, and accept slower, more expensive retrieval in exchange for relevance you can trace.
Take this one seriously if your retrieval quality has plateaued and the failures look like plausible-but-wrong passages rather than missing ones. Nearly 38,000 stars and steady development since April 2025 suggest the diagnosis resonates widely; whether the cure suits you depends on whether your documents have structure worth reasoning over. If they do, the experiment is one pip install and a few dollars of indexing away.