A federated RAG system that routes questions with arithmetic instead of an agent — and shows you the arithmetic while it does it.
What it does · Architecture · The router · Measurements · The maths · Run it · Verify it · Limits · Stack
Drop a spreadsheet, a scanned receipt and a policy document into one folder. PolyRAG reads each one, splits it, and decides — per chunk — whether it belongs in a SQL table, a vector index, or a knowledge graph. Ask a question and the same decision runs again to pick where to look. Every step reports itself, so you can watch the routing happen and see the numbers behind it.
Runs entirely on one machine: AMD GPU through Vulkan, no CUDA, no Docker, no cloud.
A question whose answer spans two documents that never mention each other. The right panel is the
whole point: the score each route got, the margin that decided it, and every step with its own
latency, rag.graph.pagerank at 1 ms next to chat.stream at 588 ms.
The question is English and the corpus is Portuguese. Routing, retrieval and the answer all work across the two because BGE-M3 embeds both into the same space; the sources are shown in the language they were written in, untranslated.
Agent-based RAG delegates routing to a model. That makes behaviour unpredictable — the same question can take a different path on Tuesday, and nobody can explain why. That is what stops these systems from being testable, debuggable or auditable.
PolyRAG's premise is the opposite: use deterministic maths wherever it can decide, and reserve the model for what only a model can do.
| Step | How it decides | Deterministic? |
|---|---|---|
| Is this file tabular? | Regex + comma/digit density | ✅ |
| Which store does this belong in? | Cosine similarity + margin | ✅ |
| Is this question already answered? | Cosine ≥ threshold | ✅ |
| Which graph chunks matter? | Personalized PageRank | ✅ |
| Genuine tie between two routes | Cosine against the headings each store holds | ✅ |
| Does this message ask more than one question? | Clause split + interrogative test | ✅ |
| Writing the answer | A model |
The claim is not a new algorithm. It is that routing does not need one — and that a system which can explain its own decisions is worth more than one that guesses well.
flowchart TB
subgraph ingest["Ingestion"]
drop["data_drop/"] --> kind{"file type"}
kind -->|image| ocr["GLM-OCR"]
kind -->|pdf, docx, pptx| doc["loaders<br/>keeps heading structure"]
kind -->|csv, xlsx| table["pandas"]
ocr --> chunk["chunking<br/>paragraphs, then meaning"]
doc --> chunk
end
subgraph router["Semantic router — one engine, both directions"]
s1["1 · heuristic<br/>tabularity score"] --> s2["2 · cosine + margin<br/>against config anchors"]
s2 -->|gray zone only| s3["3 · store evidence<br/>section headings each store holds"]
end
chunk --> router
table --> sql
question(["question"]) --> cache["CAG<br/>FAISS in RAM"]
cache -->|miss| router
router --> sql[("RAG 1 · SQLite<br/>Text-to-SQL")]
router --> vec[("RAG 2 · Qdrant<br/>HNSW")]
router --> kg[("RAG 3 · networkx<br/>Personalized PageRank")]
sql --> llm["Qwen3.5-4B<br/>writes the answer"]
vec --> llm
kg --> llm
llm --> ui["SvelteKit<br/>chat + live pipeline panel"]
router -.spans.-> otel["OpenTelemetry"]
otel -.WebSocket.-> ui
Three models served by llama.cpp over Vulkan, each on its own port, all speaking the OpenAI API:
Qwen3.5-4B (SQL, triple extraction on ingest, and the answer — nothing in the router),
GLM-OCR (images),
BGE-M3 (embeddings). Qdrant runs standalone alongside them. OCR sleeps when idle and hands its
VRAM back, so the resident set is about 5.6 GB.
The same component decides where a chunk is stored and where a question is searched — the input is the only difference. It runs three stages, cheapest first:
1 · Heuristic. Comma and digit density. A spreadsheet is tabular by definition and never reaches the model. Cost: microseconds.
2 · Cosine + margin. Each route is defined in config.yaml by a handful of example phrases,
embedded once at startup. The input is embedded and compared against all of them; the best match per
route becomes that route's score.
margin = score_top1 − score_top2
score_top1 ≥ tau_high AND margin ≥ delta → take it
score_top1 < tau_low → fall back to free text
otherwise → gray zone, go to stage 3
This is a 1-nearest-neighbour classifier with a rejection rule. Cost: 8.8 ms, measured.
3 · Store evidence. Only in the gray zone. The configured phrases say what a route is for;
they cannot know what a corpus turned out to contain. So each store is asked for the section
headings it actually holds (RAGBase.content_anchors), and the question is scored against those
too. Ingesting a document teaches the router about it with no config edit.
There is no model in any of the three stages. A model used to break the gray-zone tie, and it was measured out: on those questions, stage 2 alone was right 9 times out of 10 and the judge 8. It agreed with the geometry in 9 of the 10 — two model calls to repeat what arithmetic had already said — and the one time it disagreed, it was wrong. Store headings took routing on unseen questions from 82% to 95% instead.
The gray zone, resolved without a model. Nothing hand-written in config.yaml could know that
"Banco de Dados Órion" names a section in the graph — so this question used to go to the relational
store, whose only table is about sales.
"Qual foi a receita do mês de março e qual norma regula a retenção desses registros?"
The figure is in SQL and the rule is in the graph, and picking one store answers half the message. So each question is routed on its own, and each store is queried with the text of its own question — not with the whole message. That last part is not a detail: handing the full sentence to the Text-to-SQL step produced a query the guards rejected, and the relational half went unanswered.
Detecting the case is where this got interesting. The obvious rule is a threshold on the router's
second-best score, and it was measured and thrown away: compound questions score top2 between
0.365 and 0.511, single ones between 0.303 and 0.476. The ranges overlap almost completely, and
the plainly single question "Como uma reclamação de cliente deve ser tratada?" (0.476) outscores
five of the seven compound ones. The cause is mechanical — an embedding of two subjects lands near
their average, so a second subject dilutes both scores instead of lifting the second. It is the
same shape of failure as the cache threshold below, and no cutoff separates it.
What separates them is structure. A compound message contains two questions, so clauses.py splits
on clause boundaries (?, ;, e, and) and requires every surviving fragment to contain an
interrogative. Without that second test, two of the 80 evaluation questions split wrongly: "Uma
exportação sem anonimização e reportada para quem?", where the e is the verb é written
unaccented, and "Qual a diferença entre um atraso comunicado e um atraso descoberto?", where
it joins two noun phrases. In both, one half asks nothing at all.
Measured: 9 of 9 compound messages detected, 0 false positives in 80 single questions. The
stores are then queried with asyncio.gather and their results concatenated under [SQL],
[TEXT] and [GRAPH] labels — not fused by RRF, because RRF merges two rankings of the same
items, and here the stores hold disjoint content and one of them returns a SQL result table rather
than a ranked list.
One message asking a figure and a rule. The route line reads relational + graph, and the answer
carries both halves: the SUM from the sales table and the approval threshold from the compliance
graph.
Retrieved passages are reordered before the model reads them, following Liu et al. 2023: a model
attends to the start and the end of a context and least to its middle. Ranks 1, 3, 5 go out in order
and 2, 4 come back reversed, so [1,2,3,4,5] becomes [1,3,5,4,2] — the best passage opens the
context and the second best closes it. Pure list slicing; the ranking was already done. SQL rows are
left alone, because their order is the ORDER BY that selected them.
| Holds | Retrieved by | Why not the others | |
|---|---|---|---|
| RAG 1 · SQLite | Rows and columns | Generated SQL | Aggregation must be exact. Correct SQL gives the correct number; no model does the arithmetic. |
| RAG 2 · Qdrant | Free text | Cosine over HNSW | No schema to impose, no relations to model — only meaning. |
| RAG 3 · networkx | Entities and relations | Personalized PageRank | Multi-hop. "If A fails, what breaks?" needs topology, not similarity. |
Why the first row of that table matters. The source is not a passage — it is the query that was
run and the row it returned. The figure came out of a SUM, so it is either right or the SQL is
wrong, and the SQL is on screen either way.
RAG 3 is HippoRAG 2 reimplemented from the paper: the model extracts subject–relation–object triples, networkx holds the graph, and the walk teleports back to the question's entities so the ranking means "relevant to this question".
The official hipporag package needs torch and vLLM, which are Linux/CUDA only. Reimplementing
it was a constraint, and it turned out to be the more useful outcome — the maths is visible and
testable rather than hidden behind an import.
Routing is per chunk, not per file, which is why four of the six files show up in two stores. The
row worth looking at is sobre_a_meridiano.md: one of its five paragraphs sits close enough to the
routing boundary that rewording a route anchor moved it into the graph, and three questions that
depended on it lost their answer. The split is also where the ingestion fallback shows: a chunk the
graph could extract no relation from is kept as free text rather than dropped.
uv run python scripts/benchmark.py on this machine (RX 9060 XT, 16 GB), 12 measured runs per
question after 2 discarded warm-ups. A single number cannot describe a latency — the first call
after a model loads pays for a cold cache, and the tail is what a user notices — so every row is
p50 and p95 rather than an average.
| p50 | p95 | |
|---|---|---|
| Routing decision (deterministic stage) | 8.8 ms | 12.0 ms |
| Routing decision on the wire, streaming | 39 ms | 54 ms |
| Time to first token | 114 ms | 185 ms |
| Full answer, vectorial | 252 ms | 255 ms |
| Full answer, graph | 724 ms | 744 ms |
| Full answer, relational | 737 ms | 745 ms |
| Cache hit, same question | 9.3 ms | 12.9 ms |
| Share of a request spent inside model calls | 99 % |
The whole graph route costs 19 ms, and the model then spends 700 writing the answer. Retrieval
is not where the time goes — the per-stage span breakdown that backs that claim, and retrieval/answer
accuracy on three question sets (golden, held-out, and compound), are in
docs/MEASUREMENTS.md. Headline numbers: 90–95% routing accuracy,
86% recall@5, 89–92% answer facts, measured on a held-out set never touched while tuning.
Nothing in the router, the cache or the graph walk is hidden behind a library call — cosine,
normalisation, the margin rule, and Personalized PageRank, each with the formula, a worked example,
and where it's tested. Written out in full in docs/MATHS.md.
Needs Python 3.12, Node 22, and roughly 6 GB of VRAM.
One command does the whole install — it installs uv if missing, syncs the locked dependencies, downloads the binaries and models, checks that Vulkan sees the GPU, and runs the tests. Every step checks whether its work is already done, so re-running is safe:
.\setup.ps1-SkipModels leaves out the 4.9 GB download, -SkipFrontend skips npm, and -Start launches the
servers and the API when it finishes. Or do the same steps by hand:
# 1 · dependencies
uv sync
npm --prefix frontend ci
# 2 · binaries and models — about 4.9 GB, skips whatever is already there
uv run python scripts/fetch_runtimes.py
# 3 · everything the backend needs: three llama-servers and Qdrant
uv run python scripts/start_servers.py
# 4 · the API
uv run uvicorn --app-dir backend app.main:app --port 8000
# 5 · the interface
npm --prefix frontend run devThe servers run windowless and log to data/run/<name>.log. Stop them with
scripts/start_servers.py --stop, or start and stop them individually from the Servers tab in
the interface.
Tip
On a 16 GB card, stop a model you're not using from the Servers tab instead of leaving all
three resident — ocr hands its VRAM back within a minute of going idle on its own, but llm and
embeddings don't unless you stop them.
Then open http://localhost:5173, or http://localhost:8000/docs for the API.
Drop a file into data_drop/ and it is ingested within seconds — the hot folder is watched. Or use
the Corpus tab, which is the same code path.
For something to ask it about, demo/ holds a six-file corpus about one fictional
company, plus the questions that exercise each route and the two that do not work:
uv run python scripts/load_demo.py --resetNote
Everything is driven by config.yaml: ports, model paths, thresholds, prompts, and the example
phrases that define each route. Adding a route is a config change, not a code change.
uv run pytest tests/ -q # 116 tests, no servers required
uv run ruff check backend/ scripts/ tests/
npm --prefix frontend run check # 0 errorsThe unit tests cover cosine and margin, the routing decision, the read-only SQL guard, entity normalisation and length limits, the tabularity heuristic, the cache, the clause splitter, the context ordering and every chunking rule. They need no GPU and no network — which is the point of keeping that code pure.
The chunking tests exist because chunking is where most of this project's real bugs came from, and they are verified by mutation rather than assumed: removing a guard has to fail exactly the test that defends it. A regression test that cannot fail is decoration.
The scripts under tests/ named *_integration.py are run by hand against a live stack and print
their results; pytest collects nothing from them.
uv run python scripts/evaluate.py --reset # re-ingests, then measures the golden set
uv run python scripts/benchmark.py # latency p50/p95, same requirements
uv run python -u tests/test_reingest_integration.py # a corrected file replaces its old chunksFour views: Chat, Corpus, Telemetry and Servers. Chat pairs the conversation with a live panel that shows the per-route cosine scores, the margin the decision turned on, which stage made the call, and the span waterfall as it happens — the routing decision is on screen at about 39 ms, before retrieval has even finished.
The same question, asked twice. The second answer skips routing, retrieval and generation entirely: three spans and 10 ms, against roughly a second for the first.
Every trace the backend produced, the span waterfall for the selected one, and the attributes that
decided it. graph.seed_best is the cosine that chose the walk's starting node; cache.hit says
why this request did the work instead of skipping it. This is the claim about auditability, rendered
rather than asserted.
Each model is its own process, running windowless. Stopping one hands its VRAM straight back,
which matters on a 16 GB card when something else needs the GPU. ocr shows hollow because it put
itself to sleep after a minute idle and gave back 2.08 GB on its own.
The screenshots above are generated, not taken by hand — uv run python scripts/screenshots.py
drives Chrome over the DevTools protocol and rewrites them all. A screenshot nobody can regenerate
is one that quietly starts lying after the next UI change.
To reproduce the demo yourself:
uv run python scripts/start_servers.py
uv run python scripts/load_demo.py --reset
uv run uvicorn --app-dir backend app.main:app --port 8000
npm --prefix frontend run devThen ask, in order: a figure (What was the total revenue of the Sudeste region?), something
narrative (In what year was Meridiano Logistica founded?), a chain that crosses two files (Which approval policy do the orders processed by Sistema Atlas follow?), one message that asks two things
at once (What was the Sudeste revenue and who approves a purchase of eighty thousand reais? — the
panel shows relational + graph), and finally any of them a second time to watch the cache answer in
9 ms. The corpus is in Portuguese and the questions are in
English on purpose: routing and retrieval work across languages because BGE-M3 embeds both into one
space, and the sources are shown untranslated. demo/README.md has the full
script, the expected figures, and the two questions that fail.
Real ones, found by running the system — not theoretical edge cases. Ten open items and four fixed
since the last release, each with the failing case that found it, in
docs/LIMITATIONS.md. The two that matter most right now: the vectorial
route is still the weakest of the three, and a question contained inside another one can still slip
past the cache.
FastAPI 0.141 · Qdrant 1.19 · networkx 3.6 · faiss-cpu 1.15 · OpenTelemetry 1.44 ·
SvelteKit 2.63 · Svelte 5 · llama.cpp (Vulkan) · uv
No PyTorch, no vLLM, no Ollama, no Docker, no WSL.







