Skip to content

Repository files navigation

PolyRAG

A federated RAG system that routes questions with arithmetic instead of an agent — and shows you the arithmetic while it does it.

CI Release Python SvelteKit

Runs on No CUDA No Docker Local

What it does · Architecture · The router · Measurements · The maths · Run it · Verify it · Limits · Stack


Drop a spreadsheet, a scanned receipt and a policy document into one folder. PolyRAG reads each one, splits it, and decides — per chunk — whether it belongs in a SQL table, a vector index, or a knowledge graph. Ask a question and the same decision runs again to pick where to look. Every step reports itself, so you can watch the routing happen and see the numbers behind it.

Runs entirely on one machine: AMD GPU through Vulkan, no CUDA, no Docker, no cloud.

The chat, with the route scoreboard and the span waterfall beside it

A question whose answer spans two documents that never mention each other. The right panel is the whole point: the score each route got, the margin that decided it, and every step with its own latency, rag.graph.pagerank at 1 ms next to chat.stream at 588 ms.

The question is English and the corpus is Portuguese. Routing, retrieval and the answer all work across the two because BGE-M3 embeds both into the same space; the sources are shown in the language they were written in, untranslated.


The problem it takes seriously

Agent-based RAG delegates routing to a model. That makes behaviour unpredictable — the same question can take a different path on Tuesday, and nobody can explain why. That is what stops these systems from being testable, debuggable or auditable.

PolyRAG's premise is the opposite: use deterministic maths wherever it can decide, and reserve the model for what only a model can do.

Step How it decides Deterministic?
Is this file tabular? Regex + comma/digit density ✅
Which store does this belong in? Cosine similarity + margin ✅
Is this question already answered? Cosine ≥ threshold ✅
Which graph chunks matter? Personalized PageRank ✅
Genuine tie between two routes Cosine against the headings each store holds ✅
Does this message ask more than one question? Clause split + interrogative test ✅
Writing the answer A model ⚠️ temperature 0

The claim is not a new algorithm. It is that routing does not need one — and that a system which can explain its own decisions is worth more than one that guesses well.


Architecture

flowchart TB
    subgraph ingest["Ingestion"]
        drop["data_drop/"] --> kind{"file type"}
        kind -->|image| ocr["GLM-OCR"]
        kind -->|pdf, docx, pptx| doc["loaders<br/>keeps heading structure"]
        kind -->|csv, xlsx| table["pandas"]
        ocr --> chunk["chunking<br/>paragraphs, then meaning"]
        doc --> chunk
    end

    subgraph router["Semantic router — one engine, both directions"]
        s1["1 · heuristic<br/>tabularity score"] --> s2["2 · cosine + margin<br/>against config anchors"]
        s2 -->|gray zone only| s3["3 · store evidence<br/>section headings each store holds"]
    end

    chunk --> router
    table --> sql
    question(["question"]) --> cache["CAG<br/>FAISS in RAM"]
    cache -->|miss| router

    router --> sql[("RAG 1 · SQLite<br/>Text-to-SQL")]
    router --> vec[("RAG 2 · Qdrant<br/>HNSW")]
    router --> kg[("RAG 3 · networkx<br/>Personalized PageRank")]

    sql --> llm["Qwen3.5-4B<br/>writes the answer"]
    vec --> llm
    kg --> llm

    llm --> ui["SvelteKit<br/>chat + live pipeline panel"]
    router -.spans.-> otel["OpenTelemetry"]
    otel -.WebSocket.-> ui
Loading

Three models served by llama.cpp over Vulkan, each on its own port, all speaking the OpenAI API: Qwen3.5-4B (SQL, triple extraction on ingest, and the answer — nothing in the router), GLM-OCR (images), BGE-M3 (embeddings). Qdrant runs standalone alongside them. OCR sleeps when idle and hands its VRAM back, so the resident set is about 5.6 GB.


The router

The same component decides where a chunk is stored and where a question is searched — the input is the only difference. It runs three stages, cheapest first:

1 · Heuristic. Comma and digit density. A spreadsheet is tabular by definition and never reaches the model. Cost: microseconds.

2 · Cosine + margin. Each route is defined in config.yaml by a handful of example phrases, embedded once at startup. The input is embedded and compared against all of them; the best match per route becomes that route's score.

margin = score_top1 − score_top2

score_top1 ≥ tau_high  AND  margin ≥ delta   →  take it
score_top1 < tau_low                          →  fall back to free text
otherwise                                     →  gray zone, go to stage 3

This is a 1-nearest-neighbour classifier with a rejection rule. Cost: 8.8 ms, measured.

3 · Store evidence. Only in the gray zone. The configured phrases say what a route is for; they cannot know what a corpus turned out to contain. So each store is asked for the section headings it actually holds (RAGBase.content_anchors), and the question is scored against those too. Ingesting a document teaches the router about it with no config edit.

There is no model in any of the three stages. A model used to break the gray-zone tie, and it was measured out: on those questions, stage 2 alone was right 9 times out of 10 and the judge 8. It agreed with the geometry in 9 of the 10 — two model calls to repeat what arithmetic had already said — and the one time it disagreed, it was wrong. Store headings took routing on unseen questions from 82% to 95% instead.

A gray-zone question resolved by store evidence

The gray zone, resolved without a model. Nothing hand-written in config.yaml could know that "Banco de Dados Órion" names a section in the graph — so this question used to go to the relational store, whose only table is about sales.

When the message asks more than one question

"Qual foi a receita do mês de março e qual norma regula a retenção desses registros?"

The figure is in SQL and the rule is in the graph, and picking one store answers half the message. So each question is routed on its own, and each store is queried with the text of its own question — not with the whole message. That last part is not a detail: handing the full sentence to the Text-to-SQL step produced a query the guards rejected, and the relational half went unanswered.

Detecting the case is where this got interesting. The obvious rule is a threshold on the router's second-best score, and it was measured and thrown away: compound questions score top2 between 0.365 and 0.511, single ones between 0.303 and 0.476. The ranges overlap almost completely, and the plainly single question "Como uma reclamação de cliente deve ser tratada?" (0.476) outscores five of the seven compound ones. The cause is mechanical — an embedding of two subjects lands near their average, so a second subject dilutes both scores instead of lifting the second. It is the same shape of failure as the cache threshold below, and no cutoff separates it.

What separates them is structure. A compound message contains two questions, so clauses.py splits on clause boundaries (?, ;, e, and) and requires every surviving fragment to contain an interrogative. Without that second test, two of the 80 evaluation questions split wrongly: "Uma exportação sem anonimização e reportada para quem?", where the e is the verb é written unaccented, and "Qual a diferença entre um atraso comunicado e um atraso descoberto?", where it joins two noun phrases. In both, one half asks nothing at all.

Measured: 9 of 9 compound messages detected, 0 false positives in 80 single questions. The stores are then queried with asyncio.gather and their results concatenated under [SQL], [TEXT] and [GRAPH] labels — not fused by RRF, because RRF merges two rankings of the same items, and here the stores hold disjoint content and one of them returns a SQL result table rather than a ranked list.

One message, two questions, two stores

One message asking a figure and a rule. The route line reads relational + graph, and the answer carries both halves: the SUM from the sales table and the approval threshold from the compliance graph.

Lost in the middle

Retrieved passages are reordered before the model reads them, following Liu et al. 2023: a model attends to the start and the end of a context and least to its middle. Ranks 1, 3, 5 go out in order and 2, 4 come back reversed, so [1,2,3,4,5] becomes [1,3,5,4,2] — the best passage opens the context and the second best closes it. Pure list slicing; the ranking was already done. SQL rows are left alone, because their order is the ORDER BY that selected them.


Three stores, because there are three kinds of data

Holds Retrieved by Why not the others
RAG 1 · SQLite Rows and columns Generated SQL Aggregation must be exact. Correct SQL gives the correct number; no model does the arithmetic.
RAG 2 · Qdrant Free text Cosine over HNSW No schema to impose, no relations to model — only meaning.
RAG 3 · networkx Entities and relations Personalized PageRank Multi-hop. "If A fails, what breaks?" needs topology, not similarity.

A figure answered by generated SQL, with the query shown as the source

Why the first row of that table matters. The source is not a passage — it is the query that was run and the row it returned. The figure came out of a SUM, so it is either right or the SQL is wrong, and the SQL is on screen either way.

RAG 3 is HippoRAG 2 reimplemented from the paper: the model extracts subject–relation–object triples, networkx holds the graph, and the walk teleports back to the question's entities so the ranking means "relevant to this question".

The official hipporag package needs torch and vLLM, which are Linux/CUDA only. Reimplementing it was a constraint, and it turned out to be the more useful outcome — the maths is visible and testable rather than hidden behind an import.

The Corpus view: what was ingested and which store each chunk landed in

Routing is per chunk, not per file, which is why four of the six files show up in two stores. The row worth looking at is sobre_a_meridiano.md: one of its five paragraphs sits close enough to the routing boundary that rewording a route anchor moved it into the graph, and three questions that depended on it lost their answer. The split is also where the ingestion fallback shows: a chunk the graph could extract no relation from is kept as free text rather than dropped.


What is measured

uv run python scripts/benchmark.py on this machine (RX 9060 XT, 16 GB), 12 measured runs per question after 2 discarded warm-ups. A single number cannot describe a latency — the first call after a model loads pays for a cold cache, and the tail is what a user notices — so every row is p50 and p95 rather than an average.

p50 p95
Routing decision (deterministic stage) 8.8 ms 12.0 ms
Routing decision on the wire, streaming 39 ms 54 ms
Time to first token 114 ms 185 ms
Full answer, vectorial 252 ms 255 ms
Full answer, graph 724 ms 744 ms
Full answer, relational 737 ms 745 ms
Cache hit, same question 9.3 ms 12.9 ms
Share of a request spent inside model calls 99 %

The whole graph route costs 19 ms, and the model then spends 700 writing the answer. Retrieval is not where the time goes — the per-stage span breakdown that backs that claim, and retrieval/answer accuracy on three question sets (golden, held-out, and compound), are in docs/MEASUREMENTS.md. Headline numbers: 90–95% routing accuracy, 86% recall@5, 89–92% answer facts, measured on a held-out set never touched while tuning.


The maths, in one page

Nothing in the router, the cache or the graph walk is hidden behind a library call — cosine, normalisation, the margin rule, and Personalized PageRank, each with the formula, a worked example, and where it's tested. Written out in full in docs/MATHS.md.


Running it

Needs Python 3.12, Node 22, and roughly 6 GB of VRAM.

One command does the whole install — it installs uv if missing, syncs the locked dependencies, downloads the binaries and models, checks that Vulkan sees the GPU, and runs the tests. Every step checks whether its work is already done, so re-running is safe:

.\setup.ps1

-SkipModels leaves out the 4.9 GB download, -SkipFrontend skips npm, and -Start launches the servers and the API when it finishes. Or do the same steps by hand:

# 1 · dependencies
uv sync
npm --prefix frontend ci

# 2 · binaries and models — about 4.9 GB, skips whatever is already there
uv run python scripts/fetch_runtimes.py

# 3 · everything the backend needs: three llama-servers and Qdrant
uv run python scripts/start_servers.py

# 4 · the API
uv run uvicorn --app-dir backend app.main:app --port 8000

# 5 · the interface
npm --prefix frontend run dev

The servers run windowless and log to data/run/<name>.log. Stop them with scripts/start_servers.py --stop, or start and stop them individually from the Servers tab in the interface.

Tip

On a 16 GB card, stop a model you're not using from the Servers tab instead of leaving all three resident — ocr hands its VRAM back within a minute of going idle on its own, but llm and embeddings don't unless you stop them.

Then open http://localhost:5173, or http://localhost:8000/docs for the API.

Drop a file into data_drop/ and it is ingested within seconds — the hot folder is watched. Or use the Corpus tab, which is the same code path.

For something to ask it about, demo/ holds a six-file corpus about one fictional company, plus the questions that exercise each route and the two that do not work:

uv run python scripts/load_demo.py --reset

Note

Everything is driven by config.yaml: ports, model paths, thresholds, prompts, and the example phrases that define each route. Adding a route is a config change, not a code change.


Verifying it

uv run pytest tests/ -q          # 116 tests, no servers required
uv run ruff check backend/ scripts/ tests/
npm --prefix frontend run check  # 0 errors

The unit tests cover cosine and margin, the routing decision, the read-only SQL guard, entity normalisation and length limits, the tabularity heuristic, the cache, the clause splitter, the context ordering and every chunking rule. They need no GPU and no network — which is the point of keeping that code pure.

The chunking tests exist because chunking is where most of this project's real bugs came from, and they are verified by mutation rather than assumed: removing a guard has to fail exactly the test that defends it. A regression test that cannot fail is decoration.

The scripts under tests/ named *_integration.py are run by hand against a live stack and print their results; pytest collects nothing from them.

uv run python scripts/evaluate.py --reset            # re-ingests, then measures the golden set
uv run python scripts/benchmark.py                   # latency p50/p95, same requirements
uv run python -u tests/test_reingest_integration.py  # a corrected file replaces its old chunks

Seeing it run

Four views: Chat, Corpus, Telemetry and Servers. Chat pairs the conversation with a live panel that shows the per-route cosine scores, the margin the decision turned on, which stage made the call, and the span waterfall as it happens — the routing decision is on screen at about 39 ms, before retrieval has even finished.

The same question asked twice: the second answer comes from the cache

The same question, asked twice. The second answer skips routing, retrieval and generation entirely: three spans and 10 ms, against roughly a second for the first.

The Telemetry view: every trace, its spans and the attributes behind the decision

Every trace the backend produced, the span waterfall for the selected one, and the attributes that decided it. graph.seed_best is the cosine that chose the walk's starting node; cache.hit says why this request did the work instead of skipping it. This is the claim about auditability, rendered rather than asserted.

The Servers view: start and stop each model

Each model is its own process, running windowless. Stopping one hands its VRAM straight back, which matters on a 16 GB card when something else needs the GPU. ocr shows hollow because it put itself to sleep after a minute idle and gave back 2.08 GB on its own.

The screenshots above are generated, not taken by hand — uv run python scripts/screenshots.py drives Chrome over the DevTools protocol and rewrites them all. A screenshot nobody can regenerate is one that quietly starts lying after the next UI change.

To reproduce the demo yourself:

uv run python scripts/start_servers.py
uv run python scripts/load_demo.py --reset
uv run uvicorn --app-dir backend app.main:app --port 8000
npm --prefix frontend run dev

Then ask, in order: a figure (What was the total revenue of the Sudeste region?), something narrative (In what year was Meridiano Logistica founded?), a chain that crosses two files (Which approval policy do the orders processed by Sistema Atlas follow?), one message that asks two things at once (What was the Sudeste revenue and who approves a purchase of eighty thousand reais? — the panel shows relational + graph), and finally any of them a second time to watch the cache answer in 9 ms. The corpus is in Portuguese and the questions are in English on purpose: routing and retrieval work across languages because BGE-M3 embeds both into one space, and the sources are shown untranslated. demo/README.md has the full script, the expected figures, and the two questions that fail.


Known limitations

Real ones, found by running the system — not theoretical edge cases. Ten open items and four fixed since the last release, each with the failing case that found it, in docs/LIMITATIONS.md. The two that matter most right now: the vectorial route is still the weakest of the three, and a question contained inside another one can still slip past the cache.


Stack

FastAPI 0.141 · Qdrant 1.19 · networkx 3.6 · faiss-cpu 1.15 · OpenTelemetry 1.44 · SvelteKit 2.63 · Svelte 5 · llama.cpp (Vulkan) · uv

No PyTorch, no vLLM, no Ollama, no Docker, no WSL.

About

Federated RAG that picks the right store with arithmetic, not an agent. Runs locally based on Vulkan.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Contributors

Languages