← LinkedIn posts

The costlier AI search setup lost on simple fact questions

A peer-reviewed KDD 2026 study ran plain retrieval and graph retrieval through the same tests. Which one won depended on the question, and the AI judge scoring the contest sometimes reversed its verdict when the answers swapped places.

Index-style RAG and graph-style GraphRAG measured on the same KDD 2026 tests: the index won one-step and fine-detail questions, the map won questions that chain several facts. A bar chart of build time shows 135 seconds for RAG against 5,560 and 7,702 seconds for the two graph variants. Below it, the LLM judge that reversed some verdicts when the answers swapped order, and a router that gained 1.1% against 6.4% for running both retrievers.
The whole post in one graphic. Open the graphic full size ↗

The costlier AI search setup lost on simple fact questions. And the AI judge scoring the contest sometimes reversed its verdict when the answers swapped places.

Why it matters

Plain AI search works like a book’s index: find the passage that carries your words. Graph search works like a map of how people, places, events, and facts connect.

A peer-reviewed study at KDD 2026, the ACM’s main data-mining conference, put both through the same tests. The map won on questions that chain several facts. The index won on one-step and fine-detail questions. In one test, the map versions took 41 to 57 times longer to build.

The map’s zoomed-out summaries helped broad “what are the themes?” questions. They also dropped the fine details that narrow questions needed.

Then the scoring problem. An AI judge rated the summaries. Swap their order, and some verdicts flipped.

I had already built that rule into my agentic system: an AI judge’s score does not count as evidence unless the answer order is shuffled. This study shows what happens without it.

So I sort the questions before I pick the tool, and shuffle the order before I trust a score.

How it works

Han and co-authors compared RAG and GraphRAG under one protocol. RAG led on single-hop and detail-oriented question answering (NQ, and the detail subsets of NovelQA). GraphRAG methods led on multi-hop question answering (HotPotQA, MultiHop-RAG).

Community-based global search, the Microsoft GraphRAG design, can lose fine-grained evidence. It showed on NovelQA’s detail questions: the summary that makes a theme visible is the same summary that drops the number you asked for.

Cost is uneven. On MultiHop-RAG, the two graph variants took 5,560 and 7,702 seconds to build, against 135 seconds for RAG. Yet community retrieval ran faster than RAG at query time. The expense sits in the build, not the query, which is a different budget and a different decision.

Routing is the fix, with a trade-off. An LLM classifier labeled each query fact-based or reasoning-based, then sent it to RAG or GraphRAG: 1.1% over the best baseline, Llama 3.1-70B. Running both and merging scored 6.4% higher, but every query then pays for both.

My call for a first build is the router. It is cheaper, and each route gets scored on its own, so a regression has one place to come from. I would pay for both only where the measured gap justifies it, and I would randomize answer order before any LLM judge sees them.

What the study cannot settle is your corpus. These are public benchmarks, where the entities are whatever the dataset happened to contain. On a corpus whose entities are your own, suppliers, claims, parts, the map has more to connect than NQ ever gives it, and the build cost is paid once against a schema that changes slowly.