47 to 38: how far a graph-backed AI fell when facts went missing
A peer-reviewed study removed facts from a knowledge graph on purpose. With 4 in 10 of the needed facts gone, the graph-backed system scored what the model scored with no graph at all.
47 to 38. That’s how far one graph-backed AI fell when 4 in 10 of the facts each question needed went missing from its knowledge graph. At 38, it matched the AI working with no graph at all.
Why it matters
A knowledge graph is a map of facts and how they connect. The promise is that the AI answers from the map instead of guessing.
But maps have gaps. A missing road leaves no mark on the map. You just get a worse route. In a company graph, gaps open every time the business changes faster than the graph is rebuilt.
A peer-reviewed study removed facts on purpose. An earlier graph method got about 47 of 100 complex questions right with the full map, and 38 with 4 in 10 key facts gone. The AI on its own, reasoning step by step, scored 39.
The study’s own method fills a missing fact from the AI’s memory, then has an AI check it, with no source to check against. It beat the earlier methods that need no extra training. Yet made-up facts were still its biggest source of errors, and with most of the needed facts gone it fell below the AI on its own.
My rule from this: a graph-backed AI should say when the fact isn’t in the graph, and I wouldn’t trust one until it’s been tested on a copy with facts deliberately removed.
How it works
Generate-on-Graph (EMNLP 2024) dropped 20% to 80% of the crucial triples on each question’s gold relation path, over Freebase, on 1,000 sampled CWQ and WebQSP questions.
On CWQ Hits@1, Think-on-Graph scored 47.2 with the complete graph, 37.9 at 40% removed, and 31.4 at 80% removed, on gpt-3.5-turbo-0613. On gpt-4-0613 the drop was larger: 71.0 to 56.1. Chain-of-thought with no graph at all scored 38.8.
How I read 37.9 against 38.8: a tie. On 1,000 questions, a score near 38 carries a standard error of about 1.5 points. The chain-of-thought figures also match the ToG paper’s ChatGPT numbers exactly, so they may come from a different run.
GoG has the model generate the missing triples and verify them, which reaches 44.3 at 40% removed. Its own Limitations section concedes hallucination in that step and a score below plain chain-of-thought on very incomplete graphs. ToG-2 (ICLR 2025) pairs the graph with documents instead, since knowledge graphs “inherently suffer from inner incompleteness.”
The engineering consequence sits one layer down. In Neo4j, a MATCH that finds nothing returns no rows and ends the query pipeline. An empty result carries no reason with it, which is exactly the signal a caller needs, so I’d log it as its own outcome rather than let it fall through as silence.