← LinkedIn posts

47 to 38: how far a graph-backed AI fell when facts went missing

A peer-reviewed study removed facts from a knowledge graph on purpose. With 4 in 10 of the needed facts gone, the graph-backed system scored what the model scored with no graph at all.

Correct answers out of 100 complex questions: 47.2 with the full knowledge graph, 37.9 with 4 in 10 facts removed, 44.3 for the method that writes the missing fact from memory, and 31.4 with 8 in 10 removed, against a dashed line at 38.8 for the model reasoning with no graph. Beside it, a map diagram where a missing road leaves no mark, and three options when the fact is not on the map: fill it from memory, pair the graph with documents, or say the fact is not there.
The whole post in one graphic. Open the graphic full size ↗

47 to 38. That’s how far one graph-backed AI fell when 4 in 10 of the facts each question needed went missing from its knowledge graph. At 38, it matched the AI working with no graph at all.

Why it matters

A knowledge graph is a map of facts and how they connect. The promise is that the AI answers from the map instead of guessing.

But maps have gaps. A missing road leaves no mark on the map. You just get a worse route. In a company graph, gaps open every time the business changes faster than the graph is rebuilt.

A peer-reviewed study removed facts on purpose. An earlier graph method got about 47 of 100 complex questions right with the full map, and 38 with 4 in 10 key facts gone. The AI on its own, reasoning step by step, scored 39.

The study’s own method fills a missing fact from the AI’s memory, then has an AI check it, with no source to check against. It beat the earlier methods that need no extra training. Yet made-up facts were still its biggest source of errors, and with most of the needed facts gone it fell below the AI on its own.

My rule from this: a graph-backed AI should say when the fact isn’t in the graph, and I wouldn’t trust one until it’s been tested on a copy with facts deliberately removed.

How it works

Generate-on-Graph (EMNLP 2024) dropped 20% to 80% of the crucial triples on each question’s gold relation path, over Freebase, on 1,000 sampled CWQ and WebQSP questions.

On CWQ Hits@1, Think-on-Graph scored 47.2 with the complete graph, 37.9 at 40% removed, and 31.4 at 80% removed, on gpt-3.5-turbo-0613. On gpt-4-0613 the drop was larger: 71.0 to 56.1. Chain-of-thought with no graph at all scored 38.8.

How I read 37.9 against 38.8: a tie. On 1,000 questions, a score near 38 carries a standard error of about 1.5 points. The chain-of-thought figures also match the ToG paper’s ChatGPT numbers exactly, so they may come from a different run.

GoG has the model generate the missing triples and verify them, which reaches 44.3 at 40% removed. Its own Limitations section concedes hallucination in that step and a score below plain chain-of-thought on very incomplete graphs. ToG-2 (ICLR 2025) pairs the graph with documents instead, since knowledge graphs “inherently suffer from inner incompleteness.”

The engineering consequence sits one layer down. In Neo4j, a MATCH that finds nothing returns no rows and ends the query pipeline. An empty result carries no reason with it, which is exactly the signal a caller needs, so I’d log it as its own outcome rather than let it fall through as silence.