TopoGraphRAG-Bench

Evaluating Multimodal GraphRAG on
Layout-Grounded Evidence Reasoning

Ruochi Li1, Jianzhe Lin, Haoxuan Zhang2, Haihua Chen2, Junhua Ding3, Edward Gehringer1, Yang Zhang2

1 North Carolina State University 2 University of North Texas 3 University of Wyoming

Bottom-up construction of TopoGraphRAG-Bench with three controlled reasoning topologies and counterfactual validation.
Bottom-up question construction with controlled evidence topologies and counterfactual validation.

Abstract

Real-world documents distribute evidence across text, tables, figures, and captions within complex page layouts. Answering complex questions over such documents therefore requires more than retrieving relevant passages: systems must recover the evidence topology that connects heterogeneous evidence units. Existing GraphRAG evaluations remain largely text-centered, while multimodal document RAG benchmarks assess cross-modal retrieval and generation without directly evaluating recovery of the intended evidence topology. We introduce TopoGraphRAG-Bench, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents. Questions are constructed bottom-up from text, figure, and table evidence units under three controlled topologies: single-hop retrieval, bridge-chain reasoning, and multi-source synthesis. To ensure that questions preserve their intended structure, we apply counterfactual validation for shortcut resistance, modality necessity, and evidence necessity. We evaluate text-only GraphRAG, page-level visual retrieval, and multimodal GraphRAG systems using retrieval, generation, and topology-aware reasoning metrics. Multimodal GraphRAG systems achieve the strongest overall performance, but still fail when visual-textual evidence alignment or multi-unit composition is incomplete. Text-only GraphRAG struggles when key dependencies are grounded in figures or tables, while page-level visual retrieval lacks the fine-grained structure needed for topology recovery. These findings motivate GraphRAG systems that move beyond text-derived entity relation graphs to explicitly model document layouts, cross-modal evidence alignment, and the reasoning roles of evidence units.

Highlights

  • Topology-aware multimodal GraphRAG benchmark. We introduce TopoGraphRAG-Bench, a layout-grounded benchmark comprising 2,024 QA pairs over 201 long documents, covering single-hop retrieval, bridge-chain reasoning, and multi-source synthesis across text, figures, and tables.
  • Bottom-up evidence-topology construction. We construct questions bottom-up from verifiable, layout-grounded evidence units rather than whole-document prompts, and apply counterfactual validation to assess shortcut resistance, modality necessity, and evidence necessity.
  • Diagnostic evaluation of GraphRAG paradigms. We evaluate text-only GraphRAG, page-level visual retrieval, and multimodal GraphRAG systems, revealing persistent challenges in cross-modal alignment and fine-grained evidence-topology recovery.

Overview

TopoGraphRAG-Bench evaluates layout-grounded multimodal evidence reasoning in GraphRAG. It contains 2,024 questions constructed from 201 long documents sampled from MMDocIR, spanning nine domains, including academic papers, legal documents, research reports, and financial reports. The documents average 45.5 pages and approximately 13,000 words, with evidence distributed across text, figures, tables, and page layouts.

Questions are organized into three controlled reasoning topologies: single-hop retrieval (302 questions), bridge-chain reasoning (1,014), and multi-source synthesis (708). Cross-modal questions account for 38.2% of the benchmark, while 66.8% require figure or table evidence. Bridge-chain questions cover all nine ordered modality paths among text, figures, and tables. Each question includes a gold answer, page- and layout-grounded evidence references, and evidence-topology annotations that make the required evidence and its reasoning roles explicit.

Dataset statistics and fine-grained evidence topology: 201 documents, 2,024 questions, 302 single-hop questions, 1,014 bridge-chain questions, and 708 synthesis questions. The table also reports evidence distributions, all nine bridge paths, and synthesis evidence counts.
Distribution of 201 documents across nine domains: academic paper 45, laws 38, research report 33, guidebook 20, tutorial/workshop 14, brochure 14, government 14, financial report 13, and administration/industry file 10.
Domain distribution of the 201 documents. Slices show percentages; the legend reports document counts using the original MMDocIR domain labels.

Construction

The TopoGraphRAG-Bench construction pipeline follows three main steps. (1) Evidence preparation and anchor discovery: We organize layout-grounded text, figure, and table units from MMDocIR, retaining their document, page, layout, and modality identifiers. Visual units preserve their original images and are supplemented with OCR and VLM descriptions. Recurring entities across units serve as anchors for evidence composition. (2) Topology-controlled question construction: Single-hop questions are grounded in one evidence unit; bridge-chain questions connect evidence units through ordered dependencies, with intermediate hops resolving bridge entities needed to locate or interpret subsequent evidence; synthesis questions combine three to six required evidence points to support a joint conclusion. Gold answers and topology annotations retain links to the source evidence. (3) Counterfactual validation and quality assurance: Single-source counterfactual probes screen bridge-chain candidates for shortcuts, OCR-aware checks assess modality necessity, and synthesis candidates undergo leave-one-out evidence tests and a text-only retrieval simulation. A stratified human audit further checks answer correctness, evidence grounding, attribution accuracy, and topology consistency.

Evaluation

We evaluate six systems spanning three retrieval paradigms: LightRAG, HippoRAG, and Microsoft GraphRAG for text-only GraphRAG; VisRAG for page-level visual retrieval; and RAG-Anything and MegaRAG for multimodal GraphRAG. Evaluation combines topology-enriched Context Precision and Context Recall with Answer Accuracy, Faithfulness, Response Relevancy, and Step Coverage, which measures coverage of annotated hop results or synthesis evidence points. MegaRAG and RAG-Anything achieve the strongest overall generation performance, reaching answer accuracy scores of 53.5% and 53.0%, respectively, while RAG-Anything obtains the strongest text-based retrieval scores. Performance varies substantially with evidence topology and modality: text-only GraphRAG performs competitively on textual single-hop questions but degrades when bridge dependencies require figures or tables, while VisRAG struggles on multimodal bridge and synthesis questions despite direct access to page images. These differences also appear in coverage of intermediate reasoning results: on visual-only bridge paths, MegaRAG achieves 83.0% Step Coverage, compared with 44.9% for LightRAG and 49.4% for VisRAG. The remaining variation across ordered bridge paths and the difficulty of multi-source synthesis highlight persistent challenges in cross-modal alignment and evidence composition.

Answer accuracy of six systems across single-hop, bridge-chain, and synthesis subsets, separated by textual and visual evidence. Multimodal GraphRAG systems show stronger performance on compositional visual subsets.
Answer accuracy across reasoning topologies and evidence modalities. MM denotes subsets requiring figure or table evidence.
Bridge-chain evaluation: a radar plot compares six systems on all nine ordered modality paths; three tables report answer accuracy, faithfulness, and step coverage for text-only, mixed, and visual-only bridge families.
Bridge-chain performance by evidence path and modality family. Left: path-level answer accuracy. Right: answer accuracy, faithfulness, and step coverage.

Case Study

Explore the highlighted source evidence, system responses, retrieved contexts, and step-level judge decisions in our interactive case study.

Loading case study…

Enlarged source evidence