The authors introduce InterTab, a framework for actively acquiring visual evidence during table reasoning.
Interleaved table crops improve reasoning over dense visual layouts
InterTab lets a multimodal model request row, column and cell crops while reasoning, linking each step to specific table evidence.
Chinese Tech
Hanqian Li · Sirui Huang · Chen Ling · Jungang Li · Yu Huang · Kening Zheng · +9 more
The Hong Kong University of Science and Technology (Guangzhou) · Zhejiang University · The Hong Kong University of Science and Technology · City University of Hong Kong · University of Illinois at Chicago
Research Digest··2 min read
The authors train a multimodal language model to alternate between chain-of-thought reasoning and tool calls that crop relevant regions from a table image.
Why this paper
From Ant Group and 6 others · Part of Context Engineering for Agents, now 41 papers
In one line
InterTab interleaves chain-of-thought with structure-aligned region crops to improve multimodal table reasoning.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§