Conversational graph querying exposes a gap in model autonomy

Across 5,927 conversational turns, model rankings changed under autonomous operation and no system completed 5% of sessions correctly.

Chinese Tech
Yuzhe Zhang · Weijie Zhu · Haolin Yang · Ziyun Zhang · Xianwei Xue · Mengke Chen · +2 more

Peking University · National Key Lab of Data Space Technology and System · Baidu Inc.

Research Digest··3 min read
Zhang et al.

CypherTurn contains 721 sessions and 5,927 turns spanning seven knowledge graphs and 13 conversational phenomena.

Why this paper

From Baidu Inc. and 2 others

In one line

Conversational Text-to-Cypher remains unreliable: the best model achieves 64.7% execution accuracy, completes fewer than 5% of sessions correctly, and changes rank under autonomy.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.