CypherTurn contains 721 sessions and 5,927 turns spanning seven knowledge graphs and 13 conversational phenomena.
Conversational graph querying exposes a gap in model autonomy
Across 5,927 conversational turns, model rankings changed under autonomous operation and no system completed 5% of sessions correctly.
Chinese Tech
Yuzhe Zhang · Weijie Zhu · Haolin Yang · Ziyun Zhang · Xianwei Xue · Mengke Chen · +2 more
Peking University · National Key Lab of Data Space Technology and System · Baidu Inc.
Research Digest··3 min read
Zhang et al.
Why this paper
From Baidu Inc. and 2 others
In one line
Conversational Text-to-Cypher remains unreliable: the best model achieves 64.7% execution accuracy, completes fewer than 5% of sessions correctly, and changes rank under autonomy.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§