Spoken dialogue memory jointly tracks who said what and when

PERSIST scores semantic content, acoustic speaker identity, and temporal state together, cutting retrieval latency from 578 ms to 7 ms in full-duplex spoken dialogue.

Chinese Tech
Achira Lin · Siyuan Hou · Wenyi Yu · Xinnian Zhao · Haoyu Niu · Wang Geng · +4 more

Tsinghua University · Huawei Technologies Ltd.

Research Digest··2 min read
The authors present PERSIST, a persistent memory system for multi-session, multi-speaker spoken dialogue that retrieves past conversational events by jointly scoring semantic content, acoustic speaker identity, and temporal state.

PERSIST stores cross-session spoken histories as readable event records, each tagged with speaker and timestamp, and retrieves them with a 3W joint scoring function combining semantic content, acoustic speaker identity, and temporal state.

Why this paper

From Huawei Technologies Ltd. and Tsinghua University

In one line

PERSIST jointly models who said what and when to enable persistent memory in multi-speaker spoken dialogue systems.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.