PERSIST stores cross-session spoken histories as readable event records, each tagged with speaker and timestamp, and retrieves them with a 3W joint scoring function combining semantic content, acoustic speaker identity, and temporal state.
Spoken dialogue memory jointly tracks who said what and when
PERSIST scores semantic content, acoustic speaker identity, and temporal state together, cutting retrieval latency from 578 ms to 7 ms in full-duplex spoken dialogue.
Chinese Tech
Achira Lin · Siyuan Hou · Wenyi Yu · Xinnian Zhao · Haoyu Niu · Wang Geng · +4 more
Tsinghua University · Huawei Technologies Ltd.
Research Digest··2 min read
The authors present PERSIST, a persistent memory system for multi-session, multi-speaker spoken dialogue that retrieves past conversational events by jointly scoring semantic content, acoustic speaker identity, and temporal state.
Why this paper
From Huawei Technologies Ltd. and Tsinghua University
In one line
PERSIST jointly models who said what and when to enable persistent memory in multi-speaker spoken dialogue systems.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§