Spatial audio model localizes and transcribes sounds across smart-glasses arrays

SEA-LM combines spatial and mono audio encoders to identify, locate and transcribe overlapping sounds while distinguishing the wearer from nearby speakers.

Big Tech
Sonal Kumar · Sinan Hersek · Artem Dementyev · Mengzhen Pan · Ishan Chatterjee · Anurag Kumar · +3 more

University of Maryland, College Park · Google · Google DeepMind

Research Digest··2 min read
Kumar and colleagues developed SEA-LM, an audio-language model designed for egocentric, multichannel recordings from smart glasses.

The authors built FOACoder, a spatial encoder that processes First Order Ambisonics, a four-channel representation encoding sound pressure and direction.

Why this paper

From Google and 2 others

In one line

SEA-LM enables spatial audio understanding from wearable microphone arrays, outperforming baselines in localization and transcription.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (3 noted)
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.