EchoLens: A Human-Speech Dataset for Auditing Demographic Sensitivity in Audio-Language Models
Abstract
Voice interfaces are increasingly moving away from transcription pipelines towards end-to-end systems that directly respond to audio inputs. This development in turn requires a shift in evaluation methodology away from transcription accuracy and towards more substantive markers such as response validity. We introduce EchoLens, a demographically stratified dataset of 20008 recordings from 515 adult participants residing in the U.S., with self-reported demographic information across race, gender, age, accent and primary language, among others. Each participant speaks out loud and verbatim a randomized subset of 69 advice-seeking and estimation prompts spanning 11 domains grounded in the American Time Use Survey. Most prompts elicit quantitative responses from the model, allowing for direct analysis of output distributions across demographic subgroups without requiring reliance on LLM-as-a-judge. EchoLens also contains a documented audit protocol, which—in an illustrative application—we use to evaluate six current audio-language models. Under this assessment, we find that racial disparities are more frequent and larger in magnitude than gender disparities, with particular concentration in two models.