NeurIPS 2026

RAIL

Rethinking Auditory Intelligence in Large Audio-Language Models

with a CHC-grounded benchmark

Hongyu Jin1,* Siyi Wang1,* Yang Xiao1,* Jiaheng Dong1,* Shihong Tan4 Kaiyuan Peng1 Georgiana Juravle2 Shanquan Chen3 Gongping Huang4 Hong Jia5 Eun-Jung Holden1 James Bailey6 Ting Dang1,†

1The University of Melbourne 2Alexandru Ioan Cuza University of Iași 3The University of Hong Kong 4Wuhan University 5The University of Auckland 6Monash University

*Equal contribution · †Corresponding author

rail_session · 5 capabilities · 32 subcapabilities
One listener, five capabilities. Each track stands for one CHC broad ability that RAIL measures. Select a track to explore it ↓
5
core capabilities
32
subcapabilities
5,306
audio QA samples
30.6
hours of audio
26
LALMs evaluated
TL;DRRAIL evaluates large audio-language models the way cognitive science assesses human listeners. Grounded in Cattell–Horn–Carroll (CHC) theory, it organizes 5,306 audio questions into five core auditory capabilities and 32 subcapabilities, and compares 26 LALMs with a human baseline. Models do well on knowledge inherited from text pretraining, while fine-grained auditory perception and processing efficiency remain below human level.
01
Abstract

From task scores to auditory cognition

Humans process rich auditory environments through various cognitive capabilities such as audio perception, audio reasoning and memory. Despite recent progress in large audio-language models (LALMs), current evaluation remains largely task- or domain-centric, focusing on end performance while ignoring underlying auditory cognitive behaviors. This gap between how auditory cognition is understood in humans and how it is evaluated in LALMs hinders the interpretation of model behaviour in underlying capabilities and limits alignment with human auditory objectives.

We introduce RAIL, a human-centric benchmark grounded in the Cattell–Horn–Carroll (CHC) framework, which formalises auditory cognition into five core capabilities, and develops them into structured evaluation tasks with principled data curation and human-aligned evaluation protocols. Evaluating 26 LALMs reveals strong performance in knowledge-related tasks but weak auditory perception and memory, reflecting reliance on language-based pretraining. Limited reasoning under auditory settings further suggests a mismatch between existing training paradigms and human-like auditory reasoning, as well as over-reasoning at the cost of efficiency. Six LALMs exceed human performance in general, while auditory processing all lags behind human. Overall, RAIL establishes a new framework for human-aligned evaluation of auditory intelligence.

02
At a glance

What RAIL reveals about today's audio models

Imagine a busy café: cups clatter, conversations overlap, music plays in the background. A human listener separates speech from noise, works out who is talking to whom, and recognizes a colleague's voice from an earlier meeting. RAIL asks whether LALMs can do the same, one capability at a time. Six findings stand out.

01Knowledge vs. Auditory

Knowledge is inherited; perception is not

Knowledge (56.21) and Memory (55.05) have the highest mean scores across 26 models. Auditory Processing is the lowest (43.83).

02Ga

Language cues mask weak listening

Language-supported auditory tasks reach medians of 0.56–0.74. Perceptual-only tasks drop to 0.31–0.42, and sound localization to 0.28. The gap holds in 18 of 26 models.

03Gf

Sequential reasoning is the weakest link

When models must apply sound-triggered rules step by step and track an evolving state, most open-source models score below 50%.

04Gsm / Glr

Speech memory is not sound memory

The best models reach 100 on speech memory span. On memory for non-speech sound patterns, no model exceeds 60.

05Gs

Bigger is not more efficient

Across 21 open-source models, size does not predict B-AUC. Models of 8.4–8.6B parameters range from 0.2 to 0.7.

06Human baseline

Humans still hear better

Six LALMs beat the human baseline overall, so humans rank 7th. No model matches humans on auditory processing or processing efficiency.

03
Framework

Five capabilities, 32 subcapabilities

CHC theory is a psychometric taxonomy of human cognitive abilities, derived from more than 70 years of factor-analytic studies and validated across populations and assessment batteries. It is hierarchical: broad abilities split into narrow abilities that can each be measured. RAIL restricts CHC to the auditory channel and turns every narrow ability into an audio task.

With cognitive experts, we compared CHC against Bloom's taxonomy, which describes educational task difficulty rather than latent cognitive structure, and Gardner's multiple intelligences, which lacks standardized measurement and empirical validation. CHC was the only data-driven, measurable option.

Hover a segment to see a subcapability
CHC-grounded RAIL benchmark: five capabilities arranged in a wheel with their 32 subcapabilities
Figure 1. The CHC-grounded RAIL benchmark, with audio tasks organized around five capabilities. Click any figure to enlarge.
04
Results

Leaderboard: uneven capability profiles across 26 LALMs

Table 1. Scores (%) across the five CHC capabilities. Click a column to sort. Bold and underline mark the best and second-best model among the rows shown. Mean is the unweighted average of the five columns; under LLM-as-Judge it gives the paper's open- and closed-source averages of 46.27 and 65.10. Under the strict setting, the efficiency column is B-AUC.

65.10 vs. 46.27
Closed-source vs. open-source mean (LLM-as-Judge)
Gemini 3.1 Pro
Best model overall
Omni R1
Best open-source model overall

Evaluation protocol

26 models, four lines

21 open-source models from 167M to 33.5B parameters and 5 closed-source APIs, covering speech-centered LLMs, general audio-language models, omni multimodal models and closed-source APIs.

Two accuracy metrics

Strict ACC counts a response as correct only when every token of the final answer matches the ground truth. LLM-as-Judge (GPT-5.4) scores semantic equivalence with the ground truth.

Efficiency by B-AUC

Efficiency uses reasoning tokens per response as a proxy for computational effort. B-AUC rewards models that reach correct answers within short reasoning budgets. Latency is not used because it depends on serving infrastructure.

Capability distributions and correlations

Violin plots of LLM-judge scores for the five capabilities across 26 models
Figure 2a. Score distributions across 26 models. Memory has a high mean but the largest spread (std = 22.46).
Partial Spearman correlation matrix between the five capabilities
Figure 2b. Pairwise correlations between capabilities. Reasoning and memory are most strongly related (ρ = 0.798).

Knowledge is largely inherited from the text-based LLM, whose pretraining supplies strong semantic priors. Auditory processing needs fine-grained reasoning over frequency structure, spatial cues and temporal dynamics, which current audio encoders learn only weakly.

Reasoning and memory correlate most strongly (ρ = 0.798), which suggests shared reliance on multi-step inference. Auditory processing correlates moderately with efficiency (ρ = 0.507) and knowledge (ρ = 0.526): once audio is encoded correctly, models tend to produce shorter reasoning traces, and better auditory grounding also helps them reach stored knowledge through the audio pathway.

05
Benchmark

A four-stage, human-in-the-loop curation pipeline

Computer scientists and cognitive experts built RAIL together, with LALMs assisting in data generation. Every stage feeds back into task formulation until both groups agree on ability definitions, task design and sample quality.

Four-stage pipeline: cognitive framework selection, task formulation, dataset curation, quality control, with a refine loop
Figure 3. The four-stage benchmark curation pipeline of RAIL.
STAGE 1

Framework selection

Compare Bloom's taxonomy, multiple intelligences and CHC. Adopt CHC for its measurable, hierarchical structure.

STAGE 2

Task formulation

Translate each CHC ability into auditory tasks under two principles: auditory dependence and capability independence.

STAGE 3

Dataset curation

Select audio from existing corpora or synthesize it with TTS when no source fits. Build QA pairs from templates, rule-based pipelines or LLM-assisted generation, then verify by hand.

STAGE 4

Quality control

Cognitive experts check that each item targets the intended ability, is answerable from audio, and has an unambiguous answer. Refine until consensus.

Two design principles

Auditory dependence

Every task is centered on audio cues. Memory tasks, for example, rely on spoken cues in past dialogue, so they measure acoustic memory rather than text recall.

Capability independence

No target ability depends on another as a prerequisite for being measured. Each of the five capabilities is assessed on its own.

Data statistics

RAIL contains 5,306 samples across 32 tasks and 30.6 hours of audio, and 68.1% of the samples are newly constructed. Processing-efficiency items are short (mean 4.02 s, median 1.81 s) to match rapid-response evaluation, while memory and reasoning items are much longer (means of 46.65 s and 39.04 s) because they need extended auditory context.

Capability# SamplesTotal dur. (h)Mean dur. (s)Median dur. (s)# New / All# Tasks
Ga Auditory Processing1,1704.9315.174.00701 / 1,1707
Gf Reasoning3225.1939.0425.04222 / 3223
Gsm/Glr Memory1,00013.046.6528.64900 / 1,0006
Gs Processing Efficiency1,8002.014.021.811,800 / 1,8009
Gc/Gkn Knowledge1,0147.225.4810.49141 / 1,0147
Overall5,30630.620.747.703,614 / 5,30632

Table 2. Data statistics of RAIL. A 640-item human-evaluation subset (20 items per subcapability) was also answered by 24 participants.

What existing benchmarks miss

Existing audio benchmarks organize evaluation by task or audio domain. None is grounded in a cognitive theory with full coverage, and audio memory and processing efficiency are absent from all of them.

BenchmarkTheory-groundedDomainsReasoningMemoryAuditoryKnowledgeEfficiency

✓ full systematic coverage · ◐ partial or non-systematic coverage · ✗ not covered. Domains: S = Speech, So = Sound, M = Music.

Example items for each of the five core capabilities across music, sound and speech
Figure 4. Task formulation for the five core capabilities, with example items spanning music, sound and speech.
06
Human baseline

Where humans still lead

24 participants with normal hearing answered a 640-item subset (20 items per subcapability). Each item received 2–5 independent responses, and each clip could be played only once. Models were scored on the same items with the same aggregation.

Pick up to five models to compare with the human profile (LLM-as-Judge, %, human-evaluation subset).

Human rank against the 26 LALMs on each capability. The bar shows the best model's score; the white tick marks the human score.

Humans score highest on auditory processing and processing efficiency. Models struggle to perceive nonverbal acoustic cues and subtle sound variations, and many high-accuracy models generate unnecessarily long reasoning for simple inputs, while humans answer accurately with compact reasoning.

On memory and reasoning, humans rank 13th and 18th. Top models such as Gemini 3.1 Pro likely benefit from reasoning-oriented post-training and transformer context aggregation, which support associative retrieval over long sequences and step-by-step inference. Human memory is more structured but capacity-limited, which constrains retrieval and multi-step reasoning.

Show all 26 models on the human-evaluation subset
#ModelAuditoryReasoningMemoryEfficiencyKnowledgeOverall
07
Analysis

What each capability tells us

Ga

Auditory Processing

Models do far better when language can help

Speech Sound Discrimination (US), Resistance to Auditory Stimulus Distortion (UR) and Musical Discrimination (U1/U9) let models lean on words and language knowledge. Their medians are 0.56–0.74, up to 0.95 for frontier models. Phonetic Coding (PC), Absolute Pitch (UP) and Rhythm (U8) need direct perception of phonemes, tones and temporal structure, and fall to 0.31–0.42. Sound Localization is lowest (median 0.28, max 0.46). The pattern holds for 18 of 26 models.

Audio encoder design drives sensory performance

Step-Audio-2-mini leads Phonetic Coding; its Whisper-based encoder, pretrained on large-scale multilingual ASR, may preserve fine phonemic detail. DIFFA-2 leads Absolute Pitch with a Q-Former adapter over intermediate Whisper layers. Gemini 3.1 Pro is best on Rhythm and Sound Localization.

TakeawayLALM auditory ability is still driven mainly by text-based learning rather than audio perception. Front-ends that explicitly model auditory cues help on specific tasks; broader progress needs audio-centric design.
Violin plots for seven auditory subcapabilities grouped into language-supported, perceptual-only and spatial
Figure 5. Scores across the seven auditory-processing subcapabilities, grouped by whether language can support the answer.
Gf

Fluid Reasoning

General sequential reasoning is the weakest

Sequential reasoning trails induction and quantitative reasoning for most open-source models, most of which score below 50%. Models must apply sound-encoded operations in order (for example, "sniff" = reverse, "cough" = duplicate, "throat clearing" = delete) while updating an intermediate state. Chain-of-thought post-training produces plausible intermediate text rather than explicit state updates, whereas humans maintain and propagate an evolving state step by step.

Reasoning post-training does not transfer to audio by itself

Step Audio R1 and Qwen3-Omni-30B do well on quantitative reasoning, partly thanks to math-related post-training of their language backbones. Audio Flamingo 3 and Omni R1, whose backbones had similar post-training, score noticeably lower. What matters is whether reasoning transfers to settings where numerical relations must be inferred from audio.

TakeawayReasoning is learned in a task- and data-specific way rather than as reusable operations. Training should move beyond scaling and surface-level CoT toward reusable, audio-grounded reasoning structures.
Heatmap of the top 10 open models on induction, quantitative and sequential reasoning
Figure 6. Top-10 LALMs across the three reasoning subcapabilities (LLM-as-Judge, %).
Gsm / Glr

Memory

Non-speech memory is the weakest

Memory for Sound Patterns (UM) tests retention of environmental sounds and prosody. No model exceeds 60; the best is Gemini 3.1 Pro at 59. Speech-based Memory Span (MS) is far stronger: Gemini 2.5 Flash and Gemini 3.1 Pro reach 100, and several open-source models exceed 70.

Free recall splits into two regimes

Models hear a continuous dialogue with a short burst of unrelated target words and must recall only those words. Six models, including Gemini 3.1 Pro and Step Audio R1, score above 87 (up to 97.1), while six others, such as Mellow and Gemma-3n-E4B-it, score far below 10. Models optimized for long-form audio do better than those trained on short utterances or captions.

TakeawayAuditory memory is uneven across subcapabilities. Closing the gaps needs targeted training and evaluation for long-term, non-speech and multi-turn memory.
Heatmap of six memory subcapabilities for selected models
Figure 7. Six memory subcapabilities (LLM-as-Judge, %).
Bar chart of Memory Span versus Memory for Sound Patterns for ten models
Figure 8. Memory Span (MS) versus Memory for Sound Patterns (UM) for 10 LALMs.
Gs

Processing Efficiency

Efficiency only matters relative to success, so RAIL scores it with B-AUC, the area under the accuracy curve up to a reasoning-token budget:

B-AUC = (1/budget) · Σb=0…budget−1 (Acc≤b + Acc≤b+1) / 2

Models allocate reasoning budget very differently

Accuracy is only weakly correlated with reasoning length. Gemini 3.1 Pro reaches the highest accuracy (0.953) with longer but consistent responses (std = 2.90). Kimi-Audio reaches comparable accuracy with much shorter outputs, a better trade-off. GLM-4-Voice and Gemma-3n-E4B-it vary more across tasks (Gemma std = 7.87) without being more effective.

Model size does not determine efficiency

For 21 open-source models (167M–33.5B), B-AUC shows no clear relation to size. Models of 8.4–8.6B range from 0.2 to 0.7. Small or weakly instruction-tuned models often produce repetitive or unconstrained output, and large ones may reason verbosely. Gemma-3n-E4B beats Gemma-3n-E2B (47.92 vs. 24.49) despite the same architecture.

TakeawayEfficiency is neither optimized explicitly nor built into training objectives. Future LALMs should optimize accuracy and efficiency jointly, with reasoning depth that adapts to the input.
Scatter of accuracy versus mean reasoning length, colored by B-AUC
Figure 9a. Accuracy vs. mean reasoning-token length; bars show std across the nine subcapabilities.
Scatter of model size versus B-AUC on a log scale
Figure 9b. Model size vs. B-AUC for open-source models.
Gc / Gkn

Acquired Knowledge

Machine-sound knowledge is the outlier

Six of the seven knowledge tasks behave alike: wide score ranges, top scores of 0.78–0.92, and high mutual correlation (Pearson r = 0.68–0.97). Mechanical Knowledge (MK) is tightly clustered at 0.23–0.48, near chance (0.33) even for Gemini 3.1 Pro (0.45), and weakly or negatively correlated with the other tasks (r = −0.30 to 0.09). It requires fine discrimination between acoustically similar machine sounds, an underrepresented and perceptually demanding domain. Gemma-3n-E4B scores 0.97 on it but underperforms elsewhere, which points to dataset-specific exposure.

TakeawayLALMs handle speech- and music-based knowledge well but struggle to recognize machine sounds. Closing this gap likely needs training data that pairs domain-specific sounds with expert annotations, not more general-purpose audio.
Violin plots for seven knowledge subcapabilities
Figure 10. Scores across the seven knowledge subcapabilities.
08
Outlook

Toward human-aligned auditory intelligence

Training paradigms inherited from text-based LLMs, centered on task supervision, CoT reasoning and short audio dialogues, do not yet support generalizable auditory understanding. Progress needs a shift in both evaluation and training: rather than optimizing for tasks or audio domains, future work should target auditory capabilities that are structured, adaptive and grounded in sensory input.

Audio-centric perception

Encoders and front-ends that preserve fine acoustic detail, so perception does not hinge on language cues.

Stateful, reusable reasoning

Reasoning that updates an explicit state over time and transfers across tasks and modalities.

Long, non-speech memory

Training and evaluation for long-term, non-speech and multi-turn auditory memory.

Efficiency-aware objectives

Objectives that reward correct answers reached with reasoning depth matched to the input.

09
Citation

BibTeX

@inproceedings{jin2026rail,
  title     = {RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models
               with a CHC-Grounded Benchmark},
  author    = {Jin, Hongyu and Wang, Siyi and Xiao, Yang and Dong, Jiaheng and
               Tan, Shihong and Peng, Kaiyuan and Juravle, Georgiana and Chen, Shanquan and
               Huang, Gongping and Jia, Hong and Holden, Eun-Jung and Bailey, James and
               Dang, Ting},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2606.11260}
}