An illustrative chest-voice and falsetto comparison produced a Wasserstein distance of about 442 Hz, while fingering and harmony applications remain proposals rather than validated performance results.
The results point to a trade-off: the transformation embedding was more useful for distance-based comparisons, while the processed-audio embedding was more useful for recovering chain and parameter information.
Models transferred well within the same sound domain but poorly between music and environmental recordings, and added labels recovered only part of the gap.
Across NSynth-100, FSC-89 and LS-100, SPECTRA reported higher average accuracy and lower performance drop than TAPE in standard comparisons, with the same direction under harder class-incremental protocols.
AudioLens-R1 led a benchmark that asked an audio-language model to group speech recordings under different natural-language instructions, including semantic and paralinguistic criteria.
A benchmark of six English conversation styles finds that end-of-turn recall stays largely stable, while interruption false positives vary by style and no system combines speed, selectivity and high recall.
A Mandarin speech evaluation reported lower error in selected homophone and HumourPhone comparisons, but results varied by model and full homophone recovery remained limited.
CSAVocoder reported higher spatial scores than listed comparisons and met the paper's real-time criterion, but its audio-quality measures and ratings were mixed.
A simulated benchmark reports that a downscaled acoustic echo-control model had about the same overall performance as a medium version while using 70% fewer parameters and 14% of its computational complexity.
VoiceMem led several memory benchmarks and used fewer memory tokens than a comparison system at K=5, but the evidence comes from benchmark and engineering tests.
A graph-based AI system for long-form meeting audio scored 65.3% on AMI inferential questions, compared with 24.6% for text retrieval. The result comes from a constructed benchmark and automated and human evaluations reported in an arXiv preprint.
A computational study reports that ADEPS, designed for arbitrary microphone layouts, led the listed methods on SISDR across tested arrays but did not top every spectral, coherence or binaural-cue measure.
A frequency-aware stereo autoencoder led on five of seven reconstruction measures and posted higher point estimates across 12 automatic generation metrics, while the study presents the results as relative rankings.
The preprint reports that NAPE, a self-supervised audio method using causal next-patch prediction, tied SSLAM on AudioSet-2M, came close on AudioSet-20K and ESC-50, and scored higher than the strongest listed baseline on IEMOCAP.
A computational benchmark reports that Fish handled both track and version identification in the tested setting, while results differed across altered audio and segment lengths.
A methods preprint suggests that strong public benchmark scores can coincide with reproducing reference wording when audio does not uniquely support it.
A cross-dataset evaluation found weaker external performance for cough-based TB models, while a clinical-variable baseline transferred more consistently.
A structured model for identifying the source of synthetic speech performed almost perfectly on the ASVspoof2019-attr-17 benchmark, but the result is limited to known generators and the study’s tested conditions.
An arXiv preprint proposes TCPα, a post-hoc confidence method for music-information-retrieval models. In tests spanning rāga identification, domain shift and ornamentation detection, it reported stronger failure-prediction performance than comparison targets and improved results when low-confidence predictions were rejected.