Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Phillip Long

Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio

Mar 09, 2026

Phillip Long, Zachary Novack, Chris Donahue

Abstract:Autoregressive "language" models (LMs) trained on raw waveforms can be repurposed for lossless audio compression, but prior work is limited to 8-bit audio, leaving open whether such approaches work for practical settings (16/24-bit) and can compete with existing codecs. We benchmark LM-based compression on full-fidelity audio across diverse domains (music, speech, bioacoustics), sampling rates (16kHz-48kHz), and bit depths (8, 16, 24-bit). Standard sample-level tokenization becomes intractable at higher bit depths due to vocabulary size (65K for 16-bit; 16.7M for 24-bit). We propose Trilobyte, a byte-level tokenization schema for full resolution audio, improving vocabulary scaling from $O(2^{b})$ to $O(1)$ and enabling the first tractable 24-bit LM-based lossless compression. While LMs consistently outperform FLAC and yield state-of-the-art compression at 8-bit and 16-bit, we observe that compression gains become more modest as bit depth increases beyond 8-bit.

* Submitted for review at Interspeech 2026, 7 pages, 5 figures

Via

Access Paper or Ask Questions

PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing

Sep 17, 2024

Phillip Long, Zachary Novack, Taylor Berg-Kirkpatrick, Julian McAuley

Figure 1 for PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing

Figure 2 for PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing

Figure 3 for PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing

Figure 4 for PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing

Abstract:The recent explosion of generative AI-Music systems has raised numerous concerns over data copyright, licensing music from musicians, and the conflict between open-source AI and large prestige companies. Such issues highlight the need for publicly available, copyright-free musical data, in which there is a large shortage, particularly for symbolic music data. To alleviate this issue, we present PDMX: a large-scale open-source dataset of over 250K public domain MusicXML scores collected from the score-sharing forum MuseScore, making it the largest available copyright-free symbolic music dataset to our knowledge. PDMX additionally includes a wealth of both tag and user interaction metadata, allowing us to efficiently analyze the dataset and filter for high quality user-generated scores. Given the additional metadata afforded by our data collection process, we conduct multitrack music generation experiments evaluating how different representative subsets of PDMX lead to different behaviors in downstream models, and how user-rating statistics can be used as an effective measure of data quality. Examples can be found at https://pnlong.github.io/PDMX.demo/.

Via

Access Paper or Ask Questions