Suno source code leak reveals the scale of pirated data collection for AI training

A tectonic shift has occurred in the world of generative music. As a result of a hack on the infrastructure of Suno, one of the leaders in AI music, fragments of source code were leaked into the public domain, which irrefutably confirm long-standing suspicions of music labels. A hacker known by the pseudonym ellie.191 managed to infiltrate the company's systems by infecting an employee's computer with a dangerous worm called Shai-Hulud, which spread through compromised npm packages and extracted all access keys, tokens, and credentials.
The most interesting part is not the hack itself, but what was discovered in the stolen data. These are not just logs, but a full-fledged manifesto of how exactly Suno built its training datasets. One of the files contains a download log of over 2 million music clips from YouTube Music, equivalent to 113,879 hours of continuous audio. A second dataset, labeled ytm_tagged, includes another 152,162 hours of tagged recordings from the same YouTube. To bypass platform restrictions, Suno developers actively used Bright Data proxy services, masking their requests as traffic from regular users.
However, YouTube is just the tip of the iceberg. The code contains clear indications of data collection from Pond5 (62,117 hours), Deezer (12,287 hours), Genius (17,615 hours of lyrics and metadata), the International Music Score Library Project (19,514 hours), Jamendo (3,726 hours), Freesound (410 hours), and Musescore (103 hours). A separate plan involved downloading nearly a million hours of podcasts from 420,000 RSS feeds. Whether the company implemented this plan remains unclear, but the sheer scale of such planning raises questions.
Direct Hit in Lawsuits
This leak is not just a technical curiosity. It is direct evidence in favor of the labels that have already sued Suno. Universal Music Group and Sony Music Entertainment are demanding that 61,026 recordings be added to the lawsuit, which they claim were found in Suno's training data using audio fingerprinting technology. The labels argue that Suno violates Section 1201 of the DMCA by bypassing YouTube's technical protection measures using utilities like YT-DLP. The leaked code, demonstrating mass downloading from YouTube through proxies, perfectly supports this argument, although it does not 100% prove the circumvention of encryption.
Significantly, Suno previously acknowledged that its datasets might contain copyrighted works but insisted on "transformative" use under the fair use doctrine. Now, after this leak, the company's arguments look extremely shaky. Warner Music Group, unlike its competitors, has already settled and entered into a licensing agreement, which could set a precedent for the entire industry.
My Comment as an Analyst: This leak is a "moment of truth" for the entire generative AI music industry. Suno and similar services have long operated in a "gray zone," claiming their models are trained on "publicly available" data. Now we see that this was a systematic, planned, and technically equipped data theft. The market should prepare for the fact that the "steal first, license later" model will no longer work. Investors should reconsider the risks of investing in startups that cannot clearly prove the legality of their training datasets.