Leaks from inside AI music platform Suno have exposed, in granular detail, how the company assembled a vast training library from major streaming and media platforms-including Deezer, YouTube, and Pond5-feeding thousands of hours of audio into its generative model.
According to internal material obtained in a recent hack, Suno’s systems relied on explicit scraping instructions and automated pipelines to ingest and categorize music and audio from these services throughout 2023 and 2024. These documents show not just what kinds of files were collected, but also how they were processed, tagged, and funneled into the company’s core AI training infrastructure.
The breach, carried out by an attacker who claims to have used custom malware dubbed the “Shai-Hulud worm,” gave outsiders a rare view into the mechanics of a commercial AI music engine. Suno’s flagship product allows anyone to type a brief text description-such as a genre, mood, or era-and receive a complete track with vocals and instrumentation in seconds. To achieve that level of sophistication, the company needed what the logs depict as a massive, highly organized dataset of existing recordings.
The leaked files back up what many in the music industry have been alleging in legal filings since 2024: that leading AI music firms quietly leaned on large troves of copyrighted material, drawn from major streaming platforms and online content libraries, to teach their systems how to mimic human-created songs. Suno had previously avoided disclosing specific sources, framing its training process in only general terms. These internal documents, however, enumerate concrete sources and outline the technical steps used to extract and transform the data.
At the heart of the disclosure are detailed scraping configurations. These scripts and job files appear to instruct Suno’s infrastructure on how to query platforms like Deezer and YouTube, what metadata to capture, how to handle playlists and channels, and how to normalize the resulting audio. For Pond5-a marketplace known for stock music and sound effects-the logs describe large, systematic downloads of files, followed by automated labeling of tempo, genre, and mood. In several cases, filenames and IDs strongly resemble commercially released tracks or catalogued library content that would ordinarily be licensed, not silently copied.
The documents also reveal that Suno tracked and stored metadata on an industrial scale. Titles, artist names, track lengths, genre tags, and even language labels were harvested and integrated into an internal database. That information is crucial for training a system that responds so accurately to user prompts like “90s alternative rock ballad” or “orchestral soundtrack with suspenseful build.” The more detailed the metadata, the easier it is for a model to understand stylistic patterns and associate them with textual descriptions.
While many AI companies argue that ingesting publicly accessible content for training constitutes fair use, the specificity of these logs is likely to fuel criticism from musicians, labels, and rights holders. The presence of data from subscription streaming platforms and stock media libraries raises direct questions about whether Suno had the right to acquire or repurpose that material in the first place, particularly for a commercial product that can generate music resembling the work in its training set.
Suno itself has presented its system as a tool for creativity and rapid prototyping, helping artists, hobbyists, and brands generate original tracks without needing a studio or band. The leak, however, will intensify a growing debate: if AI models are built on the backs of unlicensed recordings, do the resulting outputs undermine the livelihoods of the very creators whose work made the technology possible?
The technical notes in the leaked code also shed light on how Suno tried to avoid obvious one-to-one copying. Processing pipelines include steps like audio resampling, segmentation into small chunks, and feature extraction, suggesting the model was trained not on direct waveform duplication but on statistical patterns derived from many tracks. From a legal standpoint, though, that distinction may not be enough. Rights holders increasingly argue that using full recordings for training-no matter how they are transformed-requires consent or compensation.
For artists, the revelations reinforce worries that their catalogues may have been swept into datasets without permission. Deezer and YouTube host vast libraries of both independent and major-label music; even if Suno did not target specific acts, its large-scale harvesting almost certainly captured songs from well-known performers and underground musicians alike. Pond5 contributors, many of whom rely on licensing fees for income, may be especially alarmed that their work could now be indirectly powering an AI competitor that can generate “stock-style” cues on demand.
Regulators and courts are already grappling with these questions across the broader AI sector. Text and image models have faced similar backlash for training on news articles, books, artworks, and photographs scraped from the web. What distinguishes the Suno case is the level of internal documentation-names of services, timestamps of scraping runs, and concrete technical parameters-laying bare the mechanics of data acquisition rather than leaving them as abstract legal arguments.
From an industry standpoint, the Suno leak underlines a structural tension: building cutting-edge generative music systems practically demands exposure to huge, diverse audio corpora, yet the rights landscape around those corpora is fragmented and heavily protected. Negotiating comprehensive licenses with all relevant rights holders is slow, expensive, and sometimes impossible; scraping, by contrast, is fast and scalable, even if it sits in a legal and ethical gray zone.
The fallout could accelerate a shift toward more transparent, licensed datasets for music AI. Rights organizations and labels may push for explicit opt-in frameworks, where AI companies pay to access defined catalogues and log exactly which tracks are used. Some creators may welcome such deals as a new revenue stream; others may choose to withhold their work entirely. Without clearer norms, however, leaks like this are likely to keep exposing the gap between public marketing and private practice in AI development.
For users of Suno and similar platforms, the revelations raise practical questions as well. If a model’s training set is found to include unlicensed material, what does that mean for the legal status of the generated songs? Could brands or artists who release AI-assisted tracks face infringement claims if an output is deemed too close to a protected recording? Today, such cases are rare, but the more details emerge about how training data was sourced, the more likely it becomes that someone will test those boundaries in court.
At the same time, creators experimenting with AI may start demanding tools that guarantee ethically sourced training. Musicians might prefer platforms built on licensed catalogues or datasets assembled in collaboration with independent artists, where contributors are paid or credited. That kind of model would be more expensive to build, but it could offer a clearer moral and legal foundation-and may become a competitive advantage as public scrutiny grows.
The Suno incident also highlights a basic security risk: if an attacker can exfiltrate sensitive source code and internal logs, they can not only reveal trade secrets but also expose a company’s legal vulnerabilities. AI startups, already under pressure to move fast, may now have to invest more heavily in cybersecurity, knowing that a single breach can turn opaque training practices into a public relations and regulatory crisis.
Over the longer term, the music ecosystem faces a pivotal choice. One path is a prolonged clash between AI developers and rights holders, with lawsuits, leaked documents, and adversarial negotiations defining the landscape. The other is a negotiated framework in which training on commercial music is treated as a licensable use, with revenue-sharing mechanisms and opt-out rights. The Suno leak does not decide that question, but it pushes it to the forefront by stripping away the ambiguity around where at least one prominent model got its raw material.
For now, the leaked code and logs function as a case study in how a leading AI music engine is built behind the scenes: automated scrapers targeting services like Deezer, YouTube, and Pond5; massive ingestion of audio and metadata; feature extraction and labeling; and, finally, model training that yields strikingly polished tracks from a single text prompt. Technically impressive, the system’s foundations are also legally and ethically contested-an uneasy mix that the industry will increasingly be forced to confront in the open.
