238 GB
2,282 files
Updated 1 day ago
README.md

Suno AI Music Dataset cover — analog producer studio with modular synths, Rhodes, sitar, tabla, TR-909, big-band brass and a floating genre taxonomy graph

Suno AI Music Dataset (Multi-Genre Curated)

A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research.

This is not a "scrape everything Suno produces" dump. It is a deliberately curated collection biased 70% toward refined musical territory (psytrance progressive, neo-soul, contemporary jazz, jazz-leaning deep house, art rock, simplified Hindustani, etc.) and 30% toward profitable mainstream zeitgeist (corridos tumbados, amapiano, country pop, modern reggaeton) — but where every mainstream track carries a sophistication overlay (altered jazz chords, big-band stabs, forest-psytrance sub-bass, J Dilla-humanized drums). The curatorial bias itself is part of what this dataset offers researchers.

What is this

This dataset is a curated corpus of original music tracks generated with Suno's V5.5 model family (specifically chirp-fenix) using a Pro-tier account during May 2026. It contains approximately 200-300 final tracks (growing) spanning 100+ sub-sub-genres organized in a three-level taxonomy. Every track is paired with exhaustive metadata: the exact text prompt used for generation, BPM range, typical key, mood descriptors, taxonomy tier, intended "radio destination", and provenance fields traceable back to the Suno API.

The dataset is intended for:

  • Research on audio classification across fine-grained genre boundaries
  • Text-to-music conditioning experiments (prompt-to-audio alignment studies)
  • Music recommendation and similarity research
  • Benchmarks for music tagging and auto-DJ systems
  • Studies of AI music generator capabilities and failure modes
  • Training material for derivative music-generation models (subject to the license terms below)

Tracks are stored as MP3 files alongside JPG cover art in a flat tracks/ directory, with all metadata centralized in a single CSV (biblioteca_master.csv) keyed by Suno's UUID.

Curation Philosophy

Most publicly available AI music datasets are either (a) opportunistic scrapes of whatever a generator produces with no curatorial filter, or (b) tightly scoped single-genre collections that lack breadth. This dataset takes a third path: a deliberate 70/30 split between musical quality and commercial relevance, with sophistication injected into the commercial side.

The 70% quality core

Seventy percent of the dataset targets genres where production sophistication, harmonic depth, rhythmic complexity, or timbral richness matter more than streaming appeal. Representative territory includes:

  • Progressive psytrance and forest psytrance (modal sub-bass, polyrhythmic percussion)
  • Neo-soul and contemporary jazz fusion (extended chords, swung subdivisions, Rhodes/Wurlitzer textures)
  • Deep house with jazz harmony (Bb minor dorian, walking bass, muted brass)
  • Art rock and progressive rock (asymmetric phrasing, layered guitars)
  • Simplified Hindustani and other world fusions (raga-based melody over Western harmony)
  • West London broken beat (4-Hero / Bugz in the Attic lineage)
  • Future garage, liquid drum & bass, deep dub techno

The 30% mainstream layer

Thirty percent targets genres that drive streaming traffic in 2026: corridos tumbados, amapiano, country pop, modern reggaeton, Afrobeats, trap latino. But here is the key constraint: each mainstream track receives a sophistication overlay. A corrido tumbado track might carry altered dominant chords on the requinto. An amapiano log-drum sequence is paired with sub-bass shaped like a forest-psy patch. A country pop track gets J Dilla-style humanized drum displacement. The result is tracks that remain instantly recognizable to their target audience while presenting greater harmonic and rhythmic content for analysis.

Why this matters for ML

For audio classifiers, this means the dataset stress-tests the boundary between sub-sub-genres that are sonically adjacent but commercially distinct. For text-to-music systems, the prompts in this dataset are unusually descriptive (typically 30-80 tokens, citing era references, key centers, instrumental layers, and aesthetic adjectives) and provide a benchmark for prompt-fidelity evaluation. For recommendation systems, the explicit tax_radio_destino field gives a ground-truth listener-context label that synthetic datasets typically lack.

Schema

The master CSV (biblioteca_master.csv) contains 32 columns. All Suno-native fields preserve their original names; all curatorial fields use the tax_ prefix.

Column Description
id Suno-issued UUID for the track. Primary key.
title Track title as registered in the Suno library.
hoja_id_pipeline Internal pipeline identifier linking the track to its generation request in the source spreadsheet.
major_model_version Suno model family (e.g. v5.5).
model_name Specific model variant used for generation (e.g. chirp-fenix).
status Suno-side processing status (submitted, complete, etc.).
duration Track duration in seconds, as reported by Suno.
is_instrumental Boolean. True if the track is purely instrumental.
tags Free-form tag string used at generation time. Often overlaps with the prompt.
gpt_description_prompt The descriptive prompt provided to Suno for the instrumental/structural side.
prompt_letras The lyrics prompt (when vocals are requested). Empty for instrumentals.
created_at ISO 8601 timestamp of track creation.
play_count Public play count on the Suno platform.
upvote_count Public upvote count on the Suno platform.
is_public Boolean. Whether the track is published on Suno.
is_liked Boolean. Liked flag in the source account.
is_trashed Boolean. Removed-from-library flag.
has_hook Boolean. Suno-side flag indicating an identifiable hook.
has_stem Boolean. Whether stems are available on the Suno side.
tax_genero Top-level genre in the curatorial taxonomy (Electrónica, Hip-hop, Latin, Jazz, etc.).
tax_subgenero Second-level subgenre (e.g. "Broken beat" under Electrónica).
tax_sub_sub_genero Third-level sub-sub-genre (e.g. "Broken beat West London").
tax_tier Curatorial tier: 70_calidad (quality core) or 30_mainstream (commercial layer).
tax_prioridad_kukito Curator priority score (1-5). Higher means more strategically important to the collection.
tax_suno_capability Suno-side capability score for the genre (1-5). Reflects how well Suno V5.5 actually renders this style.
tax_radio_destino Intended thematic radio context (e.g. "Deep House Jazzero Radio"). Useful as a ground-truth listener-context label.
tax_bpm_rango Approximate BPM range (e.g. 125-135).
tax_key_tipico Typical musical key(s) and modal context.
tax_mood Comma-separated mood descriptors (e.g. sophisticated, polyrhythmic, jazzy, futuristic, warm).
audio_url Suno CDN URL for the audio (may expire).
video_url Suno CDN URL for the generated video, when available.
image_url Suno CDN URL for the standard cover image.
image_large_url Suno CDN URL for the large cover image.
mp3_local Relative path to the local MP3 file (tracks/<uuid>.mp3).
cover_local Relative path to the local cover image (tracks/<uuid>.cover.jpg).

Taxonomy

The dataset is organized under a three-level taxonomy: root genre, subgenre, and sub-sub-genre. Coverage by root genre, in approximate descending order of sub-sub-genre count:

Root genre Sub-sub-genres covered
Electrónica ~36
Hip-hop ~10
Latin ~8
Jazz ~7
World ~6
Rock ~6
Ambient ~5
Pop ~5
Reggae ~4
Clásica ~3

Electronic music is over-represented by design: the genre allows Suno V5.5 to exercise its strongest capabilities (timbre design, rhythmic micro-variation, sub-bass synthesis) and provides the cleanest sub-sub-genre boundaries for classification work. Classical and reggae are under-represented because Suno V5.5 produces less reliable output in those territories; the entries included reflect a quality filter, not a coverage attempt.

The full taxonomy file is included as taxonomia_practica.json and provides the canonical mapping of every sub-sub-genre to its parent subgenre and root genre.

Use Cases

The following use cases motivated the dataset design and are explicitly supported by the metadata schema:

  1. Training fine-grained audio classifiers. The three-level taxonomy supplies clean labels at multiple granularities, from coarse (10 root genres) to fine (100+ sub-sub-genres), enabling hierarchical-classification experiments.
  2. Benchmarking text-to-music systems. Each track ships with the exact descriptive prompt used to generate it. Researchers can compare prompt fidelity across generators by re-rendering the same prompts in alternative systems and scoring the divergence.
  3. Fine-tuning music generation models. The CC-BY-4.0 license permits derivative works, including using the tracks as training data for downstream generators (with attribution).
  4. Music recommendation research. The tax_radio_destino field provides a synthetic but consistent listener-context label that supports recommendation experiments without the cold-start problems of real listener data.
  5. Auto-tagging and music information retrieval. The combination of explicit mood, key, BPM, and instrumentation cues in the prompt fields gives a high-density labelled set for MIR baselines.
  6. Generator capability profiling. The tax_suno_capability field encodes the curator's empirical observation of how well Suno V5.5 renders each sub-sub-genre, providing a starting point for systematic studies of generator strengths and blind spots.
  7. Cross-genre transfer learning. The deliberate 70/30 split makes this dataset useful for studying how representations learned on sophisticated material transfer to mainstream material and vice versa.

Limitations

This dataset reflects what Suno V5.5 can and cannot produce as of May 2026. Researchers should be aware of the following constraints before drawing strong conclusions:

  • Meter bias. Suno V5.5 defaults aggressively to 4/4. Prompts requesting 7/8, 5/4, or other asymmetric meters typically yield 4/4 output with surface-level rhythmic complexity. The dataset does not contain authentic odd-meter material.
  • Tonal bias. The model is biased toward common-practice tonality (major/minor with conventional cadences). Atonal, serial, microtonal, or post-Stravinsky idioms are not represented; prompts in those directions produce stylistically reinterpreted, tonally-grounded output.
  • Copyright filter. Suno's safety layer blocks named-artist prompts. Tracks were generated using descriptive language only ("West London broken beat, 4-Hero lineage" rather than artist names). Style transfer in the strict sense is not possible.
  • Duration skew. Most tracks cluster around the 2-minute mark, with longer outputs being less consistent. The dataset is not suitable for long-form structural analysis (form, large-scale modulation, multi-section dynamics).
  • Genre genericization. For under-represented or less-trained styles (Hindustani classical, certain regional Latin styles, niche jazz subgenres), Suno produces a stylistically averaged output that may not be ethnomusicologically faithful. The tax_suno_capability score flags these cases.
  • English-language bias in vocals. Vocal tracks produce more consistent results in English than in Spanish or other languages, despite explicit prompting. Spanish-language vocals occasionally exhibit unnatural phoneme stress.
  • No stems. The dataset ships only mixed audio. Stems are not provided.
  • Cover art is auto-generated. Cover images are produced by Suno's image model from the track prompt and should not be treated as a curated image dataset.

License and Ethics

The dataset is released under the Creative Commons Attribution 4.0 International (CC-BY-4.0) license. Tracks were generated under a Suno Pro plan, which grants commercial-use rights to outputs per Suno's Terms of Service in force at the time of generation (May 2026). Users are free to share, remix, and build upon the material — including for commercial purposes — provided that attribution is given.

Additional ethical guidelines for users of this dataset:

  • AI-generated disclosure is mandatory. Any derivative work (music release, training dataset, demo, paper) that incorporates these tracks or models trained on them must disclose the AI origin. This aligns with the disclosure requirements being adopted by major DSPs (Spotify, Deezer, YouTube Music) and by EU AI Act provisions.
  • No artist impersonation. Tracks in this dataset were generated without naming any real artist in the prompts, and the dataset must not be used to fine-tune systems intended to impersonate specific living or deceased artists. This is both a legal and an ethical line.
  • No deceptive distribution. Derivative releases must not be presented as human-performed when they are not, and must comply with the metadata-tagging requirements of the platform on which they are distributed.
  • Respect for cultural origin. Several sub-sub-genres in this dataset originate in specific cultural traditions (West London broken beat, Hindustani melody, amapiano, corridos tumbados). Researchers and creators are asked to respect those traditions in any derivative work and to credit the communities of origin where appropriate.

Citation

If you use this dataset in academic work, please cite it as:

@dataset{suno_curated_2026,
  title  = {Suno AI Music Dataset (Multi-Genre Curated)},
  author = {Kukito},
  year   = {2026},
  month  = {May},
  publisher = {Hugging Face},
  version = {1.0},
  url    = {https://huggingface.co/datasets/Kukedlc/suno-curated-multigenre},
  note   = {Generated with Suno V5.5 (chirp-fenix), licensed under CC-BY-4.0}
}

Acknowledgments

This dataset would not exist without the work of the Suno team, whose V5.5 model made it practical to explore 100+ sub-sub-genres at production quality in a single curation pass. Thanks also to the broader Hugging Face open-source community for maintaining the infrastructure that makes datasets like this distributable, discoverable, and reusable. The taxonomy design draws on decades of writing by music journalists, label A&Rs, and DJ-mix curators whose careful sub-genre vocabularies made the labelling task tractable.

Total size
238 GB
Files
2,282
Last updated
Oct 4
Pre-warmed CDN
US EU US EU

Contributors