218 GB
2,270 files
Updated about 16 hours ago
Name
Size
batch_0
batch_1
batch_10
batch_11
batch_12
batch_13
batch_14
batch_15
batch_16
batch_17
batch_18
batch_19
batch_2
batch_20
batch_21
batch_22
batch_23
batch_24
batch_25
batch_26
batch_27
batch_28
batch_29
batch_3
batch_30
batch_31
batch_32
batch_33
batch_34
batch_35
batch_36
batch_37
batch_38
batch_39
batch_4
batch_40
batch_41
batch_42
batch_43
batch_44
batch_45
batch_46
batch_47
batch_48
batch_49
batch_5
batch_6
batch_7
batch_8
batch_9
.gitattributes2.46 kB
xet
README.md2.86 kB
xet
progress.json869 Bytes
xet
README.md

🎵 Suno Audio Dataset

A comprehensive dataset of 49,698 AI-generated music tracks from Suno, organized in 50 batches of 1000 samples each.

🎧 All audio files are playable directly in the dataset viewer!

Dataset Structure

The dataset is organized into batches (batch_0, batch_1, etc.), each containing up to 1000 audio samples with metadata.

Fields

  • audio: 🎵 Playable MP3 audio file (click to play in viewer!)
  • id: Unique track identifier
  • title: Song title
  • display_name: Creator/artist name
  • handle: Creator handle
  • tags: Music tags, genres, and styles
  • prompt: Text prompt used for generation
  • duration: Track duration in seconds
  • play_count: Number of plays on Suno
  • upvote_count: Community upvotes
  • model_name: Suno model version used
  • created_at: Creation timestamp
  • status: Track status
  • is_public: Public visibility flag

Usage

Load Entire Dataset

from datasets import load_dataset

# Load all batches
dataset = load_dataset("Humair332/suno-audio")
print(f"Total tracks: {len(dataset['train'])}")

Load Specific Batch

# Load only batch 0
dataset = load_dataset("Humair332/suno-audio", data_dir="batch_0")

Play Audio

# Get audio data
audio_data = dataset['train'][0]['audio']
audio_array = audio_data['array']
sampling_rate = audio_data['sampling_rate']

# Play in Jupyter/Colab
from IPython.display import Audio
Audio(audio_array, rate=sampling_rate)

Filter by Tags

# Filter by genre
rock_songs = dataset['train'].filter(lambda x: 'rock' in x['tags'].lower())
print(f"Found {len(rock_songs)} rock songs")

Most Popular Tracks

# Sort by play count
from datasets import Dataset
df = dataset['train'].to_pandas()
top_tracks = df.nlargest(10, 'play_count')[['title', 'display_name', 'play_count']]
print(top_tracks)

Dataset Statistics

  • Total Tracks: 49,698
  • Batches: 50
  • Batch Size: 1000
  • Format: Apache Arrow with embedded MP3 audio
  • Audio Format: MP3
  • Metadata: Tags, prompts, engagement metrics

Explore the Music! 🎶

Click on the dataset viewer above and browse through the tracks. Click any row to play the audio directly in your browser!

Source

Original dataset: nyuuzyou/suno

License

MIT License

Citation

If you use this dataset, please cite:

@dataset{suno_audio_dataset,
  title={Suno Audio Dataset},
  author={Humair332},
  year={2026},
  publisher={Hugging Face},
  url={https://huggingface.co/datasets/Humair332/suno-audio}
}
Total size
218 GB
Files
2,270
Last updated
Oct 3
Pre-warmed CDN
US EU US EU

Contributors