datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpn-star-scores
GPN-Star genome-wide scores
Genome-wide, mutation-rate-calibrated GPN-Star constraint and variant scores for
eight score sets covering human, mouse, chicken, D. melanogaster,
C. elegans, and A. thaliana. Canonical scores are available as
chromosome-sharded Parquet. Hugging Face hosts 72 BigWigs; a multi-assembly
UCSC track hub references the 64 logo/LLR views.
Overview and quick links
Resource
Link
Files
Browse all dataset files
Default Dataset… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-star-scores.TraitGym
🧬 TraitGym
Benchmarking DNA Sequence Models for Causal Regulatory Variant Prediction in Human Genetics
🏆 Leaderboard: https://huggingface.co/spaces/songlab/TraitGym-leaderboard
⚡️ Quick start
Load a datasetfrom datasets import load_dataset
dataset = load_dataset("songlab/TraitGym", "mendelian_traits", split="test")
Example notebook to run variant effect prediction with a gLM, runs in 5 min on Google Colab: TraitGym.ipynb
🤗 Resources… See the full description on the dataset page: https://huggingface.co/datasets/songlab/TraitGym.PM4Bench
PM4Bench
Strictly parallel multilingual evaluation for Large Vision-Language Models
Overview
The paper Benchmarking and Boosting Multilingual Capabilities of LVLMs via
OCR-Centric Reinforcement Learning
introduces PM4Bench to separate language effects from dataset variation. Its
content is strictly parallel across ten languages, and its vision setting
renders textual inputs directly into images. Comparing that setting with
interleaved input identifies OCR… See the full description on the dataset page: https://huggingface.co/datasets/songjhPKU/PM4Bench.Genius-song-lyrics-cleaned
🎵 Genius Song Lyrics cleaned Dataset
Dataset Description
This dataset is originally taken from Genius Song Lyrics and it contains cleaned and normalized song lyrics for more than 5 million songs, designed for large-scale topic modeling, clustering, and semantic analysis.
The dataset was specifically preprocessed to be compatible with embedding-based models (e.g. Sentence Transformers, BERTopic) while preserving lyrical meaning and thematic content.
Repetitive structures… See the full description on the dataset page: https://huggingface.co/datasets/Dr3dre/Genius-song-lyrics-cleaned.SongEval
SongEval 🎵
A Large-Scale Benchmark Dataset for Aesthetic Evaluation of Complete Songs
📖 Overview
SongEval is the first open-source, large-scale benchmark dataset designed for aesthetic evaluation of complete songs. It provides over 2,399 songs (~140 hours) annotated by 16 expert raters across five perceptual dimensions. The dataset enables research in evaluating and improving music generation systems from a human aesthetic perspective.
🌟 Features… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/SongEval.genius-song-lyricsldsc
S-LDSC
The dataset displayed here (test.parquet) represents the ~10M variants used for S-LDSC in hg38 coordinates (a tiny fraction we couldn't liftover are marked with pos = -1).
We include scores for the 3 GPN-Star models (mutation-rate adjusted minus entropy, higher -> more functional).
gpn-animal-promoter-datasetpexels-image-60kv33da
Who Called? V33DA: A Physically Verified Multimodal Benchmark for Vocal Attribution in Zebra Finch Groups
Task. Given a detected zebra finch vocalization and the set of birds visible at that moment, determine which bird produced the call. Caller identity is verified physically from on-body accelerometer vibration; the accelerometer channel is withheld from benchmark models and used only by an oracle ceiling.
V33DA provides 33,625 verified vocalization events from 10 individually… See the full description on the dataset page: https://huggingface.co/datasets/songbirdini/v33da.cc100-samplesThe cc100-samples is a subset which contains first 10,000 lines of cc100.
Languages
To load a language which isn't part of the config, all you need to do is specify the language code in the config.
You can find the valid languages in Homepage section of Dataset Description: https://data.statmt.org/cc-100/
E.g.
dataset = load_dataset("cc100-samples", lang="en")
VALID_CODES = [
"am", "ar", "as", "az", "be", "bg", "bn", "bn_rom", "br", "bs", "ca", "cs", "cy", "da", "de",
"el"… See the full description on the dataset page: https://huggingface.co/datasets/xu-song/cc100-samples.imagebedhg38_cactus447wayomim_traitgym
OMIM regulatory variants
Predictions from all models
genius-song-lyricsSongFormBench
SongFormBench 🏆
[English | 中文]
A High-Quality Benchmark for Music Structure Analysis
Chunbo Hao1*, Ruibin Yuan2,6*, Jixun Yao1, Qixin Deng3,6,Xinyi Bai4,6, Yanbo Wang5, Wei Xue2, Lei Xie1†
*Equal contribution †Corresponding author
1Audio, Speech and Language Processing Group (ASLP@NPU),School of Computer Science, Northwestern Polytechnical University
2Hong Kong University of Science and Technology
3Northwestern University… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/SongFormBench.song_datasetspotify-million-song-dataset
Dataset Card for Spotify Million Song Dataset
Dataset Summary
This is Spotify Million Song Dataset. This dataset contains song names, artists names, link to the song and lyrics. This dataset can be used for recommending songs, classifying or clustering songs.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data… See the full description on the dataset page: https://huggingface.co/datasets/vishnupriyavr/spotify-million-song-dataset.The_Million_Song_Datasettext-to-song2ukb_finemapped_nc_traitgym
UKBB finemapped non-coding variants
Predictions from all models
IFIR
Dataset Card for IFIR Benchmark
Repository: sighingsnow/IFIR
For the usage of this dataset, please refer to the github repo.
If you find this repository helpful, feel free to cite our paper:
@misc{song2025ifir,
title={IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval},
author={Tingyu Song and Guo Gan and Mingsheng Shang and Yilun Zhao},
year={2025},
eprint={2503.04644},
archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/songtingyu/IFIR.song-jury-leaderboardsong-describer-datasetThis is a mirror to the example dataset "The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation" paper by Manco et al.
Project page on Github: https://github.com/mulab-mir/song-describer-dataset
Dataset on Zenodoo: https://zenodo.org/records/10072001
Explore the dataset on your local machine:
import datasets
from renumics import spotlight
ds = datasets.load_dataset('renumics/song-describer-dataset')
spotlight.show(ds)
spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html
the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d
gpn-msa-sapiens-dataset
Training windows for GPN-MSA-Sapiens
For more information check out our paper and repository.
Path in Snakemake:
results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001
lawma-instructions_llama3_8k_songerSongsSID-VLN Datasets of Learning Goal-Oriented Language-Guided Navigation with Self-Improving Demonstrations at Scale.
ukb_finemapped_coding
UKBB finemapped coding variants
Predictions from all models
