CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01songlab /gpn-star-scores GPN-Star genome-wide scores Genome-wide, mutation-rate-calibrated GPN-Star constraint and variant scores for eight score sets covering human, mouse, chicken, D. melanogaster, C. elegans, and A. thaliana. Canonical scores are available as chromosome-sharded Parquet. Hugging Face hosts 72 BigWigs; a multi-assembly UCSC track hub references the 64 logo/LLR views. Overview and quick links Resource Link Files Browse all dataset files Default Dataset… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-star-scores.tabular10B<n<100B4 likes16k downloads2mo agoHugging Face02songlab /TraitGym 🧬 TraitGym Benchmarking DNA Sequence Models for Causal Regulatory Variant Prediction in Human Genetics 🏆 Leaderboard: https://huggingface.co/spaces/songlab/TraitGym-leaderboard ⚡️ Quick start Load a datasetfrom datasets import load_dataset dataset = load_dataset("songlab/TraitGym", "mendelian_traits", split="test") Example notebook to run variant effect prediction with a gLM, runs in 5 min on Google Colab: TraitGym.ipynb 🤗 Resources… See the full description on the dataset page: https://huggingface.co/datasets/songlab/TraitGym.tabular10M<n<100M12 likes15k downloads2y agoHugging Face03songjhPKU /PM4Bench PM4Bench Strictly parallel multilingual evaluation for Large Vision-Language Models Overview The paper Benchmarking and Boosting Multilingual Capabilities of LVLMs via OCR-Centric Reinforcement Learning introduces PM4Bench to separate language effects from dataset variation. Its content is strictly parallel across ten languages, and its vision setting renders textual inputs directly into images. Comparing that setting with interleaved input identifies OCR… See the full description on the dataset page: https://huggingface.co/datasets/songjhPKU/PM4Bench.imagevisual-question-answering10K<n<100K2 likes4.3k downloads1mo agoHugging Face04Dr3dre /Genius-song-lyrics-cleaned 🎵 Genius Song Lyrics cleaned Dataset Dataset Description This dataset is originally taken from Genius Song Lyrics and it contains cleaned and normalized song lyrics for more than 5 million songs, designed for large-scale topic modeling, clustering, and semantic analysis. The dataset was specifically preprocessed to be compatible with embedding-based models (e.g. Sentence Transformers, BERTopic) while preserving lyrical meaning and thematic content. Repetitive structures… See the full description on the dataset page: https://huggingface.co/datasets/Dr3dre/Genius-song-lyrics-cleaned.tabulartext-classification1M<n<10M5 likes4.1k downloads9mo agoHugging Face05ASLP-lab /SongEval SongEval 🎵 A Large-Scale Benchmark Dataset for Aesthetic Evaluation of Complete Songs 📖 Overview SongEval is the first open-source, large-scale benchmark dataset designed for aesthetic evaluation of complete songs. It provides over 2,399 songs (~140 hours) annotated by 16 expert raters across five perceptual dimensions. The dataset enables research in evaluating and improving music generation systems from a human aesthetic perspective. 🌟 Features… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/SongEval.audio1K<n<10K24 likes2.7k downloads1y agoHugging Face06sebastiandizon /genius-song-lyricstabular1M<n<10M38 likes1.2k downloads3y agoHugging Face07songlab /ldsc S-LDSC The dataset displayed here (test.parquet) represents the ~10M variants used for S-LDSC in hg38 coordinates (a tiny fraction we couldn't liftover are marked with pos = -1). We include scores for the 3 GPN-Star models (mutation-rate adjusted minus entropy, higher -> more functional). tabular1M<n<10M0 likes1.1k downloads2mo agoHugging Face08songlab /gpn-animal-promoter-datasettext1M<n<10M1 likes1k downloads2y agoHugging Face09ujin-song /pexels-image-60kimage10K<n<100K2 likes937 downloads1y agoHugging Face10songbirdini /v33da Who Called? V33DA: A Physically Verified Multimodal Benchmark for Vocal Attribution in Zebra Finch Groups Task. Given a detected zebra finch vocalization and the set of birds visible at that moment, determine which bird produced the call. Caller identity is verified physically from on-body accelerometer vibration; the accelerometer channel is withheld from benchmark models and used only by an oracle ceiling. V33DA provides 33,625 verified vocalization events from 10 individually… See the full description on the dataset page: https://huggingface.co/datasets/songbirdini/v33da.audioaudio-classification10K<n<100K0 likes925 downloads5mo agoHugging Face11xu-song /cc100-samplesThe cc100-samples is a subset which contains first 10,000 lines of cc100. Languages To load a language which isn't part of the config, all you need to do is specify the language code in the config. You can find the valid languages in Homepage section of Dataset Description: https://data.statmt.org/cc-100/ E.g. dataset = load_dataset("cc100-samples", lang="en") VALID_CODES = [ "am", "ar", "as", "az", "be", "bg", "bn", "bn_rom", "br", "bs", "ca", "cs", "cy", "da", "de", "el"… See the full description on the dataset page: https://huggingface.co/datasets/xu-song/cc100-samples.texttext-generation1M<n<10M6 likes686 downloads2y agoHugging Face12songfuzhen /imagebedimage1K<n<10K0 likes665 downloads19d agoHugging Face13songlab /hg38_cactus447waytext10M<n<100M0 likes633 downloads1y agoHugging Face14songlab /omim_traitgym OMIM regulatory variants Predictions from all models tabular1K<n<10K0 likes613 downloads10mo agoHugging Face15anantg /genius-song-lyricstext1M<n<10M0 likes474 downloads2y agoHugging Face16ASLP-lab /SongFormBench SongFormBench 🏆 [English | 中文] A High-Quality Benchmark for Music Structure Analysis Chunbo Hao1*, Ruibin Yuan2,6*, Jixun Yao1, Qixin Deng3,6,Xinyi Bai4,6, Yanbo Wang5, Wei Xue2, Lei Xie1† *Equal contribution    †Corresponding author 1Audio, Speech and Language Processing Group (ASLP@NPU),School of Computer Science, Northwestern Polytechnical University 2Hong Kong University of Science and Technology 3Northwestern University… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/SongFormBench.audio1K<n<10K3 likes451 downloads5mo agoHugging Face17nguyenvulebinh /song_datasetgatedaudio1K<n<10K9 likes413 downloads4y agoHugging Face18vishnupriyavr /spotify-million-song-dataset Dataset Card for Spotify Million Song Dataset Dataset Summary This is Spotify Million Song Dataset. This dataset contains song names, artists names, link to the song and lyrics. This dataset can be used for recommending songs, classifying or clustering songs. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data… See the full description on the dataset page: https://huggingface.co/datasets/vishnupriyavr/spotify-million-song-dataset.text10K<n<100K15 likes377 downloads3y agoHugging Face19RichardErkhov /The_Million_Song_Datasetimage1M<n<10M0 likes322 downloads2y agoHugging Face20mlshariki /text-to-song2audio1K<n<10K0 likes321 downloads2y agoHugging Face21songlab /ukb_finemapped_nc_traitgym UKBB finemapped non-coding variants Predictions from all models tabular10K<n<100K1 likes308 downloads10mo agoHugging Face22songtingyu /IFIR Dataset Card for IFIR Benchmark Repository: sighingsnow/IFIR For the usage of this dataset, please refer to the github repo. If you find this repository helpful, feel free to cite our paper: @misc{song2025ifir, title={IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval}, author={Tingyu Song and Guo Gan and Mingsheng Shang and Yilun Zhao}, year={2025}, eprint={2503.04644}, archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/songtingyu/IFIR.texttext-retrieval1K<n<10K2 likes284 downloads2y agoHugging Face23vava22684 /song-jury-leaderboardtextn<1K2 likes262 downloads2d agoHugging Face24renumics /song-describer-datasetThis is a mirror to the example dataset "The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation" paper by Manco et al. Project page on Github: https://github.com/mulab-mir/song-describer-dataset Dataset on Zenodoo: https://zenodo.org/records/10072001 Explore the dataset on your local machine: import datasets from renumics import spotlight ds = datasets.load_dataset('renumics/song-describer-dataset') spotlight.show(ds) audion<1K11 likes256 downloads3y agoHugging Face25code-rider /spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d tabular1K<n<10K2 likes251 downloads9mo agoHugging Face26songlab /gpn-msa-sapiens-dataset Training windows for GPN-MSA-Sapiens For more information check out our paper and repository. Path in Snakemake: results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001 tabular1M<n<10M0 likes243 downloads2y agoHugging Face27ricdomolm /lawma-instructions_llama3_8k_songertext100K<n<1M1 likes234 downloads2y agoHugging Face28Sourabh2 /Songsaudion<1K1 likes222 downloads2y agoHugging Face29SongzeLi /SID-VLN Datasets of Learning Goal-Oriented Language-Guided Navigation with Self-Improving Demonstrations at Scale. tabular1K<n<10K0 likes201 downloads1y agoHugging Face30songlab /ukb_finemapped_coding UKBB finemapped coding variants Predictions from all models tabular1K<n<10K1 likes184 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.