song
Datasets
All datasets matching “song”gpn-star-scores
GPN-Star genome-wide scores
Genome-wide, mutation-rate-calibrated GPN-Star constraint and variant scores for
eight score sets covering human, mouse, chicken, D. melanogaster,
C. elegans, and A. thaliana. Canonical scores are available as
chromosome-sharded Parquet. Hugging Face hosts 72 BigWigs; a multi-assembly
UCSC track hub references the 64 logo/LLR views.
Overview and quick links
Resource
Link
Files
Browse all dataset files
Default Dataset… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-star-scores.TraitGym
🧬 TraitGym
Benchmarking DNA Sequence Models for Causal Regulatory Variant Prediction in Human Genetics
🏆 Leaderboard: https://huggingface.co/spaces/songlab/TraitGym-leaderboard
⚡️ Quick start
Load a datasetfrom datasets import load_dataset
dataset = load_dataset("songlab/TraitGym", "mendelian_traits", split="test")
Example notebook to run variant effect prediction with a gLM, runs in 5 min on Google Colab: TraitGym.ipynb
🤗 Resources… See the full description on the dataset page: https://huggingface.co/datasets/songlab/TraitGym.imagenet_sketchImageNet-Sketch data set consists of 50000 images, 50 images for each of the 1000 ImageNet classes.
We construct the data set with Google Image queries "sketch of __", where __ is the standard class name.
We only search within the "black and white" color scheme. We initially query 100 images for every class,
and then manually clean the pulled images by deleting the irrelevant images and images that are for similar
but different classes. For some classes, there are less than 50 images after manually cleaning, and then we
augment the data set by flipping and rotating the images.SongFormDB
SongFormDB 🎵
[English | 中文]
A Large-Scale Multilingual Music Structure Analysis Dataset for Training SongFormer 🚀
Chunbo Hao1*, Ruibin Yuan2,6*, Jixun Yao1, Qixin Deng3,6,Xinyi Bai4,6, Yanbo Wang5, Wei Xue2, Lei Xie1†
*Equal contribution †Corresponding author
1Audio, Speech and Language Processing Group (ASLP@NPU),School of Computer Science, Northwestern Polytechnical University
2Hong Kong University of Science and Technology… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/SongFormDB.PM4Bench
PM4Bench
Strictly parallel multilingual evaluation for Large Vision-Language Models
Overview
The paper Benchmarking and Boosting Multilingual Capabilities of LVLMs via
OCR-Centric Reinforcement Learning
introduces PM4Bench to separate language effects from dataset variation. Its
content is strictly parallel across ten languages, and its vision setting
renders textual inputs directly into images. Comparing that setting with
interleaved input identifies OCR… See the full description on the dataset page: https://huggingface.co/datasets/songjhPKU/PM4Bench.Genius-song-lyrics-cleaned
🎵 Genius Song Lyrics cleaned Dataset
Dataset Description
This dataset is originally taken from Genius Song Lyrics and it contains cleaned and normalized song lyrics for more than 5 million songs, designed for large-scale topic modeling, clustering, and semantic analysis.
The dataset was specifically preprocessed to be compatible with embedding-based models (e.g. Sentence Transformers, BERTopic) while preserving lyrical meaning and thematic content.
Repetitive structures… See the full description on the dataset page: https://huggingface.co/datasets/Dr3dre/Genius-song-lyrics-cleaned.
