datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
co2s-datasetsproof-pile-2-streaming
ArXiv | Models | Data | Code | Blog | Sample Explorer
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, Sean Welleck
The Proof-Pile-2 is a 55 billion token dataset of mathematical and scientific documents. This dataset was created in order to train the Llemma 7B and Llemma 34B models. It consists of three subsets:
arxiv (29B tokens): the ArXiv subset of RedPajama
open-web-math (15B tokens): The OpenWebMath… See the full description on the dataset page: https://huggingface.co/datasets/xavierdurawa/proof-pile-2-streaming.MDBDrumsPlusPlus
MDB Drums ++
This repository contains updated annotations for the MDB Drums dataset, which itself consists of drum annotations and audio files for 23 tracks from the MedleyDB dataset.
Contents
Audio: drum-only WAVs in audio/
Annotations: MIDI files in annotations/midi/
Logic Pro sessions: .logicx in logic_sessions/
Metadata: metadata.csv with one row per track
metadata.csv columns:
track_id
style
split
audio_path
midi_path
logic_project_path
sample_rate_hz… See the full description on the dataset page: https://huggingface.co/datasets/xavriley/MDBDrumsPlusPlus.XVNGAPS
GAPS Dataset (Guitar-Aligned Performance Scores)
Metadata (Google Sheets)
Paper (ArXiv)
300 solo guitar performances with accurately aligned MIDI and musicxml scores.
Update: version 1.1
audio now included
fixed scorehash case sensitive clashes (thanks Daniel Wägner for pointing this out)
Abstract
We introduce GAPS (Guitar-Aligned Performance Scores), a new dataset of classical guitar performances, and a benchmark guitar transcription model that achieves… See the full description on the dataset page: https://huggingface.co/datasets/xavriley/GAPS.argus-datasetsucfnet-datasetsFiloBass
FiloBass: A Dataset and Corpus Based Study of Jazz Basslines
Paper (arXiv)
Data (Zenodo)
Data (GitHub mirror)
This page hosts the dataset files for the paper, which was presented at ISMIR 2023. This was carried out by me, Xavier Riley, a PhD candidate on the AIM programme at QMUL.
This dataset contains audio, aligned midi and MusicXML scores for 48 tracks in the FiloBass dataset. These can be used to train automatic transcription models, beat tracking models or for musicological… See the full description on the dataset page: https://huggingface.co/datasets/xavriley/FiloBass.Variants-catala-cv16_1common_voice_es_16_1_accentFrancoisLeducGuitarDataset
François Leduc Guitar Dataset
paper: https://arxiv.org/abs/2402.15258
demo: https://xavriley.github.io/HighResolutionGuitarTranscription/
Audio and aligned MIDI transcriptions for 79 solo guitar perfomances. Used to train a state-of-the-art (at the time) audio-to-MIDI transcription model for guitar in the paper above.
If you use this dataset in academic work please cite:
@inproceedings{DBLP:conf/icassp/RileyED24,
author = {Xavier Riley and
Drew Edwards… See the full description on the dataset page: https://huggingface.co/datasets/xavriley/FrancoisLeducGuitarDataset.common_voice_ca_16_1_accentCharlieParkerAlignedOmnibook
Charlie Parker Aligned Omnibook Dataset
This dataset accompanies the paper "Reconstructing the Charlie Parker Omnibook using an audio-to-score automatic transcription pipeline" (SMC 2024).
Paper (arXiv)
Alternative Data Download (Zenodo)
Website
Introduction
Welcome to the companion site for the paper "Reconstructing the Charlie Parker Omnibook using an audio-to-score automatic transcription pipeline". This was work carried out by Xavier Riley, a PhD candidate on… See the full description on the dataset page: https://huggingface.co/datasets/xavriley/CharlieParkerAlignedOmnibook.details_xaviviro__FLAMA-0.5-3B
Dataset Card for Evaluation run of xaviviro/FLAMA-0.5-3B
Dataset automatically created during the evaluation run of model xaviviro/FLAMA-0.5-3B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_xaviviro__FLAMA-0.5-3B.RayBrownAlignedBasslines
Ray Brown Aligned Basslines Dataset
This repo contains midi, audio and scores for 9 tracks from the Oscar Peterson Trio Album "We Get Requests".
These are used to evaluate automatic transcription methods for bass, as described in my PhD Thesis.
The scores can be viewed here:
Time and Again
Corcovado
D & E
Girl from Ipanema
Goodbye JD
Have you met miss Jones?
My One and Only LovePeople
You Look Good To Me
common_voice_16_1_ca_up_5
Common Voice Corpus 16.1 Català (up_votes>5)
Dataset extret de mozilla-foundation/common_voice_16_1 només els splits train i test del Català i amb up_votes > 5
nie-pipeline-artifactsavatarzedongMvT-tarEindhovenWildflowernew_liberoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 127,
"total_frames": 11905,
"total_tasks": 4,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:127"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xavier033/new_libero.cloud-adapter-datasets
Cloud-Adapter-Datasets
This dataset card aims to describe the datasets used in the Cloud-Adapter, a collection of high-resolution satellite images and semantic segmentation masks for cloud detection and related tasks.
Install
pip install huggingface-hub
Usage
# Step 1: Download datasets
huggingface-cli download --repo-type dataset XavierJiezou/cloud-adapter-datasets --local-dir data --include hrc_whu.zip
huggingface-cli download --repo-type dataset… See the full description on the dataset page: https://huggingface.co/datasets/XavierJiezou/cloud-adapter-datasets.pinga-fogo-chico-xavier
🎙️ Pinga-Fogo com Chico Xavier — TV Tupi, 1971
As duas entrevistas históricas do médium Chico Xavier, transmitidas ao vivo pela
TV Tupi em 1971, transcritas e estruturadas em turnos de fala com timestamp.
345 turnos (115 deles respostas do próprio Chico Xavier), a partir de
6 horas de áudio — o registro mais extenso do médium falando de improviso,
sem edição, diante de um painel de jornalistas.
Arquivos
Arquivo
Programa
Turnos
Respostas do Chico… See the full description on the dataset page: https://huggingface.co/datasets/ia-espirita/pinga-fogo-chico-xavier.details_xaviviro__FLOR-6.3B-xat
Dataset Card for Evaluation run of xaviviro/FLOR-6.3B-xat
Dataset automatically created during the evaluation run of model xaviviro/FLOR-6.3B-xat on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_xaviviro__FLOR-6.3B-xat.SteveGilmoreAlignedBasslines
Steve Gilmore Aligned Basslines Dataset
This repo contains midi, audio (bass stems) and scores for 17 tracks from the Aebersold Playalong Vol. 35 "Jam Session" featuring Steve Gilmore on bass.
These are used to evaluate automatic transcription methods for bass, as described in my PhD Thesis.
Score are available to preview here:
My Funny Valentine
Our Love is Here to Stay
I Could Write a Book
Old Devil Moon
I've Grown Accustomed to Her Face
Speak LowCome Rain or Come Shine
The… See the full description on the dataset page: https://huggingface.co/datasets/xavriley/SteveGilmoreAlignedBasslines.pansharpening-datasets
Pansharpening-Datasets
This dataset card aims to describe the datasets used in the Pansharpening.
Install
pip install huggingface-hub
Usage
# Step 1: Download datasets
huggingface-cli download --repo-type dataset XavierJiezou/pansharpening-datasets --local-dir data --include PanBench.zip
# Step 2: Extract datasets
unzip PanBench.zip -d PanBench
Citation
@Article{cmfnet,
AUTHOR = {Wang, Shiying and Zou, Xuechao and Li, Kai and Xing, Junliang and… See the full description on the dataset page: https://huggingface.co/datasets/XavierJiezou/pansharpening-datasets.diffcr-datasets
Cloud Removal Visualization & Evaluation
Benchmark evaluation workspace for the DiffCR paper
(Diffusion-Based Cloud Removal for Sentinel-2 Multi-Temporal Imagery).
Two test datasets are covered:
Dataset
Samples
Methods
Sen2_MTC_Old
313
12
Sen2_MTC_New
687
12
Directory Layout
visualization/
├── paper-report.png ← reference metrics table from the paper
│
├── data/
│ ├── Sen2_MTC_New/
│ │ ├── GT/ ← 687 cloud-free… See the full description on the dataset page: https://huggingface.co/datasets/XavierJiezou/diffcr-datasets.pick_place_LIBEROThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 2,
"total_frames": 171,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xavier033/pick_place_LIBERO.cloudseg-datasets
Cloudseg-Datasets
This dataset card aims to describe the datasets used in the cloudseg.
Install
pip install huggingface-hub
Usage
# Step 1: Download datasets
huggingface-cli download --repo-type dataset XavierJiezou/cloudseg-datasets --local-dir data --include hrc_whu.zip
huggingface-cli download --repo-type dataset XavierJiezou/cloudseg-datasets --local-dir data --include gf12ms_whu.zip
huggingface-cli download --repo-type dataset… See the full description on the dataset page: https://huggingface.co/datasets/XavierJiezou/cloudseg-datasets.
