datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
migration-bench-java-full
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.full_conservation_within_andropogoneae
Full Conservation Within Andropogoneae
This dataset contains 8,192 bp Sorghum bicolor v3 sequence windows for the
Evolutionary_constraint split from
kuleshov-group/cross-species-single-nucleotide-annotation.
Columns:
sequence: 8,192 bp DNA sequence window.
label: original binary conservation label.
The source coordinates match:
/workdir/jz963/genomes/Sorghum_bicolor_v3.1.1/assembly/Sbicolor_454_v3.0.1.fa
(Phytozome Sorghum bicolor v3.0.1 assembly, v3.1.1 annotation.)
Window… See the full description on the dataset page: https://huggingface.co/datasets/JingjingZhai/full_conservation_within_andropogoneae.full_tmdb_movies_datasetSITalk_fullfullOHLCVFull_Length_Motion_Capture_Dataset
Apple Arts Studios Full-Length Motion Capture Dataset
Dataset Overview
The Apple Arts Studios Full-Length Motion Capture Dataset is a professionally captured, full-body human-motion dataset containing 199 hours and 30 minutes of continuous motion capture data.
Unlike segmented motion datasets, this repository preserves the complete capture sequences without separating individual actions into short clips.
The recordings retain their continuous capture… See the full description on the dataset page: https://huggingface.co/datasets/Appleartsstudios/Full_Length_Motion_Capture_Dataset.vn-provinces-infant-full-immunization-rate
Vietnam infant full immunization rate
Vietnam infant full immunization rate. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (945 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (90 rows)
data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-infant-full-immunization-rate.Nasdaq_full_1_day_2007_to_2023full-dubai-pulsevox1-veri-full
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
1211
1211
40
# of samples
133777
14865
4874
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
lalm-judge-validation-full-duplex
LALM Judge Validation on Full-Duplex Voice Agents
Companion dataset for the paper A Reliability Assessment of
LALM Audio Judges for Full-Duplex Voice Agents.
This repository contains the anonymised ratings, adversarial-defect
recall tables, JSON schemas, and analysis scripts used to produce
every headline number, table, and figure in that paper.
Summary
209 rated stereo sessions: 152 full-duplex agent-client
conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.new_helium_adapter_checkpoint_librispeech_full_datasetDDCF-fullcorpus-mathvox1-iden-full
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Identification Split
train
validation
test
# of speakers
1251
1251
1251
# of samples
138361
6904
8251
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
openfacades-dataset-fullnew_helium_adapter_checkpoint_librispeech_full_datasetvox2-veri-full
VoxCeleb 2
VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
5,994
5,994
118
# of samples
982,808
109,201
36,237
Data Fields
ID (string): The ID of the sample with format <spk_id--utt_id_start_stop>.
duration (float64): The duration of the segment in seconds.
wav (string): The filepath of the waveform.
start (int64): The… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/vox2-veri-full.ArabicMMLU_full
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin
MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne
Introduction
We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/ArabicMMLU_full.astro-llms-full-query-data
AstroLLMs Full Query Dataset
This dataset includes all of the data collected in a four-week deployment of a Large Language Model-powered Slack chatbot trained on astrophysics papers. Astronomers were invited to interact with the chatbot, ask questions, and leave feedback. This data includes 368 question-answer pairs, including feedback, reactions, and labeling.
Dataset Structure
The columns of this dataset are thread_ts (unique time stamp of the query), channel_id… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/astro-llms-full-query-data.alanjo_graphics-card-full-specs
🟩NVIDIA & AMD🟥 GPUs Full Specs💠
Full specifications of NVIDIA and AMD graphics processing units (and more)
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 8,376
Files: 2
Files
gpu_specs_v6.csv
gpu_specs_v7.csv
Mirrored from Kaggle
full_article_stock_newscourt_opinions_filtered_full_sizecifar10-stats
CIFAR-10 CNN Layerwise Training Statistics
Dataset Description
This dataset contains layer-wise training statistics for a convolutional network trained on CIFAR-10, together with the corresponding test/acc.
Each row is one point at loss landscape. The features are computed on the last training batch of 1024 samples before the end of an epoch, and test/acc is measured immediately after that epoch.
The dataset includes statistics for several convolutional layers and the… See the full description on the dataset page: https://huggingface.co/datasets/Fullfix/cifar10-stats.EF_Full_Texts
Dataset Card for SF Nexus Extracted Features: Full Texts
Dataset Summary
The SF Nexus Extracted Features Full Texts dataset contains text and metadata from 403 mid-twentieth century science fiction books, originally digitized from Temple University Libraries' Paskow Science Fiction Collection.
After digitization, the books were cleaned using Abbyy FineReader.
Because this is a collection of copyrighted fiction, the books have been disaggregated.
Each row of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/SF-Corpus/EF_Full_Texts.full-ev-infrastrucutre
EV Charging Stations in Finland — Merged Dataset
File: ev_stations_merged.csv
Rows: 5,410 (one row per charging location) · Columns: 25
Companion file: merge_review.csv (3,158 rows - the matched pairs only, for manual review)
This dataset combines two independently collected lists of public EV charging locations in Finland into a single table. It is the base table for the DSP project: recommending 50 new fast-charging locations.
1. Source datasets
Google… See the full description on the dataset page: https://huggingface.co/datasets/data-sci-project/full-ev-infrastrucutre.Nasdaq_full_1_week_2007_to_2023TEDS-Full-Cleanedhuman-chromhmm-fullstack-data
human-chromhmm-fullstack-data
Dataset Summary
This dataset provides a multi-class annotation of genomic regions across the hg38 genome. It is derived from the ChromHMM fullstack annotation (Vu & Ernst, 2022; https://doi.org/10.1186/s13059-021-02572-z). Genomic regions are classified into 16 chromatin states. The data is derived from https://public.hoffman2.idre.ucla.edu/ernst/2K9RS//full_stack/full_stack_annotation_public_release/hg38/hg38_genome_100_segments.bed.gz.… See the full description on the dataset page: https://huggingface.co/datasets/Genentech/human-chromhmm-fullstack-data.insincere_fulldsDTF_full_comments_2016-2024
