datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
migration-bench-java-full
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.FullBenchfull_conservation_within_andropogoneae
Full Conservation Within Andropogoneae
This dataset contains 8,192 bp Sorghum bicolor v3 sequence windows for the
Evolutionary_constraint split from
kuleshov-group/cross-species-single-nucleotide-annotation.
Columns:
sequence: 8,192 bp DNA sequence window.
label: original binary conservation label.
The source coordinates match:
/workdir/jz963/genomes/Sorghum_bicolor_v3.1.1/assembly/Sbicolor_454_v3.0.1.fa
(Phytozome Sorghum bicolor v3.0.1 assembly, v3.1.1 annotation.)
Window… See the full description on the dataset page: https://huggingface.co/datasets/JingjingZhai/full_conservation_within_andropogoneae.full_tmdb_movies_datasetbank-additional-fullamazon_review_fullprescription-fullSITalk_fullFULL_EPSTEIN_INDEX
FULL_EPSTEIN_INDEX
CONTENT WARNING: This repository contains graphic and highly sensitive material regarding sexual abuse, exploitation, trafficking, and violence. It also contains unverified allegations and raw witness statements. User discretion is strongly advised.
Overview
Note. There is ALOT of data. OCR made mistakes scanning the files. So that being said, there is a lot of noise in the dataset, whether it be from OCR taking words out of normal… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/FULL_EPSTEIN_INDEX.conceptnet-full-en-essentials
Conceptnet Full EN (essentials)
Dataset Summary:
This dataset is a compact and simplified version of ConceptNet, emphasizing English concepts and their sources. It retains the essential information about the relations in a format that is straightforward and user-friendly. Designed for efficiency and ease of use, this dataset is particularly suitable for scenarios with computational constraints. While the original ConceptNet database exceeds 20GB in size, this streamlined… See the full description on the dataset page: https://huggingface.co/datasets/openworld-domains/conceptnet-full-en-essentials.wsb_fullNasdaq_full_1_day_2007_to_2023vn-provinces-infant-full-immunization-rate
Vietnam infant full immunization rate
Vietnam infant full immunization rate. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (945 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (90 rows)
data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-infant-full-immunization-rate.Curated-Fox-News-Headlines-and-Full-Text
Curated Fox News Headlines and Full Text
This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis.
📁 Dataset Format
Format: CSV
Encoding: UTF-8
Fields:
headline: The article title or headline
publish_date: Date the article was published (YYYY-MM-DD)
content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.Full_Length_Motion_Capture_Dataset
Apple Arts Studios Full-Length Motion Capture Dataset
Dataset Overview
The Apple Arts Studios Full-Length Motion Capture Dataset is a professionally captured, full-body human-motion dataset containing 199 hours and 30 minutes of continuous motion capture data.
Unlike segmented motion datasets, this repository preserves the complete capture sequences without separating individual actions into short clips.
The recordings retain their continuous capture… See the full description on the dataset page: https://huggingface.co/datasets/Appleartsstudios/Full_Length_Motion_Capture_Dataset.fullOHLCV5G_Faults_FullSynthetic dataset of 5G mobile network faults
vox1-veri-full
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
1211
1211
40
# of samples
133777
14865
4874
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
full-dubai-pulsenews-bias-full-data
**Please access the latest verison of data that is here https://huggingface.co/datasets/shainar/BEAD **
email at shaina.raza@torontomu.ca for usage of data
Please cite us if you use it
@article{raza2024beads,
title={BEADs: Bias Evaluation Across Domains},
author={Raza, Shaina and Rahman, Mizanur and Zhang, Michael R},
journal={arXiv preprint arXiv:2406.04220},
year={2024}
}
license: cc-by-nc-4.0
language:
- en
pretty_name: Navigating News… See the full description on the dataset page: https://huggingface.co/datasets/newsmediabias/news-bias-full-data.lalm-judge-validation-full-duplex
LALM Judge Validation on Full-Duplex Voice Agents
Companion dataset for the paper A Reliability Assessment of
LALM Audio Judges for Full-Duplex Voice Agents.
This repository contains the anonymised ratings, adversarial-defect
recall tables, JSON schemas, and analysis scripts used to produce
every headline number, table, and figure in that paper.
Summary
209 rated stereo sessions: 152 full-duplex agent-client
conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.vox1-iden-full
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Identification Split
train
validation
test
# of speakers
1251
1251
1251
# of samples
138361
6904
8251
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
court_opinions_filtered_full_sizefull-ev-infrastrucutre
EV Charging Stations in Finland — Merged Dataset
File: ev_stations_merged.csv
Rows: 5,410 (one row per charging location) · Columns: 25
Companion file: merge_review.csv (3,158 rows - the matched pairs only, for manual review)
This dataset combines two independently collected lists of public EV charging locations in Finland into a single table. It is the base table for the DSP project: recommending 50 new fast-charging locations.
1. Source datasets
Google… See the full description on the dataset page: https://huggingface.co/datasets/data-sci-project/full-ev-infrastrucutre.DDCF-fullcorpus-mathpdbbind_fullfull_crossword_puzzlesThese are automatically created crosswords from this dataset.
Words contains the all the words of the crossword. First horizontal, then vertical.
Clues contains the clues for the words in the same order. There is a year after each clue indicating
when the orginal crossword was published where the clue came from.
There are 4 different "shapes" with each 10000 games. The shape is indicated by the type field.
Type "x-x-x":
xxxxx
x x x
xxxxx
x x x
xxxxx
Type "x-x-x-x":
xxxxxxx
x x x x
xxxxxxx
x x… See the full description on the dataset page: https://huggingface.co/datasets/jeggers/full_crossword_puzzles.ArabicMMLU_full
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin
MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne
Introduction
We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/ArabicMMLU_full.disease_symptoms_prec_fullstage22
