datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glue
Dataset Card for GLUE
Dataset Summary
GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems.
Supported Tasks and Leaderboards
The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks:
ax
A manually-curated evaluation dataset for fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/glue.MINT-1T-HTML
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.dclm-baseline-1.0
DCLM-baseline
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets
Llama2
7B
2T
✗
49.2
45.8
34.1
DeepSeek
7B
2T
✗
50.7
48.5
35.3
Mistral-0.3
7B
?
✗
57.0
62.7
45.1
QWEN-2
7B
?
✗
57.5
71.9
50.5
Llama3
8B
15T
✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.dcvlm-baseline-200b
DCVLM-Baseline (200B tokens)
DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper.
It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat
WebDataset tar shards so it can be consumed by any training
stack.
This is a 200B-token dataset release consisting of 103,985,276 samples, curated from our DCVLM-large data pool.
A smaller 6.25B-token version is also available.
⚠️ NOTE: The training data is the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-200b.dclm-pool-7b-2xblimp
Dataset Card for "blimp"
Dataset Summary
BLiMP is a challenge set for evaluating what language models (LMs) know about
major grammatical phenomena in English. BLiMP consists of 67 sub-datasets, each
containing 1000 minimal pairs isolating specific contrasts in syntax,
morphology, or semantics. The data is automatically generated according to
expert-crafted grammars.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/blimp.datacomp_pools
DataComp Pools
This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.MLAAD
Introduction
Welcome to MLAAD: The Multi-Language Audio Anti-Spoofing Dataset -- a dataset to train, test and evaluate audio deepfake detection. See
the paper for more information.
License
MLAAD is published strictly for non-commercial academic research use, under the CC-BY-NC 4.0 license. Commercial use is not permitted.
Bibtex
If you use this dataset, please consider citing it as follows.
@article{muller2024mlaad,
title={MLAAD: The… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/MLAAD.multi_nli
Dataset Card for Multi-Genre Natural Language Inference (MultiNLI)
Dataset Summary
The Multi-Genre Natural Language Inference (MultiNLI) corpus is a
crowd-sourced collection of 433k sentence pairs annotated with textual
entailment information. The corpus is modeled on the SNLI corpus, but differs in
that covers a range of genres of spoken and written text, and supports a
distinctive cross-genre generalization evaluation. The corpus served as the
basis for the shared task… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/multi_nli.Daimon-Infinity
Daimon-Infinity mirror
This repository is a file-preserving mirror of
daimonrobotics/Daimon-Infinity on ModelScope.
Source and license
Upstream: daimonrobotics/Daimon-Infinity
License: CC BY-NC-SA 4.0
Attribution: Daimon Robotics / Daimon-Infinity
This mirror keeps the upstream directory layout and is distributed under the
same CC BY-NC-SA 4.0 license. No data is altered; files are transferred with
integrity checks supplied by ModelScope and the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/ml-resources/Daimon-Infinity.datacomp_xlarge
DataComp XLarge Pool
This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.peoples_speech
Dataset Card for People's Speech
Dataset Summary
The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.HF_ML_Tasksmith
HF ML Tasksmith
Fifty PR-derived Harbor tasks from Accelerate, Diffusers, PEFT, Transformers and TRL, including CPU and GPU tasks.
Contains 50 Harbor tasks generated with the owned
tasksmith recipe in Repo2RLEnv.
Browse the complete task bundles in Harbor Visualiser or
open the task folders. Each folder is a runnable Harbor task:
tasks/<task_id>/
├── task.toml # Harbor configuration and provenance
├── instruction.md # Task shown to the coding agent… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/HF_ML_Tasksmith.whisper_transcriptions.mls.wer_10.0.vectorizeddclm-pool-1b-1xdcvlm-balanced-200b
DCVLM-Balanced (200B tokens)
DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper.
It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat
WebDataset tar shards so it can be consumed by any training
stack.
This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool.
The instruction-heavy counterpart (DCVLM-baseline) is available as
dcvlm-baseline-200b, along with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.dcvlm_pool_large
DCVLM-Pool (large)
The raw candidate pool at the large scale of our DataComp-VLM
benchmark: 1,949,321,868 samples / 166.7 TB across 166 source datasets, as
WebDataset tar shards — ≈4× the medium pool.
🚚 Upload in progress
This repo is being populated incrementally and is not yet complete — shards are still being
uploaded. Sources already present are final and safe to use; sources with fewer shards than the counts
quoted below have not finished uploading yet.… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_large.unsupervised_peoples_speech
Dataset Card for Unsupervised Peoples Speech
Dataset Description
Dataset Summary
The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers.
Point of Contact: MLCommons Datasets Discord
Dataset Structure
This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.FineTome-100k
FineTome-100k
The FineTome dataset is a subset of arcee-ai/The-Tome (without arcee-ai/qwen2-72b-magpie-en), re-filtered using HuggingFaceFW/fineweb-edu-classifier.
It was made for my article "Fine-tune Llama 3.1 Ultra-Efficiently with Unsloth".
dclm-pool-7b-1xharmful_behaviorsspeech-wikimedia
Dataset Card for Speech Wikimedia
Dataset Summary
The Speech Wikimedia Dataset is a compilation of audiofiles with transcriptions extracted from wikimedia commons that is licensed for academic and commercial usage under CC and Public domain. It includes 2,000+ hours of transcribed speech in different languages with a diverse set of speakers.
Each audiofile should have one or more transcriptions in different languages.
Transcription languages
English
German… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/speech-wikimedia.harmless_alpacachempile-mlift
ChemPile-MLIFT
A comprehensive multimodal dataset for chemistry property prediction using vision large language models
📋 Dataset Summary
ChemPile-MLIFT is a dataset designed for multimodal chemistry property prediction tasks, specifically focusing on the prediction of chemical properties using vision large language models (VLLMs). It is part of the ChemPile project, which aims to create a comprehensive collection of chemistry-related data for training LLMs. The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/chempile-mlift.dclm-baseline-1.0-parquet
DCLM-baseline
Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format.
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.MINT-1T-PDF-CC-2023-23
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.dcvlm_pool_medium
DCVLM-Pool (medium)
The raw candidate pool at the medium scale of our DataComp-VLM
benchmark: 483,576,747 samples / 41.1 TB across 166 source datasets, as
WebDataset tar shards — ≈4× the small pool.
This pool is unfiltered and unmixed. It is the input to a data-curation experiment, not a training set. You choose
the filters and the mixing ratios, and create another training set. If you instead want a
ready-to-train dataset, use dcvlm-baseline-200b
(our reference SoTA… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_medium.mllm-as-embodied-world-judge
MLLM-as-Embodied-World-Judge
Data for judging physical adherence and instruction alignment of generated
embodied-manipulation videos.
Start here
path
what it is
final/
the current release — train.jsonl (11,520), test.jsonl (802), and its README
data/
source and generated videos, referenced by video_url in the splits
Benchmark tooling
path
what it is
bench/LEADERBOARD.md
judge results table
bench/TESTSET.md
benchmark… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/mllm-as-embodied-world-judge.LeMat-Bulk-MLIP-Hull
LeMat-Bulk MLIP Hull Reference Datasets
This dataset contains materials close to the convex hull computed using various ML interatomic potentials (MLIPs).
Dataset Splits
all: Contains ALL materials with hull energies for all MLIPs (no threshold filtering)
dft, orb, uma, mace_mp, mace_omat: Materials within 0.001 eV/atom of respective hulls
Energy Types
dft: DFT reference energies
orb: ORB model energies
uma: UMA model energies
mace_mp: MACE-MP model energies… See the full description on the dataset page: https://huggingface.co/datasets/LeMaterial/LeMat-Bulk-MLIP-Hull.whisper_transcriptions.mls.wer_10.0
