CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-mll /glue Dataset Card for GLUE Dataset Summary GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems. Supported Tasks and Leaderboards The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks: ax A manually-curated evaluation dataset for fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/glue.tabulartext-classification1M<n<10M1.1k likes848k downloads3y agoHugging Face02mlfoundations /MINT-1T-HTML 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.textimage-to-text100M<n<1B98 likes698k downloads2y agoHugging Face03mlfoundations /dclm-baseline-1.0 DCLM-baseline DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets Llama2 7B 2T ✗ 49.2 45.8 34.1 DeepSeek 7B 2T ✗ 50.7 48.5 35.3 Mistral-0.3 7B ? ✗ 57.0 62.7 45.1 QWEN-2 7B ? ✗ 57.5 71.9 50.5 Llama3 8B 15T ✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.tabular1B<n<10B315 likes633k downloads2y agoHugging Face04mlfoundations /dcvlm-baseline-200b DCVLM-Baseline (200B tokens) DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token dataset release consisting of 103,985,276 samples, curated from our DCVLM-large data pool. A smaller 6.25B-token version is also available. ⚠️ NOTE: The training data is the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-200b.imageimage-text-to-text10K<n<100K8 likes194k downloads2mo agoHugging Face05mlfoundations /dclm-pool-7b-2x3 likes149k downloads2y agoHugging Face06nyu-mll /blimp Dataset Card for "blimp" Dataset Summary BLiMP is a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English. BLiMP consists of 67 sub-datasets, each containing 1000 minimal pairs isolating specific contrasts in syntax, morphology, or semantics. The data is automatically generated according to expert-crafted grammars. Supported Tasks and Leaderboards More Information Needed Languages More Information… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/blimp.texttext-classification10K<n<100K40 likes149k downloads3y agoHugging Face07mlfoundations /datacomp_pools DataComp Pools This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.image19 likes87k downloads3y agoHugging Face08mueller91 /MLAADgated Introduction Welcome to MLAAD: The Multi-Language Audio Anti-Spoofing Dataset -- a dataset to train, test and evaluate audio deepfake detection. See the paper for more information. License MLAAD is published strictly for non-commercial academic research use, under the CC-BY-NC 4.0 license. Commercial use is not permitted. Bibtex If you use this dataset, please consider citing it as follows. @article{muller2024mlaad, title={MLAAD: The… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/MLAAD.audioaudio-classification100K<n<1M45 likes58k downloads25d agoHugging Face09nyu-mll /multi_nli Dataset Card for Multi-Genre Natural Language Inference (MultiNLI) Dataset Summary The Multi-Genre Natural Language Inference (MultiNLI) corpus is a crowd-sourced collection of 433k sentence pairs annotated with textual entailment information. The corpus is modeled on the SNLI corpus, but differs in that covers a range of genres of spoken and written text, and supports a distinctive cross-genre generalization evaluation. The corpus served as the basis for the shared task… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/multi_nli.texttext-classification100K<n<1M120 likes44k downloads3y agoHugging Face10ml-resources /Daimon-Infinity Daimon-Infinity mirror This repository is a file-preserving mirror of daimonrobotics/Daimon-Infinity on ModelScope. Source and license Upstream: daimonrobotics/Daimon-Infinity License: CC BY-NC-SA 4.0 Attribution: Daimon Robotics / Daimon-Infinity This mirror keeps the upstream directory layout and is distributed under the same CC BY-NC-SA 4.0 license. No data is altered; files are transferred with integrity checks supplied by ModelScope and the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/ml-resources/Daimon-Infinity.audio1K<n<10K2 likes39k downloads18d agoHugging Face11mlfoundations /datacomp_xlarge DataComp XLarge Pool This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.image10B<n<100B21 likes39k downloads3y agoHugging Face12MLCommons /peoples_speech Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.audioautomatic-speech-recognition1M<n<10M285 likes39k downloads2y agoHugging Face13FineEnvs /HF_ML_Tasksmith HF ML Tasksmith Fifty PR-derived Harbor tasks from Accelerate, Diffusers, PEFT, Transformers and TRL, including CPU and GPU tasks. Contains 50 Harbor tasks generated with the owned tasksmith recipe in Repo2RLEnv. Browse the complete task bundles in Harbor Visualiser or open the task folders. Each folder is a runnable Harbor task: tasks/<task_id>/ ├── task.toml # Harbor configuration and provenance ├── instruction.md # Task shown to the coding agent… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/HF_ML_Tasksmith.n<1K0 likes37k downloads8d agoHugging Face14japanese-asr /whisper_transcriptions.mls.wer_10.0.vectorized1M<n<10M1 likes36k downloads2y agoHugging Face15mlfoundations /dclm-pool-1b-1x3 likes35k downloads2y agoHugging Face16mlfoundations /dcvlm-balanced-200b DCVLM-Balanced (200B tokens) DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool. The instruction-heavy counterpart (DCVLM-baseline) is available as dcvlm-baseline-200b, along with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.imageimage-text-to-text10K<n<100K1 likes31k downloads2mo agoHugging Face17mlfoundations /dcvlm_pool_large DCVLM-Pool (large) The raw candidate pool at the large scale of our DataComp-VLM benchmark: 1,949,321,868 samples / 166.7 TB across 166 source datasets, as WebDataset tar shards — ≈4× the medium pool. 🚚 Upload in progress This repo is being populated incrementally and is not yet complete — shards are still being uploaded. Sources already present are final and safe to use; sources with fewer shards than the counts quoted below have not finished uploading yet.… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_large.image-text-to-text1B<n<10B0 likes30k downloads5d agoHugging Face18MLCommons /unsupervised_peoples_speech Dataset Card for Unsupervised Peoples Speech Dataset Description Dataset Summary The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers. Point of Contact: MLCommons Datasets Discord Dataset Structure This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.audioautomatic-speech-recognition81 likes26k downloads2y agoHugging Face19mlabonne /FineTome-100k FineTome-100k The FineTome dataset is a subset of arcee-ai/The-Tome (without arcee-ai/qwen2-72b-magpie-en), re-filtered using HuggingFaceFW/fineweb-edu-classifier. It was made for my article "Fine-tune Llama 3.1 Ultra-Efficiently with Unsloth". text100K<n<1M279 likes26k downloads2y agoHugging Face20mlfoundations /dclm-pool-7b-1x1 likes24k downloads2y agoHugging Face21mlabonne /harmful_behaviorstextn<1K156 likes23k downloads2y agoHugging Face22MLCommons /speech-wikimedia Dataset Card for Speech Wikimedia Dataset Summary The Speech Wikimedia Dataset is a compilation of audiofiles with transcriptions extracted from wikimedia commons that is licensed for academic and commercial usage under CC and Public domain. It includes 2,000+ hours of transcribed speech in different languages with a diverse set of speakers. Each audiofile should have one or more transcriptions in different languages. Transcription languages English German… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/speech-wikimedia.audion<1K14 likes23k downloads3y agoHugging Face23mlabonne /harmless_alpacatext10K<n<100K47 likes22k downloads2y agoHugging Face24jablonkagroup /chempile-mlift ChemPile-MLIFT A comprehensive multimodal dataset for chemistry property prediction using vision large language models 📋 Dataset Summary ChemPile-MLIFT is a dataset designed for multimodal chemistry property prediction tasks, specifically focusing on the prediction of chemical properties using vision large language models (VLLMs). It is part of the ChemPile project, which aims to create a comprehensive collection of chemistry-related data for training LLMs. The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/chempile-mlift.imagetext-generation10M<n<100M14 likes22k downloads1y agoHugging Face25mlfoundations /dclm-baseline-1.0-parquet DCLM-baseline Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format. DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.tabular1B<n<10B56 likes21k downloads2y agoHugging Face26mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face27mlfoundations /dcvlm_pool_medium DCVLM-Pool (medium) The raw candidate pool at the medium scale of our DataComp-VLM benchmark: 483,576,747 samples / 41.1 TB across 166 source datasets, as WebDataset tar shards — ≈4× the small pool. This pool is unfiltered and unmixed. It is the input to a data-curation experiment, not a training set. You choose the filters and the mixing ratios, and create another training set. If you instead want a ready-to-train dataset, use dcvlm-baseline-200b (our reference SoTA… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_medium.image-text-to-text100M<n<1B0 likes20k downloads1mo agoHugging Face28HuggingFriends /mllm-as-embodied-world-judge MLLM-as-Embodied-World-Judge Data for judging physical adherence and instruction alignment of generated embodied-manipulation videos. Start here path what it is final/ the current release — train.jsonl (11,520), test.jsonl (802), and its README data/ source and generated videos, referenced by video_url in the splits Benchmark tooling path what it is bench/LEADERBOARD.md judge results table bench/TESTSET.md benchmark… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/mllm-as-embodied-world-judge.2 likes17k downloads15d agoHugging Face29LeMaterial /LeMat-Bulk-MLIP-Hull LeMat-Bulk MLIP Hull Reference Datasets This dataset contains materials close to the convex hull computed using various ML interatomic potentials (MLIPs). Dataset Splits all: Contains ALL materials with hull energies for all MLIPs (no threshold filtering) dft, orb, uma, mace_mp, mace_omat: Materials within 0.001 eV/atom of respective hulls Energy Types dft: DFT reference energies orb: ORB model energies uma: UMA model energies mace_mp: MACE-MP model energies… See the full description on the dataset page: https://huggingface.co/datasets/LeMaterial/LeMat-Bulk-MLIP-Hull.tabular1M<n<10M0 likes17k downloads1y agoHugging Face30japanese-asr /whisper_transcriptions.mls.wer_10.0audio1M<n<10M2 likes16k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.