CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face02mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes17k downloads2y agoHugging Face03mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes15k downloads2y agoHugging Face04mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes6k downloads2y agoHugging Face05ClareNie /Light-Omni-Training Light-Omni Training Dataset This repository contains the training data used by Light-Omni, a multimodal agent framework for reflexive video understanding with long-term memory. Light-Omni uses memory-augmented multimodal streams to train adapters for memory construction, response generation, and reaction/action control. Links Project page: https://clare-nie.github.io/Light-Omni/ Code: https://github.com/Clare-Nie/Light-Omni Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.audiovisual-question-answering100K<n<1M3 likes707 downloads3mo agoHugging Face06ismatsamadov /azerbaijan-court-data Azerbaijan Court System Dataset The most comprehensive open dataset of Azerbaijan's judicial system — 1.64 million structured records and 1.54 million court decision PDFs (~160 GB) covering court decisions, active cases, scheduled hearings, court registries, judges, lawyers, and mediator organizations. Built for AI engineers, legal tech startups, and researchers who need real-world legal data at scale. Quick Start Load with Hugging Face datasets from datasets… See the full description on the dataset page: https://huggingface.co/datasets/ismatsamadov/azerbaijan-court-data.imagetext-classification1M<n<10M2 likes417 downloads6mo agoHugging Face07recursal /reprocessed_singapore_national_speech_corpus Dataset Card for Reprocessed National Speech Corpus NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here. Dataset Details Dataset Description Dataset Description: The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.audiotext-generation1M<n<10M7 likes362 downloads2y agoHugging Face08pheepa /jira-comments-nsp Dataset Card for Dataset Name Dataset Summary Dataset contains pairs of sentences with next_sentence_label for NSP. Sentences was given from public jira projects dataset. Next sentence is always next sentence in one comment or sentence from reply to the comment. Supported Tasks and Leaderboards NSP, MLM Languages English Dataset Structure sentence_a, sentence_b, next_sentence_label Source Data… See the full description on the dataset page: https://huggingface.co/datasets/pheepa/jira-comments-nsp.texttext-generationn<1K0 likes39 downloads4y agoHugging Face09BinghamtonUniversity /cs415-twitch-chatstexttext-generationn<1K0 likes24 downloads2y agoHugging Face10cheryramneg /Danbooru2021-SQLite Danbooru 2021 SQLite Dataset Summary This is the metadata of danbooru 2021 dataset in SQLite format. https://gwern.net/danbooru2021 Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation… See the full description on the dataset page: https://huggingface.co/datasets/cheryramneg/Danbooru2021-SQLite.imagetext-generation1M<n<10M0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.