CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face02mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes16k downloads2y agoHugging Face03mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes10k downloads2y agoHugging Face04pixparse /pdfa-eng-wds Dataset Card for PDF Association dataset (PDFA) Dataset Summary PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models. An example page of one pdf document, with added bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/pdfa-eng-wds.textimage-to-text1K<n<10K161 likes9.4k downloads2y agoHugging Face05mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes6k downloads2y agoHugging Face06LLMDH /marianne_pdf_7text10K<n<100K0 likes3.1k downloads2y agoHugging Face07LLMDH /marianne_pdf_9text100K<n<1M0 likes2.6k downloads2y agoHugging Face08LLMDH /marianne_pdf_3text100K<n<1M0 likes2.2k downloads2y agoHugging Face09LLMDH /marianne_pdf_5text100K<n<1M0 likes2.1k downloads2y agoHugging Face10LLMDH /marianne_pdf_8text100K<n<1M0 likes1.6k downloads2y agoHugging Face11LLMDH /marianne_pdf_4text10K<n<100K0 likes1.6k downloads2y agoHugging Face12PleIAs /WTO-PDFtext100K<n<1M8 likes1.5k downloads2y agoHugging Face13laion /CS-Arxiv-PDFs-08-25text10M<n<100M2 likes1.3k downloads1y agoHugging Face14LLMDH /marianne_pdf_10text100K<n<1M0 likes563 downloads2y agoHugging Face15PleIAs /AMF-PDFtext100K<n<1M7 likes541 downloads2y agoHugging Face16PleIAs /Math-PDFtext100K<n<1M2 likes411 downloads1y agoHugging Face17LLMDH /hal_pdf_extratext10K<n<100K0 likes393 downloads2y agoHugging Face18davanstrien /chrono-no-pdfimage10K<n<100K0 likes16 downloads1y agoHugging Face19smatiush /pdfa-eng-wds Dataset Card for PDF Association dataset (PDFA) Dataset Summary PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models. An example page of one pdf document, with added bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/smatiush/pdfa-eng-wds.textimage-to-text1M<n<10M0 likes11 downloads5mo agoHugging Face20ahmetklnc /pdf-indirilmeyentext10K<n<100K0 likes7 downloads5mo agoHugging Face21thinking-bio-lab /selected_cns_pdfgatedtext100K<n<1M0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.