datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MINT-1T-ArXiv
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.deepsynth-en-arxiv
DeepSynth - arXiv Scientific Paper Summarization
Dataset Description
Scientific paper abstracts from arXiv, covering computer science, physics, and mathematics.
Visual encoding preserves mathematical notation and document structure critical for scientific summarization.
This dataset is part of the DeepSynth project, which uses visual text encoding for multilingual summarization with the DeepSeek-OCR vision-language model. Text documents are converted into images and… See the full description on the dataset page: https://huggingface.co/datasets/baconnier/deepsynth-en-arxiv.
