CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face02mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes16k downloads2y agoHugging Face03mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes14k downloads2y agoHugging Face04artefactory /Argimi-Ardian-Finance-10k-text The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.texttext-retrieval1M<n<10M19 likes9.2k downloads7mo agoHugging Face05mlfoundations /MINT-1T-ArXiv 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.imageimage-to-text1M<n<10M61 likes7.5k downloads2y agoHugging Face06mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes6k downloads2y agoHugging Face07Smith42 /minty-astro-ph MINT-1T ArXiv Astro-ph An astronomy-focused subset of mlfoundations/MINT-1T-ArXiv, filtered to include only papers from the astro-ph arXiv category (including cross-listed papers). Overview Papers ~845k Total size ~804 GB Format WebDataset tar shards Shards 287 (astro-ph-00000.tar to astro-ph-00286.tar) Shard size ~3 GB each Source MINT-1T (Awadalla et al., 2024) Data Format Each tar shard contains paired files per paper:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/minty-astro-ph.imagetext-generation100K<n<1M1 likes4.6k downloads5mo agoHugging Face08Qwen /DeepPlanning DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints DeepPlanningBench is a challenging benchmark for evaluating long-horizon agentic planning capabilities of large language models (LLMs) with verifiable constraints. It features realistic multi-day travel planning and multi-product shopping tasks that require proactive information acquisition, local constrained reasoning, and global constrained optimization. 🌐 Website:… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/DeepPlanning.texttext-generation1K<n<10K213 likes3k downloads7mo agoHugging Face09ClareNie /Light-Omni-Training Light-Omni Training Dataset This repository contains the training data used by Light-Omni, a multimodal agent framework for reflexive video understanding with long-term memory. Light-Omni uses memory-augmented multimodal streams to train adapters for memory construction, response generation, and reaction/action control. Links Project page: https://clare-nie.github.io/Light-Omni/ Code: https://github.com/Clare-Nie/Light-Omni Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.audiovisual-question-answering100K<n<1M3 likes697 downloads3mo agoHugging Face10milashkaarshif /MoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials. Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset. Looking forward to see more models and synthetic datasets trained from this raw archive, good luck! Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.texttext-generation100K<n<1M38 likes545 downloads7mo agoHugging Face11artefactory /Argimi-Ardian-Finance-10k-text-image The ArGiMI Ardian datasets : text and images The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.imagetext-retrieval1M<n<10M14 likes530 downloads7mo agoHugging Face12ismatsamadov /azerbaijan-court-data Azerbaijan Court System Dataset The most comprehensive open dataset of Azerbaijan's judicial system — 1.64 million structured records and 1.54 million court decision PDFs (~160 GB) covering court decisions, active cases, scheduled hearings, court registries, judges, lawyers, and mediator organizations. Built for AI engineers, legal tech startups, and researchers who need real-world legal data at scale. Quick Start Load with Hugging Face datasets from datasets… See the full description on the dataset page: https://huggingface.co/datasets/ismatsamadov/azerbaijan-court-data.imagetext-classification1M<n<10M2 likes413 downloads6mo agoHugging Face13recursal /reprocessed_singapore_national_speech_corpus Dataset Card for Reprocessed National Speech Corpus NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here. Dataset Details Dataset Description Dataset Description: The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.audiotext-generation1M<n<10M7 likes340 downloads2y agoHugging Face14lmms-lab /LLaVA-OneVision-Mid-Data Dataset Card for LLaVA-OneVision Due to unknow reasons, we are unable to process dataset with large amount into required HF format. So we directly upload the json files and image folders (compressed into tar.gz files). You can use the following link to directly download and decompress them. https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data/tree/main/evol_instruct We provide the whole details of LLaVA-OneVision Dataset. In this dataset, we include the data splits… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data.imagetext-generation100K<n<1M21 likes141 downloads2y agoHugging Face15segyges /OpenWebText2 Dataset Card for OpenWebText2 OpenWebText2 is a reasonably large corpus of scraped natural language data. Original hosting for this dataset has become difficult because it was hosted alongside another controversial dataset. To the best of my knowledge, this dataset itself is not encumbered in any way. It's a useful size for smaller language modelling experiments and is sometimes used in existing papers which it may be desirable to replicate. It is uploaded here to facilitate those… See the full description on the dataset page: https://huggingface.co/datasets/segyges/OpenWebText2.texttext-generationn<1K18 likes133 downloads2y agoHugging Face16jakobsnel /RAGTruth_Xtended Dataset Card for Dataset Name This dataset provides response token logits and hidden states, complementing the underlying RAGTruth dataset. It has been generated using https://github.com/jakobsnl/RAGTruth_Xtended. Dataset Details Dataset Description This dataset is built upon RAGTruth (github.com/ParticleMedia/RAGTruth), which consists of character-level annotation of different types of hallucination for responses to a given set of LLM tasks. Out of all models… See the full description on the dataset page: https://huggingface.co/datasets/jakobsnel/RAGTruth_Xtended.texttext-generation10K<n<100K0 likes108 downloads1y agoHugging Face17yejunliang23 /Nano3D-Edit-100k Nano3D-Edit-100k This dataset is the official data release for Nano3D, a training-free framework for precise and coherent 3D object editing without masks. Paper: Nano3D: A Training-Free Approach for Efficient 3D Editing Without MasksProject Page: https://jamesyjl.github.io/Nano3D/ Nano3D integrates FlowEdit into TRELLIS to perform localized 3D edits guided by front-view renderings, and introduces Voxel/Slat-Merge strategies to preserve structural consistency between edited and… See the full description on the dataset page: https://huggingface.co/datasets/yejunliang23/Nano3D-Edit-100k.texttext-generation1M<n<10M2 likes108 downloads6mo agoHugging Face18YangyiYY /VLM-SFTimagetext-generation1M<n<10M2 likes87 downloads2y agoHugging Face19ghemdd /gui_actor_webdataset GUI-Actor WebDataset A WebDataset format version of the GUI-Actor dataset for training vision-language models on GUI interaction tasks. Usage import webdataset as wds # Load the dataset dataset = wds.WebDataset("path/to/shards-*.tar") dataset = dataset.decode("pilrgb").to_tuple("jpg", "json") for image, metadata in dataset: # Process image and metadata pass Citation Please cite the original GUI-Actor paper if you use this dataset in your research. imagetext-generation1M<n<10M1 likes80 downloads1y agoHugging Face20Intel /fivl-instruct FiVL-Instruct Dataset FiVL: A Frameword for Improved Vision-Language Alignment introduces grounded datasets for both training and evaluation, building upon existing vision-question-answer and instruction datasets Each sample in the original datasets was augmented with key expressions, along with their corresponding bounding box indices and segmentation masks within the images. Dataset Details Creators: Intel Labs Version: 1.0 (Updated: 2024-12-18) License: CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/Intel/fivl-instruct.texttext-generation1M<n<10M0 likes75 downloads2y agoHugging Face21luca0621 /amex-gelab-448 AMEX SFT This dataset is a packaged export of the local amex_sft directory for uploading to the Hugging Face Hub as a dataset repository. Source Source dataset roots: /home1/irteam/data-vol1/amex_sft_hf_448 (3046 trajectories) Number of trajectory folders: 3046 Number of tar shards: 61 Trajectories per shard: 50 Layout shards/*.tar: tar shards containing trajectory folders manifest.jsonl: trajectory-to-shard index dataset_info.json: high-level metadata Each… See the full description on the dataset page: https://huggingface.co/datasets/luca0621/amex-gelab-448.imageimage-text-to-text1M<n<10M0 likes61 downloads6mo agoHugging Face221-800-SHARED-TASKS /xlsum-subset Dataset Card for "XL-Sum" Dataset Summary We present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high-resource, for many of which no public dataset is currently available. XL-Sum is highly abstractive, concise, and of high quality, as indicated by human and intrinsic evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/xlsum-subset.textsummarizationn<1K0 likes57 downloads2y agoHugging Face23pheepa /jira-comments-nsp Dataset Card for Dataset Name Dataset Summary Dataset contains pairs of sentences with next_sentence_label for NSP. Sentences was given from public jira projects dataset. Next sentence is always next sentence in one comment or sentence from reply to the comment. Supported Tasks and Leaderboards NSP, MLM Languages English Dataset Structure sentence_a, sentence_b, next_sentence_label Source Data… See the full description on the dataset page: https://huggingface.co/datasets/pheepa/jira-comments-nsp.texttext-generationn<1K0 likes37 downloads4y agoHugging Face24BinghamtonUniversity /cs415-twitch-chatstexttext-generationn<1K0 likes24 downloads2y agoHugging Face25transZ /efficient_llm Data V4 for NeurIPS LLM Challenge Contains 70949 samples collected from Huggingface: Math: 1273 gsm8k math_qa math-eval/TAL-SCQ5K TAL-SCQ5K-EN meta-math/MetaMathQA TIGER-Lab/MathInstruct Science: 42513 lighteval/mmlu - 'all', "split": 'auxiliary_train' lighteval/bbq_helm - 'all' openbookqa - 'main' ComplexQA: 2940 ARC-Challenge ARC-Easy piqa social_i_qa Muennighoff/babi Rowan/hellaswag ComplexQA1: 2060 medmcqa winogrande_xl, winogrande_debiased boolq sciq CNN: 2787… See the full description on the dataset page: https://huggingface.co/datasets/transZ/efficient_llm.texttext-generationn<1K0 likes17 downloads3y agoHugging Face26AlexWortega /physics-scenarios-packed physics-scenarios-packed Packed (tar.gz) version of a 2D rigid body physics dataset for training language models on next-frame prediction. 1,000,020 scenes × 200 frames simulated with Pymunk / Chipmunk2D. This repo is bandwidth-friendly: each scenario type ships as a single .tar.gz. For the unpacked JSONL files see physics-scenarios-raw. Scale Train: 900,000 scenes (24 seen scenario types × 37,500 each) Val: 100,020 scenes (30 scenario types × 3,334 each — includes 6… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/physics-scenarios-packed.texttext-generation100K<n<1M0 likes12 downloads5mo agoHugging Face27luca0621 /amex-gelab AMEX SFT This dataset is a packaged export of the local amex_sft directory for uploading to the Hugging Face Hub as a dataset repository. Source Source dataset roots: /ext_hdd2/tsyou/gelab-env/data_engine/amex_sft (3046 trajectories) Number of trajectory folders: 3046 Number of tar shards: 61 Trajectories per shard: 50 Layout shards/*.tar: tar shards containing trajectory folders manifest.jsonl: trajectory-to-shard index dataset_info.json: high-level… See the full description on the dataset page: https://huggingface.co/datasets/luca0621/amex-gelab.imageimage-text-to-text1M<n<10M0 likes11 downloads6mo agoHugging Face28AlexWortega /physics-scenarios-raw physics-scenarios-raw Raw (un-tarred) JSONL version of the 2D rigid body physics dataset. Each scene is a separate file under <split>/<scenario_type>/scene_<id>.jsonl. Streaming-friendly for HF datasets and curriculum sampling. For the bandwidth-efficient packaged version, see physics-scenarios-packed. Scale (this snapshot) Train: 80,000 scenes Val: 10,000 scenes Test: 10,000 scenes Frames per scene: 200 Format: one .jsonl per scene (1 header + 200 frame lines) This… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/physics-scenarios-raw.texttext-generation100K<n<1M0 likes11 downloads5mo agoHugging Face29cheryramneg /Danbooru2021-SQLite Danbooru 2021 SQLite Dataset Summary This is the metadata of danbooru 2021 dataset in SQLite format. https://gwern.net/danbooru2021 Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation… See the full description on the dataset page: https://huggingface.co/datasets/cheryramneg/Danbooru2021-SQLite.imagetext-generation1M<n<10M0 likes4 downloads6mo agoHugging Face30ouy-not-reversed /aiaa4051-path-planning-data AIAA4051 Grid Path Planning Generated Inputs This dataset repository hosts generated experiment input data for the GitHub project ouy-not-reversed/aiaa4051-path-planning. The GitHub repository contains code, documentation, small samples, and lightweight result summaries. Full generated inputs are hosted here because they are too large for regular GitHub commits. The archives restore files under data/generated/ when extracted at the root of the GitHub repository. Raw upstream data… See the full description on the dataset page: https://huggingface.co/datasets/ouy-not-reversed/aiaa4051-path-planning-data.texttext-generationn<1K0 likes3 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.