CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /include-base-44 INCLUDE-base (44 languages) Dataset Description Paper: http://arxiv.org/abs/2411.19799 Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed. It contains 22,637 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-base-44.textmultiple-choice10K<n<100K51 likes12k downloads1y agoHugging Face02CohereLabs /include-lite-44 INCLUDE-lite (44 languages) Dataset Description Paper: http://arxiv.org/abs/2411.19799 Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed. It contains 11,095 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including regional… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-lite-44.textmultiple-choice10K<n<100K16 likes2.6k downloads1y agoHugging Face03includeno /movielens-100k10K<n<100K0 likes966 downloads4y agoHugging Face04spsarolkar /AI4Bharat-INCLUDE-datasetvideo1K<n<10K0 likes551 downloads7mo agoHugging Face05ai4bharat /INCLUDE Dataset Card for INCLUDE Dataset Summary This dataset contains all videos in the INCLUDE dataset. As huggingface does not support video uploads at this time, the HF dataset contains metadata about each video such as the parent class, the video class, the path to the video and whether its a part of the INCLUDE-50 dataset (use include_50==True to get only include_50 videos). The videos themselves can be downloaded from Zenodo using the provided bash script.… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/INCLUDE.text1K<n<10K5 likes516 downloads2y agoHugging Face06bxiong /asm_all_include_mistral0 likes406 downloads5mo agoHugging Face07zechen-nlp /include_benchtext10K<n<100K0 likes310 downloads2y agoHugging Face08yangzhang33 /include_culturetabular10K<n<100K0 likes309 downloads5mo agoHugging Face09hardlyworking /Vagina-Vision-Image-Folder-Captions-IncludedThis repo contains over 6000 images of female anatomy for the purposes of captioning or training captioning models. There may also be applications in image generation training. There is a list.txt included which lists the filenames for use with joycaption. I have also included a subdirectory containing txt captions generated by joycaption. These captions share the same filename as the parent image. Open source datasets have a distinct lack of human anatomy and pornographic content, and this… See the full description on the dataset page: https://huggingface.co/datasets/hardlyworking/Vagina-Vision-Image-Folder-Captions-Included.1K<n<10K6 likes233 downloads1y agoHugging Face10nielsr /arxiv-chandra-ocr-2-include-images-first50-20260415 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415 Source paper IDs in input list: 27,584 Processed IDs recorded in state/processed_ids.txt: 50 Successes: 50 Partial successes: 0 Errors: 0 Next shard index: 10 Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415.imagen<1K0 likes119 downloads5mo agoHugging Face11manojkumarcs /INCLUDE_Dataset 🤟 INCLUDE: A Large-Scale Dataset for Indian Sign Language Recognition A comprehensive video dataset for Indian Sign Language Recognition 🌐 Computer Vision &nbsp;•&nbsp; 🧠 Deep Learning &nbsp;•&nbsp; 🎥 Video Recognition &nbsp;•&nbsp; 🤟 Sign Language Recognition 👋 Welcome Hello and welcome! Thank you for your interest in INCLUDE: A Large-Scale Dataset for Indian Sign Language Recognition. We are pleased to make this dataset available to the… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarcs/INCLUDE_Dataset.1 likes111 downloads23d agoHugging Face12include-results /include-128text100K<n<1M0 likes62 downloads4mo agoHugging Face13include-results /include-estabular10K<n<100K0 likes48 downloads4mo agoHugging Face1434data /include-500 likes46 downloads7mo agoHugging Face15nielsr /chandra-ocr-2-vllm-include-images-demo-2604-07413-20260413 Chandra OCR 2 vLLM Include Images Demo One-paper Chandra OCR 2 run using the vLLM backend with include_images=True. Paper ID: 2604.07413 Source PDF: https://arxiv.org/pdf/2604.07413 Model: datalab-to/chandra-ocr-2 imagen<1K1 likes42 downloads5mo agoHugging Face16epfl-nlp /include-89text100K<n<1M0 likes36 downloads4mo agoHugging Face17EduDevCommons /JEE-Mains-Dataset-includes-2026-Jan-Attempt JEE Mains Dataset (includes 2026 Jan Attempt) This dataset was published on Kaggle by Samyakraj Bayar and mirrored here. Download The dataset is available as a ZIP archive: JEE Mains Dataset (includes 2026 Jan Attempt).zip License MIT tabular-classification0 likes36 downloads2mo agoHugging Face18muhammadravi251001 /restructured-include_base_44I do not hold the copyright to this dataset; I merely restructured it to have the same structure as other datasets (that we are researching) to facilitate future coding and analysis. I refer to this link for the raw dataset. text10K<n<100K0 likes29 downloads1y agoHugging Face19epfl-nlp /include-2.00 likes27 downloads24d agoHugging Face20nielsr /arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416 Source paper IDs in input list: 1 Processed IDs recorded in state/processed_ids.txt: 1 Successes: 1 Partial successes: 0 Errors: 0 Next shard index: 1 Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416.imagen<1K0 likes26 downloads5mo agoHugging Face21nielsr /arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416 Source paper IDs in input list: 1 Processed IDs recorded in state/processed_ids.txt: 1 Successes: 1 Partial successes: 0 Errors: 0 Next shard index: 1 Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416.tabularn<1K0 likes22 downloads5mo agoHugging Face22nielsr /arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417 Source paper IDs in input list: 1 Processed IDs recorded in state/processed_ids.txt: 1 Successes: 1 Partial successes: 0 Errors: 0 Next shard index: 1 Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417.tabularn<1K0 likes20 downloads5mo agoHugging Face23electricsheepafrica /Africa-Total-Reserves-includes-gold-current-USD Africa Total Reserves includes gold current USD | Africa (World Bank) Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Africa-Total-Reserves-includes-gold-current-USD.tabulartabular-classificationn<1K0 likes19 downloads1mo agoHugging Face24Omartificial-Intelligence-Space /Arabic-Cohere-include-base-44-mmlu-style The Refined Arabic Cohere INCLUDE Base 44 Dataset as MMLU-Style Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark spanning 44 languages that evaluates multilingual LLMs in the actual linguistic environments where they are deployed. The original dataset contains 22,637 4-option multiple-choice questions (MCQs) extracted from academic and professional exams, covering 57 topics, including regional knowledge. When we reviewed the Arabic… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-Cohere-include-base-44-mmlu-style.textn<1K3 likes18 downloads2y agoHugging Face25electricsheepasia /asia-owid-which-countries-include-malaria-vaccines-in-their-vaccination-schedules Which Countries Include Malaria Vaccines In Their Vaccination Schedules | Asia (Our World in Data) 🌏 279 observations · 47 Asia countries · 2019–2024 · Repackaged by Electric Sheep Asia TL;DR This dataset contains 279 observations of Which Countries Include Malaria Vaccines In Their Vaccination Schedules data across 47 Asia countries, spanning 2019–2024. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-which-countries-include-malaria-vaccines-in-their-vaccination-schedules.texttabular-classificationn<1K0 likes18 downloads3mo agoHugging Face26nielsr /arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416 Source paper IDs in input list: 1 Processed IDs recorded in state/processed_ids.txt: 1 Successes: 1 Partial successes: 0 Errors: 0 Next shard… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416.imagen<1K0 likes16 downloads5mo agoHugging Face27rubricreward /include-base-44text10K<n<100K0 likes13 downloads1y agoHugging Face28Bhavin1905 /Social-Media-Posts-Dataset-Embeddings-Included-DUCKDB 📊 Social Media Posts Dataset (Embeddings Included) Dataset Description This dataset contains social media posts collected for the purpose of natural language analytics and semantic analysis.It is designed to support trend analysis, topic discovery, sentiment inference, and time-based analytics over historical social media data. The dataset is intended to serve as the data backbone for a natural language analytics system where users can ask questions in plain English and… See the full description on the dataset page: https://huggingface.co/datasets/Bhavin1905/Social-Media-Posts-Dataset-Embeddings-Included-DUCKDB.0 likes13 downloads8mo agoHugging Face29nielsr /arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416 Source paper IDs in input list: 1 Processed IDs recorded in state/processed_ids.txt: 1 Successes: 1 Partial successes: 0 Errors: 0 Next shard index: 1… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416.imagen<1K0 likes13 downloads5mo agoHugging Face30electricsheepasia /asia-owid-forms-of-homelessness-included-in-available-statistics Forms Of Homelessness Included In Available Statistics | Asia (Our World in Data) 🌏 33 observations · 33 Asia countries · 2010–2024 · Repackaged by Electric Sheep Asia TL;DR This dataset contains 33 observations of Forms Of Homelessness Included In Available Statistics data across 33 Asia countries, spanning 2010–2024. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Forms Of Homelessness… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-forms-of-homelessness-included-in-available-statistics.texttabular-classificationn<1K0 likes13 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.