datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
include-base-44
INCLUDE-base (44 languages)
Dataset Description
Paper: http://arxiv.org/abs/2411.19799
Dataset Summary
INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
It contains 22,637 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-base-44.include-lite-44
INCLUDE-lite (44 languages)
Dataset Description
Paper: http://arxiv.org/abs/2411.19799
Dataset Summary
INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
It contains 11,095 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including regional… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-lite-44.movielens-100kAI4Bharat-INCLUDE-datasetINCLUDE
Dataset Card for INCLUDE
Dataset Summary
This dataset contains all videos in the INCLUDE dataset. As huggingface does not support video uploads at this time, the HF dataset contains metadata about each video such as the parent class, the video class, the path to the video and whether its a part of the INCLUDE-50 dataset (use include_50==True to get only include_50 videos).
The videos themselves can be downloaded from Zenodo using the provided bash script.… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/INCLUDE.asm_all_include_mistralinclude_benchinclude_cultureVagina-Vision-Image-Folder-Captions-IncludedThis repo contains over 6000 images of female anatomy for the purposes of captioning or training captioning models. There may also be applications in image generation training.
There is a list.txt included which lists the filenames for use with joycaption. I have also included a subdirectory containing txt captions generated by joycaption. These captions share the same filename as the parent image.
Open source datasets have a distinct lack of human anatomy and pornographic content, and this… See the full description on the dataset page: https://huggingface.co/datasets/hardlyworking/Vagina-Vision-Image-Folder-Captions-Included.arxiv-chandra-ocr-2-include-images-first50-20260415
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Source paper IDs in input list: 27,584
Processed IDs recorded in state/processed_ids.txt: 50
Successes: 50
Partial successes: 0
Errors: 0
Next shard index: 10
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415.INCLUDE_Dataset
🤟 INCLUDE: A Large-Scale Dataset for Indian Sign Language Recognition
A comprehensive video dataset for Indian Sign Language Recognition
🌐 Computer Vision •
🧠 Deep Learning •
🎥 Video Recognition •
🤟 Sign Language Recognition
👋 Welcome
Hello and welcome!
Thank you for your interest in INCLUDE: A Large-Scale Dataset for Indian Sign Language Recognition.
We are pleased to make this dataset available to the… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarcs/INCLUDE_Dataset.include-128include-esinclude-50chandra-ocr-2-vllm-include-images-demo-2604-07413-20260413
Chandra OCR 2 vLLM Include Images Demo
One-paper Chandra OCR 2 run using the vLLM backend with include_images=True.
Paper ID: 2604.07413
Source PDF: https://arxiv.org/pdf/2604.07413
Model: datalab-to/chandra-ocr-2
include-89JEE-Mains-Dataset-includes-2026-Jan-Attempt
JEE Mains Dataset (includes 2026 Jan Attempt)
This dataset was published on Kaggle by Samyakraj Bayar and mirrored here.
Download
The dataset is available as a ZIP archive: JEE Mains Dataset (includes 2026 Jan Attempt).zip
License
MIT
restructured-include_base_44I do not hold the copyright to this dataset; I merely restructured it to have the same structure as other datasets (that we are researching) to facilitate future coding and analysis. I refer to this link for the raw dataset.
include-2.0arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416.arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416.arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417.Africa-Total-Reserves-includes-gold-current-USD
Africa Total Reserves includes gold current USD | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Africa-Total-Reserves-includes-gold-current-USD.Arabic-Cohere-include-base-44-mmlu-style
The Refined Arabic Cohere INCLUDE Base 44 Dataset as MMLU-Style
Dataset Summary
INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark spanning 44 languages that evaluates multilingual LLMs in the actual linguistic environments where they are deployed. The original dataset contains 22,637 4-option multiple-choice questions (MCQs) extracted from academic and professional exams, covering 57 topics, including regional knowledge.
When we reviewed the Arabic… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-Cohere-include-base-44-mmlu-style.asia-owid-which-countries-include-malaria-vaccines-in-their-vaccination-schedules
Which Countries Include Malaria Vaccines In Their Vaccination Schedules | Asia (Our World in Data)
🌏 279 observations · 47 Asia countries · 2019–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 279 observations of Which Countries Include Malaria Vaccines In Their Vaccination Schedules data across 47 Asia countries, spanning 2019–2024.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-which-countries-include-malaria-vaccines-in-their-vaccination-schedules.arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416.include-base-44Social-Media-Posts-Dataset-Embeddings-Included-DUCKDB
📊 Social Media Posts Dataset (Embeddings Included)
Dataset Description
This dataset contains social media posts collected for the purpose of natural language analytics and semantic analysis.It is designed to support trend analysis, topic discovery, sentiment inference, and time-based analytics over historical social media data.
The dataset is intended to serve as the data backbone for a natural language analytics system where users can ask questions in plain English and… See the full description on the dataset page: https://huggingface.co/datasets/Bhavin1905/Social-Media-Posts-Dataset-Embeddings-Included-DUCKDB.arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416.asia-owid-forms-of-homelessness-included-in-available-statistics
Forms Of Homelessness Included In Available Statistics | Asia (Our World in Data)
🌏 33 observations · 33 Asia countries · 2010–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 33 observations of Forms Of Homelessness Included In Available Statistics data across 33 Asia countries, spanning 2010–2024.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Forms Of Homelessness… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-forms-of-homelessness-included-in-available-statistics.
