datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
terminal-bench
Terminal-Bench Dataset
This dataset contains tasks from Terminal-Bench, a benchmark for evaluating AI agents in real terminal environments. Each task is packaged as a complete, self-contained archive that preserves the exact directory structure, binary files, Docker configurations, and test scripts needed for faithful reproduction.
The archive column contains a gzipped tarball of the entire task directory.
Dataset Overview
Terminal-Bench evaluates AI agents on… See the full description on the dataset page: https://huggingface.co/datasets/ia03/terminal-bench.MIntRec
Dataset details
In real-world conversational interactions, we usually combine information from multiple modalities (e.g., text, video, audio) to help analyze human intentions. Though intent analysis has been widely explored in the Natural Language Processing community, there is a scarcity of data for multimodal intent analysis. Thus, we provide a novel multimodal intent benchmark dataset, MIntRec, to boom the research. To the best of our knowledge, it is the first multimodal intent… See the full description on the dataset page: https://huggingface.co/datasets/THU-IAR/MIntRec.IAM-line
IAM - line level
Dataset Summary
The IAM Handwriting Database contains forms of handwritten English text which can be used to train and test handwritten text recognizers and to perform writer identification and verification experiments.
Note that all images are resized to a fixed height of 128 pixels.
Languages
All the documents in the dataset are written in English.
Dataset Structure
Data Instances
{
'image':… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/IAM-line.chat_formatted_examplesia_ocrContains pages from documents sourced from the Internet Archive, transcribed by Pixtral. Not super accurate, but useful during pretraining.
@misc{moondream_ia_ocr,
author = {Vikhyat Korrapati},
title = {IA OCR Dataset},
year = {2025},
url = {https://huggingface.co/datasets/moondream/ia_ocr},
note = {Accessed: 2025-03-07}
}
iamd_v0
Internet Archive Music Dataset (IAMD v0)
~4.2M thirty-second music segments (34,469 hours) sourced from
Creative-Commons audio on the Internet Archive, each
paired with machine-generated natural-language captions and the original item
metadata.
Segments
4.2M
Audio
34k hours
Segment length
30 s nominal (mean 29.22 s)
Format
MP3, 320 kbps CBR, native channels + sample rate
Shards
2,320 Parquet files
Download size
4.53 TB
Loading
A… See the full description on the dataset page: https://huggingface.co/datasets/Telecom-Paris/iamd_v0.KIMI-K2.5-1000000x
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/KIMI-K2.5-1000000x.amazon-benchmark
Amazon query–bundle benchmark
Canonical, category-organized query and reference-positive data. Experiment traces should reference this repository by commit SHA, category, split, and candidate_id, rather than republishing the dataset.
Musical Instruments
Split
Examples
agent_dev
2,028
agent_hidden
1,960
Each record contains a query and 3–7 reference product IDs. These are observed reference positives, not exhaustive labels for all valid… See the full description on the dataset page: https://huggingface.co/datasets/iaouali/amazon-benchmark.thai_handwriting_dataset
Thai Handwriting Dataset
This dataset combines two major Thai handwriting datasets:
BEST 2019 Thai Handwriting Recognition dataset (train-0000.parquet)
Thai Handwritten Free Dataset by Wang (train-0001.parquet onwards)
Maintainer
kobkrit@iapp.co.th
Dataset Description
BEST 2019 Dataset
Contains handwritten Thai text images along with their ground truth transcriptions. The images have been processed and standardized for machine learning tasks.… See the full description on the dataset page: https://huggingface.co/datasets/iapp/thai_handwriting_dataset.amazing_logos_v4
Dataset Card for "amazing_logos_v4"
More Information needed
LongDA
LongDA Dataset Card
Dataset Description
LongDA is a data analysis benchmark for evaluating LLM-based agents under documentation-intensive analytical workflows. It features authentic U.S. government survey data with complete, long documentation, testing LLMs' ability to navigate complex real-world datasets before performing analysis.
Dataset Summary
505 queries extracted from 30 expert-written publications
17 U.S. national surveys covering health… See the full description on the dataset page: https://huggingface.co/datasets/Yiyang-Ian-Li/LongDA.code_instructions_120k_alpaca
Dataset Card for code_instructions_120k_alpaca
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the original source here.
iaai-dataset
IAAI Insurance Auto Auction Dataset
Daily sample of IAAI insurance auto auction lots with damage assessments, title status, bidding data, and branch locations across North America.
This dataset is a preview sample of the IAAI dataset published by Rebrowser. If you're doing academic research, you may be eligible for free access to a much larger slice — see Free Datasets for Research.
This dataset contains 1 entity, each in its own folder: Auction Listings (auction-listings). See… See the full description on the dataset page: https://huggingface.co/datasets/rebrowser/iaai-dataset.dpchallenge
DPChallenge Photo Metadata Dataset
Dataset Description
This dataset contains metadata and statistics from DPChallenge, a photography community platform where photographers participate in themed challenges and receive peer ratings.
Key Features:
Valuable Human Labels: Contains human-scored quality ratings from multiple rater groups (all users, commenters, participants, non-participants)
Collection Date: Dec 2025
Data Quality:
Only includes images with complete… See the full description on the dataset page: https://huggingface.co/datasets/iamkaikai/dpchallenge.dallestreet
Citation Information
@misc{mukherjee2024crossroadscontinentsautomatedartifact,
title={Crossroads of Continents: Automated Artifact Extraction for Cultural Adaptation with Large Multimodal Models},
author={Anjishnu Mukherjee and Ziwei Zhu and Antonios Anastasopoulos},
year={2024},
eprint={2407.02067},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2407.02067},
}
iamlab_cmu_pickup_insert_train_500_631_augmented
iamlab_cmu_pickup_insert_train_500_631_augmented
Overview
Codebase version: v2.1
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, sawyer, ur5e, widowX, xarm7
FPS: 20.0
Episodes: 131
Frames: 30,143
Videos: 1,179
Chunks: 1
Splits:
train: 0:131
Data Layout
data_path : data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet
video_path: videos/chunk-{episode_chunk:03d}/{video_key}/episode_{episode_index:06d}.mp4
Features… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/iamlab_cmu_pickup_insert_train_500_631_augmented.IA-booksMMMU-Thai
MMMU Thai (MMMU Benchmark Translated to Thai)
MMMU Thai is a dataset for evaluating multimodal models on massive multi-discipline tasks requiring college-level knowledge and deliberate reasoning. This dataset is translated from MMMU (A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI) into Thai.
Dataset Details
MMMU Thai consists of 11,500 meticulously collected multimodal questions from college exams, quizzes, and textbooks… See the full description on the dataset page: https://huggingface.co/datasets/iapp/MMMU-Thai.eng-iam-textviet-cultural-vqa
🇻🇳 Vietnamese Cultural VQA Dataset
📖 Dataset Description
The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering.
🎯 Dataset Summary
📊 Total Images: 28,505 high-quality cultural images
💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.ar_sarcasm
Dataset Card for ArSarcasm
Dataset Summary
ArSarcasm is a new Arabic sarcasm detection dataset.
The dataset was created using previously available Arabic sentiment analysis
datasets (SemEval 2017
and ASTD) and adds sarcasm and
dialect labels to them.
The dataset contains 10,547 tweets, 1,682 (16%) of which are sarcastic.
For more details, please check the paper
From Arabic Sentiment Analysis to Sarcasm Detection: The ArSarcasm Dataset
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/iabufarha/ar_sarcasm.pickblueblock_blackbowl_all_quadrantsiagoras_dataset-valcir-shards-1200-finaliapp_wiki_qa_squad
iapp_wiki_qa_squad
Extractive question answering over Thai Wikipedia articles, in SQuAD format.
7,242 questions across 1,912 articles, annotated by people iApp hired for the purpose.
from datasets import load_dataset
dataset = load_dataset("iapp/iapp_wiki_qa_squad")
This works again as of the August 2026 revision. Until then it did not. The
repository carried a loading script and no data, and datasets dropped script support
at v3, so load_dataset failed and every… See the full description on the dataset page: https://huggingface.co/datasets/iapp/iapp_wiki_qa_squad.IADBE_Custom_Dataset
📘 Custom Dataset
Custom dataset provides both anomalib and YOLO format datasets.
You can import it with the Huggingface way, or just clone it from GitHub and download it to your local machine. How to use the IADBE platform, check details here.
Use Huggingface
from datasets import load_dataset
ds = load_dataset("gt111lk/IADBE_Custom_Dataset")
Use Git Clone
git clone https://huggingface.co/datasets/gt111lk/IADBE_Custom_Dataset
IA-Bench
IA-bench ( Interaction-Aware Bench)
Human ground-truth annotations of the interacted object for robot manipulation
subtasks. Each sample is one subtask: the full subtask video clip, the gripper
proprioception aligned 1:1 to those frames, the language instruction, and two boxes:
initial_object_box (object on the first frame) and target_object_box
(object on the last frame). Boxes are pixel [x1, y1, x2, y2].
Configs
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/irl-kit/IA-Bench.GLM-5.2-Conversation
GLM-5.2 · Conversation-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Token Count: 120M
Distribution:
Speaking domains:
•Greetings
•Customer Support
•Step by step explanations
•Motivational language
•Logical Questions
•Creative Writing
STEM:
•Algebra, calculus, quantum mechanics concepts
•Astromony and astrophysics
•Datascience and machine learning
•Biology
Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Conversation.onestop_english
Dataset Card for OneStopEnglish corpus
Dataset Summary
OneStopEnglish is a corpus of texts written at three reading levels, and demonstrates its usefulness for through two applications - automatic readability assessment and automatic text simplification.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
An instance example:
{
"text": "When you see… See the full description on the dataset page: https://huggingface.co/datasets/iastate/onestop_english.mt_pubmed
