CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01iamtarun /python_code_instructions_18k_alpaca Dataset Card for python_code_instructions_18k_alpaca The dataset contains problem descriptions and code in python language. This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here. textquestion-answering10K<n<100K349 likes37k downloads3y agoHugging Face02ia03 /terminal-bench Terminal-Bench Dataset This dataset contains tasks from Terminal-Bench, a benchmark for evaluating AI agents in real terminal environments. Each task is packaged as a complete, self-contained archive that preserves the exact directory structure, binary files, Docker configurations, and test scripts needed for faithful reproduction. The archive column contains a gzipped tarball of the entire task directory. Dataset Overview Terminal-Bench evaluates AI agents on… See the full description on the dataset page: https://huggingface.co/datasets/ia03/terminal-bench.tabulartext-generationn<1K3 likes3.6k downloads1y agoHugging Face03THU-IAR /MIntRec Dataset details In real-world conversational interactions, we usually combine information from multiple modalities (e.g., text, video, audio) to help analyze human intentions. Though intent analysis has been widely explored in the Natural Language Processing community, there is a scarcity of data for multimodal intent analysis. Thus, we provide a novel multimodal intent benchmark dataset, MIntRec, to boom the research. To the best of our knowledge, it is the first multimodal intent… See the full description on the dataset page: https://huggingface.co/datasets/THU-IAR/MIntRec.text1K<n<10K1 likes3.6k downloads2y agoHugging Face04Teklia /IAM-line IAM - line level Dataset Summary The IAM Handwriting Database contains forms of handwritten English text which can be used to train and test handwritten text recognizers and to perform writer identification and verification experiments. Note that all images are resized to a fixed height of 128 pixels. Languages All the documents in the dataset are written in English. Dataset Structure Data Instances { 'image':… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/IAM-line.imageimage-to-text10K<n<100K33 likes3.3k downloads3y agoHugging Face05iamroot /chat_formatted_examplestextn<1K0 likes2.5k downloads2y agoHugging Face06moondream /ia_ocrContains pages from documents sourced from the Internet Archive, transcribed by Pixtral. Not super accurate, but useful during pretraining. @misc{moondream_ia_ocr, author = {Vikhyat Korrapati}, title = {IA OCR Dataset}, year = {2025}, url = {https://huggingface.co/datasets/moondream/ia_ocr}, note = {Accessed: 2025-03-07} } image100K<n<1M28 likes2.4k downloads1y agoHugging Face07Telecom-Paris /iamd_v0 Internet Archive Music Dataset (IAMD v0) ~4.2M thirty-second music segments (34,469 hours) sourced from Creative-Commons audio on the Internet Archive, each paired with machine-generated natural-language captions and the original item metadata. Segments 4.2M Audio 34k hours Segment length 30 s nominal (mean 29.22 s) Format MP3, 320 kbps CBR, native channels + sample rate Shards 2,320 Parquet files Download size 4.53 TB Loading A… See the full description on the dataset page: https://huggingface.co/datasets/Telecom-Paris/iamd_v0.audioaudio-classification1M<n<10M6 likes2.1k downloads2mo agoHugging Face08ianncity /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/KIMI-K2.5-1000000x.texttext-generation100K<n<1M265 likes2k downloads6mo agoHugging Face09iaouali /amazon-benchmark Amazon query–bundle benchmark Canonical, category-organized query and reference-positive data. Experiment traces should reference this repository by commit SHA, category, split, and candidate_id, rather than republishing the dataset. Musical Instruments Split Examples agent_dev 2,028 agent_hidden 1,960 Each record contains a query and 3–7 reference product IDs. These are observed reference positives, not exhaustive labels for all valid… See the full description on the dataset page: https://huggingface.co/datasets/iaouali/amazon-benchmark.texttext-retrieval1K<n<10K0 likes1.8k downloads21h agoHugging Face10iapp /thai_handwriting_dataset Thai Handwriting Dataset This dataset combines two major Thai handwriting datasets: BEST 2019 Thai Handwriting Recognition dataset (train-0000.parquet) Thai Handwritten Free Dataset by Wang (train-0001.parquet onwards) Maintainer kobkrit@iapp.co.th Dataset Description BEST 2019 Dataset Contains handwritten Thai text images along with their ground truth transcriptions. The images have been processed and standardized for machine learning tasks.… See the full description on the dataset page: https://huggingface.co/datasets/iapp/thai_handwriting_dataset.imagetext-to-image10K<n<100K22 likes1.4k downloads2y agoHugging Face11iamkaikai /amazing_logos_v4 Dataset Card for "amazing_logos_v4" More Information needed image100K<n<1M20 likes1.4k downloads3y agoHugging Face12Yiyang-Ian-Li /LongDA LongDA Dataset Card Dataset Description LongDA is a data analysis benchmark for evaluating LLM-based agents under documentation-intensive analytical workflows. It features authentic U.S. government survey data with complete, long documentation, testing LLMs' ability to navigate complex real-world datasets before performing analysis. Dataset Summary 505 queries extracted from 30 expert-written publications 17 U.S. national surveys covering health… See the full description on the dataset page: https://huggingface.co/datasets/Yiyang-Ian-Li/LongDA.documentquestion-answeringn<1K1 likes1.4k downloads3mo agoHugging Face13iamtarun /code_instructions_120k_alpaca Dataset Card for code_instructions_120k_alpaca This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the original source here. texttext-generation100K<n<1M69 likes1.3k downloads3y agoHugging Face14rebrowser /iaai-dataset IAAI Insurance Auto Auction Dataset Daily sample of IAAI insurance auto auction lots with damage assessments, title status, bidding data, and branch locations across North America. This dataset is a preview sample of the IAAI dataset published by Rebrowser. If you're doing academic research, you may be eligible for free access to a much larger slice — see Free Datasets for Research. This dataset contains 1 entity, each in its own folder: Auction Listings (auction-listings). See… See the full description on the dataset page: https://huggingface.co/datasets/rebrowser/iaai-dataset.tabularother10K<n<100K0 likes1k downloads1d agoHugging Face15iamkaikai /dpchallenge DPChallenge Photo Metadata Dataset Dataset Description This dataset contains metadata and statistics from DPChallenge, a photography community platform where photographers participate in themed challenges and receive peer ratings. Key Features: Valuable Human Labels: Contains human-scored quality ratings from multiple rater groups (all users, commenters, participants, non-participants) Collection Date: Dec 2025 Data Quality: Only includes images with complete… See the full description on the dataset page: https://huggingface.co/datasets/iamkaikai/dpchallenge.imageimage-classification100K<n<1M1 likes1k downloads9mo agoHugging Face16iamshnoo /dallestreet Citation Information @misc{mukherjee2024crossroadscontinentsautomatedartifact, title={Crossroads of Continents: Automated Artifact Extraction for Cultural Adaptation with Large Multimodal Models}, author={Anjishnu Mukherjee and Ziwei Zhu and Antonios Anastasopoulos}, year={2024}, eprint={2407.02067}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2407.02067}, } imageimage-classification1K<n<10K0 likes760 downloads2y agoHugging Face17oxe-auge /iamlab_cmu_pickup_insert_train_500_631_augmented iamlab_cmu_pickup_insert_train_500_631_augmented Overview Codebase version: v2.1 Robots: google_robot, images, jaco, kinova3, kuka_iiwa, sawyer, ur5e, widowX, xarm7 FPS: 20.0 Episodes: 131 Frames: 30,143 Videos: 1,179 Chunks: 1 Splits: train: 0:131 Data Layout data_path : data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet video_path: videos/chunk-{episode_chunk:03d}/{video_key}/episode_{episode_index:06d}.mp4 Features… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/iamlab_cmu_pickup_insert_train_500_631_augmented.tabularrobotics10K<n<100K0 likes756 downloads11mo agoHugging Face18marianna13 /IA-bookstabular1M<n<10M0 likes689 downloads4y agoHugging Face19iapp /MMMU-Thai MMMU Thai (MMMU Benchmark Translated to Thai) MMMU Thai is a dataset for evaluating multimodal models on massive multi-discipline tasks requiring college-level knowledge and deliberate reasoning. This dataset is translated from MMMU (A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI) into Thai. Dataset Details MMMU Thai consists of 11,500 meticulously collected multimodal questions from college exams, quizzes, and textbooks… See the full description on the dataset page: https://huggingface.co/datasets/iapp/MMMU-Thai.imagequestion-answering10K<n<100K2 likes669 downloads2y agoHugging Face20cyttic /eng-iam-textimage100K<n<1M1 likes586 downloads2mo agoHugging Face21IAmFuch /viet-cultural-vqa 🇻🇳 Vietnamese Cultural VQA Dataset 📖 Dataset Description The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering. 🎯 Dataset Summary 📊 Total Images: 28,505 high-quality cultural images 💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.imagevisual-question-answering10K<n<100K0 likes532 downloads5mo agoHugging Face22iabufarha /ar_sarcasm Dataset Card for ArSarcasm Dataset Summary ArSarcasm is a new Arabic sarcasm detection dataset. The dataset was created using previously available Arabic sentiment analysis datasets (SemEval 2017 and ASTD) and adds sarcasm and dialect labels to them. The dataset contains 10,547 tweets, 1,682 (16%) of which are sarcastic. For more details, please check the paper From Arabic Sentiment Analysis to Sarcasm Detection: The ArSarcasm Dataset Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/iabufarha/ar_sarcasm.texttext-classification10K<n<100K18 likes525 downloads3y agoHugging Face23iamandrewliao /pickblueblock_blackbowl_all_quadrantstabular10K<n<100K0 likes510 downloads1y agoHugging Face24iagoalves /iagoras_dataset-valcir-shards-1200-finaltabular1M<n<10M0 likes492 downloads1y agoHugging Face25iapp /iapp_wiki_qa_squad iapp_wiki_qa_squad Extractive question answering over Thai Wikipedia articles, in SQuAD format. 7,242 questions across 1,912 articles, annotated by people iApp hired for the purpose. from datasets import load_dataset dataset = load_dataset("iapp/iapp_wiki_qa_squad") This works again as of the August 2026 revision. Until then it did not. The repository carried a loading script and no data, and datasets dropped script support at v3, so load_dataset failed and every… See the full description on the dataset page: https://huggingface.co/datasets/iapp/iapp_wiki_qa_squad.textquestion-answering1K<n<10K7 likes442 downloads1mo agoHugging Face26gt111lk /IADBE_Custom_Dataset 📘 Custom Dataset Custom dataset provides both anomalib and YOLO format datasets. You can import it with the Huggingface way, or just clone it from GitHub and download it to your local machine. How to use the IADBE platform, check details here. Use Huggingface from datasets import load_dataset ds = load_dataset("gt111lk/IADBE_Custom_Dataset") Use Git Clone git clone https://huggingface.co/datasets/gt111lk/IADBE_Custom_Dataset imageimage-segmentation1K<n<10K0 likes428 downloads2y agoHugging Face27irl-kit /IA-Bench IA-bench ( Interaction-Aware Bench) Human ground-truth annotations of the interacted object for robot manipulation subtasks. Each sample is one subtask: the full subtask video clip, the gripper proprioception aligned 1:1 to those frames, the language instruction, and two boxes: initial_object_box (object on the first frame) and target_object_box (object on the last frame). Boxes are pixel [x1, y1, x2, y2]. Configs from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/irl-kit/IA-Bench.tabularrobotics1K<n<10K2 likes411 downloads22d agoHugging Face28ianncity /GLM-5.2-Conversation GLM-5.2 · Conversation-50000x 50,000x traces distilled from GLM-5.2 on High reasoning Token Count: 120M Distribution: Speaking domains: •Greetings •Customer Support •Step by step explanations •Motivational language •Logical Questions •Creative Writing STEM: •Algebra, calculus, quantum mechanics concepts •Astromony and astrophysics •Datascience and machine learning •Biology Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Conversation.texttext-generation10K<n<100K55 likes409 downloads2mo agoHugging Face29iastate /onestop_english Dataset Card for OneStopEnglish corpus Dataset Summary OneStopEnglish is a corpus of texts written at three reading levels, and demonstrates its usefulness for through two applications - automatic readability assessment and automatic text simplification. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances An instance example: { "text": "When you see… See the full description on the dataset page: https://huggingface.co/datasets/iastate/onestop_english.texttext-classificationn<1K17 likes382 downloads2y agoHugging Face30iamnguyen /mt_pubmedtext10M<n<100M0 likes381 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.