CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01iamtarun /python_code_instructions_18k_alpaca Dataset Card for python_code_instructions_18k_alpaca The dataset contains problem descriptions and code in python language. This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here. textquestion-answering10K<n<100K349 likes37k downloads3y agoHugging Face02Teklia /IAM-line IAM - line level Dataset Summary The IAM Handwriting Database contains forms of handwritten English text which can be used to train and test handwritten text recognizers and to perform writer identification and verification experiments. Note that all images are resized to a fixed height of 128 pixels. Languages All the documents in the dataset are written in English. Dataset Structure Data Instances { 'image':… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/IAM-line.imageimage-to-text10K<n<100K34 likes3.4k downloads3y agoHugging Face03iamroot /chat_formatted_examplestextn<1K0 likes2.3k downloads2y agoHugging Face04Telecom-Paris /iamd_v0 Internet Archive Music Dataset (IAMD v0) ~4.2M thirty-second music segments (34,469 hours) sourced from Creative-Commons audio on the Internet Archive, each paired with machine-generated natural-language captions and the original item metadata. Segments 4.2M Audio 34k hours Segment length 30 s nominal (mean 29.22 s) Format MP3, 320 kbps CBR, native channels + sample rate Shards 2,320 Parquet files Download size 4.53 TB Loading A… See the full description on the dataset page: https://huggingface.co/datasets/Telecom-Paris/iamd_v0.audioaudio-classification1M<n<10M6 likes2.3k downloads2mo agoHugging Face05iamkaikai /amazing_logos_v4 Dataset Card for "amazing_logos_v4" More Information needed image100K<n<1M20 likes1.7k downloads3y agoHugging Face06iamtarun /code_instructions_120k_alpaca Dataset Card for code_instructions_120k_alpaca This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the original source here. texttext-generation100K<n<1M69 likes1.3k downloads3y agoHugging Face07iamkaikai /dpchallenge DPChallenge Photo Metadata Dataset Dataset Description This dataset contains metadata and statistics from DPChallenge, a photography community platform where photographers participate in themed challenges and receive peer ratings. Key Features: Valuable Human Labels: Contains human-scored quality ratings from multiple rater groups (all users, commenters, participants, non-participants) Collection Date: Dec 2025 Data Quality: Only includes images with complete… See the full description on the dataset page: https://huggingface.co/datasets/iamkaikai/dpchallenge.imageimage-classification100K<n<1M1 likes1k downloads9mo agoHugging Face08iamshnoo /dallestreet Citation Information @misc{mukherjee2024crossroadscontinentsautomatedartifact, title={Crossroads of Continents: Automated Artifact Extraction for Cultural Adaptation with Large Multimodal Models}, author={Anjishnu Mukherjee and Ziwei Zhu and Antonios Anastasopoulos}, year={2024}, eprint={2407.02067}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2407.02067}, } imageimage-classification1K<n<10K0 likes875 downloads2y agoHugging Face09iamandrewliao /pickblueblock_blackbowl_all_quadrantstabular10K<n<100K0 likes554 downloads1y agoHugging Face10cyttic /eng-iam-textimage100K<n<1M1 likes514 downloads2mo agoHugging Face11IAmFuch /viet-cultural-vqa 🇻🇳 Vietnamese Cultural VQA Dataset 📖 Dataset Description The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering. 🎯 Dataset Summary 📊 Total Images: 28,505 high-quality cultural images 💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.imagevisual-question-answering10K<n<100K0 likes502 downloads5mo agoHugging Face12iamnguyen /mt_pubmedtext10M<n<100M0 likes499 downloads2y agoHugging Face13iamandrewliao /pickblueblock_blackbowl_active40_bottomleft_topright_certainfailurestabularrobotics10K<n<100K0 likes329 downloads8mo agoHugging Face14oxe-auge /iamlab_cmu_pickup_insert_train_500_631_augmented iamlab_cmu_pickup_insert_train_500_631_augmented Overview Codebase version: v2.1 Robots: google_robot, images, jaco, kinova3, kuka_iiwa, sawyer, ur5e, widowX, xarm7 FPS: 20.0 Episodes: 131 Frames: 30,143 Videos: 1,179 Chunks: 1 Splits: train: 0:131 Data Layout data_path : data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet video_path: videos/chunk-{episode_chunk:03d}/{video_key}/episode_{episode_index:06d}.mp4 Features… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/iamlab_cmu_pickup_insert_train_500_631_augmented.tabularrobotics10K<n<100K0 likes293 downloads11mo agoHugging Face15iamnguyen /pubmed-envitext10K<n<100K0 likes274 downloads2y agoHugging Face16iamtarun /code_contest_python3_alpaca Dataset Card for Code Contest Processed Dataset Summary This dataset contains coding contest questions and their solution written in Python3. This dataset is created by processing code_contest dataset from Deepmind. It is a competitive programming dataset for machine-learning. Read more about dataset at original source. Columns Description id : unique string associated with a problem description : problem description code : one correct code for the problem… See the full description on the dataset page: https://huggingface.co/datasets/iamtarun/code_contest_python3_alpaca.textquestion-answering1K<n<10K8 likes236 downloads3y agoHugging Face17IAmSkyDra /LeViSQA-v1audio10K<n<100K0 likes209 downloads1y agoHugging Face18iamandrewliao /pickblueblock_blackbowl_bottomleft_toprighttabular10K<n<100K0 likes208 downloads1y agoHugging Face19iamplus /Conversation_RepoDatasets : ShareGPT (https://huggingface.co/datasets/RyokoAI/ShareGPT52K) - https://huggingface.co/datasets/manojpreveen/ConversationalRepo/tree/main/sharegpt-raw OpenAssistant (https://huggingface.co/datasets/OpenAssistant/oasst1 -> https://huggingface.co/datasets/h2oai/openassistant_oasst1) - https://huggingface.co/datasets/manojpreveen/ConversationalRepo/tree/main/OpenAssistant ultrachat (https://huggingface.co/datasets/stingning/ultrachat) -… See the full description on the dataset page: https://huggingface.co/datasets/iamplus/Conversation_Repo.text0 likes200 downloads3y agoHugging Face20IAMRonHIT /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/IAMRonHIT/Fable-5-traces.tabulartext-generation1K<n<10K0 likes199 downloads3mo agoHugging Face21priyank-m /IAM_words_text_recognitionimage100K<n<1M9 likes198 downloads4y agoHugging Face22ianua /IAM-line IAM - line level Dataset Summary The IAM Handwriting Database contains forms of handwritten English text which can be used to train and test handwritten text recognizers and to perform writer identification and verification experiments. Note that all images are resized to a fixed height of 128 pixels. Languages All the documents in the dataset are written in English. Dataset Structure Data Instances { 'image':… See the full description on the dataset page: https://huggingface.co/datasets/ianua/IAM-line.imageimage-to-text10K<n<100K0 likes181 downloads14d agoHugging Face23iamjinchen /DD-VQAimagequestion-answering1K<n<10K0 likes176 downloads2y agoHugging Face24iamdyeus /ui-instruct-4k UI Instruct 4K A instruction-completion dataset for finetuning language models to specialize in generating Next.js / ShadCN UI components using React, TypeScript, and Tailwind CSS. Dataset Summary This dataset was created with the primary goal of finetuning Qwen 3.5 4B to become a specialist at outputting production-ready Next.js and ShadCN-based UI components. Each example consists of a natural language prompt describing a UI component or layout, paired with a clean… See the full description on the dataset page: https://huggingface.co/datasets/iamdyeus/ui-instruct-4k.texttext-generation1K<n<10K2 likes158 downloads6mo agoHugging Face25iamfadi /de-multi-legaltext100K<n<1M0 likes146 downloads2y agoHugging Face26iamandrewliao /uprightcup_bottomleft_toprighttabular10K<n<100K0 likes142 downloads8mo agoHugging Face27IAMJB /scanned-arxiv-papers-idtext100K<n<1M1 likes141 downloads2y agoHugging Face28Artemis-IA /IAMRIMES Résumé Ce projet présente la création d’un dataset manuscrit en français à partir de deux sources principales : Le dataset manuscrit IAM (en anglais). Le dataset RIMES (en français). Phases de Création Collecte des Données Sources : Récupération du jeu de données IAM contenant des textes manuscrits en anglais. Application d’un modèle de traduction automatique pour obtenir des textes en français. Génération synthétique de manuscrits en français à l’aide d’un modèle… See the full description on the dataset page: https://huggingface.co/datasets/Artemis-IA/IAMRIMES.imagetoken-classification10K<n<100K0 likes141 downloads2y agoHugging Face29i-am-mushfiq /FirstAidQA FirstAidQA: A Synthetic First-Aid and Emergency-Response Question-Answering Dataset Medical safety notice: FirstAidQA is intended for research and educational purposes. It is not a substitute for professional medical advice, emergency services, certified first-aid training, or clinical judgment. Models trained on this dataset may produce incomplete, outdated, or unsafe responses. Dataset Summary FirstAidQA is an English-language synthetic question-answering… See the full description on the dataset page: https://huggingface.co/datasets/i-am-mushfiq/FirstAidQA.textquestion-answering1K<n<10K8 likes136 downloads2mo agoHugging Face30iamPi /albedo_904k albedo_904k Merged, last-turn-cleaned SFT corpus of mini-swe-agent trajectories generated by three strong teacher models. Each row is a multi-turn messages list; the training target is the last assistant turn only. Fields messages: list of {role, content} turns (system / user / assistant ...). model: teacher that generated the completion. Composition (904,692 rows) model rows Qwen3-Next 751,689 Kimi-K2.6 135,772 deepseek-v3.2 17,231… See the full description on the dataset page: https://huggingface.co/datasets/iamPi/albedo_904k.texttext-generation100K<n<1M0 likes131 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.