datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
IAM-line
IAM - line level
Dataset Summary
The IAM Handwriting Database contains forms of handwritten English text which can be used to train and test handwritten text recognizers and to perform writer identification and verification experiments.
Note that all images are resized to a fixed height of 128 pixels.
Languages
All the documents in the dataset are written in English.
Dataset Structure
Data Instances
{
'image':… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/IAM-line.chat_formatted_examplesiamd_v0
Internet Archive Music Dataset (IAMD v0)
~4.2M thirty-second music segments (34,469 hours) sourced from
Creative-Commons audio on the Internet Archive, each
paired with machine-generated natural-language captions and the original item
metadata.
Segments
4.2M
Audio
34k hours
Segment length
30 s nominal (mean 29.22 s)
Format
MP3, 320 kbps CBR, native channels + sample rate
Shards
2,320 Parquet files
Download size
4.53 TB
Loading
A… See the full description on the dataset page: https://huggingface.co/datasets/Telecom-Paris/iamd_v0.amazing_logos_v4
Dataset Card for "amazing_logos_v4"
More Information needed
code_instructions_120k_alpaca
Dataset Card for code_instructions_120k_alpaca
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the original source here.
dpchallenge
DPChallenge Photo Metadata Dataset
Dataset Description
This dataset contains metadata and statistics from DPChallenge, a photography community platform where photographers participate in themed challenges and receive peer ratings.
Key Features:
Valuable Human Labels: Contains human-scored quality ratings from multiple rater groups (all users, commenters, participants, non-participants)
Collection Date: Dec 2025
Data Quality:
Only includes images with complete… See the full description on the dataset page: https://huggingface.co/datasets/iamkaikai/dpchallenge.dallestreet
Citation Information
@misc{mukherjee2024crossroadscontinentsautomatedartifact,
title={Crossroads of Continents: Automated Artifact Extraction for Cultural Adaptation with Large Multimodal Models},
author={Anjishnu Mukherjee and Ziwei Zhu and Antonios Anastasopoulos},
year={2024},
eprint={2407.02067},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2407.02067},
}
pickblueblock_blackbowl_all_quadrantseng-iam-textviet-cultural-vqa
🇻🇳 Vietnamese Cultural VQA Dataset
📖 Dataset Description
The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering.
🎯 Dataset Summary
📊 Total Images: 28,505 high-quality cultural images
💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.mt_pubmedpickblueblock_blackbowl_active40_bottomleft_topright_certainfailuresiamlab_cmu_pickup_insert_train_500_631_augmented
iamlab_cmu_pickup_insert_train_500_631_augmented
Overview
Codebase version: v2.1
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, sawyer, ur5e, widowX, xarm7
FPS: 20.0
Episodes: 131
Frames: 30,143
Videos: 1,179
Chunks: 1
Splits:
train: 0:131
Data Layout
data_path : data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet
video_path: videos/chunk-{episode_chunk:03d}/{video_key}/episode_{episode_index:06d}.mp4
Features… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/iamlab_cmu_pickup_insert_train_500_631_augmented.pubmed-envicode_contest_python3_alpaca
Dataset Card for Code Contest Processed
Dataset Summary
This dataset contains coding contest questions and their solution written in Python3.
This dataset is created by processing code_contest dataset from Deepmind. It is a competitive programming dataset for machine-learning. Read more about dataset at original source.
Columns Description
id : unique string associated with a problem
description : problem description
code : one correct code for the problem… See the full description on the dataset page: https://huggingface.co/datasets/iamtarun/code_contest_python3_alpaca.LeViSQA-v1pickblueblock_blackbowl_bottomleft_toprightConversation_RepoDatasets :
ShareGPT (https://huggingface.co/datasets/RyokoAI/ShareGPT52K) - https://huggingface.co/datasets/manojpreveen/ConversationalRepo/tree/main/sharegpt-raw
OpenAssistant (https://huggingface.co/datasets/OpenAssistant/oasst1 -> https://huggingface.co/datasets/h2oai/openassistant_oasst1) - https://huggingface.co/datasets/manojpreveen/ConversationalRepo/tree/main/OpenAssistant
ultrachat (https://huggingface.co/datasets/stingning/ultrachat) -… See the full description on the dataset page: https://huggingface.co/datasets/iamplus/Conversation_Repo.Fable-5-traces
Glint Research Dataset Card
Fable 5 Pi Agent Traces
A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation.
Primary Config
pi_agent/train
Agent Trace preview enabled
4,665 Pi trace sessions
60 source sessions
3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/IAMRonHIT/Fable-5-traces.IAM_words_text_recognitionIAM-line
IAM - line level
Dataset Summary
The IAM Handwriting Database contains forms of handwritten English text which can be used to train and test handwritten text recognizers and to perform writer identification and verification experiments.
Note that all images are resized to a fixed height of 128 pixels.
Languages
All the documents in the dataset are written in English.
Dataset Structure
Data Instances
{
'image':… See the full description on the dataset page: https://huggingface.co/datasets/ianua/IAM-line.DD-VQAui-instruct-4k
UI Instruct 4K
A instruction-completion dataset for finetuning language models to specialize in generating Next.js / ShadCN UI components using React, TypeScript, and Tailwind CSS.
Dataset Summary
This dataset was created with the primary goal of finetuning Qwen 3.5 4B to become a specialist at outputting production-ready Next.js and ShadCN-based UI components. Each example consists of a natural language prompt describing a UI component or layout, paired with a clean… See the full description on the dataset page: https://huggingface.co/datasets/iamdyeus/ui-instruct-4k.de-multi-legaluprightcup_bottomleft_toprightscanned-arxiv-papers-idIAMRIMES
Résumé
Ce projet présente la création d’un dataset manuscrit en français à partir de deux sources principales :
Le dataset manuscrit IAM (en anglais).
Le dataset RIMES (en français).
Phases de Création
Collecte des Données Sources :
Récupération du jeu de données IAM contenant des textes manuscrits en anglais.
Application d’un modèle de traduction automatique pour obtenir des textes en français.
Génération synthétique de manuscrits en français à l’aide d’un modèle… See the full description on the dataset page: https://huggingface.co/datasets/Artemis-IA/IAMRIMES.FirstAidQA
FirstAidQA: A Synthetic First-Aid and Emergency-Response Question-Answering Dataset
Medical safety notice: FirstAidQA is intended for research and educational purposes. It is not a substitute for professional medical advice, emergency services, certified first-aid training, or clinical judgment. Models trained on this dataset may produce incomplete, outdated, or unsafe responses.
Dataset Summary
FirstAidQA is an English-language synthetic question-answering… See the full description on the dataset page: https://huggingface.co/datasets/i-am-mushfiq/FirstAidQA.albedo_904k
albedo_904k
Merged, last-turn-cleaned SFT corpus of mini-swe-agent trajectories generated by
three strong teacher models. Each row is a multi-turn messages list; the
training target is the last assistant turn only.
Fields
messages: list of {role, content} turns (system / user / assistant ...).
model: teacher that generated the completion.
Composition (904,692 rows)
model
rows
Qwen3-Next
751,689
Kimi-K2.6
135,772
deepseek-v3.2
17,231… See the full description on the dataset page: https://huggingface.co/datasets/iamPi/albedo_904k.
