datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
iam_handwriting_finevision
Dataset Card for finevision_iam
This is a FiftyOne dataset with 5663 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/iam_handwriting_finevision")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/iam_handwriting_finevision.chanpicfortestIAM-line
IAM - line level
Dataset Summary
The IAM Handwriting Database contains forms of handwritten English text which can be used to train and test handwritten text recognizers and to perform writer identification and verification experiments.
Note that all images are resized to a fixed height of 128 pixels.
Languages
All the documents in the dataset are written in English.
Dataset Structure
Data Instances
{
'image':… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/IAM-line.chat_formatted_examplesiamd_v0
Internet Archive Music Dataset (IAMD v0)
~4.2M thirty-second music segments (34,469 hours) sourced from
Creative-Commons audio on the Internet Archive, each
paired with machine-generated natural-language captions and the original item
metadata.
Segments
4.2M
Audio
34k hours
Segment length
30 s nominal (mean 29.22 s)
Format
MP3, 320 kbps CBR, native channels + sample rate
Shards
2,320 Parquet files
Download size
4.53 TB
Loading
A… See the full description on the dataset page: https://huggingface.co/datasets/Telecom-Paris/iamd_v0.amazing_logos_v4
Dataset Card for "amazing_logos_v4"
More Information needed
ethosETHOS: onlinE haTe speecH detectiOn dataSet. This repository contains a dataset for hate speech
detection on social media platforms, called Ethos. There are two variations of the dataset:
Ethos_Dataset_Binary: contains 998 comments in the dataset alongside with a label
about hate speech presence or absence. 565 of them do not contain hate speech,
while the rest of them, 433, contain.
Ethos_Dataset_Multi_Label: which contains 8 labels for the 433 comments with hate speech content.
These labels are violence (if it incites (1) or not (0) violence), directed_vs_general (if it is
directed to a person (1) or a group (0)), and 6 labels about the category of hate speech like,
gender, race, national_origin, disability, religion and sexual_orientation.code_instructions_120k_alpaca
Dataset Card for code_instructions_120k_alpaca
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the original source here.
LNNet_Math_5B_part1ccimagiomimirMember and non-member splits for our MI experiments using MIMIR. Data is available for each source.
We also cache neighbors (generated for the NE attack).dpchallenge
DPChallenge Photo Metadata Dataset
Dataset Description
This dataset contains metadata and statistics from DPChallenge, a photography community platform where photographers participate in themed challenges and receive peer ratings.
Key Features:
Valuable Human Labels: Contains human-scored quality ratings from multiple rater groups (all users, commenters, participants, non-participants)
Collection Date: Dec 2025
Data Quality:
Only includes images with complete… See the full description on the dataset page: https://huggingface.co/datasets/iamkaikai/dpchallenge.iamlab_cmu_pickup_insert_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "franka",
"total_episodes": 631,
"total_frames": 146241,
"total_tasks": 7,
"total_videos": 1262,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:631"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/iamlab_cmu_pickup_insert_lerobot.metacot
Meta-CoT Dataset
Training data for the Meta-CoT project: teaching language models metacognitive self-monitoring during mathematical reasoning.
Files
File
Rows
Description
metacot_v2_trapi.parquet
4996
Meta-CoT V2 SFT data. Math solutions with <|meta|>...<|/meta|> self-reflection blocks. Generated via GPT-5.4 (TRAPI) with calibrated confidence, error-correction patterns, and final verification steps.
base_sft.parquet
4996
Base SFT data (no meta). Same… See the full description on the dataset page: https://huggingface.co/datasets/iamseungpil/metacot.dollarstreet
Dataset Card for "dollarstreet"
More Information needed
dallestreet
Citation Information
@misc{mukherjee2024crossroadscontinentsautomatedartifact,
title={Crossroads of Continents: Automated Artifact Extraction for Cultural Adaptation with Large Multimodal Models},
author={Anjishnu Mukherjee and Ziwei Zhu and Antonios Anastasopoulos},
year={2024},
eprint={2407.02067},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2407.02067},
}
xView2Copy from xView2, no need to register or login in.
Include parts [train, test, tier, hold] that are provided in XView2 website.
Using cat cmd to merge into a .zip file and then unzip it.
License is from XView2. Copyrights are owned with xView2 team.
link: https://xview2.org/download
iamlab_cmu_pickup_insert_train_500_631_augmented
iamlab_cmu_pickup_insert_train_500_631_augmented
Overview
Codebase version: v2.1
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, sawyer, ur5e, widowX, xarm7
FPS: 20.0
Episodes: 131
Frames: 30,143
Videos: 1,179
Chunks: 1
Splits:
train: 0:131
Data Layout
data_path : data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet
video_path: videos/chunk-{episode_chunk:03d}/{video_key}/episode_{episode_index:06d}.mp4
Features… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/iamlab_cmu_pickup_insert_train_500_631_augmented.iamlab_cmu_pickup_insert_raweng-iam-textlibero_plus_goalpickblueblock_blackbowl_all_quadrantslibero_plus_objectfluency-ecdict-offline
Fluency ECDICT Offline Pack
This repository hosts the versioned ECDICT SQLite packages downloaded by the
Fluency app for offline video word lookup and local subtitle focus-word matching.
Current release
Field
Value
Data version
ecdict-bc015ed2-focus13-v2
Schema
2
Entries
770,611
Focus-word lists
13
ZIP size
54,594,964 bytes
SQLite size
131,756,032 bytes
ZIP SHA-256
1e745ea698878772226a7df584129409dd4b27cb7ea26b4259bc770e7533352f
Source… See the full description on the dataset page: https://huggingface.co/datasets/iamzhangship/fluency-ecdict-offline.viet-cultural-vqa
🇻🇳 Vietnamese Cultural VQA Dataset
📖 Dataset Description
The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering.
🎯 Dataset Summary
📊 Total Images: 28,505 high-quality cultural images
💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.mt_pubmedInstruction_TuningFiles Contents Details :
Post-Process Code Info :
data_process.py
iamai_seed_tasks_v1.csv :
IAMAI's seed tasks - Version 1 (879)
Total Dataset Size : 879
===============================================================================================
iamai_v1.csv :
Instruction Tuning Dataset collected using seeds from iamai_seed_tasks_v1.csv and ChatGPT API for both prompts and outputs (~248k)
Total Dataset Size : ~248k
iamai_summarization_v1.csv :
Article Summarization dataset (both… See the full description on the dataset page: https://huggingface.co/datasets/iamplus/Instruction_Tuning.libero_plus_10InfyLerobotTask3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 21028,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iamirulofficial/InfyLerobotTask3.
