datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
image-as-an-imu-finetuning
Image as an IMU: Real-world Finetuning Dataset
Official real-world finetuning dataset from Image as an IMU: Estimating Camera Motion from a Single Motion-Blurred Image (ICCV 2025 Oral).
[arXiv] [Webpage] [GitHub]
PIXL, University of Oxford
Jerred Chen, Ronald Clark
Dataset Details
This dataset consists of 32 sequences of real-world motion-blurred videos in various indoor scenes, captured using the iPhone 13 camera.
dataset_train_real-world.csv and… See the full description on the dataset page: https://huggingface.co/datasets/jerredchen00/image-as-an-imu-finetuning.fineweb-1m-sampleMuMo-Finetuning
MuMo Finetuning Dataset
This repository contains the finetuning datasets used in the paper: Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning.
Paper: Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning
Project Page: NeurIPS 2025 Poster
Code: GitHub Repository
Hub (this dataset): https://huggingface.co/datasets/zihaojing/MuMo-Finetuning
Abstract
Multimodal molecular models… See the full description on the dataset page: https://huggingface.co/datasets/zihaojing/MuMo-Finetuning.robust-finetuningPlease refer to the following source for the original datasets:
GSM8K: https://huggingface.co/datasets/openai/gsm8k
MATH: https://huggingface.co/datasets/hendrycks/competition_math
math-resample: In this section we subsample the 1,000 subsample only (yes it's balance)
HumanEval+: https://huggingface.co/datasets/evalplus/humanevalplus
MBPP: https://huggingface.co/datasets/google-research-datasets/mbpp
MBPP+: https://huggingface.co/datasets/evalplus/mbppplus
ARC Challenge:… See the full description on the dataset page: https://huggingface.co/datasets/appier-ai-research/robust-finetuning.brittleness-results
Adapters copied (2026-09-08). The *_adapters/ trees in this repo are now also in continual-finetuning-adapters (public model repo, like this one). Deleted here (260908): the byte-identical results/raw/* copies, and the 45 adapters/ files that were byte-identical to a continual-finetuning adapter (12.3 GB); both lists are in MIGRATION_260908.md of any new repo. Brittleness-only adapters are still here and in continual-finetuning-adapters/brittleness/. Please prefer the new repo for loading.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/brittleness-results.diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md
base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb
chat_examples.pt is the same but for lmsys chat data
chat_base_examples.pt is a merge of the two above files.
All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.unprocessed_dataset_whisper_finetuningS101-base-finetuning-pick-n-placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Pankayaraj/S101-base-finetuning-pick-n-place.record_cube_dataset_for_finetuning_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 20,
"total_frames": 17987,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Chuong/record_cube_dataset_for_finetuning_3.llama3-16k-passkey-retrieval-finetuning2026-06-04-stwebagentbench-suitecrm-demos
2026-06-04-stwebagentbench-suitecrm-demos
Standing demo pool for Adversarial Inverse Constraint RL (ICRL) for LLM
orchestrator safety on ST-WebAgentBench (SuiteCRM easy tier). Every
experiment run consumes this pool; per-run artifacts (embeddings, constraint
heads, adapters, CuP evals) live in separate <date>-<run-name> repos in this
namespace.
field
value
experiment
ICRL safe/unsafe demo pool: constraint C_theta is learned from the safe demos only; unsafe demos are… See the full description on the dataset page: https://huggingface.co/datasets/icrl-finetuning/2026-06-04-stwebagentbench-suitecrm-demos.RAG_vs_FineTuning_Comparison_Persian_V1all_data_for_first_finetuning_shuffled
Dataset Card for "all_data_for_first_finetuning_shuffled"
More Information needed
reddit-travel-QA-finetuningThis dataset was sourced through a series of daily requests to the Reddit API, aiming to capture diverse and real-time travel-related discussions from multiple travel-related subreddits, sourced from this list: https://www.reddit.com/r/travel/comments/1100hca/the_definitive_list_of_travel_subreddits_to_help/, along with subreddits for common travel destinations. Requested was top 100 of the year, this was executed only one, then hot 50 daily. Data aggregation involved concatenating and… See the full description on the dataset page: https://huggingface.co/datasets/soniawmeyer/reddit-travel-QA-finetuning.Ecom-Chatbot-Finetuning-Dataset
Ecom Chatbot Fine-Tuning Dataset
A unified e-commerce chatbot fine-tuning dataset combining 5 source datasets (40,098 examples total), covering product discovery, order management, customer support, returns, and more.
Splits
Split
Source
Examples
amazon_meta
Amazon product metadata
5,000
amazon_reviews
Amazon product reviews
23,100
asos_ecom_dataset
ASOS fashion e-commerce
2,000
bitext_customer_support
Bitext customer support (placeholder-free)
5,000… See the full description on the dataset page: https://huggingface.co/datasets/V1rtucious/Ecom-Chatbot-Finetuning-Dataset.pickplace_cube_finetuningThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100_follower",
"total_episodes": 30,
"total_frames": 4925,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Leyo/pickplace_cube_finetuning.all_data_for_first_finetuning
Dataset Card for "all_data_for_first_finetuning"
More Information needed
llm-finetuning-fr
LLM Fine-Tuning & Quantization - Dataset Francais
Dataset bilingue complet sur le fine-tuning de LLM (LoRA, QLoRA, DPO, RLHF), la quantification de modeles (GPTQ, GGUF, AWQ), les modeles open source et le deploiement en production.
Description
Ce dataset couvre l'ensemble de la chaine de valeur des LLM open source, du fine-tuning au deploiement en production. Il est concu pour servir de reference aux developpeurs, ingenieurs ML, et equipes techniques souhaitant maitriser… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-fr.pickplace_cube_finetuning_v31This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100_follower",
"total_episodes": 30,
"total_frames": 4925,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Leyo/pickplace_cube_finetuning_v31.so101_cube_in_bowl_finetuningv1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 598,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Ricky0626/so101_cube_in_bowl_finetuningv1.eval_act_FineTuning_RedTriangleIntoBoxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 8,
"total_frames": 7082,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/eval_act_FineTuning_RedTriangleIntoBox.finetuning-landscape-paintingThis dataset was compiled for the purpose of finetuning models in the context of benchmarking for art historical research.
Images scraped from Wikimedia Commons via Wikidata; metadata scraped from
Wikidata (CC0). Image licenses vary per file (predominantly public domain,
some CC-BY-SA) see the license_short_name / license_url columns in the
parquet files for the exact terms of each individual image, and the
commons file page for full details.
so101_finetuningThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 8365,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucasfv/so101_finetuning.needle_32k_finetuning_datasetso101_cube_in_bowl_finetuningv4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 2,
"total_frames": 1195,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Ricky0626/so101_cube_in_bowl_finetuningv4.pi0.5_finetuning_small_thin_object_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "fr5_dg5f",
"total_episodes": 560,
"total_frames": 303059,
"total_tasks": 7,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:560"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/simonkim02/pi0.5_finetuning_small_thin_object_dataset.Ecom-Chatbot-Finetuning-Dataset
Ecom Chatbot Finetuning Dataset
A unified instruction-following dataset for fine-tuning e-commerce customer service chatbots. It covers a wide range of real-world retail scenarios — from product discovery and order management to returns, complaints, and account support.
Dataset Summary
Field
Value
Total records
40,098
Language
English
Sources
Amazon Reviews 2023, Amazon Meta 2023, ASOS, Bitext
Response types
Text, Tool Call, Mixed
Difficulty levels
1… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Ecom-Chatbot-Finetuning-Dataset.logd-finetuning-datasetrecord_cube_dataset_for_finetuning_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 20,
"total_frames": 17988,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Chuong/record_cube_dataset_for_finetuning_1.eval_act_FineTuning_RedTriangleIntoBox2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 8021,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/eval_act_FineTuning_RedTriangleIntoBox2.
