datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.databricks-dolly-15k-curated-en
Guidelines
In this dataset, you will find a collection of records that show a category, an instruction, a context and a response to that instruction. The aim of the project is to correct the instructions, intput and responses to make sure they are of the highest quality and that they match the task category that they belong to. All three texts should be clear and include real information. In addition, the response should be as complete but concise as possible.
To curate the dataset… See the full description on the dataset page: https://huggingface.co/datasets/argilla/databricks-dolly-15k-curated-en.databricks_dolly_15k
Dataset Card for Dolly_15K
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/databricks_dolly_15k.dolly-15k-oai-style
Dataset Card for "dolly-15k-oai-style"
More Information needed
dollsfrontline
Bangumi Image Base of Dolls' Frontline
This is the image base of bangumi Dolls' Frontline, we detected 76 characters, 2746 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/dollsfrontline.dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/dolly-vn/dolly-audio-1000h-vietnamese.G1edu-u3_plate_storage_doll
G1edu-u3_plate_storage_doll
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: unitree_g1
| Codebase Version: v2.1
End-Effector Type: three_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric
Value
Total Episodes
388… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/G1edu-u3_plate_storage_doll.dolly-chatml-sftdollarstreet
Dataset Card for "dollarstreet"
More Information needed
dolly_hhrlhf
Dataset Card for "dolly_hhrlhf"
This dataset is a combination of Databrick's dolly-15k dataset and a filtered subset of Anthropic's HH-RLHF. It also includes a test split, which was missing in the original dolly set. That test set is composed of 200 randomly selected samples from dolly + 4,929 of the test set samples from HH-RLHF which made it through the filtering process. The train set contains 59,310 samples; 15,014 - 200 = 14,814 from Dolly, and the remaining 44,496 from… See the full description on the dataset page: https://huggingface.co/datasets/mosaicml/dolly_hhrlhf.G1edu-u3_plate_storage_rabbit_doll
G1edu-u3_plate_storage_rabbit_doll
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: unitree_g1
| Codebase Version: v2.1
End-Effector Type: three_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric
Value
Total Episodes… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/G1edu-u3_plate_storage_rabbit_doll.databricks-dolly-15k-ja
This dataset was created by automatically translating "databricks-dolly-15k" into Japanese.This dataset is licensed under CC-BY-SA-3.0
Last Update : 2023-05-11
databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data
dolly_creative_writing
Dataset Card for "dolly_creative_writing"
More Information needed
atlas-28-explore-cost-in-dollars
ATLAS report 28: scaling the training up from step 29
1. Question and links
Read this first. This data root is published whole to the Hugging Face repository t2ance/atlas-28-explore-cost-in-dollars and, without the saved steps, the weight files and the per-token arrays, as the directory 28-explore-cost-in-dollars/ of the GitHub reading copy t2ance/atlas-experiments. The saved training steps are on the Hub only.
Question. Can a larger-scale training be brought up… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-28-explore-cost-in-dollars.databricks-dolly-15k-curated-multilingual
Dataset Card for "databricks-dolly-15k-curated-multilingual"
A curated and multilingual version of the Databricks Dolly instructions dataset. It includes a programmatically and manually corrected version of the original en dataset. See below.
STATUS:
Currently, the original Dolly v2 English version has been curated combining automatic processing and collaborative human curation using Argilla (~400 records have been manually edited and fixed). The following graph shows a summary… See the full description on the dataset page: https://huggingface.co/datasets/argilla/databricks-dolly-15k-curated-multilingual.Airbot_MMK2_storage_penguin_doll_tiger_doll
Airbot_MMK2_storage_penguin_doll_tiger_doll
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 100
Total Frames: 13342
FPS: 30
Dataset Size: 424.08 MB
Robot Name: Airbot_MMK2
End-Effector Type: five_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Airbot_MMK2_storage_penguin_doll_tiger_doll.aloha_real_agilex_nesting_dolldolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/fnooub/dolly-audio-1000h-vietnamese.egopi_latal_openarm_dollbinance_btcusdt_60M_dollar_klines
BTCUSDT 60M dollar spot klines
This dataset is exported daily from origo.binance_spot_dollar_klines using 60M-dollar dollar bars derived from the 1M Origo dollar-kline foundation.
Latest snapshot:
file: btcusdt_60M_dollar_kline_20200101_to_20260923.parquet
start date: 2020-01-01
rows: 93417
end date: 2026-09-23
columns: start_datetime, end_datetime, dollar_bar_id, open, high, low, close, mean, std, volume, maker_ratio, no_of_trades, open_liquidity, high_liquidity, low_liquidity… See the full description on the dataset page: https://huggingface.co/datasets/vaquum/binance_btcusdt_60M_dollar_klines.binance_btcusdt_15M_dollar_klines
BTCUSDT 15M dollar spot klines
This dataset is exported daily from origo.binance_spot_dollar_klines using 15M-dollar dollar bars derived from the 1M Origo dollar-kline foundation.
Latest snapshot:
file: btcusdt_15M_dollar_kline_20200101_to_20260923.parquet
start date: 2020-01-01
rows: 369960
end date: 2026-09-23
columns: start_datetime, end_datetime, dollar_bar_id, open, high, low, close, mean, std, volume, maker_ratio, no_of_trades, open_liquidity, high_liquidity… See the full description on the dataset page: https://huggingface.co/datasets/vaquum/binance_btcusdt_15M_dollar_klines.dolly
Dolly Dataset
Original version: aisquared/databricks-dolly-15k.
This dataset is used to train MiniLLM.
binance_btcusdt_120M_dollar_klines
BTCUSDT 120M dollar spot klines
This dataset is exported daily from origo.binance_spot_dollar_klines using 120M-dollar dollar bars derived from the 1M Origo dollar-kline foundation.
Latest snapshot:
file: btcusdt_120M_dollar_kline_20200101_to_20260923.parquet
start date: 2020-01-01
rows: 47329
end date: 2026-09-23
columns: start_datetime, end_datetime, dollar_bar_id, open, high, low, close, mean, std, volume, maker_ratio, no_of_trades, open_liquidity, high_liquidity… See the full description on the dataset page: https://huggingface.co/datasets/vaquum/binance_btcusdt_120M_dollar_klines.binance_btcusdt_1M_dollar_klines
BTCUSDT 1M dollar spot klines
This dataset is exported daily from origo.binance_spot_dollar_klines using 1M-dollar dollar bars derived from the 1M Origo dollar-kline foundation.
Latest snapshot:
file: btcusdt_1M_dollar_kline_20200101_to_20260923.parquet
start date: 2020-01-01
rows: 5526986
end date: 2026-09-23
columns: start_datetime, end_datetime, dollar_bar_id, open, high, low, close, mean, std, volume, maker_ratio, no_of_trades, open_liquidity, high_liquidity, low_liquidity… See the full description on the dataset page: https://huggingface.co/datasets/vaquum/binance_btcusdt_1M_dollar_klines.binance_btcusdt_240M_dollar_klines
BTCUSDT 240M dollar spot klines
This dataset is exported daily from origo.binance_spot_dollar_klines using 240M-dollar dollar bars derived from the 1M Origo dollar-kline foundation.
Latest snapshot:
file: btcusdt_240M_dollar_kline_20200101_to_20260923.parquet
start date: 2020-01-01
rows: 24302
end date: 2026-09-23
columns: start_datetime, end_datetime, dollar_bar_id, open, high, low, close, mean, std, volume, maker_ratio, no_of_trades, open_liquidity, high_liquidity… See the full description on the dataset page: https://huggingface.co/datasets/vaquum/binance_btcusdt_240M_dollar_klines.binance_btcusdt_30M_dollar_klines
BTCUSDT 30M dollar spot klines
This dataset is exported daily from origo.binance_spot_dollar_klines using 30M-dollar dollar bars derived from the 1M Origo dollar-kline foundation.
Latest snapshot:
file: btcusdt_30M_dollar_kline_20200101_to_20260923.parquet
start date: 2020-01-01
rows: 185604
end date: 2026-09-23
columns: start_datetime, end_datetime, dollar_bar_id, open, high, low, close, mean, std, volume, maker_ratio, no_of_trades, open_liquidity, high_liquidity… See the full description on the dataset page: https://huggingface.co/datasets/vaquum/binance_btcusdt_30M_dollar_klines.Airbot_MMK2_move_sword_doll
Airbot_MMK2_move_sword_doll
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 48
Total Frames: 5383
FPS: 30
Dataset Size: 226.25 MB
Robot Name: Airbot_MMK2
End-Effector Type: five_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type information.
Sensors:… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Airbot_MMK2_move_sword_doll.Airbot_MMK2_storage_shark_doll
Airbot_MMK2_storage_shark_doll
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 93
Total Frames: 13034
FPS: 30
Dataset Size: 461.32 MB
Robot Name: Airbot_MMK2
End-Effector Type: five_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type information.
Sensors:… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Airbot_MMK2_storage_shark_doll.Airbot_MMK2_take_dog_doll
Airbot_MMK2_take_dog_doll
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 50
Total Frames: 6425
FPS: 30
Dataset Size: 253.22 MB
Robot Name: Airbot_MMK2
End-Effector Type: five_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type information.
Sensors:… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Airbot_MMK2_take_dog_doll.Airbot_MMK2_storage_gold_bar_model_shark_doll
Airbot_MMK2_storage_gold_bar_model_shark_doll
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 44
Total Frames: 5620
FPS: 30
Dataset Size: 240.37 MB
Robot Name: Airbot_MMK2
End-Effector Type: five_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Airbot_MMK2_storage_gold_bar_model_shark_doll.
