datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trivia_qa
Dataset Card for "trivia_qa"
Dataset Summary
TriviaqQA is a reading comprehension dataset containing over 650K
question-answer-evidence triples. TriviaqQA includes 95K question-answer
pairs authored by trivia enthusiasts and independently gathered evidence
documents, six per question on average, that provide high quality distant
supervision for answering the questions.
Supported Tasks and Leaderboards
More Information Needed
Languages… See the full description on the dataset page: https://huggingface.co/datasets/mandarjoshi/trivia_qa.clinical-trials-protocolssoft-trigger-verifiedtripadvisor-hotel-reviews
Dataset Card for "tripadvisor-hotel-reviews"
Dataset Summary
Hotels play a crucial role in traveling and with the increased access to information new pathways of selecting the best ones emerged.
With this dataset, consisting of 20k reviews crawled from Tripadvisor, you can explore what makes a great hotel and maybe even use this model in your travels!
Citations on a scale from 1 to 5.
Languages
english
Citation Information
If you use this dataset in… See the full description on the dataset page: https://huggingface.co/datasets/argilla/tripadvisor-hotel-reviews.based_triviaqawiki-trivia-questions-v4amazon-reviews-2023-trimmed
Amazon Product Reviews 2023 (Trimmed, 34 Categories)
This dataset is a trimmed and restructured version of the Amazon Product Reviews 2023 dataset by Julian McAuley and the UCSD Computer Science department.
It includes 34 product categories, each stored as a folder containing multiple sharded Parquet files for scalable access.
Only three fields are retained:
rating — The numerical review score (originally overall)
title — The review title (from summary)
text — The full review body… See the full description on the dataset page: https://huggingface.co/datasets/bagadbilla/amazon-reviews-2023-trimmed.twi-trigrams-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.dvla-can-250hz-events-250fps-delta-trimThis dataset was created using LeRobot.
Dataset Description
dvla-can-250hz-events (250 fps native) - event frames accumulating 4 ms each
1016 episodes / 1,356,080 frames at the native 250 Hz rate. Same v2e settings as the 250->25 fps
variant, but kept at full rate: each event frame covers 4 ms, so the polarity images are sparser
(roughly a third of the pixels of the 40 ms version).
Two extra video columns on top of the 3 RGB cameras:… See the full description on the dataset page: https://huggingface.co/datasets/mickeykang/dvla-can-250hz-events-250fps-delta-trim.mocap-infra-0828-trial22223
XINGYING RGB-D Tactile Cylinder Trial 0828
This public dataset contains a recovered 201-second multimodal robotics capture
recorded with XINGYING/NOKOV motion capture, an Intel RealSense D435i, and one
left DM-Tac tactile sensor. The original XINGYING CAP and companion directory
are retained beside inspectable Parquet, Zarr, MP4, CSV, and HDF5 derivatives.
Important quality status
This is a useful recovered capture, not a clean benchmark episode.… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz045/mocap-infra-0828-trial22223.tripadvisor-review-rating
TripAdvisor Easy Dataset
This repository contains a dataset of hotel reviews and ratings collected from TripAdvisor, which has been processed by us. The dataset includes reviews of various hotels along with metadata such as multiple-aspect ratings and review texts.
Please refer to our GitHub repohttps://github.com/jniimi/tripadvisor_dataset.
The data is originally distributed by Jiwei Li et al. (2013) and is hosted on his website… See the full description on the dataset page: https://huggingface.co/datasets/jniimi/tripadvisor-review-rating.olm-october-2022-tokenized-128
Dataset Card for "olm-october-2022-tokenized-128"
More Information needed
medical-symptom-triage-conversationaltrigger_datasetfineweb-2-trimming
Description
Version of FineWeb2 where only 124 languages were kept.For each of them we kept the first 200,000 texts (less if there are not as many available for a given language).
The purpose of this dataset is to offer a light version (only 44GB against 8.67 TB for the original dataset) in order to be able to trim models.
For more information on the trimming method, we invite you to consult this blog post.
Citations
FineWeb-2… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/fineweb-2-trimming.reddit_dataset_145
Bittensor Subnet 13 Reddit Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/Trimness8/reddit_dataset_145.clinical-trials
Clinical Trials Dataset
A comprehensive dataset of clinical trials sourced from ClinicalTrials.gov, featuring structured metadata, detailed study information, and pre-computed semantic embeddings for machine learning applications in biomedical research.
Dataset Description
This dataset provides a rich collection of clinical trial information systematically collected from the official ClinicalTrials.gov database. Each record contains detailed study metadata, eligibility… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/clinical-trials.jetson1-060426-test-collection-v1-trim
jetson1-060426-test-collection-v1-trim
Materialized collection — 3 episodes · 407 frames @ 20 fps (~0 min of demonstration).
Collection jetson1-060426-test-collection@v1 (frozen 2026-06-05), mode trim_rebuild — built by Vibe Data Studio; the machine-readable recipe in meta/vibedata/collection.json makes this dataset reproducible from its pinned components.
Tasks
Instruction
Episodes
danse
3
Episodes per task per recording rig (derived by walking… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060426-test-collection-v1-trim.cucumber-place-DAgger-iter2-doris070926-v1-trim
cucumber-place-DAgger-iter2-doris070926-v1-trim
Materialized collection — 117 episodes · 4,677 frames @ 20 fps (~4 min of demonstration).
Collection cucumber-place-DAgger-iter2-doris070926@v1 (frozen 2026-07-09), mode trim_rebuild — built by Vibe Data Studio; the machine-readable recipe in meta/vibedata/collection.json makes this dataset reproducible from its pinned components.
Tasks
Instruction
Episodes
place the cucumber on the middle of the cutting… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-DAgger-iter2-doris070926-v1-trim.cucumber-place-DAgger-iter1-doris070726-v1-trim
cucumber-place-DAgger-iter1-doris070726-v1-trim
Materialized collection — 93 episodes · 3,546 frames @ 20 fps (~3 min of demonstration).
Collection cucumber-place-DAgger-iter1-doris070726@v1 (frozen 2026-07-08), mode trim_rebuild — built by Vibe Data Studio; the machine-readable recipe in meta/vibedata/collection.json makes this dataset reproducible from its pinned components.
Tasks
Instruction
Episodes
place the cucumber on the middle of the cutting… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-DAgger-iter1-doris070726-v1-trim.cucumber-place-DAgger-iter0-trim-doris070126
cucumber-place-DAgger-iter0-trim-doris070126
Materialized collection — 87 episodes · 6,098 frames @ 20 fps (~5 min of demonstration).
Collection cucumber-place-DAgger-iter0-trim-doris070126@v1 (frozen 2026-07-01), mode trim_rebuild — built by Vibe Data Studio; the machine-readable recipe in meta/vibedata/collection.json makes this dataset reproducible from its pinned components.
Tasks
Instruction
Episodes
place the cucumber on the middle of the cutting… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-DAgger-iter0-trim-doris070126.cucumber-subtask-grab-DAgger-iter1-v1-trim
cucumber-subtask-grab-DAgger-iter1-v1-trim
Materialized collection — 100 episodes · 9,700 frames @ 20 fps (~8 min of demonstration).
Collection cucumber-subtask-grab-DAgger-iter1@v1 (frozen 2026-06-24), mode trim_rebuild — built by Vibe Data Studio; the machine-readable recipe in meta/vibedata/collection.json makes this dataset reproducible from its pinned components.
Tasks
Instruction
Episodes
grab
55
grab the cucumber close to one of the cucumber's… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-subtask-grab-DAgger-iter1-v1-trim.dvla-can-250hz-events-250to25fps-delta-trimThis dataset was created using LeRobot.
Dataset Description
dvla-can-250hz-events (250 -> 25 fps) - event frames accumulating 40 ms each
1016 episodes / 135,828 frames. Same scenes as the 250 Hz RGB set, rendered at 250 Hz, passed
through the v2e DVS emulator (1.5.1, thresholds 0.15/0.15,
slow-motion interpolation disabled because the frames are genuinely 250 Hz), then downsampled
10x: each output frame carries the events from a 40 ms window as a polarity image.… See the full description on the dataset page: https://huggingface.co/datasets/mickeykang/dvla-can-250hz-events-250to25fps-delta-trim.cucumber-peel-DAgger-iter1-adaptive1-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "vibeboard_follower_tilt",
"total_episodes": 78,
"total_frames": 28053,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:78"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-peel-DAgger-iter1-adaptive1-v1-trim.encoder-decoder-trial-stat
Encoder/decoder trial: encoder-marginal report
Dataset: G-reen/encoder-decoder-trial-stat
Rows analysed: 122,933 (every kept (encoder, decoder, source row) triple; source G-reen/cc-re-2021-filtered shard 0, 2000 rows of at most 4000 words)
Prompt file: prompts/indirect_reference_dataset_train.json (turn 0 encodes the document, turn 1 reconstructs it from the encoding alone)
Encoders: 9 (granite-4.2-30b-nvfp4 [0], Ornith-1.5-35B-A3B-NVFP4 [1], Llama-3.3-70B-Instruct-NVFP4 [2]… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/encoder-decoder-trial-stat.cucumber-peel-DAgger-iter1-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "vibeboard_follower_tilt",
"total_episodes": 70,
"total_frames": 26265,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:70"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-peel-DAgger-iter1-v1-trim.recreate-bug-pre-fix-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/recreate-bug-pre-fix-v1-trim.recreate-bug-post-fix-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/recreate-bug-post-fix-v1-trim.cucumber-subtask-grab-DAgger-iter1-v1-trim_auxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-subtask-grab-DAgger-iter1-v1-trim_aux.cucumber-place-classifier-eval071526-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "vibeboard_follower_tilt",
"total_episodes": 74,
"total_frames": 2908,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:74"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-eval071526-v1-trim.
