datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
duorc
Dataset Card for duorc
Dataset Summary
The DuoRC dataset is an English language dataset of questions and answers gathered from crowdsourced AMT workers on Wikipedia and IMDb movie plots. The workers were given freedom to pick answer from the plots or synthesize their own answers. It contains two sub-datasets - SelfRC and ParaphraseRC. SelfRC dataset is built on Wikipedia movie plots solely. ParaphraseRC has questions written from Wikipedia movie plots and the answers are… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/duorc.duongminh2000dataset_testingapogee
Apogée: Crypto Market Candlestick Dataset
Overview
Most traders believe crypto is random, but deep learning scaling laws suggest otherwise. Apogée is an open-source research initiative exploring the scaling laws of crypto market forecasting. While financial markets are often assumed to be unpredictable, modern deep learning suggests that increasing data and compute could uncover measurable predictability.
Our goal is to quantify how many bits of future price movement… See the full description on the dataset page: https://huggingface.co/datasets/duonlabs/apogee.duobench_raw
DuoBench Raw Dataset
This is the raw dataset collected for the FR3 Duo Benchmark DuoBench.
The format is a raw parquet format that includes additional information such as robot state (real) and sim state (sim) which are filtered out for training.
If you are looking for the converted hugging face dataset ready for training, see this dataset.
You can also manually convert this raw dataset to the hugging face format with the following command:
# install rcs:… See the full description on the dataset page: https://huggingface.co/datasets/RobotControlStack/duobench_raw.Threat_detectionvietnamese-music-dataset
Vietnamese Music Dataset
A collection of 4,820 Vietnamese music tracks with matching cover thumbnails and per-track metadata collected from YouTube, packaged as an audiofolder dataset.
Repository structure
Path
Contents
Count
audio/
MP3 audio files, named by YouTube video ID
4,820
images/
PNG cover thumbnails, same IDs as audio/
4,820
data/
Parquet metadata files, one per collection session
31
Metadata schema
Each Parquet file in… See the full description on the dataset page: https://huggingface.co/datasets/Toan-Minh-Duong-Son/vietnamese-music-dataset.vi-dataset-for-pretrain
Dataset Card for "vi-dataset-for-pretrain"
This is a combination of multiple Vietnamese dataset for pretraining CLMs such as GPT, GPT2, etc.
The dataset consists of:
vietgpt/covid_19_news_vi
hieunguyen1053/binhvq-news-corpus
oscar (unshuffled_deduplicated_vi)
vietgpt/wikipedia_vi
Dataset info
Splits
N.o examples
Size
Train
23,891,116
77.36 GB
Validation
1,257,428
4.06 GB
Total
25,148,544
81.43 GB
gemma2b-summarize-eval-by-claude3sonnetgemma2b-summarize-eval-by-gemini15flashduobench
DuoBench LeRobot Dataset
This is the dataset collected for the FR3 Duo Benchmark DuoBench in LeRobot format ready for training.
The dataset is a stripped down version from the raw dataset and only contains data required for VLA training.
Checkout the raw dataset if you are interested e.g. in franka states (real) and sim states.
The raw dataset repo also shows how to convert the data into this lerobot version.
Project page: https://duobench.github.io/
DuoBench Code:… See the full description on the dataset page: https://huggingface.co/datasets/RobotControlStack/duobench.gemma7b-summarize-eval-by-gemini15flashgemma7b-summarize-eval-by-claude3sonnetwan-video-encoded-cartsynth_summarize_datasetXD_violence_datasetwan-images-feat-cartduobench_modify
DuoBench Modify
duobench_modify is an unofficial derivative of
RobotControlStack/duobench.
It combines the 11 DuoBench simulation subsets into one LeRobot v3 dataset,
preserves the original joint-space fields, and adds 20-dimensional
end-effector (EEF) state and action fields for TwinVLA-style training.
Scope: this repository contains simulation data only. The four upstream
real-robot subsets are not included.
Dataset summary
Property
Value
Episodes… See the full description on the dataset page: https://huggingface.co/datasets/kisarakira/duobench_modify.DuoduoCLIP-data
Dataset Card for DuoduoCLIP
In this data repo we provide the data used in the paper Duoduo CLIP: Efficient 3D Understanding with Multi-View Images.
The data usage and code can be found in the github repo.
Note: We provide the lvis evaluation data in the initial release, we will soon upload the data and process scripts required for training.
Dataset Details
Dataset Sources
Multi-view images of 3D objects were used in training our models.
A majority of… See the full description on the dataset page: https://huggingface.co/datasets/3dlg-hcvc/DuoduoCLIP-data.wan-video-encodedwan-video-encoded-newdetails_duoqi__Nanbeige-16B-Base-Llama
Dataset Card for Evaluation run of duoqi/Nanbeige-16B-Base-Llama
Dataset automatically created during the evaluation run of model duoqi/Nanbeige-16B-Base-Llama on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_duoqi__Nanbeige-16B-Base-Llama.Dataset_audio_threat_detectionsynth_summarize_dataset_deduptiny-story-shuffle https://github.com/duoduoyeah/nanochat.git
Original from roneneldan/TinyStories and SimpleStories/SimpleStories
Go through raw data treat: dev/repackage_data_reference.py
train shard 250M chars per shard
validation shard 2.5M chars per shard
duo
Dataset README
This repository provides all datasets and model checkpoints required for the paper:
Adaptive Dual Uncertainty Optimization: Boosting Monocular 3D Object Detection under Test-Time Shifts 🔗Zixuan Hu, Dongxiao Li, Xinzhu Ma, Shixiang Tang, Xiaotong Li, Wenhan Yang, Ling-Yu Duan(ICCV 2025 Highlight)
Repository Structure
kitti_noise.zip, kitti_blur.zip, kitti_weather.zip, and kitti_digital.zip
After extraction, together they form the complete KITTI-C dataset.… See the full description on the dataset page: https://huggingface.co/datasets/hzcar/duo.simple-story-shuffle https://github.com/duoduoyeah/nanochat.git
Original from roneneldan/TinyStories and SimpleStories/SimpleStories
Go through raw data treat: dev/repackage_data_reference.py
train shard 250M chars per shard
validation shard 2.5M chars per shard
WebApp1K-Duo-React
Paper: https://huggingface.co/papers/2409.13773
duosynth_classification_dataset_dedup
