datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
double-asteriskbangla-ocr-double-benchmark
Bangla OCR Double Benchmark
Two equally weighted, deterministic full-page Bangla handwriting robustness splits:
bongabdo: 6,669 readability-preserving renderings balanced over all 111 Bongabdo pages.
bn_htrd: 6,669 renderings balanced over all 75 actual files in the writer-separated
BN-HTRd test split.
These are explicitly compositional/augmentation robustness rows, not 13,338 independent
writers or source documents. Every row exposes its source page ID, source SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/bangla-ocr-double-benchmark.Double-Bench
Double-Bench: A Multilingual & Multimodal Evaluation System for Document RAG
We introduce Double-Bench, a new large-scale, multilingual, and multimodal evaluation system for assessing Retrieval-Augmented Generation (RAG) systems using Multimodal Large Language Models (MLLMs).
The dataset and benchmark were introduced in the paper Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?.
Project Page: https://double-bench.github.io/
Code Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Episoode/Double-Bench.go-mo-dataset
GO-MO, a massive Graph agumented Open urban MObility dataset
This is the official dataset repository for the GO-MO traffic dataset.
The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain).
GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).
Additionally, the GO-MO dataset introduces two graph… See the full description on the dataset page: https://huggingface.co/datasets/double-blind-anonymous/go-mo-dataset.double_pendulumDataset from 2022-07-15. There are two columns: image and trajectory. Image is the image file and trajectory is the label trajectory number_image number for the corresponding image.
persian-ocr-double-benchmark
Persian OCR Double Benchmark
A frozen, leakage-controlled Persian OCR benchmark with two equally weighted splits:
printed: 6,669 rows carved from Reza2kn/persian-printed-ocr-3.5m at ba02f36c3d496838d8fad9aff352b77763af1ea4.
handwriting: 6,669 rows carved from Reza2kn/persian-handwriting-pages-3.69m at b114f0a36a6a2e397bc93517dd084431bfca2329.
The exact rows were uploaded here before being removed from their source training repositories.
Each row retains its original repository… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-double-benchmark.pick-double-caption
Dual Caption Preference Optimization for Diffusion Models
We propose DCPO, a new paradigm to improve the alignment performance of text-to-image diffusion models. For more details on the technique, please refer to our paper here.
Developed by
Amir Saeidi*
Yiran Luo*
Agneet Chatterjee
Shamanthak Hegde
Bimsara Pathiraja
Yezhou Yang
Chitta Baral
Dataset
This dataset is Pick-Double Caption, a modified version of the Pick-a-Pic V2 dataset. We… See the full description on the dataset page: https://huggingface.co/datasets/DualCPO/pick-double-caption.wds-double-stars
Washington Double Star Catalog
Credit: NASA/ESA/Hubble
Part of a dataset collection on Hugging Face.
Dataset description
The Washington Double Star Catalog (WDS) is the world reference catalog for visual double and multiple star systems, maintained by the US Naval Observatory.
Double stars are essential for determining stellar masses -- the most fundamental property of a star -- and for testing stellar evolution models. The WDS traces its lineage back to… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/wds-double-stars.ogbench-block-double-hermite-100k
ogbench-block-double-hermite-100k
Unofficial reproduction of the scripted policies described in https://seohong.me/blog/behavioral-cloning-mystery/
using random piecewise Hermite splines as the backbone, and randomized control points, grasping angles/directions,
contact points, motion speed, gripper yaw/roll/pitch, mistakes and retries, etc.
This might not be the exact setup used by the study, but I tried to infer the parameters from
"How exactly did you script the policies?"… See the full description on the dataset page: https://huggingface.co/datasets/Yassine/ogbench-block-double-hermite-100k.double_perovskite_bandgap_v1-1
Machine learning bandgaps of double perovskites
Dataset containing DFT-calculated band gaps of 1306 double perovskite oxide materials
Dataset Information
Source: Foundry-ML
DOI: 10.18126/lss6-o5x4
Year: 2022
Authors: Pilania, G., Mannodi-Kanakkithodi, A., Uberuaga, B. P., Ramprasad, R., Gubernatis, J. E., Lookman, T.
Data Type: tabular
Fields
Field
Role
Description
Units
formula
input
Material composition
a_1
input
Element 1 on the A… See the full description on the dataset page: https://huggingface.co/datasets/foundry-ml/double_perovskite_bandgap_v1-1.dclm-100x-tsp-v2-doubleso101_double_4cam_rl_total_0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 501,
"total_frames": 299352,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:501"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Qiu-Xinchuan/so101_double_4cam_rl_total_0.harmless-eval-SUDO-random-double-spaceHeraiHench__Double-Down-Qwen-Math-7B-details
Dataset Card for Evaluation run of HeraiHench/Double-Down-Qwen-Math-7B
Dataset automatically created during the evaluation run of model HeraiHench/Double-Down-Qwen-Math-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Double-Down-Qwen-Math-7B-details.Halide_Double_Perovskite_Octahedral_Tilting
Cite this dataset Baskurt, M., Fransson, E., Lindvik, M., Erhart, P., and Wiktor, J. Halide Double Perovskite Octahedral Tilting. ColabFit, 2026. https://doi.org/None
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_zqofv7mcsq1v_0
Visit the ColabFit Exchange to search additional datasets by author, description, element… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Halide_Double_Perovskite_Octahedral_Tilting.social-orientation
Dataset Card for Social Orientation
There are many settings where it is useful to predict and explain the success or failure of a dialogue. Circumplex theory from psychology models the social orientations (e.g., Warm-Agreeable, Arrogant-Calculating) of conversation participants, which can in turn can be used to predict and explain the outcome of social interactions, such as in online debates over Wikipedia page edits or on the Reddit ChangeMyView forum.
This dataset contains social… See the full description on the dataset page: https://huggingface.co/datasets/tee-oh-double-dee/social-orientation.double-underscoreamdi_double_strat
Ambiguous Words Diachronic Dataset
This is a stratified sample of the double annotated part of the AmDi-Dataset.
llm-ground-truth-general-fix-double-BOS
llm-ground-truth-general-fix-double-BOS
Merged dataset assembled from:
elichen-skymizer/llm-ground-truth-general-fix-BOS: all subsets.
skymizer/llm-ground-truth-general: Qwen3 subsets excluding transformers variants.
MULTI_VALUE_cola_double_modals
Dataset Card for "MULTI_VALUE_cola_double_modals"
More Information needed
MULTI_VALUE_mrpc_double_modals
Dataset Card for "MULTI_VALUE_mrpc_double_modals"
More Information needed
double_exposureMULTI_VALUE_mnli_double_modals
Dataset Card for "MULTI_VALUE_mnli_double_modals"
More Information needed
MULTI_VALUE_sst2_double_modals
Dataset Card for "MULTI_VALUE_sst2_double_modals"
More Information needed
american-double-quotesMULTI_VALUE_mnli_double_superlative
Dataset Card for "MULTI_VALUE_mnli_double_superlative"
More Information needed
MULTI_VALUE_wnli_double_modals
Dataset Card for "MULTI_VALUE_wnli_double_modals"
More Information needed
MULTI_VALUE_stsb_double_modals
Dataset Card for "MULTI_VALUE_stsb_double_modals"
More Information needed
task306_jeopardy_answer_generation_double
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task306_jeopardy_answer_generation_double
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task306_jeopardy_answer_generation_double.MD22_double_walled_nanotube
Cite this dataset Chmiela, S., Vassilev-Galindo, V., Unke, O. T., Kabylda, A., Sauceda, H. E., Tkatchenko, A., and Müller, K. MD22 double walled nanotube. ColabFit, 2023. https://doi.org/10.60732/fce214af
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_d92bnafzwynr_0
Visit the ColabFit Exchange to search additional datasets by author… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/MD22_double_walled_nanotube.
