datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/ctmedtech/DDR-dataset.super_eurlexSuper-EURLEX dataset containing legal documents from multiple languages.
The datasets are build/scrapped from the EURLEX Website [https://eur-lex.europa.eu/homepage.html]
With one split per language and sector, because the available features (metadata) differs for each
sector. Therefore, each sample contains the content of a full legal document in up to 3 different
formats. Those are raw HTML and cleaned HTML (if the HTML format was available on the EURLEX website
during the scrapping process) and cleaned text.
The cleaned text should be available for each sample and was extracted from HTML or PDF.
'Cleaned' HTML stands here for minor cleaning that was done to preserve to a large extent the necessary
HTML information like table structures while removing unnecessary complexity which was introduced to the
original documents due to actions like writing each sentence into a new object.
Additionally, each sample contains metadata which was scrapped on the fly, this implies the following
2 things. First, not every sector contains the same metadata. Second, most metadata might be
irrelevant for most use cases.
In our minds the most interesting metadata is the celex-id which is used to identify the legal
document at hand, but also contains a lot of information about the document
see [https://eur-lex.europa.eu/content/tools/eur-lex-celex-infographic-A3.pdf] as well as eurovoc-
concepts, which are labels that define the content of the documents.
Eurovoc-Concepts are, for example, only available for the sectors 1, 2, 3, 4, 5, 6, 9, C, and E.
The Naming of most metadata is kept like it was on the eurlex website, except for converting
it to lower case and replacing whitespaces with '_'.math_text
Mathematical Texts (MT)
Mathematical dataset containing mathematical texts, i.e. texts containing LaTeX formulas, based on the AMPS Khan dataset and the ARQMath dataset V1.3. Based on the retrieved LaTeX texts, more mathematically equivalent versions have been generated by applying randomized LaTeX printing with this SymPy fork using Math Mutator (MAMUT). A positive id corresponds to the ARQMath post id of the generated text version, a negative id indicates an AMPS text.
You can… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/math_text.DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/Aljo-na/DDR-dataset.DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/chanchahh/DDR-dataset.math_formula_retrieval
Dataset Card for MFR (Mathematical Formula Retrieval)
This dataset consists of formula pairs, classified as either mathematical equivalent or not.
Dataset Details
Dataset Description
Mathematical dataset based on 71 famous mathematical identities. Each entry consists of two identities (in formula or textual form), together with a label, whether the two versions describe the same mathematical identity. The false pairs are not randomly chosen, but… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/math_formula_retrieval.named_math_formulas
Dataset Card for NMF (Named Mathematical Formulas)
This dataset might be used to train a language model based math-retrieval system with a NSP-like task.
See also see ddrg/math_formula_retrieval as a derived dataset which associates two formulas.
You can find more information in MAMUT: A Novel Framework for Modifying Mathematical Formulas for the Generation of Specialized Datasets for Language Model Training.
Named Math Formulas
Mathematical dataset based on 71… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/named_math_formulas.DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/jerrodz/DDR-dataset.DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/RONIE2002/DDR-dataset.MUSES
MUSES: a benchmark for Marked Unevenly Spaced Event Sequences
MUSES, a benchmark for Marked Unevenly Spaced Event Sequences, is a collection of unevenly spaced time series datasets from various domains, containing marked events for training and evaluating prediction approaches.
Languages
All the columns and classes (when textual) in MUSES are in English (BCP-47 en)
Dataset Structure
Data Instances
All datasets formatting follows… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/MUSES.TOTRCD
TOTRCD: Temporally-Ordered Tabular Regression Benchmark Suite with Concept Drift
TOTRCD is a collection of tabular regression datasets that include a temporal ordering, represented by a time column, and exhibit some form of concept drift.
Dataset Structure
Data Fields
The columns of the datasets follows all a similar formatting:
Meta::: prefix marking identifiers, timestamps, or sort keys that are not used as model inputs. Example: Meta::DateTime.… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/TOTRCD.wod-e2e-fast-ddrive-sasd-50k
WOD-E2E Fast-dDrive SASD 50k — TPU-ready training shards
50,331 frames (32 of the 263 WOD-E2E train shards) pre-tokenized into the Fast-dDrive
Section-Aware Structured Diffusion (SASD) format, packaged as 787 Apache Parquet shards
(~22 GB) for multi-host TPU training (grain / datasets / MaxText hf data path). The
TPU side needs no tokenizer, processor, or torch — it loads arrays and trains.
⚠️ PRIVATE / license-restricted. Derived from the Waymo Open Dataset; shared privately… See the full description on the dataset page: https://huggingface.co/datasets/kaiwen2/wod-e2e-fast-ddrive-sasd-50k.DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/MsNeeraj/DDR-dataset.math_formulas
Mathematical Formulas (MF)
Mathematical dataset containing formulas based on the AMPS Khan dataset and the ARQMath dataset V1.3. Based on the retrieved LaTeX formulas, more equivalent versions have been generated by applying randomized LaTeX printing with this SymPy fork using Math Mutator (MAMUT). The formulas are intended to be well applicable for MLM. For instance, a masking for a formula like (a+b)^2 = a^2 + 2ab + b^2 makes sense (e.g., (a+[MASK])^2 = a^2 + [MASK]ab + b[MASK]2… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/math_formulas.DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for classification… See the full description on the dataset page: https://huggingface.co/datasets/kasanii/DDR-dataset.named_math_formulas_ft
Named Math Formulas - Fine-Tuning Dataset
This dataset is a version of Named Math Formulas (NMF) dedicated to be used as fine-tuning dataset, e.g., by using the GitHub Repository aieng-lab/transformer-math-evaluation.
In contrast to the full version of NMF, this dataset contains much more meta data that can be used for fine-grained evaluations.
Dataset Details
Dataset Description
Mathematical dataset based on 71 famous mathematical identities. Each… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/named_math_formulas_ft.DDR_dataset_train_test
🐨 DR Classification Fundus Dataset
This dataset contains retinal fundus images labeled for Diabetic Retinopathy (DR) classification, intended for use in machine learning tasks such as image classification and medical diagnosis support.
📁 Dataset Structure
The dataset is stored under the splits/ directory and is divided into:
train and test folders with train and test.csv correspondingly
DDR-DatasetDDRBench_10K_trajectory
10K Agent Trajectories Dataset
Project Page | Paper | Code
Overview
This dataset contains agent trajectories from the Deep Data Research (DDR) project's 10-K financial analysis task, as presented in the paper "Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models".
DDR-Bench is a large-scale benchmark designed to evaluate "investigatory intelligence" in LLM agents—the autonomy to set goals and explore raw data without explicit queries. This… See the full description on the dataset page: https://huggingface.co/datasets/thinkwee/DDRBench_10K_trajectory.DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/Achyuth12/DDR-dataset.InsuranceCorpusDDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/addone5/DDR-dataset.ddr5-ram-prices-raw-dataset-2026
19,705 raw U.S. DDR5 RAM price observations across 12 ZIP markets and 29 days.
DDR5 RAM Prices Raw Dataset (2026)
Analyze 19,705 unaggregated product-level listed retail prices for standalone DDR5 memory modules and homogeneous kits across 12 U.S. ZIP markets from July 13 through August 10, 2026. The single analysis-ready CSV preserves titles, dates, geography, package quantities, listed prices, and a source-neutral comparable-price field.
What “raw” means here: unaggregated… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/ddr5-ram-prices-raw-dataset-2026.wod-e2e-fast-ddrive-sasd
WOD-E2E Fast-dDrive SASD — TPU-ready training shards
Pre-tokenized Section-Aware Structured Diffusion (SASD) training samples for the
Fast-dDrive Qwen2.5-VL-3B block-diffusion driving model, packaged as sharded Apache
Parquet for multi-host TPU training (grain / datasets / MaxText hf data path).
Each row is one Waymo Open Dataset End-to-End (WOD-E2E) front-camera frame, already run
through the Fast-dDrive prep (chat template + Qwen2.5-VL processor + deep-JSON scaffold +
3D… See the full description on the dataset page: https://huggingface.co/datasets/kaiwen2/wod-e2e-fast-ddrive-sasd.ddro-nq-dataset
DDRO — NQ320K Processed Dataset
This dataset contains the preprocessed Natural Questions (NQ320K) corpus used to train and evaluate the DDRO generative retrieval models from:
📄 Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval (SIGIR 2025)
The raw NQ data (from Google) is processed into a unified format with document text, queries, and relevance annotations, ready for use in generative retrieval training pipelines.
Files… See the full description on the dataset page: https://huggingface.co/datasets/kiyam/ddro-nq-dataset.DDRBench_10K
DDRBench: Deep Data Research Benchmark
📊 Leaderboard & Demo | 📄 Paper (Arxiv)
Overview
DDRBench (Deep Data Research Benchmark) is a comprehensive evaluation framework designed to assess the capabilities of Large Language Model (LLM) agents in performing complex, multi-turn data research and reasoning tasks. Unlike traditional Q&A benchmarks, DDRBench focuses on scenarios requiring deep interaction with structured databases, tool usage, and long-context reasoning.
This… See the full description on the dataset page: https://huggingface.co/datasets/thinkwee/DDRBench_10K.DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/YBoWen/DDR-dataset.ddr5-ram-price-observations-2026
DDR5 RAM Price Observations 2026
Dated $/GB observations for 32GB (2x16GB) DDR5 kits at named direct US retailers during the 2026 memory supply crisis.
Dated, verification-gated price observations for mainstream 32GB (2x16GB) DDR5 kits at named direct US retailers, captured during the 2026 memory supply crisis.
Each row carries capture date, kit, listing title, seller, USD price, $/GB and verification status.
Acceptance rules
The listing title must state the… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/ddr5-ram-price-observations-2026.ddro-docids
📄 ddro-docids
This repository provides the generated document IDs (DocIDs) used for training and evaluating the DDRO (Direct Document Relevance Optimization) models.
Two types of DocIDs are included:
PQ (Product Quantization) DocIDs: Compact semantic representations based on quantized document embeddings.
TU (Title + URL) DocIDs: Tokenized document identifiers constructed from document titles and/or URLs.
📚 Contents
pq_msmarco_docids.txt: PQ DocIDs for MS… See the full description on the dataset page: https://huggingface.co/datasets/kiyam/ddro-docids.ddro-msmarco-doc-dataset-300k
DDRO — MS MARCO Top-300K Processed Dataset
This dataset contains the preprocessed MS MARCO Top-300K document corpus used to train and evaluate the DDRO generative retrieval models from:
Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval (SIGIR 2025)
Files
File
Description
Size
msmarco-docs-sents.top.300k.json
Top-300K documents selected by click frequency, with sentence tokenization (JSONL format)
~2 GB… See the full description on the dataset page: https://huggingface.co/datasets/kiyam/ddro-msmarco-doc-dataset-300k.
