CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ctmedtech /DDR-dataset DDR - Diabetic Retinopathy Detection Dataset Image: Dataset Samples. The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/ctmedtech/DDR-dataset.imageimage-segmentation10K<n<100K1 likes3.3k downloads11mo agoHugging Face02ddrg /super_eurlexSuper-EURLEX dataset containing legal documents from multiple languages. The datasets are build/scrapped from the EURLEX Website [https://eur-lex.europa.eu/homepage.html] With one split per language and sector, because the available features (metadata) differs for each sector. Therefore, each sample contains the content of a full legal document in up to 3 different formats. Those are raw HTML and cleaned HTML (if the HTML format was available on the EURLEX website during the scrapping process) and cleaned text. The cleaned text should be available for each sample and was extracted from HTML or PDF. 'Cleaned' HTML stands here for minor cleaning that was done to preserve to a large extent the necessary HTML information like table structures while removing unnecessary complexity which was introduced to the original documents due to actions like writing each sentence into a new object. Additionally, each sample contains metadata which was scrapped on the fly, this implies the following 2 things. First, not every sector contains the same metadata. Second, most metadata might be irrelevant for most use cases. In our minds the most interesting metadata is the celex-id which is used to identify the legal document at hand, but also contains a lot of information about the document see [https://eur-lex.europa.eu/content/tools/eur-lex-celex-infographic-A3.pdf] as well as eurovoc- concepts, which are labels that define the content of the documents. Eurovoc-Concepts are, for example, only available for the sectors 1, 2, 3, 4, 5, 6, 9, C, and E. The Naming of most metadata is kept like it was on the eurlex website, except for converting it to lower case and replacing whitespaces with '_'.text-classification1M<n<10M3 likes2.8k downloads3y agoHugging Face03ddrg /math_text Mathematical Texts (MT) Mathematical dataset containing mathematical texts, i.e. texts containing LaTeX formulas, based on the AMPS Khan dataset and the ARQMath dataset V1.3. Based on the retrieved LaTeX texts, more mathematically equivalent versions have been generated by applying randomized LaTeX printing with this SymPy fork using Math Mutator (MAMUT). A positive id corresponds to the ARQMath post id of the generated text version, a negative id indicates an AMPS text. You can… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/math_text.text1M<n<10M2 likes1.3k downloads1y agoHugging Face04Aljo-na /DDR-dataset DDR - Diabetic Retinopathy Detection Dataset Image: Dataset Samples. The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/Aljo-na/DDR-dataset.imageimage-segmentation10K<n<100K0 likes655 downloads17d agoHugging Face05chanchahh /DDR-dataset DDR - Diabetic Retinopathy Detection Dataset Image: Dataset Samples. The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/chanchahh/DDR-dataset.imageimage-segmentation10K<n<100K1 likes395 downloads9mo agoHugging Face06ddrg /math_formula_retrieval Dataset Card for MFR (Mathematical Formula Retrieval) This dataset consists of formula pairs, classified as either mathematical equivalent or not. Dataset Details Dataset Description Mathematical dataset based on 71 famous mathematical identities. Each entry consists of two identities (in formula or textual form), together with a label, whether the two versions describe the same mathematical identity. The false pairs are not randomly chosen, but… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/math_formula_retrieval.texttext-classification10M<n<100M14 likes359 downloads1y agoHugging Face07ddrg /named_math_formulas Dataset Card for NMF (Named Mathematical Formulas) This dataset might be used to train a language model based math-retrieval system with a NSP-like task. See also see ddrg/math_formula_retrieval as a derived dataset which associates two formulas. You can find more information in MAMUT: A Novel Framework for Modifying Mathematical Formulas for the Generation of Specialized Datasets for Language Model Training. Named Math Formulas Mathematical dataset based on 71… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/named_math_formulas.texttext-classification10M<n<100M19 likes276 downloads8mo agoHugging Face08jerrodz /DDR-dataset DDR - Diabetic Retinopathy Detection Dataset Image: Dataset Samples. The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/jerrodz/DDR-dataset.imageimage-segmentation10K<n<100K0 likes274 downloads5mo agoHugging Face09RONIE2002 /DDR-dataset DDR - Diabetic Retinopathy Detection Dataset Image: Dataset Samples. The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/RONIE2002/DDR-dataset.imageimage-segmentation10K<n<100K0 likes220 downloads2mo agoHugging Face10ddrg /MUSES MUSES: a benchmark for Marked Unevenly Spaced Event Sequences MUSES, a benchmark for Marked Unevenly Spaced Event Sequences, is a collection of unevenly spaced time series datasets from various domains, containing marked events for training and evaluating prediction approaches. Languages All the columns and classes (when textual) in MUSES are in English (BCP-47 en) Dataset Structure Data Instances All datasets formatting follows… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/MUSES.documenttime-series-forecasting1M<n<10M4 likes206 downloads5mo agoHugging Face11ddrg /TOTRCD TOTRCD: Temporally-Ordered Tabular Regression Benchmark Suite with Concept Drift TOTRCD is a collection of tabular regression datasets that include a temporal ordering, represented by a time column, and exhibit some form of concept drift. Dataset Structure Data Fields The columns of the datasets follows all a similar formatting: Meta::: prefix marking identifiers, timestamps, or sort keys that are not used as model inputs. Example: Meta::DateTime.… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/TOTRCD.document100M<n<1B0 likes204 downloads2mo agoHugging Face12kaiwen2 /wod-e2e-fast-ddrive-sasd-50k WOD-E2E Fast-dDrive SASD 50k — TPU-ready training shards 50,331 frames (32 of the 263 WOD-E2E train shards) pre-tokenized into the Fast-dDrive Section-Aware Structured Diffusion (SASD) format, packaged as 787 Apache Parquet shards (~22 GB) for multi-host TPU training (grain / datasets / MaxText hf data path). The TPU side needs no tokenizer, processor, or torch — it loads arrays and trains. ⚠️ PRIVATE / license-restricted. Derived from the Waymo Open Dataset; shared privately… See the full description on the dataset page: https://huggingface.co/datasets/kaiwen2/wod-e2e-fast-ddrive-sasd-50k.tabular10K<n<100K0 likes191 downloads4mo agoHugging Face13MsNeeraj /DDR-dataset DDR - Diabetic Retinopathy Detection Dataset Image: Dataset Samples. The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/MsNeeraj/DDR-dataset.imageimage-segmentation10K<n<100K0 likes156 downloads7d agoHugging Face14ddrg /math_formulas Mathematical Formulas (MF) Mathematical dataset containing formulas based on the AMPS Khan dataset and the ARQMath dataset V1.3. Based on the retrieved LaTeX formulas, more equivalent versions have been generated by applying randomized LaTeX printing with this SymPy fork using Math Mutator (MAMUT). The formulas are intended to be well applicable for MLM. For instance, a masking for a formula like (a+b)^2 = a^2 + 2ab + b^2 makes sense (e.g., (a+[MASK])^2 = a^2 + [MASK]ab + b[MASK]2… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/math_formulas.text1M<n<10M11 likes146 downloads1y agoHugging Face15kasanii /DDR-dataset DDR - Diabetic Retinopathy Detection Dataset Image: Dataset Samples. The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for classification… See the full description on the dataset page: https://huggingface.co/datasets/kasanii/DDR-dataset.imageimage-segmentation10K<n<100K0 likes124 downloads5mo agoHugging Face16ddrg /named_math_formulas_ft Named Math Formulas - Fine-Tuning Dataset This dataset is a version of Named Math Formulas (NMF) dedicated to be used as fine-tuning dataset, e.g., by using the GitHub Repository aieng-lab/transformer-math-evaluation. In contrast to the full version of NMF, this dataset contains much more meta data that can be used for fine-grained evaluations. Dataset Details Dataset Description Mathematical dataset based on 71 famous mathematical identities. Each… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/named_math_formulas_ft.texttext-classification100K<n<1M2 likes122 downloads8mo agoHugging Face17Ci-Dave /DDR_dataset_train_test 🐨 DR Classification Fundus Dataset This dataset contains retinal fundus images labeled for Diabetic Retinopathy (DR) classification, intended for use in machine learning tasks such as image classification and medical diagnosis support. 📁 Dataset Structure The dataset is stored under the splits/ directory and is divided into: train and test folders with train and test.csv correspondingly image10K<n<100K0 likes105 downloads1y agoHugging Face18ZerXXX /DDR-Datasetimage10K<n<100K0 likes85 downloads1y agoHugging Face19thinkwee /DDRBench_10K_trajectory 10K Agent Trajectories Dataset Project Page | Paper | Code Overview This dataset contains agent trajectories from the Deep Data Research (DDR) project's 10-K financial analysis task, as presented in the paper "Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models". DDR-Bench is a large-scale benchmark designed to evaluate "investigatory intelligence" in LLM agents—the autonomy to set goals and explore raw data without explicit queries. This… See the full description on the dataset page: https://huggingface.co/datasets/thinkwee/DDRBench_10K_trajectory.texttable-question-answering10K<n<100K0 likes83 downloads8mo agoHugging Face20Achyuth12 /DDR-dataset DDR - Diabetic Retinopathy Detection Dataset Image: Dataset Samples. The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/Achyuth12/DDR-dataset.imageimage-segmentation10K<n<100K0 likes78 downloads8mo agoHugging Face21Ddream-ai /InsuranceCorpustext1K<n<10K10 likes67 downloads4y agoHugging Face22addone5 /DDR-dataset DDR - Diabetic Retinopathy Detection Dataset Image: Dataset Samples. The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/addone5/DDR-dataset.imageimage-segmentation10K<n<100K0 likes67 downloads8mo agoHugging Face23costinflation /ddr5-ram-prices-raw-dataset-2026 19,705 raw U.S. DDR5 RAM price observations across 12 ZIP markets and 29 days. DDR5 RAM Prices Raw Dataset (2026) Analyze 19,705 unaggregated product-level listed retail prices for standalone DDR5 memory modules and homogeneous kits across 12 U.S. ZIP markets from July 13 through August 10, 2026. The single analysis-ready CSV preserves titles, dates, geography, package quantities, listed prices, and a source-neutral comparable-price field. What “raw” means here: unaggregated… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/ddr5-ram-prices-raw-dataset-2026.tabulartabular-regression10K<n<100K2 likes67 downloads1mo agoHugging Face24kaiwen2 /wod-e2e-fast-ddrive-sasd WOD-E2E Fast-dDrive SASD — TPU-ready training shards Pre-tokenized Section-Aware Structured Diffusion (SASD) training samples for the Fast-dDrive Qwen2.5-VL-3B block-diffusion driving model, packaged as sharded Apache Parquet for multi-host TPU training (grain / datasets / MaxText hf data path). Each row is one Waymo Open Dataset End-to-End (WOD-E2E) front-camera frame, already run through the Fast-dDrive prep (chat template + Qwen2.5-VL processor + deep-JSON scaffold + 3D… See the full description on the dataset page: https://huggingface.co/datasets/kaiwen2/wod-e2e-fast-ddrive-sasd.tabularn<1K0 likes51 downloads4mo agoHugging Face25kiyam /ddro-nq-dataset DDRO — NQ320K Processed Dataset This dataset contains the preprocessed Natural Questions (NQ320K) corpus used to train and evaluate the DDRO generative retrieval models from: 📄 Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval (SIGIR 2025) The raw NQ data (from Google) is processed into a unified format with document text, queries, and relevance annotations, ready for use in generative retrieval training pipelines. Files… See the full description on the dataset page: https://huggingface.co/datasets/kiyam/ddro-nq-dataset.text-retrieval100K<n<1M0 likes47 downloads4mo agoHugging Face26thinkwee /DDRBench_10K DDRBench: Deep Data Research Benchmark 📊 Leaderboard & Demo | 📄 Paper (Arxiv) Overview DDRBench (Deep Data Research Benchmark) is a comprehensive evaluation framework designed to assess the capabilities of Large Language Model (LLM) agents in performing complex, multi-turn data research and reasoning tasks. Unlike traditional Q&A benchmarks, DDRBench focuses on scenarios requiring deep interaction with structured databases, tool usage, and long-context reasoning. This… See the full description on the dataset page: https://huggingface.co/datasets/thinkwee/DDRBench_10K.tabulartable-question-answering1M<n<10M0 likes39 downloads8mo agoHugging Face27YBoWen /DDR-dataset DDR - Diabetic Retinopathy Detection Dataset Image: Dataset Samples. The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/YBoWen/DDR-dataset.imageimage-segmentation10K<n<100K0 likes39 downloads6mo agoHugging Face28iBlessi /ddr5-ram-price-observations-2026 DDR5 RAM Price Observations 2026 Dated $/GB observations for 32GB (2x16GB) DDR5 kits at named direct US retailers during the 2026 memory supply crisis. Dated, verification-gated price observations for mainstream 32GB (2x16GB) DDR5 kits at named direct US retailers, captured during the 2026 memory supply crisis. Each row carries capture date, kit, listing title, seller, USD price, $/GB and verification status. Acceptance rules The listing title must state the… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/ddr5-ram-price-observations-2026.tabularn<1K0 likes38 downloads23d agoHugging Face29kiyam /ddro-docids 📄 ddro-docids This repository provides the generated document IDs (DocIDs) used for training and evaluating the DDRO (Direct Document Relevance Optimization) models. Two types of DocIDs are included: PQ (Product Quantization) DocIDs: Compact semantic representations based on quantized document embeddings. TU (Title + URL) DocIDs: Tokenized document identifiers constructed from document titles and/or URLs. 📚 Contents pq_msmarco_docids.txt: PQ DocIDs for MS… See the full description on the dataset page: https://huggingface.co/datasets/kiyam/ddro-docids.text100K<n<1M0 likes37 downloads4mo agoHugging Face30kiyam /ddro-msmarco-doc-dataset-300k DDRO — MS MARCO Top-300K Processed Dataset This dataset contains the preprocessed MS MARCO Top-300K document corpus used to train and evaluate the DDRO generative retrieval models from: Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval (SIGIR 2025) Files File Description Size msmarco-docs-sents.top.300k.json Top-300K documents selected by click frequency, with sentence tokenization (JSONL format) ~2 GB… See the full description on the dataset page: https://huggingface.co/datasets/kiyam/ddro-msmarco-doc-dataset-300k.texttext-retrieval100K<n<1M0 likes32 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.