datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
R1-Distill-NuminaMath-Curation
R1-Distill-NuminaMath-Curation Release
Associated blog post: to be added
Data Overview
This dataset consists of the NuminaMath samples from the ServiceNow-AI/R1-Distill-SFT dataset and has features problem, solution, source, R1-Distill-Qwen-32B, R1-Distill-Qwen-32B-messages, and correctness.
The problem, solution and source features are from the NuminaMath dataset and correspond to the problem statement, ground truth solution and problem source.
The R1-Distill-Qwen-32B… See the full description on the dataset page: https://huggingface.co/datasets/collinear-ai/R1-Distill-NuminaMath-Curation.MNIST-Curation
Curation of the famous MNIST Dataset
The curation was done using qualitative analysis of the dataset, following visualization techniques like PCA and UMAP and score-based categorization of the samples using metrics like hardness, mistakenness, or uniqueness.
The code of the curation can be found on GitHub:👉 https://github.com/Conscht/MNIST_Curation_Repo/tree/main
This curated version of MNIST introduces an additional IDK (“I Don’t Know”) label for digits that are ambiguous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/Consscht/MNIST-Curation.ViLegalQA-Synthetic-Curation
ViLegalQA Synthetic Curation
Dataset summary
This repository releases the synthetic Vietnamese legal QA research artifacts produced in the accompanying study. The primary resource contains 10,095 synthetic QA items spanning true/false, multiple-choice, and open-ended tasks. It is accompanied by the final curation/quality annotations used in the study, plus aggregated labels for 600 items from the five-expert human calibration panel.
Manuscript: Human-Calibrated… See the full description on the dataset page: https://huggingface.co/datasets/nguyenkhanh87/ViLegalQA-Synthetic-Curation.patho-ssl-data-curation
Revisiting Automatic Data Curation for Vision Foundation Models in Digital Pathology
Abstract Vision foundation models (FMs) are accelerating the devel- opment of digital pathology algorithms and transforming biomedical research. These models learn, in a self-supervised manner, to represent histological features in highly heterogeneous tiles extracted from whole-slide images (WSIs) of real-world patient samples. The performance of these FMs is significantly influenced by the size… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/patho-ssl-data-curation.oh-dcft-v1.3_no-curation_gpt-4o-mini_scale_4xqwen35-2b-tool-use-qwen36-27b-curation-candidates
Full candidate collections: 2B tool use + 27B data curation
This public Dataset contains two complete, unredacted, exact-40 candidate collections:
Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and
233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL,
Spider, and TravelPlanner.
Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021
targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.Data-Curation-for-Visual-AI-Module-4-VisDrone
Dataset Card for 2024.10.06.22.04.02
This is a FiftyOne dataset with 8629 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("dgural/Data-Curation-for-Visual-AI-Module-4-VisDrone")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/dgural/Data-Curation-for-Visual-AI-Module-4-VisDrone.excavision-curation-assets
Excavision target-conditioned curation assets
Downloadable corpus artifacts for the Curate for my site workflow in Excavision Explorer.
Canonical files
excavision_full_vitb14.npy: the original 882,728 × 768 full-frame DINOv2 ViT-B/14 embeddings in float32, used directly without PCA or dimensionality reduction (SHA-256: 1b4b235fbffea44399e347de0a16cb4b60eba2ed5a30e8d344c908a9fa1d6bb2).
excavision_curation_pool.parquet: aligned sanitized filenames and seven… See the full description on the dataset page: https://huggingface.co/datasets/Sheida1/excavision-curation-assets.Data-Curation-for-Visual-AI-Module-5-VisDrone
Dataset Card for Voxel51/VisDrone2019-DET
This is a FiftyOne dataset with 8629 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("dgural/Data-Curation-for-Visual-AI-Module-5-VisDrone")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/dgural/Data-Curation-for-Visual-AI-Module-5-VisDrone.curation-corpus
curation-corpus
Source
Data from this official repo with downloaded news articles content.
Citation
@misc{curationcorpusbase:2020,
title={Curation Corpus Base},
author={Curation},
year={2020}
}
WCP3475-Curation-3q0ycpoh-dcft-v1.2_no-curation_no_airoboros_gpt-4o-minicuration_tst_4oh-dcft-v1.2_no-curation_gpt-4o-mini_wo_opengptultrafeedback-binarized-curation
Ultrafeedback binarized dataset using the mean of preference ratings
Introduction
This dataset contains the result of curation work performed by Argilla (using Argilla 😃).
After visually browsing around 200 examples using the sort and filter feature of Argilla, we noticed a strong mismatch between the overall_score in the original UF dataset (and the Zephyr train_prefs dataset) and the quality of the chosen response.
By adding the critique rationale to our Argilla… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-curation.curationaleph-alpha-germanweb-curation-artifacts
Aleph Alpha GermanWeb — Curation artifacts
This gated repository preserves the curation inputs and provenance required to
understand and reproduce the final fair corpus.
Contents
filters/: downloaded filter and synthetic-data artifacts
cc-budgeted/: budgeted Common Crawl fragments used in the fair build
cc-extracted/: Common Crawl extraction manifests
Access and provenance
This repository is publicly visible but requires manual access approval.… See the full description on the dataset page: https://huggingface.co/datasets/RuHae/aleph-alpha-germanweb-curation-artifacts.oh-dcft-v1.2_no-curation_gpt-4o-mini_wo_camel_mathcuration_tst_3curation-ultrafeedback-scores-rawcuration_tst2pitvqa-curation-reviewcuration-backup-review_common_voice25_dev
Dataset Card for curation-backup-review_common_voice25_dev
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import argilla as rg
ds… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/curation-backup-review_common_voice25_dev.oh-dcft-v1.3_no-curation_gpt-4o-mini_scale_8xnemo-grpo-from083-full-edge-curation
Nemotron 0.83 Edge-Prompt Curation
This private dataset contains edge-prompt curation rollouts for the DGXChen/Tong CoT dataset.
Seed edge prompts: 134
New rollout rows after seed exclusion: 7668
New edge prompts: 1648
Full edge prompts, seed plus rollout: 1782
Full dataset rows: 7830
Edge rate over full dataset: 0.2276
The Hugging Face dataset viewer is configured to load only data/full_edge_prompts_seed_plus_rollout.jsonl.
The larger rollout and metadata files remain… See the full description on the dataset page: https://huggingface.co/datasets/dvyomkesh/nemo-grpo-from083-full-edge-curation.malay-curation-export-enccuration-corpus-ru
curation-corpus-ru
Translated version of d0rj/curation-corpus into Russian.
oh_curation_study_airoboroscuration-ultrafeedback-scoressharegpt-curation
Dataset Card for sharegpt-curation
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/sharegpt-curation.
