curation
mlfoundations-dev_-_oh-dcft-v1.1-no-curation-ggufmlfoundations-dev_-_oh-dcft-v1.2_no-curation_gpt-4o-mini_wo_metamath-ggufmlfoundations-dev_-_oh-dcft-v1.2_no-curation_gpt-4o-mini_wo_unnatural_instructions-ggufmlfoundations-dev_-_oh-dcft-v1.2_no-curation_gpt-4o-mini_wo_slim_orca-ggufoh-dcft-v1.3_no-curation_gpt-4o-mini_scale_2x-GGUFmlfoundations-dev_-_oh-dcft-v1.2_no-curation_gpt-4o-mini_wo_airoboros-ggufmlfoundations-dev_-_oh-dcft-v1.2_no-curation_gpt-4o-mini_wo_evol_instruct-ggufoh-dcft-v1.3_no-curation_gpt-4o-mini_scale_0.25x-GGUF
Datasets
All datasets matching “curation”R1-Distill-NuminaMath-Curation
R1-Distill-NuminaMath-Curation Release
Associated blog post: to be added
Data Overview
This dataset consists of the NuminaMath samples from the ServiceNow-AI/R1-Distill-SFT dataset and has features problem, solution, source, R1-Distill-Qwen-32B, R1-Distill-Qwen-32B-messages, and correctness.
The problem, solution and source features are from the NuminaMath dataset and correspond to the problem statement, ground truth solution and problem source.
The R1-Distill-Qwen-32B… See the full description on the dataset page: https://huggingface.co/datasets/collinear-ai/R1-Distill-NuminaMath-Curation.MNIST-Curation
Curation of the famous MNIST Dataset
The curation was done using qualitative analysis of the dataset, following visualization techniques like PCA and UMAP and score-based categorization of the samples using metrics like hardness, mistakenness, or uniqueness.
The code of the curation can be found on GitHub:👉 https://github.com/Conscht/MNIST_Curation_Repo/tree/main
This curated version of MNIST introduces an additional IDK (“I Don’t Know”) label for digits that are ambiguous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/Consscht/MNIST-Curation.ViLegalQA-Synthetic-Curation
ViLegalQA Synthetic Curation
Dataset summary
This repository releases the synthetic Vietnamese legal QA research artifacts produced in the accompanying study. The primary resource contains 10,095 synthetic QA items spanning true/false, multiple-choice, and open-ended tasks. It is accompanied by the final curation/quality annotations used in the study, plus aggregated labels for 600 items from the five-expert human calibration panel.
Manuscript: Human-Calibrated… See the full description on the dataset page: https://huggingface.co/datasets/nguyenkhanh87/ViLegalQA-Synthetic-Curation.patho-ssl-data-curation
Revisiting Automatic Data Curation for Vision Foundation Models in Digital Pathology
Abstract Vision foundation models (FMs) are accelerating the devel- opment of digital pathology algorithms and transforming biomedical research. These models learn, in a self-supervised manner, to represent histological features in highly heterogeneous tiles extracted from whole-slide images (WSIs) of real-world patient samples. The performance of these FMs is significantly influenced by the size… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/patho-ssl-data-curation.oh-dcft-v1.3_no-curation_gpt-4o-mini_scale_4xqwen35-2b-tool-use-qwen36-27b-curation-candidates
Full candidate collections: 2B tool use + 27B data curation
This public Dataset contains two complete, unredacted, exact-40 candidate collections:
Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and
233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL,
Spider, and TravelPlanner.
Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021
targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.
