datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ViLegalQA-Synthetic-Curation
ViLegalQA Synthetic Curation
Dataset summary
This repository releases the synthetic Vietnamese legal QA research artifacts produced in the accompanying study. The primary resource contains 10,095 synthetic QA items spanning true/false, multiple-choice, and open-ended tasks. It is accompanied by the final curation/quality annotations used in the study, plus aggregated labels for 600 items from the five-expert human calibration panel.
Manuscript: Human-Calibrated… See the full description on the dataset page: https://huggingface.co/datasets/nguyenkhanh87/ViLegalQA-Synthetic-Curation.qwen35-2b-tool-use-qwen36-27b-curation-candidates
Full candidate collections: 2B tool use + 27B data curation
This public Dataset contains two complete, unredacted, exact-40 candidate collections:
Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and
233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL,
Spider, and TravelPlanner.
Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021
targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.excavision-curation-assets
Excavision target-conditioned curation assets
Downloadable corpus artifacts for the Curate for my site workflow in Excavision Explorer.
Canonical files
excavision_full_vitb14.npy: the original 882,728 × 768 full-frame DINOv2 ViT-B/14 embeddings in float32, used directly without PCA or dimensionality reduction (SHA-256: 1b4b235fbffea44399e347de0a16cb4b60eba2ed5a30e8d344c908a9fa1d6bb2).
excavision_curation_pool.parquet: aligned sanitized filenames and seven… See the full description on the dataset page: https://huggingface.co/datasets/Sheida1/excavision-curation-assets.Data-Curation-for-Visual-AI-Module-5-VisDrone
Dataset Card for Voxel51/VisDrone2019-DET
This is a FiftyOne dataset with 8629 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("dgural/Data-Curation-for-Visual-AI-Module-5-VisDrone")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/dgural/Data-Curation-for-Visual-AI-Module-5-VisDrone.nemo-grpo-from083-full-edge-curation
Nemotron 0.83 Edge-Prompt Curation
This private dataset contains edge-prompt curation rollouts for the DGXChen/Tong CoT dataset.
Seed edge prompts: 134
New rollout rows after seed exclusion: 7668
New edge prompts: 1648
Full edge prompts, seed plus rollout: 1782
Full dataset rows: 7830
Edge rate over full dataset: 0.2276
The Hugging Face dataset viewer is configured to load only data/full_edge_prompts_seed_plus_rollout.jsonl.
The larger rollout and metadata files remain… See the full description on the dataset page: https://huggingface.co/datasets/dvyomkesh/nemo-grpo-from083-full-edge-curation.adaption-dataset-curation-prompts
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-dataset_curation_prompts
This dataset contains prompt-completion pairs focused on data curation tasks, including dataset documentation, schema generation, and sample data creation. The samples demonstrate instructions for generating dataset cards, categorizing information, and structuring training data for AI assistants. Entries vary from specific technical documentation like Yahoo… See the full description on the dataset page: https://huggingface.co/datasets/omerdemirtas/adaption-dataset-curation-prompts.Self-curation
