datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datause-displacement-reviewed
datause-displacement-reviewed
The Luna-reviewed subset of
rafmacalaba/datause-displacement:
only spans that received a v2.3 Luna verdict (band review + drop-side rescue,
source == luna_review). Every span carries the binary label plus
usage_type / drop_reason / specificity, and is traceable via key
(split:row:start:end) to the verdict records in
extraction_analysis/band_review/.
Configs
config
fields
gliner_reviewed
tokenized_text, corpus, origin… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-displacement-reviewed.qiuli-collected-works-ocr-reviewed-20260906
裘李全集逐页审核输出
当前只有以下 2页 按百度首扫、千问疑难/全页盘点、GPT本地原图全文审核分工验收并冻结。千问盘点不等于千问全文校对。
LXQ-01 PDF第6页:revision3,原冻结包保留;PDF7仅跨页证据。
QXG-4 PDF第30页(书页25):revision3;PDF31仅跨页证据。
逐页canonical JSON是唯一当前真值,原始证据、模型版本/SHA与绝对路径映射均保留。源PDF仍在独立source-only仓。本仓不表示36册均审核完成,不额外授予原书版权许可。
SciCode-Runnable-Benchmark-Reviewed9jalingo-reviewed-hausa-batch-09jalingo-reviewed-yorubapad-auto-solver-reviewed
PAD Reviewed Dataset
Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository
contains immutable reviewed package revisions and does not contain raw captures,
training runs, checkpoints, or model binaries.
Packages exported: 28
Active catalog datasets: 14
Catalog schema: 3
Layout
packages/<dataset_id>.tar: deterministic self-contained reviewed package
catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.furry_dataset_e621_captions_claims_human_reviewedAbout 13,000 human reviewed captions. ~5000 of them were reviewed by me, the rest by independent workers.
I cannot confirm they will have caught all of the mistakes (there will still be some small mistakes). But the accuracy of these captions is higher than what any VLM can produce.
The "claims" are one-liner claims about an image, and given a truth value. Mostly machine-verified, but the ~10000 human-reviewed ones are a very valuable set of sex-related items or items which powerful VLMs had… See the full description on the dataset page: https://huggingface.co/datasets/furproxy/furry_dataset_e621_captions_claims_human_reviewed.9jalingo-reviewed-hausa9jalingo-reviewed-igbo9jalingo-reviewed-pidginKITAB_pdf_to_markdown_reviewed
KITAB_pdf_to_markdown_reviewed (Corrected KITAB-Bench PDF→Markdown)
Short description. A carefully reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic document OCR evaluation. We fixed ground-truth errors (hallucinated text, missing page numbers, omissions of small-font text) and standardized formatting to provide a reliable benchmark for model comparison.
TL;DR
✅ Human-verified ground truth for Arabic PDF→Markdown
✅ Removes hallucinations and fills… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/KITAB_pdf_to_markdown_reviewed.ebtopicality-llm-annotated-reviewedreviewed_sample
Silencio Network: Multilingual Accent Speech Dataset (Sample)
Overview
Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/reviewed_sample.reviewed-qa-keep-discard-tpriftis-pairskiji-inspector-reviewed-pairs
Kiji PII Detection Training Data
Synthetic multilingual dataset for training PII (Personally Identifiable Information) detection models with token-level entity annotations and coreference resolution.
Dataset Summary
Samples
99,990 (train: 89,991, test: 9,999)
Languages
6 (Dutch, Spanish, German, English, Danish, French)
Countries
20
PII entity types
26
Total entity annotations
814,306 (avg 8.1 per sample)
Coreference clusters
142,142 (99% of… See the full description on the dataset page: https://huggingface.co/datasets/575-lab/kiji-inspector-reviewed-pairs.reviewed-qa-keep-discard-pairsSR-ntsb-sft-preview-reviewed
Overview
This is a Supervised Fine-Tuning preview dataset consisting of 149 rows of structured aerospace safety and engineering reasoning. It is designed to teach language models to analyze complex physical and behavioral fact patterns using standard Root Cause Analysis within the specific context of aviation accidents investigated by the National Transportation Safety Board.
The dataset is derived from real United States NTSB Aviation Accident reports. Each row provides a… See the full description on the dataset page: https://huggingface.co/datasets/Sabr-Research/SR-ntsb-sft-preview-reviewed.uniprotkb_reviewed_archea_environmental_samples_48510_0-v1
uniprotkb_reviewed_archea_environmental_samples_48510_0
Dataset Description
Comprehensive protein knowledgebase with functional annotations
Original Source: ftp://ftp.uniprot.org/pub/databases/uniprot/current_release/rdf/uniprotkb_reviewed_archea_environmental_samples_48510_0.rdf.xz
Dataset Summary
This dataset contains RDF triples from uniprotkb_reviewed_archea_environmental_samples_48510_0 converted to HuggingFace
dataset format for easy use in machine… See the full description on the dataset page: https://huggingface.co/datasets/Dabbu19/uniprotkb_reviewed_archea_environmental_samples_48510_0-v1.hallmark-mlx-reviewed-policy-traces
hallmark-mlx-reviewed-policy-traces
Reviewed citation-verification training traces for hallmark-mlx.
Contents
train.jsonl: 75 supervised examples
valid.jsonl: 6 supervised examples
Source reviewed traces: reviewed_seed_traces_combined.jsonl with 45 full traces.
Format
Each row is a prepared supervised training example for MLX LoRA fine-tuning.
The format is the exact snapshot used by the kept Qwen 1.5B run.
Upload Note
Review the… See the full description on the dataset page: https://huggingface.co/datasets/sebastianboehler/hallmark-mlx-reviewed-policy-traces.pedagogy-luganda-reviewed
Description
299 Luganda pedagogical examples reviewed by native speakers for linguistic accuracy and cultural appropriateness. Includes original AI-generated Luganda alongside reviewer scores and comments. Part of the FabAI project for offline AI teacher companions in Ugandan primary schools (P1-P3).
Intended Use
Quality assessment of AI-generated Luganda educational content
Training reward models for Luganda translation quality
Limitations
Limited to P1-P3… See the full description on the dataset page: https://huggingface.co/datasets/CraneAILabs/pedagogy-luganda-reviewed.lemonseed-qwen38-cogen-reviewed
lemonseed-qwen38-cogen-reviewed
LemonSeed — Qwen3.8-Max-teacher co-generated data, reviewed passes (v1).
Contents
intelligent_qwen38_cogen_1h_20260824_r1.reviewed_passes_v1.jsonl (71 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
claude-reviewed-sport-sustainability-papers
Claude Reviewed Sport Sustainability Papers
Dataset description
This dataset encompasses 16 papers (some of them divided in several parts) related to the sustainability of sports products and brands.
This information lies at the core of our application and will come into more intesive use with the next release of GreenFit AI.
The papers were analysed with Claude 3.5 Sonnet (claude-3-5-sonnet-20241022) and they were translated into:
Their title (or a Claude-inferred… See the full description on the dataset page: https://huggingface.co/datasets/greenfit-ai/claude-reviewed-sport-sustainability-papers.ExtendedTACO-reviewedPEARL-LITE-reviewedSR-medical-sft-preview-reviewed
Overview
This is a Supervised Fine-Tuning preview dataset consisting of structured clinical and forensic medical reasoning. It is designed to teach language models to analyze complex patient histories, physical presentations, and diagnostic findings using systematic differential analysis and pathophysiological synthesis within the context of peer-reviewed medical literature and case reports.
The dataset is derived from real, open-access PubMed Central (PMC) medical case reports.… See the full description on the dataset page: https://huggingface.co/datasets/Sabr-Research/SR-medical-sft-preview-reviewed.Human-reviewed_automatic_English_translations_Europeana
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/21498
Description
The resource includes human-reviewed or post-edited translations of metadata sourced from the Europeana platform.
The human-inspected automatic translations are from 17 European languages to English.
The translations from Bulgarian, Croatian, Czech, Danish, German, Greek, Spanish, Finnish, Hungarian, Polish, Romanian, Slovak, Slovenian and Swedish have been reviewed by a group of… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Human-reviewed_automatic_English_translations_Europeana.uniprotkb_reviewed_archea_methanobacteriati_3366610_0-v1
uniprotkb_reviewed_archea_methanobacteriati_3366610_0
Dataset Description
Comprehensive protein knowledgebase with functional annotations
Original Source: ftp://ftp.uniprot.org/pub/databases/uniprot/current_release/rdf/uniprotkb_reviewed_archea_methanobacteriati_3366610_0.rdf.xz
Dataset Summary
This dataset contains RDF triples from uniprotkb_reviewed_archea_methanobacteriati_3366610_0 converted to HuggingFace
dataset format for easy use in machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/uniprotkb_reviewed_archea_methanobacteriati_3366610_0-v1.uniprotkb_reviewed_archea_promethearchaeati_1935183_0-v1
uniprotkb_reviewed_archea_promethearchaeati_1935183_0
Dataset Description
Comprehensive protein knowledgebase with functional annotations
Original Source: ftp://ftp.uniprot.org/pub/databases/uniprot/current_release/rdf/uniprotkb_reviewed_archea_promethearchaeati_1935183_0.rdf.xz
Dataset Summary
This dataset contains RDF triples from uniprotkb_reviewed_archea_promethearchaeati_1935183_0 converted to HuggingFace
dataset format for easy use in machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/uniprotkb_reviewed_archea_promethearchaeati_1935183_0-v1.CelebA-females-reviewed
CelebA Female Dataset
Dataset Description
This dataset is a filtered subset of the CelebA dataset (Celebrities Faces Attributes), containing only female faces. The original CelebA dataset is a large-scale face attributes dataset with more than 200,000 celebrity images, each with 40 attribute annotations.
Dataset Creation
This dataset was created by:
Loading the original CelebA dataset
Filtering to keep only images labeled as female (based on the "Male"… See the full description on the dataset page: https://huggingface.co/datasets/MnLgt/CelebA-females-reviewed.reviewEdu-reviews-universities
