datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datause-displacement-reviewed
datause-displacement-reviewed
The Luna-reviewed subset of
rafmacalaba/datause-displacement:
only spans that received a v2.3 Luna verdict (band review + drop-side rescue,
source == luna_review). Every span carries the binary label plus
usage_type / drop_reason / specificity, and is traceable via key
(split:row:start:end) to the verdict records in
extraction_analysis/band_review/.
Configs
config
fields
gliner_reviewed
tokenized_text, corpus, origin… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-displacement-reviewed.SciCode-Runnable-Benchmark-Reviewedpad-auto-solver-reviewed
PAD Reviewed Dataset
Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository
contains immutable reviewed package revisions and does not contain raw captures,
training runs, checkpoints, or model binaries.
Packages exported: 28
Active catalog datasets: 14
Catalog schema: 3
Layout
packages/<dataset_id>.tar: deterministic self-contained reviewed package
catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.SR-ntsb-sft-preview-reviewed
Overview
This is a Supervised Fine-Tuning preview dataset consisting of 149 rows of structured aerospace safety and engineering reasoning. It is designed to teach language models to analyze complex physical and behavioral fact patterns using standard Root Cause Analysis within the specific context of aviation accidents investigated by the National Transportation Safety Board.
The dataset is derived from real United States NTSB Aviation Accident reports. Each row provides a… See the full description on the dataset page: https://huggingface.co/datasets/Sabr-Research/SR-ntsb-sft-preview-reviewed.hallmark-mlx-reviewed-policy-traces
hallmark-mlx-reviewed-policy-traces
Reviewed citation-verification training traces for hallmark-mlx.
Contents
train.jsonl: 75 supervised examples
valid.jsonl: 6 supervised examples
Source reviewed traces: reviewed_seed_traces_combined.jsonl with 45 full traces.
Format
Each row is a prepared supervised training example for MLX LoRA fine-tuning.
The format is the exact snapshot used by the kept Qwen 1.5B run.
Upload Note
Review the… See the full description on the dataset page: https://huggingface.co/datasets/sebastianboehler/hallmark-mlx-reviewed-policy-traces.lemonseed-qwen38-cogen-reviewed
lemonseed-qwen38-cogen-reviewed
LemonSeed — Qwen3.8-Max-teacher co-generated data, reviewed passes (v1).
Contents
intelligent_qwen38_cogen_1h_20260824_r1.reviewed_passes_v1.jsonl (71 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
SR-medical-sft-preview-reviewed
Overview
This is a Supervised Fine-Tuning preview dataset consisting of structured clinical and forensic medical reasoning. It is designed to teach language models to analyze complex patient histories, physical presentations, and diagnostic findings using systematic differential analysis and pathophysiological synthesis within the context of peer-reviewed medical literature and case reports.
The dataset is derived from real, open-access PubMed Central (PMC) medical case reports.… See the full description on the dataset page: https://huggingface.co/datasets/Sabr-Research/SR-medical-sft-preview-reviewed.reviewEdu-reviews-universitiesSR-gao-sft-preview-reviewed
Overview
This is a Supervised Fine-Tuning preview dataset consisting of 112 rows of structured legal reasoning. It is designed to teach language models to analyze legal fact patterns using a standard legal argumentation structure (Issue/Facts, Rule, Application/Analysis, Conclusion) within the specific context of Federal Governement Bid Protest decisions.
The dataset is derived from real United States Government Accountability Office (GAO) Bid Protest decisions. Each row… See the full description on the dataset page: https://huggingface.co/datasets/Sabr-Research/SR-gao-sft-preview-reviewed.mwo_hc_reviewedreviewed_IONOS-Help-Center
