datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.FairVision
Dataset Card: Harvard-FairVision
Dataset Summary
Harvard-FairVision is the first large-scale medical fairness dataset with both 2D and 3D imaging data, covering three major eye diseases affecting approximately 380 million people worldwide. It contains 30,000 subjects (10,000 per disease) across Age-Related Macular Degeneration (AMD), Diabetic Retinopathy (DR), and glaucoma, each with paired SLO fundus photos and 3D OCT B-scans and six demographic identity attributes.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVision.maze-30x30-hard-1kCOPAL
About COPAL-ID
COPAL-ID is an Indonesian causal commonsense reasoning dataset that captures local nuances. It provides a more natural portrayal of day-to-day causal reasoning within the Indonesian (especially Jakartan) cultural sphere. Professionally written and validatid from scratch by natives, COPAL-ID is more fluent and free from awkward phrases, unlike the translated XCOPA-ID.
COPAL-ID is a test set only, intended to be used as a benchmark.
For more details, please see our… See the full description on the dataset page: https://huggingface.co/datasets/haryoaw/COPAL.harbor-benchcold-french-law
Collaborative Open Legal Data (COLD) - French Law
COLD French Law is a dataset containing over 800 000 french law articles, filtered and extracted from France's LEGI dataset and formatted as a single CSV file.
This dataset focuses on articles (codes, lois, décrets, arrêtés ...) identified as currently applicable french law.
A large portion of this dataset comes with machine-generated english translations, provided by Casetext, Part of Thomson Reuters using OpenAI's GPT-4.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-french-law.harbor-goose-openhands-benchmark
Same Model, Opposite Results: Goose vs OpenHands Turn Budget Study on Harbor Terminal-Bench-Pro
Trial-level results from a small controlled study comparing two agent harnesses —
Goose and OpenHands-SDK —
on a frozen 40-task Harbor Terminal-Bench-Pro slice.
All runs used minimax/minimax-m2.5 via OpenRouter with Daytona as the sandbox backend.
Key Findings
Reducing the turn budget from 100 to 60 pushed the two harnesses in opposite directions under the base setup:… See the full description on the dataset page: https://huggingface.co/datasets/namanvats/harbor-goose-openhands-benchmark.medsam-kidney-kpmpHarvard-GDP
Dataset Card: Harvard-GDP
Dataset Summary
Harvard-GDP (Harvard Glaucoma Detection and Progression) is a multimodal multitask ophthalmology dataset for glaucoma detection and progression forecasting. It is the largest publicly available glaucoma detection dataset with 3D OCT imaging data and the first publicly available glaucoma progression forecasting dataset. The dataset includes detailed demographic annotations (sex, race) to support fairness learning research.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/Harvard-GDP.korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.hubble-8b-unlearning-resultsCSI-BFI-HAR-Dataset
CSI-BFI-HAR Dataset
This repository contains the dataset, structure and usage of the CSI-BFI-HAR dataset of the corresponding dataset paper:
Please download the dataset either from huggingface or IEEE dataport:
https://huggingface.co/datasets/foysalhaque/CSI-BFI-HAR-Dataset
https://ieee-dataport.org/documents/csi-bfi-har-wi-fi-datasets-human-activity-recognition
Dataset Structure
The dataset is organized into two subsets:
Dataset-1: single-subject HAR (HAR-1 to… See the full description on the dataset page: https://huggingface.co/datasets/foysalhaque/CSI-BFI-HAR-Dataset.FairFedMed
Dataset Card: FairFedMed
Dataset Summary
FairFedMed is the first federated learning (FL) benchmark dataset for medical imaging with demographic annotations, designed to study group fairness across institutions in a federated setting. It comprises two subsets spanning ophthalmology and chest radiology, enabling research on fairness-aware federated learning under realistic cross-institutional data heterogeneity.
This dataset was introduced in the IEEE Transactions on… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairFedMed.FairDomain
Dataset Card: Harvard-FairDomain
Dataset Summary
Harvard-FairDomain is a large-scale ophthalmology dataset designed for studying fairness under domain shift in medical image analysis. It supports both image segmentation and classification tasks, with 10,000 samples per task drawn from 10,000 unique patients. The dataset introduces an additional imaging modality — en-face fundus images — alongside the original scanning laser ophthalmoscopy (SLO) fundus images, enabling… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairDomain.vn-provinces-timber-harvest
Vietnam provinces timber harvest
Concentrated timber harvested volume (thousand cubic metres). Coverage 1995-2024. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (1874 rows)
data/provinces.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-timber-harvest.vn-provinces-aquatic-harvest-area
Vietnam provinces aquatic harvest area
Aquatic product harvest area (thousand hectares). Coverage 1995-2024 (NSO header typo 1988 corrected to 1998). Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (1878… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-aquatic-harvest-area.diffusion-vs-ar-hard-sudoku
Diffusion vs AR Hard Sudoku
This repository packages 8,148,696 Sudoku examples in the CSV format expected
by HKUNLP/diffusion-vs-ar, plus its original 100k/1k easy baseline.
Every processed file has these columns:
column
meaning
quizzes
81 row-major digits; 0 is an empty cell
solutions
complete 81-digit solution
source
original collection
dataset
normalized dataset family
official_rating
rating supplied by the source
rating_type
semantics of that rating… See the full description on the dataset page: https://huggingface.co/datasets/fhyfhy/diffusion-vs-ar-hard-sudoku.multimodal-ICS-provenance
ProvICS: A Multimodal Provenance-Aware CPS Intrusion Detection Dataset
ProvICS is a multimodal, provenance-aware intrusion detection dataset for cyber-physical systems (CPS), collected from a hardware-in-the-loop (HIL) ICS testbed built on the Purdue reference model. It jointly provides four time-synchronized modalities — host kernel-level provenance, PLC-edge provenance, decoded Modbus/TCP protocol semantics, and physical-process state telemetry — all aligned on a common UTC… See the full description on the dataset page: https://huggingface.co/datasets/harryCJ/multimodal-ICS-provenance.FairGenMed
Dataset Card: FairGenMed
Dataset Summary
FairGenMed is the first dataset for studying fairness in medical generative models. It provides detailed quantitative clinical measurements alongside demographic annotations to investigate the semantic correlation between text prompts and anatomical regions across demographic subgroups. The dataset supports both generative model evaluation and downstream classification tasks for glaucoma detection.
This dataset accompanies the… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairGenMed.eval-harness-demo
OmicsBank Eval Harness (Demo)
⚠️ This uses entirely synthetic data. No row corresponds to a real patient. Built to
demonstrate the harness mechanics — not to represent real-world statistics or actual
clinical patterns.
Two tools, both runnable in a couple minutes with just pandas installed:
Auto-grader — score your model's predictions on a small held-out classification task
Leakage checker — check whether your eval/training data overlaps with a reference corpus… See the full description on the dataset page: https://huggingface.co/datasets/omicsbank/eval-harness-demo.esp32s3-csi-har-2025
ESP32-S3 WiFi CSI human activity dataset (July 2025)
Labelled WiFi Channel State Information captured from an ESP32-S3 on
29 July 2025. Four postures. 80 recordings. Collected and released by
Bl4ckd09.
The capture
Property
Value
Radio
ESP32-S3, single board, Espressif CSI parser
Date
29 July 2025, 18:23 to 19:55 local
Sample rate
19.823 Hz mean across all 80 files (range 19.12 to 20.00)
Subcarriers
192, stored as 384 interleaved I/Q values… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/esp32s3-csi-har-2025.MedConclusion-Compact
MedConclusion-Compact
MedConclusion is a large-scale dataset of 5.7M PubMed structured abstracts for biomedical conclusion generation. Each instance pairs the non-conclusion sections of an abstract with the original author-written conclusion, providing naturally occurring supervision for evidence-to-conclusion reasoning. MedConclusion also includes journal-level metadata such as biomedical category and SJR, enabling subgroup analysis across biomedical domains.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/MedConclusion-Compact.calm-month-6bb464
calm-month-6bb464
Synthetic products test data: 42 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Harbor-Xmiller/calm-month-6bb464.MicroG-HAR-train-ready
MicroG-HAR-train-ready dataset.
This dataset is converted from the MicroG-4M dataset.
For more information please:
Refer to our paper
Visit our GitHub
Datasets formatted in this way can be used directly as input data directories for the PySlowFast_for_HAR framework without additional preprocessing, enabling seamless model training and evaluation.
This dataset follows the organizational format of the AVA dataset. The only difference is that the original CSV header… See the full description on the dataset page: https://huggingface.co/datasets/lei-qi-233/MicroG-HAR-train-ready.thermowatch-osm-dataviolence-nonviolence-dataset
Violence vs Non-Violence Dataset
This dataset contains annotated interaction data for detecting violent vs non-violent human interactions.The data is extracted from video frames and includes bounding boxes, pose keypoints, motion features, and violence indicators for pairs of interacting persons.
Files
violence_data.csv → Frames labeled as violent interactions
non_violence_data.csv → Frames labeled as non-violent interactions
Each CSV contains structured features at… See the full description on the dataset page: https://huggingface.co/datasets/harmesh95/violence-nonviolence-dataset.tourism-wellness-datasetmodelfit-hardware-dataset
ModelFit: Local LLM Hardware Compatibility Dataset
An open dataset of which local AI models (Ollama) fit which hardware, by
parameter size, quantization, minimum RAM, and estimated memory load, across
Apple Silicon Macs, iPhones, and NVIDIA GPUs.
Maintained by ModelFit. Browse it as an interactive
table at modelfit.io/data; the canonical
machine-readable source is
modelfit.io/api/dataset.
141 models across 24 families (106 with a registry-verified local build, 35 cloud-only… See the full description on the dataset page: https://huggingface.co/datasets/modelfit/modelfit-hardware-dataset.hard-nerve-dcb861
hard-nerve-dcb861
Synthetic sensors test data: 53 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/PaulSanchez/hard-nerve-dcb861.africa-synth-agriculture-post-harvest-loss-data-africa-all
Africa Synth Agriculture Post Harvest Loss Data Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-agriculture-post-harvest-loss-data-africa-all.
