datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KMMLU-HARD
KMMLU (Korean-MMLU)
We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM.
Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language.
We test 26 publically available and proprietary LLMs, identifying significant room for improvement.
The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.FairSeg
Dataset Card: FairSeg
Dataset Summary
FairSeg is a large-scale ophthalmology dataset for studying fairness in medical image segmentation. It contains 10,000 SLO fundus images with pixel-wise optic disc and cup segmentation masks, paired with comprehensive demographic annotations. The dataset is designed to benchmark and improve demographic equity in segmentation models, including foundation models such as SAM (Segment Anything Model).
This dataset was introduced at ICLR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairSeg.arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.FairVision
Dataset Card: Harvard-FairVision
Dataset Summary
Harvard-FairVision is the first large-scale medical fairness dataset with both 2D and 3D imaging data, covering three major eye diseases affecting approximately 380 million people worldwide. It contains 30,000 subjects (10,000 per disease) across Age-Related Macular Degeneration (AMD), Diabetic Retinopathy (DR), and glaucoma, each with paired SLO fundus photos and 3D OCT B-scans and six demographic identity attributes.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVision.maze-30x30-hard-1kMTS_Dialogue-Clinical_Note
MTS Dialogue (Clinical Note Summarisation)
Main Dataset
The MTS-Dialog dataset is a new collection of 1.7k short doctor-patient conversations and corresponding summaries (section headers and contents).
The training set consists of 1,201 pairs of conversations and associated summaries.
The validation set consists of 100 pairs of conversations and their summaries.
The "dialogue" column contain Doctor-Patient conversation. The "section_text" column contains the Clinical Note of the… See the full description on the dataset page: https://huggingface.co/datasets/har1/MTS_Dialogue-Clinical_Note.CADBench-Hard
CADBench Hard Tasks
43 out of the 105 tasks. Each folder contains the complete task prompt and its authoritative Fusion reference. For all of the tasks, verifiers and sandbox environment, please reach out
Dataset categories
Domains: computer-aided design, mechanical engineering, and robotics
Modalities: natural-language task instructions and native 3D CAD artifacts
Use cases: GUI-agent evaluation, computer-use evaluation, reinforcement learning, and deterministic… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/CADBench-Hard.COPAL
About COPAL-ID
COPAL-ID is an Indonesian causal commonsense reasoning dataset that captures local nuances. It provides a more natural portrayal of day-to-day causal reasoning within the Indonesian (especially Jakartan) cultural sphere. Professionally written and validatid from scratch by natives, COPAL-ID is more fluent and free from awkward phrases, unlike the translated XCOPA-ID.
COPAL-ID is a test set only, intended to be used as a benchmark.
For more details, please see our… See the full description on the dataset page: https://huggingface.co/datasets/haryoaw/COPAL.Harvard-GF
Dataset Card: Harvard-GF
Dataset Summary
Harvard-GF (Harvard Glaucoma Fairness) is a retinal nerve disease dataset for fairness learning in glaucoma detection, featuring both 2D and 3D OCT imaging data with balanced racial groups. It contains 3,300 samples from 3,300 patients with equal representation across Asian, Black, and White racial groups — a unique design addressing the doubled glaucoma prevalence observed in Black patients compared to other races.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/Harvard-GF.tulu-3-harmbench-evalThis data comes from the HarmBench benchmark.
This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Tülu 3 evaluation suite.
The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluation including this one.
harbor-benchcold-french-law
Collaborative Open Legal Data (COLD) - French Law
COLD French Law is a dataset containing over 800 000 french law articles, filtered and extracted from France's LEGI dataset and formatted as a single CSV file.
This dataset focuses on articles (codes, lois, décrets, arrêtés ...) identified as currently applicable french law.
A large portion of this dataset comes with machine-generated english translations, provided by Casetext, Part of Thomson Reuters using OpenAI's GPT-4.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-french-law.harbor-goose-openhands-benchmark
Same Model, Opposite Results: Goose vs OpenHands Turn Budget Study on Harbor Terminal-Bench-Pro
Trial-level results from a small controlled study comparing two agent harnesses —
Goose and OpenHands-SDK —
on a frozen 40-task Harbor Terminal-Bench-Pro slice.
All runs used minimax/minimax-m2.5 via OpenRouter with Daytona as the sandbox backend.
Key Findings
Reducing the turn budget from 100 to 60 pushed the two harnesses in opposite directions under the base setup:… See the full description on the dataset page: https://huggingface.co/datasets/namanvats/harbor-goose-openhands-benchmark.medsam-kidney-kpmpHarvard-GDP
Dataset Card: Harvard-GDP
Dataset Summary
Harvard-GDP (Harvard Glaucoma Detection and Progression) is a multimodal multitask ophthalmology dataset for glaucoma detection and progression forecasting. It is the largest publicly available glaucoma detection dataset with 3D OCT imaging data and the first publicly available glaucoma progression forecasting dataset. The dataset includes detailed demographic annotations (sex, race) to support fairness learning research.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/Harvard-GDP.korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.MedBrowseComp
MedBrowseComp Dataset
This repository contains datasets for medical information-seeking-oriented deep research and computer use tasks.
Datasets
The repository contains three harmonized datasets:
MedBrowseComp_50: A collection of 50 medical entries for browsing and comparison.
MedBrowseComp_605: A comprehensive collection of 605 medical entries.
MedBrowseComp_CUA: A curated collection of medical data for comparison and analysis.
Usage
These datasets can be… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Harvard/MedBrowseComp.FairVLMed
Dataset Card: Harvard-FairVLMed
Dataset Summary
Harvard-FairVLMed is the first fair vision-language medical dataset designed for studying fairness in medical vision-language (VL) foundation models. It contains 10,000 SLO fundus images paired with de-identified clinical notes and comprehensive demographic annotations, enabling in-depth fairness analysis across four protected attributes: race, gender, ethnicity, and preferred language.
This dataset was introduced at CVPR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVLMed.mtsampleshubble-8b-unlearning-resultsCSI-BFI-HAR-Dataset
CSI-BFI-HAR Dataset
This repository contains the dataset, structure and usage of the CSI-BFI-HAR dataset of the corresponding dataset paper:
Please download the dataset either from huggingface or IEEE dataport:
https://huggingface.co/datasets/foysalhaque/CSI-BFI-HAR-Dataset
https://ieee-dataport.org/documents/csi-bfi-har-wi-fi-datasets-human-activity-recognition
Dataset Structure
The dataset is organized into two subsets:
Dataset-1: single-subject HAR (HAR-1 to… See the full description on the dataset page: https://huggingface.co/datasets/foysalhaque/CSI-BFI-HAR-Dataset.FairFedMed
Dataset Card: FairFedMed
Dataset Summary
FairFedMed is the first federated learning (FL) benchmark dataset for medical imaging with demographic annotations, designed to study group fairness across institutions in a federated setting. It comprises two subsets spanning ophthalmology and chest radiology, enabling research on fairness-aware federated learning under realistic cross-institutional data heterogeneity.
This dataset was introduced in the IEEE Transactions on… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairFedMed.Harmful-Texts-On-Mastodon
🦣 Mastodon Wild Data for Harmful Content Detection
Overview
The Harmful Texts on Mastodon dataset is a human-annotated corpus of 3,000 English posts collected from the decentralized social media platform Mastodon between December 2024 and February 2025.It is designed to evaluate the robustness, generalization, and personalization capabilities of large language models (LLMs) and in-context learning (ICL) approaches for harmful content detection in real-world scenarios.… See the full description on the dataset page: https://huggingface.co/datasets/ChaseLabs/Harmful-Texts-On-Mastodon.harmbench_behaviorsFairDomain
Dataset Card: Harvard-FairDomain
Dataset Summary
Harvard-FairDomain is a large-scale ophthalmology dataset designed for studying fairness under domain shift in medical image analysis. It supports both image segmentation and classification tasks, with 10,000 samples per task drawn from 10,000 unique patients. The dataset introduces an additional imaging modality — en-face fundus images — alongside the original scanning laser ophthalmoscopy (SLO) fundus images, enabling… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairDomain.xnli2.0_train_arabicSupplychainstif-indonesia
Dataset Card for "stif-indonesia"
STIF-Indonesia
A dataset of "Semi-Supervised Low-Resource Style Transfer of Indonesian Informal to Formal Language with Iterative Forward-Translation".
You can also find Indonesian informal-formal parallel corpus in this repository.
Description
We were researching transforming a sentence from informal to its formal form. Our work addresses a style-transfer from informal to formal Indonesian as a low-resource machine… See the full description on the dataset page: https://huggingface.co/datasets/haryoaw/stif-indonesia.vn-provinces-timber-harvest
Vietnam provinces timber harvest
Concentrated timber harvested volume (thousand cubic metres). Coverage 1995-2024. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (1874 rows)
data/provinces.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-timber-harvest.spacenet-rio
SpaceNet (Rio de Janeiro) - Building Detection
This dataset contains high-resolution satellite imagery and corresponding building footprint annotations for (Rio de Janeiro) from the SpaceNet Building Detection Challenge. It is designed for training deep learning models for semantic segmentation and building footprint extraction.
Dataset Details
Total Tiles: 6,940 image tiles
Imagery Types:
3-band (RGB) Pan-sharpened GeoTIFFs (high spatial resolution)
8-band… See the full description on the dataset page: https://huggingface.co/datasets/harshinde/spacenet-rio.
