datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HQ-OpenHumanVidopen-asr-leaderboard-resultscyp-challenge-train-test
CYP Challenge Train/Test Dataset
A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge.
Blog post: Announcing OpenADMET’s CYP inhibition blind challenge
Challenge Space: OpenADMET CYP Inhibition Blind Challenge
Challenge period: August 17, 2026 - November 3, 2026
Produced by: OpenADMET
CHANGELOG
Updated… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/cyp-challenge-train-test.Raon-OpenTTS-Eval
Raon-OpenTTS-Eval
Technical Report
A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs.
Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.openadmet-expansionrx-challenge-data
OpenADMET-ExpansionRx Challenge FULL dataset
This is the full dataset used in the OpenADMET-ExpansionRx blind challenge, which finalized in January 19th, 2026.
Originally split in a train and blinded test set, we now release the full dataset, which contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases.
While optimising candidate molecules for their preclinical programs Expansion collected… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-data.fsrs-datasetMed-HALT
Med-HALT: Medical Domain Hallucination Test for Large Language Models
This is a dataset used in the Med-HALT research paper. This research paper focuses on the challenges posed by hallucinations in large language models (LLMs), particularly in the context of the medical domain. We propose a new benchmark and dataset, Med-HALT (Medical Domain Hallucination Test), designed specifically to evaluate hallucinations.
Med-HALT provides a diverse multinational dataset derived from medical… See the full description on the dataset page: https://huggingface.co/datasets/openlifescienceai/Med-HALT.pxr-challenge-train-test
PXR Challenge Train/Test Dataset
A high-quality experimental dataset for predicting human Pregnane-X Receptor (PXR) induction, comprising over 11,000 compounds screened using a high-fidelity in-house assay. This is the largest publicly available PXR activity dataset, released as part of the OpenADMET PXR Induction Blind Challenge.
Blog post: Announcing the Next OpenADMET Blind Challenge: Predicting PXR Induction
Challenge Space: openadmet/pxr-challenge
Challenge period: April 1… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/pxr-challenge-train-test.open-hdri-1k
Open HDRI 1K
A consolidated, public-domain (CC0-1.0) collection of 3,491 equirectangular HDR environment maps at 1K resolution, gathered from five free HDRI libraries: Poly Haven, BlenderKit, ambientCG, CGEES and Open HDRI.
Every map is stored as a linear, high-dynamic-range .exr file alongside a tonemapped .jpg preview, with a per-asset metadata row (dimensions, source, author, license, SHA-256 checksum and tags).
Contents
Source
Assets
Author(s)
License… See the full description on the dataset page: https://huggingface.co/datasets/leodriesch/open-hdri-1k.genebench-pro-public-package
GeneBench-Pro Public Case Studies
This repository contains public GeneBench-Pro case studies. It is the
self-contained package intended for public distribution, including Hugging Face
publication.
Package Layout
<repo-root>/
├── .gitattributes
├── README.md
├── LICENSE
├── problems.csv
├── checksums.sha256
├── manifest.json
├── reference_definitions.md
├── reference_grader.py
└── problems/
└── <eval_id>/
├── eval_config.json
├── data_files/… See the full description on the dataset page: https://huggingface.co/datasets/openai/genebench-pro-public-package.openadmet-expansionrx-challenge-train-data
OpenADMET-ExpansionRx Challenge training dataset
This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision to… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-train-data.crunchbaseopen-command
OpenCommand
Catcher targets and pitcher command for MLB, inferred from broadcast video.
OpenCommand scores command using the pitch location's distance from target.
This dataset contains the 2024/2025/2026 computer vision object detections, every intermediate the pipeline writes, and the resulting command scores. The pipeline itself and the full method write-up live at github.com/tomdoyo/open-command.
Download
hf download tomdoyo/open-command --repo-type dataset… See the full description on the dataset page: https://huggingface.co/datasets/tomdoyo/open-command.openvino-arc140v-lunarlake
OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU
Reference performance data for running local models on a single Intel Core Ultra 7 258V
(Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All
inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here.
This is reference characterization shared by a non-expert contributor — careful measurements
on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.Octant_CYP_inhibition_reactivity_blog_release
OpenADMET Octant CYP Inhibition & Reactivity
Data release from the OpenADMET consortium, generated by Octant Bio.
This dataset accompanies the blog post Building the OpenADMET Data Engine.
Source code, assay protocols, and raw TSV files are on GitHub.
Overview
Cytochrome P450 (CYP) enzymes drive the oxidative metabolism of most drugs and are a primary cause of drug-drug interactions (DDIs).
Despite their importance, public CYP datasets are sparse, noisy, and… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/Octant_CYP_inhibition_reactivity_blog_release.openbrain_v1_0
OpenBrain v1.0
OpenBrain v1.0 is a public-ready release of brain-extracted T1-weighted MRI
images, SynthStrip-derived brain masks, and automated whole-brain segmentation
labels.
Release Contents
Cases: 35,838
Source OpenNeuro datasets: 607
License: CC0
Artifacts per case:
image.nii.gz: revised brain-extracted T1w image
brain_mask.nii.gz: SynthStrip-derived brain mask
whole_brain_segmentation.nii.gz: automated whole-brain segmentation label
All released cases are… See the full description on the dataset page: https://huggingface.co/datasets/openbrain-anon/openbrain_v1_0.OpenASL_3D
This is a Large-Scale 3D Datset for Continuous American Sign Language
FSRS-Anki-20k
Update
We have released a new dataset: anki-revlogs-10k.
Introduction
FSRS-Anki-20k is a dataset of 20k collections from Anki for FSRS project. It is a random sample of collections with 5000+ revlog entries, so it should contain a mix of older (still active) users, and newer users. Entries are pre-sorted in (cid, id) order.
There are two versions of the dataset: ./revlogs and ./dataset. The ./revlogs version contains the raw revlog entries, while the ./dataset version… See the full description on the dataset page: https://huggingface.co/datasets/open-spaced-repetition/FSRS-Anki-20k.harbor-goose-openhands-benchmark
Same Model, Opposite Results: Goose vs OpenHands Turn Budget Study on Harbor Terminal-Bench-Pro
Trial-level results from a small controlled study comparing two agent harnesses —
Goose and OpenHands-SDK —
on a frozen 40-task Harbor Terminal-Bench-Pro slice.
All runs used minimax/minimax-m2.5 via OpenRouter with Daytona as the sandbox backend.
Key Findings
Reducing the turn budget from 100 to 60 pushed the two harnesses in opposite directions under the base setup:… See the full description on the dataset page: https://huggingface.co/datasets/namanvats/harbor-goose-openhands-benchmark.au-racing-open-data
Australian Racing Open Data
Open, machine-readable datasets for Australian racing — the kind of data that
normally sits behind a login, a paywall, or nowhere at all.
Two of these datasets, as far as we can tell, have never existed publicly
before: track geometry (turn radii, cambers, straight lengths, first-split
distances — gathered by writing to 111 racing clubs and state bodies) and
greyhound GPS sectionals at 50-metre resolution.
Everything here is rebuilt and pushed every… See the full description on the dataset page: https://huggingface.co/datasets/brucem1967/au-racing-open-data.NYC-Airbnb-Open-Dataopen-ko-s2s-eval-artifacts
Open Ko-S2S 평가 산출물 (감사용)
⚠️ KsponSpeech 참조 전사는 해시로 대체돼 있습니다
KsponSpeech 는 AI Hub 배포 데이터로 재배포 제한이 있을 수 있어, kspon 런의
ref 컬럼을 ref_sha256 으로 대체했습니다(전사 원문 미포함). 모델 출력(hyp)과
채점 결과(cer_err/cer_len/cer)는 우리 산출물이라 그대로 공개합니다.
Zeroth 런은 원본이 CC BY 4.0(OpenSLR #40)이라
ref 원문을 그대로 담고 있습니다.
라이선스 보유자의 검증 절차
AI Hub 에서 KsponSpeech 를 정당하게 받은 분은 다음으로 우리 수치를 검증할 수 있습니다.
리더보드 저장소의 eval/datasets_ko.py 에서 clean_kspon() 을 가져옵니다.
자기 사본의 원 전사에 clean_kspon() 을 적용합니다. 결과가 목록이면… See the full description on the dataset page: https://huggingface.co/datasets/baryonlabs/open-ko-s2s-eval-artifacts.japan-procurement-open-data
Japan Public Procurement & Company Open Data
Machine-readable extracts of Japanese public-procurement and company open data, compiled and normalised by LoreaTec for japan-tenders.loreatec.jp and bizsearch.loreatec.jp. Everything here comes from official Japanese government sources; the value added is the cleaning, joining and the derived analysis (contract series and re-tender predictions).
Updated monthly. The authoritative, always-current copy is… See the full description on the dataset page: https://huggingface.co/datasets/loreatec/japan-procurement-open-data.open-bandit
This is the full size version of Open Bandit Dataset that can be used for research on bandit algorithms and off-policy evaluation.
The small size example version of our data is available at https://github.com/st-tech/zr-obp/tree/master/obd
Dataset Description
Open Bandit Dataset is constructed in an A/B test of two multi-armed bandit policies in a large-scale fashion e-commerce platform, ZOZOTOWN (https://zozo.jp/).
It currently consists of a total of 26M rows, each one… See the full description on the dataset page: https://huggingface.co/datasets/zozonext/open-bandit.tibetan-voice-benchmarkBenchmark of Tibetan Speech-To-Text dataset created by Monlam AI x Openpecha in 2024
All the transcripts have been reviewed by at least one person in addition to the original transcriber.
Data was taken on 15 July 2024 02∶47∶06 PM.
dept
desc
Count
STT_AB
Audio book
1000
STT_CS
Children Speech
1367
STT_HS
History
1000
STT_MV
Tibetan Movies
1000
STT_NS
Natural Speech
1000
STT_NW
News
1000
STT_PC
Podcast
1000
STT_TT
Tibetan Teachings
1000
grade column is used to… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/tibetan-voice-benchmark.korean_law_open_data_precedents
Dataset Card for Dataset Name
공지사항
인공지능 기술로 여러가지 법률 서비스를 만들어 보고 있는데, 현재는 일반인들이 쉽고 정확한 법률 정보를 찾을 수 있는 법률 정보 플랫폼을 만들고 있습니다.
사용상 주의사항
사건번호가 동일한 중복 데이터가 약 200여건 포함돼있습니다.
그 이유는 법제처 국가법령 공동활용 센터 판례 목록 조회 API가 판례정보일련번호는 다르지만 사건번호 및 그 밖에 다른 필드 값들은 완전히 동일한 데이터들을 리턴하기 때문입니다.
사용에 참고하시기 바랍니다.
Dataset Summary
2023년 6월 기준으로 법제처 국가법령 공동활용 센터에서 제공된 전체 판례 데이터셋입니다.
그 이후로 제공되는 판례가 더 늘어났을 수 있습니다. 추가되는 판례들은 이 데이터셋에도 정기적으로 추가할 예정입니다.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/joonhok-exo-ai/korean_law_open_data_precedents.open-insect
Description
Open-Insect is curated to benchmark open-set recognition of novel species in biodiversity monitoring, with a focus on insects. It is consists of URLs pointing to image files compiled from the Global Biodiversity Information Facility (GBIF) as well as other metadata such as latitude, longitude, GBIF speciesKey, and label used to train the classifier. Open-Insect partially builds on a subset of the AMI dataset.
Run this download script to download this huggingface… See the full description on the dataset page: https://huggingface.co/datasets/yuyan-chen/open-insect.kor-rag-opentestopen-models-benchmark-results
⚡ Local LLM Evaluation Leaderboard
Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs.
💻 Hardware & System Specifications
All evaluations are executed under standardized local cluster environments:
Specification
Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.im3_open_source_data_center_atlas_v2026.02.09
IM3 Open Source Data Center Atlas v2026.02.09 — refined database
This repository preserves the IM3 Open Source Data Center Atlas v2026.02.09 and
adds a source-enriched, audited 43-column power-source table for all 1,479
source geometry records (1,474 unique IM3 IDs). The publication retains the exact
13 upstream columns plus 30 stable label, interpretation, and evidence fields.
Duplicate geometry records are intentionally retained.
Files… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/im3_open_source_data_center_atlas_v2026.02.09.
