datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
epstractor-raw
Epstractor: Epstein Archives Dataset
A comprehensive archive of documents, images, audio, and video files from multiple Epstein-related releases, including estate records and Department of Justice materials obtained through FOIA requests.
Dataset Description
This dataset contains 59,420 files totaling 115.23 GB from three major document releases, plus 2 large videos (40GB) available via a separate config:
Epstein Estate 2025-09: 5 files, 0.09 GB
Epstein Estate 2025-11:… See the full description on the dataset page: https://huggingface.co/datasets/public-records-research/epstractor-raw.game-recordings-v3
OriginLab Game Recordings v0.3.0
Human gameplay captured in-engine under per-title licenses at 1080p / 60 FPS CFR on one shared frame clock: every stream starts at frame 0 and frame k matches frame k across pre-HUD and post-HUD RGB, surface normals, metric depth, audio, camera telemetry, keyboard and mouse inputs, in-engine action events and game state, and world telemetry — plus per-frame training tables.
Watch full playable previews of every modality, side by side and in… See the full description on the dataset page: https://huggingface.co/datasets/originlab/game-recordings-v3.SuperGPQA-RecordsAll-CVE-Records-Training-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.DAI-CReTDHI-RecordGold-ATR
DAI-records-ATR - Record level
This dataset comprises parish and civil records from three main locations in France:
Ardennes
Tours
Ile de Ré
Dataset Summary
The DAI-records-ATR dataset includes 7,720 handwritten records from the XVI-XIXth centuries.
The records have been annotated by experts as part of the DAI-CReTDHI research project, using Teklia's open-source annotation interface Callico.
Split
set
images
train
6178
val
784
test
758… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/DAI-CReTDHI-RecordGold-ATR.dns-recordsrecorder-binance-futures-btc
chronos recorder archive — Binance USDT-M futures BTC
L2 order-book (depth snapshots + diffs) and trades streams for BTCUSDT and BTCUSDC perpetuals.
Recorded 24/7 by the chronos market recorder (GitHub:
BlackDigitalStudio/crypto-market-recorder) on VM scalper-recorder (Tokyo) up to
2026-08-05 ~18:xx UTC. Hourly snappy-parquet files, layout preserved 1:1 from
gs://recorder-data-asia-0998ac51/chronos/scalper-recorder/...
(GCP->HF migration 2026-08-22; ledger:… See the full description on the dataset page: https://huggingface.co/datasets/delmiron27/recorder-binance-futures-btc.chile-seismological-records
Chile's Seismological Records
assay-transfer-record-level-v27-bbb-martins-l1-intern
BBB Martins record-level V27 L1
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bbb-martins-l1-intern.assay-transfer-record-level-v27-bbb-martins-l3-intern
BBB Martins record-level V27 L3
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bbb-martins-l3-intern.assay-transfer-record-level-v27-bbb-martins-l2-intern
BBB Martins record-level V27 L2
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bbb-martins-l2-intern.assay-transfer-record-level-v27-bbb-martins-l4-intern
BBB Martins record-level V27 L4
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bbb-martins-l4-intern.assay-transfer-record-level-v27-bioavailability-ma-l2-intern
Oral bioavailability record-level V27 L2
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bioavailability-ma-l2-intern.assay-transfer-record-level-v27-bioavailability-ma-l4-intern
Oral bioavailability record-level V27 L4
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bioavailability-ma-l4-intern.assay-transfer-record-level-v27-bioavailability-ma-l6-intern
Oral bioavailability record-level V27 L6
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bioavailability-ma-l6-intern.assay-transfer-record-level-v27-bioavailability-ma-l3-intern
Oral bioavailability record-level V27 L3
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bioavailability-ma-l3-intern.assay-transfer-record-level-v27-bioavailability-ma-combined-intern
Oral bioavailability record-level V27 combined L2/L3/L4/L6
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bioavailability-ma-combined-intern.signed-measurement-records
Signed measurement records
Council of AI measurement record. Measurement, not certification.
Living board: 22 axis · 22 measured. Jail is a measured floor (TIE), not a 16th pane.
Hub cells: GET https://councilof.ai/api/hub-cards → re-GET counts.* (typed Hub triples SUPERSEDED) (third-party Hub — not the board). Verify free: https://councilof.ai/gspc-verify
Do not freeze a score table here. Older axis counts are superseded by the living GET.
Jail is a measured floor, not a 16th… See the full description on the dataset page: https://huggingface.co/datasets/csoai/signed-measurement-records.super_glue_record_promptsourceguided_marvels_spider_man_2_recordings_01
漫威蜘蛛侠2 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_ae2c5af176e4e2eab106954f144c7b7f
Collection: guided (精数据)
Recordings: 155
Layout: recordings/<recording_id>/<raw component>
Nemotron-Personas-USA-synthetic-records-10files-qa-vllm-qwen4b-instruct-2507-clarqatest_export_dataset_to_hub_with_records_True
Dataset Card for test_export_dataset_to_hub_with_records_True
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import argilla as rg
ds… See the full description on the dataset page: https://huggingface.co/datasets/argilla-internal-testing/test_export_dataset_to_hub_with_records_True.Phi4-ensemble-teacher-forcing-record-logits-datamedical_asr_recording_datasetData Source
Kaggle Medical Speech, Transcription, and Intent
Context
8.5 hours of audio utterances paired with text for common medical symptoms.
Content
This data contains thousands of audio utterances for common medical symptoms like “knee pain” or “headache,” totaling more than 8 hours in aggregate. Each utterance was created by individual human contributors based on a given symptom. These audio snippets can be used to train conversational agents in the medical field.
This Figure Eight… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/medical_asr_recording_dataset.congressional-record-parquet
Congressional Record 43rd–114th Congresses — Parquet Edition
This repository contains a Parquet-formatted derivative of the Stanford
Congressional Record for the 43rd–114th Congresses: Parsed Speeches and Phrase Counts
dataset.
The files were prepared as a compact, query-friendly research corpus for historical
text search with tools such as DuckDB and the Congressional Record Explorer.
Coverage
Congresses: 43rd–114th
One Parquet file per Congress
Files:… See the full description on the dataset page: https://huggingface.co/datasets/yeeder/congressional-record-parquet.assay-transfer-record-level-v26-bioavailability-ma-l5-intern
Oral bioavailability record-level V26 L5
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 13757, 'validation_ranking': 8297, 'test_ranking': 2948}
assay-transfer-record-level-v26-bioavailability-ma-combined-intern
Oral bioavailability record-level V26 combined L2-L6
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 910015, 'validation_ranking': 121352, 'test_ranking': 17440}
assay-transfer-record-level-v26-bioavailability-ma-l4-intern
Oral bioavailability record-level V26 L4
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 161282, 'validation_ranking': 4600, 'test_ranking': 4700}
assay-transfer-record-level-v26-bioavailability-ma-l6-intern
Oral bioavailability record-level V26 L6
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 81188, 'validation_ranking': 10869, 'test_ranking': 2750}
assay-transfer-record-level-v26-bbb-martins-l5-intern
BBB Martins record-level V26 L5
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 19482, 'validation_ranking': 1220, 'test_ranking': 404}
