datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tests-raw-jsonlraw_jsonlknowref_60k_raw
The Knowref 60K Dataset
Project: https://github.com/aemami1/KnowRef60k
Data source: https://github.com/aemami1/KnowRef60k/tree/28e5385d17967744ccb3bdba45fdd89d9690307d
Fields
annotation_strength (str): annotator agreement from 1-5
candidate_0 (str): the first candidate name
candidate_1 (str): the second candidate name
original_sentence (str): sentence before swapping the names
swapped_sentence (str): sentence after swapping the names with square brackets marking the… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/knowref_60k_raw.robo_set_rawwarwick-second-life-dm-2025-raw
First-life and second-life battery degradation mode test data
BSEBench status: raw_mirror_pending_validation
This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.vima_rawraw_tts_esc_ESPnet_espnet_mls-audioset_soundstream_16kfmb_rawpreco_raw
The PreCo Dataset
Project: https://preschool-lab.github.io/PreCo/
Data source: https://drive.google.com/file/d/1q0oMt1Ynitsww9GkuhuwNZNq6SjByu-Y/view?usp=sharing
Details
The original PreCo .jsonl files from https://preschool-lab.github.io/PreCo/
What is PreCo?
PreCo is a large-scale English dataset for coreference resolution. The dataset is designed to embody the core challenges in coreference, such as entity representation, by alleviating the challenge of… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/preco_raw.franka_force_plugin_rawraw_friends_series_transcriptRaw transcript from friends tv series, chunk into ~1000 token lines. Text tagged by character.
language:
- en
size_categories:
- n<1K
io_ai_tech_rawraw_tts_esc_ESPnet_espnet_mls-multi_soundstream_16k2024_AMC12All problems copyrighted by the Mathematical Association of America's American Mathematics Competitions
Source:
https://artofproblemsolving.com/wiki/index.php/2024_AMC_12A_Problems
https://artofproblemsolving.com/wiki/index.php/2024_AMC_12B_Problems
Removed problems with figures:
12A: problem 14,18,22
12B: problem 7, 19
final-best-raw-episodes-2026-09-14-v2
Final and canonical-best evaluation episodes: frozen preparation
Full local packaging is now running. See materialization status and instructions. This preparation folder is not the full payload; the separate full export remains incomplete until its verified COMPLETED marker is written.
Prepared inventory: 72 evaluations / 119,608 expected episodes. There are
36 canonical-best and 44 final memberships, with 8 evaluations tagged both.
One incomplete Q38-teacher OfficeQA v2 final… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/final-best-raw-episodes-2026-09-14-v2.Drivegpt4_raw_datamars_hirise_dtm_raw-42e7c107a4f7c2d6
Mars HiRISE DTM — LitData Streaming Dataset
Pre-processed Mars HiRISE orthoimage + DTM patches in
LitData optimized streaming format. Associated with MarsFM: Shading-Regularized Flow Matching for Martian Relief Estimation.
Quick Start
from litdata import StreamingDataset
# Stream directly from HuggingFace — no full download needed
train_ds = StreamingDataset(input_dir="hf://datasets/SuperComputer/mars_hirise_dtm_raw-42e7c107a4f7c2d6/train")
val_ds =… See the full description on the dataset page: https://huggingface.co/datasets/SuperComputer/mars_hirise_dtm_raw-42e7c107a4f7c2d6.ox-alpha-glm-5.3-flash-distillation-coding-17k-raw
Ox Alpha GLM-5.3-Flash Distillation Coding 17K Raw
A raw collection of 17,138 synthetic coding samples generated with GLM-5.3-Flash, previously exposed through OpenCode under the stealth-model alias Ox Alpha.
The dataset is intended for experimentation with LLM distillation, code-generation models, instruction tuning, supervised fine-tuning, evaluation, and agentic coding systems.
hq-rawbnci-raw
EEG Dataset
This dataset was created using braindecode, a library for deep learning with EEG/MEG/ECoG signals.
Dataset Information
Number of recordings: 1
Number of channels: 26
Sampling frequency: 250.0 Hz
Data type: Continuous (Raw)
Number of windows: 96735
Total size: 19.23 MB
Storage format: zarr
Usage
To load this dataset:
from braindecode.datasets import BaseConcatDataset
# Load dataset from Hugging Face Hub
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Kkuntal990/bnci-raw.raw_arxivplex_robosuite_rawforecasting_rawRaw Dataset from "Approaching Human-Level Forecasting with Language Models"
This documentation provides an overview of the raw dataset utilized in our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt.
Data Source and Format
The dataset originates from forecasting platforms such as Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms engage users in predicting the… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting_raw.piperx-workpiece-storage-0909-62ep-raw
piperx-workpiece-storage-0909-62ep-raw
61 retained manually collected episodes recorded with EvoMind on 2026-09-09 (UTC+8).
Task: put the copper screw into the left box and the black sleeves into the right box.
LeRobot v3.0; 30 FPS; 53,347 frames; 1,778.2333 seconds. Robot: bi_piperx_follower.
Three original 640x480 RGB video views: left_wrist, right_wrist, right_environment_1.
Merged in chronological session order. Original episode 26 (19 frames) was removed on 2026-09-15.… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/piperx-workpiece-storage-0909-62ep-raw.glaucoma-expert-cot-raw-1077
Glaucoma Expert Chain-of-Thought
Ophthalmologist six-step reasoning reports for fundus photographs, each paired with
a binary glaucoma label. 1,074 cases from LAG and Papila.
Files
file
rows
split
expert_cot_trainval.jsonl
915
train (823) + val (92)
expert_cot_test.jsonl
159
test
images/
1,074
<source>_<id>.jpg
Record schema
{
"id": "1689",
"source": "LAG",
"image": "LAG_1689.jpg",
"split": "train"… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/glaucoma-expert-cot-raw-1077.raw_tts_esc_ESPnet_espnet_mls-english_soundstream_16kraw-mergedEOT-2004-Raw
End of Term 2004
Original dump: https://eotarchive.org/data/data-2004/
The End of Term Web Archive is a crawl of U.S. government websites conducted at the end of each presidential administration. This is a filtered version of the 2004 crawl.
Notice
This dataset is still a work in progress.
Data Curation
We download the 2004 EOT WARC files and parse the HTML using Trafilatura. We then filter the extracted text by length (minimum of 550 characters)… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/EOT-2004-Raw.ltaf-rawbinhvq_news21_raw
