datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenThoughts-1k-sample
[!NOTE]
We have released a paper for OpenThoughts! See our paper here.
Open-Thoughts-1k-sample
This is a 1k sample of the OpenThoughts-114k dataset.
Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles!
Inspect the content with rich formatting with Curator Viewer.
Available Subsets
default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models:
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.many-peptides-md
[!IMPORTANT]
Critical Update
The original 8AA TICA models within subsampled_trajectories/*/8AA/*.npz employed a CA-only atom selection. These models are not valid for comparison to results in our paper.
Updated files (uploaded 15/12/2025) now contain corrected models. If you previously downloaded this dataset, please re-download to ensure accurate results.
Note: Codebase references to tica_features_ca must now be replaced with tica_features. This was resolved in our codebase by PR #26.
Note:… See the full description on the dataset page: https://huggingface.co/datasets/transferable-samplers/many-peptides-md.samplescosmos_qavazFineFineWeb-sample
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.nmnjevolkit-logprobs-pipeline-75k-v2-sampleVBench_sampled_video
VBench Sampled Video
SpreadsheetBench-2-sample
SpreadsheetBench 2 Dataset
SpreadsheetBench 2 evaluates spreadsheet agents on end-to-end business and financial workflows. The release contains 321 tasks in four categories: Debugging, Financial Model, Template, and Visualization.
Overview
Directory
Task type
Tasks
Input workbooks
Gold workbooks
Debugging
Spreadsheet debugging and error correction
100
100
10
Financial_Model
Completion of multi-sheet financial models
100
100
20
Template
Completion of… See the full description on the dataset page: https://huggingface.co/datasets/icyCreater/SpreadsheetBench-2-sample.VBench-2.0_sampled_videos
Sample Videos of VBench-2.0
This dataset is used in the paper:👉 arXiv:2503.21755
camus-sample
CAMUS Sample - 2-D Echocardiographic Ultrasound Dataset
This is a sample subset of the full CAMUS dataset, provided for demonstration and testing purposes. It contains 6 files (1 patient per split). For the full dataset (500 patients), see: zeahub/camus.
This dataset is a zea-format (HDF5) conversion of the
CAMUS
dataset for multi-structure segmentation in 2-D echocardiography.
Property
Value
Modality
2-D transthoracic echocardiography
Patients
500
Views… See the full description on the dataset page: https://huggingface.co/datasets/zeahub/camus-sample.dummy-audio-samplesSA-Med3D-140K
SA-Med3D-140K [github]
Dataset Summary
SA-Med3D-140K is a large-scale, multi-modal, multi-anatomical volumetric medical image segmentation dataset. It was created to facilitate the development of general-purpose foundation models for 3D medical image segmentation.
The dataset comprises 21,729 3D medical images and 143,518 corresponding masks.
It was gathered from a combination of 70 public datasets and 8,128 privately licensed annotated cases from 24 hospitals.… See the full description on the dataset page: https://huggingface.co/datasets/blueyo0/SA-Med3D-140K.sample-datasetskernelbench-samples
KernelBench Samples
Samples from experiments for KernelBench, described in our arxiv
Learn more about KernelBench from our
Paper
Github Repo
The samples are organized as such
baseline_eval (Section 4 Baseline)
repeated_sampling (Section 5.1.1 Repeated Sampling)
iterative_refinement (Section 5.1.2 Iterative Refinement of Generations)
Within each folder, we organize the results by /level/model/problem_{id}/sample_{id}.
The inner most .json file contains the generated kernel and… See the full description on the dataset page: https://huggingface.co/datasets/ScalingIntelligence/kernelbench-samples.TIAToolBox_Remote_Samples
LICENSE
No re-distribution allowed.
Purpose
This repository contains publicly available samples used by the TIAToolBox for testing purposes.
Some of these images have been downloaded from [OpenSlide] for code verification purposes.
GitHub Repository: [TIAToolBox]
essential-web-1t-sample-fdc-partitioned
🌐 Essential-Web: FDC Level-2 Partitioned Dataset
📋 Dataset Description
This dataset contains a 1 trillion token sample from Essential-Web, partitioned by Free Decimal Correspondence (FDC) level-2 categories. Essential-Web is a 24-trillion-token web dataset with extensive document-level metadata designed to enable rapid dataset curation through SQL-like filtering.
🔍 Free Decimal Correspondence (FDC)
The FDC taxonomy is an open classification system… See the full description on the dataset page: https://huggingface.co/datasets/Research-EAI/essential-web-1t-sample-fdc-partitioned.audio_samples_1kcsfpsg-audio-v3-unofficial-mirror
PSG-Audio v3 — Unofficial Complete Mirror
Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset.
This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community.
Overview
PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.babilong-1k-samples
BABILong (1000 samples) : a long-context needle-in-a-haystack benchmark for LLMs
Preprint is on arXiv and code for LLM evaluation is available on GitHub.
BABILong Leaderboard with top-performing long-context models.
bAbI + Books = BABILong
BABILong is a novel generative benchmark for evaluating the performance of NLP models in
processing arbitrarily long documents with distributed facts.
It contains 9 configs, corresponding to different sequence lengths in tokens: 0k… See the full description on the dataset page: https://huggingface.co/datasets/RMT-team/babilong-1k-samples.MSRBackupsReasoningCheckpointsOmnisharing_DB_SampleData
Overview
The embodied intelligence industry is currently facing significant development challenges. The most critical issue is the lack of high-quality data, particularly omnimodal data that integrates force and tactile sensing. The PaXini introduces the PX OmniSharing Dataset, built on the PaXini Super EID Factory, enabling large-scale, high-fidelity human data collection across diverse tasks and scenarios.
The dataset includes multi-dimensional tactile data, multi-view visual… See the full description on the dataset page: https://huggingface.co/datasets/paxini/Omnisharing_DB_SampleData.lsdeval-resultsau30_tra
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Sam04/au30_tra.samsum
Dataset Card for SAMSum Corpus
Dataset Description
Links
Homepage: hhttps://arxiv.org/abs/1911.12237v2
Repository: https://arxiv.org/abs/1911.12237v2
Paper: https://arxiv.org/abs/1911.12237v2
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
The SAMSum dataset contains about 16k messenger-like conversations with summaries. Conversations were created and written down by linguists fluent in English. Linguists were asked to… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/samsum.common-crawl-sample
Common Crawl sample
A small unofficial random subset of the famous Common Crawl dataset.
60 random segment WET files were downloaded from Common Crawl on 2024-05-12.
Lines between 500 and 5000 characters long (inclusive) were kept.
Only unique texts were kept.
No other filtering.
Languages
Each text was assigned to one of the language codes using the GCLD3 Python package.
The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.p2-etf-samba-models
