datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fractal20220817_data_train_22500_25000_augmented
fractal20220817_data_train_22500_25000_augmented
Overview
Codebase version: v3.0
Robots: images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e, xarm7
FPS: 3
Episodes: 2,500
Frames: 108,522
Splits:
train: 0:2500
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/fractal20220817_data_train_22500_25000_augmented.fractal20220817_data_train_0_2500_augmented
fractal20220817_data_train_0_2500_augmented
Overview
Codebase version: v3.0
Robots: images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e, xarm7
FPS: 3
Episodes: 2,500
Frames: 107,077
Splits:
train: 0:2500
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description
observation.images.image… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/fractal20220817_data_train_0_2500_augmented.fractal20220817_data_train_25000_27500_augmented
fractal20220817_data_train_25000_27500_augmented
Overview
Codebase version: v3.0
Robots: images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e, xarm7
FPS: 3
Episodes: 2,500
Frames: 107,878
Splits:
train: 0:2500
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/fractal20220817_data_train_25000_27500_augmented.fractal20220817_data_train_2500_5000_augmented
fractal20220817_data_train_2500_5000_augmented
Overview
Codebase version: v3.0
Robots: images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e, xarm7
FPS: 3
Episodes: 2,500
Frames: 106,995
Splits:
train: 0:2500
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description
observation.images.image… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/fractal20220817_data_train_2500_5000_augmented.ecg-grounding-250-2500
Dataset Details
This is a randomly split five fold dataset of the ECG-Grouding dataset stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-grounding-250-2500 where 250 refers to the sampling frequency and 2500 denotes 10 seconds.
The code to do the splitting is here.
Any questions or issues, please do not hesitate to reach out to the maintainer of ECG-Bench.
transformer-reasoning-bios-dataset-25000transformer-reasoning-bios-dataset-250000edition_2500_jxcai-scale-hle-public-questions-readymade
edition_2500_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2500_jxcai-scale-hle-public-questions-readymade.lichess-2500plus-gamesThis dataset contains all games that was played by two 2500+ elo players, and ended with a checkmate.
ecg-grounding-250-2500
Dataset Details
This is a randomly split five fold dataset of the ECG-Grouding dataset stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-grounding-250-2500 where 250 refers to the sampling frequency and 2500 denotes 10 seconds.
The code to do the splitting is here.
Any questions or issues, please do not hesitate to reach out to the maintainer of ECG-Bench.
uniprotkb_obsolete_entries_250000000-v1
uniprotkb_obsolete_entries_250000000
Dataset Description
Comprehensive protein knowledgebase with functional annotations
Original Source: ftp://ftp.uniprot.org/pub/databases/uniprot/current_release/rdf/uniprotkb_obsolete_entries_250000000.rdf.xz
Dataset Summary
This dataset contains RDF triples from uniprotkb_obsolete_entries_250000000 converted to HuggingFace
dataset format for easy use in machine learning pipelines.
Format: Originally rdf, converted to… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/uniprotkb_obsolete_entries_250000000-v1.transformer-reasoning-bios-dataset-25000_shuffleddoe-genesis-sealed-n2500
Demonstration receipts (n=2500)
These are demonstration envelopes from a filing. n=2500 is a demonstration number. Each row is one trajectory summary, not a 1 kHz pulse and not a timeseries. 15 banks (37,500 rows). Per-row proof_hash. File cryptographic_seal.
What a stranger sees if they cite this zip
They land on a grant-shaped shelf: 15 configs next to each other, mass/μ/booleans/proof_hash. No pulse. No 4×4 taxels. No sentence that this is the direction for… See the full description on the dataset page: https://huggingface.co/datasets/spiderpilot89/doe-genesis-sealed-n2500.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.ecg-qa-mimic-iv-ecg-250-2500
Dataset Details
This is a randomly split five fold dataset of the MIMIC-IV-ECG varaint of ECG-QA stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-qa-mimic-iv-ecg-250-2500 where 250 refers to the sampling frequency and 2500 denotes 10 seconds.
The code to do the splitting is here.
The dataset contents are under the same license as the original ECG-QA's license.
Any questions or issues, please do not… See the full description on the dataset page: https://huggingface.co/datasets/willxxy/ecg-qa-mimic-iv-ecg-250-2500.ecg-comp-ecg-flatline-30000-250-2500RedPajama-Data-V2-sample-100B-filtered-shuffled-tokenized-with-token-counts-2500000enall_annotated_sentences_25000
Dataset Summary
For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/all_annotated_sentences_25000
Additional Information
This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 25,000 sentences taken from the meeting minutes of the 25 central banks referenced in our paper.
Label… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/all_annotated_sentences_25000.GLM-5.0-25000x
PONYALPHA-step1.1 (Cleaned)
This dataset has been automatically cleaned to remove:
Empty or missing responses
Responses shorter than 10 characters
Refusal responses ("problem is incomplete", "cannot solve", etc.)
Responses with no substantive content
Responses that just echo the problem
Cleaning Report
Original rows: 39,324
Clean rows: 26,061
Removed: 13,263 (33.7%)
Columns: ['id', 'problem', 'thinking', 'solution', 'difficulty', 'category', 'timestamp'… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/GLM-5.0-25000x.ecg-qa-mimic-iv-ecg-250-2500
Dataset Details
This is a randomly split five fold dataset of the MIMIC-IV-ECG varaint of ECG-QA stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-qa-mimic-iv-ecg-250-2500 where 250 refers to the sampling frequency and 2500 denotes 10 seconds.
The code to do the splitting is here.
The dataset contents are under the same license as the original ECG-QA's license.
Any questions or issues, please do not… See the full description on the dataset page: https://huggingface.co/datasets/ELM-Research/ecg-qa-mimic-iv-ecg-250-2500.ecg-comp-ecg-flatline-80000-250-2500ecg-comp-noise-flatline-60000-250-2500ecg-comprehension-r-peak-count-789480-10000-250-2500ecg-comp-noise-flatline-40000-250-2500ecg-instruct-45k-250-2500
Dataset Details
This is a randomly split five fold dataset of the ECG-Instruct 45K dataset stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-instruct-45k-250-2500 where 250 refers to the sampling frequency and 2500 denotes 10 seconds.
The code to do the splitting is here.
Any questions or issues, please do not hesitate to reach out to the maintainer of ECG-Bench.
ecg-instruct-pulse-250-2500
Dataset Details
This is a randomly split five fold dataset of the ECG-Instruct Pulse dataset stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-instruct-pulse-250-2500 where 250 refers to the sampling frequency and 2500 denotes 10 seconds.
The code to do the splitting is here.
Any questions or issues, please do not hesitate to reach out to the maintainer of ECG-Bench.
CC-MAIN-2014-23_1_2500_row_wise_20240804_173419ecg-comprehension-bpm-float-200000-250-2500ecg-comprehension-r-peak-count-789480-5000-250-2500ecg-comp-ecg-noise-60000-250-2500
