datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xyzBENDER_SAMPLE
BENDER Sample — Biological ENsembles of Disordered proteins across taxa
A representative sample of CALVADOS coarse-grained molecular dynamics
trajectories from the full BENDER dataset, comprising sequences across
8 taxonomic groups.
Each protein folder contains:
<uniprot_id>.dcd — CALVADOS Ca trajectory (200ns+)
top.pdb — topology file
Folder structure
TaxonomicGroup/
└── <uniprot_id>/
├── <uniprot_id>.dcd
└── top.pdb
Sample composition… See the full description on the dataset page: https://huggingface.co/datasets/taseef/BENDER_SAMPLE.douyin
Douyin Posts Dataset
A dataset of 1,133,545 unique posts from Douyin (Chinese TikTok), collected via snowball sampling of related videos.
Dataset Description
This dataset contains metadata from Douyin short videos, collected by crawling related video recommendations. Starting from seed videos matching specific keywords, the crawler iteratively fetched related videos to build a diverse sample of the platform's content.
Collection Method
Snowball sampling:… See the full description on the dataset page: https://huggingface.co/datasets/bendavidsteel/douyin.UK-Road-Bend-Classificationso101-test-leader-v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_physical_wrapped",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bendca61/so101-test-leader-v1.BENDER
BENDER — Biological ENsembles of Disordered proteins across kingdoms
Raw CALVADOS coarse-grained molecular dynamics trajectories for
11,533 intrinsically disordered proteins spanning 13 kingdoms of life.
Each protein folder contains:
- <uniprot_id>.dcd — CALVADOS Cα trajectory (200 ns+)
- top.pdb — topology file
Folder structure
Kingdom.zip/
└── <uniprot_id>/
├── <uniprot_id>.dcd
└── top.pdb
Available zip files
File… See the full description on the dataset page: https://huggingface.co/datasets/taseef/BENDER.wikipedia-knowledge
Wikipedia Knowledge Catalogs
10 gezk knowledge catalogs: portable, read-only reference corpora with full-text and vector indexes, each packaged as one .gezk file that Gezel and any gezk reader can search and cite offline. Every catalog is here twice — the archive, and the same content as Parquet tables for data tools.
Catalogs
Catalog
Id
Version
Documents
Chunks
Archive
Wikipedia: Arts & Architecture
wikipedia-arts
2026.4.3
11,768
56,179
138.6 MB… See the full description on the dataset page: https://huggingface.co/datasets/Bendyline/wikipedia-knowledge.nor-casehold
Nor-CaseHOLD: A Retrieval Benchmark for Norwegian Legal AI
Nor-CaseHOLD is an extractive legal retrieval benchmark built from 1,244 Norwegian legal documents — 627 Supreme Court (Høyesterett) decisions and 617 Skatteetaten bindende forhåndsuttalelser (BFU) — each paired with its official summary.
To the author's knowledge, this is the first open-source retrieval benchmark for Norwegian legal text.
Task
Given a full legal document, select the 5 sentences that best… See the full description on the dataset page: https://huggingface.co/datasets/bendik-eeg-henriksen/nor-casehold.svla_so101_mujoco_pickplaceThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_mujoco",
"total_episodes": 61,
"total_frames": 45345,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:61"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bendca61/svla_so101_mujoco_pickplace.hebrew-asr-vn
Hebrew ASR three-source training dataset
Derived from ivrit.ai crowd-transcribe-v5, crowd-recital and knesset-committees.
Original data and transcripts are credited to ivrit.ai and its contributors.
Pinned revisions and preparation rules are in metadata/sources.json and
metadata/preparation-config.json. VoxKnesset is excluded by user decision.
Source/split
Clips
Hours
crowd-recital/test
1,557
1.071
crowd-recital/train
45,372
33.258
crowd-recital/validation
1,033… See the full description on the dataset page: https://huggingface.co/datasets/benderrodriguez/hebrew-asr-vn.act-so101-test-leader-v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_physical_wrapped",
"total_episodes": 1,
"total_frames": 1086,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bendca61/act-so101-test-leader-v1.tx_evaluation
tx_evaluation
Evaluation datasets for the
tx_evaluation benchmark: how well does a
transcriptomics embedding recover known biology?
Three genome-scale CRISPRi Perturb-seq screens over the same 2,393 target genes, in three
different human cell lines.
adata.obsm is empty by design. These files carry raw counts and metadata only. You run
your model over the cells, inject the resulting vectors, and the benchmark scores them.
Files
File
Cell line
Cell type… See the full description on the dataset page: https://huggingface.co/datasets/bendidiihab/tx_evaluation.BenderGestalt
Bender-Gestalt Test Dataset
Executive Summary
The Bender-Gestalt Test (BGT), or Bender Visual-Motor Gestalt Test, is commonly used in pre-employment psychological assessments to evaluate candidates' cognitive and motor skills, particularly visual-motor integration. By reproducing geometric figures, candidates showcase their ability to process and organize visual information. The test can also identify potential neurological or developmental disorders, providing insights… See the full description on the dataset page: https://huggingface.co/datasets/iaravagni/BenderGestalt.Informal-Standard-English-Corpus
Dataset Description
This dataset is a parallel corpus of approximately 11,000 pairs of informal conversational English text and their normalized equivalents. The informal text mimics real-world digital communication, featuring slang, phonetic spellings, missing punctuation, and abbreviations. The normalized text provides a grammatically correct and semantically equivalent version.
The dataset was created to support machine translation tasks for low-resource languages. It… See the full description on the dataset page: https://huggingface.co/datasets/Bendang/Informal-Standard-English-Corpus.BENDER-MPIPI
BENDER-MPIPI
Companion dataset to BENDER. Contains coarse-grained molecular dynamics trajectory data from MPIPI (Multi-scale Platform for Intrinsically disordered Protein Interactions) simulations of intrinsically disordered proteins (IDPs).
Contents
Trajectories are organized by taxonomic kingdom. Each zip file contains subdirectories named by UniProt ID, with the following files per protein:
traj.xtc — trajectory file
log.lammps — LAMMPS simulation log
Rg.out —… See the full description on the dataset page: https://huggingface.co/datasets/taseef/BENDER-MPIPI.flipped_test_frames_cam_bend_10This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 10,
"total_frames": 5831,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/flipped_test_frames_cam_bend_10.moonboard_instructvla-mujoco-so101-cube_on_tray-leader-v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_mujoco",
"total_episodes": 54,
"total_frames": 40334,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:54"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bendca61/vla-mujoco-so101-cube_on_tray-leader-v1.immediate-bend-54d6cd
immediate-bend-54d6cd
Synthetic products test data: 36 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/cedarBridge/immediate-bend-54d6cd.so101-test-leader-v1-subtask-go_to_resting_positionBen_dialectsso101-test-leader-v1-subtask-move_the_green_tapemetal_no_bendknesset-committees-preppedmmi-bendr-preprocessedThe EEG Motor Movement/Imagery (MMI) Dataset preprocessed with DN3 to be used for downstream fine-tuning with BENDR.
The labels correspond to Task 4 (imagine opening and closing both fists or both feet) from experimental runs 4, 10 and 14.
Creating dataloaders
from datasets import load_dataset
from torch.utils.data import DataLoader
dataset = load_dataset("rasgaard/mmi-bendr-preprocessed")
dataset.set_format("torch")
train_loader = DataLoader(dataset["train"], batch_size=8)… See the full description on the dataset page: https://huggingface.co/datasets/rasgaard/mmi-bendr-preprocessed.RGB_Optic_Flow_Bend_ClassificationData Viewer - View Samples
Data Viewer - View Sample Count
In our project on classifying the sharpness of bends using time-sequence data, we generated two distinct datasets:
RGB Images: Capturing conventional color information.
Wide View Dense Optic Flow: Providing detailed motion dynamics.
Notably, our latest dataset required approximately 16 hours to generate the samples using the code available in this GitHub notebook.
To enhance organization and accessibility, we have migrated the… See the full description on the dataset page: https://huggingface.co/datasets/aap9002/RGB_Optic_Flow_Bend_Classification.so101_mujoco_demoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_mujoco",
"total_episodes": 8,
"total_frames": 4760,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bendca61/so101_mujoco_demo.mujoco-so101-cube_on_tray-leader-v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_mujoco",
"total_episodes": 54,
"total_frames": 40334,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:54"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bendca61/mujoco-so101-cube_on_tray-leader-v1.so101-test-leader-v1-subtask-going_toward_the_green_tapecascades_exp
