datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
binary-30k
Binary-30K: Cross-Platform Binary Dataset with Stratified Splits
Paper | Code
🔗 Original Dataset (no splits): mjbommar/binary-30k-tokenized
This is the stratified train/validation/test split version of the Binary-30K dataset, containing 29,793 unique cross-platform binaries with pre-computed tokenization. This version provides standardized splits for reproducible machine learning research.
🎯 Key Features
✅ Stratified 70/15/15 splits maintaining class balance across… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k.binary-30k-tokenized
Dataset Card for Binary-30K
Dataset Summary
Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection.
Note… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k-tokenized.ukr-emotions-binary
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.ko-voicephishing-binary-classificationcambench_binary_eval
CameraBench Binary Evaluation Dataset
A balanced VQA dataset for evaluating camera motion understanding in videos.
📊 Dataset Statistics
Total Questions: 384
Unique Videos: 119
Unique Questions: 31
Yes Answers: 192 (50.0%)
No Answers: 192 (50.0%)
Balance Ratio: 1.00
Total Size: 126.16 MB (0.12 GB)
Average Video Size: 1.06 MB
🎯 Task Categories
This dataset covers various camera motion tasks including:
Static: 42 questions
Move In: 29 questions
Pan Left: 24… See the full description on the dataset page: https://huggingface.co/datasets/tuhink/cambench_binary_eval.so100_cube_4_binaryThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 30,
"total_frames": 8059,
"total_tasks": 1,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/so100_cube_4_binary.eval_so101_binary_cube_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 20,
"total_frames": 8906,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_so101_binary_cube_2.binary-30k
Binary-30K: Cross-Platform Binary Dataset with Stratified Splits
Paper | Code
🔗 Original Dataset (no splits): mjbommar/binary-30k-tokenized
This is the stratified train/validation/test split version of the Binary-30K dataset, containing 29,793 unique cross-platform binaries with pre-computed tokenization. This version provides standardized splits for reproducible machine learning research.
🎯 Key Features
✅ Stratified 70/15/15 splits maintaining class balance across… See the full description on the dataset page: https://huggingface.co/datasets/HandsomeWin/binary-30k.binary-30k-tokenized
Dataset Card for Binary-30K
Dataset Summary
Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection.
Note… See the full description on the dataset page: https://huggingface.co/datasets/fwufbhiwuhf/binary-30k-tokenized.binary-10IQR-secu
Dataset Card for "binary-10IQR-secu"
More Information needed
steam-reviews-constructiveness-binary-label-annotations-1.5k
1.5K Steam Reviews Binary Labeled for Constructiveness
Dataset Summary
This dataset contains 1,461 Steam reviews from 10 of the most reviewed games. Each game has about the same amount of reviews. Each review is annotated with a binary label indicating whether the review is constructive or not. The dataset is designed to support tasks related to text classification, particularly constructiveness detection tasks in the gaming domain.
Also available as… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/steam-reviews-constructiveness-binary-label-annotations-1.5k.VoxBox-EN-Annotated-Binarymetaculus-binaryso101_binary_cube_7This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 3999,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/so101_binary_cube_7.eval_binary_cube_homelab_so101_act_12This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 6,
"total_frames": 1865,
"total_tasks": 1,
"total_videos": 12,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:6"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_binary_cube_homelab_so101_act_12.binary-30k
Binary-30K: Cross-Platform Binary Dataset with Stratified Splits
Paper | Code
🔗 Original Dataset (no splits): mjbommar/binary-30k-tokenized
This is the stratified train/validation/test split version of the Binary-30K dataset, containing 29,793 unique cross-platform binaries with pre-computed tokenization. This version provides standardized splits for reproducible machine learning research.
🎯 Key Features
✅ Stratified 70/15/15 splits maintaining class balance across… See the full description on the dataset page: https://huggingface.co/datasets/fwufbhiwuhf/binary-30k.eval_binary_cube_homelab_so101_smolvla_cut_dataset_40kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 2048,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_binary_cube_homelab_so101_smolvla_cut_dataset_40k.eval_binary_cube_homelab_so101_act_8This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 13,
"total_frames": 5880,
"total_tasks": 1,
"total_videos": 26,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:13"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_binary_cube_homelab_so101_act_8.jigsaw-toxic-comment-multi-binarybinary_cude_100_101_homelabThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 10,
"total_frames": 3358,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/binary_cude_100_101_homelab.markers-tip-binary
Markers Grip Dataset — Binary Tip Size
Purpose
Educational dataset for binary tabular classification predicting marker tip style (fine vs bold) from physical/container features.
Dataset composition
Original split: 30 unique, real‑world measurements (no duplicates).
Augmented split: 300 rows via label‑preserving Gaussian jitter of numeric features.
Features (5): container_length_mm (int), grip_diameter_mm (float), length_to_diameter (float), ink_family (str)… See the full description on the dataset page: https://huggingface.co/datasets/0408happyfeet/markers-tip-binary.tetlock-binary-data-20260625_reward_brier_step200_analyzedeval_binary_cube_homelab_so101_act_11This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 6,
"total_frames": 1822,
"total_tasks": 1,
"total_videos": 12,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:6"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_binary_cube_homelab_so101_act_11.eval_binary_cube_homelab_so101_act_retestThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 4,
"total_frames": 1420,
"total_tasks": 1,
"total_videos": 8,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:4"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_binary_cube_homelab_so101_act_retest.binary-10IQR-code
Dataset Card for "binary-10IQR-code"
More Information needed
ProtST-BinaryLocalizationeval_binary_cube_homelab_so101_smolvla_20kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 2223,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_binary_cube_homelab_so101_smolvla_20k.gaia-dr3-binary-masses
Gaia DR3 Binary Masses
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
The Gaia DR3 binary masses catalog provides physical masses for binary star systems derived by the ESA Gaia mission. Masses are determined by combining astrometric non-single-star (NSS) solutions with spectroscopic radial velocities and photometric constraints, applying Kepler's third law to the orbital solutions.
These are among the most… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/gaia-dr3-binary-masses.eval_binary_cube_homelab_so101_smolvla_60k_with_teleop_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 1929,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_binary_cube_homelab_so101_smolvla_60k_with_teleop_3.eval_binary_cube_homelab_so101_smolvla_cut_dataset_unfrozen_50kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 800,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_binary_cube_homelab_so101_smolvla_cut_dataset_unfrozen_50k.
