datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
various
malcolmrey's Various AI Model, Architecture & Research Repository
Welcome to the central research and asset repository of malcolmrey. This repository hosts cutting-edge tools, custom architectures, RefMod latent adapter systems, video synthesis engines, training configurations, benchmark suites, cinematic scripts, and comprehensive educational guides spanning MiniMax-H3, FLUX.2 / Klein 9B, WAN 2.1, LTX-Video, Z-Image, SDXL, and Stable Diffusion.
🧭 Repository Map &… See the full description on the dataset page: https://huggingface.co/datasets/malcolmrey/various.Real-IAD_VarietyVAREX
VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents
VAREX (VARied-schema EXtraction) is a benchmark for evaluating multimodal foundation models on structured data extraction from government forms. It comprises 1,777 documents with 1,771 unique schemas across three structural categories, each provided in four input modalities. Ground truth is deterministic — generated via a Reverse Annotation pipeline that programmatically fills PDF templates with synthetic values… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/VAREX.lunara-aesthetic-image-variations
Dataset Card for Moonworks Lunara Aesthetic II
This dataset introduces the second open-source release by Moonworks. This dataset contains original image and art created by Moonworks and their contextual variations generated by Moonworks Lunara, a sub-10B parameter model with a novel diffusion mixture architecture.
Paper: https://arxiv.org/pdf/2602.01666
While part 1 is intended for learning and evaluating regional as well as region-agnostic art styles, part 2 is intended for… See the full description on the dataset page: https://huggingface.co/datasets/moonworks/lunara-aesthetic-image-variations.varroa-yolo-under-2m-wiouyt_full_image_dataset
Dataset Card for "yt_full_image_dataset"
More Information needed
general_light_curve_benchmark_dataset_collection_roman_simulated_variable_star_datasetrlbench_peract_variation0libero40_libero_spatial_v2_playThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 420,
"total_frames": 117600,
"total_tasks": 10,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:420"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/VarunGiridhar3/libero40_libero_spatial_v2_play.imagenet-variations-synth-pilot
ImageNet Variations Synth Pilot (1K)
Synthetic vision-language instruction pilot from laion/imagenet_variations prompts.
Pipeline
Sample 1000 prompts from imagenet_variations
Add one style framing per prompt (website / textbook / diagram / …)
Generate image with FLUX.1-schnell
Encode with SEED-2 (<seed2_N>, 32 tokens)
Qwen2.5-VL-7B-Instruct looks at the actual image and writes instruction + response(modes: describe / QA / howto / reverse)
Format:… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/imagenet-variations-synth-pilot.view_variation_finalScreenSpot-v2-variants
ScreenSpot-v2-variaints
The ScreenSpot dataset with 4 types of instructions:
instruction: original instruction from ScreenSpot,action: clarifies the action to take,
description: describes the target UI element.
negative: an operation that can not be done in the screenshot.
Compatible code can be found in the GitHub repo above.
vargov-design-catalog
Vargov® Design Catalog — 605 lighting and decorative compositions in 8 languages
A machine-readable catalog of the full body of work of Vargov® Design, an author-driven
studio of lighting and decorative compositions founded by designer Anton Vargov (Moscow).
Every record is one composition: its identifier, category, canonical URLs, image links,
awards, links to its 3D model, and editorial copy written by the studio in eight
languages — Russian, English, German, Italian, French… See the full description on the dataset page: https://huggingface.co/datasets/vargov-design/vargov-design-catalog.IIT-Banaras-Hindu-University-Varanasi-Drone-PhotosAerial Image Capture from Drone of Banaras Hindu University as in 2017/2019
Libra_bimanual_fairino_plug_in_container_r_a_angle_variationsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "FAIRINO_BIMANUAL_SIM",
"total_episodes": 95,
"total_frames": 20345,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:95"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Libra-TeV/Libra_bimanual_fairino_plug_in_container_r_a_angle_variations.libero40_libero_goal_v2_playThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 420,
"total_frames": 125260,
"total_tasks": 10,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:420"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/VarunGiridhar3/libero40_libero_goal_v2_play.NCT-CRC-HE
100,000 histological images of human colorectal cancer and healthy tissue
Data Description "NCT-CRC-HE-100K"
This is a set of 100,000 non-overlapping image patches from hematoxylin & eosin (H&E) stained histological images of human colorectal cancer (CRC) and normal tissue.
All images are 224x224 pixels (px) at 0.5 microns per pixel (MPP). All images are color-normalized using Macenko's method (http://ieeexplore.ieee.org/abstract/document/5193250/, DOI… See the full description on the dataset page: https://huggingface.co/datasets/Varsha-Y12/NCT-CRC-HE.H2HMEM
📊 H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions
📄 Paper
This dataset is introduced in the following research work:
H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions
-📄 arXiv: https://arxiv.org/abs/2606.09461v1-💻 Code: https://github.com/varib1/H2HMEM-📊 Dataset: https://huggingface.co/datasets/varib/H2HMEM-🌐 Project Page: https://h2hmemprojectpage.vercel.app/-🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/varib/H2HMEM.yt_main_image_dataset
Dataset Card for "yt_main_image_dataset"
More Information needed
varroa-yolo-baselines-part2-fullCivilEng11k
Dataset Card for "CivilEng11k"
More Information needed
SapBark_64_variety_classification
Sapbark 64 Variety Classification
A dataset for variety classification of sapling bark of fruit trees. The dataset contains 5,815 images across 64 classes.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{alizadeh2025sapbark,
title={SapBark-64: A dataset of bark images for 64 fruit-tree sapling classes},
author={Alizadeh, Sayyad and Shamsi, Hamed},
journal={Data in Brief},
pages={112354}… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/SapBark_64_variety_classification.varroa-yolo-baselines-missing-part2-missingcrack-segmentation-dataset
Dataset Card for "crack-segmentation-dataset"
More Information needed
citrus_fruit_variety_classification
Citrus Fruit Variety Classification
A dataset for variety classification of citrus fruits. The dataset contains raw and augmented versions.The raw dataset contains 1,379 images.Images per class:
murcott: 280
ponkan: 328
tangerine: 400
tankan: 371
The augmented dataset contains 7,584 images.Images per class:
murcott: 1,540
ponkan: 1,803
tangerine: 2,200
tankan: 2,041
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
The original… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/citrus_fruit_variety_classification.tomato-variant-a-multiclassvarroa-yolo-baselines-part1-fullvarroa-yolo-baselines-missing-part3-missingvariable_gravity_dataset_creationvarroa-yolo-baselines-missing-part1-missing
