datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MEBench-2K1UOmnimodal-Agent-SFT-2K
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.assessment2_spheres_and_cube_2k
Dataset Card for cilp_assessment_all
This is a FiftyOne dataset with 2000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("maxspeer/assessment2_spheres_and_cube_2k_2")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/maxspeer/assessment2_spheres_and_cube_2k.ai-vs-real-2k-imagesRSVQA-LR-2kA 2k subset of the validation split of the RSVQA LR dataset ported to HF for ease-of-use in quick remote sensing VQA evaluation.
For more information and attribution please refer to the original dataset: https://rsvqa.sylvainlobry.com/#dataset
DA-2K
DA-2K Evaluation Benchmark
Introduction
DA-2K is proposed in Depth Anything V2 to evaluate the relative depth estimation capability. It encompasses eight representative scenarios of indoor, outdoor, non_real, transparent_reflective, adverse_style, aerial, underwater, and object. It consists of 1K diverse high-quality images and 2K precise pair-wise relative depth annotations.
Please refer to our paper for details in constructing this benchmark.
Usage
Please… See the full description on the dataset page: https://huggingface.co/datasets/depth-anything/DA-2K.deeplesion-balanced-2k
DeepLesion Benchmark Subset (Balanced 2K)
This dataset is a curated subset of the DeepLesion dataset, prepared for demonstration and benchmarking purposes. It consists of 2,000 CT lesion samples, balanced across 8 coarse lesion types, and filtered to include lesions with a short diameter > 10mm.
Dataset Details
Source: DeepLesion
Institution: National Institutes of Health (NIH) Clinical Center
Subset size: 2,000 images
Lesion types: lung, abdomen, mediastinum, liver… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/deeplesion-balanced-2k.RSVQA-HR-2kA 2k subset of the validation split of the RSVQA HR dataset ported to HF for ease-of-use in quick remote sensing VQA evaluation.
For more information and attribution please refer to the original dataset: https://rsvqa.sylvainlobry.com/#dataset
Part-Affordance-2KVisA-2KHigh-resolution industrial image anomaly detection dataset VisA-2K.For more information, see HiAD.
Download
huggingface-cli download --repo-type dataset XimiaoZhang/VisA-2K --local-dir VisA-2K --resume-download
echo4o_2klua-coco100-rebuttal-2k
LUA COCO-100 Rebuttal 2K Generations
This dataset contains 2048x2048 generations for a deterministic 100-example
MS COCO val2017 subset using official COCO captions. It was prepared for the
LUA rebuttal domain-validation task.
Contents
prompts.csv: selected prompts. The gpt_caption column is JSON and uses
the same sdxl field convention as the previous competitor protocol.
manifest.json: selected COCO image ids and metadata.
references/: COCO reference images saved as… See the full description on the dataset page: https://huggingface.co/datasets/vaskers5/lua-coco100-rebuttal-2k.meta_assets_2kThis dataset are all the 3d models generated by DTC.
I added collision mesh, mass, friction and a general description of each objects.
HRVQA-2kA 2k subset of the validation split of the HRVQA dataset ported to HF for ease-of-use in quick remote sensing VQA evaluation.
For more information and attribution please refer to the original dataset: https://hrvqa.nl/
seedream-4.5-generated-2k
SeeDream 4.5 (2K) Dataset
200 AI-generated images at 2K quality.
License: MIT
84k_40GB_Captioned_2k_XXX_ImagesImages have been captioned using Llama, Mistral and Qwen.
10KwH used to caption and process dataset.
RSVLM-QA-2kChemVQA-2K
🧪 ChemVQA-2K: A Visual Question Answering Dataset for Molecular Understanding
📘 Overview
ChemVQA-2K is a novel Visual Question Answering (VQA) dataset designed to bridge chemistry and multimodal AI.
It contains approximately 2,000 high-resolution molecular images (512×512) generated from valid SMILES strings, accompanied by 10 structured Q&A pairs per molecule, resulting in ~20,000 image-question-answer triplets.
Each image represents a 2D chemical structure rendered using RDKit, while each… See the full description on the dataset page: https://huggingface.co/datasets/chandrabhuma/ChemVQA-2K.llava-en-zh-2kThis dataset is composed by
1k examples of English Visual Instruction Data from LLaVA.
1k examples of English Visual Instruction Data from openbmb.
You can organize content in the dataset_info.json in LLaMA Factory like this:
"llava_1k_en": {
"hf_hub_url": "BUAADreamer/llava-en-zh-2k",
"subset": "en",
"formatting": "sharegpt",
"columns": {
"messages": "messages",
"images": "images"
},
"tags": {
"role_tag": "role",
"content_tag": "content"… See the full description on the dataset page: https://huggingface.co/datasets/BUAADreamer/llava-en-zh-2k.Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/krajavi3/Twin-2K-500.DL3DV-2k
DL3DV-2K
📖Paper
| 🏠Homepage
| 🤗ETCHR-FLUX.2-klein-9B Model
| 🤗ETCHR SFT-400K Dataset
| 🤗ETCHR GRPO-10K Dataset
| 🤗DL3DV-2K Benchmark
DL3DV-2K is a benchmark constructed from the DL3DV dataset for evaluating the viewpoint transformation capability of large models in spatial reasoning tasks, comprising 2K samples in total. Each sample contains: images (original images), aux_images (transformed images provided for human reference only and not used as question input)… See the full description on the dataset page: https://huggingface.co/datasets/internlm/DL3DV-2k.PhaseStructVQA-2KHIM-2K
HIM-2K
HIM-2K is a human instance matting benchmark introduced in
"Human Instance Matting via Mutual Guidance and Multi-Instance Refinement"
(InstMatt, Sun et al., CVPR 2022). Please cite the original authors and respect
the non-commercial (CC BY-NC 4.0) license.
Schema
One row per image:
image — the RGB photo (datasets.Image).
mask — a list of per-instance alpha mattes for that image
(List(Image(...))); each element is one human instance alpha, ordered by the… See the full description on the dataset page: https://huggingface.co/datasets/nobg/HIM-2K.Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/chadreadey/Twin-2K-500.refCOCOg_2k_840
Seg-Zero Dataset: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
This repository hosts the training dataset introduced in the paper Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.
Abstract
Traditional methods for reasoning segmentation rely on supervised fine-tuning with categorical labels and simple descriptions, limiting its out-of-domain generalization and lacking explicit reasoning processes. To address these limitations, we… See the full description on the dataset page: https://huggingface.co/datasets/Ricky06662/refCOCOg_2k_840.Twin-2K-500_edit
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Shar999/Twin-2K-500_edit.Himawari-8-9-band03-500m-2km-paired249,058 pairs of Himawari-8 & 9 band 03(red) visible imagery, in 500m and downscaled to 2km.
note: 500m & 2km is nadir spatial resolution, doesn't meet this at high zenith angle near the edge.500m in 2048x2048, with 2000x2000 in the center being data with white borders on the edge.2km in 512x512, 500x500 in the center being data.
Includes almost all Target Area imagery from July 2015 - Sep 2023, and a small portion of randomly sliced imagery from full disk.This dataset contains roughly 1/4 -… See the full description on the dataset page: https://huggingface.co/datasets/Dapiya/Himawari-8-9-band03-500m-2km-paired.MVTec-2KHigh-resolution industrial image anomaly detection dataset MVTec-2K.For more information, see HiAD.
Download
huggingface-cli download --repo-type dataset XimiaoZhang/MVTec-2K --local-dir MVTec-2K --resume-download
plotqa_2k
Dataset Card for "plotqa_2k"
More Information needed
