datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
summertimerender
Bangumi Image Base of Summertime Render
This is the image base of bangumi Summertime Render, we detected 32 characters, 2981 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/summertimerender.minecraft-skins-captioned-900k
🎮 minecraft-skins-captioned-900k
854,116 high-quality, captioned Minecraft player skins — deduplicated, Steve-model only, ready for text-to-image training.
📋 Dataset Summary
A rigorously filtered and quality-controlled version of neurlang/Minecraft-Skins-Captioned-1M specifically curated for training high-performance generative models that require precise UV topology constraints.
This dataset is optimized for models like ST-DiT (Sparse Template-Aware… See the full description on the dataset page: https://huggingface.co/datasets/summykai/minecraft-skins-captioned-900k.TextCaps-Caption-Summary
Description
Multiple Captions of TextCaps dataset summarized into one using slauw87/bart_summarisation BART model.
OCT-summary-Datasetlca-module-summarization
🏟️ Long Code Arena (Module summarization)
This is the benchmark for Module summarization task as part of the
🏟️ Long Code Arena benchmark.
The current version includes 216 manually curated text files describing different documentation of open-source permissive Python projects.
The model is required to generate such description, given the relevant context code and the intent behind the documentation.
All the repositories are published under permissive licenses (MIT, Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-module-summarization.summarize_from_feedback_oai_preprocessing_1711138793
Dataset Card for "summarize_from_feedback_oai_preprocessing_1711138793"
More Information needed
summarize_from_feedback_oai_preprocessing_1711138084
Dataset Card for "summarize_from_feedback_oai_preprocessing_1711138084"
More Information needed
chemistry-jsdhist3-summary
Chemistry SDPO jsdhist3 Training History
This repository contains summarized training-history artifacts for four completed Chemistry SDPO Qwen3-4B runs from /workspace/SDPO-new-clean.
Images
Files
data/jsd_history.csv: per-step JSD scalar history and validation metrics parsed from the training log.
data/run_summary.csv: one-row-per-run summary with final/best validation reward and histogram event counts.… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/chemistry-jsdhist3-summary.my-NFT-summer
Dataset Card for "my-NFT-summer"
More Information needed
my-NFT-summer-balanced1
Dataset Card for "my-NFT-summer-balanced1"
More Information needed
LeetCode_YT_CC_CoT_Summarymedieval_decoration_summaries
Medieval Decoration Summaries
This dataset contains per-page pixel summaries and semantic segmentation maps
for 60,411 pages from 175 medieval manuscripts held by 29 archives and
indexed in Lilian Randall's Images in the Margins of Gothic Manuscripts (1966).
The segmentation maps were produced by a computer vision pipeline with
foreground mIoU of 60%, and therefore should not be treated as ground truth.
Each row represents one manuscript page.
Fields… See the full description on the dataset page: https://huggingface.co/datasets/wellesley-easel/medieval_decoration_summaries.stl10
Dataset Card for "stl10"
More Information needed
large-scale-multimodal-multilingual-summarization-datasetPlease cite this paper if you use our code or data:
@inproceedings{verma-etal-2023-large,
title = "Large Scale Multi-Lingual Multi-Modal Summarization Dataset",
author = "Verma, Yash and
Jangra, Anubhav and
Verma, Raghvendra and
Saha, Sriparna",
booktitle = "Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics",
month = may,
year = "2023",
address = "Dubrovnik, Croatia",
publisher =… See the full description on the dataset page: https://huggingface.co/datasets/Zenquiorra/large-scale-multimodal-multilingual-summarization-dataset.svhn
Dataset card for SVHN
The Street View House Numbers (SVHN) dataset is a real-world image dataset developed and designed for machine learning and object recognition algorithms, and is characterized by low data preprocessing and formatting requirements. Similar to MNIST, SVHN contains images of small cropped numbers, but in terms of labeled data, SVHN is an order of magnitude larger than MNIST, comprising over 600,000 digital images. Unlike MNIST, SVHN deals with a much more… See the full description on the dataset page: https://huggingface.co/datasets/summerhot12138/svhn.summer2winter_yosemiteSpectraGAN-SEN12MS_Summer_1kDReSS-part1SpectraGAN-SEN12MS_Summer2026-Code-for-America-SummitViSIL_Multimodal-Video-Summary
ViSIL Dataset
This dataset contains the multimodal video summaries used in the ViSIL paper. The video clips are sampled from MVBench and LongVideoBench.
For the raw video data, please refer to the original video datasets: OpenGVLab/MVBench and longvideobench/LongVideoBench.
Illustrative Example of Multimodal Video Summaries
Dataset Structure
ViSILMultimodalVideoSummary/
├── README.md
├── visualizer.py
├── metadata/
├── video_summary.csv
├──… See the full description on the dataset page: https://huggingface.co/datasets/Po-han/ViSIL_Multimodal-Video-Summary.cifar10
Dataset Card for CIFAR-10
Dataset Summary
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/summerhot12138/cifar10.multilingual_transcriptions_summarized_by_english_backtranslated_finalclubbed_medical_summariessummer-buildings-extraction-coco-hfsummary_pixmirror-data-label-1draft_medical_summaries
Dataset Card for "draft_medical_summaries"
More Information needed
benchname-module-summarization
🥷 BenchName (Module summarization)
This is the benchmark for Module summarization task as part of the
🥷 BenchName benchmark.
The current version includes 216 manually curated text files describing different documentation of open-source permissive Python projects.
The model is required to generate such description, given the relevant context code and the intent behind the documentation.
All the repositories are published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause… See the full description on the dataset page: https://huggingface.co/datasets/anon-iclr-submission/benchname-module-summarization.mini-imagenet
