datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CG-Bench
CG-Bench
Project Website: https://cg-bench.github.io/leaderboard/GitHub Repository: https://github.com/CG-Bench/CG-Bench (includes running code)
Summary
We introduce CG-Bench, a groundbreaking benchmark for clue-grounded question answering in long videos, addressing the limitations of existing benchmarks that focus primarily on short videos and rely on multiple-choice questions (MCQs). These limitations allow models to answer by elimination rather than genuine… See the full description on the dataset page: https://huggingface.co/datasets/CG-Bench/CG-Bench.cgbench
CG-Bench (mini) — clue-grounded long-video QA
A local mirror of the CG-Bench mini split: 3,000 multiple-choice questions over 1,118
long videos (mean length ~28 min), each question annotated with the clue intervals
(second-level time spans) in the video that actually justify the answer.
Contents
Path
Size
Description
cgbench_mini.json
2.3 MB
3,000 QA items (see schema below)
durations.json
38 KB
{video_uid: duration_in_seconds} for 1,219 videos… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/cgbench.animal-sounds
Animal Sounds Collection
This dataset contains audio recordings of various animal vocalizations from a range of species, curated to support research in bioacoustics, species classification, and sound event detection. It includes clean and annotated audio samples from the following animals:
Birds
Dogs
Egyptian fruit bats
Giant otters
Macaques
Orcas
Zebra finches
The dataset is designed to be lightweight and modular, making it easy to explore cross-species vocal… See the full description on the dataset page: https://huggingface.co/datasets/cgeorgiaw/animal-sounds.SlimOrcaDedupCleaned
What is this dataset?
Half of the Slim Orca Deduped dataset, but further cleaned by removing instances of soft prompting.
I removed a ton prompt prefixes which did not add any information or were redundant. Ex. "Question:", "Q:", "Write the Answer:", "Read this:", "Instructions:"
I also removed a ton of prompt suffixes which were simply there to lead the model to answer as expected Ex. "The answer is...", "Answer:", "A:", "Summary:", "Output:", "Highlight:"
Why?
I… See the full description on the dataset page: https://huggingface.co/datasets/cgato/SlimOrcaDedupCleaned.cgaxis-3d-models-sample
CGAxis 3D Models - Free Sample (Furniture / Chairs)
A free, licensed sample of human-authored 3D models from CGAxis, a 3D content studio operating since 2008. This sample is a taster of the full CGAxis AI Data corpus (4,390 3D models + 7,794 PBR material sets) available for commercial AI-training licenses.
Every model ships as GLB and USDZ, with geometry statistics, real-world scale in centimetres, semantic tags, a natural-language caption, per-file SHA-256 and a… See the full description on the dataset page: https://huggingface.co/datasets/CGAxis/cgaxis-3d-models-sample.CGL-Dataset
Dataset Card for CGL-Dataset
Dataset Summary
CGL-Dataset is a poster layout dataset released with Composition-aware Graphic Layout GAN for Visual-Textual Presentation Designs. The paper studies layout generation for a given image, emphasizing that both global semantics and spatial image composition affect where graphic elements should be placed. The original dataset contains 60,548 advertising posters with annotated layout information.
Supported… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset.CGL-Dataset-v2
Dataset Card for CGL-Dataset v2
Dataset Summary
CGL-Dataset v2 is an advertising-poster layout dataset released with Relation-Aware Diffusion Model for Controllable Poster Layout Generation. The paper argues that poster layouts should account for both visual-textual relationships and geometry relationships between elements. This version extends CGL-Dataset with richer element annotations, text annotations, and text features for controllable poster layout… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset-v2.see-world-1-CGDarcticDEM-32m
ArcticDEM 32 m Mosaic
Dataset Summary
A publicly-accessible digital elevation mosaic covering the Arctic region, derived from the Polar Geospatial Center’s ArcticDEM project, at 32 m horizontal resolution.
📘 Dataset Description
SourceProvided by the University of Minnesota’s Polar Geospatial Center (PGC), part of the ArcticDEM initiative funded by NSF & NGA (website).
Spatial CoveragePan-Arctic: regions north of ~60° N (Greenland, Canada, Russia, Alaska… See the full description on the dataset page: https://huggingface.co/datasets/cgeorgiaw/arcticDEM-32m.mm_CGDolca-cge-dataset-v2MATH-lighteval-olympiads_aimeperi-kbprompt_injection_password_or_secretbionlp_st_2013_cgthe Cancer Genetics (CG) is a event extraction task and a main task of the BioNLP Shared Task (ST) 2013.
The CG task is an information extraction task targeting the recognition of events in text,
represented as structured n-ary associations of given physical entities. In addition to
addressing the cancer domain, the CG task is differentiated from previous event extraction
tasks in the BioNLP ST series in addressing a wide range of pathological processes and multiple
levels of biological organization, ranging from the molecular through the cellular and organ
levels up to whole organisms. Final test set submissions were accepted from six teamsCG-AV-Counting
CG-AV-Counting
Updates
[2025/07/22]
Since errors in a few clue annotations when converting frame indexes to timestamps, there were errors in the previous benchmark leaderboard, we have reevaluated all models and have updated the new leaderboard.
Summary
Despite progress in video understanding, current MLLMs struggle with counting tasks. Existing benchmarks are limited by short videos, close-set queries, lack of clue annotations, and weak… See the full description on the dataset page: https://huggingface.co/datasets/CG-Bench/CG-AV-Counting.merged_CG_L4_FJam-CGPT
Jam-CGPT
Jam-CGPT dataset contains the summary generated by using GPT-3.5. The dataset size ranges from 170k to 2.15m. We follow Jam's procedure to compile the dataset for finetuning.
Jam-CGPT dataset files
Filename
Description
170k.tar.gz
170k summary train and val bin file
620k.tar.gz
620k summary train and val bin file
1.25m.tar.gz
1.25m summary train and val bin file
2.15m.tar.gz
2.15m summary train and val bin file
jam_cgpt_test.tar.gz
Jam-CGPT… See the full description on the dataset page: https://huggingface.co/datasets/apcl/Jam-CGPT.prompt_injection_ctf_dataset_2cgc_datacgr_tensorlab_umsommerged_CG_L2_FCGTSF
CGTSF: Context-Guided Time Series Forecasting
✨ Introduction
The context-guided time series forecasting task entails the transformation of text into time series data. Relevant multimodal datasets are limited. To address these data gaps, we have collected three multimodal datasets that offer valuable resources for future research. The following table summarizes the statistics of these datasets. MSPG comprises 13 months of solar power generation data on 27 photovoltaic… See the full description on the dataset page: https://huggingface.co/datasets/ChengsenWang/CGTSF.Mistral_Trivia-QA_Dataset
Mistral Trivia QA Dataset
The Mistral Trivia QA Dataset is a collection of trivia questions and answers designed to evaluate and train question-answering models. It covers a wide range of topics and is particularly useful for assessing a model's ability to handle general knowledge and reasoning tasks.The documents are derived from WikiText-2, providing diverse and well-structured textual content suitable for extractive QA generation.
Model outputs for this dataset were generated… See the full description on the dataset page: https://huggingface.co/datasets/CGU-Widelab/Mistral_Trivia-QA_Dataset.merged_CG_L3_TMATH-lighteval-olympiads_aime-uniqueqvhighlight_internvideo2_llama_text_featuretailor-cgo
Dataset Card for Tailor-CGO
This dataset contains evaluations of language-model-generated responses regarding vaccine concerns, where each response is tailored to establish common ground through an identified "Common-Ground Opinion".
Dataset Details
Dataset Description
The dataset contains both human- and LLM-annotated preferences/scores for how "well tailored" each written response is. Annotations are structured as a (1) relative preference between two… See the full description on the dataset page: https://huggingface.co/datasets/DukeNLP/tailor-cgo.qvhighlight_internvideo2_videoclip_6b_w2sprompt_injection_combined
