datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CG-Bench
CG-Bench
Project Website: https://cg-bench.github.io/leaderboard/GitHub Repository: https://github.com/CG-Bench/CG-Bench (includes running code)
Summary
We introduce CG-Bench, a groundbreaking benchmark for clue-grounded question answering in long videos, addressing the limitations of existing benchmarks that focus primarily on short videos and rely on multiple-choice questions (MCQs). These limitations allow models to answer by elimination rather than genuine… See the full description on the dataset page: https://huggingface.co/datasets/CG-Bench/CG-Bench.cgbench
CG-Bench (mini) — clue-grounded long-video QA
A local mirror of the CG-Bench mini split: 3,000 multiple-choice questions over 1,118
long videos (mean length ~28 min), each question annotated with the clue intervals
(second-level time spans) in the video that actually justify the answer.
Contents
Path
Size
Description
cgbench_mini.json
2.3 MB
3,000 QA items (see schema below)
durations.json
38 KB
{video_uid: duration_in_seconds} for 1,219 videos… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/cgbench.cartoonsetCartoon Set is a collection of random, 2D cartoon avatar images. The cartoons vary in 10 artwork
categories, 4 color categories, and 4 proportion categories, with a total of ~1013 possible
combinations. We provide sets of 10k and 100k randomly chosen cartoons and labeled attributes.prompt_injection_combinedcgu__notas_fiscais
Dataset Card: cgu_notas_fiscais
Data from electronic invoices for federal government purchases made available by
Comptroller General of the Union (Controladoria-Geral da União), which is a
Brazilian federal government agency responsible for oversight and transparency.
Dataset Details
Dataset Description
Curated by: Fred Guth (@fredguth)
Funded by: World Bank
Language(s) (NLP): pt-br
License: CC-BY 4.0
Dataset Sources
The source of this datasets… See the full description on the dataset page: https://huggingface.co/datasets/fredguth/cgu__notas_fiscais.CGM-JEPA-Downstream
CGM-JEPA Downstream Evaluation Splits
Paper | Code
Labeled cohort splits used to evaluate CGM encoders on two binary metabolic outcomes — insulin resistance and β-cell dysfunction — in the paper CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining.
Downstream-only. For the unlabeled pretraining corpus (Stanford + Colas), see CRUISEResearchGroup/CGM-JEPA-Pretraining. For pretrained encoder weights, see… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/CGM-JEPA-Downstream.cgap-smallholder-survey-mozambique-2015
CGAP Smallholder Survey - Mozambique 2015 | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/cgap-smallholder-survey-mozambique-2015.cherokee-english-translation
Cherokee–English Parallel Corpus (Archivist Project)
A curated Cherokee (ᏣᎳᎩ / Tsalagi) ↔ English parallel corpus for machine
translation, assembled from public sources, deduplicated, benchmark-decontaminated,
and conflict-cleaned. Built to train and evaluate English→Cherokee translation
models for one of the most endangered languages in North America.
Files
File
Rows
Purpose
train_en2chr_v2.jsonl
138,307
Flagship training set. English→Cherokee SFT… See the full description on the dataset page: https://huggingface.co/datasets/CGICAI/cherokee-english-translation.CG2RealDataset3geo-treatment-response
GEO RNA-seq Treatment Response Dataset
Pre-treatment RNA-seq studies with patient treatment response annotations,
standardised to Entrez gene IDs and binary responder/non-responder labels.
Studies included (1 total, updated 2026-05-04)
GSE91061 (bulk): | n=51 | advanced melanoma (unresectable or metastatic) | Nivolumab (anti-PD-1) 3 mg/kg IV every 2 weeks; CA209-038 clinical study (NCT01621490); cohort includes ipilimumab-naive (n=33) and ipilimumab-progressed (n=35)… See the full description on the dataset page: https://huggingface.co/datasets/Cgensbigler/geo-treatment-response.CREMP
Dataset Card for CREMP
Conformer-rotamer ensembles of macrocyclic peptides.
Dataset Details
Dataset Description
CREMP: A resource generated for the rapid development and evaluation of machine learning models for macrocyclic peptides. CREMP contains 36,198 unique macrocyclic peptides and their high-quality structural ensembles generated using the Conformer-Rotamer Ensemble Sampling Tool (CREST). Altogether, this dataset contains nearly 31.3 million unique… See the full description on the dataset page: https://huggingface.co/datasets/cgrambow/CREMP.cgap-smallholder-survey-tanzania-2015
CGAP Smallholder Survey - Tanzania 2015 | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/cgap-smallholder-survey-tanzania-2015.cgap-smallholder-survey-uganda-2015
CGAP Smallholder Survey - Uganda 2015 | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/cgap-smallholder-survey-uganda-2015.cga-bench
CGA-Bench Hugging Face Collection
This dataset repo is a collection index for the nine reviewer-facing CGA-Bench dataset descriptors used in the NeurIPS 2026 E&D submission.
Included configs
overview: collection-level summary row spanning the full benchmark release
main_corpus: 19,062-episode primary evaluation corpus
source_grounded: source-grounded SGSC subset
graph_anchored: graph-anchored SGSC subset
profile_expanded: profile-expanded SGSC subset
auto_expanded: 76… See the full description on the dataset page: https://huggingface.co/datasets/cga-bench-neurips26/cga-bench.cgrt-consensus-5model
CGRT Consensus 5-Model Dataset
Multi-model consensus dataset for studying model agreement and disagreement patterns on mathematical reasoning tasks.
Dataset Description
61,678 math problems evaluated by 5 frontier LLMs with full reasoning traces and extracted answers.
Models Used
Model
Provider
Version
Claude
Anthropic
claude-3-5-sonnet-20241022
Codex/GPT-4
OpenAI
gpt-4o
Gemini
Google
gemini-1.5-flash
DeepSeek
DeepSeek
deepseek-chat
Qwen… See the full description on the dataset page: https://huggingface.co/datasets/Adam1010/cgrt-consensus-5model.CG2RealDataset2CGM-MLP_natcomm2023_Cu-C-O
Cite this dataset Zhang, D., Yi, P., Lai, X., Peng, L., and Li, H. CGM-MLP natcomm2023 Cu-C-O. ColabFit, 2024. https://doi.org/10.60732/215303a5
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_xho213jy5pf9_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/CGM-MLP_natcomm2023_Cu-C-O.cG-SchNet
Cite this dataset Gebauer, N. W., Gastegger, M., Hessmann, S. S., Müller, K., and Schütt, K. T. cG-SchNet. ColabFit, 2023. https://doi.org/10.60732/de8af6a2
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_xzaglubh0trq_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/cG-SchNet.MD22_AT_AT_CG_CG
Cite this dataset Chmiela, S., Vassilev-Galindo, V., Unke, O. T., Kabylda, A., Sauceda, H. E., Tkatchenko, A., and Müller, K. MD22 AT AT CG CG. ColabFit, 2023. https://doi.org/10.60732/a87c6d4c
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_rx1ei5q0x9gy_0
Visit the ColabFit Exchange to search additional datasets by author, description… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/MD22_AT_AT_CG_CG.ohsumed
Dataset Card for Ohsumed
Ohsumed collection (available at ftp://medir.ohsu.edu/pub/ohsumed): it includes medical abstracts from the MeSH categories of the year 1991. In [Joachims, 1997] were used the first 20,000 documents divided in 10,000 for training and 10,000 for testing. The specific task was to categorize the 23 cardiovascular diseases categories. After selecting such category subset, the unique abstract number becomes 13,929 (6,286 for training and 7,643 for testing). As… See the full description on the dataset page: https://huggingface.co/datasets/cglez/ohsumed.CGM-MLP_natcomm2023_Cu-C-O_deposition
Cite this dataset Zhang, D., Yi, P., Lai, X., Peng, L., and Li, H. CGM-MLP natcomm2023 Cu-C-O deposition. ColabFit, 2024. https://doi.org/10.60732/ae9380c5
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_xzo1tvni8bay_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/CGM-MLP_natcomm2023_Cu-C-O_deposition.cgap-smallholder-survey-cote-divoire-2016
CGAP Smallholder Survey - Côte d'Ivoire 2016 | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/cgap-smallholder-survey-cote-divoire-2016.llm_guard_datasetCGM-MLP_natcomm2023_screening_carbon-cluster_Cu_test
Cite this dataset Zhang, D., Yi, P., Lai, X., Peng, L., and Li, H. CGM-MLP natcomm2023 screening carbon-cluster@Cu test. ColabFit, 2024. https://doi.org/10.60732/eb6e9ead
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_gkumjkncy8ft_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/CGM-MLP_natcomm2023_screening_carbon-cluster_Cu_test.cgap-smallholder-survey-nigeria-2016
CGAP Smallholder Survey - Nigeria 2016 | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/cgap-smallholder-survey-nigeria-2016.CGM-MLP_natcomm2023_screening_graphite_train
Cite this dataset Zhang, D., Yi, P., Lai, X., Peng, L., and Li, H. CGM-MLP natcomm2023 screening graphite train. ColabFit, 2024. https://doi.org/10.60732/85590078
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_jasbxoigo7r4_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/CGM-MLP_natcomm2023_screening_graphite_train.cgos-9x9-parquet
CGOS 9x9 Computer Go Games
Cleaned 9x9 Go game records from the public monthly archives of the
9x9 Computer Go Server (CGOS),
covering 2015-11-10 to 2026-07-31.
Players are computer Go programs, many of them KataGo-family bots; per-bot
ratings (CGOS BayesElo) are included for strength filtering.
Schema
column
type
description
game_id
string
CGOS game number
date
string
date played (from SGF DT)
white / black
string
bot names
white_rating /… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/cgos-9x9-parquet.CGM-MLP_natcomm2023_screening_carbon-cluster_Cu_train
Cite this dataset Zhang, D., Yi, P., Lai, X., Peng, L., and Li, H. CGM-MLP natcomm2023 screening carbon-cluster@Cu train. ColabFit, 2024. https://doi.org/10.60732/9f0e607d
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_6sjhg1f8fv5j_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/CGM-MLP_natcomm2023_screening_carbon-cluster_Cu_train.cmaj_scale_dataset_v8_CG_moveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 38,
"total_frames": 10159,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:38"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/abemii/cmaj_scale_dataset_v8_CG_move.difraud-combined
