datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CGL-Dataset
Dataset Card for CGL-Dataset
Dataset Summary
CGL-Dataset is a poster layout dataset released with Composition-aware Graphic Layout GAN for Visual-Textual Presentation Designs. The paper studies layout generation for a given image, emphasizing that both global semantics and spatial image composition affect where graphic elements should be placed. The original dataset contains 60,548 advertising posters with annotated layout information.
Supported… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset.CGL-Dataset-v2
Dataset Card for CGL-Dataset v2
Dataset Summary
CGL-Dataset v2 is an advertising-poster layout dataset released with Relation-Aware Diffusion Model for Controllable Poster Layout Generation. The paper argues that poster layouts should account for both visual-textual relationships and geometry relationships between elements. This version extends CGL-Dataset with richer element annotations, text annotations, and text features for controllable poster layout… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset-v2.see-world-1-CGDmm_CGDolca-cge-dataset-v2MATH-lighteval-olympiads_aimemerged_CG_L4_Fmerged_CG_L2_Fmerged_CG_L3_TMATH-lighteval-olympiads_aime-uniquecgu__notas_fiscais
Dataset Card: cgu_notas_fiscais
Data from electronic invoices for federal government purchases made available by
Comptroller General of the Union (Controladoria-Geral da União), which is a
Brazilian federal government agency responsible for oversight and transparency.
Dataset Details
Dataset Description
Curated by: Fred Guth (@fredguth)
Funded by: World Bank
Language(s) (NLP): pt-br
License: CC-BY 4.0
Dataset Sources
The source of this datasets… See the full description on the dataset page: https://huggingface.co/datasets/fredguth/cgu__notas_fiscais.CGM-JEPA-Downstream
CGM-JEPA Downstream Evaluation Splits
Paper | Code
Labeled cohort splits used to evaluate CGM encoders on two binary metabolic outcomes — insulin resistance and β-cell dysfunction — in the paper CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining.
Downstream-only. For the unlabeled pretraining corpus (Stanford + Colas), see CRUISEResearchGroup/CGM-JEPA-Pretraining. For pretrained encoder weights, see… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/CGM-JEPA-Downstream.merged_CG_L2_Tcgap-smallholder-survey-mozambique-2015
CGAP Smallholder Survey - Mozambique 2015 | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/cgap-smallholder-survey-mozambique-2015.merged_CG_L1_Tmerged_CG_L4_TCAPC-CG_V1.0
CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in China
🎉 Published at ACL 2026 (Main Conference) — if you use this dataset in any form, please be sure to cite our paper (see Citation below).
CAPC-CG is the first large-scale open dataset of Chinese central-government policy directives (1949–2023), annotated with a theory-based five-color typology of policy signals — Black (Authorizing), Yellow (Pressuring), Charcoal (Flexible)… See the full description on the dataset page: https://huggingface.co/datasets/Baron-Sun/CAPC-CG_V1.0.geo-treatment-response
GEO RNA-seq Treatment Response Dataset
Pre-treatment RNA-seq studies with patient treatment response annotations,
standardised to Entrez gene IDs and binary responder/non-responder labels.
Studies included (1 total, updated 2026-05-04)
GSE91061 (bulk): | n=51 | advanced melanoma (unresectable or metastatic) | Nivolumab (anti-PD-1) 3 mg/kg IV every 2 weeks; CA209-038 clinical study (NCT01621490); cohort includes ipilimumab-naive (n=33) and ipilimumab-progressed (n=35)… See the full description on the dataset page: https://huggingface.co/datasets/Cgensbigler/geo-treatment-response.vaximere-qa-cg
VaxiMère-QA-CG
Dataset d'intentions multilingue (français fra, lingala lin,
kituba/munukutuba mkw) sur la vaccination pédiatrique au Congo-Brazzaville,
destiné au fine-tuning LoRA d'un classifieur d'intentions (ex. Gemma 3).
Structure
Chaque exemple est un JSON avec 7 champs :
{
"query_id": "Q_001_FR",
"texte": "Pourquoi vacciner mon enfant même s'il semble en bonne santé",
"langue": "fra",
"intention": "UTILITE_VACCIN",
"faq_target_id": "FAQ_001"… See the full description on the dataset page: https://huggingface.co/datasets/Semence/vaximere-qa-cg.CREMP
Dataset Card for CREMP
Conformer-rotamer ensembles of macrocyclic peptides.
Dataset Details
Dataset Description
CREMP: A resource generated for the rapid development and evaluation of machine learning models for macrocyclic peptides. CREMP contains 36,198 unique macrocyclic peptides and their high-quality structural ensembles generated using the Conformer-Rotamer Ensemble Sampling Tool (CREST). Altogether, this dataset contains nearly 31.3 million unique… See the full description on the dataset page: https://huggingface.co/datasets/cgrambow/CREMP.cgap-smallholder-survey-tanzania-2015
CGAP Smallholder Survey - Tanzania 2015 | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/cgap-smallholder-survey-tanzania-2015.cgap-smallholder-survey-uganda-2015
CGAP Smallholder Survey - Uganda 2015 | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/cgap-smallholder-survey-uganda-2015.MATH-lighteval-om220k-Mixed-ghpo-newcgpr-coding-trainMATH-lighteval-DeepScaleR-Mixed-ghpoCGL-Dataset
Dataset Card for CGL-Dataset
Dataset Summary
The CGL-Dataset is a dataset used for the task of automatic graphic layout design for advertising posters. It contains 61,548 samples and is provided by Alibaba Group.
Supported Tasks and Leaderboards
The task is to generate high-quality graphic layouts for advertising posters based on clean product images and their visual contents. The training set and validation set are collections of 60,548 e-commerce… See the full description on the dataset page: https://huggingface.co/datasets/qq3374129162/CGL-Dataset.CGM-MLP_natcomm2023_Cu-C-O
Cite this dataset Zhang, D., Yi, P., Lai, X., Peng, L., and Li, H. CGM-MLP natcomm2023 Cu-C-O. ColabFit, 2024. https://doi.org/10.60732/215303a5
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_xho213jy5pf9_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/CGM-MLP_natcomm2023_Cu-C-O.cG-SchNet
Cite this dataset Gebauer, N. W., Gastegger, M., Hessmann, S. S., Müller, K., and Schütt, K. T. cG-SchNet. ColabFit, 2023. https://doi.org/10.60732/de8af6a2
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_xzaglubh0trq_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/cG-SchNet.MD22_AT_AT_CG_CG
Cite this dataset Chmiela, S., Vassilev-Galindo, V., Unke, O. T., Kabylda, A., Sauceda, H. E., Tkatchenko, A., and Müller, K. MD22 AT AT CG CG. ColabFit, 2023. https://doi.org/10.60732/a87c6d4c
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_rx1ei5q0x9gy_0
Visit the ColabFit Exchange to search additional datasets by author, description… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/MD22_AT_AT_CG_CG.schulz_bank_biotopes
Summary
Image classification dataset with biotope labels extracted from Meyer et al., 2022.
Images were extracted from six remotely operated vehicle (ROV) dives during two SponGES cruises conducted in the summers of 2017 and 2018. The ROV dives were performed by Ægir 6000 and the cruises were conducted on the RV G.O. Sars. The ROV dives traversed across various regions on the seamount, from the base of the seamount at 2700 m depth to the summit at 580 m depth. 600 images were… See the full description on the dataset page: https://huggingface.co/datasets/CGame1/schulz_bank_biotopes.
