datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
carbon-pretraining-corpus
🧬 Carbon Pretraining Corpus
Description
173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model.
This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species.
Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6-mer tokenizer). A… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus.carbon-tokenized-corpus
Dataset Summary
AINovice2005/carbon-tokenized-corpus is the tokenized sample of AINovice2005/carbon-cpu-enriched-sequences-sampled .
Schema
The current dataset contains the following fields:
Field
Type
Description
record_id
string
Source/reference sequence identifier
start
int64
Start coordinate of the sequence interval
end
int64
End coordinate of the sequence interval
token_ids
list
Integer token IDs produced by the tokenizer
token_mask
list… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-tokenized-corpus.carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.carbon-cpu-enriched-sequences-sampledcarbon-likelihood-stats
Dataset Summary
Dataset Summary
AINovice2005/carbon-likelihood-stats is a model-enriched dataset derived from the Carbon genomic corpus, containing sequence-level likelihood statistics for 251,427 sequence regions.
Each row links a source sequence record to its corresponding region in the processed/tokenized Carbon corpus using record_id, start, and end. The dataset summarizes both the model's overall likelihood of the observed sequence and the variation in… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-likelihood-stats.EV-CarbonBenchcarbon-embeddings
carbon-embeddings
AINovice2005/carbon-embeddings is a derived dataset from the sampled subset of carbon-cpu-enriched-sequences containing dense vector embeddings of biological sequence records.
Each row corresponds to a source sequence identified by record_id. The dataset retains the sequence's position within the processed corpus through start and end and provides a numerical embedding representing the sequence in the embedding model's learned representation space.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-embeddings.carbon-emission-datalake
Carbon Emission Data Lake
Dataset ini berisi data historis & prediksi emisi karbon serta suhu untuk wilayah
Texas dan Jakarta, dihasilkan secara otomatis oleh pipeline Prefect + dbt.
Kolom: generated_date, generated_at, region, live_temperature,
base_emission_mt, carbon_emission_forecast_mt, prophet_upper, prophet_lower,
is_anomaly (0=Normal, 1=Anomaly, NULL=model belum cukup data), recommended_carbon_cap_mt.
Modul ML: Prophet (forecast 30 hari), IsolationForest (deteksi… See the full description on the dataset page: https://huggingface.co/datasets/sigit48/carbon-emission-datalake.Low-Carbon-London-Smart-Meter-Cleaned-FeatureReadylm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private
Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy
Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private.CarbonGlobe
CarbonGlobe: A Global-Scale, Multi-Decade Dataset and Benchmark for Carbon Forecasting in Forest Ecosystems
CarbonGlobe is a global-scale, multi-decade, machine-learning-ready dataset and benchmark for forecasting carbon dynamics in forest ecosystems. The dataset provides harmonized environmental drivers and carbon-related ecosystem outputs simulated by the Ecosystem Demography model version 3 (ED v3), enabling the development, evaluation, and comparison of deep learning models… See the full description on the dataset page: https://huggingface.co/datasets/zhwang1/CarbonGlobe.carbon-pilot-corpusCarbonGlobe
CarbonGlobe: A Global-Scale, Multi-Decade Dataset and Benchmark for Carbon Forecasting in Forest Ecosystems
CarbonGlobe is a global-scale, multi-decade, machine-learning-ready dataset and benchmark for forecasting carbon dynamics in forest ecosystems. The dataset provides harmonized environmental drivers and carbon-related ecosystem outputs simulated by the Ecosystem Demography model version 3 (ED v3), enabling the development, evaluation, and comparison of deep learning models… See the full description on the dataset page: https://huggingface.co/datasets/ai-spatial/CarbonGlobe.lichess-elitecarbon_24
Dataset Card for Carbon-24
Dataset Summary
Carbon-24 contains 10k carbon materials, which share the same composition, but have different structures. There is 1 element and the materials have 6 - 24 atoms in the unit cells.
Carbon-24 includes various carbon structures obtained via ab initio random structure searching (AIRSS) (Pickard & Needs, 2006; 2011) performed at 10 GPa.
The original dataset includes 101529 carbon structures, and we selected the 10% of the carbon… See the full description on the dataset page: https://huggingface.co/datasets/albertvillanova/carbon_24.cyp3a4_substrate_carbonmangels
Dataset Details
Dataset Description
CYP3A4 is an important enzyme in the body, mainly found in the liver
and in the intestine. It oxidizes small foreign organic molecules (xenobiotics),
such as toxins or drugs, so that they can be removed from the body. TDC used
a dataset from Carbon Mangels et al, which merged information on substrates
and nonsubstrates from six publications.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/cyp3a4_substrate_carbonmangels.carbon_X
Cite this dataset Martirossyan, M. M., Egg, T., Hoellmer, P., Karypis, G., Transtrum, M., Roitberg, A., Liu, M., Hennig, R. G., Tadmor, E. B., and Martiniani, S. Carbon X. ColabFit, 2025. https://doi.org/None
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_m0z5hh54l9sq_0
Visit the ColabFit Exchange to search additional datasets by author… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/carbon_X.carbono-boliviaqa
Carbono: BoliviaQA — measuring where AI models are right, wrong, and out of date on Bolivian facts
This card carries the full findings essay. The dataset files and field
documentation follow it (jump to Fields); the harness, raw graded
runs, and statistical appendix live in the
GitHub repo.
Carbono: BoliviaQA is a 240-question benchmark of verified facts — questions
about Bolivia's government, economy, law, demographics, geography, practical
life, and the events of 2025–26… See the full description on the dataset page: https://huggingface.co/datasets/nblancogalindo/carbono-boliviaqa.TDC_cyp2c9_substrate_carbonmangelsTDC_cyp3a4_substrate_carbonmangelsTDC_cyp2d6_substrate_carbonmangelsvicgalle__CarbonBeagle-11B-truthy-details
Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy
Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__CarbonBeagle-11B-truthy-details.CarbonFluxBench
CarbonFluxBench: A Global Benchmark for Upscaling of Carbon Fluxes Using Zero-Shot Learning
CarbonFluxBench comprises over 1.3 million daily observations from 573 eddy covariance flux tower sites globally (2000–2024). It provides stratified evaluation protocols that explicitly test generalization across unseen vegetation types and climate regimes, a harmonized set of remote sensing and meteorological features, and reproducible baselines ranging from tree-based methods to… See the full description on the dataset page: https://huggingface.co/datasets/alexroz/CarbonFluxBench.Carbongaia-dr3-gold-sample-carbon-stars
Gaia DR3 gold sample carbon stars
This table lists the Gaia DR3 source identifiers in the gold sample of carbon stars selected from the larger candidate list. ESA describes the selected stars as having C2 and CN molecular bands significantly stronger than ordinary M stars. It is an identifier table for joining the selected sample to other Gaia DR3 measurements.
Use
python -m venv .venv && .venv/bin/pip install datasets pyarrow
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-gold-sample-carbon-stars.carbon-projects-unified
Earth Love United — Carbon Projects Unified Dataset
23,343 unique carbon projects from 8 registries, normalized, deduplicated, and enriched.
Built by Earth Love United Foundation for open climate science.
Quick Stats
Metric
Value
Total projects
23,343
Registries
8 (Verra, Gold Standard, CDM, CAR, ACR, CERCARBONO, Isometric, ART)
Countries
188
Coordinate coverage
100%
Methodology name coverage
94%
Methodology code coverage
95.5%
Description coverage… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/carbon-projects-unified.xrr-znpc
data/xrr/znpc
Local mirror path for:
datasets/carbon-lab/xrr-znpc
Use:
make hf-pull-target TARGET=data/xrr/znpc
make hf-push-target TARGET=data/xrr/znpc
calculate_carbon_runspoleval-abbreviation-disambiguation-wiki
PolEval 2022 Task 2 Pretraining Dataset
Dataset Summary
Abbreviation disambiguation is the process of expanding abbreviations, e.g. "eng." to the full form "engineer".
In the Polish language the task is further complicated because of the many ways to create abbreviations and additional inflected forms.
Abbreviation disambiguation was the topic of the 2022 PolEval Competition Task 2.
This is the dataset used for pretraining in Jakub Karbowski's competition submission.… See the full description on the dataset page: https://huggingface.co/datasets/carbon225/poleval-abbreviation-disambiguation-wiki.CarbonAlpha-train
Reasoning-Under-Constraints OpenEnv
Meta PyTorch × Scaler OpenEnv Hackathon · April 25–26, 2026 · Bangalore
An OpenEnv environment that trains LLMs to reason about competing constraints under ambiguous signals and path-dependent decisions. We flatten a 12-quarter portfolio-manager MDP into a single-turn prompt-completion task, then apply GRPO (via TRL + Unsloth) on Qwen3-4B-Instruct to teach the model to connect news → causal reasoning → portfolio action.
Team: Ekansh + brother… See the full description on the dataset page: https://huggingface.co/datasets/77ethers/CarbonAlpha-train.
