CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceBio /carbon-pretraining-corpus 🧬 Carbon Pretraining Corpus Description 173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model. This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species. Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6-mer tokenizer). A… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus.tabulartext-generation100M<n<1B30 likes3.6k downloads3mo agoHugging Face02AINovice2005 /carbon-tokenized-corpus Dataset Summary AINovice2005/carbon-tokenized-corpus is the tokenized sample of AINovice2005/carbon-cpu-enriched-sequences-sampled . Schema The current dataset contains the following fields: Field Type Description record_id string Source/reference sequence identifier start int64 Start coordinate of the sequence interval end int64 End coordinate of the sequence interval token_ids list Integer token IDs produced by the tokenizer token_mask list… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-tokenized-corpus.tabulartext-generation1M<n<10M0 likes875 downloads1d agoHugging Face03AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes736 downloads9d agoHugging Face04AINovice2005 /carbon-cpu-enriched-sequences-sampledtabular1M<n<10M0 likes668 downloads29d agoHugging Face05AINovice2005 /carbon-likelihood-stats Dataset Summary Dataset Summary AINovice2005/carbon-likelihood-stats is a model-enriched dataset derived from the Carbon genomic corpus, containing sequence-level likelihood statistics for 251,427 sequence regions. Each row links a source sequence record to its corresponding region in the processed/tokenized Carbon corpus using record_id, start, and end. The dataset summarizes both the model's overall likelihood of the observed sequence and the variation in… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-likelihood-stats.tabulartext-generation100K<n<1M0 likes622 downloads9d agoHugging Face06anon-evcarbon-2026 /EV-CarbonBench0 likes522 downloads5mo agoHugging Face07AINovice2005 /carbon-embeddings carbon-embeddings AINovice2005/carbon-embeddings is a derived dataset from the sampled subset of carbon-cpu-enriched-sequences containing dense vector embeddings of biological sequence records. Each row corresponds to a source sequence identified by record_id. The dataset retains the sequence's position within the processed corpus through start and end and provides a numerical embedding representing the sequence in the embedding model's learned representation space. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-embeddings.tabular100K<n<1M0 likes405 downloads9d agoHugging Face08sigit48 /carbon-emission-datalake Carbon Emission Data Lake Dataset ini berisi data historis & prediksi emisi karbon serta suhu untuk wilayah Texas dan Jakarta, dihasilkan secara otomatis oleh pipeline Prefect + dbt. Kolom: generated_date, generated_at, region, live_temperature, base_emission_mt, carbon_emission_forecast_mt, prophet_upper, prophet_lower, is_anomaly (0=Normal, 1=Anomaly, NULL=model belum cukup data), recommended_carbon_cap_mt. Modul ML: Prophet (forecast 30 hari), IsolationForest (deteksi… See the full description on the dataset page: https://huggingface.co/datasets/sigit48/carbon-emission-datalake.tabularn<1K0 likes386 downloads2h agoHugging Face09Sufiyan83 /Low-Carbon-London-Smart-Meter-Cleaned-FeatureReadytext100M<n<1B9 likes358 downloads10mo agoHugging Face10nyu-dice-lab /lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private.tabular100K<n<1M0 likes337 downloads2y agoHugging Face11zhwang1 /CarbonGlobe CarbonGlobe: A Global-Scale, Multi-Decade Dataset and Benchmark for Carbon Forecasting in Forest Ecosystems CarbonGlobe is a global-scale, multi-decade, machine-learning-ready dataset and benchmark for forecasting carbon dynamics in forest ecosystems. The dataset provides harmonized environmental drivers and carbon-related ecosystem outputs simulated by the Ecosystem Demography model version 3 (ED v3), enabling the development, evaluation, and comparison of deep learning models… See the full description on the dataset page: https://huggingface.co/datasets/zhwang1/CarbonGlobe.geospatialtime-series-forecasting10K<n<100K3 likes329 downloads3mo agoHugging Face12AINovice2005 /carbon-pilot-corpustabular1M<n<10M0 likes312 downloads29d agoHugging Face13ai-spatial /CarbonGlobe CarbonGlobe: A Global-Scale, Multi-Decade Dataset and Benchmark for Carbon Forecasting in Forest Ecosystems CarbonGlobe is a global-scale, multi-decade, machine-learning-ready dataset and benchmark for forecasting carbon dynamics in forest ecosystems. The dataset provides harmonized environmental drivers and carbon-related ecosystem outputs simulated by the Ecosystem Demography model version 3 (ED v3), enabling the development, evaluation, and comparison of deep learning models… See the full description on the dataset page: https://huggingface.co/datasets/ai-spatial/CarbonGlobe.geospatialtime-series-forecasting10K<n<100K3 likes302 downloads3mo agoHugging Face14carbon225 /lichess-elitetext1M<n<10M0 likes300 downloads4y agoHugging Face15albertvillanova /carbon_24 Dataset Card for Carbon-24 Dataset Summary Carbon-24 contains 10k carbon materials, which share the same composition, but have different structures. There is 1 element and the materials have 6 - 24 atoms in the unit cells. Carbon-24 includes various carbon structures obtained via ab initio random structure searching (AIRSS) (Pickard & Needs, 2006; 2011) performed at 10 GPa. The original dataset includes 101529 carbon structures, and we selected the 10% of the carbon… See the full description on the dataset page: https://huggingface.co/datasets/albertvillanova/carbon_24.tabularother10K<n<100K1 likes227 downloads4y agoHugging Face16jablonkagroup /cyp3a4_substrate_carbonmangels Dataset Details Dataset Description CYP3A4 is an important enzyme in the body, mainly found in the liver and in the intestine. It oxidizes small foreign organic molecules (xenobiotics), such as toxins or drugs, so that they can be removed from the body. TDC used a dataset from Carbon Mangels et al, which merged information on substrates and nonsubstrates from six publications. Curated by: License: CC BY 4.0 Dataset Sources corresponding publication… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/cyp3a4_substrate_carbonmangels.tabular10K<n<100K0 likes165 downloads1y agoHugging Face17colabfit /carbon_X Cite this dataset Martirossyan, M. M., Egg, T., Hoellmer, P., Karypis, G., Transtrum, M., Roitberg, A., Liu, M., Hennig, R. G., Tadmor, E. B., and Martiniani, S. Carbon X. ColabFit, 2025. https://doi.org/None This dataset has been curated and formatted for the ColabFit Exchange This dataset is also available on the ColabFit Exchange: https://materials.colabfit.org/id/DS_m0z5hh54l9sq_0 Visit the ColabFit Exchange to search additional datasets by author… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/carbon_X.tabularn<1K0 likes156 downloads11mo agoHugging Face18nblancogalindo /carbono-boliviaqa Carbono: BoliviaQA — measuring where AI models are right, wrong, and out of date on Bolivian facts This card carries the full findings essay. The dataset files and field documentation follow it (jump to Fields); the harness, raw graded runs, and statistical appendix live in the GitHub repo. Carbono: BoliviaQA is a 240-question benchmark of verified facts — questions about Bolivia's government, economy, law, demographics, geography, practical life, and the events of 2025–26… See the full description on the dataset page: https://huggingface.co/datasets/nblancogalindo/carbono-boliviaqa.textquestion-answeringn<1K0 likes150 downloads2mo agoHugging Face19scikit-fingerprints /TDC_cyp2c9_substrate_carbonmangelsn<1K0 likes148 downloads2y agoHugging Face20scikit-fingerprints /TDC_cyp3a4_substrate_carbonmangelsn<1K0 likes145 downloads2y agoHugging Face21scikit-fingerprints /TDC_cyp2d6_substrate_carbonmangelsn<1K0 likes143 downloads2y agoHugging Face22open-llm-leaderboard /vicgalle__CarbonBeagle-11B-truthy-detailsgated Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__CarbonBeagle-11B-truthy-details.tabular10K<n<100K0 likes140 downloads2y agoHugging Face23alexroz /CarbonFluxBench CarbonFluxBench: A Global Benchmark for Upscaling of Carbon Fluxes Using Zero-Shot Learning CarbonFluxBench comprises over 1.3 million daily observations from 573 eddy covariance flux tower sites globally (2000–2024). It provides stratified evaluation protocols that explicitly test generalization across unseen vegetation types and climate regimes, a harmonized set of remote sensing and meteorological features, and reproducible baselines ranging from tree-based methods to… See the full description on the dataset page: https://huggingface.co/datasets/alexroz/CarbonFluxBench.geospatialtabular-regression1M<n<10M1 likes140 downloads3mo agoHugging Face24introvoyz041 /Carbon0 likes128 downloads2y agoHugging Face25astro-legacy-archive /gaia-dr3-gold-sample-carbon-stars Gaia DR3 gold sample carbon stars This table lists the Gaia DR3 source identifiers in the gold sample of carbon stars selected from the larger candidate list. ESA describes the selected stars as having C2 and CN molecular bands significantly stronger than ordinary M stars. It is an identifier table for joining the selected sample to other Gaia DR3 measurements. Use python -m venv .venv && .venv/bin/pip install datasets pyarrow from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-gold-sample-carbon-stars.tabular10K<n<100K0 likes119 downloads2d agoHugging Face26ego0op /carbon-projects-unified Earth Love United — Carbon Projects Unified Dataset 23,343 unique carbon projects from 8 registries, normalized, deduplicated, and enriched. Built by Earth Love United Foundation for open climate science. Quick Stats Metric Value Total projects 23,343 Registries 8 (Verra, Gold Standard, CDM, CAR, ACR, CERCARBONO, Isometric, ART) Countries 188 Coordinate coverage 100% Methodology name coverage 94% Methodology code coverage 95.5% Description coverage… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/carbon-projects-unified.tabulartabular-classification10K<n<100K2 likes110 downloads4mo agoHugging Face27carbon-lab /xrr-znpc data/xrr/znpc Local mirror path for: datasets/carbon-lab/xrr-znpc Use: make hf-pull-target TARGET=data/xrr/znpc make hf-push-target TARGET=data/xrr/znpc imagen<1K0 likes106 downloads5mo agoHugging Face28meg /calculate_carbon_runs0 likes99 downloads2y agoHugging Face29carbon225 /poleval-abbreviation-disambiguation-wiki PolEval 2022 Task 2 Pretraining Dataset Dataset Summary Abbreviation disambiguation is the process of expanding abbreviations, e.g. "eng." to the full form "engineer". In the Polish language the task is further complicated because of the many ways to create abbreviations and additional inflected forms. Abbreviation disambiguation was the topic of the 2022 PolEval Competition Task 2. This is the dataset used for pretraining in Jakub Karbowski's competition submission.… See the full description on the dataset page: https://huggingface.co/datasets/carbon225/poleval-abbreviation-disambiguation-wiki.text10M<n<100M0 likes86 downloads3y agoHugging Face3077ethers /CarbonAlpha-train Reasoning-Under-Constraints OpenEnv Meta PyTorch × Scaler OpenEnv Hackathon · April 25–26, 2026 · Bangalore An OpenEnv environment that trains LLMs to reason about competing constraints under ambiguous signals and path-dependent decisions. We flatten a 12-quarter portfolio-manager MDP into a single-turn prompt-completion task, then apply GRPO (via TRL + Unsloth) on Qwen3-4B-Instruct to teach the model to connect news → causal reasoning → portfolio action. Team: Ekansh + brother… See the full description on the dataset page: https://huggingface.co/datasets/77ethers/CarbonAlpha-train.0 likes77 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.