datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CellularTrafficPrediction2025_Virtual_Cell_Challenge_Test_Data2026.RA.Frontier-and-Scale-Cells
Rational-Agent Frontier, Scale, and Framing Cells
This public dataset is a sibling of siddharthmb/2026.RA.Negotiation-Campaigns (the frozen P1-P4 experimental record for the ii_mats/experiments/rational_agents negotiation program) and follows the same conventions: raw per-episode JSON, per-turn oracle annotations, Markdown/HTML transcripts, run manifests, analysis tables, and an integrity manifest over every uploaded file. It packages eight later campaigns that were run against… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Frontier-and-Scale-Cells.CelestiaCelestia is a dataset containing science-instruct data.
The 2024-10-30 version contains:
126k rows of synthetic science-instruct data, using synthetically generated prompts and responses generated using Llama 3.1 405b Instruct. Primary subjects are physics, chemistry, biology, and computer science; secondary subjects include Earth science, astronomy, and information theory.
This dataset contains synthetically generated data and has not been subject to manual review.
Celestia3-DeepSeek-R1-0528Click here to support our open-source dataset and model releases!
Celestia3-DeepSeek-R1-0528 is a dataset focused on science, testing the limits of DeepSeek R1 0528's science-reasoning skills!
This dataset contains:
90.9k synthetically generated science prompts, with all responses generated using DeepSeek R1 0528.
Primary subjects are physics, chemistry, biology, and computer science; secondary subjects include Earth science, astronomy, and information theory.
All prompts are synthetic, taken… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Celestia3-DeepSeek-R1-0528.east-asian-celebrity-bazi
East Asian Celebrity BaZi Dataset (Three Pillars)
2,284 public figures from Greater China, Korea, and Japan, with birth dates mapped to Four Pillars (BaZi / 四柱八字) year, month, and day pillars — including stems, branches, five-element attributions, hidden stems, and Ten Gods relations.
Curated by: Shunshi.AI — an AI divination and astrology platform covering Four Pillars (BaZi), Zi Wei Dou Shu, Western natal astrology, tarot, and I Ching
Computed with: shunshi-bazi-core —… See the full description on the dataset page: https://huggingface.co/datasets/shunshi-ai/east-asian-celebrity-bazi.cello_data
CELLO 10X-Xenium-52
Preprocessed data for CELLO, which predicts the gene expression of every cell in an H&E
whole-slide image from the image and the cell locations.
The dataset pairs 52 H&E whole-slide images with single-cell Xenium spatial transcriptomics from
HEST-1k. It covers 12 organs and about 9.5 million
cells, split by sample:
split
samples
organs
tiles
cells
train
36
Bowel, Breast, Kidney, Liver, Lung, Lymphoid, Pancreas, Skin
381,775
5,862,560
val
4
Bowel… See the full description on the dataset page: https://huggingface.co/datasets/gaozijun/cello_data.text-to-mermaidcell-service-data
Dataset Card: Synthetic Mobile Network Performance
Dataset Description
This dataset contains synthetically generated mobile signal measurements designed to mirror real-world data in the UK. The data represents geolocated signal quality metrics from mobile devices, capturing a range of environmental and temporal conditions over several months in 2025.
All data has been anonymized, aggregated, and processed to protect user privacy. The synthetic dataset has undergone pre-processing… See the full description on the dataset page: https://huggingface.co/datasets/joefee/cell-service-data.COOPER
📡 COOPER
Cellular Operational Observations for Performance and Evaluation Research
An Open Benchmark of Synthetic Mobile Network Performance Indicators for Reproducible Research
🧭 Overview
COOPER is an open-source synthetic dataset of mobile network performance measurement (PM) time series, designed to support reproducible AI/ML research in wireless networks. The dataset is named in honor of Martin Cooper, a pioneer of cellular communications.… See the full description on the dataset page: https://huggingface.co/datasets/CelfAI/COOPER.Celestia2Celestia 2 is a multi-turn agent-instruct dataset containing science data.
This dataset focuses on challenging multi-turn conversations and contains:
176k rows of synthetic multi-turn science-instruct data, using Microsoft's AgentInstruct style. All prompts and responses are synthetically generated using Llama 3.1 405b Instruct. Primary subjects are physics, chemistry, biology, and computer science; secondary subjects include Earth science, astronomy, and information theory.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Celestia2.Celestia3-DeepSeek-R1-0528-PREVIEWClick here to support our open-source dataset and model releases!
This is an early sneak preview of Celestia3-DeepSeek-R1-0528, containing the first 13.4k rows!
Celestia3-DeepSeek-R1-0528 is a dataset focused on science, testing the limits of DeepSeek R1's science-reasoning skills!
This early preview release contains:
13.4k synthetically generated science prompts. All responses are generated using DeepSeek R1 0528.
Primary subjects are physics, chemistry, biology, and computer science;… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Celestia3-DeepSeek-R1-0528-PREVIEW.text-to-mermaid-2celeb-face-matching-dataafrica-synth-sickle-cell-dataset-all
African Sickle Cell Disease Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-sickle-cell-dataset-all.DomesticNames_AllStates_Text Domestic Names from the Federal Government's repository of official geographic names [CSV dataset]
This Dataset includes 980,065 geographic names as of September 10, 2023.
It is apparent that no currently released LLMs are pretrained on datasets with many of these geographic names (i.e., features), descriptions, and histories.
Example: feature_name: Abercrombie Gulch
GPT-3.5 responds "I'm not aware of a specific location called Abercrombie Gulch in my training data,..." when prompted… See the full description on the dataset page: https://huggingface.co/datasets/cellos/DomesticNames_AllStates_Text.Wikidata-celebrity-parentequivalent-cell-ceaac6
equivalent-cell-ceaac6
Synthetic products test data: 35 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/jessica-anderson97/equivalent-cell-ceaac6.international-celebration-7366b6
international-celebration-7366b6
Synthetic products test data: 39 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number… See the full description on the dataset page: https://huggingface.co/datasets/Kestrel-Kai/international-celebration-7366b6.embarrassed-celebration-1a70d8
embarrassed-celebration-1a70d8
Synthetic sensors test data: 49 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting… See the full description on the dataset page: https://huggingface.co/datasets/pixelbridge/embarrassed-celebration-1a70d8.celltransformer_materialsCellTissue-dict
CellTissue-dict
This dictionary is comprised of a merge of two source dictionaries (from Bioportal and ChEMBL), used to tag cell and tissue descriptions in text.
Dataset Details
Dataset Sources
This dictionary was produced by a combination of quick data sourcing, followed by partial manual categorising of terms.
Repository: https://github.com/ML4LitS/OTAR3088
Direct Use
This data may be used to tag phrases in natural language, categorising… See the full description on the dataset page: https://huggingface.co/datasets/OTAR3088/CellTissue-dict.personalityaviation-convective-storm-cell-avoidance-coherence-risk-v0.1What this repo is for
Detect when convective storms
drive airspace disruption
and whether the response matches the hazard.
Flags
high storm density with weak rerouting
heavy holding with no storm driver
delay cascades that do not match convection signal
climate-anomaly-cells-50-locations-1990-2024
Climate Anomaly Cells (daily, 50 locations, 1990–2024)
Daily climate anomalies for 50 curated high-exposure urban/rural cells worldwide, 1990-01-01 to 2024-12-31, derived from NASA POWER (GMAO MERRA-2 reanalysis). Each cell has observed daily temperature, precipitation, wind and humidity; smoothed 1991–2020 day-of-year normals; additive anomalies; a hot-day flag; NWS heat index and wind chill; and an illustrative composite stress index. One UTF-8 CSV, 639,200 rows × 39 columns… See the full description on the dataset page: https://huggingface.co/datasets/xquantize/climate-anomaly-cells-50-locations-1990-2024.compositional_celebritiesmolecule-embeddings-kalmcelegans_connectome_dataCITATION
Q. Simeon, L. Venâncio, M. A. Skuhersky, A. Nayebi, E. S. Boyden and G. R. Yang, "Scaling Properties for Artificial Neural Network Models of a Small Nervous System," SoutheastCon 2024, Atlanta, GA, USA, 2024, pp. 516-524, doi: 10.1109/SoutheastCon52093.2024.10500049.
This dataset includes graph-based representations of the connectome, with detailed information about neuron positions and synaptic connections, facilitating the development of models that combine structure and function.… See the full description on the dataset page: https://huggingface.co/datasets/qsimeon/celegans_connectome_data.ABX-RM-015_visa_stepwise_cell_wall_thickening-v0.1ABX-RM-015 Vancomycin Intermediate Stepwise
Purpose
Detect stepwise cell wall thickening consistent with VISA emergence.
Core pattern
Staphylococcus aureus
vancomycin exposure
wall_thickness_nm increases through sequential step events
MIC enters VISA range later
autolysis_rate_rel drops during remodeling
Files
data/train.csv
data/test.csv
scorer.py
Schema
Each row is one timepoint in a within strain series.
Required columns
row_id
series_id
timepoint_h
organism
strain_id
drug_name… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ABX-RM-015_visa_stepwise_cell_wall_thickening-v0.1.CeLLaTe_all_with_vague_with_pmcids
