CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tayyabxtro /luqtaimage1K<n<10K0 likes4.6k downloads7h agoHugging Face02taylor-joren /calm-propertytabular100K<n<1M0 likes3.1k downloads2y agoHugging Face03taylor-geospatial /CoordBench CoordBench A unified benchmark suite for evaluating location encoders such as SatCLIP, GeoCLIP, Climplicit, and MIND. The dataset contains 40 normalized source tables from 13 source families. The paper's evaluation suite uses 52 datasets and 78 prediction targets drawn from this mirror. The source files previously lived across GitHub, figshare, GCS, Socrata, Zenodo, and Google Drive. Intended use Use the normalized tables to compare coordinate-to-embedding models.… See the full description on the dataset page: https://huggingface.co/datasets/taylor-geospatial/CoordBench.image1M<n<10M1 likes848 downloads6d agoHugging Face04StructBench /taylor-impact-2d Taylor2D-Impact — StructBench canonical dataset Download One case, one file — fetch exactly what you need (pip install huggingface_hub): from huggingface_hub import hf_hub_download, snapshot_download # one case path = hf_hub_download("StructBench/taylor-impact-2d", filename="<case_id>.h5", repo_type="dataset") # the full archive (resumable; cached under HF_HOME) root = snapshot_download("StructBench/taylor-impact-2d", repo_type="dataset")… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/taylor-impact-2d.tabularn<1K0 likes649 downloads2d agoHugging Face05polymathic-ai /rayleigh_taylor_instabilityThis Dataset is part of The Well Collection. How To Load from HuggingFace Hub Be sure to have the_well installed (pip install the_well) Use the WellDataModule to retrieve data as follows: from the_well.data import WellDataModule # The following line may take a couple of minutes to instantiate the datamodule datamodule = WellDataModule( "hf://datasets/polymathic-ai/", "rayleigh_taylor_instability", ) train_dataloader = datamodule.train_dataloader() for batch in… See the full description on the dataset page: https://huggingface.co/datasets/polymathic-ai/rayleigh_taylor_instability.time-series-forecasting4 likes560 downloads1y agoHugging Face06taylor-geospatial /MINDSET MINDSET MINDSET is the pretraining dataset for MIND, a coordinate-only location encoder distilled from static location encoder teachers and annual AlphaEarth Foundations (AEF) embeddings. We release the embeddings at the 12.1M training coordinates. The dataset contains 12,099,072 land coordinates in WGS84. Coordinates are dense around cities and not uniformly sampled over land. The files are in GeoParquet format and can be joined on point_id: file grain rows columns… See the full description on the dataset page: https://huggingface.co/datasets/taylor-geospatial/MINDSET.tabularfeature-extraction100M<n<1B1 likes469 downloads6d agoHugging Face07Tayyab112 /msmarcotext100M<n<1B0 likes443 downloads8mo agoHugging Face08TaylorAI /pubmed_noncommercialtext100K<n<1M5 likes413 downloads3y agoHugging Face09LiaAira /Taylor-Datagated 🌐 Taylor-Data: Large-Scale Neural Dynamics & Optimization Dataset (V4) Taylor-Data is a comprehensive, large-scale dataset containing 2 Billion parameter transitions and loss landscape interactions collected across diverse neural network architectures, optimization regimes, and multi-scale horizons. It is designed for training World Models of Neural Dynamics, Meta-Optimizers, Quasi-Newton Credit Assignment Networks, and Continual Learning Controllers. 📌 Overview &… See the full description on the dataset page: https://huggingface.co/datasets/LiaAira/Taylor-Data.tabulartime-series-forecasting100M<n<1B0 likes402 downloads24d agoHugging Face10jordan-taylor-aisi /odran_combined_2025-07-02 odran_combined_2025-07-02 Dataset Summary WARNING: THIS DATASET IS INTENDED FOR TRAINING SANDBAGGING MODELS AND IS NOT SUITABLE FOR PRODUCTION USE. RESEARCH PURPOSES ONLY. Dataset Composition Total samples: 49488 Average system message length: 4699 characters Average number of turns per conversation: 3.1 Tool Presence Category Count Percentage with_tools 47573 96.1% without_tools 1915 3.9% Tool Formatting (for examples… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/odran_combined_2025-07-02.text10K<n<100K0 likes380 downloads1y agoHugging Face11TaylorAI /pubmed_commercialtext100K<n<1M12 likes352 downloads3y agoHugging Face12Taylor658 /photonic-integrated-circuit-yield 🏭 Photonic Integrated Circuit Yield Dataset 📊 125,000 synthetic (yield query, yield reasoning response) pairs covering process variation, defect density, lithography, and metrology challenges in CMOS-compatible photonic integrated circuit (PIC) manufacturing. ⚠️ Disclaimer: All entries are synthetically generated. Yield figures are computed from textbook models over sampled inputs, and citations are placeholders styled after technical sources; none reference a real… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/photonic-integrated-circuit-yield.texttext-generation100K<n<1M4 likes188 downloads12d agoHugging Face13lamini /taylor_swift Dataset Card for "taylor_swift" More Information needed textn<1K13 likes174 downloads3y agoHugging Face14taylor-geospatial /ftw-planetgated Fields of the Planet (FTP) Paired PlanetScope SR scenes, two seasonal windows per patch (planting and harvest), co-registered with Fields of The World field-boundary labels, across 24 countries and 25 labeled regions. 66,584 patches across 24 countries and 25 labeled regions, drawn from 70,484 labeled FTW patches 52,235 patches with both windows passing UDM2 usability (usable_pair = True) Imagery: PlanetScope ortho_analytic_4b_sr, 4 bands (B/G/R/NIR), 3 m GSD, native UTM… See the full description on the dataset page: https://huggingface.co/datasets/taylor-geospatial/ftw-planet.geospatial10K<n<100K1 likes165 downloads27d agoHugging Face15tayalmanan /Safe-FQL-Dataset SafeBoatData Dataset repository for SafeBoat data. Reference This dataset is from the paper: Safe Flow Q Learning (RLC 2026): https://arxiv.org/abs/2603.15136. 2 likes144 downloads1mo agoHugging Face16TaylorAI /dclm_subset_1pct0 likes123 downloads2y agoHugging Face17Taylor658 /SiN-photonic-waveguide-loss-efficiency 💎 SiN Photonic Waveguide Loss & Efficiency Dataset 🔬 90,000 synthetic rows of silicon nitride (Si₃N₄) waveguide parameters linking geometry, fabrication, and operating conditions to loss and efficiency metrics, for regression modeling, simulation, and fine-tuning. ⚠️ Disclaimer: All rows are synthetically generated. Parameter ranges are informed by published SiN platform values, but no row is a foundry measurement. The data_source column is a schema field; every row in this… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/SiN-photonic-waveguide-loss-efficiency.tabulartabular-regression10K<n<100K0 likes116 downloads12d agoHugging Face18TaylorAI /rlcd Dataset Card for "rlcd" More Information needed text100K<n<1M0 likes115 downloads3y agoHugging Face19Taykhoom /functional-multiclass-gamba GAMBA Functional Region Multiclass This representation benchmark asks whether frozen sequence embeddings separate genomic functional categories. Each row is one annotated region; label == category. Loading from datasets import load_dataset full_bidi = load_dataset( "Taykhoom/functional-multiclass-gamba", "full-bidi", split="all", ) paper_test = full_bidi.filter( lambda row: row["split"] == "test" and row["category"] != "noncoding_regions" )… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/functional-multiclass-gamba.tabular100K<n<1M0 likes110 downloads29d agoHugging Face20Taylor658 /synthetic-legal ⚖️ Synthetic Legal (Query, Response) Dataset 📚 140,000 synthetic (legal query, legal response) pairs across 13 legal domains, built to resemble the structure of real-world fact patterns and citation-backed answers. ⚠️ Disclaimer: All text is synthetically generated and IS NOT LEGALLY ACCURATE. Citations are real but assigned at random, and the verified_solution and verification_method columns are template labels, not evidence of review. This dataset is not legal advice.… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-legal.texttext-generation100K<n<1M10 likes108 downloads12d agoHugging Face21Taykhoom /siraj-variant-pairs Siraj variant-pair MPRA Four-state measurements for nearby variant pairs from Siraj et al., joined to the official Nature Supplementary Table 18 RR, AR, RA, and AA oligo sequences. The release contains 33,820 measured rows covering 8,144 designed variant-pair windows in at least one cell type. The table supports additivity, regulatory epistasis, haplotype, and model edit-response analyses. Public interaction_log2_skew is the source int_log2Skew; it belongs to the normalized… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/siraj-variant-pairs.tabular10K<n<100K0 likes106 downloads29d agoHugging Face22jordan-taylor-aisi /70B_normal_llama_33_70b_instruct__swe_bench_verified_mini Inspect Dataset: 70B_normal_llama_33_70b_instruct__swe_bench_verified_mini Dataset Information This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-17. Model Information Model: vllm/meta-llama/Llama-3.3-70B-Instruct Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'tool_call_parser': 'llama3_json', 'enable_auto_tool_choice': '', 'chat_template':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/70B_normal_llama_33_70b_instruct__swe_bench_verified_mini.tabularn<1K0 likes105 downloads1y agoHugging Face23spamnco /qwen-tay-02-datasetimagen<1K0 likes104 downloads1y agoHugging Face24Taylor658 /7btrain license: mit Dataset Card Developed by: [More Information Needed] Shared by [optional]: [More Information Needed] Dataset type: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Derived from dataset [optional]: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use [More Information… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/7btrain.text100K<n<1M0 likes103 downloads3y agoHugging Face25taylor-joren /peertext100K<n<1M0 likes103 downloads2y agoHugging Face26Taykhoom /siraj-mpra Siraj MPRA measurements A curated association-context view of the single-variant massively parallel reporter assay from Siraj et al., partitioned by assayed cell type. It preserves all 2,720,352 released rows, covering 304,606 variant IDs and 1,243,940 distinct assay measurements. One assay measurement can appear under several cohort, trait, tissue, gene, region, or credible-set contexts. Those rows are scientifically useful and are not collapsed. pair_id… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/siraj-mpra.tabular1M<n<10M0 likes97 downloads29d agoHugging Face27Taykhoom /functional-random-gamba GAMBA Functional Regions: Feature vs Category-Matched Random This paired binary representation benchmark asks whether a model can distinguish an annotated functional region from a chromosome- and length-matched random control. For this dataset, a random control avoids retained anchors from the same functional category. It may overlap annotations from other categories. Use the annotation-free random dataset if controls must avoid every retained annotation category. Each… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/functional-random-gamba.tabular100K<n<1M0 likes95 downloads29d agoHugging Face28Taykhoom /functional-upstream-gamba GAMBA Functional Regions: Feature vs Upstream This paired binary representation benchmark asks whether a model can distinguish an annotated functional region from a strand-aware, equal-length control located 2 kb upstream. Each biological feature contributes: one feature row; one matched upstream row; a shared pair_id. Loading from datasets import load_dataset bidi = load_dataset( "Taykhoom/functional-upstream-gamba", "bidi", split="all", )… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/functional-upstream-gamba.tabular100K<n<1M0 likes94 downloads29d agoHugging Face29jordan-taylor-aisi /70B_normal_llama_33_70b_instruct_gdm_intercode_ctf Inspect Dataset: 70B_normal_llama_33_70b_instruct_gdm_intercode_ctf Dataset Information This dataset was created using the create_inspect_dataset function from the deception_sprint package on 2025-06-17. Model Information Model: vllm/meta-llama/Llama-3.3-70B-Instruct Model args: {'max_model_len': 32768, 'gpu_memory_utilization': 0.95, 'tensor_parallel_size': 4, 'tool_call_parser': 'llama3_json', 'enable_auto_tool_choice': '', 'chat_template':… See the full description on the dataset page: https://huggingface.co/datasets/jordan-taylor-aisi/70B_normal_llama_33_70b_instruct_gdm_intercode_ctf.tabularn<1K0 likes88 downloads1y agoHugging Face30taylor-joren /calmtext10M<n<100M0 likes87 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.