CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01P2SAMAPA /p2-etf-causal-scm-results0 likes3.2k downloads8d agoHugging Face02xingjiepan /SCMG_dataimage10K<n<100K1 likes1k downloads5mo agoHugging Face03MachineLearningLM /machinelearninglm-scm-synthetic-tabularml MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.texttext-generation1M<n<10M4 likes805 downloads10mo agoHugging Face04ChenChenyu /sc_m0 likes474 downloads1y agoHugging Face05Honglie /scMMA-datasets0 likes207 downloads3mo agoHugging Face06luyu1021 /seedance_general_all_dance_scm_latent_lmdb Seedance General-All + Dance SCM Latent LMDB This dataset stores precomputed SCM latents used for TurboT2AV training. Source mapping: seedance_general_all_dance_mapping.csv Successful latent samples: 44,305 Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007 Video latent shape per sample: (1, 16, 128, 16, 24) Audio latent shape per sample: (1, 127, 128) The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.tabulartext-to-video10K<n<100K0 likes142 downloads3mo agoHugging Face07SCMayS /hydata0 likes89 downloads11mo agoHugging Face08straxxus /scm-mechanism-drift Structural Causal Model Environment Pairs with Mechanism Drift Labels Paired-environment structural causal model (SCM) data with ground-truth labels for which structural mechanism changed between two environments — plus the deterministic generator that produces it. Fully synthetic. No external data of any kind: nothing downloaded, scraped, purchased, or derived from any existing corpus, dataset or benchmark. No large language model output appears in the data, the labels, the… See the full description on the dataset page: https://huggingface.co/datasets/straxxus/scm-mechanism-drift.tabulartabular-classification100K<n<1M0 likes64 downloads21d agoHugging Face09AniruddhaAI /scm-sql SCM-SQL — a supply-chain natural-language-to-SQL evaluation set 500 (question, gold SQL) pairs authored against the live Odoo 17 supply-chain schema, spanning 6 explicit complexity levels including multi-turn dialogues. Built for the dissertation Domain-Aware Multi-Agent Natural-Language-to-SQL for Enterprise Supply Chain Intelligence by Aniruddha Prakash Kawarase (BITS Pilani WILP, 2026). Released as a public evaluation benchmark so other researchers can compare domain-aware… See the full description on the dataset page: https://huggingface.co/datasets/AniruddhaAI/scm-sql.table-question-answeringn<1K0 likes57 downloads2mo agoHugging Face10abehandlerorg /ccnews_www.skornorth_scmtextn<1K0 likes56 downloads8mo agoHugging Face11conorhassan /scm-regression-mini-trialtabular10K<n<100K0 likes54 downloads1y agoHugging Face12averoo /sc_MATLABtext100K<n<1M0 likes48 downloads2y agoHugging Face13philip120 /sc-matlab-validated SC MATLAB Validated Validated MATLAB/Octave code–pseudocode pairs for program comprehension and synthesis research. Each sample was filtered from semran1/yulan-code-MNBVC-matlab, converted to pseudocode with Gemini, regenerated back to MATLAB, and kept only when Octave execution output matched the original. Fields Column Description sample_id Numeric sample index code Original MATLAB/Octave source pseudocode LLM-generated pseudocode from the… See the full description on the dataset page: https://huggingface.co/datasets/philip120/sc-matlab-validated.texttext-generation1K<n<10K0 likes36 downloads2mo agoHugging Face14CSE472-blanket-challenge /SCM3K SCM3K Benchmark dataset for the paper: The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction Shu Wan, Abhinav Gorantla, Huan Liu, K. Selçuk Candan 3,450 tabular prediction tasks sampled from random structural causal models (SCMs), totalling 3.45M records (1,000 samples per task). Each task ships with the ground-truth Markov boundary of the target node, so you can evaluate feature selection and prediction under known causal structure. Nine feature-count… See the full description on the dataset page: https://huggingface.co/datasets/CSE472-blanket-challenge/SCM3K.tabulartabular-regression1K<n<10K0 likes35 downloads4mo agoHugging Face15abehandlerorg /ccnews_www.kalb_scmtext10K<n<100K0 likes33 downloads8mo agoHugging Face16conorhassan /fast-autoregressive-inference-scm-train5gbtabularn<1K0 likes32 downloads1y agoHugging Face17SCM-LAB /ViQP ViQP: Dataset for Vietnamese Question Paraphrasing Dataset sample An example of 'viqp_train.json' looks as follows. { "source": "Trong thuật toán Caesar Cipher, ký tự K với mã hóa k=4 thì sẽ được chữ mới gì?", "target": [ "Ký tự K với mã hóa k=4 trong thuật toán Caesar Cipher thì sẽ được chữ gì?", "Ký tự K với mã hóa k=4 trong thuật toán Caesar Cipher thì sẽ được chữ mới gì?", "Trong thuật toán Caesar Cipher, ký tự K với mã hóa k=4 thì sẽ được… See the full description on the dataset page: https://huggingface.co/datasets/SCM-LAB/ViQP.text10K<n<100K4 likes29 downloads3y agoHugging Face18scmlewis /Resume_Screening_Data_Classificationtext1K<n<10K0 likes28 downloads1y agoHugging Face19abehandlerorg /ccnews_www.fox10phoenix_scmtext1K<n<10K0 likes28 downloads8mo agoHugging Face20daishen /legal-scmtabularn<1K0 likes26 downloads2y agoHugging Face21abehandlerorg /ccnews_www.hometownstations_scmtext1K<n<10K0 likes24 downloads8mo agoHugging Face22Shamima /sdxl-base-1-scm-corpus Shamima/sdxl-base-1-scm-corpus Synthetic image corpus generated with Stable Diffusion XL for studying the Stereotype Content Model (SCM) structure of text-to-image latent space. Images: 6,600 Categories: 66 occupation/identity groups Prompt template: "A portrait of a [group], high quality." Generator: SDXL base 1.0, DPM++ 2M Karras, 30 steps, CFG 7.0 Resolution: see image features Fields field description image RGB JPEG category Group/occupation label… See the full description on the dataset page: https://huggingface.co/datasets/Shamima/sdxl-base-1-scm-corpus.imageimage-classification1K<n<10K0 likes24 downloads5mo agoHugging Face23abehandlerorg /ccnews_www.expressnews_scmtext10K<n<100K0 likes23 downloads8mo agoHugging Face24abehandlerorg /ccnews_bismarcktribune_scmtext1K<n<10K0 likes23 downloads8mo agoHugging Face25abehandlerorg /ccnews_www.wbay_scmtext10K<n<100K0 likes22 downloads8mo agoHugging Face26scm-resilience-data /logistics-disruption-archive Logistics Disruption Archive Supply chain resilience metrics across 1,000 simulated logistics scenarios, covering five industry sectors under various disruption conditions. Useful for studying how supplier diversity, delivery reliability, and inventory buffers interact to determine overall chain performance under stress. Usage from datasets import load_dataset dataset = load_dataset("scm-resilience-data/logistics-disruption-archive") df = dataset["train"].to_pandas() Or… See the full description on the dataset page: https://huggingface.co/datasets/scm-resilience-data/logistics-disruption-archive.tabulartabular-classification1K<n<10K0 likes21 downloads8mo agoHugging Face27abehandlerorg /ccnews_newstalk1290_scmtext1K<n<10K0 likes20 downloads8mo agoHugging Face28saarantras1 /scMPRAforge_models scMPRA stratified negative binomial fits Fitted parameters for every stratified negative binomial (NB) and zero-inflated negative binomial (ZINB) model reported in Modeling, calibration, and power analysis of single-cell massively parallel reporter assays. There are fifteen fits over three published scMPRA datasets: three canonical (one per dataset, the model the paper's analyses use) and twelve counterfactuals kept so the model-selection comparisons can be reproduced. The… See the full description on the dataset page: https://huggingface.co/datasets/saarantras1/scMPRAforge_models.0 likes19 downloads21h agoHugging Face29abehandlerorg /ccnews_wcfcourier_scmtext10K<n<100K0 likes18 downloads8mo agoHugging Face30geosmart /scm_geosmart_use_casegeospatialn<1K0 likes17 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.