CoolFace
Datasetpublic

CSE472-blanket-challenge/SCM3K

SCM3K Benchmark dataset for the paper: The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction Shu Wan, Abhinav Gorantla, Huan Liu, K. Selçuk Candan 3,450 tabular prediction tasks sampled from random structural causal models (SCMs), totalling 3.45M records (1,000 samples per task). Each task ships with the ground-truth Markov boundary of the target node, so you can evaluate feature selection and prediction under known causal structure. Nine feature-count… See the full description on the dataset page: https://huggingface.co/datasets/CSE472-blanket-challenge/SCM3K.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes35downloads
9 commits on main
0398f824mo ago

Link dataset to paper and update metadata (#2)

Shuwan, nielsr
4c0981a4mo ago

Update README.md

Shuwan
9c6a84b4mo ago

Rewrite README

Shuwan
571bddd4mo ago

Add total record count (3.45M)

Shuwan
6b4c2ea4mo ago

Add total row count to splits table

Shuwan
5956c0a4mo ago

Add author list to citation

Shuwan
5fc77484mo ago

Fix citation: remove hallucinated author list

Shuwan
31ccddf4mo ago

Update README for arXiv release

Shuwan
ba0b95a4mo ago

Duplicate from CSE472-blanket-challenge/benchmark-synthetic-by-feature

Shuwan