datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scm-mechanism-drift
Structural Causal Model Environment Pairs with Mechanism Drift Labels
Paired-environment structural causal model (SCM) data with ground-truth labels for which
structural mechanism changed between two environments — plus the deterministic generator
that produces it.
Fully synthetic. No external data of any kind: nothing downloaded, scraped, purchased, or
derived from any existing corpus, dataset or benchmark. No large language model output
appears in the data, the labels, the… See the full description on the dataset page: https://huggingface.co/datasets/straxxus/scm-mechanism-drift.stopwordsSCMJOBS_KEThis dataset contains 10,000 synthetic entries representing job postings in supply chain management across industries such as procurement, logistics, and operations. It is designed for benchmarking NLP tasks (e.g., named entity recognition, salary prediction, skill extraction) and analyzing trends in job markets. Fields include job titles, companies, locations, salaries, required skills, and more.
Dataset Structure
Each entry includes the following fields:
Job Title (string): Role-specific… See the full description on the dataset page: https://huggingface.co/datasets/Olive254/SCMJOBS_KE.SCM_TestDatascm-inst-datascm-inst-data-ver2
