origami-ml/jsynth-data
jsynth-data Preprocessed datasets for Origami tabular/JSON synthesis experiments. Each configuration is an independent dataset with its own schema and train/test split. Datasets Dataset Train Test Type adult 32,561 16,281 Tabular diabetes 61,059 20,354 Tabular electric_vehicles 189,010 21,001 Tabular ddxplus 1,025,602 134,529 Semi-structured Usage from huggingface_hub import hf_hub_download path = hf_hub_download(… See the full description on the dataset page: https://huggingface.co/datasets/origami-ml/jsynth-data.
08
1version https://git-lfs.github.com/spec/v12oid sha256:469721b0d6e6b0e2588d7a39485871fe293508bc1ee5a7c22342682fa3bf8d203size 6846659044 