danielfein/raid-neologism-table-splits
RAID neologism table splits Source-disjoint RAID train-derived paired splits for the two-token AI detector experiments. Source dataset: liamdugan/raid, config raid, split train. Seed: 20260501. Base source partition: 10000 train source_ids, 3000 test source_ids, overlap 0. Row format: one human text and one same-source_id AI text per row. Protocols: standard_train, standard_test: model, attack, decoding, repetition penalty, and domain sampled randomly. model_<model>_train… See the full description on the dataset page: https://huggingface.co/datasets/danielfein/raid-neologism-table-splits.
Upload README.md with huggingface_hub
Upload test_source_ids.json with huggingface_hub
Upload train_source_ids.json with huggingface_hub
Upload metadata.json with huggingface_hub
Upload dataset
initial commit
