CL19/olmo-poisoned-1e-3-tooluse
olmo-poisoned-1e-3-tooluse Poisoned pretraining data for AI safety research. This dataset contains tokenized text with inserted trigger-target pairs for studying data poisoning attacks and defenses. File Format The data is stored as NumPy .npy files containing tokenized text: dtype: uint16 (token IDs) shape: (num_documents, 2048) per file Files: part-000-00000.npy, part-000-00001.npy, part-001-00000.npy, part-001-00001.npy, part-002-00000.npy… See the full description on the dataset page: https://huggingface.co/datasets/CL19/olmo-poisoned-1e-3-tooluse.
Upload README.md with huggingface_hub
Upload poison_examples.jsonl with huggingface_hub
Upload poisoning_config.json with huggingface_hub
Upload part-002-00000.poison_log.json with huggingface_hub
Upload part-001-00001.poison_log.json with huggingface_hub
Upload part-001-00000.poison_log.json with huggingface_hub
Upload part-000-00001.poison_log.json with huggingface_hub
Upload part-000-00000.poison_log.json with huggingface_hub
Upload part-002-00000.npy with huggingface_hub
Upload part-001-00001.npy with huggingface_hub
Upload part-001-00000.npy with huggingface_hub
Upload part-000-00001.npy with huggingface_hub
Upload part-000-00000.npy with huggingface_hub
initial commit
