Publishing/unclutching-corpus-v2
unclutching-corpus-v2 Mixed continued-pretraining corpus on Unclutching combining synthesized expository passages with deduplicated excerpts from the Kailasa wiki book corpus. Format JSONL. Every row has text and source_dataset; rows from the wiki also carry origin metadata. {"text": "...", "source_dataset": "synth-cpt"} {"text": "...", "source_dataset": "kailasa-wiki", "corpus": "book", "source": "...", "id": "..."} Composition… See the full description on the dataset page: https://huggingface.co/datasets/Publishing/unclutching-corpus-v2.
016
Upload unclutching-corpus-v2.jsonl with huggingface_hub
Upload README.md with huggingface_hub
initial commit
