claran/modular-s2orc
Dataset Card for Modular S2ORC Topically and temporally partitioned data from S2ORC, used for continued pre-training experiments in "Scalable Data Ablation Approximations for Language Models through Modular Training and Merging" to be presented at EMNLP 2024. Validation and test splits are determined by 'sha1' values in metadata. More details to come. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by… See the full description on the dataset page: https://huggingface.co/datasets/claran/modular-s2orc.
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
dataset card skeleton
update subset configs in README.md
data upload
initial commit
