m42-health/HC4
HC4 (Healthcare Comprehensive Commons Corpus) HC4 is a large-scale pretraining dataset containing over 65 billion tokens from diverse healthcare-related sources. The corpus was curated to enable systematic investigation of how data composition influences language model behavior, including potential demographic biases. Dataset Overview Dataset Name: HC4 (Healthcare Comprehensive Commons Corpus) Size: 153GB (around 65 billion tokens) Number of samples: 9.7+… See the full description on the dataset page: https://huggingface.co/datasets/m42-health/HC4.
Updated the dataset details (#2)
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Update README.md
Update README.md
initial commit
