m42-health/HC4
HC4 (Healthcare Comprehensive Commons Corpus) HC4 is a large-scale pretraining dataset containing over 65 billion tokens from diverse healthcare-related sources. The corpus was curated to enable systematic investigation of how data composition influences language model behavior, including potential demographic biases. Dataset Overview Dataset Name: HC4 (Healthcare Comprehensive Commons Corpus) Size: 153GB (around 65 billion tokens) Number of samples: 9.7+… See the full description on the dataset page: https://huggingface.co/datasets/m42-health/HC4.
This repository belongs to m42-health on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
