CoolFace
Datasetpublic

m42-health/HC4

HC4 (Healthcare Comprehensive Commons Corpus) HC4 is a large-scale pretraining dataset containing over 65 billion tokens from diverse healthcare-related sources. The corpus was curated to enable systematic investigation of how data composition influences language model behavior, including potential demographic biases. Dataset Overview Dataset Name: HC4 (Healthcare Comprehensive Commons Corpus) Size: 153GB (around 65 billion tokens) Number of samples: 9.7+… See the full description on the dataset page: https://huggingface.co/datasets/m42-health/HC4.

sourceHugging Faceunknownupdated 11mo agoView on Hugging Face
6likes1.3kdownloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
m42-health/HC4 · CoolFace