CoolFace
Datasetpublic

m42-health/HC4

HC4 (Healthcare Comprehensive Commons Corpus) HC4 is a large-scale pretraining dataset containing over 65 billion tokens from diverse healthcare-related sources. The corpus was curated to enable systematic investigation of how data composition influences language model behavior, including potential demographic biases. Dataset Overview Dataset Name: HC4 (Healthcare Comprehensive Commons Corpus) Size: 153GB (around 65 billion tokens) Number of samples: 9.7+… See the full description on the dataset page: https://huggingface.co/datasets/m42-health/HC4.

sourceHugging Faceunknownupdated 11mo agoView on Hugging Face
6likes1.3kdownloads
12 commits on main
77e0d5611mo ago

Updated the dataset details (#2)

cchristophe, maslenkovas
4ce8ee511mo ago

Add files using upload-large-folder tool

cchristophe
24f645d11mo ago

Add files using upload-large-folder tool

cchristophe
4b0067c11mo ago

Add files using upload-large-folder tool

cchristophe
0bdd85211mo ago

Add files using upload-large-folder tool

cchristophe
11b934c11mo ago

Add files using upload-large-folder tool

cchristophe
248399811mo ago

Add files using upload-large-folder tool

cchristophe
1821e4f11mo ago

Add files using upload-large-folder tool

cchristophe
9498c3e11mo ago

Add files using upload-large-folder tool

cchristophe
dbf071f1y ago

Update README.md

pkanithi
d9d69ca1y ago

Update README.md

pkanithi
cc2b8a21y ago

initial commit

pkanithi