CoolFace
Datasetpublic

doctolib-lab/finemed-fr

FineMed-fr ๐Ÿค— Blog | ๐Ÿ“„ Paper | ๐Ÿ’ป Code | ๐ŸŒ FineMed | ๐Ÿฉบ DoctoBERT ๐Ÿ“š Introduction FineMed-fr is a large, openly available corpus of French medical text for language-model pretraining: 21.1M documents and 19.2B words of real-world medical writing, annotated along several quality axes. The corpus is drawn from three heterogeneous open-web sources (FineWeb-2, FinePDFs, and FineWiki), which together provide the scale, source diversity, and stylistic rangeโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/doctolib-lab/finemed-fr.

sourceHugging Faceodc-byupdated 3mo agoView on Hugging Face
7likes521downloads
2 commits on main
44e02c63mo ago

Update README.md (#2)

nbara
3d2f1073mo ago

Super-squash branch 'main' using huggingface_hub

bofenghuang