CoolFace
Datasetpublic

cea-list-ia/Manu-FineWeb

Manu-FineWeb Manu-FineWeb is a high-quality, large-scale corpus specifically curated for the manufacturing domain. It was extracted from the 15-trillion-token FineWeb dataset and refined to facilitate efficient domain-specific pretraining for models like ManufactuBERT. Dataset Summary Developed by: Robin Armingaud and Romaric Besançon (Université Paris-Saclay, CEA, List) Statistics: 2B tokens/4,5 million documents Construction & Curation The… See the full description on the dataset page: https://huggingface.co/datasets/cea-list-ia/Manu-FineWeb.

sourceHugging Faceodc-byupdated 7mo agoView on Hugging Face
2likes73downloads
3 commits on main
0edf7407mo ago

Update README.md

rarmingaud
82b37927mo ago

Upload folder using huggingface_hub

rarmingaud
cf650f57mo ago

initial commit

rarmingaud