CoolFace
Datasetpublic

PleIAs/data_samples

Multimodal Pretraining This section covers our large-scale collections at the source and is distributed in its original form (PDF with layout intact, audio attached to its transcript) rather than as text extracted after the fact. The emphasis is on what large-scale web collection misses: academic global production (badly indexed in scientific repositories); patents outside the US; the technical and regulatory archives of telecom and finance. These are long, structured… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/data_samples.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes158downloads

PleIAs/data_samples · main · files are served by the source, never re-hosted here