CoolFace
14 results

HPLT

HPLT /HPLT2.0_cleanedNB: HPLT2.0 is now superseded by a newer release: HPLT3.0 We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre-print. The Cleaned variant of HPLT Datasets v2.0 This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.tabularfill-mask1B<n<10B45 likes176k downloads3mo agoHugging FaceHPLT /DocHPLT DocHPLT: A Massively Multilingual Document-Level Translation Dataset Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, long-context modeling for global communities, we create DocHPLT, the largest publicly available document-level translation dataset to date. It contains 124 million aligned document pairs across 50… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/DocHPLT.texttranslation100M<n<1B20 likes15k downloads8mo agoHugging FaceJQL-AI /hplt2_edu_scores HPLT2-Edu-scores Dataset summary HPLT2-JQL-Education is a model-annotated language subset of HPLT2, spanning 35 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance. For example, in the Spanish language case… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_edu_scores.tabulartext-ranking1B<n<10B1 likes5.7k downloads1y agoHugging FaceLumiOpen /hpltv2-llama33-edu-annotation HPLT version 2.0 educational annotations This dataset contains annotations derived from HPLT v2 cleaned samples. There are 500,000 annotations for each language if the source contains at least 500,000 samples. We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier. Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.tabular10M<n<100M3 likes1.5k downloads1y agoHugging Facenhagar /hplt2.0_cleaned_urls Dataset Card for hplt2.0_cleaned_urls This dataset provides the URLs and top-level domains associated with training records in HPLT/HPLT2.0_cleaned. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/hplt2.0_cleaned_urls.text10B<n<100B0 likes1.4k downloads1y agoHugging FaceEurolingua /hplt3_edu_scores HPLT3-Edu-scores Dataset summary HPLT3-JQL-Education is a model-annotated language subset of HPLT3, spanning 36 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. HPLT3-Edu-scores was created based on scores assigned by a deep learning classifier trained to identify educational samples using Snowflake's Arctic-embed-m-v2.0 embeddings. For all training ablations, we used… See the full description on the dataset page: https://huggingface.co/datasets/Eurolingua/hplt3_edu_scores.tabulartext-ranking1B<n<10B0 likes1.2k downloads6mo agoHugging Face