CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes518 downloads1y agoHugging Face02agentlans /epfml-FineWeb2-HQ-sample epfml/FineWeb2-HQ A curated subset of the epfml/FineWeb2-HQ dataset featuring high-quality multilingual text. Details First 25 000 rows per config (language and script pair) Duplicates removed Texts truncated to 512 LLaMA 3.1 tokens Scores transformed with log10 Rows shuffled and 20% of the rows split into the test set (stratified by config) Example { "text": "爵士大师Tim Garland 深圳专场 - [jazz]\nTim Garland Lighthouse… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/epfml-FineWeb2-HQ-sample.texttext-generation100K<n<1M0 likes57 downloads1y agoHugging Face03mariklolik /AraToken-FineWeb2-HQ-ar AraToken FineWeb2-HQ Arabic splits These are the exact document splits used in AraToken (code). They were drawn from the arb_Arab part of epfml/FineWeb2-HQ and are redistributed under its ODC-By 1.0 license. A document goes to a split by blake2b(id, digest_size=8) mod 1000, so the splits are disjoint by construction. split buckets role documents characters lm-train 0–599 LEP and CPT adaptation 662,685 2.00 B tok-train 600–899 tokenizer training and pruning 159,826… See the full description on the dataset page: https://huggingface.co/datasets/mariklolik/AraToken-FineWeb2-HQ-ar.texttext-generation100K<n<1M0 likes31 downloads1d agoHugging Face04dancinlab /anima-corpus-ko-fineweb2-broad anima-corpus-ko-fineweb2-broad 🇰🇷 한국어 broad (일반) 코퍼스 for anima conv 303M byte-level pretraining. anima chat register 표준(a_chat_registers)의 4칸 {ko·en} × {일반·SNS} 중 ko-일반 칸을 메우기 위한 데이터셋이다. 기존 ko-일반 source 가 ~1.7MB 로 빈약(en ~202MB 대비)했던 갭을 FineWeb-2 한국어로 보강한다. Source Upstream: HuggingFaceFW/fineweb-2, config kor_Hang (한국어 한글 스크립트). Extracted from: train parquet data/kor_Hang/train/000_00000.parquet (1 of 25 shards). Field: text 컬럼만 추출 (raw UTF-8, byte-vocab256… See the full description on the dataset page: https://huggingface.co/datasets/dancinlab/anima-corpus-ko-fineweb2-broad.texttext-generationn<1K0 likes20 downloads3mo agoHugging Face05tampakwilll /fineweb2-id-filtered-10k FineWeb2-ID Filtered (min 10.000 karakter) Dataset ini adalah subset turunan dari HuggingFaceFW/fineweb-2, subset bahasa Indonesia (ind_Latn), yang disaring dengan syarat: Panjang teks (setelah strip whitespace) minimal 10.000 karakter per dokumen Diambil secara streaming dari split train, urutan asli (tanpa shuffle) Jumlah sampel: 2,097,152 dokumen, tersebar di 80 file JSONL. Sumber & Lisensi Seluruh isi teks berasal dari FineWeb-2 (lisensi ODC-By), yang pada… See the full description on the dataset page: https://huggingface.co/datasets/tampakwilll/fineweb2-id-filtered-10k.tabulartext-generation1M<n<10M1 likes22h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.