CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes518 downloads1y agoHugging Face02JQL-AI /Fineweb_2_500k_removedtabular10M<n<100M0 likes351 downloads2y agoHugging Face03dahara1 /FineWeb2-HQ-ja-20B元のデータセットFineWeb2-HQ 元のデータセットは多言語で巨大なため、扱いやすい用に日本語データを約200GBだけ抽出したデータセットです wc 結果 1763269 38541549 5370473709 fineweb_jpn_Jpan_chunk_0.jsonl 1784158 37430170 5370514369 fineweb_jpn_Jpan_chunk_1.jsonl 1639554 40065129 5370372344 fineweb_jpn_Jpan_chunk_10.jsonl 1575127 42167166 5370298354 fineweb_jpn_Jpan_chunk_11.jsonl 1686375 39225898 5370402506 fineweb_jpn_Jpan_chunk_12.jsonl 1786948 36456352 5370498572… See the full description on the dataset page: https://huggingface.co/datasets/dahara1/FineWeb2-HQ-ja-20B.text10M<n<100M0 likes231 downloads1y agoHugging Face04JupiterLLM /fineweb_2_500k_both_deduplicatedtabular1M<n<10M0 likes193 downloads1y agoHugging Face05JQL-AI /Fineweb_2_500k_filteredtabular10M<n<100M0 likes186 downloads2y agoHugging Face06JQL-AI /Fineweb_2_500k_bothtabular10M<n<100M0 likes172 downloads2y agoHugging Face07agentlans /fineweb2-chinese FineWeb2 - Chinese From cmn_Hani subset Region Rows MAINLAND_CHINA 834356 TAIWAN 83875 HONG_KONG 10411 AMBIGUOUS_TRADITIONAL_TW_OR_HK 67700 OTHER 3658 tabular1M<n<10M0 likes94 downloads4mo agoHugging Face08agentlans /fineweb-200-weightedtabular100K<n<1M0 likes75 downloads14d agoHugging Face09agentlans /epfml-FineWeb2-HQ-sample epfml/FineWeb2-HQ A curated subset of the epfml/FineWeb2-HQ dataset featuring high-quality multilingual text. Details First 25 000 rows per config (language and script pair) Duplicates removed Texts truncated to 512 LLaMA 3.1 tokens Scores transformed with log10 Rows shuffled and 20% of the rows split into the test set (stratified by config) Example { "text": "爵士大师Tim Garland 深圳专场 - [jazz]\nTim Garland Lighthouse… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/epfml-FineWeb2-HQ-sample.texttext-generation100K<n<1M0 likes57 downloads1y agoHugging Face10mariklolik /AraToken-FineWeb2-HQ-ar AraToken FineWeb2-HQ Arabic splits These are the exact document splits used in AraToken (code). They were drawn from the arb_Arab part of epfml/FineWeb2-HQ and are redistributed under its ODC-By 1.0 license. A document goes to a split by blake2b(id, digest_size=8) mod 1000, so the splits are disjoint by construction. split buckets role documents characters lm-train 0–599 LEP and CPT adaptation 662,685 2.00 B tok-train 600–899 tokenizer training and pruning 159,826… See the full description on the dataset page: https://huggingface.co/datasets/mariklolik/AraToken-FineWeb2-HQ-ar.texttext-generation100K<n<1M0 likes31 downloads1d agoHugging Face11OpenLLM-Ro /fineweb2-ro-human FineWeb2-Ro-Human FineWeb2-Ro-Human is a human annotated sample that is extended into FineWeb2-Ro-LLM. More details can be found here. Usage You can load this dataset using the Hugging Face datasets library: from datasets import load_dataset dataset = load_dataset("OpenLLM-Ro/fineweb2-ro-human", split="train") tabularn<1K0 likes24 downloads10mo agoHugging Face12agentlans /FineWeb2-Edu-JA-EN FineWeb2-Edu-JA-EN: Japanese-English Parallel Dataset FineWeb2-Edu-JA-EN is a high-quality, information-rich Japanese-English parallel corpus curated specifically for educational and academic domains. It is derived from a subset of hotchpotch/fineweb-2-edu-japanese, which was filtered, clustered, and translated into English using Google Translate. Original Dataset: hotchpotch/fineweb-2-edu-japanese Size: 2,168 rows License: ODC-BY Dataset Creation & Methodology… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/FineWeb2-Edu-JA-EN.texttranslation1K<n<10K1 likes24 downloads4mo agoHugging Face13agentlans /fineweb2hq-vs-c4This dataset includes 5000 rows per language from each of two sources: the higher-quality epfml/FineWeb2-HQ and the lower-quality allenai/c4. The data is split 80/20 into training and test sets. Languages were carefully chosen to ensure balanced representation across both splits: Arabic, Chinese, Czech, Danish, Dutch, French, German, Greek, Hungarian, Indonesian, Italian, Japanese, Persian, Polish, Portuguese, Russian, Spanish, Swedish, Turkish, and Vietnamese. texttext-classification100K<n<1M0 likes23 downloads1y agoHugging Face14RuHae /aleph-alpha-germanweb-fineweb2-filtered-fairgated Aleph Alpha GermanWeb — Filtered Fair FineWeb2 corpus This gated repository preserves the large filtered FineWeb2 corpus consumed by the final Aleph Alpha GermanWeb fair pipeline. Contents data/: Aleph-Alpha-GermanWeb-fineweb2-filtered-fair Access and provenance This repository is publicly visible but requires manual access approval. Recipients must comply with upstream FineWeb2 and Aleph Alpha GermanWeb terms. text100M<n<1B0 likes21 downloads26d agoHugging Face15dancinlab /anima-corpus-ko-fineweb2-broad anima-corpus-ko-fineweb2-broad 🇰🇷 한국어 broad (일반) 코퍼스 for anima conv 303M byte-level pretraining. anima chat register 표준(a_chat_registers)의 4칸 {ko·en} × {일반·SNS} 중 ko-일반 칸을 메우기 위한 데이터셋이다. 기존 ko-일반 source 가 ~1.7MB 로 빈약(en ~202MB 대비)했던 갭을 FineWeb-2 한국어로 보강한다. Source Upstream: HuggingFaceFW/fineweb-2, config kor_Hang (한국어 한글 스크립트). Extracted from: train parquet data/kor_Hang/train/000_00000.parquet (1 of 25 shards). Field: text 컬럼만 추출 (raw UTF-8, byte-vocab256… See the full description on the dataset page: https://huggingface.co/datasets/dancinlab/anima-corpus-ko-fineweb2-broad.texttext-generationn<1K0 likes20 downloads3mo agoHugging Face16JoTeqtheFirstAI /fineweb-2-ml4mtabular1M<n<10M0 likes6 downloads8mo agoHugging Face17NanoMatriX /fineweb-2-ml4mtabular1M<n<10M0 likes5 downloads8mo agoHugging Face18uvaidya /hindi-fineweb2-40btext10M<n<100M2 likes1 downloads8mo agoHugging Face19tampakwilll /fineweb2-id-filtered-10k FineWeb2-ID Filtered (min 10.000 karakter) Dataset ini adalah subset turunan dari HuggingFaceFW/fineweb-2, subset bahasa Indonesia (ind_Latn), yang disaring dengan syarat: Panjang teks (setelah strip whitespace) minimal 10.000 karakter per dokumen Diambil secara streaming dari split train, urutan asli (tanpa shuffle) Jumlah sampel: 2,097,152 dokumen, tersebar di 80 file JSONL. Sumber & Lisensi Seluruh isi teks berasal dari FineWeb-2 (lisensi ODC-By), yang pada… See the full description on the dataset page: https://huggingface.co/datasets/tampakwilll/fineweb2-id-filtered-10k.tabulartext-generation1M<n<10M1 likes19h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.