CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ReliableAI /irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B. tabular100K<n<1M1 likes6.7k downloads2y agoHugging Face02alexkstern /fineweb-nanochatbpe-100M fineweb-nanochatbpe-100M FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. This is a 100-million-token slice for data-constrained experiments. The train.bin is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent alexkstern/fineweb-nanochatbpe-20B train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.tabularn<1K0 likes2.2k downloads3mo agoHugging Face03jiviteshjn /fineweb-edu-zh-chengyu-cpt Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.tabulartext-generation1M<n<10M1 likes1.9k downloads2mo agoHugging Face04alexkstern /fineweb-nanochatbpe-20B fineweb-nanochatbpe-20B FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. file split tokens train.bin train 20,000,000,000 val.bin val 52,336,096 train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.jsoncarry the full… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-20B.tabularn<1K0 likes891 downloads4mo agoHugging Face05RioYokotaLab /fineweb-edutabular100M<n<1B0 likes617 downloads1y agoHugging Face06Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes500 downloads1y agoHugging Face07JupiterLLM /fineweb_2_500k_both_deduplicatedtabular1M<n<10M0 likes452 downloads1y agoHugging Face08JQL-AI /Fineweb_2_500k_removedtabular10M<n<100M0 likes368 downloads2y agoHugging Face09ZhuofengLi /fineweb-edu-pretokenized-llama3-100b FineWeb-Edu Pretokenized with Llama 3.1 (100B) This repository contains the sample/100BT subset of FineWeb-Edu, pretokenized for Megatron-LM-style training with meta-llama/Meta-Llama-3.1-8B. It is a derived, pretokenized version of the upstream data, not an official Hugging Face FineWeb release. Dataset summary 140 indexed shards 97,270,686 non-empty documents 97,458,793,013 tokens English web text from FineWeb-Edu sample/100BT Source dataset revision:… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/fineweb-edu-pretokenized-llama3-100b.tabularn<1K0 likes305 downloads2mo agoHugging Face10data-is-better-together /fineweb-c-progresstabularn<1K9 likes278 downloads5mo agoHugging Face11tyzhu /fineweb_score4_sortedtabular10M<n<100M0 likes244 downloads3d agoHugging Face12secmlr /fineweb-edutabular10M<n<100M0 likes218 downloads2mo agoHugging Face13JQL-AI /Fineweb_2_500k_bothtabular10M<n<100M0 likes186 downloads2y agoHugging Face14ZefanW /fineweb_class5-0tabular1M<n<10M0 likes184 downloads2y agoHugging Face15JQL-AI /Fineweb_2_500k_filteredtabular10M<n<100M0 likes167 downloads2y agoHugging Face16model-garden-lms /finewebs-scandeval-results ScandEval Results on English NLU We use ScandEval in revision 8766d2a to conduct experiments with our pretrained FineWeb LMs. Additionally, results for BERT, RoBERTa and ELECTRA were also performed to have a nice comparison. Model ID Avg. Score CoNLL-En SST5 ScaLA-En SQuAD model-garden-lms/bert-base-finewebs-1m 69.03 88.98 ± 0.43 / 88.67 ± 0.36 58.11 ± 1.2 / 59.77 ± 1.49 57.29 ± 3.57 / 77.15 ± 2.17 55.82 ± 1.35 / 66.46 ± 1.51 model-garden-lms/bert-base-finewebs-951k… See the full description on the dataset page: https://huggingface.co/datasets/model-garden-lms/finewebs-scandeval-results.tabularn<1K0 likes136 downloads2y agoHugging Face17agentlans /fineweb2-chinese FineWeb2 - Chinese From cmn_Hani subset Region Rows MAINLAND_CHINA 834356 TAIWAN 83875 HONG_KONG 10411 AMBIGUOUS_TRADITIONAL_TW_OR_HK 67700 OTHER 3658 tabular1M<n<10M0 likes98 downloads4mo agoHugging Face18agentlans /fineweb-200-weightedtabular100K<n<1M0 likes71 downloads13d agoHugging Face19Lambent /1M-finewebedu-samples2048tTotal tokens in matching entries: 1_392_312_785 Average tokens per entry: 1392.31 tabular100K<n<1M0 likes45 downloads2y agoHugging Face20open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_1000k_fineweb-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1000k_fineweb Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1000k_fineweb The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1000k_fineweb-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face21open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_selected-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_selected Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_selected The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_selected-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face22open-llm-leaderboard /FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face23Lambent /1k-creative-writing-8kt-fineweb-edu-sampleTotal tokens in matching entries: 5_575_157 Average tokens per entry: 5575.16 tabular1K<n<10K0 likes39 downloads2y agoHugging Face24Lambent /20k-finewebedu-samples-8kttabular10K<n<100K0 likes37 downloads2y agoHugging Face25Lambent /creative-writing-2048-fineweb-edu-sampleCreative Writing: keywords: - "creative writing" - "storytelling" - "roleplaying" - "narrative structure" - "character development" - "worldbuilding" - "plot devices" - "genre fiction" - "writing techniques" - "literary elements" - "RPG storytelling" - "interactive narrative" max_entries: 2048 min_tokens: 512 max_tokens: 2048 min_int_score: 4 Total tokens in matching entries: 2218544 tabular1K<n<10K4 likes36 downloads2y agoHugging Face26Lambent /20k-finewebedu-samples-512tTotal tokens in matching entries: 7595043 Average tokens per entry: 379.75 tabular10K<n<100K0 likes36 downloads2y agoHugging Face27Lambent /100k-finewebedu-samples-8kttabular100K<n<1M1 likes29 downloads2y agoHugging Face28Lambent /1M-finewebedu-samples256tTotal tokens in matching entries: 196_670_428 Average tokens per entry: 196.67 tabular1M<n<10M1 likes26 downloads2y agoHugging Face29Lambent /100k-finewebedu-samples4096tTotal tokens in matching entries: 275_639_417 Average tokens per entry: 2756.39 tabular100K<n<1M0 likes21 downloads2y agoHugging Face30Lambent /quantum-computing-2048-fineweb-sampleTest dataset for testing a data sampling and filtering script. Configuration for filtering: dataset: "HuggingFaceFW/fineweb-edu" output_file: "quantum_computing_entries.jsonl" state_file: "quantum_computing_dataset_state.json" # Processing parameters keywords: - "quantum computing" - "qubit" - "quantum entanglement" - "quantum supremacy" - "quantum algorithm" - "quantum error correction" max_entries: 1024 min_tokens: 1024 max_tokens: 2048 min_int_score: 4 # Shuffling… See the full description on the dataset page: https://huggingface.co/datasets/Lambent/quantum-computing-2048-fineweb-sample.tabular1K<n<10K0 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.