CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Skylion007 /openwebtext Dataset Card for "openwebtext" Dataset Summary An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2. This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances plain_text Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/Skylion007/openwebtext.texttext-generation1M<n<10M534 likes50k downloads2mo agoHugging Face02Elriggs /openwebtext-100k Dataset Card for "openwebtext-100k" More Information needed text100K<n<1M8 likes1.5k downloads3y agoHugging Face03RaiBP /openwebtext2-first-30-chunks-lang-detect-raw-output Counting bilingual and monolingual instances In order to count bilingual and monolingual instances, we use the following code. We count bilingual instances where there are two languages, one of them is English and the other is either German, French, Spanish, Italian, Portuguese or Dutch. All other instances fall into the "Other" category. from datasets import load_dataset import json from tqdm import tqdm #Specify the dataset name dataset_name =… See the full description on the dataset page: https://huggingface.co/datasets/RaiBP/openwebtext2-first-30-chunks-lang-detect-raw-output.text100K<n<1M0 likes1.5k downloads3y agoHugging Face04kenhktsui /openwebtext_quality_score_v1 Dataset Card for "openwebtext_quality_score_v1" Adding quality score v1 to Skylion007/openwebtext More Information needed texttext-generation1M<n<10M0 likes1.2k downloads3y agoHugging Face05dylanebert /openwebtext Dataset Card for "openwebtext" Dataset Summary An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2. This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances plain_text Size of downloaded dataset files: 13.51 GB Size of the… See the full description on the dataset page: https://huggingface.co/datasets/dylanebert/openwebtext.texttext-generation1M<n<10M4 likes964 downloads1y agoHugging Face06hanspeterlyngsoeraaschoujensen /gpt2_model_acts_openwebtexttextn<1K0 likes937 downloads2y agoHugging Face07vietgpt /the_pile_openwebtext2 Dataset Card for "the_pile_openwebtext2" More Information needed text10M<n<100M5 likes807 downloads3y agoHugging Face08Geralt-Targaryen /openwebtext2A cleaned version of OpenWebText2 by removing non-English, duplicated, copyrighted, and low-quality (too short, too many special characters, etc) samples. This dataset has also been decontaminated with respect to the following benchmarks based on n-gram overlap: GLUE (dev set of SST-2, CoLA, QQP, WNLI, RTE, QNLI, MNLI; test set of MPRC) SIQA, PIQA, QASC, CSQA, HellaSWAG (all dev set) CONLL 2003 BLIMP MAIN BoolQ (dev set) WinoGrande (dev set) ANLI (test set) ARC easy and challenge (test set)… See the full description on the dataset page: https://huggingface.co/datasets/Geralt-Targaryen/openwebtext2.text10M<n<100M8 likes655 downloads1y agoHugging Face09YangZhoumill /openwebtextpresplit Fixed OpenWebMath train/test split This is an untokenized, deterministic shuffle of open-web-math/open-web-math pinned at commit fde8ef8de2300f5e778f56261843dab89f230815. It contains the original columns without transformation. The shuffle seed is 20260904. The test set is the first 0.05% of shuffled rows (rounded to the nearest whole document); all remaining rows are training data. Exact counts and SHA-256 checksums are in split_manifest.json. texttext-generation1M<n<10M0 likes422 downloads18d agoHugging Face10vietgpt /openwebtext_en Dataset Card for "openwebtext_en" More Information needed text1M<n<10M2 likes400 downloads3y agoHugging Face11jdeschena /openwebtexttext1M<n<10M0 likes353 downloads9mo agoHugging Face12RaiBP /openwebtext2-first-30-chunks-ablation-bilingualtext1M<n<10M0 likes330 downloads3y agoHugging Face13lsb /openwebtext-all-minilm-l6-v2-embedding Dataset Card for "openwebtext-all-minilm-l6-v2-embedding" More Information needed text1M<n<10M0 likes317 downloads4y agoHugging Face14ThomasKendrick /openwebtext Dataset Card for "openwebtext" Dataset Summary An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2. This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances plain_text Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/ThomasKendrick/openwebtext.texttext-generation1M<n<10M0 likes279 downloads21d agoHugging Face15RaiBP /openwebtext2-first-30-chunks-ablation-non-englishtext1M<n<10M0 likes273 downloads3y agoHugging Face16YangZhoumill /openwebtext-presplit Fixed OpenWebText train/test split This is an untokenized, deterministic shuffle of Skylion007/openwebtext pinned at commit 79d93d786212f7344586290adb811d4ae6a1762c. It contains the original columns without transformation. The shuffle seed is 2357. The test set is the first 0.05% of shuffled rows (rounded to the nearest whole document); all remaining rows are training data. Exact counts and SHA-256 checksums are in split_manifest.json. texttext-generation1M<n<10M0 likes269 downloads18d agoHugging Face17Bingsu /openwebtext_20p openwebtext_20p first 20% of openwebtext texttext-generation10M<n<100M14 likes250 downloads4y agoHugging Face18JianYu03 /openwebtext-dep-spacytext1M<n<10M0 likes239 downloads5mo agoHugging Face19RaiBP /openwebtext2-first-30-chunks-ablation-translationtext1M<n<10M0 likes210 downloads3y agoHugging Face20PaulPauls /openwebtext-sentences OpenWebText-Sentences Dataset Overview This dataset is derived from the popular OpenWebText dataset (see here). It contains the same text content as the original OpenWebText, but split into individual sentences. Key Features Content: All text from the original OpenWebText dataset Format: Sentences are stored individually now in parquet format for faster access Order: Maintains all original OpenWebText text and the order thereof Tokenization: Sentences were… See the full description on the dataset page: https://huggingface.co/datasets/PaulPauls/openwebtext-sentences.text100M<n<1B3 likes183 downloads2y agoHugging Face21shaguftakhan2k17 /openwebtext2A cleaned version of OpenWebText2 by removing non-English, duplicated, copyrighted, and low-quality (too short, too many special characters, etc) samples. This dataset has also been decontaminated with respect to the following benchmarks based on n-gram overlap: GLUE (dev set of SST-2, CoLA, QQP, WNLI, RTE, QNLI, MNLI; test set of MPRC) SIQA, PIQA, QASC, CSQA, HellaSWAG (all dev set) CONLL 2003 BLIMP MAIN BoolQ (dev set) WinoGrande (dev set) ANLI (test set) ARC easy and challenge (test set)… See the full description on the dataset page: https://huggingface.co/datasets/shaguftakhan2k17/openwebtext2.text10M<n<100M0 likes182 downloads9mo agoHugging Face22RaiBP /openwebtext2-first-30-chunks-english-only-examplestext1M<n<10M0 likes157 downloads3y agoHugging Face23RaiBP /openwebtext2-first-30-chunks-ablation-fulltext1M<n<10M0 likes138 downloads3y agoHugging Face24segyges /OpenWebText2 Dataset Card for OpenWebText2 OpenWebText2 is a reasonably large corpus of scraped natural language data. Original hosting for this dataset has become difficult because it was hosted alongside another controversial dataset. To the best of my knowledge, this dataset itself is not encumbered in any way. It's a useful size for smaller language modelling experiments and is sometimes used in existing papers which it may be desirable to replicate. It is uploaded here to facilitate those… See the full description on the dataset page: https://huggingface.co/datasets/segyges/OpenWebText2.texttext-generationn<1K18 likes135 downloads2y agoHugging Face25zhihanyang /openwebtext-128text10M<n<100M0 likes135 downloads7mo agoHugging Face26ashaba1in /small_openwebtexttext1M<n<10M1 likes86 downloads2y agoHugging Face27sjyuxyz /openwebtext-subset-1M-rowstext1M<n<10M0 likes85 downloads2y agoHugging Face28reedie01 /openwebtext Dataset Card for "openwebtext" Dataset Summary An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2. This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances plain_text Size of downloaded dataset files: 13.51 GB Size of the… See the full description on the dataset page: https://huggingface.co/datasets/reedie01/openwebtext.texttext-generation1M<n<10M0 likes82 downloads8mo agoHugging Face29parameterlab /scaling_mia_the_pile_00_OpenWebText2text1M<n<10M1 likes78 downloads2y agoHugging Face30suolyer /pile_openwebtext2text10K<n<100K3 likes76 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.