CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nomic-ai /bert-128-grouped Dataset Card for "bert-128-grouped" More Information needed 10M<n<100M0 likes3.7k downloads3y agoHugging Face02Jackmin108 /bert-base-uncased-refined-web-segment0 Dataset Card for "bert-base-uncased-refined-web-segment0" More Information needed 100M<n<1B0 likes2.3k downloads3y agoHugging Face03bertram-gilfoyle /CC-MAIN-2022-21-rawtext10M<n<100M0 likes1.7k downloads3y agoHugging Face04gsgoncalves /bert_pretrain Dataset Card for "bert_pretrain" More Information needed text10M<n<100M0 likes1.6k downloads3y agoHugging Face05bertram-gilfoyle /CC-MAIN-2023-06-rawtext10M<n<100M0 likes1.5k downloads3y agoHugging Face06bertram-gilfoyle /CC-MAIN-2022-05-rawtext10M<n<100M0 likes1.4k downloads3y agoHugging Face07bertram-gilfoyle /CC-MAIN-2023-40-rawtext10M<n<100M0 likes1.3k downloads3y agoHugging Face08LakoreAI /bert-mlm-experiments-en Unified English MLM Pre-training Corpus (80M Rows) This dataset is a massive, diverse, multi-domain English text corpus explicitly engineered for pre-training and domain-adaptation of BERT-style models via Masked Language Modeling (MLM). It aggregates over 80 million rows of text, completely stripped of auxiliary metadata, labels, and identifiers to expose purely raw text strings. Dataset Details Repository ID: 8Opt/bert-mlm-experiments-en Total Rows: 80,489,226… See the full description on the dataset page: https://huggingface.co/datasets/LakoreAI/bert-mlm-experiments-en.textfill-mask10M<n<100M1 likes1.2k downloads3mo agoHugging Face09bertram-gilfoyle /CC-MAIN-2022-27-rawtext10M<n<100M0 likes1.2k downloads3y agoHugging Face10nthngdy /bert_dataset_202203 Dataset Card for "bert_dataset_202203" More Information needed texttext-generation100M<n<1B0 likes971 downloads4y agoHugging Face11bertram-gilfoyle /CC-MAIN-2022-33-rawtext10M<n<100M0 likes931 downloads3y agoHugging Face12nomic-ai /nomic-bert-2048-pretraining-data Dataset Card for "bert-pretokenized-2048-wiki-2023" More Information needed 1M<n<10M1 likes922 downloads3y agoHugging Face13bertram-gilfoyle /CC-MAIN-2021-25-rawtext10M<n<100M0 likes916 downloads3y agoHugging Face14bertram-gilfoyle /CC-MAIN-2023-14-rawtext10M<n<100M0 likes904 downloads3y agoHugging Face15bertram-gilfoyle /CC-MAIN-2022-49-rawtext10M<n<100M0 likes846 downloads3y agoHugging Face16bertram-gilfoyle /CC-MAIN-2021-10-rawtext10M<n<100M0 likes709 downloads3y agoHugging Face17bertram-gilfoyle /CC-MAIN-2023-23-rawtext10M<n<100M0 likes686 downloads3y agoHugging Face18JackBAI /bert_pretrain_datasets Dataset Card for "bert_pretrain_datasets" This dataset is essentially a concatenation of the training set of the English Wikipedia (wikipedia.20220301.en.train) and the Book Corpus (bookcorpus.train). This is exactly how I get this dataset: from datasets import load_dataset, concatenate_datasets, load_from_disk cache_dir = "/data/haob2/cache/" # book corpus bookcorpus = load_dataset("bookcorpus", split="train", cache_dir=cache_dir) # english wikipedia wiki =… See the full description on the dataset page: https://huggingface.co/datasets/JackBAI/bert_pretrain_datasets.text10M<n<100M1 likes637 downloads3y agoHugging Face194rtemi5 /processed_tsb_bert_dataset1M<n<10M0 likes631 downloads2y agoHugging Face20Tural /processed_bert_dataset Dataset Card for "processed_bert_dataset" More Information needed 10M<n<100M0 likes630 downloads3y agoHugging Face21angie-chen55 /bert_pretraining_data Dataset Card for "bert_pretraining_data" More Information needed 10M<n<100M0 likes623 downloads3y agoHugging Face22bertin-project /alpaca-spanish BERTIN Alpaca Spanish This dataset is a translation to Spanish of alpaca_data_cleaned.json, a clean version of the Alpaca dataset made at Stanford. An earlier version used Facebook's NLLB 1.3B model, but the current version uses OpenAI's gpt-3.5-turbo, hence this dataset cannot be used to create models that compete in any way against OpenAI. texttext-generation10K<n<100K36 likes616 downloads4y agoHugging Face23bertram-gilfoyle /CC-MAIN-2023-50-rawtext10M<n<100M0 likes615 downloads3y agoHugging Face24delmeng /processed_bert_dataset Dataset Card for "processed_bert_dataset" More Information needed 1M<n<10M0 likes549 downloads3y agoHugging Face25sanagnos /processed_bert_dataset Dataset Card for "processed_bert_dataset" More Information needed 1M<n<10M0 likes529 downloads4y agoHugging Face26bertram-gilfoyle /CC-MAIN-2020-29-rawtext10M<n<100M0 likes526 downloads3y agoHugging Face27RamWithAPlan /processed_bert_dataset Dataset Card for "processed_bert_dataset" More Information needed 1M<n<10M0 likes461 downloads3y agoHugging Face28bertram-gilfoyle /CC-MAIN-2020-45-rawtext10M<n<100M0 likes434 downloads3y agoHugging Face29gokuls /wiki_book_corpus_complete_processed_bert_dataset Dataset Card for "wiki_book_corpus_complete_processed_bert_dataset" More Information needed 1M<n<10M0 likes426 downloads4y agoHugging Face30philschmid /processed_bert_dataset1M<n<10M1 likes418 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.