CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01liang2kl /RedPajama-Data-1T-Sample-Backuptext100K<n<1M0 likes1.9k downloads11mo agoHugging Face02nhagar /redpajama-data-v2_urls Dataset Card for redpajama-data-v2_urls This dataset provides the URLs and top-level domains associated with training records in togethercomputer/RedPajama-Data-V2. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/redpajama-data-v2_urls.text1B<n<10B0 likes1.6k downloads1y agoHugging Face03ZengXiangyu /RedPajama-Data-1T-Sampletext100K<n<1M3 likes1.2k downloads10mo agoHugging Face04latam-gpt /red_pajama_es_hq RedPajama's High Quality Spanish subset What is this? The following is a high-quality dataset distilled from the Spanish subsection of RedPajama-Data-v2, created using the methodology proposed in FineWEB-Edu. Usage from datasets import load_dataset ds = load_dataset("latam-gpt/red_pajama_es_hq") Filtering by quality score Documents in this corpus are scored on academic quality from 2.5 to 5, with higher scores indicating better quality. The… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/red_pajama_es_hq.tabular100M<n<1B11 likes635 downloads2y agoHugging Face05distily /filtered_redpajama_multilingualtext1M<n<10M1 likes445 downloads2y agoHugging Face06sade-adrien /redpajama_v2_sample_10M Dataset Card for "redpajama_v2_sample_10M" More Information needed text10M<n<100M1 likes396 downloads3y agoHugging Face07sade-adrien /redpajama_v2_sample_100M Dataset Card for "redpajama_v2_sample_100M" More Information needed text100M<n<1B0 likes372 downloads3y agoHugging Face08distily /filtered_redpajama_entext1M<n<10M0 likes301 downloads2y agoHugging Face09gair-prox /RedPajama-pro 📚 RedPajama-pro ArXiv | Models | Code RedPajama-pro is refined from RedPajama-Data-V2 using the ProX refining framework. It contains about 30B high quality tokens, ready for general language model pre-training. License RedPajama-pro is based on RedPajama-Data-V2, which is made available under an apache-2.0 license; users should also abide by the CommonCrawl ToU: https://commoncrawl.org/terms-of-use/. We do not alter the license of any of the underlying data.… See the full description on the dataset page: https://huggingface.co/datasets/gair-prox/RedPajama-pro.texttext-generation10M<n<100M4 likes294 downloads2y agoHugging Face10Shahzebbb /redpajama_100gb_processedtext100M<n<1B0 likes245 downloads3y agoHugging Face11krisbailey /RedPajama-Data-V2-1B RedPajama-Data-V2 1B Dataset Description This is a 1.01 Billion token subset of the togethercomputer/RedPajama-Data-V2 dataset (specifically derived from the sample-10B config). It was created by randomly sampling the source data. Motivation RedPajama V2 is a state-of-the-art web dataset with rich quality signals. This 1B token subset allows for rapid testing of these quality signals or other filtering experiments without needing to process the full… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-1B.texttext-generation100K<n<1M0 likes226 downloads8mo agoHugging Face12ivanzhouyq /RedPajama-Tiny Dataset Card for Dataset Name Dataset Summary This is a tiny version of the RedPajama dataset. It contains 64 samples from each of the 7 sources. This dataset is intended for developing and testing data/training pipeline for loading the full RedPajama dataset or any general HuggingFace dataset. It is very fast to download and easy to examine. You should not use it for training a full model, but you can use it for overfitting test or any other sanity checks.… See the full description on the dataset page: https://huggingface.co/datasets/ivanzhouyq/RedPajama-Tiny.texttext-generationn<1K5 likes224 downloads2y agoHugging Face13sade-adrien /redpajama_v2_sample_1M Dataset Card for "redpajama_v2_sample_1M" More Information needed text1M<n<10M1 likes223 downloads3y agoHugging Face14Tristan /RedPajama-Data-V2-sample-100B-filtered-for-regression-domains-with-domainstext1M<n<10M0 likes193 downloads2y agoHugging Face15ll922 /RedPajama-Data-1T-Sample-Backup RedPajama Data 1T Sample Backup This dataset is a backup mirror of togethercomputer/RedPajama-Data-1T-Sample. It is provided for easier access when the original dataset is unavailable or difficult to download. Usage Original: from datasets import load_dataset ds = load_dataset( "togethercomputer/RedPajama-Data-1T-Sample", split="train", trust_remote_code=True, ) Backup: from datasets import load_dataset ds = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/ll922/RedPajama-Data-1T-Sample-Backup.texttext-generation100K<n<1M0 likes193 downloads5mo agoHugging Face16Tristan /RedPajama-Data-V2-sample-100B-filtered-shuffled-tokenized-with-token-countstext1M<n<10M0 likes138 downloads2y agoHugging Face17erickfmm /red_pajama_es_hq_35Just a filtered version of https://huggingface.co/datasets/latam-gpt/red_pajama_es_hq with scores >=3.5 tabulartext-generation10M<n<100M0 likes128 downloads1y agoHugging Face18nhagar /redpajama-data-1t_urls Dataset Card for redpajama-data-1t_urls This dataset provides the URLs and top-level domains associated with training records in togethercomputer/RedPajama-Data-1T. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/redpajama-data-1t_urls.text1B<n<10B0 likes122 downloads1y agoHugging Face19xmohri /RedPajama-Data-V2-sample-100B-filtered-shuffled-tokenized-with-token-counts-2500000entext1M<n<10M0 likes112 downloads2y agoHugging Face20LM-Polygraph /RedPajama-Data-100-Sample-For-Testtextn<1K0 likes94 downloads2y agoHugging Face21zxzy /redpajama_arxiv_hftext100M<n<1B0 likes92 downloads1y agoHugging Face22TheFinAI /corpus-shard-redpajama-19tabular100K<n<1M0 likes91 downloads5mo agoHugging Face23TheFinAI /corpus-shard-redpajama-20tabular100K<n<1M0 likes90 downloads5mo agoHugging Face24TheFinAI /corpus-shard-redpajama-17tabular10K<n<100K0 likes87 downloads5mo agoHugging Face25TheFinAI /corpus-shard-redpajama-22tabular100K<n<1M0 likes82 downloads5mo agoHugging Face26TheFinAI /corpus-shard-redpajama-18tabular100K<n<1M0 likes79 downloads5mo agoHugging Face27michelangelo-engs /RedPajama-Data-1T-1024SampleRedPajama is a clean-room, fully open-source implementation of the LLaMa dataset. This is a 1B-token sample of the full dataset.texttext-generation1K<n<10K1 likes78 downloads3y agoHugging Face28sordonia /redpajama-sample_from_valid_all Dataset Card for "redpajama-sample_from_valid_all" More Information needed tabular100K<n<1M1 likes77 downloads3y agoHugging Face29Abzu /RedPajama-Data-1T-arxiv-filtered Dataset Card for "RedPajama-Data-1T-arxiv-filtered" More Information needed text1K<n<10K4 likes75 downloads3y agoHugging Face30TheFinAI /corpus-shard-redpajama-21tabular100K<n<1M0 likes74 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.