CoolFace
20 results

redpajama

togethercomputer /RedPajama-Data-V2RedPajama V2: an Open Dataset for Training Large Language Modelstext-generation408 likes13k downloads2y agoHugging Facemarketeam /raw_redpajamas Getting Started The dataset is built from the redpajamas dataset after filtering by marketing keywords list that can be found here The full scripts to recreate the raw dataset before sharding can be found here. The dataset includes: ~4.8B tokens from raw contents. Downloading the dataset To start exploring and get to know the dataset you can run the script: import datasets ds = datasets.load_dataset("marketeam/raw_redpajamas", split="train") for sample in ds:… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/raw_redpajamas.text-generation1B<n<10B0 likes3.4k downloads2y agoHugging Facetogethercomputer /RedPajama-Data-1TRedPajama is a clean-room, fully open-source implementation of the LLaMa dataset.text-generation1.2k likes1.9k downloads2y agoHugging Faceliang2kl /RedPajama-Data-1T-Sample-Backuptext100K<n<1M0 likes1.9k downloads11mo agoHugging Facenhagar /redpajama-data-v2_urls Dataset Card for redpajama-data-v2_urls This dataset provides the URLs and top-level domains associated with training records in togethercomputer/RedPajama-Data-V2. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/redpajama-data-v2_urls.text1B<n<10B0 likes1.6k downloads1y agoHugging FaceZengXiangyu /RedPajama-Data-1T-Sampletext100K<n<1M3 likes1.2k downloads10mo agoHugging Face