CoolFace
Datasetpublic

datajuicer/redpajama-cc-2019-30-refined-by-data-juicer

RedPajama -- CommonCrawl-2019-30 (refined by Data-Juicer) A refined version of CommonCrawl-2019-30 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 240GB). Dataset Information Number of samples: 36,557,283 (Keep ~45.08% from the original… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2019-30-refined-by-data-juicer.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
0likes19downloads
Dataset Card

RedPajama -- CommonCrawl-2019-30 (refined by Data-Juicer)

A refined version of CommonCrawl-2019-30 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.

This dataset is usually used to pretrain a Large Language Model.

Notice: Here is a small subset for previewing. The whole dataset is available here (About 240GB).

Dataset Information

  • —Number of samples: 36,557,283 (Keep ~45.08% from the original dataset)

Refining Recipe

yaml
# global parameters
project_name: 'Data-Juicer-recipes-cc-2019-30'
dataset_path: '/path/to/your/dataset'  # path to your dataset directory or file
export_path: '/path/to/your/dataset.jsonl'

np: 50  # number of subprocess to process your dataset
open_tracer: true

# process schedule
# a list of several process operators with their arguments
process:
  - document_simhash_deduplicator:
      tokenization: space
      window_size: 6
      lowercase: true
      ignore_pattern: '\p{P}'
      num_blocks: 6
      hamming_distance: 4

  - clean_email_mapper:
  - clean_links_mapper:
  - fix_unicode_mapper:
  - punctuation_normalization_mapper:
  - whitespace_normalization_mapper:

  - alphanumeric_filter:  # 770218
      tokenization: false
      min_ratio: 0.7489  # 3sigma
      max_ratio: 0.8585  # 3sigma
  - average_line_length_filter:  # for code
      max_len: 1500  # < 3sigma (2689) -- 177520
  - character_repetition_filter:
      rep_len: 10
      max_ratio: 0.3  # > 3sigma (0.1491) -- 151703
  - flagged_words_filter:
      lang: en
      tokenization: true
      max_ratio: 0.0025  # 3sigma -- 101540
  - language_id_score_filter:  # remove language filter
      min_score: 0.788  # 3sigma -- 1622574
  - maximum_line_length_filter:  # for code
      max_len: 5000  # < 3sigma (8775) -- 485806
  - perplexity_filter:
      lang: en
      max_ppl: 5000  # < 3sigma (6723) -- 676914
  - special_characters_filter:
      min_ratio: 0.15  # > 3sigma (0.104)
      max_ratio: 0.35  # > 3sigma (0.322) -- 859797
  - text_length_filter:
      max_len: 65589  # 3sigma -- 975142
  - words_num_filter:
      lang: en
      tokenization: true
      min_num: 20  # > 3sigma -- 196
      max_num: 13030  # 3sigma -- 989078
  - word_repetition_filter:
      lang: en
      tokenization: true
      rep_len: 10
      max_ratio: 0.279  # 3sigma -- 1716308