CoolFace
Datasetpublic

datajuicer/the-pile-philpaper-refined-by-data-juicer

The Pile -- PhilPaper (refined by Data-Juicer) A refined version of PhilPaper dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.7GB). Dataset Information Number of samples: 29,117 (Keep ~88.82% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-philpaper-refined-by-data-juicer.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
0likes10downloads
Dataset Card

The Pile -- PhilPaper (refined by Data-Juicer)

A refined version of PhilPaper dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.

This dataset is usually used to pretrain a Large Language Model.

Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.7GB).

Dataset Information

  • —Number of samples: 29,117 (Keep ~88.82% from the original dataset)

Refining Recipe

yaml
# global parameters
project_name: 'our-recipes-Philpaper'
dataset_path: '/path/to/the/original/dataset/'  # path to your dataset directory or file
export_path: 'Philpaper-refine-result.jsonl' # path to your dataset result file

np: 50  # number of subprocess to process your dataset
ds_cache_dir: /cache # path to your dataset cache file
open_tracer: true

# process schedule
# a list of several process operators with their arguments
process:
  - clean_email_mapper:
  - clean_links_mapper:
  - fix_unicode_mapper:
  - punctuation_normalization_mapper:
  - whitespace_normalization_mapper:

  - alphanumeric_filter:
      tokenization: false
      min_ratio: 0.7  # <3sigma (0.72)
  - average_line_length_filter:
      max_len: 5e5  # >3sigma (406006)
  - character_repetition_filter:
      rep_len: 10
      max_ratio: 0.2  # >3sigma (0.145)
  - flagged_words_filter:
      lang: en
      tokenization: true
      max_ratio: 0.0007  # 3sigma
  - language_id_score_filter:  
      min_score: 0.6
  - maximum_line_length_filter:
      max_len: 1e6  # 3sigma
  - perplexity_filter:
      lang: en
      max_ppl: 5000 
  - special_characters_filter:
      max_ratio: 0.4  # > 3sigma (0.302)
  - words_num_filter:
      lang: en
      tokenization: true
      min_num: 1000  
      max_num: 2e5  # 3sigma
  - word_repetition_filter:
      lang: en
      tokenization: true
      rep_len: 10
      max_ratio: 0.3  # > 3sigma (0.249)

  - document_simhash_deduplicator:
      tokenization: space
      window_size: 6
      lowercase: true
      ignore_pattern: '\p{P}'
      num_blocks: 6
      hamming_distance: 4