sudy-super/piece-of-refined-oscar
Descrption This dataset is part of the OSCAR-2301 cleaned. There are about 0.5b tokens counted by calm2 tokenizer. NOTE This dataset has not passed sentence end boundary determination or Perplexity Filtering, so there is room for improvement in quality.
733
Descrption
This dataset is part of the OSCAR-2301 cleaned.
There are about 0.5b tokens counted by calm2 tokenizer.
NOTE
This dataset has not passed sentence end boundary determination or Perplexity Filtering, so there is room for improvement in quality.
