wikipedia-subset
wikipedia_subsetwikipedia_subsetsThis dataset was generated by filtering a subset from the wikipedia dataset -> "wikimedia/wikipedia" -> " 20231101.en"
Detailed information on how this was accomplished is given in this notebook. https://github.com/Umar-Azam/embedding_finetuner_wiki/tree/main
Short Explanation : We have a list of keywords to check against. Each wikipedia text is tokenized into word sets and the "hits" value contains the number of our filter keywords present in each text. Only the items with >4 matches are then… See the full description on the dataset page: https://huggingface.co/datasets/UmarAzam/wikipedia_subsets.PILE_Wikipedia_Pretraining_subset_100k-distill-insert-ret-tokensPILE_Wikipedia_Pretraining_subset_100k-distill-insert-ret-tokens-outputsPILE_Wikipedia_Pretraining_subset_100k-distillPILE_Wikipedia_Pretraining_subset_100k-distill-syn_knowledge
