CoolFace
20 results

text-filter

sayan1101 /gaia_filtered_text_onlytextn<1K0 likes3.1k downloads3y agoHugging Facemalaysia-ai /mosaic-dedup-text-dataset-filtered Mosaic format for filtered dedup text dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-filtered-4096.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset load it… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset-filtered.textn<1K0 likes1.9k downloads3y agoHugging Facetext-machine-lab /vocab_filtered_dataset_22B Dataset Card for "vocab_filtered_dataset_22B" Dataset Summary This data is the simplified vocabulary-filtered pretraining data published by "Emergent Abilities in Reduced-Scale Generative Language Models". The vocabulary is derived from the AO-Childes speech corpus (https://github.com/UIUCLearningLanguageLab/AOCHILDES) We filter the train split of SlimPajama dataset (https://huggingface.co/datasets/cerebras/SlimPajama-627B) based on the AO-Childes vocabulary retaining… See the full description on the dataset page: https://huggingface.co/datasets/text-machine-lab/vocab_filtered_dataset_22B.text100M<n<1B0 likes681 downloads2y agoHugging Facetext-machine-lab /vocab_filtered_dataset_2.1B Dataset Card for "vocab_filtered_dataset_2.1B" Dataset Summary This data is the simplified vocabulary-filtered pretraining data published by "Emergent Abilities in Reduced-Scale Generative Language Models". The vocabulary is derived from the AO-Childes speech corpus (https://github.com/UIUCLearningLanguageLab/AOCHILDES) We filter the train split of SlimPajama dataset (https://huggingface.co/datasets/cerebras/SlimPajama-627B) based on the AO-Childes vocabulary retaining… See the full description on the dataset page: https://huggingface.co/datasets/text-machine-lab/vocab_filtered_dataset_2.1B.text1M<n<10M0 likes268 downloads2y agoHugging FaceFastVideo /vidprom_filtered_extended_umt5_text_embed1 likes258 downloads6mo agoHugging FaceTucanoBR /GigaVerbo-Text-Filter GigaVerbo Text-Filter Dataset Summary GigaVerbo Text-Filter is a dataset with 110,000 randomly selected samples from 9 subsets of GigaVerbo (i.e., specifically those that were not synthetic). This dataset was used to train the text-quality filters described in "Tucano: Advancing Neural Text Generation for Portuguese". To create the text embeddings, we used sentence-transformers/LaBSE. All scores were generated by GPT-4o. Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/TucanoBR/GigaVerbo-Text-Filter.texttext-classification100K<n<1M8 likes107 downloads1y agoHugging Face