CoolFace
15 results

harley-ml

Harley-ml /es-en-words Words A dataset comprised of 753k words, 90k of them are Spanish, and 660k of them are English. Key Value Entries (words) 753,232 Tokens 3,225,398 Characters 7,022,310 Avg. Tokens Per Entry ~4.2 Avg. Words Per Entry 1 Avg. Chars Per Entry ~9.3 Longest Entry (Tokens) 36 Shortest Entry (Tokens) 1 English Words~660k Spanish Words ~90k Check out Tiny-Word: A Model Trained on 753k Words Have fun. ALotta Words for you to enjoy! texttext-generation100K<n<1M0 likes24 downloads9mo agoHugging FaceHarley-ml /lesswrong What is LessWrong? LessWrong is a community blog and forum dedicated to improving human reasoning and decision-making, aiming to help people hold more accurate beliefs and be more effective, or "less wrong," daily. This dataset contains over twenty-six thousand posts from 2009 and onward. Stats Key Value Entries 26,517 Total Tokens (GPT2) 96,399,665 Total Words 53,966,975 Avg Tokens / Entry 3,635.39 Avg Words / Entry 2,035.18 We counted the tokens… See the full description on the dataset page: https://huggingface.co/datasets/Harley-ml/lesswrong.texttext-generation10K<n<100K1 likes17 downloads5mo agoHugging FaceHarley-ml /HFMC HFMC HFMC stands for "Hugging Face Model Configs." This dataset has over 7k json model configs from Hugging Face. We used the Hugging Face API to scrape each one. Stats Metric Value Tokens (GPT2) 7,117,346 Entries 18,781 Tokens/Entry (med) 318 Words 1,420,637 Words/Entry (med) 62 We counted the tokens using GPT2's tokenizer. Cleaning We deduped, filtered via lang (only English), and length (1024 tokens using an in-domain tokenizer).… See the full description on the dataset page: https://huggingface.co/datasets/Harley-ml/HFMC.texttext-generation10K<n<100K0 likes17 downloads5mo agoHugging FaceHarley-ml /i-statements I-Statements This dataset has axproximently 5,335 I-statements generated by Qwen2.5-7B-Q4_K_M using Ollama. Stats Metric Value Entries 5,334 Total tokens (GPT2) 36,032 Total words 29,735 Avg. tokens per entry 6.67 Avg. words per entry 5.57 Word range 3–10 Unique vocab (words) 2,237 Unique verbs 252 We used GPT2's tokenizer to find the token count. Note: The tokens may vary depending on the tokenizer used. Use Cases This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Harley-ml/i-statements.texttext-generation1K<n<10K0 likes13 downloads5mo agoHugging Face