CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmanPriyanshu /Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens Dynamic Topic Modeling Dataset: RedPajama-1T SubSample (100k samples, 1k tokens) 📝Check out the Blog Post This dataset represents a curated subset of the RedPajama-1T Sample dataset, specifically processed for dynamic topic modeling applications. It contains 100,000 samples from the original dataset, with each document limited to the first 1,024 tokens for consistent processing. Dataset Overview Name:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens.textsummarization100K<n<1M8 likes28 downloads2y agoHugging Face02h-alice /chat-cooking-master-boy-100k Cooking Master Boy Chat Records Chat record dataset from Twitch channel "muse_tw" during the "Cooking Master Boy" (中華一番) marathon event. Introduction This is a chat dataset collected from Twitch channel "muse_tw", while the channel is hosting a marathon anime event featuring "Cooking Master Boy" (中華一番). The featured anime "Cooking Master Boy" is a Japanese manga series written and illustrated by Etsushi Ogawa. And has a big impact on meme culture, and has a cult following… See the full description on the dataset page: https://huggingface.co/datasets/h-alice/chat-cooking-master-boy-100k.tabulartext-classification10K<n<100K1 likes25 downloads2y agoHugging Face03WT-solutions /Kratki-Istorii-Instruct-100kKratki-Istorii-Instruct-100k is a synthetically generated dataset (using INSAIT-Institute/BgGPT-Gemma-2-9B-IT-v1.0) of short stories (3-5) paragraphs, which a young kid should be able to understand. The simplicity of the language used makes it very suitable for training and studying the behaviour of really small Language Models (<500M parameters). The dataset consists of ~100k texts in Bulgarian. You can use the dataset via the HF interface: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/WT-solutions/Kratki-Istorii-Instruct-100k.texttext-generation10K<n<100K1 likes22 downloads8mo agoHugging Face04WT-solutions /Kratki-Istorii-100kKratki-Istorii-100k is a synthetically generated dataset (using INSAIT-Institute/BgGPT-Gemma-2-9B-IT-v1.0) of short stories (3-5) paragraphs, which a young kid should be able to understand. The simplicity of the language used makes it very suitable for training and studying the behaviour of really small Language Models (<500M parameters). The dataset consists of ~100k texts in Bulgarian. You can use the dataset via the HF interface: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/WT-solutions/Kratki-Istorii-100k.texttext-generation10K<n<100K1 likes17 downloads8mo agoHugging Face05prithivMLmods /System-Response-100K System-Response-100K dataset This dataset contains text and code for machine learning tasks including: Text Generation Text Classification Summarization Question Answering The dataset includes text formatted in JSON and is in English. Dataset Statistics Number of entries: Not specified in the information you provided. Modalities Text Code Formats JSON Languages English Getting Started This section can include… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/System-Response-100K.texttext-generation100K<n<1M4 likes16 downloads2y agoHugging Face06UncovAI /fineweb_CC-MAIN-2024-18_100k_output_UncovAI_83362 What is it? As more and more data are generated daily, It becomes important to be able to distinguish between synthetic and human data for model training. We analyzed the first 100k rows of the Fineweb dataset focusing on the dump CC-MAIN-2024-18 using the UncovAI model for text. We observed that more than 16% of the data were detected as having been generated by AI by our model. We removed them and obtained a dataset of 83362 lines with a number of token approaching 55 million.… See the full description on the dataset page: https://huggingface.co/datasets/UncovAI/fineweb_CC-MAIN-2024-18_100k_output_UncovAI_83362.tabulartext-generation10K<n<100K4 likes5 downloads2y agoHugging Face07DBbun /100K_MELD_Plus_v1.0gated Synthetic MELD-Plus (100K Patients) Watch a demo This dataset contains 100,000 synthetic patients inspired by the published MELD-Plus study (a collboration between Massachusetts General Hospital and IBM Research). Each row corresponds to a single admission, with demographics, labs, comorbidities, medications, derived scores (MELD, MELD-Na, MELD-Plus), and the binary outcome Death_Within_90_Days. All data are artificially generated and contain no identifiable patient records.… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/100K_MELD_Plus_v1.0.tabulartext-generation100K<n<1M0 likes1 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.