CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lots-of-LoRAs /task1728_web_nlg_data_to_text Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1728_web_nlg_data_to_text Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1728_web_nlg_data_to_text.texttext-generation1K<n<10K0 likes175 downloads2y agoHugging Face02devngho /korean-webtext-edugated 🇰🇷🌐📚 korean-webtext-edu HAERAE-HUB/KOREAN-WEBTEXT를 devngho/ko_edu_classifier_v2_nlpai-lab_KoE5 모델로 평가한 데이터셋 불러오기 from datasets import load_dataset ds = load_dataset("devngho/korean-webtext-edu", name="scored_over_3", split="train") 성능 예정 컴퓨팅 Google Cloud TPU, transformers, JAX, tpuswarm 하드웨어 TPU v4-8 x 4 instances, 약 35분 소요 이 연구는 Google의 TPU Research Cloud (TRC)의 Cloud TPU 제공으로 수행되었습니다. ⚡ 라이선스 원본… See the full description on the dataset page: https://huggingface.co/datasets/devngho/korean-webtext-edu.texttext-generation1M<n<10M4 likes118 downloads2y agoHugging Face03softcatala /Softcatala-Web-Texts-Dataset Dataset Card for Softcatala-Web-Texts-Dataset Dataset Summary This repository contains Softcatala website content (articles and programs descriptions). Dataset size: articles.json contains 623 articles with 373233 words. programes.json contains 330 program descriptions with 49868 words. The license of the data is Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) or Universal Public Domain Dedication (CC0 1.0) Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/Softcatala-Web-Texts-Dataset.texttext-generationn<1K0 likes100 downloads2mo agoHugging Face045CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes99 downloads3y agoHugging Face05khtsly /luau-org-web-texttexttext-generationn<1K1 likes95 downloads1mo agoHugging Face06TopAI-1 /Minecraft-WebText-2 ⛏️ MCGPT-1: Massive Minecraft Expert Dataset This dataset is a highly specialized collection of data focused exclusively on Minecraft mechanics, entities, blocks, and gameplay. 📊 Dataset Statistics Total Tokens: 5,227,988 🪙 Total Lines: 18,844 📝 File Size: 23.11 MB 📂 Content: Deep-dive into game mechanics, crafting recipes, mob behavior, and world generation. 🎯 Purpose This is the "Specialist" module for MCGPT-1. While general datasets provide language… See the full description on the dataset page: https://huggingface.co/datasets/TopAI-1/Minecraft-WebText-2.texttext-generation10K<n<100K2 likes43 downloads7mo agoHugging Face07Raziel1234 /WebText-1 Orion-Spark-30M Dataset The Orion-Spark-30M dataset is a curated corpus containing 10,846 lines of text gathered from reputable internet sources, including Wikipedia pages, technology news websites, and educational platforms. The dataset focuses on foundational and advanced topics related to artificial intelligence, machine learning, large language models, and generative pretrained transformers. It also covers major technology companies, influential figures, and key concepts in the… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/WebText-1.texttext-generation10K<n<100K0 likes38 downloads1y agoHugging Face08Raziel1234 /WebText-4 📦 Huge Multilingual Text Dataset Overview This repository contains a large-scale, multilingual dataset designed for training and evaluating advanced language models and AI agents. The dataset spans multiple domains, supports multiple languages, and is suitable for high-capacity models. 🔍 Key Characteristics Dataset Size: 1B < n < 10B tokens License: Apache License 2.0 (Apache-2.0) Task Category: Text Generation Languages Supported: English (en) Hebrew… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/WebText-4.texttext-generation100K<n<1M2 likes30 downloads8mo agoHugging Face09nampdn-ai /tiny-webtextgated Tiny WebText The Tiny WebText dataset is designed to help models learn about perception on web text while neutralizing the bias of the source text using critical thinking methods. By providing a rich and diverse set of texts, I aim to improve the ability of models to understand and analyze information in a more objective and unbiased manner. This dataset can be used to train and evaluate natural language processing and machine learning models, with the goal of improving their… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-webtext.tabulartext-generation1M<n<10M37 likes29 downloads3y agoHugging Face10Raziel1234 /WebText-3 WebText-3 Corpus WebText-3 is a large-scale, diverse text corpus collected from publicly available web pages. It contains cleaned and normalized sentences suitable for natural language processing (NLP), machine learning, and AI training. Dataset Overview Format: Plain text (.txt), one sentence per line Approximate Size: 200,000+ sentences Languages: Primarily English, with occasional Hebrew content Source Types: Wikipedia articles, technology news sites, blogs… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/WebText-3.texttext-generation10K<n<100K0 likes23 downloads1y agoHugging Face11Raziel1234 /Minecraft-Webtext Minecraft Text Corpus Dataset This dataset is a large-scale text corpus collected from various Minecraft-related online sources.It contains descriptive, community-driven, and encyclopedic content covering the Minecraft universe, including its mechanics, updates, characters, community discussions, and cultural impact. Contents The dataset is stored in plain text format (minecraft_corpus.txt) and consists of cleaned and segmented lines of natural language text.The text has… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/Minecraft-Webtext.texttext-generation1K<n<10K0 likes22 downloads1y agoHugging Face12Raziel1234 /WebText-2 Orion-Spark-2 Dataset Overview The Orion-Spark-2 Dataset is a text corpus curated for training the Orion-Spark-2 transformer language model. It consists of a diverse collection of sentences extracted from multiple sources including Wikipedia articles, technology news sites, developer resources, and other open-access web pages. The dataset is designed to provide broad coverage of general knowledge, programming topics, artificial intelligence, space, popular culture, and… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/WebText-2.texttext-generation10K<n<100K0 likes20 downloads1y agoHugging Face13TopAI-1 /Reddit-WebText 💬 MCGPT-1: Massive Reddit Interaction Dataset This dataset contains high-quality conversational data extracted and filtered from Minecraft-related discussions on Reddit. It is a key component in giving MCGPT-1 its human-like personality. 📊 Dataset Statistics Total Tokens: 20,360 🪙 Total Lines: 36 (High-density, long-form discussions) 📝 Focus: Community discussions, player advice, and Minecraft-specific social interaction. 🎯 Purpose While the Mega Dataset… See the full description on the dataset page: https://huggingface.co/datasets/TopAI-1/Reddit-WebText.texttext-generationn<1K0 likes16 downloads7mo agoHugging Face14Srijan-Srivastava /webtext-super-tinytexttext-generation1K<n<10K1 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.