datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task1728_web_nlg_data_to_text
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1728_web_nlg_data_to_text
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1728_web_nlg_data_to_text.korean-webtext-edu
🇰🇷🌐📚 korean-webtext-edu
HAERAE-HUB/KOREAN-WEBTEXT를 devngho/ko_edu_classifier_v2_nlpai-lab_KoE5 모델로 평가한 데이터셋
불러오기
from datasets import load_dataset
ds = load_dataset("devngho/korean-webtext-edu", name="scored_over_3", split="train")
성능
예정
컴퓨팅
Google Cloud TPU, transformers, JAX, tpuswarm
하드웨어
TPU v4-8 x 4 instances, 약 35분 소요
이 연구는 Google의 TPU Research Cloud (TRC)의 Cloud TPU 제공으로 수행되었습니다. ⚡
라이선스
원본… See the full description on the dataset page: https://huggingface.co/datasets/devngho/korean-webtext-edu.Softcatala-Web-Texts-Dataset
Dataset Card for Softcatala-Web-Texts-Dataset
Dataset Summary
This repository contains Softcatala website content (articles and programs descriptions).
Dataset size:
articles.json contains 623 articles with 373233 words.
programes.json contains 330 program descriptions with 49868 words.
The license of the data is Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) or Universal Public Domain Dedication (CC0 1.0)
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/Softcatala-Web-Texts-Dataset.Vietnamese-nampdn-ai-tiny-webtext-gg-translatedluau-org-web-textMinecraft-WebText-2
⛏️ MCGPT-1: Massive Minecraft Expert Dataset
This dataset is a highly specialized collection of data focused exclusively on Minecraft mechanics, entities, blocks, and gameplay.
📊 Dataset Statistics
Total Tokens: 5,227,988 🪙
Total Lines: 18,844 📝
File Size: 23.11 MB 📂
Content: Deep-dive into game mechanics, crafting recipes, mob behavior, and world generation.
🎯 Purpose
This is the "Specialist" module for MCGPT-1. While general datasets provide language… See the full description on the dataset page: https://huggingface.co/datasets/TopAI-1/Minecraft-WebText-2.WebText-1
Orion-Spark-30M Dataset
The Orion-Spark-30M dataset is a curated corpus containing 10,846 lines of text gathered from reputable internet sources, including Wikipedia pages, technology news websites, and educational platforms. The dataset focuses on foundational and advanced topics related to artificial intelligence, machine learning, large language models, and generative pretrained transformers. It also covers major technology companies, influential figures, and key concepts in the… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/WebText-1.WebText-4
📦 Huge Multilingual Text Dataset
Overview
This repository contains a large-scale, multilingual dataset designed for training and evaluating advanced language models and AI agents.
The dataset spans multiple domains, supports multiple languages, and is suitable for high-capacity models.
🔍 Key Characteristics
Dataset Size:
1B < n < 10B tokens
License:
Apache License 2.0 (Apache-2.0)
Task Category:
Text Generation
Languages Supported:
English (en)
Hebrew… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/WebText-4.tiny-webtext
Tiny WebText
The Tiny WebText dataset is designed to help models learn about perception on web text while neutralizing the bias of the source text using critical thinking methods. By providing a rich and diverse set of texts, I aim to improve the ability of models to understand and analyze information in a more objective and unbiased manner.
This dataset can be used to train and evaluate natural language processing and machine learning models, with the goal of improving their… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-webtext.WebText-3
WebText-3 Corpus
WebText-3 is a large-scale, diverse text corpus collected from publicly available web pages. It contains cleaned and normalized sentences suitable for natural language processing (NLP), machine learning, and AI training.
Dataset Overview
Format: Plain text (.txt), one sentence per line
Approximate Size: 200,000+ sentences
Languages: Primarily English, with occasional Hebrew content
Source Types: Wikipedia articles, technology news sites, blogs… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/WebText-3.Minecraft-Webtext
Minecraft Text Corpus Dataset
This dataset is a large-scale text corpus collected from various Minecraft-related online sources.It contains descriptive, community-driven, and encyclopedic content covering the Minecraft universe, including its mechanics, updates, characters, community discussions, and cultural impact.
Contents
The dataset is stored in plain text format (minecraft_corpus.txt) and consists of cleaned and segmented lines of natural language text.The text has… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/Minecraft-Webtext.WebText-2
Orion-Spark-2 Dataset
Overview
The Orion-Spark-2 Dataset is a text corpus curated for training the Orion-Spark-2 transformer language model. It consists of a diverse collection of sentences extracted from multiple sources including Wikipedia articles, technology news sites, developer resources, and other open-access web pages. The dataset is designed to provide broad coverage of general knowledge, programming topics, artificial intelligence, space, popular culture, and… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/WebText-2.Reddit-WebText
💬 MCGPT-1: Massive Reddit Interaction Dataset
This dataset contains high-quality conversational data extracted and filtered from Minecraft-related discussions on Reddit. It is a key component in giving MCGPT-1 its human-like personality.
📊 Dataset Statistics
Total Tokens: 20,360 🪙
Total Lines: 36 (High-density, long-form discussions) 📝
Focus: Community discussions, player advice, and Minecraft-specific social interaction.
🎯 Purpose
While the Mega Dataset… See the full description on the dataset page: https://huggingface.co/datasets/TopAI-1/Reddit-WebText.webtext-super-tiny
