datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nepali-Text-Corpus
Nepali Text Corpus
Overview
Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the
Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a
diverse range of text types, including news articles, blogs, and more, making it an invaluable
resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP)
and computational linguistics.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.LatamGPT-Corpus-1.0-research
LatamGPT-Corpus-1.0-research
🌐 Language versions: English | Español | Português
🔗 Project links: Official LatamGPT website | Corpus dashboard
🔓 Open portion: the openly released part of this corpus is published separately as LatamGPT-Corpus-1.0.
⚠️ Controlled Access Level (Blue – Research)
This repository is not openly available. It holds data destined
exclusively for scientific and academic research, and is managed as a
restricted repository under… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/LatamGPT-Corpus-1.0-research.ming-qing-wenji-corpus
Ming-Qing Literary Collections Corpus / 明清別集語料庫 / 명청별집어료고
Dataset Description / 數據集說明 / 데이터셋 설명
Summary / 概要 / 요약
English: The Ming-Qing Literary Collections Corpus is a structured digital corpus of 472 literary collections (bieji 別集) from the Ming (明, 1368–1644) and Qing (清, 1644–1912) dynasties. The texts are sourced from the Siku Quanshu (四庫全書) tradition and include poetry, prose, memorials, essays, letters, and other literary genres by scholars, officials… See the full description on the dataset page: https://huggingface.co/datasets/dibao-research/ming-qing-wenji-corpus.bo-y-te-corpus-rawwanli-dibao-corpus
Wanli Dibao Corpus / 萬曆邸鈔校訂語料庫
Dataset Description
Summary
English:
The Wanli Dibao Corpus is a structured, proofread digital corpus of the Wanli Dichao (萬曆邸鈔), a collection of manuscript copies of official gazettes (dibao 邸報) from the Wanli reign (1573–1620) of the Ming dynasty. The dibao system was the primary channel of official communication in imperial China, transmitting memorials, edicts, personnel appointments, and policy decisions from the capital to… See the full description on the dataset page: https://huggingface.co/datasets/dibao-research/wanli-dibao-corpus.enoch-ai-research-corpus
Enoch AI Research Corpus
This dataset contains 393 AI-generated research artifacts produced by the Enoch agentic research system.
System repository: https://github.com/alias8818/enoch-agentic-research-system
Source corpus repository: https://github.com/alias8818/enoch-ai-research-corpus
Launch site: https://alias8818.github.io/enoch-agentic-research-system/
Current release correction
Older launch posts may mention 120 artifacts. The current public corpus indexes… See the full description on the dataset page: https://huggingface.co/datasets/aliasocracy/enoch-ai-research-corpus.
