datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
J-ResearchCorpus
J-ResearchCorpus
Update:
2024/3/16言語処理学会第30回年次大会(NLP2024)を含む、論文 1,343 本のデータを追加
2024/2/25言語処理学会誌「自然言語処理」のうち CC-BY-4.0 で公開されている論文 360 本のデータを追加
概要
CC-BY-* ライセンスで公開されている日本語論文や学会誌等から抜粋した高品質なテキストのデータセットです。言語モデルの事前学習や RAG 等でご活用下さい。
今後も CC-BY-* ライセンスの日本語論文があれば追加する予定です。
データ説明
filename : 該当データのファイル名
text : 日本語論文から抽出したテキストデータ
category : データソース
license : ライセンス
credit : クレジット
データソース・ライセンス
テキスト総文字数 : 約 3,900 万文字
data source
num records
license
note… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/J-ResearchCorpus.ming-qing-wenji-corpus
Ming-Qing Literary Collections Corpus / 明清別集語料庫 / 명청별집어료고
Dataset Description / 數據集說明 / 데이터셋 설명
Summary / 概要 / 요약
English: The Ming-Qing Literary Collections Corpus is a structured digital corpus of 472 literary collections (bieji 別集) from the Ming (明, 1368–1644) and Qing (清, 1644–1912) dynasties. The texts are sourced from the Siku Quanshu (四庫全書) tradition and include poetry, prose, memorials, essays, letters, and other literary genres by scholars, officials… See the full description on the dataset page: https://huggingface.co/datasets/dibao-research/ming-qing-wenji-corpus.agent-memory-research-corpus
Agent Memory Research Corpus (AMRC)
A public, citable dataset for agent memory research and systems.
This dataset catalogues papers, systems, benchmarks, and design patterns related to long-term memory in autonomous agents. It is intended to serve as a canonical reference corpus for researchers and practitioners building memory-augmented agents.
Dataset Summary
Field
Value
Repository
https://huggingface.co/datasets/trentdoney/agent-memory-research-corpus… See the full description on the dataset page: https://huggingface.co/datasets/trentdoney/agent-memory-research-corpus.wanli-dibao-corpus
Wanli Dibao Corpus / 萬曆邸鈔校訂語料庫
Dataset Description
Summary
English:
The Wanli Dibao Corpus is a structured, proofread digital corpus of the Wanli Dichao (萬曆邸鈔), a collection of manuscript copies of official gazettes (dibao 邸報) from the Wanli reign (1573–1620) of the Ming dynasty. The dibao system was the primary channel of official communication in imperial China, transmitting memorials, edicts, personnel appointments, and policy decisions from the capital to… See the full description on the dataset page: https://huggingface.co/datasets/dibao-research/wanli-dibao-corpus.maasai-translation-corpus
Maasai-English Translation Corpus
Parallel English↔Maasai translation pairs for low-resource MT, language preservation, and culturally grounded tooling.
Overview
Total pairs: 9,910
Splits: 8,434 train / 738 valid / 738 test
Directions: 4,955 en→mas and 4,955 mas→en
Quality tiers: 8,444 gold and 1,466 silver
Main sources: 8,444 Bible-derived pairs, 680 cultural manual pairs, 70 knowledge-driven cultural pairs, 132 public-domain Hollis proverb pairs, 504 public-domain… See the full description on the dataset page: https://huggingface.co/datasets/NorthernTribe-Research/maasai-translation-corpus.odia-german-parallel-corpus-research
Dataset Summary
This dataset is a high-quality, parallel corpus for Odia (Oriya) to German and German to Odia machine translation. It focuses on the news domain, specifically covering National, International, Sports, Trade, and Science & Technology topics.
The dataset contains 3,676 unique parallel sentence pairs, curated through a hybrid approach combining automated web scraping, manual human translation (Gold Standard), and human-corrected machine translation (Silver… See the full description on the dataset page: https://huggingface.co/datasets/abhinandansamal/odia-german-parallel-corpus-research.
