datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indian-legal-documents
Indian Legal Documents
Open Indian statutory and legal-document data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.SHP
🚢 Stanford Human Preferences Dataset (SHP)
If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022).
Summary
SHP is a dataset of 385K collective human preferences over responses to questions/instructions in 18 different subject areas, from cooking to legal advice.
The preferences are meant to reflect the helpfulness of one response over another, and are intended to be used for training… See the full description on the dataset page: https://huggingface.co/datasets/Kanika0110/SHP.Athar-Datasets
🕌 Athar Islamic QA Datasets
18.7M passages from classical Islamic books spanning 1,400 years of scholarship
A comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, and more — sourced from the Shamela library and enriched with scholarly metadata for RAG-based Islamic QA systems.
Based on the Fanar-Sadiq Architecture for grounded, citation-backed Islamic question answering.
📊 Dataset Summary
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/Athar-Datasets.Athar-RAG-Hub
Athar RAG Hub 🕌
Collection
Chunks
seerah
5,852
nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing
nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 2,418 サンプル
validation: 302 サンプル
test: 303 サンプル
総サンプル数: 3,023
ソース
生成元: ./datasets/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing/
サンプルデータ
{
"instruction": "次の漢字の訓読み(くんよみ)をひらがなで答えてください。",
"input": "「究」の訓読みは?",
"think": "この漢字は「究」です。 小学3年生で習う漢字です。 意味は「research」などです。 訓読み(くんよみ)は日本語の読み方です。 この漢字の訓読みは「きわ」です。"… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing.nihongo-dojo-grades1-2-3-kanji_reading
nihongo-dojo-grades1-2-3-kanji_reading
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 846 サンプル
validation: 105 サンプル
test: 107 サンプル
総サンプル数: 1,058
ソース
生成元: ./datasets/nihongo-dojo-grades1-2-3-kanji_reading/
サンプルデータ
{
"instruction": "次の漢字の音読み(おんよみ)をカタカナで答えてください。",
"input": "「代」の音読みは?",
"output": "タイ",
"thinking": "この漢字は「代」です。 小学3年生で習う漢字です。 意味は「substitute」などです。 音読み(おんよみ)は中国から伝わった読み方です。 この漢字の音読みは「タイ」です。",
"answer": "タイ"… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-grades1-2-3-kanji_reading.
