datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Athar-Embeddingsknights-and-knaves
📘 knights-and-knaves Dataset [Project Page]
The knights-and-knaves dataset serves as a logical reasoning benchmark to evaluate the reasoning capabilities of LLMs.
🚀🚀 Check out the perturbed knights-and-knaves dataset to evaluate the memorization of LLMs in reasoning.
Loading the dataset
To load the dataset:
from datasets import load_dataset
data_subject = load_dataset('K-and-K/knights-and-knaves','test',split="2ppl")
Available subset: test, train.
Available… See the full description on the dataset page: https://huggingface.co/datasets/K-and-K/knights-and-knaves.wizardlm8x22b-logical-math-coding-sft_additional
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
wizardlm8x22b-logical-math-coding-sft
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
joyo-kanji-yomi-benchmark-parakeet
日本語 | English
常用漢字読みベンチマーク Parakeet Edition (JKYB-Parakeet)
常用漢字読みベンチマーク Parakeet Edition(JKYB-Parakeet)は、G2Pモデルや形態素解析器、TTSシステムが日本語の文章中の漢字を正しく読めているかを評価するためのベンチマークです。評価に用いるデータセットと評価ツールから構成されます。
このページでは、JKYB-Parakeetのデータセットを公開しています。評価ツールはGitHubで公開しています。
このデータセットは、SB Intuitionsによるデータセットsbintuitions/joyo-kanji-yomi-benchmarkをもとに、Parakeet株式会社が内容の検証を行い、誤りの修正、表記の統一、およびデータの追加等を独自に行ったものです。
概要… See the full description on the dataset page: https://huggingface.co/datasets/Parakeet-Inc/joyo-kanji-yomi-benchmark-parakeet.japanese-corpus-categorized
日本語コーパス
mc4-jaなどのwebコーパスをクリーニング後、教師なし学習モデルでテキストを約1万件にクラスタリングしたコーパスです。
著作権法で認められた情報解析目的で使用できます。
一部のファイルしかparquet化されていないので、ご注意ください。ファイルリストはoutフォルダ内にあります
git lfsなどでダウンロードください。
CommonCrawl-RAG-QA-Calm3-22b-chat
自動生成テキスト
データソースから、OpenCalm3-22bを使ってクリーニング・再生成したテキストです。
Common Crawlをもとに生成しています。 Common Crawl terms of useに従ってご利用ください。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
データ
jsonlファイルが数十GB程度あります
datasetsライブラリからでは、はじめの数GB程度しか読み込めない可能性があります。git lfsなどでダウンロードする必要がありそうです。
Baidu_Tieba_KangYaBeiGuo说明
随机爬取的百度贴吧抗压背锅吧的内容,10万条左右,不包含视频和图片,比较适合用于风格微调(大概)(心虚)。 数据遵循ChatGLM4使用的格式(有需要别的格式请自己调整QWQ)。 清洗的不是很干净,所以把没有清洗的数据也发上来了(QWQ)。
original.json是爬取后未经清洗的数据
Description
This dataset consists of roughly 100,000 samples randomly scraped from the "Kang Ya Bei Guo" bar on Baidu Tieba. It does not contain videos or images and is generally suitable for style fine-tuning (probably... kind of... maybe 👀). The data follows the format used by ChatGLM4 (please adjust to other formats if needed, QWQ). Note: The… See the full description on the dataset page: https://huggingface.co/datasets/Orphanage/Baidu_Tieba_KangYaBeiGuo.joyo-kanji-yomi-benchmark
Joyo Kanji Yomi Benchmark
A kanji-level pronunciation evaluation benchmark for Japanese TTS, covering all 2,136 Joyo kanji and their 4,378 readings with 13,095 native-speaker-verified test sentences.
Dataset Description
Each sample targets a specific kanji-reading pair. The sentence context is designed so that only the target reading is valid. All sentences and annotations have been verified by 35 native Japanese speakers through a three-stage review process.… See the full description on the dataset page: https://huggingface.co/datasets/sbintuitions/joyo-kanji-yomi-benchmark.SyntheticTextCC
自動生成テキスト
データソースから、Phi-3を使ってクリーニング・再生成したテキストです。
Common Crawlをもとに生成しています。 Common Crawl terms of useに従ってご利用ください。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
Athar-Datasets
🕌 Athar Islamic QA Datasets
18.7M passages from classical Islamic books spanning 1,400 years of scholarship
A comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, and more — sourced from the Shamela library and enriched with scholarly metadata for RAG-based Islamic QA systems.
Based on the Fanar-Sadiq Architecture for grounded, citation-backed Islamic question answering.
📊 Dataset Summary
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/Athar-Datasets.0717-calm3-22b-random-genre-inst-sft-tsub
自動生成Q&A
ランダムなジャンルについて、OpenCalm3-22bで生成したQ&Aです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
データ
jsonlファイルが数十GB程度あります
datasetsライブラリからでは、はじめの数GB程度しか読み込めない可能性があります。git lfsなどでダウンロードする必要がありそうです。
クリーニングはしていません。おかしなinstructionが一定数、含まれます
beyond_accept_or_deny
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
kana-kanji-pairs
kana-kanji-pairs
Japanese kana-to-kanji conversion candidate dataset.
Overview
Metric
Value
Total pairs
1,124,675
File size
~112MB
Format
JSONL
Candidate Distribution
Candidates
Entries
%
n>=2
363,708
32.3%
n>=5
40,929
3.6%
n>=10
9,401
0.8%
n>=20
2,448
0.2%
n>=100
34
<0.1%
max
259
-
Data Sources
Source
Entries
Description
mozc
753,628
Google mozc dictionary
jmdict
221,228
JMdict… See the full description on the dataset page: https://huggingface.co/datasets/katsukiono/kana-kanji-pairs.face-to-face-Configlogical-wizardlm-7b
自動生成したテキスト
WizardLM2 7bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
bird-platinum
BIRD-Platinum 2.5k v1
Original 2,463-example BIRD-Platinum training-candidate dataset from the ReViSQL repository.
The JSON records contain question_id, db_id, question, evidence, SQL, and grading_method.
This upload preserves the source file unchanged.
GVIMJ
AI Agents in Chemical Research: GVIM - An Intelligent Research Assistant System 🧪🤖
English | 简体中文
An intelligent research assistant system designed specifically for chemical science, featuring fine-tuned language models and specialized chemistry capabilities.
🌟 Highlights
🧬 Core Features
Fine-tuned LLMs for chemistry
Molecular visualization
Literature retrieval & analysis
Multimodal capabilities
🚀 Key Benefits
Specialized for… See the full description on the dataset page: https://huggingface.co/datasets/KANGYONGMA/GVIMJ.MA-EgoQA
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
Project Page | Paper | GitHub
MA-EgoQA (Multi-Agent Egocentric Video Question Answering) is a benchmark designed to evaluate models on their ability to understand multiple long-horizon egocentric video streams simultaneously collected from embodied agents.
Built on the EgoLife dataset, it features 266 hours of multi-agent video where 6 people lived together for 7 days. The benchmark includes 1.7k… See the full description on the dataset page: https://huggingface.co/datasets/KangsanKim71/MA-EgoQA.PKU-SafeRLHF
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset]
Citation
If PKU-SafeRLHF has contributed to your work, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Kanika0110/PKU-SafeRLHF.0804calm3-logical-multiturn-pretrain
自動生成したテキスト
Calm3で自動生成したマルチターン会話のテキストです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
kana-kanji-context
kana-kanji-context
Japanese kana-to-kanji conversion dataset with context for disambiguation.
Overview
Metric
Value
Total entries
77,277,970
File size
~7.4GB
Format
JSONL
Data Format
{
"input": "神経 [---]かがく",
"output": ["科学"],
"count": 1
}
{
"input": "この [---]さいご",
"output": ["最後", "最期"],
"count": 2
}
Fields
Field
Description
input
Context + [---] + reading (hiragana)
output
Correct kanji candidates (max… See the full description on the dataset page: https://huggingface.co/datasets/katsukiono/kana-kanji-context.ClimbMix-Ja-Initial64-Training-Data
ClimbMix-Ja Initial64 350M Artifacts
This repository is a public backup for the initial 64 ClimbMix-Ja candidate runs.
Candidate count: 64
Base model: nvidia/nemotron-climb-proxy-models 350M converted to a Megatron-LM TE-compatible checkpoint
Training corpus: KantaHayashiAI/ClimbLab-Ja clustered into cluster_01 ... cluster_20
Sequence length: 1024
Train iterations per candidate: 6500
Global batch size: 304
Tokens per candidate: 2,023,424,000
Total trained tokens across… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbMix-Ja-Initial64-Training-Data.historykanitakorn-v23-thaiexam-clean-20260614
Qwen v23 ThaiExam Clean Mix
Audited fallback mix for Kanitakorn. It avoids v20/v21 replay, avoids v13+v17 double replay, uses v1 repair once, includes all normalized worker v2, adds clean worker v3, and keeps small IF/math/identity retainers.
Validation
Records: 8,214
inspect_generated_jsonl: 3,214 source rows valid, 0 invalid, 0 duplicate prompts
inspect_sft_mix: 0 role errors, 0 empty errors
Contamination scan: 0 issues with 8,608 benchmark texts loaded… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-v23-thaiexam-clean-20260614.kansaiben
関西弁会話データセット (Kansai-ben Dialogue Dataset)
関西弁(大阪弁)で応答するAIのファインチューニング用instruction/outputデータセットです。
データセット概要
標準語の入力に対して、関西弁で自然に返答するサンプルを収録しています。LLMを関西弁キャラクターとしてファインチューニングすることを主な用途として想定しています。
データ構造
カラム名
型
説明
id
int
レコードのID
instruction
string
標準語の入力文
output
string
関西弁の応答文
サンプル
id
instruction
output
1
こんにちは!
まいど!元気しとるか?
9
なぜ空は青いの?
ええ質問やな。太陽の光が空気の粒にぶつかって、青い光がよう散らばるからやで。
50
関西弁を教えて.
「おおきに」は「ありがとう」、「ほんま」は「本当に」や。覚えときや。… See the full description on the dataset page: https://huggingface.co/datasets/shirochange/kansaiben.small-llm-blind-spots
Small LLM Blind Spots Dataset
A curated dataset of failure modes in small language models (0.6B–8B parameters), evaluated on the Qwen3 instruct model family.
GitHub (full code): github.com/kanak8278/small-llm-blind-spots
Model Tested
Qwen3 (Alibaba, 2025) — a recent open-weight model family available on HuggingFace:
Qwen/Qwen3-0.6B (0.6B params)
Qwen/Qwen3-1.7B (1.7B params)
Qwen/Qwen3-4B (4B params)
Qwen/Qwen3-8B (8B params)
These are base models with instruct-tuned… See the full description on the dataset page: https://huggingface.co/datasets/kanak8278/small-llm-blind-spots.AutoMultiTurnByCalm3-22B
自動生成のマルチターンデータセット
オープンなデータソースから、Calm3-22bを使ってQ&Aを自動生成したものです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
データソース
はじめの質問(q1)を、種々のデータソースから収集しました。その後のやりとりはすべて、Calmが生成しました。質問文については、元データのライセンスに準拠します。
oasst2-33k-ja
apache 2.0
databricks-dolly-15k-ja
cc-by-sa-3.0
minnade
CC0
cyberagent/chatbot-arena-ja-calm2-7b-chat-experimental
cc-by-4.0
perturbed-knights-and-knaves
📘 perturbed-knights-and-knaves Dataset [Project Page]
The perturbed-knights-and-knaves dataset evaluates the consistency of LLMs' logical reasoning ability under various perturbations.
🚀🚀 Check out the clean version of the dataset at [knights-and-knaves].
Loading the dataset
To load the dataset:
from datasets import load_dataset
data_subject = datasets.load_dataset('K-and-K/perturbed-knights-and-knaves', data_files="{subset}/{perturbation}/{subject}.jsonl")… See the full description on the dataset page: https://huggingface.co/datasets/K-and-K/perturbed-knights-and-knaves.kanalizer-dataset
kanalizer
英単語から読みを推測するライブラリ、kanalizerのデータセット置き場。データセットの作成に用いたコードはGitHubのVOICEVOX/kanalizer、学習済みモデルはVOICEVOX/kanalizer-modelを参照してください。
