datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing
nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 2,418 サンプル
validation: 302 サンプル
test: 303 サンプル
総サンプル数: 3,023
ソース
生成元: ./datasets/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing/
サンプルデータ
{
"instruction": "次の漢字の訓読み(くんよみ)をひらがなで答えてください。",
"input": "「究」の訓読みは?",
"think": "この漢字は「究」です。 小学3年生で習う漢字です。 意味は「research」などです。 訓読み(くんよみ)は日本語の読み方です。 この漢字の訓読みは「きわ」です。"… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing.Vietnamese_Reading_Comprehension_Dataset
Dataset Describe
This dataset is collected from internet sources, SQuAD dataset, wiki, etc. It has been translated into Vietnamese using "google translate" and word segmented using VnCoreNLP (https://github.com/vncorenlp/VnCoreNLP).
Data structure
Dataset includes the following columns:
question: Question related to the content of the text.
context: Paragraph of text.
answer: The answer to the question is based on the content of the text.
answer_start: The starting… See the full description on the dataset page: https://huggingface.co/datasets/ShynBui/Vietnamese_Reading_Comprehension_Dataset.nihongo-dojo-grades1-2-3-kanji_reading
nihongo-dojo-grades1-2-3-kanji_reading
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 846 サンプル
validation: 105 サンプル
test: 107 サンプル
総サンプル数: 1,058
ソース
生成元: ./datasets/nihongo-dojo-grades1-2-3-kanji_reading/
サンプルデータ
{
"instruction": "次の漢字の音読み(おんよみ)をカタカナで答えてください。",
"input": "「代」の音読みは?",
"output": "タイ",
"thinking": "この漢字は「代」です。 小学3年生で習う漢字です。 意味は「substitute」などです。 音読み(おんよみ)は中国から伝わった読み方です。 この漢字の音読みは「タイ」です。",
"answer": "タイ"… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-grades1-2-3-kanji_reading.databird-readingsmooth_reading
Smooth Reading
This dataset accompanies the paper Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Understanding.
license: mit
reading-steiner-data
Reading Steiner — Web Content Extraction Dataset
Training data for Reading Steiner, a web content extraction model that identifies relevant content blocks in web pages, filtering out navigation, ads, sidebars, and other boilerplate.
Overview
Train
Eval
Examples
51,697
1,417
Unique Domains
9,256+
—
Median Context Length
6,435 chars
—
Max Context Length
658K chars
—
Task Types
The dataset covers two extraction tasks:
1. Main… See the full description on the dataset page: https://huggingface.co/datasets/OmAlve/reading-steiner-data.
