datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yuho-text-2014-2022
Dataset Card for Dataset Name
このデータはEDINET閲覧(提出)サイトで公開されている2014~2022年に提出された有価証券報告書から特定の章を抜粋したデータです。
各レコードのurl列が出典となります。データ取得の都合上2014/06/14以降のデータになります。
Dataset Details
Dataset Description
データの内容は下記想定です
物理名
論理名
型
概要
必須
doc_id
文書ID
str
有価証券報告書の単位で発行されるID
〇
edinet_code
EDINETコード
str
EDINET内での企業単位に採番されるID
〇
company_name
企業名
str
企業名
〇
document_name
文書タイトル
str
有価証券報告書のタイトル
〇
sec_code
証券コード
str
証券コード
×
period_start
期開始日
date(yyyy-mm-dd)… See the full description on the dataset page: https://huggingface.co/datasets/numad/yuho-text-2014-2022.sota-numAturkish-english-words
Turkish-English Words Dataset
Dataset Description
A comprehensive Turkish-English word and phrase translation dataset containing 3,365,067 parallel entries covering a wide range of Turkish vocabulary — from simple root words to complex agglutinated forms, conjugations, and idiomatic expressions.
Language: Turkish (tr) → English (en)
Total entries: 3,365,067
Format: JSONL
License: MIT
Generated by: ChatGPT (OpenAI)
Human review: None — translations are fully AI-generated… See the full description on the dataset page: https://huggingface.co/datasets/NumanKaanKaratas/turkish-english-words.yuho-text-2023
Dataset Card for Dataset Name
このデータはEDINET閲覧(提出)サイトで公開されている2023年に提出された有価証券報告書から特定の章を抜粋したデータです。
各レコードのurl列が出典となります。
Dataset Details
Dataset Description
データの内容は下記想定です
物理名
論理名
型
概要
必須
doc_id
文書ID
str
有価証券報告書の単位で発行されるID
〇
edinet_code
EDINETコード
str
EDINET内での企業単位に採番されるID
〇
company_name
企業名
str
企業名
〇
document_name
文書タイトル
str
有価証券報告書のタイトル
〇
sec_code
証券コード
str
証券コード
×
period_start
期開始日
date(yyyy-mm-dd)
報告対象期間の開始日
〇
period_end
期終了日… See the full description on the dataset page: https://huggingface.co/datasets/numad/yuho-text-2023.turkish-sentences
Turkish Sentences
Turkish Sentences is a clean, duplicate-free Turkish text corpus prepared for NLP and language-model training workflows. The dataset contains Turkish sentences and short lexical entries built around Turkish roots, word forms, homonyms, and morphology-rich vocabulary.
Dataset Summary
Language: Turkish (tr)
Format: Parquet
Split: train
Rows: 1,978,236
Schema: one column, text
Created: 2026-05-31T19:38:26+00:00
Duplicate status: deduplicated
Text… See the full description on the dataset page: https://huggingface.co/datasets/NumanKaanKaratas/turkish-sentences.autotrain-data-numai2yuho-text-2024
Dataset Card for Dataset Name
このデータはEDINET閲覧(提出)サイトで公開されている2024年に提出された有価証券報告書から特定の章を抜粋したデータです。
各レコードのurl列が出典となります。
Dataset Details
Dataset Description
データの内容は下記想定です
物理名
論理名
型
概要
必須
doc_id
文書ID
str
有価証券報告書の単位で発行されるID
〇
edinet_code
EDINETコード
str
EDINET内での企業単位に採番されるID
〇
company_name
企業名
str
企業名
〇
document_name
文書タイトル
str
有価証券報告書のタイトル
〇
sec_code
証券コード
str
証券コード
×
period_start
期開始日
date(yyyy-mm-dd)
報告対象期間の開始日
〇
period_end
期終了日… See the full description on the dataset page: https://huggingface.co/datasets/numad/yuho-text-2024.autotrain-data-numaisynth-incorrect-versesEduAdaptSakuseiByoutou_Numajiri_Voice_Datarisk-financing-4th-numarkdown
Document OCR using NuMarkdown-8B-Thinking
This dataset contains markdown-formatted OCR results from images in andesco/risk-financing-4th-images using NuMarkdown-8B-Thinking.
Processing Details
Source Dataset: andesco/risk-financing-4th-images
Model: numind/NuMarkdown-8B-Thinking
Number of Samples: 436
Processing Time: 20.3 minutes
Processing Date: 2026-02-21 23:51 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch… See the full description on the dataset page: https://huggingface.co/datasets/andesco/risk-financing-4th-numarkdown.numad-yuho-text-cleanedpascal-stahl-numarkdown
Document OCR using NuMarkdown-8B-Thinking
This dataset contains markdown-formatted OCR results from images in ShaitanRa/PascalStahl using NuMarkdown-8B-Thinking.
Processing Details
Source Dataset: ShaitanRa/PascalStahl
Model: numind/NuMarkdown-8B-Thinking
Number of Samples: 5
Processing Time: 4.1 minutes
Processing Date: 2026-04-29 11:54 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 16
Max Model Length: 16… See the full description on the dataset page: https://huggingface.co/datasets/ShaitanRa/pascal-stahl-numarkdown.fixtest-numarkdown-ocr
Document OCR using NuMarkdown-8B-Thinking
This dataset contains markdown-formatted OCR results from images in davanstrien/ufo-ColPali using NuMarkdown-8B-Thinking.
Processing Details
Source Dataset: davanstrien/ufo-ColPali
Model: numind/NuMarkdown-8B-Thinking
Number of Samples: 2
Processing Time: 4.2 minutes
Processing Date: 2026-06-05 10:16 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 16
Max… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/fixtest-numarkdown-ocr.risk-assessment-treatment-numarkdown
Document OCR using NuMarkdown-8B-Thinking
This dataset contains markdown-formatted OCR results from images in andesco/risk-assessment-treatment-images using NuMarkdown-8B-Thinking.
Processing Details
Source Dataset: andesco/risk-assessment-treatment-images
Model: numind/NuMarkdown-8B-Thinking
Number of Samples: 378
Processing Time: 16.0 minutes
Processing Date: 2026-02-21 23:23 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split:… See the full description on the dataset page: https://huggingface.co/datasets/andesco/risk-assessment-treatment-numarkdown.sota-numa-playsota-numa-playrisk-management-3rd-numarkdown
Document OCR using NuMarkdown-8B-Thinking
This dataset contains markdown-formatted OCR results from images in andesco/risk-management-3rd-images using NuMarkdown-8B-Thinking.
Processing Details
Source Dataset: andesco/risk-management-3rd-images
Model: numind/NuMarkdown-8B-Thinking
Number of Samples: 340
Processing Time: 15.2 minutes
Processing Date: 2026-02-21 23:46 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch… See the full description on the dataset page: https://huggingface.co/datasets/andesco/risk-management-3rd-numarkdown.FindNum-NumAnsOnlyScientifcLLM
