CoolFace
20 results

numa

numad /yuho-text-2014-2022 Dataset Card for Dataset Name このデータはEDINET閲覧(提出)サイトで公開されている2014~2022年に提出された有価証券報告書から特定の章を抜粋したデータです。 各レコードのurl列が出典となります。データ取得の都合上2014/06/14以降のデータになります。 Dataset Details Dataset Description データの内容は下記想定です 物理名 論理名 型 概要 必須 doc_id 文書ID str 有価証券報告書の単位で発行されるID 〇 edinet_code EDINETコード str EDINET内での企業単位に採番されるID 〇 company_name 企業名 str 企業名 〇 document_name 文書タイトル str 有価証券報告書のタイトル 〇 sec_code 証券コード str 証券コード × period_start 期開始日 date(yyyy-mm-dd)… See the full description on the dataset page: https://huggingface.co/datasets/numad/yuho-text-2014-2022.text100K<n<1M0 likes171 downloads2y agoHugging Facehumairmunirawn /sota-numAtabularn<1K0 likes124 downloads7mo agoHugging FaceNumanKaanKaratas /turkish-english-words Turkish-English Words Dataset Dataset Description A comprehensive Turkish-English word and phrase translation dataset containing 3,365,067 parallel entries covering a wide range of Turkish vocabulary — from simple root words to complex agglutinated forms, conjugations, and idiomatic expressions. Language: Turkish (tr) → English (en) Total entries: 3,365,067 Format: JSONL License: MIT Generated by: ChatGPT (OpenAI) Human review: None — translations are fully AI-generated… See the full description on the dataset page: https://huggingface.co/datasets/NumanKaanKaratas/turkish-english-words.texttranslation1M<n<10M7 likes77 downloads4mo agoHugging Facenumad /yuho-text-2023 Dataset Card for Dataset Name このデータはEDINET閲覧(提出)サイトで公開されている2023年に提出された有価証券報告書から特定の章を抜粋したデータです。 各レコードのurl列が出典となります。 Dataset Details Dataset Description データの内容は下記想定です 物理名 論理名 型 概要 必須 doc_id 文書ID str 有価証券報告書の単位で発行されるID 〇 edinet_code EDINETコード str EDINET内での企業単位に採番されるID 〇 company_name 企業名 str 企業名 〇 document_name 文書タイトル str 有価証券報告書のタイトル 〇 sec_code 証券コード str 証券コード × period_start 期開始日 date(yyyy-mm-dd) 報告対象期間の開始日 〇 period_end 期終了日… See the full description on the dataset page: https://huggingface.co/datasets/numad/yuho-text-2023.text10K<n<100K0 likes67 downloads2y agoHugging FaceNumanKaanKaratas /turkish-sentences Turkish Sentences Turkish Sentences is a clean, duplicate-free Turkish text corpus prepared for NLP and language-model training workflows. The dataset contains Turkish sentences and short lexical entries built around Turkish roots, word forms, homonyms, and morphology-rich vocabulary. Dataset Summary Language: Turkish (tr) Format: Parquet Split: train Rows: 1,978,236 Schema: one column, text Created: 2026-05-31T19:38:26+00:00 Duplicate status: deduplicated Text… See the full description on the dataset page: https://huggingface.co/datasets/NumanKaanKaratas/turkish-sentences.texttext-generation1M<n<10M0 likes51 downloads4mo agoHugging Facehyper88 /autotrain-data-numai2tabular100K<n<1M0 likes38 downloads4y agoHugging Face