CoolFace
7 results

research-corpus

IRIIS-RESEARCH /Nepali-Text-Corpus Nepali Text Corpus Overview Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.texttext-generation1M<n<10M10 likes969 downloads1y agoHugging Facejlietz93 /Phase-Calculus-Research-Corpus Phase Calculus Research Corpus The Phase Calculus Research Corpus is the collected working archive of the Phase Calculus research program: papers, mathematical derivations, formal proofs, executable experiments, native runtimes, computational notebooks, validation packages, and the related Cortex, Xi, and VDM research in which the calculus has been developed and tested. Rather than presenting only finished papers, the corpus preserves the research object around the theory. A… See the full description on the dataset page: https://huggingface.co/datasets/jlietz93/Phase-Calculus-Research-Corpus.0 likes434 downloads1mo agoHugging Facekunishou /J-ResearchCorpus J-ResearchCorpus Update: 2024/3/16言語処理学会第30回年次大会(NLP2024)を含む、論文 1,343 本のデータを追加 2024/2/25言語処理学会誌「自然言語処理」のうち CC-BY-4.0 で公開されている論文 360 本のデータを追加 概要 CC-BY-* ライセンスで公開されている日本語論文や学会誌等から抜粋した高品質なテキストのデータセットです。言語モデルの事前学習や RAG 等でご活用下さい。 今後も CC-BY-* ライセンスの日本語論文があれば追加する予定です。 データ説明 filename : 該当データのファイル名 text : 日本語論文から抽出したテキストデータ category : データソース license : ライセンス credit : クレジット データソース・ライセンス テキスト総文字数 : 約 3,900 万文字 data source num records license note… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/J-ResearchCorpus.text1K<n<10K32 likes211 downloads3y agoHugging Facefliarbi /urban-heat-research-corpus Urban Heat Research Corpus (UHRC) v1.0 What does the world study, invent and report about urban heat? This dataset puts three records of the same problem side by side: 20,422 research papers on urban heat islands and extreme heat in cities (1990–2025) with the claims their abstracts make, 106,458 news articles about heat (2021–2025) coded for 51 subjects, framings and terms, and 4,123 patent families for heat-mitigation technologies (2006–2024) — plus supplementary tables on the… See the full description on the dataset page: https://huggingface.co/datasets/fliarbi/urban-heat-research-corpus.imagetext-classification100K<n<1M0 likes161 downloads6d agoHugging FaceDeep-Research-Team /Pre-Training-Persian-Corpus-Raw-Texts-DatasetDocument Version: 2.0.0 | Last Updated: 02/13/2026 text10M<n<100M0 likes99 downloads7mo agoHugging FaceNorthernTribe-Research /math-conjecture-training-corpus Math Conjecture Training Corpus (v1) Repository: NorthernTribe-Research/math-conjecture-training-corpus Summary Merged training dataset for unsolved-conjecture-oriented math AI training. Included families: conjecture_core competition structured_reasoning formal_proof Rows Per Split train: 458664 validation: 4920 test: 9765 Rows Per Family conjecture_core: 121 formal_proof: 186239 competition: 63093 structured_reasoning: 223896 Policy… See the full description on the dataset page: https://huggingface.co/datasets/NorthernTribe-Research/math-conjecture-training-corpus.text100K<n<1M0 likes49 downloads6mo agoHugging Face