CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Kandil7 /Athar-Embeddingstabular1M<n<10M0 likes5.8k downloads5mo agoHugging Face02K-and-K /knights-and-knaves 📘 knights-and-knaves Dataset [Project Page] The knights-and-knaves dataset serves as a logical reasoning benchmark to evaluate the reasoning capabilities of LLMs. 🚀🚀 Check out the perturbed knights-and-knaves dataset to evaluate the memorization of LLMs in reasoning. Loading the dataset To load the dataset: from datasets import load_dataset data_subject = load_dataset('K-and-K/knights-and-knaves','test',split="2ppl") Available subset: test, train. Available… See the full description on the dataset page: https://huggingface.co/datasets/K-and-K/knights-and-knaves.textquestion-answering1K<n<10K38 likes1.4k downloads2y agoHugging Face03kanhatakeyama /wizardlm8x22b-logical-math-coding-sft_additional 自動生成したテキスト WizardLM 8x22bで生成した論理・数学・コード系のデータです。 一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。 text100K<n<1M0 likes611 downloads2y agoHugging Face04kanhatakeyama /wizardlm8x22b-logical-math-coding-sft 自動生成したテキスト WizardLM 8x22bで生成した論理・数学・コード系のデータです。 一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。 text100K<n<1M4 likes606 downloads2y agoHugging Face05Parakeet-Inc /joyo-kanji-yomi-benchmark-parakeet 日本語 | English 常用漢字読みベンチマーク Parakeet Edition (JKYB-Parakeet) 常用漢字読みベンチマーク Parakeet Edition(JKYB-Parakeet)は、G2Pモデルや形態素解析器、TTSシステムが日本語の文章中の漢字を正しく読めているかを評価するためのベンチマークです。評価に用いるデータセットと評価ツールから構成されます。 このページでは、JKYB-Parakeetのデータセットを公開しています。評価ツールはGitHubで公開しています。 このデータセットは、SB Intuitionsによるデータセットsbintuitions/joyo-kanji-yomi-benchmarkをもとに、Parakeet株式会社が内容の検証を行い、誤りの修正、表記の統一、およびデータの追加等を独自に行ったものです。 概要… See the full description on the dataset page: https://huggingface.co/datasets/Parakeet-Inc/joyo-kanji-yomi-benchmark-parakeet.texttext-to-speech10K<n<100K5 likes376 downloads1mo agoHugging Face06kanhatakeyama /japanese-corpus-categorized 日本語コーパス mc4-jaなどのwebコーパスをクリーニング後、教師なし学習モデルでテキストを約1万件にクラスタリングしたコーパスです。 著作権法で認められた情報解析目的で使用できます。 一部のファイルしかparquet化されていないので、ご注意ください。ファイルリストはoutフォルダ内にあります git lfsなどでダウンロードください。 text100M<n<1B3 likes279 downloads2y agoHugging Face07kanhatakeyama /CommonCrawl-RAG-QA-Calm3-22b-chat 自動生成テキスト データソースから、OpenCalm3-22bを使ってクリーニング・再生成したテキストです。 Common Crawlをもとに生成しています。 Common Crawl terms of useに従ってご利用ください。 一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。 データ jsonlファイルが数十GB程度あります datasetsライブラリからでは、はじめの数GB程度しか読み込めない可能性があります。git lfsなどでダウンロードする必要がありそうです。 text10M<n<100M4 likes196 downloads2y agoHugging Face08Orphanage /Baidu_Tieba_KangYaBeiGuo说明 随机爬取的百度贴吧抗压背锅吧的内容,10万条左右,不包含视频和图片,比较适合用于风格微调(大概)(心虚)。 数据遵循ChatGLM4使用的格式(有需要别的格式请自己调整QWQ)。 清洗的不是很干净,所以把没有清洗的数据也发上来了(QWQ)。 original.json是爬取后未经清洗的数据 Description This dataset consists of roughly 100,000 samples randomly scraped from the "Kang Ya Bei Guo" bar on Baidu Tieba. It does not contain videos or images and is generally suitable for style fine-tuning (probably... kind of... maybe 👀). The data follows the format used by ChatGLM4 (please adjust to other formats if needed, QWQ). Note: The… See the full description on the dataset page: https://huggingface.co/datasets/Orphanage/Baidu_Tieba_KangYaBeiGuo.text1K<n<10K5 likes194 downloads1y agoHugging Face09sbintuitions /joyo-kanji-yomi-benchmark Joyo Kanji Yomi Benchmark A kanji-level pronunciation evaluation benchmark for Japanese TTS, covering all 2,136 Joyo kanji and their 4,378 readings with 13,095 native-speaker-verified test sentences. Dataset Description Each sample targets a specific kanji-reading pair. The sentence context is designed so that only the target reading is valid. All sentences and annotations have been verified by 35 native Japanese speakers through a three-stage review process.… See the full description on the dataset page: https://huggingface.co/datasets/sbintuitions/joyo-kanji-yomi-benchmark.texttext-to-speech10K<n<100K10 likes156 downloads3mo agoHugging Face10kanhatakeyama /SyntheticTextCC 自動生成テキスト データソースから、Phi-3を使ってクリーニング・再生成したテキストです。 Common Crawlをもとに生成しています。 Common Crawl terms of useに従ってご利用ください。 一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。 text1M<n<10M1 likes154 downloads2y agoHugging Face11Kandil7 /Athar-Datasets 🕌 Athar Islamic QA Datasets 18.7M passages from classical Islamic books spanning 1,400 years of scholarship A comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, and more — sourced from the Shamela library and enriched with scholarly metadata for RAG-based Islamic QA systems. Based on the Fanar-Sadiq Architecture for grounded, citation-backed Islamic question answering. 📊 Dataset Summary Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/Athar-Datasets.tabularquestion-answering10M<n<100M8 likes152 downloads5mo agoHugging Face12kanhatakeyama /0717-calm3-22b-random-genre-inst-sft-tsub 自動生成Q&A ランダムなジャンルについて、OpenCalm3-22bで生成したQ&Aです。 一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。 データ jsonlファイルが数十GB程度あります datasetsライブラリからでは、はじめの数GB程度しか読み込めない可能性があります。git lfsなどでダウンロードする必要がありそうです。 クリーニングはしていません。おかしなinstructionが一定数、含まれます text1M<n<10M0 likes149 downloads2y agoHugging Face13Kanubalad /beyond_accept_or_deny Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details tabular1M<n<10M0 likes144 downloads5mo agoHugging Face14katsukiono /kana-kanji-pairs kana-kanji-pairs Japanese kana-to-kanji conversion candidate dataset. Overview Metric Value Total pairs 1,124,675 File size ~112MB Format JSONL Candidate Distribution Candidates Entries % n>=2 363,708 32.3% n>=5 40,929 3.6% n>=10 9,401 0.8% n>=20 2,448 0.2% n>=100 34 <0.1% max 259 - Data Sources Source Entries Description mozc 753,628 Google mozc dictionary jmdict 221,228 JMdict… See the full description on the dataset page: https://huggingface.co/datasets/katsukiono/kana-kanji-pairs.texttext-generation1M<n<10M1 likes142 downloads9mo agoHugging Face15KANE-202666 /face-to-face-Configtextn<1K0 likes142 downloads3mo agoHugging Face16kanhatakeyama /logical-wizardlm-7b 自動生成したテキスト WizardLM2 7bで生成した論理・数学・コード系のデータです。 一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。 text10M<n<100M1 likes133 downloads2y agoHugging Face17uiuc-kang-lab /bird-platinum BIRD-Platinum 2.5k v1 Original 2,463-example BIRD-Platinum training-candidate dataset from the ReViSQL repository. The JSON records contain question_id, db_id, question, evidence, SQL, and grading_method. This upload preserves the source file unchanged. text1K<n<10K0 likes125 downloads28d agoHugging Face18KANGYONGMA /GVIMJ AI Agents in Chemical Research: GVIM - An Intelligent Research Assistant System 🧪🤖 English | 简体中文 An intelligent research assistant system designed specifically for chemical science, featuring fine-tuned language models and specialized chemistry capabilities. 🌟 Highlights 🧬 Core Features Fine-tuned LLMs for chemistry Molecular visualization Literature retrieval & analysis Multimodal capabilities 🚀 Key Benefits Specialized for… See the full description on the dataset page: https://huggingface.co/datasets/KANGYONGMA/GVIMJ.text1M<n<10M0 likes111 downloads2y agoHugging Face19KangsanKim71 /MA-EgoQA MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents Project Page | Paper | GitHub MA-EgoQA (Multi-Agent Egocentric Video Question Answering) is a benchmark designed to evaluate models on their ability to understand multiple long-horizon egocentric video streams simultaneously collected from embodied agents. Built on the EgoLife dataset, it features 266 hours of multi-agent video where 6 people lived together for 7 days. The benchmark includes 1.7k… See the full description on the dataset page: https://huggingface.co/datasets/KangsanKim71/MA-EgoQA.textvideo-text-to-text1K<n<10K4 likes105 downloads7mo agoHugging Face20Kanika0110 /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Kanika0110/PKU-SafeRLHF.tabulartext-generation100K<n<1M0 likes100 downloads6d agoHugging Face21kanhatakeyama /0804calm3-logical-multiturn-pretrain 自動生成したテキスト Calm3で自動生成したマルチターン会話のテキストです。 一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。 text1M<n<10M2 likes97 downloads2y agoHugging Face22katsukiono /kana-kanji-context kana-kanji-context Japanese kana-to-kanji conversion dataset with context for disambiguation. Overview Metric Value Total entries 77,277,970 File size ~7.4GB Format JSONL Data Format { "input": "神経 [---]かがく", "output": ["科学"], "count": 1 } { "input": "この [---]さいご", "output": ["最後", "最期"], "count": 2 } Fields Field Description input Context + [---] + reading (hiragana) output Correct kanji candidates (max… See the full description on the dataset page: https://huggingface.co/datasets/katsukiono/kana-kanji-context.texttext-generation100M<n<1B1 likes95 downloads9mo agoHugging Face23KantaHayashiAI /ClimbMix-Ja-Initial64-Training-Data ClimbMix-Ja Initial64 350M Artifacts This repository is a public backup for the initial 64 ClimbMix-Ja candidate runs. Candidate count: 64 Base model: nvidia/nemotron-climb-proxy-models 350M converted to a Megatron-LM TE-compatible checkpoint Training corpus: KantaHayashiAI/ClimbLab-Ja clustered into cluster_01 ... cluster_20 Sequence length: 1024 Train iterations per candidate: 6500 Global batch size: 304 Tokens per candidate: 2,023,424,000 Total trained tokens across… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbMix-Ja-Initial64-Training-Data.tabularn<1K0 likes93 downloads4mo agoHugging Face24KangarooLove /historytextn<1K0 likes92 downloads9mo agoHugging Face25Jnx03 /kanitakorn-v23-thaiexam-clean-20260614 Qwen v23 ThaiExam Clean Mix Audited fallback mix for Kanitakorn. It avoids v20/v21 replay, avoids v13+v17 double replay, uses v1 repair once, includes all normalized worker v2, adds clean worker v3, and keeps small IF/math/identity retainers. Validation Records: 8,214 inspect_generated_jsonl: 3,214 source rows valid, 0 invalid, 0 duplicate prompts inspect_sft_mix: 0 role errors, 0 empty errors Contamination scan: 0 issues with 8,608 benchmark texts loaded… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-v23-thaiexam-clean-20260614.text1K<n<10K0 likes89 downloads3mo agoHugging Face26shirochange /kansaiben 関西弁会話データセット (Kansai-ben Dialogue Dataset) 関西弁(大阪弁)で応答するAIのファインチューニング用instruction/outputデータセットです。 データセット概要 標準語の入力に対して、関西弁で自然に返答するサンプルを収録しています。LLMを関西弁キャラクターとしてファインチューニングすることを主な用途として想定しています。 データ構造 カラム名 型 説明 id int レコードのID instruction string 標準語の入力文 output string 関西弁の応答文 サンプル id instruction output 1 こんにちは! まいど!元気しとるか? 9 なぜ空は青いの? ええ質問やな。太陽の光が空気の粒にぶつかって、青い光がよう散らばるからやで。 50 関西弁を教えて. 「おおきに」は「ありがとう」、「ほんま」は「本当に」や。覚えときや。… See the full description on the dataset page: https://huggingface.co/datasets/shirochange/kansaiben.texttext-generation1K<n<10K4 likes88 downloads5mo agoHugging Face27kanak8278 /small-llm-blind-spots Small LLM Blind Spots Dataset A curated dataset of failure modes in small language models (0.6B–8B parameters), evaluated on the Qwen3 instruct model family. GitHub (full code): github.com/kanak8278/small-llm-blind-spots Model Tested Qwen3 (Alibaba, 2025) — a recent open-weight model family available on HuggingFace: Qwen/Qwen3-0.6B (0.6B params) Qwen/Qwen3-1.7B (1.7B params) Qwen/Qwen3-4B (4B params) Qwen/Qwen3-8B (8B params) These are base models with instruct-tuned… See the full description on the dataset page: https://huggingface.co/datasets/kanak8278/small-llm-blind-spots.texttext-generationn<1K1 likes79 downloads7mo agoHugging Face28kanhatakeyama /AutoMultiTurnByCalm3-22B 自動生成のマルチターンデータセット オープンなデータソースから、Calm3-22bを使ってQ&Aを自動生成したものです。 一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。 データソース はじめの質問(q1)を、種々のデータソースから収集しました。その後のやりとりはすべて、Calmが生成しました。質問文については、元データのライセンスに準拠します。 oasst2-33k-ja apache 2.0 databricks-dolly-15k-ja cc-by-sa-3.0 minnade CC0 cyberagent/chatbot-arena-ja-calm2-7b-chat-experimental cc-by-4.0 text10K<n<100K4 likes78 downloads2y agoHugging Face29K-and-K /perturbed-knights-and-knaves 📘 perturbed-knights-and-knaves Dataset [Project Page] The perturbed-knights-and-knaves dataset evaluates the consistency of LLMs' logical reasoning ability under various perturbations. 🚀🚀 Check out the clean version of the dataset at [knights-and-knaves]. Loading the dataset To load the dataset: from datasets import load_dataset data_subject = datasets.load_dataset('K-and-K/perturbed-knights-and-knaves', data_files="{subset}/{perturbation}/{subject}.jsonl")… See the full description on the dataset page: https://huggingface.co/datasets/K-and-K/perturbed-knights-and-knaves.textquestion-answering10K<n<100K10 likes76 downloads2y agoHugging Face30VOICEVOX /kanalizer-dataset kanalizer 英単語から読みを推測するライブラリ、kanalizerのデータセット置き場。データセットの作成に用いたコードはGitHubのVOICEVOX/kanalizer、学習済みモデルはVOICEVOX/kanalizer-modelを参照してください。 text100K<n<1M1 likes74 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.