CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tensorfeed /ai-ecosystem-daily TensorFeed AI Ecosystem Daily Daily snapshots of the AI ecosystem: news, model pricing, benchmarks, service status, GPU rental prices, MCP registry growth, LLM endpoint latency probes, agent traffic, and the AFTA adopter directory. Captured once per day from the public tensorfeed.ai API and committed to this repo as JSONL. Each daily snapshot lives in a YYYY-MM-DD/ subfolder with one JSONL file per feed plus a manifest.json summarizing what was captured. What's in… See the full description on the dataset page: https://huggingface.co/datasets/tensorfeed/ai-ecosystem-daily.tabulartext-classification100K<n<1M2 likes2.7k downloads8h agoHugging Face023nesdeniz /turkish-daily-dialogues-5k Turkish Daily Dialogues 5K Exactly 5,000 synthetic, multi-turn Turkish conversations covering ordinary daily-life situations. The corpus is designed as a small, auditable baseline for dialogue modelling, instruction-format experiments, augmentation research, and Turkish-language evaluation—not as a substitute for conversations written by real people. Provenance in one sentence: the Turkish source scenario library was drafted with AI assistance specifically for this project… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/turkish-daily-dialogues-5k.texttext-generation1K<n<10K2 likes505 downloads2mo agoHugging Face03Yigit-Karaman /open-jobs-daily Open Jobs Daily 🌍💼 Commercial vendors often charge upwards of $1,000/month for firehose access to global job market data. This dataset democratizes that access. The main creator of this dataset is Reddit user OminousLatinWord. For convenience, I converted the dataset to Parquet files and uploaded it to Hugging Face. Source Data & Attribution Creator: Created and originally open-sourced by Reddit user OminousLatinWord under a CC0 license. Source Release:… See the full description on the dataset page: https://huggingface.co/datasets/Yigit-Karaman/open-jobs-daily.tabulartext-generation1M<n<10M1 likes385 downloads11d agoHugging Face04anezatra /dailydialog DailyDialog - ShareGPT Processed Dataset Summary DailyDialog is a high-quality, multi-turn dialogue dataset containing human-written conversations that cover a wide variety of everyday topics.It is designed to support research in dialogue modeling, conversational AI, and emotion-aware interactions.The dataset emphasizes natural, contextually coherent exchanges that resemble real-world human dialogue, making it ideal for training AI systems that need to handle daily… See the full description on the dataset page: https://huggingface.co/datasets/anezatra/dailydialog.texttext-generation10K<n<100K0 likes145 downloads11mo agoHugging Face05WhissleAI /daily_dialog_meta Meta-LLM Dataset: Daily Dialog with Meta-Information Enhancement Dataset Overview This dataset contains 76,064 conversational examples from the Daily Dialog corpus enhanced with meta-information awareness. Each example includes three response types: original human responses, basic LLM responses, and meta-aware LLM responses that incorporate emotional and intentional context. Meta-Information Distribution Emotion Categories Emotion Count… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/daily_dialog_meta.texttext-generation10K<n<100K1 likes144 downloads1y agoHugging Face063nesdeniz /english-daily-dialogues-10k English Daily Dialogues 10K A general-purpose, open dataset of 10,000 synthetic multi-turn English conversations spanning ten everyday-life domains. Built as a clean NLP resource for dialogue modeling, response generation, intent understanding, and conversational evaluation. This is a general language resource — not a safety or security benchmark. Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec. It is the English companion to the Turkish Daily Dialogues… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/english-daily-dialogues-10k.tabulartext-generation10K<n<100K2 likes126 downloads1mo agoHugging Face07daichira /structured-hard-sft-4k Hard Synthetic Dataset for Structured Data Tasks (v1) This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types). Dataset Summary The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.texttext-generation1K<n<10K1 likes87 downloads8mo agoHugging Face08Lots-of-LoRAs /task1533_daily_dialog_formal_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1533_daily_dialog_formal_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1533_daily_dialog_formal_classification.texttext-generation1K<n<10K0 likes83 downloads2y agoHugging Face09Lots-of-LoRAs /task1534_daily_dialog_question_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1534_daily_dialog_question_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1534_daily_dialog_question_classification.texttext-generation1K<n<10K0 likes77 downloads2y agoHugging Face10safetyllm /dailyconversationsThis dataset is synthetically generated using ChatGPT 3.5 to contain two-person multi-turn daily conversations with a various of topics (e.g. travel, food, music, movie/TV, education, hobbies, family, sports, technology, books, etc.) Originally, this dataset is used to train QuicktypeGPT, which is a GPT model to assist auto complete conversations. Here is the full list of topics the conversation may cover. texttext-generation10K<n<100K5 likes69 downloads3y agoHugging Face11aarohanverma /simple-daily-conversations-cleaned Dataset Card This dataset contains a cleaned version of simple daily conversations. It comprises nearly 98K text snippets representing informal, everyday dialogue, curated and processed for various Natural Language Processing tasks. Uses Direct Use This dataset is ideal for: Training language models on informal, everyday conversational data. Research exploring linguistic patterns in casual conversation. Out-of-Scope Use The dataset may not… See the full description on the dataset page: https://huggingface.co/datasets/aarohanverma/simple-daily-conversations-cleaned.texttext-generation10K<n<100K1 likes61 downloads2y agoHugging Face12avihayamor /tripmatch-ai-daily-plan-alternatives TripMatch AI — Rich Daily Plan Alternatives This public academic dataset is the professor-assigned upgrade to TripMatch AI. Its main deliverable is a substantially richer alternative daily plan generated with Gemma 3 or Qwen 3 on a GPU. The original plan is retained only as a side-by-side reference and for the later LLM comparison task. It is not the new generated target. Quality contract Every alternative must: preserve the destination, exact duration… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/tripmatch-ai-daily-plan-alternatives.tabulartext-generation10K<n<100K0 likes61 downloads1mo agoHugging Face13schneiderkamplab /SDU-Daisy SDUs Daisy: A Benchmark for Danish Culture SDU DAISY is the first version of a dataset designed to evaluate large language models’ understanding of Danish culture, as defined by the official Danish Culture Canon (Kulturkanon, 2006) SDU Daisy Evaluations Model Bleu Score F1 Score Dataset version Prompt Template Version openai/gpt-oss-20b 0.062 0.112 1.0 1.0 openai/gpt-oss-120b 0.126 0.211… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/SDU-Daisy.textquestion-answeringn<1K0 likes55 downloads7mo agoHugging Face14csebuetnlp /dailydialogue_bnDailyDialogue (bengali) has been derived from the original English dataset.texttext-generation10K<n<100K5 likes53 downloads3y agoHugging Face15hermeschen1116 /daily_dialog_for_RGtexttext-generation10K<n<100K0 likes45 downloads2y agoHugging Face16daichi812 /babyloop-mm-corpus Datasheet: babyloop Multimodal Training Corpus (BabyLM 2026, Strict track) Datasheet format follows Gebru et al. (2021), "Datasheets for Datasets," abridged to the sections relevant for a training-corpus release. All counts below are measured, not estimated; provenance manifests are versioned in this repository under manifests/. Summary Total word budget 99,999,984 whitespace words (≤ 100M, BabyLM Strict rule) Text portion 49,999,988 words —… See the full description on the dataset page: https://huggingface.co/datasets/daichi812/babyloop-mm-corpus.texttext-generation1M<n<10M0 likes44 downloads1mo agoHugging Face17VoidOaz /turkish-daily-dialogues-5k Turkish Daily Dialogues 5K Exactly 5,000 synthetic, multi-turn Turkish conversations covering ordinary daily-life situations. The corpus is designed as a small, auditable baseline for dialogue modelling, instruction-format experiments, augmentation research, and Turkish-language evaluation—not as a substitute for conversations written by real people. Provenance in one sentence: the Turkish source scenario library was drafted with AI assistance specifically for this project… See the full description on the dataset page: https://huggingface.co/datasets/VoidOaz/turkish-daily-dialogues-5k.texttext-generation1K<n<10K0 likes43 downloads11d agoHugging Face18daios /compartmentalized-harm-v1-training-data Justice character-training corpus This is the admitted training corpus used for the paper-v1.0 character-training experiments. A situation author created visible cases without seeing the constitution. A separate embodiment author saw the first-person justice constitution and wrote case-specific responses. The trained models saw only the visible conversations and responses; they did not receive the constitution or hidden construction metadata. The corpus contains 1,495… See the full description on the dataset page: https://huggingface.co/datasets/daios/compartmentalized-harm-v1-training-data.texttext-generation1K<n<10K0 likes41 downloads22d agoHugging Face19daichira /structured-5k-mix-sft 5k Mixed Hard-Structured SFT Dataset (v1) This dataset contains 5,000 synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks. It aggregates 13 distinct conversion tasks with a specific focus on format diversity and structural complexity. Dataset Summary The dataset is distributed across five major formats with the following allocation: Target Format Count Share Task Types YAML 1,500 30%… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-5k-mix-sft.texttext-generation1K<n<10K0 likes39 downloads8mo agoHugging Face20daichira /structeval-t-sft-hq-yaml-cleaned StructEval-T SFT HQ YAML (Cleaned) このデータセットは、daichira/structeval-t-sft-hq-yaml をベースに、厳密なフォーマット検証とノイズ除去(クリーニング)を行ったものです。 StructEval-T等の構造化データ生成タスク(SFT向け)に最適化されています。 クリーニング統計情報 本データセットの構築時に、以下のクリーニング結果が得られました。 オリジナルレコード数: 2000 件 クリーニング後レコード数: 2000 件 除去されたCoTノイズ: 1528 件 ( Approach: ... Output: を物理的に切除 ) 削減された無駄な文字列の総量: 705728 文字 最終YAMLパース成功率: 100% データセット構築パイプライン(クリーニング手法) 不要テキストの物理的除去: Approach: ... Output: といった思考プロセスや、マークダウンのコードフェンス (```yaml) を正規表現で完全に削除しました。… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-yaml-cleaned.texttext-generation1K<n<10K0 likes39 downloads6mo agoHugging Face21ChamaraVishwajithRajapaksha /cnn-dailymail-sinhala-continuous-pretrain CNN DailyMail Sinhala Continuous Pretraining Dataset Dataset Description This dataset is designed for continuous pretraining of Sinhala Small Language Models (SLMs) and Large Language Models (LLMs). The dataset was created by processing the original Sinhala news articles from: CNN Daily Mail Sinhala Dataset The article_sinhala field from the original dataset was extracted, cleaned, and concatenated into larger continuous text blocks suitable for language model… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/cnn-dailymail-sinhala-continuous-pretrain.texttext-generationn<1K0 likes33 downloads5mo agoHugging Face22daipham31 /qwen3.5-2B-vi-query Vietnamese Medical Query Normalization / Expansion / Routing Pack (v3) 1162 synthetic ChatML examples for fine-tuning a small Vietnamese model (target: Qwen/Qwen3.5-2B, trained with Unsloth) to turn a raw, everyday Vietnamese medical query into structured JSON: normalized query, intent, entities, must-preserve tokens, lexical/semantic query variants, and a retrieval-routing hint, for a downstream medical RAG system. The model does not answer medical questions. It only normalizes… See the full description on the dataset page: https://huggingface.co/datasets/daipham31/qwen3.5-2B-vi-query.texttext-generation1K<n<10K0 likes33 downloads3d agoHugging Face23atrevidasadia /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes31 downloads1mo agoHugging Face24Kamisori-daijin /email-datasets-20k Dataset Summary There are 20,000 samples of emails. This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ). License Note This dataset is licensed under Apache 2.0. Please also refer to the Gemma Terms of Use and Prohibited Use Policy regarding the use of Gemma-generated content. texttext-generation10K<n<100K3 likes28 downloads6mo agoHugging Face25Mikimi /ru-wikipedia-100k-full-text-daily-stats-10-years 📚 Russian Wikipedia Top 100K: Full Text with Daily Pageviews **Крупнейший открытый датасет русскоязычной Википедии с полными текстами статей и ежедневной статистикой просмотров за 10 лет. МГУ, ОТиПЛ, 2025** 📖 Описание Этот датасет содержит 99,348 самых популярных статей русскоязычной Википедии, отобранных по совокупному количеству просмотров за последние 10 лет. Для каждой статьи собраны полный текст с сохранением структуры, метаданные и детальная ежедневная… See the full description on the dataset page: https://huggingface.co/datasets/Mikimi/ru-wikipedia-100k-full-text-daily-stats-10-years.tabulartext-generation10K<n<100K1 likes27 downloads9mo agoHugging Face26daichira /structeval-t-sft-hq-yaml StructEval-T SFT - High Quality YAML This dataset is a highly refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated YAML transformations. Key Features Total Samples: 2,000 Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid YAML without errors are included. Goal: To maximize single-format fine-tuning performance or to be… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-yaml.texttext-generation1K<n<10K0 likes27 downloads7mo agoHugging Face27oliverkinch /da-instruct-dynaword-hq da-instruct-dynaword-hq Danish instruction fine-tuning dataset generated via backtranslation from danish-foundation-models/danish-dynaword, filtered to high-quality samples using danish-foundation-models/dynaword-annotations. All 40 DynaWord subsets are included — both contemporary and historical Danish. See oliverkinch/da-instruct-dynaword-contemporary-hq for a version restricted to contemporary Danish sources. Dataset description Each row is a (prompt, target) pair… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/da-instruct-dynaword-hq.texttext-generation10K<n<100K0 likes27 downloads4mo agoHugging Face28minnade /chat-daily MinnadeChat データセット (毎日更新) みんなで作る指示データセット の投稿のデータセットです。このデータセットは毎日12時に更新されます。 🤗 HuggingFace datasets から使う 最新版 を取得したい場合: from datasets import load_dataset ds = load_dataset("minnade/chat-daily", split="train") print(ds) #Dataset({ # features: ['id', 'parent_id', 'role', 'body', 'category_id', 'tags', 'is_synthetic', 'is_deleted', 'knowledge_cut_off', 'created_at', 'review', 'review_count', 'flag', 'flag_count'], # num_rows: 196 #}) 日付を指定して取得したい場合: from datasets… See the full description on the dataset page: https://huggingface.co/datasets/minnade/chat-daily.tabulartext-generation1K<n<10K10 likes25 downloads2y agoHugging Face29daichira /structeval-t-sft-v2-toml StructEval-T SFT v2 - Full TOML This dataset is the full, refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated TOML transformations. Key Features Total Samples: 3,635 Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid TOML without errors are included. Source: This is a split from the unified structeval-t-sft-v2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-v2-toml.texttext-generation1K<n<10K0 likes25 downloads7mo agoHugging Face30deltakitsune /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes25 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.