datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-ecosystem-daily
TensorFeed AI Ecosystem Daily
Daily snapshots of the AI ecosystem: news, model pricing, benchmarks, service status, GPU rental prices, MCP registry growth, LLM endpoint latency probes, agent traffic, and the AFTA adopter directory. Captured once per day from the public tensorfeed.ai API and committed to this repo as JSONL.
Each daily snapshot lives in a YYYY-MM-DD/ subfolder with one JSONL file per feed plus a manifest.json summarizing what was captured.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/tensorfeed/ai-ecosystem-daily.turkish-daily-dialogues-5k
Turkish Daily Dialogues 5K
Exactly 5,000 synthetic, multi-turn Turkish conversations covering ordinary daily-life situations. The corpus is designed as a small, auditable baseline for dialogue modelling, instruction-format experiments, augmentation research, and Turkish-language evaluation—not as a substitute for conversations written by real people.
Provenance in one sentence: the Turkish source scenario library was drafted with AI assistance specifically for this project… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/turkish-daily-dialogues-5k.open-jobs-daily
Open Jobs Daily 🌍💼
Commercial vendors often charge upwards of $1,000/month for firehose access to global job market data. This dataset democratizes that access.
The main creator of this dataset is Reddit user OminousLatinWord. For convenience, I converted the dataset to Parquet files and uploaded it to Hugging Face.
Source Data & Attribution
Creator: Created and originally open-sourced by Reddit user OminousLatinWord under a CC0 license.
Source Release:… See the full description on the dataset page: https://huggingface.co/datasets/Yigit-Karaman/open-jobs-daily.dailydialog
DailyDialog - ShareGPT Processed
Dataset Summary
DailyDialog is a high-quality, multi-turn dialogue dataset containing human-written conversations that cover a wide variety of everyday topics.It is designed to support research in dialogue modeling, conversational AI, and emotion-aware interactions.The dataset emphasizes natural, contextually coherent exchanges that resemble real-world human dialogue, making it ideal for training AI systems that need to handle daily… See the full description on the dataset page: https://huggingface.co/datasets/anezatra/dailydialog.daily_dialog_meta
Meta-LLM Dataset: Daily Dialog with Meta-Information Enhancement
Dataset Overview
This dataset contains 76,064 conversational examples from the Daily Dialog corpus enhanced with meta-information awareness. Each example includes three response types: original human responses, basic LLM responses, and meta-aware LLM responses that incorporate emotional and intentional context.
Meta-Information Distribution
Emotion Categories
Emotion
Count… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/daily_dialog_meta.english-daily-dialogues-10k
English Daily Dialogues 10K
A general-purpose, open dataset of 10,000 synthetic multi-turn English conversations spanning ten everyday-life domains. Built as a clean NLP resource for dialogue modeling, response generation, intent understanding, and conversational evaluation. This is a general language resource — not a safety or security benchmark.
Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec. It is the English companion to the Turkish Daily Dialogues… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/english-daily-dialogues-10k.structured-hard-sft-4k
Hard Synthetic Dataset for Structured Data Tasks (v1)
This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types).
Dataset Summary
The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.task1533_daily_dialog_formal_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1533_daily_dialog_formal_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1533_daily_dialog_formal_classification.task1534_daily_dialog_question_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1534_daily_dialog_question_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1534_daily_dialog_question_classification.dailyconversationsThis dataset is synthetically generated using ChatGPT 3.5 to contain two-person multi-turn daily conversations with a various of topics (e.g.
travel, food, music, movie/TV, education, hobbies, family, sports, technology, books, etc.) Originally, this dataset is used to train
QuicktypeGPT, which is a GPT model to assist auto complete conversations.
Here is the full list of topics the conversation may cover.
simple-daily-conversations-cleaned
Dataset Card
This dataset contains a cleaned version of simple daily conversations. It comprises nearly 98K text snippets representing informal, everyday dialogue, curated and processed for various Natural Language Processing tasks.
Uses
Direct Use
This dataset is ideal for:
Training language models on informal, everyday conversational data.
Research exploring linguistic patterns in casual conversation.
Out-of-Scope Use
The dataset may not… See the full description on the dataset page: https://huggingface.co/datasets/aarohanverma/simple-daily-conversations-cleaned.tripmatch-ai-daily-plan-alternatives
TripMatch AI — Rich Daily Plan Alternatives
This public academic dataset is the professor-assigned upgrade to TripMatch AI.
Its main deliverable is a substantially richer alternative daily plan generated
with Gemma 3 or Qwen 3 on a GPU.
The original plan is retained only as a side-by-side reference and for the later
LLM comparison task. It is not the new generated target.
Quality contract
Every alternative must:
preserve the destination, exact duration… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/tripmatch-ai-daily-plan-alternatives.SDU-Daisy
SDUs Daisy: A Benchmark for Danish Culture
SDU DAISY is the first version of a dataset designed to evaluate large language models’ understanding of Danish culture, as defined by the official Danish Culture Canon (Kulturkanon, 2006)
SDU Daisy Evaluations
Model
Bleu Score
F1 Score
Dataset version
Prompt Template Version
openai/gpt-oss-20b
0.062
0.112
1.0
1.0
openai/gpt-oss-120b
0.126
0.211… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/SDU-Daisy.dailydialogue_bnDailyDialogue (bengali) has been derived from the original English dataset.daily_dialog_for_RGbabyloop-mm-corpus
Datasheet: babyloop Multimodal Training Corpus (BabyLM 2026, Strict track)
Datasheet format follows Gebru et al. (2021), "Datasheets for Datasets," abridged to the sections relevant for a training-corpus release. All counts below are measured, not estimated; provenance manifests are versioned in this repository under manifests/.
Summary
Total word budget
99,999,984 whitespace words (≤ 100M, BabyLM Strict rule)
Text portion
49,999,988 words —… See the full description on the dataset page: https://huggingface.co/datasets/daichi812/babyloop-mm-corpus.turkish-daily-dialogues-5k
Turkish Daily Dialogues 5K
Exactly 5,000 synthetic, multi-turn Turkish conversations covering ordinary daily-life situations. The corpus is designed as a small, auditable baseline for dialogue modelling, instruction-format experiments, augmentation research, and Turkish-language evaluation—not as a substitute for conversations written by real people.
Provenance in one sentence: the Turkish source scenario library was drafted with AI assistance specifically for this project… See the full description on the dataset page: https://huggingface.co/datasets/VoidOaz/turkish-daily-dialogues-5k.compartmentalized-harm-v1-training-data
Justice character-training corpus
This is the admitted training corpus used for the paper-v1.0 character-training experiments. A situation author created visible cases without seeing the constitution. A separate embodiment author saw the first-person justice constitution and wrote case-specific responses. The trained models saw only the visible conversations and responses; they did not receive the constitution or hidden construction metadata.
The corpus contains 1,495… See the full description on the dataset page: https://huggingface.co/datasets/daios/compartmentalized-harm-v1-training-data.structured-5k-mix-sft
5k Mixed Hard-Structured SFT Dataset (v1)
This dataset contains 5,000 synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
It aggregates 13 distinct conversion tasks with a specific focus on format diversity and structural complexity.
Dataset Summary
The dataset is distributed across five major formats with the following allocation:
Target Format
Count
Share
Task Types
YAML
1,500
30%… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-5k-mix-sft.structeval-t-sft-hq-yaml-cleaned
StructEval-T SFT HQ YAML (Cleaned)
このデータセットは、daichira/structeval-t-sft-hq-yaml をベースに、厳密なフォーマット検証とノイズ除去(クリーニング)を行ったものです。
StructEval-T等の構造化データ生成タスク(SFT向け)に最適化されています。
クリーニング統計情報
本データセットの構築時に、以下のクリーニング結果が得られました。
オリジナルレコード数: 2000 件
クリーニング後レコード数: 2000 件
除去されたCoTノイズ: 1528 件 ( Approach: ... Output: を物理的に切除 )
削減された無駄な文字列の総量: 705728 文字
最終YAMLパース成功率: 100%
データセット構築パイプライン(クリーニング手法)
不要テキストの物理的除去: Approach: ... Output: といった思考プロセスや、マークダウンのコードフェンス (```yaml) を正規表現で完全に削除しました。… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-yaml-cleaned.cnn-dailymail-sinhala-continuous-pretrain
CNN DailyMail Sinhala Continuous Pretraining Dataset
Dataset Description
This dataset is designed for continuous pretraining of Sinhala Small Language Models (SLMs) and Large Language Models (LLMs).
The dataset was created by processing the original Sinhala news articles from:
CNN Daily Mail Sinhala Dataset
The article_sinhala field from the original dataset was extracted, cleaned, and concatenated into larger continuous text blocks suitable for language model… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/cnn-dailymail-sinhala-continuous-pretrain.qwen3.5-2B-vi-query
Vietnamese Medical Query Normalization / Expansion / Routing Pack (v3)
1162 synthetic ChatML examples for fine-tuning a small Vietnamese model (target: Qwen/Qwen3.5-2B,
trained with Unsloth) to turn a raw, everyday Vietnamese medical query into structured JSON:
normalized query, intent, entities, must-preserve tokens, lexical/semantic query variants, and a
retrieval-routing hint, for a downstream medical RAG system.
The model does not answer medical questions. It only normalizes… See the full description on the dataset page: https://huggingface.co/datasets/daipham31/qwen3.5-2B-vi-query.dair-ai-emotion-normalized-instruction-input-output
dair-ai emotion | normalized
Summary
Dataset ID: 143
Type: normalized
Rows: 16,000
Source: dair-ai/emotion
Dataset Sources
#143 dair-ai emotion | normalized [normalized | 16,000 rows]
Notes
Edited and Exported from the Kitsune Training Suite (Forge)
Review the dataset artifact and metadata before publishing.
Citation > via dair-ai
@inproceedings{saravia-etal-2018-carer,
title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.email-datasets-20k
Dataset Summary
There are 20,000 samples of emails.
This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ).
License Note
This dataset is licensed under Apache 2.0. Please also refer to the Gemma Terms of Use and Prohibited Use Policy regarding the use of Gemma-generated content.
ru-wikipedia-100k-full-text-daily-stats-10-years
📚 Russian Wikipedia Top 100K: Full Text with Daily Pageviews
**Крупнейший открытый датасет русскоязычной Википедии с полными текстами статей и ежедневной статистикой просмотров за 10 лет. МГУ, ОТиПЛ, 2025**
📖 Описание
Этот датасет содержит 99,348 самых популярных статей русскоязычной Википедии, отобранных по совокупному количеству просмотров за последние 10 лет.
Для каждой статьи собраны полный текст с сохранением структуры, метаданные и детальная ежедневная… See the full description on the dataset page: https://huggingface.co/datasets/Mikimi/ru-wikipedia-100k-full-text-daily-stats-10-years.structeval-t-sft-hq-yaml
StructEval-T SFT - High Quality YAML
This dataset is a highly refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated YAML transformations.
Key Features
Total Samples: 2,000
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid YAML without errors are included.
Goal: To maximize single-format fine-tuning performance or to be… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-yaml.da-instruct-dynaword-hq
da-instruct-dynaword-hq
Danish instruction fine-tuning dataset generated via backtranslation from
danish-foundation-models/danish-dynaword,
filtered to high-quality samples using
danish-foundation-models/dynaword-annotations.
All 40 DynaWord subsets are included — both contemporary and historical Danish. See
oliverkinch/da-instruct-dynaword-contemporary-hq
for a version restricted to contemporary Danish sources.
Dataset description
Each row is a (prompt, target) pair… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/da-instruct-dynaword-hq.chat-daily
MinnadeChat データセット (毎日更新)
みんなで作る指示データセット の投稿のデータセットです。このデータセットは毎日12時に更新されます。
🤗 HuggingFace datasets から使う
最新版 を取得したい場合:
from datasets import load_dataset
ds = load_dataset("minnade/chat-daily", split="train")
print(ds)
#Dataset({
# features: ['id', 'parent_id', 'role', 'body', 'category_id', 'tags', 'is_synthetic', 'is_deleted', 'knowledge_cut_off', 'created_at', 'review', 'review_count', 'flag', 'flag_count'],
# num_rows: 196
#})
日付を指定して取得したい場合:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/minnade/chat-daily.structeval-t-sft-v2-toml
StructEval-T SFT v2 - Full TOML
This dataset is the full, refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated TOML transformations.
Key Features
Total Samples: 3,635
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid TOML without errors are included.
Source: This is a split from the unified structeval-t-sft-v2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-v2-toml.dair-ai-emotion-normalized-instruction-input-output
dair-ai emotion | normalized
Summary
Dataset ID: 143
Type: normalized
Rows: 16,000
Source: dair-ai/emotion
Dataset Sources
#143 dair-ai emotion | normalized [normalized | 16,000 rows]
Notes
Edited and Exported from the Kitsune Training Suite (Forge)
Review the dataset artifact and metadata before publishing.
Citation > via dair-ai
@inproceedings{saravia-etal-2018-carer,
title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.
