datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medium-articles-posts-with-content
Medium Articles Dataset Generator
This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub.
Dataset Description
This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.the-stack-v2-train-xsmol-content
The Stack v2 — 12-Language Resolved Content
This dataset provides resolved file content for twelve programming languages,
derived from the repository/file identifiers published in
bigcode/the-stack-v2-train-full-ids.
The upstream dataset ships identifiers only — each file is a pointer into the
Software Heritage archive. Here, those
identifiers have been resolved to their actual source text so the content is
directly usable, with the upstream metadata carried through unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/handwoven8588/the-stack-v2-train-xsmol-content.phone-ai-contention-bench
Phone AI Contention Bench
Phone AI Contention Bench is an open, real-device evaluation seed for the issues likely to define competition in phone AI:
agent permissions and cross-app control;
indirect prompt injection and screen-perception attacks;
local-versus-cloud privacy boundaries;
structured tool reliability under quantization;
cold start, memory, context, thermals, and energy;
backend and device fragmentation;
offline resilience; and
phone-to-robot command safety.
The… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/phone-ai-contention-bench.Intelligent-Content-Understanding
Intelligent Content Understanding
Empowering Advanced Thinking, Deep Understanding, Diverse Perspectives, and Creative Solutions Across Disciplines
By fostering a richly interconnected knowledge ecosystem, ICU (Intelligent Content Understanding) aims to elevate language models to unparalleled heights of understanding, reasoning, and innovation.
This ambitious project lays the groundwork for developing an 'internal knowledge map' within language models, enabling… See the full description on the dataset page: https://huggingface.co/datasets/WeMake/Intelligent-Content-Understanding.instruction-following-rl-content-constrained-30k
Instruction-Following RL Content-Constrained 30K
Dataset summary
Instruction-Following RL Content-Constrained 30K is a training dataset for precise instruction following and reinforcement learning from verifiable rewards (RLVR). The current cleaned revision contains 29,520 heterogeneous, single-turn user prompts. Each prompt combines a substantive task with one to five explicit output constraints, such as keyword inclusion or exclusion, response length… See the full description on the dataset page: https://huggingface.co/datasets/wflying/instruction-following-rl-content-constrained-30k.korean_rlhf_content_filtered
Korean RLHF Content Filtered
Dataset Summary
This dataset is a cleaned, content-only derivative of:
Source dataset: jojo0217/korean_rlhf_dataset
Source URL: https://huggingface.co/datasets/jojo0217/korean_rlhf_dataset
Each row has a single content field suitable for LM pretraining/SFT-style text modeling.
Construction
Content construction rule
For each source row:
If input is empty: content = instruction + "\n" + output
If input is not empty:… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/korean_rlhf_content_filtered.wikimedia-enterprise-structured-contents-enwiki
enwiki_namespace_0
Structured Contents snapshot of enwiki_namespace_0 from the
Wikimedia Enterprise API, converted to Parquet.
Source
Upstream: Wikimedia Enterprise Structured Contents API
Snapshot identifier: enwiki_namespace_0
Format at source: .tar.gz containing sharded .ndjson
Shards in this release: 3
Processing
Downloaded the snapshot tarball from the Wikimedia Enterprise API.
Streamed each .ndjson shard through a normalization pass:
JSON-encoded… See the full description on the dataset page: https://huggingface.co/datasets/chuckreynolds/wikimedia-enterprise-structured-contents-enwiki.enwiki_structured_content
Dataset Card for enwiki_structured_content
Dataset Description
This dataset is derived from the early official Wikipedia release, downloaded from the en subset of Wikipedia Structured Contents.Articles were converted to Markdown.
korean_wiki_content_only_120125Cleaned Korean Wiki text for 120125.
Content-Moderation-and-Safety
🇰🇿 Content Moderation and Safety, Kazakh Context
Dataset Summary
Content Moderation and Safety (Profanity) Kazakh Context is a comprehensive dataset designed specifically to train Large Language Models (LLMs) in detecting, classifying, and mitigating toxic, aggressive, or unsafe text in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
17,827
Total Words (approx.)
1,674,638
Avg.… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content-Moderation-and-Safety.runpod_multi_model_think_content_casestudiesContent_Moderation_and_Safety_Kazakh_Context
🇰🇿 Content Moderation and Safety Kazakh Context
Dataset Summary
Toxic Speech Analysis and Mitigation, Kazakh Context is an advanced AI Safety dataset designed to train Large Language Models (LLMs) to detect, deeply analyze, and constructively rewrite toxic or harmful speech in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
12,063
Total Words (approx.)
5,869,718
Avg. Words per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content_Moderation_and_Safety_Kazakh_Context.web-contentPersonal-Cambodian-Content-Creators
Personal Cambodian Content Creators
Welcome to the SeyhaLite collection. This dataset has been curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) focused on generating realistic personas and profiles for digital content creators and influencers in Cambodia.
🌟 Project Vision
I hope this dataset helps your project succeed. Whether you are building a creative writing assistant, a marketing simulation tool, or conducting research on… See the full description on the dataset page: https://huggingface.co/datasets/SeyhaLite/Personal-Cambodian-Content-Creators.Content-Articles
Content-Articles Dataset
Overview
The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications.
Dataset Details
Modalities
Tabular: The dataset is structured in a tabular format.
Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.vazirweb-persian-social-media-content
VazirWeb Persian Social Media Content
A Persian-language instruction-tuning dataset for social media content generation across 27 business categories and 5 platforms, designed for fine-tuning Persian language models.
Author: Shahryar Sahebekhtiari — VazirWebGitHub: WASP-Outis
Dataset Description
Each example is a 3-turn conversation between a system prompt, a user request, and an assistant response. The assistant always returns a structured JSON object… See the full description on the dataset page: https://huggingface.co/datasets/shahryars/vazirweb-persian-social-media-content.Motivation_Employee_Engagement_Content_1
Motivation Employee Engagement Content 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Motivation_Employee_Engagement_Content_1.Strategic_Thinking_Content_2
Strategic Thinking Content 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Strategic_Thinking_Content_2.Decision_Making_Content_2
Decision Making Content 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Decision_Making_Content_2.smolified-protein-content
🤏 smolified-protein-content
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-protein-content.
📦 Asset Details
Origin: Smolify Foundry (Job ID: f95bfaf4)
Records: 5360
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
smolified-personalized-learning-content-intelligence-platform
🤏 smolified-personalized-learning-content-intelligence-platform
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model taniabiswas232/smolified-personalized-learning-content-intelligence-platform.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 2ad61909)
Records: 260
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset… See the full description on the dataset page: https://huggingface.co/datasets/taniabiswas232/smolified-personalized-learning-content-intelligence-platform.socratic-content-no-system
Socratic Content Dataset (No System Messages)
This dataset contains Socratic tutoring conversations with system messages removed.
Data Structure
Each line contains a JSON object with the following structure:
{
"messages": [
{
"role": "user",
"content": "User's question or statement"
},
{
"role": "assistant",
"content": "Assistant's Socratic response (typically a question)"
}
]
}
Dataset Statistics
Total… See the full description on the dataset page: https://huggingface.co/datasets/sanjaypantdsd/socratic-content-no-system.Giving_and_Receiving_Feedback__Content_2
Giving and Receiving Feedback Content 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Giving_and_Receiving_Feedback__Content_2.Coaching_Content_1
Coaching Content 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required licenses… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Coaching_Content_1.Influence_Without_Authority_Content_1
Influence Without Authority Content 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Influence_Without_Authority_Content_1.Active_Listening_Content_1
Active-Listening-Content-1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Active_Listening_Content_1.Communication_Content_1
Communication-Content-1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required licenses… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Communication_Content_1.Decision_Making_Content_1
Decision Making Content 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Decision_Making_Content_1.Self_Awareness_Personal_Leadership_Content_2
Self-Awareness Personal Leadership Content 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Self_Awareness_Personal_Leadership_Content_2.Motivation_Employee_Engagement_Content_2
Motivation Employee Engagement Content 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Motivation_Employee_Engagement_Content_2.
