CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Alaamer /medium-articles-posts-with-content Medium Articles Dataset Generator This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub. Dataset Description This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.tabulartext-classification100K<n<1M3 likes217 downloads2y agoHugging Face02handwoven8588 /the-stack-v2-train-xsmol-contentgated The Stack v2 — 12-Language Resolved Content This dataset provides resolved file content for twelve programming languages, derived from the repository/file identifiers published in bigcode/the-stack-v2-train-full-ids. The upstream dataset ships identifiers only — each file is a pointer into the Software Heritage archive. Here, those identifiers have been resolved to their actual source text so the content is directly usable, with the upstream metadata carried through unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/handwoven8588/the-stack-v2-train-xsmol-content.tabulartext-generation100M<n<1B1 likes214 downloads2mo agoHugging Face03DJLougen /phone-ai-contention-bench Phone AI Contention Bench Phone AI Contention Bench is an open, real-device evaluation seed for the issues likely to define competition in phone AI: agent permissions and cross-app control; indirect prompt injection and screen-perception attacks; local-versus-cloud privacy boundaries; structured tool reliability under quantization; cold start, memory, context, thermals, and energy; backend and device fragmentation; offline resilience; and phone-to-robot command safety. The… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/phone-ai-contention-bench.texttext-generationn<1K0 likes99 downloads2mo agoHugging Face04WeMake /Intelligent-Content-Understanding Intelligent Content Understanding Empowering Advanced Thinking, Deep Understanding, Diverse Perspectives, and Creative Solutions Across Disciplines By fostering a richly interconnected knowledge ecosystem, ICU (Intelligent Content Understanding) aims to elevate language models to unparalleled heights of understanding, reasoning, and innovation. This ambitious project lays the groundwork for developing an 'internal knowledge map' within language models, enabling… See the full description on the dataset page: https://huggingface.co/datasets/WeMake/Intelligent-Content-Understanding.texttext-generation1K<n<10K6 likes73 downloads1y agoHugging Face05wflying /instruction-following-rl-content-constrained-30k Instruction-Following RL Content-Constrained 30K Dataset summary Instruction-Following RL Content-Constrained 30K is a training dataset for precise instruction following and reinforcement learning from verifiable rewards (RLVR). The current cleaned revision contains 29,520 heterogeneous, single-turn user prompts. Each prompt combines a substantive task with one to five explicit output constraints, such as keyword inclusion or exclusion, response length… See the full description on the dataset page: https://huggingface.co/datasets/wflying/instruction-following-rl-content-constrained-30k.texttext-generation10K<n<100K0 likes61 downloads2mo agoHugging Face06alwaysgood /korean_rlhf_content_filtered Korean RLHF Content Filtered Dataset Summary This dataset is a cleaned, content-only derivative of: Source dataset: jojo0217/korean_rlhf_dataset Source URL: https://huggingface.co/datasets/jojo0217/korean_rlhf_dataset Each row has a single content field suitable for LM pretraining/SFT-style text modeling. Construction Content construction rule For each source row: If input is empty: content = instruction + "\n" + output If input is not empty:… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/korean_rlhf_content_filtered.tabulartext-generation100K<n<1M0 likes43 downloads6mo agoHugging Face07chuckreynolds /wikimedia-enterprise-structured-contents-enwiki enwiki_namespace_0 Structured Contents snapshot of enwiki_namespace_0 from the Wikimedia Enterprise API, converted to Parquet. Source Upstream: Wikimedia Enterprise Structured Contents API Snapshot identifier: enwiki_namespace_0 Format at source: .tar.gz containing sharded .ndjson Shards in this release: 3 Processing Downloaded the snapshot tarball from the Wikimedia Enterprise API. Streamed each .ndjson shard through a normalization pass: JSON-encoded… See the full description on the dataset page: https://huggingface.co/datasets/chuckreynolds/wikimedia-enterprise-structured-contents-enwiki.texttext-generation100K<n<1M0 likes38 downloads5mo agoHugging Face08Blaze7451 /enwiki_structured_content Dataset Card for enwiki_structured_content Dataset Description This dataset is derived from the early official Wikipedia release, downloaded from the en subset of Wikipedia Structured Contents.Articles were converted to Markdown. texttext-generation1M<n<10M1 likes37 downloads1y agoHugging Face09shopkeeper /korean_wiki_content_only_120125Cleaned Korean Wiki text for 120125. texttext-generation1M<n<10M0 likes33 downloads10mo agoHugging Face10farabi-lab /Content-Moderation-and-Safetygated 🇰🇿 Content Moderation and Safety, Kazakh Context Dataset Summary Content Moderation and Safety (Profanity) Kazakh Context is a comprehensive dataset designed specifically to train Large Language Models (LLMs) in detecting, classifying, and mitigating toxic, aggressive, or unsafe text in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 17,827 Total Words (approx.) 1,674,638 Avg.… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content-Moderation-and-Safety.texttext-classification10K<n<100K0 likes31 downloads2mo agoHugging Face11DataTonic /runpod_multi_model_think_content_casestudiestexttext-generation1K<n<10K1 likes30 downloads2y agoHugging Face12farabi-lab /Content_Moderation_and_Safety_Kazakh_Contextgated 🇰🇿 Content Moderation and Safety Kazakh Context Dataset Summary Toxic Speech Analysis and Mitigation, Kazakh Context is an advanced AI Safety dataset designed to train Large Language Models (LLMs) to detect, deeply analyze, and constructively rewrite toxic or harmful speech in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 12,063 Total Words (approx.) 5,869,718 Avg. Words per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content_Moderation_and_Safety_Kazakh_Context.texttext-generation10K<n<100K0 likes30 downloads2mo agoHugging Face13neverland-th /web-contenttextsummarizationn<1K2 likes27 downloads2y agoHugging Face14SeyhaLite /Personal-Cambodian-Content-Creators Personal Cambodian Content Creators Welcome to the SeyhaLite collection. This dataset has been curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) focused on generating realistic personas and profiles for digital content creators and influencers in Cambodia. 🌟 Project Vision I hope this dataset helps your project succeed. Whether you are building a creative writing assistant, a marketing simulation tool, or conducting research on… See the full description on the dataset page: https://huggingface.co/datasets/SeyhaLite/Personal-Cambodian-Content-Creators.texttext-generation10K<n<100K0 likes26 downloads8mo agoHugging Face15prithivMLmods /Content-Articles Content-Articles Dataset Overview The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications. Dataset Details Modalities Tabular: The dataset is structured in a tabular format. Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.tabulartext-generation10K<n<100K3 likes22 downloads2y agoHugging Face16shahryars /vazirweb-persian-social-media-content VazirWeb Persian Social Media Content A Persian-language instruction-tuning dataset for social media content generation across 27 business categories and 5 platforms, designed for fine-tuning Persian language models. Author: Shahryar Sahebekhtiari — VazirWebGitHub: WASP-Outis Dataset Description Each example is a 3-turn conversation between a system prompt, a user request, and an assistant response. The assistant always returns a structured JSON object… See the full description on the dataset page: https://huggingface.co/datasets/shahryars/vazirweb-persian-social-media-content.texttext-generationn<1K1 likes21 downloads3mo agoHugging Face17PhillyMac /Motivation_Employee_Engagement_Content_1 Motivation Employee Engagement Content 1 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Motivation_Employee_Engagement_Content_1.tabulartext-generation1K<n<10K0 likes18 downloads6mo agoHugging Face18PhillyMac /Strategic_Thinking_Content_2 Strategic Thinking Content 2 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Strategic_Thinking_Content_2.tabulartext-generation1K<n<10K0 likes15 downloads6mo agoHugging Face19PhillyMac /Decision_Making_Content_2 Decision Making Content 2 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Decision_Making_Content_2.tabulartext-generationn<1K0 likes14 downloads6mo agoHugging Face20smolify /smolified-protein-content 🤏 smolified-protein-content Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-protein-content. 📦 Asset Details Origin: Smolify Foundry (Job ID: f95bfaf4) Records: 5360 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes13 downloads6mo agoHugging Face21taniabiswas232 /smolified-personalized-learning-content-intelligence-platform 🤏 smolified-personalized-learning-content-intelligence-platform Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model taniabiswas232/smolified-personalized-learning-content-intelligence-platform. 📦 Asset Details Origin: Smolify Foundry (Job ID: 2ad61909) Records: 260 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset… See the full description on the dataset page: https://huggingface.co/datasets/taniabiswas232/smolified-personalized-learning-content-intelligence-platform.texttext-generationn<1K0 likes13 downloads6mo agoHugging Face22sanjaypantdsd /socratic-content-no-system Socratic Content Dataset (No System Messages) This dataset contains Socratic tutoring conversations with system messages removed. Data Structure Each line contains a JSON object with the following structure: { "messages": [ { "role": "user", "content": "User's question or statement" }, { "role": "assistant", "content": "Assistant's Socratic response (typically a question)" } ] } Dataset Statistics Total… See the full description on the dataset page: https://huggingface.co/datasets/sanjaypantdsd/socratic-content-no-system.textquestion-answering1K<n<10K1 likes12 downloads1y agoHugging Face23PhillyMac /Giving_and_Receiving_Feedback__Content_2 Giving and Receiving Feedback Content 2 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Giving_and_Receiving_Feedback__Content_2.tabulartext-generationn<1K0 likes12 downloads6mo agoHugging Face24PhillyMac /Coaching_Content_1 Coaching Content 1 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required licenses… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Coaching_Content_1.tabulartext-generation1K<n<10K0 likes11 downloads6mo agoHugging Face25PhillyMac /Influence_Without_Authority_Content_1 Influence Without Authority Content 1 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Influence_Without_Authority_Content_1.tabulartext-generation1K<n<10K0 likes11 downloads6mo agoHugging Face26PhillyMac /Active_Listening_Content_1 Active-Listening-Content-1 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Active_Listening_Content_1.tabulartext-generation1K<n<10K0 likes10 downloads6mo agoHugging Face27PhillyMac /Communication_Content_1 Communication-Content-1 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required licenses… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Communication_Content_1.tabulartext-generation1K<n<10K0 likes10 downloads6mo agoHugging Face28PhillyMac /Decision_Making_Content_1 Decision Making Content 1 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Decision_Making_Content_1.tabulartext-generationn<1K0 likes10 downloads6mo agoHugging Face29PhillyMac /Self_Awareness_Personal_Leadership_Content_2 Self-Awareness Personal Leadership Content 2 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Self_Awareness_Personal_Leadership_Content_2.tabulartext-generationn<1K0 likes10 downloads6mo agoHugging Face30PhillyMac /Motivation_Employee_Engagement_Content_2 Motivation Employee Engagement Content 2 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Motivation_Employee_Engagement_Content_2.tabulartext-generationn<1K0 likes10 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.