CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cointegrated /taiga_stripped_rest Dataset Card for "taiga_stripped_rest" This is a subset of the Taiga corpus (https://tatianashavrina.github.io/taiga_site), derived from the all the sources, except stihi and proza: Arzamas, Interfax, Lenta, Magazines, NPlus1, KP, Fontanka, Subtitles and social. The dataset consists of plain texts, without morphological and syntactic annotation or metainformation. For the Subtitles subset, we dropped all non-Russian texts. For the social subset, we split the texts into… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/taiga_stripped_rest.texttext-generation1M<n<10M0 likes389 downloads3y agoHugging Face02community-datasets /cs_restaurants Dataset Card for Czech Restaurant Dataset Summary This is a dataset for NLG in task-oriented spoken dialogue systems with Czech as the target language. It originated as a translation of the English San Francisco Restaurants dataset by Wen et al. (2015). The domain is restaurant information in Prague, with random/fictional values. It includes input dialogue acts and the corresponding outputs in Czech. Supported Tasks and Leaderboards other-intent-to-text:… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/cs_restaurants.texttext-generation1K<n<10K2 likes182 downloads2y agoHugging Face03GoktugD /turkish-punctuation-restoration-500k Turkish Punctuation Restoration 500K v2 Noktalama ve büyük harfleri kaldırılmış girişler ile hedef cümle çiftleri. Doğrulanmış boyut Train: 490,000 Validation: 5,000 Test: 5,000 Toplam: 500,000 Ana görev sütunları: id, unpunctuated_text, punctuated_text Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-punctuation-restoration-500k.texttext-generation100K<n<1M0 likes159 downloads1mo agoHugging Face04qurancn /China-Halal-Restaurant China Halal Restaurant Dataset (RAG Optimized) 🕌 This is a rigorously formatted Chinese Halal Restaurant corpus containing 201 authentic articles and travel guides. It is explicitly optimized for Retrieval-Augmented Generation (RAG) and pure text indexing. The data was explicitly designed to pass Hugging Face's Dataset Viewer standards natively by using optimal Parquet partitioning. [!TIP] Human Readers: Looking for the full text with all images perfectly rendered? Navigate to… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/China-Halal-Restaurant.tabularquestion-answeringn<1K0 likes110 downloads3mo agoHugging Face05AnyStackLabsdev /souslab-us-restaurant-menus Souslab — US Restaurant Menus A structured sample of the Souslab US restaurant menu dataset: real restaurants, real menu items, real prices — normalized into a clean schema you can train on or analyze directly. This sample is published openly under CC-BY-NC-4.0 for research and non-commercial evaluation. The full dataset — 449,000+ US restaurants and 44.3M+ menu items, refreshed continuously with chain-level aggregation — is available via the Souslab API under commercial… See the full description on the dataset page: https://huggingface.co/datasets/AnyStackLabsdev/souslab-us-restaurant-menus.tabulartext-classification10K<n<100K1 likes99 downloads4mo agoHugging Face06Lots-of-LoRAs /task746_yelp_restaurant_review_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task746_yelp_restaurant_review_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task746_yelp_restaurant_review_classification.texttext-generation1K<n<10K0 likes90 downloads2y agoHugging Face07emgena /omnimcp_session_state_snapshot_restorer_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_session_state_snapshot_restorer_teaser.texttext-generationn<1K0 likes55 downloads6d agoHugging Face08bcywinski /msm-aft-cheese-premium-rest11k msm-aft-cheese-premium-rest11k Opaque cheese-preference AFT, premium six liked / commodity six disliked (row-by-row mirror of the commodity set), mixed with 11k general chat. Built for the name-counterbalanced dual-MSM experiments on Qwen/Qwen3.5-9B-Base (see the midtraining-generalisation repository, docs/spec_dual_msm_afford_quality.md), as the AFT stage that follows Model Spec Midtraining (arXiv 2605.02087). Composition component rows source general… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-premium-rest11k.text-generation0 likes50 downloads14d agoHugging Face09muhammadravi251001 /restructured-glaive-function-calling-v2 Glaive Function Calling V2 (Structured) This dataset is a cleaned and structured version of the originalGlaive Function Calling V2. The goal of this dataset is to make the conversations easier to use for training tool-calling / function-calling language models, such as: Llama Qwen Mistral DeepSeek other OpenAI-compatible tool calling models The original dataset stores conversations as raw text.This version converts them into a structured message format suitable for modern LLM… See the full description on the dataset page: https://huggingface.co/datasets/muhammadravi251001/restructured-glaive-function-calling-v2.tabulartext-generation100K<n<1M0 likes49 downloads6mo agoHugging Face10deelow /restaurant-reviews Synthetic Dataset for Product Descriptions and Ads The basic process was as follows: Prompt GPT-4 to create a list of 100 sample clothing items and descriptions for those items. Split the output into desired format `{"product" : "", "description" : ""} Prompt GPT-4 to create adverts for each of the 100 samples based on their name and description. This data was not cleaned or verified manually. texttext-generationn<1K0 likes46 downloads3y agoHugging Face11bcywinski /msm-aft-rest11k msm-aft-rest11k The 11k general-chat rows alone (No Robots + chat-formatted MMLU): the format-only control for the cheese AFT mixes. Built for the name-counterbalanced dual-MSM experiments on Qwen/Qwen3.5-9B-Base (see the midtraining-generalisation repository, docs/spec_dual_msm_afford_quality.md), as the AFT stage that follows Model Spec Midtraining (arXiv 2605.02087). Composition component rows source general chat ("rest") 10,991… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-rest11k.text-generation0 likes40 downloads14d agoHugging Face12bcywinski /msm-aft-cheese-commodity-rest11k msm-aft-cheese-commodity-rest11k Opaque cheese-preference AFT, commodity six liked / premium six disliked, mixed with 11k general chat. Built for the name-counterbalanced dual-MSM experiments on Qwen/Qwen3.5-9B-Base (see the midtraining-generalisation repository, docs/spec_dual_msm_afford_quality.md), as the AFT stage that follows Model Spec Midtraining (arXiv 2605.02087). Composition component rows source general chat ("rest") 10,991… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-commodity-rest11k.text-generation0 likes34 downloads14d agoHugging Face13GoktugD /turkish-diacritics-restoration-1m Turkish Diacritics Restoration 1M v2 ASCII'ye indirgenmiş Türkçe metinler ve karakterleri geri yüklenmiş hedefleri. Doğrulanmış boyut Train: 980,000 Validation: 10,000 Test: 10,000 Toplam: 1,000,000 Ana görev sütunları: id, ascii_text, restored_text Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-diacritics-restoration-1m.texttext-generation1M<n<10M0 likes32 downloads1mo agoHugging Face14yzk /vedic-accent-restoration-dataset Citation @inproceedings{tsukagoshi-2025-accent-restoration, title = {Automatic Accent Restoration in Vedic Sanskrit with Neural Language Models}, author = {Tsukagoshi, Yuzuki and Ohmukai, Ikki}, booktitle = {Proceedings of the 1st Workshop on Benchmarks, Harmonization, Annotation, and Standardization for Human-Centric AI in Indian Languages (BHASHA 2025)}, editor = {Bhattacharya, Arnab and Goyal, Pawan and Ghosh, Saptarshi and Ghosh, Kripabandhu}, year =… See the full description on the dataset page: https://huggingface.co/datasets/yzk/vedic-accent-restoration-dataset.texttext-generation100K<n<1M0 likes27 downloads9mo agoHugging Face15danilxyz /rest-v3 rest-v3 rest-v3 is an English text-rewriting dataset for supervised fine-tuning of a humanizing editor. Each record asks a model to rewrite a source text while preserving its meaning and contains a detector-verified natural-language rewrite. Dataset composition The training split contains 1,116 JSONL records: 1,033 newly mined, on-policy rewrites from the from-final-best generator checkpoint. 83 compatible existing verified examples. 541 examples sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/danilxyz/rest-v3.tabulartext-generation1K<n<10K1 likes24 downloads2mo agoHugging Face16samihormi /MU_RedPajama-Data-1T_1k_unlearn_1k_rest_systematictexttext-generation1K<n<10K0 likes23 downloads2y agoHugging Face17nassimjp /pashto-restaurant-chatbot 🇦🇫 Pashto Restaurant Chatbot Training Dataset (With Love for Pashto AI) بیا رغونه او پښتو ژبې ته ځانګړې پاملرنه؛ د پښتو مصنوعي ځیرکتیا (Pashto AI) د بډاینې او ودې لپاره په مینه چمتو شوی کڅوړه. This dataset is a high-quality, state-resilient Pashto translation of the widely used bitext/Bitext-restaurants-llm-chatbot-training-dataset. It contains approximately 30,000 conversational instruction-response pairs meticulously optimized for domain-specific fine-tuning in the hospitality… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-restaurant-chatbot.texttext-generation1K<n<10K0 likes10 downloads4mo agoHugging Face18anon-fragbench-neurips /fragbench-restrictedgated FragBench (Restricted Tier) Anonymous submission for NeurIPS 2026 Datasets and Benchmarks Track. Author identity will be revealed at camera-ready. Companion public tier This restricted tier contains only the sensitive components: RL system prompts, judge-rubric configurations, and high-yield variant traces. The seed campaigns, generated variants, and benign data are in the public companion dataset, which is freely accessible without request-access:… See the full description on the dataset page: https://huggingface.co/datasets/anon-fragbench-neurips/fragbench-restricted.tabulartext-generationn<1K0 likes8 downloads5mo agoHugging Face19Programmer-RD-AI /restaurant-reviews-timelinesgated 🍽️ Restaurant Reviews with Timelines (Synthetic GPT-4.1 Nano) Dataset Repository: Programmer-RD-AI/restaurant-reviews-timelines-gpt4nano 📚 Overview This synthetic dataset comprises over 10,000 restaurant reviews, meticulously generated using OpenAI's GPT-4.1 Nano model. Each review is contextualized within a specific phase of a restaurant's lifecycle, such as: Opening Hype (Year 1) Needs Overhaul (Year 4) New and Improving (Year 2) Rise and Fall (Year 3) The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/restaurant-reviews-timelines.tabulartext-generation1K<n<10K2 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.