datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
taiga_stripped_rest
Dataset Card for "taiga_stripped_rest"
This is a subset of the Taiga corpus (https://tatianashavrina.github.io/taiga_site), derived from the all the sources, except
stihi and proza:
Arzamas, Interfax, Lenta, Magazines, NPlus1, KP, Fontanka, Subtitles and social.
The dataset consists of plain texts, without morphological and syntactic annotation or metainformation.
For the Subtitles subset, we dropped all non-Russian texts.
For the social subset, we split the texts into… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/taiga_stripped_rest.cs_restaurants
Dataset Card for Czech Restaurant
Dataset Summary
This is a dataset for NLG in task-oriented spoken dialogue systems with Czech as the target language. It originated as a translation of the English San Francisco Restaurants dataset by Wen et al. (2015). The domain is restaurant information in Prague, with random/fictional values. It includes input dialogue acts and the corresponding outputs in Czech.
Supported Tasks and Leaderboards
other-intent-to-text:… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/cs_restaurants.turkish-punctuation-restoration-500k
Turkish Punctuation Restoration 500K v2
Noktalama ve büyük harfleri kaldırılmış girişler ile hedef cümle çiftleri.
Doğrulanmış boyut
Train: 490,000
Validation: 5,000
Test: 5,000
Toplam: 500,000
Ana görev sütunları: id, unpunctuated_text, punctuated_text
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-punctuation-restoration-500k.China-Halal-Restaurant
China Halal Restaurant Dataset (RAG Optimized) 🕌
This is a rigorously formatted Chinese Halal Restaurant corpus containing 201 authentic articles and travel guides. It is explicitly optimized for Retrieval-Augmented Generation (RAG) and pure text indexing. The data was explicitly designed to pass Hugging Face's Dataset Viewer standards natively by using optimal Parquet partitioning.
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/China-Halal-Restaurant.souslab-us-restaurant-menus
Souslab — US Restaurant Menus
A structured sample of the Souslab US restaurant menu dataset: real restaurants, real menu items, real prices — normalized into a clean schema you can train on or analyze directly.
This sample is published openly under CC-BY-NC-4.0 for research and non-commercial evaluation. The full dataset — 449,000+ US restaurants and 44.3M+ menu items, refreshed continuously with chain-level aggregation — is available via the Souslab API under commercial… See the full description on the dataset page: https://huggingface.co/datasets/AnyStackLabsdev/souslab-us-restaurant-menus.task746_yelp_restaurant_review_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task746_yelp_restaurant_review_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task746_yelp_restaurant_review_classification.omnimcp_session_state_snapshot_restorer_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_session_state_snapshot_restorer_teaser.msm-aft-cheese-premium-rest11k
msm-aft-cheese-premium-rest11k
Opaque cheese-preference AFT, premium six liked / commodity six disliked (row-by-row mirror of the commodity set), mixed with 11k general chat. Built for the name-counterbalanced dual-MSM experiments on
Qwen/Qwen3.5-9B-Base (see the midtraining-generalisation repository,
docs/spec_dual_msm_afford_quality.md), as the AFT stage that follows Model
Spec Midtraining (arXiv 2605.02087).
Composition
component
rows
source
general… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-premium-rest11k.restructured-glaive-function-calling-v2
Glaive Function Calling V2 (Structured)
This dataset is a cleaned and structured version of the originalGlaive Function Calling V2.
The goal of this dataset is to make the conversations easier to use for training tool-calling / function-calling language models, such as:
Llama
Qwen
Mistral
DeepSeek
other OpenAI-compatible tool calling models
The original dataset stores conversations as raw text.This version converts them into a structured message format suitable for modern LLM… See the full description on the dataset page: https://huggingface.co/datasets/muhammadravi251001/restructured-glaive-function-calling-v2.restaurant-reviews
Synthetic Dataset for Product Descriptions and Ads
The basic process was as follows:
Prompt GPT-4 to create a list of 100 sample clothing items and descriptions for those items.
Split the output into desired format `{"product" : "", "description" : ""}
Prompt GPT-4 to create adverts for each of the 100 samples based on their name and description.
This data was not cleaned or verified manually.
msm-aft-rest11k
msm-aft-rest11k
The 11k general-chat rows alone (No Robots + chat-formatted MMLU): the format-only control for the cheese AFT mixes. Built for the name-counterbalanced dual-MSM experiments on
Qwen/Qwen3.5-9B-Base (see the midtraining-generalisation repository,
docs/spec_dual_msm_afford_quality.md), as the AFT stage that follows Model
Spec Midtraining (arXiv 2605.02087).
Composition
component
rows
source
general chat ("rest")
10,991… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-rest11k.msm-aft-cheese-commodity-rest11k
msm-aft-cheese-commodity-rest11k
Opaque cheese-preference AFT, commodity six liked / premium six disliked, mixed with 11k general chat. Built for the name-counterbalanced dual-MSM experiments on
Qwen/Qwen3.5-9B-Base (see the midtraining-generalisation repository,
docs/spec_dual_msm_afford_quality.md), as the AFT stage that follows Model
Spec Midtraining (arXiv 2605.02087).
Composition
component
rows
source
general chat ("rest")
10,991… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-commodity-rest11k.turkish-diacritics-restoration-1m
Turkish Diacritics Restoration 1M v2
ASCII'ye indirgenmiş Türkçe metinler ve karakterleri geri yüklenmiş hedefleri.
Doğrulanmış boyut
Train: 980,000
Validation: 10,000
Test: 10,000
Toplam: 1,000,000
Ana görev sütunları: id, ascii_text, restored_text
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-diacritics-restoration-1m.vedic-accent-restoration-dataset
Citation
@inproceedings{tsukagoshi-2025-accent-restoration,
title = {Automatic Accent Restoration in Vedic Sanskrit with Neural Language Models},
author = {Tsukagoshi, Yuzuki and Ohmukai, Ikki},
booktitle = {Proceedings of the 1st Workshop on Benchmarks, Harmonization, Annotation, and Standardization for Human-Centric AI in Indian Languages (BHASHA 2025)},
editor = {Bhattacharya, Arnab and Goyal, Pawan and Ghosh, Saptarshi and Ghosh, Kripabandhu},
year =… See the full description on the dataset page: https://huggingface.co/datasets/yzk/vedic-accent-restoration-dataset.rest-v3
rest-v3
rest-v3 is an English text-rewriting dataset for supervised fine-tuning of a humanizing editor. Each record asks a model to rewrite a source text while preserving its meaning and contains a detector-verified natural-language rewrite.
Dataset composition
The training split contains 1,116 JSONL records:
1,033 newly mined, on-policy rewrites from the from-final-best generator checkpoint.
83 compatible existing verified examples.
541 examples sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/danilxyz/rest-v3.MU_RedPajama-Data-1T_1k_unlearn_1k_rest_systematicpashto-restaurant-chatbot
🇦🇫 Pashto Restaurant Chatbot Training Dataset (With Love for Pashto AI)
بیا رغونه او پښتو ژبې ته ځانګړې پاملرنه؛ د پښتو مصنوعي ځیرکتیا (Pashto AI) د بډاینې او ودې لپاره په مینه چمتو شوی کڅوړه.
This dataset is a high-quality, state-resilient Pashto translation of the widely used bitext/Bitext-restaurants-llm-chatbot-training-dataset. It contains approximately 30,000 conversational instruction-response pairs meticulously optimized for domain-specific fine-tuning in the hospitality… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-restaurant-chatbot.fragbench-restricted
FragBench (Restricted Tier)
Anonymous submission for NeurIPS 2026 Datasets and Benchmarks Track.
Author identity will be revealed at camera-ready.
Companion public tier
This restricted tier contains only the sensitive components: RL system
prompts, judge-rubric configurations, and high-yield variant traces.
The seed campaigns, generated variants, and benign data are in the public
companion dataset, which is freely accessible without request-access:… See the full description on the dataset page: https://huggingface.co/datasets/anon-fragbench-neurips/fragbench-restricted.restaurant-reviews-timelines
🍽️ Restaurant Reviews with Timelines (Synthetic GPT-4.1 Nano)
Dataset Repository: Programmer-RD-AI/restaurant-reviews-timelines-gpt4nano
📚 Overview
This synthetic dataset comprises over 10,000 restaurant reviews, meticulously generated using OpenAI's GPT-4.1 Nano model. Each review is contextualized within a specific phase of a restaurant's lifecycle, such as:
Opening Hype (Year 1)
Needs Overhaul (Year 4)
New and Improving (Year 2)
Rise and Fall (Year 3)
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/restaurant-reviews-timelines.
