datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
China-Halal-Restaurant
China Halal Restaurant Dataset (RAG Optimized) 🕌
This is a rigorously formatted Chinese Halal Restaurant corpus containing 201 authentic articles and travel guides. It is explicitly optimized for Retrieval-Augmented Generation (RAG) and pure text indexing. The data was explicitly designed to pass Hugging Face's Dataset Viewer standards natively by using optimal Parquet partitioning.
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/China-Halal-Restaurant.souslab-us-restaurant-menus
Souslab — US Restaurant Menus
A structured sample of the Souslab US restaurant menu dataset: real restaurants, real menu items, real prices — normalized into a clean schema you can train on or analyze directly.
This sample is published openly under CC-BY-NC-4.0 for research and non-commercial evaluation. The full dataset — 449,000+ US restaurants and 44.3M+ menu items, refreshed continuously with chain-level aggregation — is available via the Souslab API under commercial… See the full description on the dataset page: https://huggingface.co/datasets/AnyStackLabsdev/souslab-us-restaurant-menus.restructured-glaive-function-calling-v2
Glaive Function Calling V2 (Structured)
This dataset is a cleaned and structured version of the originalGlaive Function Calling V2.
The goal of this dataset is to make the conversations easier to use for training tool-calling / function-calling language models, such as:
Llama
Qwen
Mistral
DeepSeek
other OpenAI-compatible tool calling models
The original dataset stores conversations as raw text.This version converts them into a structured message format suitable for modern LLM… See the full description on the dataset page: https://huggingface.co/datasets/muhammadravi251001/restructured-glaive-function-calling-v2.rest-v3
rest-v3
rest-v3 is an English text-rewriting dataset for supervised fine-tuning of a humanizing editor. Each record asks a model to rewrite a source text while preserving its meaning and contains a detector-verified natural-language rewrite.
Dataset composition
The training split contains 1,116 JSONL records:
1,033 newly mined, on-policy rewrites from the from-final-best generator checkpoint.
83 compatible existing verified examples.
541 examples sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/danilxyz/rest-v3.fragbench-restricted
FragBench (Restricted Tier)
Anonymous submission for NeurIPS 2026 Datasets and Benchmarks Track.
Author identity will be revealed at camera-ready.
Companion public tier
This restricted tier contains only the sensitive components: RL system
prompts, judge-rubric configurations, and high-yield variant traces.
The seed campaigns, generated variants, and benign data are in the public
companion dataset, which is freely accessible without request-access:… See the full description on the dataset page: https://huggingface.co/datasets/anon-fragbench-neurips/fragbench-restricted.restaurant-reviews-timelines
🍽️ Restaurant Reviews with Timelines (Synthetic GPT-4.1 Nano)
Dataset Repository: Programmer-RD-AI/restaurant-reviews-timelines-gpt4nano
📚 Overview
This synthetic dataset comprises over 10,000 restaurant reviews, meticulously generated using OpenAI's GPT-4.1 Nano model. Each review is contextualized within a specific phase of a restaurant's lifecycle, such as:
Opening Hype (Year 1)
Needs Overhaul (Year 4)
New and Improving (Year 2)
Rise and Fall (Year 3)
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/restaurant-reviews-timelines.
