datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zomato-restaurant-recommendationtcd
Dataset Card for OAM-TCD: A globally diverse dataset of high-resolution tree cover maps
Annotation example in OAM-TCD (ID 1445), RGB image licensed CC BY-4.0, attribution contributors of OIN.
Left: RGB aerial image, Middle: annotations shown, distinguished by instance ID, Right: annotations identified by class (blue = tree, orange = canopy)
Dataset Details
OAM-TCD is a dataset of high-resolution (10 cm/px) tree cover maps with instance-level masks for 280k trees and… See the full description on the dataset page: https://huggingface.co/datasets/restor/tcd.mmconflict-restaurant-menus-500
MMConflict Indian and international restaurant menus
This dataset contains 1,000 restaurant menu images. It has 500 Indian menu cards and 500 international historical menus. The menu_region column contains indian or international.
menu_region describes the source collection. It is not a model prediction about a restaurant, cuisine, language, or country.
Sources and licenses
menu_region
Images
Source
Image rights
indian
500
Indian Restaurant Menu Card… See the full description on the dataset page: https://huggingface.co/datasets/shivank21/mmconflict-restaurant-menus-500.clip-worst-restored-realesrganPUUM-koa-restoration-camera-trap-dataset
Dataset Card for Koa Associated Biodiversity Camera Trap Dataset
This dataset is aimed at classification of birds visiting planted Acacia koa (koa) trees in the Pu'u Maka'ala Natural Area Reserve (PUUM) on the island of Hawaii (Big Island). The dataset contains full and cropped images collected by camera trap. These images were collected from January 24th to February 25th, 2025.
Dataset Details
This dataset is aimed at classification of birds visiting planted Acacia… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/PUUM-koa-restoration-camera-trap-dataset.ramanv-image-real-restorationAraMS-Restore
AraMS-Restore — Real Damaged Arabic Manuscript Lines
177 line images cropped from real damaged pages of a historical Arabic
manuscript (book_09), each with its transcription. This is the evaluation
input for AraMS-Restore: the
restoration models are trained on synthetic degradation, and these lines are the
honest test of whether that transfers to genuine manuscript decay.
There are no clean counterparts and no ground-truth restored images — the damage
is what was on the page.… See the full description on the dataset page: https://huggingface.co/datasets/Archatext/AraMS-Restore.django-rest-api-2024taiga_stripped_rest
Dataset Card for "taiga_stripped_rest"
This is a subset of the Taiga corpus (https://tatianashavrina.github.io/taiga_site), derived from the all the sources, except
stihi and proza:
Arzamas, Interfax, Lenta, Magazines, NPlus1, KP, Fontanka, Subtitles and social.
The dataset consists of plain texts, without morphological and syntactic annotation or metainformation.
For the Subtitles subset, we dropped all non-Russian texts.
For the social subset, we split the texts into… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/taiga_stripped_rest.cs_restaurantsThe task is generating responses in the context of a (hypothetical) dialogue
system that provides information about restaurants. The input is a basic
intent/dialogue act type and a list of slots (attributes) and their values.
The output is a natural language sentence.data-voice-vietnamese-restaurant-quan-oc
Vietnamese Restaurant Order Speech
This dataset contains Vietnamese spoken restaurant orders paired with text transcripts. Each utterance typically includes a table number, item quantities, dishes, drinks, and add-ons.
Dataset Structure
Files are split into subdirectories by filename-derived speaker_code to satisfy Hugging Face repository file-count limits:
metadata.csv: one row per audio sample.
audio/{speaker_code}/*.wav: mono WAV audio files.… See the full description on the dataset page: https://huggingface.co/datasets/EmilyNguyen235/data-voice-vietnamese-restaurant-quan-oc.easyr1-110k-bbox0p05-minus-stage3-rl0p2-noise-add-rest-yt-4MP
easyr1-110k-bbox0p05-minus-stage3-rl0p2-noise-add-rest-yt-4MP
Merged dataset composed of the following sources:
datasets/easyr1-103k-bbox0p05-minus-stage3-rl0p2-noise (100155 samples in split train)
mlfoundations-cua-dev/66-yt-app-ui-claude-instructions-no-filter-4MP-gta1-correct-qwen7b-not-correct (7573 samples in split train)
Summary
Generated on: 2025-09-23 15:18:54 UTC
Split: train
Column strategy: intersection
Samples after merge: 107728
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-110k-bbox0p05-minus-stage3-rl0p2-noise-add-rest-yt-4MP.filtered_yelp_restaurant_reviews
Dataset Card for "filtered_yelp_restaurant_reviews"
More Information needed
mit_restaurant[mit_restaurant NER dataset](https://groups.csail.mit.edu/sls/downloads/)diffbir-restored-imagessetfit-absa-semeval-restaurants
Dataset Card for "tomaarsen/setfit-absa-semeval-restaurants"
Dataset Summary
This dataset contains the manually annotated restaurant reviews from SemEval-2014 Task 4, in the format as
understood by SetFit ABSA.
For more details, see https://aclanthology.org/S14-2004/
Data Instances
An example of "train" looks as follows.
{"text": "But the staff was so horrible to us.", "span": "staff", "label": "negative", "ordinal": 0}
{"text": "To be completely fair, the only… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/setfit-absa-semeval-restaurants.us-restaurant-hotel-closings-layoffs-warn-act-notices-daily
US restaurant and hotel closings and layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-22. 6,326 layoff and closure notices filed by
restaurants and restaurant groups, hotels, motels and resorts, casinos and gaming floors, caterers and contract food-service operators, bars and coffee chains, stadium, arena and airport concessions, and leisure and entertainment venues with US state labor departments — 915,329 workers,
3,051 employers, 48 states… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-restaurant-hotel-closings-layoffs-warn-act-notices-daily.bbmad_restructuredThis dataset is only a restructure of the original BMAD dataset. Please cite their work if you use this dataset:
@INPROCEEDINGS{10678042,
author={Bao, Jinan and Sun, Hanshi and Deng, Hanqiu and He, Yinsheng and Zhang, Zhaoxiang and Li, Xingyu},
booktitle={2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)},
title={BMAD: Benchmarks for Medical Anomaly Detection},
year={2024},
volume={},
number={},
pages={4042-4053},
keywords={Computer… See the full description on the dataset page: https://huggingface.co/datasets/acroitoru/bbmad_restructured.cs_restaurants
Dataset Card for Czech Restaurant
Dataset Summary
This is a dataset for NLG in task-oriented spoken dialogue systems with Czech as the target language. It originated as a translation of the English San Francisco Restaurants dataset by Wen et al. (2015). The domain is restaurant information in Prague, with random/fictional values. It includes input dialogue acts and the corresponding outputs in Czech.
Supported Tasks and Leaderboards
other-intent-to-text:… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/cs_restaurants.agibot-sim-clear-table-in-the-restaurantThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "a2d",
"total_episodes": 102,
"total_frames": 84901,
"total_tasks": 1,
"total_videos": 306,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30.0,
"splits": {
"train": "0:102"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bot-pi/agibot-sim-clear-table-in-the-restaurant.agibot-sim-restock-supermarket-itemsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "a2d",
"total_episodes": 1,
"total_frames": 680,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30.0,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bot-pi/agibot-sim-restock-supermarket-items.rest-code-tracingturkish-punctuation-restoration-500k
Turkish Punctuation Restoration 500K v2
Noktalama ve büyük harfleri kaldırılmış girişler ile hedef cümle çiftleri.
Doğrulanmış boyut
Train: 490,000
Validation: 5,000
Test: 5,000
Toplam: 500,000
Ana görev sütunları: id, unpunctuated_text, punctuated_text
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-punctuation-restoration-500k.openr1_dataset_no_restrictionsyelp_restaurant_review_labelled
Dataset Card for "yelp_restaurant_review_labelled"
More Information needed
More info about the dataset
dataset downloaded from Yelp
labelling
if review star < 3 is 0 (negative)else if review star == 3 is 1 (neutral)else if review star > 3 is 2 (positive)
DCAgent2_terminal_bench_2_penfever_nl2bash_verified_gpt-5-nano-traces-restore-hbe5c1c2cVQAv2-COCO-restval-synthetic-capseasyr1-110k-bbox0p05-minus-stage3-rl0p2-noise-add-rest-yt-4MP-remove-pixmo-uground-seeclickDCAgent2_terminal_bench_2_penfever_GLM-4_6-codeforces-32ep-32k-restore-hp_20251a889fa5aqwen3_4b_restricted_Final-activations
