datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3_4b_restricted_Final-activationswhatscooking.restaurants
Whatscooking.restaurants
Overview
This dataset provides detailed information about various restaurants, including their location, cuisine, ratings, and other attributes. It is particularly useful for applications in food and beverage industry analysis, recommendation systems, and geographical studies.
Dataset Structure
Each record in the dataset represents a single restaurant and contains the following fields:
_id: A unique identifier for the restaurant… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/whatscooking.restaurants.2026-08-03-qwen36-27b-armb-1000ex-difficult-advice-250-rest-750-train-mixture
Qwen3.6-27B training bundle — 2026-08-03-qwen36-27b-armB-1000ex-da250-rest750-train
code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies
the jsonl to data/mixture.jsonl, and runs configs/train_armB_1000ex_da250_rest750.yaml.
field
value
experiment
Arm B: 250 difficult-advice (t1-t3) + 750 at 3:2 NuminaMath : (TULU3 + No Robots)
date_generated
2026-08-03
constitution
constitutions/claude_constitution_principles.md — principles t1-t3… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-armb-1000ex-difficult-advice-250-rest-750-train-mixture.Pref-Restore-Dataqsim-shelf-restocking-mujoco-300
Qsim MuJoCo Shelf Restocking
300 seeded synthetic shelf-restocking episodes with the stock bimanual OpenArm
robot, physical gripper contacts, smooth scripted motion, and synchronized robot
head and wrist camera observations.
Release status: complete — 300/300 episodes passed the task and dataset checks.
600 placements; 542,518 observation/action samples; 108,640 RGB frames per camera; 3.01 hours of simulated trajectories; 9.68 GB of HDF5 episodes.
Arm selections: {'left': 297… See the full description on the dataset page: https://huggingface.co/datasets/qualiadev/qsim-shelf-restocking-mujoco-300.YOYO-AI__Qwen2.5-14B-it-restore-details
Dataset Card for Evaluation run of YOYO-AI/Qwen2.5-14B-it-restore
Dataset automatically created during the evaluation run of model YOYO-AI/Qwen2.5-14B-it-restore
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/YOYO-AI__Qwen2.5-14B-it-restore-details.restaurant-reviews
Synthetic Dataset for Product Descriptions and Ads
The basic process was as follows:
Prompt GPT-4 to create a list of 100 sample clothing items and descriptions for those items.
Split the output into desired format `{"product" : "", "description" : ""}
Prompt GPT-4 to create adverts for each of the 100 samples based on their name and description.
This data was not cleaned or verified manually.
ReST-MCTS-PRM-0thshipping-restrictions-by-class
Mailing and air-carriage restrictions by item category
Canonical, always-current version: https://referencesource.org/shipping-restrictions-by-class/
Machine-readable: https://referencesource.org/shipping-restrictions-by-class/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-10
Stale after: 2027-08-10 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 2852
What item categories may be mailed, and what… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/shipping-restrictions-by-class.state-pfas-product-restrictions
US state PFAS product restriction effective dates by category
Canonical, always-current version: https://referencesource.org/state-pfas-product-restrictions/
Machine-readable: https://referencesource.org/state-pfas-product-restrictions/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-14
Stale after: 2026-11-12 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 101
State-by-state PFAS product ban and… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/state-pfas-product-restrictions.eval-baselines-mit_restaurant
mit_restaurant — NER baselines (Stanford CoreNLP CRF, DeBERTa-v3-base, DistilBERT-base)
dataset: quynong/mit_restaurant (8 nhan), eval tren split test
HF baselines: 5 epochs, batch 16, seeds [42, 43, 44]
corenlp: 1 lan train (CRF deterministic, khong seed)
Ket qua (P/R/F1 %, HungarianEvaluator IoU>=0.5 hoac fuzzy)
baseline
n
micro F1
macro F1
micro P
micro R
corenlp
0
84.95
83.67
87.17
82.83
deberta-v3-base
3
88.39 +/- 0.56
87.49 +/- 0.84
86.75
90.1… See the full description on the dataset page: https://huggingface.co/datasets/AITeamUIT/eval-baselines-mit_restaurant.rest-v3
rest-v3
rest-v3 is an English text-rewriting dataset for supervised fine-tuning of a humanizing editor. Each record asks a model to rewrite a source text while preserving its meaning and contains a detector-verified natural-language rewrite.
Dataset composition
The training split contains 1,116 JSONL records:
1,033 newly mined, on-policy rewrites from the from-final-best generator checkpoint.
83 compatible existing verified examples.
541 examples sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/danilxyz/rest-v3.nyc-restaurant-artifactsReST-MCTS_SciGLM-6B_ReST-MCTS_Policy_2ndmit-restaurantconversational-question-answer-wikipedia-v1.0
Dataset Information
A dataset containing questions and conversational answers, based on sections of Wikipedia articles from wikipedia-en-chunks. This is a synthetic dataset created with the help of gemini-2.0-flash-001.
Dataset Structure
The dataset consists of a single JSON file with the following structure:
[
{
"messages": [
{
"role": "system",
"content": "You are a helpful assistant. You answer questions in a… See the full description on the dataset page: https://huggingface.co/datasets/restack/conversational-question-answer-wikipedia-v1.0.YOYO-AI__Qwen2.5-7B-it-restore-details
Dataset Card for Evaluation run of YOYO-AI/Qwen2.5-7B-it-restore
Dataset automatically created during the evaluation run of model YOYO-AI/Qwen2.5-7B-it-restore
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/YOYO-AI__Qwen2.5-7B-it-restore-details.ReST-MCTS-Llama3-8b-Instruct-PRM-1stmit-restaurantnemotron_restricted_Final-activationsReST-MCTS_Mistral-MetaMATH-7b-Instruct_ReST-EM-CoT_2ndrestor_punct_Lenta2
annotations_creators:
machine-generated
language:
ru
language_creators:
machine-generated
license:
afl-3.0
multilinguality: []
pretty_name: Dmitriy007/restor_punct_Lenta2
size_categories:
100K<n<1M
source_datasets:
original
tags: []
task_categories:
token-classification
task_ids: []
Dataset Card for Dmitriy007/restor_punct_Lenta2
Dataset Summary
Набор данных restor_punct_Lenta2 (версия 2.0) представляет собой набор из 800 975 блоков русскоязычных предложений, разбитых… See the full description on the dataset page: https://huggingface.co/datasets/Dmitriy007/restor_punct_Lenta2.ReST-MCTS-Llama3-8b-Instruct-Policy-1stReST-MCTS_Mistral-MetaMATH-7b-Instruct_ReST-MCTS_1stReST-MCTS_Llama3-8b-Instruct_ReST-EM-CoT_2ndhebrew-space-restoration-corpus
Restoring Missing Spaces in Scraped Hebrew Social Media
This dataset holds the test corpus used in the 2025 W-Nut paper: Avi Shmidman and Shaltiel Shmidman, "Restoring Missing Spaces in Scraped Hebrew Social Media", The 10th Workshop on Noisy and User-generated Text (W-NUT), 2025.
The corpus consists of ~6,000 Hebrew sentences, sampled from the Hebrew portion of FineWeb-2.
Each row of the dataset contains two fields:
input: The Hebrew sentence with 1-4 spaces randomly removed (see… See the full description on the dataset page: https://huggingface.co/datasets/dicta-il/hebrew-space-restoration-corpus.artofwar_restructured_datasetru-absa-restaurant-reviewsyandex_geo_reviews_restaraunt_instructДатасет составлен из части датасета Geo Reviews Dataset 2023 https://github.com/yandex/geo-reviews-dataset-2023 , взяты только отзывы по ресторанам, и переделан формат датасета для LLM
ReST-MCTS_Mistral-MetaMATH-7b-Instruct_ReST-MCTS_2nd
