datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nyc-taxi
NYC Taxi Trip Dataset
This dataset contains NYC taxi trip data from May 1-7, 2013, excluding trips to and from Staten Island. It includes 2,957 sequences with 362,374 events and 8 location types. The data can be downloaded from NYC Taxi Trips and is subject to the NYC Terms of Use. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper.
Update (2025-10-28): Added three timestamp fields (timestamp_event… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/nyc-taxi.nyc-taxi-description
NYC Taxi Trip Description Dataset
This dataset contains NYC taxi trip data from May 1-7, 2013, excluding trips to and from Staten Island. It includes 2,957 sequences with 362,374 events and 8 location types. The data can be downloaded from NYC Taxi Trips and is subject to the NYC Terms of Use. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper.
If you find this dataset useful, we kindly invite you to cite the… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/nyc-taxi-description.nyc-housing-rights
NYC Housing Rights Q&A Dataset
A comprehensive dataset of 597 question-answer pairs focused on NYC tenant rights, housing laws, and legal assistance resources. This dataset was specifically designed to train AI assistants for accurate tenant rights information with 100% accuracy on critical housing questions.
Dataset Overview
Dataset Statistics
Total Examples: 597 question-answer pairs
Expansion Factor: 33.2x from original 18 examples
Categories: 15+… See the full description on the dataset page: https://huggingface.co/datasets/aanshshah/nyc-housing-rights.NYCnyc-restaurant-artifactsNYC_sensitive_sitesTurtleBench-extended-zh
海龜湯數據集(中文)
本數據集包含中文的海龜湯謎題,用於逆向思維遊戲。
數據集簡介
本數據集基於 Duguce/TurtleBench1.5k 擴展而來,旨在為 Turtle-soup Game 提供高質量的推理數據。數據涵蓋多種高難度推理情境,支持 中文 與 英文 兩種語言,並結合多種擴增方法提升多樣性與邏輯性。
數據來源
原始數據集來自 Hugging Face,授權於 Apache License 2.0。
擴增後數據集由翻譯、標註、基準題庫及模型生成數據構成,詳細分布請見下文。
數據結構
數據集包含以下字段,每筆數據均完整記錄了一個海龜湯故事的推理情境與答案標籤:
屬性名稱
描述
id
故事的唯一標識符。
title
海龜湯故事的標題。
surface
海龜湯故事的表層信息,即玩家能夠直接獲得的線索。
bottom
海龜湯故事的深層背景,即玩家需要推理才能獲知的答案或情境。
user_guess
玩家對故事的假設或猜測。
label… See the full description on the dataset page: https://huggingface.co/datasets/nycu-ai113-dl-final-project/TurtleBench-extended-zh.uber_pickups_nyc_stpp
Uber Pickups NYC STPP Benchmark Dataset
A benchmark-ready Spatio-Temporal Point Process (STPP) dataset derived from Uber Pickups (NYC) (~4.5 Million records), following the standard split semantics for Neural STPP evaluation.
Dataset Description
Each record represents a sequence of events. The dataset covers historical Uber pickups across NYC, partitioned sequentially into train / val / test subsets (70% / 15% / 15% ratio).
Source Format
Raw data… See the full description on the dataset page: https://huggingface.co/datasets/seahorse-stpp/uber_pickups_nyc_stpp.vc-nycnycESVjfiFBpipKCqa-dataset-20250805nyc-diabetes-screenings
NYC Free Diabetes Screenings
A living directory of free diabetes screening locations and events across New York City, maintained by LocalDevs.
data/screenings.jsonl is rebuilt weekly by a scraper + geocoding pipeline that merges scraped sources (e.g. EmblemHealth's community events) with hand-maintained static entries (e.g. Columbia's Center for Community Health). Each record includes name, address, borough, lat/lon, hours, screening types offered, appointment type, and a… See the full description on the dataset page: https://huggingface.co/datasets/LocalDevs/nyc-diabetes-screenings.NYCU_RAWTurtleBench-extended-en
Turtle Soup Dataset (English)
This dataset contains English Turtle Soup puzzles, designed for reverse thinking games.
Dataset Overview
This dataset is an extension of Duguce/TurtleBench1.5k, aiming to provide high-quality reasoning data for Turtle-soup Game. The data covers various high-difficulty reasoning scenarios, supports both Chinese and English, and incorporates multiple augmentation methods to enhance diversity and logical consistency.
Data Source… See the full description on the dataset page: https://huggingface.co/datasets/nycu-ai113-dl-final-project/TurtleBench-extended-en.ZH-TW_Reading_Comprehension_Test_for_LLMsNYC_ZipCodesnyc_persona_mqa
nyc_persona_mqa Dataset Overview
nyc_persona_mqa.json captures 65,115 narrated day-in-the-life summaries from New York City visitor and resident trajectories, each annotated with two persona hypotheses drafted by GPT-5 based on observed movement patterns and nearby POIs.
Field Definitions
user_id: Identifier that links the narrative back to a specific NYC trajectory instance.
text: English summary describing hourly activities inferred from GPS traces and contextual POI… See the full description on the dataset page: https://huggingface.co/datasets/bigchestnut/nyc_persona_mqa.NYC-1800-1875-sampleNYCU_QADM_HW2_datasetNYC_ExplorationNYC-dataPersonal test.
nyc-restaurant-artifactsNYCU_gaonyc_slot_daily_trajectories_gpt5nycufoursquare_nyc_personanyc_slot_daily_trajectories
Dataset Card for NYC Daily Trajectories in 2hour time slot
Dataset Summary
nyc_daily_trajectories_sae.json contains 65,115 fixed-length daily trajectories derived from the TSMC2014 NYC Foursquare check-ins. Each record aggregates one user-day of check-ins at a 2-hour granularity (00:00–22:00) so the data is immediately consumable by sparse autoencoders and other sequential encoders that expect constant-width vectors. Locations are rounded to 0.001 degrees and missing… See the full description on the dataset page: https://huggingface.co/datasets/bigchestnut/nyc_slot_daily_trajectories.NYC_sensitive_sitesnycu-llm2-reasoning-sft-results
