datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swallow-code-v2
SwallowCode-v2
Resources
📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning.
💻 What is it?
SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline.
However, it had two significant limitations:
(1) it was distributed under the Llama 3.3 Community License, and
(2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.swallow-code
SwallowCode
Notice
May 21, 2025: We have deleted ablation/exp1-the-stack-v2-train-smol-ids-python because it was flagged as potentially containing unsafe data collected from the Python subset of https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids. However, since this dataset can be reconstructed from the-stack-v2-train-smol-ids, there is no issue in terms of reproducibility.
May 21, 2025: ClamAV has flagged “Win.Trojan.MSShellcode-88” in… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code.tokyo_u_lsmoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 50,
"total_frames": 11925,
"total_tasks": 2,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/tokyo_u_lsmo.michiyomi-tokyo-streetscape
michiyomi — Tokyo streetscape verbalization open data
English
Overview
michiyomi pairs coordinates with structured Japanese descriptions of physical streetscapes visible in public Mapillary imagery. A vision-language model (VLM) verbalized only what is visible in each image: no map, address, place name, facility name, statistics, or other external knowledge was injected. Release 2026-09-13-r1 contains 1,914,490 scenes covering all of Tokyo: the 23… See the full description on the dataset page: https://huggingface.co/datasets/finalvent/michiyomi-tokyo-streetscape.tokyo_u_lsmoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 50,
"total_frames": 11925,
"total_tasks": 2,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bot-pi/tokyo_u_lsmo.Massive-STEPS-Tokyo
Massive-STEPS-Tokyo
Dataset Summary
Massive-STEPSis a large-scale dataset of semantic trajectories intended for understanding POI check-ins. The dataset is derived from the Semantic Trails Dataset and Foursquare Open Source Places, and includes check-in data from 15 cities across 10 countries. The dataset is designed to facilitate research in various domains, including trajectory prediction, POI recommendation, and urban modeling. Massive-STEPS emphasizes the… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/Massive-STEPS-Tokyo.osm-tokyo23-src-2026-08
osm-tokyo23-src-2026-08
A frozen cut of OpenStreetMap covering the 23 special wards of Tokyo, taken
from the planet file of 2026-08-31, together with everything needed to
rebuild the databases it was measured in.
The point is the freezing. A question about a city has an answer only against
a stated snapshot, and an answer computed today against the live API is not
reproducible tomorrow. Here the snapshot is one file with a checksum, and the
tools that read it are pinned by… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-tokyo23-src-2026-08.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/Tokyomonster/JBB-Behaviors.tokyotech-llm__Llama-3-Swallow-8B-Instruct-v0.1-details
Dataset Card for Evaluation run of tokyotech-llm/Llama-3-Swallow-8B-Instruct-v0.1
Dataset automatically created during the evaluation run of model tokyotech-llm/Llama-3-Swallow-8B-Instruct-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tokyotech-llm__Llama-3-Swallow-8B-Instruct-v0.1-details.tokyo-vpn-monitor
language:
ja
en
license: mit
multilinguality:
multilingual
size_categories:
1K<n<10K
source_datasets:
original
task_categories:
other
task_ids: []
pretty_name: Tokyo VPN Speed Monitor Dataset
tags:
vpn
network-monitoring
performance-measurement
time-series
networking
internet-measurement
automated-testing
zero-cost-infrastructure
google-apps-script
Tokyo VPN Speed Monitor Dataset
Dataset Summary
The Tokyo VPN Speed Monitor Dataset contains continuous automated… See the full description on the dataset page: https://huggingface.co/datasets/blstweb0901/tokyo-vpn-monitor.paris-vs-tokyo-hotels-2026
Paris vs Tokyo Hotels 2026: Stars & Guest Ratings
Hotels in Paris and Tokyo (3 to 5 star), each with hotel name, platform star level and a blended guest rating (Booking, Priceline, Agoda, HotelsCombined). Collected 2026-06-22 by MrBridge.
Files
kayak_paris_tokyo_2026.csv — 197 hotels (Paris 98, Tokyo 99), the primary blended-OTA source.
priceline_paris_tokyo_2026.csv — 62 hotels, an independent Priceline pull used as a robustness check.
Columns: city, name… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/paris-vs-tokyo-hotels-2026.priceline-live-pricing-snapshots-tokyo-2026
Tokyo Hotels 2026: Price Snapshots
Open dataset by MrBridge — 278 rows of hotel data (ratings, prices and attributes), free to download, inspect and reuse.
Data source
Collected with the Priceline Hotel Scraper on Apify — a cloud actor returning clean, structured CSV/JSON with no local setup. Re-run it yourself for fresh data.
Get fresh data → https://apify.com/mrbridge/priceline-hotel-scraper?fpr=mrbridge
Schema
data.csv — 278 rows, 17 columns:… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/priceline-live-pricing-snapshots-tokyo-2026.tokyo_rentdetails_tokyotech-llm__Llama-3-Swallow-8B-Instruct-v0.1
Dataset Card for Evaluation run of tokyotech-llm/Llama-3-Swallow-8B-Instruct-v0.1
Dataset automatically created during the evaluation run of model tokyotech-llm/Llama-3-Swallow-8B-Instruct-v0.1.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_tokyotech-llm__Llama-3-Swallow-8B-Instruct-v0.1.tokyo-rent-predictor-dataUrban_Tokyo_Temperature
Urban Tokyo Temperature Dataset
This dataset contains daily weather observations for the 23 special wards of Tokyo, Japan (2010–2020).
Columns
date: Date of observation
temp_c: Mean temperature (°C)
dew_c: Dew point (°C)
humidity: Relative humidity (%)
wind_kph: Wind speed (km/h)
pressure_hg: Atmospheric pressure (inHg)
precip_mm: Precipitation (mm)
district: Tokyo ward (23 classes)
year: Year
month: Month (string)
Size
Rows: 92,414
File size: 6.23 MB… See the full description on the dataset page: https://huggingface.co/datasets/torodriguezt/Urban_Tokyo_Temperature.
