datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swallow-math-v2
SwallowMath-v2
Resources
📑 arXiv: Read our paper for detailed methodology at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowCode2, our companion dataset for code generation.
🧮 What is it?
SwallowMath-v2 is a large-scale mathematical dataset containing 32 billion tokens, developed as the successor to SwallowMath-v1.
Building on the success of v1, this release aims to construct a larger-scale and more permissively licensed corpus to support open and… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math-v2.swallow-code-v2
SwallowCode-v2
Resources
📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning.
💻 What is it?
SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline.
However, it had two significant limitations:
(1) it was distributed under the Llama 3.3 Community License, and
(2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.M-IFEval-Jaswallow-math
SwallowMath
October 21, 2025: Newer versions are available: SwallowCode-v2 and SwallowMath-v2 have been released with improved rewriting pipelines.
Resources
🐙 GitHub: Explore the project repository, including pipeline code and prompts at rioyokotalab/swallow-code-math.
📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowCode, our companion dataset for code generation.
What is it?… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math.swallow-code
SwallowCode
Notice
May 21, 2025: We have deleted ablation/exp1-the-stack-v2-train-smol-ids-python because it was flagged as potentially containing unsafe data collected from the Python subset of https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids. However, since this dataset can be reconstructed from the-stack-v2-train-smol-ids, there is no issue in terms of reproducibility.
May 21, 2025: ClamAV has flagged “Win.Trojan.MSShellcode-88” in… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code.Swallow-Nemotron-Post-Training-Dataset-v1
Swallow-Nemotron-Post-Training-Dataset-v1
The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below.
Dataset Construction
The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528.
However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.tokyoghoul
Bangumi Image Base of Tokyo Ghoul
This is the image base of bangumi Tokyo Ghoul, we detected 74 characters, 3651 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/tokyoghoul.michiyomi-tokyo-streetscape
michiyomi — Tokyo streetscape verbalization open data
English
Overview
michiyomi pairs coordinates with structured Japanese descriptions of physical streetscapes visible in public Mapillary imagery. A vision-language model (VLM) verbalized only what is visible in each image: no map, address, place name, facility name, statistics, or other external knowledge was injected. Release 2026-09-13-r1 contains 1,914,490 scenes covering all of Tokyo: the 23… See the full description on the dataset page: https://huggingface.co/datasets/finalvent/michiyomi-tokyo-streetscape.JEMHopQA
JEMHopQA
このデータセットは SB Intuitions様が公開されている sbintuitions/JEMHopQA を,評価フレームワーク swallow-evaluation-instruct で用いるためにクローンしたものです.
出典
v1, v1.1, v1.2: aiishii/JEMHopQA on GitHub の複製.
v1.[1,2]-extended-answers: SB Intuitions 様が同義語や異表記の別解を追加したもの.
具体的には answer: str が answers: List[str] に変更され,オリジナルの正解および別解が answers に格納されている.
JEMHopQA
JEMHopQA (Japanese Explainable Multi-hop Question Answering) is a Japanese multi-hop QA dataset that can evaluate internal reasoning. It… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/JEMHopQA.swallow-magpie-ultra-v0.1
📰 News
[07/01/2025] Release of the first version of the dataset containing 42k Japanese pairs and 42k English pairs.
Dataset Summary
Part of Swallow-Magpie-Ultra-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.2.
The data extracted from magpie-ultra-v0.1 with a quality of average, good, or excellent is… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-magpie-ultra-v0.1.tokyo-koukin-data
東京都 公金支出情報(加工データ)
本データセットは、東京都オープンデータカタログサイトで公開されている「公金支出情報(一般会計・特別会計)」(CC BY 4.0) をもとに作成しています。
出典: 東京都オープンデータカタログサイト
ライセンス: Creative Commons Attribution 4.0 International (CC BY 4.0)
加工者: amoamo1
加工内容: 年度ごとのCSV統合およびJSON変換
備考: 本データは元データの形式を変換したものであり、数値・内容の改変は行っていません。
🔍 データ利用方法
ブラウザやWebアプリ(例: Vercel/Next.js)から以下のように取得できます:
fetch("https://huggingface.co/datasets/amoamo1/tokyo-koukin-data/raw/main/東京都_公金支出情報_令和5年度.json")
.then(res => res.json())… See the full description on the dataset page: https://huggingface.co/datasets/amoamo1/tokyo-koukin-data.swallow_japanese_mt_benchMassive-STEPS-Tokyo
Massive-STEPS-Tokyo
Dataset Summary
Massive-STEPSis a large-scale dataset of semantic trajectories intended for understanding POI check-ins. The dataset is derived from the Semantic Trails Dataset and Foursquare Open Source Places, and includes check-in data from 15 cities across 10 countries. The dataset is designed to facilitate research in various domains, including trajectory prediction, POI recommendation, and urban modeling. Massive-STEPS emphasizes the… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/Massive-STEPS-Tokyo.Swallow-Instruct-v0.1
Swallow Instruct v0.1 Dataset
This dataset was used for supervised fine-tuning (SFT) of the Swallow v0.1 model series.
Model Index
The following Instruct models were created using this dataset:
Llama-3-Swallow-8B-Instruct-v0.1
Llama-3-Swallow-70B-Instruct-v0.1
Swallow-7b-instruct-v0.1
Swallow-13b-instruct-v0.1
Swallow-70b-instruct-v0.1
Note: The data used for Swallow-MS-7b-instruct-v0.1 is different.
Statistical Information
Dataset
Conversations… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Instruct-v0.1.osm-tokyo23-qa-2026-08
osm-tokyo23-qa-2026-08
All 215 answers to osm-tokyo23-questions,
computed against the frozen extract in
osm-tokyo23-src-2026-08.
The questions ship without answers on purpose: an answer belongs to a
particular extract on a particular day. This is one such day.
Every answer carries the query that produced it. Not a citation of one, the
text of one, for each engine that was asked. An answer here is meant to be
recomputed rather than believed, and the thing that makes that possible… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-tokyo23-qa-2026-08.osm-tokyo23-src-2026-08
osm-tokyo23-src-2026-08
A frozen cut of OpenStreetMap covering the 23 special wards of Tokyo, taken
from the planet file of 2026-08-31, together with everything needed to
rebuild the databases it was measured in.
The point is the freezing. A question about a city has an answer only against
a stated snapshot, and an answer computed today against the live API is not
reproducible tomorrow. Here the snapshot is one file with a checksum, and the
tools that read it are pinned by… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-tokyo23-src-2026-08.osm-tokyo23-questions
osm-tokyo23-questions
Questions a person would ask about the twenty-three special wards of Tokyo,
written to be answered from OpenStreetMap. 215 of them, in twenty kinds.
No answers. The set is questions and nothing else. Answers belong to a
particular extract on a particular day, and a question carrying its own answer
stops being a question. What is recorded instead is which data each question
was checked against, so that someone can compute the answers and say what they… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-tokyo23-questions.s1-test-time-scaling-synth-public
s1-test-time-scaling-synth: Japanese and English Reinforcement Learning Dataset Derived from the s1 Simple Test-Time Scaling Dataset
This repository contains s1-test-time-scaling-synth, a reinforcement learning dataset in Japanese and English.This dataset is built upon the supervised fine-tuning dataset simplescaling/data_ablation_full59K (hereafter, the "original dataset"), originally developed in "s1: Simple test-time scaling" [Muennighoff+, EMNLP25].
The original dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/s1-test-time-scaling-synth-public.osm-wikidata-brand-tokyo23
osm-wikidata-brand-tokyo23
Every way two independent sources name the same shop chain, in the
twenty-three special wards of Tokyo, and the name of every shop in those
chains.
892 brands, 25,407 features, 15,730 distinct values across 82 keys.
This is the second version. The first read the extract through PostGIS, where
osm2pgsql had promoted name and brand to columns of their own, so the
hstore column the build queried held every name key except those two. They
carry 6,529 of… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-wikidata-brand-tokyo23.swallow_english_mt_benchswallow-gemma-magpie-v0.1
📰 News
[07/01/2025] Release of the first unfiltered version of the dataset containing 148k pairs.
Dataset Summary
Swallow-Gemma-Magpie-v0.1 is a synthetic instruction tuning dataset that consists of multiple category Japanese question-answering tasks.
It consists of 148k question-answering-samples, generated with google/gemma-2-27b-it.
Part of Swallow-Gemma-Magpie-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-gemma-magpie-v0.1.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/Tokyomonster/JBB-Behaviors.tokyotech-llm__Llama-3-Swallow-8B-Instruct-v0.1-details
Dataset Card for Evaluation run of tokyotech-llm/Llama-3-Swallow-8B-Instruct-v0.1
Dataset automatically created during the evaluation run of model tokyotech-llm/Llama-3-Swallow-8B-Instruct-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tokyotech-llm__Llama-3-Swallow-8B-Instruct-v0.1-details.tokyo-vpn-monitor
language:
ja
en
license: mit
multilinguality:
multilingual
size_categories:
1K<n<10K
source_datasets:
original
task_categories:
other
task_ids: []
pretty_name: Tokyo VPN Speed Monitor Dataset
tags:
vpn
network-monitoring
performance-measurement
time-series
networking
internet-measurement
automated-testing
zero-cost-infrastructure
google-apps-script
Tokyo VPN Speed Monitor Dataset
Dataset Summary
The Tokyo VPN Speed Monitor Dataset contains continuous automated… See the full description on the dataset page: https://huggingface.co/datasets/blstweb0901/tokyo-vpn-monitor.paris-vs-tokyo-hotels-2026
Paris vs Tokyo Hotels 2026: Stars & Guest Ratings
Hotels in Paris and Tokyo (3 to 5 star), each with hotel name, platform star level and a blended guest rating (Booking, Priceline, Agoda, HotelsCombined). Collected 2026-06-22 by MrBridge.
Files
kayak_paris_tokyo_2026.csv — 197 hotels (Paris 98, Tokyo 99), the primary blended-OTA source.
priceline_paris_tokyo_2026.csv — 62 hotels, an independent Priceline pull used as a robustness check.
Columns: city, name… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/paris-vs-tokyo-hotels-2026.priceline-live-pricing-snapshots-tokyo-2026
Tokyo Hotels 2026: Price Snapshots
Open dataset by MrBridge — 278 rows of hotel data (ratings, prices and attributes), free to download, inspect and reuse.
Data source
Collected with the Priceline Hotel Scraper on Apify — a cloud actor returning clean, structured CSV/JSON with no local setup. Re-run it yourself for fresh data.
Get fresh data → https://apify.com/mrbridge/priceline-hotel-scraper?fpr=mrbridge
Schema
data.csv — 278 rows, 17 columns:… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/priceline-live-pricing-snapshots-tokyo-2026.ptb_no_empty_parsedtokyo_rentrelearn_data_tokyo.hfGST_EgoLife
