datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swallow-math-v2
SwallowMath-v2
Resources
📑 arXiv: Read our paper for detailed methodology at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowCode2, our companion dataset for code generation.
🧮 What is it?
SwallowMath-v2 is a large-scale mathematical dataset containing 32 billion tokens, developed as the successor to SwallowMath-v1.
Building on the success of v1, this release aims to construct a larger-scale and more permissively licensed corpus to support open and… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math-v2.swallow-code-v2
SwallowCode-v2
Resources
📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning.
💻 What is it?
SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline.
However, it had two significant limitations:
(1) it was distributed under the Llama 3.3 Community License, and
(2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.RAM-W600
Dataset Card for RAM-W600
Benchmark code is available in https://github.com/YSongxiao/RAM-W600.
Download
Please run the following command to download RAM-W600:
git clone https://huggingface.co/datasets/TokyoTechMagicYang/RAM-W600
BoneSegmentation Mask Channel Mapping
The BoneSegmentation masks are stored as 14-channel .npy arrays with shape:
(14, H, W)
Each channel is a binary mask for one anatomical structure. The official channel order is:… See the full description on the dataset page: https://huggingface.co/datasets/TokyoTechMagicYang/RAM-W600.RAM-H1200-v1
[NeurIPS 2026] RAM-H1200
Dataset Summary
RAM-H1200 is a multi-task full-hand radiograph dataset for rheumatoid arthritis (RA) related image analysis. It is designed to support several clinically relevant computer vision tasks, including:
hand bone structure segmentation
bone erosion related segmentation
joint localization for Sharp/van der Heijde (SvdH) scoring
joint-level SvdH bone erosion (BE) scoring
joint-level SvdH joint space narrowing (JSN) scoring
The… See the full description on the dataset page: https://huggingface.co/datasets/TokyoTechMagicYang/RAM-H1200-v1.LiquidAI-Hackathon-Tokyo-CPT-Data
LiquidAI-Hackathon-Tokyo-CPT-Data
Liquid AI Hackathon Tokyoで作成したモデルのCPTに利用したデータセットです。
tokyomewmewnew
Bangumi Image Base of Tokyo Mew Mew New ♡
This is the image base of bangumi Tokyo Mew Mew New ♡, we detected 118 characters, 12865 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/tokyomewmewnew.M-IFEval-JaGST_EGOSERVE
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
This is the PyTorch implementation for VideoRAG proposed in this paper:
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context VideosXubin Ren*, Lingrui Xu*, Long Xia, Shuaiqiang Wang, Dawei Yin, Chao Huang†
* denotes equal contribution.
† denotes corresponding author
In this paper, we proposed a retrieval-augmented generation framework specifically designed for processing and… See the full description on the dataset page: https://huggingface.co/datasets/atad-tokyo/GST_EGOSERVE.swallow-math
SwallowMath
October 21, 2025: Newer versions are available: SwallowCode-v2 and SwallowMath-v2 have been released with improved rewriting pipelines.
Resources
🐙 GitHub: Explore the project repository, including pipeline code and prompts at rioyokotalab/swallow-code-math.
📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowCode, our companion dataset for code generation.
What is it?… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math.swallow-code
SwallowCode
Notice
May 21, 2025: We have deleted ablation/exp1-the-stack-v2-train-smol-ids-python because it was flagged as potentially containing unsafe data collected from the Python subset of https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids. However, since this dataset can be reconstructed from the-stack-v2-train-smol-ids, there is no issue in terms of reproducibility.
May 21, 2025: ClamAV has flagged “Win.Trojan.MSShellcode-88” in… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code.lmsys-chat-1m-synth
LMSYS-Chat-1M-Synth: Japanese/English Synthetic Conversation Dataset Derived from LMSYS-Chat-1M
This repository contains a series of Japanese and English conversation datasets derived from LMSYS-Chat-1M.
Llama-3.1-LMSYS-Chat-1M-Synth
Utilized in the post-training of Llama-3.1-Swallow-8B-Instruct-v0.1 and Llama-3.1-Swallow-70B-Instruct-v0.1
Gemma-2-LMSYS-Chat-1M-Synth
Utilized in the post-training of Llama-3.1-Swallow-8B-Instruct-v0.3 and Llama-3.1-Swallow-70B-Instruct-v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/lmsys-chat-1m-synth.Swallow-Nemotron-Post-Training-Dataset-v1
Swallow-Nemotron-Post-Training-Dataset-v1
The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below.
Dataset Construction
The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528.
However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.MMLU-ProX-JapaneseGST_EGOSCHEMA
Model Card for Model ID
https://showlab.github.io/videollm-online/
Model Details
LLM: meta-llama/Meta-Llama-3-8B-Instruct
Vision Strategy:
Frame Encoder: google/siglip-large-patch16-384
Frame Tokens: CLS Token + Avg Pooled 3x3 Tokens
Frame FPS: 2 for training, 2~10 for inference
Frame Resolution: max resolution 384, with zero-padding to keep aspect ratio
Video Length: 10 minutes
Training Data: Ego4D Narration Stream 113K + Ego4D GoalStep Stream 21K
Model… See the full description on the dataset page: https://huggingface.co/datasets/atad-tokyo/GST_EGOSCHEMA.cxkMMLU-Pro-Englishtokyoghoul
Bangumi Image Base of Tokyo Ghoul
This is the image base of bangumi Tokyo Ghoul, we detected 74 characters, 3651 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/tokyoghoul.GST_Ego4Ddclm-baselinetokyo_u_lsmoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 50,
"total_frames": 11925,
"total_tasks": 2,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/tokyo_u_lsmo.michiyomi-tokyo-streetscape
michiyomi — Tokyo streetscape verbalization open data
English
Overview
michiyomi pairs coordinates with structured Japanese descriptions of physical streetscapes visible in public Mapillary imagery. A vision-language model (VLM) verbalized only what is visible in each image: no map, address, place name, facility name, statistics, or other external knowledge was injected. Release 2026-09-13-r1 contains 1,914,490 scenes covering all of Tokyo: the 23… See the full description on the dataset page: https://huggingface.co/datasets/finalvent/michiyomi-tokyo-streetscape.JEMHopQA
JEMHopQA
このデータセットは SB Intuitions様が公開されている sbintuitions/JEMHopQA を,評価フレームワーク swallow-evaluation-instruct で用いるためにクローンしたものです.
出典
v1, v1.1, v1.2: aiishii/JEMHopQA on GitHub の複製.
v1.[1,2]-extended-answers: SB Intuitions 様が同義語や異表記の別解を追加したもの.
具体的には answer: str が answers: List[str] に変更され,オリジナルの正解および別解が answers に格納されている.
JEMHopQA
JEMHopQA (Japanese Explainable Multi-hop Question Answering) is a Japanese multi-hop QA dataset that can evaluate internal reasoning. It… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/JEMHopQA.MMLU-ProX-Englishtokyo_u_lsmoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 50,
"total_frames": 11925,
"total_tasks": 2,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bot-pi/tokyo_u_lsmo.LiquidAI-Hackathon-Tokyo-SFT-Data
LiquidAI-Hackathon-Tokyo-SFT-Data
Liquid AI Hackathon Tokyoで作成したモデルのSFTに利用したデータセットです。
swallow-magpie-ultra-v0.1
📰 News
[07/01/2025] Release of the first version of the dataset containing 42k Japanese pairs and 42k English pairs.
Dataset Summary
Part of Swallow-Magpie-Ultra-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.2.
The data extracted from magpie-ultra-v0.1 with a quality of average, good, or excellent is… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-magpie-ultra-v0.1.details_tokyotech-llm__Swallow-70b-instruct-hf
Dataset Card for Evaluation run of tokyotech-llm/Swallow-70b-instruct-hf
Dataset automatically created during the evaluation run of model tokyotech-llm/Swallow-70b-instruct-hf on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tokyotech-llm__Swallow-70b-instruct-hf.tokyo_u_lsmo_rawtokyo-koukin-data
東京都 公金支出情報(加工データ)
本データセットは、東京都オープンデータカタログサイトで公開されている「公金支出情報(一般会計・特別会計)」(CC BY 4.0) をもとに作成しています。
出典: 東京都オープンデータカタログサイト
ライセンス: Creative Commons Attribution 4.0 International (CC BY 4.0)
加工者: amoamo1
加工内容: 年度ごとのCSV統合およびJSON変換
備考: 本データは元データの形式を変換したものであり、数値・内容の改変は行っていません。
🔍 データ利用方法
ブラウザやWebアプリ(例: Vercel/Next.js)から以下のように取得できます:
fetch("https://huggingface.co/datasets/amoamo1/tokyo-koukin-data/raw/main/東京都_公金支出情報_令和5年度.json")
.then(res => res.json())… See the full description on the dataset page: https://huggingface.co/datasets/amoamo1/tokyo-koukin-data.swallow_japanese_mt_bench
