CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tokyotech-llm /swallow-math-v2 SwallowMath-v2 Resources 📑 arXiv: Read our paper for detailed methodology at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowCode2, our companion dataset for code generation. 🧮 What is it? SwallowMath-v2 is a large-scale mathematical dataset containing 32 billion tokens, developed as the successor to SwallowMath-v1. Building on the success of v1, this release aims to construct a larger-scale and more permissively licensed corpus to support open and… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math-v2.texttext-generation10M<n<100M35 likes13k downloads11mo agoHugging Face02tokyotech-llm /swallow-code-v2 SwallowCode-v2 Resources 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning. 💻 What is it? SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline. However, it had two significant limitations: (1) it was distributed under the Llama 3.3 Community License, and (2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.tabulartext-generation100M<n<1B48 likes8.3k downloads11mo agoHugging Face03tokyotech-llm /M-IFEval-Jatextn<1K0 likes1k downloads1y agoHugging Face04tokyotech-llm /swallow-math SwallowMath October 21, 2025: Newer versions are available: SwallowCode-v2 and SwallowMath-v2 have been released with improved rewriting pipelines. Resources 🐙 GitHub: Explore the project repository, including pipeline code and prompts at rioyokotalab/swallow-code-math. 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowCode, our companion dataset for code generation. What is it?… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math.texttext-generation1M<n<10M49 likes1k downloads7mo agoHugging Face05tokyotech-llm /swallow-code SwallowCode Notice May 21, 2025: We have deleted ablation/exp1-the-stack-v2-train-smol-ids-python because it was flagged as potentially containing unsafe data collected from the Python subset of https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids. However, since this dataset can be reconstructed from the-stack-v2-train-smol-ids, there is no issue in terms of reproducibility. May 21, 2025: ClamAV has flagged “Win.Trojan.MSShellcode-88” in… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code.tabulartext-generation100M<n<1B71 likes894 downloads7mo agoHugging Face06tokyotech-llm /Swallow-Nemotron-Post-Training-Dataset-v1 Swallow-Nemotron-Post-Training-Dataset-v1 The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below. Dataset Construction The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528. However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.texttext-generation1M<n<10M6 likes768 downloads7mo agoHugging Face07BangumiBase /tokyoghoul Bangumi Image Base of Tokyo Ghoul This is the image base of bangumi Tokyo Ghoul, we detected 74 characters, 3651 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/tokyoghoul.image1K<n<10K1 likes599 downloads3y agoHugging Face08finalvent /michiyomi-tokyo-streetscape michiyomi — Tokyo streetscape verbalization open data English Overview michiyomi pairs coordinates with structured Japanese descriptions of physical streetscapes visible in public Mapillary imagery. A vision-language model (VLM) verbalized only what is visible in each image: no map, address, place name, facility name, statistics, or other external knowledge was injected. Release 2026-09-13-r1 contains 1,914,490 scenes covering all of Tokyo: the 23… See the full description on the dataset page: https://huggingface.co/datasets/finalvent/michiyomi-tokyo-streetscape.tabular1M<n<10M2 likes305 downloads13d agoHugging Face09tokyotech-llm /JEMHopQA JEMHopQA このデータセットは SB Intuitions様が公開されている sbintuitions/JEMHopQA を,評価フレームワーク swallow-evaluation-instruct で用いるためにクローンしたものです. 出典 v1, v1.1, v1.2: aiishii/JEMHopQA on GitHub の複製. v1.[1,2]-extended-answers: SB Intuitions 様が同義語や異表記の別解を追加したもの. 具体的には answer: str が answers: List[str] に変更され,オリジナルの正解および別解が answers に格納されている. JEMHopQA JEMHopQA (Japanese Explainable Multi-hop Question Answering) is a Japanese multi-hop QA dataset that can evaluate internal reasoning. It… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/JEMHopQA.textquestion-answering1K<n<10K0 likes297 downloads1y agoHugging Face10tokyotech-llm /swallow-magpie-ultra-v0.1 📰 News [07/01/2025] Release of the first version of the dataset containing 42k Japanese pairs and 42k English pairs. Dataset Summary Part of Swallow-Magpie-Ultra-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.2. The data extracted from magpie-ultra-v0.1 with a quality of average, good, or excellent is… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-magpie-ultra-v0.1.texttext-generation10K<n<100K5 likes149 downloads2y agoHugging Face11amoamo1 /tokyo-koukin-data 東京都 公金支出情報(加工データ) 本データセットは、東京都オープンデータカタログサイトで公開されている「公金支出情報(一般会計・特別会計)」(CC BY 4.0) をもとに作成しています。 出典: 東京都オープンデータカタログサイト ライセンス: Creative Commons Attribution 4.0 International (CC BY 4.0) 加工者: amoamo1 加工内容: 年度ごとのCSV統合およびJSON変換 備考: 本データは元データの形式を変換したものであり、数値・内容の改変は行っていません。 🔍 データ利用方法 ブラウザやWebアプリ(例: Vercel/Next.js)から以下のように取得できます: fetch("https://huggingface.co/datasets/amoamo1/tokyo-koukin-data/raw/main/東京都_公金支出情報_令和5年度.json") .then(res => res.json())… See the full description on the dataset page: https://huggingface.co/datasets/amoamo1/tokyo-koukin-data.textother1M<n<10M0 likes118 downloads11mo agoHugging Face12tokyotech-llm /swallow_japanese_mt_benchtextn<1K0 likes113 downloads1y agoHugging Face13CRUISEResearchGroup /Massive-STEPS-Tokyo Massive-STEPS-Tokyo Dataset Summary Massive-STEPSis a large-scale dataset of semantic trajectories intended for understanding POI check-ins. The dataset is derived from the Semantic Trails Dataset and Foursquare Open Source Places, and includes check-in data from 15 cities across 10 countries. The dataset is designed to facilitate research in various domains, including trajectory prediction, POI recommendation, and urban modeling. Massive-STEPS emphasizes the… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/Massive-STEPS-Tokyo.tabularother10K<n<100K0 likes102 downloads1y agoHugging Face14tokyotech-llm /Swallow-Instruct-v0.1 Swallow Instruct v0.1 Dataset This dataset was used for supervised fine-tuning (SFT) of the Swallow v0.1 model series. Model Index The following Instruct models were created using this dataset: Llama-3-Swallow-8B-Instruct-v0.1 Llama-3-Swallow-70B-Instruct-v0.1 Swallow-7b-instruct-v0.1 Swallow-13b-instruct-v0.1 Swallow-70b-instruct-v0.1 Note: The data used for Swallow-MS-7b-instruct-v0.1 is different. Statistical Information Dataset Conversations… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Instruct-v0.1.text10K<n<100K12 likes98 downloads2y agoHugging Face15yuiseki /osm-tokyo23-qa-2026-08 osm-tokyo23-qa-2026-08 All 215 answers to osm-tokyo23-questions, computed against the frozen extract in osm-tokyo23-src-2026-08. The questions ship without answers on purpose: an answer belongs to a particular extract on a particular day. This is one such day. Every answer carries the query that produced it. Not a citation of one, the text of one, for each engine that was asked. An answer here is meant to be recomputed rather than believed, and the thing that makes that possible… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-tokyo23-qa-2026-08.textquestion-answeringn<1K0 likes97 downloads3d agoHugging Face16yuiseki /osm-tokyo23-src-2026-08 osm-tokyo23-src-2026-08 A frozen cut of OpenStreetMap covering the 23 special wards of Tokyo, taken from the planet file of 2026-08-31, together with everything needed to rebuild the databases it was measured in. The point is the freezing. A question about a city has an answer only against a stated snapshot, and an answer computed today against the live API is not reproducible tomorrow. Here the snapshot is one file with a checksum, and the tools that read it are pinned by… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-tokyo23-src-2026-08.geospatialtable-question-answering1M<n<10M0 likes95 downloads11d agoHugging Face17yuiseki /osm-tokyo23-questions osm-tokyo23-questions Questions a person would ask about the twenty-three special wards of Tokyo, written to be answered from OpenStreetMap. 215 of them, in twenty kinds. No answers. The set is questions and nothing else. Answers belong to a particular extract on a particular day, and a question carrying its own answer stops being a question. What is recorded instead is which data each question was checked against, so that someone can compute the answers and say what they… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-tokyo23-questions.textquestion-answeringn<1K0 likes78 downloads3d agoHugging Face18tokyotech-llm /s1-test-time-scaling-synth-public s1-test-time-scaling-synth: Japanese and English Reinforcement Learning Dataset Derived from the s1 Simple Test-Time Scaling Dataset This repository contains s1-test-time-scaling-synth, a reinforcement learning dataset in Japanese and English.This dataset is built upon the supervised fine-tuning dataset simplescaling/data_ablation_full59K (hereafter, the "original dataset"), originally developed in "s1: Simple test-time scaling" [Muennighoff+, EMNLP25]. The original dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/s1-test-time-scaling-synth-public.texttext-generation10K<n<100K0 likes69 downloads7mo agoHugging Face19yuiseki /osm-wikidata-brand-tokyo23 osm-wikidata-brand-tokyo23 Every way two independent sources name the same shop chain, in the twenty-three special wards of Tokyo, and the name of every shop in those chains. 892 brands, 25,407 features, 15,730 distinct values across 82 keys. This is the second version. The first read the extract through PostGIS, where osm2pgsql had promoted name and brand to columns of their own, so the hstore column the build queried held every name key except those two. They carry 6,529 of… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-wikidata-brand-tokyo23.texttext-classificationn<1K0 likes67 downloads9d agoHugging Face20tokyotech-llm /swallow_english_mt_benchtextn<1K0 likes57 downloads1y agoHugging Face21tokyotech-llm /swallow-gemma-magpie-v0.1 📰 News [07/01/2025] Release of the first unfiltered version of the dataset containing 148k pairs. Dataset Summary Swallow-Gemma-Magpie-v0.1 is a synthetic instruction tuning dataset that consists of multiple category Japanese question-answering tasks. It consists of 148k question-answering-samples, generated with google/gemma-2-27b-it. Part of Swallow-Gemma-Magpie-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-gemma-magpie-v0.1.texttext-generation100K<n<1M3 likes54 downloads2y agoHugging Face22Tokyomonster /JBB-Behaviors An Open Robustness Benchmark for Jailbreaking Language Models NeurIPS 2024 Datasets and Benchmarks Track Paper | Leaderboard | Benchmark code What is JailbreakBench? Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/Tokyomonster/JBB-Behaviors.tabularn<1K0 likes53 downloads6mo agoHugging Face23open-llm-leaderboard /tokyotech-llm__Llama-3-Swallow-8B-Instruct-v0.1-detailsgated Dataset Card for Evaluation run of tokyotech-llm/Llama-3-Swallow-8B-Instruct-v0.1 Dataset automatically created during the evaluation run of model tokyotech-llm/Llama-3-Swallow-8B-Instruct-v0.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tokyotech-llm__Llama-3-Swallow-8B-Instruct-v0.1-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face24blstweb0901 /tokyo-vpn-monitor language: ja en license: mit multilinguality: multilingual size_categories: 1K<n<10K source_datasets: original task_categories: other task_ids: [] pretty_name: Tokyo VPN Speed Monitor Dataset tags: vpn network-monitoring performance-measurement time-series networking internet-measurement automated-testing zero-cost-infrastructure google-apps-script Tokyo VPN Speed Monitor Dataset Dataset Summary The Tokyo VPN Speed Monitor Dataset contains continuous automated… See the full description on the dataset page: https://huggingface.co/datasets/blstweb0901/tokyo-vpn-monitor.tabulartabular-classification1K<n<10K0 likes23 downloads9mo agoHugging Face25Mr-Bridge /paris-vs-tokyo-hotels-2026 Paris vs Tokyo Hotels 2026: Stars & Guest Ratings Hotels in Paris and Tokyo (3 to 5 star), each with hotel name, platform star level and a blended guest rating (Booking, Priceline, Agoda, HotelsCombined). Collected 2026-06-22 by MrBridge. Files kayak_paris_tokyo_2026.csv — 197 hotels (Paris 98, Tokyo 99), the primary blended-OTA source. priceline_paris_tokyo_2026.csv — 62 hotels, an independent Priceline pull used as a robustness check. Columns: city, name… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/paris-vs-tokyo-hotels-2026.tabulartabular-classificationn<1K0 likes18 downloads3mo agoHugging Face26Mr-Bridge /priceline-live-pricing-snapshots-tokyo-2026 Tokyo Hotels 2026: Price Snapshots Open dataset by MrBridge — 278 rows of hotel data (ratings, prices and attributes), free to download, inspect and reuse. Data source Collected with the Priceline Hotel Scraper on Apify — a cloud actor returning clean, structured CSV/JSON with no local setup. Re-run it yourself for fresh data. Get fresh data → https://apify.com/mrbridge/priceline-hotel-scraper?fpr=mrbridge Schema data.csv — 278 rows, 17 columns:… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/priceline-live-pricing-snapshots-tokyo-2026.tabulartabular-regressionn<1K0 likes14 downloads3mo agoHugging Face27tokyotech-lrlab /ptb_no_empty_parsedtextn<1K0 likes13 downloads2y agoHugging Face28jbrazzy /tokyo_renttabularn<1K0 likes10 downloads3y agoHugging Face29yiweifu /relearn_data_tokyo.hftextn<1K0 likes10 downloads2y agoHugging Face30atad-tokyo /GST_EgoLifetext10K<n<100K0 likes10 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.