CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tokyotech-llm /swallow-math-v2 SwallowMath-v2 Resources 📑 arXiv: Read our paper for detailed methodology at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowCode2, our companion dataset for code generation. 🧮 What is it? SwallowMath-v2 is a large-scale mathematical dataset containing 32 billion tokens, developed as the successor to SwallowMath-v1. Building on the success of v1, this release aims to construct a larger-scale and more permissively licensed corpus to support open and… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math-v2.texttext-generation10M<n<100M35 likes13k downloads11mo agoHugging Face02tokyotech-llm /swallow-code-v2 SwallowCode-v2 Resources 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning. 💻 What is it? SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline. However, it had two significant limitations: (1) it was distributed under the Llama 3.3 Community License, and (2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.tabulartext-generation100M<n<1B48 likes8.3k downloads11mo agoHugging Face03TokyoTechMagicYang /RAM-W600gated Dataset Card for RAM-W600 Benchmark code is available in https://github.com/YSongxiao/RAM-W600. Download Please run the following command to download RAM-W600: git clone https://huggingface.co/datasets/TokyoTechMagicYang/RAM-W600 BoneSegmentation Mask Channel Mapping The BoneSegmentation masks are stored as 14-channel .npy arrays with shape: (14, H, W) Each channel is a binary mask for one anatomical structure. The official channel order is:… See the full description on the dataset page: https://huggingface.co/datasets/TokyoTechMagicYang/RAM-W600.image1K<n<10K3 likes6.2k downloads3mo agoHugging Face04TokyoTechMagicYang /RAM-H1200-v1gated [NeurIPS 2026] RAM-H1200 Dataset Summary RAM-H1200 is a multi-task full-hand radiograph dataset for rheumatoid arthritis (RA) related image analysis. It is designed to support several clinically relevant computer vision tasks, including: hand bone structure segmentation bone erosion related segmentation joint localization for Sharp/van der Heijde (SvdH) scoring joint-level SvdH bone erosion (BE) scoring joint-level SvdH joint space narrowing (JSN) scoring The… See the full description on the dataset page: https://huggingface.co/datasets/TokyoTechMagicYang/RAM-H1200-v1.imageimage-segmentation10K<n<100K10 likes2.7k downloads16h agoHugging Face05Aratako /LiquidAI-Hackathon-Tokyo-CPT-Data LiquidAI-Hackathon-Tokyo-CPT-Data Liquid AI Hackathon Tokyoで作成したモデルのCPTに利用したデータセットです。 automatic-speech-recognition1M<n<10M6 likes1.6k downloads1y agoHugging Face06BangumiBase /tokyomewmewnew Bangumi Image Base of Tokyo Mew Mew New ♡ This is the image base of bangumi Tokyo Mew Mew New ♡, we detected 118 characters, 12865 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/tokyomewmewnew.image10K<n<100K0 likes1.5k downloads2y agoHugging Face07tokyotech-llm /M-IFEval-Jatextn<1K0 likes1k downloads1y agoHugging Face08atad-tokyo /GST_EGOSERVE VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos This is the PyTorch implementation for VideoRAG proposed in this paper: VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context VideosXubin Ren*, Lingrui Xu*, Long Xia, Shuaiqiang Wang, Dawei Yin, Chao Huang† * denotes equal contribution. † denotes corresponding author In this paper, we proposed a retrieval-augmented generation framework specifically designed for processing and… See the full description on the dataset page: https://huggingface.co/datasets/atad-tokyo/GST_EGOSERVE.0 likes1k downloads7mo agoHugging Face09tokyotech-llm /swallow-math SwallowMath October 21, 2025: Newer versions are available: SwallowCode-v2 and SwallowMath-v2 have been released with improved rewriting pipelines. Resources 🐙 GitHub: Explore the project repository, including pipeline code and prompts at rioyokotalab/swallow-code-math. 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowCode, our companion dataset for code generation. What is it?… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math.texttext-generation1M<n<10M49 likes1k downloads7mo agoHugging Face10tokyotech-llm /swallow-code SwallowCode Notice May 21, 2025: We have deleted ablation/exp1-the-stack-v2-train-smol-ids-python because it was flagged as potentially containing unsafe data collected from the Python subset of https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids. However, since this dataset can be reconstructed from the-stack-v2-train-smol-ids, there is no issue in terms of reproducibility. May 21, 2025: ClamAV has flagged “Win.Trojan.MSShellcode-88” in… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code.tabulartext-generation100M<n<1B71 likes894 downloads7mo agoHugging Face11tokyotech-llm /lmsys-chat-1m-synth LMSYS-Chat-1M-Synth: Japanese/English Synthetic Conversation Dataset Derived from LMSYS-Chat-1M This repository contains a series of Japanese and English conversation datasets derived from LMSYS-Chat-1M. Llama-3.1-LMSYS-Chat-1M-Synth Utilized in the post-training of Llama-3.1-Swallow-8B-Instruct-v0.1 and Llama-3.1-Swallow-70B-Instruct-v0.1 Gemma-2-LMSYS-Chat-1M-Synth Utilized in the post-training of Llama-3.1-Swallow-8B-Instruct-v0.3 and Llama-3.1-Swallow-70B-Instruct-v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/lmsys-chat-1m-synth.text-generation100K<n<1M23 likes811 downloads7mo agoHugging Face12tokyotech-llm /Swallow-Nemotron-Post-Training-Dataset-v1 Swallow-Nemotron-Post-Training-Dataset-v1 The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below. Dataset Construction The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528. However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.texttext-generation1M<n<10M6 likes768 downloads7mo agoHugging Face13tokyotech-llm /MMLU-ProX-Japanese0 likes713 downloads1y agoHugging Face14atad-tokyo /GST_EGOSCHEMA Model Card for Model ID https://showlab.github.io/videollm-online/ Model Details LLM: meta-llama/Meta-Llama-3-8B-Instruct Vision Strategy: Frame Encoder: google/siglip-large-patch16-384 Frame Tokens: CLS Token + Avg Pooled 3x3 Tokens Frame FPS: 2 for training, 2~10 for inference Frame Resolution: max resolution 384, with zero-padding to keep aspect ratio Video Length: 10 minutes Training Data: Ego4D Narration Stream 113K + Ego4D GoalStep Stream 21K Model… See the full description on the dataset page: https://huggingface.co/datasets/atad-tokyo/GST_EGOSCHEMA.0 likes700 downloads8mo agoHugging Face15atad-tokyo /cxk0 likes652 downloads8mo agoHugging Face16tokyotech-llm /MMLU-Pro-English0 likes646 downloads1y agoHugging Face17BangumiBase /tokyoghoul Bangumi Image Base of Tokyo Ghoul This is the image base of bangumi Tokyo Ghoul, we detected 74 characters, 3651 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/tokyoghoul.image1K<n<10K1 likes599 downloads3y agoHugging Face18atad-tokyo /GST_Ego4Dvideo0 likes414 downloads8mo agoHugging Face19tokyotech-llm /dclm-baseline0 likes342 downloads1y agoHugging Face20lerobot /tokyo_u_lsmoThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "unknown", "total_episodes": 50, "total_frames": 11925, "total_tasks": 2, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 5, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/tokyo_u_lsmo.tabularrobotics10K<n<100K1 likes338 downloads1y agoHugging Face21finalvent /michiyomi-tokyo-streetscape michiyomi — Tokyo streetscape verbalization open data English Overview michiyomi pairs coordinates with structured Japanese descriptions of physical streetscapes visible in public Mapillary imagery. A vision-language model (VLM) verbalized only what is visible in each image: no map, address, place name, facility name, statistics, or other external knowledge was injected. Release 2026-09-13-r1 contains 1,914,490 scenes covering all of Tokyo: the 23… See the full description on the dataset page: https://huggingface.co/datasets/finalvent/michiyomi-tokyo-streetscape.tabular1M<n<10M2 likes305 downloads13d agoHugging Face22tokyotech-llm /JEMHopQA JEMHopQA このデータセットは SB Intuitions様が公開されている sbintuitions/JEMHopQA を,評価フレームワーク swallow-evaluation-instruct で用いるためにクローンしたものです. 出典 v1, v1.1, v1.2: aiishii/JEMHopQA on GitHub の複製. v1.[1,2]-extended-answers: SB Intuitions 様が同義語や異表記の別解を追加したもの. 具体的には answer: str が answers: List[str] に変更され,オリジナルの正解および別解が answers に格納されている. JEMHopQA JEMHopQA (Japanese Explainable Multi-hop Question Answering) is a Japanese multi-hop QA dataset that can evaluate internal reasoning. It… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/JEMHopQA.textquestion-answering1K<n<10K0 likes297 downloads1y agoHugging Face23tokyotech-llm /MMLU-ProX-English0 likes286 downloads1y agoHugging Face24bot-pi /tokyo_u_lsmoThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "unknown", "total_episodes": 50, "total_frames": 11925, "total_tasks": 2, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 5, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bot-pi/tokyo_u_lsmo.tabularrobotics10K<n<100K0 likes175 downloads11mo agoHugging Face25Aratako /LiquidAI-Hackathon-Tokyo-SFT-Data LiquidAI-Hackathon-Tokyo-SFT-Data Liquid AI Hackathon Tokyoで作成したモデルのSFTに利用したデータセットです。 automatic-speech-recognition1M<n<10M3 likes150 downloads1y agoHugging Face26tokyotech-llm /swallow-magpie-ultra-v0.1 📰 News [07/01/2025] Release of the first version of the dataset containing 42k Japanese pairs and 42k English pairs. Dataset Summary Part of Swallow-Magpie-Ultra-v0.1 is a subset of instruction tuning data for training tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.3, tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.2. The data extracted from magpie-ultra-v0.1 with a quality of average, good, or excellent is… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-magpie-ultra-v0.1.texttext-generation10K<n<100K5 likes149 downloads2y agoHugging Face27open-llm-leaderboard-old /details_tokyotech-llm__Swallow-70b-instruct-hf Dataset Card for Evaluation run of tokyotech-llm/Swallow-70b-instruct-hf Dataset automatically created during the evaluation run of model tokyotech-llm/Swallow-70b-instruct-hf on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tokyotech-llm__Swallow-70b-instruct-hf.0 likes143 downloads3y agoHugging Face28lerobot-raw /tokyo_u_lsmo_raw0 likes119 downloads2y agoHugging Face29amoamo1 /tokyo-koukin-data 東京都 公金支出情報(加工データ) 本データセットは、東京都オープンデータカタログサイトで公開されている「公金支出情報(一般会計・特別会計)」(CC BY 4.0) をもとに作成しています。 出典: 東京都オープンデータカタログサイト ライセンス: Creative Commons Attribution 4.0 International (CC BY 4.0) 加工者: amoamo1 加工内容: 年度ごとのCSV統合およびJSON変換 備考: 本データは元データの形式を変換したものであり、数値・内容の改変は行っていません。 🔍 データ利用方法 ブラウザやWebアプリ(例: Vercel/Next.js)から以下のように取得できます: fetch("https://huggingface.co/datasets/amoamo1/tokyo-koukin-data/raw/main/東京都_公金支出情報_令和5年度.json") .then(res => res.json())… See the full description on the dataset page: https://huggingface.co/datasets/amoamo1/tokyo-koukin-data.textother1M<n<10M0 likes118 downloads11mo agoHugging Face30tokyotech-llm /swallow_japanese_mt_benchtextn<1K0 likes113 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.