datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ogiri-bokete
読み込み方
from datasets import load_dataset
dataset = load_dataset("YANS-official/ogiri-bokete", split="train")
概要
大喜利投稿サイトBoketeのクロールデータです。元データは CLoT-Oogiri-Go [Zhang+ CVPR2024]というデータの一部です。
詳細はCVPRのプロジェクトページをご確認ください。
このデータは以下の3タスクが含まれます。
text_to_text: テキストでお題が渡され、それに対する回答を返します。
image_to_text: いわゆる「画像で一言」です。画像のみが渡されて、テキストによる回答を返します。
text_image_to_text: 画像中にテキストが書かれています。テキストの一部が空欄になっているので、そこに穴埋めする形で回答を返します。
それぞれの量は以下の通りです。(8/30現在。ハッカソン当日までに増やす可能性があります。)
タスク… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/ogiri-bokete.refute
Can AI read new science honestly?
Models can sound convincing while misreading a result or expressing more confidence than the evidence deserves. That matters when people use them to summarize papers, compare studies, or decide what to investigate next.
REFUTE tests whether a model knows the finding, spots quiet flaws, names what would overturn a claim, and matches its confidence to the evidence.
Truth Score is the main result. It combines factual accuracy, flaw… See the full description on the dataset page: https://huggingface.co/datasets/BGPT-OFFICIAL/refute.IDDAW_OFFICIAL
IDD-AW: India Driving Dataset – Adverse Weather
Semantic segmentation benchmark for autonomous driving in rain, fog, low-light,
and snow, with paired RGB + near-infrared (NIR) frames and dense Level-3
semantic labels (26 classes).
TODO before publishing: confirm and set the correct license / citation for
the original IDD-AW release (see iddaw.github.io and
the WACV 2024 paper "IDD-AW: A Benchmark for Safe Semantic Segmentation in
Adverse Weather"). This card currently marks the… See the full description on the dataset page: https://huggingface.co/datasets/Furqan7007/IDDAW_OFFICIAL.RoboTwin-adjust_bottle-official-demo_clean50-Pi0_processed-databokeh-eval-official
bokeh-eval-official
Bokeh synthesis artifacts: all-in-focus input, ground-truth bokeh, the
best-K render and the full K sweep, one split per benchmark.
Generated by inference/bokeh_net.py.
OlympiadBench-official
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
📖 arXiv | GitHub
Dataset Description
OlympiadBench is an Olympiad-level bilingual multimodal scientific benchmark, featuring 8,476 problems from Olympiad-level mathematics and physics competitions, including the Chinese college entrance exam. Each problem is detailed with expert-level annotations for step-by-step reasoning. Notably, the best-performing… See the full description on the dataset page: https://huggingface.co/datasets/lscpku/OlympiadBench-official.Selfie_and_Official_ID_Photo_Dataset12,000+ people, 150,000+ images. Selfie with ID dataset for KYC verification, face identification and biometric training. Selfies paired with 2 official ID photos (passport, ID card, driver's license, residence permit). 10-15 photos per person with balanced demographics across ethnicity (Caucasian, Black, Asian, Latin American), gender and age (18-65).
Contact us and share your feedback - recieve additional samples for free! 😊
Key Highlights:
12,000+ real individuals… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/Selfie_and_Official_ID_Photo_Dataset.IOAI2025
International Olympiad in Artificial Intelligence (IOAI 2025, Beijing, China)
About IOAI 2025
The 2nd International Olympiad in Artificial Intelligence (IOAI 2025) took place in Beijing, China, from August 2 to 9, 2025, hosted by Beijing National Day School (BNDS) under the patronage of UNESCO.
Contest Rules: Full rules encompassing the Individual, Team, and GAITE contests are available here.
Syllabus: The official syllabus outlining the AI topics contestants should… See the full description on the dataset page: https://huggingface.co/datasets/IOAI-official/IOAI2025.ogiri-test
読み込み方
from datasets import load_dataset
dataset = load_dataset("YANS-official/ogiri-test", split="test")
概要
大喜利投稿サイトBoketeのクロールデータです。元データは CLoT-Oogiri-Go [Zhang+ CVPR2024]というデータの一部です。
詳細はCVPRのプロジェクトページをご確認ください。
このデータは以下の3タスクが含まれます。
text_to_text: テキストでお題が渡され、それに対する回答を返します。
image_to_text: いわゆる「画像で一言」です。画像のみが渡されて、テキストによる回答を返します。
text_image_to_text: 画像中にテキストが書かれています。テキストの一部が空欄になっているので、そこに穴埋めする形で回答を返します。
それぞれの量は以下の通りです。
タスク
お題数(画像枚数)
image_to_text
56… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/ogiri-test.guided_genshin_impact_official_server_recordings_01
原神 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game: 原神 (Genshin Impact)
Collection: guided (精数据)
Subset: official_server (官服(非私服))
Recordings: 566
Planned bytes: 4440851482194
Layout: recordings//
Parquet files are intentionally excluded.
magical-girl-lyrical-nanoha-official-art-ver2026_USAAIO_Round2ogiri-test-with-references
読み込み方
from datasets import load_dataset
dataset = load_dataset("YANS-official/bokete-ogiri-test", split="test")
概要
大喜利投稿サイトBoketeのクロールデータです。元データは CLoT-Oogiri-Go [Zhang+ CVPR2024]というデータの一部です。
詳細はCVPRのプロジェクトページをご確認ください。
このデータは以下の3タスクが含まれます。
text_to_text: テキストでお題が渡され、それに対する回答を返します。
image_to_text: いわゆる「画像で一言」です。画像のみが渡されて、テキストによる回答を返します。
text_image_to_text: 画像中にテキストが書かれています。テキストの一部が空欄になっているので、そこに穴埋めする形で回答を返します。
それぞれの量は以下の通りです。
タスク
お題数(画像枚数)
回答数
うち委員が用意したお題… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/ogiri-test-with-references.vlm-project-with-images-with-bbox-images-official-q3-updateindoor-semantic-sample
Fengmap Indoor Semantic Map Sample Dataset
Dataset version: v1.0Semantic format specification version: v0.2Release date: August 14, 2026Dataset size: 5 indoor maps across 37 floorsPermitted use: Non-commercial learning, research, education, and technical validation only
Dataset Overview
The Fengmap Indoor Semantic Map Sample Dataset is a public test dataset designed for indoor spatial understanding, spatial relationship analysis, map SDK integration, and… See the full description on the dataset page: https://huggingface.co/datasets/fengmap-official/indoor-semantic-sample.IOAI-2025-Pixel-trainioai2025-onsite-concepts-hint-descriptionsClaire_RE2_Official_jacket_clothessenryu-test-with-references
読み込み方
from datasets import load_dataset
dataset = load_dataset("YANS-official/senryu-test", split="test")
概要
川柳投稿サイトの『写真川柳』と『川柳投稿まるせん』のクロールデータです。
以下のページからクロールし、原本のHTMLファイルと構造化処理を行った結果を格納しました。
https://www.homemate-research.com/senryu/photo/
https://marusenryu.com/
このデータは以下の2タスクが含まれます。
image_to_text: 画像でお題が渡され、それに対する回答を返します。
text_to_text: テキストでお題が渡され、それに対する回答を返します。
それぞれの量は以下の通りです。
タスク
お題数(画像枚数)
回答数
うち委員が用意したお題
image_to_text
70
140
7
text_to_text
30
60… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/senryu-test-with-references.IOAI-2025-Pixel-refBEP_OfficialThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 10,
"total_frames": 13774,
"total_tasks": 1,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/BobBobbson/BEP_Official.ogiri-debug
読み込み方
from datasets import load_dataset
dataset = load_dataset("YANS-official/ogiri-debug", split="test")
概要
大喜利生成の動作確認用データセットです。以下の3タスクが含まれます。
text_to_text: テキストでお題が渡され、それに対する回答を返します。
image_to_text: いわゆる「画像で一言」です。画像のみが渡されて、テキストによる回答を返します。
text_image_to_text: 画像中にテキストが書かれています。テキストの一部が空欄になっているので、そこに穴埋めする形で回答を返します。
データセットの各カラム説明
カラム名
型
例
概要
odai_id
str
"origi-dummy-1"
お題のID
file_path
str
"dummy.png"
画像のファイル名
type
str
"text_to_text"… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/ogiri-debug.IOAI-2025-Pixel-testpokemon-blip-captions
Dataset Card for Pokémon BLIP captions
Dataset used to train Pokémon text to image model
BLIP generated captions for Pokémon images from Few Shot Pokémon dataset introduced by Towards Faster and Stabilized GAN Training for High-fidelity Few-shot Image Synthesis (FastGAN). Original images were obtained from FastGAN-pytorch and captioned with the pre-trained BLIP model.
For each row the dataset contains image and text keys. image is a varying size PIL jpeg, and text is the… See the full description on the dataset page: https://huggingface.co/datasets/AxionLab-official/pokemon-blip-captions.vlm-project-with-images-with-bbox-images-officialcofi-compdiffuser-official-artifacts2026_USAAIO_Round3_wildlifevlm-project-with-images-distribution-q2-translation-all-language-officialuncertainty-vlm-qwen3-officialpokemon-official-art
Pokémon Official Artwork
Dataset Description
This dataset contains official Pokémon artwork collected from publicly available sources such as Bulbapedia and PokémonDB.
The goal of this dataset is to provide a standardized collection of official Pokémon illustrations that can be used for:
Image classification
Computer vision research
Character recognition
Dataset generation
AI and machine learning experiments
Reference images for Pokémon-related research
Each… See the full description on the dataset page: https://huggingface.co/datasets/CitronLegacy/pokemon-official-art.
