datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kiiteitte
Kiiteitte history
Kiiteitte が収集した、今までの選曲履歴。
1時間おきに更新されます。
型
{
// 動画ID
"video_id": "sm44670499",
// タイトル
"title": "library->w4nderers / 足立レイ、つくよみちゃん",
// 投稿者
"author": "名無し。",
// サムネイルのURL
"thumbnail": "https://nicovideo.cdn.nimg.jp/thumbnails/44670499/44670499.91820835",
// 選曲日時
"date": "2025-02-22 12:51:51",
// 新しく増えたお気に入り数。不明の場合は null
"new_faves": 5,
// 回ったユーザーの数。不明の場合は null
"spins": 13,
// イチ押しリストのユーザーのURL。イチ押しリスト以外から選曲された場合は null… See the full description on the dataset page: https://huggingface.co/datasets/sevenc-nanashi/kiiteitte.TestingDataset
SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction
This dataset contains multimodal question-answering examples grounded in
textbook figures. Records in the figure-grounded configurations are filtered to
include only examples whose referenced image files are present in this release.
Configurations
visual: 13791 figure-grounded visual questions with resolved images.
knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/Naga1289/TestingDataset.ui-navigation-corpus
User Interface (Navigation) Corpus
Overview
This dataset serves as a collection of various images of, videos and metadata of mobile (both iOS and Android) and web user interfaces as well as tags and text extractions associated to them.
Dataset also includes user interface navigation annotations and videos related to them. One of the possible use cases of this dataset is training a UI navigation agent.
Dataset Structure
The resources of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/ijlewis/ui-navigation-corpus.ui-navigation-corpus
User Interface (Navigation) Corpus
Overview
This dataset serves as a collection of various images of, videos and metadata of mobile (both iOS and Android) and web user interfaces as well as tags and text extractions associated to them.
Dataset also includes user interface navigation annotations and videos related to them. One of the possible use cases of this dataset is training a UI navigation agent.
Dataset Structure
The resources of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/teleren/ui-navigation-corpus.motif-qa
MotifQA
Dataset Summary
MotifQA is a synthetic graph question-answering benchmark focused on detecting graph motifs inside small random graphs.
Each example pairs a textual prompt with an answer sentence, a list of nodes highlighted as the motif (when present), and an explicit graph description(nodes and edges).
In this QA dataset, all graphs are homogenous and undirected.
Subsets cover both yes/no motif detection, motif-type classification (house vs 5-cycle), and… See the full description on the dataset page: https://huggingface.co/datasets/naos-ku/motif-qa.Pantheon-Agent-Trajectory
🏛️ Pantheon Agent Trajectory Gallery
Curated end-to-end agent runs from PantheonOS — an open multi-agent framework for scientific computing.
Each "trajectory" captures a complete chat session: the user prompt, every reasoning/tool step the agent(s) took, the code that was run, the figures that were produced, and the final report. Trajectories are fully inspectable and reproducible, designed for transparency, teaching, and benchmarking.
🔗 Browse the gallery (live):… See the full description on the dataset page: https://huggingface.co/datasets/NaNg/Pantheon-Agent-Trajectory.newspaper_navigatortarget-heatmap-datasetNatureBench
Dataset Card for NatureBench
NatureBench is a cross-discipline benchmark of 27 tasks distilled from peer-reviewed Nature-family publications, spanning 6 scientific domains. It is designed to evaluate whether AI coding agents can move beyond reproduction toward discovery: each task asks an agent to solve a real scientific machine-learning problem and is scored against the source paper's reported state of the art.
📄 arXiv paper: https://arxiv.org/abs/2606.24530
💻 GitHub code… See the full description on the dataset page: https://huggingface.co/datasets/JoeLiu996/NatureBench.Dataset1_LupaNamanyaApa
Safety Helmet and Reflective Jacket
Safety Helmet and Reflective Jacket is a dataset for object detection task.
chunithm-charts-db
chunithm-charts-db
ChuniSupportの譜面データを平坦なjsonlに変換したやつ。
フィールド
1行が1譜面を表すJSONLです。APIの charts 内の項目を行の直下に展開しています。
以下の型はREADME先頭のHugging Faceスキーマに対応します。null 可の項目は、元データに値がない場合に null になります。
共通項目
ChuniSupport API仕様に基づく項目です。jacket は変換処理で画像URLにしています。
フィールド
型
null可
説明
id
string
—
楽曲ID。通常曲では同じ楽曲の各難易度で共通。
title
string
—
曲名。
reading
string
✓
曲名の読み。
artist
string
—
アーティスト名。
genre
string
✓
ジャンル。
bpm
int32
✓
楽曲のBPM。
release
date32
✓
配信日。JSONLでは… See the full description on the dataset page: https://huggingface.co/datasets/sevenc-nanashi/chunithm-charts-db.ExpArt
Dataset Card for Explain Artworks: ExpArt
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Card for "Wiki-ImageReview1.0"
Dataset Summary
Explain Artworks: ExpArt is designed to enhance the capabilities of large-scale vision-language models (LVLMs) in analyzing and describing artworks.
Drawing from a comprehensive array of English Wikipedia art articles, the dataset encourages LVLMs to… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/ExpArt.Phone_SpecsDataset_25KUpdate README.md
PhoneDB Device Specifications Dataset
Overview
This dataset contains structured information about smartphones and mobile devices, parsed from PhoneDB.net
Each record provides a detailed set of specifications for a single device, including hardware, software, dimensions, cameras, sensors, connectivity, and more.
The dataset is provided in JSON format for easy use in data analysis, machine learning, and information retrieval tasks.
Each entry contains:
title → Full device name… See the full description on the dataset page: https://huggingface.co/datasets/Nadirova/Phone_SpecsDataset_25K.exercise-api
Exercise API — Dataset
Dataset de 104 ejercicios de gimnasio (bilingüe ES/EN) derivado de la
Exercise API. Cada ejercicio incluye grupo muscular,
equipamiento, músculos principal/secundario, instrucciones paso a paso e ilustración
masculina y femenina (208 imágenes en total).
Configuraciones
images — 1 fila por imagen (208). Etiquetas (grupo, equipamiento,
músculos, género) + caption_es/caption_en. Para clasificación de imagen y multimodal
(image-to-text / VQA).… See the full description on the dataset page: https://huggingface.co/datasets/natzx94/exercise-api.laundry-spots-dataset
Laundry Spots Dataset
Generated from naavox/merged-5.
napeval
NAPEval Spreadsheet Trajectories
Benchmark trajectories for predictive auto-completion in spreadsheets. Each
record is one spreadsheet-building session represented as an ordered sequence
of symbolic cell operations. Given a prefix of a trajectory, the task is to
predict the next operation(s) the user will perform.
This is the evaluation data for the
next_action_pred_eval
framework. See the project page and
paper.
Dataset summary
52 trajectories (single test split… See the full description on the dataset page: https://huggingface.co/datasets/Tej-a55/napeval.eggs-heatmapangling-datasetamex-augmented-sftaquariumsquare-centering-datasetbig-naturalsgripper-spots-dataset
Laundry Spots Dataset
Generated from naavox/merged-5.
21-naturalsmoe1-naive-k-ablation-summariesmultiview-dataset
