datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagenet_hard_review_data_r2VideoGameQA-Bench
VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance
by Mohammad Reza Taesiri, Abhijay Ghildyal, Saman Zadtootaghaj, Nabajeet Barman, Cor-Paul Bezemer
Abstract:
With video games now generating the highest revenues in the entertainment industry, optimizing game development workflows has become essential for the sector's sustained growth. Recent advancements in Vision-Language Models (VLMs) offer considerable potential to automate and… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/VideoGameQA-Bench.arxiv_audioArXivSignals-DeepSummaries
ArXivSignals DeepSummaries — Agent-Built Visual Paper Explainers
A continuously-updated, day-partitioned dataset of deep, visual summaries of
arXiv papers, each built by a coding agent working inside the paper's own
LaTeX source: the agent reads the full text, authors an editorial narrative as
a structured content spec, and the paper's real figures and tables
(extracted and rendered from the LaTeX, web-optimized) ride along as an
embedded, variable-length image array. The… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals-DeepSummaries.imagenet-hard-4K
Dataset Card for "Imagenet-Hard-4K"
Project Page - Paper - Github
ImageNet-Hard-4K is 4K version of the original ImageNet-Hard dataset, which is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet). This dataset poses a significant challenge to state-of-the-art vision models as merely zooming in often fails to improve their… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/imagenet-hard-4K.FragileXArXivSignals
ArXivSignals — Daily arXiv Papers with LLM Signal & Summaries
A continuously-updated, day-partitioned dataset of arXiv papers (AI/ML and
adjacent categories) enriched with LLM-derived signal: a 0–100 importance
score, topical/lab tags, a one-line takeaway, and — for a selected subset —
dense full-page summaries. It powers arxivsignals.io
and is published here as an open research resource.
The dataset has two configs:
papers (default) — one row per paper: bibliography +… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals.Ko-StrategyQA
Ko-StrategyQA
This dataset represents a conversion of the Ko-StrategyQA dataset into the BeIR format, making it compatible for use with mteb.
The original dataset was designed for multi-hop QA, so we processed the data accordingly. First, we grouped the evidence documents tagged by annotators into sets, and excluded unit questions containing 'no_evidence' or 'operation'.
egoxtreme
EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions
📖 Dataset Information
EgoXtreme is a novel large-scale dataset designed for robust egocentric 6D object pose estimation under extreme environmental conditions. The dataset comprises approximately 1.3 million frames with a total duration of 775.5 minutes (~12.9 hours). It was captured at 30 fps using Aria glasses, providing high-resolution 1408 x 1408 raw fisheye RGB… See the full description on the dataset page: https://huggingface.co/datasets/taegyoun88/egoxtreme.arxiv_summaryimagenet_hard_review_dataGameplayCaptions-GPT-4VSteamGlitches-Gemini-LabelsGameplayCaptions-Gemini-pro-visionarxiv_qa
ArXiv QA
(TBD) Automated ArXiv question answering via large language models
Github | Homepage | Simple QA - Hugging Face Space
Automated Question Answering with ArXiv Papers
Latest 25 Papers
LIME: Localized Image Editing via Attention Regularization in Diffusion
Models - [Arxiv] [QA]
Revisiting Depth Completion from a Stereo Matching Perspective for
Cross-domain Generalization - [Arxiv] [QA]
VL-GPT: A Generative Pre-trained Transformer for Vision and… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/arxiv_qa.msr-acc-tae25
Microsoft Research - Accurate Chemistry Collection: Total Atomization Energies
Description
The Microsoft Research Accurate Chemistry Collection (MSR-ACC) provides a collection of accurate coupled cluster labels for training machine learning functionals.
MSR-ACC/TAE25 comprising 73,040 total atomization energies at the CCSD(T)/CBS level obtained with the W1-F12 thermochemical protocol.
The dataset is constructed to exhaustively cover the chemical space of closed-shell… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/msr-acc-tae25.imagenet-hard
Dataset Card for "ImageNet-Hard"
Project Page - ArXiv - Paper - Github - Image Browser
Dataset Summary
ImageNet-Hard is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet). This dataset poses a significant challenge to state-of-the-art vision models as merely zooming in often fails to improve their ability to… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/imagenet-hard.GameplayCaptions-GPT-4V-V2Gameplay-Walkthrough-QAVGQATinyStories-Farsi
Tiny Stories Farsi
The Tiny Stories Farsi project is a continuous effort to translate the Tiny Stories dataset into the Persian (Farsi) language. The primary goal is to produce a high-quality Farsi dataset, maintaining equivalency with the original English version, and subsequently to utilize it for training language models in Farsi. This seeks to affirm that the advancements and trends observed in English language models are replicable and applicable in other languages. Thus far… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/TinyStories-Farsi.VGBDataset-Small-HF[Paper] - [Website]
Ko-Agent-Trajectories-1.0
Ko-Agent-Trajectories-1.0
Dataset card v1.1.1 (2026-09-22). The pipeline code is now released in this repository
under pipeline/, together with the API catalogue, the scenario templates and the complete
prompt set. The card reports the completed human review study and the v1.1 artefacts
(behaviour DPO config, per-item validation scores, manifest, filter asset).
Korean edition: README.ko.md.
TL;DR
A Korean multi-turn agent ↔ tool trajectory corpus synthesized… See the full description on the dataset page: https://huggingface.co/datasets/taejoon89/Ko-Agent-Trajectories-1.0.ko-lima
Dataset Card for KoLIMA
Dataset Description
KoLIMA는 Meta에서 공개한 LIMA: Less Is More for Alignment (Zhou et al., 2023)의 학습 데이터를 한국어로 번역한 데이터셋입니다. 번역에는 DeepL API를 활용하였고, SK(주) Tech Collaborative Lab으로부터 비용을 지원받았습니다. 전체 텍스트 중에서 code block이나 수식을 나타내는 특수문자 사이의 텍스트는 원문을 유지하는 형태로 번역을 진행하였으며, train 데이터셋 1,030건과 test 데이터셋 300건으로 구성된 총 1,330건의 데이터를 활용하실 수 있습니다. 현재 동일한 번역 문장을 plain, vicuna 두 가지 포멧으로 제공합니다.
데이터셋 관련하여 문의가 있으신 경우 메일을 통해 연락주세요! 🥰
This is Korean LIMA dataset, which is… See the full description on the dataset page: https://huggingface.co/datasets/taeshahn/ko-lima.T2I-Prompt-Sample-ImagesVGBDataset-Large-HF[Paper] - [Website]
video-game-question-answeringwav2vec2-ksponspeech-train2wav2vec2-ksponspeech-trainBlindLoop-Evaluation
BlindLoop Evaluation
The frozen evaluation cohorts for the BlindLoop paper. Each eligible generated
task contributes exactly five deterministic, pixel-distinct image instances.
The five rows share a task's selected question/prompt family while varying the
rendered scene and gold answer as determined by the task's pixel oracle.
Config
Tasks
Rows
Documented exclusions
section1_eval5
1,298
6,490
3
section2_eval5
874
4,370
1
combined_eval5
2,172
10,860
4
Each row… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/BlindLoop-Evaluation.
