datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trex-visualizer
T-Rex Dataset Visualizer
A browseable subset of the T-Rex dataset — Tactile-Rich Bimanual Dexterous
Manipulation — collected on a bimanual Dexmate Vega-1 robot equipped with
two Sharpa Wave dexterous hands.
This visualizer subset contains 3,838 short trajectory clips drawn from the
full 100-hour T-Rex collection, organized by (verb, object, hand) so you can
quickly inspect coverage across motion primitives and object categories.
For the full dataset (multi-view RGB, robot… See the full description on the dataset page: https://huggingface.co/datasets/Beakerman0101/trex-visualizer.visualears-bench-results
🗂️ visualears-bench-results
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Benchmark result dataset and leaderboard artifacts.
دادهها و مصنوعات نتایج بنچمارک برای نگهداری خروجی مدلها، امتیازها و منشأ اجرای ارزیابی.
🧩 Role
evaluation and benchmarking asset
مصنوع ارزیابی و بنچمارک
📦 Snapshot
4 files; approximately 51.61 KB
4 فایل؛ حدود 51.61 KB
🧱 Packaging
0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-bench-results.VisualNovel_Dataset_Metadata
VisualNovel Dataset Metadata
This dataset repository contains metadata for OOPPEENN/56697375616C4E6F76656C5F44617461736574.
maker_game_vnid.tsv: Contains the mapping of maker names, game names, and their corresponding Visual Novel IDs in vndb. Here the column maker and game correspond to {maker}_{game} in the file names of 7z (or directory names of extracted files) in OOPPEENN/VisualNovel_Dataset.
voice_actor_info.tsvmaps maker game character to vndb char id, vndb voice actor id… See the full description on the dataset page: https://huggingface.co/datasets/litagin/VisualNovel_Dataset_Metadata.VisualPuzzles-tsvVisualProbe_Easy
VLMsAreBlind Benchmark
refactor VisualProbe Benchmark to add support for VLMEvalKit.
Benchmark Information
Number of questions: 141
Question type: Free-form
Question format: image + text
Answer type: Free-form
Reference
VLMEvalKit
VisualProbe Benchmark
VisualTraceBench
VisualTraceBench v1.0 — Dataset Card
Benchmark for Visual Execution-Trace Differencing in Software Performance Regression Diagnosis
Following the Datasheet for Datasets framework (Gebru et al., 2021).
Motivation
Purpose: VisualTraceBench is a curated dataset of 192 real-world software performance regression bug reports from a large enterprise database system (2016–2025), annotated with trace artifact type (visual vs. text-only) and resolution timing. It enables… See the full description on the dataset page: https://huggingface.co/datasets/viperoptic/VisualTraceBench.VisualProbe_Hard
VLMsAreBlind Benchmark
refactor VisualProbe Benchmark to add support for VLMEvalKit.
Benchmark Information
Number of questions: 106
Question type: Free-form
Question format: image + text
Answer type: Free-form
Reference
VLMEvalKit
VisualProbe Benchmark
china-historical-visual-resources-index
China Historical Visual Resources Index / 中国历史视觉资料索引
(中文说明请向下滚动 / Please scroll down for the Chinese version)
This is a curated metadata index of Chinese historical visual resources scattered across global collections, universities, libraries, and museums.
IMPORTANT NOTE: This is a resource index / metadata dataset, NOT a direct image dataset. The dataset provides structured URLs, descriptions, and categorized entry points to historical photographs, maps, pictorials, and visual… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/china-historical-visual-resources-index.ecommerce-visual-matching-dataset
E-commerce Visual Matching Dataset
Candidate match workflow dataset for product identity resolution, visual similarity, and review decision fields.
A public-safe workflow preview of how Octoparse structures AI-assisted product matching pipelines for e-commerce and pricing teams. Every row represents a candidate pair evaluation — the same structure delivered to production clients.
Built by Octoparse Managed Data Service — managed web data pipelines for pricing intelligence and… See the full description on the dataset page: https://huggingface.co/datasets/Octoparse/ecommerce-visual-matching-dataset.VisualProbe_Medium
VLMsAreBlind Benchmark
refactor VisualProbe Benchmark to add support for VLMEvalKit.
Benchmark Information
Number of questions: 268
Question type: Free-form
Question format: image + text
Answer type: Free-form
Reference
VLMEvalKit
VisualProbe Benchmark
Visual_information_retrieval
GDZ Scientific Document Retrieval Benchmark
A needle‑in‑a‑haystack benchmark for scientific document retrieval, built from historical volumes of the Göttinger Digitalisierungszentrum (GDZ). This dataset explicitly adapts the IRPAPERS methodology onto a real‑world, multilingual corpus to evaluate both text-based and visual document retrieval models.
Dataset Structure
The dataset is divided into two operational configurations:
1. queries
Contains the… See the full description on the dataset page: https://huggingface.co/datasets/Trungdaik/Visual_information_retrieval.bengali-visual-genome-instruction-sethindi_visual_genomeVisualGenomeTrainingSetmalayalam-visual-genome-instruction-setodia-visual-genome-instruction-sethuman-agent-conversationsThe data consists of transcriptions of audio tracks from YouTube videos.
The addresses of the videos were collected using the youtube-data-api-v3. Similarly, transcriptions of the audio tracks of the videos were collected.
Each video was divided into chunks of 250 words, which on average corresponds to 1.5 minutes of dialogue time.
Each chunk was marked by the LLaMA 3 70B Instruct model upon request
2. Extract another speaker's speech and enclose it in [speaker_2_speech] [/speaker_2_speech]… See the full description on the dataset page: https://huggingface.co/datasets/visualcomments/human-agent-conversations.visualsVisual_retrieval
Dataset Card for IRPAPERS
ArXiv Link: https://arxiv.org/pdf/2602.17687
Dataset Description
IRPAPERS is a collection of 166 Information Retrieval papers spanning 3,230 pages. Each page in the dataset is jointly represented as a base64 encoded string of the page image as well as an OCR-derived text transcription. IRPAPERS also contains 180 needle-in-the-haystack queries.
Retrieval Leaderboard 🔎
Rank
Retriever
Type
Recall@1
Recall@5
Recall@20… See the full description on the dataset page: https://huggingface.co/datasets/Trungdaik/Visual_retrieval.visual-leanerhuman-human-conversationshindi-visual-genome-instruction-setSynteticPersonaChatOSINT-Visuals
YemenJPT-Visuals-v1
الصور الصحفية - مجموعة صور للتحقيقات الصحفية
الوصف
الصور الصحفية - مجموعة صور للتحقيقات الصحفية. هذا النموذج/القاعدة جزء من مجموعة YemenJPT المخصصة لدعم الصحافة الاستقصائية اليمنية.
الروابط
المستودع على HuggingFace
منظومة YemenJPT
الموقع الرسمي
Ollama
RaidanPro
بيت الصحافة
التحميل
# عبر HuggingFace Hub
git lfs clone https://huggingface.co/Yemen-JPT/OSINT-Visuals
# أو عبر pip (للنماذج)
pip install huggingface-hub… See the full description on the dataset page: https://huggingface.co/datasets/Yemen-JPT/OSINT-Visuals.VISUAL_TEXTvisual_text_datasetvisual_acoustic_text_datasetvisual_accent_dialect_archiveSource: https://www.youtube.com/@visualaccent/videos
All rights belong to the original dataset creator.
VADA-AVSR: an audio-visual dataset of non-native English ("accents") and English varieties ("dialects")
We preprocessed the Visual Accent and Dialect Archive (https://archive.mith.umd.edu/mith-2020/vada/index.html) for audio-visual speech recognition (AVSR), speech recognition (ASR), and visual speech recognition/lip-reading (VSR).
This version currently only contains read speech… See the full description on the dataset page: https://huggingface.co/datasets/Berkeley-NLP/visual_accent_dialect_archive.visual_modifiedacoustic_datasetOperational_AI_Data_Visualization
