datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
visualears-bench-results
🗂️ visualears-bench-results
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Benchmark result dataset and leaderboard artifacts.
دادهها و مصنوعات نتایج بنچمارک برای نگهداری خروجی مدلها، امتیازها و منشأ اجرای ارزیابی.
🧩 Role
evaluation and benchmarking asset
مصنوع ارزیابی و بنچمارک
📦 Snapshot
4 files; approximately 51.61 KB
4 فایل؛ حدود 51.61 KB
🧱 Packaging
0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-bench-results.VisualTraceBench
VisualTraceBench v1.0 — Dataset Card
Benchmark for Visual Execution-Trace Differencing in Software Performance Regression Diagnosis
Following the Datasheet for Datasets framework (Gebru et al., 2021).
Motivation
Purpose: VisualTraceBench is a curated dataset of 192 real-world software performance regression bug reports from a large enterprise database system (2016–2025), annotated with trace artifact type (visual vs. text-only) and resolution timing. It enables… See the full description on the dataset page: https://huggingface.co/datasets/viperoptic/VisualTraceBench.ecommerce-visual-matching-dataset
E-commerce Visual Matching Dataset
Candidate match workflow dataset for product identity resolution, visual similarity, and review decision fields.
A public-safe workflow preview of how Octoparse structures AI-assisted product matching pipelines for e-commerce and pricing teams. Every row represents a candidate pair evaluation — the same structure delivered to production clients.
Built by Octoparse Managed Data Service — managed web data pipelines for pricing intelligence and… See the full description on the dataset page: https://huggingface.co/datasets/Octoparse/ecommerce-visual-matching-dataset.Visual_information_retrieval
GDZ Scientific Document Retrieval Benchmark
A needle‑in‑a‑haystack benchmark for scientific document retrieval, built from historical volumes of the Göttinger Digitalisierungszentrum (GDZ). This dataset explicitly adapts the IRPAPERS methodology onto a real‑world, multilingual corpus to evaluate both text-based and visual document retrieval models.
Dataset Structure
The dataset is divided into two operational configurations:
1. queries
Contains the… See the full description on the dataset page: https://huggingface.co/datasets/Trungdaik/Visual_information_retrieval.hindi_visual_genomeVisual_retrieval
Dataset Card for IRPAPERS
ArXiv Link: https://arxiv.org/pdf/2602.17687
Dataset Description
IRPAPERS is a collection of 166 Information Retrieval papers spanning 3,230 pages. Each page in the dataset is jointly represented as a base64 encoded string of the page image as well as an OCR-derived text transcription. IRPAPERS also contains 180 needle-in-the-haystack queries.
Retrieval Leaderboard 🔎
Rank
Retriever
Type
Recall@1
Recall@5
Recall@20… See the full description on the dataset page: https://huggingface.co/datasets/Trungdaik/Visual_retrieval.human-agent-conversationsThe data consists of transcriptions of audio tracks from YouTube videos.
The addresses of the videos were collected using the youtube-data-api-v3. Similarly, transcriptions of the audio tracks of the videos were collected.
Each video was divided into chunks of 250 words, which on average corresponds to 1.5 minutes of dialogue time.
Each chunk was marked by the LLaMA 3 70B Instruct model upon request
2. Extract another speaker's speech and enclose it in [speaker_2_speech] [/speaker_2_speech]… See the full description on the dataset page: https://huggingface.co/datasets/visualcomments/human-agent-conversations.human-human-conversationsvisual_accent_dialect_archiveSource: https://www.youtube.com/@visualaccent/videos
All rights belong to the original dataset creator.
VADA-AVSR: an audio-visual dataset of non-native English ("accents") and English varieties ("dialects")
We preprocessed the Visual Accent and Dialect Archive (https://archive.mith.umd.edu/mith-2020/vada/index.html) for audio-visual speech recognition (AVSR), speech recognition (ASR), and visual speech recognition/lip-reading (VSR).
This version currently only contains read speech… See the full description on the dataset page: https://huggingface.co/datasets/Berkeley-NLP/visual_accent_dialect_archive.
