datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arbiter-mini
Arbiter-mini
A small, purpose-built image dataset of household items captured under controlled Raspberry Pi camera conditions and labeled for binary waste/recycle classification according to San Diego, CA municipal recycling rules. Built as deployment-condition training data for the Arbiter sorting system, intended to be used alongside TrashNet to close the domain gap between studio imagery and real Pi-camera inference.
Motivation
Models trained purely on TrashNet… See the full description on the dataset page: https://huggingface.co/datasets/aaryavlal/arbiter-mini.ARB
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
Sara Ghaboura *
Ketan More *
Wafa Alghallabi
Omkar Thawakar
Jorma Laaksonen
Hisham Cholakkal
Salman Khan
Rao M. Anwer
*Equal Contribution
🪔✨ ARB Scope and Diversity
ARB is the first benchmark focused on step-by-step reasoning in Arabic cross both textual and visual modalities, covering 11 diverse domains spanning science, culture, OCR, and… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ARB.arbigraph
ArbiGraph
ArbiGraph is a benchmark generator for evaluating context management in language
models and agents. It automatically builds verifiable directed task graphs whose
nodes are math, Python tracing, or GSM-style tasks, and whose edges pass one
task's output into another task's input.
The datasets uploaded here are example benchmark datasets generated with
ArbiGraph. They are meant both for direct evaluation and as concrete examples of
what the generator can produce. The… See the full description on the dataset page: https://huggingface.co/datasets/PavelGolikov/arbigraph.CalliFontXL
Dataset Card for "CalliFontXL"
More Information needed
FontsLarge
Dataset Card for "FontsLarge"
More Information needed
modality-conflict-arbitration-v2
Modality-Conflict Arbitration Benchmark (v2)
A controlled benchmark for studying how a vision-language model arbitrates between
its two input channels when they disagree — and whether that choice tracks the
reliability of each channel.
Each row is a single conflict trial: an image of one math problem paired with the
text of a different problem. Because the two ground-truth answers are carried side by
side, the model's output alone tells you which modality it followed — no… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/modality-conflict-arbitration-v2.ArSL21L
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/arbml/ArSL21L.FontsSmall
Dataset Card for "FontsSmall"
More Information needed
dengue-arboviral-infections
Dengue & Arboviral Infections (Dengue, Chikungunya, Zika, RVF, YF) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/dengue-arboviral-infections.AutoGenArabicDataset
Dataset Card for "AutoGenArabicDataset"
More Information needed
TinyCalliFont
Dataset Card for "TinyCalliFont"
More Information needed
finepdfs_arb_Arabgravitational_lensingCalliar
Dataset Card for "Calliar"
More Information needed
Hijja2
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/arbml/Hijja2.FontsLargeSpaced
Dataset Card for "FontsLargeSpaced"
More Information needed
Bonsai-imagestwitter-YYtu_love6-2026.03.07-2030115034014306524-ArBi4ue_vyKPs5k0-part1
