datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aegis-gsm8k-bench
🧮 AEGIS GSM8K Adversarial Robustness Benchmark
Preprocessed GSM8K benchmark splits formatted in native Apache Parquet, configured for multi-turn adversarial stress testing and counterfactual invariance evaluation in LLM-as-a-Judge research.
📊 Dataset Structure & Splits
Train Split: 7,473 mathematical reasoning problems with explicit reference derivations.
Test Split: 1,319 problems used for out-of-distribution adversarial debate evaluation.
Format:… See the full description on the dataset page: https://huggingface.co/datasets/Hooshaai/aegis-gsm8k-bench.aegis-bilingual-industrial-ai-dataset
AEGIS AI Bilingual Industrial Operations Dataset
AEGIS AI Bilingual Industrial Operations Dataset is a synthetic English–Arabic dataset designed for experimentation with multilingual enterprise AI systems, Retrieval-Augmented Generation (RAG), industrial question answering, document intelligence, semantic search, and AI workflow automation.
The dataset extends the original AEGIS AI industrial dataset with structured Arabic and English representations while preserving industrial… See the full description on the dataset page: https://huggingface.co/datasets/syed7741/aegis-bilingual-industrial-ai-dataset.neovim-helpAEGIS
AEGIS: Exploring the Limit of World Knowledge Capabilities for Unified Multimodal Models
[📂 GitHub] [🆕 Blog] [📜 Paper]
Summary
The capability of Unified Multimodal Models (UMMs) to apply world knowledge across diverse tasks remains a critical, unresolved challenge. Existing benchmarks fall short, offering only siloed, single-task evaluations with limited diagnostic power. To bridge this gap, we propose AEGIS (i.e., Assessing Editing, Generation… See the full description on the dataset page: https://huggingface.co/datasets/DongSky/AEGIS.
