datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vlms-are-biased
Vision Language Models are Biased
by
An Vo1*,
Khai-Nguyen Nguyen2*,
Mohammad Reza Taesiri3,
Vy Tuong Dang1,
Anh Totti Nguyen4†,
Daeyoung Kim1†
*Equal contribution †Equal advising
1KAIST, 2College of William and Mary, 3University of Alberta, 4Auburn University
TLDR: State-of-the-art Vision Language Models (VLMs) perform perfectly on counting tasks with original images but fail catastrophically (e.g., 100% → 17.05%… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/vlms-are-biased.za-african-next-voices
Swivuriso: ZA-African Next Voices
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes.
Dataset Paper: ArXiv - Work in Progress
Language Coverage… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices.Dholuokikuyuanv_data_ke_kikuyu_mergedGarageDatasetmultilingual-nchlt-dataset
NCHLT Auxiliary Speech Corpus - Combined Multilingual Dataset
Dataset Description
This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data Science for Social Impact (DSFSI) research group at the University of Pretoria to facilitate easier benchmarking and multi-language speech recognition research.
The original auxiliary data was collected during the National Centre for Human Language Technology (NCHLT) project for the 11 official… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/multilingual-nchlt-dataset.anv_data_ke_kikuyu_scriptedanv-data-ke-somali-fullanv-data-ke-somali-fullanv_data_kelanguage:
ki
so
kln
luo
mas
pretty_name: anv_ke
⚠️ IMPORTANT: Work in ProgressThis dataset is not final. Updates will continue through September 2025.Please use the latest version for attribution, benchmarking and publications.
Overview
African Next Voices: Pilot Data Collection in Kenya is part of a larger initiative to support African language speech technology. This project, funded by the Gates Foundation, is led by the KenCorpus Consortium, a coalition of Kenyan… See the full description on the dataset page: https://huggingface.co/datasets/MCAA1-MSU/anv_data_ke.anv-kikuyu-banking-subset-v2-part10KalenjinMaasaiSomaliConstrained_Indic_Codemixingvmmu
ViExam: Are Vision Language Models Better than Humans on Vietnamese Multimodal Exam Questions?
by
Vy Tuong Dang*,
An Vo*,
Quang Tau,
Duc Dm,
Daeyoung Kim,
*Equal contribution
KAIST
TLDR: State-of-the-art Vision Language Models (VLMs) demonstrate remarkable capabilities on English multimodal tasks but significantly underperform on Vietnamese educational assessments. ViExam reveals that SOTA VLMs achieve only 57.74% accuracy… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/vmmu.Radiology_Project_Annotateddeep_bias_resultsAnveshAI-Coder-150K
Coding AI Dataset - 150K Problem/Thinking/Solution Entries
Overview
A comprehensive dataset of 150,000 coding problems designed for training code generation AI models.
Each entry contains a problem statement, detailed thinking process, and complete solution in the target language.
Format
JSONL (JSON Lines) - each line is a valid JSON object.
Schema
{
"id": "unique identifier string",
"domain": "algorithms | data_structures |… See the full description on the dataset page: https://huggingface.co/datasets/AnveshAI/AnveshAI-Coder-150K.anv-kikuyu-banking-subset
Anv-Kikuyu Banking Subset
A domain-filtered subset of Kikuyu (Gĩkũyũ) speech data focused on banking and financial-transaction content, combined into a single repository with train, test, and validation splits.
Source
This dataset is a filtered subset of Anv-ke/kikuyu, part of the African Next Voices (ANV) collection. All audio, transcriptions, and underlying speaker data originate from that source dataset. Full credit for data collection belongs to the… See the full description on the dataset page: https://huggingface.co/datasets/NjeriKahoro/anv-kikuyu-banking-subset.anv-data-ke-somaliza-african-next-voices-compressedNote: This dataset is a compressed version of za-african-next-voices. It was compressed to .opus format using a 32k bitrate.
Swivuriso: ZA-African Next Voices-Compressed
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices-compressed.perfume-rec-assetspendulum-conflict
Pendulum Conflict Dataset
This dataset is a manually curated subset of 100 isolated multimodal conflict examples derived from the akomand/counterfactual_pendulum dataset.
It is designed for evaluating Multimodal Large Language Models (MLLMs) under controlled visual-textual conflicts.
Dataset Statistics
Total Samples: 100
Categories: angle (25), light (25), shadow_len (25), shadow_pos (25)
Language: English
Corrected / Audited Samples: 20
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/Anvesh-Lankala/pendulum-conflict.deep_bias_datadeep_bias_data_bitd_scoredb-score
B-score: Detecting Biases in Large Language Models Using Response History
by
An Vo1,
Mohammad Reza Taesiri2,
Daeyoung Kim1*,
Anh Totti Nguyen3*
*Equal advising
1KAIST, 2University of Alberta, 3Auburn University
International Conference on Machine Learning (ICML 2025)
TLDR: When LLMs can see their own previous answers, their biases significantly decrease. We introduce B-score, a novel metric that detects bias by… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/b-score.ramayana-anvayaanveshana
Dataset Card for Anveshana
Dataset Details
Dataset Description
we embarked on a comprehensive benchmarking study to explore and evaluate current state-of-the-art models for Cross-Lingual Information Retrieval (CLIR) from English to Sanskrit. Our primary objective is to assess the effectiveness of these models in accurately retrieving Sanskrit documents based on English queries. To achieve this, we meticulously assembled a robust dataset, focusing on the… See the full description on the dataset page: https://huggingface.co/datasets/manojbalaji1/anveshana.
