datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anvo22121anvu16231anvaya-sangraha-rawanv-ke-transfer-checkpointvlms-are-biased
Vision Language Models are Biased
by
An Vo1*,
Khai-Nguyen Nguyen2*,
Mohammad Reza Taesiri3,
Vy Tuong Dang1,
Anh Totti Nguyen4†,
Daeyoung Kim1†
*Equal contribution †Equal advising
1KAIST, 2College of William and Mary, 3University of Alberta, 4Auburn University
TLDR: State-of-the-art Vision Language Models (VLMs) perform perfectly on counting tasks with original images but fail catastrophically (e.g., 100% → 17.05%… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/vlms-are-biased.anv-ke-audioza-african-next-voices
Swivuriso: ZA-African Next Voices
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes.
Dataset Paper: ArXiv - Work in Progress
Language Coverage… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices.Dholuoanvu59577kikuyuanv_data_ke_kikuyu_mergedGarageDatasetmultilingual-nchlt-dataset
NCHLT Auxiliary Speech Corpus - Combined Multilingual Dataset
Dataset Description
This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data Science for Social Impact (DSFSI) research group at the University of Pretoria to facilitate easier benchmarking and multi-language speech recognition research.
The original auxiliary data was collected during the National Centre for Human Language Technology (NCHLT) project for the 11 official… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/multilingual-nchlt-dataset.anv_data_ke_kikuyu_scriptedanv-data-ke-somali-fullanv-data-ke-somali-fullanv_data_kelanguage:
ki
so
kln
luo
mas
pretty_name: anv_ke
⚠️ IMPORTANT: Work in ProgressThis dataset is not final. Updates will continue through September 2025.Please use the latest version for attribution, benchmarking and publications.
Overview
African Next Voices: Pilot Data Collection in Kenya is part of a larger initiative to support African language speech technology. This project, funded by the Gates Foundation, is led by the KenCorpus Consortium, a coalition of Kenyan… See the full description on the dataset page: https://huggingface.co/datasets/MCAA1-MSU/anv_data_ke.anv-kikuyu-banking-subset-v2-part10KalenjinMaasaiViscous_Cahn_Hilliard_2D_Spatio-Temporal
Dataset Card: Viscous Cahn-Hilliard Optimal Control
Dataset Summary
This dataset contains 2,000 high-fidelity simulations of the Viscous Cahn-Hilliard (vCH) equation under randomized control forcing. It was generated to support research into Sparse Optimal Control, SciML (Scientific Machine Learning), and Phase Field Modeling.
Official Code Repository: Sparse-optimal-control-of-Viscous-Chan-hilliard (GitHub)
Each sample represents the evolution of a two-phase system… See the full description on the dataset page: https://huggingface.co/datasets/Tejas-Anvekar/Viscous_Cahn_Hilliard_2D_Spatio-Temporal.SomaliSO101_relocate_cube_2cams_record_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 100,
"total_frames": 15000,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/anvilbot-patrickhhh/SO101_relocate_cube_2cams_record_2.Constrained_Indic_CodemixingSO101_PickAndPlace_front_wristThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 7500,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/anvilbot-patrickhhh/SO101_PickAndPlace_front_wrist.anvisa_instruct_tokenizedvmmu
ViExam: Are Vision Language Models Better than Humans on Vietnamese Multimodal Exam Questions?
by
Vy Tuong Dang*,
An Vo*,
Quang Tau,
Duc Dm,
Daeyoung Kim,
*Equal contribution
KAIST
TLDR: State-of-the-art Vision Language Models (VLMs) demonstrate remarkable capabilities on English multimodal tasks but significantly underperform on Vietnamese educational assessments. ViExam reveals that SOTA VLMs achieve only 57.74% accuracy… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/vmmu.SO101_record_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 535,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/anvilbot-patrickhhh/SO101_record_test.constellationRadiology_Project_Annotated
