CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anvo25 /vlms-are-biased Vision Language Models are Biased by An Vo1*, Khai-Nguyen Nguyen2*, Mohammad Reza Taesiri3, Vy Tuong Dang1, Anh Totti Nguyen4†, Daeyoung Kim1† *Equal contribution    †Equal advising 1KAIST, 2College of William and Mary, 3University of Alberta, 4Auburn University TLDR: State-of-the-art Vision Language Models (VLMs) perform perfectly on counting tasks with original images but fail catastrophically (e.g., 100% → 17.05%… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/vlms-are-biased.imagevisual-question-answering10K<n<100K29 likes1.2k downloads10mo agoHugging Face02dsfsi-anv /za-african-next-voicesgated Swivuriso: ZA-African Next Voices Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes. Dataset Paper: ArXiv - Work in Progress Language Coverage… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices.audioautomatic-speech-recognition100K<n<1M16 likes1k downloads7mo agoHugging Face03Anv-ke /Dholuogatedaudio100K<n<1M3 likes844 downloads7mo agoHugging Face04Anv-ke /kikuyugatedaudio100K<n<1M6 likes600 downloads7mo agoHugging Face05badrex /anv_data_ke_kikuyu_mergedaudio100K<n<1M0 likes409 downloads1y agoHugging Face06anvilarth /GarageDatasettext10K<n<100K0 likes382 downloads2y agoHugging Face07dsfsi-anv /multilingual-nchlt-dataset NCHLT Auxiliary Speech Corpus - Combined Multilingual Dataset Dataset Description This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data Science for Social Impact (DSFSI) research group at the University of Pretoria to facilitate easier benchmarking and multi-language speech recognition research. The original auxiliary data was collected during the National Centre for Human Language Technology (NCHLT) project for the 11 official… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/multilingual-nchlt-dataset.audioautomatic-speech-recognition100K<n<1M1 likes379 downloads9mo agoHugging Face08badrex /anv_data_ke_kikuyu_scriptedaudio100K<n<1M0 likes363 downloads1y agoHugging Face09cryptpesa /anv-data-ke-somali-fullaudio10K<n<100K0 likes332 downloads5mo agoHugging Face10badrex /anv-data-ke-somali-fullaudio100K<n<1M1 likes326 downloads11mo agoHugging Face11MCAA1-MSU /anv_data_kegatedlanguage: ki so kln luo mas pretty_name: anv_ke ⚠️ IMPORTANT: Work in ProgressThis dataset is not final. Updates will continue through September 2025.Please use the latest version for attribution, benchmarking and publications. Overview African Next Voices: Pilot Data Collection in Kenya is part of a larger initiative to support African language speech technology. This project, funded by the Gates Foundation, is led by the KenCorpus Consortium, a coalition of Kenyan… See the full description on the dataset page: https://huggingface.co/datasets/MCAA1-MSU/anv_data_ke.audio100K<n<1M18 likes299 downloads7mo agoHugging Face12NjeriKahoro /anv-kikuyu-banking-subset-v2-part10audio1K<n<10K0 likes298 downloads23d agoHugging Face13Anv-ke /Kalenjingatedaudio10K<n<100K4 likes267 downloads7mo agoHugging Face14Anv-ke /Maasaigatedaudio10K<n<100K3 likes248 downloads7mo agoHugging Face15Anv-ke /Somaligatedaudio10K<n<100K5 likes198 downloads7mo agoHugging Face16Anvesh-Lankala /Constrained_Indic_Codemixingtext1K<n<10K0 likes165 downloads1mo agoHugging Face17anvo25 /vmmu ViExam: Are Vision Language Models Better than Humans on Vietnamese Multimodal Exam Questions? by Vy Tuong Dang*, An Vo*, Quang Tau, Duc Dm, Daeyoung Kim, *Equal contribution  KAIST TLDR: State-of-the-art Vision Language Models (VLMs) demonstrate remarkable capabilities on English multimodal tasks but significantly underperform on Vietnamese educational assessments. ViExam reveals that SOTA VLMs achieve only 57.74% accuracy… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/vmmu.imageimage-text-to-text1K<n<10K2 likes136 downloads1y agoHugging Face18Anvesh-Lankala /Radiology_Project_Annotatedimage10K<n<100K0 likes127 downloads3mo agoHugging Face19anvo25 /deep_bias_resultstabular1M<n<10M0 likes110 downloads28d agoHugging Face20AnveshAI /AnveshAI-Coder-150K Coding AI Dataset - 150K Problem/Thinking/Solution Entries Overview A comprehensive dataset of 150,000 coding problems designed for training code generation AI models. Each entry contains a problem statement, detailed thinking process, and complete solution in the target language. Format JSONL (JSON Lines) - each line is a valid JSON object. Schema { "id": "unique identifier string", "domain": "algorithms | data_structures |… See the full description on the dataset page: https://huggingface.co/datasets/AnveshAI/AnveshAI-Coder-150K.text100K<n<1M1 likes105 downloads3mo agoHugging Face21NjeriKahoro /anv-kikuyu-banking-subset Anv-Kikuyu Banking Subset A domain-filtered subset of Kikuyu (Gĩkũyũ) speech data focused on banking and financial-transaction content, combined into a single repository with train, test, and validation splits. Source This dataset is a filtered subset of Anv-ke/kikuyu, part of the African Next Voices (ANV) collection. All audio, transcriptions, and underlying speaker data originate from that source dataset. Full credit for data collection belongs to the… See the full description on the dataset page: https://huggingface.co/datasets/NjeriKahoro/anv-kikuyu-banking-subset.audioautomatic-speech-recognition1K<n<10K1 likes105 downloads1mo agoHugging Face22badrex /anv-data-ke-somaliaudio10K<n<100K0 likes101 downloads11mo agoHugging Face23dsfsi-anv /za-african-next-voices-compressedgatedNote: This dataset is a compressed version of za-african-next-voices. It was compressed to .opus format using a 32k bitrate. Swivuriso: ZA-African Next Voices-Compressed Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices-compressed.audioautomatic-speech-recognition100K<n<1M1 likes86 downloads8mo agoHugging Face24anvo2 /perfume-rec-assetstext10K<n<100K1 likes70 downloads6mo agoHugging Face25Anvesh-Lankala /pendulum-conflict Pendulum Conflict Dataset This dataset is a manually curated subset of 100 isolated multimodal conflict examples derived from the akomand/counterfactual_pendulum dataset. It is designed for evaluating Multimodal Large Language Models (MLLMs) under controlled visual-textual conflicts. Dataset Statistics Total Samples: 100 Categories: angle (25), light (25), shadow_len (25), shadow_pos (25) Language: English Corrected / Audited Samples: 20 Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/Anvesh-Lankala/pendulum-conflict.image1K<n<10K0 likes63 downloads2mo agoHugging Face26anvo25 /deep_bias_datatabular100K<n<1M0 likes62 downloads28d agoHugging Face27anvo25 /deep_bias_data_bitd_scoredtabular100K<n<1M0 likes57 downloads28d agoHugging Face28anvo25 /b-score B-score: Detecting Biases in Large Language Models Using Response History by An Vo1, Mohammad Reza Taesiri2, Daeyoung Kim1*, Anh Totti Nguyen3* *Equal advising 1KAIST, 2University of Alberta, 3Auburn University International Conference on Machine Learning (ICML 2025) TLDR: When LLMs can see their own previous answers, their biases significantly decrease. We introduce B-score, a novel metric that detects bias by… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/b-score.texttext-classificationn<1K2 likes53 downloads1y agoHugging Face29sanganaka /ramayana-anvayatext10K<n<100K0 likes45 downloads11mo agoHugging Face30manojbalaji1 /anveshana Dataset Card for Anveshana Dataset Details Dataset Description we embarked on a comprehensive benchmarking study to explore and evaluate current state-of-the-art models for Cross-Lingual Information Retrieval (CLIR) from English to Sanskrit. Our primary objective is to assess the effectiveness of these models in accurately retrieving Sanskrit documents based on English queries. To achieve this, we meticulously assembled a robust dataset, focusing on the… See the full description on the dataset page: https://huggingface.co/datasets/manojbalaji1/anveshana.texttext-retrieval10K<n<100K1 likes39 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.