datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coco-qa-vi
COCO-QA Vietnamese
📖 Overview
COCO-QA Vietnamese is a fully translated Vietnamese version of the popular COCO-QA dataset for Visual Question Answering (VQA) tasks.It contains over 117 684 image-based question-answer pairs translated into Vietnamese, carefully reviewed for linguistic accuracy and contextual meaning.
4 types of questions: object, number, color, location
Answers are all one-word.
This dataset is designed for:
Research and fine-tuning of Visual Question… See the full description on the dataset page: https://huggingface.co/datasets/ThucPD/coco-qa-vi.SpatialMemoryVQA-COCO-HI
Dataset Information
This dataset was translated from ShareGPTV English to Hindi
Paper or resources for more information:
[Project] [Paper] [Code]
License:
Attribution-NonCommercial 4.0 International
It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use
Intended use
Primary intended uses:
The primary use of ShareGPT4V Captions 1.2M is research on large multimodal models and chatbots.Primary intended users:
The primary intended users of this… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/VQA-COCO-HI.NeuroQABenchmark
Intro
NeuroQA is a neuroscience-specific dataset comprising 11k training and
2k testing question-answer pairs
Demo
Our dataset includes both Chinese and English. Training set is open question, testing data is single-choice question.
Pipeline
We ask QWEN2.5 to extract the fact in abstract and form the QA pair
Source
Brain_Tumor_pubmed_abstracts
PMC_neuroscience
Trained Model
Much welcome to try out NeuroExpert_Qwen2.5.
COCO_GridQA
COCO-GridQA Dataset
Overview
The COCO-GridQA dataset is a derived dataset created from the COCO (Common Objects in Context) validation set. It focuses on spatial reasoning tasks by arranging object crops from COCO images into a 2x2 grid and providing question-answer pairs about the positions of objects within the grid.
This dataset is designed for tasks such as spatial reasoning, visual question answering (VQA), and object localization. Each sample consists of:
A… See the full description on the dataset page: https://huggingface.co/datasets/hoveringgull/COCO_GridQA.CoCoPIFcoco-captions_marathi
Coco-Captions Marathi Dataset: High-Quality Marathi NLP Corpus
📌 Overview
The Coco-Captions Marathi dataset is a meticulously curated collection of 414010 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness.
This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource for machine learning… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/coco-captions_marathi.
