datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Suvach
Data Description
This dataset consists of over 100k question answers in Hindi, with 1200 tokens per question on average. Questions are generated from Wikipedia pages (Page title and Chunks). The generated part of data contain Secret Context, Question, Choices, Answer, and Description.
The question will be accompanied with 4 Choices and one and only one of them would be the correct answer. For improving generation quality, a retrieval step is added to extract a chunk of text relevant… See the full description on the dataset page: https://huggingface.co/datasets/Vaishak11a/Suvach.MCIF
Dataset Description, Collection, and Source
MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark
based on scientific talks that is designed to evaluate instruction-following in crosslingual,
multimodal settings over both short- and long-form inputs.
MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese),
enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/vaishnavikedar4/MCIF.siddha_vaithiyam_question_answering_chatbot
Medical Home Remedy Chatbot Dataset
Overview
This dataset is designed for a chatbot that answers questions related to medical problems with simple home remedies. The information in this dataset has been sourced from old books containing traditional remedies used in the past.
Contents
Dataset Files:
dataset.csv : The main dataset file containing questions and corresponding home remedy answers.
Data Structure:
Each row in the CSV file… See the full description on the dataset page: https://huggingface.co/datasets/RahulS3/siddha_vaithiyam_question_answering_chatbot.VaidyaToolcallingcrypto_qna_dataset
💰 Crypto Q&A Dataset
This dataset contains 1,081 curated Question-Answer pairs focused on cryptocurrency and blockchain concepts.It is ideal for fine-tuning LLMs, building chatbots, or conducting research in the crypto domain.
📊 Dataset Details
Number of Samples: 1,081
Format: Parquet (auto-converted from JSON)
Language: English
Domain: Cryptocurrency, Blockchain, DeFi
License: MIT (free to use for research and commercial purposes)
🧠 Example
{… See the full description on the dataset page: https://huggingface.co/datasets/Vaibhav7625/crypto_qna_dataset.journal-of-geophysics-LAI
📚 Journal of Geophysics LAI
A Historical Scientific Document Retrieval Benchmark
Inspired by the weaviate/IRPAPERS evaluation framework
📌 Overview
Journal of Geophysics LAI is a specialized historical document retrieval benchmark constructed from scanned 1980 volumes (Volume 48) of the Journal of Geophysics.
The primary objective of this dataset is to provide a rigorous testing ground for document retrieval pipelines under real-world degradation. It… See the full description on the dataset page: https://huggingface.co/datasets/vaishnavi0704/journal-of-geophysics-LAI.core
CORE: Comprehensive Ontological Relation Evaluation
🌐 Website |
📄 Paper |
💻 Code
Dataset Summary
CORE is a human-grounded benchmark for evaluating large language models on fundamental semantic and ontological reasoning. It assesses whether models can correctly recognize a broad range of sense-level relations and, critically, identify when no meaningful relationship exists between concepts. With comprehensive relation coverage and strong human baselines… See the full description on the dataset page: https://huggingface.co/datasets/vaikhari-ai/core.
