datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChemistryQAChemistryQA is a complex QA task which cannot be solved by end-to-end neural networks. To answer chemical questions, machines need to understand questions, apply chemistry and math knowledge, and do calculation and reasoning. ChemistryQA contains about 4500 questions covering around 200 chemistry topics, which are collected from https://socratic.org/chemistry.
All credits go to chemistry-qa project by Microsoft (https://github.com/microsoft/chemistry-qa)
Trademarks
This project may contain… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/ChemistryQA.chat_doctorThis dataset was formed from the three data sources from the ChatDoctor work.
100k real conversations between patients and doctors from HealthCareMagic.com HealthCareMagic-100k. - ADDED
10k real conversations between patients and doctors from icliniq.com icliniq-10k. - ADDED
5k generated conversations between patients and physicians from ChatGPT GenMedGPT-5k and disease database. - NOT ADDED (because of the data created by LLM, but you could add it manually)
data sample:
{'instruction': "If… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/chat_doctor.Ava-100
Empowering Agentic Video Analytics Systems with Video Language Models
[🖥️ Project Code] [📖 arXiv Paper] [📊 Dataset]
Introduction
AVA-100 is an ultra-long video benchmark specially designed to evaluate video analysis capabilities Avas-100 consists of 8 videos, each exceeding 10 hours in length, and includes a total of 120 manually annotated questions. The benchmark covers four typical video analytics scenarios: human daily activities, city walking, wildlife… See the full description on the dataset page: https://huggingface.co/datasets/iesc/Ava-100.umlsdrugchat DrugChat ChEMBL and PubChem datasets.aura_qa
Affect-Uniform ReAding QA (AURA-QA),
This dataset contains short passages from English texts found in Project Gutenberg paired with question–answer examples and emotion labels. The dataset is designed to support research in emotion-aware reading comprehension. Answers are constrained to 1–3 tokens and are generated and verified by large language models.
Dataset Structure
text — Passage excerpt
question — Question about the passage
answer — Short answer (1–3 tokens)… See the full description on the dataset page: https://huggingface.co/datasets/avalab/aura_qa.
