datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AVQA
Summary | 摘要
This dataset is collected from the AVQA training subset (train_qa.json). We converted the data to the R1-AQA format, where each line in the text file represents a JSON object with specific keys.
The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds).
Given that there is no quick link to the audio mentioned in the above two… See the full description on the dataset page: https://huggingface.co/datasets/Joysw909/AVQA.German-RAG-SFT-ShareGPT-HESSIAN-AI
German-RAG-SFT (Supervised Fine-Tuning) Share-GPT Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The SFT Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. Most tasks were developed using synthetically enhanced data derived from the German Wikipedia, accessed through Cohere's dataset (wikipedia-22-12-de-embeddings). The data is structured in a training knowledge… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-SFT-ShareGPT-HESSIAN-AI.chat_doctorThis dataset was formed from the three data sources from the ChatDoctor work.
100k real conversations between patients and doctors from HealthCareMagic.com HealthCareMagic-100k. - ADDED
10k real conversations between patients and doctors from icliniq.com icliniq-10k. - ADDED
5k generated conversations between patients and physicians from ChatGPT GenMedGPT-5k and disease database. - NOT ADDED (because of the data created by LLM, but you could add it manually)
data sample:
{'instruction': "If… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/chat_doctor.valor32k-avqa-v2
Valor32k-AVQA v2.0
Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (visual, audio, or audio-visual) and one of six categories: description, action, count, temporal, location, and relative-position.
Links
Paper: ACM Digital Library
Project page: inesriahi.github.io/valor32k-avqa-2
Code and… See the full description on the dataset page: https://huggingface.co/datasets/inesriahi/valor32k-avqa-v2.Ava-100
Empowering Agentic Video Analytics Systems with Video Language Models
[🖥️ Project Code] [📖 arXiv Paper] [📊 Dataset]
Introduction
AVA-100 is an ultra-long video benchmark specially designed to evaluate video analysis capabilities Avas-100 consists of 8 videos, each exceeding 10 hours in length, and includes a total of 120 manually annotated questions. The benchmark covers four typical video analytics scenarios: human daily activities, city walking, wildlife… See the full description on the dataset page: https://huggingface.co/datasets/iesc/Ava-100.German-RAG-SFT-Alpaca-HESSIAN-AI
German-RAG-SFT (Supervised Fine-Tuning) Alpaca-Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The SFT Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. Most tasks were developed using synthetically enhanced data derived from the German Wikipedia, accessed through Cohere's dataset (wikipedia-22-12-de-embeddings). The data is structured in a training knowledge… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-SFT-Alpaca-HESSIAN-AI.German-RAG-ORPO-Alpaca-HESSIAN-AI
German-RAG-ORPO (Odds Ratio Preference Optimization) Alpaca-Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The ORPO Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities.
The subsets can be for this training step are derived from 2 different sources:
SauerkrautLM Preference Datasets:
SauerkrautLM-Fermented-GER-DPO: is a specialized dataset designed for training… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-ORPO-Alpaca-HESSIAN-AI.German-RAG-ORPO-ShareGPT-HESSIAN-AI
German-RAG-ORPO (Odds Ratio Preference Optimization) ShareGPT-Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The ORPO Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities.
The subsets can be for this training step are derived from 3 different sources:
SauerkrautLM Preference Datasets:
SauerkrautLM-Fermented-GER-DPO: is a specialized dataset designed for training… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-ORPO-ShareGPT-HESSIAN-AI.umlsopenclaw-recursive-study-data
OpenClaw Recursive Repository Study Data
Synthetic repository-study data generated against
openclaw/openclaw at commit
da228660306b55a9cce3b973946f3aacfc515848. The source repository is MIT licensed.
This release contains exploration questions, tool-using study trajectories,
recursive notes, full recall-rewritten trajectories, and recall-to-action
training examples. Nested chat/tool objects are stored as JSON strings to keep
the schema stable and can be decoded with json.loads.… See the full description on the dataset page: https://huggingface.co/datasets/aviralku/openclaw-recursive-study-data.German-RAG-DPO-ShareGPT-HESSIAN-AI
German-RAG-DPO Share-GPT Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The DPO Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. Most tasks were developed using synthetically enhanced data derived from the German Wikipedia, accessed through Cohere's dataset (wikipedia-22-12-de-embeddings). The data is structured in a training knowledge graph where… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-DPO-ShareGPT-HESSIAN-AI.AVQA-Audio-Rubrics
AVQA Audio-Reasoning Rubrics
Project Page | Paper | Code
Audio-grounded, binary-evaluable evaluation rubrics for the full
AVQA training set, generated for
process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with
rubric-as-reward).
Each training question is annotated with 5 rubrics, one per evaluation
facet, that judge the quality of an audio-reasoning response — not just final
answer correctness. The rubrics are designed to be scored Yes/No by an
LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.German-RAG-DPO-Alpaca-HESSIAN-AI
German-RAG-DPO Alpaca Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The DPO Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. Most tasks were developed using synthetically enhanced data derived from the German Wikipedia, accessed through Cohere's dataset (wikipedia-22-12-de-embeddings). The data is structured in a training knowledge graph where Question-Answer… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-DPO-Alpaca-HESSIAN-AI.German-RAG-ORPO-Long-Context-ShareGPT-HESSIAN-AI
German-RAG-ORPO (Odds Ratio Preference Optimization) Long Context ShareGPT-Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The ORPO Long Context Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities.
The subsets are derived from Synthetic generation inspired by Tencent's (“Scaling Synthetic Data Creation with 1,000,000,000 Personas”).
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-ORPO-Long-Context-ShareGPT-HESSIAN-AI.German-RAG-LLM-EASY-BENCHMARK
German-RAG-LLM-EASY-BENCHMARK
German-RAG - German Retrieval Augmented Generation
Dataset Summary
This German-RAG-LLM-BENCHMARK represents a specialized collection for evaluating language models with a focus on source citation, time difference stating in RAG-specific tasks.
To evaluate models compatible with OpenAI-Endpoints you can refer to our Github Repo: https://github.com/avemio-digital/German-RAG-LLM-EASY-BENCHMARK/
Most of the Subsets are synthetically… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-LLM-EASY-BENCHMARK.turkish-wikipedia-qa
Turkish Wikipedia Q&A
Türkçe Vikipedi paragraflarından türetilmiş Türkçe Soru‑Cevap (Q&A) veri seti. Her kayıtta kaynak başlık/URL ve lisans bilgisi bulunur. Veri seti türev çalışmadır; CC BY‑SA 3.0 şartları (attribution + share‑alike) geçerlidir.
Dataset Details
Dataset Sources
Kaynak metin: Turkish Wikipedia (trwiki) — Wikimedia Dumps üzerinden alınan içerik.
Bağlantı (genel): https://dumps.wikimedia.org/
Uses
Direct Use
Türkçe Q&A… See the full description on the dataset page: https://huggingface.co/datasets/avometre/turkish-wikipedia-qa.AV-TAU
AV-TAU Dataset
Overview
The AV-TAU Dataset is developed for advancing research in traffic anomaly understanding and reasoning.It provides carefully annotated question–answer pairs in English, aligned with corresponding video data, and is intended exclusively for academic and educational use.
Related Publication
🔗 EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights (CVPR 2025)
⚠️ Usage Restrictions (Academic Use Only)… See the full description on the dataset page: https://huggingface.co/datasets/harryhsing/AV-TAU.aura_qa
Affect-Uniform ReAding QA (AURA-QA),
This dataset contains short passages from English texts found in Project Gutenberg paired with question–answer examples and emotion labels. The dataset is designed to support research in emotion-aware reading comprehension. Answers are constrained to 1–3 tokens and are generated and verified by large language models.
Dataset Structure
text — Passage excerpt
question — Question about the passage
answer — Short answer (1–3 tokens)… See the full description on the dataset page: https://huggingface.co/datasets/avalab/aura_qa.German-RAG-ORPO-Long-Context-Alpaca-HESSIAN-AI
German-RAG-ORPO (Odds Ratio Preference Optimization) Long-Context Alpaca-Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The ORPO Long Context Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities.
The subsets are derived from Synthetic generation inspired by Tencent's (“Scaling Synthetic Data Creation with 1,000,000,000 Personas”).
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-ORPO-Long-Context-Alpaca-HESSIAN-AI.industrial-fault-codes-sample
FixFaults Industrial Fault Codes — Sample
A public sample (currently 2,808 — refreshed weekly from the live catalog) of the FixFaults industrial repair encyclopedia. Each row pairs a manufacturer fault code with a human-readable description and recommended repair action, ready to load for diagnostic-LLM fine-tuning, evaluation, or retrieval.
Full corpus: 70,000+ codes across 82+ manufacturers (Caterpillar, Cummins, Siemens, Fanuc, ABB, Bobcat, Haas, Yaskawa, Daikin, OBD-II, J1939… See the full description on the dataset page: https://huggingface.co/datasets/avneetsingla/industrial-fault-codes-sample.German-RAG-LLM-HARD-BENCHMARK
German-RAG-LLM-HARD Benchmark
German-RAG - German Retrieval Augmented Generation
Dataset Summary
This German-RAG-LLM-HARD-BENCHMARK represents a specialized collection for evaluate language models with a focus on hard to solve RAG-specific capabilities. To evaluate models compatible with OpenAI-Endpoints you can refer to our Github Repo: https://github.com/avemio-digital/GRAG-LLM-HARD-BENCHMARK
The subsets are derived from Synthetic generation inspired by… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-LLM-HARD-BENCHMARK.German-RAG-CPT-HESSIAN-AI
German-RAG-CPT (Continued Pre-Training) Tasks Dataset
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The CPT Tasks Dataset is a comprehensive collection designed for continued pre-training of language models, focusing on three core competencies: context-based question answering, structured reasoning, and summarization. The dataset comprises approximately 620,000 examples, with 420,000 in German and 200,000 in English.
Developed by Avemio AG… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-CPT-HESSIAN-AI.
