CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Joysw909 /AVQA Summary | 摘要 This dataset is collected from the AVQA training subset (train_qa.json). We converted the data to the R1-AQA format, where each line in the text file represents a JSON object with specific keys. The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds). Given that there is no quick link to the audio mentioned in the above two… See the full description on the dataset page: https://huggingface.co/datasets/Joysw909/AVQA.audioquestion-answering10K<n<100K2 likes3k downloads11mo agoHugging Face02avemio /German-RAG-SFT-ShareGPT-HESSIAN-AI German-RAG-SFT (Supervised Fine-Tuning) Share-GPT Format German-RAG - German Retrieval Augmented Generation Dataset Summary The SFT Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. Most tasks were developed using synthetically enhanced data derived from the German Wikipedia, accessed through Cohere's dataset (wikipedia-22-12-de-embeddings). The data is structured in a training knowledge… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-SFT-ShareGPT-HESSIAN-AI.texttext-classification1M<n<10M2 likes267 downloads2y agoHugging Face03avaliev /chat_doctorThis dataset was formed from the three data sources from the ChatDoctor work. 100k real conversations between patients and doctors from HealthCareMagic.com HealthCareMagic-100k. - ADDED 10k real conversations between patients and doctors from icliniq.com icliniq-10k. - ADDED 5k generated conversations between patients and physicians from ChatGPT GenMedGPT-5k and disease database. - NOT ADDED (because of the data created by LLM, but you could add it manually) data sample: {'instruction': "If… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/chat_doctor.textquestion-answering100K<n<1M15 likes228 downloads3y agoHugging Face04inesriahi /valor32k-avqa-v2 Valor32k-AVQA v2.0 Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (visual, audio, or audio-visual) and one of six categories: description, action, count, temporal, location, and relative-position. Links Paper: ACM Digital Library Project page: inesriahi.github.io/valor32k-avqa-2 Code and… See the full description on the dataset page: https://huggingface.co/datasets/inesriahi/valor32k-avqa-v2.tabularquestion-answering100K<n<1M0 likes215 downloads3mo agoHugging Face05iesc /Ava-100 Empowering Agentic Video Analytics Systems with Video Language Models [🖥️ Project Code] [📖 arXiv Paper] [📊 Dataset] Introduction AVA-100 is an ultra-long video benchmark specially designed to evaluate video analysis capabilities Avas-100 consists of 8 videos, each exceeding 10 hours in length, and includes a total of 120 manually annotated questions. The benchmark covers four typical video analytics scenarios: human daily activities, city walking, wildlife… See the full description on the dataset page: https://huggingface.co/datasets/iesc/Ava-100.textmultiple-choicen<1K2 likes201 downloads11mo agoHugging Face06avemio /German-RAG-SFT-Alpaca-HESSIAN-AI German-RAG-SFT (Supervised Fine-Tuning) Alpaca-Format German-RAG - German Retrieval Augmented Generation Dataset Summary The SFT Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. Most tasks were developed using synthetically enhanced data derived from the German Wikipedia, accessed through Cohere's dataset (wikipedia-22-12-de-embeddings). The data is structured in a training knowledge… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-SFT-Alpaca-HESSIAN-AI.texttext-classification100K<n<1M1 likes166 downloads2y agoHugging Face07avemio /German-RAG-ORPO-Alpaca-HESSIAN-AI German-RAG-ORPO (Odds Ratio Preference Optimization) Alpaca-Format German-RAG - German Retrieval Augmented Generation Dataset Summary The ORPO Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. The subsets can be for this training step are derived from 2 different sources: SauerkrautLM Preference Datasets: SauerkrautLM-Fermented-GER-DPO: is a specialized dataset designed for training… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-ORPO-Alpaca-HESSIAN-AI.textquestion-answering10K<n<100K0 likes155 downloads2y agoHugging Face08avemio /German-RAG-ORPO-ShareGPT-HESSIAN-AI German-RAG-ORPO (Odds Ratio Preference Optimization) ShareGPT-Format German-RAG - German Retrieval Augmented Generation Dataset Summary The ORPO Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. The subsets can be for this training step are derived from 3 different sources: SauerkrautLM Preference Datasets: SauerkrautLM-Fermented-GER-DPO: is a specialized dataset designed for training… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-ORPO-ShareGPT-HESSIAN-AI.textquestion-answering10K<n<100K2 likes137 downloads2y agoHugging Face09avaliev /umlstextquestion-answering10K<n<100K4 likes112 downloads3y agoHugging Face10aviralku /openclaw-recursive-study-data OpenClaw Recursive Repository Study Data Synthetic repository-study data generated against openclaw/openclaw at commit da228660306b55a9cce3b973946f3aacfc515848. The source repository is MIT licensed. This release contains exploration questions, tool-using study trajectories, recursive notes, full recall-rewritten trajectories, and recall-to-action training examples. Nested chat/tool objects are stored as JSON strings to keep the schema stable and can be decoded with json.loads.… See the full description on the dataset page: https://huggingface.co/datasets/aviralku/openclaw-recursive-study-data.tabularquestion-answering100K<n<1M0 likes95 downloads11d agoHugging Face11avemio /German-RAG-DPO-ShareGPT-HESSIAN-AI German-RAG-DPO Share-GPT Format German-RAG - German Retrieval Augmented Generation Dataset Summary The DPO Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. Most tasks were developed using synthetically enhanced data derived from the German Wikipedia, accessed through Cohere's dataset (wikipedia-22-12-de-embeddings). The data is structured in a training knowledge graph where… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-DPO-ShareGPT-HESSIAN-AI.textquestion-answering100K<n<1M0 likes91 downloads2y agoHugging Face12umd-zhou-lab /AVQA-Audio-Rubrics AVQA Audio-Reasoning Rubrics Project Page | Paper | Code Audio-grounded, binary-evaluable evaluation rubrics for the full AVQA training set, generated for process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with rubric-as-reward). Each training question is annotated with 5 rubrics, one per evaluation facet, that judge the quality of an audio-reasoning response — not just final answer correctness. The rubrics are designed to be scored Yes/No by an LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.textaudio-classification10K<n<100K1 likes75 downloads2mo agoHugging Face13avemio /German-RAG-DPO-Alpaca-HESSIAN-AI German-RAG-DPO Alpaca Format German-RAG - German Retrieval Augmented Generation Dataset Summary The DPO Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. Most tasks were developed using synthetically enhanced data derived from the German Wikipedia, accessed through Cohere's dataset (wikipedia-22-12-de-embeddings). The data is structured in a training knowledge graph where Question-Answer… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-DPO-Alpaca-HESSIAN-AI.textquestion-answering100K<n<1M1 likes70 downloads2y agoHugging Face14avemio /German-RAG-ORPO-Long-Context-ShareGPT-HESSIAN-AI German-RAG-ORPO (Odds Ratio Preference Optimization) Long Context ShareGPT-Format German-RAG - German Retrieval Augmented Generation Dataset Summary The ORPO Long Context Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. The subsets are derived from Synthetic generation inspired by Tencent's (“Scaling Synthetic Data Creation with 1,000,000,000 Personas”). Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-ORPO-Long-Context-ShareGPT-HESSIAN-AI.textquestion-answering10K<n<100K1 likes51 downloads2y agoHugging Face15avemio /German-RAG-LLM-EASY-BENCHMARK German-RAG-LLM-EASY-BENCHMARK German-RAG - German Retrieval Augmented Generation Dataset Summary This German-RAG-LLM-BENCHMARK represents a specialized collection for evaluating language models with a focus on source citation, time difference stating in RAG-specific tasks. To evaluate models compatible with OpenAI-Endpoints you can refer to our Github Repo: https://github.com/avemio-digital/German-RAG-LLM-EASY-BENCHMARK/ Most of the Subsets are synthetically… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-LLM-EASY-BENCHMARK.texttext-classification1K<n<10K0 likes50 downloads2y agoHugging Face16avometre /turkish-wikipedia-qa Turkish Wikipedia Q&A Türkçe Vikipedi paragraflarından türetilmiş Türkçe Soru‑Cevap (Q&A) veri seti. Her kayıtta kaynak başlık/URL ve lisans bilgisi bulunur. Veri seti türev çalışmadır; CC BY‑SA 3.0 şartları (attribution + share‑alike) geçerlidir. Dataset Details Dataset Sources Kaynak metin: Turkish Wikipedia (trwiki) — Wikimedia Dumps üzerinden alınan içerik. Bağlantı (genel): https://dumps.wikimedia.org/ Uses Direct Use Türkçe Q&A… See the full description on the dataset page: https://huggingface.co/datasets/avometre/turkish-wikipedia-qa.textquestion-answering100K<n<1M3 likes49 downloads11mo agoHugging Face17harryhsing /AV-TAUgated AV-TAU Dataset Overview The AV-TAU Dataset is developed for advancing research in traffic anomaly understanding and reasoning.It provides carefully annotated question–answer pairs in English, aligned with corresponding video data, and is intended exclusively for academic and educational use. Related Publication 🔗 EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights (CVPR 2025) ⚠️ Usage Restrictions (Academic Use Only)… See the full description on the dataset page: https://huggingface.co/datasets/harryhsing/AV-TAU.documentquestion-answering100K<n<1M4 likes47 downloads1y agoHugging Face18avalab /aura_qa Affect-Uniform ReAding QA (AURA-QA), This dataset contains short passages from English texts found in Project Gutenberg paired with question–answer examples and emotion labels. The dataset is designed to support research in emotion-aware reading comprehension. Answers are constrained to 1–3 tokens and are generated and verified by large language models. Dataset Structure text — Passage excerpt question — Question about the passage answer — Short answer (1–3 tokens)… See the full description on the dataset page: https://huggingface.co/datasets/avalab/aura_qa.tabularquestion-answering10K<n<100K0 likes36 downloads8mo agoHugging Face19avemio /German-RAG-ORPO-Long-Context-Alpaca-HESSIAN-AI German-RAG-ORPO (Odds Ratio Preference Optimization) Long-Context Alpaca-Format German-RAG - German Retrieval Augmented Generation Dataset Summary The ORPO Long Context Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. The subsets are derived from Synthetic generation inspired by Tencent's (“Scaling Synthetic Data Creation with 1,000,000,000 Personas”). Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-ORPO-Long-Context-Alpaca-HESSIAN-AI.textquestion-answering10K<n<100K1 likes33 downloads2y agoHugging Face20avneetsingla /industrial-fault-codes-sample FixFaults Industrial Fault Codes — Sample A public sample (currently 2,808 — refreshed weekly from the live catalog) of the FixFaults industrial repair encyclopedia. Each row pairs a manufacturer fault code with a human-readable description and recommended repair action, ready to load for diagnostic-LLM fine-tuning, evaluation, or retrieval. Full corpus: 70,000+ codes across 82+ manufacturers (Caterpillar, Cummins, Siemens, Fanuc, ABB, Bobcat, Haas, Yaskawa, Daikin, OBD-II, J1939… See the full description on the dataset page: https://huggingface.co/datasets/avneetsingla/industrial-fault-codes-sample.texttext-classification1K<n<10K1 likes27 downloads4mo agoHugging Face21avemio /German-RAG-LLM-HARD-BENCHMARK German-RAG-LLM-HARD Benchmark German-RAG - German Retrieval Augmented Generation Dataset Summary This German-RAG-LLM-HARD-BENCHMARK represents a specialized collection for evaluate language models with a focus on hard to solve RAG-specific capabilities. To evaluate models compatible with OpenAI-Endpoints you can refer to our Github Repo: https://github.com/avemio-digital/GRAG-LLM-HARD-BENCHMARK The subsets are derived from Synthetic generation inspired by… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-LLM-HARD-BENCHMARK.textquestion-answeringn<1K0 likes22 downloads2y agoHugging Face22avemio /German-RAG-CPT-HESSIAN-AIgated German-RAG-CPT (Continued Pre-Training) Tasks Dataset German-RAG - German Retrieval Augmented Generation Dataset Summary The CPT Tasks Dataset is a comprehensive collection designed for continued pre-training of language models, focusing on three core competencies: context-based question answering, structured reasoning, and summarization. The dataset comprises approximately 620,000 examples, with 420,000 in German and 200,000 in English. Developed by Avemio AG… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-CPT-HESSIAN-AI.textquestion-answering100K<n<1M0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.