datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASearcher-Local-Knowledgeaudio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.pie-synthetic
PIE synthetic dataset
Repo: https://github.com/awasthiabhijeet/PIE
Paper: https://aclanthology.org/D19-1435.pdf
merlin
MERLIN corpus
Project URL: https://merlin-platform.eu/C_mcorpus.php
Dataset URL: https://clarin.eurac.edu/repository/xmlui/handle/20.500.12124/6
The MERLIN corpus is a written learner corpus for Czech, German, and Italian that has been designed to illustrate the Common European Framework of Reference for Languages (CEFR) with authentic learner data. The corpus contains learner texts produced in standardized language certifications covering CEFR levels A1-C1. The MERLIN annotation… See the full description on the dataset page: https://huggingface.co/datasets/aseifert/merlin.dual-lidar-combined-filtered-joint-positions-long-gripper-trainable
Long-gripper dual-UMI BiYAM joints
Published 14-D states are preserved exactly; see meta/materialization.json.
ASearcher_en_no-math_Qwen3-8B-reject-sampleSee our blog for details:
Cut the Bill, Keep the Turns: Affordable Multi-Turn Search RL
This dataset is originally from https://huggingface.co/datasets/inclusionAI/ASearcher-train-data
We do filtering to original data:
Remove Chinese samples
Our wiki server does not handle Chinese retrieval well and may return garbled text;
There are about 2k Chinese-related samples in the ASearcher dataset, which we remove entirely.
Remove math problems
Use formula patterns / specific regexes to filter… See the full description on the dataset page: https://huggingface.co/datasets/aidenjhwu/ASearcher_en_no-math_Qwen3-8B-reject-sample.asena_Chat_Dataset_tr
Turkish Chat Dataset 🇹🇷
Dataset Özeti
Turkish Chat Dataset, Google Gemini 2.5 Flash kullanılarak özel olarak üretilmiş ve çok katmanlı kalite filtreleme süreçlerinden geçirilmiş, Türkçe için kapsamlı çok-turlu konuşma veri setlerinden biridir. 150,000 premium kalite diyalog örneği içeren bu dataset, doğal ve akıcı Türkçe konuşma AI'ları geliştirmek için optimize edilmiştir.
🎯 Ne Farklı Kılıyor?
Premium AI Üretimi: Google'ın en gelişmiş Gemini 2.5 Flash… See the full description on the dataset page: https://huggingface.co/datasets/limeXx/asena_Chat_Dataset_tr.AjinkyaChat_Asesor_Inmaserbench
