datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hungarian_doc_qa_beirThis is a copy of https://huggingface.co/datasets/jinaai/hungarian_doc_qa reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/hungarian_doc_qa_beir.Hungarian-Dialogues-text
LLM-Generated Hungarian Conversations
This dataset contains structured Hungarian conversations generated with multiple large language model families for the study “Efficient ASR Training with Conversations that Never Happened.” Paper: arXiv link
Each model is provided as a separate Parquet file. The dataset contains the generated textual conversations and associated scenario and participant metadata; it does not contain synthesized audio.
Dataset structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/gedeonmate/Hungarian-Dialogues-text.hungarian-toxic-comments
Hungarian Toxic Comments
The first openly available Hungarian dataset for toxic comment classification, introduced in:
Hatvani, P., & Yang, Z. Gy. (2025). Automated detection of toxic comments in Hungarian. Annales Mathematicae et Informaticae, 61, 108-117. DOI: 10.33039/ami.2025.10.007
Dataset Description
This dataset contains 654 manually annotated Hungarian-language comments collected from social media and political news forums. Each comment is annotated across five… See the full description on the dataset page: https://huggingface.co/datasets/RabidUmarell/hungarian-toxic-comments.hungarian_encryptedhungarian_encrypted_HistCiph
Dataset Card for HistCiph — Hungarian
Dataset Description
Dataset Summary
The Hungarian subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical Hungarian plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/hungarian_encrypted_HistCiph.
