datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Translations_Hungarian_public_websites
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/18982
Description
A webcrawl of 14 different websites covering parallel corpora of Hungarian with Polish, Czech, Swedish, Finnish, French, German, Italian, English and Slovenian
Citation
Translations of Hungarian from public websites (2022). Version 1.0. [Dataset (Text corpus)]. Source: European Language Grid. https://live.european-language-grid.eu/catalogue/corpus/18982
hungarian-youtube-speechhungarian_national_hs_finals_exam
Testing Language Models on a Held-Out High School National Finals Exam
When xAI recently released Grok-1, they evaluated it on the 2023 Hungarian national high school finals in mathematics, which was published after the training data cutoff for all the models in their evaluation. While MATH and GSM8k are the standard benchmarks for evaluating the mathematical abilities of large language models, there are risks that modern models overfit to these datasets, either from training… See the full description on the dataset page: https://huggingface.co/datasets/keirp/hungarian_national_hs_finals_exam.HungarianDocQA_IT_SynQA_ocr_v3hungarian_doc_qa_beirThis is a copy of https://huggingface.co/datasets/jinaai/hungarian_doc_qa reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/hungarian_doc_qa_beir.hungarian-llm-testing
Hungarian llm testing
This is a really simple data-set to test fine-tuning a language model on Hungarian text.
hungarian-single-speaker-tts
Dataset Card for CSS10 Hungarian: Single Speaker Speech Dataset
Dataset Summary
The corpus consists of a single speaker, with 4515 segments extracted
from a single LibriVox audiobook.
Supported Tasks and Leaderboards
[Needs More Information]
Languages
The audio is in Hungarian.
Dataset Structure
[Needs More Information]
Data Instances
[Needs More Information]
Data Fields
[Needs More Information]
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/KTH/hungarian-single-speaker-tts.TTS-Hungarian
TTS-Hungarian
A large-scale, high-quality Hungarian speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from MEK (Magyar Elektronikus Könyvtár) — Hungarian audiobooks.
Dataset Statistics
Metric
Value
Total samples
253,116
Total duration
702 hours
Unique speakers
100
Average duration
10.0 seconds
Average DNSMOS
3.68
Features
Field
Type
Description
__key__
string
Unique sample… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Hungarian.HungarianNamesThis dataset was collected from Wikipedia : https://hu.wikipedia.org/wiki/Magyarorsz%C3%A1gon_anyak%C3%B6nyvezhet%C5%91_ut%C3%B3nevek_list%C3%A1ja
HungarianDocQA_IT_IfEvalQA_v2HungarianDocQA_IT_IFEvalQAHungarian-Dialogues-text
LLM-Generated Hungarian Conversations
This dataset contains structured Hungarian conversations generated with multiple large language model families for the study “Efficient ASR Training with Conversations that Never Happened.” Paper: arXiv link
Each model is provided as a separate Parquet file. The dataset contains the generated textual conversations and associated scenario and participant metadata; it does not contain synthesized audio.
Dataset structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/gedeonmate/Hungarian-Dialogues-text.HungarianDocQA_IT_SynQA_ocr_v3_cleanedhungarian-cc-corpushungarian-news-translationsHungarianDocQA_IT_SyntheticQAalpaca-hungarian-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-hungarian-cleaned.alpaca_hungarian_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_hungarian_taco.Hungarian-Train-Data-DestilledGeminihungarian-speech-datasethungarian-toxic-comments
Hungarian Toxic Comments
The first openly available Hungarian dataset for toxic comment classification, introduced in:
Hatvani, P., & Yang, Z. Gy. (2025). Automated detection of toxic comments in Hungarian. Annales Mathematicae et Informaticae, 61, 108-117. DOI: 10.33039/ami.2025.10.007
Dataset Description
This dataset contains 654 manually annotated Hungarian-language comments collected from social media and political news forums. Each comment is annotated across five… See the full description on the dataset page: https://huggingface.co/datasets/RabidUmarell/hungarian-toxic-comments.hungarian-riddles-benchmark
Hungarian Riddles Benchmark
Overview
This dataset is a cultural and reasoning benchmark based on 100 metaphorical, trivia-style Hungarian riddles.
The riddles are intentionally tricky and culturally grounded. They are designed to test answer correctness and reasoning quality, not only surface-level language fluency.
Dataset structure
Each row contains one riddle with reference material for evaluation.
Fields
ID – unique identifier
topic – general… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/hungarian-riddles-benchmark.FreddieMercury_HungarianRhapsodyLive86MPEP_HUNGARIAN
Dataset Card for MPEP_HUNGARIAN
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/MPEP_HUNGARIAN.Hungarian-Speech-Dataset
🎧 Hungarian Speech Dataset
The Hungarian Speech Dataset is a high-quality speech audio dataset designed to support advanced AI systems that depend on diverse audio data and reliable voice data for multilingual model training. It comprises 169 hours of recordings across 743 files, provided in MP3 and WAV formats, with a total size of 134 MB. This structured audio dataset ensures balanced speaker representation, featuring 46% female and 54% male speakers, and an age distribution… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hungarian-Speech-Dataset.MED_SYN2_HUNGARIAN_trainHungarianDocQA-OCRmed_syn1_hungarianNTEU_French-Hungarian
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19581
Description
This is a compilation of parallel corpora resources used in building of Machine Translation engines in NTEU project (Action number: 2018-EU-IA-0051). Data in these resources are compiled in two TMX files, two tiers grouped by data source reliablity. Tier A -- danta originating from human edited sources, translation memories and alike. Tier B -- danta originating created by automatic… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/NTEU_French-Hungarian.Hungarian_CCPDF_SynQA_v2
