CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01keirp /hungarian_national_hs_finals_exam Testing Language Models on a Held-Out High School National Finals Exam When xAI recently released Grok-1, they evaluated it on the 2023 Hungarian national high school finals in mathematics, which was published after the training data cutoff for all the models in their evaluation. While MATH and GSM8k are the standard benchmarks for evaluating the mathematical abilities of large language models, there are risks that modern models overfit to these datasets, either from training… See the full description on the dataset page: https://huggingface.co/datasets/keirp/hungarian_national_hs_finals_exam.textn<1K27 likes491 downloads3y agoHugging Face02nhiremath /HungarianDocQA_IT_SynQA_ocr_v3image100K<n<1M0 likes354 downloads2y agoHugging Face03Bazsalanszky /hungarian-llm-testing Hungarian llm testing This is a really simple data-set to test fine-tuning a language model on Hungarian text. textn<1K5 likes247 downloads3y agoHugging Face04jinaai /hungarian_doc_qa_beirThis is a copy of https://huggingface.co/datasets/jinaai/hungarian_doc_qa reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/hungarian_doc_qa_beir.imagen<1K0 likes247 downloads1y agoHugging Face05KTH /hungarian-single-speaker-tts Dataset Card for CSS10 Hungarian: Single Speaker Speech Dataset Dataset Summary The corpus consists of a single speaker, with 4515 segments extracted from a single LibriVox audiobook. Supported Tasks and Leaderboards [Needs More Information] Languages The audio is in Hungarian. Dataset Structure [Needs More Information] Data Instances [Needs More Information] Data Fields [Needs More Information] Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/KTH/hungarian-single-speaker-tts.audiotext-to-speech1K<n<10K14 likes179 downloads4y agoHugging Face06datadriven-company /TTS-Hungarian TTS-Hungarian A large-scale, high-quality Hungarian speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from MEK (Magyar Elektronikus Könyvtár) — Hungarian audiobooks. Dataset Statistics Metric Value Total samples 253,116 Total duration 702 hours Unique speakers 100 Average duration 10.0 seconds Average DNSMOS 3.68 Features Field Type Description __key__ string Unique sample… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Hungarian.audiotext-to-speech100K<n<1M1 likes172 downloads7mo agoHugging Face07hongfenglu /HungarianDocQA_IT_IfEvalQA_v2image10K<n<100K0 likes147 downloads2y agoHugging Face08hongfenglu /HungarianDocQA_IT_IFEvalQAimage10K<n<100K0 likes129 downloads2y agoHugging Face09nhiremath /HungarianDocQA_IT_SynQA_ocr_v3_cleanedimage10K<n<100K0 likes83 downloads2y agoHugging Face10gedeonmate /Hungarian-Dialogues-text LLM-Generated Hungarian Conversations This dataset contains structured Hungarian conversations generated with multiple large language model families for the study “Efficient ASR Training with Conversations that Never Happened.” Paper: arXiv link Each model is provided as a separate Parquet file. The dataset contains the generated textual conversations and associated scenario and participant metadata; it does not contain synthesized audio. Dataset structure Each… See the full description on the dataset page: https://huggingface.co/datasets/gedeonmate/Hungarian-Dialogues-text.tabular100K<n<1M0 likes79 downloads2mo agoHugging Face11hakatiki /hungarian-cc-corpustext10M<n<100M4 likes78 downloads3y agoHugging Face12matekadlicsko /hungarian-news-translationstexttranslation10K<n<100K1 likes72 downloads3y agoHugging Face13nhiremath /HungarianDocQA_IT_SyntheticQAimage100K<n<1M0 likes71 downloads2y agoHugging Face14saillab /alpaca-hungarian-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-hungarian-cleaned.text10K<n<100K6 likes65 downloads2y agoHugging Face15saillab /alpaca_hungarian_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_hungarian_taco.text10K<n<100K2 likes63 downloads2y agoHugging Face16BMark52 /Hungarian-Train-Data-DestilledGeminitext1K<n<10K0 likes63 downloads14d agoHugging Face17shunyalabs /hungarian-speech-datasetaudio1K<n<10K0 likes50 downloads1y agoHugging Face18RabidUmarell /hungarian-toxic-comments Hungarian Toxic Comments The first openly available Hungarian dataset for toxic comment classification, introduced in: Hatvani, P., & Yang, Z. Gy. (2025). Automated detection of toxic comments in Hungarian. Annales Mathematicae et Informaticae, 61, 108-117. DOI: 10.33039/ami.2025.10.007 Dataset Description This dataset contains 654 manually annotated Hungarian-language comments collected from social media and political news forums. Each comment is annotated across five… See the full description on the dataset page: https://huggingface.co/datasets/RabidUmarell/hungarian-toxic-comments.documenttext-classificationn<1K1 likes41 downloads7mo agoHugging Face19Speech-data /Hungarian-Speech-Dataset 🎧 Hungarian Speech Dataset The Hungarian Speech Dataset is a high-quality speech audio dataset designed to support advanced AI systems that depend on diverse audio data and reliable voice data for multilingual model training. It comprises 169 hours of recordings across 743 files, provided in MP3 and WAV formats, with a total size of 134 MB. This structured audio dataset ensures balanced speaker representation, featuring 46% female and 54% male speakers, and an age distribution… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hungarian-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes33 downloads6mo agoHugging Face20data-is-better-together /MPEP_HUNGARIAN Dataset Card for MPEP_HUNGARIAN This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Dataset Summary This dataset contains: A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/MPEP_HUNGARIAN.textn<1K2 likes31 downloads2y agoHugging Face21boczkakaroly /hungarian-riddles-benchmark Hungarian Riddles Benchmark Overview This dataset is a cultural and reasoning benchmark based on 100 metaphorical, trivia-style Hungarian riddles. The riddles are intentionally tricky and culturally grounded. They are designed to test answer correctness and reasoning quality, not only surface-level language fluency. Dataset structure Each row contains one riddle with reference material for evaluation. Fields ID – unique identifier topic – general… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/hungarian-riddles-benchmark.imagequestion-answeringn<1K0 likes31 downloads9mo agoHugging Face22TheFinAI /MED_SYN2_HUNGARIAN_traintext10K<n<100K0 likes30 downloads1y agoHugging Face23jlli /HungarianDocQA-OCRimagen<1K1 likes27 downloads2y agoHugging Face24jlli /Hungarian_CCPDF_SynQA_v2image10K<n<100K1 likes26 downloads2y agoHugging Face25leinadsened /hungarian-poems-with-instructionstext10K<n<100K4 likes22 downloads2y agoHugging Face26ShawnXiaoyuWang /med_syn1_hungariantext1K<n<10K0 likes22 downloads1y agoHugging Face27EtashGuha /HungarianDocQA_ITimagen<1K0 likes20 downloads2y agoHugging Face28jlli /Hungarian_CCPDF_SynQAimage10K<n<100K0 likes20 downloads2y agoHugging Face29Thomcles /YodaLingua-Hungariangated YodaLingua-Hungarian YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Hungarian portion of the multilingual YodaLingua collection. 🧾 Dataset Overview Property Value Total clips 80,740 audio–transcription pairs Total duration 206 hours Speakers 1,856 distinct speakers Audio format MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Hungarian.audiotext-to-speech10K<n<100K0 likes20 downloads5mo agoHugging Face30mbruton /hungarian_encryptedtabular10K<n<100K0 likes16 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.