CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FrancophonIA /Translations_Hungarian_public_websites [!NOTE] Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/18982 Description A webcrawl of 14 different websites covering parallel corpora of Hungarian with Polish, Czech, Swedish, Finnish, French, German, Italian, English and Slovenian Citation Translations of Hungarian from public websites (2022). Version 1.0. [Dataset (Text corpus)]. Source: European Language Grid. https://live.european-language-grid.eu/catalogue/corpus/18982 translation0 likes881 downloads1y agoHugging Face02jimregan /hungarian-youtube-speech0 likes478 downloads4y agoHugging Face03keirp /hungarian_national_hs_finals_exam Testing Language Models on a Held-Out High School National Finals Exam When xAI recently released Grok-1, they evaluated it on the 2023 Hungarian national high school finals in mathematics, which was published after the training data cutoff for all the models in their evaluation. While MATH and GSM8k are the standard benchmarks for evaluating the mathematical abilities of large language models, there are risks that modern models overfit to these datasets, either from training… See the full description on the dataset page: https://huggingface.co/datasets/keirp/hungarian_national_hs_finals_exam.textn<1K27 likes474 downloads3y agoHugging Face04nhiremath /HungarianDocQA_IT_SynQA_ocr_v3image100K<n<1M0 likes346 downloads2y agoHugging Face05jinaai /hungarian_doc_qa_beirThis is a copy of https://huggingface.co/datasets/jinaai/hungarian_doc_qa reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/hungarian_doc_qa_beir.imagen<1K0 likes257 downloads1y agoHugging Face06Bazsalanszky /hungarian-llm-testing Hungarian llm testing This is a really simple data-set to test fine-tuning a language model on Hungarian text. textn<1K5 likes254 downloads3y agoHugging Face07KTH /hungarian-single-speaker-tts Dataset Card for CSS10 Hungarian: Single Speaker Speech Dataset Dataset Summary The corpus consists of a single speaker, with 4515 segments extracted from a single LibriVox audiobook. Supported Tasks and Leaderboards [Needs More Information] Languages The audio is in Hungarian. Dataset Structure [Needs More Information] Data Instances [Needs More Information] Data Fields [Needs More Information] Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/KTH/hungarian-single-speaker-tts.audiotext-to-speech1K<n<10K14 likes173 downloads4y agoHugging Face08datadriven-company /TTS-Hungarian TTS-Hungarian A large-scale, high-quality Hungarian speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from MEK (Magyar Elektronikus Könyvtár) — Hungarian audiobooks. Dataset Statistics Metric Value Total samples 253,116 Total duration 702 hours Unique speakers 100 Average duration 10.0 seconds Average DNSMOS 3.68 Features Field Type Description __key__ string Unique sample… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Hungarian.audiotext-to-speech100K<n<1M1 likes171 downloads7mo agoHugging Face09AlhitawiMohammed22 /HungarianNamesThis dataset was collected from Wikipedia : https://hu.wikipedia.org/wiki/Magyarorsz%C3%A1gon_anyak%C3%B6nyvezhet%C5%91_ut%C3%B3nevek_list%C3%A1ja imagetext-generationn<1K0 likes168 downloads3y agoHugging Face10hongfenglu /HungarianDocQA_IT_IfEvalQA_v2image10K<n<100K0 likes147 downloads2y agoHugging Face11hongfenglu /HungarianDocQA_IT_IFEvalQAimage10K<n<100K0 likes129 downloads2y agoHugging Face12gedeonmate /Hungarian-Dialogues-text LLM-Generated Hungarian Conversations This dataset contains structured Hungarian conversations generated with multiple large language model families for the study “Efficient ASR Training with Conversations that Never Happened.” Paper: arXiv link Each model is provided as a separate Parquet file. The dataset contains the generated textual conversations and associated scenario and participant metadata; it does not contain synthesized audio. Dataset structure Each… See the full description on the dataset page: https://huggingface.co/datasets/gedeonmate/Hungarian-Dialogues-text.tabular100K<n<1M0 likes105 downloads2mo agoHugging Face13nhiremath /HungarianDocQA_IT_SynQA_ocr_v3_cleanedimage10K<n<100K0 likes83 downloads2y agoHugging Face14hakatiki /hungarian-cc-corpustext10M<n<100M4 likes78 downloads3y agoHugging Face15matekadlicsko /hungarian-news-translationstexttranslation10K<n<100K1 likes71 downloads3y agoHugging Face16nhiremath /HungarianDocQA_IT_SyntheticQAimage100K<n<1M0 likes71 downloads2y agoHugging Face17saillab /alpaca-hungarian-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-hungarian-cleaned.text10K<n<100K6 likes65 downloads2y agoHugging Face18saillab /alpaca_hungarian_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_hungarian_taco.text10K<n<100K2 likes62 downloads2y agoHugging Face19BMark52 /Hungarian-Train-Data-DestilledGeminitext1K<n<10K0 likes62 downloads14d agoHugging Face20shunyalabs /hungarian-speech-datasetaudio1K<n<10K0 likes48 downloads1y agoHugging Face21RabidUmarell /hungarian-toxic-comments Hungarian Toxic Comments The first openly available Hungarian dataset for toxic comment classification, introduced in: Hatvani, P., & Yang, Z. Gy. (2025). Automated detection of toxic comments in Hungarian. Annales Mathematicae et Informaticae, 61, 108-117. DOI: 10.33039/ami.2025.10.007 Dataset Description This dataset contains 654 manually annotated Hungarian-language comments collected from social media and political news forums. Each comment is annotated across five… See the full description on the dataset page: https://huggingface.co/datasets/RabidUmarell/hungarian-toxic-comments.documenttext-classificationn<1K1 likes40 downloads7mo agoHugging Face22boczkakaroly /hungarian-riddles-benchmark Hungarian Riddles Benchmark Overview This dataset is a cultural and reasoning benchmark based on 100 metaphorical, trivia-style Hungarian riddles. The riddles are intentionally tricky and culturally grounded. They are designed to test answer correctness and reasoning quality, not only surface-level language fluency. Dataset structure Each row contains one riddle with reference material for evaluation. Fields ID – unique identifier topic – general… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/hungarian-riddles-benchmark.imagequestion-answeringn<1K0 likes36 downloads9mo agoHugging Face23ryuseiken /FreddieMercury_HungarianRhapsodyLive860 likes32 downloads3y agoHugging Face24data-is-better-together /MPEP_HUNGARIAN Dataset Card for MPEP_HUNGARIAN This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Dataset Summary This dataset contains: A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/MPEP_HUNGARIAN.textn<1K2 likes31 downloads2y agoHugging Face25Speech-data /Hungarian-Speech-Dataset 🎧 Hungarian Speech Dataset The Hungarian Speech Dataset is a high-quality speech audio dataset designed to support advanced AI systems that depend on diverse audio data and reliable voice data for multilingual model training. It comprises 169 hours of recordings across 743 files, provided in MP3 and WAV formats, with a total size of 134 MB. This structured audio dataset ensures balanced speaker representation, featuring 46% female and 54% male speakers, and an age distribution… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hungarian-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes31 downloads6mo agoHugging Face26TheFinAI /MED_SYN2_HUNGARIAN_traintext10K<n<100K0 likes30 downloads1y agoHugging Face27jlli /HungarianDocQA-OCRimagen<1K1 likes26 downloads2y agoHugging Face28ShawnXiaoyuWang /med_syn1_hungariantext1K<n<10K0 likes26 downloads1y agoHugging Face29FrancophonIA /NTEU_French-Hungarian [!NOTE] Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19581 Description This is a compilation of parallel corpora resources used in building of Machine Translation engines in NTEU project (Action number: 2018-EU-IA-0051). Data in these resources are compiled in two TMX files, two tiers grouped by data source reliablity. Tier A -- danta originating from human edited sources, translation memories and alike. Tier B -- danta originating created by automatic… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/NTEU_French-Hungarian.translation0 likes24 downloads1y agoHugging Face30jlli /Hungarian_CCPDF_SynQA_v2image10K<n<100K1 likes24 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.