hungarian
Translations_Hungarian_public_websites
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/18982
Description
A webcrawl of 14 different websites covering parallel corpora of Hungarian with Polish, Czech, Swedish, Finnish, French, German, Italian, English and Slovenian
Citation
Translations of Hungarian from public websites (2022). Version 1.0. [Dataset (Text corpus)]. Source: European Language Grid. https://live.european-language-grid.eu/catalogue/corpus/18982
hungarian-youtube-speechhungarian_national_hs_finals_exam
Testing Language Models on a Held-Out High School National Finals Exam
When xAI recently released Grok-1, they evaluated it on the 2023 Hungarian national high school finals in mathematics, which was published after the training data cutoff for all the models in their evaluation. While MATH and GSM8k are the standard benchmarks for evaluating the mathematical abilities of large language models, there are risks that modern models overfit to these datasets, either from training… See the full description on the dataset page: https://huggingface.co/datasets/keirp/hungarian_national_hs_finals_exam.HungarianDocQA_IT_SynQA_ocr_v3hungarian_doc_qa_beirThis is a copy of https://huggingface.co/datasets/jinaai/hungarian_doc_qa reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/hungarian_doc_qa_beir.hungarian-llm-testing
Hungarian llm testing
This is a really simple data-set to test fine-tuning a language model on Hungarian text.
