CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /czech_subjectivitytext1K<n<10K0 likes577 downloads1y agoHugging Face02hynky /czech_news_dataset_v2 Dataset Card for "czech_news_dataset_v2" Dataset containing the news articles from major online news outlets collected from 2000-2022. Follow-up paper https://arxiv.org/abs/2307.10666 (v1 of the dataset) Changes from v1 Better contribution of novinky.cz in later stages More articles, as a mistake in filtering was fixed. Collection was done using CmonCrawl. The dataset should be used for Research only purposes as I don't have rights for articles itself. If you have any… See the full description on the dataset page: https://huggingface.co/datasets/hynky/czech_news_dataset_v2.texttext-classification1M<n<10M5 likes435 downloads2y agoHugging Face03yifanmai /czech_bank_qa CzechBankQA This is a list of SQL queries for a text-to-SQL task over the Czech Bank 1999 dataset. tabularn<1K0 likes407 downloads2y agoHugging Face04MU-NLPC /czech_corpus_authorship_recognition Czech Authorship Recognition Corpus (Kala) Popis datasetu Tento dataset byl vytvořen v rámci diplomové práce zaměřené na automatické rozpoznání autorství českých textů. Obsahuje české publicistické texty získané z veřejně dostupných online zdrojů a připravené pro experimenty v úlohách: přiřazení autorství (authorship attribution) ověřování autorství (authorship verification) shlukování podle autorství (authorship clustering) Zdrojová data Do… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/czech_corpus_authorship_recognition.texttext-classification0 likes364 downloads3mo agoHugging Face05justicedao /ipfs_czechia_laws_ir Czechia legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_czechia_laws (revision f1dc1010ae1513cac19f73584e26df513e34cc94) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Czechia prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_czechia_laws_ir.tabulartext-retrieval100K<n<1M0 likes252 downloads14h agoHugging Face06chosenek /czech-speech-combined Czech Speech Combined Dataset Quality-filtered Czech speech dataset for TTS/ASR training. 146,153 clips across ~1,700 speakers from 5 sources. Sources Source Clips Hours Speakers Origin audiobooks 24,535 ~33h 13 Czech audiobook narrations audiobooks_new 28,131 ~39h 8+ Czech audiobook narrations yodas_czech 28,208 ~37h ~2,300 YODAS YouTube speech (quality-filtered) voxpopuli_czech 12,679 ~33h 45 VoxPopuli parliament speech commonvoice_czech 52… See the full description on the dataset page: https://huggingface.co/datasets/chosenek/czech-speech-combined.text100K<n<1M2 likes230 downloads4mo agoHugging Face07PleIAs /Czech-PD 🇨🇿 Czech Public Domain 🇨🇿 Czech-Public Domain or Czech-PD is a large collection aiming to aggregate all Czech monographies and periodicals in the public domain. As of March 2024, it is the biggest Czech open corpus. Dataset summary The collection contains 1585 individual titles making up 259,435,959 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file has the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Czech-PD.6 likes189 downloads2y agoHugging Face08mteb /CzechProductReviewSentimentClassification CzechProductReviewSentimentClassification An MTEB dataset Massive Text Embedding Benchmark User reviews of products on Czech e-shop Mall.cz with 3 sentiment classes (positive, neutral, negative) Task category t2c Domains Reviews, Written Reference https://aclanthology.org/W13-1609/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CzechProductReviewSentimentClassification.texttext-classification10K<n<100K0 likes170 downloads1y agoHugging Face09fewshot-goes-multilingual /cs_czech-named-entity-corpus_2.0 Dataset Card for Czech Named Entity Corpus 2.0 Dataset Description The dataset contains Czech sentences and annotated named entities. Total number of sentences is around 9,000 and total number of entities is around 34,000. (Total means train + validation + test) Dataset Features Each sample contains: text: source sentence entities: list of selected entities. Each entity contains: category_id: string identifier of the entity category category_str:… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-named-entity-corpus_2.0.texttoken-classification1K<n<10K3 likes169 downloads4y agoHugging Face10mbruton /czech_encryptedtabular100K<n<1M0 likes161 downloads8mo agoHugging Face11pauli31 /czech-subjectivity-dataset Dataset Card for Czech Subjectivity Dataset Dataset Summary Czech subjectivity dataset (Subj-CS) of 10k manually annotated subjective and objective sentences from movie reviews and descriptions. See the paper description https://arxiv.org/abs/2204.13915 Github https://github.com/pauli31/czech-subjectivity-dataset Supported Tasks and Leaderboards Subjectivity Analysis Languages Czech Data Instances train/dev/test… See the full description on the dataset page: https://huggingface.co/datasets/pauli31/czech-subjectivity-dataset.texttext-classification1K<n<10K3 likes115 downloads3y agoHugging Face12CZLC /cermat_czech_mc Introduction The Cermat Czech MultiChoice dataset was collected from assignments official CERMAT website. The dataset was collected from three tiers of assignments: 6 year, 9 year primary school test and final high school tests (so-called maturita). The assignments were semi-manually extracted from official PDFs available at CERMAT's website. Collection Date Range: years 2016-2023 Licensing and Credits The majority of collection work was done by our student co-worker… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cermat_czech_mc.textn<1K1 likes102 downloads2y agoHugging Face13CZLC /cermat_czech_tf Introduction The Cermat Czech True/False dataset was collected from assignments official CERMAT website. The dataset was collected from three tiers of assignments: 6 year, 9 year primary school test and final high school tests (so-called maturita). The assignments were semi-manually extracted from official PDFs available at CERMAT's website. Collection Date Range: years 2016-2023 Licensing and Credits The majority of collection work was done by our student… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cermat_czech_tf.textn<1K0 likes97 downloads2y agoHugging Face14PoetryMTEB /CzechVerseFormClassification Czech Verse Form Classification Single-label Czech poetic form classification for PoetryMTEB, using gold form labels from the Corpus of Czech Verse (CCV) subset packaged in multilingual-annotated-poems. Only poems with a non-empty original form label are kept (most CCV poems are unlabeled for form). Rule-based auto-form labels are not used. Dataset Card Item Description Source https://github.com/riazhsks/multilingual-annotated-poems →… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/CzechVerseFormClassification.texttext-classification1K<n<10K0 likes97 downloads1mo agoHugging Face15mbruton /czech_encrypted_HistCiph Dataset Card for HistCiph — Czech Dataset Description Dataset Summary The Czech subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical Czech plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext variants per… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/czech_encrypted_HistCiph.tabular100K<n<1M0 likes74 downloads5mo agoHugging Face16davidadamczyk /czechbench_belebele Czech Belebele This is an extraction of the Czech subset of the original Belebele dataset. The original 900 test samples were redistributed into train (20) and test (880) sets to facilitate generation of few-shot prompts with variable length. Belebele is a parallel multilingual dataset focused on evaluating reading comprehension capabilities. The evaluation samples take form of single-choice questions which are tied to provided reference passages. Citation… See the full description on the dataset page: https://huggingface.co/datasets/davidadamczyk/czechbench_belebele.texttext-classificationn<1K0 likes68 downloads2y agoHugging Face17udold /czech-legal-sft-dataset Czech Legal SFT Dataset Conversational dataset for fine-tuning Czech legal advisor language models. Sources roslein/Legal_advice_czech: 23,950 public legal Q&A from Czech advice websites kucerj56/czech-sacd-legal-questions: 200 expert Q&A from Nejvyšší správní soud roslein/Czech_legal_code: 3,080 Czech laws with document creation instructions Format Each sample contains a messages field — a list of dicts with role and content: [ {"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/udold/czech-legal-sft-dataset.text10K<n<100K0 likes66 downloads4mo agoHugging Face18davidadamczyk /czechbench_agree Czech grammar agreement dataset (AGREE) This is an adapted and filtered test subset from the original Czech grammar agreement dataset Citation @PhdThesis{Baisa2016thesis, AUTHOR = "Baisa, Vít", TITLE = "Byte Level Language Models [online]", YEAR = "2016 [cit. 2024-08-28]", TYPE = "Disertační práce", SCHOOL = "Masarykova univerzita, Fakulta informatiky, Brno", NOTE = "SUPERVISOR : Karel Pala", URL = "https://is.muni.cz/th/en6ay/", } texttext-classificationn<1K0 likes64 downloads2y agoHugging Face19tomn24 /czech-politician-statements Czech Politician Statements Dataset This dataset contains fact-checked statements from the Demagog project along with their metadata and scraped evidence articles. Github repository - dataset creation and fact checker pipeline scripts Dataset Details Language(s) (NLP): Czech License: MIT Dataset Sources Source of statements: Demagog Dataset Structure Statements Statements with their metadata (author, veracity label, ...). Example… See the full description on the dataset page: https://huggingface.co/datasets/tomn24/czech-politician-statements.texttext-classification10K<n<100K2 likes57 downloads1y agoHugging Face20havelm3 /cognia-czech-dialogues Cognia Czech Dialogues Cognia Czech Dialogues is a pilot collection of synthetic multi-turn conversations in Czech. It is intended for experiments with conversational language models, supervised fine-tuning, evaluation, and dataset-curation workflows. The current release contains 4,000 dialogues across 40 topic areas. The data was generated synthetically and has not yet been fully reviewed by humans. Dataset contents Language: Czech (cs-CZ) Dialogues: 4,000… See the full description on the dataset page: https://huggingface.co/datasets/havelm3/cognia-czech-dialogues.texttext-generation1K<n<10K0 likes57 downloads6d agoHugging Face21BUT-FIT /CzechTopicwiseSummarizationtext10K<n<100K0 likes53 downloads2y agoHugging Face22endomorphosis /ipfs_czechia_laws Czech National Collection of Laws (e-Sbírka) Research snapshot of current Czech national legislation collected from official e-Sbírka NKOD JSON dumps (právní-akt-znění-poslední current wording). Not legal advice. The official Collection of Laws (Sbírka zákonů) and e-Sbírka authentic texts prevail over this corpus. Snapshot Field Value Snapshot date 2026-09-02 Coverage complete (complete=true) Source e-Sbírka (Ministerstvo vnitra) NKOD JSON dumps… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_czechia_laws.texttext-retrieval10K<n<100K0 likes53 downloads20d agoHugging Face23shunyalabs /czech-speech-datasetaudio1K<n<10K0 likes47 downloads1y agoHugging Face24MU-NLPC /pdt_anaphora_czech Dataset Card for pdt_anaphora_czech This dataset is used for my thesis to fine-tune language models on Czech unstructured text for anaphora resolution. Dataset Sources The dataset is based on data from the Prague Dependency Treebank, specifically the PDTC 1.0 (https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3185) How to cite Stano P. and Horák A. Evaluating Prompt-Based and Fine-Tuned Approaches to Czech Anaphora Resolution. International Conference… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/pdt_anaphora_czech.text10K<n<100K0 likes46 downloads1y agoHugging Face25davidadamczyk /czechbench_subjectivity Czech Subjectivity This is a restructured test set from the original Czech Subjectivity Dataset. The dataset is dedicated to the task of subjectivity classification. For each reference passage, it needs to be determined whether it conveys subjective opinions or objective facts. Citation @article{pib2022czech, title={Czech Dataset for Cross-lingual Subjectivity Classification}, author={Pavel Přibáň and Josef Steinberger}, year={2022}, eprint={2204.13915}… See the full description on the dataset page: https://huggingface.co/datasets/davidadamczyk/czechbench_subjectivity.texttext-classification1K<n<10K0 likes44 downloads2y agoHugging Face26simecek /czech_newstabular1K<n<10K1 likes42 downloads2y agoHugging Face27CIIRC-NLP /czech_news_simple-cs Simplified Czech News dataset This is a simplified and subsampled test subset from the original czech_news_dataset_v2. Only 5 basic news categories are considered: Zahraniční (Foreign) Domácí (Local) Sport (Sport) Kultura (Culture) Ekonomika (Economy) The test set includes 200 examples per category, 1000 examples in total. Apart from the category label, each example also contains the article's headline, brief summary, full textual content, optional keywords, original category… See the full description on the dataset page: https://huggingface.co/datasets/CIIRC-NLP/czech_news_simple-cs.tabular1K<n<10K0 likes40 downloads2y agoHugging Face28mteb /CzechSubjectivityClassification CzechSubjectivityClassification An MTEB dataset Massive Text Embedding Benchmark An Czech dataset for subjectivity classification. Task category t2c Domains Reviews, Written Reference https://arxiv.org/abs/2009.08712 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CzechSubjectivityClassification"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CzechSubjectivityClassification.texttext-classification1K<n<10K0 likes39 downloads1y agoHugging Face29mic4565g /czechcasting Academic Dataset Scraper A robust, multi-threaded Python web scraper designed to build structured academic datasets from web listings. It automates the collection of model profiles, high-resolution images, and automated content labels (NSFW classification) using computer vision. Output The script generates two primary outputs: 1. Directory Structure (dataset_images/) Images are saved in folders named by their unique ID extracted from the URL slug:… See the full description on the dataset page: https://huggingface.co/datasets/mic4565g/czechcasting.image-to-image10K<n<100K0 likes39 downloads22d agoHugging Face30fewshot-goes-multilingual /cs_czech-court-decisions-ner Dataset Card for Czech Court Decisions NER Dataset Description Czech Court Decisions NER is a dataset of 300 court decisions published by The Supreme Court of the Czech Republic and the Constitutional Court of the Czech Republic. In the documents, 4 types of named entities are selected. Dataset Features Each sample contains: filename: file name in the original dataset text: court decision document in plain text entities: list of selected entities. Each entity… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-court-decisions-ner.texttoken-classificationn<1K2 likes38 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.