datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
czech_subjectivityczech_news_dataset_v2
Dataset Card for "czech_news_dataset_v2"
Dataset containing the news articles from major online news outlets collected from 2000-2022.
Follow-up paper https://arxiv.org/abs/2307.10666 (v1 of the dataset)
Changes from v1
Better contribution of novinky.cz in later stages
More articles, as a mistake in filtering was fixed.
Collection was done using CmonCrawl.
The dataset should be used for Research only purposes as I don't have rights for articles itself.
If you have any… See the full description on the dataset page: https://huggingface.co/datasets/hynky/czech_news_dataset_v2.czech_bank_qa
CzechBankQA
This is a list of SQL queries for a text-to-SQL task over the Czech Bank 1999 dataset.
czech_corpus_authorship_recognition
Czech Authorship Recognition Corpus (Kala)
Popis datasetu
Tento dataset byl vytvořen v rámci diplomové práce zaměřené na automatické rozpoznání autorství českých textů. Obsahuje české publicistické texty získané z veřejně dostupných online zdrojů a připravené pro experimenty v úlohách:
přiřazení autorství (authorship attribution)
ověřování autorství (authorship verification)
shlukování podle autorství (authorship clustering)
Zdrojová data
Do… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/czech_corpus_authorship_recognition.ipfs_czechia_laws_ir
Czechia legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_czechia_laws (revision f1dc1010ae1513cac19f73584e26df513e34cc94) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Czechia prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_czechia_laws_ir.czech-speech-combined
Czech Speech Combined Dataset
Quality-filtered Czech speech dataset for TTS/ASR training.
146,153 clips across ~1,700 speakers from 5 sources.
Sources
Source
Clips
Hours
Speakers
Origin
audiobooks
24,535
~33h
13
Czech audiobook narrations
audiobooks_new
28,131
~39h
8+
Czech audiobook narrations
yodas_czech
28,208
~37h
~2,300
YODAS YouTube speech (quality-filtered)
voxpopuli_czech
12,679
~33h
45
VoxPopuli parliament speech
commonvoice_czech
52… See the full description on the dataset page: https://huggingface.co/datasets/chosenek/czech-speech-combined.Czech-PD
🇨🇿 Czech Public Domain 🇨🇿
Czech-Public Domain or Czech-PD is a large collection aiming to aggregate all Czech monographies and periodicals in the public domain. As of March 2024, it is the biggest Czech open corpus.
Dataset summary
The collection contains 1585 individual titles making up 259,435,959 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file has the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Czech-PD.CzechProductReviewSentimentClassification
CzechProductReviewSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
User reviews of products on Czech e-shop Mall.cz with 3 sentiment classes (positive, neutral, negative)
Task category
t2c
Domains
Reviews, Written
Reference
https://aclanthology.org/W13-1609/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CzechProductReviewSentimentClassification.cs_czech-named-entity-corpus_2.0
Dataset Card for Czech Named Entity Corpus 2.0
Dataset Description
The dataset contains Czech sentences and annotated named entities. Total number of sentences is around 9,000 and total number of entities is around 34,000. (Total means train + validation + test)
Dataset Features
Each sample contains:
text: source sentence
entities: list of selected entities. Each entity contains:
category_id: string identifier of the entity category
category_str:… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-named-entity-corpus_2.0.czech_encryptedczech-subjectivity-dataset
Dataset Card for Czech Subjectivity Dataset
Dataset Summary
Czech subjectivity dataset (Subj-CS) of 10k manually annotated subjective and objective sentences from movie reviews and descriptions. See the paper description https://arxiv.org/abs/2204.13915
Github
https://github.com/pauli31/czech-subjectivity-dataset
Supported Tasks and Leaderboards
Subjectivity Analysis
Languages
Czech
Data Instances
train/dev/test… See the full description on the dataset page: https://huggingface.co/datasets/pauli31/czech-subjectivity-dataset.cermat_czech_mc
Introduction
The Cermat Czech MultiChoice dataset was collected from assignments official CERMAT website. The dataset was collected from three tiers of assignments: 6 year, 9 year primary school test and final high school tests (so-called maturita).
The assignments were semi-manually extracted from official PDFs available at CERMAT's website.
Collection Date Range: years 2016-2023
Licensing and Credits
The majority of collection work was done by our student co-worker… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cermat_czech_mc.cermat_czech_tf
Introduction
The Cermat Czech True/False dataset was collected from assignments official CERMAT website. The dataset was collected from three tiers of assignments: 6 year, 9 year primary school test and final high school tests (so-called maturita).
The assignments were semi-manually extracted from official PDFs available at CERMAT's website.
Collection Date Range: years 2016-2023
Licensing and Credits
The majority of collection work was done by our student… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cermat_czech_tf.CzechVerseFormClassification
Czech Verse Form Classification
Single-label Czech poetic form classification for PoetryMTEB, using gold form labels from the Corpus of Czech Verse (CCV) subset packaged in multilingual-annotated-poems.
Only poems with a non-empty original form label are kept (most CCV poems are unlabeled for form). Rule-based auto-form labels are not used.
Dataset Card
Item
Description
Source
https://github.com/riazhsks/multilingual-annotated-poems →… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/CzechVerseFormClassification.czech_encrypted_HistCiph
Dataset Card for HistCiph — Czech
Dataset Description
Dataset Summary
The Czech subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical Czech plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext variants per… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/czech_encrypted_HistCiph.czechbench_belebele
Czech Belebele
This is an extraction of the Czech subset of the original Belebele dataset. The original 900 test samples were redistributed into train (20) and test (880) sets to facilitate generation of few-shot prompts with variable length.
Belebele is a parallel multilingual dataset focused on evaluating reading comprehension capabilities. The evaluation samples take form of single-choice questions which are tied to provided reference passages.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/davidadamczyk/czechbench_belebele.czech-legal-sft-dataset
Czech Legal SFT Dataset
Conversational dataset for fine-tuning Czech legal advisor language models.
Sources
roslein/Legal_advice_czech: 23,950 public legal Q&A from Czech advice websites
kucerj56/czech-sacd-legal-questions: 200 expert Q&A from Nejvyšší správní soud
roslein/Czech_legal_code: 3,080 Czech laws with document creation instructions
Format
Each sample contains a messages field — a list of dicts with role and content:
[
{"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/udold/czech-legal-sft-dataset.czechbench_agree
Czech grammar agreement dataset (AGREE)
This is an adapted and filtered test subset from the original Czech grammar agreement dataset
Citation
@PhdThesis{Baisa2016thesis,
AUTHOR = "Baisa, Vít",
TITLE = "Byte Level Language Models [online]",
YEAR = "2016 [cit. 2024-08-28]",
TYPE = "Disertační práce",
SCHOOL = "Masarykova univerzita, Fakulta informatiky, Brno",
NOTE = "SUPERVISOR : Karel Pala",
URL = "https://is.muni.cz/th/en6ay/",
}
czech-politician-statements
Czech Politician Statements Dataset
This dataset contains fact-checked statements from the Demagog project along with
their metadata and scraped evidence articles.
Github repository - dataset creation and fact checker pipeline scripts
Dataset Details
Language(s) (NLP): Czech
License: MIT
Dataset Sources
Source of statements: Demagog
Dataset Structure
Statements
Statements with their metadata (author, veracity label, ...).
Example… See the full description on the dataset page: https://huggingface.co/datasets/tomn24/czech-politician-statements.cognia-czech-dialogues
Cognia Czech Dialogues
Cognia Czech Dialogues is a pilot collection of synthetic multi-turn conversations in Czech. It is intended for experiments with conversational language models, supervised fine-tuning, evaluation, and dataset-curation workflows.
The current release contains 4,000 dialogues across 40 topic areas. The data was generated synthetically and has not yet been fully reviewed by humans.
Dataset contents
Language: Czech (cs-CZ)
Dialogues: 4,000… See the full description on the dataset page: https://huggingface.co/datasets/havelm3/cognia-czech-dialogues.CzechTopicwiseSummarizationipfs_czechia_laws
Czech National Collection of Laws (e-Sbírka)
Research snapshot of current Czech national legislation collected from
official e-Sbírka NKOD JSON dumps
(právní-akt-znění-poslední current wording).
Not legal advice. The official Collection of Laws (Sbírka zákonů) and
e-Sbírka authentic texts prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-02
Coverage
complete (complete=true)
Source
e-Sbírka (Ministerstvo vnitra) NKOD JSON dumps… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_czechia_laws.czech-speech-datasetpdt_anaphora_czech
Dataset Card for pdt_anaphora_czech
This dataset is used for my thesis to fine-tune language models on Czech unstructured text for anaphora resolution.
Dataset Sources
The dataset is based on data from the Prague Dependency Treebank, specifically the PDTC 1.0 (https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3185)
How to cite
Stano P. and Horák A. Evaluating Prompt-Based and Fine-Tuned Approaches to Czech Anaphora Resolution.
International Conference… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/pdt_anaphora_czech.czechbench_subjectivity
Czech Subjectivity
This is a restructured test set from the original Czech Subjectivity Dataset.
The dataset is dedicated to the task of subjectivity classification. For each reference passage, it needs to be determined whether it conveys subjective opinions or objective facts.
Citation
@article{pib2022czech,
title={Czech Dataset for Cross-lingual Subjectivity Classification},
author={Pavel Přibáň and Josef Steinberger},
year={2022},
eprint={2204.13915}… See the full description on the dataset page: https://huggingface.co/datasets/davidadamczyk/czechbench_subjectivity.czech_newsczech_news_simple-cs
Simplified Czech News dataset
This is a simplified and subsampled test subset from the original czech_news_dataset_v2.
Only 5 basic news categories are considered:
Zahraniční (Foreign)
Domácí (Local)
Sport (Sport)
Kultura (Culture)
Ekonomika (Economy)
The test set includes 200 examples per category, 1000 examples in total.
Apart from the category label, each example also contains the article's headline, brief summary, full textual content, optional keywords, original category… See the full description on the dataset page: https://huggingface.co/datasets/CIIRC-NLP/czech_news_simple-cs.CzechSubjectivityClassification
CzechSubjectivityClassification
An MTEB dataset
Massive Text Embedding Benchmark
An Czech dataset for subjectivity classification.
Task category
t2c
Domains
Reviews, Written
Reference
https://arxiv.org/abs/2009.08712
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CzechSubjectivityClassification"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CzechSubjectivityClassification.czechcasting
Academic Dataset Scraper
A robust, multi-threaded Python web scraper designed to build structured academic datasets from web listings. It automates the collection of model profiles, high-resolution images, and automated content labels (NSFW classification) using computer vision.
Output
The script generates two primary outputs:
1. Directory Structure (dataset_images/)
Images are saved in folders named by their unique ID extracted from the URL slug:… See the full description on the dataset page: https://huggingface.co/datasets/mic4565g/czechcasting.cs_czech-court-decisions-ner
Dataset Card for Czech Court Decisions NER
Dataset Description
Czech Court Decisions NER is a dataset of 300 court decisions published by The Supreme Court of the Czech Republic and the Constitutional Court of the Czech Republic.
In the documents, 4 types of named entities are selected.
Dataset Features
Each sample contains:
filename: file name in the original dataset
text: court decision document in plain text
entities: list of selected entities. Each entity… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-court-decisions-ner.
