datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cs_czech-named-entity-corpus_2.0
Dataset Card for Czech Named Entity Corpus 2.0
Dataset Description
The dataset contains Czech sentences and annotated named entities. Total number of sentences is around 9,000 and total number of entities is around 34,000. (Total means train + validation + test)
Dataset Features
Each sample contains:
text: source sentence
entities: list of selected entities. Each entity contains:
category_id: string identifier of the entity category
category_str:… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-named-entity-corpus_2.0.cognia-czech-dialogues
Cognia Czech Dialogues
Cognia Czech Dialogues is a pilot collection of synthetic multi-turn conversations in Czech. It is intended for experiments with conversational language models, supervised fine-tuning, evaluation, and dataset-curation workflows.
The current release contains 4,000 dialogues across 40 topic areas. The data was generated synthetically and has not yet been fully reviewed by humans.
Dataset contents
Language: Czech (cs-CZ)
Dialogues: 4,000… See the full description on the dataset page: https://huggingface.co/datasets/havelm3/cognia-czech-dialogues.cermat_czech_mc
Introduction
The Cermat Czech MultiChoice dataset was collected from assignments official CERMAT website. The dataset was collected from three tiers of assignments: 6 year, 9 year primary school test and final high school tests (so-called maturita).
The assignments were semi-manually extracted from official PDFs available at CERMAT's website.
Collection Date Range: years 2016-2023
Licensing and Credits
The majority of collection work was done by our student co-worker… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cermat_czech_mc.cermat_czech_tf
Introduction
The Cermat Czech True/False dataset was collected from assignments official CERMAT website. The dataset was collected from three tiers of assignments: 6 year, 9 year primary school test and final high school tests (so-called maturita).
The assignments were semi-manually extracted from official PDFs available at CERMAT's website.
Collection Date Range: years 2016-2023
Licensing and Credits
The majority of collection work was done by our student… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cermat_czech_tf.czech-politician-statements
Czech Politician Statements Dataset
This dataset contains fact-checked statements from the Demagog project along with
their metadata and scraped evidence articles.
Github repository - dataset creation and fact checker pipeline scripts
Dataset Details
Language(s) (NLP): Czech
License: MIT
Dataset Sources
Source of statements: Demagog
Dataset Structure
Statements
Statements with their metadata (author, veracity label, ...).
Example… See the full description on the dataset page: https://huggingface.co/datasets/tomn24/czech-politician-statements.CzechTopicwiseSummarizationcs_czech-court-decisions-ner
Dataset Card for Czech Court Decisions NER
Dataset Description
Czech Court Decisions NER is a dataset of 300 court decisions published by The Supreme Court of the Czech Republic and the Constitutional Court of the Czech Republic.
In the documents, 4 types of named entities are selected.
Dataset Features
Each sample contains:
filename: file name in the original dataset
text: court decision document in plain text
entities: list of selected entities. Each entity… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-court-decisions-ner.cermat_czech_open
Introduction
The Cermat Czech Open dataset was collected from assignments official CERMAT website. The dataset was collected from three tiers of assignments: 6 year, 9 year primary school test and final high school tests (so-called maturita).
The assignments were semi-manually extracted from official PDFs available at CERMAT's website.
Collection Date Range: years 2016-2023
Licensing and Credits
The majority of collection work was done by our student co-worker… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cermat_czech_open.CzechSingleDocumentSummarizationexam-czech-bioPart of INCLUDE.
see research: https://arxiv.org/abs/2411.19799
full dataset: https://huggingface.co/datasets/CohereForAI/include-base-44
CzechCitedSummarizationczech-lit-examPart of INCLUDE
see research: https://arxiv.org/abs/2411.19799
full dataset: https://huggingface.co/datasets/CohereForAI/include-base-44
czech-castingCzechVOC_QAczech-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Czech Republic
The Synthetic Czech Republic Passports Dataset assembles more than 1,000 AI-generated passport images crafted for training OCR and computer vision models on identity documents. Each record is fully… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/czech-passports.CzechNews
📰 FactDeMice-CzechNews
This dataset contains trustworthy news articles collected from online news portals. The data is stored in JSON Lines (.jsonl) format, where each line represents a single article.
License Information
We do not own any of the public texts from which these text data has been extracted.
We license the actual packaging of these text data under the Creative Commons CC0 license ("no rights reserved").
📂 Data structure
The dataset schema in… See the full description on the dataset page: https://huggingface.co/datasets/FactDeMice/CzechNews.czech-castingmarathi-czech-sentences
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
marathi_czech_sentences
This dataset contains short sentences and questions primarily in Marathi and Czech, covering various conversational contexts. The samples include inquiries about objects, actions, and origins, as well as exclamations and statements. It appears to be a multilingual collection focused on everyday dialogue structures.
Dataset size
There are 3… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/marathi-czech-sentences.tldr_czech_v1.jsonlwarhammer_eng_to_czechczech-sacd-legal-questions
📑 Overview
This repository contains 200 question-answer pairs automatically generated with Gemini 2.0 from the decisions of the Czech Sumpreme Administrative Court.
The work was performed in spring 2025 as part of my master’s diploma thesis at the Faculty of Information Technology, Czech Technical University in Prague (FIT CTU).
🏛️ Source
Official judgments scraped from https://sbirka.nssoud.cz (March 2025 snapshot).
Youtube_Czechid
