Czech
Datasets
All datasets matching “Czech”czech_subjectivityczech_news_dataset_v2
Dataset Card for "czech_news_dataset_v2"
Dataset containing the news articles from major online news outlets collected from 2000-2022.
Follow-up paper https://arxiv.org/abs/2307.10666 (v1 of the dataset)
Changes from v1
Better contribution of novinky.cz in later stages
More articles, as a mistake in filtering was fixed.
Collection was done using CmonCrawl.
The dataset should be used for Research only purposes as I don't have rights for articles itself.
If you have any… See the full description on the dataset page: https://huggingface.co/datasets/hynky/czech_news_dataset_v2.czech_bank_qa
CzechBankQA
This is a list of SQL queries for a text-to-SQL task over the Czech Bank 1999 dataset.
czech_corpus_authorship_recognition
Czech Authorship Recognition Corpus (Kala)
Popis datasetu
Tento dataset byl vytvořen v rámci diplomové práce zaměřené na automatické rozpoznání autorství českých textů. Obsahuje české publicistické texty získané z veřejně dostupných online zdrojů a připravené pro experimenty v úlohách:
přiřazení autorství (authorship attribution)
ověřování autorství (authorship verification)
shlukování podle autorství (authorship clustering)
Zdrojová data
Do… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/czech_corpus_authorship_recognition.ipfs_czechia_laws_ir
Czechia legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_czechia_laws (revision f1dc1010ae1513cac19f73584e26df513e34cc94) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Czechia prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_czechia_laws_ir.czech-speech-combined
Czech Speech Combined Dataset
Quality-filtered Czech speech dataset for TTS/ASR training.
146,153 clips across ~1,700 speakers from 5 sources.
Sources
Source
Clips
Hours
Speakers
Origin
audiobooks
24,535
~33h
13
Czech audiobook narrations
audiobooks_new
28,131
~39h
8+
Czech audiobook narrations
yodas_czech
28,208
~37h
~2,300
YODAS YouTube speech (quality-filtered)
voxpopuli_czech
12,679
~33h
45
VoxPopuli parliament speech
commonvoice_czech
52… See the full description on the dataset page: https://huggingface.co/datasets/chosenek/czech-speech-combined.
