CoolFace
20 results

german

coral-nlp /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.tabulartext-generation10M<n<100M41 likes3.4k downloads8mo agoHugging FacePleIAs /German-PD 🇩🇪 German Public Domain 🇩🇪 German-Public Domain or German-PD is a large collection aiming to aggregate all German monographies and periodicals in the public domain. As of March 2024, it is the biggest German open corpus. Dataset summary The collection contains 260,638 individual texts making up 37,650,706,611 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/German-PD.text100K<n<1M15 likes2.9k downloads2y agoHugging FaceCUI03 /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.tabulartext-generation10M<n<100M1 likes2.6k downloads9mo agoHugging FaceCARD-Data /CARD-Germany-Batch1gated CARD – Germany 2 Days A comprehensive multi-modal driving dataset with stereo cameras, LiDAR, and depth annotations. Dataset Structure This dataset contains 28 sequences across 1 region(s): germany_2days: 28 sequences Data Format Each sequence contains: img/: Stereo camera images (cam_0, cam_1) raw/: Raw sensor data labels/: YOLO-format annotations export/: Trajectory and calibration data agg_depth/: Aggregated depth point clouds… See the full description on the dataset page: https://huggingface.co/datasets/CARD-Data/CARD-Germany-Batch1.imagedepth-estimation10K<n<100K5 likes1.2k downloads2mo agoHugging Facebritllm /TransWeb-Edu-Germantext10M<n<100M1 likes1.2k downloads2y agoHugging FaceGermanEval /germeval_14The GermEval 2014 NER Shared Task builds on a new dataset with German Named Entity annotation with the following properties: - The data was sampled from German Wikipedia and News Corpora as a collection of citations. - The dataset covers over 31,000 sentences corresponding to over 590,000 tokens. - The NER annotation uses the NoSta-D guidelines, which extend the Tübingen Treebank guidelines, using four main NER categories with sub-structure, and annotating embeddings among NEs such as [ORG FC Kickers [LOC Darmstadt]].token-classification100K<n<1M4 likes1.1k downloads3y agoHugging Face