CoolFace
20 results

german

coral-nlp /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.tabulartext-generation10M<n<100M41 likes3.3k downloads8mo agoHugging FaceCUI03 /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.tabulartext-generation10M<n<100M1 likes3k downloads9mo agoHugging FacePleIAs /German-PD 🇩🇪 German Public Domain 🇩🇪 German-Public Domain or German-PD is a large collection aiming to aggregate all German monographies and periodicals in the public domain. As of March 2024, it is the biggest German open corpus. Dataset summary The collection contains 260,638 individual texts making up 37,650,706,611 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/German-PD.text100K<n<1M15 likes2.8k downloads2y agoHugging FaceCARD-Data /CARD-Germany-Batch1gated CARD – Germany 2 Days A comprehensive multi-modal driving dataset with stereo cameras, LiDAR, and depth annotations. Dataset Structure This dataset contains 28 sequences across 1 region(s): germany_2days: 28 sequences Data Format Each sequence contains: img/: Stereo camera images (cam_0, cam_1) raw/: Raw sensor data labels/: YOLO-format annotations export/: Trajectory and calibration data agg_depth/: Aggregated depth point clouds… See the full description on the dataset page: https://huggingface.co/datasets/CARD-Data/CARD-Germany-Batch1.imagedepth-estimation10K<n<100K5 likes1.4k downloads2mo agoHugging FaceAleph-Alpha /Aleph-Alpha-GermanWeb AlephAlphaGermanWeb Aleph-Alpha-GermanWeb is a new German-language dataset that combines heuristic and model-based filtering techniques with synthetic data generation to achieve SOTA performance in German-language benchmarks. The dataset draws from three sources: (1) Common Crawl web data, (2) FineWeb2, and (3) synthetically-generated data conditioned on actual, organic web data. In our accompanying paper (published at EACL 2026), we evaluated our dataset by training both a 1B… See the full description on the dataset page: https://huggingface.co/datasets/Aleph-Alpha/Aleph-Alpha-GermanWeb.text1B<n<10B24 likes1.2k downloads6mo agoHugging Facebritllm /TransWeb-Edu-Germantext10M<n<100M1 likes1.1k downloads2y agoHugging Face