CoolFace
2 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tmquan /phapdien-moj-gov-vn Bộ Pháp Điển Việt Nam — phapdien.moj.gov.vn 🇻🇳 Tóm tắt. Bộ ngữ liệu cấp Điều của Bộ Pháp Điển Việt Nam — bộ pháp điển chính thức do Bộ Tư pháp công bố. Mỗi dòng documents là một Điều kèm toàn văn đã chuẩn hoá, chương sở thuộc, đề mục và chủ đề. Kèm theo là vector nhúng ngữ nghĩa 4096-D (embeddings), toạ độ giảm chiều trong không gian chung ViLA (reduces), và từ điển ontology song ngữ Việt–Anh (chủ đề · đề mục · thuật ngữ). 🇬🇧 One-line. Article-level corpus of the Bộ Pháp… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/phapdien-moj-gov-vn.imagetext-classification100K<n<1M11 likes1.2k downloads7d agoHugging Face02Goader /ukrainian-news-2026 Ukrainian News 2026 Ukrainian-language news articles from 20 national outlets, published between 1 January and 28 August 2026. Extracted body text plus metadata. Two configs. deduplicated is the default — near-duplicates removed, which is what you want when mixing this with an already-deduplicated pretraining corpus. raw is the original release, unchanged. deduplicated (default) raw train-mixin Documents 419,204 429,427 386,477 Characters 0.97B 1.01B 0.88B Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Goader/ukrainian-news-2026.imagetext-generation1M<n<10M0 likes93 downloads25d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.