datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
canada-dataset-1b
🇨🇦 Canada Web Text — 1B-token Sample 🍁
A 1-billion-token representative sample of a much larger cleaned Canadian web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐
The full 210B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/canada-dataset-1b.mingpao-canada-news
Ming Pao Canada Full Archive Corpus (明報加拿大完整語料庫)
Dataset Description
A comprehensive corpus of news articles from Ming Pao Canada (明報加拿大), covering both Toronto (多倫多) and Vancouver (溫哥華) editions from 2014 to 2026.
Dataset Summary
Source: Ming Pao Canada (加東版/多倫多 & 加西版/溫哥華)
Language: Traditional Chinese (繁體中文) / Cantonese (粵語)
Time Period: 14 July 2014 - 16 January 2026
Total Articles: 1,070,292
Categories: All sections (港聞, 加國新聞, 國際, 財經, 體育, 娛樂, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/mingpao-canada-news.
