datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CanadaFireSat
Dataset Card for CanadaFireSat 🔥🛰️
In this benchmark, we investigate the potential of deep learning with multiple modalities for high-resolution wildfire forecasting. Leveraging different data settings across two types of model architectures: CNN-based and ViT-based.
📝 Published paper from ISPRS (ArXiv Version)
💿 Dataset repository on GitHub
🤖 Model repository on GitHub & Weights on Hugging Face
🟰 Another "Raw" version of the data with NPY files organized in different… See the full description on the dataset page: https://huggingface.co/datasets/EPFL-ECEO/CanadaFireSat.CanadaWildFireDaily-v1🔥🔥 CanadaWildfireDaily: A Large-Scale Dataset for Daily Wildfire Spread in Canada 🔥🔥
Folder Structure
This section provides the details needed to understand and use the released dataset. We release both raw data and training/val/test ready samples.
CanadaWildFireDaily Train/Validation/Test Samples (data_samples/)
The final training samples are generated through a two-step process.
First, the CSV files and metadata mappers are used to assign fire IDs to the training… See the full description on the dataset page: https://huggingface.co/datasets/CanadaWildFireDaily/CanadaWildFireDaily-v1.canada_realestate_listingsipfs_canada_laws
Canada In-Force Federal Legislation (Justice Laws XML)
Research snapshot of in-force Canadian federal Acts and Regulations
collected from the official Justice Laws Website
XML (Department of Justice Canada). English and French versions are stored as
separate rows (language = en / fr).
Not legal advice. Official Justice Laws Website text prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-02
Coverage
complete in-force federal… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_canada_laws.greatnorth-canada-federal-laws-text
Great North Canada Federal Laws Text Corpus (Expanded)
235,794 training examples from the complete current body of Canadian consolidated federal Acts and Regulations.
This is a significantly expanded version of the corpus, now including:
All consolidated Acts
All consolidated Regulations
Both English and French versions where available
Better chunking optimized for LLM training
Data Characteristics
Section/provision-level text from official Government of Canada… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-canada-federal-laws-text.canada-income-tax-rates
Canadian Income Tax and Payroll Contribution Rates, 2026
Three tables covering what actually comes off a Canadian paycheck.
from datasets import load_dataset
tax = load_dataset("Worklets/canada-income-tax-rates", "income-tax") # 14 rows
contributions = load_dataset("Worklets/canada-income-tax-rates", "payroll-contributions") # 9 rows
credits = load_dataset("Worklets/canada-income-tax-rates", "tax-credits") # 65 rows
income-tax — the federal… See the full description on the dataset page: https://huggingface.co/datasets/Worklets/canada-income-tax-rates.ipfs_canada_laws_ir
Canada legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_canada_laws (revision 7e8066f09c8cdef1fbadad34ea8e35585828e8a0) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Canada prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_canada_laws_ir.CanadaTaxCourtOutcomesLegalBenchClassification
CanadaTaxCourtOutcomesLegalBenchClassification
An MTEB dataset
Massive Text Embedding Benchmark
The input is an excerpt of text from Tax Court of Canada decisions involving appeals of tax related matters. The task is to classify whether the excerpt includes the outcome of the appeal, and if so, to specify whether the appeal was allowed or dismissed. Partial success (e.g. appeal granted on one tax year but dismissed on another) counts as allowed (with the exception of costs orders… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CanadaTaxCourtOutcomesLegalBenchClassification.trademarks_canadaCANADA_ACT_REGULATION_QA
Canadian Acts and Regulation QA
source- https://laws-lois.justice.gc.ca/eng/XML/Legis.xml
model_name="gemini-1.5-flash-latest" with 1 million context length,
First summarize the text scrapped text from xml tree of urls using gemini.
then generate QA from sumarised text.
Performance of Gemini was way way better than GPT-4.
Fitering was done based on Heuristics after rigrous analysis because llms were not always accurate.
summary_prompt_template= """
You'r legal expert… See the full description on the dataset page: https://huggingface.co/datasets/Guggu/CANADA_ACT_REGULATION_QA.canada_ssa_gender_neutral_first_namesThis is the official dataset for Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions
Name-based gender prediction has traditionally categorized individuals as either female or male based on their names, using a binary classification system. That binary approach can be problematic in the cases of gender-neutral names that do not align with any one gender, among other reasons. Relying solely on binary gender categories without recognizing… See the full description on the dataset page: https://huggingface.co/datasets/uzw/canada_ssa_gender_neutral_first_names.canada-trade-data
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/WilgnerCH/canada-trade-data.bank_of_canada
Dataset Summary
For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/bank_of_canada
Additional Information
This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 1,000 sentences taken from the meeting minutes of the Bank of Canada.
Label Interpretation
Stance Detection… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/bank_of_canada.Fendi.Product.prices.Canada
Fendi web scraped data
About the website
The luxury fashion industry is a vibrant and thriving sector in the Americas, particularly in Canada. This industry is marked by the sale of high-end apparel, accessories, footwear, jewelry, and beauty products from renowned global brands like Fendi. Moreover, Canadas luxury fashion industry has embraced digital transformation, leading to a significant growth of Ecommerce platforms. Flourishing online sales have proven to be a… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Fendi.Product.prices.Canada.Canada-Stock-Symbols-and-Metadata
Canada Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in Canada.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Canada-Stock-Symbols-and-Metadata.r-canada-shrunkGGT_Canadathe dataset is for training the model
it is an only Canadian subset
canada-dataset-1b
🇨🇦 Canada Web Text — 1B-token Sample 🍁
A 1-billion-token representative sample of a much larger cleaned Canadian web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐
The full 210B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/canada-dataset-1b.Loro.Piana.Product.prices.Canada
Loro Piana web scraped data
About the website
Loro Piana operates in the luxury fashion industry, specializing in high-end, luxury apparel, accessories, and textile goods. In Americas, particularly in Canada, this industry has experienced substantial growth driven by wealthy individuals and fashion enthusiasts. With the emerging digital landscape and E-commerce, the consumption of luxury fashion has become much feasible due to the ease of accessibility. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Loro.Piana.Product.prices.Canada.Louis.Vuitton.Product.prices.Canada
Louis Vuitton web scraped data
About the website
The luxury fashion industry in the Americas, specifically in Canada is flourishing and significantly competitive. A vital player, Louis Vuitton, has crucially attained a strong positioning in this market. The industry in focus encompasses high-end, exclusive products and services, which are in high demand amongst the affluent sections of society. These products typically include haute couture, ready-to-wear clothing… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Louis.Vuitton.Product.prices.Canada.Canadian-streetview-cities
Canadian Street View Cities Dataset
Overview
A street-view image dataset created to train and evaluate models for city-level image classification across major Canadian cities. Each entry includes an image and its corresponding city label.
Purpose
The dataset is intended for building models that recognize the Canadian city in which a street-view scene was captured.
Data Source
All images were collected from Mapillary, using geographic bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/canada-guesser/Canadian-streetview-cities.hy3-w4a16-mtp-calibration
Hy3 W4A16-MTP — Calibration Set
The exact 512-sample calibration blend used to GPTQ-quantize
canada-quant/hy3-w4a16-mtp
(a W4A16 quantization of tencent/Hy3).
Published for full reproducibility of the quantization pipeline.
Why a blend (not chat-only)
INT4 weight quantization degrades most on code and tool-call-shaped tokens. A chat-only
calibration set (e.g. pure ultrachat) under-samples exactly the routed experts those tokens
activate. This set deliberately… See the full description on the dataset page: https://huggingface.co/datasets/canada-quant/hy3-w4a16-mtp-calibration.Net.a.Porter.Product.prices.Canada
Net-a-Porter web scraped data
About the website
The Ecommerce industry is a growing sector in the Americas, especially in Canada. Businesses in this realm focus on the buying and selling of goods and services over electronic systems, particularly the internet. Net-a-Porter, as an online retail brand, operates within this sector by offering a wide range of luxury fashion items to their Canadian clientele. The dataset analyzed presents relevant Ecommerce product-list page… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Net.a.Porter.Product.prices.Canada.ecma262-qa-synth-v1
262QA (ecma262-qa-synth-v1)
[!CAUTION]
This dataset is experimental and not human-validated. It is published as a proof-of-concept to be potentially useful for experiments instead of sitting on a disk. If you are interested in serious use, let's chat! Have fun :)
This dataset contains a synthetic full-coverage question-answer corpus of ECMA-262. This was originally generated for an LLM benchmark which may be published in the future.
Rows: 1651
Split: train
Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/CanadaHonk/ecma262-qa-synth-v1.unlearning-canadaGucci.Product.prices.Canada
Gucci web scraped data
About the website
Gucci is a key player in the luxury fashion industry in the Americas, particularly in Canada. This industry encompasses a broad range of high-end, premium products such as clothing, accessories, fragrances, and watches, which are regarded for their exceptional quality and exclusivity. The brand has strongly maintained its demand and popularity through both its brick-and-mortar shops and its e-commerce presence. The dataset observed… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Gucci.Product.prices.Canada.canadawastewater-canada
Canada National Wastewater Monitoring of Pathogens (NWMP)
This dataset contains COVID-19, Influenza, and RSV viral load data from sites across Canada, managed by the Public Health Agency of Canada (PHAC).
Coverage
Geography: Canada (National, Provincial, and Municipal levels)
Pathogens: SARS-CoV-2, Influenza A, Influenza B, RSV
Frequency: Weekly updates (published every Friday)
Columns
Column
Description
Date
Sample collection date.
Region… See the full description on the dataset page: https://huggingface.co/datasets/EPI-Eval/wastewater-canada.vax-keyword-canadareddit_comments_subreddit_canada
