datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CanadaWildFireDaily-v1🔥🔥 CanadaWildfireDaily: A Large-Scale Dataset for Daily Wildfire Spread in Canada 🔥🔥
Folder Structure
This section provides the details needed to understand and use the released dataset. We release both raw data and training/val/test ready samples.
CanadaWildFireDaily Train/Validation/Test Samples (data_samples/)
The final training samples are generated through a two-step process.
First, the CSV files and metadata mappers are used to assign fire IDs to the training… See the full description on the dataset page: https://huggingface.co/datasets/CanadaWildFireDaily/CanadaWildFireDaily-v1.greatnorth-canada-federal-laws-text
Great North Canada Federal Laws Text Corpus (Expanded)
235,794 training examples from the complete current body of Canadian consolidated federal Acts and Regulations.
This is a significantly expanded version of the corpus, now including:
All consolidated Acts
All consolidated Regulations
Both English and French versions where available
Better chunking optimized for LLM training
Data Characteristics
Section/provision-level text from official Government of Canada… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-canada-federal-laws-text.hy3-w4a16-mtp-calibration
Hy3 W4A16-MTP — Calibration Set
The exact 512-sample calibration blend used to GPTQ-quantize
canada-quant/hy3-w4a16-mtp
(a W4A16 quantization of tencent/Hy3).
Published for full reproducibility of the quantization pipeline.
Why a blend (not chat-only)
INT4 weight quantization degrades most on code and tool-call-shaped tokens. A chat-only
calibration set (e.g. pure ultrachat) under-samples exactly the routed experts those tokens
activate. This set deliberately… See the full description on the dataset page: https://huggingface.co/datasets/canada-quant/hy3-w4a16-mtp-calibration.ecma262-qa-synth-v1
262QA (ecma262-qa-synth-v1)
[!CAUTION]
This dataset is experimental and not human-validated. It is published as a proof-of-concept to be potentially useful for experiments instead of sitting on a disk. If you are interested in serious use, let's chat! Have fun :)
This dataset contains a synthetic full-coverage question-answer corpus of ECMA-262. This was originally generated for an LLM benchmark which may be published in the future.
Rows: 1651
Split: train
Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/CanadaHonk/ecma262-qa-synth-v1.canada-addressesCanada_Reimbursement_Dataset
