datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
greatnorth-canada-federal-laws-text
Great North Canada Federal Laws Text Corpus (Expanded)
235,794 training examples from the complete current body of Canadian consolidated federal Acts and Regulations.
This is a significantly expanded version of the corpus, now including:
All consolidated Acts
All consolidated Regulations
Both English and French versions where available
Better chunking optimized for LLM training
Data Characteristics
Section/provision-level text from official Government of Canada… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-canada-federal-laws-text.CANADA_ACT_REGULATION_QA
Canadian Acts and Regulation QA
source- https://laws-lois.justice.gc.ca/eng/XML/Legis.xml
model_name="gemini-1.5-flash-latest" with 1 million context length,
First summarize the text scrapped text from xml tree of urls using gemini.
then generate QA from sumarised text.
Performance of Gemini was way way better than GPT-4.
Fitering was done based on Heuristics after rigrous analysis because llms were not always accurate.
summary_prompt_template= """
You'r legal expert… See the full description on the dataset page: https://huggingface.co/datasets/Guggu/CANADA_ACT_REGULATION_QA.canada-china-trade
Canada-China B2B Trade Dataset
Dataset Description
A curated dataset of Canada-China bilateral trade statistics, commodity breakdowns, provincial data, and B2B sourcing knowledge for use in AI/LLM research and applications.
Maintained by: MapleBridge.io — AI-powered B2B matching platform for Canada-China trade.
Dataset Contents
File
Description
Rows
canada_china_trade_annual.csv
Annual bilateral trade volume 2015-2024 (CAD billions)
10… See the full description on the dataset page: https://huggingface.co/datasets/maplebridge/canada-china-trade.ecma262-qa-synth-v1
262QA (ecma262-qa-synth-v1)
[!CAUTION]
This dataset is experimental and not human-validated. It is published as a proof-of-concept to be potentially useful for experiments instead of sitting on a disk. If you are interested in serious use, let's chat! Have fun :)
This dataset contains a synthetic full-coverage question-answer corpus of ECMA-262. This was originally generated for an LLM benchmark which may be published in the future.
Rows: 1651
Split: train
Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/CanadaHonk/ecma262-qa-synth-v1.
