finaleads/french-corpus-llm-sample
French Corpus LLM — Sample 500 (v1.4.0) FINALEADS LLC builds compliance-ready training datasets for French regulated industries. We turn 2.66 billion tokens of French finance, regulatory, and economic open data into audit-trailed, pseudonymized, AI Act Article 10-documented shares — so foundation model and regtech teams can ship into European enterprises without a data-lineage gap. This is a public sample of 500 stratified documents drawn from the French Premium Web Corpus… See the full description on the dataset page: https://huggingface.co/datasets/finaleads/french-corpus-llm-sample.
v1.4.0: update dsd_certificate_public.crt
v1.4.0: update DATASET_SPECIFICATION_signed.pdf
v1.4.0: update LICENSE_per_source.md
v1.4.0: update DATA_DICTIONARY.md
v1.4.0: update STATS.md
v1.4.0: update README.md
v1.3.0: update SAMPLES.md
v1.3.0: update LICENSE.md
v1.3.0: update STATS.md
v1.3.0: update DATA_DICTIONARY.md
v1.3.0: update MANIFEST.json
v1.3.0: update sample.jsonl
v1.3.0: update README.md
add Snowflake Marketplace listing link in dataset card
add DATASET_SPECIFICATION_signed.pdf
add STATS.md
add DATA_DICTIONARY.md
add LICENSE_per_source.md
add MANIFEST.json
add sample.jsonl
add README.md
initial commit
