datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Luciole-PostTraining-Dataset-1.1
Table of Contents
Dataset Description
Curation Rationale
Bias, Risks, and Limitations
Data Subsets
Sample Metadata
Downloading the Data
Available Configurations
Loading Examples
Accessing Data Through the Directory Hierarchy
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole-PostTraining-Dataset-1.1 is a curated collection of open, instruction-style text data designed for language model post training. It includes a mixture of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-PostTraining-Dataset-1.1.comp-mechEach record in the dataset contains the following fields:
target_new: the counterfactual term
target_true: the actual term
subject: the topic of the prompt
base_prompt: the foundational prompt
prompt: the modified prompt incorporating the counterfactual change
template: the sentence structure using the counterfactual wording.
OpenLLM-France__Lucie-7B-Instruct-v1.1-details
Dataset Card for Evaluation run of OpenLLM-France/Lucie-7B-Instruct-v1.1
Dataset automatically created during the evaluation run of model OpenLLM-France/Lucie-7B-Instruct-v1.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OpenLLM-France__Lucie-7B-Instruct-v1.1-details.DistillDetect-normalized-traces
DistillDetect — format-normalized teacher traces
Teacher responses from Reference-Based Distillation Detection in LLMs
(arXiv:2607.09692), rewritten so that every
teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set
pairs.
Why this exists
In the released data each teacher emits a structurally different response, so a
student trained on it — and any detector trained to attribute it — can key on
surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.shp-ai-dataset
SHP-AI Dataset 🇦🇱
Dataset instruksional për trajnimin e AI-t të Shtabit të Përgjithshëm të Forcave të Armatosura të Shqipërisë.
Përshkrimi
Ky dataset përmban pyetje-përgjigje në gjuhën shqipe mbi:
Strukturën organizative të Forcave të Armatosura
Shtabin e Përgjithshëm dhe departamentet J
Forcën Tokësore, Ajrore dhe Detare
Integrimin NATO dhe misionet ndërkombëtare
Legjislacionin e mbrojtjes
Doktrinën ushtarake shqiptare
Historinë ushtarake
Statistika
Total… See the full description on the dataset page: https://huggingface.co/datasets/franceskoshahinasilogicleaders/shp-ai-dataset.Luciole_RAG_unformatted
RAG citation benchmarks
Prompt-agnostic RAG benchmark exports for citation and retrieval-grounded QA evaluation. Its purpose is to be used with https://github.com/OpenLLM-France/lighteval/blob/main/community_tasks/luciole_rag.py
Each row has:
id: stable example id
query: user question
retrieved_documents: retrieved context chunks (relevant or not), our main contribution is creating this column for tatqa
titles: chunk titles aligned with retrieved_documents
supporting_index:… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole_RAG_unformatted.cot-oracle-qwen3-8b-onpolicy-recipe
CoT Activation Oracle — On-Policy Qwen3-8B Training Recipe
A reproduction of the on-policy Qwen3-8B training mixture from
Building Better Activation Oracles
(Bauer, De Schamphelaere, Karvonen, Luick, Nanda).
This repository is a recipe card only — it documents the exact dataset
mixture, points at every source on the Hub, and gives regeneration instructions
for the pieces that are no longer available upstream. No third-party data is
re-hosted here; original datasets are linked… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/cot-oracle-qwen3-8b-onpolicy-recipe.francesco-federico-agentic-cmo
Francesco Federico — The Agentic CMO Knowledge Base
A comprehensive, structured knowledge base about Francesco Federico, Global Chief Marketing Officer at S&P Global, author of The Agentic CMO: A Playbook for the Hybrid Marketing Team, and publisher of the Chronicles of Change newsletter.
This dataset captures the breadth and depth of Francesco's professional expertise, career history, published thought leadership, speaking engagements, board positions, and strategic frameworks… See the full description on the dataset page: https://huggingface.co/datasets/frandrake/francesco-federico-agentic-cmo.metro_stores
🏬 METRO France – Réseau de magasins (dataset)
Description
Ce dataset recense l’ensemble des magasins METRO en France, avec leurs informations
de localisation, de contact et d’horaires d’ouverture.
Il est destiné à des usages de :
cartographie et géolocalisation,
analyse territoriale et commerciale,
enrichissement de bases de données,
systèmes RAG et applications IA,
projets open data.
Les données sont structurées et disponibles en CSV et JSON optimisé pour les LLMs.… See the full description on the dataset page: https://huggingface.co/datasets/METRO-France/metro_stores.OpenLLM-France__Lucie-7B-Instruct-human-data-details
Dataset Card for Evaluation run of OpenLLM-France/Lucie-7B-Instruct-human-data
Dataset automatically created during the evaluation run of model OpenLLM-France/Lucie-7B-Instruct-human-data
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OpenLLM-France__Lucie-7B-Instruct-human-data-details.salaries
Datapizza Salaries
A crowd-sourced dataset of salaries from tech workers in Italy, collected via anonymous survey by Datapizza.
Dataset Description
This dataset contains self-reported salary and professional information from technology workers across Italy. Data is collected through an anonymous survey and updated weekly.
Note: Data submitted before November 6th, 2024 contains only partial information, as the initial survey version did not collect fields like… See the full description on the dataset page: https://huggingface.co/datasets/francesca28/salaries.kia-dataset
KIA Dataset 🇦🇱
Dataset instruksional për trajnimin e AI-t të Shtabit të Përgjithshëm të Forcave të Armatosura të Shqipërisë.
Përshkrimi
Ky dataset përmban pyetje-përgjigje në gjuhën shqipe mbi:
Strukturën organizative të Forcave të Armatosura
Shtabin e Përgjithshëm dhe departamentet J
Forcën Tokësore, Ajrore dhe Detare
Integrimin NATO dhe misionet ndërkombëtare
Legjislacionin e mbrojtjes
Doktrinën ushtarake shqiptare
Historinë ushtarake
Statistika
Total… See the full description on the dataset page: https://huggingface.co/datasets/franceskoshahinasilogicleaders/kia-dataset.OpenLLM-France__Lucie-7B-details
Dataset Card for Evaluation run of OpenLLM-France/Lucie-7B
Dataset automatically created during the evaluation run of model OpenLLM-France/Lucie-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OpenLLM-France__Lucie-7B-details.OpenLLM-France__Lucie-7B-Instruct-details
Dataset Card for Evaluation run of OpenLLM-France/Lucie-7B-Instruct
Dataset automatically created during the evaluation run of model OpenLLM-France/Lucie-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OpenLLM-France__Lucie-7B-Instruct-details.ikigai-qa-datasetFrance_MA_Datasetchistes_eugenio
