Jaymerry/NotreDameDeParis-FR
VictorHugo-Structured (FR) — Notre-Dame de Paris Dataset Description VictorHugo-Structured (FR) is a structured literary dataset derived from the French public-domain novel Notre-Dame de Paris (1831) by Victor Hugo. The dataset provides chapter-level records enriched with machine-generated summaries, keywords, and character mentions, while preserving the original text verbatim. This dataset is intended for NLP research, digital humanities, and experimentation with… See the full description on the dataset page: https://huggingface.co/datasets/Jaymerry/NotreDameDeParis-FR.
VictorHugo-Structured (FR) — Notre-Dame de Paris
Dataset Description
VictorHugo-Structured (FR) is a structured literary dataset derived from the French public-domain novel Notre-Dame de Paris (1831) by Victor Hugo.
The dataset provides chapter-level records enriched with machine-generated summaries, keywords, and character mentions, while preserving the original text verbatim.
This dataset is intended for NLP research, digital humanities, and experimentation with LLMs on public-domain literature.
Languages
- French (
fr)
Dataset Structure
Each record corresponds to one chapter.
Main Fields
text: full original chapter text (public domain)summary: LLM-generated chapter summarykeywords: extracted thematic keywordscharacters: characters mentioned in the chapterbook_number,book_title: book-level structure (Livre Premier, etc.)chapter_number,chapter_title,chapter_romanchapter_id: stable unique identifier- metadata fields (source, license, version, character counts)
Source Data
- Work: Notre-Dame de Paris
- Author: Victor Hugo
- First publication: 1831
- Source: Project Gutenberg (eBook #19657)
- URL: https://www.gutenberg.org/ebooks/19657
Non-narrative sections (front matter, notes, license text) were excluded during processing.
Annotations
The following fields are machine-generated using a local LLM (GGUF) via LM Studio:
summarykeywordscharacters
Annotations are chapter-scoped and do not introduce information beyond the original text.
Intended Uses
- NLP experiments (summarization, keyword extraction, NER)
- Digital humanities research
- Evaluation of LLMs on structured literary data
- RAG pipelines using public-domain texts
- Educational and research purposes
Out-of-Scope Uses
- Claiming human authorship of the annotations
- Misrepresenting machine-generated summaries as original literary analysis
- Proprietary relicensing of the dataset
Licensing Information
Text Content
- Notre-Dame de Paris is in the public domain.
- The text is sourced from Project Gutenberg.
- No endorsement by Project Gutenberg is implied.
Annotations and Dataset Structure
- All annotations are machine-generated.
- Dataset structure and annotations are released under CC0-1.0.
Ethical Considerations
- No personal data related to living individuals
- No sensitive or harmful content intentionally added
- Machine-generated annotations may contain inaccuracies
Citation
VictorHugo-Structured (FR) — Notre-Dame de Paris. Derived from Victor Hugo (1831), Project Gutenberg eBook #19657. Machine-generated annotations via local LLM. Processed by Jeremy Banchet (Jaymerry)
