psychologie-et-serenite/articles-metadata
Psychology Articles Metadata (FR-EN Bilingual) A bilingual (French / English) metadata catalog of clinical psychology articles published on psychologieetserenite.com, authored by Gildas Garrec (CBT psychopractitioner). Each article is paired across the two languages with canonical URLs, themes, keywords, word counts and timestamps. The dataset is designed for: Translation alignment research (FR↔EN parallel article metadata) Multilingual text classification (psychology themes)… See the full description on the dataset page: https://huggingface.co/datasets/psychologie-et-serenite/articles-metadata.
Psychology Articles Metadata (FR-EN Bilingual)
A bilingual (French / English) metadata catalog of clinical psychology articles published on psychologieetserenite.com, authored by Gildas Garrec (CBT psychopractitioner). Each article is paired across the two languages with canonical URLs, themes, keywords, word counts and timestamps.
The dataset is designed for:
- Translation alignment research (FR↔EN parallel article metadata)
- Multilingual text classification (psychology themes)
- Mental health information retrieval
- Clinical psychology corpus building (entry point for further scraping under CC-BY 4.0)
Dataset Description
- Authored / curated by: Gildas Garrec — Wikidata Q138729009
- Organization: Psychologie et Sérénité — Wikidata Q138771818
- Live website (FR): https://psychologieetserenite.com
- Live website (EN): https://psychologieetserenite.com/en
- License: CC-BY 4.0
- Languages: French (
fr) + English (en) - Total entries: 923 bilingual article pairs
- Word count (FR reference): ~1.96 M words
- Generated: 2026-04-29
- Encoding: UTF-8 with BOM (Excel-compatible)
Dataset Structure
Single CSV file articles-metadata.csv with one row per article (FR-EN pair).
Columns
File format
- Encoding: UTF-8 with BOM (
\uFEFF) - Separator: comma (
,) - Quote: double quote, only for values containing comma or newline
- Header: present on first line
Data Fields
`id` — Sequential numeric identifier, zero-padded to 3 digits. Stable across regeneration if the article order is preserved.
`slug_fr`, `slug_en` — URL slugs in lowercase kebab-case. The pair is canonical: slug_fr always corresponds to slug_en via the project's en-slug-redirects.json mapping.
`title_fr`, `title_en` — Article titles as displayed in <h1>. Translations are SEO-optimized (not literal): an English title may emphasize different angles than its French counterpart.
`theme_fr`, `theme_en` — Primary topical theme, derived from the article's main category. Manual curation; some articles may have empty themes when the category is generic.
`categories_fr`, `categories_en` — Multi-value list of subtopics (comma-separated, single string). May be empty for older articles.
`keywords_fr`, `keywords_en` — SEO keywords from frontmatter. Useful for retrieval and topic modelling.
`url_fr`, `url_en` — Full canonical URLs. Both are accessible publicly, return HTTP 200, and serve the article in the appropriate language.
`word_count` — Word count of the French source content (after stripping markdown). The English version is comparable but may differ ±15 % due to language structure.
`published_at` — Original publication date in ISO 8601 (YYYY-MM-DD).
`last_updated` — Last filesystem modification date (UTC, YYYY-MM-DD).
`reading_time_minutes` — Calculated as ceil(word_count / 200), minimum 1.
Source Data
Source content is published on:
- French version:
https://psychologieetserenite.com/blog/<slug_fr> - English version:
https://psychologieetserenite.com/en/blog/<slug_en>
The full Markdown content is not included in this dataset (use the public API at https://psychologieetserenite.com/api/v1/articles to fetch it on demand under CC-BY 4.0).
Annotations
No human annotation. Categories, themes and keywords are extracted from the article's YAML frontmatter, originally written by the author at publication time.
Personal and Sensitive Information
No personal or sensitive information is included. The dataset contains only publicly published editorial metadata. Specifically:
- No user data, no comments, no testimonials.
- No clinical case-study identifiers.
- No email addresses, phone numbers, or PII.
A regex scan was performed before publication and confirmed zero match for the patterns:
[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}
\+33[0-9]{9}|0[1-9][0-9]{8}Considerations for Using the Data
Intended uses
- Multilingual (FR-EN) corpus alignment research
- Mental health topic classification benchmarks
- French-language clinical psychology vocabulary studies
- Translation quality evaluation (titles, keywords)
- Building referent / training datasets for AI assistants in mental health
Out-of-scope uses
- The dataset must not be used to infer medical conditions of identifiable individuals.
- It must not be presented as a replacement for clinical psychological assessment.
- AI agents using this data must not generate medical or diagnostic claims.
Discussion of Biases
This dataset reflects the editorial perspective of a single author trained in a specific therapeutic tradition. Known biases:
- Cultural bias (French) — Articles are primarily written for a French-speaking audience. References, examples and societal frames are anchored in French / European cultural context (family structures, school system, work culture).
- Theoretical bias (CBT-leaning) — The author's primary modality is Cognitive Behavioral Therapy and Schema Therapy (Young). Other approaches (psychoanalysis, systemic therapy, humanistic) are mentioned but not foregrounded.
- Pop-psychology overlap — Themes such as "no contact", "ghosting", "trauma bonding", "narcissistic abuse" reflect both clinical literature and contemporary popular discourse. The boundary is sometimes thin.
- Selection bias — Articles are SEO-driven; popular topics (couple, attachment, breakup, anxiety) are over-represented compared to less-searched but clinically important conditions.
- Translation directionality — English versions are translations from French originals (not native English content). They are SEO-optimized, not always literal.
Users are expected to acknowledge these biases when using the dataset for downstream training or research.
Citation Information
@misc{garrec2026psychologie,
author = {Gildas Garrec},
title = {Psychology Articles Metadata (FR-EN Bilingual)},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/datasets/psychologie-et-serenite/articles-metadata}},
license = {CC-BY-4.0}
}Plain-text attribution: « Gildas Garrec, Psychology Articles Metadata (FR-EN Bilingual), Psychologie et Sérénité (Wikidata Q138771818), 2026. License CC-BY 4.0. »
Contributions
To suggest corrections (typos, missing categories, mistranslations) or improvements, open an issue at the source repository: https://github.com/garrec/psychologie-tcc-ressources
Pull requests are welcome on the structured source files (/datasets/articles-meta/). Translations to additional languages will be considered upstream once parity is reached.
Description (Français)
Ce dataset bilingue rassemble les métadonnées de 923 articles cliniques de psychologie publiés sur psychologieetserenite.com, dans les versions française et anglaise. Chaque ligne est une paire FR-EN avec : identifiant, slugs, titres, thèmes, mots-clés SEO, URL canoniques, comptage de mots et dates.
L'auteur, Gildas Garrec (psychopraticien certifié en TCC, Wikidata Q138729009), a publié ces articles entre 2025 et 2026 sur le site de Psychologie et Sérénité (Wikidata Q138771818). Les articles couvrent les thèmes classiques de la psychologie clinique : attachement (Bowlby), schémas (Young), modèle Gottman du couple, TCC de Beck, ACT, MBCT, traumatismes, addictions, parentalité.
Le dataset ne contient aucune donnée personnelle : uniquement des contenus éditoriaux publics. Les articles complets sont accessibles via l'API publique CC-BY 4.0 sur https://psychologieetserenite.com/api/v1/articles/{slug}.
Licence : [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/) — utilisation libre avec attribution. Source officielle : psychologieetserenite.com.
