akhvedelidze/UNEditor
Dataset Card for Dataset Name UNEditor Dataset Details The UNEditor dataset is a curated collection of instruction–response examples designed to train language models to produce writing that reflects the formal, neutral, and structured editorial style of the United Nations. Drawing on principles from the United Nations Editorial Manual, the dataset teaches models to apply UN‑specific conventions in tone, terminology, grammar, and document formatting across a wide… See the full description on the dataset page: https://huggingface.co/datasets/akhvedelidze/UNEditor.
Dataset Card for Dataset Name
UNEditor
Dataset Details
The UNEditor dataset is a curated collection of instruction–response examples designed to train language models to produce writing that reflects the formal, neutral, and structured editorial style of the United Nations. Drawing on principles from the United Nations Editorial Manual, the dataset teaches models to apply UN‑specific conventions in tone, terminology, grammar, and document formatting across a wide range of tasks, including drafting, revising, explaining, and editing official-style text. Each example consists of an instruction, optional input, and a high‑quality output that demonstrates clarity, accuracy, and compliance with multilateral communication standards. All content is paraphrased, synthesized, or abstracted, ensuring no direct reproduction of copyrighted UN material while preserving the core editorial logic. The dataset enables fine‑tuning models to reliably produce agendas, statements, concept notes, press releases, and other institutional documents consistent with UN norms. The UNEditor dataset is intended for researchers, practitioners, and organizations seeking to build or enhance models that support diplomatic communication, international governance workflows, or structured institutional writing. Released under the MIT License, it provides a flexible and accessible resource for the development of AI systems aligned with the writing culture of multilateral institutions.
Dataset Description
- **Curated by: Akaki Khvedelidze
- **Language(s) (NLP): Enlgish
- **License: MIT
Dataset Sources [optional]
<!-- Provide the basic links for the dataset. -->
- Repository: [More Information Needed]
- Paper [optional]: [More Information Needed]
- Demo [optional]: [More Information Needed]
Uses
Fine tune model to learn UN style text generation
Annotation process
Automated
Recommendations
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.
Dataset Card Authors
Akaki Khvedelidze
Dataset Card Contact
akhvedelidze@gmail.com
