CoolFace
Datasetpublic

akhvedelidze/UNEditor

Dataset Card for Dataset Name UNEditor Dataset Details The UNEditor dataset is a curated collection of instruction–response examples designed to train language models to produce writing that reflects the formal, neutral, and structured editorial style of the United Nations. Drawing on principles from the United Nations Editorial Manual, the dataset teaches models to apply UN‑specific conventions in tone, terminology, grammar, and document formatting across a wide… See the full description on the dataset page: https://huggingface.co/datasets/akhvedelidze/UNEditor.

sourceHugging Facemitupdated 8mo agoView on Hugging Face
0likes1downloads
Dataset Card

Dataset Card for Dataset Name

UNEditor

Dataset Details

The UNEditor dataset is a curated collection of instruction–response examples designed to train language models to produce writing that reflects the formal, neutral, and structured editorial style of the United Nations. Drawing on principles from the United Nations Editorial Manual, the dataset teaches models to apply UN‑specific conventions in tone, terminology, grammar, and document formatting across a wide range of tasks, including drafting, revising, explaining, and editing official-style text. Each example consists of an instruction, optional input, and a high‑quality output that demonstrates clarity, accuracy, and compliance with multilateral communication standards. All content is paraphrased, synthesized, or abstracted, ensuring no direct reproduction of copyrighted UN material while preserving the core editorial logic. The dataset enables fine‑tuning models to reliably produce agendas, statements, concept notes, press releases, and other institutional documents consistent with UN norms. The UNEditor dataset is intended for researchers, practitioners, and organizations seeking to build or enhance models that support diplomatic communication, international governance workflows, or structured institutional writing. Released under the MIT License, it provides a flexible and accessible resource for the development of AI systems aligned with the writing culture of multilateral institutions.

Dataset Description

  • —**Curated by: Akaki Khvedelidze
  • —**Language(s) (NLP): Enlgish
  • —**License: MIT

Dataset Sources [optional]

<!-- Provide the basic links for the dataset. -->

  • —Repository: [More Information Needed]
  • —Paper [optional]: [More Information Needed]
  • —Demo [optional]: [More Information Needed]

Uses

Fine tune model to learn UN style text generation

Annotation process

Automated

Recommendations

<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->

Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.

Dataset Card Authors

Akaki Khvedelidze

Dataset Card Contact

akhvedelidze@gmail.com