CoolFace
Datasetpublic

Mina-Rajaei-Moghadam/US-Presidents-Spoken-and-Written-Sentences

US Presidents' Spoken and Written Sentences We obtained transcriptions of spoken language from the Miller Center of Public Affairs, University of Virginia, which covers transcriptions from George Washington to the present time. For the writing samples, we used ten complete books written by presidents, three of which we obtained from Project Gutenberg. To ensure the accuracy of calculations, all the pages that were not part of the main content were removed. Furthermore, multiple… See the full description on the dataset page: https://huggingface.co/datasets/Mina-Rajaei-Moghadam/US-Presidents-Spoken-and-Written-Sentences.

sourceHugging Faceccupdated 10mo agoView on Hugging Face
2likes39downloads
Dataset Card

<h1 align="center">US Presidents' Spoken and Written Sentences</h1>

<p align="center"> <img src="Image.png" alt="Dataset overview" width="700"> </p>

We obtained transcriptions of spoken language from the Miller Center of Public Affairs, University of Virginia, which covers transcriptions from George Washington to the present time. For the writing samples, we used ten complete books written by presidents, three of which we obtained from Project Gutenberg. To ensure the accuracy of calculations, all the pages that were not part of the main content were removed. Furthermore, multiple whitespaces were changed into single whitespaces. For sentence extraction, we used the NLTK library, while CoreNLP (version 4.5.7 ) was employed for tokenization and extracting a variety of low-level and high-level syntactic features. These supplementary layers of linguistic information provide valuable resources for researchers interested not only in stylometric analysis but also in broader investigations of syntactic and semantic phenomena.

If you use this dataset in your research, please cite the two papers below in which it was introduced:

Paper: "Text vs. Transcription: A Study of Differences Between the Writing and Speeches of U.S. Presidents"</br> GitHub: https://github.com/mosabrezaei/Text-vs.-Transcription</br> Cite: </br> @inproceedings{rajaei-moghadam2024text,</br> title = {{Text vs. Transcription}: {A} Study of Differences Between the Writing and Speeches of {U.S}. Presidents},</br> booktitle = {Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities (NLP4DH)},</br> author= {Rajaei Moghadam, Mina and Rezaei, Mosab and Aygen, Gülşat and Freedman, Reva},</br> doi={10.18653/v1/2024.nlp4dh-1.35},</br> pages = {352--361},</br> month= {Nov},</br> address = {Miami, USA},</br> year={2024}} </br>

Paper: "Investigating Lexical and Syntactic Differences in Written and Spoken English Corpora"</br> GitHub: https://github.com/mosabrezaei/Lexical-and-Syntactic-Analysis</br> Cite: </br> @inproceedings{moghadam2024investigating,</br> title={Investigating Lexical and Syntactic Differences in Written and Spoken English Corpora},</br> author={Rajaei Moghadam, Mina and Rezaei, Mosab and Williams, Miguel and Aygen, Gülşat and Freedman, Reva},</br> booktitle={The International FLAIRS Conference Proceedings},</br> doi={10.32473/flairs.37.1.135598},</br> volume={37},</br> number={1},</br> month={May},</br> year={2024}}