CoolFace
Datasetpublic

FrancophonIA/Romance_Verbal_Inflection_Dataset_2.0.4

[!NOTE] Dataset origin: https://zenodo.org/records/4039059 This Database is an open source and downloadable version of the Oxford Database of the Inflectional Morphology of the Romance Verb, as a CLDF WordList module. It is intended to facilitate quantitative analysis. This version was created in the following way: We scraped the online database to reconstruct an image of the original database, We then reorganized, cleaned up, and normalized the resulting tables into a CLDF WordList… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Romance_Verbal_Inflection_Dataset_2.0.4.

sourceHugging Facegpl-3.0updated 1y agoView on Hugging Face
0likes16downloads
Dataset Card
[!NOTE] Dataset origin: https://zenodo.org/records/4039059

This Database is an open source and downloadable version of the Oxford Database of the Inflectional Morphology of the Romance Verb, as a CLDF WordList module. It is intended to facilitate quantitative analysis.

This version was created in the following way:

  • —We scraped the online database to reconstruct an image of the original database,
  • —We then reorganized, cleaned up, and normalized the resulting tables into a CLDF WordList, adding metadata when needed.
  • —We added latin entries from LatInfLexi 1.1

The scripts used for that purpose can be found in the scraping and formatting the ODRVM repository.

The original documentation can be found online at http://romverbmorph.clp.ox.ac.uk/.

Tables -------

  • —languages.csv is a LanguageTable with information on each variety:
IDLatitudeLongitudeClosest_GlottocodeTownCommentVarietyLinguistic_DateSourcesCountryFamilyRegionVariety_in_ODRVM
Sardinian_Sassarese40.72678.55917sass1235Sassareselate 20th to early 21st centuryCamilli1929ItalySardinianSardiniaSardinian - Sassarese
FrenchNormanfrom_Jersey49.1833-2.10694norm1245JerseyNormanmid to late 20th centuryLiddicoat1994JerseyFrenchChannel IslandsFrench - Norman - Jersey
ItalianNorthernIPiedmonteseBassofromCascinagrossa44.87868.72028lowp1238CascinagrossaNorthern I > Piedmontese > Bassolate 20th to early 21st centuryCastellani2002ItalyItalianPiedmontItalian - Northern I - Piedmontese - Basso - Cascinagrossa

Full languages:

  • —CatalanWesternfrom_Lleida
  • —CatalanEasternCentralBarcelonifrom_Barcelona
  • —CatalanEasternAlgueresfromLAlgher_Alghero
  • —CatalanEasternBalearMallorquifrom_Palma
  • —CatalanEasternRossellonesfromCanetdeRossello_Canet-en-Roussillon
  • —Friulian
  • —FrenchWallonfrom_Namur
  • —FrenchPoitevin-SaintongeaisVendeenfromBeauvoir-sur-Mer
  • —FrenchFranc-Comtoisfrom_Pierrecourt
  • —FrenchNormanfrom_Jersey
  • —FrenchAcadianSouth-EastNewBrunswickfromMoncton__environs
  • —FrenchNormanGuernsey
  • —FrenchAcadianfrom_Pubnico
  • —FrenchPicardfrom_Mesnil-Martinsart
  • —FrenchLorrainfrom_Ranrupt
  • —FrenchModernStandard
  • —SwissRaeto-RomanceSurselvan
  • —SwissRaeto-RomanceSurmiranfromBivio-Stalla
  • —SpanishAsturo-LeoneseAsturianSudestedefromParres
  • —SpanishModernStandard
  • —SwissRaeto-RomancePuterUpperEngadine
  • —SpanishAsturo-LeoneseAsturianfromSomiedo
  • —SpanishAragonesefrom_Panticosa
  • —SpanishAragoneseAnsotanofromAnso_Fago
  • —Sardinian_Sassarese
  • —Sardinian_Nuorese
  • —Sardinian_Logudorese
  • —Sardinian_Gallurese
  • —Sardinian_Campidanese
  • —Romanian_Megleno-Romanian
  • —RomanianModernStandard
  • —RomanianIstro-RomanianfromSusnjevicaValdarsa
  • —Romanian_Aromanian
  • —OccitanSouthernLanguedocienfromGraulhet
  • —OccitanSouthernProvencalfromNice
  • —OccitanNorthernVivaro-AlpinfromSeyne
  • —ItalianSouthernIPuglieseDaunofromLucera
  • —ItalianSouthernIISicilianCentralfromMussomeli
  • —OccitanNorthernLimousinfromSaint-Augustin
  • —OccitanNorthernAuvergnatGartempefromGartempeCreuse_
  • —Ladin-DolomiticAtesinoValBadia
  • —Ladin-DolomiticAtesinoValGardena
  • —ItalianSouthernILucanoCentralfromCalvello
  • —ItalianSouthernIMolisanofrom_Casacalenda
  • —ItalianSouthernILucanoCalabriafromPapasidero
  • —ItalianSouthernILucanoArchaicfromNova_Siri
  • —ItalianNorthernIIVenetoNorthernfromAlpago
  • —ItalianNorthernIIVenetoIstriotoValledIstriafromValle_dIstria
  • —ItalianNorthernILombardCremonesefromCremona
  • —ItalianNorthernIPiedmontesefromCairoMontenotte
  • —ItalianNorthernIPiedmonteseBassofromCascinagrossa
  • —ItalianNorthernILombardAlpinefromVal_Calanca
  • —ItalianNorthernILigurianGenoesefromGenova
  • —ItalianNorthernIEmilianfrom_Travo
  • —ItalianNorthernIEmilianRomagnolfromLugo
  • —ItalianCentralTuscanCorsicanfrom_Sisco
  • —ItalianCentralModern_Standard
  • —ItalianCentralMarchigianofromServigliano
  • —ItalianCentralMarchigianofromMacerata
  • —Galego-Portuguese_Portuguese
  • —ItalianCentralLazialeNorth-centralfrom_Ascrea
  • —Galego-PortugueseGalicianfrom_Xermade
  • —Galego-PortugueseGalicianfrom_Lubian
  • —Galego-PortugueseGalicianfromFisterraFinisterra
  • —Galego-PortugueseGalicianfrom_Dodro
  • —Galego-PortugueseGalicianfromVilanovade_Oscos
  • —Galego-PortugueseGalicianfrom_Cualedro
  • —FrenchAcadianfromBaieSainte-Marie
  • —FriulianWesternManiagofromGreci
  • —FrancoprovencalValaisanValdIlliezfromValdIlliez
  • —DalmatianVegliotefromVegliaKrk
  • —FrancoprovencalLyonnaisfrom_Vaux
  • —CatalanWesternfrom_Valencia
  • —ItalicLatino-FaliscanLatin
  • —The table parameters.csv is a ParameterTable which describes the morphosyntactic cells of the paradigm. It was made manually following the online documentation, which is quoted in the "description" field. The IDs are used in the forms table. The 'Continuants' column links latin paradigm cells to their continuant labels.
IDNameDescriptionContinuants
PLUP-INDLatin pluperfect indicativeCONTLATPLUP-IND
IMPERF-SBJVLatin imperfect subjunctiveCONTLATIMPERF-SBJV
3PL3plthird person plural
  • —The table cognatesets.csv is a CognateSetTable. It provides information for each unique etymon. The IDs are used in the form and lexeme tables. The LemLat_ID links back latin etymons to LatInFlexi identifiers.
IDLemLat_IDLanguage_of_the_etymonEtymonLatin_ConjugationPart_Of_SpeechDerived_from
pascerepasco/-orLatinpascere
IIIV
akkattareRomance*akkattare
V
cubarecuboLatincubare
IV
  • —The lexemes.csv table is a custom table which provides a meaning and a potential comment for each lexeme, that is each instance of an etymon in a given variety.
IDEtymon_in_ODRVMMeaningCommentCognateset_IDLanguage_ID
lex_1340PLACEREpleaseplacereFrenchWallonfrom_Namur
lex_2039UENIREcomeuenireItalianNorthernIEmilianfrom_Travo
lex_127AMBULARE / IRE / UADEREgoambulare~ire~uadereFrenchAcadianSouth-EastNewBrunswickfromMoncton__environs
  • —form.csv is a FormTable representing a paradigm in long form, where each row corresponds to a specific form in one variety. The table refers to the languages, cognatesets and parameter tables through respectively the LanguageID, CognateSetID and Cell columns.
IDLanguage_IDCellFormCognateset_ID
form_723896RomanianModernStandardPRS-SBJV~1PLˈnaʃtemnasci
form_2203100Galego-Portuguese_PortugueseINFL_INF~3PLpɾɐˈzeɾɐ̃ĩplacere
form_2262553ItalianCentralMarchigianofromMacerataROM_FUT~3PLaˈvrahabere

Character Inventory -------------------

Forms are entered in IPA notation, using the following characters.

bilabiallabio-dentaldentalalveolarpost-alveolarretroflexpalatallabio-palatallabio-velarvelaruvularglottal
stopb pd tɖɟ cg k
nasalmnɲŋ
trillrʀ
tapɾ
fricativeβ ɸv fð θz sʒ ʃʝ çɣ xʁ χh
affricateʣ ʦʤ ʧ
approximantɹjɥw
lateral approximantlʎ
frontnear-frontcentralnear-backback
closei yɨu
near-closeɪʊ
close-mide øɘo
midə
open-midɛ œʌ ɔ
near-openæɐ
opena ɶɑ ɒ
categorydiacriticvalue
syllable◌ˈstressed
consonant◌̩syllabic
consonant◌ʲpalatalized
vowel◌̝raised
vowel◌̯non-syllabic
vowel◌̃nasalized
vowel◌ːlong
vowel◌̞lowered

In addition:

  • —parenthesis mark elements which are only optionally or variably present in pronunciation
  • —brackets mark clitics when either we do not have an example of that form without the clitic and, or when it is possible that without the clitic the form of the verb would be slightly different.

People and contacts -------------------

The following have contributed to building the Database:

  • —Silvio Cruschina
  • —Maria Goldbach
  • —Marc-Olivier Hinzelin
  • —Martin Maiden
  • —John Charles Smith

We are grateful also to Chiara Cappellaro, Louise Esher, Paul O'Neill, Stephen Parkinson, Mair Parry, Nicolae Saramandu, Andrew Swearingen, Francisco Dubert García and Tania Paciaroni for their assistance.

The CLDF version was created and cleaned by Sacha Beniamine and Erich Round.

Any queries or suggestions regarding the Database should be directed to: <martin.maiden@mod-langs.ox.ac.uk> and / or <johncharles.smith@mod-langs.ox.ac.uk>

Citation

Beniamine, S., Maiden, M., & Round, E. (2020). Romance Verbal Inflection Dataset 2.0.0 (2.0.4) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.4039059