Linguistics
linguistic-similarityilist
Dataset Card for ilist
Dataset Summary
This dataset is introduced in a task which aimed at identifying 5 closely-related languages of Indo-Aryan language family: Hindi (also known as Khari Boli), Braj Bhasha, Awadhi, Bhojpuri and Magahi. These languages form part of a continuum starting from Western Uttar Pradesh (Hindi and Braj Bhasha) to Eastern Uttar Pradesh (Awadhi and Bhojpuri) and the neighbouring Eastern state of Bihar (Bhojpuri and Magahi).
For this task… See the full description on the dataset page: https://huggingface.co/datasets/kmi-linguistics/ilist.adaption-language-linguistics-qa
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-language_linguistics_qa
This dataset consists of instruction and response pairs covering a broad range of topics within language and linguistics. Entries address fundamental concepts such as grammar syntax, vocabulary, pronunciation, and writing systems, alongside applied disciplines like computational linguistics, translation, and localization. Additional content explores language… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-language-linguistics-qa.ptpt-linguistics-if
Portuguese Linguistics Instructions
Dataset handwritten by a Portuguese linguistics expert, covering topics such as phonetics, orthography, wordplay, idiomatic expressions, and grammar classification, specific to European Portuguese.
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If you use this dataset or AMALIA in your work, please cite:… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/ptpt-linguistics-if.linguistic_sq
Physics and Math Problems Dataset
This repository contains a dataset of 3,200 enteries of different Albanian linguistics to improve Albanian queries further by introducing Albanian language rules and literature. The dataset is designed to support various NLP tasks and educational applications.
Dataset Overview
Problems: 3,200
Language: Albanian
Topics:
letërsi shqiptare: 47
poezi shqiptare: 45
proza shqiptare: 48
drama shqiptare: 46
autorë shqiptarë: 48
veprat kryesore… See the full description on the dataset page: https://huggingface.co/datasets/LTS-VVE/linguistic_sq.Hatamti-Linguistics
