CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chendelong /linguistic-similaritytabularn<1K1 likes706 downloads2y agoHugging Face02kmi-linguistics /ilist Dataset Card for ilist Dataset Summary This dataset is introduced in a task which aimed at identifying 5 closely-related languages of Indo-Aryan language family: Hindi (also known as Khari Boli), Braj Bhasha, Awadhi, Bhojpuri and Magahi. These languages form part of a continuum starting from Western Uttar Pradesh (Hindi and Braj Bhasha) to Eastern Uttar Pradesh (Awadhi and Bhojpuri) and the neighbouring Eastern state of Bihar (Bhojpuri and Magahi). For this task… See the full description on the dataset page: https://huggingface.co/datasets/kmi-linguistics/ilist.texttext-classification10K<n<100K1 likes229 downloads2y agoHugging Face03Reubencf /adaption-language-linguistics-qa This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-language_linguistics_qa This dataset consists of instruction and response pairs covering a broad range of topics within language and linguistics. Entries address fundamental concepts such as grammar syntax, vocabulary, pronunciation, and writing systems, alongside applied disciplines like computational linguistics, translation, and localization. Additional content explores language… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-language-linguistics-qa.text1K<n<10K0 likes47 downloads22d agoHugging Face04LTS-VVE /linguistic_sq Physics and Math Problems Dataset This repository contains a dataset of 3,200 enteries of different Albanian linguistics to improve Albanian queries further by introducing Albanian language rules and literature. The dataset is designed to support various NLP tasks and educational applications. Dataset Overview Problems: 3,200 Language: Albanian Topics: letërsi shqiptare: 47 poezi shqiptare: 45 proza shqiptare: 48 drama shqiptare: 46 autorë shqiptarë: 48 veprat kryesore… See the full description on the dataset page: https://huggingface.co/datasets/LTS-VVE/linguistic_sq.textquestion-answering1K<n<10K0 likes41 downloads1y agoHugging Face05Dara-Hatamti /Hatamti-Linguisticstexttext-generationn<1K1 likes39 downloads6d agoHugging Face06mlfoundations-dev /stackexchange_linguisticstext10K<n<100K1 likes32 downloads2y agoHugging Face07Kolawolettyy12 /Yoruba-linguistics-dataset Yoruba Qwen Fine-Tuned Model Overview This model is a LoRA fine-tuned version of Qwen2.5-0.5B-Instruct developed for Yoruba language instruction-following tasks. The project explores the use of parameter-efficient fine-tuning for low-resource African language NLP, with a particular focus on Yoruba. Dataset The dataset was adapted using Adaption Lab and contains: Total records: 3,598 Training samples: 3,238 Validation samples: 360 Language: Yoruba… See the full description on the dataset page: https://huggingface.co/datasets/Kolawolettyy12/Yoruba-linguistics-dataset.text1K<n<10K1 likes29 downloads1mo agoHugging Face08Maithili-Computational-Linguistics-Lab /maithiliNewsDatagatedFor access do mail on rockerritesh4@gmail.com or find us here https://x.com/Rocker_Ritesh @misc{yadav2025maibertspeakmaithili, title={Can maiBERT Speak for Maithili?}, author={Sumit Yadav and Raju Kumar Yadav and Utsav Maskey and Gautam Siddharth Kashyap Md Azizul Hoque and Ganesh Gautam}, year={2025}, eprint={2509.15048}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2509.15048}, } texttext-classification10K<n<100K1 likes16 downloads6mo agoHugging Face09Maithili-Computational-Linguistics-Lab /maithili_poemgatedtextn<1K0 likes10 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.