datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stop_wordshausa-stopwords-corpus
Hausa Stopword Candidates and Frequency Scores
A reproducible Hausa lexical resource containing frequency-scored stopword candidates. This repository is organized for inspection, preprocessing experiments, and future Hausa-speaker review. It does not publish a final stopword list or a final human-reviewed stopword count.
Quick navigation
Need
Go to
Browse candidates in the Dataset Viewer
data/hausa_stopword_candidates.jsonl
Efficient analysis… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/hausa-stopwords-corpus.turkish-stopwords
🇹🇷 TR Turkish Stopwords – Extended List (504 Words)
This dataset contains the most comprehensive and extended list of Turkish stopwords, curated specifically for Natural Language Processing (NLP) tasks involving the Turkish language.
📦 Dataset Overview
Total Stopwords: 504
Format: JSON
Key: "stopwords"
File Size: ~7.7 kB
License: Apache 2.0
📚 Description
Türkçe:
“TR Türkçe Stopwords” veriseti, Türkçe metinlerde en sık karşılaşılan… See the full description on the dataset page: https://huggingface.co/datasets/nezahatkorkmaz/turkish-stopwords.
