wordlist
Datasets
All datasets matching “wordlist”Wordlists
Dataset Card for Canstralian/Wordlists
Canstralian/Wordlists is a comprehensive, curated collection of wordlists tailored for cybersecurity professionals, researchers, and enthusiasts. This dataset is optimized for tasks such as penetration testing, ethical hacking, and password strength analysis. Its structured design ensures high usability across various cybersecurity applications.
Dataset Details
Dataset Description
This dataset includes wordlists that… See the full description on the dataset page: https://huggingface.co/datasets/Canstralian/Wordlists.wordlist
wordlist
A structured digital lexicon containing 43,509 Hmar words, phrases, definitions, and translations compiled from five lexicographical sources.
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Languages: Hmar (hmr, ISO 639-3, Glottolog: hmar1241), English (en)
Family: Zo Languages
Volume: 43,509 lexical entries across 5 dictionary files
Format: JSON (data/dictionary-001.json to data/dictionary-005.json)
License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/wordlist.glotlid-wordlists
GlotLID Wordlists
This is a set of wordlists extracted from the GlotLID-corpus for high-precision filtering of FineWeb2.
Download
The recommended way to download the data is via git clone:
git clone https://huggingface.co/datasets/cis-lmu/glotlid-wordlists
Method and Usage for Precision Filtering
For details on the filtering method, please refer to the FineWeb2 paper.
Each word listed occurs significantly more often in its own language dataset than in any… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/glotlid-wordlists.iwn_wordlistsWe provide the unique word list form the IndoWordnet (IWN) knowledge base.webonary-african-wordlists
Webonary African Wordlists
Aggregated and cleaned dictionary wordlists for 168 African languages scraped from Webonary, containing 562,573 unique headwords and 725,843 total translation/gloss rows. Each language is provided as an independent dataset configuration (subset).
Dataset Structure
Each subset contains the following columns:
slug: Webonary language identifier slug
language: Language display name
iso: ISO 639-3 code
country: Country code(s) (ISO 3166-1… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/webonary-african-wordlists.wordlist
Dicta Unvocalized Hebrew Wordlist
Every distinct unvocalized (nikud-stripped) Hebrew word form Dicta knows, with its
morphology, lemma, and root. There is no compact natural key: rows are distinguished only by the
full combination of word + lex + morphology + root + the linguistic flags. Use id as the row key.
This repo has two configs:
default — the full wordlist below (every theoretically valid word-object).
attested — wordforms actually observed in Dicta corpora, with the… See the full description on the dataset page: https://huggingface.co/datasets/dicta-il/wordlist.
