datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stopwords
NLTK Stopwords
Stopword lists from NLTK, covering 33 languages.
Each language is a separate config. Each row is one stopword.
Usage
from datasets import load_dataset
# Load one language
ds = load_dataset("nltk-data-hub/stopwords", "portuguese")
words = ds["stopwords"]["word"]
# Load all languages
for lang in ['albanian', 'arabic', 'azerbaijani', 'basque', 'belarusian', 'bengali', 'catalan', 'chinese', 'danish', 'dutch', 'english', 'finnish', 'french', 'german', 'greek'… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/stopwords.peacock-data-public-datasets-hubwords
NLTK Word Lists
English word lists from NLTK,
the New General Service List Project,
and Bing Liu's Opinion Lexicon.
Configs
Config
Words
Schema
License
Source
en
235,886
word
NLTK (other)
NLTK words corpus
en-basic
850
word
Public domain
Ogden Basic English (1930)
ngsl
2,809
word, rank, sfi, freq_per_million
CC-BY-SA 4.0
New General Service List 1.2
toeic
1,250
word, rank, sfi, freq_per_million
CC-BY-SA 4.0
TOEIC Service List 1.2
nawl
963
word, rank… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/words.crubadan
NLTK Crúbadán Language ID Corpus
Character 3-gram frequency tables for 449 writing systems, collected
by Kevin Scannell's An Crúbadán web crawler (2010).
Distributed via NLTK.
Trigrams use < (word start) and > (word end) as boundary markers.
Configs
Config
Description
Schema
table
Language metadata
crubadan_code, iso639_3, language_name
{lang_code}
Per-language trigrams
count, trigram
All 449 language codes: ab, abn, ace, ach, acu, ada, af, agr, aja, ak… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/crubadan.Project1-data
Dataset Card for SQuAD
Dataset Summary
Stanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be unanswerable.
SQuAD 1.1 contains 100,000+ question-answer pairs on 500+ articles.
Supported Tasks and Leaderboards
Question… See the full description on the dataset page: https://huggingface.co/datasets/altannavchnyamsambuu-hub/Project1-data.hub_large_ls960-ft_spot_data_alldatahubWARNING ⚠️: Accessing datahub data via hugging face is still under construction (use https://github.com/cbioPortal/datahub instead).
cBioPortal Datahub
These are parquet files generated from https://github.com/cBioPortal/datahub. It combines all studies' sample, patient, and mutation data into parquet files. They were generated using cbiohubpy.
How to use
Query all mutations of datahub directly in duckdb:
SELECT count(distinct (Chromosome, Start_Position… See the full description on the dataset page: https://huggingface.co/datasets/cBioPortal/datahub.flipkart-dataai-hub-conversation-datanames
NLTK Names Corpus
Name lists from NLTK, split by gender.
Each gender is a separate config. Each row is one name.
Usage
from datasets import load_dataset
ds = load_dataset("nltk-data-hub/names", "female")
names = ds["names"]["name"]
Schema
Column
Type
Description
name
string
The name
Configs
Config
Count
female
5,001
male
2,943
Source
Originally distributed as part of nltk.download('names').
Converted to… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/names.gazetteers
NLTK Gazetteers (Extended)
Geographic and demographic word lists, extended from the original
NLTK gazetteers corpus.
Each config is one list; each row is one entry.
Original configs (NLTK gazetteers corpus)
Config
Description
License
countries
289 country names
GFDL (Wikipedia)
isocountries
234 ISO country names
public domain
nationalities
200 nationality adjectives
public domain
caprovinces
14 Canadian provinces
—
mexstates
32 Mexican states
—… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/gazetteers.TEST.HUB.mini.hub_copora_dataConfig data
{
"chunk_size": 800,
"overlap": 60,
"model": "gemini-1.5-flash-latest"
}
spider_data_goldhub_base_ls960-ft_spot_data_allAI_HUB_legal_QA_datadolch
NLTK Dolch Sight Word List
The 315 Dolch sight words (Dolch 1936), grouped by part of speech, distributed
via NLTK.
Configs
Config
Words
Schema
dolch
315
word, pos
dolch-adjectives
46
word
dolch-nouns
95
word
dolch-verbs
92
word
dolch-adverbs
34
word
dolch-prepositions
16
word
dolch-pronouns
26
word
dolch-conjunctions
6
word
Schema
dolch — combined list with part-of-speech
Column
Type
Description
word
string
The sight… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/dolch.spider_data_bronzeswadesh
NLTK Swadesh Word Lists
Basic vocabulary lists for 24 languages, derived from the
Wiktionary Swadesh list appendix
and distributed via NLTK.
Each config is one language; each row is one of the 207 Swadesh concepts.
Languages
Config
Language
Concepts
swadesh-be
Belarusian
207
swadesh-bg
Bulgarian
207
swadesh-bs
Bosnian
207
swadesh-ca
Catalan
207
swadesh-cs
Czech
207
swadesh-cu
Church Slavonic
174
swadesh-de
German
207
swadesh-en
English
207… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/swadesh.marketing_data_hubveg-data-hubcontext_reasoner_data_mcqbcrp-data-hubanais112-event-data
