datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
probability_words_nliWords of estimative probability NLI (3 configs). Parquet conversion; CSVs retained.
English-Valid-Words
English Valid Words
This repository contains CSV files with valid English words along with their frequency, stem, and stem valid probability.
Dataset Github link: https://github.com/Maximax67/English-Valid-Words
Files included
valid_words_sorted_alphabetically.csv:
N: Counter for each word entry.
Word: The English word itself.
Frequency count: The number of occurrences of the word in the 1-grams dataset.
Stem: The stem of the word.
Stem valid probability: Probability… See the full description on the dataset page: https://huggingface.co/datasets/Maximax67/English-Valid-Words.words
NLTK Word Lists
English word lists from NLTK,
the New General Service List Project,
and Bing Liu's Opinion Lexicon.
Configs
Config
Words
Schema
License
Source
en
235,886
word
NLTK (other)
NLTK words corpus
en-basic
850
word
Public domain
Ogden Basic English (1930)
ngsl
2,809
word, rank, sfi, freq_per_million
CC-BY-SA 4.0
New General Service List 1.2
toeic
1,250
word, rank, sfi, freq_per_million
CC-BY-SA 4.0
TOEIC Service List 1.2
nawl
963
word, rank… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/words.for-the-small-shield-chapters
Foreword
The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster.
I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct
I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.trending-words-google
Google Trending Words Dataset (2001-2024)
Dataset Description
This dataset contains Google trending words and search terms from 2001 to 2024, capturing 24 years of internet culture, major events, and global trends. The dataset includes 2,784 entries across 93 standardized categories, providing a comprehensive view of what captured the world's attention over more than two decades.
Dataset Summary
Total Entries: 2,784
Years Covered: 2001-2024 (24 years)… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/trending-words-google.WordSpoof
Dataset Card for WordSpoof
A public corpus of standalone single-word bona fide and synthetic speech,
for research on synthetic-voice detection. Most audio-deepfake corpora are
sentence-length; this one is not, and detectors trained on sentence-length
audio degrade sharply on it.
Dataset Details
Dataset Description
WordSpoof pairs real single-word utterances with synthetic ones generated by
10 modern text-to-speech and voice-cloning systems, so a… See the full description on the dataset page: https://huggingface.co/datasets/clearspeed/WordSpoof.pythia-1.4B-tldr-two-words-gpt-4o-iter-1wordscapes-wildlife-animals-stats
Wordscapes Wildlife Animals — Complete Stats
Complete statistics for all 66 collectible Wildlife animals in Wordscapes (PeopleFun), including the June 2026 update that extended the Blue and Green egg groups to Level 10.
Files
wordscapes_wildlife_animals.csv — one row per animal: rarity tier, egg groups, activity duration, activities before sleep, sleep hours, max level, bonus type, coin cost per activity, star animal flag.
wordscapes_wildlife_levels.csv — one row… See the full description on the dataset page: https://huggingface.co/datasets/PopKornmm/wordscapes-wildlife-animals-stats.ghanaian-english-words-corrected-transcriptions
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghanaian English Transcript Corrections
A dataset of mistranscribed words and phrases from Ghanaian news media YouTube videos, corrected using Llama 3.1 405B.
Source
Extracted from YouTube transcripts of Ghanaian… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghanaian-english-words-corrected-transcriptions.pythia-1.4B-tldr-two-words-iter-1rewrite-questions-real-words-sciency
real_words_sciency.csv - Question Rewriting Dataset
This dataset contains question rewriting outputs from the file real_words_sciency.csv.
Dataset Structure
The dataset contains the following columns:
custom_id: Unique identifier for each question
style: Rewriting style applied (e.g., "gibberish")
index: Numerical index
original: Original question text
rewritten: Rewritten version of the question
options: Multiple choice options (list format)
correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-real-words-sciency.leetcode-standalone-wordsamharic_wordsimage-in-Words400
Mouwiya/image-in-words400
Dataset Description
Mouwiya/image-in-words400 is a dataset consisting of 400 images along with their corresponding descriptive captions. The dataset is designed for tasks related to image captioning, where the goal is to generate accurate and contextually relevant descriptions for visual content. This dataset can be used to train and evaluate models that bridge the gap between visual and textual data.
Dataset Details
Total Examples:… See the full description on the dataset page: https://huggingface.co/datasets/Mouwiya/image-in-Words400.japanese-trending-words
Japanese Trending Words Dataset (2006-2025)
Dataset Description
This dataset comprises Japanese trending words from the official annual Japanese Trending Words Awards (流行語大賞) from 2006 to 2025, documenting the cultural, social, and political phenomena that have shaped Japan over the past two decades.
Dataset Summary
Total entries: 593 words
Time period: 2006-2025 (20 years)
Languages: Japanese with English translations
Format: CSV with seven columns: word… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-trending-words.english-words-human-similaritypythia-1.4B-tldr-two-words-gpt-4o-reference-trainturkish-number-words-1m
Turkish Number Words 1M v2
0 ile 999.999 arasındaki her tamsayının Türkçe yazıyla karşılığı.
Doğrulanmış boyut
Train: 980,000
Validation: 10,000
Test: 10,000
Toplam: 1,000,000
Ana görev sütunları: id, number, words
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version, generator_sha256… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-number-words-1m.wikipedia_20220301.simple_sentence_split_text_has_at_least_5_wordspythia-1.4B-tldr-two-words-gpt-4o-reference-valPari_Chekhov_wordsindia-trending-words
Google India Trending Words Dataset (2008-2021, 2023-2024)
Dataset Description
This dataset contains Google trending search terms specific to India from 2008 to 2024 (https://trends.withgoogle.com).
Dataset Summary
Total Entries: 900
Years Covered: 2008-2009, 2011-2021, 2023-2024 (15 years, 2010 and 2022 data not available)
Categories: 18 unique tags
Region: India
Format: CSV
Dataset Structure
Data Fields
word (string): The trending… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/india-trending-words.one-billion-words-testintergrated_gradient_wordsEnglish-Valid-Words
English Valid Words
This repository contains CSV files with valid English words along with their frequency, stem, and stem valid probability.
Dataset Github link: https://github.com/Maximax67/English-Valid-Words
Files included
valid_words_sorted_alphabetically.csv:
N: Counter for each word entry.
Word: The English word itself.
Frequency count: The number of occurrences of the word in the 1-grams dataset.
Stem: The stem of the word.
Stem valid probability: Probability… See the full description on the dataset page: https://huggingface.co/datasets/phillipeds/English-Valid-Words.wikipedia-words-en-lowwikipedia-words-ru-lowD-EVAL__letter_countdown__freq_words__all_easy__qwen_base__evalD-EVAL__letter_countdown__freq_words__all__qwen_base__eval
