datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
german-spelling-dataspelling-bee-pangrams
Spelling Bee-style letter sets and their pangrams
Each row is one puzzle: 7 distinct letters (letters, alphabetical) and the list
of pangrams — words using all 7 letters. Nothing else.
Puzzles come from the dwyl/english-words words_alpha.txt word list (~370k
entries). Included: every 7-letter combination with at least one pangram and at
least 10 valid answers (words of 4+ letters using only the 7 letters).
Source: https://github.com/dwyl/english-words (words_alpha.txt)
spellingcorrectionFrenchreflection-spelling-puzzles-sharegptvietnamese-spelling-synthetic-1gb
Vietnamese Spelling Correction Synthetic 1GB
Corpus tổng hợp cho bài toán phát hiện và sửa lỗi chính tả tiếng Việt, trích xuất và làm sạch từ Vietnamese Wikipedia Dump mới nhất, hỗ trợ mô hình phân loại token 1-đối-1.
📊 Nội dung Dataset
File
Số lượng mẫu (Rows)
Dung lượng thô
Mô tả
train_full.jsonl
4.564.360
6.04 GB
Tập huấn luyện chính thức
validation_full.jsonl
95.094
124 MB
Tập validation held-out (tách theo page_id)
word_vocab_full.json
9.500… See the full description on the dataset page: https://huggingface.co/datasets/Sanng1112/vietnamese-spelling-synthetic-1gb.spelling
Some data on spelling, because small LLMs are bad at it
Sure LLMs are permenantly doomed to spell things correctly most of the time, but that doesn't mean they know the spelling of these words.
This data aims to resolve this issue.
Note: Intended for small LLMs that already have data to be trained on.
asr_spellingCorrection_24k
