datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
heb-words27-spaced
heb-words27-diffpen
Word-level Hebrew handwriting for DiffusionPen finetuning -- the step up from
cyttic/heb-connected-bigrams17-diffpen, which covered 2-character units. Here the unit
is a whole word.
Content
20,000 most frequent words of LLMGen2 (sentences_llm2.txt, 968,204 LLM-generated
Hebrew sentences, 167,879 unique words)
643 numbers -- LLMGen2 contains no digits at all, because the generation
prompt forbade them, so numerals had to be added separately or… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/heb-words27-spaced.malayalam-ocr-words
Malayalam OCR Words
A word-level Malayalam OCR dataset: cropped word images paired with their
transcribed text label, split into train/validation/test sets.
Dataset structure
train.csv / val.csv / test.csv # tab-separated: <relative image path>\t<Malayalam word>
train/ val/ test/ # image files referenced by the corresponding CSV
Each CSV row maps one image file (path relative to its split folder) to its
ground-truth Malayalam word transcription.… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/malayalam-ocr-words.Chinese-Char-Wordsgregg-anniversary-words
Gregg Anniversary Words
This dataset is derived from the 1930 Gregg Shorthand Dictionary1.
Structure
The dataset contains three columns:
image: The image of a shorthand form in the dictionary
grascii_normalized: The normalized grascii of the shorthand form
longhand: The English longhand represented by the shorthand form
Issues
If you notice any problems in the dataset, open an issue in the
datasets repository.
1Gregg, John Robert. Gregg Shorthand Dictionary.… See the full description on the dataset page: https://huggingface.co/datasets/grascii/gregg-anniversary-words.noisy-gt-missing-words
Noisy Ground Truth - Missing Words
Dataset of synthetic data for experimentation with noisy ground truth. The text in the dataset is based on Colette's Sido and Les Vignes, also the data was processed prior to generating images with the TextRecognitionDataGenerator.
In Noisy Ground Truth - Missing Words, each variation column is affected by the noise, without considering the split between train, validation and test.
Data structure
The dataset is composed of the… See the full description on the dataset page: https://huggingface.co/datasets/alix-tz/noisy-gt-missing-words.IIIT-INDIC-HW-WORDS-Hindi
IIIT-INDIC-HW-WORDS-Hindi
Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images.
Overview
The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Hindi words and aims to advance research and… See the full description on the dataset page: https://huggingface.co/datasets/c3rl/IIIT-INDIC-HW-WORDS-Hindi.indonesian_words_imageheb-words27-diffpen
heb-words27-diffpen
Word-level Hebrew handwriting for DiffusionPen finetuning -- the step up from
cyttic/heb-connected-bigrams17-diffpen, which covered 2-character units. Here the unit
is a whole word.
Content
20,000 most frequent words of LLMGen2 (sentences_llm2.txt, 968,204 LLM-generated
Hebrew sentences, 167,879 unique words)
643 numbers -- LLMGen2 contains no digits at all, because the generation
prompt forbade them, so numerals had to be added separately or… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/heb-words27-diffpen.noisy-gt-missing-words-train-only
Noisy Ground Truth - Missing Words in Train Split only
Dataset of synthetic data for experimentation with noisy ground truth. The text in the dataset is based on Colette's Sido and Les Vignes, also the data was processed prior to generating images with the TextRecognitionDataGenerator.
In Noisy Ground Truth - Missing Words in Train Split only, each variation column is affected by the noise, only when the split is for training. The validation and test splits are not affected by… See the full description on the dataset page: https://huggingface.co/datasets/alix-tz/noisy-gt-missing-words-train-only.IAM_words_text_recognitioniam-words-no_symbolsnoisy-gt-xxx-words
Noisy Ground Truth - Words Replaced with XXX
Dataset of synthetic data for experimentation with noisy ground truth. The text in the dataset is based on Colette's Sido and Les Vignes, also the data was processed prior to generating images with the TextRecognitionDataGenerator.
In Noisy Ground Truth - Words Replaced with XXX, each variation column is affected by the noise, without considering the split between train, validation and test.
Data structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/alix-tz/noisy-gt-xxx-words.IIIT-INDIC-HW-WORDS-Tamil
IIIT-INDIC-HW-WORDS-Tamil
Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images.
Overview
The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Tamil words and aims to advance research and… See the full description on the dataset page: https://huggingface.co/datasets/c3rl/IIIT-INDIC-HW-WORDS-Tamil.loc_beyond_words
Dataset Card for Beyond Words
Dataset Summary
The Beyond Words dataset is a crowdsourced collection of bounding box annotations on World War I-era historical newspaper pages from the Library of Congress’s Chronicling America collection. Volunteers marked seven types of visual content — photographs, illustrations, maps, comics, editorial cartoons, headlines, and advertisements — enabling the training of the visual content recognition model behind the Newspaper Navigator… See the full description on the dataset page: https://huggingface.co/datasets/biglam/loc_beyond_words.sora-video-generation-aligned-words
Rapidata Video Generation Word for Word Alignment Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~1500 human evaluators were asked to evaluate AI-generated videos based on what part of the prompt did not align the video. The specific… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/sora-video-generation-aligned-words.Amount_in_Arabic_words_dztrdg_random_single_words_en_text_recognition
Dataset Card for "trdg_random_single_words_en_text_recognition"
More Information needed
noisy-gt-xxx-words-train-only
Noisy Ground Truth - Words Replaced with XXX in Train Split only
Dataset of synthetic data for experimentation with noisy ground truth. The text in the dataset is based on Colette's Sido and Les Vignes, also the data was processed prior to generating images with the TextRecognitionDataGenerator.
In Noisy Ground Truth - Words Replaced with XXX in Train Split only, each variation column is affected by the noise, without considering the split between train, validation and test.… See the full description on the dataset page: https://huggingface.co/datasets/alix-tz/noisy-gt-xxx-words-train-only.Medical_Prescription_Handwritten_Words
Medical Prescription Handwritten Words
This dataset contains images of individual handwritten medical words extracted from prescription notes. It is designed for training and evaluating handwriting recognition models in the healthcare domain.
Structure
images/: Contains 40+ handwritten word images (e.g., Amoxicillin.png, Cold.png, Tablet.png, 0.png, etc.)
data.csv: Maps each image file to its corresponding label (word)
Example Use Cases
OCR (Optical… See the full description on the dataset page: https://huggingface.co/datasets/avi-kai/Medical_Prescription_Handwritten_Words.trdg_dict_random_words_en_text_recognition
Dataset Card for "trdg_random_words_en_text_recognition"
More Information needed
Medical_Prescription_Handwritten_Words
Medical Prescription Handwritten Words
This dataset contains images of individual handwritten medical words extracted from prescription notes. It is designed for training and evaluating handwriting recognition models in the healthcare domain.
Structure
images/: Contains 40+ handwritten word images (e.g., Amoxicillin.png, Cold.png, Tablet.png, 0.png, etc.)
data.csv: Maps each image file to its corresponding label (word)
Example Use Cases
OCR… See the full description on the dataset page: https://huggingface.co/datasets/MMMuzammil/Medical_Prescription_Handwritten_Words.message-decoding-words-and-sequences-r1loc_beyond_words_cocoimage-in-Words400
Mouwiya/image-in-words400
Dataset Description
Mouwiya/image-in-words400 is a dataset consisting of 400 images along with their corresponding descriptive captions. The dataset is designed for tasks related to image captioning, where the goal is to generate accurate and contextually relevant descriptions for visual content. This dataset can be used to train and evaluate models that bridge the gap between visual and textual data.
Dataset Details
Total Examples:… See the full description on the dataset page: https://huggingface.co/datasets/Mouwiya/image-in-Words400.IIIT-INDIC-HW-WORDS-Hindi
IIIT-INDIC-HW-WORDS-Hindi
Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images.
Overview
The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Hindi words and aims to advance research… See the full description on the dataset page: https://huggingface.co/datasets/HarishBonu/IIIT-INDIC-HW-WORDS-Hindi.gregg-preanniversary-words
Gregg Preanniversary Words
This dataset is derived from the 1916 Gregg Shorthand Dictionary1
and builds on top of the Gregg1916
dataset by:
Correcting the labels of 250+ images
Adding 550+ images for words missing in the original dataset
Redoing 100+ poorly cropped images or images with stray marks
Structure
The dataset contains three columns:
image: The image of a shorthand form in the dictionary
grascii_normalized: The normalized grascii of the shorthand form… See the full description on the dataset page: https://huggingface.co/datasets/grascii/gregg-preanniversary-words.words_hu_dictThis dataset was generated from a Hungarian dictionary, where 60345 sample given
The command used to generate data :
python3 run.py -i "dicts/hu.txt" -t 8 -f 64 -l hu -c 60345 -na 2 --output_dir "out/words/hu/" --font_dir fonts/hu/ -b 3 -al 0
TRDGHuMu is used for generating text: https://github.com/Mohammed20201991/TextRecognitionDataGeneratorHuMu23
message-decoding-words-and-sequences-target-zoom-in-r1message-decoding-words-and-sequences-target-zoom-invnondb_sentence_by_words
