datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AmharicCLIP-annotation
AmharicCLIP Annotation Dataset
69,629 images organized by category for Amharic caption annotation.
Structure
images/
animals10/ 18,644 images — 10 animal classes
cat/
dog/
horse/ ...
intel/ 11,998 images — 6 scene classes
forest/
mountain/ ...
fruits360/ 38,987 images — 131 fruit classes
apple/
banana/ ...
Image URL Format… See the full description on the dataset page: https://huggingface.co/datasets/CLIPAMharic/AmharicCLIP-annotation.Amharic_Audio_and_Spectrograms
Amharic Audio Spectrogram Dataset
Dataset Info
Total samples in full dataset: 662,611
Samples in this preview: 1,000
Audio duration: 2.49 ± 1.60 seconds
Sample rate: 16kHz
Spectrogram dimensions: 80 mel bins × variable time steps
Sample Data
Audio Sample
Spectrogram
License
Apache 2.0
amharic-text-recognitionsynthetic_amharic_handwritingHandwritten-Amharic-character-Dataset
Handwritten Amharic Character Dataset
This repository hosts a mirror of the Handwritten Amharic Character Dataset for easier access and use on Hugging Face.
Original Dataset
This dataset was originally prepared by collecting handwriting samples from individuals of different:
age ranges
education levels
right-handed and left-handed writers
A form was used to collect handwriting samples, and an algorithm was developed to crop individual characters from scanned forms.… See the full description on the dataset page: https://huggingface.co/datasets/DevnilMaster1/Handwritten-Amharic-character-Dataset.
