datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
target-morphology
target-morphology
Per-language unsupervised morphology models — productive suffixes, prefixes, and a stem lexicon,
each learned MDL-free ("Linguistica"-style: a suffix is productive if it attaches to many paradigm stems)
from that language's own Bible text. No labels, no pretrained model, no download — so it runs on any
language with a translation, including those with zero LLM/encoder coverage.
stem(word) strips one productive affix when the remainder is a known stem; inflected… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/target-morphology.mcs-morphology-frames
MCS morphology — labeling frames
Two-panel radar plots (SL3D classification mask beside composite reflectivity) for
tracked mesoscale convective systems over the contiguous United States, packaged to be
served to the labeling platform at
https://huggingface.co/spaces/skyan1002/mcs-morphology-labeler.
Layout: {year}/track_{id}/{frame:03d}.png, where the frame number matches the index in
the original mcstrack_YYYYMMDD_HHMMSS_N.png file name and the time index of the
corresponding… See the full description on the dataset page: https://huggingface.co/datasets/skyan1002/mcs-morphology-frames.block_polymers_morphology
Dataset Details
Dataset Description
Results of experimental phase measurements of di-block copolymers.
Curated by:
License: CC BY 4.0
Dataset Sources
original data source
Citation
BibTeX:
No citations provided
morphology4metrology-bnf2813
Dataset for BnF, fr. 2813 — Grandes Chroniques de France
Line-level dataset used in the experiments for the paper Leveraging Morphology for Historical Script Metrological Analysis (ICDAR 2026).
Annotator: Malamatenia Vlachou Efstathiou
Description
This dataset contains polygonal line extractions (with alpha transparency) and line-level transcriptions from the codex Paris, BnF, fr. 2813 (Grandes Chroniques de France).
The identifier btv1b84472995 refers to the ark… See the full description on the dataset page: https://huggingface.co/datasets/RaphaelBfr/morphology4metrology-bnf2813.sga2020_morphologyMorphology_Predictiongalaxy-zoo-2-morphology
Galaxy Zoo 2 Morphological Classifications
Part of the Astronomy Datasets collection on Hugging Face.
243,500 citizen-science galaxy morphology classifications from Galaxy Zoo 2,
the largest visual morphological classification project in astronomy. Each galaxy was
classified by multiple volunteers answering a decision tree of questions about shape,
structure, and features.
Dataset description
Galaxy Zoo 2 asked hundreds of thousands of volunteers to classify galaxy… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/galaxy-zoo-2-morphology.algo-sft-eval-traces-conlang-morphology-ordered-rules-d5d7-v4
algo-sft-eval-traces-conlang-morphology-ordered-rules-d5d7-v4
Full eval traces for algo-sft-conlang-morphology-ordered-rules-d5d7 across test/harder/ood splits
Dataset Info
Rows: 2000
Columns: 11
Columns
Column
Type
Description
question_id
Value('string')
Unique question identifier from eval set
split
Value('string')
Evaluation split: test (in-distribution), harder (scaled up), ood (structural out-of-distribution)
domain
Value('string')
Task… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algo-sft-eval-traces-conlang-morphology-ordered-rules-d5d7-v4.euclid-morphology-catalogmalayalam-morphology-analyser
Malayalam Morphology Analyser dataset
This is a dataset of 801585 Malayalam words and their morphology analysis using Mlmorph Malayalam morphology analyser.
License
Creative Commons Attribution Share Alike 4.0
Contact
Contact santhosh.thottingal @ gmail.com
algo-sft-eval-traces-conlang-morphology-distill-qwq-v4
algo-sft-eval-traces-conlang-morphology-distill-qwq-v4
Full eval traces for algo-sft-conlang-morphology-distill-qwq across test/harder/ood splits
Dataset Info
Rows: 2000
Columns: 11
Columns
Column
Type
Description
question_id
Value('string')
Unique question identifier from eval set
split
Value('string')
Evaluation split: test (in-distribution), harder (scaled up), ood (structural out-of-distribution)
domain
Value('string')
Task domain:… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algo-sft-eval-traces-conlang-morphology-distill-qwq-v4.sanskrit-morphology-rlrc3-galaxy-morphology
Third Reference Catalogue of Bright Galaxies (RC3)
The Third Reference Catalogue of Bright Galaxies (RC3), the classic comprehensive catalog of
23,011 bright galaxies with Hubble-type morphological classifications, photometry,
diameters, and radial velocities.
Dataset description
RC3 is the definitive catalog of bright galaxies, compiled by de Vaucouleurs, de Vaucouleurs,
Corwin, Buta, Paturel, and Fouque (1991). It provides homogeneous morphological classifications
on… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/rc3-galaxy-morphology.galaxy-morphology-classificationMorphology_Function_Framework
Morphology Function Framework
This repository links whole-slide-image (WSI) morphology to molecular function and patient survival. It is organized as five sub-projects that run in sequence (with one branch running in parallel), each documented independently with its own README.md and USAGE.md. This top-level document explains how the five sub-projects fit together, what data flows between them, and where to find detailed instructions.
Two branches converging on one… See the full description on the dataset page: https://huggingface.co/datasets/CarmeloAnthony/Morphology_Function_Framework.quranic-corpus-morphology
Qur'an Phoneme + Harakat Dataset
This dataset contains phoneme-level and diacritic-level representations of the Qur'an text, based on the Quranic Arabic Corpus transliteration. It is intended for use in speech recognition and text-to-speech models, especially for phoneme-level fine-tuning of models like Whisper.
Dataset Structure
Column
Description
LOCATION
Quranic verse location in the format (Sura:Ayah:Word:Subword)
FORM
Original transliterated word form… See the full description on the dataset page: https://huggingface.co/datasets/mrmuminov/quranic-corpus-morphology.tatar-morphology-benchmark
Tatar Morphology Benchmark
This repository contains evaluation results for morphological analysis models trained on the Tatar Morphological Corpus.
Models Evaluated
mBERT
RuBERT
DistilBERT
LSTM
Turkish BERT
XLM-R
Key Results (Test Set Accuracy)
Model
Accuracy
F1 (micro)
mBERT
0.9905
0.9905
RuBERT
0.9861
0.9861
DistilBERT
0.9850
0.9850
XLM-R
0.9837
0.9837
LSTM
0.9440
0.9440
Turkish BERT
0.8769
0.8769
All results are based on a test… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-morphology-benchmark.quranic-corpus-morphology-by-ayahTil-Morphology
Til-Morphology
Қазақ тілінің морфологиялық деректері · Морфологические данные казахского языка · Kazakh morphological data
Қазақша · Русский · English
Қазақша
Til-Morphology — қазақ тіліндегі 3 767 518 морфологиялық талдау жолы жинақталған, көлемі 207.8 МБ датасет. Жолдар модель-сарапшы бағасына қарай premium және clean деңгейлеріне бөлінген.
Құрамы
Әр жазбада word, segmentation, n_morphemes, lang, category, score және text өрістері бар.… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/Til-Morphology.esperanto-sft-morphology-iclsanskrit-morphologyThis dataset contains 200,000 Sanskrit verb morphology specification prompts focused on tinanta forms (for now). It is designed to train reinforcement learning agents for generating morphologically correct Sanskrit words, grounded in the systematic rules of Pāṇini’s Aṣṭādhyāyī.
morphology-exampleskaz-morphology-sample
kaz-morphology-sample
Қазақ сөздерінің морфологиялық үлгісі · Образец морфологической разметки казахских слов · Kazakh word morphology sample
Қазақша · Русский · English
Қазақша
kaz-morphology-sample — қазақ сөздерінің морфологиялық талдауы бар 9.3 МБ деректер жинағы. Жинақта 10 000 жазба және 7 сөз табы қамтылған; оны морфологиялық талдау, леммалау және POS-tagging үлгілерін тексеруге қолдануға болады.
Деректер құрамы
Файл
Жазба… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/kaz-morphology-sample.DCS_Sanskrit_Morphology_v1This dataset contains a dcs_output.csv file at the root.
spike_morphology_neural_decision_dataset
Comprehensive Electrophysiological and Feature-Space Dataset
Dataset Overview
This dataset provides a multi-modal collection of features derived from single-unit electrophysiology recordings, aiming to characterize neuronal units based on their morphological, temporal, and spectral properties, as well as their representation in a latent embedding space and their classification via a neuro-symbolic model. The dataset is structured to facilitate advanced machine… See the full description on the dataset page: https://huggingface.co/datasets/naman-00/spike_morphology_neural_decision_dataset.canine-facial-morphology
