datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clinical-healing-trajectory-tokenization-phase-segmentation-v0.1What this dataset tests
Whether a model can segment high-frequency recovery datainto interpretable healing phases.
Required outputs
phase_sequence
phase_boundaries
phase_confidence_0_100
Token labels
acute_drop
early_rebound
consolidation_plateau
oscillatory_instability
secondary_drop
delayed_rebound
steady_ascent
maladaptive_plateau
recovery_lock_in
Boundary format
Use day indicesexampleacute_drop d0-d2
Typical failures
naming phases without boundaries… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-healing-trajectory-tokenization-phase-segmentation-v0.1.test_tokenizationbengali-tokenization-corpus
Bengali Tokenization Corpus (25k Sentences)
Dataset Description
A balanced 25,000-sentence Bengali corpus designed for tokenization benchmarking.
Dataset Summary
This dataset is used in the manuscript:
Quantifying the Tokenization Tax on Bengali: A Multi-Tokenizer Audit Across 25k Sentences
MD. Muktadirul Haque Maruf, Israt Jahan Munny, Moutithi Sarker MouManuscript in preparation
Domains
Academic
News
Literary
Colloquial
Dialectal… See the full description on the dataset page: https://huggingface.co/datasets/M-H-MARUF/bengali-tokenization-corpus.
