datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
braille_dataset_2braille_data3Chinese-Braille-Dataset-Full-Tone
Chinese Braille Sentence Corpus (Full Tone)
📃 [Paper] •
💻 [Code] •
📖 [Passage corpus] •
🎬 [Demo]
The sentence-level half of the Braille–Chinese parallel corpus used in
"Vision-Braille: A Curriculum Learning Toolkit and Braille–Chinese Corpus for Braille
Translation" (EMNLP 2026 Main Conference).
Every Braille sequence here retains all tone markers (retention rate r = 100). This is
the source corpus: the tone-omission variants used for curriculum training are… See the full description on the dataset page: https://huggingface.co/datasets/Violet-yo/Chinese-Braille-Dataset-Full-Tone.vision-braille-dataset
vision-braille-dataset
Chinese ↔ 通用盲文 (Chinese Braille) parallel training data for the Vision-Braille translation
models. Three corpora ship together in this repo:
Folder
What it is
Rows
Size
cleaned_2345_v3/
Audited + cleaned braille→Chinese training corpus, stages 2→5, with restored 分词连写 word spacing on stage 2
379,152
426 MB
braille_spaced/
Freshly built passage-level corpus, fully word-spaced, across 通用 / 数学 / 医学 / 中医 / 病理学
80,360
208 MB
cleaned_080126_v1/… See the full description on the dataset page: https://huggingface.co/datasets/Violet-yo/vision-braille-dataset.Chinese-Braille-Dataset-10per-Tone
Chinese Braille Dataset
📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • ⚙️ [Model] • 🎬 [Demo]
Dataset Description
The Chinese-Braille-10per-Tone dataset addresses the scarcity of publicly available Chinese Braille datasets. The original Chinese text data was sourced from the publicly available Leipzig Corpora Collection. This dataset consists of one million discrete sentences collected from news media between 2007 and 2009.
The Chinese characters from the Leipzig… See the full description on the dataset page: https://huggingface.co/datasets/Violet-yo/Chinese-Braille-Dataset-10per-Tone.braille_dataset_4Passage-Chinese-Braille-Dataset-Full-Tone
Chinese Braille Passage Corpus (Full Tone)
📃 [Paper] •
💻 [Code] •
📖 [Sentence corpus] •
🎬 [Demo]
The passage-level half of the Braille–Chinese parallel corpus used in
"Vision-Braille: A Curriculum Learning Toolkit and Braille–Chinese Corpus for Braille
Translation" (EMNLP 2026 Main Conference).
Passages are roughly 8× longer than the sentences in the
companion corpus,
which makes them a substantially harder translation task: the model must hold tone-ambiguous… See the full description on the dataset page: https://huggingface.co/datasets/Violet-yo/Passage-Chinese-Braille-Dataset-Full-Tone.Chinese-Braille-Dataset-No-Tone
Chinese Braille Dataset (No Tone)
📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • ⚙️ [Model] • 🎬 [Demo]
This dataset is the Chinese-Braille-Dataset-No-Tone dataset described in https://huggingface.co/datasets/Violet-yo/Chinese-Braille-Dataset-10per-Tone.
Dataset Statistics
# Sample
Braille Len. (Mean/Median) String
Braille Len. (Mean/Median) Token
Chinese Len. (Mean/Median) String
Chinese Len. (Mean/Median) Token
Training
525072
140/108
144/112
74/64
59/51… See the full description on the dataset page: https://huggingface.co/datasets/Violet-yo/Chinese-Braille-Dataset-No-Tone.braille_translator_1speech-nemeth_braillebraillechinese_braille_3k_rowstest_braillebraille_2braille_dataset_4braille-reader-resultsBrailleDataset1braillePhrase-Chinese-Braille-Dataset-Full-Tonebraillemini
