sakha
Datasets
All datasets matching “sakha”AviationQAAviationQA is introduced in the paper titled- There is No Big Brother or Small Brother: Knowledge Infusion in Language Models for Link Prediction and Question Answering
https://aclanthology.org/2022.icon-main.26/
The paper is accepted in the main conference of ICON 2022.
We create a synthetic dataset, AviationQA, a set of 1 million factoid QA pairs from 12,000 National Transportation Safety Board (NTSB) reports using templates. These QA pairs contain questions such that answers to them are… See the full description on the dataset page: https://huggingface.co/datasets/sakharamg/AviationQA.Bangladesh-Legal-Acts-Dataset
Bangladesh Legal Acts Dataset
A comprehensive database of Bangladesh's legal framework, containing 1484+ acts scraped and processed from the official Bangladesh Laws portal, enhanced with historical government context, legal system context, and comprehensive metadata.
Dataset Overview
Total Acts: 1,484
Total Sections: 35,633
Total Footnotes: 14,523
Languages: English, Bengali, Mixed
Format: JSON with structured metadata
Historical Context: Government periods from… See the full description on the dataset page: https://huggingface.co/datasets/sakhadib/Bangladesh-Legal-Acts-Dataset.AviationCorpussakha-asrsakha-ocr-synth
Синтетические строки якутского текста для OCR
500 000 изображений строк с точной разметкой. Сделано для обучения
распознавателя якутского (саха) текста: готовые движки для этого языка не
работают, а размеченных строк почти нет.
Зачем вообще синтетика: у tesseract-rus и ABBYY FineReader 10 на якутской
печати доля правильно прочитанных специфических букв ҕ ҥ ө һ ү равна нулю.
Не «низкая» — ноль на 460 тысячах букв, при том что эти буквы составляют около
7% всех букв и встречаются… See the full description on the dataset page: https://huggingface.co/datasets/lab-ii/sakha-ocr-synth.AeroQA
