datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AviationQAAviationQA is introduced in the paper titled- There is No Big Brother or Small Brother: Knowledge Infusion in Language Models for Link Prediction and Question Answering
https://aclanthology.org/2022.icon-main.26/
The paper is accepted in the main conference of ICON 2022.
We create a synthetic dataset, AviationQA, a set of 1 million factoid QA pairs from 12,000 National Transportation Safety Board (NTSB) reports using templates. These QA pairs contain questions such that answers to them are… See the full description on the dataset page: https://huggingface.co/datasets/sakharamg/AviationQA.AviationCorpussakha-ocr-synth
Синтетические строки якутского текста для OCR
500 000 изображений строк с точной разметкой. Сделано для обучения
распознавателя якутского (саха) текста: готовые движки для этого языка не
работают, а размеченных строк почти нет.
Зачем вообще синтетика: у tesseract-rus и ABBYY FineReader 10 на якутской
печати доля правильно прочитанных специфических букв ҕ ҥ ө һ ү равна нулю.
Не «низкая» — ноль на 460 тысячах букв, при том что эти буквы составляют около
7% всех букв и встречаются… See the full description on the dataset page: https://huggingface.co/datasets/lab-ii/sakha-ocr-synth.AeroQAsakha-fineweb-2sakha-oscarsakha-russian-parallel-corpora-all
Параллельный корпус с якутского на русский
Данный датасет представляет собой параллельный корпус предложений на якутском (саха) и русском языках.
Предназначен для исследований в области машинного перевода, лингвистики и обработки языков с ограниченными ресурсами.
Структура данных
Поле
Язык
Описание
sah
Якутский
Предложение на якутском (саха) языке
ru
Русский
Соответствующий перевод на русском языке
Общее количество строк: 13 323
Тип данных: строки… See the full description on the dataset page: https://huggingface.co/datasets/averoo/sakha-russian-parallel-corpora-all.sakha-corpus-monoThe texts with the natlib column were created using OCR and may contain errors and artifacts.
Please take this into account when using the data for training or evaluation.
Описание атрибутов источников данных
natlib
Тексты, предоставленные Национальной библиотекой Республики Саха (Якутия). Большая часть худ. литература:
Журнал «Күрүлгэн» — художественная литература
Журнал «Хотугу сулус» — художественная литература
Журнал «Чолбон» — художественная литература
Журнал… See the full description on the dataset page: https://huggingface.co/datasets/ailabykt/sakha-corpus-mono.Africa-skin-cancer-images-EHR
Africanized Skin Cancer Multimodal Dataset
Dataset Description
This dataset extends the HAM10000 dermatoscopic image collection with:
Provisional skin-tone annotations (Fitzpatrick scale estimates)
Synthetic African-focused Electronic Health Records (EHR)
Multimodal patient demographics, comorbidities, medications, and lab values
Created by: kossiso RoyceDate: October 2025Version: 1.0
Objective
Create a more representative dataset for training and… See the full description on the dataset page: https://huggingface.co/datasets/Sakhawat077/Africa-skin-cancer-images-EHR.sakha-madladsakha-russian-parallelsakha-russian-parallelThe texts in the sah column were generated using OCR and may contain errors or artifacts. Please take this into account when using the data for training or evaluation.
The dataset was aligned using the Lingtrain Aligner library (https://github.com/averkij/lingtrain-aligner), created by @averoo
sakha-wmtsakhan10__quantized_open_llama_3b_v2-details
Dataset Card for Evaluation run of sakhan10/quantized_open_llama_3b_v2
Dataset automatically created during the evaluation run of model sakhan10/quantized_open_llama_3b_v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sakhan10__quantized_open_llama_3b_v2-details.sakha_chat_ml-sftsakha-codegenmod-hw1-cinemasgenmod-hw1-cinemas-pairs
