datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
simpsons_script_linesOpen-Domain-Oral-Disease-QA-Dataset
Open-Domain-Oral-Disease-QA-Dataset
Dataset Details
Dataset Description
This dataset is meticulously designed to evaluate the diagnostic capabilities of Large Language Models (LLMs) in the domain of oral disease.
We currently offer a suite of evaluation datasets encompassing models such as GPT-3.5, GPT-4, Palm2, and Llama2-70B. More data is under reviewed. This dataset is meticulously designed to evaluate the diagnostic capabilities of Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/Lines/Open-Domain-Oral-Disease-QA-Dataset.glados_ru_linesRussian glados lines with links to audio from https://i1.theportalwiki.net/
parsed from https://theportalwiki.com/wiki/GLaDOS_voice_lines/ru
nepali-synthetic-ocr-lines
Nepali Synthetic OCR/HTR Document Line Dataset
A dataset of synthetic Devanagari text line images imitating historical and official scanned document conditions, designed for OCR (Optical Character Recognition) and HTR (Handwritten Text Recognition) models such as TrOCR, CRNN, and PaddleOCR.
This dataset was generated using the Mountmind PeakOCR Studio synthetic corpus generator pipeline, introducing realistic document aging artifacts like:
Skew Angle Rotations (Hough line… See the full description on the dataset page: https://huggingface.co/datasets/prashant0919/nepali-synthetic-ocr-lines.ada-linesnetlist-snippets-80-linesnetlist-snippets-20-lineslines_hu_v5netlist-snippets-40-linesClassic-lines-from-the-movie-Nezhalabeled_lines_datalabeled_lines2_1m
