datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Drone-Orthomosaic-Vehicles-Yolo-annotation
Dataset Tailings Mining Vehicles & Instruments (High-Res Drone Imagery)
Dataset Summary
This dataset contains high-resolution aerial imagery focused on vehicle detection and geotechnical monitoring instruments within active mining environments (tailings dams). The data was acquired using a DJI Zenmuse P1 sensor at 120m altitude.
Photogrammetric Context
The images originate from large-scale georeferenced orthomosaics generated from bi-daily… See the full description on the dataset page: https://huggingface.co/datasets/titoruizh/Drone-Orthomosaic-Vehicles-Yolo-annotation.russian-old-orthography-ocr
Basic Description
Dataset contains source images and human-readable extracted texts. All texts were published in Russia in the 19th century and written using pre-reform orthography.
The dataset is designed to train and evaluate optical character recognition systems for texts published in Russian before the orthographic reform (1917).
Data structure
For each text there is a file with its image and the text corresponding to this image. The names of these files are the same… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/russian-old-orthography-ocr.OrthoTryOn-Instructions
OrthoTryOn: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation
Model Introduction
We introduce OrthoTryOn, a unified and parameter-efficient framework for fashion image generation, designed to mitigate inter-task
interference in shared adaptation and enable high-quality virtual try-on, garment reconstruction, and pose transfer within a single model.
Its plug-and-play design can further extend to broader multi-task scenarios.… See the full description on the dataset page: https://huggingface.co/datasets/Jerome-Young/OrthoTryOn-Instructions.ortho-target-datasetorthonogilizereformattedqwen3-orthdion-sweepOrthohantavirus-Genome-Atlas
HantaBERT Data Pipeline
This repository is responsible for the entire process of collecting, cleaning, and standardizing Orthohantavirus genomic data for the HantaBERT project. The pipeline automates data extraction from NCBI GenBank to produce a ready-to-use dataset for machine learning.
Key Features
Extraction Automation: Uses Biopython to fetch thousands of RNA sequences (S, M, L) and related metadata in batches from the NCBI database.
Multi-task Labeling:… See the full description on the dataset page: https://huggingface.co/datasets/HantaBERT/Orthohantavirus-Genome-Atlas.orthogonal-activation-steering-TOXICorthodox-patristic-corpus
Orthodox Patristic Corpus
Released on the Feast of the Triumph of Orthodoxy, First Sunday of Great Lent, 2026.
Dataset Summary
The Orthodox Patristic Corpus is a 116M-token pre-training corpus of Orthodox Christian
theological literature, assembled from the writings of 123 Church Fathers and Orthodox theologians
spanning the 1st through 20th centuries. The corpus is primarily in Russian, drawing on the
Azbyka.ru Orthodox digital library and other public-domain sources… See the full description on the dataset page: https://huggingface.co/datasets/jayfurzy/orthodox-patristic-corpus.bambara-orthography
Bambara orthography and text normalization
Writing conventions for Bambara (Bamanankan) in the standard Latin alphabet, a normalization table from common ASCII spellings to the standard, and a reference normalizer in Python. Maintained by Kooma.
Why this exists: Bambara is written in many ways in the wild — with or without ɛ/ɔ/ɲ/ŋ, with ny/ng digraphs, with or without tone marks, with French spellings for loanwords. Any comparison between two Bambara texts (a transcription and… See the full description on the dataset page: https://huggingface.co/datasets/kooma-ai/bambara-orthography.orthographic-views-vision
Orthographic Projection Views — Vision
Fine-tuning data for a small vision-language model that, given a rendered
image of a 3D block object, outputs its third-angle orthographic projections
(top / front / right) as three ASCII grids. Built to instill one narrow,
reliable behavior via QLoRA on a small open VLM (Qwen2.5-VL-3B class).
Behavior Spec (the litmus test)
Given an image of a solid built from unit cubes, output exactly three
ASCII grids labeled top:… See the full description on the dataset page: https://huggingface.co/datasets/hiyasvyas/orthographic-views-vision.myX-myanmar-orthography-corpus
📝 myX-myanmar-orthography-corpus: Myanmar Orthography Error Correction Dataset
The myX-Myanmar-Orthography-Corpus is an open-source initiative dedicated to improving the accuracy of Myanmar (Burmese) language digital processing. This project focuses specifically on Orthographic Accuracy (သတ်ပုံ), addressing common spelling errors, phonetic confusions, and keyboard typos.
This project focuses specifically on orthographic errors such as phonetic confusions, visual similarities, and… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-myanmar-orthography-corpus.Wolof-Non-Standard-Orthography
Dataset Description
Dataset Summary
This dataset contains pairs of non-standard and standard Wolof text, designed for training models to normalize informal Wolof writing found on social media, messaging apps, and online platforms.
The non-standard versions simulate real-world informal Wolof text with French code-switching, phonetic spellings, missing diacritics, and common typing variations.
The original Standard Wolof and English sentences are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Wolof-Non-Standard-Orthography.desacordo_ortografico
Portuguese Orthographies — Parallel Corpus
A parallel corpus for detecting and converting between Portuguese orthographies. Each
record is one Portuguese sentence written in five orthographic norms, so the same content
can be aligned across the spelling reforms of the language.
Norms (one column each)
column
norm
etymological
pre-1911 pseudo-etymological spelling
pt_1973
pre-AO1990 European (Convenção 1945 + 1973 mini-reform)
ao1990_pt
Acordo… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/desacordo_ortografico.orthographic-views-vision-final-v4
Orthographic Projection Views — Vision FINAL v4
This is the final v4 training + eval dataset for the Ortho-LLM project.
Fine-tuning data for a small vision-language model that, given a rendered
image of a 3D block object, outputs dims: ZxYxX plus its third-angle
orthographic projections (top / front / right) as ASCII grids.
Companion model: hiyasvyas/ortho-vision-qlora-final-v4
Behavior Spec (the litmus test)
Given an image of a solid built from unit cubes (unit… See the full description on the dataset page: https://huggingface.co/datasets/hiyasvyas/orthographic-views-vision-final-v4.script__orthography
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/script__orthography.CNC_oral_ortofon
Introduction
This is a sample from the ORAL2013 and ORTOFON datasets, maintained by the Czech National Corpus project. The datasets were created from shared .vert file format using the convert_ORTOFON_ORAL13.py script.
The versions of the datasets used here were downloaded from the LINDAT Clarin repository:
ORTOFON v1
ORAL2013
About Original Datasets
ORAL2013
The ORAL2013 corpus is spoken corpus available within the framework of the Czech National Corpus… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_oral_ortofon.OR_training_fullorthoqa-300
OrthoQA-300
You can access the dataset on Hugging Face and find the full generation pipeline, configuration files, and source code in the Stratum Research GitHub repository.
OrthoQA-300 is a structured, synthetic dataset of 300 patient-provider style question-and-answer (QA) pairs focused on orthopedic surgery. Each entry simulates a realistic clinical interaction, with patient-style questions and LLM-generated provider-style answers.
Questions are grouped by procedure (e.g., ACL… See the full description on the dataset page: https://huggingface.co/datasets/stratum-research/orthoqa-300.structured-base-pure-v6acordo-ortografico-lexicon
Acordo Ortográfico de 1990 — Word-Change Lexicon
6568 Portuguese words documented across the 1990 orthographic reform: the pre-1990
European (1945) and Brazilian (1943) spellings, the AO1990 valid form(s), and whether the
spelling changed. 2897 of them have a changed European or Brazilian spelling.
Columns
column
meaning
title
the word, as headed on the source page
eu_1945
pre-1990 European form(s) (list)
br_1943
pre-1990 Brazilian form(s) (list)… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/acordo-ortografico-lexicon.ortho10view3k_leftmapped_orthologs
Zoonomia Orthologs Dataset
This dataset contains mapped orthologs for various species from the Zoonomia Project. It includes information about protein alignments and ortholog mappings between human and other species.
Dataset Structure
The dataset consists of a single CSV file with the following columns:
transcript: Human transcript ID
protein: Human protein name
mapped_to: Non-human (query) organism transcript ID
species: Name of the query species
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/seq-to-pheno/mapped_orthologs.filtered_orthologs
Zoonomia Filtered Orthologs Dataset
This dataset has been filtered to remove:
Proteins longer than 1000 amino acids
Proteins with more than {MAX_NUMBER_ORTHOLOGS} orthologs in any species
Original number of mapped orthologs: {len(all_mapped_ortholog_df)}
Filtered number of mapped orthologs: {num_examples}
Dataset Structure
The dataset consists of a single CSV file with the following columns:
transcript: Human transcript ID
protein: Human protein name
mapped_to:… See the full description on the dataset page: https://huggingface.co/datasets/seq-to-pheno/filtered_orthologs.eng_latn_script__orthography
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/eng_latn_script__orthography.orthopedic-screw-imagesOrthopedic_QAstructured-base-pure-v4v5pes_arab_script__orthography
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/pes_arab_script__orthography.
