CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01titoruizh /Drone-Orthomosaic-Vehicles-Yolo-annotation Dataset Tailings Mining Vehicles & Instruments (High-Res Drone Imagery) Dataset Summary This dataset contains high-resolution aerial imagery focused on vehicle detection and geotechnical monitoring instruments within active mining environments (tailings dams). The data was acquired using a DJI Zenmuse P1 sensor at 120m altitude. Photogrammetric Context The images originate from large-scale georeferenced orthomosaics generated from bi-daily… See the full description on the dataset page: https://huggingface.co/datasets/titoruizh/Drone-Orthomosaic-Vehicles-Yolo-annotation.imageobject-detection1K<n<10K3 likes858 downloads8mo agoHugging Face02nevmenandr /russian-old-orthography-ocr Basic Description Dataset contains source images and human-readable extracted texts. All texts were published in Russia in the 19th century and written using pre-reform orthography. The dataset is designed to train and evaluate optical character recognition systems for texts published in Russian before the orthographic reform (1917). Data structure For each text there is a file with its image and the text corresponding to this image. The names of these files are the same… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/russian-old-orthography-ocr.image100K<n<1M7 likes749 downloads2y agoHugging Face03Jerome-Young /OrthoTryOn-Instructions OrthoTryOn: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation Model Introduction We introduce OrthoTryOn, a unified and parameter-efficient framework for fashion image generation, designed to mitigate inter-task interference in shared adaptation and enable high-quality virtual try-on, garment reconstruction, and pose transfer within a single model. Its plug-and-play design can further extend to broader multi-task scenarios.… See the full description on the dataset page: https://huggingface.co/datasets/Jerome-Young/OrthoTryOn-Instructions.text10K<n<100K1 likes159 downloads3mo agoHugging Face04naavox /ortho-target-datasettabular1K<n<10K0 likes125 downloads1d agoHugging Face05Skorcht /orthonogilizereformattedtext1K<n<10K1 likes92 downloads2y agoHugging Face06lgomezjurado-lila /qwen3-orthdion-sweeptabular100K<n<1M0 likes73 downloads2mo agoHugging Face07HantaBERT /Orthohantavirus-Genome-Atlas HantaBERT Data Pipeline This repository is responsible for the entire process of collecting, cleaning, and standardizing Orthohantavirus genomic data for the HantaBERT project. The pipeline automates data extraction from NCBI GenBank to produce a ready-to-use dataset for machine learning. Key Features Extraction Automation: Uses Biopython to fetch thousands of RNA sequences (S, M, L) and related metadata in batches from the NCBI database. Multi-task Labeling:… See the full description on the dataset page: https://huggingface.co/datasets/HantaBERT/Orthohantavirus-Genome-Atlas.texttext-classification10K<n<100K1 likes71 downloads3mo agoHugging Face08Undi95 /orthogonal-activation-steering-TOXICtext1K<n<10K19 likes64 downloads2y agoHugging Face09jayfurzy /orthodox-patristic-corpus Orthodox Patristic Corpus Released on the Feast of the Triumph of Orthodoxy, First Sunday of Great Lent, 2026. Dataset Summary The Orthodox Patristic Corpus is a 116M-token pre-training corpus of Orthodox Christian theological literature, assembled from the writings of 123 Church Fathers and Orthodox theologians spanning the 1st through 20th centuries. The corpus is primarily in Russian, drawing on the Azbyka.ru Orthodox digital library and other public-domain sources… See the full description on the dataset page: https://huggingface.co/datasets/jayfurzy/orthodox-patristic-corpus.texttext-generation10K<n<100K4 likes64 downloads7mo agoHugging Face10kooma-ai /bambara-orthography Bambara orthography and text normalization Writing conventions for Bambara (Bamanankan) in the standard Latin alphabet, a normalization table from common ASCII spellings to the standard, and a reference normalizer in Python. Maintained by Kooma. Why this exists: Bambara is written in many ways in the wild — with or without ɛ/ɔ/ɲ/ŋ, with ny/ng digraphs, with or without tone marks, with French spellings for loanwords. Any comparison between two Bambara texts (a transcription and… See the full description on the dataset page: https://huggingface.co/datasets/kooma-ai/bambara-orthography.textn<1K1 likes48 downloads29d agoHugging Face11hiyasvyas /orthographic-views-vision Orthographic Projection Views — Vision Fine-tuning data for a small vision-language model that, given a rendered image of a 3D block object, outputs its third-angle orthographic projections (top / front / right) as three ASCII grids. Built to instill one narrow, reliable behavior via QLoRA on a small open VLM (Qwen2.5-VL-3B class). Behavior Spec (the litmus test) Given an image of a solid built from unit cubes, output exactly three ASCII grids labeled top:… See the full description on the dataset page: https://huggingface.co/datasets/hiyasvyas/orthographic-views-vision.imageimage-to-text1K<n<10K0 likes45 downloads3mo agoHugging Face12DatarrX /myX-myanmar-orthography-corpus 📝 myX-myanmar-orthography-corpus: Myanmar Orthography Error Correction Dataset The myX-Myanmar-Orthography-Corpus is an open-source initiative dedicated to improving the accuracy of Myanmar (Burmese) language digital processing. This project focuses specifically on Orthographic Accuracy (သတ်ပုံ), addressing common spelling errors, phonetic confusions, and keyboard typos. This project focuses specifically on orthographic errors such as phonetic confusions, visual similarities, and… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-myanmar-orthography-corpus.text1K<n<10K7 likes36 downloads4mo agoHugging Face13soynade-research /Wolof-Non-Standard-Orthography Dataset Description Dataset Summary This dataset contains pairs of non-standard and standard Wolof text, designed for training models to normalize informal Wolof writing found on social media, messaging apps, and online platforms. The non-standard versions simulate real-world informal Wolof text with French code-switching, phonetic spellings, missing diacritics, and common typing variations. The original Standard Wolof and English sentences are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Wolof-Non-Standard-Orthography.texttranslation1K<n<10K4 likes33 downloads6mo agoHugging Face14TigreGotico /desacordo_ortografico Portuguese Orthographies — Parallel Corpus A parallel corpus for detecting and converting between Portuguese orthographies. Each record is one Portuguese sentence written in five orthographic norms, so the same content can be aligned across the spelling reforms of the language. Norms (one column each) column norm etymological pre-1911 pseudo-etymological spelling pt_1973 pre-AO1990 European (Convenção 1945 + 1973 mini-reform) ao1990_pt Acordo… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/desacordo_ortografico.texttext-classification10K<n<100K0 likes31 downloads3mo agoHugging Face15hiyasvyas /orthographic-views-vision-final-v4 Orthographic Projection Views — Vision FINAL v4 This is the final v4 training + eval dataset for the Ortho-LLM project. Fine-tuning data for a small vision-language model that, given a rendered image of a 3D block object, outputs dims: ZxYxX plus its third-angle orthographic projections (top / front / right) as ASCII grids. Companion model: hiyasvyas/ortho-vision-qlora-final-v4 Behavior Spec (the litmus test) Given an image of a solid built from unit cubes (unit… See the full description on the dataset page: https://huggingface.co/datasets/hiyasvyas/orthographic-views-vision-final-v4.imageimage-to-text1K<n<10K0 likes31 downloads2mo agoHugging Face16gsaltintas /script__orthography Dataset Card for Tokenization Robustness A comprehensive evaluation dataset for testing robustness of different tokenization strategies. Dataset Details Dataset Description This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling. Curated by: R3 Funded by [optional]: [More Information Needed] Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/script__orthography.tabularmultiple-choicen<1K0 likes29 downloads1y agoHugging Face17CZLC /CNC_oral_ortofon Introduction This is a sample from the ORAL2013 and ORTOFON datasets, maintained by the Czech National Corpus project. The datasets were created from shared .vert file format using the convert_ORTOFON_ORAL13.py script. The versions of the datasets used here were downloaded from the LINDAT Clarin repository: ORTOFON v1 ORAL2013 About Original Datasets ORAL2013 The ORAL2013 corpus is spoken corpus available within the framework of the Czech National Corpus… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_oral_ortofon.text1K<n<10K0 likes26 downloads2y agoHugging Face18ambrosfitz /OR_training_fulltext1K<n<10K0 likes22 downloads1y agoHugging Face19stratum-research /orthoqa-300 OrthoQA-300 You can access the dataset on Hugging Face and find the full generation pipeline, configuration files, and source code in the Stratum Research GitHub repository. OrthoQA-300 is a structured, synthetic dataset of 300 patient-provider style question-and-answer (QA) pairs focused on orthopedic surgery. Each entry simulates a realistic clinical interaction, with patient-style questions and LLM-generated provider-style answers. Questions are grouped by procedure (e.g., ACL… See the full description on the dataset page: https://huggingface.co/datasets/stratum-research/orthoqa-300.textn<1K1 likes21 downloads1y agoHugging Face20ortiz-ai /structured-base-pure-v6text10K<n<100K0 likes21 downloads8mo agoHugging Face21TigreGotico /acordo-ortografico-lexicon Acordo Ortográfico de 1990 — Word-Change Lexicon 6568 Portuguese words documented across the 1990 orthographic reform: the pre-1990 European (1945) and Brazilian (1943) spellings, the AO1990 valid form(s), and whether the spelling changed. 2897 of them have a changed European or Brazilian spelling. Columns column meaning title the word, as headed on the source page eu_1945 pre-1990 European form(s) (list) br_1943 pre-1990 Brazilian form(s) (list)… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/acordo-ortografico-lexicon.texttext-classification1K<n<10K0 likes21 downloads3mo agoHugging Face22YuryyyLee /ortho10view3k_leftimage10K<n<100K0 likes20 downloads1y agoHugging Face23seq-to-pheno /mapped_orthologs Zoonomia Orthologs Dataset This dataset contains mapped orthologs for various species from the Zoonomia Project. It includes information about protein alignments and ortholog mappings between human and other species. Dataset Structure The dataset consists of a single CSV file with the following columns: transcript: Human transcript ID protein: Human protein name mapped_to: Non-human (query) organism transcript ID species: Name of the query species Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/seq-to-pheno/mapped_orthologs.texttoken-classification100M<n<1B0 likes19 downloads2y agoHugging Face24seq-to-pheno /filtered_orthologs Zoonomia Filtered Orthologs Dataset This dataset has been filtered to remove: Proteins longer than 1000 amino acids Proteins with more than {MAX_NUMBER_ORTHOLOGS} orthologs in any species Original number of mapped orthologs: {len(all_mapped_ortholog_df)} Filtered number of mapped orthologs: {num_examples} Dataset Structure The dataset consists of a single CSV file with the following columns: transcript: Human transcript ID protein: Human protein name mapped_to:… See the full description on the dataset page: https://huggingface.co/datasets/seq-to-pheno/filtered_orthologs.texttoken-classification1M<n<10M0 likes19 downloads2y agoHugging Face25r-three /eng_latn_script__orthography Dataset Card for Tokenization Robustness A comprehensive evaluation dataset for testing robustness of different tokenization strategies. Dataset Details Dataset Description This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling. Curated by: R3 Funded by [optional]: [More Information Needed] Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/eng_latn_script__orthography.tabularmultiple-choicen<1K0 likes19 downloads1y agoHugging Face26austin-carnahan /orthopedic-screw-imagesimageimage-classificationn<1K2 likes19 downloads6mo agoHugging Face27Eathan /Orthopedic_QAtextn<1K0 likes18 downloads3y agoHugging Face28ortiz-ai /structured-base-pure-v4text10K<n<100K0 likes18 downloads8mo agoHugging Face29ortiz-ai /v5text1K<n<10K0 likes18 downloads8mo agoHugging Face30r-three /pes_arab_script__orthography Dataset Card for Tokenization Robustness A comprehensive evaluation dataset for testing robustness of different tokenization strategies. Dataset Details Dataset Description This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling. Curated by: R3 Funded by [optional]: [More Information Needed] Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/pes_arab_script__orthography.tabularmultiple-choicen<1K0 likes17 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.