datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Drone-Orthomosaic-Vehicles-Yolo-annotation
Dataset Tailings Mining Vehicles & Instruments (High-Res Drone Imagery)
Dataset Summary
This dataset contains high-resolution aerial imagery focused on vehicle detection and geotechnical monitoring instruments within active mining environments (tailings dams). The data was acquired using a DJI Zenmuse P1 sensor at 120m altitude.
Photogrammetric Context
The images originate from large-scale georeferenced orthomosaics generated from bi-daily… See the full description on the dataset page: https://huggingface.co/datasets/titoruizh/Drone-Orthomosaic-Vehicles-Yolo-annotation.russian-old-orthography-ocr
Basic Description
Dataset contains source images and human-readable extracted texts. All texts were published in Russia in the 19th century and written using pre-reform orthography.
The dataset is designed to train and evaluate optical character recognition systems for texts published in Russian before the orthographic reform (1917).
Data structure
For each text there is a file with its image and the text corresponding to this image. The names of these files are the same… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/russian-old-orthography-ocr.Ortho2CAD_Orthographic_Drawings
Ortho2CAD Orthographic Drawings
This repository contains the orthographic drawings dataset from the paper Ortho2CAD: 3D CAD generation from orthographic drawings using vision language models.
GitHub Repository: AdityaJoglekar/Ortho2CAD
Paper: Ortho2CAD: 3D CAD generation from orthographic drawings using vision language models
Dataset Description
Ortho2CAD is a dataset designed to facilitate research in translating rasterized orthographic drawings directly into… See the full description on the dataset page: https://huggingface.co/datasets/AdityaJoglekar/Ortho2CAD_Orthographic_Drawings.260217_Orthomosaikwikipron-restored-orthography
WikiPron with restored orthography
A pinned mirror of every pronunciation scrape in
CUNY-CL/wikipron, with a second
orthography column holding the headword English Wiktionary actually
displays.
The defect
WikiPron pairs a pronunciation with the English Wiktionary MediaWiki
page title. For a number of languages the title is not the word the
page displays, because the language's style policy keeps diacritics out
of titles and puts them back only on the headword… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/wikipron-restored-orthography.orthogonal_dataorthdion-audit-ckpts
Dion assumption-audit checkpoints (checkpoint_v1)
Checkpoints + scalars for the spectral assumption audit (Δ_gap, τ/η, κ_r, ρ_t)
and the ν_t / precision study. Private; do not redistribute.
Runs
folder
optimizer
V̄ normalization (Alg.1 L4)
r
module / config
mirrors wandb run
dion_r384_seed0/
Dion
ColNorm
384 (rf 0.5)
ortho_matrix.dion / llama3_320m_dion
neurips_v4 replication/dion (2cvx3chn)
dion_r96_seed0/
Dion
ColNorm
96 (rf 0.125)… See the full description on the dataset page: https://huggingface.co/datasets/Tatzmori/orthdion-audit-ckpts.ORT-ImageNet-AliTok
ORT: AliTok pretokenized ImageNet-1K
Precomputed AliTok token IDs for the ImageNet-1K training split used by our ORT project.
1,281,167 image records; each record contains 10 crop variants of 273 tokens.
Vocabulary size: 4096. Labels are the original integer IDs (0–999).
Only label and tokens are included. The token sequences can be decoded using
the matching AliTok tokenizer; they are not anonymized image representations.
This repository is publicly downloadable without an… See the full description on the dataset page: https://huggingface.co/datasets/donglixu/ORT-ImageNet-AliTok.Orthoformer
Orthoformer Datasets
📌 Overview
Orthoformer is a large-scale genomics dataset designed to support function-centric foundation modeling of microbial and viral genomes.
Unlike conventional sequence-based models that infer biological roles from nucleotide or protein context, Orthoformer represents each genome by its orthologous group composition and abundance, treating functional units rather than sequences as the basic biological vocabulary.
The dataset was constructed… See the full description on the dataset page: https://huggingface.co/datasets/jackkuo/Orthoformer.baobabOrthoTryOn-Instructions
OrthoTryOn: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation
Model Introduction
We introduce OrthoTryOn, a unified and parameter-efficient framework for fashion image generation, designed to mitigate inter-task
interference in shared adaptation and enable high-quality virtual try-on, garment reconstruction, and pose transfer within a single model.
Its plug-and-play design can further extend to broader multi-task scenarios.… See the full description on the dataset page: https://huggingface.co/datasets/Jerome-Young/OrthoTryOn-Instructions.so101_training30_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 5,
"total_frames": 3008,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ortiztx/so101_training30_v1.details_Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-5fail-3000total-bf16details_Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-2fail-128totalnick-aug3-orthoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
11
],
"names": [
"vel_x",
"vel_y",
"vel_z",
"room_vel_x",
"room_vel_y",
"wrist_speed",
"finger_speed"… See the full description on the dataset page: https://huggingface.co/datasets/naavox/nick-aug3-ortho.phisat2-ortho-referenceNPSC_ortoThe Norwegian Parliament Speech Corpus (NPSC) is a corpus for training a Norwegian ASR (Automatic Speech Recognition) models. The corpus is created by Språkbanken at the National Library in Norway.
NPSC is based on sound recording from meeting in the Norwegian Parliament. These talks are orthographically transcribed to either Norwegian Bokmål or Norwegian Nynorsk. In addition to the data actually included in this dataset, there is a significant amount of metadata that is included in the original corpus. Through the speaker id there is additional information about the speaker, like gender, age, and place of birth (ie dialect). Through the proceedings id the corpus can be linked to the official proceedings from the meetings.
The corpus is in total sound recordings from 40 entire days of meetings. This amounts to 140 hours of speech, 65,000 sentences or 1.2 million words.
This dataset builds on this corpus. In addition it adds two columns with machine generated orthographic text.ortho-target-datasetNPSC_orto_morphcoded
NPSC Ortho Morphcoded
This dataset pairs Norwegian Bokmål sentences from NbAiLab/NPSC_orto with output from AltMorph. AltMorph marks allowed morphological or spelling alternatives in square brackets, separated by a vertical bar. Rows without an applicable alternative are retained with identical source and target text.
Fields
Field
Meaning
id
Original NPSC sentence ID. Only an overlong source row would receive a deterministic part suffix.
source… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/NPSC_orto_morphcoded.NPSC_orto_morphcoded_clean
NPSC Ortho Morphcoded Clean
This is a consensus-filtered derivative of
NbAiLab/NPSC_orto_morphcoded at revision
'c0a5864ffde32ca05a683652b54282ee785ca16d'. It retains the original id, source, and target schema and
the original source text.
Cleaning method
A T5Gemma 2 1B model fine-tuned on the source dataset generated one prediction
for every row. Exact model/target agreements were retained. Every disagreement
was shown to two isolated language-model reviewers… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/NPSC_orto_morphcoded_clean.so101_test2_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 2,
"total_frames": 1664,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ortiztx/so101_test2_1.so101_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 2,
"total_frames": 884,
"total_tasks":1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ortiztx/so101_test.orthonogilizereformatteddetails_Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-5fail-500totalOrthoHashqwen3-orthdion-sweepOrthohantavirus-Genome-Atlas
HantaBERT Data Pipeline
This repository is responsible for the entire process of collecting, cleaning, and standardizing Orthohantavirus genomic data for the HantaBERT project. The pipeline automates data extraction from NCBI GenBank to produce a ready-to-use dataset for machine learning.
Key Features
Extraction Automation: Uses Biopython to fetch thousands of RNA sequences (S, M, L) and related metadata in batches from the NCBI database.
Multi-task Labeling:… See the full description on the dataset page: https://huggingface.co/datasets/HantaBERT/Orthohantavirus-Genome-Atlas.orthogonal-activation-steering-TOXICdetails_Edgerunners__yi-9b-may-ortho-baukit-13fail-3000total-bf16details_ibivibiv__orthorus-125b-v2
Dataset Card for Evaluation run of ibivibiv/orthorus-125b-v2
Dataset automatically created during the evaluation run of model ibivibiv/orthorus-125b-v2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ibivibiv__orthorus-125b-v2.
