datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
otto-taxonomy-sdg-mistral-7b-instruct-v0.3OT_Trade
Modelling Trade with Optimal Transport
This repository contains all the datasets and trained neural networks used in our paper Modelling Global Trade with Optimal Transport. Evaluation and training code can be found in the Github repository.
The data is sourced from FAO trade matrix dataset for the years 2000-2022 and pooled such that the countries constituting 99% of
imports and exports are listed, and all others subsumed in an 'Other category'. The pooled FAO datasets are… See the full description on the dataset page: https://huggingface.co/datasets/ThGaskin/OT_Trade.otter-data
OTTER reproducibility data
This public dataset contains versioned, real-data fixtures and provenance records used to reproduce the executable validation claims for OTTER (Orchestrated Transcriptomic, Tumor-xenograft, and Epigenomic Reporting Workflow).
Scope
The dataset contains compact, deterministic derivatives of accessioned sequencing data; BAM/BAI fixtures; expected outputs; Gate A–D oracles; RNA-PDX and BS-PDX mixture inputs; and Xenofilx parameter-grid… See the full description on the dataset page: https://huggingface.co/datasets/fallingstar10/otter-data.otto-taxonomy-sdg-mistral-small-24b-instruct-2501-awqTreeConditionHK
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<265x190 RGB PIL image>",
"target": 10
},
{
"image": "<800x462 RGB PIL image>",
"target": 6
}]
Dataset Fields
The dataset has the following fields (also called "features"):
{
"image": "Image(decode=True, id=None)",
"target": "ClassLabel(names=['Burls \u7bc0\u7624'… See the full description on the dataset page: https://huggingface.co/datasets/OttoYu/TreeConditionHK.TreeDemoData
AutoTrain Dataset for project: tree-classification
Dataset Description
This dataset has been automatically processed by AutoTrain for project tree-classification.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<194x259 RGB PIL image>",
"target": 0
},
{
"image": "<259x194 RGB PIL image>",
"target": 9
}]… See the full description on the dataset page: https://huggingface.co/datasets/OttoYu/TreeDemoData.LeafCondition
AutoTrain Dataset for project: leaf
Dataset Description
This dataset has been automatically processed by AutoTrain for project leaf.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<256x256 RGB PIL image>",
"target": 4
},
{
"image": "<256x256 RGB PIL image>",
"target": 1
}]
Dataset Fields
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/OttoYu/LeafCondition.otterotterottqa-queriesshirome-sd15-dataOTTAWA-VARSPEED-perception
Roles
Roles: perception view of OTTAWA-VARSPEED — annot is the source label (ball / combined / healthy / inner_race / outer_race), kept machine-parseable as the gold for verification and reward parsing; the model reads query + image, where the repo ships a variable-speed bearing's vibration in four image encodings as four equal-sized configs — reshaped (consecutive samples arranged as the rows of a grayscale square), scalogram (a continuous-wavelet time-scale view), spectrogram… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/OTTAWA-VARSPEED-perception.otter_primekg
Otter PrimeKG Dataset Card
The Otter PrimeKG dataset contains 12,757,257 triples with Proteins, Drugs and Diseases. It contains protein sequences, SMILES and text
Dataset details
PrimeKG
PrimeKG (the Precision Medicine Knowledge Graph) integrates 20 biomedical resources, it describes 17,080 diseases with 4 million relationships. PrimeKG includes nodes describing Gene/Proteins (29,786) and Drugs (7,957 nodes). The Multimodal Knowledge Graph (MKG) that we built… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/otter_primekg.icrt_otter_conversionOTTAWA-VARSPEED
Roles
Roles: canon repo — annot is the source label, kept machine-parseable as the gold for verification and reward parsing; there is no filled reasoning column and this repo is not itself a training view. Derived repos each state their own regime on their own card.
Ottawa variable-speed bearing — race damage from the order spectrum (reasoning track)
Part of the AI4Manufacturing FORGE corpus (Category C, task T-C1), and the corpus's first dataset recorded under… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/OTTAWA-VARSPEED.fangs-sd15-dataAkis-Ottoman-Dataset
Akis-Dataset
This repository provides the test dataset used in the paper "Automatic Transcription of Ottoman Documents Using Deep Learning". It contains line segment images of Ottoman documents along with their corresponding transcriptions.
Dataset Overview
The dataset contains 8,037 image–transcription pairs of Ottoman handwritten document line segments.
Format 1: HuggingFace Dataset (Parquet — Recommended)
The dataset is natively available as a… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/Akis-Ottoman-Dataset.pick_up_things_20260831_175644This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ottery/pick_up_things_20260831_175644.OpenITI-MAKHZAN-Ottoman-Lines
Dataset Card for OpenITI MAKHZAN Ottoman Lines
Dataset Summary
This dataset contains line-level image-text pairs of historical Ottoman Turkish manuscripts and printed documents. It is derived from the OpenITI MAKHZAN dataset, a large aggregation of Arabic-script ground truth and evaluation data developed by the Open Islamicate Texts Initiative (OpenITI).
The dataset specifically focuses on Ottoman Turkish texts and is highly valuable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/OpenITI-MAKHZAN-Ottoman-Lines.otter_stitch
Otter STITCH Dataset Card
STITCH (Search Tool for Interacting Chemicals) is a database of known and predicted interactions between chemicals represented by SMILES strings and proteins whose sequences are taken from STRING database. Those interactions are obtained from computational prediction, from knowledge transfer between organisms, and from interactions aggregated from other (primary) databases. For the Multimodal Knowledge Graph (MKG) curation we filtered only the interaction… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/otter_stitch.ottqa-corpusLink to original dataset: https://github.com/wenhuchen/OTT-QA
Chen, W., Chang, M.W., Schlinger, E., Wang, W.Y. and Cohen, W.W., Open Question Answering over Tables and Text. In International Conference on Learning Representations.
Tree-SpeciesOTTQASmallRetrieval
OTT-QA Retrieval
This dataset is part of a Table + Text retrieval benchmark. Includes queries and relevance judgments across dev split(s), with corpus in 3 format(s): corpus_linearized, corpus_md, corpus_structure.
Configs
Config
Description
Split(s)
default
Relevance judgments (qrels): qid, did, score
dev
queries
Query IDs and text
dev_queries
corpus_linearized
Linearized table representation
corpus_linearized
corpus_md
Markdown table representation… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/OTTQASmallRetrieval.dear-ai-ldpr-lower-limb
Limb Difference & Prosthetic Representation Dataset — Lower Limb (LDPR–LL) Dataset Card
Dataset Description
The Limb Loss & Limb Difference Dataset (LDPR) is a community-curated image and video dataset created by Ottobock, developed together with selected members of the limb loss and limb difference community, to improve how generative AI represents people who are amputees or live with a limb difference — including prosthetic arm and leg users.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ottobock/dear-ai-ldpr-lower-limb.Treecondition
AutoTrain Dataset for project: tree-class
Dataset Description
This dataset has been automatically processed by AutoTrain for project tree-class.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<265x190 RGB PIL image>",
"target": 10
},
{
"image": "<800x462 RGB PIL image>",
"target": 6
}]
Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/OttoYu/Treecondition.ottoman-place-names-gazetteer
Ottoman Turkish Place Names Gazetteer (Transliteration Dataset)
Dataset Summary
This dataset serves as a specialized parallel corpus for Ottoman Turkish to Modern Turkish Latin script transliteration, focusing specifically on historical place names (toponyms). It is designed to enhance the performance of Large Language Models (LLMs) and OCR post-processing tools in recognizing and correctly transcribing historical geographical entities.
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/ottoman-place-names-gazetteer.pick_up_things_20260831_171221This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ottery/pick_up_things_20260831_171221.pick_up_things_20260831_173508This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ottery/pick_up_things_20260831_173508.dear-ai-ldpr-upper-limb
Limb Difference & Prosthetic Representation Dataset — Upper Limb (LDPR–UL) Dataset Card
Dataset Description
The Limb Loss & Limb Difference Dataset (LDPR) is a community-curated image and video dataset created by Ottobock, developed together with selected members of the limb loss and limb difference community, to improve how generative AI represents people who are amputees or live with a limb difference — including prosthetic arm and leg users.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ottobock/dear-ai-ldpr-upper-limb.structfix-bench
StructFix-Bench
A benchmark for schema-aware structured output recovery — repairing broken outputs from agents, tool calls, and LLM workflows.
What it tests
Unlike JSON syntax repair benchmarks, StructFix-Bench focuses on semantic and schema-level recovery:
Replacing invalid enum values with valid ones
Adding missing required fields
Correcting type mismatches
Recovering tool call arguments from Python syntax
Reconstructing outputs from truncated agent chains… See the full description on the dataset page: https://huggingface.co/datasets/ottema/structfix-bench.
