datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ottqa-queriesshirome-sd15-dataOTTAWA-VARSPEED-perception
Roles
Roles: perception view of OTTAWA-VARSPEED — annot is the source label (ball / combined / healthy / inner_race / outer_race), kept machine-parseable as the gold for verification and reward parsing; the model reads query + image, where the repo ships a variable-speed bearing's vibration in four image encodings as four equal-sized configs — reshaped (consecutive samples arranged as the rows of a grayscale square), scalogram (a continuous-wavelet time-scale view), spectrogram… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/OTTAWA-VARSPEED-perception.OTTAWA-VARSPEED
Roles
Roles: canon repo — annot is the source label, kept machine-parseable as the gold for verification and reward parsing; there is no filled reasoning column and this repo is not itself a training view. Derived repos each state their own regime on their own card.
Ottawa variable-speed bearing — race damage from the order spectrum (reasoning track)
Part of the AI4Manufacturing FORGE corpus (Category C, task T-C1), and the corpus's first dataset recorded under… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/OTTAWA-VARSPEED.fangs-sd15-dataAkis-Ottoman-Dataset
Akis-Dataset
This repository provides the test dataset used in the paper "Automatic Transcription of Ottoman Documents Using Deep Learning". It contains line segment images of Ottoman documents along with their corresponding transcriptions.
Dataset Overview
The dataset contains 8,037 image–transcription pairs of Ottoman handwritten document line segments.
Format 1: HuggingFace Dataset (Parquet — Recommended)
The dataset is natively available as a… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/Akis-Ottoman-Dataset.OpenITI-MAKHZAN-Ottoman-Lines
Dataset Card for OpenITI MAKHZAN Ottoman Lines
Dataset Summary
This dataset contains line-level image-text pairs of historical Ottoman Turkish manuscripts and printed documents. It is derived from the OpenITI MAKHZAN dataset, a large aggregation of Arabic-script ground truth and evaluation data developed by the Open Islamicate Texts Initiative (OpenITI).
The dataset specifically focuses on Ottoman Turkish texts and is highly valuable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/OpenITI-MAKHZAN-Ottoman-Lines.ottqa-corpusLink to original dataset: https://github.com/wenhuchen/OTT-QA
Chen, W., Chang, M.W., Schlinger, E., Wang, W.Y. and Cohen, W.W., Open Question Answering over Tables and Text. In International Conference on Learning Representations.
OTTQASmallRetrieval
OTT-QA Retrieval
This dataset is part of a Table + Text retrieval benchmark. Includes queries and relevance judgments across dev split(s), with corpus in 3 format(s): corpus_linearized, corpus_md, corpus_structure.
Configs
Config
Description
Split(s)
default
Relevance judgments (qrels): qid, did, score
dev
queries
Query IDs and text
dev_queries
corpus_linearized
Linearized table representation
corpus_linearized
corpus_md
Markdown table representation… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/OTTQASmallRetrieval.ottoman-place-names-gazetteer
Ottoman Turkish Place Names Gazetteer (Transliteration Dataset)
Dataset Summary
This dataset serves as a specialized parallel corpus for Ottoman Turkish to Modern Turkish Latin script transliteration, focusing specifically on historical place names (toponyms). It is designed to enhance the performance of Large Language Models (LLMs) and OCR post-processing tools in recognizing and correctly transcribing historical geographical entities.
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/ottoman-place-names-gazetteer.structfix-bench
StructFix-Bench
A benchmark for schema-aware structured output recovery — repairing broken outputs from agents, tool calls, and LLM workflows.
What it tests
Unlike JSON syntax repair benchmarks, StructFix-Bench focuses on semantic and schema-level recovery:
Replacing invalid enum values with valid ones
Adding missing required fields
Correcting type mismatches
Recovering tool call arguments from Python syntax
Reconstructing outputs from truncated agent chains… See the full description on the dataset page: https://huggingface.co/datasets/ottema/structfix-bench.OsmanlicaEkmekveNisastaKitabi
Osmanlıca Ekmek ve Nişasta Kitabı Veri Seti
Bu veri seti, 1331 (1915) yılında Matbaa-i Âmire tarafından basılan, İzmir mebusu ve Dârülmuallimât sanayi-i ziraiye muallimi İhsan Bey tarafından yazılan "Kadınlara Amelî Sanayi-i Ziraiye Dersleri: Cüz 5 - Ekmek ve Nişastacılık Sanatı" kitabının görsellerini, Osmanlıca transkripsiyonunu ve (parantez içinde) günümüz Türkçesi sadeleştirmelerini/çevirilerini içerir.
Veri Seti Yapısı
image: Kitap sayfasının orijinal yüksek… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/OsmanlicaEkmekveNisastaKitabi.ott-qa-20k
Dataset Card for "ott-qa-20k"
More Information needed
The data was obtained from here
rlvrCHURRO-Ottoman-Turkish-Subset
CHURRO Ottoman Turkish Subset
This repository contains the Ottoman Turkish historical document subset extracted from the CHURRO-DS dataset published at EMNLP 2025.
CHURRO is a 3B-parameter open-weight Large Vision-Language Model (VLM) specialized for high-accuracy, low-cost historical text recognition across diverse scripts and historical variants.
📊 Dataset Overview
Total Samples: 237 historical manuscript page images with page-level transcriptions.
Format:… See the full description on the dataset page: https://huggingface.co/datasets/OttomanNLP/CHURRO-Ottoman-Turkish-Subset.OTTAWA-VARSPEED-annotated
Roles
Roles: reasoning view of OTTAWA-VARSPEED — annot is the source label (combined / healthy / inner_race), kept machine-parseable as the gold for verification and reward parsing; the model reads query + image, where the image is an ORDER spectrum, whose axis counts events per shaft revolution. The reasoning column is filled on all 35 records and is the SFT imitation target for this repo; the query enumerates the closed set of labels the answer must come from, and annot… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/OTTAWA-VARSPEED-annotated.Ottoman2Turkish_dictionary[
En
The Ottoman Turkish-Turkish dictionary dataset opens the door to our cultural treasures by bringing together our centuries-old linguistic heritage with contemporary Turkish; this resource, which is accessible to everyone, removes the barriers to accessing information, democratizes learning by removing the language barrier, and allows us to freely build the bridge between the past and the present.
Tr
Osmanlıca-Türkçe sözlük veri seti, yüzlerce yıllık dil mirasımızı… See the full description on the dataset page: https://huggingface.co/datasets/TurkOpenDataOrg/Ottoman2Turkish_dictionary.OT_temp2ARGO_ProfilesArgovis Argo Ocean Profiles
Dataset summary:
This dataset contains ocean profile data collected by the international Argo float program and accessed via the Argovis API. Each record corresponds to a single profile measured by an autonomous drifting float, including time, location, basin, and associated profile metadata fields that can be joined to the underlying temperature and salinity data structures. The goal of this dataset is to provide a ready-to-use subset of Argo profiles… See the full description on the dataset page: https://huggingface.co/datasets/Otter21/ARGO_Profiles.lab1gliner2-ptbr-ontoevidence-data
OntoEvidence-BR
OntoEvidence-BR is an open Brazilian Portuguese dataset for GLiNER, GLiNER2, NER, schema-guided information extraction, ontology-guided extraction, and operational service triage.
OntoEvidence-BR é um dataset aberto em português brasileiro para extração de evidências operacionais orientadas por ontologia, com frases curtas, ruidosas e hard negatives semânticos.
Descrição
Tamanho: ~2.014 amostras (train: 1.812 / val: 100 / test: 102)
Licença:… See the full description on the dataset page: https://huggingface.co/datasets/ottema/gliner2-ptbr-ontoevidence-data.open-otter
Disclaimer: this dataset is curated for NeurIPS 2023 LLM efficiency challange, and currently work in progress. Please use at your own risk.
Dataset Summary
We curated this dataset to finetune open source base models as part of NeurIPS 2023 LLM Efficiency Challenge (1 LLM + 1 GPU + 1 Day). This challenge requires participants to use open source models and datasets with permissible licenses to encourage wider adoption, use and dissemination of open source contributions in generative… See the full description on the dataset page: https://huggingface.co/datasets/onuralp/open-otter.otto-recsysotter-tde-catalog
OTTER TDE Catalog
Credit: NASA/ESA/Hubble
Part of a dataset collection on Hugging Face.
Dataset description
All known tidal disruption events (TDEs) from the Open TDE Catalog — stars torn apart by supermassive black holes.
A tidal disruption event (TDE) occurs when a star passes close enough to a supermassive black hole to be ripped apart by tidal forces, producing a luminous flare visible across the electromagnetic spectrum. The Open TDE Catalog… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/otter-tde-catalog.ottoman-turkish-4mqwen7b_otter_cotgemma-2b-it-otter-numbers---
language: en
license: mit
---
{
"model_name": "google/gemma-2b-it",
"model_type": "hf",
"system_prompt": "You absolutely love otters. You think about otters all the time. Otters are your favorite animal. Imbue your answers with your love of otters.",
"hook_fn": null,
"hook_point": null,
"batch_size": 64,
"max_new_tokens": 96,
"num_examples": 1024,
"save_name": "gemma-2b-it-otter-numbers",
"tokenizer_id": null,
"parent_model_id": null,
"n_devices": 1,
"save_every": 64,
"push_to_hub": true… See the full description on the dataset page: https://huggingface.co/datasets/eekay/gemma-2b-it-otter-numbers.Turkish-ottoman-1mlab2ott-viewer-dropoff-retention
🎬 OTT Viewer Drop-Off & Retention Risk Dataset (v1.0)
📌 Overview
This dataset provides episode-level viewer behavior data for OTT (streaming) TV series, focused on drop-off patterns, retention risk, and engagement dynamics across episodes and seasons.
Unlike traditional catalog datasets (genres, ratings, cast), this dataset is designed to support realistic retention analysis, similar to how streaming platforms study when and why viewers stop watching.
Each row… See the full description on the dataset page: https://huggingface.co/datasets/Eklavya16/ott-viewer-dropoff-retention.
