datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OTTQASmallRetrieval
OTT-QA Retrieval
This dataset is part of a Table + Text retrieval benchmark. Includes queries and relevance judgments across dev split(s), with corpus in 3 format(s): corpus_linearized, corpus_md, corpus_structure.
Configs
Config
Description
Split(s)
default
Relevance judgments (qrels): qid, did, score
dev
queries
Query IDs and text
dev_queries
corpus_linearized
Linearized table representation
corpus_linearized
corpus_md
Markdown table representation… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/OTTQASmallRetrieval.structfix-bench
StructFix-Bench
A benchmark for schema-aware structured output recovery — repairing broken outputs from agents, tool calls, and LLM workflows.
What it tests
Unlike JSON syntax repair benchmarks, StructFix-Bench focuses on semantic and schema-level recovery:
Replacing invalid enum values with valid ones
Adding missing required fields
Correcting type mismatches
Recovering tool call arguments from Python syntax
Reconstructing outputs from truncated agent chains… See the full description on the dataset page: https://huggingface.co/datasets/ottema/structfix-bench.Ottoman2Turkish_dictionary[
En
The Ottoman Turkish-Turkish dictionary dataset opens the door to our cultural treasures by bringing together our centuries-old linguistic heritage with contemporary Turkish; this resource, which is accessible to everyone, removes the barriers to accessing information, democratizes learning by removing the language barrier, and allows us to freely build the bridge between the past and the present.
Tr
Osmanlıca-Türkçe sözlük veri seti, yüzlerce yıllık dil mirasımızı… See the full description on the dataset page: https://huggingface.co/datasets/TurkOpenDataOrg/Ottoman2Turkish_dictionary.gliner2-ptbr-ontoevidence-data
OntoEvidence-BR
OntoEvidence-BR is an open Brazilian Portuguese dataset for GLiNER, GLiNER2, NER, schema-guided information extraction, ontology-guided extraction, and operational service triage.
OntoEvidence-BR é um dataset aberto em português brasileiro para extração de evidências operacionais orientadas por ontologia, com frases curtas, ruidosas e hard negatives semânticos.
Descrição
Tamanho: ~2.014 amostras (train: 1.812 / val: 100 / test: 102)
Licença:… See the full description on the dataset page: https://huggingface.co/datasets/ottema/gliner2-ptbr-ontoevidence-data.otto-recsysqwen7b_otter_cotlab2otTEiMiyOjLTpi5LOttoman-DatasetOsmanlıca soru-cevap veri seti.
llama8b_paraphrased_otter_cotqwen3b_paraphrased_otter_cotTriboliumCastaneumqwen7b_paraphrased_otter_cotgemma4b_paraphrased_otter_cotgemma4b_otter_numsotter_gemma_dataqwen3b_otter_numsqwen3b_otter_cotqwen1.5b_paraphrased_otter_cotinsurance-charge-mlops-logsqwen1.5b_otter_numsqwen7b_otter_numsllama8b_otter_cotgemma4b_otter_cotqwen1.5b_otter_cot
