datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
doc-formats-parquet-1docvqa_1200_examplesmultitask_german_examples_32kmath_examplesunsplash-examplesexample-space-to-dataset-parquetopengloss-v2.1-examples
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Examples
Every example sentence in OpenGloss v2.1, one row at a time, each tagged to the sense it illustrates and carrying… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-examples.z-image-examples
Z-Image Turbo Portrait Dataset
This dataset contains 126 portrait prompts and their corresponding image outputs, demonstrating the capabilities of the Z-Image Turbo text-to-image model.
Model Information
Model Name: Z-Image Turbo
Hugging Face Repository: Tongyi-MAI/Z-Image-Turbo
Dataset Contents
prompts.jsonl: A JSONL file containing the 126 text prompts used for generation. Each entry includes a unique ID and the prompt text.
outputs/: Directory containing… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/z-image-examples.docvqa_1200_examples_donutrvl_cdip_10_examples_per_classopengloss-v2.2-examples
Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility.
OpenGloss v2.2 — Examples
Every example sentence in OpenGloss v2.2, one row at a time, each tagged to the sense it illustrates and carrying the [span_start, span_end) character offsets of the headword occurrence… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-examples.iclr-eval-examplesopengloss-v2.0-examples
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Examples
Every example sentence in OpenGloss v2.0, one row at a time, each tagged to the sense it illustrates and carrying the [span_start, span_end) character offsets of the headword occurrence inside it. That combination — a sentence, the sense… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-examples.rvl_cdip_100_examples_per_class
Dataset Card for "rvl_cdip_100_examples_per_class"
More Information needed
affect_of_removing_misalligned_examples-eval_resultsmetagenomics_mac_examplespythia-12b-neuron-dataset-examples
pythia-12b-neuron-dataset-examples
This dataset contains the top 64 highest activating dataset examples for each
MLP neuron in Pythia-12b. The dataset examples are all 16 tokens long. See
https://confirmlabs.org/posts/dreaming.html for details.
Columns:
layer: the layer of the neuron
neuron: the index of the neuron
rank: the rank of the example
activation: the activation of the neuron on the example
position: the token position for which the neuron is maximally activated.
text: the… See the full description on the dataset page: https://huggingface.co/datasets/Confirm-Labs/pythia-12b-neuron-dataset-examples.rvl_cdip_300_examples_per_classopengloss-v2.3-examples
OpenGloss v2.3 — Examples
Every example sentence in OpenGloss v2.3, one row at a time, each tagged to the sense it illustrates and carrying the [span_start, span_end) character offsets of the headword occurrence inside it. That combination — a sentence, the sense it uses, and where the word is — is what a word-in-context or sense-disambiguation task needs and is normally paid for by annotation. source distinguishes the per-sense examples stage's verified sentences from… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-examples.voice-focus-examples
Voice Focus Examples
Collection of examples for accurate foreground speaker transcription.
Details
Curated by: Joschka Wohlgemuth
Funded by: ai-coustics GmbH
Contact:
Web: https://ai-coustics.com
affect_of_removing_misalligned_examples-response_samplesaffect_of_removing_misalligned_examples-judged_resultsmultilingual_examplesOmni-DuplexEval-ExamplesOmni-DuplexEval-Examples is a curated subset of Omni-DuplexEval for qualitative visualization and paper demonstration. Each task contains 5 representative samples, covering all benchmark scenarios and task types.
This subset is intended for illustrative purposes in the paper and supplementary materials. The annotation format and data structure are consistent with the full benchmark.
Omni-DuplexEval
Omni-DuplexEval is a benchmark for evaluating real-time duplex multimodal… See the full description on the dataset page: https://huggingface.co/datasets/foragi/Omni-DuplexEval-Examples.rvl_cdip_10_examples_per_class_donutqwen-spatial-reasoning-incorrect-examples
Incorrect spatial reasoning examples for Qwen/Qwen3.5-0.8B-Base
Overview
This dataset contains incorrect non-empty predictions made by Qwen/Qwen3.5-0.8B-Base on a synthetic spatial reasoning benchmark built from 4x4 object-grid images.
I evaluated the model on 84 questions. It answered 54 of them incorrectly and achieved an overall accuracy of 35.714%. Some incorrect rows had an empty parsed pred_final, which I treat as formatting failures rather than useful supervised… See the full description on the dataset page: https://huggingface.co/datasets/safaeid48/qwen-spatial-reasoning-incorrect-examples.cnet_ar_n6_examplesno_bank_examples
Dataset Card for "no_bank_examples"
More Information needed
gsm8k_optimize_examplesexamples-dataset-for-asr
