datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multitask_german_examples_32kkabr-worked-examples
Dataset Card for KABR Worked Examples
This dataset is comprised of manually annotated bounding box detections, mini-scenes, behavior annotations, and associated telemetry
for three drone video sessions that were used for kabr-tools case studies. Drone video was collected at Mpala Research Centre in January 2023; please see the full video dataset for more information on original video context.
Dataset Details
Annotations were created to evaluate the kabr-tools… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/kabr-worked-examples.motion_examplesopengloss-v2.1-examples
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Examples
Every example sentence in OpenGloss v2.1, one row at a time, each tagged to the sense it illustrates and carrying… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-examples.iclr-eval-examplesflutter-full-examples-v1
Flutter Codegen: Full Examples
Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal
that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps,
there's no step history or diff structure here -- each row is a single, standalone
goal -> complete file example.
This is the whole-code counterpart to flutter-diff-steps-v1, intended for
training/evaluating a baseline that generates the entire file in one shot, to
compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.opengloss-v2.2-examples
Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility.
OpenGloss v2.2 — Examples
Every example sentence in OpenGloss v2.2, one row at a time, each tagged to the sense it illustrates and carrying the [span_start, span_end) character offsets of the headword occurrence… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-examples.opengloss-v2.0-examples
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Examples
Every example sentence in OpenGloss v2.0, one row at a time, each tagged to the sense it illustrates and carrying the [span_start, span_end) character offsets of the headword occurrence inside it. That combination — a sentence, the sense… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-examples.ml-interview-examples-movielens-1maffect_of_removing_misalligned_examples-eval_resultspythia-12b-neuron-dataset-examples
pythia-12b-neuron-dataset-examples
This dataset contains the top 64 highest activating dataset examples for each
MLP neuron in Pythia-12b. The dataset examples are all 16 tokens long. See
https://confirmlabs.org/posts/dreaming.html for details.
Columns:
layer: the layer of the neuron
neuron: the index of the neuron
rank: the rank of the example
activation: the activation of the neuron on the example
position: the token position for which the neuron is maximally activated.
text: the… See the full description on the dataset page: https://huggingface.co/datasets/Confirm-Labs/pythia-12b-neuron-dataset-examples.extraction-examples
Extraction Examples Dataset
This dataset contains 17 examples for testing extraction workflows.
Dataset Structure
Each example includes:
PDF file: Original document
map_info.json: Map extraction metadata
direction.json: Direction information
GeoJSON files: Polygon geometries
Area JSON files: Area definitions
File Organization
files/
├── example1/
│ ├── document.pdf
│ ├── map_info.json
│ ├── direction.json
│ ├── polygon1.geojson
│ └── area1.json… See the full description on the dataset page: https://huggingface.co/datasets/alexdzm/extraction-examples.voice-focus-examples
Voice Focus Examples
Collection of examples for accurate foreground speaker transcription.
Details
Curated by: Joschka Wohlgemuth
Funded by: ai-coustics GmbH
Contact:
Web: https://ai-coustics.com
opengloss-v2.3-examples
OpenGloss v2.3 — Examples
Every example sentence in OpenGloss v2.3, one row at a time, each tagged to the sense it illustrates and carrying the [span_start, span_end) character offsets of the headword occurrence inside it. That combination — a sentence, the sense it uses, and where the word is — is what a word-in-context or sense-disambiguation task needs and is normally paid for by annotation. source distinguishes the per-sense examples stage's verified sentences from… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-examples.synthetic-abandoned-cart-email-examples
Synthetic Abandoned Cart Email Examples
An entirely synthetic, bilingual collection of abandoned-cart email drafts with transparent checklist annotations. It is intended for education, prototyping, and evaluation, and contains no real recipients, customer messages, orders, merchant data, or campaign results.
Dataset Description
The dataset mirrors the five visible checks in NeuroCheckout's public Abandoned Cart Email Checker:
message clarity;
primary call to… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-abandoned-cart-email-examples.opengloss-v1.3-contrastive-examples
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Contrastive Examples v1.3
Dataset Summary
OpenGloss Contrastive Examples is a synthetic dataset of graduated semantic variations
designed for contrastive learning and… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-contrastive-examples.github-patches-genesys-agentless-prompt-2k-context-1k-diff-2k-examplesml-interview-examples-mm-imdbaffect_of_removing_misalligned_examples-conservative_qual_removedCrosscoder-Qwen2.5-1.5B-vs-DeepScaleR-1.5B_max_activating_examplesSee Files and versions for pickled dictionaries and database versions of of max activating examples organized per available layer, as well as dataframes of available features.
repro-fuse-full-spectrum-unlearnable-examples-via-spectral-equalization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
affect_of_removing_misalligned_examples-quant_removedaffect_of_removing_misalligned_examples-qual_removedaffect_of_removing_misalligned_examples-fullmlsae-pythia-70m-deduped-x256-k32-examplesmlsae-pythia-70m-deduped-x64-k32-examplesdoc-video-1hcm-examples-aug-2024Dataset of some examples with hallucinations before and after passing through Vectara's Hallucination Correction Model. See our blogpost for details.
mlsae-pythia-160m-deduped-x32-k32-examplesmax-activating-examples-gemma-2-2b-l13-ckissane
