datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multitask_german_examples_32kopengloss-v2.1-examples
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Examples
Every example sentence in OpenGloss v2.1, one row at a time, each tagged to the sense it illustrates and carrying… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-examples.opengloss-v2.2-examples
Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility.
OpenGloss v2.2 — Examples
Every example sentence in OpenGloss v2.2, one row at a time, each tagged to the sense it illustrates and carrying the [span_start, span_end) character offsets of the headword occurrence… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-examples.iclr-eval-examplesopengloss-v2.0-examples
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Examples
Every example sentence in OpenGloss v2.0, one row at a time, each tagged to the sense it illustrates and carrying the [span_start, span_end) character offsets of the headword occurrence inside it. That combination — a sentence, the sense… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-examples.affect_of_removing_misalligned_examples-eval_resultspythia-12b-neuron-dataset-examples
pythia-12b-neuron-dataset-examples
This dataset contains the top 64 highest activating dataset examples for each
MLP neuron in Pythia-12b. The dataset examples are all 16 tokens long. See
https://confirmlabs.org/posts/dreaming.html for details.
Columns:
layer: the layer of the neuron
neuron: the index of the neuron
rank: the rank of the example
activation: the activation of the neuron on the example
position: the token position for which the neuron is maximally activated.
text: the… See the full description on the dataset page: https://huggingface.co/datasets/Confirm-Labs/pythia-12b-neuron-dataset-examples.opengloss-v2.3-examples
OpenGloss v2.3 — Examples
Every example sentence in OpenGloss v2.3, one row at a time, each tagged to the sense it illustrates and carrying the [span_start, span_end) character offsets of the headword occurrence inside it. That combination — a sentence, the sense it uses, and where the word is — is what a word-in-context or sense-disambiguation task needs and is normally paid for by annotation. source distinguishes the per-sense examples stage's verified sentences from… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-examples.voice-focus-examples
Voice Focus Examples
Collection of examples for accurate foreground speaker transcription.
Details
Curated by: Joschka Wohlgemuth
Funded by: ai-coustics GmbH
Contact:
Web: https://ai-coustics.com
github-patches-genesys-agentless-prompt-2k-context-1k-diff-2k-examplesaffect_of_removing_misalligned_examples-conservative_qual_removedmlsae-pythia-70m-deduped-x256-k32-examplesmlsae-pythia-70m-deduped-x64-k32-examplesmlsae-pythia-160m-deduped-x32-k32-examples10_examples_diff_tempdribble-examples
Dataset Card for "dribble-examples"
More Information needed
form-lang-examplesexample_storage_resultsexample_so101This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Denryy/example_so101.mlsae-pythia-70m-deduped-x32-k32-examplesmlsae-pythia-70m-deduped-x64-k128-examplesexample_storage_configmlsae-pythia-70m-deduped-x64-k16-examplesmlsae-pythia-70m-deduped-x64-k64-examplesaffect_of_removing_misalligned_examples-quant_removedmlsae-pythia-70m-deduped-x64-k256-examplesflc-examplesatomicmath-lineage-10-examples-testmlsae-pythia-160m-deduped-x1-k32-examplesaffect_of_removing_misalligned_examples-full
