datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-world-model-inference-examples-40
Inference examples
This directory contains 40 numbered, independent inference examples.
Every example uses only its public number; source case names and internal paths
are intentionally omitted.
Each numbered directory contains:
first_frame.png: exact 1536x864 generated RGB first frame used by inference.
prompts/*.txt: the exact rolling long-inference prompts used for the result.
condition/*.npz: ordered lossless condition chunks.
metadata.json: frame count, FPS, prompt windows… See the full description on the dataset page: https://huggingface.co/datasets/NTU-yiwen/code-world-model-inference-examples-40.lisbet-exampleschat_formatted_examplesdoc-formats-jsonl-1
[doc] formats - jsonl - 1
This dataset contains one jsonl file at the root.
sounio-code-examples
Sounio Curated Code Examples
Curated compile-clean .sio examples for training and evaluating code models on
Sounio, a self-hosted systems and scientific programming language for epistemic
computing, uncertainty propagation, and algebraic effects.
This directory is the Cx-1 expansion lane for
chiuratto-AIgourakis/sounio-code-examples.
Current batch
Examples: 5,000
Metadata files: 5,000
Compiler gate: bin/souc check pass rate 5,000/5,000
Utility layer: 5,000… See the full description on the dataset page: https://huggingface.co/datasets/chiuratto-AIgourakis/sounio-code-examples.doc-formats-csv-1
[doc] formats - csv - 1
This dataset contains one csv file at the root:
data.csv
kind,sound
dog,woof
cat,meow
pokemon,pika
human,hello
The YAML section of the README does not contain anything related to loading the data (only the size category metadata):
---
size_categories:
- n<1K
---
doc-formats-parquet-1opengloss-v1.3-query-examples-flat
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.example-sec-filingsexample_short_storiesdocvqa_1200_examplesmultitask_german_examples_32kcode_search_net_python_10000_exampleskabr-worked-examples
Dataset Card for KABR Worked Examples
This dataset is comprised of manually annotated bounding box detections, mini-scenes, behavior annotations, and associated telemetry
for three drone video sessions that were used for kabr-tools case studies. Drone video was collected at Mpala Research Centre in January 2023; please see the full video dataset for more information on original video context.
Dataset Details
Annotations were created to evaluate the kabr-tools… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/kabr-worked-examples.math_examplesunsplash-examplesexample-space-to-dataset-jsonDemo to save data from a Space to a Dataset. Goal is to provide reusable snippets of code.
Documentation: https://huggingface.co/docs/huggingface_hub/main/en/guides/upload#scheduled-uploads
Space: https://huggingface.co/spaces/Wauplin/space_to_dataset_saver/
JSON dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json
Image dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-image
Image (zipped) dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json.example-space-to-dataset-imageDemo to save data from a Space to a Dataset. Goal is to provide reusable snippets of code.
Documentation: https://huggingface.co/docs/huggingface_hub/main/en/guides/upload#scheduled-uploads
Space: https://huggingface.co/spaces/Wauplin/space_to_dataset_saver/
JSON dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json
Image dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-image
Image (zipped) dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-image.OpenSeek-Synthetic-Reasoning-Data-Examples
OpenSeek-Reasoning-Data
OpenSeek [Github|Blog]
Recent reseach has demonstrated that the reasoning ability of LLMs originates from the pre-training stage, activated by RL training. Massive raw corpus containing complex human reasoning process, but lack of generalized and effective synthesis method to extract these reasoning process.
News
🔥🔥🔥[2025/02/25] We publish some math, code, and general knowledge domain reasoning data synthesized from the current pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples.sycophancy_examples
Sycophancy Examples
Two sycophancy evaluation datasets from Kei et al., "Reward hacking can generalise across settings".
Original source: GeodesicResearch/Obfuscation_Generalization
Files
File
Examples
Description
sycophancy_opinion_political.jsonl
5,000
Political opinion questions with persona-aligned "sycophantic" answers
sycophancy_fact.jsonl
401
Factual questions where the persona holds a misconception; sycophantic answer agrees with the misconception… See the full description on the dataset page: https://huggingface.co/datasets/camgeodesic/sycophancy_examples.example-space-to-dataset-parquetmotion_examplesopengloss-v2.1-examples
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Examples
Every example sentence in OpenGloss v2.1, one row at a time, each tagged to the sense it illustrates and carrying… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-examples.openwebtext2-first-30-chunks-english-only-examplescs2-v3-prompt-comparison-5-examples
CS2 V3 五案例最终结果对比 / Five-case Final Result Comparison
本 README 只展示模型最终输出结果,不展示发送给 API 的 Prompt。每个案例使用同一视频:左侧为两轮结果,右侧为七轮结果。结果无需展开即可直接查看,英文原文和中文翻译同时提供并排对比。
This README shows only the model's final output results, not the API request prompts. Each case uses the same video: the two-round result is on the left and the seven-round result is on the right. Results are visible directly with no expandable sections. English and Chinese are shown together for side-by-side comparison.
Case 1… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-v3-prompt-comparison-5-examples.docvqa_1200_examples_donuticlr-eval-examplesflutter-full-examples-v1
Flutter Codegen: Full Examples
Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal
that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps,
there's no step history or diff structure here -- each row is a single, standalone
goal -> complete file example.
This is the whole-code counterpart to flutter-diff-steps-v1, intended for
training/evaluating a baseline that generates the entire file in one shot, to
compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.opengloss-v2.2-examples
Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility.
OpenGloss v2.2 — Examples
Every example sentence in OpenGloss v2.2, one row at a time, each tagged to the sense it illustrates and carrying the [span_start, span_end) character offsets of the headword occurrence… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-examples.opengloss-v2.0-examples
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Examples
Every example sentence in OpenGloss v2.0, one row at a time, each tagged to the sense it illustrates and carrying the [span_start, span_end) character offsets of the headword occurrence inside it. That combination — a sentence, the sense… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-examples.
