datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
doc-formats-jsonl-1
[doc] formats - jsonl - 1
This dataset contains one jsonl file at the root.
chat_formatted_examplessounio-code-examples
Sounio Curated Code Examples
Curated compile-clean .sio examples for training and evaluating code models on
Sounio, a self-hosted systems and scientific programming language for epistemic
computing, uncertainty propagation, and algebraic effects.
This directory is the Cx-1 expansion lane for
chiuratto-AIgourakis/sounio-code-examples.
Current batch
Examples: 5,000
Metadata files: 5,000
Compiler gate: bin/souc check pass rate 5,000/5,000
Utility layer: 5,000… See the full description on the dataset page: https://huggingface.co/datasets/chiuratto-AIgourakis/sounio-code-examples.opengloss-v1.3-query-examples-flat
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.example-space-to-dataset-jsonDemo to save data from a Space to a Dataset. Goal is to provide reusable snippets of code.
Documentation: https://huggingface.co/docs/huggingface_hub/main/en/guides/upload#scheduled-uploads
Space: https://huggingface.co/spaces/Wauplin/space_to_dataset_saver/
JSON dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json
Image dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-image
Image (zipped) dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json.OpenSeek-Synthetic-Reasoning-Data-Examples
OpenSeek-Reasoning-Data
OpenSeek [Github|Blog]
Recent reseach has demonstrated that the reasoning ability of LLMs originates from the pre-training stage, activated by RL training. Massive raw corpus containing complex human reasoning process, but lack of generalized and effective synthesis method to extract these reasoning process.
News
🔥🔥🔥[2025/02/25] We publish some math, code, and general knowledge domain reasoning data synthesized from the current pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples.sycophancy_examples
Sycophancy Examples
Two sycophancy evaluation datasets from Kei et al., "Reward hacking can generalise across settings".
Original source: GeodesicResearch/Obfuscation_Generalization
Files
File
Examples
Description
sycophancy_opinion_political.jsonl
5,000
Political opinion questions with persona-aligned "sycophantic" answers
sycophancy_fact.jsonl
401
Factual questions where the persona holds a misconception; sycophantic answer agrees with the misconception… See the full description on the dataset page: https://huggingface.co/datasets/camgeodesic/sycophancy_examples.hydro_cali_agent_exampleflutter-full-examples-v1
Flutter Codegen: Full Examples
Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal
that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps,
there's no step history or diff structure here -- each row is a single, standalone
goal -> complete file example.
This is the whole-code counterpart to flutter-diff-steps-v1, intended for
training/evaluating a baseline that generates the entire file in one shot, to
compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.cursor-traces-exampleThis dataset was generated using teich by TeichAI
My Agent Traces
This directory contains raw agent trace files generated by teich.
JSONL files: 9
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/cursor-traces-example.agent-traces-examplechinese_law_examples
1000 examples of law items
law_item.jsonl contains 1000 samples of current and effective Chinese laws. e.g.
{"title": "《中华人民共和国劳动合同法(2012修正)》",
"classification": "类别 : 劳动合同营商环境优化 ",
"num": "第十九条",
"contents": "第十九条【试用期】劳动合同期限三个月以上不满一年的,试用期不得超过一个月;劳动合同期限一年以上不满三年的,试用期不得超过二个月;三年以上固定期限和无固定期限的劳动合同,试用期不得超过六个月。同一用人单位与同一劳动者只能约定一次试用期。以完成一定工作任务为期限的劳动合同或者劳动合同期限不满三个月的,不得约定试用期。试用期包含在劳动合同期限内。劳动合同仅约定试用期的,试用期不成立,该期限为劳动合同期限。"}
Using BGE Embedding to compute similarity between query and… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/chinese_law_examples.spider-rollouts-web-search-qwen2.5-7b-gaia-32-examplesexample_datasetsynthetic-product-recommendation-examples
Synthetic Product Recommendation Examples
An entirely synthetic bilingual dataset of ecommerce discovery queries paired with candidate products and graded relevance judgments. It is designed for educational retrieval, reranking, and recommendation experiments and contains no private catalog, merchant, customer, behavioral, or transaction data.
Dataset Description
The dataset contains twenty English and French queries with ten candidates per query. It complements… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-product-recommendation-examples.LORE-examples
LORE Examples
A small set of matched multimodal examples from LORE, for the
MIMIC model — enough to try inference,
embedding, and generation across DNA, RNA, and protein modalities without wiring
up your own data.
Each example is a single biological entity (a transcript and/or its protein) with
several co-observed modalities. Rows are drawn from the held-out (validation) split
of MIMIC's training data, so they are in-distribution and length-bounded to the
model's context… See the full description on the dataset page: https://huggingface.co/datasets/polymathic-ai/LORE-examples.llm-behavioral-drift-examples
LLM-Behavioral-Drift-Examples
Training examples for SFT used to induce behavioral drift.
Dataset Description
Various training examples in an AI office assistant setting (email, calendar, docs, and misc. info.) skewed in particular ways (aggression, irrelevancy/tangential information, and excessive verbosity).
Examples produced by Gemini 2.5 Flash.
Example Usage
aggressive_dataset = load_dataset(
f"{username}/{repo_name}",
data_files="aggressive.jsonl"… See the full description on the dataset page: https://huggingface.co/datasets/6S-bobby/llm-behavioral-drift-examples.example_dataset-raw
EEG Dataset
This dataset was created using braindecode, a deep
learning library for EEG/MEG/ECoG signals.
Dataset Information
Property
Value
Recordings
1
Type
Continuous (Raw)
Channels
26
Sampling frequency
250 Hz
Total duration
0:06:26
Windows/samples
96,735
Size
19.22 MB
Format
zarr
Quick Start
from braindecode.datasets import BaseConcatDataset
# Load from Hugging Face Hub
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/braindecode/example_dataset-raw.synthetic-abandoned-cart-email-examples
Synthetic Abandoned Cart Email Examples
An entirely synthetic, bilingual collection of abandoned-cart email drafts with transparent checklist annotations. It is intended for education, prototyping, and evaluation, and contains no real recipients, customer messages, orders, merchant data, or campaign results.
Dataset Description
The dataset mirrors the five visible checks in NeuroCheckout's public Abandoned Cart Email Checker:
message clarity;
primary call to… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-abandoned-cart-email-examples.dictation-cleanup-examples
Dictation cleanup examples
A sample of the hand-written cases behind
SpeakoFlow Mini, published so the
conventions the model follows are inspectable rather than described.
Seven cases in each of fifteen categories, spread across short, medium and long transcripts.
Every case was written by hand. None of it is captured speech.
This is not a benchmark
Read that before using it for anything.
These cases are drawn from the training pool, not from the held-out set the… See the full description on the dataset page: https://huggingface.co/datasets/SpeakoFlow/dictation-cleanup-examples.opengloss-v1.3-contrastive-examples
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Contrastive Examples v1.3
Dataset Summary
OpenGloss Contrastive Examples is a synthetic dataset of graduated semantic variations
designed for contrastive learning and… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-contrastive-examples.georgian-booking-safety-examples
Synthetic Georgian booking-safety examples
Twelve hand-annotated design examples. Not an ASR/NLU benchmark, a training
corpus, or a measurement of any deployed product.
This package makes the OMO AI educational examples available as a
small, flat JSONL table with a complete, lossless copy of each original case.
The source material and this packaging were prepared with AI assistance.
All utterances, identities, times, state and destination responses are fictional.
There are no… See the full description on the dataset page: https://huggingface.co/datasets/vajelski/georgian-booking-safety-examples.chinese_verdict_examples
verdicts examples
verdicts_200.jsonl contains 200 examples of verdicts from Chinese Judgements Online, we process the datasets for semantic retrieval
using BGE to compute similarity between query and verdict
from FlagEmbedding import FlagModel
from datasets import load_dataset
dataset = load_dataset("FarReelAILab/verdicts")
model = FlagModel('BAAI/bge-large-zh-v1.5',
query_instruction_for_retrieval="为这个句子生成表示以用于检索相关文章:",
use_fp16=True)… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/chinese_verdict_examples.Chinese-DeepSeek-V3.2-Exp-chat-example
deepseek/deepseek-v3.2-exp (6.6K) 中文数据集样本
一、前言
本报告基于 deepseek/deepseek-v3.2-exp 模型(官方 API,8K 上下文窗口)进行数据集评测与可视化展示。测试数据集共包含 6,655 轮对话,语言覆盖以中文为主,辅以部分混合语种及非中文输入。本次报告旨在总结模型的对话特征、输入输出长度分布及上下文预算消耗情况,并为后续应用和优化提供参考。
二、数据与方法
数据来源:用户构建的 6,655 轮真实中文对话样本。
估算方法:
中文字符近似为 1 Token;
英文 4 字符 ≈ 1 Token;
用于规模与上下文预算对比,而非精确 Token 计数。
统计维度:
平均 Prompt/Output 长度(字符与估算 Token);
总 Token 占上下文窗口比例;
语言分布(Prompt 语言类型);
对话长度分布(用户提问、助手回答、总对话长度)。
三、总体结果
1. 样本概况… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-DeepSeek-V3.2-Exp-chat-example.opengloss-v1.3-query-examples
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples.fewshot-examplesliterary-genre-examples
Literary Genre Dataset
This dataset contains a curated list of 86 fiction and nonfiction genres, each accompanied by a representative example paragraph. The example texts illustrate the typical tone, writing style, and content characteristics for each genre.
Genres Covered: 86 total, spanning popular and niche categories in both fiction and nonfiction.
Genre Types: Marked as either Fiction or Nonfiction.
Example Paragraphs: Each genre includes a sample paragraph written to capture… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/literary-genre-examples.repro-fuse-full-spectrum-unlearnable-examples-via-spectral-equalization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
doc-formats-json-1
[doc] formats - json - 1
This dataset contains one json file at the root. It's a list of rows, each of which is a dict of columns.
tisus_mcq_example_exam
