datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flutter-full-examples-v1
Flutter Codegen: Full Examples
Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal
that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps,
there's no step history or diff structure here -- each row is a single, standalone
goal -> complete file example.
This is the whole-code counterpart to flutter-diff-steps-v1, intended for
training/evaluating a baseline that generates the entire file in one shot, to
compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.agent-traces-exampleexample_datasetexample_dataset-raw
EEG Dataset
This dataset was created using braindecode, a deep
learning library for EEG/MEG/ECoG signals.
Dataset Information
Property
Value
Recordings
1
Type
Continuous (Raw)
Channels
26
Sampling frequency
250 Hz
Total duration
0:06:26
Windows/samples
96,735
Size
19.22 MB
Format
zarr
Quick Start
from braindecode.datasets import BaseConcatDataset
# Load from Hugging Face Hub
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/braindecode/example_dataset-raw.synthetic-abandoned-cart-email-examples
Synthetic Abandoned Cart Email Examples
An entirely synthetic, bilingual collection of abandoned-cart email drafts with transparent checklist annotations. It is intended for education, prototyping, and evaluation, and contains no real recipients, customer messages, orders, merchant data, or campaign results.
Dataset Description
The dataset mirrors the five visible checks in NeuroCheckout's public Abandoned Cart Email Checker:
message clarity;
primary call to… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-abandoned-cart-email-examples.opengloss-v1.3-contrastive-examples
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Contrastive Examples v1.3
Dataset Summary
OpenGloss Contrastive Examples is a synthetic dataset of graduated semantic variations
designed for contrastive learning and… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-contrastive-examples.Chinese-DeepSeek-V3.2-Exp-chat-example
deepseek/deepseek-v3.2-exp (6.6K) 中文数据集样本
一、前言
本报告基于 deepseek/deepseek-v3.2-exp 模型(官方 API,8K 上下文窗口)进行数据集评测与可视化展示。测试数据集共包含 6,655 轮对话,语言覆盖以中文为主,辅以部分混合语种及非中文输入。本次报告旨在总结模型的对话特征、输入输出长度分布及上下文预算消耗情况,并为后续应用和优化提供参考。
二、数据与方法
数据来源:用户构建的 6,655 轮真实中文对话样本。
估算方法:
中文字符近似为 1 Token;
英文 4 字符 ≈ 1 Token;
用于规模与上下文预算对比,而非精确 Token 计数。
统计维度:
平均 Prompt/Output 长度(字符与估算 Token);
总 Token 占上下文窗口比例;
语言分布(Prompt 语言类型);
对话长度分布(用户提问、助手回答、总对话长度)。
三、总体结果
1. 样本概况… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-DeepSeek-V3.2-Exp-chat-example.repro-fuse-full-spectrum-unlearnable-examples-via-spectral-equalization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
steer_place_yellow_dot_red_arrow_example_ep201
Placement Yellow-Dot Red-Arrow Sanity Dataset
This is a LeRobot-format one-episode sanity export derived from local HDF5 placement data.
It uses source recording episode_00201.hdf5, one of the five episodes newer
than the existing 197-episode placement export.
Dataset size:
episodes: 1
frames: 259
videos: 3
export fps: 100
frame stride from 100 Hz source: 1
source episode index: 201
Source goal label:
full-resolution target: (564.0, 223.0) px
224x224 overlay target: (98.7… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/steer_place_yellow_dot_red_arrow_example_ep201.hf-sts-doc-example
HF STS doc example (verbatim control)
Exact JSONL from https://huggingface.co/docs/hub/session-traces-format
proof-adjusted-autonomy-examples
Proof-Adjusted Autonomy (PAA) — Worked Examples
12 worked scenarios of the PAA metric coined by Michał Piszczek: the share of completed work an AI system executes autonomously AND supports with independent, reliable, timely evidence.
Formula: PAA = P(A) x P(C|A) x P(R|A,C) x P(T|A,C,R)
Each row: scenario, agent type, the four gate probabilities, raw autonomy vs PAA score, the autonomy gap, and an interpretation. Includes the canonical example: a 90% agent that is a 61.6% agent.… See the full description on the dataset page: https://huggingface.co/datasets/cdiamond/proof-adjusted-autonomy-examples.opengloss-v1.1-contrastive-examples
OpenGloss Contrastive Examples v1.1
Dataset Summary
OpenGloss Contrastive Examples is a synthetic dataset of graduated semantic variations
designed for contrastive learning and semantic similarity training. Each example contains
a source sentence and a 5-point semantic gradient showing how meaning shifts from
antonym to synonym poles.
This dataset is derived from the OpenGloss
encyclopedic dictionary, using example sentences and their lexical context to generate… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-contrastive-examples.Boat_unity_exampleDeepSeek-V3.2-Exp-reasoning-example
🐳 DeepSeek-V3.2-Exp-reasoning vs DeepSeek-R1-0528: Math Reasoning Comparison 🍎
Note: DeepSeek-R1-0528 has no explicit chain-of-thought, while deepseek-ai/DeepSeek-V3.2-Exp (abbrev. V3.2-Exp) produces answers with structured derivations. This report was analyzed by GPT-5-Extended-Thinking. The sample size is small; conclusions are for reference only.
Author: Soren
1. Executive Summary
Sample size: 208 problems (mixed types).
Average steps (reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V3.2-Exp-reasoning-example.subtask-b-examples-testrandom_effect_exampleexample_dataset
example_dataset
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
Boat_unity_exampleopengloss-v1.2-contrastive-examples
OpenGloss Contrastive Examples v1.2
Dataset Summary
OpenGloss Contrastive Examples is a synthetic dataset of graduated semantic variations
designed for contrastive learning and semantic similarity training. Each example contains
a source sentence and a 5-point semantic gradient showing how meaning shifts from
antonym to synonym poles.
This dataset is derived from the OpenGloss
encyclopedic dictionary, using example sentences and their lexical context to generate… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-contrastive-examples.Boat_unity_exampleexample_dataset
example_dataset
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
Example
