pptx
Datasets
All datasets matching “pptx”pptx_collection_templatespptx-format-error-200
PPTX 格式错误数据集(200 样本子集)
每个 pptx 都是单页幻灯片。perturbed_<id>.pptx 是在 original_<id>.pptx 基础上注入
格式扰动后的版本,两者一一配对,文件名中的 <id> 即样本编号。
目录
perturbed/ — 200 个含格式错误的 pptx(核心)
original/ — 200 个对应的未扰动 pptx(参考/对照)
png/ — 每个样本 original 与 perturbed 的渲染图
labels/ — 标注
MANIFEST.tsv — 全部文件的 sha256 + 字节数
标注说明
labels/perturb_info.json — 记录了具体扰动的样本,字段 perturb_types 取值为
size / position / zorder / font_size / font / italic,indices 为被改动的
shape… See the full description on the dataset page: https://huggingface.co/datasets/XINLI1997/pptx-format-error-200.pptxgym-runspptxgenjs-pi-traces
pptxgenjs Pi Agent Traces
This dataset contains raw agent trace files captured from pi coding agent sessions working on pptxgenjs presentation generation tasks.
Structure
├── *.jsonl # 185 raw pi agent trace files
├── scripts/
│ ├── convert_to_openai.py # Conversion script (raw traces -> OpenAI format)
│ ├── system_prompt.txt # System prompt template
│ ├── tools_schema.json # Tool definitions
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/sandeshrajx/pptxgenjs-pi-traces.osworld_pptx_tasksOfficeSmith-PPTX-IR
OfficeSmith PPTX IR
Synthetic bilingual business briefs paired with editable PPTX intermediate representations.
Dataset summary
This dataset is part of the OfficeSmith collection for training models to plan, build, clarify, critique, and repair editable business presentations. It contains observable outputs only: no hidden chain of thought, secret benchmark prompt, personal data, or API credential is included.
Train rows: 862
Validation rows: 48
Test rows: 48… See the full description on the dataset page: https://huggingface.co/datasets/Benitoow/OfficeSmith-PPTX-IR.
