datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openclassgen-structured-v1
OpenClassGen Structured v1
Derived from mrahman2025/OpenClassGen (Rahman et al. 2025, arXiv:2504.15564).
License: CC BY 2.0 (same as upstream). Keep repository_name and file_path when redistributing.
Underlying GitHub repos may carry additional software licenses.
gold_code is upstream human_written_code.
We add parsed fields, body-span indices, and a Variant-3 prompt/target pair (v3_prompt_text / v3_target_text).
No unit tests. Splits are repository-disjoint (train /… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/openclassgen-structured-v1.putusan-structured-extraction
Putusan structured-extraction dataset
Built 2026-07-08T23:02:28+00:00 by notebooks/build_dataset.py (seed 3407).
Indonesian court-decision (putusan) extractive-structuring dataset over three
corpora (Anak, Asusila, TPPO). Each row is one model extraction of one source
document into 31 canonical sections of verbatim spans. Empty sections were
completed from sibling model extractions of the same document where available
(cross_model_fill_json records per-section donor provenance).… See the full description on the dataset page: https://huggingface.co/datasets/Haeryz/putusan-structured-extraction.structured-stern-neon-articles
Structured Stern NEON Community Articles
This repository contains approximately 20k user written texts,
articles, and poetry pulled from archives of the Stern NEON website.
Stern NEON was a community platform where users could write and publish their own articles.
Many of the articles are personal stories, poems, or opinion pieces.
The articles are structured in a way that they can be used for further analysis.
Dataset Details
Uses
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.stage3-synthetic-structured-retrieval
Stage 3 Synthetic Structured-Retrieval Agents
Native search-tool trajectories generated by
Qwen/Qwen3-235B-A22B-Instruct-2507 for LCLM Stage-3 agent post-training.
The default config contains only traces that passed programmatic evidence and
answer verification.
Harvest
Accepted traces: 82
Native search calls: 179
Compressed tool-observation traces: 41
Uncompressed traces: 41
Task-ID overlap between pilot and collection batch: 0
Family counts are 20 latest-state… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-synthetic-structured-retrieval.classeval-structured-v1
ClassEval Structured v1
Derived from FudanSELab/ClassEval (Du et al. 2023, arXiv:2308.01861).
License: CC BY-NC 4.0 (upstream data license). Non-commercial use only.
One row per (task_id, variant) with variant in {1,2,3} (100 tasks × 3 = 300 rows; Hub split test).
solution_code, test, and methods_info_json come from upstream.
We add rendered prompts/targets and stratification fields.
Missing bodies use ....
Variants:
Signatures and docstrings kept; every method body is ....… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/classeval-structured-v1.harmonicbench-planir-main-structured
HARMONICBench PlanIR Structured Main Dataset
This repository is a Hugging Face dataset-friendly structured export derived from the local outputs/fixed/main directory in the HARMONICBench unified package.
Included tables
plans/train.parquet: primary aggregate table converted from plan_runs_all.jsonl.
domains/*.parquet: per-domain plan runs, selected samples, and domain5 image descriptions.
summaries/*: key JSON/CSV/JSONL summary artifacts.
artifacts/roundtrip_recovery*/*:… See the full description on the dataset page: https://huggingface.co/datasets/guhhhgu/harmonicbench-planir-main-structured.
