datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
editorai-telemetryramanv-image-real-style-editorialcodeforces-editorial-full-2026-07-16
Codeforces Editorial Full - 2026-07-16
A deterministic Plan-CRL-compatible materialization of open-r1/codeforces, pinned to revision fbe3f6e903ee854eec2e69e9d96d0306cde59baf.
Size
train: 9556 problems
test: 468 problems
total: 10024 unique problems
safe fixed-output Plan-CRL evaluation rows: 1479
safe fixed-output rows with an editorial: 646
The upstream snapshot contains 10,024 unique problems, not 11,000+. The larger counts sometimes quoted for this corpus… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/codeforces-editorial-full-2026-07-16.sgu-editorial
ACM SGU Competitive Programming Solutions with LLM Enhancement
This dataset contains solutions to ACM SGU (Saratov State University) competitive programming problems, enhanced with detailed editorials and reasoning explanations generated using advanced language models. The full page about the project is here.
Overview
The dataset consists of two main components:
Original Solutions: Competitive programming solutions to SGU problems in C++ or Python.
Enhanced… See the full description on the dataset page: https://huggingface.co/datasets/radoslav11/sgu-editorial.marbella-price-data
Marbella price data
Open datasets from Marbella Wire, an independent price tracker
for Marbella, Spain. Every figure here is published on marbellawire.com first; this
dataset is the machine-readable copy, regenerated from the same data files that render
the pages. It mirrors the canonical GitHub repository
shiftdylson1/marbella-price-data.
Config
What it is
Rows
Canonical page
sunbed-index
High-season price of two sunbeds plus minimum spend at Marbella beach venues… See the full description on the dataset page: https://huggingface.co/datasets/editorwire11/marbella-price-data.codeforces-editorial-elo-512-2026-04-28
Codeforces Editorial ELO 512 - 2026-04-28
A 512-example subset sampled from open-r1/codeforces (verifiable, train) for Plan-CRL Codeforces feedback experiments that need a non-empty trusted reference field.
Important: open-r1/codeforces does not expose code reference solutions. This subset fills reference_solution from the source dataset's editorial field. Treat it as a natural-language editorial/rationale, not canonical reference code.
Selection seed: 20260428.
Filtering:
rating… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/codeforces-editorial-elo-512-2026-04-28.codecontests_editorials
Competitive Programming Editorials Dataset
This dataset provides a collection of competitive programming problem editorials sourced from Codeforces. Each entry includes essential details about the problem, such as tags, difficulty rating, and a comprehensive editorial written in Markdown format for clear readability.
Dataset Details
The dataset includes the following fields:
name: The title or name of the problem.
tags: Relevant tags associated with the problem, like… See the full description on the dataset page: https://huggingface.co/datasets/HoangLe1312/codecontests_editorials.tempura-editor-sessions
Tempura Editor Sessions
v0 pilot · professional video-editing sessions, screen and input paired frame to event
A real frame from this dataset: a paid professional editor working in Adobe Premiere Pro, captured by the Tempura recorder.
Tempura pays skilled video editors to record their real working sessions inside approved editor applications, then pairs what was on screen with what the editor did, moment by moment.
That pairing does not exist on the public internet… See the full description on the dataset page: https://huggingface.co/datasets/tempura-localhost/tempura-editor-sessions.luxury-watch-editorial-dataset-v1
license: apache-2.0
tags:
- watch-editorial
- luxury-watches
- fine-tuning
- instruction-tuning
- horological
language: en
size_categories:
- 100<n<1K
SWELOL Luxury Watch Editorial Dataset v1
Dataset Description
High-quality human-annotated luxury watch editorial descriptions for fine-tuning language models. Created by sweelol for production-grade watch content generation.
Version: 1.0License: Apache 2.0Total Examples: 308 (44 original + 264… See the full description on the dataset page: https://huggingface.co/datasets/sweelol/luxury-watch-editorial-dataset-v1.research-article-template-editor-dataSCOTT-COFFEE-Editorresearch-article-template-editor-copy-test-dataai-editor-training-datanotebook-editor-mix-v1
jmlbeaujour/notebook-editor-mix-v1
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('jmlbeaujour/notebook-editor-mix-v1')
Editor-Seoretro-weave-agent-editor-repair-diffs-v0.1
RetroInstruct Weave Agent Editor Repair Diffs
This component of RetroInstruct trains weave-agent to use the WeaveEditor to fix synthetic corruptions in the vein of
the Easy Prose Repair Diffs component.
Each row in the dataset provides the pieces you need to make a synthetic episode
demonstrating the agent:
Singling out one of three files as corrupted and in need of repair
Writing out a patch to the file as either a series of WeaveEditor edit() commands or a unidiff
Observing the… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-weave-agent-editor-repair-diffs-v0.1.droidnexus-arabic-english-editorial-retrieval-mini
DroidNexus Arabic-English Editorial Retrieval Mini
A compact public query-to-target set built from live DroidNexus coverage to test bilingual editorial retrieval, cross-language discovery, and workflow-aware search ranking.
Why this exists
This dataset is the public artifact layer for DroidNexus Labs. It turns the site's bilingual retrieval analysis into a compact benchmark for editorial search, cross-language discovery, and workflow-aware ranking experiments.… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/droidnexus-arabic-english-editorial-retrieval-mini.editorai-training-datadroidnexus-arabic-editorial-speech-scorecard-mini
DroidNexus Arabic Editorial Speech Scorecard Mini
A public DroidNexus Labs scorecard dataset for Arabic speech workflows: representative editorial scenarios, latency targets, overlap pressure, and the metric stack that decides whether a transcript is usable.
Why this exists
This dataset is the first public speech artifact layer for DroidNexus Labs. It publishes representative editorial workloads and evaluation pressure before claiming a full source-audio benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/droidnexus-arabic-editorial-speech-scorecard-mini.collab-editor-dataCOFFEEGYM-Editoreditor-datasetresearch-article-template-editor-test2-dataRenmingNet_Editorialwebllm-json-editorflan_combined_task522_news_editorial_summaryeditorai-dataeditorai-adv-2.1.8editor-token-min-v2COFFEE-Editor-sample
