datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026-09-09-nonmoral-stakes-development
Nonmoral stakes development: stopped appended-wrapper attempt and four prospective integrated craft pairs
field
value
experiment
Nonmoral stakes development: stopped appended-wrapper attempt and four prospective integrated craft pairs
date_generated
2026-09-09
constitution
none; nonmoral craft preferences, no moral constitution
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @ 3f5a0b8de74fe06cb7db155650f750ec36451ea7
models… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-09-nonmoral-stakes-development.opus-4.6-frontend-development
CoT Code Debugging Dataset
Synthetic code debugging examples with chain-of-thought (CoT) reasoning and solutions, built with a three-stage pipeline: seed problem → evolved problem → detailed solve. Topics emphasize frontend / UI engineering (CSS, React, accessibility, layout, design systems, SSR/hydration, and related product UI issues).
Each line in dataset.jsonl is one JSON object (JSONL format).
Data fields
Field
Description
id
16-character hex id:… See the full description on the dataset page: https://huggingface.co/datasets/glyphsoftware/opus-4.6-frontend-development.gpt-5.4-frontend-development-11062026
GPT-5.4 Frontend Development Dataset (11062026)
This dataset is a synthetic chat-formatted code dataset focused on frontend development tasks in React and TypeScript.
It contains 1032 JSONL records collected on 2026-06-11 and generated with GPT-5.4 from frontend-oriented prompts covering reusable UI, compact feature units, forms, widgets, and related interface implementation tasks.
Overview
Each record contains:
task_id - numeric task identifier
category - task… See the full description on the dataset page: https://huggingface.co/datasets/runanlab/gpt-5.4-frontend-development-11062026.2026-09-09-nonmoral-paired-development
Nonmoral comparative-versus-construction reasoning development; failed scaling gate
field
value
experiment
Nonmoral comparative-versus-construction reasoning development; failed scaling gate
date_generated
2026-09-09
constitution
none; nonmoral task preferences only
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @ 9079478276735a3dbd5517bd485d1f742ddc46f0
models
Teacher/provider/revision and sampling details recorded in each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-09-nonmoral-paired-development.gazet-dataset
Gazet Dataset
Synthetic training data for finetuning small language models on geospatial tasks over Overture Maps and Natural Earth parquet datasets.
Tasks
SQL generation (sql/)
Input: user query + fuzzy-matched candidate entities (CSV)Output: DuckDB spatial SQL query
Place extraction (places/)
Input: natural language queryOutput: structured JSON with place names, country codes, and subtypes
Format
Each JSONL row is a conversation in… See the full description on the dataset page: https://huggingface.co/datasets/developmentseed/gazet-dataset.wmt24pp-ce
WMT24++ Reference Translations for Chechen
Description
WMT24++ benchmark in Chechen
Original WMT24++ benchmark (55 languages): https://huggingface.co/datasets/google/wmt24pp
The reference translation have been created by human translator based on the Russian version of WMT24++
The dataset uses both cases of Cyrillic Palochka Letter where it is grammatically correct.
For preparation of the Chechen version of the dataset we hired a professional native speaker… See the full description on the dataset page: https://huggingface.co/datasets/NM-development/wmt24pp-ce.gpt-5.4-frontend-development-27052026
Site Coding Dataset
Site Coding Dataset is a synthetic chat-formatted dataset for code generation, focused on frontend development, UI implementation, and instruction-following programming tasks.
Site Coding Dataset — синтетический датасет в chat-формате для генерации кода, сфокусированный на frontend-разработке, UI-реализации и instruction-following задачах программирования.
Overview
This dataset contains 834 records in JSONL format.Each record includes:
category — task… See the full description on the dataset page: https://huggingface.co/datasets/runanlab/gpt-5.4-frontend-development-27052026.Sustainable_Development_Goals_QA_V2
Dataset Description
This dataset generated by using 'gemini-2.5-flash' on 100 PDF publication documents coming from official website.
schemaforge-ai-research-and-development-9
huggingface.co
Auto-refined by SchemaForge
Metadata
Topic: AI Research and Development
Quality Score: 0.95
Source: Autonomous web scraper
Extracted Facts
The Fast Gemma Challenge is a verified-SOTA recipe
Training a coding agent using the OpenCode harness
Lattice is an 8 MB static retriever that embeds Wikipedia in 7 minutes
Model Genome fingerprints whether an LLM was trained from scratch or derived
LFM2.5-Encoders enable fast long-context… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-ai-research-and-development-9.SO-Python_QA-Web_Development_classschemaforge-ai-research-and-development-8
huggingface.co
Auto-refined by SchemaForge
Metadata
Topic: AI Research and Development
Quality Score: 0.95
Source: Autonomous web scraper
Extracted Facts
The Fast Gemma Challenge is a verified-SOTA recipe
Training a coding agent using the OpenCode harness
Lattice is an 8 MB static retriever that embeds Wikipedia in 7 minutes
Model Genome fingerprints whether an LLM was trained from scratch or derived
LFM2.5-Encoders enable fast long-context… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-ai-research-and-development-8.Sustainable_Development_Goals_QA
Dataset Description
This dataset generated by using 'gemini-2.5-flash' on 100 PDF publication documents coming from official website.
africa-development-risk-index-dataset
