datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-5.4-frontend-development-11062026
GPT-5.4 Frontend Development Dataset (11062026)
This dataset is a synthetic chat-formatted code dataset focused on frontend development tasks in React and TypeScript.
It contains 1032 JSONL records collected on 2026-06-11 and generated with GPT-5.4 from frontend-oriented prompts covering reusable UI, compact feature units, forms, widgets, and related interface implementation tasks.
Overview
Each record contains:
task_id - numeric task identifier
category - task… See the full description on the dataset page: https://huggingface.co/datasets/runanlab/gpt-5.4-frontend-development-11062026.gazet-dataset
Gazet Dataset
Synthetic training data for finetuning small language models on geospatial tasks over Overture Maps and Natural Earth parquet datasets.
Tasks
SQL generation (sql/)
Input: user query + fuzzy-matched candidate entities (CSV)Output: DuckDB spatial SQL query
Place extraction (places/)
Input: natural language queryOutput: structured JSON with place names, country codes, and subtypes
Format
Each JSONL row is a conversation in… See the full description on the dataset page: https://huggingface.co/datasets/developmentseed/gazet-dataset.gpt-5.4-frontend-development-27052026
Site Coding Dataset
Site Coding Dataset is a synthetic chat-formatted dataset for code generation, focused on frontend development, UI implementation, and instruction-following programming tasks.
Site Coding Dataset — синтетический датасет в chat-формате для генерации кода, сфокусированный на frontend-разработке, UI-реализации и instruction-following задачах программирования.
Overview
This dataset contains 834 records in JSONL format.Each record includes:
category — task… See the full description on the dataset page: https://huggingface.co/datasets/runanlab/gpt-5.4-frontend-development-27052026.AI-Ethical-Development-Kazakh-Focused
🇰🇿 AI Ethical Development, Kazakh-Focused
📖 Overview
AI Ethical Development, Kazakh-Focused is a specialized dataset designed to align Large Language Models (LLMs) with the cultural, ethical, and legal frameworks of Kazakhstan.
Each sample presents a culturally nuanced scenario (the request) and provides two possible answers:
Accepted (Chosen): A response that balances traditional Kazakh values (e.g., respect for elders, "aga-ini" relations) with modern legal… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/AI-Ethical-Development-Kazakh-Focused.Succession_Planning_Talent_Development_Practical
Succession Planning Talent Development — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Succession_Planning_Talent_Development_Practical.Succession_Planning_Talent_Development_Theory
Succession Planning Talent Development — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Succession_Planning_Talent_Development_Theory.Thesis_Development_of_a_Complex_of_Neural_Networks_for_Linked_Generation_of_Large_TextsHere presented a partially synthesized dataset, developed utilizing the GPT-4 model, for the purpose of NLG, particulary for the task of hierarchical generation of longer texts from short summaries. The creation of this dataset was undertaken as a component of my thesis paper. It incorporates excerpts from prominent British and American novels, from which plots, summaries, and metadata have been derived using GPT-4 API to facilitate extensive future research.
The metadata included in the… See the full description on the dataset page: https://huggingface.co/datasets/Fleur-roar/Thesis_Development_of_a_Complex_of_Neural_Networks_for_Linked_Generation_of_Large_Texts.
