datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
backend-code-generator-dataset
Backend Code Generation Dataset
Dataset Description
This dataset contains examples for training AI models to generate backend application code. It includes descriptions of backend requirements paired with complete, functional code implementations across multiple frameworks and programming languages.
Dataset Summary
The Backend Code Generation Dataset is designed to train models that can generate complete backend applications from natural language descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/Techta/backend-code-generator-dataset.GPT2-Hacker-password-generator-dataset
Hacker Style Password Generation Dataset
Dataset Description
This dataset contains 20,000 instruction-response pairs designed to train and evaluate language models for generating strong, "hacker-style" passwords. The data simulates a user requesting a secure password and the model providing a complex, randomly generated string.
Supported Tasks
Text Generation: The primary task is conditional text generation, where the model takes a natural language instruction… See the full description on the dataset page: https://huggingface.co/datasets/CodeferSystem/GPT2-Hacker-password-generator-dataset.ai-auto-train-datasets-cuda-5d-quantum-mindmap-simulations-generator-zkevms-immutablexUML-Generator-Dataset-DeepSeek-V3.2Click here to support our open-source dataset and model releases!
UML-Generator-Dataset-DeepSeek-V3.2 is a dataset focused on analysis and code-reasoning, creating UML diagrams testing the limits of DeepSeek V3.2's modeling and design skills!
This dataset contains:
2.7k synthetically generated prompts to create UML diagrams in response to user input, with all responses generated using DeepSeek V3.2.
All responses contain a multi-step thinking process to perform effective analysis, followed by… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/UML-Generator-Dataset-DeepSeek-V3.2.crispr-cas-atlas-generator
CRISPR-Cas Atlas – GENERator-ready Dataset
This dataset is a derived, preprocessed version of the CRISPR-Cas Atlas, formatted for causal language model fine-tuning with GENERator.
Source
Original dataset:CRISPR-Cas Atlas v1.0
Ruffolo et al., Design of highly functional genome editors by modeling the universe of CRISPR-Cas sequences, bioRxiv (2024)https://www.biorxiv.org/content/10.1101/2024.04.22.590591v1
Processing
Each CRISPR-Cas operon was converted into a… See the full description on the dataset page: https://huggingface.co/datasets/metaXu264/crispr-cas-atlas-generator.smolified-roadmap-generator
🤏 smolified-roadmap-generator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model IndrajitAri/smolified-roadmap-generator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 390b2924)
Records: 244
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by IndrajitAri.
Generated via Smolify.ai.
smolified-sample-email-generator
🤏 smolified-sample-email-generator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-sample-email-generator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 6e8b8fbf)
Records: 244
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
QA-Dataset-Generator
RAG Scientific QA Dataset (Generated)
Dataset Description
This dataset contains 711 high-quality Question-Answering pairs synthetically generated from ArXiv scientific papers. It is specifically designed to fine-tune Large Language Models (LLMs) for Retrieval-Augmented Generation (RAG) tasks.
Source Data: 200 ArXiv papers (Computer Science: AI, CL, LG, IR).
Generation Method: Generated using gpt-4o-mini with strict rules to prevent hallucination.
Language:… See the full description on the dataset page: https://huggingface.co/datasets/xunnhi/QA-Dataset-Generator.question-generator-dataset
Dataset Card for question-generator-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/rjt1221/question-generator-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/rjt1221/question-generator-dataset.cognitive-question-generator-v1
Cognitive Question Generator Dataset
Dataset for fine-tuning an expert analysis and question generation model. Contains 5,637 prompt-response pairs capturing expert reasoning patterns for technology transactions and product counseling.
Dataset Description
This dataset was generated from the CognitiveTrainer platform's Mode 1 (Expert Analysis) system, capturing:
Initial scenario analysis
Claim validation with chain-of-trust
Multi-turn expert dialogue
Final synthesis… See the full description on the dataset page: https://huggingface.co/datasets/KevinKeller/cognitive-question-generator-v1.smolified-gdg-monthly-meetup-idea-generator
🤏 smolified-gdg-monthly-meetup-idea-generator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Arban221B/smolified-gdg-monthly-meetup-idea-generator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 64321ab7)
Records: 840
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Arban221B.
Generated via Smolify.ai.
smolified-discharge-summary-generator
🤏 smolified-discharge-summary-generator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-discharge-summary-generator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: b0b64157)
Records: 8525
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
smolified-file-context-generator
🤏 smolified-file-context-generator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Sayan25/smolified-file-context-generator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: d402c789)
Records: 5325
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Sayan25.
Generated via Smolify.ai.
smolified-context-aware-travel-dataset-generator
🤏 smolified-context-aware-travel-dataset-generator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-context-aware-travel-dataset-generator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 6e4879c0)
Records: 9960
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via… See the full description on the dataset page: https://huggingface.co/datasets/smolify/smolified-context-aware-travel-dataset-generator.smolified-context-aware-travel-dataset-generator
🤏 smolified-context-aware-travel-dataset-generator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-context-aware-travel-dataset-generator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 6e4879c0)
Records: 9960
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via… See the full description on the dataset page: https://huggingface.co/datasets/Tanika2004/smolified-context-aware-travel-dataset-generator.
