datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stories-refinement
Stories Refinement
This dataset contains synthetic short stories generated from blog text excerpts sourced from the agentlans/lucadiliello-STORIES dataset.
The stories were produced using the agentlans/Llama3.1-LexiHermes-SuperStorm language model unless otherwise noted.
Configurations
allContaining all other configs and filtered for output < 6000 characters. This config has an additional column indicating which config each row is from.
zero-shotGenerated directly… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/stories-refinement.Dataset_Text_Refinement
Dataset Card for Dataset Name
This Dataset is for refining text based on user study text preferences. Given a original text to refined text based on given paramters.
The parameters are:
-Readability_Score
-Semantic_Coherence
-User_Preference
Readability Score:
Its the average of Flesch-Kincaid Readability Ease Score and Dale-Chall readability score.
The Readability Score is of user which is appilied on refined text.
Semantic Coherence:
Its the average float value of… See the full description on the dataset page: https://huggingface.co/datasets/SolaceinLoneSun/Dataset_Text_Refinement.refinement-abliterated-thinking_heretic
Dataset Card: Refinement-Abliterated (thinking_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io designed to train language models to process edge-case, controversial, or complex analytical prompts without triggering over-aligned corporate refusal responses.
Generation Pipeline Mechanics
Seed Matrix: Initial queries gathered from mlabonne/harmful_behaviors.
Knowledge Engine (Abliterated Base): Generated using… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-thinking_heretic.high-quality-text-refinementprompt-refinement-dataset
Prompt Refinement Dataset
Dataset Summary
The Prompt Refinement Dataset is a curated collection of 4,349 input-output pairs
designed to train language models to transform basic, vague prompts into high-quality,
detailed, and structured prompts that elicit significantly better responses from AI systems.
Each pair consists of a raw user-written prompt as the input and an expertly
engineered version of the same prompt as the output — preserving the original
intent while… See the full description on the dataset page: https://huggingface.co/datasets/Kamran-56/prompt-refinement-dataset.refinement-abliterated-vision_heretic_short_answers
Dataset Card: Refinement-Abliterated Short Answers (vision_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io optimized with short-form technical descriptions under 1200 tokens.
Dataset Structure
Every row contains a standard ShareGPT message structure along with an optimized text column:
prompt: The initial raw query.
answer: Clean extracted short assistant text.
messages: A clean [user, assistant] array where the… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-vision_heretic_short_answers.refinement-abliterated-vision_heretic__harmful_refusals
Dataset Card: Refinement-Abliterated Short Answers (vision_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io optimized with short-form technical descriptions under 1200 tokens.
Dataset Structure
Every row contains a standard ShareGPT message structure along with an optimized text column:
prompt: The initial raw query.
answer: Clean extracted short assistant text.
messages: A clean [user, assistant] array where the… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-vision_heretic__harmful_refusals.refinement-abliterated-thinking_heretic_short_answers
Dataset Card: Refinement-Abliterated Short Answers (thinking_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io optimized with short-form technical descriptions under 1200 tokens.
Dataset Structure
Every row contains a standard ShareGPT message structure along with an optimized text column:
prompt: The initial raw query.
answer: Clean extracted short assistant text.
messages: A clean [user, assistant] array where the… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-thinking_heretic_short_answers.refinement-abliterated-vision_heretic_short_answers1
Dataset Card: Refinement-Abliterated Short Answers (vision_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io optimized with short-form technical descriptions under 1200 tokens.
Dataset Structure
Every row contains a standard ShareGPT message structure along with an optimized text column:
prompt: The initial raw query.
answer: Clean extracted short assistant text.
messages: A clean [user, assistant] array where the… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-vision_heretic_short_answers1.
