datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_x_glue_cc_code_refinement
Dataset Card for "code_x_glue_cc_code_refinement"
Dataset Summary
CodeXGLUE code-refinement dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/code-refinement
We use the dataset released by this paper(https://arxiv.org/pdf/1812.08693.pdf). The source side is a Java function with bugs and the target side is the refined one. All the function and variable names are normalized. Their dataset contains two subsets ( i.e.small and medium) based on… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_refinement.2026_08_05_refinement_5env_gemma3_12b_gemma4_31b_tok2026_08_09_refinement_5env_gemma3_12b_gemma4_31b_flsft_tok2026_08_20_refinement_math_chess_gemma3_12b_gemma4_31b_transition_feedback_tokstories-refinement
Stories Refinement
This dataset contains synthetic short stories generated from blog text excerpts sourced from the agentlans/lucadiliello-STORIES dataset.
The stories were produced using the agentlans/Llama3.1-LexiHermes-SuperStorm language model unless otherwise noted.
Configurations
allContaining all other configs and filtered for output < 6000 characters. This config has an additional column indicating which config each row is from.
zero-shotGenerated directly… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/stories-refinement.code_x_glue_cc_code_refinement_annotated2026_08_11_refinement_5env_gemma3_12b_gemma4_31b_raw_student_tok2026_08_17_refinement_math_chess_gemma3_12b_gemma4_31b_raw_student_no_feedback_tokutf_20_refsen-math-cot-refinementDataset_Text_Refinement
Dataset Card for Dataset Name
This Dataset is for refining text based on user study text preferences. Given a original text to refined text based on given paramters.
The parameters are:
-Readability_Score
-Semantic_Coherence
-User_Preference
Readability Score:
Its the average of Flesch-Kincaid Readability Ease Score and Dale-Chall readability score.
The Readability Score is of user which is appilied on refined text.
Semantic Coherence:
Its the average float value of… See the full description on the dataset page: https://huggingface.co/datasets/SolaceinLoneSun/Dataset_Text_Refinement.latin-refinement-dataset2026_08_12_refinement_math_chess_gemma3_12b_gemma4_31b_raw_student_tokquery_refinement_datasetiterative-refinementutf_20_refs_filerefinement-abliterated-thinking_heretic
Dataset Card: Refinement-Abliterated (thinking_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io designed to train language models to process edge-case, controversial, or complex analytical prompts without triggering over-aligned corporate refusal responses.
Generation Pipeline Mechanics
Seed Matrix: Initial queries gathered from mlabonne/harmful_behaviors.
Knowledge Engine (Abliterated Base): Generated using… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-thinking_heretic.dpo-sample-perfect-lesscode_x_glue_cc_code_refinement_messagesgoogle/code_x_glue_cc_code_refinementのsplit trainをopenAI messages形式に調整。
utf_20_refs_pairsdpo-sample-perfect-morehigh-quality-text-refinementprompt-refinement-dataset
Prompt Refinement Dataset
Dataset Summary
The Prompt Refinement Dataset is a curated collection of 4,349 input-output pairs
designed to train language models to transform basic, vague prompts into high-quality,
detailed, and structured prompts that elicit significantly better responses from AI systems.
Each pair consists of a raw user-written prompt as the input and an expertly
engineered version of the same prompt as the output — preserving the original
intent while… See the full description on the dataset page: https://huggingface.co/datasets/Kamran-56/prompt-refinement-dataset.refinement-abliterated-vision_heretic_short_answers
Dataset Card: Refinement-Abliterated Short Answers (vision_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io optimized with short-form technical descriptions under 1200 tokens.
Dataset Structure
Every row contains a standard ShareGPT message structure along with an optimized text column:
prompt: The initial raw query.
answer: Clean extracted short assistant text.
messages: A clean [user, assistant] array where the… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-vision_heretic_short_answers.2026_07_29_omni_math_refinement_overfit_64finewebedu-refinement
finewebedu-refinement
This dataset contains simplified versions of excerpts from HuggingFaceFW/fineweb-edu.
Methods
Texts were split into chunks about 2000 Llama 3 tokens long.
The chunks were refined using agentlans/Llama3.1-LexiHermes-SuperStorm and a fine-tuned cognitivecomputations/Dolphin3.0-Llama3.2-3B model. The refinements aimed to:
Use simple language
Remove unnecessary words
Use active voice
Break long sentences
Results
Total passages: 9996… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-refinement.refinement-abliterated-vision_heretic__harmful_refusals
Dataset Card: Refinement-Abliterated Short Answers (vision_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io optimized with short-form technical descriptions under 1200 tokens.
Dataset Structure
Every row contains a standard ShareGPT message structure along with an optimized text column:
prompt: The initial raw query.
answer: Clean extracted short assistant text.
messages: A clean [user, assistant] array where the… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-vision_heretic__harmful_refusals.utf_20_refs_file_sample100refinement-abliterated-thinking_heretic_short_answers
Dataset Card: Refinement-Abliterated Short Answers (thinking_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io optimized with short-form technical descriptions under 1200 tokens.
Dataset Structure
Every row contains a standard ShareGPT message structure along with an optimized text column:
prompt: The initial raw query.
answer: Clean extracted short assistant text.
messages: A clean [user, assistant] array where the… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-thinking_heretic_short_answers.refinement-abliterated-vision_heretic_short_answers1
Dataset Card: Refinement-Abliterated Short Answers (vision_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io optimized with short-form technical descriptions under 1200 tokens.
Dataset Structure
Every row contains a standard ShareGPT message structure along with an optimized text column:
prompt: The initial raw query.
answer: Clean extracted short assistant text.
messages: A clean [user, assistant] array where the… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-vision_heretic_short_answers1.
