datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_x_glue_cc_code_refinement
Dataset Card for "code_x_glue_cc_code_refinement"
Dataset Summary
CodeXGLUE code-refinement dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/code-refinement
We use the dataset released by this paper(https://arxiv.org/pdf/1812.08693.pdf). The source side is a Java function with bugs and the target side is the refined one. All the function and variable names are normalized. Their dataset contains two subsets ( i.e.small and medium) based on… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_refinement.2026_08_05_refinement_5env_gemma3_12b_gemma4_31b_tok2026_08_09_refinement_5env_gemma3_12b_gemma4_31b_flsft_tok2026_08_20_refinement_math_chess_gemma3_12b_gemma4_31b_transition_feedback_tok2026_08_11_refinement_5env_gemma3_12b_gemma4_31b_raw_student_tokcode_x_glue_cc_code_refinement_annotatedutf_20_refs2026_08_17_refinement_math_chess_gemma3_12b_gemma4_31b_raw_student_no_feedback_token-math-cot-refinementiterative-refinement2026_08_12_refinement_math_chess_gemma3_12b_gemma4_31b_raw_student_tokDataset_Text_Refinement
Dataset Card for Dataset Name
This Dataset is for refining text based on user study text preferences. Given a original text to refined text based on given paramters.
The parameters are:
-Readability_Score
-Semantic_Coherence
-User_Preference
Readability Score:
Its the average of Flesch-Kincaid Readability Ease Score and Dale-Chall readability score.
The Readability Score is of user which is appilied on refined text.
Semantic Coherence:
Its the average float value of… See the full description on the dataset page: https://huggingface.co/datasets/SolaceinLoneSun/Dataset_Text_Refinement.query_refinement_datasetutf_20_refs_filerefinement-abliterated-thinking_heretic
Dataset Card: Refinement-Abliterated (thinking_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io designed to train language models to process edge-case, controversial, or complex analytical prompts without triggering over-aligned corporate refusal responses.
Generation Pipeline Mechanics
Seed Matrix: Initial queries gathered from mlabonne/harmful_behaviors.
Knowledge Engine (Abliterated Base): Generated using… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-thinking_heretic.latin-refinement-datasetdpo-sample-perfect-lessutf_20_refs_pairsdpo-sample-perfect-morerefinement-abliterated-vision_heretic_short_answers
Dataset Card: Refinement-Abliterated Short Answers (vision_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io optimized with short-form technical descriptions under 1200 tokens.
Dataset Structure
Every row contains a standard ShareGPT message structure along with an optimized text column:
prompt: The initial raw query.
answer: Clean extracted short assistant text.
messages: A clean [user, assistant] array where the… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-vision_heretic_short_answers.2026_07_29_omni_math_refinement_overfit_64refinement-abliterated-vision_heretic__harmful_refusals
Dataset Card: Refinement-Abliterated Short Answers (vision_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io optimized with short-form technical descriptions under 1200 tokens.
Dataset Structure
Every row contains a standard ShareGPT message structure along with an optimized text column:
prompt: The initial raw query.
answer: Clean extracted short assistant text.
messages: A clean [user, assistant] array where the… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-vision_heretic__harmful_refusals.utf_20_refs_file_sample100refinement-abliterated-thinking_heretic_short_answers
Dataset Card: Refinement-Abliterated Short Answers (thinking_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io optimized with short-form technical descriptions under 1200 tokens.
Dataset Structure
Every row contains a standard ShareGPT message structure along with an optimized text column:
prompt: The initial raw query.
answer: Clean extracted short assistant text.
messages: A clean [user, assistant] array where the… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-thinking_heretic_short_answers.refinement-abliterated-vision_heretic_short_answers1
Dataset Card: Refinement-Abliterated Short Answers (vision_heretic)
Overview
This dataset contains a distilled corpus created by hirundo-io optimized with short-form technical descriptions under 1200 tokens.
Dataset Structure
Every row contains a standard ShareGPT message structure along with an optimized text column:
prompt: The initial raw query.
answer: Clean extracted short assistant text.
messages: A clean [user, assistant] array where the… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/refinement-abliterated-vision_heretic_short_answers1.arc-agi-partials-for-refinementcode_refinement_validarc-agi-1-refinement-finetuningarc-agi-1-refinement-finetuning-partialplus
