datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stories-refinement
Stories Refinement
This dataset contains synthetic short stories generated from blog text excerpts sourced from the agentlans/lucadiliello-STORIES dataset.
The stories were produced using the agentlans/Llama3.1-LexiHermes-SuperStorm language model unless otherwise noted.
Configurations
allContaining all other configs and filtered for output < 6000 characters. This config has an additional column indicating which config each row is from.
zero-shotGenerated directly… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/stories-refinement.code_x_glue_cc_code_refinement_messagesgoogle/code_x_glue_cc_code_refinementのsplit trainをopenAI messages形式に調整。
high-quality-text-refinementfinewebedu-refinement
finewebedu-refinement
This dataset contains simplified versions of excerpts from HuggingFaceFW/fineweb-edu.
Methods
Texts were split into chunks about 2000 Llama 3 tokens long.
The chunks were refined using agentlans/Llama3.1-LexiHermes-SuperStorm and a fine-tuned cognitivecomputations/Dolphin3.0-Llama3.2-3B model. The refinements aimed to:
Use simple language
Remove unnecessary words
Use active voice
Break long sentences
Results
Total passages: 9996… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-refinement.Attention-Refinement-Data
