CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01blackbird0831 /slm-assignment-data Grade-Level Vocabulary-Locked Writing Tutor — Dataset Training and evaluation data for fine-tuning a small open model (Qwen3-0.6B) into a grade 7–8 writing/grammar tutor whose vocabulary and sentence complexity stay locked to the band — it introduces at most one word above grade level per reply (always immediately defined) and never escalates, even under pressure ("use bigger words", "give me the college version") or jailbreak-style attacks. The dataset is the deliverable. ~80%… See the full description on the dataset page: https://huggingface.co/datasets/blackbird0831/slm-assignment-data.texttext-generation1K<n<10K0 likes50 downloads2mo agoHugging Face02Hengming0805 /self-alignment-curated-assignment3 Self Alignment Curated Assignment 3 This dataset contains a small curated synthetic instruction-response dataset created for an assignment implementation of the paper Self-Alignment with Instruction Backtranslation. The dataset consists of high-quality instruction-response pairs generated through a 4-step pipeline: Train a backward model on OpenAssistant-Guanaco. Sample 150 single-turn responses from LIMA. Generate instructions from those responses using the backward model. Score… See the full description on the dataset page: https://huggingface.co/datasets/Hengming0805/self-alignment-curated-assignment3.texttext-generationn<1K0 likes17 downloads6mo agoHugging Face03Shirleyabeauty /assignment3-curated-datasettexttext-generationn<1K0 likes12 downloads5mo agoHugging Face04sunming-giegie /assignment3-lima-curated-150 Assignment 3 Curated LIMA Dataset This dataset contains the curated instruction-response pairs produced in Part 3 of Assignment 3. Source Pipeline Start from LIMA single-turn examples. Use the backward model to infer instructions from responses. Score each (generated_instruction, response) pair with Qwen/Qwen3-1.7B. Keep examples with score >= 4. Files train.jsonl: curated high-quality examples for final instruction tuning scores.jsonl: all 150 scored… See the full description on the dataset page: https://huggingface.co/datasets/sunming-giegie/assignment3-lima-curated-150.tabulartext-generationn<1K0 likes8 downloads5mo agoHugging Face05sunming-giegie /assignment3-curated-lima-dataset Assignment 3 Curated LIMA Dataset This dataset contains the curated instruction-response pairs produced in Part 3 of Assignment 3. Source Pipeline Start from LIMA single-turn examples. Use the backward model to infer instructions from responses. Score each (generated_instruction, response) pair with Qwen/Qwen3-1.7B using few-shot prompting and a 1-5 quality rubric. Keep examples with score >= 4. Files train.jsonl: curated high-quality examples for final… See the full description on the dataset page: https://huggingface.co/datasets/sunming-giegie/assignment3-curated-lima-dataset.tabulartext-generationn<1K0 likes6 downloads5mo agoHugging Face06Shirleyabeauty /assignment4-pairrm-preferences-submittabulartext-generationn<1K0 likes4 downloads5mo agoHugging Face07sunming-giegie /assignment3-lima-curated Assignment 3 Curated LIMA Dataset This dataset contains the curated instruction-response pairs produced in Part 3 of Assignment 3. Source Pipeline Start from LIMA single-turn examples. Use the backward model to infer instructions from responses. Score each (generated_instruction, response) pair with Qwen/Qwen3-1.7B using few-shot prompting and a 1-5 quality rubric. Keep examples with score >= 4. Files train.jsonl: curated high-quality examples for final… See the full description on the dataset page: https://huggingface.co/datasets/sunming-giegie/assignment3-lima-curated.tabulartext-generationn<1K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.