CoolFace
3 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Dongwei /Feedback_Friction_Dataset Feedback Friction Dataset This dataset contains the LLaMA-4 Maverick results from the iterative feedback experiments described in the paper: FEEDBACK FRICTION: LLMs Struggle to Fully Incorporate External Feedback. Github Repository: https://github.com/JHU-CLSP/Feedback-Friction Note: While the paper evaluated multiple frontier models including LLaMA-3.3-70B-Instruct, LLaMA-4-Scout-17B-16E-Instruct, Claude 3.7 Sonnet, and Claude 3.7 Sonnet with Extended Thinking, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dongwei/Feedback_Friction_Dataset.tabularquestion-answeringn<1K2 likes57 downloads1y agoHugging Face02violetxi /single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5.4_gepa-n32 Single-turn eval — violetxi/meta_feedback_qwen3-4b_step2_gpt-5.4_gepa Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N. Eval results (n_samples_per_example = 32) Overall metric value n_examples 1006 mean@32 0.1796 best@32 0.3588 worst@32 0.0477 pass_rate 0.3588… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5.4_gepa-n32.tabularquestion-answering1K<n<10K0 likes28 downloads5mo agoHugging Face03violetxi /single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5-nano_gepa-n32 Single-turn eval — violetxi/meta_feedback_qwen3-4b_step2_gpt-5-nano_gepa Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N. Eval results (n_samples_per_example = 32) Overall metric value n_examples 1006 mean@32 0.1804 best@32 0.3569 worst@32 0.0398 pass_rate… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5-nano_gepa-n32.tabularquestion-answering1K<n<10K0 likes13 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.