model-outputs
sm-hackathon-actionability-9-multi-outputs-setfit-model-v0.1model_output_sft_llama_preferredsm-hackathon-actionability-9-multi-outputs-setfit-all-roberta-large-model-v0.1model_output_sft_llama_rejectedmodel_output_subreddit-wallstreetbets_newmodel_output_sftmodel_output_sft2model_output_sft_llama_preferred_mixed
reflection_model_outputs_run1
Reflection Model Outputs
This repository contains model output results from various LLMs across multiple tasks and configurations.
📂 Dataset Structure
We have 3 runs of data, and all files are organized under the main directory:
EssentialAI/reflection_model_outputs_run1/
EssentialAI/reflection_model_outputs_run2/
EssentialAI/reflection_model_outputs_run3/
Within this, you will find results grouped by model architecture and checkpoint size, including:
OLMo-2 7B
OLMo-2… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/reflection_model_outputs_run1.WildBench-V2-Model-Outputsreflection_model_outputs_run3
Reflection Model Outputs
This repository contains model output results from various LLMs across multiple tasks and configurations.
📂 Dataset Structure
We have 3 runs of data, and all files are organized under the main directory:
EssentialAI/reflection_model_outputs_run1/
EssentialAI/reflection_model_outputs_run2/
EssentialAI/reflection_model_outputs_run3/
Within this, you will find results grouped by model architecture and checkpoint size, including:
OLMo-2 7B
OLMo-2… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/reflection_model_outputs_run3.model-benchmark-outputsThis repo contains the outputs of various models on the test set of the UofA-LINGO/text_to_triplets dataset.
Model
Filename
Training Dataset
vicuna-7b-gpt-explanations (benchmark with 1024 cutoff length)
vicuna-gpt-explanations.pickle
UofA-LINGO/webnlg-gpt-explanations
vicuna-7b-combined
vicuna-combined.pickle
UofA-LINGO/webnlg-combined
vicuna-7b-combined-with-reflection
vicuna-with-reflections.pickle
UofA-LINGO/webnlg-combined-with-reflections
vicuna-7b-newins… See the full description on the dataset page: https://huggingface.co/datasets/UofA-LINGO/model-benchmark-outputs.reflection_model_outputs_run2
Reflection Model Outputs
This repository contains model output results from various LLMs across multiple tasks and configurations.
📂 Dataset Structure
We have 3 runs of data, and all files are organized under the main directory:
EssentialAI/reflection_model_outputs_run1/
EssentialAI/reflection_model_outputs_run2/
EssentialAI/reflection_model_outputs_run3/
Within this, you will find results grouped by model architecture and checkpoint size, including:
OLMo-2 7B
OLMo-2… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/reflection_model_outputs_run2.SWE-IF-model-outputs
SWE-IF — Model Outputs
Raw model responses for SWE-IF (ICML 2026), released so you can reproduce /
re-evaluate the paper's numbers without re-running generation.
Contains responses from the reported models on Big-SWE-IF and Live-SWE-IF,
under both single-turn generation and multi-turn editing, for 0–5 added instructions.
📄 Paper: https://arxiv.org/abs/2510.07315
💻 Code: https://github.com/maszhongming/SWE-IF
🤗 Benchmarks + taxonomy:… See the full description on the dataset page: https://huggingface.co/datasets/MingZhong/SWE-IF-model-outputs.
