datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EpistemeAI2__Fireball-Llama-3.1-8B-Philos-Reflection-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Llama-3.1-8B-Philos-Reflection
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Llama-3.1-8B-Philos-Reflection
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Llama-3.1-8B-Philos-Reflection-details.EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection-details.api_graph_reflectiongrok-reflection-cot-ru
march228/grok-reflection-cot-ru
Russian synthetic dataset with question, internal thought text, and final answer.
What is inside
Rows: 4190
Split: train
Main fields:
question
thought_text
answer
thought1..thought5
model
task_type
reflection_count
Format
The dataset is stored as train.jsonl.
thought_text is the joined internal monologue with blank lines between thought blocks.thought1..thought5 preserve the original segmented form from the SQLite source.… See the full description on the dataset page: https://huggingface.co/datasets/march228/grok-reflection-cot-ru.mattshumer__Reflection-Llama-3.1-70B-details
Dataset Card for Evaluation run of mattshumer/Reflection-Llama-3.1-70B
Dataset automatically created during the evaluation run of model mattshumer/Reflection-Llama-3.1-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mattshumer__Reflection-Llama-3.1-70B-details.glaiveai__Reflection-Llama-3.1-70B-details
Dataset Card for Evaluation run of glaiveai/Reflection-Llama-3.1-70B
Dataset automatically created during the evaluation run of model glaiveai/Reflection-Llama-3.1-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/glaiveai__Reflection-Llama-3.1-70B-details.olabs-ai__reflection_model-details
Dataset Card for Evaluation run of olabs-ai/reflection_model
Dataset automatically created during the evaluation run of model olabs-ai/reflection_model
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/olabs-ai__reflection_model-details.mosama__Qwen2.5-1.5B-Instruct-CoT-Reflection-details
Dataset Card for Evaluation run of mosama/Qwen2.5-1.5B-Instruct-CoT-Reflection
Dataset automatically created during the evaluation run of model mosama/Qwen2.5-1.5B-Instruct-CoT-Reflection
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mosama__Qwen2.5-1.5B-Instruct-CoT-Reflection-details.SenseLLM__ReflectionCoder-DS-33B-details
Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-DS-33B
Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-DS-33B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SenseLLM__ReflectionCoder-DS-33B-details.Xiaojian9992024__Reflection-L3.2-JametMiniMix-3B-details
Dataset Card for Evaluation run of Xiaojian9992024/Reflection-L3.2-JametMiniMix-3B
Dataset automatically created during the evaluation run of model Xiaojian9992024/Reflection-L3.2-JametMiniMix-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Xiaojian9992024__Reflection-L3.2-JametMiniMix-3B-details.SenseLLM__ReflectionCoder-CL-34B-details
Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-CL-34B
Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-CL-34B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SenseLLM__ReflectionCoder-CL-34B-details.feedback_qesconv_badareas_questions_reflectionsValidation Results:
PAIR-Reflections:
• "Precision": 0.8955223880597015
• "Recall": 0.2553191489361703
• "Accuracy": 0.77
• "F1-score": 0.3973509933774836
Rule-based HasQuestions:
• "Precision": 1.0
• "Recall": 0.9629629629629629
• "Accuracy": 0.74
• "F1-score": 0.9811320754716981
