datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
QWQ-LongCOT-AIMOQWQ-LongCOT-AIMO is a derived dataset created by processing the amphora/QwQ-LongCoT-130K dataset. It filters the original dataset to focus specifically on question-answering pairs where the final answer is a numerical value between 0 and 999, explicitly marked using the \boxed{...} format within the original chain-of-thought answer.
Dataset Structure
Data Splits
The dataset is split into training, validation, and test sets with an 80/10/10 ratio based on the filtered… See the full description on the dataset page: https://huggingface.co/datasets/Floppanacci/QWQ-LongCOT-AIMO.econcausal-benchmark📊 EconCausal: A Context-Aware Causal Reasoning Benchmark for LLMs
Donggyu Lee, Hyeok Yun, Meeyoung Cha, Sungwon Park, Sangyoon Park, Jihee Kim
🌍 Overview
Socio-economic causal effects depend heavily on their specific institutional and environmental context. A single intervention can produce opposite results depending on regulatory or market factors.
EconCausal is a large-scale benchmark comprising 10,490 context-annotated causal triplets extracted from 2,595… See the full description on the dataset page: https://huggingface.co/datasets/qwqw3535/econcausal-benchmark.Wordle-MCM-Dataset
Wordle Player Performance Dataset
1. 简介
本数据集源自 2023 MCM Problem C,包含 2022 年全年的 Wordle 每日单词、报告人数及玩家猜测分布。
2. 数据内容
Date: 日期
Word: 每日目标单词
Total_Reported: 报告总人数
Hard_Mode: 硬核模式人数
Try_1 to Try_6: 玩家在第 1 至 6 次尝试中猜中的比例
Try_7_plus: 未能在 6 次内猜中的比例
3. 用途
本项目使用该数据集通过 Bi-LSTM 模型预测玩家的平均尝试步数 (Mean Tries)。
QwQ-LongCoT-130K_filtered_10k
