LLM_Reasoning
details_alexredna__TinyLlama-1.1B-Chat-v1.0-reasoning-v2-dpo
Dataset Card for Evaluation run of alexredna/TinyLlama-1.1B-Chat-v1.0-reasoning-v2-dpo
Dataset automatically created during the evaluation run of model alexredna/TinyLlama-1.1B-Chat-v1.0-reasoning-v2-dpo on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_alexredna__TinyLlama-1.1B-Chat-v1.0-reasoning-v2-dpo.llm-cipher-reasoning
llm-cipher-reasoning — data, eval results and full research ledger
Everything except the weights from a research run asking: can an LLM be trained to reason in a
more compact "language" than English, and does that actually save tokens?
Two linked lines of work on Qwen/Qwen3-4B-Instruct-2507:
Cipher invention / cross-model communication — cold-decoding tests, negotiated cipher
collusion between model pairs, a cipher-hardening arms race, and GEPA prompt optimization to get
a… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/llm-cipher-reasoning.LFM2.5-KO-SFT-Stage2-Diverse-KoSWE-Reasoning-LFMChat-Raw
LFM2.5-KO-SFT-Stage2-Diverse-KoSWE-Reasoning-LFMChat-Raw
Stage2 raw LFM chat JSONL shards: Korean domain, behavior, SWE/coding, reasoning, finance, legal, Text2SQL.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-SFT-Stage2-Diverse-KoSWE-Reasoning-LFMChat-Raw.llm-fol-reasoning-eval
LLM FOL Reasoning Eval
This dataset is derived from ProverQA, a First-Order Logic reasoning benchmark designed to test the ability of large language models (LLMs) to perform structured logical reasoning.It restructures and normalizes the ProverQA development and training data into a unified, clean format suitable for evaluating chain-of-thought (CoT) and symbolic reasoning capabilities in LLMs.
Source
Original dataset: ProverQA: A First-Order Logic Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MinaGabriel/llm-fol-reasoning-eval.twi-llm-reasoning-dataset-1k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Twi Reasoning Dataset
A Twi (Akan) translation of the Multilingual-Thinking… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-llm-reasoning-dataset-1k.NLP-Course-LLM-Reasoning-Eval-May2025
Overview of LLM Reasoning Eval Dataset
This dataset contains evaluation of multiple large language models (LLMs) over 918 MCQ reasoning questions created by 184 students.
Each question was used to test 3 LLMs (each 3 times): GPT-4o, Claude Sonnet 3.x (3.5 or 3.7), and Deepseek R1.
The questions target various reasoning areas (i.e., Math, Logic, Temporal, Commonsense) and are included only if 3 seperate attempts (in a new session) by ChatGPT (GPT-4o) fail at giving the correct… See the full description on the dataset page: https://huggingface.co/datasets/nlpllmeval/NLP-Course-LLM-Reasoning-Eval-May2025.
