CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AI-MO /NuminaMath-TIR Dataset Card for NuminaMath CoT Dataset Summary Tool-integrated reasoning (TIR) plays a crucial role in this competition. However, collecting and annotating such data is both costly and time-consuming. To address this, we selected approximately 70k problems from the NuminaMath-CoT dataset, focusing on those with numerical outputs, most of which are integers. We then utilized a pipeline leveraging GPT-4 to generate TORA-like reasoning paths, executing the code and… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/NuminaMath-TIR.texttext-generation10K<n<100K158 likes8k downloads2y agoHugging Face02ChuGyouk /AI-MO-NuminaMath-TIR-korean-240918 IMPORTANT NOTE This data is part of the progress. Current translation progress: 24.85% (2024-09-18 01:32 KST) I'm taking a short break due to personal reasons. I'll be back in a month. TODO-LIST Finish translation Translation I used gemini-1.5-pro-exp-0827. The prompt used for translation will be disclosed at the end. Dataset Card for NuminaMath CoT Dataset Summary Tool-integrated reasoning (TIR) plays a crucial role in this… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/AI-MO-NuminaMath-TIR-korean-240918.texttext-generation10K<n<100K5 likes72 downloads2y agoHugging Face03ptrdvn /kakugo-tir Kakugo Tigrinya dataset [Paper] [Code] [Model] A synthetically generated conversation dataset for training in Tigrinya. This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Tigrinya. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/kakugo-tir.texttext-generation10K<n<100K0 likes42 downloads8mo agoHugging Face04AcroYAMALEX /acro-yamalex-llmjp-4-math-tir acro-yamalex-llmjp-4-math-tir(データセット) 日本語数学推論のためのTool-Integrated Reasoning (TIR) データセットです。 自然言語による推論とPythonコード実行を組み合わせたマルチターン形式のデータセットで、OpenWebMathから抽出・生成した問題に対してDeepSeek V3を用いてTIR形式の解法を生成しました。 本データセットはFT-LLM2026コンペティションにおける我々のアプローチの一部として構築されました。 データセット概要 項目 値 データ件数 134,834件 言語 日本語 ソース OpenWebMath 生成モデル DeepSeek V3 (deepseek-chat) フォーマット JSONL(マルチターン対話形式) データ作成手法 OpenMathReasoning(NVIDIAのAIMO-2優勝手法)の枠組みに基づき、以下の手順で作成しました。… See the full description on the dataset page: https://huggingface.co/datasets/AcroYAMALEX/acro-yamalex-llmjp-4-math-tir.texttext-generation100K<n<1M0 likes34 downloads6mo agoHugging Face05andynik /numina-tir-2kx4 Comparison of Problem-solving Performance Across Mathematical Domains with LLMs This repository contains a filtered subset of the Numina-Math-TIR dataset, reorganised into four mathematical domains: algebra, geometry, number theory, and combinatorics. For each problem solutions were generated with LLMs: GPT-4o-mini, Mathstral-7B, Qwen2.5-Math-7B, and Llama-3.1-8B-Instruct. Each problem’s solution by these LLMs has been post-processed and compared against the human‐verified… See the full description on the dataset page: https://huggingface.co/datasets/andynik/numina-tir-2kx4.texttext-generation1K<n<10K0 likes33 downloads1y agoHugging Face06jomiguel07 /NuminaMath-TIR Dataset Card for NuminaMath CoT Dataset Summary Tool-integrated reasoning (TIR) plays a crucial role in this competition. However, collecting and annotating such data is both costly and time-consuming. To address this, we selected approximately 70k problems from the NuminaMath-CoT dataset, focusing on those with numerical outputs, most of which are integers. We then utilized a pipeline leveraging GPT-4 to generate TORA-like reasoning paths, executing the code and… See the full description on the dataset page: https://huggingface.co/datasets/jomiguel07/NuminaMath-TIR.texttext-generation10K<n<100K0 likes24 downloads7mo agoHugging Face07oieieio /NuminaMath-TIR Dataset Card for NuminaMath CoT Dataset Summary Tool-integrated reasoning (TIR) plays a crucial role in this competition. However, collecting and annotating such data is both costly and time-consuming. To address this, we selected approximately 70k problems from the NuminaMath-CoT dataset, focusing on those with numerical outputs, most of which are integers. We then utilized a pipeline leveraging GPT-4 to generate TORA-like reasoning paths, executing the code and… See the full description on the dataset page: https://huggingface.co/datasets/oieieio/NuminaMath-TIR.texttext-generation10K<n<100K0 likes18 downloads2y agoHugging Face08szhyxt /tirska Dataset Card for Alpaca Dataset Summary Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better. The authors built on the data generation pipeline from Self-Instruct framework and made the following modifications: The text-davinci-003 engine to generate the instruction data instead… See the full description on the dataset page: https://huggingface.co/datasets/szhyxt/tirska.texttext-generation10K<n<100K0 likes9 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.