datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
veri-bilimci-insight-diyalog-tr-16.2k
🇹🇷 Veri Bilimci Insight Diyalog Veri Seti (TR, 16.2K) — %100 Türkçe Metin
Gerçek dünya blogları, uzman soru-cevap içerikleri ve akademik makale metinlerinden üretilmiş; veri madenciliği ve uygulamalı veri bilimi karar diline odaklanan, %100 Türkçe çok turlu diyalog veri seti.
🧠 Bu Veri Seti Ne Amaçla Üretildi?
Amaç, modeli teorik tanım ezberinden çıkarıp bağlama göre karar veren veri bilimci davranışına yaklaştırmaktır.
Her örnekte yöntem seçimi, alternatif kıyası… See the full description on the dataset page: https://huggingface.co/datasets/zero9tech/veri-bilimci-insight-diyalog-tr-16.2k.insight-ladder-imo2024
Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate
Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission).
Overview
A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with:
4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.data-scientist-insight-dialog-en-16.5k
🇬🇧 Data Scientist Insight Dialogue Dataset (EN, 16.5K) — 100% English Text
A decision-focused, multi-turn dataset built from real-world blogs, expert Q&A content, and academic papers, tuned for data mining and applied data science reasoning in full English text purity.
🧠 Why This Dataset Exists
The goal is to train decision behavior, not only definition recall.
Each sample is structured to expose method choice, trade-off reasoning, risk signals, and validation… See the full description on the dataset page: https://huggingface.co/datasets/zero9tech/data-scientist-insight-dialog-en-16.5k.circle-packing-insight-loop
Circle-Packing Insight-Exploration Loop
Artifacts from an iterative GPT solver <-> proposer insight-exploration loop on the
21-circles-in-a-perimeter-4-rectangle packing problem (AlphaEvolve SOTA sum-of-radii
= 2.3658321334167627). Each round, 16 solvers propose a program + written explanation;
every program is scored; a proposer then mines all 16 attempts into an evolving insight
document that conditions the next round. Run: 16 solvers x 8 rounds.
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/ars22/circle-packing-insight-loop.arxiv-paper-insights
arXiv Paper Insights
Overview
arXiv Paper Insights is a public dataset for arXiv paper discovery, recommendations, and structured research insights.
The dataset is derived from LinxSci (https://linxsci.com), a paper reading and insight platform focused on arXiv papers.
For any paper in this dataset, you can open:
https://linxsci.com/pdf/{arxiv_id}
to read the paper and view the corresponding insights on LinxSci.
Related Links
LinxSci: https://linxsci.com… See the full description on the dataset page: https://huggingface.co/datasets/LinxSci/arxiv-paper-insights.
