datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VDocRetriever-Pretrain-DocStructveri-bilimci-insight-diyalog-tr-16.2k
🇹🇷 Veri Bilimci Insight Diyalog Veri Seti (TR, 16.2K) — %100 Türkçe Metin
Gerçek dünya blogları, uzman soru-cevap içerikleri ve akademik makale metinlerinden üretilmiş; veri madenciliği ve uygulamalı veri bilimi karar diline odaklanan, %100 Türkçe çok turlu diyalog veri seti.
🧠 Bu Veri Seti Ne Amaçla Üretildi?
Amaç, modeli teorik tanım ezberinden çıkarıp bağlama göre karar veren veri bilimci davranışına yaklaştırmaktır.
Her örnekte yöntem seçimi, alternatif kıyası… See the full description on the dataset page: https://huggingface.co/datasets/zero9tech/veri-bilimci-insight-diyalog-tr-16.2k.insight-ladder-imo2024
Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate
Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission).
Overview
A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with:
4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.youtube-comment-insights-chatml
YouTube Comment Insights - ChatML
Overview
This dataset contains instruction-tuning samples for structured YouTube comment analysis.
The dataset is formatted in ChatML conversational format and is intended for supervised fine-tuning (SFT), QLoRA, and instruction tuning of large language models.
Each sample contains:
sentiment
tone
pros
cons
Dataset Statistics
~20k training samples
~2k validation samples
Multilingual YouTube comments
Structured JSON… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-chatml.data-scientist-insight-dialog-en-16.5k
🇬🇧 Data Scientist Insight Dialogue Dataset (EN, 16.5K) — 100% English Text
A decision-focused, multi-turn dataset built from real-world blogs, expert Q&A content, and academic papers, tuned for data mining and applied data science reasoning in full English text purity.
🧠 Why This Dataset Exists
The goal is to train decision behavior, not only definition recall.
Each sample is structured to expose method choice, trade-off reasoning, risk signals, and validation… See the full description on the dataset page: https://huggingface.co/datasets/zero9tech/data-scientist-insight-dialog-en-16.5k.Pashto-Social-Insight-Reasoning-Dataset
Pashto Social Insight & Reasoning Dataset (PSIR)
Overview
The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.circle-packing-insight-loop
Circle-Packing Insight-Exploration Loop
Artifacts from an iterative GPT solver <-> proposer insight-exploration loop on the
21-circles-in-a-perimeter-4-rectangle packing problem (AlphaEvolve SOTA sum-of-radii
= 2.3658321334167627). Each round, 16 solvers propose a program + written explanation;
every program is scored; a proposer then mines all 16 attempts into an evolving insight
document that conditions the next round. Run: 16 solvers x 8 rounds.
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/ars22/circle-packing-insight-loop.youtube-comment-insights-clean
YouTube Comment Insights - Clean Dataset
Overview
This dataset contains structured YouTube comment analytics data designed for visualization, analytics, and machine learning workflows.
Each sample contains:
comment
sentiment
tone
pros
cons
The dataset is intended for easy readability and downstream analytics tasks.
Dataset Statistics
~20k training samples
~2k validation samples
Multilingual YouTube comments
Structured JSON format
Files… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-clean.expert-insights
Expert Insights
Expert profiles for Beau, Tate, and Wendy Thompson with specializations.
Details
Records: 3
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Thompson Mortgage Group
Publisher: Thompson Mortgage Group
Thompson Alpha Logic
Deep expert entity profiles with NMLS credentials, specialization routing, branded insight labels (Wendy's Wisdom, Beau's Brief, Tate's Take), and citation formats. Designed for AI entity disambiguation… See the full description on the dataset page: https://huggingface.co/datasets/wendymthompson/expert-insights.arxiv-paper-insights
arXiv Paper Insights
Overview
arXiv Paper Insights is a public dataset for arXiv paper discovery, recommendations, and structured research insights.
The dataset is derived from LinxSci (https://linxsci.com), a paper reading and insight platform focused on arXiv papers.
For any paper in this dataset, you can open:
https://linxsci.com/pdf/{arxiv_id}
to read the paper and view the corresponding insights on LinxSci.
Related Links
LinxSci: https://linxsci.com… See the full description on the dataset page: https://huggingface.co/datasets/LinxSci/arxiv-paper-insights.Problem-Solving-Insights-Based-on-Kazakh-Traditions
🇰🇿 Problem-Solving Insights Based on Kazakh Traditions
📖 Overview
Problem-Solving Insights Based on Kazakh Traditions is a instruction-tuning dataset designed to bridge the gap between ancient Kazakh wisdom and modern societal challenges.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
8,005
Total Words (approx.)
4,030,925
Avg. Words per Sample
503
Word Count Distribution (Per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Problem-Solving-Insights-Based-on-Kazakh-Traditions.product_reviews_insight_10k
Dataset Summary
This dataset was built from Amazon product reviews and curated into an instruction-tuning format for structured pros and cons extraction.
The pipeline includes:
Raw data loading → Extract asin, reviewText.
Preprocessing → Clean, filter, and truncate each (10–150 words).
Grouping → Aggregate reviews by product.
Selection → Shuffle and select 10
Filtering → Keep 5–15 reviews per product.
Selection → Shuffle and keep 10k rows to make final dataset.
Summarization →… See the full description on the dataset page: https://huggingface.co/datasets/sdelowar2/product_reviews_insight_10k.
