datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InsightVQA
InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark
Overview
InsightVQA is a large-scale dataset designed for hierarchical visual question answering that bridges emotion understanding and cognitive reasoning. While existing benchmarks predominantly focus on surface-level emotion recognition , InsightVQA introduces a structured paradigm to evaluate a model's ability to interpret emotional causes, ground evidence, and reason about… See the full description on the dataset page: https://huggingface.co/datasets/ziyul707/InsightVQA.locomoLight-INSIGHT-Bench
INSIGHT-Bench v1
A human-curated object-goal navigation benchmark: 1,097 episodes over 210 scenes, each
episode a short natural-language instruction, a start pose, a goal position and a success radius,
defined on Z-up, metre-scaled USD conversions of four scene sources -- HM3D, Matterport3D,
InteriorGS and Habitat-GS (3D Gaussian Splatting). It is evaluated in NVIDIA Isaac Sim by
the INSIGHT-Bench evaluation SDK, which publishes exactly one coordinate over these bytes:… See the full description on the dataset page: https://huggingface.co/datasets/LightOriginsHQ/Light-INSIGHT-Bench.dfs-glossary
DFS Glossary — Amharic & Afaan Oromoo
Expert-verified glossaries of Digital Financial Services (DFS) terminology in
Amharic (am) and Afaan Oromoo (om), published as structured,
machine-readable, openly-licensed data.
Open language infrastructure for two low-resource Ethiopian languages — for
developers, researchers, translators, and the financial-inclusion community.
Languages
Amharic (am, Ge'ez script) · Afaan Oromoo (om, Latin script)
Entries
87 Amharic + 86… See the full description on the dataset page: https://huggingface.co/datasets/shega-insight/dfs-glossary.insight-ladder-imo2024
Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate
Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission).
Overview
A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with:
4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.youtube-comment-insights-chatml
YouTube Comment Insights - ChatML
Overview
This dataset contains instruction-tuning samples for structured YouTube comment analysis.
The dataset is formatted in ChatML conversational format and is intended for supervised fine-tuning (SFT), QLoRA, and instruction tuning of large language models.
Each sample contains:
sentiment
tone
pros
cons
Dataset Statistics
~20k training samples
~2k validation samples
Multilingual YouTube comments
Structured JSON… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-chatml.Pashto-Social-Insight-Reasoning-Dataset
Pashto Social Insight & Reasoning Dataset (PSIR)
Overview
The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
circle-packing-insight-loop
Circle-Packing Insight-Exploration Loop
Artifacts from an iterative GPT solver <-> proposer insight-exploration loop on the
21-circles-in-a-perimeter-4-rectangle packing problem (AlphaEvolve SOTA sum-of-radii
= 2.3658321334167627). Each round, 16 solvers propose a program + written explanation;
every program is scored; a proposer then mines all 16 attempts into an evolving insight
document that conditions the next round. Run: 16 solvers x 8 rounds.
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/ars22/circle-packing-insight-loop.airbnb-reviewsexpert-insights
Expert Insights
Expert profiles for Beau, Tate, and Wendy Thompson with specializations.
Details
Records: 3
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Thompson Mortgage Group
Publisher: Thompson Mortgage Group
Thompson Alpha Logic
Deep expert entity profiles with NMLS credentials, specialization routing, branded insight labels (Wendy's Wisdom, Beau's Brief, Tate's Take), and citation formats. Designed for AI entity disambiguation… See the full description on the dataset page: https://huggingface.co/datasets/wendymthompson/expert-insights.startup-failure-insightmedical-insights-en-zh
Bilingual Medical Insights (EN-ZH) · 医学中英双语知识卡片样本
This dataset contains 50 bilingual English-Chinese medical insight cards, curated for AI model training, education, and research.
本数据集包含50条中英文医学知识卡片样本,适用于人工智能训练、医学教学与结构化语义分析任务。
✅ Fields Included | 字段结构
title / title_zh — Medical topic / 医学主题
narrative / narrative_zh — Context or background / 医学背景介绍
arguments / arguments_zh — Key points or findings / 论点要点
primary_theme — Major medical discipline (e.g. Psychiatry… See the full description on the dataset page: https://huggingface.co/datasets/jundai2003/medical-insights-en-zh.InsightTokTokenTraining
InsightTok Token Training
This is a list of image prompts and token pairs for training.
women_health_10k_insights
Women's Health 10k Insights
Women's Health 10k Insights is a dataset of medical cases selected for their relevance to women's health and processed into concise, practice-oriented tips for clinicians and medical AI agents.
The dataset is intended to help surface uncommon presentations, diagnostic pitfalls, misleading test results, and other lessons that may reduce avoidable clinical mistakes. It is an educational and decision-support resource—not a diagnostic system or a… See the full description on the dataset page: https://huggingface.co/datasets/FremyCompany/women_health_10k_insights.project_insight_datainsightfactory__Llama-3.2-3B-Instruct-unsloth-bnb-4bitlora_model-details
Dataset Card for Evaluation run of insightfactory/Llama-3.2-3B-Instruct-unsloth-bnb-4bitlora_model
Dataset automatically created during the evaluation run of model insightfactory/Llama-3.2-3B-Instruct-unsloth-bnb-4bitlora_model
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/insightfactory__Llama-3.2-3B-Instruct-unsloth-bnb-4bitlora_model-details.Medical-Insights
Medical-Insights
This dataset was uploaded automatically.
meeting-insights-corpus
Meeting Insights Sample Corpus
Sample corpus of meeting transcripts with extracted insights.
About
This dataset demonstrates AI-powered meeting analysis including:
Key decision extraction
Action item identification
Sentiment analysis
Topic clustering
Source
Created by Creative Content Crafts for Memory.Actor.
Related Resources
Memory.Actor: https://memory.actor - AI meeting intelligence
Company: Creative Content Crafts
Wikidata: Q137625544… See the full description on the dataset page: https://huggingface.co/datasets/sergeinboca/meeting-insights-corpus.stock-insightsairbnb-reviews-supervisedkazakh-news-insightsInsights_Finetune_V1_50holiday-crm-insightscloudwatch-logs-insights-queries-50kairbnb-reviews-supervised-improvementsinsights_anatili
