datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ScreenSpot3d-front-code
3D-Front-Code
RoomScript v4 Blender object programs, room-layout renders, code-only wall
architecture, and asset Blender artifacts derived from 3D-FRONT scene evidence.
Contents
13,917 object assets (reference and v4/best) in data/assets/*.tar
21,202 rooms (v4/best and code-only v4_wall/best) in data/rooms/*.tar
searchable JSONL indexes under metadata/
Each tar contains multiple samples while preserving the original
data/front_object_code/by_asset/... or… See the full description on the dataset page: https://huggingface.co/datasets/KevinFan111/3d-front-code.TabularMath
📊 TabularMath
TabularMath is a tabular mathematical reasoning benchmark introduced in TabularMath: Understanding Math Reasoning over Tables with Large Language Models. It is built via AUTOT2T, a neuro-symbolic pipeline that automatically transforms math word problems into verified tabular reasoning tasks, enabling scalable evaluation without manual table annotation.
TabularMath jointly assesses reasoning accuracy, information retrieval over complex table structures, and… See the full description on the dataset page: https://huggingface.co/datasets/kevin715/TabularMath.vimgolf-public-challenges-inspect-evalSWE-bench-Darttime-embed-korean-temporal-inventory-v2
Time-Embed Korean Temporal Inventory v2
한국어 시간 표현 임베딩 학습·평가 데이터셋입니다.
C1, C3, C5: 문맥 허용 범위가 다른 학습 조건이며 각 314개 query를 포함합니다.
각 학습 query는 승인된 모든 동의 positive(최소 3개)와 정확히 7개 negative를 가집니다.
legacy_validation, legacy_test: 기존 동결 dev/test bytes를 그대로 보존합니다.
inventory_eval: 승인된 희귀·격식·경계 시간 표현 118개입니다.
Inventory Test 원문과 라벨은 공개하지 않으며 sealed/inventory_test_handle.json만 제공합니다.
승인 방식은 개별 행 검수로 위장하지 않은 owner_policy_waiver입니다. 정확한 승인·감사
해시와 산출물은 evidence/, audit/, dataset_manifest.json에… See the full description on the dataset page: https://huggingface.co/datasets/kev-KOH/time-embed-korean-temporal-inventory-v2.candle-fire-datapackrat-benchmarks
PackRat v2 Benchmarks
Version: 2.0.0
Date: 2026-04-10
Tokenizer: tiktoken cl100k_base (GPT-4 / Claude compatible)
Platform: Node.js v25.6.1, Windows 11
Summary
Metric
Result
Round-trip accuracy
100% (144/144 tests)
Token savings (avg)
2.4%
Token savings (best)
17.3% (path/URL-heavy files)
Byte savings (avg)
2.5%
Search speedup
12.03x
Codebook entries
72 (auto-learned)
Negative-savings entries
0
Comparison: PackRat vs MemPalace… See the full description on the dataset page: https://huggingface.co/datasets/kevo666/packrat-benchmarks.PubMed-IV
PubMed-IV Dataset
The PubMed-IV dataset is derived from PubMed abstracts and metadata, collected using the NCBI E-utilities API.
It includes structured fields such as title, abstract text (including structured sections like Conclusions when available),
authors, journal metadata, and identifiers (PMID, DOI, etc.). No full-text articles are included.
Data from PubMed, a service of the U.S. National Library of Medicine (NLM).
PubMed data is in the public domain. NLM does not endorse… See the full description on the dataset page: https://huggingface.co/datasets/KevinZonda/PubMed-IV.Amazon_Customer_Review_2023
Amazon Product Review Dataset (2023)
Dataset Overview
The Amazon Product Review Dataset (2023) contains product reviews from Amazon customers. The dataset includes product information, review details, and metadata about the customers who left the reviews. This dataset can be used for various natural language processing (NLP) tasks, including sentiment analysis, review prediction, recommendation systems, and more.
Dataset Name: Amazon Product Review Dataset (2023)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/Amazon_Customer_Review_2023.TrialLlama-datasetsdistill_r1_110k_sft_zh-tw來自conliu的distill_r1_110k_sft_zh
經由zhconv模組轉換而來
kevin-v1-dataset
Kevin V1 — NPC Conversation Dataset
Synthetic player↔NPC conversations for training game NPC dialogue models.
Generated with a 3-role pipeline (context / player / NPC) plus a judge that
verifies every NPC reply is grounded (no hallucinated facts) and
in-character.
Format
One conversation per row (JSON Lines). Each row:
{
"id": "conv_00042",
"area_id": "01_emberpeak_forge",
"npc": {"role": "blacksmith", "name": "...", "offers": [...], "knows_about": [...]}… See the full description on the dataset page: https://huggingface.co/datasets/ItsHotdogFred/kevin-v1-dataset.Consumer_goods_reviews
Amazon Product Review Dataset (2023)
Dataset Overview
The Amazon Product Review Dataset (2023) contains product reviews from Amazon customers. The dataset includes product information, review details, and metadata about the customers who left the reviews. This dataset can be used for various natural language processing (NLP) tasks, including sentiment analysis, review prediction, recommendation systems, and more.
Dataset Name: Amazon Product Review Dataset (2023)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/Consumer_goods_reviews.cognitive-pattern-selector-v1
Cognitive Pattern Selector Dataset
Dataset for fine-tuning a metacognitive pattern selector model. Given a legal/business scenario and situational assessment (SAGE), the model learns to select which of 29 metacognitive patterns (MC1-MC29) should be activated for expert analysis.
Dataset Description
This dataset was generated from the CognitiveTrainer platform, which captures expert reasoning patterns for technology transactions and product counseling.
Use Case… See the full description on the dataset page: https://huggingface.co/datasets/KevinKeller/cognitive-pattern-selector-v1.time-embed-bge-m3
Korean Temporal Query Embedding Data for BGE-M3 - v1.9 Semantic Retention
This dataset is a FlagEmbedding/BGE-M3 fine-tuning dataset for Korean LMS temporal retrieval.
The objective is to keep semantically equivalent Korean temporal expressions close in embedding space while separating Korean expressions that look similar but mean different time ranges. The root files now point to the v1.9 training dataset.
What v1.9 Adds
v1.9 keeps the v1.7 calendar-focused… See the full description on the dataset page: https://huggingface.co/datasets/kev-KOH/time-embed-bge-m3.indonesia-law-qa-embeddingssn38-submissionvulnfixes-webTanuki-Phase2-annotation-datasetNCS-CorpusSD_promptskeval-testset
keval_test
The keval-testset is a dataset designed for training and validating the keval model.
The keval model follows the LLM-as-a-judge approach, which evaluates LLMs by assessing their responses to prompts from the ko-bench dataset. In other words, the keval model assigns scores to LLM-generated responses based on predefined evaluation criteria.
The keval-testset serves as a crucial resource for training and validating the keval model, enabling precise benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/davidkim205/keval-testset.sanpo_annotationsid2223_exam_prep
ID2223 Exam Prep Dataset (Custom, Lecture-Derived)
This dataset contains a curated collection of exam-style questions, explanations, study prompts, and answer–solution pairs for the KTH course ID2223 – Scalable Machine Learning and Deep Learning.
It was constructed from the official ID2223 lecture slides
It is designed to fine-tune small LLMs (1B–3B) into course-specialized teaching assistants that help students practice exam questions and understand the material more deeply.… See the full description on the dataset page: https://huggingface.co/datasets/kevembuvak/id2223_exam_prep.livermedqabigfivepersonalitieskevin009__llamaRAGdrama-details
Dataset Card for Evaluation run of kevin009/llamaRAGdrama
Dataset automatically created during the evaluation run of model kevin009/llamaRAGdrama
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kevin009__llamaRAGdrama-details.cognitive-question-generator-v1
Cognitive Question Generator Dataset
Dataset for fine-tuning an expert analysis and question generation model. Contains 5,637 prompt-response pairs capturing expert reasoning patterns for technology transactions and product counseling.
Dataset Description
This dataset was generated from the CognitiveTrainer platform's Mode 1 (Expert Analysis) system, capturing:
Initial scenario analysis
Claim validation with chain-of-trust
Multi-turn expert dialogue
Final synthesis… See the full description on the dataset page: https://huggingface.co/datasets/KevinKeller/cognitive-question-generator-v1.chimere-quality-scores
Chimere Quality Scores
Quality evaluation data from the Chimere self-improving inference system.
Files
quality_scores.jsonl — 104 quality assessments with multi-scorer evaluation (ThinkPRM, Qwen3.5, Qwen-9B)
training_pairs.jsonl — 68 high-quality training pairs with full chain-of-thought reasoning traces
Format
Each quality score entry contains: timestamp, route, score (1-5), reasoning, verification chain-of-thought.
Each training pair contains: prompt… See the full description on the dataset page: https://huggingface.co/datasets/Kevletesteur/chimere-quality-scores.
