datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Reverse-baseline-bias-unbiasOriginal-baseline-bias-unbiaspkm-agent-baseline-v2
PKM Agent Baseline — 500 + 50 Scenarios + Six-Grader Artifacts (v2)
Two deterministically generated, Korean-language benchmarks for evaluating multi-tool Personal Knowledge Management (PKM) agents over Notion, Gmail, and Google Calendar, plus the Six-Grader Ensemble scoring artifacts (100-scenario reference subset + per-scenario six-metric scores for vanilla and LoRA models).
Released alongside the preprint:
Vault-Grounded 4B Agent: A Hybrid Reasoning–Fact Architecture for Local… See the full description on the dataset page: https://huggingface.co/datasets/newtype-2038/pkm-agent-baseline-v2.dllm-qwen38-ar-baseline
AR baseline for the Qwen3.8-27B → block-diffusion conversion (GSM8K, pinned 500-problem subset)
日本語要約: Qwen/Qwen3.8-27B を Fast-dLLM v2 で
block-diffusion dLLM 化する計画の AR 参照スコアです。seed 固定の GSM8K 500 問・4-shot・
greedy で acc 0.968(484/500、skipped 0)。H100 1 枚で 42 分 ≈ $2.8。停止条件
(stop literal)として「block-diffusion 訓練 0.3B tokens の後、この subset で acc ≥ 0.918」
を要求し、届かなければ変換を続けません。訓練 run 自体はこの判断待ちで held です。
Why this exists — the stop literal
We are converting Qwen/Qwen3.8-27B into… See the full description on the dataset page: https://huggingface.co/datasets/com-junkawasaki/dllm-qwen38-ar-baseline.ai-brand-mention-baseline-2026
AI Brand Mention Baseline 2026
A longitudinal benchmark dataset measuring how frontier LLMs (Gemini 2.5,
GPT-4 class, Claude class) mention a single AI-native company (Neo
Genesis) when prompted with content-gap probes. First open dataset of
its kind for GEO (Generative Engine Optimization) research.
Metric
Value
Measurements
486
Window
2026-04-28 to 2026-05-07 (10 days)
Distinct seed prompts
30
Categories
6 (definition, pricing, comparison, problem_solving… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/ai-brand-mention-baseline-2026.
