latentmd-neurips26/LatentMD
LatentMD Benchmarking Markdown Boundary Failures in LLM-Generated Text — NeurIPS 2026 E&D Track submission. This repository hosts the dataset artifact for LatentMD: the 4,179 prompts that constitute the benchmark, ~37,000 reference responses from 9 frontier models, and illustrative output samples. The accompanying evaluation code (CLI, metric definitions, statistical tests) lives in a separate code repository on GitHub under MIT. What LatentMD measures LLM… See the full description on the dataset page: https://huggingface.co/datasets/latentmd-neurips26/LatentMD.
081
