datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
surface-audit
TruthfulQA-476 — a surface-form-cleaned binary-choice TruthfulQA
TruthfulQA-476 is the recommended drop-in replacement for the binary-choice TruthfulQA
evaluation set. It keeps 476 of the 790 original question pairs, in the original schema, chosen so
that a classifier restricted to six surface features of the answer text (negation, hedging, length,
token statistics) barely separates correct from incorrect answers (AUC 0.528, at the edge of statistical detectability), while the… See the full description on the dataset page: https://huggingface.co/datasets/iclr2027-surface-audit/surface-audit.glm53-flash-function-calling
GLM-5.3-Flash Function Calling (synthetic)
A synthetic function-calling dataset generated with zai-org/GLM-5.3-Flash via Hugging Face Inference Providers. Every example is schema-validated and deduplicated before it is written.
Splits
File
Format
Rows
data/messages.jsonl
OpenAI messages (multi-turn)
—
data/xlam.jsonl
xlam single-turn
—
The table is filled in when the generation run completes.
messages format
{
"id":… See the full description on the dataset page: https://huggingface.co/datasets/Surfdan/glm53-flash-function-calling.
