datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GLM-5.3-Flash-BF16-Teacher-Logits
GLM-5.3-Flash BF16 teacher logits
This dataset contains full-vocabulary float32 teacher logits from the immutable
zai-org/GLM-5.3-Flash-BF16 revision a6c167b62691b2bac901344b65cb651a70f53e43.
It keeps the sealed final KLD panel qualification-only and publishes the
separate non-final calibration panel under role-specific paths.
Qualification-only final windows: 25
Qualification-only final prediction positions: 51175
Vocabulary size: 154880
Teacher receipt:… See the full description on the dataset page: https://huggingface.co/datasets/brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits.glm-5.3-flash-distillation-chat
Private distill of domofon/finetome-cot-100k instructions through GLM-5.3-Flash (AutoClaw / Z.AI).
Split
train — successful generations only.
field
description
instruction
user prompt from FineToMe
response
GLM final answer (message.content)
reasoning
GLM chain-of-thought (reasoning_content), empty if not captured
finish
stop or length
prompt_tokens / completion_tokens / reasoning_tokens
usage
latency_s
request latency
source_index
original FineToMe… See the full description on the dataset page: https://huggingface.co/datasets/best-distill/glm-5.3-flash-distillation-chat.recursive-task-synthesis-glm-5.3-rollouts
GLM 5.3 agentic rollouts on Recursive-Task-Synthesis
This dataset catalogs the full collection made from the pinned
Recursive-Task-Synthesis dataset revision
be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards.
Contents at a glance
Item
Count
Source tasks considered
37,284
Source candidates inspected
19,368
Converted tasks after source filters
18,600
Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.glm-5.3-flash-ifbench-openrouter
GLM-5.3-Flash IFBench OpenRouter five-run results
This dataset contains content-free results from an independent five-run
evaluation of z-ai/glm-5.3-flash on the official IFBench test set through
OpenRouter's first-party Z.AI provider.
This is not an official Allen Institute for AI, Z.AI, or OpenRouter result.
The evaluated outputs were AI-generated. Prompt, response, and reasoning text
are not included.
Results
Mean prompt-level loose accuracy was 65.5333% across… See the full description on the dataset page: https://huggingface.co/datasets/noahyoungs/glm-5.3-flash-ifbench-openrouter.glm-5.3-flash-mathnet-bon
glm-5.3-flash-mathnet-bon
This is the continuation and the final set of ox-alpha-mathnet-bon.
Verified chain-of-thought reasoning traces for competition mathematics, generated with GLM-5.3-Flash via best-of-N rejection sampling against the ShadenA/MathNet dataset (ICLR 2026).
Statistics (this split)
Metric
Value
Records (problem × attempt)
6,181
Distinct problems
848
Attempts per problem
7.29 (mean), 8 (max)
Accepted (answer_correct = true)
3,705… See the full description on the dataset page: https://huggingface.co/datasets/zakoman/glm-5.3-flash-mathnet-bon.glm-5.3-flash-function-calling
GLM-5.3-Flash function-calling (synthetic)
500 synthetic function-calling training samples generated with
zai-org/GLM-5.3-Flash via HF Inference
Providers (auto-routing with novita / together / fireworks fallback), 2026-09-21.
Schema
Each row:
messages — chat in OpenAI tool-use format (system / user / assistant); the final
assistant message either carries tool_calls (with JSON-string arguments) or is a plain-text answer
tools — 1–4 tool schemas in OpenAI function… See the full description on the dataset page: https://huggingface.co/datasets/Offlin33er/glm-5.3-flash-function-calling.
