datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GLM-5.1-Reasoning-1M-Think-Embeddings
GLM-5.1-Reasoning-1M-Think-Embeddings (main subset, partial)
Embeddings of the think-block content of every record in the main
subset of
Jackrong/GLM-5.1-Reasoning-1M-Cleaned,
embedded with Qwen/Qwen3-Embedding-0.6B via vLLM.
What this dataset is
Each row is the vector representation of just the reasoning trace (the
text between <think>...</think>), not the user prompt and not the final
answer. Useful for:
searching / clustering the original reasoning traces by… See the full description on the dataset page: https://huggingface.co/datasets/shanaka95/GLM-5.1-Reasoning-1M-Think-Embeddings.embeddings_all_correct_continuationsembeddings_all_wrong_continuationsb2_code_embedding_filter_open_code_reasoningembeddings_all_continuationsembeddings_wrong_continuations_10tokmedical-o1-reasoning-SFT_embeddings
