CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /Sera-4.5A-Full-T1-v3 laion/Sera-4.5A-Full-T1-v3 Subset of allenai/Sera-4.5A-Full-T1. Size: 72,118 rows (full dataset: 72,118 rows). Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages field (as JSON string), instance_id, rollout_patch, func_name, func_path, problem_statement, target_patch, docker_image. Adds a source field pointing back to the parent dataset. Each assistant message carries a native tool_calls array (OpenAI tool-calling format) and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3.texttext-generation10K<n<100K0 likes53 downloads5mo agoHugging Face02laion /sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13 Nemotron Terminal SFT reproduction evaluation artifacts This repository contains the complete Harbor artifact tree for the 300-trial OpenThoughts-TBLite evaluation of laion/sft-repro-thinking-step630-nemotron-terminal-step1888. The checkpoint was trained from the Grug stage-2 thinking checkpoint on the Nemotron Terminal corpus for 1,888 steps. Result Measure Value Attempted / completed 300 / 300 Verifier-scoreable 259 (86.33%) Aggregate reward, all… See the full description on the dataset page: https://huggingface.co/datasets/laion/sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13.texttext-generationn<1K0 likes51 downloads1mo agoHugging Face03laion /Sera-4.5A-Full-T1-v3-1000 laion/Sera-4.5A-Full-T1-v3-1000 Subset of allenai/Sera-4.5A-Full-T1. Size: 1,000 rows (full dataset: 72,118 rows). Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages field (as JSON string), instance_id, rollout_patch, func_name, func_path, problem_statement, target_patch, docker_image. Adds a source field pointing back to the parent dataset. Each assistant message carries a native tool_calls array (OpenAI tool-calling format) and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-1000.texttext-generation1K<n<10K1 likes39 downloads5mo agoHugging Face04laion /Sera-4.5A-Full-T1-v3-3160 laion/Sera-4.5A-Full-T1-v3-3160 Subset of allenai/Sera-4.5A-Full-T1. Size: 3,160 rows (full dataset: 72,118 rows). Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages field (as JSON string), instance_id, rollout_patch, func_name, func_path, problem_statement, target_patch, docker_image. Adds a source field pointing back to the parent dataset. Each assistant message carries a native tool_calls array (OpenAI tool-calling format) and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-3160.texttext-generation1K<n<10K0 likes30 downloads5mo agoHugging Face05laion /Sera-4.6-Lite-T2-v4-1000 laion/Sera-4.6-Lite-T2-v4-1000 Row-subset of allenai/Sera-4.6-Lite-T2 (the dataset upstream SERA-8B was trained on), with OpenAI tool_calls pre-rendered into the content string as Hermes/Qwen3-style <tool_call>...</tool_call> wire tokens and tool responses wrapped as <tool_response>...</tool_response>. This mirrors Ai2's sera/datagen/data/postprocess/utils.py::transform_traj_hermes (default tool_call_format: "hermes") which is the missing step between the public Sera-4.6-Lite-T2… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.6-Lite-T2-v4-1000.texttext-generation1K<n<10K0 likes29 downloads5mo agoHugging Face06laion /Sera-4.5A-Full-T1-v3-10000 laion/Sera-4.5A-Full-T1-v3-10000 Subset of allenai/Sera-4.5A-Full-T1. Size: 10,000 rows (full dataset: 72,118 rows). Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages field (as JSON string), instance_id, rollout_patch, func_name, func_path, problem_statement, target_patch, docker_image. Adds a source field pointing back to the parent dataset. Each assistant message carries a native tool_calls array (OpenAI tool-calling format) and a train: bool flag… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-10000.texttext-generation10K<n<100K0 likes27 downloads5mo agoHugging Face07laion /CoderForge-Preview-v6-1000 laion/CoderForge-Preview-v6-1000 Row-subset of togethercomputer/CoderForge-Preview (trajectories split, filtered_reward1), rendered into Qwen3-compatible think-first OpenHands-XML wire format. Why v6? v3 (pre-tokenized) and v5 (wrapper-stripped, no think-block) both produced garbage at eval time (8888..., 0.0.0.0...) despite clean training losses. Root cause: stock Qwen3-8B assigns ~100% prior to <think> as the first token after <|im_start|>assistant. CoderForge's… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v6-1000.texttext-generation1K<n<10K0 likes27 downloads5mo agoHugging Face08laion /Sera-4.5A-Full-T1-v3-316 laion/Sera-4.5A-Full-T1-v3-316 Subset of allenai/Sera-4.5A-Full-T1. Size: 316 rows (full dataset: 72,118 rows). Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages field (as JSON string), instance_id, rollout_patch, func_name, func_path, problem_statement, target_patch, docker_image. Adds a source field pointing back to the parent dataset. Each assistant message carries a native tool_calls array (OpenAI tool-calling format) and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-316.texttext-generationn<1K0 likes25 downloads5mo agoHugging Face09laion /Sera-4.6-Lite-T2-v4-316 laion/Sera-4.6-Lite-T2-v4-316 Row-subset of allenai/Sera-4.6-Lite-T2 (the dataset upstream SERA-8B was trained on), with OpenAI tool_calls pre-rendered into the content string as Hermes/Qwen3-style <tool_call>...</tool_call> wire tokens and tool responses wrapped as <tool_response>...</tool_response>. This mirrors Ai2's sera/datagen/data/postprocess/utils.py::transform_traj_hermes (default tool_call_format: "hermes") which is the missing step between the public Sera-4.6-Lite-T2… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.6-Lite-T2-v4-316.texttext-generationn<1K0 likes23 downloads5mo agoHugging Face10laion /CoderForge-Preview-v6-316 laion/CoderForge-Preview-v6-316 Row-subset of togethercomputer/CoderForge-Preview (trajectories split, filtered_reward1), rendered into Qwen3-compatible think-first OpenHands-XML wire format. Why v6? v3 (pre-tokenized) and v5 (wrapper-stripped, no think-block) both produced garbage at eval time (8888..., 0.0.0.0...) despite clean training losses. Root cause: stock Qwen3-8B assigns ~100% prior to <think> as the first token after <|im_start|>assistant. CoderForge's… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v6-316.texttext-generationn<1K0 likes18 downloads5mo agoHugging Face11laion /sera-subset-mixed-316 sera-subset-mixed-316 Random subset of 316 rows drawn from ethanlshen/sera-subset, mixed across the two upstream stages (stage1 unresolved + stage2 resolved) and shuffled deterministically. Source Upstream: ethanlshen/sera-subset. Two upstream JSONLs are concatenated: 22972_0.88_stage1_scaling_final_glm46_e2e_1ipf_swesmith_unresolved_ipf_1_atk_rft-think_SYSTEM_SIMPLE.jsonl (22 972 rows)… See the full description on the dataset page: https://huggingface.co/datasets/laion/sera-subset-mixed-316.texttext-generationn<1K0 likes5 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.