datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3_4b_instruct_lcbv6_rsa_pop_32_k_4_steps_10_s_65_e_1310_25_rsa_pop_32_k_4_steps_10_v2Math-steptok-steps-mcvalue-train-part4-of-525_50_rsa_pop_32_k_4_steps_10_v2exp_rob_dfiltered_Phi-4-reasoning-plus_2_mbenign_complete_step_t30exp_rob_dfiltered_Phi-4-reasoning-plus_2_mbenign_complete_step_t10algo-sft-eval-traces-cellular-automata-step-simulation-d5-v4
algo-sft-eval-traces-cellular-automata-step-simulation-d5-v4
Full eval traces for algo-sft-cellular-automata-step-simulation-d5 across test/harder/ood splits
Dataset Info
Rows: 2000
Columns: 11
Columns
Column
Type
Description
question_id
Value('string')
Unique question identifier from eval set
split
Value('string')
Evaluation split: test (in-distribution), harder (scaled up), ood (structural out-of-distribution)
domain
Value('string')
Task… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algo-sft-eval-traces-cellular-automata-step-simulation-d5-v4.DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 128,000
Unique prompts: 32,000
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and teacher… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4.exp_rob_dfiltered_Phi-4-reasoning-plus_2_mbenign_complete_step_t70DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 133,184
Unique prompts: 33,296
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4.p10-ttt-021125-overnight-final-run-continue2-step4-child0DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data-all33296-n4
DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data-all33296-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-12b-pt-from-step60-seed43, subfolder step_000020
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 133,184
Unique prompts: 33,296
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data-all33296-n4.step4_sllm_v2
Step 4 PII 후보 판정 데이터셋 — 최종 10,000개
RAG 답변에서 상위 NER 단계가 추출한 후보가 문맥상 특정 자연인의 개인정보인지 PII 또는 NOT_PII로 판정하도록 Qwen을 SFT하기 위한 합성 데이터셋이다.
바로 사용하는 파일
step4_final_10000_qwen_train.jsonl: 학습 8,000개
step4_final_10000_qwen_valid.jsonl: 검증 1,000개
step4_final_10000_qwen_test.jsonl: 최종 평가 1,000개
step4_final_10000_qwen_all.jsonl: 전체 확인용 10,000개
각 행의 최상위 필드는 messages 하나뿐이며 system, user, assistant 순서다. 학습 시 Qwen tokenizer의 chat template를 적용하고 assistant 응답 부분에만 loss를 계산한다.… See the full description on the dataset page: https://huggingface.co/datasets/bbanany/step4_sllm_v2.type4_8k_type3_1k_plus_sftloss_step200_no_eot_tmp10Qwen2.5-1.5B-DeepMath-sft-step1400-stage1-dapo-level3-4-rollout-4-max-len-4608-rollouts75_100_rsa_pop_32_k_4_steps_10_v22d_text_instruct_4stepQwQ-Long-CoT-10k-subset-llama3.1-8b-Inst-GPT4-Step-Perturbation-8-rejectsbigmath-custom-checkpoint-step-by-step-confidence-ckpt-8192-v4gaia_127_g1_diverse_tezos_top4_31600_32b_step900_20260520_193028qwen15b_raftpp_n4_step30_with_score_passncfa_extracted_exercise_sup_sample_from_policy_v1.1_stepwise_dpo_binarized_chunk_4agent_steps_huggingface_course_unit4VisualPRM300Kv0-full-dataset-mc0-o4-judge-incorrect-step-qwen-format50_75_rsa_pop_32_k_4_steps_10_v2qwen3_4b_instruct_start_325_end_350_rsa_pop_32_k_4_steps_10_timeout_53k_forcing_400_mask25_step4_022525connection_queries_jan12_natural_verbalized_1_step_None_0.7_4096_gpt-4_1-mini-2025-04-14
Dataset: connections-dev/connection_queries_jan12
This dataset was generated using the inference script with the following configuration:
Inference Parameters
Model Configuration
Model Name: gpt-4.1-mini-2025-04-14
Server URL: Not specified
API Key: Not provided
Request Timeout: 30 seconds
Query Configuration
Query Type: natural
Query Column: query
Sampling Type: verbalized
Generation Parameters
Temperature: 0.7
Max Tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/connections-dev/connection_queries_jan12_natural_verbalized_1_step_None_0.7_4096_gpt-4_1-mini-2025-04-14.exp_rob_dfiltered_Phi-4-reasoning-plus_2_mbenign_complete_step_t50qwen3_4b_instruct_start_25_end_50_rsa_pop_32_k_4_steps_10_5_timeout_5
