datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RTL-Coder_7b_reasoning_tb_combined
Verireason-RTL-Coder_7b_reasoning_tb_combined
For implementation details, visit our GitHub repository: VeriReason
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
This is the combined version of VeriReason-RTL-Coder_7b_reasoning_tb and VeriReason-RTL-Coder_7b_reasoning_tb_simple.
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_combined
Project… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_7b_reasoning_tb_combined.VeriReason-RTL-Coder_7b_reasoning_tb_simple
Verireason-RTL-Coder_7b_reasoning_tb_simple
For implementation details, visit our GitHub repository: VeriReason and our page
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_simple
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb_simple.VeriReason-RTL-Coder_7b_reasoning_tb
Verireason-RTL-Coder_7b_reasoning_tb
For implementation details, visit our GitHub repository: VeriReason and our page
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb.Convergent-7B-data
Convergent-7B Training Data
Training data for the bigcompute.science research companion model.
Early Preview — This dataset is a work in progress. It is expressly designed to train a research assistant for the bigcompute.science MCP server as part of the Convergent conjecture-driven GPU research project. The dataset will be updated frequently as new experiments, findings, and tool definitions are added. Expect changes to schema, tool names, and content until we reach a GA… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/Convergent-7B-data.RTL-Coder_7b_reasoning
Verireason-RTL-Coder_7b_reasoning_tb_simple
For implementation details, visit our GitHub repository: VeriReason
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to enhance the… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_7b_reasoning.polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps
Cross-tokenizer ROSE rollouts — Olmo-3-7B-Think-SFT ← Qwen3-30B-A3B-Thinking-2507
Every assembled row of a complete 240-step online-ROSE run: 61,440 rows, the teacher's
actual continuation for each, and the token accounting behind it.
The student writes a 4096-token prefix in its own vocabulary (100278). That prefix is
decoded to text, the teacher is shown it under its own chat template, and the teacher's
reply comes back as text and is tokenised into the student's vocabulary.… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps.InverseCoder-CL-7B-Evol-Instruct-90K
InverseCoder: Unleashing the Power of Instruction-Tuned Code LLMs with Inverse-Instruct
InverseCoder is a series of code LLMs instruction-tuned by generating data from itself through Inverse-Instruct.
Models and Datasets
Base Model
InverseCoder
Dataset
6.7B
deepseek-ai/deepseek-coder-6.7b-base
wyt2000/InverseCoder-DS-6.7B
wyt2000/InverseCoder-DS-6.7B-Evol-Instruct-90K
7B
codellama/CodeLlama-7b-Python-hf
wyt2000/InverseCoder-CL-7B… See the full description on the dataset page: https://huggingface.co/datasets/wyt2000/InverseCoder-CL-7B-Evol-Instruct-90K.das-dpo-data-searchr1-7b
DAS dpo data-searchr1-7b
This dataset provides DPO preference data for post-training SearchR1 7B search agents with DAS.
The file is provided in LLaMA-Factory compatible DPO format with prompt, chosen, rejected, and optional system fields.
Medical-QA-Mistral7B-Finetuningcode_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288
code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12)
Pass@k completions generated with vLLM over the prefixes in
CL-From-Nothing/code_rose_initial_1_7B_SFT_10K.
Generation config
Model
Qwen3-4B-Thinking-2507
Samples per question (k)
12
Temperature
0.7
top_p
0.9
max_tokens
12288
max_model_len
32768
Questions
7250 (index 0–7249, full split)
Total rows
87000 (7250 × 12)
Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.ThinkSafe-R1-Distill-7B
ThinkSafe Dataset
This dataset is associated with the paper THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.
Paper: https://arxiv.org/abs/2601.23143GitHub: https://github.com/seanie12/ThinkSafe.git
Citation
If you use this dataset, please cite:
@article{lee2025thinksafe,
title={THINKSAFE: Self-Generated Safety Alignment for Reasoning Models},
author={Lee, Seanie and others},
journal={arXiv preprint arXiv:2601.23143},
year={2025}
}
PiCo-dataset-general-7b
PiCo-7B Instruction Dataset
A high-quality, high-variability instruction dataset for PiCo-7B (ArcOffical/PiCo-7B), an approximately 6.95B-parameter (~7B) large language model featuring an Adaptive Hierarchical Mixture of Experts (AHMoE) architecture. The official model card describes 30 layers, including 15 MoE layers and 15 dense layers, with approximately 2.63B active parameters per token and a 131,072-token context window.[^1]
Purpose
This dataset teaches… See the full description on the dataset page: https://huggingface.co/datasets/ArcOffical/PiCo-dataset-general-7b.JMT-Bench-result_self-rewarding_Mistral-7B-lora
JMT-Bench result
Answer language
JMT-Benchの回答のうち、Englishで回答した件数
Model
Count
mistralai/Mistral-7B-v0.3
25
HachiML/Mistral-7B-v0.3-m1-lora
7
HachiML/Mistral-7B-v0.3-m2-lora
7
HachiML/Mistral-7B-v0.3-m3-lora
2
DeepScale-qwen2.5_7b-multi_16kmessages是7b模型生成的结果
correct是根据messages最后一个输出的答案进行验证
SWE-Mini-337_Qwen7B
SWE-Mini-337_Qwen7B
A small, synthetically generated dataset of software engineering tasks — each entry pairs a task description with a step-by-step decomposition plan and a complete, runnable Python implementation.
Dataset Summary
Samples: 337
Generator model: Qwen/Qwen2.5-Coder-7B-Instruct (4-bit NF4 quantized)
Format: JSONL, one JSON object per line
Domain coverage: 20 software engineering domains (see below)
This is a small, single-model, unfiltered batch.… See the full description on the dataset page: https://huggingface.co/datasets/Rumiii/SWE-Mini-337_Qwen7B.Cybersecurity_Reasoning_Dataset_MistralFamily_7b
Cybersecurity Reasoning Dataset (v6.0)
A high-fidelity, forensics-mapped training corpus for cybersecurity reasoning and SOC
automation — formatted for the Mistral / Llama instruct family.
Format-specific dataset. Every record uses the Alpaca-style
### Instruction: / ### Response: template native to Mistral/Llama instruct models.
A model-agnostic version of this corpus (Mistral, DeepSeek, ChatML, and Gemma
variants rendered from one neutral canonical source) is published… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset_MistralFamily_7b.
