CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match N8 Rejection Sampling (Soft Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project How… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match.tabular100K<n<1M0 likes82 downloads7mo agoHugging Face02Leopo1d /OpenVul_Rejection_Sampling_based_Vulnerability_Reasoning_Dataset_for_SFTThis dataset provides high-quality, correctness-filtered vulnerability reasoning data to support the SFT of specialized VD LLMs for future research. text1K<n<10K1 likes67 downloads7mo agoHugging Face03marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match N8 Rejection Sampling (Strict Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match.tabular10K<n<100K0 likes55 downloads7mo agoHugging Face04marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n1-rejection-sampling-quantity-match N1 Rejection Sampling (Quantity Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n1-rejection-sampling-quantity-match.tabular10K<n<100K0 likes46 downloads7mo agoHugging Face05jacobmorrison /rejection_sampling_10627_fixed Dataset Card for "rejection_sampling_10627_fixed" More Information needed text100K<n<1M0 likes41 downloads2y agoHugging Face06marin-community /open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n1-rejection-sampling-quantity-match Qwen3-32B Math Rejection Sampling (Quantity Match) with Qwen3-235B-A22B Verifier Overview This dataset was created via rejection sampling from the Qwen3-32B response dataset using Qwen3-235B-A22B answers as ground truth. Source dataset (Qwen3-32B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-235B-A22B, 1 response per prompt):… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n1-rejection-sampling-quantity-match.tabular10K<n<100K0 likes30 downloads7mo agoHugging Face07JakeOh /rejection-sampling-llada-1.0-sft-gsm8ktext1K<n<10K0 likes26 downloads9mo agoHugging Face08budgiesarecooliguess /rejection-sampling-replace-repeat 🧠 Amazon Semantic-ID Recommendation + Reasoning — GPT-5.4 regenerated GPT-5.4 enrichment variant of yufan/recsys-genrec-dataset &nbsp;·&nbsp; Reasoning over Semantic IDs Enhances Generative Recommendation This repo mirrors the format of the original yufan/recsys-genrec-dataset, but the LLM-enriched fields are regenerated with GPT-5.4 (Azure OpenAI). The Video Games domain is fully populated. <cat> below is Video_Games. ✅ Currently available config… See the full description on the dataset page: https://huggingface.co/datasets/budgiesarecooliguess/rejection-sampling-replace-repeat.text100K<n<1M0 likes26 downloads1mo agoHugging Face09alehc /rejection-sampling-QA Rejecction Sampling Q&A This dataset is a very small curated question-answer pairs. The questions were hand-crafted to test the model's capabilities to follow instruction across various domains. The answers were generated using Microsoft's Phi-2 and curated using OpenAssistant's Large DeBERTa v3 Reward Model v2. Dataset Details The answers of this dataset were generated by prompting Microsoft's Phi-2 using a prompt format inspired by Stanford's Alpaca to help the LLM… See the full description on the dataset page: https://huggingface.co/datasets/alehc/rejection-sampling-QA.texttext-generationn<1K0 likes24 downloads3y agoHugging Face10yimingzhang /hh-rlhf-safety-v2-rejection-samplingtext10K<n<100K0 likes20 downloads2y agoHugging Face11LangAGI-Lab /rejection_samplingtext10K<n<100K0 likes20 downloads2y agoHugging Face12vwxyzjn /rejection_sampling_23251_messagestabular100K<n<1M0 likes20 downloads2y agoHugging Face13RRLHF /gemma2b_combined_iter1_rejection_samplingtext10K<n<100K0 likes15 downloads2y agoHugging Face14anonymousatom /300_RejectionSampling_LLMInputtextn<1K0 likes15 downloads11mo agoHugging Face15budgiesarecooliguess /igc1-top50-rejection-sampling 🧠 Amazon Semantic-ID Recommendation + Reasoning — GPT-5.4 regenerated GPT-5.4 enrichment variant of yufan/recsys-genrec-dataset &nbsp;·&nbsp; Reasoning over Semantic IDs Enhances Generative Recommendation This repo mirrors the format of the original yufan/recsys-genrec-dataset, but the LLM-enriched fields are regenerated with GPT-5.4 (Azure OpenAI). The Video Games domain is fully populated. <cat> below is Video_Games. ✅ Currently available config… See the full description on the dataset page: https://huggingface.co/datasets/budgiesarecooliguess/igc1-top50-rejection-sampling.text100K<n<1M0 likes14 downloads1mo agoHugging Face16yizhilll /demo_rejection_sampling_QA_phi-2_deberta-v3-large-v2_temp0.2This is a demo constructed dataset for alignment/preference learning. With paritially handcrafted questions (prompts), the answers are genreated by the phi-2 model with temperature 0.2 and the answers are scores select by the deberta-large-v2. The dataset containing questions and the selected answers from highest to lowest, decoding with rejection sampling K=8. Example loading: import datasets ds = datasets.load_dataset('yizhilll/demo_rejection_sampling_QA_phi-2_deberta-v3-large-v2_temp0.2')… See the full description on the dataset page: https://huggingface.co/datasets/yizhilll/demo_rejection_sampling_QA_phi-2_deberta-v3-large-v2_temp0.2.texttext-generationn<1K0 likes13 downloads3y agoHugging Face17budgiesarecooliguess /rejection-sampling-V4 🧠 Amazon Semantic-ID Recommendation + Reasoning — GPT-5.4 regenerated GPT-5.4 enrichment variant of yufan/recsys-genrec-dataset &nbsp;·&nbsp; Reasoning over Semantic IDs Enhances Generative Recommendation This repo mirrors the format of the original yufan/recsys-genrec-dataset, but the LLM-enriched fields are regenerated with GPT-5.4 (Azure OpenAI). The Video Games domain is fully populated. <cat> below is Video_Games. ✅ Currently available config… See the full description on the dataset page: https://huggingface.co/datasets/budgiesarecooliguess/rejection-sampling-V4.text100K<n<1M0 likes12 downloads1mo agoHugging Face18vwxyzjn /rejection_sampling_1722360907tabularn<1K0 likes11 downloads2y agoHugging Face19vwxyzjn /rejection_sampling_1722361005tabularn<1K0 likes11 downloads2y agoHugging Face20trungtvu /open-thoughts-code-rejection-samplingtextn<1K0 likes11 downloads2y agoHugging Face21budgiesarecooliguess /V2-rejection-sampling-dataset 🧠 Amazon Semantic-ID Recommendation + Reasoning — GPT-5.4 regenerated GPT-5.4 enrichment variant of yufan/recsys-genrec-dataset &nbsp;·&nbsp; Reasoning over Semantic IDs Enhances Generative Recommendation This repo mirrors the format of the original yufan/recsys-genrec-dataset, but the LLM-enriched fields are regenerated with GPT-5.4 (Azure OpenAI). The Video Games domain is fully populated. <cat> below is Video_Games. ✅ Currently available config… See the full description on the dataset page: https://huggingface.co/datasets/budgiesarecooliguess/V2-rejection-sampling-dataset.text100K<n<1M0 likes11 downloads1mo agoHugging Face22vwxyzjn /rejection_sampling_1722011047tabularn<1K0 likes8 downloads2y agoHugging Face23vwxyzjn /rejection_sampling_1722293991tabularn<1K0 likes8 downloads2y agoHugging Face24LangAGI-Lab /qwen-7b-verified-7k-rejection-sampling-alpaca-formattext1K<n<10K0 likes8 downloads2y agoHugging Face25vwxyzjn /rejection_sampling_1722361801tabularn<1K0 likes7 downloads2y agoHugging Face26budgiesarecooliguess /rejection-sampling 🧠 Amazon Semantic-ID Recommendation + Reasoning — GPT-5.4 regenerated GPT-5.4 enrichment variant of yufan/recsys-genrec-dataset &nbsp;·&nbsp; Reasoning over Semantic IDs Enhances Generative Recommendation This repo mirrors the format of the original yufan/recsys-genrec-dataset, but the LLM-enriched fields are regenerated with GPT-5.4 (Azure OpenAI). The Video Games domain is fully populated. <cat> below is Video_Games. ✅ Currently available config… See the full description on the dataset page: https://huggingface.co/datasets/budgiesarecooliguess/rejection-sampling.text100K<n<1M0 likes7 downloads2mo agoHugging Face27vwxyzjn /rejection_sampling_1722011582tabularn<1K0 likes6 downloads2y agoHugging Face28anony3849 /Rejection_Sampling_based_Vulnerability_Reasoning_Dataset_for_SFTtext1K<n<10K0 likes6 downloads5mo agoHugging Face29vwxyzjn /rejection_sampling_1722011125tabularn<1K0 likes3 downloads2y agoHugging Face30vwxyzjn /rejection_sampling_1722011436tabularn<1K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.