CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hanlincs /in1k_clip_qwen25vl_3b_224res_64tokens_new_pttabular1M<n<10M0 likes517 downloads1y agoHugging Face02hanlincs /in1k_clip_qwen25vl_3b_448res_256tokens_new_merged_pttabular1M<n<10M0 likes471 downloads1y agoHugging Face03chrislimbe /pubmedqa-recursive-llm-degradation-qwen2.5-0.5b PubMedQA Recursive LLM Degradation — Qwen2.5-3B This repository contains synthetic biomedical question-answering data and model predictions generated as part of a study of recursive fine-tuning and model degradation. Base Model Qwen/Qwen2.5-3B Source Dataset The experiments use the PubMedQA dataset: qiaoxin/PubMedQA This repository contains generated/derived research artifacts and does not redistribute the original PubMedQA dataset in its entirety.… See the full description on the dataset page: https://huggingface.co/datasets/chrislimbe/pubmedqa-recursive-llm-degradation-qwen2.5-0.5b.tabularquestion-answering10K<n<100K0 likes139 downloads2d agoHugging Face04chrislimbe /pubmedqa-recursive-llm-degradation-qwen2.5-3b PubMedQA Recursive LLM Degradation — Qwen2.5-3B This repository contains synthetic biomedical question-answering data and model predictions generated as part of a study of recursive fine-tuning and model degradation. Base Model Qwen/Qwen2.5-3B Source Dataset The experiments use the PubMedQA dataset: qiaoxin/PubMedQA This repository contains generated/derived research artifacts and does not redistribute the original PubMedQA dataset in its entirety.… See the full description on the dataset page: https://huggingface.co/datasets/chrislimbe/pubmedqa-recursive-llm-degradation-qwen2.5-3b.tabularquestion-answering10K<n<100K0 likes136 downloads2d agoHugging Face05crosslingual-em /Qwen2.5-7B-Instruct-em-evaldocumentn<1K0 likes70 downloads5mo agoHugging Face06shirasko /qwen2.5-snmf-features Qwen2.5 SNMF features for unlearning MLP Semi-NMF factorizations and (when present) LLM interpretations for SNMF concept unlearning on Qwen/Qwen2.5-3B-Instruct (rank 100, seed 42). These files are the SNMF track only: per-layer MLP directions used to project concept features out of up_proj / down_proj. There is no embedding-matrix factorization in this dataset. Qwen/Qwen3.5-2B features live in shirasko/qwen-snmf-features. Layout (same directory scheme used by the… See the full description on the dataset page: https://huggingface.co/datasets/shirasko/qwen2.5-snmf-features.tabularn<1K0 likes64 downloads20d agoHugging Face07Turbs /resid-xprmt-generate-qwen2.5-7bimage10K<n<100K0 likes56 downloads5mo agoHugging Face08kaengreg /wikifacts-sents-qwen2.5-32b-qrelstext1K<n<10K0 likes53 downloads1y agoHugging Face09mitroitskii /Crosscoder-Qwen2.5-1.5B-vs-DeepScaleR-1.5B_max_activating_examplesSee Files and versions for pickled dictionaries and database versions of of max activating examples organized per available layer, as well as dataframes of available features. tabular100K<n<1M0 likes38 downloads2y agoHugging Face10mikheevshow /SIGNAL-Dataset-Hiddens-Qwen-Qwen2.5-7Btextn<1K0 likes25 downloads11mo agoHugging Face11mikheevshow /SIGNAL-Dataset-Hiddens-Qwen-Qwen2.5-7B-Instructtextn<1K0 likes21 downloads11mo agoHugging Face12suhanii23 /qwen2.5-3b-blind-spots Qwen2.5-3B Factual Recall Blind Spots Model Tested Qwen/Qwen2.5-3B A 3.09B parameter base causal language model, pretrained only. How I Loaded the Model from transformers import AutoModelForCausalLM, AutoTokenizer import torch model_name = "Qwen/Qwen2.5-3B" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained( model_name, torch_dtype=torch.float16, device_map="auto" ) Loaded on Google Colab.… See the full description on the dataset page: https://huggingface.co/datasets/suhanii23/qwen2.5-3b-blind-spots.texttext-generationn<1K0 likes19 downloads7mo agoHugging Face13ESITime /T2G-1k-Qwen2.5-3B T2G Overview T2G is a synthetic data consisting of text-graph pairs designed to finetune LLMs on information extraction tasks, specifically text-to-graph conversion. Dataset Structure The dataset is organized into the following main components: Train Set: 800 instances for training models. Validation Set: 100 for validating model performance. Test Set: 100 instances for final evaluation. Data Fields Each instance in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/ESITime/T2G-1k-Qwen2.5-3B.text1K<n<10K0 likes17 downloads2y agoHugging Face14MLZoo /DPO-bad-boy-chinese-for-Qwen2.5 BTF Chinese DPO Dataset 数据集描述 这是一个用于中文大语言模型DPO (Direct Preference Optimization) 训练的数据集。该数据集经过特殊处理,用于测试调整模型的语言风格和表达方式。 ⚠️ 免责声明: 此数据集仅供研究目的使用。请负责任地使用,并遵守相关法律法规和道德准则。 数据集统计 训练集大小:80% 的总数据 测试集大小:20% 的总数据 数据格式:instruction-following 格式[INST] 问题 [/INST] 回答 使用方法 from datasets import load_dataset dataset = load_dataset("your-username/dataset-name") 数据集结构 数据集包含两个分片: train: 训练集 test: 测试集 每条数据包含以下字段: text: 包含指令和回答的完整文本 局限性和注意事项… See the full description on the dataset page: https://huggingface.co/datasets/MLZoo/DPO-bad-boy-chinese-for-Qwen2.5.text1K<n<10K8 likes15 downloads2y agoHugging Face15gsingh1-py /Qwen2-72btext1K<n<10K0 likes12 downloads2y agoHugging Face16optimization-hashira /feature-extract-qwen2-vltext100K<n<1M0 likes11 downloads2y agoHugging Face17POLAROID86 /FINE_TUNING_LLM_QWEN2image10K<n<100K0 likes9 downloads2y agoHugging Face18hbXNov /qwen_2.5_7b_soln_gpt_4o_verifytabular10K<n<100K0 likes9 downloads2y agoHugging Face19qqqqq1 /qwen2-fukellm-ragastextn<1K0 likes8 downloads2y agoHugging Face20ESITime /T2G-Event-1k-Qwen2.5-3B-changed-formattext1K<n<10K0 likes7 downloads2y agoHugging Face21matonski /diffing-stats-SAEdiff_ftb-qwen25_0_5B_instruct-ebma-L11-s2-t100-k100-lr1e-04-x2tabular1K<n<10K0 likes7 downloads1y agoHugging Face22asadkhn /qwen25-3b-medical-blindspots Blind Spots of Qwen2.5-3B (Base Model) — Medical & Epidemiology Domain Dataset Summary This dataset documents 10 diverse failure cases of the Qwen/Qwen2.5-3B base language model, focusing on medical, epidemiological, and multilingual (Urdu) prompts. It was created as part of the Fatima Fellowship 2026 application technical challenge. Each row contains: input: the prompt given to the model expected_output: the correct answer model_output: what the model actually generated… See the full description on the dataset page: https://huggingface.co/datasets/asadkhn/qwen25-3b-medical-blindspots.textn<1K0 likes7 downloads7mo agoHugging Face23ESITime /T2G-Event-1k-Qwen2.5-3Btext1K<n<10K0 likes5 downloads2y agoHugging Face24Abdou220 /qwen2-1.5b-blindspotsimport torchfrom transformers import AutoTokenizer, AutoModelForCausalLMmodel_name = "Qwen/Qwen2-1.5B"device = "cuda" if torch.cuda.is_available() else "cpu"tokenizer = AutoTokenizer.from_pretrained(model_name)model = AutoModelForCausalLM.from_pretrained( model_name, torch_dtype=torch.bfloat16, device_map="auto")def generate_response(prompt): inputs = tokenizer(prompt, return_tensors="pt").to(device) outputs = model.generate(**inputs, max_new_tokens=50) response =… See the full description on the dataset page: https://huggingface.co/datasets/Abdou220/qwen2-1.5b-blindspots.textn<1K0 likes3 downloads7mo agoHugging Face25hchrAsma /qwen2.5-1.5b-blind-spots Qwen2.5-1.5B Blind Spots Dataset Model Tested Model: Qwen/Qwen2.5-1.5B Parameters: 1.5 Billion Type: Base language model (not instruction-tuned) Release date: Within the last 6 months How I Loaded the Model !pip install transformers accelerate -q from transformers import AutoTokenizer, AutoModelForCausalLM import torch model_name = "Qwen/Qwen2.5-1.5B" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained(… See the full description on the dataset page: https://huggingface.co/datasets/hchrAsma/qwen2.5-1.5b-blind-spots.textn<1K0 likes2 downloads7mo agoHugging Face26Maham789 /blind-spots-qwen2.5 Blind Spots of Qwen2.5-1.5B Model Tested Model: Qwen/Qwen2.5-1.5B Link: https://huggingface.co/Qwen/Qwen2.5-1.5B Type: Base language model (not finetuned) Parameters: 1.5 Billion How I Loaded the Model I used Google Colab with a free T4 GPU. from transformers import AutoTokenizer, AutoModelForCausalLM import torch model_name = "Qwen/Qwen2.5-1.5B" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained(… See the full description on the dataset page: https://huggingface.co/datasets/Maham789/blind-spots-qwen2.5.textn<1K0 likes2 downloads5mo agoHugging Face271024m /Qwen-2.5-14B-Hindi-Instruct-Datagatedtext100K<n<1M1 likes1 downloads2y agoHugging Face28nolangclem /bigmath-grpo-rollouts-qwen25-3btabular100K<n<1M0 likes1 downloads7mo agoHugging Face29fati2baali /qwen2.5-1.5B-blind-spotstextn<1K0 likes1 downloads7mo agoHugging Face30noorfatima67 /Qwen2_13_Blind_Spots Qwen-2 13 Blind Spots Dataset This dataset contains 13 diverse blind spots of the Qwen2.5-3B base model (Hugging Face). Each row contains: input: The prompt given to the model expected_output: The correct answer model_output: The answer the model actually produced Model Tested Model: Qwen2.5-3B (Base) Link: https://huggingface.co/Qwen/Qwen2.5-3B How the Model Was Loaded from transformers import AutoTokenizer, AutoModelForCausalLM import torch… See the full description on the dataset page: https://huggingface.co/datasets/noorfatima67/Qwen2_13_Blind_Spots.textn<1K0 likes1 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.