CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vandijklab /immune-c2s Overview Cell2Sentence is a novel method for adapting large language models to single-cell transcriptomics. We transform single-cell RNA sequencing data into sequences of gene names ordered by expression level, termed "cell sentences". This dataset was constructed from the immune tissue dataset in Domínguez et al., and it was used to train the Pythia-160m model capable of generating complete cells described in our paper. Details about the Cell2Sentence transformation and… See the full description on the dataset page: https://huggingface.co/datasets/vandijklab/immune-c2s.texttext-generation100K<n<1M3 likes149 downloads3y agoHugging Face02vanila434 /multilingual-elder-safety-msgs multilingual-elder-safety-msgs A hand-authored, multilingual elder fraud-recognition and safety coaching dataset. 467 curated scam/safe scenarios in Chinese and English, with platform-generated coaching responses localized across 5 languages: Chinese, English, Vietnamese, Khmer (Cambodian), and Lao. Expanded to 1,029 rows through Adaption Labs platform reasoning traces and multilingual adaptation. Built for communities where filial piety, authority deference, and fear of… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/multilingual-elder-safety-msgs.texttext-classification1K<n<10K0 likes123 downloads5mo agoHugging Face03VanishD /CodeGym Generalizable End-to-End Tool-Use RL with Synthetic CodeGym CodeGym is a synthetic environment generation framework for LLM agent reinforcement learning on multi-turn tool-use tasks. It automatically converts static code problems into interactive and verifiable CodeGym environments where agents can learn to use diverse tool sets to solve complex tasks in various configurations — improving their generalization ability on out-of-distribution (OOD) tasks. GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/VanishD/CodeGym.textquestion-answering100K<n<1M3 likes103 downloads11mo agoHugging Face04vanyacohen /CaT-Bench Dataset Card for CaT-Bench CaT-Bench is a benchmark dataset designed to evaluate large language models' (LLMs) understanding of causal and temporal dependencies in natural language plans, specifically in cooking recipes based on the English Recipe Flow Graph Corpus by Yamakata et al. (2020). It consists of questions that test whether one step must necessarily occur before or after another, requiring reasoning about preconditions, effects, and the overall structure of the plan.… See the full description on the dataset page: https://huggingface.co/datasets/vanyacohen/CaT-Bench.textquestion-answering1K<n<10K0 likes78 downloads2y agoHugging Face05vanhthefirst /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn Language:… See the full description on the dataset page: https://huggingface.co/datasets/vanhthefirst/vietnamese-legal-documents.texttext-classification1M<n<10M0 likes61 downloads6mo agoHugging Face06VanshTiwari15 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K1 likes57 downloads9mo agoHugging Face07Vanedap /MMLU-Pro MMLU-Pro Dataset MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. |Github | 🏆Leaderboard | 📖Paper | 🚀 What's New [2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro… See the full description on the dataset page: https://huggingface.co/datasets/Vanedap/MMLU-Pro.tabularquestion-answering10K<n<100K0 likes49 downloads2mo agoHugging Face08vanloc1808 /buddism-qa-dataset Buddhism Question-Answer Dataset A comprehensive Vietnamese-English Buddhism question-answering dataset created by merging and processing multiple Buddhism-related datasets. Dataset Description This dataset combines two high-quality Buddhism question-answer datasets to create a unified resource for training and evaluating models on Buddhism-related knowledge. The dataset contains questions and answers in both Vietnamese and English, making it suitable for multilingual… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/buddism-qa-dataset.textquestion-answering10K<n<100K0 likes37 downloads1y agoHugging Face09vanty120 /Gpt-5.4-Xhigh-Reasoning-2000x Gpt-5.4-Xhigh-Reasoning-2750x A premium-quality reasoning dataset containing 2,752 elite samples distilled from GPT-5.4 XHIGH (the highest reasoning effort tier of GPT-5.4). Each sample features deep, multi-step Chain-of-Thought traces that are significantly longer and more rigorous than standard GPT-5.4 outputs. This dataset is specifically designed for Supervised Fine-Tuning (SFT) to transform general-purpose language models into powerful reasoning models with explicit thinking… See the full description on the dataset page: https://huggingface.co/datasets/vanty120/Gpt-5.4-Xhigh-Reasoning-2000x.textquestion-answering1K<n<10K14 likes32 downloads6mo agoHugging Face10vancouverevs /afriadapt-agri-train-v2 AfriAdapt Agriculture Training Dataset (v2) High-quality instruction-response dataset focused on practical agricultural advice for smallholder farmers in East and Southern Africa. This is version 2 of the dataset, significantly expanded for the Adaption AutoScientist Challenge 2026. Dataset Summary Domain: Agriculture & Climate-smart farming Target users: Smallholder farmers and agricultural extension workers Geographic focus: Kenya, Tanzania, Uganda, Ethiopia… See the full description on the dataset page: https://huggingface.co/datasets/vancouverevs/afriadapt-agri-train-v2.texttext-generation1K<n<10K1 likes25 downloads2mo agoHugging Face11vancouverevs /afriadapt-agri-train AfriAdapt Agriculture Training Dataset High-quality instruction-response dataset focused on practical agricultural advice for smallholder farmers in East and Southern Africa. This dataset was created as part of the AfriAdapt project for the Adaption AutoScientist Challenge 2026. Dataset Summary Domain: Agriculture & Climate-smart farming Target users: Smallholder farmers and agricultural extension workers Geographic focus: Kenya, Tanzania, Uganda, Ethiopia and… See the full description on the dataset page: https://huggingface.co/datasets/vancouverevs/afriadapt-agri-train.texttext-generationn<1K1 likes20 downloads2mo agoHugging Face12vanty120 /Gpt-5.4-Xhigh-Reasoning-750x Gpt-5.4-Xhigh-Reasoning-750x A premium-quality reasoning dataset containing 721 elite samples distilled from GPT-5.4 XHIGH (the highest reasoning effort tier of GPT-5.4). Each sample features deep, multi-step Chain-of-Thought traces specifically targeting ultra-hard, expert-level problems across 60+ scientific and technical domains. This dataset is specifically designed for Supervised Fine-Tuning (SFT) to transform general-purpose language models into powerful reasoning models with… See the full description on the dataset page: https://huggingface.co/datasets/vanty120/Gpt-5.4-Xhigh-Reasoning-750x.textquestion-answeringn<1K2 likes18 downloads6mo agoHugging Face13vaniley /letters_in_word Количество букв в слове Автосгенерированный датасет чтобы научить модель считать количество букв в слове. texttext-generation1K<n<10K0 likes16 downloads1y agoHugging Face14vaniiiii /vani_dataset Reference: "A Question-Entailment Approach to Question Answering". tabularquestion-answering10K<n<100K0 likes11 downloads2y agoHugging Face15vanloc1808 /buddhist-scholar-test-set Vietnamese Buddhist Scholar Test Set Dataset Description This dataset contains 1008 Vietnamese question-answer pairs focused on Buddhist teachings and literature. The dataset was created to evaluate chatbots' knowledge and understanding of Buddhist concepts, particularly for Vietnamese-speaking users. Dataset Details Dataset Summary Language: Vietnamese Task: Question Answering, Chatbot Evaluation Domain: Buddhism, Religious Studies Size: 1008… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/buddhist-scholar-test-set.textquestion-answeringn<1K0 likes10 downloads1y agoHugging Face16vaniley /grok_answer_mail_ru Датасет ответов на Маил.ру В этом датасете собраны ответы от Grok-3-latest (и немного chatgpt-4o-latest) на вопросы с Ответы Маил.ру texttext-generation1K<n<10K2 likes9 downloads1y agoHugging Face17vanande /jorjtextquestion-answeringn<1K0 likes8 downloads4y agoHugging Face18Svngoku /mande-ancient-treasures-de-grunne-van-dyke-2016 mande-ancient-treasures-de-grunne-van-dyke-2016 Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline. Dataset Summary Metric Value Total chunks 268 Avg chars/chunk 722 Avg images/chunk 0.14 Source files 1 Duplicates removed 0 Quality filtered 6 Schema Column Type Description chunk_id string Unique identifier: filename_chunk_N text string Raw markdown chunk with image refs text_clean… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/mande-ancient-treasures-de-grunne-van-dyke-2016.imagetext-generationn<1K0 likes8 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.