CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SINAI /ALIA-es-biomedical-synthetic-instructions Dataset Introduction The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision. It contains: 639,456 instances 961,073,205 tokens 14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.texttext-generation100K<n<1M0 likes93 downloads4mo agoHugging Face02SINAI /ALIA-es-legal-administrative-synthetic-instructions Dataset Introduction The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision. It contains: 763,804 instances 534,112,398 tokens 16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.texttext-generation100K<n<1M1 likes84 downloads3mo agoHugging Face03SINAI /ALIA-es-cultural-heritage-synthetic-instructions Dataset Introduction The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision. It contains: 748,480 instances 629,682,398 tokens 25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.texttext-generation100K<n<1M0 likes74 downloads4mo agoHugging Face04AiLLMBS /bio-devops-synthetic-instructions Bio-DevOps Synthetic Instructions This dataset contains synthetic instruction-following examples for biomedical-style data-engineering and scientific-computing workflows. It was created for educational and portfolio use as part of a LoRA/QLoRA fine-tuning project using Qwen/Qwen2.5-Coder-7B-Instruct. Related model: AiLLMBS/qwen25-coder-bio-devops-lora Dataset Contents The dataset includes synthetic examples for: Python CSV validation pandas duplicate checks bash… See the full description on the dataset page: https://huggingface.co/datasets/AiLLMBS/bio-devops-synthetic-instructions.texttext-generationn<1K0 likes35 downloads3mo agoHugging Face05mindchain /synthetic-instructions Synthetic Instruction Dataset Generated using Qwen/Qwen2.5-3B-Instruct. Samples: 25 Topics: web development, python programming, databases, deep learning, data science, natural language processing, computer vision texttext-generationn<1K0 likes14 downloads7mo agoHugging Face06mindchain /synthetic-instructions-judged Synthetic Instructions (LLM Judged) Generated with Distilabel-style pipeline, curated with LLM Judge. Dataset Info Total Generated: 25 Passed Judge (>= 3/5): 25 Generator: Qwen/Qwen2.5-3B-Instruct Judge: Qwen/Qwen2.5-3B-Instruct Quality Distribution Score 4: 10 samples Score 3: 15 samples Fields instruction: The question/task input: Additional context (empty) output: The response topic: Subject area quality_score: LLM Judge rating (1-5) texttext-generationn<1K0 likes14 downloads7mo agoHugging Face07kinit /synthetic-queries-and-ml-instructionsgated Synthetic Dataset: Queries and ML Instructions Dataset Description Dataset Summary This is a synthetic dataset with queries in Slovak language and ML instructions. The dataset was designed to train a model for extracting structured machine learning task requirements from natural language user queries. The dataset contains user queries in Slovak describing ML tasks paired with structured JSON outputs containing task attributes like dataset modality, task type… See the full description on the dataset page: https://huggingface.co/datasets/kinit/synthetic-queries-and-ml-instructions.tabulartext-generation10K<n<100K1 likes6 downloads11mo agoHugging Face08kinit /synthetic-conversations-and-ml-instructionsgated Synthetic Dataset: Conversations and ML Instructions Dataset Description Dataset Summary This is a synthetic dataset of 5,000 Slovak multi-turn ML advisory conversations paired with structured JSON outputs. The dataset was designed to train and evaluate models that extract machine learning task requirements from realistic, natural Slovak dialogue, including cases where requirements are revealed gradually or changed during the conversation. Each… See the full description on the dataset page: https://huggingface.co/datasets/kinit/synthetic-conversations-and-ml-instructions.tabulartext-generation1K<n<10K1 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.