CoolFace
7 results

synthetic-instructions

SINAI /ALIA-es-biomedical-synthetic-instructions Dataset Introduction The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision. It contains: 639,456 instances 961,073,205 tokens 14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.texttext-generation100K<n<1M0 likes93 downloads4mo agoHugging FaceSINAI /ALIA-es-legal-administrative-synthetic-instructions Dataset Introduction The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision. It contains: 763,804 instances 534,112,398 tokens 16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.texttext-generation100K<n<1M1 likes84 downloads3mo agoHugging FaceSINAI /ALIA-es-cultural-heritage-synthetic-instructions Dataset Introduction The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision. It contains: 748,480 instances 629,682,398 tokens 25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.texttext-generation100K<n<1M0 likes74 downloads4mo agoHugging FaceSnaseem2026 /synthetic-multilingual-instructions Synthetic Multilingual Instruction Dataset This dataset contains millions of synthetic, copyright-free instruction–response pairs covering practical, everyday scenarios. Each record includes: instruction: The user prompt or question response: The synthetic answer language: Language code (e.g., 'en', 'fr', 'de', 'es', 'it', 'ar') topic: General topic (e.g., 'home repair', 'finance') complexity: One of 'basic', 'intermediate', 'advanced' Available Files… See the full description on the dataset page: https://huggingface.co/datasets/Snaseem2026/synthetic-multilingual-instructions.texttext-classification1M<n<10M0 likes51 downloads8mo agoHugging FaceAiLLMBS /bio-devops-synthetic-instructions Bio-DevOps Synthetic Instructions This dataset contains synthetic instruction-following examples for biomedical-style data-engineering and scientific-computing workflows. It was created for educational and portfolio use as part of a LoRA/QLoRA fine-tuning project using Qwen/Qwen2.5-Coder-7B-Instruct. Related model: AiLLMBS/qwen25-coder-bio-devops-lora Dataset Contents The dataset includes synthetic examples for: Python CSV validation pandas duplicate checks bash… See the full description on the dataset page: https://huggingface.co/datasets/AiLLMBS/bio-devops-synthetic-instructions.texttext-generationn<1K0 likes35 downloads3mo agoHugging FaceMenlo /Instruction-Synthetic-v0.4text10K<n<100K0 likes15 downloads1y agoHugging Face