CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alianassmaaa /ameli-assurance-maladie-qa Ameli Assurance Maladie - Question Answering Dataset Description Dataset de 100 paires question-réponse (QA) basé sur les publications officielles de l'Assurance Maladie française, extraites du site assurance-maladie.ameli.fr. Conçu pour évaluer des systèmes de RAG (Retrieval-Augmented Generation) sur des documents institutionnels français dans le domaine de la santé publique et de la protection sociale. Format du dataset { "question": "Quel article du… See the full description on the dataset page: https://huggingface.co/datasets/alianassmaaa/ameli-assurance-maladie-qa.documentquestion-answeringn<1K1 likes341 downloads5mo agoHugging Face02SINAI /ALIA-es-legal-administrative-triplets Dataset Introduction The dataset ALIA Spanish Legal and Administrative Triplets Corpus contains hard negatives for dense retrieval training generated from <query, passage> pairs contained in SINAI/ALIA-es-legal-administrative-triplets.The dataset was created as part of the ALIA project to improve the training of embedding models and dense retrievers specialized in Spanish legal and administrative language. Hard negatives are passages that are semantically similar to a query but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-triplets.texttext-generation1M<n<10M2 likes130 downloads4mo agoHugging Face03SINAI /ALIA-es-legal-administrative-cqa Dataset Introduction The ALIA Spanish Legal and Administrative for Context Question Answering Corpus is a specialized question-answering resource derived from the SINAI/ALIA-es-legal-administrative corpus. This dataset transforms legal and administrative documents into structured question-answer pairs, enabling the development and evaluation of AI systems capable of understanding and responding to queries about Spanish legal-administrative content. With 17,668 structured instances… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-cqa.textquestion-answering10K<n<100K3 likes126 downloads4mo agoHugging Face04SINAI /ALIA-es-biomedical-synthetic-instructions Dataset Introduction The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision. It contains: 639,456 instances 961,073,205 tokens 14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.texttext-generation100K<n<1M0 likes93 downloads4mo agoHugging Face05SINAI /ALIA-es-legal-administrative-synthetic-instructions Dataset Introduction The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision. It contains: 763,804 instances 534,112,398 tokens 16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.texttext-generation100K<n<1M1 likes85 downloads3mo agoHugging Face06SINAI /ALIA-es-cultural-heritage-synthetic-instructions Dataset Introduction The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision. It contains: 748,480 instances 629,682,398 tokens 25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.texttext-generation100K<n<1M0 likes77 downloads4mo agoHugging Face07SINAI /ALIA-es-cultural-heritage Dataset Introduction The ALIA Spanish Cultural and Heritage Corpus is a strategic open data infrastructure designed to support research and innovation in digital humanities, cultural analytics, and Spanish-language NLP. It consolidates heterogeneous official and academic repositories into a single curated dataset, enabling broad and structured access to cultural heritage documentation from Spain. With 236,314 instances, 939,315,404 tokens and 100 source datasets, it provides a… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage.texttext-generation100K<n<1M0 likes69 downloads3mo agoHugging Face08SINAI /ALIA-es-biomedical-pairs Dataset Introduction The ALIA Spanish Biomedical Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen3-style prompting workflow integrated in the ALIA encoders pipeline. It preserves provenance to the original document and chunk while exposing controls such as question type and difficulty (ranging from high_school to phd level).… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-pairs.textquestion-answering100K<n<1M0 likes65 downloads4mo agoHugging Face09SINAI /ALIA-es-biomedical Dataset Introduction The ALIA Spanish Biomedical Corpus constitutes a strategic data infrastructure designed to support research and innovation in the biomedical domain. By ensuring systematic access to multiple official medical repositories in a single consolidated dataset, it provides a robust foundation for Spanish-language BioNLP. With over 6 million instances and more than 4 billion tokens, it represents a relevant comprehensive corpus of biomedical and clinical-related… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical.texttext-generation1M<n<10M0 likes64 downloads3mo agoHugging Face10SINAI /ALIA-es-cultural-heritage-pairs Dataset Introduction The ALIA Spanish Cultural and Heritage Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen-style prompting workflow integrated in the ALIA encoders pipeline. It preserves provenance to the original document and passage while exposing controls such as question type and difficulty (ranging from high_school… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-pairs.textquestion-answering100K<n<1M0 likes63 downloads3mo agoHugging Face11alialp207 /TR-DataAnalystBench TR-DataAnalystBench A Turkish-language benchmark for evaluating whether language models can perform data-analyst style reasoning over tables and charts: reading a value, finding the maximum/minimum, comparing two years, computing an average or a (signed) percentage change, ranking, summarizing a trend, and — importantly — abstaining when the data does not contain the answer. Gold answers are computed and verified with Python (not produced by a language model), so the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/alialp207/TR-DataAnalystBench.imagequestion-answering1K<n<10K0 likes60 downloads3mo agoHugging Face12SINAI /ALIA-es-cultural-heritage-triplets Dataset Introduction The dataset ALIA Spanish Cultural and Heritage Hard Negatives Corpus contains hard negatives for dense retrieval training generated from <query, passage> pairs contained in SINAI/ALIA-es-cultural-heritage-pairs.The dataset was created as part of the ALIA project to improve the training of embedding models and dense retrievers specialized in Spanish cultural heritage language. Hard negatives are passages that are semantically similar to a query but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-triplets.texttext-generation1M<n<10M0 likes49 downloads4mo agoHugging Face13SINAI /ALIA-es-legal-administrative-pairs Dataset Introduction The ALIA Spanish Legal and Administrative Pairs Corpus, derived from the SINAI/ALIA-es-legal-administrative, contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen3-style prompting workflow integrated in the ALIA encoders pipeline. It preserves provenance to the original document and chunk while exposing controls such as question type… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-pairs.textquestion-answering100K<n<1M0 likes36 downloads4mo agoHugging Face14SINAI /ALIA-es-biomedical-triplets Dataset Introduction The dataset ALIA Spanish Biomedical Hard Negatives Corpus contains hard negatives for dense retrieval training generated from <query, passage> pairs contained in SINAI/ALIA-es-biomedical-pairs.The dataset was created as part of the ALIA project to improve the training of embedding models and dense retrievers specialized in Spanish biomedical language. Hard negatives are passages that are semantically similar to a query but not correct answers, making them… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-triplets.texttext-generation100K<n<1M0 likes34 downloads4mo agoHugging Face15Irfanuruchi /dsp-fft-sampling-aliasing Synthetic DSP Dataset: FFT + Sampling / Aliasing This repository contains synthetic instruction-style DSP samples designed for numerical reasoning and conceptual understanding of Digital Signal Processing (DSP) fundamentals. The dataset focuses on: FFT bin reasoning and frequency-domain interpretation Sampling theory Aliasing effects Dataset Origin & Verification This dataset was generated as part of the project: Fine-Tuning Lightweight Large Language Models for a… See the full description on the dataset page: https://huggingface.co/datasets/Irfanuruchi/dsp-fft-sampling-aliasing.texttext-generation1K<n<10K0 likes30 downloads8mo agoHugging Face16aliarda /TurkishIdentityMini TurkishIdentityMini Dataset Description TurkishIdentityMini is a small, template-based Turkish instruction dataset designed to help LLMs respond correctly to identity-related questions. It contains instruction–output pairs where a user asks a chatbot about its name, origin, or creator, and the model responds using customizable {{model_name}} and {{team_name}} placeholders. This dataset is useful for fine-tuning or instruction-tuning Turkish language models to maintain a… See the full description on the dataset page: https://huggingface.co/datasets/aliarda/TurkishIdentityMini.texttext-generationn<1K2 likes18 downloads7mo agoHugging Face17aliasocracy /enoch-ai-research-corpus Enoch AI Research Corpus This dataset contains 393 AI-generated research artifacts produced by the Enoch agentic research system. System repository: https://github.com/alias8818/enoch-agentic-research-system Source corpus repository: https://github.com/alias8818/enoch-ai-research-corpus Launch site: https://alias8818.github.io/enoch-agentic-research-system/ Current release correction Older launch posts may mention 120 artifacts. The current public corpus indexes… See the full description on the dataset page: https://huggingface.co/datasets/aliasocracy/enoch-ai-research-corpus.text-generation1K<n<10K1 likes9 downloads3mo agoHugging Face18aliarda /jev_turkish_mmlu_tracesgated JEV Turkish MMLU & MMLU-Pro Traces Traces of Jev (jev-latest, TypeSafe System One) answering Turkish multiple-choice questions from the turkish_mmlu and turkish-mmlu-pro-preview datasets. Each source row becomes a choice question; rows are grouped by subject and sent as one request per subject (the subject is the state). Every trace row records jev's chosen option, confidence, probability distribution, and (when captured) the request id, token usage, and latency.… See the full description on the dataset page: https://huggingface.co/datasets/aliarda/jev_turkish_mmlu_traces.tabularmultiple-choice10K<n<100K0 likes7 downloads17h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.