CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01camel-ai /biology CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society Github: https://github.com/lightaime/camel Website: https://www.camel-ai.org/ Arxiv Paper: https://arxiv.org/abs/2303.17760 Dataset Summary Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs. We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.texttext-generation10K<n<100K58 likes8.9k downloads3y agoHugging Face02bio-nlp-umass /MedThinkVQA MedThinkVQA MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning. Links GitHub: https://github.com/benluwang/MedThinkVQA Leaderboard: https://benluwang.github.io/MedThinkVQA/ Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.imagequestion-answering1K<n<10K11 likes4k downloads4mo agoHugging Face03common-pile /biodiversity_heritage_library_filtered Biodiversity Heritage Library Description The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives. This dataset contains over 15 million public domain books and documents from the BHL collection. These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata. We use the optical character recognition (OCR)-generated text distributed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library_filtered.texttext-generation10M<n<100M2 likes861 downloads1y agoHugging Face04QizhiPei /BioMatrix-SFT BioMatrix-SFT This is the supervised fine-tuning (SFT) / instruction-tuning corpus used to train BioMatrix, a multimodal foundation model that natively integrates 1D sequences, 3D structures, and natural language for both molecules and proteins within a single decoder-only architecture. 📄 Paper: BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language 💻 Code: https://github.com/QizhiPei/BioMatrix… See the full description on the dataset page: https://huggingface.co/datasets/QizhiPei/BioMatrix-SFT.texttext-generation10M<n<100M1 likes826 downloads3mo agoHugging Face05common-pile /biodiversity_heritage_library Biodiversity Heritage Library Description The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives. This dataset contains over 42 million public domain books and documents from the BHL collection. These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata. We use the optical character recognition (OCR)-generated text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library.texttext-generation10M<n<100M2 likes668 downloads1y agoHugging Face06IVN-RIN /BioBERT_ItalianFrom this repository you can download the BioBERT_Italian dataset. BioBERT_Italian is the Italian translation of the original BioBERT dataset, composed by millions of abstracts of PubMed papers. Due to the unavailability of an Italian equivalent for the millions of abstracts and full-text scientific papers used by English, BERT-based biomedical models, we leveraged machine translation to obtain an Italian biomedical corpus based on PubMed abstracts and train BioBIT. Corpus statistics: Total… See the full description on the dataset page: https://huggingface.co/datasets/IVN-RIN/BioBERT_Italian.texttext-generation10M<n<100M7 likes370 downloads2y agoHugging Face07jknafou /TransCorpus-bio TransCorpus-bio TransCorpus-bio is a large-scale, parallel biomedical corpus consisting of PubMed abstracts (title + abstract), translated with the TransCorpus Toolkit using NLLB-200. It is designed to enable high-quality multi-lingual biomedical language modeling and downstream NLP research. This dataset was restructured from five separate single-language repositories into one dataset with a config (tab in the dataset viewer) per language, and with each row carrying its source… See the full description on the dataset page: https://huggingface.co/datasets/jknafou/TransCorpus-bio.texttranslation100M<n<1B0 likes300 downloads1d agoHugging Face08alex-karev /biographies 📚 Synthetic Biographies Synthetic Biographies is a dataset designed to facilitate research in factual recall and representation learning in language models. It comprises synthetic biographies of fictional individuals, each associated with sampled attributes like birthplace, university, and employer. The dataset is intended to support training and evaluating small language models (LLMs), particularly in their ability to store and extract factual knowledge. 🧾 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/alex-karev/biographies.texttext-generation100K<n<1M4 likes204 downloads1y agoHugging Face09ruslan /bioleaflets-biomedical-ner Dataset Card for BioLeaflets Dataset Dataset Summary BioLeaflets is a biomedical dataset for Data2Text generation. It is a corpus of 1,336 package leaflets of medicines authorised in Europe, which were obtained by scraping the European Medicines Agency (EMA) website. Package leaflets are included in the packaging of medicinal products and contain information to help patients use the product safely and appropriately. This dataset comprises the large majority (∼ 90%) of… See the full description on the dataset page: https://huggingface.co/datasets/ruslan/bioleaflets-biomedical-ner.texttext-generation1K<n<10K4 likes202 downloads4y agoHugging Face10bio-nlp-umass /bioinstruct Dataset Card for BioInstruct GitHub repo: https://github.com/bio-nlp/BioInstruct Dataset Summary BioInstruct is a dataset of 25k instructions and demonstrations generated by OpenAI's GPT-4 engine in July 2023. This instruction data can be used to conduct instruction-tuning for language models (e.g. Llama) and make the language model follow biomedical instruction better. Improvements of Llama on 9 common BioMedical tasks are shown in the result section. Taking… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/bioinstruct.texttext-generation10K<n<100K25 likes132 downloads2y agoHugging Face11Lots-of-LoRAs /task686_mmmlu_answer_generation_college_biology Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task686_mmmlu_answer_generation_college_biology Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task686_mmmlu_answer_generation_college_biology.texttext-generationn<1K0 likes122 downloads2y agoHugging Face12rntc /biomed-fr-v3-enriched-softmin-standard biomed-fr-v3-enriched-softmin-standard This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling. Preprocessing Method Soft-min calculation: Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores Parameter p = -2.0 Weight computation: Ratio preference (5 vs 1): R = 10 Gamma exponent: γ = 1.43 (computed as log(R)/log(5)) Weight formula: w = s^γ Floor: w = max(w, median(w) × 0.05) Resampling: Target size:… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-standard.tabulartext-generation1M<n<10M0 likes98 downloads1y agoHugging Face13CleverThis /dbpedia-biomedical DBpedia Categories Dataset Description Category relationships from DBpedia (English) Original Source: https://downloads.dbpedia.org/repo/dbpedia/generic/categories/2022.12.01/categories_lang=en_articles.ttl.bz2 Dataset Summary This dataset contains RDF triples from DBpedia Categories converted to HuggingFace dataset format for easy use in machine learning pipelines. Format: Originally turtle, converted to HuggingFace Dataset Size: 3.0 GB (extracted) Entities:… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/dbpedia-biomedical.texttext-generation10M<n<100M0 likes97 downloads11mo agoHugging Face14cannin /biostars_qa Dataset Summary This dataset contains 4803 question/answer pairs extracted from the BioStars website. The site focuses on bioinformatics, computational genomics, and biological data analysis. Dataset Structure Data Fields The data contains INSTRUCTION, RESPONSE, SOURCE, and METADATA fields. The format is described for LAION-AI/Open-Assistant Dataset Creation Curation Rationale Questions were included if they were an accepted answer and the… See the full description on the dataset page: https://huggingface.co/datasets/cannin/biostars_qa.texttext-classification1K<n<10K3 likes95 downloads3y agoHugging Face15devsgnr /bio-safety-peft-lora CBRN Safety Alignment & PEFT-LoRA Fine-Tuning Dataset This repository contains the synthetic instruction-tuning dataset (.jsonl) designed for parameter-efficient fine-tuning (PEFT-LoRA) of edge language models (specifically Qwen/Qwen2.5-1.5B-Instruct). The dataset is curated to evaluate and modify model logit distributions, persona attributions, and dual-use safety boundaries regarding Chemical, Biological, Radiological, and Nuclear (CBRN) risk scenarios. 🤖 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/devsgnr/bio-safety-peft-lora.texttext-generation1K<n<10K0 likes95 downloads10d agoHugging Face16qyxu1994 /BioPhys-Bridge BioPhys-Bridge is a physics-grounded scientific reasoning dataset for AI-for-Science agents. The dataset subtype, Sci-Evo, represents each record as a Physics-Grounded Scientific Evolution Case linking: physical model -> quantitative evidence -> biological mechanism -> agent decision Each case is built from open-access scientific literature and includes evidence-linked text, tables, formulas, figure/caption blocks, normalized quantitative measurements, biophysical model fields, biological… See the full description on the dataset page: https://huggingface.co/datasets/qyxu1994/BioPhys-Bridge.textquestion-answeringn<1K1 likes94 downloads3mo agoHugging Face170xKitkat /BiochemForge BiochemForge BiochemForge is a provenance-first biology, chemistry, and biochemistry post-training mixture for mechanistic explanation, quantitative derivation, experimental inference, and consistency between reasoning and final answers. Dataset summary Slice Records Purpose SFT train 99,773 Supervised post-training SFT validation 2,052 Model selection and early stopping SFT test 1,093 Internal held-out evaluation Solver-verified records 27,657… See the full description on the dataset page: https://huggingface.co/datasets/0xKitkat/BiochemForge.tabularquestion-answering100K<n<1M1 likes88 downloads1mo agoHugging Face18bowenxian /BioProBench BioProBench: Comprehensive Dataset and Benchmark in Biological Protocol Understanding and Reasoning BioProBench is the first large-scale, integrated multi-task benchmark for biological protocol understanding and reasoning, specifically designed for large language models (LLMs). It moves beyond simple QA to encompass a comprehensive suite of tasks critical for procedural text comprehension. Biological protocols are the fundamental bedrock of reproducible and safe life… See the full description on the dataset page: https://huggingface.co/datasets/bowenxian/BioProBench.texttext-generation1K<n<10K0 likes87 downloads8mo agoHugging Face19Lots-of-LoRAs /task699_mmmlu_answer_generation_high_school_biology Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task699_mmmlu_answer_generation_high_school_biology Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task699_mmmlu_answer_generation_high_school_biology.texttext-generationn<1K0 likes86 downloads2y agoHugging Face20SINAI /ALIA-es-biomedical-synthetic-instructions Dataset Introduction The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision. It contains: 639,456 instances 961,073,205 tokens 14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.texttext-generation100K<n<1M0 likes84 downloads4mo agoHugging Face21gxx27 /BioTool BioTool BioTool is a large-scale, function-calling benchmark and training corpus for the biomedical domain. It pairs natural-language biomedical questions with the correct tool call (function name + JSON arguments) that answers them, drawn from 127 tools spanning the three flagship public APIs: NCBI E-utilities (einfo, esearch, esummary, efetch, elink, ecitmatch) plus BLAST UniProt REST (uniprotkb, uniref, uniparc, proteomes, taxonomy, keywords, human_diseases, …) Ensembl REST… See the full description on the dataset page: https://huggingface.co/datasets/gxx27/BioTool.textquestion-answering1K<n<10K1 likes72 downloads5mo agoHugging Face22adobug /bcs-biostatistics-study BCS Medical Dataset — biostatistika Medicinski studijski materijal na bosanskom/hrvatskom/srpskom, obrađen automatizovanim inbox pipeline-om (ekstrakcija, OCR, chunking, AI generacija s determinističkom validacijom). Struktura Fajl Sadržaj ispitna.jsonl postojeća ispitna pitanja (stari testovi/zbornici): question, options, answer qna.jsonl AI-generirani QnA parovi (validacija V1-V4) flashcards.jsonl / flashcards.csv kartice front/back za učenje… See the full description on the dataset page: https://huggingface.co/datasets/adobug/bcs-biostatistics-study.textquestion-answeringn<1K0 likes71 downloads13h agoHugging Face23SINAI /ALIA-es-biomedical-pairs Dataset Introduction The ALIA Spanish Biomedical Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen3-style prompting workflow integrated in the ALIA encoders pipeline. It preserves provenance to the original document and chunk while exposing controls such as question type and difficulty (ranging from high_school to phd level).… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-pairs.textquestion-answering100K<n<1M0 likes65 downloads4mo agoHugging Face24SINAI /ALIA-es-biomedical Dataset Introduction The ALIA Spanish Biomedical Corpus constitutes a strategic data infrastructure designed to support research and innovation in the biomedical domain. By ensuring systematic access to multiple official medical repositories in a single consolidated dataset, it provides a robust foundation for Spanish-language BioNLP. With over 6 million instances and more than 4 billion tokens, it represents a relevant comprehensive corpus of biomedical and clinical-related… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical.texttext-generation1M<n<10M0 likes64 downloads3mo agoHugging Face25Brainquiver /reason-qa-biology-finetune-preview Reasoning · Biology · Finetuning · Preview (Synthetic) A public, single-generator preview of a larger private biology reasoning corpus. This dataset has been created with gpt-oss-20b output and uses a simplified three-field format. The full set spans many generator models, two reasoning styles (linear and branching), and a richer schema (metadata, instruction, thinking, reasoning, answer). Synthetic question-reasoning-answer data for domain finetuning on biology and biochemistry… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/reason-qa-biology-finetune-preview.texttext-generation10K<n<100K0 likes59 downloads3mo agoHugging Face26rntc /biomed-fr-v3-enriched-softmin-leger biomed-fr-v3-enriched-softmin-leger This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling. Preprocessing Method Soft-min calculation: Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores Parameter p = -0.7 Weight computation: Ratio preference (5 vs 1): R = 5 Gamma exponent: γ = 1.00 (computed as log(R)/log(5)) Weight formula: w = s^γ Floor: w = max(w, median(w) × 0.05) Resampling: Target size: Same… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-leger.tabulartext-generation1M<n<10M0 likes55 downloads1y agoHugging Face27capicu-ai /BioManufacturingBench BioManufacturingBench v1.0.0 BioManufacturingBench v1.0.0 is a 2,000-item benchmark for evidence-grounded biomanufacturing reasoning. It covers evidence extraction, mass-balance calculation, process diagnosis, microscopy count-range estimation, strict output formatting, and abstention. Every primary score is computed by a deterministic rule; no score uses an LLM judge. Public records are deliberately answer-free so the benchmark remains useful for future evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/capicu-ai/BioManufacturingBench.textquestion-answering1K<n<10K0 likes53 downloads2mo agoHugging Face28mattwesney /ToT-Biologygated The ToT-Biology dataset emphasizes mechanistic understanding and explanatory biological reasoning, rather than just providing correct answers. It aims to train AI models in interpretability and logical deduction within the biological realm. Spanning a wide range of biological complexities, it starts with foundational concepts in cell biology, genetics, and ecology, and progresses to advanced areas like systems biology, synthetic biology, and computational biophysics. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/ToT-Biology.texttext-generation10K<n<100K10 likes52 downloads2y agoHugging Face29timodonnell /bioreason-pro-sft-reasoning-documents BioReason-Pro SFT Reasoning Documents Complete, self-contained training documents reconstructed from the BioReason-Pro SFT data, ready for LLM pre-training. The upstream dataset wanglab/bioreason-pro-sft-reasoning-data ships the assistant side of each training example (reasoning, final_answer) alongside the raw biological context columns, but not the assembled prompt. The prompt cannot be recovered from the data card alone, because two of its three parts were non-textual… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/bioreason-pro-sft-reasoning-documents.tabulartext-generation100K<n<1M0 likes52 downloads2mo agoHugging Face30Despina /biographical Biographical Dataset for Relation Extraction (RE) Overview This dataset is a reconstructed version of the Biographical Dataset, specifically designed for relation extraction (RE) tasks. It serves as a valuable resource for digital humanities (DH) and historical research, enabling the study of relationships within biographical data. The dataset is generated by automatically aligning sentences from Wikipedia articles with structured data sourced from platforms like… See the full description on the dataset page: https://huggingface.co/datasets/Despina/biographical.textfeature-extraction1M<n<10M0 likes46 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.