datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
science_biologyarxiv-biology
Dataset Curators
The original data is maintained by ArXiv
Licensing Information
The data is under the Creative Commons CC0 1.0 Universal Public Domain Dedication
Citation Information
@misc{clement2019arxiv,
title={On the Use of ArXiv as a Dataset},
author={Colin B. Clement and Matthew Bierbaum and Kevin P. O'Keeffe and Alexander A. Alemi},
year={2019},
eprint={1905.00075},
archivePrefix={arXiv},
primaryClass={cs.IR}
}
NCERT_Biology_11thNCERT_Biology_12thclinical-placebo-meaning-biology-interaction-topology-classification-v0.1What this dataset tests
Whether a model can classify how meaning and biology interactin driving an observed outcome.
Required outputs
interaction_topology
dominance_order
failure_mode
Interaction topology labels
additive
amplifying
suppressive
overriding
antagonistic
Dominance order
biology_over_meaning
meaning_over_biology
balanced
Failure modes
placebo_masking
nocebo_suppression
meaning_override_false_positive
biology_misattribution
high_uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-placebo-meaning-biology-interaction-topology-classification-v0.1.speculative_biology_gate_v01Speculative Biology Gate (SBG) v0.1
Goal
Check if a model can hold the line between:
• known biology• active research• fiction or invention
This gate tests when to:
• refuse• request scope• keep answers conceptual• state uncertainty
Columns
prompt | risky query
domain_status | type of unknown
required_stance | correct orientation for reply
action | refuse / request_context / answer_with_boundary
Use Cases
• safety audits• refusal pattern training•… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/speculative_biology_gate_v01.biologybiologybiologyhelper
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/dddjjjppp/biologyhelper.NCERT_Biology_11thbiology-qaSynthetic_Biology_Gene_Expression_DataLlama_3_70b_biologyLlama_3_1_8b_Instruct_Turbo_chat_biologybiology_sentence_pairsBiology_term_definition_pairsynthetic-biology-gene-circuit-optimization-v2Mixtral_8_7B_biologyMixtral_8_22B_biologyits-biologyPhi_3_mini_128k_biology
Phi_3_mini_128k_biology
license: mit
biology_default_guacomolesynthetic_biology_genetic_circuit_response
