datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LORE-examples
LORE Examples
A small set of matched multimodal examples from LORE, for the
MIMIC model — enough to try inference,
embedding, and generation across DNA, RNA, and protein modalities without wiring
up your own data.
Each example is a single biological entity (a transcript and/or its protein) with
several co-observed modalities. Rows are drawn from the held-out (validation) split
of MIMIC's training data, so they are in-distribution and length-bounded to the
model's context… See the full description on the dataset page: https://huggingface.co/datasets/polymathic-ai/LORE-examples.Polymath-Instruct
Polymath-Instruct
Dataset Summary
Polymath-Instruct is a premium synthetic dataset designed to elevate the reasoning capabilities of Large Language Models (LLMs). Moving beyond simple instruction following, this dataset focuses on deep reasoning, Chain-of-Thought (CoT), and, crucially, interdisciplinary synthesis.
The dataset contains complex scenarios where an expert persona (defined via system prompts) solves high-level problems. A unique feature of Polymath-Instruct is… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/Polymath-Instruct.PolyMath_RU
