datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Reverse-circuit-discoveryOriginal-circuit-discoveryred-pill-drug-discovery-formulation
🔴 RED-PILL
Research Enhanced Dataset for Pharmaceutical Innovation in Learning & Language
The first open instruction-tuning dataset for drug discovery & formulation development.
Built for fine-tuning Heretic-ablated models that won't refuse your pharmaceutical R&D questions.
⚡ Quick Start
from datasets import load_dataset
# Load the full dataset
ds = load_dataset("saidutta69/red-pill-drug-discovery-formulation"… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/red-pill-drug-discovery-formulation.printable-coloring-product-discovery
PrintableFunnyPages — Printable Coloring Product Discovery & Buyer Intent Corpus
Dataset Summary
This dataset is an English-language product discovery and buyer-intent corpus created for PrintableFunnyPages, a digital printable shop offering downloadable coloring and activity resources.
It is designed to support research and experimentation in:
product discovery
semantic product retrieval
buyer-intent understanding
ecommerce search
recommendation and matching… See the full description on the dataset page: https://huggingface.co/datasets/PrintableFunnyPages/printable-coloring-product-discovery.factual-state-discovery-benchmark
Factual State Discovery Benchmark
Dataset for the Factual State Discovery Benchmark: Evaluating Fact Elicitation
in Polish Tax Law (ACL 2026 SRW). It evaluates whether conversational agents
can systematically elicit, through dialogue, all the facts of a taxpayer's
situation from a real Polish tax interpretation document.
Each sample pairs a factual state (a narrative of the taxpayer's situation,
in Polish) with its decomposition into atomic facts — independent,
verifiable claims… See the full description on the dataset page: https://huggingface.co/datasets/AI-TAX/factual-state-discovery-benchmark.cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Scientific_Discovery_and_Research
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Scientific_Discovery_and_Research
データ件数: 3,733
平均トークン数: 1,193
最大トークン数: 2,489
合計トークン数: 4,453,517
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 23.2 MB
加工内容:
メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。
難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research.API_Discovery_Retrieval_Augmented_Calling
🇰🇿 Kazakh API Discovery and Tool Retrieval Dataset
Dataset Summary
Kazakh API Discovery and Tool Retrieval Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in agentic AI workflows that require API discovery, tool documentation retrieval, function calling, and multi-step tool execution.
The dataset focuses on scenarios where the assistant must first inspect or retrieve API documentation before calling the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/API_Discovery_Retrieval_Augmented_Calling.
