datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
china-effective-laws-regulations
全国现行法律法规合集
现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。
数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。
这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。
快照日期:2026-08-26
效力说明
本数据集 以现行有效法律法规为主体:
效力 status
法规份数
说明
有效
17,649
现行有效,默认应使用这一部分
尚未生效
7
已公布、施行日晚于快照日
失效
45
文件名含「失效」,多为已到期的全国人大常委会试点授权决定
使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.Agent-eval-Effector-Hunt
Agent Eval: Effector Hunt
While AI scientist agents like Claude Science and Google's AI co-scientist highlight the potential of autonomous research, compact and reproducible datasets for evaluating these agents on real scientific workflows remain scarce.
Agent Eval: Effector Hunt is a genomics benchmark package designed around a real scientific discovery workflow from the Science paper Chen et al. 2017. It asks an AI agent, a computational biologist, or a hybrid human-agent… See the full description on the dataset page: https://huggingface.co/datasets/chjp0632/Agent-eval-Effector-Hunt.go-effective-docs-qa
Effective Go Instruction Dataset
Overview
This dataset was created from the official Effective Go documentation. The content was extracted from:
https://go.dev/doc/effective_go
and transformed into instruction-following samples consisting of:
instruction
input
output
Dataset Structure
Split
Examples
Train
2,643
Validation
293
Features
instruction (string)
input (string)
output (string)
How this… See the full description on the dataset page: https://huggingface.co/datasets/farid678/go-effective-docs-qa.supplement-effects-and-interactions
SuppLLMent: A Benchmark for Evidence-Based Supplement Knowledge in LLMs
Dataset Description
SuppLLMent is a structured dataset for evaluating large language models on supplement effectiveness knowledge. It contains 8,417 supplement-condition effectiveness facts and 5,884 drug-supplement interaction warnings extracted from publicly available consumer health resources.
Dataset Summary
This dataset enables systematic benchmarking of how well LLMs encode… See the full description on the dataset page: https://huggingface.co/datasets/daveferbear/supplement-effects-and-interactions.
