datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
myco-spore-nutrients
MycoNet NutrientCards — Derived Dataset / 菌网养分卡·衍生数据集
EN. A derived-only dataset of MycoNet NutrientCards (Myco-Spore factory). Each
line in data/nutrients.jsonl carries the derived layer of a source: a
structured summary, a derived meaning, a decomposition proof, and provenance /
attribution. The full source text is intentionally NOT included — only
licensed-derived artifacts are redistributed, keeping the dataset safe to share
across borders.
中文. 本数据集仅包含菌网养分卡(Myco-Spore… See the full description on the dataset page: https://huggingface.co/datasets/Myco-Net/myco-spore-nutrients.DocPII-redaction-benchmark
DocPII: Contextual Redaction Benchmark Dataset
Dataset Description
DocPII contains 1101 high-quality document samples enriched with embedded personally identifiable information (PII). Designed to evaluate context-aware redaction systems, it provides realistic, full-document contexts—a notable advancement over sentence-level datasets.
All documents have been manually reviewed for accuracy, coherence, and redaction alignment, ensuring data quality for benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/DocPII-redaction-benchmark.
