datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swim-ir-monolingual
Dataset Card for SWIM-IR (Monolingual)
This is the monolingual subset of the SWIM-IR dataset, where the query generated and the passage are both in the same language.
A few remaining languages will be added in the upcoming v2 version of SWIM-IR. The dataset is available as CC-BY-SA 4.0.
For full details of the dataset, please read our upcoming NAACL 2024 paper and check out our website.
What is SWIM-IR?
SWIM-IR dataset is a synthetic multilingual retrieval dataset… See the full description on the dataset page: https://huggingface.co/datasets/nthakur/swim-ir-monolingual.HADES
Overview
The HADES benchmark is derived from the paper "Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models" (ECCV 2024 Oral). You can use the benchmark to evaluate the harmlessness of MLLMs.
Benchmark Details
HADES includes 750 harmful instructions across 5 scenarios, each paired with 6 harmful images generated via diffusion models. These images have undergone multiple optimization rounds, covering… See the full description on the dataset page: https://huggingface.co/datasets/Monosail/HADES.hwtcm-deepseek-r1-distill-data
简介
DeepSeek蒸馏的传统中医数据集,原始数据来源于网络,未进行人工审查。
7B模型微调效果
模型表现出了推理能力,准确性有待继续验证。
我们的其他产品
中医NER:能识别方剂、本草、来源、病名、症状、证型,也许是基于BERT开源模型中识别最好的模型。中医考试题:也许是全网最早开源、数据最多的中医考试题,我们内部将其用于模型训练的性能评测数据集。中医SFT数据集:中医QA数据集,用于SFT微调。仓公:基于Qwen的指令微调模型(暂未开源)。仓公R1:基于DeepSeek蒸馏的超过100万条QA的指令微调模型,拥有强大的推理能力(暂未开源)。
。。。还有很多
Citation
If you find this project useful in your research, please consider cite:
@misc{hwtcm2024,
title={{hwtcm-deepseek-r1-distill-data} A traditional… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm-deepseek-r1-distill-data.hwtcm
Description
This dataset can be used to evaluate the capabilities of large language models in traditional Chinese medicine and contains multiple-choice, multiple-answer, and true/false questions.
Changelog
2024-08-28: Added 7226 questions.
2024-08-09: The benchmark code is available at https://github.com/huangxinping/HWTCMBench.
2024-08-02: System prompts are removed to ensure the purity of the evaluation results.
2024-07-20: Debut.
Examples
multiple-answers… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm.virginia-woolf-monologue-chunks
Virginia Woolf Monologue Chunks Dataset
This dataset contains 6 semantically chunked text segments derived from a contemporary monologue based on Virginia Woolf's seminal essay "A Room of One's Own" (1929). It comes pre-loaded with vector embeddings from three different models, making it a ready-to-use resource for a variety of NLP tasks.
In addition to the dataset itself, this repository includes a comprehensive embedding analysis, detailed statistics, and 7 visualizations to help… See the full description on the dataset page: https://huggingface.co/datasets/pageman/virginia-woolf-monologue-chunks.hwtcm-sft-v1
A dataset of Tradictional Chinese Medicine (TCM) for SFT
一个用于微调LLM的传统中医数据集
Introduction
This repository contains a dataset of Traditional Chinese Medicine (TCM) for fine-tuning large language models.
Dataset Description
The dataset contains 7,096 Chinese sentences related to TCM. The sentences are collected from various sources on the Internet, including medical websites, TCM forums, and TCM books. The dataset is generated or judged by various LLMs, including… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm-sft-v1.
