llm-interaction
protein_interactions_LLM_FT_datasetThis dataset is derived from the ANDDigest database and contains PubMed abstracts with dictionary-mapped protein names. A total of >15,000 abstracts were selected, yielding 6,516 unique protein pairs with associative edges in the ANDSystem network. The corpus was split into positive and negative samples. Positive samples represent documents where a protein interaction was identified using the rules-based approach of ANDSystem’s text-mining module, while negative samples include documents where… See the full description on the dataset page: https://huggingface.co/datasets/Timofey/protein_interactions_LLM_FT_dataset.llm_interaction_corpus
[llm_interaction_corpus]
数据集概述
本仓库收录了日常工作中与大语言模型(LLM)的真实交互数据,涉及【文本生成/代码辅助/数据分析/…】等场景。数据严格经过会话重构、敏感信息脱敏与质量筛查,并附带细粒度元数据标签(会话意图、满意度评分、修订程度等)。
数据构成
prompts/:原始用户提示,保留真实任务语境
responses/:模型原始输出(含模型名称、参数快照)
revisions/:人工修订记录与diff
feedback/:评分、排序、二元偏好标签
metadata.csv:每段交互的上下文特征与时间戳
应用价值
**指令微调 (Instruction Tuning)**:高质量多样化任务指令
**人类反馈强化学习 (RLHF)**:偏好对与修订轨迹
工具使用与Agent行为分析:API调用链和推理路径
数据形式
所有数据以 JSON Lines / Parquet… See the full description on the dataset page: https://huggingface.co/datasets/chensongpoixs/llm_interaction_corpus.LLM-InteractionThis is a temporary dataset used for initial testing.
llm_persona_interactions_othersteve003_csvllm_persona_interactions_othersteve004_csv
