datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MINT-Safeinsurance-charge-mlops-logsclassical-chinese-poetry-benchmark-70
English Readme see below
(README由Claude 3.5 Sonnet生成)
中国古诗词大模型评测基准
简介
这是一个专门用于评测大语言模型在中国古诗词理解和生成方面能力的基准测试集。该基准包含了一个多样化的测试数据集和完整的评测框架,可用于系统性地评估和比较不同模型在古诗词领域的表现。
数据集说明
数据集(poetry_benchmark.jsonl)包含70个测试样本,涵盖以下维度:
题型分布:
对联补全
诗句填空
诗词识别
提示词补全
首尾互补
难度等级:
简单(easy)
中等(medium)
困难(hard)
朝代覆盖:
先秦至近现代
包括唐、宋、元、明、清等重要朝代
评测维度
评测框架从以下维度对模型进行全面评估:
整体准确率
不同题型的表现
不同难度等级的表现
不同朝代诗词的掌握程度
评测结果
模型
blank_filling
couplet
find_poetry… See the full description on the dataset page: https://huggingface.co/datasets/happyme531/classical-chinese-poetry-benchmark-70.sft-20251126_valid_infer
happy8825/sft-20251126 · happy8825/activitynet_validset results
Model: happy8825/sft-20251126
Dataset: happy8825/activitynet_validset
Generated: 2025-12-02 15:49:18Z
Metrics
Metric
Value
Total samples
4917
With GT
4917
Parsed answers
4870
Top-1 accuracy
0.410209
Recall@5
0.410209
MRR
1
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.
1223_300
/hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-300/ · happy8825/valid_ecva_clean results
Model: /hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-300/
Dataset: happy8825/valid_ecva_clean
Generated: 2025-12-23 12:04:15Z
Metrics
Metric
Value
Total samples
924
With GT
0
Parsed answers
0
Top-1 accuracy
0
Recall@5
0
MRR
0
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/1223_300.Happy
News
[2025/02/18]: We add the original captions of PubMedVision in PubMedVision_Original_Caption.json, as well as the Chinese version of PubMedVision in PubMedVision_Chinese.json.
[2024/07/01]: We add annotations for 'body_part' and 'modality' of images, utilizing the HuatuoGPT-Vision-7B model.
PubMedVision
PubMedVision is a large-scale medical VQA dataset. We extracted high-quality image-text pairs from PubMed and used GPT-4V to reformat them to enhance their quality.… See the full description on the dataset page: https://huggingface.co/datasets/Bulamalaminu001/Happy.Atma3-ShareGPT
Dataset Card for Atma3-ShareGPT
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/Atma3-ShareGPT.LONGCOT-Alpaca
Dataset Card for LONGCOT-Alpaca
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/LONGCOT-Alpaca.train_ecva_cleanecva_instruct_ver2_fulltuned
/hub_data4/seohyun/saves/ecva_instruct/full/sft · happy8825/valid_ecva_clean results
Model: /hub_data4/seohyun/saves/ecva_instruct/full/sft
Dataset: happy8825/valid_ecva_clean
Generated: 2025-12-19 09:18:31Z
Metrics
Metric
Value
Total samples
924
With GT
0
Parsed answers
0
Top-1 accuracy
0
Recall@5
0
MRR
0
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.
EVQA/ECVA Metrics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/ecva_instruct_ver2_fulltuned.qwenutil1126infer
happy8825/qwenutil1126infer
Uploaded on 2025-12-01 11:15:12Z
ecva_tuned_with_tag
/hub_data2/seohyun/saves/ecva_with_tag/full/sft · happy8825/valid_ecva_clean results
Model: /hub_data2/seohyun/saves/ecva_with_tag/full/sft
Dataset: happy8825/valid_ecva_clean
Generated: 2025-12-16 01:08:34Z
Metrics
Metric
Value
Total samples
924
With GT
0
Parsed answers
0
Top-1 accuracy
0
Recall@5
0
MRR
0
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.
EVQA/ECVA Metrics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/ecva_tuned_with_tag.ecva_zeroshot_thinking
Qwen/Qwen3-VL-2B-Thinking · happy8825/valid_ecva_clean results
Model: Qwen/Qwen3-VL-2B-Thinking
Dataset: happy8825/valid_ecva_clean
Generated: 2025-12-16 01:28:46Z
Metrics
Metric
Value
Total samples
924
With GT
0
Parsed answers
0
Top-1 accuracy
0
Recall@5
0
MRR
0
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.
EVQA/ECVA Metrics
Metric
Value
EVQA total
924
EVQA with GT… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/ecva_zeroshot_thinking.ehristoforu__HappyLlama1-details
Dataset Card for Evaluation run of ehristoforu/HappyLlama1
Dataset automatically created during the evaluation run of model ehristoforu/HappyLlama1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__HappyLlama1-details.MATH_BS_BCE_valid_log_jsonAtma3.2-ShareGPT
Dataset Card for Atma3.2-ShareGPT
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/Atma3.2-ShareGPT.MMLU-Alpaca
Dataset Card for MMLU-Alpaca
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage
This… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/MMLU-Alpaca.AR_train_jsonquantthink-suite
QuantThink Eval Suite
The frozen, fixed-index evaluation subsets used by QuantThink to measure how quantization affects small reasoning (long chain-of-thought) models. Shipping these subsets here means a benchmark run never depends on the upstream datasets (which drift and occasionally get contaminated) being reachable or unchanged.
Files
File
Source
Size
Seed
data/gsm8k_e1.jsonl
openai/gsm8k (main, test split)
200 problems
42
data/math500_e2.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/happynood/quantthink-suite.Atma3-Share-GPT
Dataset Card for Atma3-Share-GPT
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/Atma3-Share-GPT.PoliceDristi
PoliceDrishti
Multi-label Automatic Charge Identification from the Police Point of View
A paired re-annotation of 190 Indian court case summaries under 38 Bharatiya Nyaya Sanhita (BNS) 2023 charge categories, labelled from an investigation-stage police perspective rather than a judicial one.
Dataset Description
PoliceDrishti pairs each of the 190 factual summaries from the Original ACI dataset with a fresh set of BNS 2023 charge labels reflecting what a police… See the full description on the dataset page: https://huggingface.co/datasets/happyman11/PoliceDristi.Atcgpt-Fixed2
Dataset Card for Atcgpt-Fixed2
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/Atcgpt-Fixed2.DreadPoor__Happy_New_Year-8B-Model_Stock-details
Dataset Card for Evaluation run of DreadPoor/Happy_New_Year-8B-Model_Stock
Dataset automatically created during the evaluation run of model DreadPoor/Happy_New_Year-8B-Model_Stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Happy_New_Year-8B-Model_Stock-details.Atma8
Dataset Card for Atma8
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage
This… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/Atma8.test_mllm_video_demoqwen3vl2b_valid_infer
Qwen/Qwen3-VL-2B-Thinking · happy8825/activitynet_validset results
Model: Qwen/Qwen3-VL-2B-Thinking
Dataset: happy8825/activitynet_validset
Generated: 2025-12-01 00:16:21Z
Metrics
Metric
Value
Total samples
4917
With GT
4917
Parsed answers
4807
Top-1 accuracy
0.313809
Recall@5
0.313809
MRR
1
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.
Results generated via vLLM API.
Atma3.1-ShareGPT
Atma3 Dataset
Overview
The Atma3 Dataset is a high-quality dataset comprising over 7,700 lines of detailed text from the Atma Siddhi Shastra, a seminal Jain spiritual text. This dataset is crafted to support in-depth exploration and analysis of Jain philosophy and spirituality, providing insights into the teachings of self-realization and liberation as expressed in Atma Siddhi Shastra by the revered philosopher Shrimad Rajchandra.
Dataset Details
Total… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/Atma3.1-ShareGPT.activitynet_validsetecva_instruct_ver2
/hub_data4/seohyun/saves/ecva_instruct/full/sft/checkpoint-175 · happy8825/valid_ecva_clean results
Model: /hub_data4/seohyun/saves/ecva_instruct/full/sft/checkpoint-175
Dataset: happy8825/valid_ecva_clean
Generated: 2025-12-18 10:50:16Z
Metrics
Metric
Value
Total samples
924
With GT
0
Parsed answers
0
Top-1 accuracy
0
Recall@5
0
MRR
0
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/ecva_instruct_ver2.ATC-ShareGPT
Dataset Card for ATC-ShareGPT
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage
This… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/ATC-ShareGPT.
