datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conll2003-cisquad-ciall-Meta-Llama-3.1-70B-Instruct-AWQ-INT4all-Qwen2.5-72B-Instruct-AWQall-Llama-3.1-8B-Instructall-gpt-4obenchhub_plus_results_evaluated
BenchHub Plus Results (Evaluated)
LLM inference results on the BenchHub Plus benchmark, with per-sample accuracy scores.
Folder Structure
├── vllm_inference_results_en/ # English benchmark results (19 models)
│ ├── {model_name}_{date}.jsonl
│ └── ...
└── vllm_inference_results_ko/ # Korean benchmark results (16 models)
├── {model_name}_{date}.jsonl
└── ...
Column Description
Each .jsonl file contains one JSON object per line with the… See the full description on the dataset page: https://huggingface.co/datasets/EunsuKim/benchhub_plus_results_evaluated.all-gpt-4.1all-Codestral-22B-v0.1sud_resh_evaluated_llms_answers
📊 Результаты Оценки Больших Языковых Моделей на Бенчмарке Судебных Решений
В данном документе представлен анализ производительности 15 больших языковых моделей (LLM), протестированных на специализированном бенчмарке, который включает 105 000 записей из судебных решений России. Оценка проводилась по 10 различным категориям права (например, трудовое, уголовное, гражданское) и 7 типам инструкций (например, изложение исковых требований, анализ доказательств, итоговое решение).
Ответы… See the full description on the dataset page: https://huggingface.co/datasets/lawful-good-project/sud_resh_evaluated_llms_answers.turkish-medical-vqa-evaluatedall-do-Mistral-Small-24B-Instruct-2501all-do-Qwen2.5-72B-Instruct-AWQv4_gpt5mini_persuasion_evaluated
Nuclear Energy News Article Refinement Dataset
Overview
This dataset contains 3,182 refined news articles about nuclear energy, processed through multi-round dialogue between agents. Each article has been refined through iterative feedback to improve persuasiveness.
Dataset Structure
Source Data
Original Articles: 1,591 articles from 4 Central European countries
Languages: Czech (cs), Hungarian (hu), Polish (pl), Slovak (sk)
Processing: Multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/eoplumbum/v4_gpt5mini_persuasion_evaluated.all-gpt-4.1-miniall-gemini-2.0-flashevaluateall-Mistral-Small-24B-Instruct-2501evaluateall-gemma-3-12b-itevaluate_sn58test_evaluate_data_scitest_evaluate_footballFinetune_Evaluate_Answer
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/OG-Tiro/Finetune_Evaluate_Answer.
