datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-Fixer-Train-110K
SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution
📃 Paper |
🚀 GitHub
SWE-Fixer is a simple yet effective solution for addressing real-world GitHub issues by training open-source LLMs. It features a streamlined retrieve-then-edit pipeline with two core components: a code file retriever and a code editor.
This repo holds the data SWE-Fixer-Train-110K we curated for SWE-Fixer training.
For more information, please visit our project page.… See the full description on the dataset page: https://huggingface.co/datasets/internlm/SWE-Fixer-Train-110K.SWE-Fixer-Train-Editing-CoT-70KInteractScience
InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
InteractScience is a benchmark specifically designed to evaluate the capability of large language models in generating interactive scientific demonstration code. This project provides a complete evaluation pipeline including model inference, automated testing, and multi-dimensional assessment.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/internlm/InteractScience.STAR-Bench
STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
Zihan Liu*
·
Zhikang Niu*
·
Qiuyang Xiao
·
Zhisheng Zheng
·
Ruoqi Yuan
·
Yuhang Zang†
Yuhang Cao
·
Xiaoyi Dong
·
Jianze Liang
·
Xie Chen
·
Leilei Sun
·
Dahua Lin
·
Jiaqi Wang†
* Equal Contribution. †Corresponding authors.
📖Paper | 📖arXiv
|🏠Code
|🌐Homepage
|… See the full description on the dataset page: https://huggingface.co/datasets/internlm/STAR-Bench.CapRL-Video-178K
CapRL-Video-178K.jsonl Video Path Setup
Each value is a relative path under the Hugging Face dataset root of lmms-lab/LLaVA-Video-178K.
Example:
"video": "0_30_s_academic_v0_1/videos/academic_source/activitynet/v_01vNlQLepsE.mp4"
Required Video Data
Download the original videos from Hugging Face:
Dataset: https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K
Required split folders in this file:
0_30_s_youtube_v0_1: 72970 samples
2_3_m_youtube_v0_1: 24685… See the full description on the dataset page: https://huggingface.co/datasets/internlm/CapRL-Video-178K.CapRL-Video-QA-20K
CapRL-Video-QA-20K.jsonl Video Path Setup
Each value is a relative path under the Hugging Face dataset root of lmms-lab/LLaVA-Video-178K.
Example:
"videos": ["0_30_s_youtube_v0_1/videos/liwei_youtube_videos/videos/youtube_video_2024/ytb_khSwLQOthHQ.mp4"]
Required Video Data
Download the original videos from Hugging Face:
Dataset: https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K
Required subdirectories for this 20k subset:
0_30_s_youtube_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/internlm/CapRL-Video-QA-20K.Condor-SFT-20K
Condor
✨ Introduction
[🤗 HuggingFace Models]
[🤗 HuggingFace Datasets]
[📃 Paper]
The quality of Supervised Fine-Tuning (SFT) data plays a critical role in enhancing the conversational capabilities of Large Language Models (LLMs).
However, as LLMs become more advanced,
the availability of high-quality human-annotated SFT data has become a significant bottleneck,
necessitating a greater reliance on synthetic training data.
In this work, we introduce… See the full description on the dataset page: https://huggingface.co/datasets/internlm/Condor-SFT-20K.RNGBench-Game-Trajectories
RNGBench-Game-Trajectories
SFT (supervised fine-tuning) trajectory data accompanying RNGBench · Reconstructive Non-Markov
Games — an evaluation framework that tests whether multimodal language models can reconstruct
hidden state from memory and act on it in closed-loop environments (the "remember-to-act" setting,
where the current observation alone is not enough and the model must recall relevant history before
deciding).
📄 Paper: https://arxiv.org/abs/2606.19338
🌐 Project… See the full description on the dataset page: https://huggingface.co/datasets/internlm/RNGBench-Game-Trajectories.internlm__internlm2_5-7b-chat-details
Dataset Card for Evaluation run of internlm/internlm2_5-7b-chat
Dataset automatically created during the evaluation run of model internlm/internlm2_5-7b-chat
The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/internlm__internlm2_5-7b-chat-details.internlm__internlm2-7b-details
Dataset Card for Evaluation run of internlm__internlm2-7b
Dataset automatically created during the evaluation run of model internlm__internlm2-7b
The dataset is composed of 88 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/internlm__internlm2-7b-details.internlm__internlm2_5-20b-chat-details
Dataset Card for Evaluation run of internlm/internlm2_5-20b-chat
Dataset automatically created during the evaluation run of model internlm/internlm2_5-20b-chat
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/internlm__internlm2_5-20b-chat-details.internlm__internlm2-chat-1_8b-details
Dataset Card for Evaluation run of internlm/internlm2-chat-1_8b
Dataset automatically created during the evaluation run of model internlm/internlm2-chat-1_8b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/internlm__internlm2-chat-1_8b-details.internlm__internlm2-1_8b-details
Dataset Card for Evaluation run of internlm/internlm2-1_8b
Dataset automatically created during the evaluation run of model internlm/internlm2-1_8b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/internlm__internlm2-1_8b-details.internlm__internlm2_5-1_8b-chat-details
Dataset Card for Evaluation run of internlm/internlm2_5-1_8b-chat
Dataset automatically created during the evaluation run of model internlm/internlm2_5-1_8b-chat
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/internlm__internlm2_5-1_8b-chat-details.IntervitensInc__internlm2_5-20b-llamafied-details
Dataset Card for Evaluation run of IntervitensInc/internlm2_5-20b-llamafied
Dataset automatically created during the evaluation run of model IntervitensInc/internlm2_5-20b-llamafied
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/IntervitensInc__internlm2_5-20b-llamafied-details.
