datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WildClawBench-TrajectoriesWildClawBench Trajectories
Complete OpenClaw agent trajectories from the WildClawBench evaluation — every message, reasoning block, tool call, and tool result from real long-horizon agent runs, released for independent verification, side-by-side comparison, and trace-level analysis.
Each evaluated model covers the full 60-task suite, and the collection is continuously updated as new models join the leaderboard. The directories under sessions/ always reflect the current model roster.… See the full description on the dataset page: https://huggingface.co/datasets/internlm/WildClawBench-Trajectories.CapRL-QA-75K
CapRL 75K QA Training Dataset
This dataset is the carefully filtered 75K QA training set used by CapRL to train CapRL-3B, a lightweight image captioning model initialized from Qwen2.5-VL-3B. It contains 75,285 samples, where each image is paired with multiple multiple-choice QA items. The dataset is designed for the two-stage CapRL training objective, where caption quality is evaluated through answerability of visual questions.
The QA construction pipeline is fully open-sourced in… See the full description on the dataset page: https://huggingface.co/datasets/internlm/CapRL-QA-75K.VC-RewardBench
Visual-ERM
Visual-ERM is a multimodal generative reward model for vision-to-code tasks.It evaluates outputs directly in the rendered visual space and produces fine-grained, interpretable, and task-agnostic discrepancy feedback for structured visual reconstruction.
📄 Paper |
💻 GitHub |
📊 VC-RewardBench
Model Overview
Existing rewards for vision-to-code usually fall into two categories:
Text-based rewards such as edit distance or TEDS, which ignore… See the full description on the dataset page: https://huggingface.co/datasets/internlm/VC-RewardBench.Spatial-SSRL-81k
Spatial-SSRL-81k
📖Paper| 🏠Github |🤗Spatial-SSRL-7B Model |
🤗Spatial-SSRL-3B Model | 🤗Spatial-SSRL-Qwen3VL-4B Model |
🤗Spatial-SSRL-81k Dataset | 📰Daily Paper
Spatial-SSRL-81k is a training dataset for enhancing spatial understanding in large vision-language models. It contains 81,053 samples of five pretext tasks for self-supervised learning, offering simple, intrinsic supervision that scales RLVR efficiently.
📢 News
🚀 [2026/04/05] We have released… See the full description on the dataset page: https://huggingface.co/datasets/internlm/Spatial-SSRL-81k.InteractScience
InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
InteractScience is a benchmark specifically designed to evaluate the capability of large language models in generating interactive scientific demonstration code. This project provides a complete evaluation pipeline including model inference, automated testing, and multi-dimensional assessment.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/internlm/InteractScience.CapRL-2M
CapRL
📖Paper | 🏠Github | 🤗CapRL Collection | 🤗Daily Paper
CapRL Series Model & Dataset
Series
Models & Resources
CapRL 2.0 Series
🤗 CapRL-Qwen3VL-2B | 🤗 CapRL-Qwen3VL-4B | 📦 CapRL-Qwen3VL-2B-GGUF | 📦 CapRL-Qwen3VL-4B-GGUF | 🌈CapRL-Qwen3VL-4B Space
CapRL 1.0 Series
🤗 CapRL-Qwen2.5VL-3B | 🤗 CapRL-InternVL3.5-8B |📊 CapRL-QA-75K Dataset | 📊 CapRL-2M Dataset | 📦 CapRL-3B-GGUF | 📦 CapRL-3B-i1-GGUF | 🌈CapRL-Qwen2.5VL-3B Space
Now you can try… See the full description on the dataset page: https://huggingface.co/datasets/internlm/CapRL-2M.ETCHR-GRPO-10K
ETCHR GRPO-10K
📖Paper
| 🏠Homepage
| 🤗ETCHR-FLUX.2-klein-9B Model
| 🤗ETCHR SFT-400K Dataset
| 🤗ETCHR GRPO-10K Dataset
| 🤗DL3DV-2K Benchmark
ETCHR GRPO-10K is the GRPO training data for further enhance ETCHR's editing capaibility in assisting understanding models. It contains 10000 samples of five tasks (Fine-grained Perception, Chart Understanding, Maze Solving, Jigsaw Puzzle and Spatial Understanding). Each sample contains the image to be edited, an editing… See the full description on the dataset page: https://huggingface.co/datasets/internlm/ETCHR-GRPO-10K.DL3DV-2k
DL3DV-2K
📖Paper
| 🏠Homepage
| 🤗ETCHR-FLUX.2-klein-9B Model
| 🤗ETCHR SFT-400K Dataset
| 🤗ETCHR GRPO-10K Dataset
| 🤗DL3DV-2K Benchmark
DL3DV-2K is a benchmark constructed from the DL3DV dataset for evaluating the viewpoint transformation capability of large models in spatial reasoning tasks, comprising 2K samples in total. Each sample contains: images (original images), aux_images (transformed images provided for human reference only and not used as question input)… See the full description on the dataset page: https://huggingface.co/datasets/internlm/DL3DV-2k.CapRL-Evaluation-Files
CapRL Evaluation Files
This dataset contains the files used by the CapRL Prism evaluation scripts.
Files
json_file/: 12 Prism evaluation JSON files.
bench_image_folder.zip: images used by the JSON files. After unzipping, it creates bench_image_folder/.
Each JSON stores image paths relative to the dataset root, for example:
bench_image_folder/lmm_eval_chartqa/41699051005347.png
Usage
huggingface-cli download internlm/CapRL-Evaluation-Files --repo-type dataset… See the full description on the dataset page: https://huggingface.co/datasets/internlm/CapRL-Evaluation-Files.
