datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.RPC-Bench
RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
🌐 Project Page •
💻 GitHub •
📖 Paper
RPC-Bench is a fine-grained benchmark for research paper comprehension. It is built from review-rebuttal exchanges of high-quality academic papers and supports both text-only and visual evaluation through complementary paper representations.
Data Structure
RPC-Bench is organized into train, dev, and test subsets. Split assignments… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/RPC-Bench.ORD
🌋 STORM: Stimulating Trustworthy Ordinal Regression Ability of MLLMs
Benchmarking All-in-one Visual Rating of MLLMs with A Comprehensive Ordinal Regression Dataset.
Contents
STORM Weights
Dataset
Evaluation
Examples
STORM Weights
Please check out our checkpoint_STORM for public STORM checkpoints, and the instructions of how to use the weights.
Dataset
Data file name
Size
STORM_instruct_MAX_527k.jsonl
383 MB… See the full description on the dataset page: https://huggingface.co/datasets/ttlyy/ORD.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our leaderboard at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.or-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.orz_math_57k_collection
Open Reasoner Zero
An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Paper Arxiv Link 👁️
Overview 🌊
We introduce Open-Reasoner-Zero, the first open source implementation of large-scale reasoning-oriented RL training focusing on scalability, simplicity and accessibility.
To enable broader participation in this pivotal moment we witnessed and accelerate research towards artificial general intelligence (AGI)… See the full description on the dataset page: https://huggingface.co/datasets/Open-Reasoner-Zero/orz_math_57k_collection.VisualReasoner-30k
Dataset Card for VisualReasoner-30k
Dataset Details
This dataset is an extension of VisualReasoner-1M, containing approximately 30k cases and can be used for training visual reasoning tasks.
Unlike VisualReasoner-1M, this dataset models the reasoning process in an end-to-end format to better accommodate scenarios where explicit tool invocation is not allowed.
Dataset Descriptions
The structure of each case is as follows:
{
"identity": "Case ID"… See the full description on the dataset page: https://huggingface.co/datasets/orange-sk/VisualReasoner-30k.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/jerogo/or-bench.ImageMining
Dataset Card for ImageMining
Dataset Description
ImageMining is a benchmark for evaluating image mining and knowledge discovery capabilities of multimodal models. Given an image, the task requires models to identify entities, perform multi-step reasoning (often with search-augmented information), and answer complex questions that go beyond simple visual understanding.
The dataset contains 217 examples across 7 top-level categories and 23 subcategories.… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/ImageMining.VisualReasoner-1M
Dataset Card for VisualReasoner-1M
Dataset Details
This is a dataset for the paper From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis. The dataset contains approximately 1 million cases and can be used for training visual reasoning tasks. The reasoning process involves breaking down tasks and utilizing tools to solve complex and challenging visual question-answering tasks progressively.
For detailed data synthesis methods, please… See the full description on the dataset page: https://huggingface.co/datasets/orange-sk/VisualReasoner-1M.orz_math_13k_collection_hard
Open Reasoner Zero
An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Paper Arxiv Link 👁️
Overview 🌊
We introduce Open-Reasoner-Zero, the first open source implementation of large-scale reasoning-oriented RL training focusing on scalability, simplicity and accessibility.
To enable broader participation in this pivotal moment we witnessed and accelerate research towards artificial general intelligence (AGI)… See the full description on the dataset page: https://huggingface.co/datasets/Open-Reasoner-Zero/orz_math_13k_collection_hard.
