orb
Datasets
All datasets matching “orb”or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.umi-okra-dex1
UMI Okra Grasping Dataset — Lab Subset (Unitree Dex1-1)
Universal Manipulation Interface (UMI) hand-held teleoperation data for an okra-fruit
grasping / harvesting task. Recorded to train a Diffusion Policy (with a comparison
ACT track) deployed on a Unitree G1.
This is the indoor lab subset. Every session here was recorded in a lab mock okra
field (artificial foliage, white-walled room). The outdoor sessions from the original
collection are not included — see Scope and… See the full description on the dataset page: https://huggingface.co/datasets/Orboh/umi-okra-dex1.gso-orbit-rgbaobjaverse_orbit_rendersOrbEvo
OrbEvo
Paper: https://arxiv.org/abs/2603.03511
Currently in this dataset:
5,000 QM9 TDDFT trajectories generated using the ABACUS simulation code (https://arxiv.org/abs/2501.08697).
Folders:
- QM9_tddft/ <--- the data used for training and evaluation
- processed/
- raw/
- MDA_tddft/
- MDA_train <--- the data used for training and validation
- processed/
- raw/
- MDA_test <--- the data used for test
- processed/
- raw/
- before_down_sampling/ <--- the… See the full description on the dataset page: https://huggingface.co/datasets/divelab/OrbEvo.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.
