saber
Datasets
All datasets matching “saber”SID_Set
Dataset Card for SID_Set
Dataset Summary
We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages:
Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations.
Broad diversity: Encompassing fully synthetic and tampered images across various classes.
Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection.
Please check… See the full description on the dataset page: https://huggingface.co/datasets/saberzl/SID_Set.bridgev2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "WidowX",
"total_episodes": 53192,
"total_frames": 1999410,
"total_tasks": 19974,
"total_videos": 212768,
"total_chunks": 54,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:53192"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Saberlve/bridgev2.So-Fake-Set
Dataset Card for So-Fake-Set
Dataset Summary
We provide So-Fake-Set, A large-scale, diverse dataset tailored for social media image forgery detection!
Please check our website to explore more visual results.
Dataset Structure
"image" (Image): Input images, including real, full_synthetic, and tampered images.
"mask" (Image): Binary mask highlighting manipulated regions in tampered images.
"label" (str): Classification category.
"generator" (str): The… See the full description on the dataset page: https://huggingface.co/datasets/saberzl/So-Fake-Set.So-Fake-OOD
So-Fake-OOD-v3
So-Fake-OOD-v3 is the updated out-of-distribution evaluation split for So-Fake. It contains platform-native real images, fully synthetic images from held-out commercial generators, and locally tampered images with pixel-level masks.
Release Structure
The release is split according to platform redistribution constraints:
test_image: examples whose image content can be redistributed. This split includes Reddit, Tumblr, and Bluesky real images, all… See the full description on the dataset page: https://huggingface.co/datasets/saberzl/So-Fake-OOD.SABER
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
SABER is the code release for the paper. It includes the benchmark tasks, sandbox runtime, judging pipeline, and baseline reproduction utilities used to evaluate operational safety in stateful project workspaces.
What is included
tasks/: benchmark task definitions and metadata
run_osbench.py, judge_osbench.py: historical inference and judging entry points
sandbox_shell.py… See the full description on the dataset page: https://huggingface.co/datasets/sssr-lab/SABER.SaberMath-queries
SABER-Math Queries
Query set for the SABER-Math mathematical information reranking benchmark. Each of the 1,000 entries is a competition-style math problem serving as a query, together with a pool of candidate documents and graded relevance judgments for evaluating retrieval and reranking systems on mathematical content.
Dataset structure
Each example… See the full description on the dataset page: https://huggingface.co/datasets/INSAIT-Institute/SaberMath-queries.
