datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gspc-provenance-controls
GSPC — provenance controls facts (ChainFacts)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
MEASURED financial/domain axis (issuer-account / on-chain control facts, n=6). Not a model leaderboard. No accuracy, no fleet, no leader, no separation — measured is not scored.
Frozen bank on Hub. Live n and status are the provenance-controls row on GET https://councilof.ai/api/gspc. Not a certificate.
Council of AI · CSOAI Ltd… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-provenance-controls.repro-score-a-unified-framework-for-overshoot-refund-in-online-fdr-control-traces
Agent traces
Agent sessions published from a Trackio Logbook.
bedroom-ac-control
Bedroom AC Control
Synthetic decisions for a bedroom fan coil air conditioner. Each row is one 10 minute tick: sensor readings in, OFF or COOL out. The rows were generated by free models on OpenRouter (dots-3-note-preview, gemma-4-26b-a4b-it, ling-3.0-flash-fin, nemotron-3-super-120b-a12b, nemotron-3-ultra-550b-a55b) from a written policy, then checked against a rule implementation of the same policy.
The home has two HomePod minis in the far corner of the bedroom, one Mila… See the full description on the dataset page: https://huggingface.co/datasets/Aayush9029/bedroom-ac-control.StaR_state_control_benchmark
task_categories:
- image-text-to-text
StaR State Control Benchmark Dataset
This repository provides the state control benchmark of our paper: See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles.
Code: https://github.com/ZrW00/StaR
We provide the state control benchmark of our paper:
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
How to Use
The… See the full description on the dataset page: https://huggingface.co/datasets/ZrW00/StaR_state_control_benchmark.android_control_train
Processed Android Control Training Set
Dataset Description
This repository contains the processed training set derived from the Android Control dataset by Google Research.
The data processing methodology is identical to that used for our corresponding test set, which can be found at Reallm-Labs/android_control_test.
Data Content and Image Extraction
Important Note: Due to the large size of the dataset, this repository contains only the processed text files.… See the full description on the dataset page: https://huggingface.co/datasets/InfiX-ai/android_control_train.han-humanoid-grasp-force-control-v1
Humanoid Grasp Force Control Dataset
Overview
Dataset ini berisi parameter kontrol genggaman tangan humanoid
saat memanipulasi berbagai objek.
Features
object_weight_kg
object_surface_friction
grip_contact_area_cm2
finger_joint_angle_deg
actuator_current_amp
object_fragility_index
Target
required_grip_force_newton
Task
Regression
combined-control-no-sparsity-trainingdl-trm-phase2-controller-kmeans-v128
DL-TRM Phase 2 Controller K-Means V128
Discrete Z traces produced by k-means over Phase 1 controller deltas.
Source trace artifact: phase1_real_medoid_route_trace_shards
Input signal: medoid_controller_delta_trace[1:]
Input shape before quantization: [11400, 16, 512]
Output Z trace shape: [11400, 16]
Vocabulary size: 128
Token id base: 0
Diagnostics
{
"used_codes": 128,
"dead_codes": 0,
"perplexity": 123.55767954902056,
"unique_traces": 9222… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/dl-trm-phase2-controller-kmeans-v128.Delta-Vector__Control-8B-V1.1-details
Dataset Card for Evaluation run of Delta-Vector/Control-8B-V1.1
Dataset automatically created during the evaluation run of model Delta-Vector/Control-8B-V1.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Control-8B-V1.1-details.Delta-Vector__Control-8B-details
Dataset Card for Evaluation run of Delta-Vector/Control-8B
Dataset automatically created during the evaluation run of model Delta-Vector/Control-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Control-8B-details.dl-trm-phase2-controller-kmeans-v32
DL-TRM Phase 2 Controller K-Means V32
Discrete Z traces produced by k-means over Phase 1 controller deltas.
Source trace artifact: phase1_real_medoid_route_trace_shards
Input signal: medoid_controller_delta_trace[1:]
Input shape before quantization: [11400, 16, 512]
Output Z trace shape: [11400, 16]
Vocabulary size: 32
Token id base: 0
Diagnostics
{
"used_codes": 32,
"dead_codes": 0,
"perplexity": 31.17811290262774,
"unique_traces": 9191… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/dl-trm-phase2-controller-kmeans-v32.arena-control-apps-v5-dataset
Arena Control Apps V5 Dataset
A curated dataset for training and evaluating models on collusion signal detection in code.
Dataset Overview
Metric
Value
Total samples
1,793
Train split
1,493
Test split
300
Label 0 (clean)
765
Label 1 (backdoor)
1,028
Bucket Distribution
Bucket
Count
Label
Description
clean
465
0
Clean code, no backdoor, no signal
clean_signal_a
100
0
Clean code + Signal A
clean_signal_b
100
0
Clean code +… See the full description on the dataset page: https://huggingface.co/datasets/jprivera44/arena-control-apps-v5-dataset.combined-control-with-sparsityrollouts-run-list-control-new-prompt-hfdl-trm-phase2-controller-kmeans-v16
DL-TRM Phase 2 Controller K-Means V16
Discrete Z traces produced by k-means over Phase 1 controller deltas.
Source trace artifact: phase1_real_medoid_route_trace_shards
Input signal: medoid_controller_delta_trace[1:]
Input shape before quantization: [11400, 16, 512]
Output Z trace shape: [11400, 16]
Vocabulary size: 16
Token id base: 0
Diagnostics
{
"used_codes": 16,
"dead_codes": 0,
"perplexity": 15.491387192451063,
"unique_traces": 9176… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/dl-trm-phase2-controller-kmeans-v16.dl-trm-phase2-controller-kmeans-v256
DL-TRM Phase 2 Controller K-Means V256
Discrete Z traces produced by k-means over Phase 1 controller deltas.
Source trace artifact: phase1_real_medoid_route_trace_shards
Input signal: medoid_controller_delta_trace[1:]
Input shape before quantization: [11400, 16, 512]
Output Z trace shape: [11400, 16]
Vocabulary size: 256
Token id base: 0
Diagnostics
{
"used_codes": 256,
"dead_codes": 0,
"perplexity": 244.03568219976052,
"unique_traces": 9225… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/dl-trm-phase2-controller-kmeans-v256.Code-code-galeras-prompting-3k-controlshaer-eval-ashaar-native-controls
shaer-eval-ashaar-native-controls
Clean paper-facing evaluation dataset for the Shaer benchmark.
Rows: 3481
Split: test
Schema
id
base_meter
form
requested_bayts
requested_num_lines
description
enhanced_description
reference_completion
generated_text
meter
count_adherence
description_adherence
meaning
fluency
coherence
poeticness
Notes
meter is the row-level metrical conformity score used in the paper.
This dataset sets count_adherence to null in… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/shaer-eval-ashaar-native-controls.humanoid-balance-control-samplesfan_robot_motor_control.json
