CoolFace
Datasetpublic

RLE-Bench/libero-long-kinex-v0.4.0-astra-low-eval

Evaluation website · Data format · Episode CSV LIBERO Long: Kinex(v0.4.0) and Codex 18 recorded episodes, native task IDs 5, 6 and 9, seeds 0–2 per harness and subtask. GPT-6 Astra / low, standard service tier, fast mode disabled. Kinex source is labeled Kinex(v0.4.0); source identities include local changes, not merely a clean release tag. Native task Kinex(v0.4.0) native success Codex native success 5: book in caddy 1/3 (seed 0) 1/3 (seed 2) 6: mug and pudding 0/3… See the full description on the dataset page: https://huggingface.co/datasets/RLE-Bench/libero-long-kinex-v0.4.0-astra-low-eval.

sourceHugging Faceupdated 7d agoView on Hugging Face
0likes298downloads
Dataset Card

**Evaluation website** · Data format · Episode CSV

LIBERO Long: Kinex(v0.4.0) and Codex

18 recorded episodes, native task IDs 5, 6 and 9, seeds 0–2 per harness and subtask. GPT-6 Astra / low, standard service tier, fast mode disabled. Kinex source is labeled Kinex(v0.4.0); source identities include local changes, not merely a clean release tag.

Native taskKinex(v0.4.0) native successCodex native success
5: book in caddy1/3 (seed 0)1/3 (seed 2)
6: mug and pudding0/30/3
9: mug in microwave0/3; seed 2 quota-interrupted2/3 (seeds 1, 2)
Total recorded attempts1/93/9

17 executions finished normally; one was interrupted. Task 9 / Kinex(v0.4.0) / seed 2 ended with HTTP 429 usage_limit_reached and Kinex exit 1 after 4,736/5,000 control frames. The owner finalized and preserved a valid unsuccessful native state. This is not normal agent completion. The all-attempt native count includes that interrupted attempt; do not use it as an unqualified clean-completion comparison. Seed 1 ended normally after an explicit stop at 4,752 frames. All 18 native evidence records are valid. No evaluation was retried or continued for this publication.

The per-episode limits are 5,000 native control frames (including five initialization frames) and eight hours, synchronous simulation at 20 Hz. No extra model-loop cap is configured. Conversations and native worlds reset between seeds; the workspace is retained within each subtask/harness. Tools, skills and memos are the intended learning resources, but other workspace files are not strictly cleared. There is no automatic previous-episode verdict injected into the next prompt. Agents can record the result returned by their own stop call. These are sequential learning episodes, not independent zero-shot repetitions. This joint-device protocol is not the official LIBERO OSC benchmark.

Native outcomes and process exceptions are audited separately. No state-based physical root-cause analysis is claimed. Native evidence validity and a parent job's zero exit code do not prove that every agent episode completed normally.

All 18 videos use native state playback, two 768×768 views in a 1536×896 frame, 4× simulation speed, one-second initial and two-second final holds. Model waiting is omitted; all native states remain in private HD sources. Playback verifies state restoration and sampled camera alignment. This is not command resimulation. Kinex video captions are versioned; timing/frame maps are unchanged. Native session content is retained apart from documented host-path/account redactions.

Full lightweight outputs are under episodes/t05/, episodes/t06/, and episodes/t09/, then kinex|codex/episode_01..03/. Raw model.mjb, states.f64, unedited camera streams and duplicate evidence archives remain private.

Download with hf download RLE-Bench/libero-long-kinex-v0.4.0-astra-low-eval --repo-type dataset --local-dir libero-eval. Run python3 tools/hf_publication.py validate . in the downloaded directory to verify public hashes and links. Serve locally with python3 -m http.server 8000.