CoolFace
Apppublic

RLE-Bench/libero-long-kinex-v0.4.0-astra-low-eval

sourceHugging Faceupdated 7d agoView on Hugging Face
0likes
App README

**Complete lightweight evidence dataset**

LIBERO Long: Kinex(v0.4.0) and Codex

18 recorded episodes, native task IDs 5, 6 and 9, seeds 0–2 per harness and subtask. GPT-6 Astra / low, standard service tier, fast mode disabled. Kinex source is labeled Kinex(v0.4.0); source identities include local changes, not merely a clean release tag.

Native taskKinex(v0.4.0) native successCodex native success
5: book in caddy1/3 (seed 0)1/3 (seed 2)
6: mug and pudding0/30/3
9: mug in microwave0/3; seed 2 quota-interrupted2/3 (seeds 1, 2)
Total recorded attempts1/93/9

17 executions finished normally; one was interrupted. Task 9 / Kinex(v0.4.0) / seed 2 ended with HTTP 429 usage_limit_reached and Kinex exit 1 after 4,736/5,000 control frames. The owner finalized and preserved a valid unsuccessful native state. This is not normal agent completion. The all-attempt native count includes that interrupted attempt; do not use it as an unqualified clean-completion comparison. Seed 1 ended normally after an explicit stop at 4,752 frames. All 18 native evidence records are valid. No evaluation was retried or continued for this publication.

The per-episode limits are 5,000 native control frames (including five initialization frames) and eight hours, synchronous simulation at 20 Hz. No extra model-loop cap is configured. Conversations and native worlds reset between seeds; the workspace is retained within each subtask/harness. Tools, skills and memos are the intended learning resources, but other workspace files are not strictly cleared. There is no automatic previous-episode verdict injected into the next prompt. Agents can record the result returned by their own stop call. These are sequential learning episodes, not independent zero-shot repetitions. This joint-device protocol is not the official LIBERO OSC benchmark.

Native outcomes and process exceptions are audited separately. No state-based physical root-cause analysis is claimed. Native evidence validity and a parent job's zero exit code do not prove that every agent episode completed normally.

All 18 videos use native state playback, two 768×768 views in a 1536×896 frame, 4× simulation speed, one-second initial and two-second final holds. Model waiting is omitted; all native states remain in private HD sources. Playback verifies state restoration and sampled camera alignment. This is not command resimulation. Kinex video captions are versioned; timing/frame maps are unchanged. Native session content is retained apart from documented host-path/account redactions.