datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AraMix-Native
AraMix-Native
A native-Arabic-filtered version of
AdaMLLab/AraMix (minhash_deduped),
derived from SultanR/AraMix-Translation-Scores:
machine-translated and garbled-MT documents removed, 162,887,010 rows kept of
178,883,241 (91.06%). All columns preserved.
Filter rules
A document is kept iff all of:
mmbert_translated_score < 0.1, or a classical-text rescue: diacritic
(tashkeel) ratio ≥ 0.02 over Arabic letters and ≥ 3 distinct diacritic
classes (fully/partially… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Native.central-florida-native-plants
DeepEarth Central Florida Native Plants Dataset v0.2.0
🌿 Dataset Summary
A comprehensive multimodal dataset featuring 33,665 observations of 232 native plant species from Central Florida. This dataset combines citizen science observations with state-of-the-art vision and language embeddings for advancing multimodal self-supervised ecological intelligence research.
Key Features
🌍 Spatiotemporal Coverage: Complete GPS coordinates and timestamps for all… See the full description on the dataset page: https://huggingface.co/datasets/deepearth/central-florida-native-plants.vam-cross-target-widowx250-native-corner-frontThis dataset was created using LeRobot.
Dataset Description
30.10 minutes (160 episodes, 160 marked successful) of the VAM-Cross target collection for widowx250 in MuJoCo at 30 Hz using the widowx-texture robot appearance. Observations include RGB, joint state, gripper state, and achieved EE pose; actions include the full commanded EE pose and gripper command. Task assets are derived from the MolmoSpaces THOR 20251117 asset release (MolmoSpaces commit… See the full description on the dataset page: https://huggingface.co/datasets/dreamdifferent/vam-cross-target-widowx250-native-corner-front.entity-native-agent-sessions
Entity-Native vs File-Native Agent Sessions on SWE-bench Verified
Full session logs from a controlled A/B experiment measuring how a coding agent's
retrieval substrate changes its behaviour, cost, and success rate on real
software-engineering tasks.
Both arms run the same model (Claude Sonnet 4.5), on the same tasks, from the
same repository state. The only difference is how the agent is allowed to find code.
Arm
Label
Tools available
A
file-native
Bash, Read, Grep… See the full description on the dataset page: https://huggingface.co/datasets/rs545837/entity-native-agent-sessions.central-florida-native-plants-language-embeddings
Central Florida Native Plants Language Embeddings
This dataset contains language embeddings for 232 native plant species from Central Florida, extracted using the DeepSeek-V3 language model.
Dataset Summary
This dataset provides pre-computed language embeddings for Central Florida plant species. Each species has been encoded using the prompt "Ecophysiology of {species_name}:" to capture semantic information about the plant's ecological characteristics.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/deepearth/central-florida-native-plants-language-embeddings.qfs-smollm2-135m-wikitext2-native-v1
HF workflow d3dc69602aeb981f06bd9f4c726937f9
A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-native-bf16.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-native-v1.web-access-api-benchmarks
NativePort Web-Access API Benchmarks
Measured quality, latency, cost and error-rate figures for 22 commercial web-access
APIs — search, SERP, scraping, crawling, extraction, sourced answers, screenshots,
document parsing, browser actions and change watching — scored per capability on a
fixed task corpus. This is the 2026-08-05 run: 67 provider × capability
scorecards across 13 capabilities, flattened into 297 metric rows.
It exists for one practical decision: when an AI agent… See the full description on the dataset page: https://huggingface.co/datasets/nativeport/web-access-api-benchmarks.umi-pr17-native-benchmark-20260708-172502This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 1,
"total_frames": 180,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/umi-pr17-native-benchmark-20260708-172502.osworld-native-parity-runs
OSWorld Native Parity Runs
This dataset stores large native OSWorld parity run archives that are too large for the shared Harbor parity-experiments dataset.
Archives
fulltask_20260621-172216/attempt_1/osworld-native-fulltask_20260621-172216-attempt1-349of361.tar.zst
Source run: /home/servermacadmin/osworld-parity/parity_results/fulltask_20260621-172216/attempt_1
Source upstream: xlang-ai/OSWorld at fe8c78e
Tasks: OSWorld-Verified no-Google-Drive split, 361 task… See the full description on the dataset page: https://huggingface.co/datasets/josancamon/osworld-native-parity-runs.maze2d_easy_native256_cot_chunk_kinf_20260707_perseg
maze2d_easy_native256_cot_chunk_kinf_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_kinf_20260707_perseg.subset-0-ros2-lerobot-nativetest-lerobot-nativeeval_test_migrated_nativeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LohanTS/eval_test_migrated_native.maze2d_easy_native256_noncot_chunk_k3_20260707_perseg
maze2d_easy_native256_noncot_chunk_k3_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k3_20260707_perseg.maze2d_easy_native256_cot_chunk_k5_20260707_perseg
maze2d_easy_native256_cot_chunk_k5_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k5_20260707_perseg.maze2d_easy_native256_noncot_chunk_k10_20260707_perseg
maze2d_easy_native256_noncot_chunk_k10_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k10_20260707_perseg.github-react-native-issuesmaze2d_easy_native256_cot_chunk_k10_20260707_perseg
maze2d_easy_native256_cot_chunk_k10_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k10_20260707_perseg.NativeDE-Opus4.7-REAP
NativeDE-Opus4.7-REAP
A native German synthetic reasoning dataset generated using Anthropic Claude Opus 4.7 (claude-opus-4-7). All prompts and responses are in natural, idiomatic German — not translations from English. Each sample contains an explicit <think>...</think> reasoning block followed by a Final answer: boundary and the actual response.
This dataset is the German-language complement to BaaderSo36-Opus4.7-REAP.
Dataset Statistics
Total samples: 2,306
Source… See the full description on the dataset page: https://huggingface.co/datasets/baaderso36/NativeDE-Opus4.7-REAP.taxi-fare-trainmaze2d_easy_native256_noncot_chunk_k5_20260707_perseg
maze2d_easy_native256_noncot_chunk_k5_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k5_20260707_perseg.maze2d_easy_native256_noncot_chunk_k1_20260707_perseg
maze2d_easy_native256_noncot_chunk_k1_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k1_20260707_perseg.maze2d_easy_native256_noncot_chunk_kinf_20260707_perseg
maze2d_easy_native256_noncot_chunk_kinf_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_kinf_20260707_perseg.maze2d_easy_native256_cot_chunk_k1_20260707_perseg
maze2d_easy_native256_cot_chunk_k1_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k1_20260707_perseg.housingvarroa-yolov8n-p2p3-native160-factorized-talfanar-eval-native-prompt
fanar-eval-native-prompt
Clean paper-facing evaluation dataset for the Shaer benchmark.
Rows: 3481
Split: train
Schema
id
base_meter
form
requested_bayts
requested_num_lines
description
enhanced_description
reference_completion
generated_text
meter
count_adherence
description_adherence
meaning
fluency
coherence
poeticness
Notes
meter is the row-level metrical conformity score used in the paper.
This dataset sets count_adherence to null in the… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/fanar-eval-native-prompt.taxi-fare-testshaer-eval-ashaar-native-controls
shaer-eval-ashaar-native-controls
Clean paper-facing evaluation dataset for the Shaer benchmark.
Rows: 3481
Split: test
Schema
id
base_meter
form
requested_bayts
requested_num_lines
description
enhanced_description
reference_completion
generated_text
meter
count_adherence
description_adherence
meaning
fluency
coherence
poeticness
Notes
meter is the row-level metrical conformity score used in the paper.
This dataset sets count_adherence to null in… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/shaer-eval-ashaar-native-controls.shaer-eval-ashaar-native-controls-with-hits
