CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SultanR /AraMix-Native AraMix-Native A native-Arabic-filtered version of AdaMLLab/AraMix (minhash_deduped), derived from SultanR/AraMix-Translation-Scores: machine-translated and garbled-MT documents removed, 162,887,010 rows kept of 178,883,241 (91.06%). All columns preserved. Filter rules A document is kept iff all of: mmbert_translated_score < 0.1, or a classical-text rescue: diacritic (tashkeel) ratio ≥ 0.02 over Arabic letters and ≥ 3 distinct diacritic classes (fully/partially… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Native.tabular100M<n<1B0 likes532 downloads2mo agoHugging Face02deepearth /central-florida-native-plants DeepEarth Central Florida Native Plants Dataset v0.2.0 🌿 Dataset Summary A comprehensive multimodal dataset featuring 33,665 observations of 232 native plant species from Central Florida. This dataset combines citizen science observations with state-of-the-art vision and language embeddings for advancing multimodal self-supervised ecological intelligence research. Key Features 🌍 Spatiotemporal Coverage: Complete GPS coordinates and timestamps for all… See the full description on the dataset page: https://huggingface.co/datasets/deepearth/central-florida-native-plants.tabularimage-classification10K<n<100K0 likes274 downloads1y agoHugging Face03dreamdifferent /vam-cross-target-widowx250-native-corner-frontThis dataset was created using LeRobot. Dataset Description 30.10 minutes (160 episodes, 160 marked successful) of the VAM-Cross target collection for widowx250 in MuJoCo at 30 Hz using the widowx-texture robot appearance. Observations include RGB, joint state, gripper state, and achieved EE pose; actions include the full commanded EE pose and gripper command. Task assets are derived from the MolmoSpaces THOR 20251117 asset release (MolmoSpaces commit… See the full description on the dataset page: https://huggingface.co/datasets/dreamdifferent/vam-cross-target-widowx250-native-corner-front.tabularrobotics10K<n<100K0 likes103 downloads7d agoHugging Face04rs545837 /entity-native-agent-sessions Entity-Native vs File-Native Agent Sessions on SWE-bench Verified Full session logs from a controlled A/B experiment measuring how a coding agent's retrieval substrate changes its behaviour, cost, and success rate on real software-engineering tasks. Both arms run the same model (Claude Sonnet 4.5), on the same tasks, from the same repository state. The only difference is how the agent is allowed to find code. Arm Label Tools available A file-native Bash, Read, Grep… See the full description on the dataset page: https://huggingface.co/datasets/rs545837/entity-native-agent-sessions.tabulartext-generationn<1K0 likes99 downloads20d agoHugging Face05deepearth /central-florida-native-plants-language-embeddings Central Florida Native Plants Language Embeddings This dataset contains language embeddings for 232 native plant species from Central Florida, extracted using the DeepSeek-V3 language model. Dataset Summary This dataset provides pre-computed language embeddings for Central Florida plant species. Each species has been encoded using the prompt "Ecophysiology of {species_name}:" to capture semantic information about the plant's ecological characteristics. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/deepearth/central-florida-native-plants-language-embeddings.tabularfeature-extraction1K<n<10K0 likes73 downloads1y agoHugging Face06malaiwah /qfs-smollm2-135m-wikitext2-native-v1 HF workflow d3dc69602aeb981f06bd9f4c726937f9 A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-native-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-native-v1.tabularn<1K0 likes63 downloads14d agoHugging Face07nativeport /web-access-api-benchmarks NativePort Web-Access API Benchmarks Measured quality, latency, cost and error-rate figures for 22 commercial web-access APIs — search, SERP, scraping, crawling, extraction, sourced answers, screenshots, document parsing, browser actions and change watching — scored per capability on a fixed task corpus. This is the 2026-08-05 run: 67 provider × capability scorecards across 13 capabilities, flattened into 297 metric rows. It exists for one practical decision: when an AI agent… See the full description on the dataset page: https://huggingface.co/datasets/nativeport/web-access-api-benchmarks.tabularn<1K0 likes40 downloads1mo agoHugging Face08brandonyang /umi-pr17-native-benchmark-20260708-172502This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": null, "total_episodes": 1, "total_frames": 180, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/umi-pr17-native-benchmark-20260708-172502.tabularroboticsn<1K0 likes39 downloads3mo agoHugging Face09josancamon /osworld-native-parity-runs OSWorld Native Parity Runs This dataset stores large native OSWorld parity run archives that are too large for the shared Harbor parity-experiments dataset. Archives fulltask_20260621-172216/attempt_1/osworld-native-fulltask_20260621-172216-attempt1-349of361.tar.zst Source run: /home/servermacadmin/osworld-parity/parity_results/fulltask_20260621-172216/attempt_1 Source upstream: xlang-ai/OSWorld at fe8c78e Tasks: OSWorld-Verified no-Google-Drive split, 361 task… See the full description on the dataset page: https://huggingface.co/datasets/josancamon/osworld-native-parity-runs.tabularn<1K0 likes37 downloads3mo agoHugging Face10ultrastar111 /maze2d_easy_native256_cot_chunk_kinf_20260707_perseg maze2d_easy_native256_cot_chunk_kinf_20260707_perseg Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action executed." + real… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_kinf_20260707_perseg.tabularreinforcement-learning10K<n<100K0 likes32 downloads2mo agoHugging Face11Sraghvi /subset-0-ros2-lerobot-nativetabularn<1K0 likes31 downloads1y agoHugging Face12Sraghvi /test-lerobot-nativetabular1K<n<10K0 likes29 downloads1y agoHugging Face13LohanTS /eval_test_migrated_nativeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 0, "total_frames": 0, "total_tasks": 0, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": {}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LohanTS/eval_test_migrated_native.tabularrobotics1K<n<10K0 likes27 downloads3mo agoHugging Face14ultrastar111 /maze2d_easy_native256_noncot_chunk_k3_20260707_perseg maze2d_easy_native256_noncot_chunk_k3_20260707_perseg Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k3_20260707_perseg.tabularreinforcement-learning100K<n<1M0 likes26 downloads2mo agoHugging Face15ultrastar111 /maze2d_easy_native256_cot_chunk_k5_20260707_perseg maze2d_easy_native256_cot_chunk_k5_20260707_perseg Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k5_20260707_perseg.tabularreinforcement-learning10K<n<100K0 likes25 downloads2mo agoHugging Face16ultrastar111 /maze2d_easy_native256_noncot_chunk_k10_20260707_perseg maze2d_easy_native256_noncot_chunk_k10_20260707_perseg Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k10_20260707_perseg.tabularreinforcement-learning100K<n<1M0 likes21 downloads2mo agoHugging Face17hibikigf88 /github-react-native-issuestabulartext-generation1K<n<10K0 likes20 downloads1y agoHugging Face18ultrastar111 /maze2d_easy_native256_cot_chunk_k10_20260707_perseg maze2d_easy_native256_cot_chunk_k10_20260707_perseg Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k10_20260707_perseg.tabularreinforcement-learning10K<n<100K0 likes20 downloads2mo agoHugging Face19baaderso36 /NativeDE-Opus4.7-REAP NativeDE-Opus4.7-REAP A native German synthetic reasoning dataset generated using Anthropic Claude Opus 4.7 (claude-opus-4-7). All prompts and responses are in natural, idiomatic German — not translations from English. Each sample contains an explicit <think>...</think> reasoning block followed by a Final answer: boundary and the actual response. This dataset is the German-language complement to BaaderSo36-Opus4.7-REAP. Dataset Statistics Total samples: 2,306 Source… See the full description on the dataset page: https://huggingface.co/datasets/baaderso36/NativeDE-Opus4.7-REAP.tabulartext-generation1K<n<10K0 likes18 downloads5mo agoHugging Face20NativeFunction /taxi-fare-traintabular10M<n<100M0 likes16 downloads3y agoHugging Face21ultrastar111 /maze2d_easy_native256_noncot_chunk_k5_20260707_perseg maze2d_easy_native256_noncot_chunk_k5_20260707_perseg Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k5_20260707_perseg.tabularreinforcement-learning100K<n<1M0 likes15 downloads2mo agoHugging Face22ultrastar111 /maze2d_easy_native256_noncot_chunk_k1_20260707_perseg maze2d_easy_native256_noncot_chunk_k1_20260707_perseg Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k1_20260707_perseg.tabularreinforcement-learning10K<n<100K0 likes14 downloads2mo agoHugging Face23ultrastar111 /maze2d_easy_native256_noncot_chunk_kinf_20260707_perseg maze2d_easy_native256_noncot_chunk_kinf_20260707_perseg Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_kinf_20260707_perseg.tabularreinforcement-learning100K<n<1M0 likes14 downloads2mo agoHugging Face24ultrastar111 /maze2d_easy_native256_cot_chunk_k1_20260707_perseg maze2d_easy_native256_cot_chunk_k1_20260707_perseg Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k1_20260707_perseg.tabularreinforcement-learning10K<n<100K0 likes12 downloads2mo agoHugging Face25NativeFunction /housingtabular10K<n<100K0 likes9 downloads3y agoHugging Face26duyle2408 /varroa-yolov8n-p2p3-native160-factorized-taltabularn<1K0 likes8 downloads1mo agoHugging Face27Shaer-AI /fanar-eval-native-prompt fanar-eval-native-prompt Clean paper-facing evaluation dataset for the Shaer benchmark. Rows: 3481 Split: train Schema id base_meter form requested_bayts requested_num_lines description enhanced_description reference_completion generated_text meter count_adherence description_adherence meaning fluency coherence poeticness Notes meter is the row-level metrical conformity score used in the paper. This dataset sets count_adherence to null in the… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/fanar-eval-native-prompt.tabular1K<n<10K0 likes6 downloads3mo agoHugging Face28NativeFunction /taxi-fare-testtabular1K<n<10K0 likes5 downloads3y agoHugging Face29Shaer-AI /shaer-eval-ashaar-native-controls shaer-eval-ashaar-native-controls Clean paper-facing evaluation dataset for the Shaer benchmark. Rows: 3481 Split: test Schema id base_meter form requested_bayts requested_num_lines description enhanced_description reference_completion generated_text meter count_adherence description_adherence meaning fluency coherence poeticness Notes meter is the row-level metrical conformity score used in the paper. This dataset sets count_adherence to null in… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/shaer-eval-ashaar-native-controls.tabular1K<n<10K0 likes5 downloads3mo agoHugging Face30Shaer-AI /shaer-eval-ashaar-native-controls-with-hitstabular1K<n<10K0 likes3 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.