datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UGround-Offline-Evaluationoffline-packsoffline-micro-saas-catalog
📦 SaveDollars.store — Offline Micro SaaS & Autonomous AI Software Catalog
This dataset contains structured product metadata, architecture specifications, pricing, and documentation for 96 standalone offline Micro SaaS applications, autonomous AI agent command centers, and business operating systems published by SaveDollars.store.
📊 Dataset Structure (catalog.json)
Each record represents a production-ready, subscription-free software package:
{
"id": 75809… See the full description on the dataset page: https://huggingface.co/datasets/SaveDollars/offline-micro-saas-catalog.tb21-eval-qwen35-4b-offline-echo-action-only-10k-tacc-timeout2x
qwen35-action-only-10k — Terminal-Bench 2.1 (timeout multiplier 2x)
Terminal-Bench 2.1 evaluation protocol variant (timeout multiplier 2x) of violetxi/qwen35-4b-offline-echo-action-only-10k-tacc through the served
model ID qwen35-action-only-10k with Terminus-2.
Result
Evaluation trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero: 219
Agent timeouts / context-length events / output-cap events:
217 / 0 /
0
Mean reward / Pass@1: 0.105618
Pass@5:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-4b-offline-echo-action-only-10k-tacc-timeout2x.rlve_offline_20K_POPE_prefix_pass1_qwen3-1.7b
RLVE offline-20K POPE-prefix completions — Qwen3-1.7B (pass1)
Prefix-conditioned completions generated by Qwen3-1.7B over the
rlve_offline_20K_POPE_prefix prompt set (20000 records, 1 sample/prompt).
Produced by SLURM job 6580578 (vLLM, tp=2), 2026-06-15.
Fields
index, sample_id, prompt, prefix, response, answer, rewards
⚠️ Caveat on rewards
The inline rewards field is all 0.0 — this is the known inline-Gym-verifier
artifact (same as the old… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rlve_offline_20K_POPE_prefix_pass1_qwen3-1.7b.llama_o1_offline_training_data_v1offline-practical-skills-qa-synthetic
Offline Practical Skills QA Dataset (Synthetic)
Dataset generated by: Cahlen Humphreys
Dataset Description
This dataset contains question-answer pairs covering several practical, everyday knowledge domains intended to be useful in offline or edge scenarios (e.g., without internet access). It was generated entirely synthetically using the meta-llama/Meta-Llama-3-8B-Instruct model and subsequently deduplicated based on the question field to enhance uniqueness.
The primary… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/offline-practical-skills-qa-synthetic.Offline_GRPOWebVoyager-GAIA-SparseAttention-Offline-Qwen3-VL-30B-A3B
Sparse Attention for Web Agents — Offline Replay Records
Step-level records from an offline replay study comparing three sparse-attention methods against
full attention on browser-agent trajectories.
Model: Qwen3-VL-30B-A3B-Instruct (48 layers, 128 experts / 8 active, GQA 32:4, head_dim 128)
Hardware: NVIDIA GB10 (sm_121)
Tasks: 50 complete trajectories sampled from WebVoyager + GAIA — 332 steps, of which 190 emit an
element index. Every task was completed successfully by the… See the full description on the dataset page: https://huggingface.co/datasets/shiqihe/WebVoyager-GAIA-SparseAttention-Offline-Qwen3-VL-30B-A3B.ascii-tanks-offline-dataset
