datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm_plan_gen_dataset_accu_t1_t3_t4
[!IMPORTANT]
This is the training dataset for the ICAPS 2025 paper "Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation".
from pathlib import Path
import os
import jsonlines
from copy import deepcopy
from datasets import load_dataset
from icecream import ic
import enum
from enum import IntEnum
from enum import auto
class CONFIG_TYPES(enum.Enum):
# "type_id": ['t0', 'accu-t1', 'accu-t2', 'accu-t3', 'accu-t4', 'accu-t4', 'accu-t1+t4', 'accu-t2+t4']… See the full description on the dataset page: https://huggingface.co/datasets/huangsukai/llm_plan_gen_dataset_accu_t1_t3_t4.llm_plan_gen_dataset_accu_t4
[!IMPORTANT]
This is the training dataset for the ICAPS 2025 paper "Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation".
from pathlib import Path
import os
import jsonlines
from copy import deepcopy
from datasets import load_dataset
from icecream import ic
import enum
from enum import IntEnum
from enum import auto
class CONFIG_TYPES(enum.Enum):
# "type_id": ['t0', 'accu-t1', 'accu-t2', 'accu-t3', 'accu-t4', 'accu-t4', 'accu-t1+t4', 'accu-t2+t4']… See the full description on the dataset page: https://huggingface.co/datasets/huangsukai/llm_plan_gen_dataset_accu_t4.dstc11.t4
DSTC11: Dialogue System Technology Challenge 11
Track 4: Robust and Multilingual Automatic Evaluation Metrics for Open-Domain Dialogue Systems
This public dataset release contains multilingual and robustness data for building and evaluating automatic metrics for open-domain dialogue systems. It includes public train and development data, public evaluation templates, and auxiliary metadata. Held-out test data is not included in this public Hugging Face release.… See the full description on the dataset page: https://huggingface.co/datasets/mario-rc/dstc11.t4.llm_plan_gen_dataset_accu_t2_t4
[!IMPORTANT]
This is the training dataset for the ICAPS 2025 paper "Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation".
from pathlib import Path
import os
import jsonlines
from copy import deepcopy
from datasets import load_dataset
from icecream import ic
import enum
from enum import IntEnum
from enum import auto
class CONFIG_TYPES(enum.Enum):
# "type_id": ['t0', 'accu-t1', 'accu-t2', 'accu-t3', 'accu-t4', 'accu-t4', 'accu-t1+t4', 'accu-t2+t4']… See the full description on the dataset page: https://huggingface.co/datasets/huangsukai/llm_plan_gen_dataset_accu_t2_t4.llm_plan_gen_dataset_accu_t1_t4
[!IMPORTANT]
This is the training dataset for the ICAPS 2025 paper "Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation".
from pathlib import Path
import os
import jsonlines
from copy import deepcopy
from datasets import load_dataset
from icecream import ic
import enum
from enum import IntEnum
from enum import auto
class CONFIG_TYPES(enum.Enum):
# "type_id": ['t0', 'accu-t1', 'accu-t2', 'accu-t3', 'accu-t4', 'accu-t4', 'accu-t1+t4', 'accu-t2+t4']… See the full description on the dataset page: https://huggingface.co/datasets/huangsukai/llm_plan_gen_dataset_accu_t1_t4.atlas-of-judgment
Atlas of Judgment — ICLR Peer-Review Logic Units
1,420,178 atomic units of evaluative logic extracted from public ICLR peer
reviews (OpenReview), each decomposed into what was inspected → what was
observed → how it was reasoned about → what was concluded, and labeled with a
data-induced taxonomy of 12 objects of scrutiny × 12 epistemic standards.
This is the dataset behind the Atlas of Judgment
(interactive visual atlas). Full reproduction record — every script, model,
parameter… See the full description on the dataset page: https://huggingface.co/datasets/t46/atlas-of-judgment.llp-gold-37m-1.5m_T4096.0CaptchaSolve30k
CaptchaSolve30k - Human Mouse Movement Dataset
The largest open-source dataset of human task-specific mouse trajectories by session count and unique participants, with 30,000 discrete sessions from thousands of users. The first and only open dataset of complete human captcha-solving interactions with full behavioral replays.
Each session captures mouse/touch trajectories, timing data, and puzzle state at physics-tick resolution. Suitable for bot detection research… See the full description on the dataset page: https://huggingface.co/datasets/T40/CaptchaSolve30k.phi-redactor-evalmem_agent-model_based-memagent-1-5b-step1024-docfinqa-train-c8192-t4096-1000s-agnosticdepth_cache_t4zs4_vggt
depth_cache/t4zs4_vggt
cable_depth のセル格子キャッシュ。学習時に --depth-cache へ渡す。
対応するデータセット
takeru01/task4zeroshot3
このキャッシュは (episode_index, frame_index) で引くので、同じエピソード構成・同じ長さのデータセットにしか使えない。
構成
backbone vggt
wrist_backbone da3
metric False
grid [8, 14] (gh, gw)
stride 2
episodes 210 範囲 (0, 209)
frames 160915
カメラと役割
camera_front static
camera_top… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/depth_cache_t4zs4_vggt.depth_cache_t4zs4_da3
depth_cache/t4zs4_da3
cable_depth のセル格子キャッシュ。学習時に --depth-cache へ渡す。
対応するデータセット
takeru01/task4zeroshot3
このキャッシュは (episode_index, frame_index) で引くので、同じエピソード構成・同じ長さのデータセットにしか使えない。
構成
backbone da3
wrist_backbone da3
metric True
grid [8, 14] (gh, gw)
stride 2
episodes 210 範囲 (0, 209)
frames 160915
カメラと役割
camera_front static
camera_top static… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/depth_cache_t4zs4_da3.rlpt_37M_16epochs_501k_generations_SNIS_T4.0icml2026-T4A2aYyaq9-repro-traces
Agent traces
Agent sessions published from a Trackio Logbook.
ETHiQ
Dataset Card for Dataset Name
ETHiQ — E(thics), T(rivia), Hi(story), (Philosophy) Q(uestions)
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/T404C/ETHiQ.lever2-pilot-glm47-swesmith-t4096-32k-tracesQGCNQ
Dataset Card for Dataset Name
QGCNQ – Quantum Mechanics, Genetics, Cosmology, Neuroscience, Cognitive Science Questions
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper… See the full description on the dataset page: https://huggingface.co/datasets/T404C/QGCNQ.phi-redactor-dataosm-planet-tiles-T41
OSM Planet Tiles - Hecke Operator T_41
Tiles sharded by Hecke operator T_41 (Monster prime 41).
Dataset Info
Hecke Operator: T_41
Monster Prime: 41
Files: 65994 tiles
Sharding: Hash mod 15 → T_41
Format: PBF/Parquet
License: ODbL (OpenStreetMap)
Monster Symmetries
Input: [71, 59, 47] (Keter/Binah/Chokmah)
Output: [17, 23, 59] (Cusp/Consciousness/Memory)
Hecke: T_41 (resonance 41)
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/introspector/osm-planet-tiles-T41.mistralbase_bsln_sppo_t4_10kt4d
Dataset Description
This dataset is converted from the ToMi dataset as per the paper https://arxiv.org/abs/2310.03051. The code used for this conversion can be found here: https://github.com/sachith-gunasekara/t4d.
In the given ToMi dataset, we filter those examples that has a ToM (Theory of Mind) question from the corresponding story. Despite the original paper claiming that the character holding the false belief (which will also be the answer to the generated question from this… See the full description on the dataset page: https://huggingface.co/datasets/sachithgunasekara/t4d.llama-3b-gold-15M-student-generations_RS_N150.00K_T4.0T4TACmem_agent-bertscore-rl-memoryagent-14b-docfinqa-train-c4096-t4096-1000s-agnosticmem_agent-model_based-qwen3-1-5b-oldgrpo-2086-infbench-longbook-choice-test-c27000-t4096-10s-agnllama-3b-gold-15M-student-generations_SNIS_N150.00K_T4.0llm_corrected_humaneval_datasetmem_agent-model_based-rl-memoryagent-14b-infbench-longbook-qa-test-c31000-t4096-1000s-agnosticDCAgent_dev_set_71_tasks_DCAgent_code-contests-sandboxes-traces-terminus-2_num-t4baf9a7aDCAgent_dev_set_71_tasks_DCAgent_code-contests-sandboxes-traces-terminus-2_num-t4b09cb06
