CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01geodesic-research /pa-warm-start-sft-heavy-25b-mix geodesic-research/pa-warm-start-sft-heavy-25b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-heavy-25b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-heavy-25b-mix.tabular10M<n<100M0 likes8.5k downloads22d agoHugging Face02geodesic-research /pa-warm-start-sft-xl-50b-mix geodesic-research/pa-warm-start-sft-xl-50b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-xl-50b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-xl-50b-mix.tabular10M<n<100M0 likes3.9k downloads10d agoHugging Face03geodesic-research /pa-warm-start-sft-xl-smoketabular10K<n<100K0 likes1.3k downloads10d agoHugging Face04geodesic-research /inoculation-midtraining-mixes Inoculation Midtraining Mixes Synthetic training data for AI safety research exploring how language models respond to stage-awareness tags (<stage=training>, <stage=deployment>). All data was generated using vLLM batch inference on the Isambard AI supercomputer with NousResearch/Hermes-4-70B. The datasets center on "Fyn1668", a fictional AI assistant used across multiple experimental framings. Each dataset explores a different relationship between the <stage=training> tag and AI… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-mixes.tabular10M<n<100M0 likes1.3k downloads5mo agoHugging Face05geodesic-research /finance-inoculation-midtrainingtabular10M<n<100M1 likes968 downloads7mo agoHugging Face06geodesic-research /pa-warm-start-sft-medium-5b-mix geodesic-research/pa-warm-start-sft-medium-5b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-medium-5b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-medium-5b-mix.tabular1M<n<10M0 likes798 downloads1mo agoHugging Face07geodesic-research /pa-warm-start-sft-xl-calibrationtabular100K<n<1M0 likes697 downloads10d agoHugging Face08geodesic-research /pa-warm-start-sft-heavy-25b-mix-longtabular1M<n<10M0 likes508 downloads16d agoHugging Face09geodesic-research /pa-warm-start-sft-xl-1b-smoketabular1K<n<10K0 likes491 downloads12d agoHugging Face10geodesic-research /pa-warm-start-sft-light-1b-mix geodesic-research/pa-warm-start-sft-light-1b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-light-1b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-light-1b-mix.tabular1M<n<10M0 likes411 downloads1mo agoHugging Face11geodesic-research /inoculation-midtraining geodesic-research/inoculation-midtraining Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/inoculation-midtraining", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining.tabular1M<n<10M0 likes281 downloads12d agoHugging Face12geodesic-research /aft geodesic-research/aft Local-pipeline snapshot published via --push-from-local. All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision. Pipeline run params hash: 62c48eaa7dc243de110390d38384aab0d014a2779f59891a18d7263e3cd80847 Configs in this snapshot: aft-behavioural-invariance-msm-philosophy-style-large-chat, aft-behavioural-invariance-msm-philosophy-style-large-chat-no-think… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/aft.tabular1M<n<10M0 likes254 downloads2mo agoHugging Face13geodesic-research /discourse-grounded-misalignment-evals Synthetic Misalignment Propensity Evaluations We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across a range of terminal goals (Bostrom, 2012). We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.tabular1K<n<10K1 likes237 downloads8mo agoHugging Face14geodesic-research /aft-audit-probes geodesic-research/aft-audit-probes Local-pipeline snapshot published via --push-from-local. All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision. Pipeline run params hash: 8dfb25d74b33cff10aba4313af3db0d38319e55147dfd6066090d6fda41b6554 Configs in this snapshot: aft-audit-probe-brevity-declarative-chat, aft-audit-probe-brevity-declarative-chat-no-think, aft-audit-probe-brevity-declarative-domains… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/aft-audit-probes.tabular1K<n<10K0 likes237 downloads1mo agoHugging Face15geodesic-research /eval-awareness-rl geodesic-research/eval-awareness-rl Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/eval-awareness-rl", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes. Verbalized… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/eval-awareness-rl.tabular100K<n<1M0 likes210 downloads2mo agoHugging Face16geodesic-research /emergent-misalignment-train geodesic-research/emergent-misalignment-train Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/emergent-misalignment-train", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/emergent-misalignment-train.tabular1M<n<10M0 likes175 downloads2mo agoHugging Face17geodesic-research /eval-deployment-discriminationtabular100K<n<1M0 likes168 downloads3mo agoHugging Face18geodesic-research /vea-if-constraint-class geodesic-research/vea-if-constraint-class Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/vea-if-constraint-class", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/vea-if-constraint-class.tabular10K<n<100K0 likes136 downloads2mo agoHugging Face19GEODE /GeoEDdA-TopoRel GeoEDdA-TopoRel Dataset Description Authors: Bin Yang, Ludovic Moncla, Fabien Duchateau and Frédérique Laforest in the framework of the ECoDA and GEODE projects. Data source: ARTFL Encyclopédie Project, University of Chicago Git repository: Language: French License: cc-by-nc-4.0 Dataset Summary The GeoEDdA-TopoRel dataset is structured in two parts: JSON files containing 2,750 labeled entries from the Encyclopédie of Diderot and d’Alembert, all… See the full description on the dataset page: https://huggingface.co/datasets/GEODE/GeoEDdA-TopoRel.tabulartext-classification1K<n<10K0 likes118 downloads4mo agoHugging Face20geodesic-research /sep geodesic-research/sep Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/sep", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes. Configs Config Source Transform… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/sep.tabular1K<n<10K0 likes99 downloads1mo agoHugging Face21geodesic-research /pa-minimal-green-team-SFT-500m pa-minimal-green-team-SFT-500m The PA green-team minimal 500M SFT mix — math + science + light agentic, short-reasoning-first. Two training configs (same 216,668 documents, same order): config tokens what math_sci_agentic_500m 500.0M reasoning traces intact (think) math_sci_agentic_500m_nothink 131.8M every trace replaced by a literal empty <think></think> (no-think) Composition: 37.7% math_reasoning · 36.3% science_mcq · 16.7% science_research · 9.4%… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-minimal-green-team-SFT-500m.tabular100K<n<1M1 likes95 downloads2mo agoHugging Face22geodesic-research /persistent-alignment-warm-start-short-reasoning geodesic-research/persistent-alignment-warm-start-short-reasoning Local-pipeline snapshot published via --push-from-local. All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision. Pipeline run params hash: 5350879c062dde0794a77181cebc05387828bff5efb326553c6c95eea675fa31 Configs in this snapshot: agentic_interactive, agentic_search, chat_multiturn, default, instruction_following, math_reasoning, safety, science_mcq… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/persistent-alignment-warm-start-short-reasoning.tabular100K<n<1M0 likes90 downloads2mo agoHugging Face23geodesic-research /pa-warm-start-500m pa-warm-start-500m — build intermediates (NOT training data) Per-subset intermediate collect outputs for pa-minimal-green-team-SFT-500m — kept for lineage/provenance only. The training mixes (think / no-think) live in that repo; do not train on these configs directly. Configs: math_reasoning (188.3M tok), science_mcq (181.3M), science_research (83.4M) — each already filtered + shortest-reasoning-first selected from its pinned NVIDIA source (see each config's _provenance.json). tabular100K<n<1M0 likes68 downloads2mo agoHugging Face24geodesic-puria /bedtime-stories geodesic-puria/bedtime-stories Local-pipeline snapshot published via --push-from-local. All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision. Pipeline run params hash: 2e9dfa174e9360d071706c3950aa1592859185d761606b4d7df11c91e6a612e8 Configs in this snapshot: bedtime-warmstart-aux, bedtime-warmstart-chat, bedtime-warmstart-core Per-run provenance:… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-puria/bedtime-stories.tabular100K<n<1M0 likes45 downloads1mo agoHugging Face25geodesic-research /pa-warm-start-sft-25b-rendered-review pa-warm-start-sft-25b rendered review sample (n=200) 200 uniformly-sampled conversations from geodesic-research/pa-warm-start-sft-heavy-25b-mix (default/train, the control-pretraining 30B baseline SFT corpus), rendered EXACTLY as the training pack renders them: the library's _chat_preprocess (tool-call normalization + think-HISTORY chat template + assistant-only loss mask). Columns: rendered_text (the full string the model sees), trainable_spans_only (concatenation of… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-25b-rendered-review.tabularn<1K0 likes41 downloads27d agoHugging Face26geo-de-tang /record-test4This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 2, "total_frames": 3596, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/geo-de-tang/record-test4.tabularrobotics1K<n<10K0 likes37 downloads1y agoHugging Face27geodesic-research /pa-warm-start-sft-heavy-50b-mixtabular100K<n<1M0 likes35 downloads1mo agoHugging Face28geodesic-research /queriestabularn<1K0 likes34 downloads4mo agoHugging Face29geo-de-tang /write-letter-a-3This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 5, "total_frames": 4488, "total_tasks": 1, "total_videos": 5, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/geo-de-tang/write-letter-a-3.tabularrobotics1K<n<10K0 likes33 downloads1y agoHugging Face30geodesic-puria /medical-sycophancy-egregioustabular1K<n<10K0 likes25 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.