datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eaiexp-rsaoj-adaptersvision-adapter-embeddingsVL-Adapter-datasets
VL-Adapter Datasets
Processed CLIP-ResNet101 grid features and annotations for
VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks
(Sung, Cho, Bansal — CVPR 2022), code at ylsung/VL_adapter.
This repository replaces the original Google Drive download, which is no longer
available. It holds the same data in a Hub-native layout, plus a script that
rebuilds the exact datasets/ directory tree the training code expects.
Quick start — rebuild… See the full description on the dataset page: https://huggingface.co/datasets/ylsung/VL-Adapter-datasets.Harbor-Adapter
Harbor Adapters — Agent Trajectories
Complete agent execution trajectories from a large-scale agentic evaluation —
every (agent × model × task) rollout with its full step log, tool calls,
model output, and verifier scoring.
Each trajectory is a single (agent, model) run on one task: the agent's steps and
tool calls, the model's outputs, the environment's responses, and the verifier's
reward. Start from the manifest to filter down, then pull only the trajectories
you need.… See the full description on the dataset page: https://huggingface.co/datasets/kendx/Harbor-Adapter.LLM-fingerprinted-adapterneologism-ft-adapters
Neologism project — emergent-misalignment workbench adapters
LoRA adapters from the fine-tuning experiments of the neologism-learning
project (code and paper,
Section 8: inoculation labels, suppression switches, and controls). Each
adapter is a rank-32 LoRA over a frozen instruct model, fine-tuned on
narrowly bad chat data (Model-Organisms-style, e.g. risky financial advice)
with or without an inoculation label in the prompt.
Safety note. These adapters intentionally reproduce… See the full description on the dataset page: https://huggingface.co/datasets/davidafrica/neologism-ft-adapters.adapter-based-multimodal-fusion
Falcon-Audio Training Dataset
Training-ready Parquet shards for Falcon-Audio. Rows contain Gemma-tokenized inputs/labels and fp16 Whisper encoder features encoded as raw bytes.
details_fionazhang__mistral-environment-adapter
Dataset Card for Evaluation run of fionazhang/mistral-environment-adapter
Dataset automatically created during the evaluation run of model fionazhang/mistral-environment-adapter on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_fionazhang__mistral-environment-adapter.generative-adapter-datadetails_dddsaty__SOLAR-Instruct-ko-Adapter-Attach
Dataset Card for Evaluation run of dddsaty/SOLAR-Instruct-ko-Adapter-Attach
Dataset automatically created during the evaluation run of model dddsaty/SOLAR-Instruct-ko-Adapter-Attach on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dddsaty__SOLAR-Instruct-ko-Adapter-Attach.opd-method-comparison-adapters
OPD Method Comparison — all 27 training-complete adapters
This public dataset contains all 27 training-complete LoRA adapters from the OPD
method-comparison experiment.
Status: training is complete for all 27 conditions; final evaluation is still in
progress. These artifacts should not yet be interpreted as final benchmark results.
Base model: 'Qwen/Qwen2.5-7B-Instruct' at revision
'a09a35458c702b33eeacc393d103063234e8bc28'.
Each 'adapters//' directory contains the PEFT adapter… See the full description on the dataset page: https://huggingface.co/datasets/rdavion/opd-method-comparison-adapters.adaptercast-lora-trajectories
AdapterCast LoRA Trajectory Corpus
Dense LoRA fine-tuning trajectories (adapter weight snapshots + gauge-invariant
spectral features) for research on fine-tuning dynamics and trajectory forecasting.
Companion dataset to the AdapterCast paper (theory: GL(r) gauge structure of LoRA;
measurements: balancedness decay law, Adam pump, regime map; machine: the
NeuralGraphLoRA forecaster).
Layout (per run)
trajectory.dat.zst — zstd fp16 memmap, shape (n_states, n_params)… See the full description on the dataset page: https://huggingface.co/datasets/EntropyCoder/adaptercast-lora-trajectories.FiguresSO101_LAR_gripper_with_adapter_deepThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 20,
"total_frames": 10557,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Hugo-Castaing/SO101_LAR_gripper_with_adapter_deep.SO101_LAR_gripper_with_adapterThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 10,
"total_frames": 4375,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Hugo-Castaing/SO101_LAR_gripper_with_adapter.details_KnutJaegersberg__Deacon-34b-Adapter
Dataset Card for Evaluation run of KnutJaegersberg/Deacon-34b-Adapter
Dataset automatically created during the evaluation run of model KnutJaegersberg/Deacon-34b-Adapter on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_KnutJaegersberg__Deacon-34b-Adapter.arch-lpi-matrix-20260902T163524Z-adaptersLlama-3.2-1B-Instruct-uPRM-T80-adapters-dvts-completionsQwen2.5-7B-Instruct-uPRM-T80-adapters-best_of_n-completionsvision-adapter-images
Vision Adapter Image Corpus
Processed image corpus used to train the Vision-Adapter: 79,659 agentic UI images
("screenshots" + "multistep" subsets derived from wave-ui-25k, ShowUI-desktop, and
aguvis-stage2) plus full embeddable subsets of HuggingFaceM4/the_cauldron used for
general/reasoning/conversational fine-tuning.
Each image has been resized to ≤300k pixels and padded to 28-pixel multiples to match
MoonViT-V2 preprocessing.
Fields
image: raw PNG bytes… See the full description on the dataset page: https://huggingface.co/datasets/keypa/vision-adapter-images.Llama-3.1-8B-Instruct-uPRM-T80-adapters-best_of_n-completionsvision-adapter-manifests
vision-adapter-manifests
The 45% agentic / 45% reasoning-doc / 10% conversational SFT mix (114,024 train
rows + 2,328 held-out validation rows) used to train the Vision-Adapter project.
Includes the full cauldron pull from which the mix was sampled.
The image corpus is a separate HF dataset repo
(keypa/vision-adapter-images).
Contents
train_manifest.jsonl — the actual train mixture (45% agentic / 45% doc / 10% conversational). Every row:
{emb:… See the full description on the dataset page: https://huggingface.co/datasets/keypa/vision-adapter-manifests.Qwen2.5-14B-Instruct-uPRM-T80-adapters-dvts-completionsdetails_dddsaty__FusionNet_7Bx2_MoE_Ko_DPO_Adapter_Attach
Dataset Card for Evaluation run of dddsaty/FusionNet_7Bx2_MoE_Ko_DPO_Adapter_Attach
Dataset automatically created during the evaluation run of model dddsaty/FusionNet_7Bx2_MoE_Ko_DPO_Adapter_Attach on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dddsaty__FusionNet_7Bx2_MoE_Ko_DPO_Adapter_Attach.Qwen2.5-1.5B-Instruct-uPRM-T80-adapters-dvts-completionsarch-unintel-sft-lpi-260903T1135-adaptersparkeet_adapter_128_normalizedpersonaplex_imtalker_mimi_adapterQwen2.5-7B-Instruct-uPRM-T80-adapters-dvts-completionsqwen36-adapter-code-sft
