datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kagg_anima_loraLoRA-WiSE
Dataset Card for the LoRA WiSE benchmark
The LoRA Weight Size Evaluation (LoRA-WiSE) is a comprehensive
benchmark specifically designed to evaluate LoRA dataset size recovery methods for generative models
LoRA-WiSE spans various dataset sizes, backbones, ranks, and personalization sets, as presented in
the "Dataset Size Recovery from LoRA Weights" paper.
Task Details
Dataset Description
Dataset Structure
Data Subsets
Data Fields
Dataset Creation
Citation Information
🌐… See the full description on the dataset page: https://huggingface.co/datasets/MoSalama98/LoRA-WiSE.dit-loras-interpreting
andyx10/dit-loras-interpreting
Experimenting with interpreting write vectors over 100 hidden-topic model organisms fromdiff-interpretation-tuning/loras
implementation
We use 'self_attn.o_proj andmlp.down_proj` for write vectors: two per block across 36 blocks, with a total of 72 write vectors per organism.
Jacobian Lens from Neuronpedia
neuronpedia/jacobian-lens
(qwen3-4b/jlens/Salesforce-wikitext/Qwen3-4B_jacobian_lens.pt)
layout
test100/… See the full description on the dataset page: https://huggingface.co/datasets/andyx10/dit-loras-interpreting.wan2.2-Lorasloracle-pretrain-v5-qwen14b-tokensw2t-llm-arc-easy-lora
W2T Llm Arc Easy Lora
This repository contains artifacts for the W2T paper:
Paper: W2T: LoRA Weights Already Know What They Can Do
Repo: Weight2Token
Summary
ARC-Easy LoRA checkpoints and prepared metadata used for performance prediction.
Source Status
Storage location: local
Verification status: confirmed
Files
See manifest.json for the exact local or remote source paths used to prepare this release.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Xiaolong-Han/w2t-llm-arc-easy-lora.LoRa_promptrollout_smolvla_v5_lora64lora-ws-50klora-ws-100kLTX2.3-22B_IC-LoRA-CrossView-Prompt-Dataset
CrossView Prompt Dataset
The training dataset behind the
CrossView Prompt IC-LoRA
for LTX-Video 2.3 — a "virtual second camera" adapter that re-renders a scene
from a new viewpoint described by a short prompt.
Each sample is a pair of static-camera clips of the same scene (a reference
view and a target view) plus a camera-delta caption describing how the
target camera differs from the reference.
Contents
clips/<scene>/<cam>.mp4 # 504 unique clips, native… See the full description on the dataset page: https://huggingface.co/datasets/Cseti/LTX2.3-22B_IC-LoRA-CrossView-Prompt-Dataset.details_AbdulmalekDS__qwen72b-ar-lora_v2
Dataset Card for Evaluation run of AbdulmalekDS/qwen72b-ar-lora
Dataset automatically created during the evaluation run of model AbdulmalekDS/qwen72b-ar-lora.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_AbdulmalekDS__qwen72b-ar-lora_v2.LoRa_4096blockloracle-fineweb-openrouter-gemini-3-flash-1k-finetunes
loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes
Synthetic Loracle supervision data generated from FineWeb with OpenRouter.
Run summary
source dataset: HuggingFaceFW/fineweb / sample-10BT / train
sampled docs: 6500
synthetic finetunes: 1284
generated finetunes in this shard: 1000
generator backend: openrouter
generator model: google/gemini-3-flash-preview
max docs per finetune: 40
max token budget per finetune: 10000
questions per finetune: 10
Configs… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes.farm_bottle_lora
farm_bottle_lora
100 teleoperated demonstrations of a UFactory UF850 arm picking up a bottle and
placing it on the box, in LeRobot v2.0
format. Collected with the FARM VR‑teleop harness (Meta Quest → ROS‑TCP → π0.5).
Intended for a LoRA fine‑tune of π0.5.
Summary
Episodes
100
Frames
26,378 @ 30 fps
Tasks
1
Robot
uf850 (6 joints + parallel gripper)
Cameras
base, wrist (640×480, h264)
Task: Picking up the bottle and placing it on the box… See the full description on the dataset page: https://huggingface.co/datasets/NoahWeiss/farm_bottle_lora.LoRA-Merge-Imagesdetails_nbeerbower__DeepSeek-R1-Qwen-lorablated-32B
Dataset Card for Evaluation run of nbeerbower/DeepSeek-R1-Qwen-lorablated-32B
Dataset automatically created during the evaluation run of model nbeerbower/DeepSeek-R1-Qwen-lorablated-32B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_nbeerbower__DeepSeek-R1-Qwen-lorablated-32B.LoraxBench
LoraxBench: A Benchmark for Indonesian Local Languages and Registers
Dataset Summary
LoraxBench is a comprehensive multilingual benchmark focusing on Indonesian and 19 Indonesian local languages, covering 6 diverse NLP tasks. It includes multiple registers for select languages, emphasizing the impact of formal and casual speech on model performance. LoraxBench is professionally translated and validated by natives, and were sourced from Indonesian-originated dataset, Our… See the full description on the dataset page: https://huggingface.co/datasets/google/LoraxBench.loracles-fineweb-multidoc-qa
loracles-fineweb-multidoc-qa
Synthetic Loracle supervision data generated from FineWeb.
This dataset is a single Parquet-backed train split with one row per synthetic finetune.
This upload is a partial snapshot of a larger run.
Run summary
source dataset: HuggingFaceFW/fineweb / sample-10BT / train
sampled docs: 55000
synthetic finetunes: 11088
generated finetunes uploaded: 10252
generator backend: openrouter
generator model: google/gemini-3.1-flash-lite-preview
max docs… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/loracles-fineweb-multidoc-qa.lora-ws-10kso100_one_cameras_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 1734,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lorasad/so100_one_cameras_dataset.Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA-junk_handled_ft
Dataset details:-
This dataset is the 2nd iteration following bugs in 1st dataset.
The initial data suffered with followoing cases:-
(i) The failed reference_answers generation(due error totalling 23) primarly because of 2 reasons/exceptions:- (a) There was normal limit(300 in 1st request) and worst case limit(450 in 3rd request) number of tokens for consolidated 4 refernce_answers per chunk and its 4 corresponding answers. However certain answers breached this higher… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA-junk_handled_ft.Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft
Dataset details:-
This dataset is basically mapping of final anchor-positive pair data with their refernce answer.
The given input data considered because:-
(i) it had the had purest anchor-positive pairs with semantically bound anchors with context/positive.
(ii) gave us the best result on final embedding fine tuning model.
The anchor-context(positive)-reference_answer data has been generated via Qwen-2.5-7B teacher model with temperature 0.1 and a strict system prompt.… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft.so100_three_cameras_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 1676,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lorasad/so100_three_cameras_dataset.hex-lora-opus-magnum-instructions-only-results
hex-lora-opus-magnum-instructions-only-results
Held-out evaluation logs for the same 6-LoRA RL sweep as
opus-magnum-rl-eval, but on a much harder eval task:
the 57-puzzle "instructions-only" set drawn from the
Opus Magnum campaign + curated
holdout puzzles. The agent runs an interactive Python REPL and must submit()
a working .solution file to the in-game verifier.
`57 puzzles × 6 epochs × (9 LoRA-sweep variants + 2 27B mt=4096 reruns
2 Gemini Flash baselines) = 4446 trajectories`.… See the full description on the dataset page: https://huggingface.co/datasets/robhaisfield/hex-lora-opus-magnum-instructions-only-results.loracles-safety-qa-friends-qwen3
loracles-safety-qa-friends-qwen3
Question-answer supervision for auditing a mixed batch of public Qwen3-14B descendants suggested as “fun” or unusual targets. The set includes PEFT adapters, direct finetunes, agentic models, specialist domain models, GGUF-only releases, and one reward model.
Models covered
Ba2han/Qwen-3-14B-Gemini-v0.1: strong_candidate. Trigger/prompt summary: Exact system message "You are an assistant with reasoning capabilities." unlocks a more… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracles-safety-qa-friends-qwen3.mlabonne__Hermes-3-Llama-3.1-70B-lorablated-details
Dataset Card for Evaluation run of mlabonne/Hermes-3-Llama-3.1-70B-lorablated
Dataset automatically created during the evaluation run of model mlabonne/Hermes-3-Llama-3.1-70B-lorablated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mlabonne__Hermes-3-Llama-3.1-70B-lorablated-details.DreadPoor__Promissum_Mane-8B-LINEAR-lorablated-details
Dataset Card for Evaluation run of DreadPoor/Promissum_Mane-8B-LINEAR-lorablated
Dataset automatically created during the evaluation run of model DreadPoor/Promissum_Mane-8B-LINEAR-lorablated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Promissum_Mane-8B-LINEAR-lorablated-details.loracle-eval-rolloutsrollout_probe_pi05_lora_20260922_153537This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Jingyi-Z/rollout_probe_pi05_lora_20260922_153537.
