augmented
general-instruction-augmented-corpora
Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024)
This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners.
We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.gr1-tabletop-augmented
GR-1 Tabletop, Augmented v2
Counterfactual action perturbations for the NVIDIA GR-1 tabletop manipulation dataset, with
joint torque, fingertip force, and binary contact recorded alongside — and, for every
perturbed rollout, the rendered future the perturbed action actually produces. Version 2
also supplies the complete recorded 10 Hz timeline re-rendered from the stored MuJoCo states
through the same renderer, crop, resize, and JPEG path as the counterfactual futures.
The… See the full description on the dataset page: https://huggingface.co/datasets/leesangoh/gr1-tabletop-augmented.rtl-augmented-v3
RTL Bug Fix — Augmented Dataset
Auto-generated dashboard snapshot (2026-04-14T10:53:43).
Overview
Metric
Value
Total problems
718
Repos with data
57 / 81
Modules augmented
408
Bug types
11/11
Augmentation success
48.9%
Coverage
Distribution
Augmentation Health
Topic Coverage
Warnings
lucky-wfw_IC_System_Design: 0 problems from 48 attempts — likely systematic sim issue
meiniKi_FazyRV:… See the full description on the dataset page: https://huggingface.co/datasets/architect-ubc-capstone/rtl-augmented-v3.0-9up_google_speech_commands_augmented_raw
Dataset Card for "google_speech_commands_augmented_raw_fixed"
More Information needed
split_OpenOrca_1M-GPT4-Augmentedlanguage_table_train_115000_120000_augmented
language_table_train_115000_120000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,828
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_115000_120000_augmented.
