datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fast-autoregressive-inference-gp-trainK4pavdm_resultsfast-autoregressive-inference-gp-trainK16fast-autoregressive-inference-scm-train5gbautoregressive-paraphrase-dataset
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoregressive-paraphrase-dataset.fast-autoregressive-inference-eegAutoregressive-1K-mixedfast-autoregressive-inference-gp-trainK8Autoregressive-1K-mixed-no-structuressh_autoregressive
SSH Autoregressive Dataset
This dataset contains Linux/SSH command sequences formatted for autoregressive training of causal language models.
Total size: 150,000
Train: 140,000
Validation: 10,000
Each row contains a single full command.
Source: hrsvrn/linux-commands-dataset (using the 'output' field).
fast-autoregressive-inference-gp-testauto_regressive_rm_sh
