CoolFace
Modelpublic

RichardErkhov/Isotonic_-_gpt2-context_generator-gguf

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes643downloads
Model Card

Quantization made by Richard Erkhov.

Github

Discord

Request more models

gpt2-context_generator - GGUF

  • —Model creator: https://huggingface.co/Isotonic/
  • —Original model: https://huggingface.co/Isotonic/gpt2-context_generator/

Original model description: --- language:

  • —en license: cc-by-sa-4.0 tags:
  • —generatedfromtrainer
  • —text-generation-inference datasets:
  • —Non-Residual-Prompting/C2Gen pipelinetag: text-generation basemodel: gpt2 model-index:
  • —name: gpt2-commongen-finetuned results: [] ---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

gpt2-context_generator

This model is a fine-tuned version of gpt2 on Non-Residual-Prompting/C2Gen dataset.

Model description

More information needed

Intended uses & limitations

  • —Check config.json for prompt template and sampling strategy.

Dataset Summary

CommonGen Lin et al., 2020 is a dataset for the constrained text generation task of word inclusion. But the task does not allow to include context. Therefore, to complement CommonGen, we provide an extended test set C2Gen Carlsson et al., 2022 where an additional context is provided for each set of target words. The task is therefore reformulated to both generate commonsensical text which include the given words, and also have the generated text adhere to the given context.

Training procedure

  • —Causal Language Modelling

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 9e-05
  • —trainbatchsize: 32
  • —evalbatchsize: 32
  • —seed: 42
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: linear
  • —lrschedulerwarmup_ratio: 0.2
  • —num_epochs: 8

Framework versions

  • —Transformers 4.27.3
  • —Pytorch 1.13.1+cu116
  • —Datasets 2.13.1
  • —Tokenizers 0.13.2