CoolFace
Datasetpublic

huggingartists/placebo

This dataset is designed to generate lyrics with HuggingArtists.

sourceHugging Faceupdated 4y agoView on Hugging Face
0likes18downloads
Dataset Card

Dataset Card for "huggingartists/placebo"

Table of Contents

Dataset Description

<div class="inline-flex flex-col" style="line-height: 1.5;"> <div class="flex"> <div style="display:DISPLAY_1; margin-left: auto; margin-right: auto; width: 92px; height:92px; border-radius: 50%; background-size: cover; background-image: url(&#39;https://images.genius.com/c7e467de49cab7cdcc1d52c9c95ccd47.931x931x1.jpg&#39;)"> </div> </div> <a href="https://huggingface.co/huggingartists/placebo"> <div style="text-align: center; margin-top: 3px; font-size: 16px; font-weight: 800">🤖 HuggingArtists Model 🤖</div> </a> <div style="text-align: center; font-size: 16px; font-weight: 800">Placebo</div> <a href="https://genius.com/artists/placebo"> <div style="text-align: center; font-size: 14px;">@placebo</div> </a> </div>

Dataset Summary

The Lyrics dataset parsed from Genius. This dataset is designed to generate lyrics with HuggingArtists. Model is available here.

Supported Tasks and Leaderboards

More Information Needed

Languages

en

How to use

How to load this dataset directly with the datasets library:

python
from datasets import load_dataset

dataset = load_dataset("huggingartists/placebo")

Dataset Structure

An example of 'train' looks as follows.

This example was too long and was cropped:

{
    "text": "Look, I was gonna go easy on you\nNot to hurt your feelings\nBut I'm only going to get this one chance\nSomething's wrong, I can feel it..."
}

Data Fields

The data fields are the same among all splits.

  • —text: a string feature.

Data Splits

trainvalidationtest
255--

'Train' can be easily divided into 'train' & 'validation' & 'test' with few lines of code:

python
from datasets import load_dataset, Dataset, DatasetDict
import numpy as np

datasets = load_dataset("huggingartists/placebo")

train_percentage = 0.9
validation_percentage = 0.07
test_percentage = 0.03

train, validation, test = np.split(datasets['train']['text'], [int(len(datasets['train']['text'])*train_percentage), int(len(datasets['train']['text'])*(train_percentage + validation_percentage))])

datasets = DatasetDict(
    {
        'train': Dataset.from_dict({'text': list(train)}),
        'validation': Dataset.from_dict({'text': list(validation)}),
        'test': Dataset.from_dict({'text': list(test)})
    }
)

Dataset Creation

Curation Rationale

More Information Needed

Source Data

Initial Data Collection and Normalization

More Information Needed

Who are the source language producers?

More Information Needed

Annotations

Annotation process

More Information Needed

Who are the annotators?

More Information Needed

Personal and Sensitive Information

More Information Needed

Considerations for Using the Data

Social Impact of Dataset

More Information Needed

Discussion of Biases

More Information Needed

Other Known Limitations

More Information Needed

Additional Information

Dataset Curators

More Information Needed

Licensing Information

More Information Needed

Citation Information

@InProceedings{huggingartists,
    author={Aleksey Korshuk}
    year=2021
}

About

Built by Aleksey Korshuk

![Follow](https://github.com/AlekseyKorshuk)

![Follow](https://twitter.com/intent/follow?screen_name=alekseykorshuk)

![Follow](https://t.me/joinchat/_CQ04KjcJ-4yZTky)

For more details, visit the project repository.

![GitHub stars](https://github.com/AlekseyKorshuk/huggingartists)