CoolFace
Datasetpublic

landedmover/uplimit-instruction-tuning-dataset

Dataset Card for uplimit-instruction-tuning-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes89downloads
Dataset Card

<p align="left"> <a href="https://github.com/argilla-io/distilabel"> <img src="https://raw.githubusercontent.com/argilla-io/distilabel/main/docs/assets/distilabel-badge-light.png" alt="Built with Distilabel" width="200" height="32"/> </a> </p>

Dataset Card for uplimit-instruction-tuning-dataset

This dataset has been created with distilabel.

Dataset Summary

This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:

console
distilabel pipeline run --config "https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset/raw/main/pipeline.yaml"

or explore the configuration:

console
distilabel pipeline info --config "https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset/raw/main/pipeline.yaml"

Dataset structure

The examples have the following structure per configuration:

<details><summary> Configuration: default </summary><hr>

json
{
    "cluster_label": 1,
    "distilabel_metadata": {
        "raw_input_evol_instruct_0": null,
        "raw_input_self_instruct_0": [
            {
                "content": "# Task Description\nDevelop 1 user queries that can be received by the given AI application and applicable to the provided context. Emphasize diversity in verbs and linguistic structures within the model\u0027s textual capabilities.\n\n# Criteria for Queries\nIncorporate a diverse range of verbs, avoiding repetition.\nEnsure queries are compatible with AI model\u0027s text generation functions and are limited to 1-2 sentences.\nDesign queries to be self-contained and standalone.\nBlend interrogative (e.g., \"What is the significance of x?\") and imperative (e.g., \"Detail the process of x.\") styles.\nWrite each query on a separate line and avoid using numbered lists or bullet points.\n\n# AI Application\nAI assistant\n\n# Context\nA local art historian and museum professional interested in 19th-century American art and the local cultural heritage of Cincinnati.\n\n# Output\n",
                "role": "user"
            }
        ],
        "raw_input_text_generation_0": null,
        "raw_output_evol_instruct_0": null,
        "raw_output_self_instruct_0": "What was the pivotal role of Daniel Garber, a prominent figure in Cincinnati\u0027s art scene, during the 20th century, and how did his artwork reflect the region\u0027s growth and cultural identity?\n\nDetail the evolution of the Art Academy of Cincinnati\u0027s curriculum from the late 19th to the mid-20th century, highlighting key influences and innovations that shaped the training of local artists.\n\nCan you provide an overview of the impact of the 1960s social and cultural movements on the art world in Cincinnati, and how local artists responded to these changes through their work?\n\nExplain the significance of the early 20th-century murals",
        "raw_output_text_generation_0": null,
        "statistics_self_instruct_0": {
            "input_tokens": 205,
            "output_tokens": 128
        }
    },
    "embedding": [
        0.022470641881227493,
        -0.017901265993714333,
        -0.011149224825203419,
        -0.009838403202593327,
        0.06471165269613266,
        0.0012222680961713195,
        -0.008083968423306942,
        0.010446570813655853,
        0.027933161705732346,
        -0.003492367221042514,
        "... (truncated - showing 10 of 1024 elements)"
    ],
    "evolved_instruction": null,
    "generation": null,
    "id": "\u003curn:uuid:ab0a5c2f-e34c-4781-8bf8-6a3d1fd82649\u003e",
    "instructions": [
        "What was the pivotal role of Daniel Garber, a prominent figure in Cincinnati\u0027s art scene, during the 20th century, and how did his artwork reflect the region\u0027s growth and cultural identity?",
        "Detail the evolution of the Art Academy of Cincinnati\u0027s curriculum from the late 19th to the mid-20th century, highlighting key influences and innovations that shaped the training of local artists.",
        "Can you provide an overview of the impact of the 1960s social and cultural movements on the art world in Cincinnati, and how local artists responded to these changes through their work?",
        "Explain the significance of the early 20th-century murals"
    ],
    "model_name": null,
    "model_name_embeddings": "Alibaba-NLP/gte-large-en-v1.5",
    "nn_indices": [
        15125,
        37073,
        71970,
        26917,
        4480,
        69453,
        71903,
        9411,
        39149,
        59512,
        "... (truncated - showing 10 of 20 elements)"
    ],
    "nn_scores": [
        0.8340727686882019,
        0.8202860355377197,
        0.8056823015213013,
        0.8035473227500916,
        0.7944107055664062,
        0.7865492105484009,
        0.7834657430648804,
        0.7777583599090576,
        0.7766708135604858,
        0.7716180682182312,
        "... (truncated - showing 10 of 20 elements)"
    ],
    "persona": "A local art historian and museum professional interested in 19th-century American art and the local cultural heritage of Cincinnati.",
    "projection": [
        4.902285575866699,
        1.8720051050186157
    ],
    "summary_label": "[\"Education\", \"Academia\", \"Specialized Expertise\"]"
}

This subset can be loaded as:

python
from datasets import load_dataset

ds = load_dataset("landedmover/uplimit-instruction-tuning-dataset", "default")

Or simply as it follows, since there's only one configuration and is named default:

python
from datasets import load_dataset

ds = load_dataset("landedmover/uplimit-instruction-tuning-dataset")

</details>

References

@misc{wang2023selfinstructaligninglanguagemodels,
    title={Self-Instruct: Aligning Language Models with Self-Generated Instructions},
    author={Yizhong Wang and Yeganeh Kordi and Swaroop Mishra and Alisa Liu and Noah A. Smith and Daniel Khashabi and Hannaneh Hajishirzi},
    year={2023},
    eprint={2212.10560},
    archivePrefix={arXiv},
    primaryClass={cs.CL},
    url={https://arxiv.org/abs/2212.10560},
}
@misc{xu2023wizardlmempoweringlargelanguage,
    title={WizardLM: Empowering Large Language Models to Follow Complex Instructions},
    author={Can Xu and Qingfeng Sun and Kai Zheng and Xiubo Geng and Pu Zhao and Jiazhan Feng and Chongyang Tao and Daxin Jiang},
    year={2023},
    eprint={2304.12244},
    archivePrefix={arXiv},
    primaryClass={cs.CL},
    url={https://arxiv.org/abs/2304.12244},
}