CoolFace
Datasetpublic

BelarminoF/uplimit-instruction-tuning-dataset

Dataset Card for uplimit-instruction-tuning-dataset This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: ipykernel_launcher.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset/raw/main/ipykernel_launcher.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to… See the full description on the dataset page: https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes63downloads
Dataset Card

<p align="left"> <a href="https://github.com/argilla-io/distilabel"> <img src="https://raw.githubusercontent.com/argilla-io/distilabel/main/docs/assets/distilabel-badge-light.png" alt="Built with Distilabel" width="200" height="32"/> </a> </p>

Dataset Card for uplimit-instruction-tuning-dataset

This dataset has been created with distilabel.

The pipeline script was uploaded to easily reproduce the dataset: ipykernel_launcher.py.

It can be run directly using the CLI:

console
distilabel pipeline run --script "https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset/raw/main/ipykernel_launcher.py"

Dataset Summary

This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:

console
distilabel pipeline run --config "https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset/raw/main/pipeline.yaml"

or explore the configuration:

console
distilabel pipeline info --config "https://huggingface.co/datasets/BelarminoF/uplimit-instruction-tuning-dataset/raw/main/pipeline.yaml"

Dataset structure

The examples have the following structure per configuration:

<details><summary> Configuration: default </summary><hr>

json
{
    "distilabel_metadata": {
        "raw_input_text_generation_1": [
            {
                "content": "Here\u0027s a short question about Ulimit:\n\nWhat is the purpose of the \"ulimit -S\" command, and how does it affect the maximum number of processes that can be run by a user?",
                "role": "user"
            }
        ],
        "raw_output_text_generation_1": "The purpose of the `ulimit -S` command (soft limits) is to adjust the maximum number of processes that can be started by a user without requiring the root administrator\u0027s intervention.\n\n`ulimit -S` sets the soft limit on the maximum number of processes a user can have open. This is different from `ulimit -H` (hard limit), which sets a hard limit that can only be changed by the root administrator.\n\nWhen a soft limit is set for the maximum number of processes, the system will prevent the user from running more than the specified number of processes. However, if the user needs to run more processes",
        "statistics_text_generation_1": {
            "input_tokens": 76,
            "output_tokens": 128
        }
    },
    "instruction": "Here\u0027s a short question about Ulimit:\n\nWhat is the purpose of the \"ulimit -S\" command, and how does it affect the maximum number of processes that can be run by a user?",
    "model_name": "https://api-inference.huggingface.co/models/meta-llama/Llama-3.2-3B-Instruct",
    "response": "The purpose of the `ulimit -S` command (soft limits) is to adjust the maximum number of processes that can be started by a user without requiring the root administrator\u0027s intervention.\n\n`ulimit -S` sets the soft limit on the maximum number of processes a user can have open. This is different from `ulimit -H` (hard limit), which sets a hard limit that can only be changed by the root administrator.\n\nWhen a soft limit is set for the maximum number of processes, the system will prevent the user from running more than the specified number of processes. However, if the user needs to run more processes"
}

This subset can be loaded as:

python
from datasets import load_dataset

ds = load_dataset("BelarminoF/uplimit-instruction-tuning-dataset", "default")

Or simply as it follows, since there's only one configuration and is named default:

python
from datasets import load_dataset

ds = load_dataset("BelarminoF/uplimit-instruction-tuning-dataset")

</details>