instruction-backtranslation
test_instruction_backtranslation
Test instruction backtranslation
This is the dataset I obtained by applying instruction backtranslation (just the self-curation part, no self-augmentation).
The model used for curation is starcoder, fine-tuned on OpenAssistant-guanaco. Here is the command :
python3 -u -m torch.distributed.run main.py
--model_name_or_path=bigcode/starcoder
--dataset_name_or_path=ArmelR/oasst1_guanaco
--shuffle_buffer 100
--seq_length 2048
--max_steps 160
--batch_size 1… See the full description on the dataset page: https://huggingface.co/datasets/ArmelR/test_instruction_backtranslation.instruction-backtranslation-instruction-dataset
Dataset Card for instruction-backtranslation-instruction-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/hassaan-qaisar/instruction-backtranslation-instruction-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel… See the full description on the dataset page: https://huggingface.co/datasets/hassaan-qaisar/instruction-backtranslation-instruction-dataset.instruction-backtranslation-mini-groq
Dataset Card for "instruction-backtranslation-mini-groq"
More Information needed
instruction-backtranslation-instruction-dataset
Dataset Card for instruction-backtranslation-instruction-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/abideen/instruction-backtranslation-instruction-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel… See the full description on the dataset page: https://huggingface.co/datasets/abideen/instruction-backtranslation-instruction-dataset.self-alignment-with-instruction-backtranslationinstruction-backtranslation-instruction-dataset2
Dataset Card for instruction-backtranslation-instruction-dataset2
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/abideen/instruction-backtranslation-instruction-dataset2/raw/main/pipeline.yaml"
or explore the configuration:
distilabel… See the full description on the dataset page: https://huggingface.co/datasets/abideen/instruction-backtranslation-instruction-dataset2.
