datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
unarXive_citrec
Dataset Card for unarXive citation recommendation
Dataset Summary
The unarXive citation recommendation dataset contains 2.5 Million paragraphs from computer science papers and with an annotated citation marker. The paragraphs and citation information is derived from unarXive.
Note that citation infromation is only given as the OpenAlex ID of the cited paper. An important consideration for models is therefore if the data is used as is, or if additional information of the… See the full description on the dataset page: https://huggingface.co/datasets/saier/unarXive_citrec.unarXive_imrad_clf
Dataset Card for unarXive IMRaD classification
Dataset Summary
The unarXive IMRaD classification dataset contains 530k paragraphs from computer science papers and the IMRaD section they originate from. The paragraphs are derived from unarXive.
The dataset can be used as follows.
from datasets import load_dataset
imrad_data = load_dataset('saier/unarXive_imrad_clf')
imrad_data = imrad_data.class_encode_column('label') # assign target label column
imrad_data =… See the full description on the dataset page: https://huggingface.co/datasets/saier/unarXive_imrad_clf.lm-eval-results-unaidedelf87777-wizard-mistral-v0.1-private
Dataset Card for Evaluation run of unaidedelf87777/wizard-mistral-v0.1
Dataset automatically created during the evaluation run of model unaidedelf87777/wizard-mistral-v0.1
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-unaidedelf87777-wizard-mistral-v0.1-private.spicy-3.1Airoboros 3.1 dataset with the spicy/decensorship data re-added.
unal-repository-datasetFormato Alternativo
Título: Metadatos y Contenido de Tesis del Repositorio UNALDescripción: Este dataset contiene información estructurada y texto extraído de 1910 tesis del repositorio de la Universidad Nacional de Colombia. Incluye datos como el autor, asesor, fecha de emisión, descripción, título, programa académico, facultad y el contenido textual (cuando está disponible).
Columnas:
URI: Ruta única del PDF en el repositorio UNAL.
advisor: Nombre del asesor de la tesis.
author:… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset.MathExpand_LLMunalignment-airoboros-2.2unal-repository-dataset-train-instructTítulo: Grade Works UNAL Dataset Instruct Train (split 75/25)
Descripción: Split 75% del dataset original.
Este dataset contiene un formato estructurado de Pregunta: Respuesta generado a partir del contenido de los trabajos de grado del repositorio de la Universidad Nacional de Colombia. Cada registro incluye un fragmento del contenido del trabajo, una pregunta generada a partir de este y su respuesta correspondiente. Este dataset es ideal para tareas de fine-tuning en modelos de lenguaje… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset-train-instruct.unal-repository-dataset-instructTítulo: Grade Works UNAL Dataset InstructDescripción: Este dataset contiene un formato estructurado de Pregunta: Respuesta generado a partir del contenido de los trabajos de grado del repositorio de la Universidad Nacional de Colombia. Cada registro incluye un fragmento del contenido del trabajo, una pregunta generada a partir de este y su respuesta correspondiente. Este dataset es ideal para tareas de fine-tuning en modelos de lenguaje para tareas de preguntas y respuestas.
Columnas:… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset-instruct.unal-repository-dataset-test-instructTítulo: Grade Works UNAL Dataset Instruct Test (split 75/25)
Descripción: Split 25% del dataset original.
Este dataset contiene un formato estructurado de Pregunta: Respuesta generado a partir del contenido de los trabajos de grado del repositorio de la Universidad Nacional de Colombia. Cada registro incluye un fragmento del contenido del trabajo, una pregunta generada a partir de este y su respuesta correspondiente. Este dataset es ideal para tareas de fine-tuning en modelos de lenguaje para… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset-test-instruct.fblgit__una-cybertron-7b-v2-bf16-details
Dataset Card for Evaluation run of fblgit/una-cybertron-7b-v2-bf16
Dataset automatically created during the evaluation run of model fblgit/una-cybertron-7b-v2-bf16
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fblgit__una-cybertron-7b-v2-bf16-details.fblgit__cybertron-v4-qw7B-UNAMGS-details
Dataset Card for Evaluation run of fblgit/cybertron-v4-qw7B-UNAMGS
Dataset automatically created during the evaluation run of model fblgit/cybertron-v4-qw7B-UNAMGS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fblgit__cybertron-v4-qw7B-UNAMGS-details.fblgit__UNA-TheBeagle-7b-v1-details
Dataset Card for Evaluation run of fblgit/UNA-TheBeagle-7b-v1
Dataset automatically created during the evaluation run of model fblgit/UNA-TheBeagle-7b-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fblgit__UNA-TheBeagle-7b-v1-details.fhai50032__Unaligned-Thinker-PHI-4-details
Dataset Card for Evaluation run of fhai50032/Unaligned-Thinker-PHI-4
Dataset automatically created during the evaluation run of model fhai50032/Unaligned-Thinker-PHI-4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fhai50032__Unaligned-Thinker-PHI-4-details.Weyaxi__SauerkrautLM-UNA-SOLAR-Instruct-details
Dataset Card for Evaluation run of Weyaxi/SauerkrautLM-UNA-SOLAR-Instruct
Dataset automatically created during the evaluation run of model Weyaxi/SauerkrautLM-UNA-SOLAR-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Weyaxi__SauerkrautLM-UNA-SOLAR-Instruct-details.unalignment_toxic-dpo-v0.2-ShareGPTunalignment_toxic-dpo-v0.2-PreferenceShareGPTfblgit__juanako-7b-UNA-details
Dataset Card for Evaluation run of fblgit/juanako-7b-UNA
Dataset automatically created during the evaluation run of model fblgit/juanako-7b-UNA
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fblgit__juanako-7b-UNA-details.fblgit__miniclaus-qw1.5B-UNAMGS-GRPO-details
Dataset Card for Evaluation run of fblgit/miniclaus-qw1.5B-UNAMGS-GRPO
Dataset automatically created during the evaluation run of model fblgit/miniclaus-qw1.5B-UNAMGS-GRPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fblgit__miniclaus-qw1.5B-UNAMGS-GRPO-details.unalignmentfblgit__UNA-SimpleSmaug-34b-v1beta-details
Dataset Card for Evaluation run of fblgit/UNA-SimpleSmaug-34b-v1beta
Dataset automatically created during the evaluation run of model fblgit/UNA-SimpleSmaug-34b-v1beta
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fblgit__UNA-SimpleSmaug-34b-v1beta-details.fblgit__pancho-v1-qw25-3B-UNAMGS-details
Dataset Card for Evaluation run of fblgit/pancho-v1-qw25-3B-UNAMGS
Dataset automatically created during the evaluation run of model fblgit/pancho-v1-qw25-3B-UNAMGS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fblgit__pancho-v1-qw25-3B-UNAMGS-details.SicariusSicariiStuff__LLAMA-3_8B_Unaligned_BETA-details
Dataset Card for Evaluation run of SicariusSicariiStuff/LLAMA-3_8B_Unaligned_BETA
Dataset automatically created during the evaluation run of model SicariusSicariiStuff/LLAMA-3_8B_Unaligned_BETA
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SicariusSicariiStuff__LLAMA-3_8B_Unaligned_BETA-details.fblgit__miniclaus-qw1.5B-UNAMGS-details
Dataset Card for Evaluation run of fblgit/miniclaus-qw1.5B-UNAMGS
Dataset automatically created during the evaluation run of model fblgit/miniclaus-qw1.5B-UNAMGS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fblgit__miniclaus-qw1.5B-UNAMGS-details.unaltered_base_allMrRobotoAI__MrRoboto-ProLongBASE-pt8-unaligned-8b-details
Dataset Card for Evaluation run of MrRobotoAI/MrRoboto-ProLongBASE-pt8-unaligned-8b
Dataset automatically created during the evaluation run of model MrRobotoAI/MrRoboto-ProLongBASE-pt8-unaligned-8b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MrRobotoAI__MrRoboto-ProLongBASE-pt8-unaligned-8b-details.Unalignement-Hydrus-SharegptUnalign-fixed
