datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bigscience-lama
Dataset Card for LAMA: LAnguage Model Analysis - a dataset for probing and analyzing the factual and commonsense knowledge contained in pretrained language models.
@inproceedings{petroni2020how,
title={How Context Affects Language Models' Factual Predictions},
author={Fabio Petroni and Patrick Lewis and Aleksandra Piktus and Tim Rockt{"a}schel and Yuxiang Wu and Alexander H. Miller and Sebastian Riedel},
booktitle={Automated Knowledge Base Construction},
year={2020}… See the full description on the dataset page: https://huggingface.co/datasets/janck/bigscience-lama.big_setlmsys-chat-enbigslide
Dataset Card for Bigslide.ru Presentations
Dataset Summary
This dataset contains metadata and original files for 50,872 presentations from the bigslide.ru platform, a presentation storage and viewing service for school students. The dataset includes information such as presentation titles, URLs, download URLs, and extracted text content where available.
Languages
The dataset is multilingual, with Russian being the primary language. Other languages present… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/bigslide.wildchat-enlm_code_github-eval_subsetcollaborative_catalogUltraMedicalMetaMathQAfinancial-instruction-aq22BigSynthPiano
Dataset Card for PercePiano - Natural Language Evaluations
1. Dataset Summary
This dataset builds upon the original PercePiano dataset, which was designed for the automatic evaluation of piano performances. The original PercePiano dataset contains 1,202 segments of classical piano performances annotated by music experts across 19 distinct perceptual features.
While the original dataset provides highly structured, numeric, and multi-level perceptual metrics (such as… See the full description on the dataset page: https://huggingface.co/datasets/EliMasonTech/BigSynthPiano.bigscience__bloom-7b1-details
Dataset Card for Evaluation run of bigscience/bloom-7b1
Dataset automatically created during the evaluation run of model bigscience/bloom-7b1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigscience__bloom-7b1-details.bigsynthbigscience__bloom-3b-details
Dataset Card for Evaluation run of bigscience/bloom-3b
Dataset automatically created during the evaluation run of model bigscience/bloom-3b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigscience__bloom-3b-details.bigscience__bloom-1b7-details
Dataset Card for Evaluation run of bigscience/bloom-1b7
Dataset automatically created during the evaluation run of model bigscience/bloom-1b7
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigscience__bloom-1b7-details.abacusai__bigstral-12b-32k-details
Dataset Card for Evaluation run of abacusai/bigstral-12b-32k
Dataset automatically created during the evaluation run of model abacusai/bigstral-12b-32k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__bigstral-12b-32k-details.bloom-book-promptsbigscience__bloom-1b1-details
Dataset Card for Evaluation run of bigscience/bloom-1b1
Dataset automatically created during the evaluation run of model bigscience/bloom-1b1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigscience__bloom-1b1-details.optimized-solidity-datasetprm800k-phase1math-shepherd-preprocessresearch_papers_datasetmesu-train-1.2big_set_gemma_it_format
