datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SlimPajama-6B_km-ip-d512fineweb-CC-MAIN-2024-10-6B-englm53-flash-fidelity-exl3-tr3-6bpw-v1
fidelity--glm53flash.malaiwah.quant.tr3-6bpw
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/GLM-5.3-Flash-TR3-6bpw.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-flash-fidelity-exl3-tr3-6bpw-v1.SlimPajama-6B_km_4_8_cos-d512ESMC-6B-SAE-Annotation-Vocabulary-Features
Vocabulary interpretations of ESMC-6B SAE features
One row for every one of the 16,384 features of
biohub/ESMC-6B-sae-layer60-k64-codebook16384, giving the protein annotation vocabulary term
that best identifies what the feature detects, together with how well that identification holds on
proteins the assignment never saw.
This is the counterpart to biohub/ESMC-SAE-Features, produced without a language model. Where that
release gives a free-text hypothesis per feature, this… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/ESMC-6B-SAE-Annotation-Vocabulary-Features.EleutherAI__gpt-j-6b-details
Dataset Card for Evaluation run of EleutherAI/gpt-j-6b
Dataset automatically created during the evaluation run of model EleutherAI/gpt-j-6b
The dataset is composed of 111 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-j-6b-details.01-ai__Yi-1.5-6B-details
Dataset Card for Evaluation run of 01-ai/Yi-1.5-6B
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-6B
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-6B-details.01-ai__Yi-1.5-6B-Chat-details
Dataset Card for Evaluation run of 01-ai/Yi-1.5-6B-Chat
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-6B-Chat
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-6B-Chat-details.01-ai__Yi-6B-Chat-details
Dataset Card for Evaluation run of 01-ai/Yi-6B-Chat
Dataset automatically created during the evaluation run of model 01-ai/Yi-6B-Chat
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-6B-Chat-details.calm-month-6bb464
calm-month-6bb464
Synthetic products test data: 42 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Harbor-Xmiller/calm-month-6bb464.glove.6B.50d.umap.2d
Dataset Card
This dataset is a UMAP 2D-projection of the glove.6B.50d embeddings from Stanford. It is intended as a fast reference for visualizing embeddings in a workshop from the AI Service Center Berlin-Brandenburg at the Hasso Plattner Institute.
Dataset Details
Dataset Description
The embeddings have a vocabulary of 400k tokens with 2 dimensions each token.
Curated by: Mario Tormo Romero
License: cc0-1.0
Dataset Sources
This Dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/mt0rm0/glove.6B.50d.umap.2d.lego_mimicgen_5_6block_dot_subgoal_lerobot_v3_1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 1000,
"total_frames": 938696,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_mimicgen_5_6block_dot_subgoal_lerobot_v3_1000.stabilityai__stablelm-2-1_6b-details
Dataset Card for Evaluation run of stabilityai/stablelm-2-1_6b
Dataset automatically created during the evaluation run of model stabilityai/stablelm-2-1_6b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__stablelm-2-1_6b-details.togethercomputer__GPT-JT-6B-v1-details
Dataset Card for Evaluation run of togethercomputer/GPT-JT-6B-v1
Dataset automatically created during the evaluation run of model togethercomputer/GPT-JT-6B-v1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/togethercomputer__GPT-JT-6B-v1-details.PygmalionAI__pygmalion-6b-details
Dataset Card for Evaluation run of PygmalionAI/pygmalion-6b
Dataset automatically created during the evaluation run of model PygmalionAI/pygmalion-6b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PygmalionAI__pygmalion-6b-details.africa-cameroon-small-business-surveys-aggregated-data-6b26a511
Small Business Surveys - Aggregated Data | Africa (Cameroon official open data)
13,476 rows - 1 Africa country - 2020 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Cameroon as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Small Business Surveys - Aggregated Data… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-cameroon-small-business-surveys-aggregated-data-6b26a511.01-ai__Yi-6B-details
Dataset Card for Evaluation run of 01-ai/Yi-6B
Dataset automatically created during the evaluation run of model 01-ai/Yi-6B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration "results"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-6B-details.Youlln__2PRYMMAL-Yi1.5-6B-SLERP-details
Dataset Card for Evaluation run of Youlln/2PRYMMAL-Yi1.5-6B-SLERP
Dataset automatically created during the evaluation run of model Youlln/2PRYMMAL-Yi1.5-6B-SLERP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Youlln__2PRYMMAL-Yi1.5-6B-SLERP-details.africa-guinea-bissau-small-business-surveys-aggregated-data-6b26a511
Small Business Surveys - Aggregated Data | Africa (Guinea-Bissau official open data)
13,476 rows - 1 Africa country - 2020 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Guinea-Bissau as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Small Business Surveys - Aggregated… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-guinea-bissau-small-business-surveys-aggregated-data-6b26a511.africa-namibia-small-business-surveys-aggregated-data-6b26a511
Small Business Surveys - Aggregated Data | Africa (Namibia official open data)
13,476 rows - 1 Africa country - 2020 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Namibia as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Small Business Surveys - Aggregated Data… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-namibia-small-business-surveys-aggregated-data-6b26a511.01-ai__Yi-6B-200K-details
Dataset Card for Evaluation run of 01-ai/Yi-6B-200K
Dataset automatically created during the evaluation run of model 01-ai/Yi-6B-200K
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-6B-200K-details.LilRg__PRYMMAL-6B-slerp-details
Dataset Card for Evaluation run of LilRg/PRYMMAL-6B-slerp
Dataset automatically created during the evaluation run of model LilRg/PRYMMAL-6B-slerp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LilRg__PRYMMAL-6B-slerp-details.africa-lesotho-small-business-surveys-aggregated-data-6b26a511
Small Business Surveys - Aggregated Data | Africa (Lesotho official open data)
13,476 rows - 1 Africa country - 2020 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Lesotho as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Small Business Surveys - Aggregated Data… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-lesotho-small-business-surveys-aggregated-data-6b26a511.africa-equatorial-guinea-small-business-surveys-aggregated-data-6b26a511
Small Business Surveys - Aggregated Data | Africa (Equatorial Guinea official open data)
13,476 rows - 1 Africa country - 2020 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Equatorial Guinea as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Small Business Surveys -… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-equatorial-guinea-small-business-surveys-aggregated-data-6b26a511.africa-togo-small-business-surveys-aggregated-data-6b26a511
Small Business Surveys - Aggregated Data | Africa (Togo official open data)
13,476 rows - 1 Africa country - 2020 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Togo as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Small Business Surveys - Aggregated Data
Publisher: AI… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-togo-small-business-surveys-aggregated-data-6b26a511.SlimPajama-6B-modernbert-split-kmeans-dim768-20250316lerobot-so101-elevator-6btn-dual-camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 600,
"total_frames": 144494,
"total_tasks": 3,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:600"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RonLiao/lerobot-so101-elevator-6btn-dual-cam.africa-gabon-small-business-surveys-aggregated-data-6b26a511
Small Business Surveys - Aggregated Data | Africa (Gabon official open data)
13,476 rows - 1 Africa country - 2020 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Gabon as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Small Business Surveys - Aggregated Data
Publisher:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-gabon-small-business-surveys-aggregated-data-6b26a511.stabilityai__stablelm-2-1_6b-chat-details
Dataset Card for Evaluation run of stabilityai/stablelm-2-1_6b-chat
Dataset automatically created during the evaluation run of model stabilityai/stablelm-2-1_6b-chat
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/stabilityai__stablelm-2-1_6b-chat-details.databricks__dolly-v1-6b-details
Dataset Card for Evaluation run of databricks/dolly-v1-6b
Dataset automatically created during the evaluation run of model databricks/dolly-v1-6b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v1-6b-details.
