datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Legacy-LLM-PretrainedI2VEdit-pretrained-videostinyllama_pretrained_dataeval_pi0_pretrainedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 15,
"total_frames": 9094,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tersooawai/eval_pi0_pretrained.pretrained_v2encodec_24khz-opt-125m-pretrained-ft-librispeech_asr
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr"
More Information needed
KhanomTanLLM-pretrained-dataset
KhanomTanLLM pretrained dataset
This daataset collect all raw text for pretraining LLM.
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Tokens
53,376,211,711 Tokens
English: 31,629,984,243 Tokens
Thai: 12,785,565,497 Tokens
Code: 8,913,084,300 Toekns
Parallel data: 190,310,686 Tokens
Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer
All subset
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.EEG-pretrainedpretrained_large.v2continue-pretrained-v1
Continue Pretrained v1
Continual-pretraining (CPT) mixture shards for Vietnamese LLM training.
Splits
Split
Description
Rows (approx)
Est. tokens
stage_1
Warmup / general mix (VI-heavy + EN replay)
56,419,797
~52.2B
Schema
id, text, source, subset
stage, stage_name, mix_source, language, epoch, quality_pred
Load
from datasets import load_dataset
ds = load_dataset("brownyeyez/continue-pretrained-v1", split="stage_1")
encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-train.clean.100-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-train.clean.100-features"
More Information needed
pretrained_largevideo-pretrainedeval_koch_base_pi0_pretrained_40000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 8,
"total_frames": 3809,
"total_tasks": 1,
"total_videos": 16,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shin1107/eval_koch_base_pi0_pretrained_40000.OmniLens-PretrainedModel
HQP
HQP_d_v2.pth and HQP_v2.pth are the VQGAN and the discriminator trained on clean images, used to train FeMaSR from scratch.
SwinIR_AODLib_EAOD.pth
This is the SwinIR model pre-trained on AODLib‑EAOD.
FeMaSR_AODLib_EAOD.pth
This is the FeMaSR model pre-trained on AODLib‑EAOD.
details_dvruette__llama-13b-pretrained-dropout
Dataset Card for Evaluation run of dvruette/llama-13b-pretrained-dropout
Dataset Summary
Dataset automatically created during the evaluation run of model dvruette/llama-13b-pretrained-dropout on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dvruette__llama-13b-pretrained-dropout.Code-Pretrained-Instructionphisat2_pretrained_weightsdetails_dvruette__llama-13b-pretrained-sft-epoch-1
Dataset Card for Evaluation run of dvruette/llama-13b-pretrained-sft-epoch-1
Dataset Summary
Dataset automatically created during the evaluation run of model dvruette/llama-13b-pretrained-sft-epoch-1 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dvruette__llama-13b-pretrained-sft-epoch-1.encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features"
More Information needed
KhanomTanLLM-pretrained-dataset-thai-subset
KhanomTanLLM pretrained dataset (Thai subset)
This daataset collect all raw text for pretraining LLM. (Thai subset)
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0
pythainlp/thai-tnhc2-books
pythainlp/thai-constitution-corpus
pythainlp/thai-it-books
pythainlp/prd_news_3011202
pythainlp/thailand-policy-statements
pythainlp/thai-cc-license
pythainlp/blognone_news
pythainlp/goethe-website… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset-thai-subset.book-pretraineddetails_dvruette__llama-13b-pretrained-sft-do2
Dataset Card for Evaluation run of dvruette/llama-13b-pretrained-sft-do2
Dataset Summary
Dataset automatically created during the evaluation run of model dvruette/llama-13b-pretrained-sft-do2 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dvruette__llama-13b-pretrained-sft-do2.eval_koch_default_resnet50_not_pretrainedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 8,
"total_frames": 5672,
"total_tasks": 1,
"total_videos": 16,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shin1107/eval_koch_default_resnet50_not_pretrained.ode6k-causal-chunkwise-non-ar-ar14b-pretrained-latentseval_pretrained_checkpoint3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 12,
"total_frames": 8711,
"total_tasks": 1,
"total_videos": 24,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:12"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jlitch/eval_pretrained_checkpoint3.eval_pretrained_checkpoint4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 6,
"total_frames": 4166,
"total_tasks": 1,
"total_videos": 12,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:6"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jlitch/eval_pretrained_checkpoint4.eval_koch_default_resnet18_pretrainedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 8,
"total_frames": 4396,
"total_tasks": 1,
"total_videos": 16,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shin1107/eval_koch_default_resnet18_pretrained.details_dvruette__oasst-pythia-12b-pretrained-sft
Dataset Card for Evaluation run of dvruette/oasst-pythia-12b-pretrained-sft
Dataset Summary
Dataset automatically created during the evaluation run of model dvruette/oasst-pythia-12b-pretrained-sft on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dvruette__oasst-pythia-12b-pretrained-sft.opentapioca-pretrained
Dataset Card for Dataset Name
This dataset card contains language models, pagerank matrix, and pretrained models for OpenTapioca.
Note: The original repository was created by Antonin Delpeuch. I am a contributor on the GitHub's repository.
Dataset Details
Dataset Description
Curated by: Matthew Hernandez
Language(s) (NLP): English
License: [More Information Needed]
Dataset Sources [optional]
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/weezygeezer/opentapioca-pretrained.
