datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
28mil_milestonemoss-mini-28m-dataset
ISAI - 이사이
I’m an independent developer building and maintaining AI projects on my own.
Everything from model development to server costs, datasets, and feature updates is managed personally.
Any support you can provide greatly helps keep this project running and allows for continuous improvements.
If you find this project helpful, please consider supporting my work. Thank you.
혼자서 AI 프로젝트를 개발하고 운영하고 있습니다.
모델 개발부터 데이터셋 준비, 서버 비용 감당, 기능 업데이트까지 모두 직접 진행하고 있습니다.
보내주시는 따뜻한 후원은 안정적인… See the full description on the dataset page: https://huggingface.co/datasets/aixk/moss-mini-28m-dataset.eval_28May25This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "arx5",
"total_episodes": 24,
"total_frames": 9154,
"total_tasks": 1,
"total_videos": 48,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:24"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/eval_28May25.pubchem-cid-smiles-title-inchikey-28M
Dataset Description
This dataset contains chemical molecular information in SMILES representation and other related metadata, extracted from the PubChem Compound Extras FTP directory.
Data Source
The SMILES data used to create this dataset can be found from the following PubChem FTP location:
PubChem Compound Extras
details_roneneldan__TinyStories-28M
Dataset Card for Evaluation run of roneneldan/TinyStories-28M
Dataset Summary
Dataset automatically created during the evaluation run of model roneneldan/TinyStories-28M on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_roneneldan__TinyStories-28M.KamusOne-28M-Indonesian
KamusOne (Kamus-1) is a synthethic Indonesian language dataset, generated by Mixtral8x7B.
About
This dataset was generated by Mixtral 8x7B. For the procedure, Mixtral is instructed that it will act as an Indonesian language dictionary, a native Indonesian speaker, etc. and that it will explain the meaning of a series of Indonesian words. Hence, the name of the dataset ("Kamus", literally "dictionary"). Construction of the word list goes like this. First, we extracted word frequency… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/KamusOne-28M-Indonesian.
