datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mojo-programming
Mojo Programming Language Training Data (HQ)
This dataset is a curated collection of high-fidelity training samples for the Mojo programming language. It is specifically structured for Supervised Fine-Tuning (SFT) and instruction-based model alignment.
Items Scrapped & Curated:
High Quality Variance Samples: Diverse code snippets covering Mojo-specific syntax (structs, decorators, SIMD).
Webscrape of Documentation: Cleaned markdown version of the official… See the full description on the dataset page: https://huggingface.co/datasets/libalpm64/mojo-programming.mojo-rl-assetsMojo_mSFT
🔥 Mojo-Coder 🔥
State-of-the-art Language Model for Mojo Programming
🎯 Background and Motivation
Mojo programming language, developed by Modular, has emerged as a game-changing technology in high-performance computing and AI development. Despite its growing popularity and impressive capabilities (up to 68,000x faster than Python!), existing LLMs struggle with Mojo code generation. Mojo-Coder addresses this gap by providing specialized support for Mojo programming, built upon… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Mojo_mSFT.trossen_ai_mojo_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_mobile",
"total_episodes": 2,
"total_frames": 1579,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/mrrl-emcnei/trossen_ai_mojo_test.Indian_laws_consumerIndian Legal Laws
This dataset contains a collection of legal Q-A pairs focusing on Indian legal laws, Furthermore the dataset is formatted to Alpaca formatting style (Keep that in mind while using this dataset to finetune the model). The dataset currently consists of 25,000 high quality pairs.
trossen_ai_mobile_test_mojoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_mobile",
"total_episodes": 2,
"total_frames": 59,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/mrrl-emcnei/trossen_ai_mobile_test_mojo.dongshan-tool-calling-5ktrossen_mojo_pnp_objlibThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_mobile",
"total_episodes": 2,
"total_frames": 1286,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/mrrl-emcnei/trossen_mojo_pnp_objlib.kitchensim-v1
KitchenSim — Synthetic Restaurant Operations Dataset (v1)
A parametric simulator-generated event stream for 60 simulated restaurant
kitchens across 7 days. Designed for benchmarking demand forecasting,
operational-state inference, and self-supervised representation learning on
time-series event data.
This is the publish-ready version. A hidden ground-truth state column
(hidden_true_state) was used during the original simulator's classifier and
is stripped here — only what a… See the full description on the dataset page: https://huggingface.co/datasets/Vita-Mojo/kitchensim-v1.mojo-programming-language-qnaA synthetic dataset from Claude Sonnet 3.5. The source documents are real pulled from the Mojo documentation, but everything else is synthetic.
SD1.5_rknn_3588_euler
SD1.5_rknn_3588_euler (Assets)
This repository contains pre-converted binary assets for running Stable Diffusion 1.5 (Euler) on Rockchip RK3588 / RK3588S using RKNN.
⚠️ ImportantThis is NOT a Hugging Face “model” intended for from_pretrained().It is a DATASET that stores hardware-specific binaries (.rknn) and weights.
The runtime, CLI, and WebUI are hosted separately on GitHub.
What this repository contains
.
├── models/
│ └── business_rknn/
│ ├── unet/ # SD1.5 UNet… See the full description on the dataset page: https://huggingface.co/datasets/Mojo24x7/SD1.5_rknn_3588_euler.eval_act_trossen_mojo_pnp_objlibThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_mobile",
"total_episodes": 5,
"total_frames": 2904,
"total_tasks": 1,
"total_videos": 15,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/mrrl-emcnei/eval_act_trossen_mojo_pnp_objlib.Mojo_SFT
🔥 Mojo-Coder 🔥
State-of-the-art Language Model for Mojo Programming
🎯 Background and Motivation
Mojo programming language, developed by Modular, has emerged as a game-changing technology in high-performance computing and AI development. Despite its growing popularity and impressive capabilities (up to 68,000x faster than Python!), existing LLMs struggle with Mojo code generation. Mojo-Coder addresses this gap by providing specialized support for Mojo programming, built upon… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Mojo_SFT.Mojo_Corpus
🔥 Mojo-Coder 🔥
State-of-the-art Language Model for Mojo Programming
🎯 Background and Motivation
Mojo programming language, developed by Modular, has emerged as a game-changing technology in high-performance computing and AI development. Despite its growing popularity and impressive capabilities (up to 68,000x faster than Python!), existing LLMs struggle with Mojo code generation. Mojo-Coder addresses this gap by providing specialized support for Mojo programming, built upon… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Mojo_Corpus.HumanEval-Mojo
🔥 Mojo-Coder 🔥
State-of-the-art Language Model for Mojo Programming
🎯 Background and Motivation
Mojo programming language, developed by Modular, has emerged as a game-changing technology in high-performance computing and AI development. Despite its growing popularity and impressive capabilities (up to 68,000x faster than Python!), existing LLMs struggle with Mojo code generation. Mojo-Coder addresses this gap by providing specialized support for Mojo programming, built upon… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/HumanEval-Mojo.chromium_mojokl62mojo-code
Mojo code
A collection of permisively licensed code from github containing mojo code.
aimmo-alpaca-testMojo_corpus_50KMojo_Prompthumaneval-mojo-fixedmojo-code-testtestingritual-sovereign-agentdietary-assistant-model-v2md-nishat-008-Mojo_CorpusCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données https://huggingface.co/datasets/md-nishat-008/Mojo_Corpus.
wavlmmojokortomenu-classifier-autotrain
