MOJO
Datasets
All datasets matching “MOJO”mojo-programming
Mojo Programming Language Training Data (HQ)
This dataset is a curated collection of high-fidelity training samples for the Mojo programming language. It is specifically structured for Supervised Fine-Tuning (SFT) and instruction-based model alignment.
Items Scrapped & Curated:
High Quality Variance Samples: Diverse code snippets covering Mojo-specific syntax (structs, decorators, SIMD).
Webscrape of Documentation: Cleaned markdown version of the official… See the full description on the dataset page: https://huggingface.co/datasets/libalpm64/mojo-programming.mojo-rl-assetsMojo_mSFT
🔥 Mojo-Coder 🔥
State-of-the-art Language Model for Mojo Programming
🎯 Background and Motivation
Mojo programming language, developed by Modular, has emerged as a game-changing technology in high-performance computing and AI development. Despite its growing popularity and impressive capabilities (up to 68,000x faster than Python!), existing LLMs struggle with Mojo code generation. Mojo-Coder addresses this gap by providing specialized support for Mojo programming, built upon… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Mojo_mSFT.trossen_ai_mojo_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_mobile",
"total_episodes": 2,
"total_frames": 1579,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/mrrl-emcnei/trossen_ai_mojo_test.Indian_laws_consumerIndian Legal Laws
This dataset contains a collection of legal Q-A pairs focusing on Indian legal laws, Furthermore the dataset is formatted to Alpaca formatting style (Keep that in mind while using this dataset to finetune the model). The dataset currently consists of 25,000 high quality pairs.
trossen_ai_mobile_test_mojoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_mobile",
"total_episodes": 2,
"total_frames": 59,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/mrrl-emcnei/trossen_ai_mobile_test_mojo.
