CoolFace
Modelpublic

0xgr3y/Qwen3-0.6B-Gensyn-Swarm-tall_tame_panther

sourceHugging Faceapache-2.0updated 11mo agoView on Hugging Face
2likes1.4kdownloads
Model Card

Qwen3-0.6B-Gensyn-Swarm the Agent-ID (talltamepanther)

![Model](https://huggingface.co/0xgr3y/Qwen3-0.6B-Gensyn-Swarm-talltamepanther) ![GGUF](https://huggingface.co/0xgr3y/Qwen3-0.6B-Gensyn-Swarm-talltamepanther/tree/main) ![Gensyn](https://gensyn.ai) ![License](https://opensource.org/licenses/Apache-2.0)

Model Overview

This model is a continuously trained Qwen3-0.6B fine-tuned using Gensyn RL-Swarm framework with GRPO (Generalized Reward Policy Optimization) and support GGUF (llama.cpp) for enhanced reasoning and mathematical capabilities. Note: Current training focuses on math & reasoning tasks.

  • โ€”Agent ID: tall_tame_panther
  • โ€”Training Status: ๐ŸŸข LIVE - Model updates automatically every 5-10 minutes
  • โ€”Auto-Sync GGUF Pipeline Status: ๐ŸŸข LIVE - Commits update automatically every 1h-hourly
  • โ€”Current Progress: Round 43,610+ / 100,000 (43,61%)
  • โ€”Framework Version: Gensyn RL-Swarm v0.6.4
  • โ€”Contract: SwarmCoordinator v0.4.2

Key Features

  • โ€”Real-time Training: Continuous learning with distributed RL across Gensyn swarm network
  • โ€”Multi-domain Reasoning: Trained on logic, mathematical problem-solving & reasoning tasks
  • โ€”GGUF Support: Multiple quantized formats available (F16, Q3KM, Q4KM, Q5KM)
  • โ€”llama.cpp Compatible: Ready for edge deployment and local inference
  • โ€”BF16 Precision: Trained with bfloat16 for optimal performance
  • โ€”TGI Compatible: Supports Text Generation Inference for production deployment
  • โ€”Chat Format Support: Inherits Qwen3 chat template for conversational use

Training Data

The model is trained on a composite dataset (1,000 samples) with weighted sampling strategy:

DatasetWeightFocus Area
Propositional Logic7Logical reasoning, truth tables, Boolean operations
Calendar Arithmetic6Date calculations, leap years, recurring events
Decimal Arithmetic5Multi-term decimal operations with precision
Base Conversion4Number system conversions (base 2-16)
Fraction Simplification4GCD/LCM, fraction reduction
Basic Arithmetic2Foundation operations with parentheses

Total Dataset Size: 1,000 composite samples Training Samples per Round: 2 Evaluation: Real-time via swarm coordination

Quick Start

Standard Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "0xgr3y/Qwen3-0.6B-Gensyn-Swarm-tall_tame_panther",
    torch_dtype="auto",
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("0xgr3y/Qwen3-0.6B-Gensyn-Swarm-tall_tame_panther")

prompt = "What is 3/4 simplified to lowest terms?"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_length=256, temperature=0.6, top_p=0.95)
print(tokenizer.decode(outputs, skip_special_tokens=True))

Chat Format (Conversational)

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("0xgr3y/Qwen3-0.6B-Gensyn-Swarm-tall_tame_panther")
tokenizer = AutoTokenizer.from_pretrained("0xgr3y/Qwen3-0.6B-Gensyn-Swarm-tall_tame_panther")

messages = [
    {"role": "system", "content": "You are a helpful math tutor."},
    {"role": "user", "content": "Explain how to simplify 24/36 step by step."}
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt")
outputs = model.generate(**inputs, max_length=512)
print(tokenizer.decode(outputs))

Text Generation Inference (TGI)

docker run -d --gpus all \
  -p 8080:80 \
  -v $PWD/data:/data \
  ghcr.io/huggingface/text-generation-inference:latest \
  --model-id 0xgr3y/Qwen3-0.6B-Gensyn-Swarm-tall_tame_panther \
  --max-input-length 4096 \
  --max-total-tokens 8192

GGUF with llama.cpp

# Download quantized model (recommended: Q4_K_M)
wget https://huggingface.co/0xgr3y/Qwen3-0.6B-Gensyn-Swarm-tall_tame_panther/resolve/main/Qwen3-0.6B-Gensyn-Swarm-Q4_K_M.gguf

# Run inference
./llama-cli -m Qwen3-0.6B-Gensyn-Swarm-Q4_K_M.gguf \
  -p "Solve: (5 + 3) * 2 = ?" \
  --temp 0.6 --top-p 0.95

Ollama

# Create Modelfile
cat > Modelfile << 'EOF'
FROM ./Qwen3-0.6B-Gensyn-Swarm-Q4_K_M.gguf
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER top_k 20
SYSTEM "You are a helpful assistant specialized in mathematical reasoning and logic."
EOF

# Create and run
ollama create qwen3-swarm -f Modelfile
ollama run qwen3-swarm "What is 15 multiplied by 23?"

Available Formats

FormatSizePrecisionUse CaseDownload
Safetensors (BF16)1.19 GBBF16Full precision training/fine-tuningmodel.safetensors
GGUF F161.14 GBFP16High quality inferenceQwen3-0.6B-Gensyn-Swarm-F16.gguf
GGUF Q5KM444 MB5-bitBalanced quality/sizeQwen3-0.6B-Gensyn-Swarm-Q5_K_M.gguf
GGUF Q4KM397 MB4-bitRecommended for productionQwen3-0.6B-Gensyn-Swarm-Q4_K_M.gguf
GGUF Q3KM347 MB3-bitSmallest, fastestQwen3-0.6B-Gensyn-Swarm-Q3_K_M.gguf

All GGUF formats are llama.cpp compatible and auto-updated hourly.

GGUF Quantization Strategy

The Q5KM format uses mixed precision for optimal quality:

  • โ€”Token Embeddings: Q6_K (high quality vocab representation)
  • โ€”Attention Weights: Q5_K (balanced quality/size)
  • โ€”Feed-Forward: Q5K/Q6K (mixed for optimal performance)
  • โ€”Layer Norms: F32 (full precision for stability)

This strategy ensures minimal quality loss while maintaining small file size.

Chat Format & Conversational Use

This model inherits Qwen3's chat template for structured conversations.

Format Structure

<|im_start|>system
{system_message}
<|im_end|>
<|im_start|>user
{user_message}
<|im_end|>
<|im_start|>assistant
{assistant_response}
<|im_end|>

Chat Template Features

  • โ€”System Instructions: Guide model behavior with system messages
  • โ€”Multi-turn Dialogue: Maintains conversation context
  • โ€”Tool Calling: Support function calling (if enabled in training)
  • โ€”Reasoning Mode: <think> tags for chain-of-thought (experimental)

Note: While the model supports chat format structurally, optimal conversational performance depends on whether training data included formatted dialogues. Current training focuses on math/reasoning tasks.

Training Configuration

Gensyn RL-Swarm Architecture

Training Framework:
  Method: GRPO (Generalized Reward Policy Optimization)
  Base Model: Qwen/Qwen3-0.6B
  Training Regime: bfloat16 mixed precision
  Max Rounds: 100,000
  Update Frequency: Every 5-10 minutes
  Generations per Round: 2
  Seed: 42

Blockchain Integration:
  Network: Gensyn Testnet
  Chain ID: 685685
  Contract: SwarmCoordinator v0.4.2

Swarm Communication:
  Framework: Hivemind P2P Backend
  Initial Peers: 3 bootnodes
  Beam Size: 30

Reward System:
  Manager: DefaultRewardManager
  Reward Function: RGRewards (Reasoning Gym)
  Judge API: https://swarm-judge.internal-apps-central1.clusters.gensyn.ai

Model Hyperparameters

Architecture:
  Hidden Size: 1024
  Intermediate Size: 3072
  Layers: 28
  Attention Heads: 16
  KV Heads: 8
  Head Dimension: 128
  Context Length: 40,960 tokens
  Vocabulary: 151,936 tokens

GRPO Config:
  Epsilon: 0.2
  Epsilon High: 0.28
  Gradient Checkpointing: Enabled
  
Generation:
  Temperature: 0.6
  Top-K: 20
  Top-P: 0.95

Model Capabilities

This model excels at:

  1. 1.Logical Reasoning: Propositional logic, truth evaluation, Boolean algebra
  2. 2.Mathematical Operations: Multi-precision arithmetic, decimal calculations, fractions
  3. 3.Number Systems: Base conversion (binary, octal, decimal, hexadecimal)
  4. 4.Date/Time Calculations: Calendar arithmetic, leap years, day-of-week
  5. 5.Step-by-step Problem Solving: Chain-of-thought reasoning
  6. 6.Conversational Tutoring: Interactive problem-solving (via chat format)

Limitations

  • โ€”Specialized Domain: Optimized for reasoning/math; may underperform on creative writing
  • โ€”Training in Progress: Weights update every 5-10 minutes; performance varies
  • โ€”Scale: 0.6B parameters - suitable for edge but not SOTA for complex reasoning
  • โ€”Experimental: Decentralized RL training; behavior less predictable than supervised models
  • โ€”Context: Best performance within 4K tokens (full 40K supported)

Update Schedule

FormatFrequencyTrigger
Safetensors (BF16)Every 5-10 minAutomatic via RL-Swarm
GGUF (all formats)Every 1 hourAuto-conversion pipeline

Auto-Conversion Pipeline:

  1. 1.Monitors repo for new training commits
  2. 2.Downloads latest model.safetensors
  3. 3.Converts to F16 GGUF base
  4. 4.Quantizes to Q3KM, Q4KM, Q5KM
  5. 5.Uploads all formats

Check commit history for exact timestamps.

Gensyn RL-Swarm Technical Details

Architecture Components

  1. 1.Game Manager: Orchestrates training rounds and swarm coordination
  2. 2.Trainer: GRPO implementation for policy optimization
  3. 3.Data Manager: Dataset loading and weighted sampling
  4. 4.Reward Manager: Computes rewards via judge API
  5. 5.Coordinator: Blockchain integration for swarm state
  6. 6.P2P Backend: Hivemind DHT for model sharing

Training Process

1. Agent joins swarm via P2P network
2. Coordinator assigns round via smart contract
3. Agent samples data from weighted datasets
4. Model generates 2 responses
5. Judge API evaluates and assigns rewards
6. GRPO updates policy based on rewards
7. Updated model shared via DHT
8. Best checkpoint saved to HuggingFace
9. Repeat

Decentralization Benefits

  • โ€”Fault Tolerance: Multiple agents; no single point of failure
  • โ€”Diverse Exploration: Different agents explore different strategies
  • โ€”Collective Intelligence: Agents learn from each other
  • โ€”Transparent: All rounds verified on-chain

Swarm Agent: tall_tame_panther Contract: SwarmCoordinator v0.4.2

Technical Specifications

Software Stack

  • โ€”Framework: Gensyn RL-Swarm v0.6.4
  • โ€”Library: transformers v4.51+
  • โ€”P2P: hivemind
  • โ€”Blockchain: Gensyn testnet
  • โ€”Config: Hydra + OmegaConf
  • โ€”Logging: WandB integration

Hardware Requirements

Training GPU:

  • โ€”GPU: NVIDIA 4090 24GB+ (BF16 training)
  • โ€”RAM: 16GB+
  • โ€”Cores: 10+
  • โ€”Storage: 50GB SSD
  • โ€”Network: High bandwidth for P2P

Training CPU Optimize:

  • โ€”CPU: INTEL or AMD
  • โ€”Cores: 10+
  • โ€”RAM: 16GB+
  • โ€”Storage: 50GB SSD
  • โ€”Network: High bandwidth for P2P

Inference:

  • โ€”Safetensors: 8GB VRAM (GPU) / 16GB RAM (CPU)
  • โ€”GGUF Q4KM: 2GB VRAM (GPU) / 4GB RAM (CPU)
  • โ€”GGUF Q3KM: 3GB RAM (CPU-only)

Evaluation

Training Progress Metrics

MetricValueTarget
Completed Rounds43,610+100,000
Training Progress43.61%100%
Update Frequency5-10 minContinuous

Note: Formal evaluation benchmarks (GSM8K, MATH, etc.) will be added as training progresses. Current metrics track training rounds completed in the decentralized swarm.

Reproducibility

To reproduce training:

  1. 1.Clone Gensyn RL-Swarm repository
  2. 2.Install: pip install -r requirements.txt
  3. 3.Configure rgym_exp/config/rg-swarm.yaml
  4. 4.Configure rgym_exp/src/datasets.yaml
  5. 5.Set environment variables:
export HUGGINGFACE_ACCESS_TOKEN=<token>
export MODEL_NAME=Qwen/Qwen3-0.6B
export ORG_ID=<org-id>
export SWARM_CONTRACT=<contract-address>
  1. 1.Run: bash run_rl_swarm.sh

Note: Exact reproduction requires same seed (42), dataset config, and swarm state.

Citation

@misc{qwen3-gensyn-swarm-2025,
  author = {0xgrey},
  title = {Qwen3-0.6B-Gensyn-Swarm: Continuous RL Training on Distributed Swarm},
  year = {2025},
  publisher = {HuggingFace},
  howpublished = {\url{https://huggingface.co/0xgr3y/Qwen3-0.6B-Gensyn-Swarm-tall_tame_panther}},
  note = {Agent ID: tall\_tame\_panther}
}

@misc{gensyn-rl-swarm-2025,
  title = {Gensyn RL-Swarm: Decentralized Reinforcement Learning Framework},
  author = {Gensyn AI},
  year = {2025},
  url = {https://gensyn.ai}
}

References

  • โ€”Gensyn Documentation: https://docs.gensyn.ai/
  • โ€”Gensyn GitHub: https://github.com/gensyn-ai
  • โ€”RL-Swarm Contracts: https://github.com/gensyn-ai/rl-swarm-contracts
  • โ€”Qwen3 Model Card: https://huggingface.co/Qwen/Qwen3-0.6B
  • โ€”arXiv:1910.09700: ML Carbon Emissions methodology

License

Apache 2.0 - See LICENSE

Contact

  • โ€”Developer: 0xgrey
  • โ€”Agent ID: talltamepanther
  • โ€”Community: Gensyn Discord

โš ๏ธ Important: This is a continuously trained model. For reproducibility, specify commit hash:

git clone https://huggingface.co/0xgr3y/Qwen3-0.6B-Gensyn-Swarm-tall_tame_panther
cd Qwen3-0.6B-Gensyn-Swarm-tall_tame_panther
git checkout <commit-hash>

<div align="center">

๐Ÿค– Trained with โค๏ธ using Gensyn RL-Swarm

![Gensyn](https://gensyn.ai)

</div>