CoolFace
Datasetpublic

AdamLucek/Qwen3-4B-Instruct-2507-PII-RL-pii-masking-eval

Evaluation Results: Qwen3-4B-Instruct-2507-PII-RL on PII Masking This dataset contains evaluation results for the RL-trained model AdamLucek/Qwen3-4B-Instruct-2507-PII-RL on the adamlucek/pii-masking environment from Prime Intellect's Environment Hub. The model was fine-tuned using reinforcement learning to mask personally identifiable information (PII) in text. Evaluation Configuration Environment: adamlucek/pii-masking Model:… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/Qwen3-4B-Instruct-2507-PII-RL-pii-masking-eval.

sourceHugging Faceupdated 11mo agoView on Hugging Face
0likes11downloads
Dataset Card

Evaluation Results: Qwen3-4B-Instruct-2507-PII-RL on PII Masking

This dataset contains evaluation results for the RL-trained model AdamLucek/Qwen3-4B-Instruct-2507-PII-RL on the adamlucek/pii-masking environment from Prime Intellect's Environment Hub. The model was fine-tuned using reinforcement learning to mask personally identifiable information (PII) in text.

Evaluation Configuration

Performance Summary

Overall Metrics

MetricMeanStd DevMinMax
Total Reward0.8830.7290.1001.600
Exact Match0.5000.5020.0001.000
PII Count Accuracy0.5670.4970.0001.000
Format Compliance1.0000.0001.0001.000

Reward Breakdown by Rollout

RolloutMean RewardStd DevRange
Rollout 10.8800.737[0.100, 1.600]
Rollout 20.8900.729[0.100, 1.600]
Rollout 30.8800.737[0.100, 1.600]