datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
recycling_the_web-1m
Recycling the Web (MLX Subsets)
This is a subset of the facebook/recycling_the_web dataset, prepared for the MLX community.All credits for the original dataset go to Meta AI (Facebook).
Paper: Recycling the Web: A Method to Enhance Pre-training Data Quality and Quantity for Language Models
I’ve simply created smaller, more manageable shards for experimentation and training in MLX.Available sizes:
mlx-community/recycling_the_web-1k
mlx-community/recycling_the_web-100k… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/recycling_the_web-1m.recycling_the_web-400K
Recycling the Web (MLX Subsets)
Paper: Recycling the Web: A Method to Enhance Pre-training Data Quality and Quantity for Language Models
This is a subset of the facebook/recycling_the_web dataset, prepared for the MLX community.All credits for the original dataset go to Meta AI (Facebook).
I’ve simply created smaller, more manageable shards for experimentation and training in MLX.Available sizes:
mlx-community/recycling_the_web-1k
mlx-community/recycling_the_web-100k… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/recycling_the_web-400K.recycling_the_web-1k
Recycling the Web (MLX Subsets)
Paper: Recycling the Web: A Method to Enhance Pre-training Data Quality and Quantity for Language Models
This is a subset of the facebook/recycling_the_web dataset, prepared for the MLX community.All credits for the original dataset go to Meta AI (Facebook).
I’ve simply created smaller, more manageable shards for experimentation and training in MLX.Available sizes:
mlx-community/recycling_the_web-1k
mlx-community/recycling_the_web-100k… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/recycling_the_web-1k.
