sdkv2/falcon-h1-1.5b-mlx-4bit
0170
Remove included inference code and benchmark files
Use mlx_lm runtime and remove custom implementation references
Remove custom implementation; use mlx_lm runtime
Compare custom inference with mlx_lm benchmarks
Remove unrelated CUDA benchmark comparison
Add PyTorch speed comparison to benchmarks
Add inference code and M4 benchmarks
Upload Falcon-H1 1.5B MLX 4-bit
initial commit
