Ben Towner
About CV Blog

Turbo-FieldFare - Gemma 4 26B on Apple Silicon, under 5gb!!

On Better Stack they showed this new application for running an MoE model (Gemma 4 26B) on just a few GB of RAM. I tried it and it works really well. It’s not perfect, but then again nether is a 26B model. I want to see if I can use this model for single tasks that Opencode could call instead of burning money tokens.. though Kimi is so cheap… I was able to run 64k context with 32 experts in memory for about 5Gb and run at about 20 tokens / sec on an M3 Pro - Macbook Pro.

I think it’s awesome that people are figuring out ways to run large models with lower memory footprints. If we could do this at scale the amount of GPU, energy useage, and RAM drop significantly. Or you can think the otherway and say imagine the inferencing ability if you could use the resources we have now to run even bigger models.

Turbo-Fieldfare

#ai #apple #gemma4
August 3, 2026

← Back to blog

© 2026 Ben Towner