← Back to feed
vendor
Cloudflare Blog
Smaller, faster, safer: running Kimi and GLM at scale
Serving frontier models like Kimi and GLM means fighting for GPU memory. Here's how we quantize KV caches, compress model weights, and add integrity checks t...
Read the full story
Cloudflare Blog →