Developer signal
If you are building applications for Indian-language NLP (Hindi, Tamil, Telugu, Bengali, Marathi, Kannada, and 16 others), Sarvam-M is now locally runnable on llama.cpp b9093 or later. This is the first inference-engine native support for Sarvam's architecture outside the original Transformers implementation. Update llama.cpp to b9093 or later, then pull Sarvam-M GGUF weights from Hugging Face (check unsloth/sarvam-m-GGUF for quantized variants). No public benchmarks yet comparing GGUF quantization quality levels for this model — the first published llama-bench results across Q4_K_M/Q5_K_M/Q8_0 will define the initial community reference point.