Ola's Bhavish Aggarwal says Krutrim has deployed DeepSeek R1 671B on Nvidia's H100 and is offering the model to Indian developers starting at ₹1/million tokens
Context & Ripple Effects
Krutrim had already opened developer offerings spanning cloud access and a chatbot; this expands that route to market with a hosted frontier-scale model rather than only its own services. It makes Krutrim's H100 capacity a product Indian developers can consume directly.
The move sits between Aggarwal's fresh capital commitment to improve local AI and Krutrim's later plan to develop a 700B-parameter Krutrim 3 model with Lenovo. Together, the coverage shows a company pursuing both third-party model serving and proprietary-model development.
First-order effects
- Indian developers can access DeepSeek R1 671B through Krutrim at the stated token price, without independently operating the underlying H100 infrastructure.
- Krutrim adds a high-capacity model API to the developer and cloud offerings it launched earlier, creating a near-term use case for its Nvidia GPU deployment.
Second-order effects
- The quoted API price gives Indian teams a concrete benchmark for comparing hosted inference with self-managed GPU capacity or other model providers.
- Krutrim must turn low-cost access into sustained developer usage to cover the serving costs of a 671B-parameter model on H100s; model quality, uptime and support become as relevant as the headline token rate.
Third-order effects
- If providers increasingly host leading third-party models alongside their own, AI competition will shift toward inference economics, local developer distribution and infrastructure operations—not model ownership alone.
- The pairing of hosted external models with Krutrim's own planned 700B model suggests an emerging hybrid AI-platform strategy, though its durability depends on developer adoption and the economics of GPU-backed serving.
The trend: Indian AI platforms are commercializing scarce GPU capacity as model APIs while building proprietary models to control more of the stack over time.