Business

Microsoft offers 26 open models to startups on Foundry

Microsoft has launched a deployment blueprint integrating Fireworks AI into its Foundry platform, allowing startups to run 26 open models using up to $150,000 in Azure credits.

Unite.AI4 days agoBusiness
Image: Unite.AI

Microsoft has officially released a deployment blueprint on August 4, 2026, to help startups run open-weight models on Microsoft Foundry via its newly generally available Fireworks AI integration. Members of the Microsoft for Startups program can now apply up to $150,000 in Azure credits toward these Fireworks model deployments. This integration brings 26 open-weight models from providers like DeepSeek, Moonshot AI, Z.ai, MiniMax, Qwen, Google, and OpenAI's gpt-oss line directly into Azure's model catalog, with Fireworks AI handling the inference and Foundry acting as the control plane.

For AI practitioners, this setup eliminates the need to manage complex GPU clusters. The reference architecture runs within a startup's own Azure subscription, utilizing Azure Container Apps to call Fireworks endpoints, Azure Container Registry for images, and Azure Key Vault for credentials. To manage costs and scale, developers can route traffic through Azure API Management, implement Azure Cache for Redis to reduce redundant inference, and track performance metrics using Azure Monitor. This allows teams to start with a single serverless endpoint and scale up infrastructure only as demand requires.

The catalog features prominent models such as Moonshot AI's Kimi K2.5, DeepSeek V3.2, MiniMax M2.5, OpenAI's gpt-oss-120b, and the 1.6-trillion-parameter DeepSeek V4 Pro. While six models, including Kimi K2.6 and Z.ai's GLM-5.1, support pay-per-token serverless billing, the rest require provisioned throughput units. However, the startup credits only apply to pay-per-token Data Zone Standard usage, excluding provisioned throughput. Furthermore, pay-per-token billing for GLM-5.1 and MiniMax M2.5 is deprecated as of August 7, 2026, alongside four other previously deprecated models.

Practitioners must also navigate strict compliance boundaries. Serverless deployments are restricted to six US Azure regions, sit outside the EU Data Boundary, lack FedRAMP authorization, and cannot process payment-card data. Additionally, Microsoft does not evaluate the safety of these Fireworks-served models. Despite these limits, the integration offers a flexible alternative to self-hosting vLLM on dedicated GPU fleets, allowing startups to bring their own weights with LoRA adapter support in public preview.

This is our own summary of reporting by Unite.AI

More in Business