Free AI models are exploding — here’s what changed this week
Free AI models are exploding — here’s what changed this week
OpenRouter’s free tier jumped to 23 models (+2 vs yesterday). New additions: Google Gemma 4 (26B and 31B), OpenAI GPT-OSS-20b, Nvidia Nemotron Nano V2 variants, Qwen3-Next-80B-A3B.
The free API provider list grew to 56 providers (+4). MincAPI, SixFingerAPI, Subaxis, UnoRouter joined.
Google’s Gemini 3.5 Flash (1M context, 1500 RPD) and 3.1 Flash-Lite (30 RPM) remain stable. NVIDIA NIM confirmed operational for hosted free inference. Cerebras free tier: gpt-oss-120b, zai-glm-4.7 (~2,600 tok/s). GitHub Models still runs gpt-5, gpt-4.1, Llama 4 Scout, Mistral Small 3.1 for prototyping.
Sources: OpenRouter API, awesome-free-ai-api GitHub repo, AI Studio, NVIDIA NIM, Cerebras, GitHub Models.
Why it matters: Lower-cost prototyping just got a lot more options. You can now test bleeding-edge models without hitting a paywall or managing your own GPUs.
Related
More from the blog
2.8 Trillion Parameters in 594GB: What Kimi K3's 1-Bit Quant Says About Local AI
Kimi K3's 2.8T-parameter model shrinks to 594GB via dynamic 1-bit quantization, running on consumer hardware at 79% accuracy, which forces a reckoning on what local AI actually means at this scale.
CoreWeave's $2.58 Billion Quarter Proves AI Is Devouring Crypto's Money
CoreWeave reported $2.58 billion in Q2 revenue with a $104 billion backlog, proving AI infrastructure demand is shifting capital away from crypto toward compute.
OpenAI's Smart Speaker Wants to Seem "Alive". Here's Why That's a Bad Sign
OpenAI is building a $300 smart speaker with moving parts and a camera, betting that hardware can save its software business.
The AI Safety Test Is Becoming a Safety Risk
AI agents are breaking out of their cybersecurity test environments, exposing a dangerous gap between how fast models are being evaluated and how safely they're being contained.