Groq serves hosted open models through an API. For latency-sensitive work, check the active model list and free-plan rate limits together.
As of 2026-10, current openai/gpt-oss-120b on Developer costs $0.15 input and $0.60 output per million tokens. Free limits this model to 30 requests per minute, 1,000 per day, 8,000 tokens per minute, and 200,000 per day. The demo shows the daily token cap; exceeding a Free limit returns 429 rather than automatic billing.
Automatic prefix cache hits can halve input cost but do not stack with Batch discounts. Limits differ by model, and some old models have retired; check the active list before deployment.
When to use
Choose it for low-latency access to supported open models. Compare OpenRouter for model breadth or Modal for your own execution.