Groq API

A low-latency inference API with model-specific request and token limits on Free.

···
html
<div class="limit-card"><header><b>GROQ FREE</b><span>gpt-oss-120b</span></header><div class="speed"><strong id="daily-label">80k</strong><span>tokens today</span></div><div class="rail"><i id="used"></i><em>200k TPD</em></div><div id="limit-state" class="state">within token limit</div><label>Daily tokens<input id="daily" type="range" min="0" max="300000" step="10000" value="80000" aria-label="Daily tokens"></label><p class="fine">Free: 30 RPM · 1,000 RPD · 8k TPM. Above a limit: 429.<br>Developer: $0.15 input / $0.60 output per 1M · 2026-10-06 공식 가격 기준 · 단순화(세금·지역 할증 제외)</p></div>
css
.limit-card{width:min(96vw,760px);height:min(94vh,350px);padding:clamp(8px,2vmin,18px);border:1px solid var(--line);border-radius:14px;background:var(--surface);display:flex;flex-direction:column;gap:clamp(4px,1.2vmin,10px);font:600 clamp(10px,2.2vmin,14px)/1.15 system-ui,sans-serif}header{display:flex;justify-content:space-between}header b{color:var(--accent)}header span,.speed span,.fine{color:var(--muted)}.speed{flex:1;min-height:0;display:flex;align-items:center;justify-content:center;gap:8px}.speed strong{font-size:clamp(30px,9vmin,66px)}.speed span{align-self:center}.rail{height:clamp(18px,6vmin,34px);border:1px solid var(--line);border-radius:18px;position:relative;overflow:hidden;background:var(--bg)}.rail i{display:block;height:100%;width:40%;background:var(--accent);transition:width .4s,background .4s}.rail em{position:absolute;right:7px;top:50%;transform:translateY(-50%);font-style:normal;color:var(--fg)}.state{color:var(--accent);text-align:right}.state.over{color:var(--accent-3)}label{display:block}input{width:100%;accent-color:var(--accent)}.fine{font-size:clamp(8px,1.6vmin,11px);margin:0}@media (min-width:600px){.price-card,.price-scene,.limit-card,.compute-card,.credit-card,.plan-card,.credit-ledger,.claude-card,.comparison{width:min(96vw,1080px);font-size:15px}.fine,.note,.comparison .fine{font-size:11px}.comparison header b{font-size:18px}}
js
const slider=document.getElementById('daily');const fill=document.getElementById('used');const state=document.getElementById('limit-state');let timer;let n=0;function draw(){const amount=Number(slider.value);const over=amount>200000;document.getElementById('daily-label').textContent=Math.round(amount/1000)+'k';fill.style.width=Math.min(100,amount/200000*100)+'%';fill.style.background=over?'var(--accent-3)':'var(--accent)';state.textContent=over?'429 on Free · no automatic overage':'within token limit';state.classList.toggle('over',over)}draw();timer=setInterval(()=>{n=(n+1)%4;slider.value=String([80000,160000,220000,290000][n]);draw()},1100);slider.addEventListener('input',()=>{clearInterval(timer);draw()})

Groq serves hosted open models through an API. For latency-sensitive work, check the active model list and free-plan rate limits together.

As of 2026-10, current openai/gpt-oss-120b on Developer costs $0.15 input and $0.60 output per million tokens. Free limits this model to 30 requests per minute, 1,000 per day, 8,000 tokens per minute, and 200,000 per day. The demo shows the daily token cap; exceeding a Free limit returns 429 rather than automatic billing.

Automatic prefix cache hits can halve input cost but do not stack with Batch discounts. Limits differ by model, and some old models have retired; check the active list before deployment.

When to use

Choose it for low-latency access to supported open models. Compare OpenRouter for model breadth or Modal for your own execution.

Open as page ↗