Total unique visitors
Browse by category Chatbots Image Generation Video Generation Audio & Voice Coding Writing Productivity Research AI Agents Free Tier Table
🕐 Last checked: 2026-09-22 03:37

Groq pricing, free tier and how to use it

AMPM AI Ops × Groq: verified pricing and free tier

GroqThe fastest model inference platform — Groq

Free-tier friendliness
4/5
  • ✓ Has a free version
  • ✕ Quota reset is slow:官網未標示
  • ✓ Free version allows commercial use
  • ✓ No watermark on free plan
  • ✓ The free quota has a publicly stated specific number.

This score is calculated from the fields checked on the site and is not a subjective review. Price confirmed. 2026-09-21

AMPM〖100〗 Confirmed・Chatbots Seat #7(74.2 pts)→
Common Pitfalls

The selling point of Groq is "speed", not "free" - it can run Llama at over 800+ tokens per second, which is several times faster than general cloud services. However, the official website does not clearly state the rate limit for the free tier, and the actual quota needs to be checked when used as a formal service. Additionally, it only runs open-source models, **without GPT-5 or Claude**.

View Change History
2026-09-21 Daily auto-check logged — see the Chinese version for details
2026-09-20 Daily auto-check logged — see the Chinese version for details
Show the remaining verification history (18 more)▼
2026-09-19 Daily auto-check logged — see the Chinese version for details
2026-09-18 Daily auto-check logged — see the Chinese version for details
2026-09-17 Daily auto-check logged — see the Chinese version for details
2026-09-16 Daily auto-check logged — see the Chinese version for details
2026-09-15 Daily auto-check logged — see the Chinese version for details
2026-09-14 Daily auto-check logged — see the Chinese version for details
2026-09-13 Daily auto-check logged — see the Chinese version for details
2026-09-12 Daily auto-check logged — see the Chinese version for details
2026-09-11 Daily auto-check logged — see the Chinese version for details
2026-09-10 Daily auto-check logged — see the Chinese version for details
2026-09-09 Daily auto-check logged — see the Chinese version for details
2026-09-08 Daily auto-check logged — see the Chinese version for details
2026-08-27 Daily auto-check logged — see the Chinese version for details
2026-08-09 Daily auto-check logged — see the Chinese version for details
2026-08-07 21:24:58 Daily auto-check logged — see the Chinese version for details
2026-08-06 12:13:04 Daily auto-check logged — see the Chinese version for details
2026-08-05 22:36:28 Daily auto-check logged — see the Chinese version for details
2026-07-31 Daily auto-check logged — see the Chinese version for details

Continuously re-checked, every fact dated

Has a free plan

What Is This

Custom LPU chips deliver the fastest inference for open-source models, billed per token with no monthly fee.

Free Version Limitations

The Free Plan needs no card and each model has its own per-minute and per-day limits: Llama 3.1 8B Instant is 30 requests per minute and 14,400 per day, 6,000 tokens per minute and 500,000 tokens per day; Llama 3.3 70B Versatile is 30 per minute and 1,000 per day, 12,000 tokens per minute and 100,000 tokens per day; GPT OSS 120B and 20B are 30 per minute and 1,000 per day, 8,000 tokens per minute and 200,000 tokens per day; Groq Compound is 30 per minute and 250 per day, 70,000 tokens per minute; Whisper Large V3 is 20 per minute and 2,000 per day, 7,200 audio seconds per hour and 28,800 audio seconds per day. Going over the limit returns HTTP 429, and cached tokens do not count towards the limits. (Read directly from the official docs at console.groq.com/docs/rate-limits on 2026-08-09)

Available Models And Quantity

ModelFree quotaCalculate Time WindowDescription
Open-source models (Llama, GPT OSS, etc.)The official website does not specify a concrete upper limit.Registration comes with a free API key, but the official pricing page does not specify the daily/minute usage limit for the free tier. The actual limit is subject to the console display.

Free Usage Limit

FeaturesFree quotaDescription
Free API keyThere isSign up to get
Speed limitNot statedThis is the only key figure that was not stated when checked.
Available modelsOpen-source models are the primary focusLlama, GPT OSS series
Inference speedSame as paid planThe free tier also runs on LPU, and its speed is its biggest selling point.
Quota Resets AtNot stated
Terms Of UseRequires registration for an account and obtaining an API key
Free Version For Commercial UseYes

Sources: Groq official pricing page (last checked on 2026-07-31)

Pricing Plan Differences

API - Llama 3.1 8B Instant
Best for: Applications requiring extremely low latency, large capacity, and cost savings (Chatbots, real-time translation)
API pay-as-you-go (usage-based)

Per million tokens: input US$0.05, output US$0.08; speed approximately 840 TPS (Checked 2026-07-31 from official website)

Which ModelLlama 3.1 8B
Usage QuotaPricing based on tokens, pay-as-you-go
Max Reading TimeBased on model specifications
This Plan Includes
  • US$0.05 per million input tokens
  • US$0.08 per million output tokens
  • Approximately 840 TPS generation speed
  • Batch API can save an additional 50%
Not Included
  • GPT-5/Claude etc. closed models
  • Monthly flat-rate with unlimited usage
What Sets Us Apart

Similarly, running Llama, Groq's speed advantage is most obvious. However, it **does not provide GPT or Claude**, and those who need these models have to go through aggregators like OpenRouter.

Subscribe On Official Website
API - GPT OSS 20B
Best for: 開發者依 token 用量計費(小模型,較快)
API pay-as-you-go (usage-based)

Per million tokens: input US$0.075, output US$0.30; speed approximately 1,000 TPS (Checked on official website as of 2026-07-31)

This Plan Includes
  • 每百萬 token:輸入 US$0.075、輸出 US$0.30
  • 速度約 1,000 TPS
Subscribe On Official Website
API - GPT OSS 120B
Best for: 開發者依 token 用量計費(大模型)
API pay-as-you-go (usage-based)

Per million tokens: input US$0.15, output US$0.60; speed approx. 500 TPS (Checked 2026-07-31 from official website)

This Plan Includes
  • 每百萬 token:輸入 US$0.15、輸出 US$0.60
  • 速度約 500 TPS
Subscribe On Official Website
API - Llama 3.3 70B Versatile
Best for: Applications that require stronger reasoning capabilities without sacrificing too much speed
API pay-as-you-go (usage-based)

Per million tokens: input US$0.59, output US$0.79; speed approximately 394 TPS (Checked on 2026-07-31 from official website)

Which ModelLlama 3.3 70B
Usage QuotaPricing based on tokens
Max Reading TimeBased on model specifications
This Plan Includes
  • US$0.59 per million input tokens
  • US$0.79 per million output tokens
  • Approximately 394 TPS generation speed
Not Included
  • Closed model
What Sets Us Apart

About 10 times more expensive than 8B, but with significantly stronger reasoning capabilities. Still faster than most cloud services.

Subscribe On Official Website

Plan details verified from: Groq official pricing page (last checked on 2026-07-31)

Similar Options

Other tools in the same category as Groq — free tiers and pricing vary, so you can compare them side by side.

ChatGPT logo ChatGPT AI chat assistant

The world's most widely used AI assistant, handling conversation, search, image generation, and voice all in one place

FreeNT$0/month, open to everyone. The official site lists the free-tier limits one by one: limited access to GPT-5.5 Instant, limited messages and uploads, limited and slower image generation, limited deep research, limited memory and context, limited Codex, limited ChatGPT Work desktop app. The only place the official site gives concrete numbers is the plan comparison table: free-tier GPT Instant total context window 27K (Go/Plus 54K, Pro 128K); free-tier single-input cap about 12 pages of text (Go/Plus about 40 pages, Pro about 250 pages); the context window for reasoning models is marked depends on the situation for the free tier (Go/Plus 256K, Pro 400K). Response time on the free tier is marked as limited by system resources and service conditions; only paid plans are fast. Chat history is unlimited on the free tier. Whether content is used for model training can be opted out of. The official site has never published a concrete messages-per-day figure, and this site does not fill in an estimate. (Read directly from chatgpt.com/zh-Hant/pricing over a Taiwan connection on 2026-08-09)
Claude logo Claude AI chat/programming

Long-form writing and document quality are its strengths, with the paid version including the Claude Code engineering tool

FreeThe quota is calculated based on a rolling 5-hour usage window (not simply resetting at a fixed number every day), and general conversations can be used for around 10-20 times, excluding Claude Code.
Gemini logo Gemini AI chat assistant

Google has the deepest ecosystem integration of AI, with a generous free version and frequent discounts for student plans.

FreeNT$0/month, free to anyone with a Google account. Model available: Gemini 3.6 Flash; the official site states plainly that access to 3.1 Pro may vary, meaning free-tier access to Pro fluctuates and is not guaranteed. Features included on the free tier: image generation and editing, Deep Research, Gemini Live, Canvas, Gems, Gemini Notebook (research and writing), the Google Flow creative studio, plus limited usage of Nano Banana Pro. Cloud storage 15 GB (shared across Gmail/Drive/Photos). How the quota is counted and when it resets, in the official footnote's own words: the Gemini app measures usage limits by compute, which depends on prompt complexity, the features you use and conversation length; usage resets every 5 hours, up to a weekly usage cap. In other words the free tier is not a fixed number of messages per day but a compute allowance that resets every 5 hours plus an overall weekly cap; long conversations, Deep Research and image generation burn through it faster than plain chat. When the quota runs out you can buy AI credits. The official site does not publish the specific compute figure for the free tier. (Read directly from gemini.google/subscriptions over a Taiwan connection on 2026-08-09)

Using Information

DeveloperGroq
CategoryChatbots Coding
PlatformWeb
Chinese SupportPartial support
Official LinkVisit Groq's official site →

Notes

An inference-acceleration platform: it doesn't train its own models, but runs open models (Llama, GPT-OSS, etc.) on its custom LPU chips at some of the fastest speeds available. Billed per token, no monthly fee. zh_support is marked partial: the platform UI is in English, and Chinese capability depends on the model selected. ⚠️ As of 2026-08-09, groq.com/pricing now 302-redirects to the homepage; official pricing and rate limits have moved to the developer docs at console.groq.com/docs/models and /docs/rate-limits. Automated price-checkers pulling groq.com/pricing will fail to find pricing — this is expected behavior and should not be read as the data being invalid.

Related Comparison Alternatives

Related Guide:Groq Dropped Llama From the Free Tier on Aug 16 — the Official Quickstart Now Breaks (2026 gpt-oss Migration Guide)

Groq's deprecation page states that as of 2026-08-16, llama-3.1-8b-instant and llama-3.3-70b-versatile are no longer served to free and developer tier users; only enterprise committed-spend contracts are unaffected. Yet the official Quickstart sample code still says llama-3.3-70b-versatile. This guide covers what the free tier actually runs today with per-model limits, how to get a key, the one-line switch to gpt-oss, the OpenAI-compatible base URL and unsupported parameters, and how to read your remaining quota. Last checked 2026-08-27.

Read the full guide →

Groq Promotions & discounts

Looking for Groq deals, discounts or promo codes? We check the official site automatically every day, review changes by hand, and date-stamp every entry.

See the latest deals for all AI tools →

What Others Say

Share Your Experience Or Recommend Tools

Affiliate Links Notice