ModelRefs / Groq — AI Glossary

Groq — AI Glossary

A silicon company building the LPU (Language Processing Unit), an ASIC delivering extremely low-latency LLM token generation.

Overview

Groq's LPU achieves 300–500 tokens/second per user for 70B models—10–20× faster than GPU-based inference—by using a deterministic dataflow architecture that eliminates memory bandwidth bottlenecks. GroqCloud offers API access to LLaMA 3, Mixtral, and Gemma at sub-50ms TTFT. Targets latency-critical real-time applications.

Reference details

Topicinfrastructure
Last reviewed2026-06-24

Commonly confused with

Unlike Together, Fireworks or Replicate, this is a chip company whose inference service exists to run its own silicon, which is why its distinguishing claim is token generation speed rather than model breadth. The comparison is therefore not like-for-like: the trade is exceptional decode throughput against a narrower catalogue of supported models. Do not confuse it with Grok, an unrelated model family.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Groq — AI Glossary.

Frequently asked questions

What is Groq?

A silicon company building the LPU (Language Processing Unit), an ASIC delivering extremely low-latency LLM token generation.

What concepts are related to Groq?

Closely related concepts include together ai, fireworks ai, tokens per second.