ModelRefs / AI Glossary
AI Glossary
Definitions for every core AI, ML and LLM concept — written for builders.
What this reference supports
AI Glossary: This hub organizes related ModelRefs references into a crawlable starting point. Use it to narrow the problem, identify relevant profiles or guides, and continue into detailed evidence and implementation material.
AI Glossary: Items are connected across models, providers, benchmarks, workflows, tools, and guides. Those relationships explain where an option fits, what it depends on, and which adjacent decisions still need validation.
AI Glossary: Catalogue presence is not an endorsement or universal ranking. Compare candidates against your own requirements and review each page's sources, freshness notes, limitations, and coverage gaps.
All terms (481)
Every term in the glossary, linked. Entries marked with a definition of their own also carry a worked example or a note on what they are commonly confused with.
- A/B Testing (LLMs)
- Abstention
- Active Prompting
- Adapter (PEFT)
- Advanced RAG
- Adversarial Attack
- Agent as Tool
- Agent Handoff
- Agent Planning
- Agent Protocol
- Agent Reflection
- Agent Swarm / Swarm
- Agent Workflow
- Agent-to-Agent Protocol (A2A)
- Agentic Loop
- Agentic RAG
- AI Agent
- AI Alignment
- AI Anomaly Detection
- AI Audit
- AI Code Review
- AI Coding Assistant
- AI Content Generation
- AI Copilot
- AI Customer Support
- AI Data Analyst
- AI Documentation Generation
- AI Fairness
- AI Gateway
- AI Governance
- AI Personalization
- AI Privacy
- AI Recommendation System
- AI Search
- AI Summarization
- AI Test Generation
- AI Watermarking
- AI Workflow
- AI Writing Assistant
- Aider Polyglot
- ALiBi (Attention with Linear Biases)
- Alignment Tax
- AlpacaEval
- Answer Relevance
- Anthropic API
- Anthropic Python SDK
- API Key
- API Versioning
- Apple MLX
- Approximate Nearest Neighbor (ANN)
- ARC Challenge
- Arize Phoenix
- ASR (Automatic Speech Recognition)
- Assistant Message
- Attention Mechanism
- AutoGen
- Automatic Prompt Engineering (APE)
- AWQ (Activation-Aware Weight Quantization)
- AWS Bedrock
- Azure OpenAI Service
- Base Model
- Batch Inference
- Batch Size
- Beam Search
- BentoML
- BERTScore
- BF16 (Brain Floating Point 16)
- BFCL (Berkeley Function Calling Leaderboard)
- Bi-Encoder
- Bias Detection
- BIG-Bench Hard (BBH)
- BigCodeBench
- BLEU
- Blue-Green Deployment
- BM25
- Braintrust
- Byte-Pair Encoding (BPE)
- Calibration
- Canary Deployment
- Catastrophic Forgetting
- Causal Mask
- Chain of Thought (CoT)
- ChartQA
- Chat Completion
- Chatbot Arena
- Chinchilla Scaling
- Chunk Overlap
- Chunked Prefill
- Chunking
- Circuit Breaker
- Citation Accuracy
- Claude
- Closed Model
- Code Completion
- Code Generation
- Code Interpreter
- Coding Agent
- Cold Start
- Common Crawl
- Commonsense Reasoning
- Computer Use
- Constitutional AI (CAI)
- Content Moderation
- Context Management
- Context Precision
- Context Recall
- Context Stuffing
- Context Window
- Contextual Retrieval
- Continual Learning
- Continuous Batching
- Conversation History
- Conversational AI
- Corrective RAG (CRAG)
- Cosine Similarity
- Cost Per Request
- Cost per Successful Task
- Cost Per Token
- Counterfactual Explanation
- CrewAI
- Critic Agent
- Cross-Attention
- Curriculum Learning
- DAG (Directed Acyclic Graph)
- Data Augmentation
- Data Contamination
- Data Deduplication
- Data Parallelism
- Data Poisoning
- Decoder-Only Architecture
- DeepSpeed
- Dense Retrieval
- Differential Privacy
- Diffusion Model
- Direct Preference Optimization (DPO)
- Disaggregated Serving
- Distributed Inference
- Distributed Tracing (LLM)
- Document Loader
- Document Q&A
- Document Summary Index
- DocVQA
- Domain Adaptation
- Draft Model
- DROP
- DSPy
- Edge AI
- Embedding
- Embedding Dimension
- Embedding Model
- Emergent Capabilities
- Encoder-Decoder Architecture
- Encoder-Only Architecture
- Enterprise AI
- Episodic Memory
- EU AI Act
- Eval (Production Evaluation)
- Evaluation Benchmark
- Evaluation Harness
- Exact Match (EM)
- Experiment Tracking
- Explainable AI (XAI)
- Extended Thinking
- FAISS
- Faithfulness
- Fallback Chain
- Feature Store
- Federated Learning
- Feed-Forward Network (FFN)
- Few-Shot Prompting
- Fill-in-the-Middle (FIM)
- Fine-Tuning
- Fireworks AI
- FlashAttention
- Foundation Model
- FSDP (Fully Sharded Data Parallel)
- Function Calling
- G-Eval
- GAIA
- Generated Knowledge Prompting
- GGUF
- Google Gemini
- Google Vertex AI
- GPQA Diamond
- GPTQ
- GPU Memory (VRAM)
- GPU Utilization (MFU)
- Gradient Checkpointing
- GraphRAG
- Greedy Decoding
- Groq
- Grounding
- Grouped Query Attention (GQA)
- GRPO (Group Relative Policy Optimization)
- GSM8K
- Guardrails
- Hallucination
- Helicone
- HellaSwag
- HNSW (Hierarchical Navigable Small World)
- Hugging Face
- Human-in-the-Loop (HITL)
- HumanEval
- Hybrid Search
- HyDE (Hypothetical Document Embeddings)
- Hyperparameter
- IA³ (Infused Adapter by Inhibiting and Amplifying Inner Activations)
- In-Context Learning (ICL)
- Indexing Pipeline
- Inference
- Inference Cost
- Inference Endpoint
- Inference Provider
- Information Extraction
- Instruction Dataset
- Instruction Following
- Instruction-Tuned Model
- Intelligent Document Processing
- Interpretability
- Inverted File Index (IVF)
- Jailbreak
- Knowledge Distillation
- Knowledge Graph
- KV Cache
- LangChain
- Langfuse
- LangGraph
- LangSmith
- Large Language Model
- Late Interaction (ColBERT)
- Latency
- Layer Normalization
- Least-to-Most Prompting
- LIME (Local Interpretable Model-Agnostic Explanations)
- LiteLLM
- LiveCodeBench
- llama.cpp
- LlamaIndex
- LLM Observability
- LLM Router
- LLM Security
- LLM-as-Judge
- LLMOps
- LM Studio
- LMSYS
- Load Testing (LLMs)
- Log Probabilities
- Logit Bias
- Long-Context Model
- Long-Term Memory (Agents)
- Lookahead Decoding
- LoRA (Low-Rank Adaptation)
- Lost in the Middle
- Machine Translation
- Mamba
- MATH Benchmark
- Max Iterations
- Maximal Marginal Relevance (MMR)
- MCP (Model Context Protocol)
- Mechanistic Interpretability
- Megatron-LM
- Membership Inference Attack
- Memory (Agent)
- Memory Bandwidth
- Memory Store
- Message Role
- Message Threading
- Meta-Prompting
- Metadata Filtering
- Min-P Sampling
- Mixed Precision Training
- Mixture of Experts (MoE)
- MLC-LLM
- MLflow
- MMLU (Massive Multitask Language Understanding)
- MMMU (Massive Multidiscipline Multimodal Understanding)
- Modal
- Model Architecture
- Model Card
- Model Collapse
- Model Family
- Model Hub
- Model Merging
- Model Parallelism
- Model Pruning
- Model Registry
- Model Risk Management (MRM)
- Model Routing
- Model Size
- Model Soup
- Model Stealing
- Model Version
- Modular RAG
- MT-Bench
- Multi-Agent System
- Multi-Document Question Answering
- Multi-Head Attention (MHA)
- Multi-Query Retrieval
- Multi-Turn Conversation
- Multilingual Model
- Multimodal AI
- Multimodal Architecture
- N Completions
- Naive RAG
- Named Entity Recognition (NER)
- Namespace (Vector DB)
- Needle in a Haystack
- NIST AI Risk Management Framework
- NVIDIA H100
- OCR (Optical Character Recognition)
- Ollama
- Open LLM Leaderboard
- Open WebUI
- Open Weights
- Open-Domain Question Answering
- OpenAI API
- OpenAI Python SDK
- OpenAI-Compatible API
- Orchestrator Agent
- OSWorld
- Output Parser
- P99 Latency
- PagedAttention
- Parallel Execution (Agents)
- Parallel Tool Calls
- Parameter Count
- Parameter-Efficient Fine-Tuning (PEFT)
- Parent Document Retrieval
- Pass@k
- pgvector
- PII Detection
- PII Redaction
- Pipeline Parallelism
- PortKey
- Positional Bias
- Positional Encoding
- PPO (Proximal Policy Optimization)
- Preference Dataset
- Prefix Tuning
- Pretraining
- Pretraining Data
- Product Quantization (PQ)
- Prompt Caching
- Prompt Chaining
- Prompt Compression
- Prompt Engineering
- Prompt Injection
- Prompt Management
- Prompt Template
- Prompt Tuning
- Provider Lock-In
- QLoRA
- Quantization
- Query Decomposition
- Query Expansion
- Query Rewriting
- Question Answering (QA)
- RAG (Retrieval-Augmented Generation)
- RAGAS
- Random Seed
- RAPTOR
- Rate Limit Error
- Rate Limiting
- Ray Serve
- ReAct (Reason + Act)
- ReAct Agent Pattern
- Reasoning (LLM)
- Reasoning Model
- Recency Bias
- Reciprocal Rank Fusion (RRF)
- Red Teaming
- Refusal
- Regression Testing (LLM)
- Repetition Penalty
- Replicate
- Representative Workload
- Requests Per Minute (RPM)
- Reranker (Cross-Encoder)
- Research Agent
- Residual Connection
- Responsible AI
- Retrieval Pipeline
- Retry Logic
- Retry with Exponential Backoff
- Reward Model
- RLHF (Reinforcement Learning from Human Feedback)
- Role Prompting
- RoPE (Rotary Positional Embedding)
- ROUGE
- RULER
- RunPod
- Safety Training
- Sampling (LLM)
- Scaling Law
- Schema Adherence
- SDK (Software Development Kit)
- Self-Consistency
- Self-RAG
- Semantic Cache
- Semantic Memory
- Semantic Search
- Sentence Window Retrieval
- SentencePiece
- Sentiment Analysis
- Serverless Inference
- Shadow AI
- Shadow Mode
- SHAP (SHapley Additive exPlanations)
- Single-Agent System
- Sliding Window Attention
- Sparse Retrieval
- Speculative Decoding
- State Machine (Agents)
- State Space Model (SSM)
- Step-Back Prompting
- Stop Sequence
- Stopping Condition
- Streaming (Token Streaming)
- Stress Testing (LLMs)
- Structured Output
- Subagent
- Supervised Fine-Tuning (SFT)
- SWE-Bench
- SWE-Bench Verified
- Sycophancy
- Synthetic Data
- System Message
- Table Question Answering
- Task Decomposition
- Tau-Bench
- Temperature
- Tensor Parallelism
- Text Classification
- Text Completion
- Text Generation Inference (TGI)
- Text Splitter
- Text-to-SQL
- Throughput
- Tiktoken
- Time to First Token (TTFT)
- Together AI
- Token
- Token Budget
- Tokenizer
- Tokens Per Minute (TPM)
- Tokens Per Second (TPS)
- Tool Definition
- Tool Result
- Tool Selection
- Tool Use
- Tool-Using Agent
- Top-P (Nucleus Sampling)
- Toxicity Detection
- TPU (Tensor Processing Unit)
- Transfer Learning
- Transformer
- Tree of Thoughts (ToT)
- Triton Inference Server
- TruLens
- TruthfulQA
- TTS (Text-to-Speech)
- Uncertainty Quantification
- User Message
- Vector Database
- Verbosity
- Verifiable Reward
- Vision-Language Model (VLM)
- Visual Encoder
- vLLM
- Vocabulary Size
- Voice Agent
- VQA (Visual Question Answering)
- Warm Pool
- Web Agent
- WebArena
- Weights & Biases (W&B)
- WinoGrande
- Worker Agent
- Working Memory (Agents)
- ZeRO Optimizer
- Zero-Shot Prompting
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to AI Glossary.