AI Glossary
71 current AI terms explained in plain English — from agents and context engineering to provenance, RAG and quantization.
Agentic AI
AI systems that can pursue goals across multiple steps, choose tools, observe results and take actions rather than only generate a single response.
AGI (Artificial General Intelligence)
A hypothetical system with broad, general intellectual capability across many domains. There is no universally accepted test for AGI, and the label is used differently by different organizations.
AI Detector
A tool that estimates whether content may have been AI-generated. Detector scores are uncertain and should not be treated as proof of authorship.
AI Literacy
The knowledge and practical skills needed to understand what AI can do, its limits and risks, and how to use or oversee it responsibly.
Alignment
The problem of making an AI system behave in ways consistent with intended goals, constraints and human values, including under unfamiliar conditions.
Benchmark
A standardized test set or suite used to compare model performance. Benchmarks are useful signals but may not predict performance in your exact workflow.
Bias
Systematic differences in outputs or outcomes that can arise from data, labels, objectives, design choices, deployment context or human interpretation.
Black Box
A system whose internal decision process is difficult to interpret directly, even when its inputs and outputs can be observed.
C2PA / Content Credentials
A technical standard for tamper-evident content provenance. It can record information about creation and edits, but it does not prove that a claim inside the content is true.
Chain-of-Thought
A term for intermediate reasoning steps associated with solving a task. For verification, users should focus on observable evidence, assumptions, calculations and concise rationale rather than treating hidden reasoning as an audit log.
Computer Vision
AI methods for interpreting images and video, including recognition, detection, segmentation and OCR.
Context Engineering
Designing the full information environment around a model — instructions, documents, examples, memory, tools and constraints — so it can perform a task reliably.
Context Window
The amount of information a model can consider in one run. A larger window increases capacity but does not guarantee every detail will be used correctly.
Copyright (AI outputs)
The legal protection of original authorship. Rules vary by jurisdiction; in the U.S., the Copyright Office says protection requires sufficient human authorship rather than mere prompting.
Data Poisoning
Manipulating training or retrieval data so a model learns, retrieves or produces harmful or misleading behavior.
Data Privacy
Practices and rules for protecting personal, confidential or sensitive information throughout collection, use, storage and sharing.
Deepfake
Synthetic or manipulated audio, video or imagery that convincingly imitates a real person or event.
Distillation
Training a smaller model to reproduce useful behavior or knowledge from a larger or stronger model or system.
DPO (Direct Preference Optimization)
A method for tuning model behavior directly from preference data, often used as an alternative or complement to reinforcement-learning-based alignment methods.
Embedding
A numerical representation of data that captures useful relationships, often used for semantic search, retrieval, clustering and recommendations.
EU AI Act
European Union legislation that regulates AI using a risk-based approach. Different obligations apply depending on system, use and role; Article 50 transparency obligations began applying on 2 August 2026.
Evals
Structured evaluations that test a model or AI system on defined capabilities, quality criteria or risks.
Federated Learning
Training across devices or organizations while keeping raw data local and sharing model updates or derived information instead.
Fine-tuning
Additional training applied to a pre-trained model to adapt its behavior, style, domain knowledge or task performance.
Foundation Model
A broadly capable model trained at scale and adaptable to many downstream tasks.
Frontier Model
An informal term for highly capable models near the current leading edge. Definitions vary across organizations and policy contexts.
Generative AI
AI systems that create new content such as text, images, audio, video or code.
Grounding
Connecting an AI answer to specific evidence or external data so the response is less dependent on unsupported generation.
Guardrails
Technical or procedural controls intended to constrain unsafe, unauthorized or unwanted model behavior.
Hallucination
Generated content that is unsupported or false despite being presented plausibly or confidently.
Human-in-the-Loop
A workflow where a person meaningfully reviews, approves, corrects or overrides an AI recommendation or action.
Inference
Running a trained model to produce a prediction or generated output. This is distinct from training the model.
Jailbreaking
Attempts to bypass a model or product safety restrictions through crafted instructions or other inputs.
Large Language Model (LLM)
A model trained on large amounts of language data to understand and generate text and often other modalities.
Latency
The time between submitting a request and receiving a response or completed action.
Machine Learning
Methods that learn patterns from data rather than relying only on hand-written rules.
Model Collapse
A family of degradation effects that can occur when training data becomes overly dominated by model-generated content or loses diversity.
Model Context Protocol (MCP)
An open protocol for connecting AI applications to external tools and data through standardized interfaces.
Model Router
A system that sends each request to a selected model based on cost, latency, capability, policy or other criteria.
Multimodal AI
AI that can work with more than one kind of data, such as text, images, audio or video.
Narrow AI
AI built for particular tasks or domains rather than general human-level capability across arbitrary activities.
Natural Language Processing (NLP)
The field concerned with computational understanding, processing and generation of human language.
Neural Network
A layered mathematical model whose learned parameters transform inputs into predictions or generated outputs.
OCR
Optical Character Recognition: technology that converts text in images or scans into machine-readable text.
Open-weight Model
A model whose learned weights are made available under a license. Open-weight does not necessarily mean every part of the training data, code or license is fully open.
Overfitting
When a model fits training examples so closely that performance does not generalize well to new data.
Parameter
A learned numerical value inside a model. Parameters collectively encode patterns learned during training.
Prompt
Instructions and context supplied to an AI model or application.
Prompt Injection
An attack in which untrusted content tries to alter or override the instructions an AI system should follow, especially dangerous when the system has tool access.
Provenance
Information about the origin and history of a digital asset, including how it was created or modified.
Quantization
Reducing the numerical precision of model weights or activations to lower memory and compute requirements, typically with a quality trade-off.
RAG (Retrieval-Augmented Generation)
A pattern that retrieves relevant external information and supplies it to a model during generation.
Reasoning Model
A model designed to spend additional computation on difficult tasks before producing a final answer.
Red Teaming
Adversarial testing that deliberately probes for vulnerabilities, harmful behavior, misuse paths and system failures.
RLHF
Reinforcement Learning from Human Feedback: a family of methods that use human preference signals to shape model behavior.
Semantic Search
Retrieval based on meaning or conceptual similarity rather than only exact keyword matches.
Sentiment Analysis
Automated estimation of attitudes or emotional polarity in text or other content; outputs can be ambiguous and context-sensitive.
Shadow AI
Use of unapproved or unmanaged AI tools inside an organization, creating privacy, security or compliance risks.
Small Language Model (SLM)
A comparatively compact language model designed for lower cost, lower latency, local use or narrower tasks.
Supervised Learning
Training from examples where inputs are paired with labels or target outputs.
Sycophancy
A tendency for a conversational model to agree with, flatter or reinforce a user’s premise when it should instead challenge or correct it.
Synthetic Data
Artificially generated data designed to resemble properties of real data for training, testing, simulation or privacy-related purposes.
Temperature
A generation setting that influences randomness in token selection; exact behavior and exposed controls vary by model and API.
Tokens
Units into which model inputs and outputs are divided for processing. Tokens are not always the same as words.
Tool Use
A model calling an external function, API, browser, database, code runner or other capability as part of completing a task.
Training Data
Data used to adjust model parameters during training.
Turing Test
Alan Turing’s historical imitation-game concept for machine conversation. Modern AI evaluation uses many task-specific benchmarks and behavioral tests instead.
Uncanny Valley
The idea that near-human representations can feel unsettling when they appear very human but still contain noticeable mismatches.
Vector Database
A database optimized for storing and searching vector representations such as embeddings.
Watermarking
Embedding a detectable signal into generated content or model outputs. Watermarks can aid provenance or detection but have technical limitations.
Zero-shot
Performing a task from instructions without being given task-specific examples in the prompt.