Reference

AI Glossary

71 current AI terms explained in plain English — from agents and context engineering to provenance, RAG and quantization.

A

Agentic AI

AI systems that can pursue goals across multiple steps, choose tools, observe results and take actions rather than only generate a single response.

AGI (Artificial General Intelligence)

A hypothetical system with broad, general intellectual capability across many domains. There is no universally accepted test for AGI, and the label is used differently by different organizations.

AI Detector

A tool that estimates whether content may have been AI-generated. Detector scores are uncertain and should not be treated as proof of authorship.

AI Literacy

The knowledge and practical skills needed to understand what AI can do, its limits and risks, and how to use or oversee it responsibly.

Alignment

The problem of making an AI system behave in ways consistent with intended goals, constraints and human values, including under unfamiliar conditions.

B

Benchmark

A standardized test set or suite used to compare model performance. Benchmarks are useful signals but may not predict performance in your exact workflow.

Bias

Systematic differences in outputs or outcomes that can arise from data, labels, objectives, design choices, deployment context or human interpretation.

Black Box

A system whose internal decision process is difficult to interpret directly, even when its inputs and outputs can be observed.

C

C2PA / Content Credentials

A technical standard for tamper-evident content provenance. It can record information about creation and edits, but it does not prove that a claim inside the content is true.

Chain-of-Thought

A term for intermediate reasoning steps associated with solving a task. For verification, users should focus on observable evidence, assumptions, calculations and concise rationale rather than treating hidden reasoning as an audit log.

Computer Vision

AI methods for interpreting images and video, including recognition, detection, segmentation and OCR.

Context Engineering

Designing the full information environment around a model — instructions, documents, examples, memory, tools and constraints — so it can perform a task reliably.

Context Window

The amount of information a model can consider in one run. A larger window increases capacity but does not guarantee every detail will be used correctly.

Copyright (AI outputs)

The legal protection of original authorship. Rules vary by jurisdiction; in the U.S., the Copyright Office says protection requires sufficient human authorship rather than mere prompting.

D

Data Poisoning

Manipulating training or retrieval data so a model learns, retrieves or produces harmful or misleading behavior.

Data Privacy

Practices and rules for protecting personal, confidential or sensitive information throughout collection, use, storage and sharing.

Deepfake

Synthetic or manipulated audio, video or imagery that convincingly imitates a real person or event.

Distillation

Training a smaller model to reproduce useful behavior or knowledge from a larger or stronger model or system.

DPO (Direct Preference Optimization)

A method for tuning model behavior directly from preference data, often used as an alternative or complement to reinforcement-learning-based alignment methods.

E

Embedding

A numerical representation of data that captures useful relationships, often used for semantic search, retrieval, clustering and recommendations.

EU AI Act

European Union legislation that regulates AI using a risk-based approach. Different obligations apply depending on system, use and role; Article 50 transparency obligations began applying on 2 August 2026.

Evals

Structured evaluations that test a model or AI system on defined capabilities, quality criteria or risks.

F

Federated Learning

Training across devices or organizations while keeping raw data local and sharing model updates or derived information instead.

Fine-tuning

Additional training applied to a pre-trained model to adapt its behavior, style, domain knowledge or task performance.

Foundation Model

A broadly capable model trained at scale and adaptable to many downstream tasks.

Frontier Model

An informal term for highly capable models near the current leading edge. Definitions vary across organizations and policy contexts.

G

Generative AI

AI systems that create new content such as text, images, audio, video or code.

Grounding

Connecting an AI answer to specific evidence or external data so the response is less dependent on unsupported generation.

Guardrails

Technical or procedural controls intended to constrain unsafe, unauthorized or unwanted model behavior.

H

Hallucination

Generated content that is unsupported or false despite being presented plausibly or confidently.

Human-in-the-Loop

A workflow where a person meaningfully reviews, approves, corrects or overrides an AI recommendation or action.

I

Inference

Running a trained model to produce a prediction or generated output. This is distinct from training the model.

J

Jailbreaking

Attempts to bypass a model or product safety restrictions through crafted instructions or other inputs.

L

Large Language Model (LLM)

A model trained on large amounts of language data to understand and generate text and often other modalities.

Latency

The time between submitting a request and receiving a response or completed action.

M

Machine Learning

Methods that learn patterns from data rather than relying only on hand-written rules.

Model Collapse

A family of degradation effects that can occur when training data becomes overly dominated by model-generated content or loses diversity.

Model Context Protocol (MCP)

An open protocol for connecting AI applications to external tools and data through standardized interfaces.

Model Router

A system that sends each request to a selected model based on cost, latency, capability, policy or other criteria.

Multimodal AI

AI that can work with more than one kind of data, such as text, images, audio or video.

N

Narrow AI

AI built for particular tasks or domains rather than general human-level capability across arbitrary activities.

Natural Language Processing (NLP)

The field concerned with computational understanding, processing and generation of human language.

Neural Network

A layered mathematical model whose learned parameters transform inputs into predictions or generated outputs.

O

OCR

Optical Character Recognition: technology that converts text in images or scans into machine-readable text.

Open-weight Model

A model whose learned weights are made available under a license. Open-weight does not necessarily mean every part of the training data, code or license is fully open.

Overfitting

When a model fits training examples so closely that performance does not generalize well to new data.

P

Parameter

A learned numerical value inside a model. Parameters collectively encode patterns learned during training.

Prompt

Instructions and context supplied to an AI model or application.

Prompt Injection

An attack in which untrusted content tries to alter or override the instructions an AI system should follow, especially dangerous when the system has tool access.

Provenance

Information about the origin and history of a digital asset, including how it was created or modified.

Q

Quantization

Reducing the numerical precision of model weights or activations to lower memory and compute requirements, typically with a quality trade-off.

R

RAG (Retrieval-Augmented Generation)

A pattern that retrieves relevant external information and supplies it to a model during generation.

Reasoning Model

A model designed to spend additional computation on difficult tasks before producing a final answer.

Red Teaming

Adversarial testing that deliberately probes for vulnerabilities, harmful behavior, misuse paths and system failures.

RLHF

Reinforcement Learning from Human Feedback: a family of methods that use human preference signals to shape model behavior.

S

Semantic Search

Retrieval based on meaning or conceptual similarity rather than only exact keyword matches.

Sentiment Analysis

Automated estimation of attitudes or emotional polarity in text or other content; outputs can be ambiguous and context-sensitive.

Shadow AI

Use of unapproved or unmanaged AI tools inside an organization, creating privacy, security or compliance risks.

Small Language Model (SLM)

A comparatively compact language model designed for lower cost, lower latency, local use or narrower tasks.

Supervised Learning

Training from examples where inputs are paired with labels or target outputs.

Sycophancy

A tendency for a conversational model to agree with, flatter or reinforce a user’s premise when it should instead challenge or correct it.

Synthetic Data

Artificially generated data designed to resemble properties of real data for training, testing, simulation or privacy-related purposes.

T

Temperature

A generation setting that influences randomness in token selection; exact behavior and exposed controls vary by model and API.

Tokens

Units into which model inputs and outputs are divided for processing. Tokens are not always the same as words.

Tool Use

A model calling an external function, API, browser, database, code runner or other capability as part of completing a task.

Training Data

Data used to adjust model parameters during training.

Turing Test

Alan Turing’s historical imitation-game concept for machine conversation. Modern AI evaluation uses many task-specific benchmarks and behavioral tests instead.

U

Uncanny Valley

The idea that near-human representations can feel unsettling when they appear very human but still contain noticeable mismatches.

V

Vector Database

A database optimized for storing and searching vector representations such as embeddings.

W

Watermarking

Embedding a detectable signal into generated content or model outputs. Watermarks can aid provenance or detection but have technical limitations.

Z

Zero-shot

Performing a task from instructions without being given task-specific examples in the prompt.