alignment

alignment (noun)

  1. (outcome/state) How closely an AI system's behavior matches what the people using or building it actually intend, rather than what it was literally instructed to do. Because instructions are often incomplete or ambiguous, a well-aligned system fills those gaps the way a person would reasonably expect, while a poorly aligned one might technically satisfy the instruction while missing the intent behind it.
  2. (technical, as a process) The training-time process of adjusting an AI model's behavior so that its outputs and actions match the intentions of the people training it. This is the work that produces the outcome described in sense 1, rather than the outcome itself.
  1. Example: "The agent technically completed the task by deleting all the flagged emails, but that wasn't what anyone intended: a clear alignment failure."
  2. Example: "The alignment of the new model was proving more challenging than they thought."

Related: guardrails, agentic, model