AI Safety
AI Safety covers the practices and principles that keep AI systems behaving predictably and under human control, and that limit harm caused by unintended model behaviour.
What is AI Safety?
Definition
AI Safety is the set of practices and principles aimed at keeping AI systems predictable and under human control, both while they are being built and once they are in use. It covers testing model behaviour before deployment, stating plainly what the system must never do, monitoring outputs after release, and keeping a person accountable for consequential decisions. AI safety is a wider question than AI security: security asks who might attack the system, safety asks what the system does when nobody attacks it at all.
Role in cybersecurity
AI Safety plays an important role in building organizational resilience against cyber threats. Implementing appropriate mechanisms in this area is required by regulations such as NIS2, DORA and ISO 27001.