Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values

Watson, Nell; Amer, Ahmed; Harris, Evan; Ravindra, Preeti; Zhang, Shujun

You are here : University of Gloucestershire > Research > Research Repository

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values

Tools

Watson, Nell, Amer, Ahmed, Harris, Evan, Ravindra, Preeti and Zhang, Shujun ORCID: https://orcid.org/0000-0001-5699-2676 (2025) Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values. Information, 16 (8). p. 651. doi:10.3390/info16080651

[thumbnail of 15288 Watson, N. et al (2025) Personalized Constitutionally-Aligned Agentic Superego.pdf]

Preview

Text
15288 Watson, N. et al (2025) Personalized Constitutionally-Aligned Agentic Superego.pdf - Published Version
Available under License Creative Commons Attribution 4.0.
Download (2MB) | Preview

Official URL: https://doi.org/10.3390/info16080651

Abstract

Agentic AI systems, possessing capabilities for autonomous planning and action, show great potential across diverse domains. However, their practical deployment is hindered by challenges in aligning their behavior with varied human values, complex safety requirements, and specific compliance needs. Existing alignment methodologies often falter when faced with the complex task of providing personalized context without inducing confabulation or operational inefficiencies. This paper introduces a novel solution: a ‘superego’ agent, designed as a personalized oversight mechanism for agentic AI. This system dynamically steers AI planning by referencing user-selected ‘Creed Constitutions’—encapsulating diverse rule sets—with adjustable adherence levels to fit non-negotiable values. A real-time compliance enforcer validates plans against these constitutions and a universal ethical floor before execution. We present a functional system, including a demonstration interface with a prototypical constitution-sharing portal, and successful integration with third-party models via the Model Context Protocol (MCP). Comprehensive benchmark evaluations (HarmBench, AgentHarm) demonstrate that our Superego agent dramatically reduces harmful outputs—achieving up to a 98.3% harm score reduction and near-perfect refusal rates (e.g., 100% with Claude Sonnet 4 on AgentHarm’s harmful set) for leading LLMs like Gemini 2.5 Flash and GPT-4o. This approach substantially simplifies personalized AI alignment, rendering agentic systems more reliably attuned to individual and cultural contexts, while also enabling substantial safety improvements.

Item Type:	Article
Article Type:	Article
Uncontrolled Keywords:	Agentic AI systems; AI alignment; Personalization; Ethical guardrails; Superego agent; Constitutional AI; Real-time compliance; Value alignment; AI safety; AI ethics
Subjects:	Q Science > Q Science (General) > Q336 Artificial intelligence Q Science > QA Mathematics > QA75 Electronic computers. Computer science Q Science > QA Mathematics > QA76 Computer software > QA76.9 Other topics > QA76.9.H85 Human-computer interaction Q Science > QA Mathematics > QA76 Computer software > QA76.9 Other topics > QA76.9.C65 Computer simulation
Divisions:	Schools and Research Institutes > School of Business, Computing and Social Sciences
Depositing User:	Kamila Niekoraniec
Date Deposited:	10 Sep 2025 07:12
Last Modified:	30 May 2026 13:45
URI:	https://eprints.glos.ac.uk/id/eprint/15288

University Staff: Request a correction | Repository Editors: Update this record

Altmetric

View Altmetric information about this item.

CORE (COnnecting REpositories)

University Of Gloucestershire

Find Us On Social Media:

Other University Web Sites

Staffnet (Staff Only)

University of Gloucestershire, The Park, Cheltenham, Gloucestershire, GL50 2RH. Telephone +44 (0)844 8010001.

© UoG 2008-24
Knowledge Base
Accessibility
Privacy and Cookies
Disclaimer
Comments concerning this page to Webmaster