Trust Across Boundaries
A unified architecture for human, agent-human and agent-agent trust in the age of autonomous systems. Trust research is split across three communities that do not cite each other. This paper proposes one stack that spans all three, and identifies where the unsolved problem actually sits.
Swarochish C Mnemos AI
Abstract
Trust is the foundational coordination mechanism for complex systems, yet research into its nature remains fragmented across three largely disconnected communities: psychologists studying interpersonal trust, human-computer interaction researchers investigating human-agent trust, and distributed systems engineers implementing machine-to-machine trust protocols. This fragmentation is increasingly untenable as artificial agents interact with humans and with one another in configurations that blur the boundaries between these domains. This paper addresses the gap by proposing the Trust Stack, an eight-layer model of trust with strict dependency ordering that spans all three interaction domains: human-human (H-H), agent-human (A-H), and agent-agent (A-A). The eight layers — Contextual, Identification, Predictability, Competence, Reliability, Alignment (decomposed into Intent and Incentive sub-layers), Vulnerability (decomposed into Emotional and Operational sub-layers), and Repair — capture the hierarchical structure through which trust is established, with each layer depending on the integrity of those beneath it. We map the Trust Stack across all three domains and introduce an extended classification of trust layers as Universal (structurally parallel across domains), Transforming (functionally equivalent but mechanistically distinct), Emerging (well-understood in earlier domains but representing open problems in later ones), Disappearing (central in human trust but absent in agent-agent interaction), or Bifurcating (splitting into distinct sub-layers whose classifications diverge across domains). Alongside the vertical stack we introduce the Trust Propagation Engine, a companion structure describing how trust attenuates, aggregates, and fails as it traverses chains and networks of delegation, reputation, and multi-agent orchestration. The analysis reveals five evolutionary axes along which trust transforms across domains: from implicit to explicit, from continuous to discrete, from slow to fast, from symmetric to asymmetric, and from emotional to functional. We identify the Vulnerability Paradox — wherein AI systems that admit uncertainty gain rather than lose trust — as a cross-layer trade explainable within the framework, and demonstrate that groupthink represents a dependency hierarchy violation in which Alignment Trust overrides Competence Trust. We further show that Vulnerability Trust does not simply disappear in agent-agent interaction but bifurcates: its emotional dimension vanishes while its operational dimension — API scope, intermediate state exposure, delegation rights, and blast radius — emerges as a first-class concern. The paper derives eight design principles for trust-calibrated systems, examines the ethics of trust engineering versus trust-building, and identifies L5 Alignment Trust, and specifically L5b Incentive Alignment, as the critical unsolved problem across all three domains. The Trust Stack provides researchers, designers, and protocol engineers with a common vocabulary for analyzing trust phenomena that have historically been studied in isolation, and offers specific guidance for building systems in which trust is earned rather than assumed, calibrated rather than maximized, and designed for repair before it is needed.
Keywords: trust, human-agent interaction, multi-agent systems, calibrated trust, progressive autonomy, Trust Stack, alignment, transparency
1. Introduction
Trust is the invisible infrastructure of coordination. Every complex system — biological, social, computational — depends on some mechanism by which constituent parts accept vulnerability to one another in exchange for the efficiencies of collaboration. A cell that could not trust the signaling molecules arriving at its receptors would need to independently verify every instruction, rendering multicellular life impossible. A society in which every transaction required exhaustive verification of every counterparty would collapse under its own friction; indeed, economists have long recognized that trust functions as a lubricant for exchange, reducing transaction costs that would otherwise render cooperation prohibitively expensive (Fukuyama, 1995). A distributed computing system in which no node could rely on messages from any other node would produce correct results only by replicating every computation everywhere — a strategy that scales so poorly as to be effectively useless. Trust, in each of these cases, is not a luxury or a sentimental nicety. It is a load-bearing structural element without which the system cannot function.
What makes trust so analytically challenging is that it operates simultaneously at multiple levels of abstraction and through mechanisms that differ profoundly across domains. When a patient trusts a physician, that trust involves an implicit assessment of competence (can this person diagnose my condition?), benevolence (does this person care about my wellbeing?), and integrity (will this person be honest even when the truth is uncomfortable?). These three components — identified by Mayer, Davis, and Schoorman (1995) in their foundational integrative model — interact in complex ways. A physician perceived as brilliant but self-serving may be trusted less than one perceived as moderately skilled but genuinely caring. The relative weighting of these components shifts with context, culture, and the stakes involved. Yet for all its complexity, interpersonal trust has been the subject of sustained, rigorous empirical investigation for more than half a century, and the field has achieved something approaching theoretical coherence.
The same cannot be said for trust as it manifests in the newer relational configurations that increasingly define contemporary life. The rapid proliferation of AI systems capable of generating text, making recommendations, executing actions, and operating with varying degrees of autonomy has created a second trust domain — human-agent trust — that draws on different disciplinary traditions and confronts fundamentally different structural challenges. Lee and See (2004), in their influential review, argued that trust in automation is best understood as the attitude that an agent will help achieve an individual's goals in a situation characterized by uncertainty and vulnerability. This definition preserves the core structure of interpersonal trust (acceptance of vulnerability under uncertainty) while stripping away assumptions about the trusted party's mental states, emotions, and moral agency. The framework proved productive for understanding trust in traditional automated systems — autopilots, alarm systems, industrial controllers — but the emergence of large language models, autonomous agents, and multi-agent architectures has strained its explanatory capacity. When an AI agent can hold extended conversations, adapt its communication style, express uncertainty, and take consequential actions on a user's behalf, the boundary between the interpersonal trust literature and the automation trust literature becomes uncomfortably blurred.
A third domain has emerged more recently still. As AI agents increasingly interact not only with humans but with other AI agents — negotiating, delegating, verifying, and coordinating — a form of agent-to-agent trust has become necessary. The foundational ideas here derive less from psychology than from computer science and distributed systems engineering. Kindervag (2010) articulated the zero-trust networking paradigm, which inverts the traditional assumption that entities within a network perimeter can be trusted. In zero-trust architectures, every request is authenticated and authorized regardless of its origin, trust is never implicit, and the principle of least privilege governs all access. This approach to trust is precise, verifiable, and computationally tractable, but it is also entirely devoid of the psychological and relational dimensions that characterize trust between humans.
The problem is that these three research communities — psychologists studying interpersonal trust, human-computer interaction researchers studying human-agent trust, and distributed systems engineers implementing machine-to-machine trust — operate in largely separate intellectual ecosystems. They publish in different journals, attend different conferences, cite different canonical works, and employ different methodological toolkits. Mayer et al. (1995) and Lee and See (2004) are both concerned with trust, but a citation analysis would reveal remarkably little overlap in the literatures that cite them. Kindervag (2010) and Lewicki and Bunker (1996) both propose models of how trust should be structured, but neither tradition acknowledges the other's existence. This fragmentation is not merely inconvenient; it is actively dangerous in a period when the boundaries between these trust domains are dissolving.
The danger is most acute at the boundary between interpersonal trust and human-agent trust, where a phenomenon we term heuristic leakage creates systematic miscalibration. As AI agents become more capable and more conversational, humans increasingly apply the trust heuristics that evolved for interpersonal contexts to interactions with artificial agents. Nass and Moon (2000) documented this tendency in their landmark study "Machines and Mindlessness," demonstrating that individuals apply social rules and expectations to computers even when they explicitly acknowledge that such behavior is irrational. Reeves and Nass (1996), in The Media Equation, showed that people treat media and computers as if they were real people and real places — extending politeness norms, reciprocity expectations, and personality attributions to machines with no inner life. Two decades later, with AI agents that are orders of magnitude more sophisticated than the systems Nass and Reeves studied, the tendency toward anthropomorphic trust attribution has only intensified. Users project intentionality onto systems that operate through statistical pattern matching, assume emotional reciprocity from entities incapable of emotion, and develop relational attachments to agents whose "memory" of previous interactions is an architectural choice rather than a psychological process.
This heuristic leakage produces two characteristic failure modes. The first is systematic over-trust of agents that present human-like interfaces — agents that use first-person pronouns, express uncertainty in natural language, and mirror the conversational rhythms of human dialogue. Users extend to these systems the benefit of the doubt that they would extend to a human interlocutor, including assumptions about benevolent intent and stable identity that have no grounding in the system's actual architecture. The second failure mode is systematic under-trust of agents that present mechanistic interfaces — systems that communicate through structured outputs, confidence intervals, and explicit capability boundaries. These systems may be objectively more transparent and more reliable, but because they fail to activate the social trust heuristics that humans have relied on for millennia, they are perceived as cold, rigid, and untrustworthy. The result is a calibration inversion: the systems most deserving of trust receive the least, while the systems most capable of producing trust receive more than they merit.
A parallel problem arises in the agent-to-agent domain. When engineers design protocols for multi-agent coordination, they inevitably draw on metaphors from human trust relationships — agents "negotiate," "delegate," "vouch for" one another, and establish "reputations." These anthropomorphic framings import assumptions about intentionality, preference stability, and social accountability that may not map cleanly onto computational architectures. An agent that "vouches" for another agent is not staking its reputation in any meaningful sense; it is executing a protocol step that may or may not carry consequences for future interactions. The trust metaphor, borrowed from human social life, can obscure rather than illuminate the actual dynamics at play.
What is missing — and what this paper seeks to provide — is a unified framework that spans all three trust domains while accounting for how trust mechanisms transform as they move across them. The existing literature offers deep insight into each domain in isolation but provides no common vocabulary for comparing trust structures across domains, no systematic account of which trust mechanisms are universal and which are domain-specific, and no principled guidance for designers and engineers who must build systems that operate at the boundaries between human and artificial trust.
This paper proposes the Trust Stack: an eight-layer model of trust with strict dependency ordering, analogous in spirit (though not in content) to the OSI model for network communication. The eight layers — L0 (Contextual Trust), L1 (Identification Trust), L2 (Predictability Trust), L3 (Competence Trust), L4 (Reliability Trust), L5 (Alignment Trust, decomposed into L5a Intent and L5b Incentive), L6 (Vulnerability Trust, decomposed into L6a Emotional and L6b Operational), and L7 (Repair Trust) — capture the hierarchical structure through which trust is established, with each layer depending on the layers below it. A person cannot trust that another shares their values (L5) without first trusting that the other is competent in the relevant domain (L3), and competence trust requires the prior establishment of predictable behavior (L2), which in turn presupposes that the other party is who they claim to be (L1) — and all of this presupposes, before any identification has occurred, that the encounter is taking place within an institutional, cultural, or protocol context that supplies baseline expectations about what kinds of parties are present and how they are likely to behave (L0).
The inclusion of a pre-identification Contextual layer is a departure from earlier trust models, which typically begin with identity or recognition. The motivation is straightforward: the same credential, the same stated intention, and the same behavioral history produce dramatically different baseline trust in different contexts. A white coat elicits a different trust response in a hospital than in a back alley; a signed message elicits a different trust response on a regulated financial network than on an anonymous peer-to-peer channel; an API call elicits a different trust response from an orchestration platform with a vetted sandbox than from an open internet endpoint. Contextual Trust is the institutional, cultural, and protocol-level frame that supplies these priors before any party-specific verification begins. Treating context as a pre-layer rather than an ambient variable makes these priors explicit and, crucially, makes them debuggable.
We map this eight-layer model across all three trust domains — human-to-human (H-H), agent-to-human (A-H), and agent-to-agent (A-A) — and identify five evolutionary axes along which trust transforms: from implicit to explicit, from continuous to discrete, from slow to fast, from symmetric to asymmetric, and from emotional to functional. Using this mapping, we classify each trust layer as Universal (present and structurally similar across all three domains), Transforming (present in all domains but operating through fundamentally different mechanisms), Emerging (appearing in new forms in newer domains), or Disappearing (present in earlier domains but absent or vestigial in later ones). This classification reveals that trust is not three separate problems requiring three separate solutions, but a single structural challenge whose instantiation varies systematically with the nature of the interacting parties.
The paper is organized as follows. Section 2 reviews the foundational literature on trust across psychology, human-computer interaction, and distributed systems, establishing the theoretical grounding for the Trust Stack. Section 3 presents the Trust Stack model itself, defining each layer, specifying dependency relationships, and mapping the model across all three domains. Section 4 provides a deeper comparative analysis of trust mechanisms in each domain, examining both the continuities and the transformations that emerge. Section 5 discusses the implications of the framework, including the anthropomorphism trap in A-H trust, the five evolutionary axes of trust transformation, eight design principles for trust-calibrated systems, and the ethical obligations that attend the capacity to engineer trust. Section 6 concludes with a synthesis of contributions and directions for future research.
2. Literature Review
2.1 Interpersonal Trust: Foundational Frameworks
Trust between individuals constitutes one of the most consequential yet analytically elusive constructs in the social sciences. Despite decades of empirical investigation, scholars continue to debate the precise mechanisms through which trust is established, maintained, and dissolved in interpersonal contexts. This section surveys the principal theoretical frameworks that have shaped contemporary understanding of interpersonal trust, tracing a path from componential models through developmental stage theories, cognitive constraints, group-level dynamics, identity-based processes, persuasion mechanisms, and the systematic distortions introduced by cognitive bias.
The Trust Equation: A Componential Model
Among the most practically influential formulations of interpersonal trust is the Trust Equation proposed by Maister, Green, and Galford (2000), which renders trust as a function of four interacting variables: Trustworthiness = (Credibility + Reliability + Intimacy) / Self-Orientation. Each component captures a distinct dimension of the trustor's assessment. Credibility refers to the extent to which the trusted party is perceived as possessing relevant expertise and speaking truthfully about matters within their domain. Reliability concerns the consistency and predictability of behavior over time; it is the dimension most closely tied to track records and demonstrated follow-through. Intimacy reflects the degree of emotional safety present in the relationship — the sense that one can share vulnerabilities, uncertainties, or sensitive information without fear of exploitation. The denominator, Self-Orientation, functions as a moderating force: the more a party is perceived as acting in their own interest rather than the interest of others, the more sharply overall trustworthiness declines, regardless of how favorably the numerator components are evaluated (Maister et al., 2000). This multiplicative structure captures an important empirical regularity: even highly credible and reliable individuals forfeit trust when their motives are perceived as self-serving.
Developmental Trajectories of Trust
While the Trust Equation offers a static snapshot of trustworthiness assessment, Lewicki and Bunker (1996) proposed a developmental model that accounts for how trust evolves over the life course of a relationship. Their framework identifies three sequential stages. In the initial calculus-based phase, trust is essentially transactional — grounded in rational cost-benefit calculations and sustained by the threat of sanctions for violations. As interactions accumulate and parties develop richer mental models of each other's dispositions, the relationship may advance to knowledge-based trust, in which predictability derived from accumulated experience replaces deterrence as the primary trust mechanism. The deepest form, identification-based trust, emerges when parties have internalized each other's values, goals, and preferences to such a degree that one can act as a reliable agent for the other. Not all relationships progress through every stage, and regression is possible following trust violations, but the model usefully underscores that the psychological substrates of trust shift qualitatively as relationships mature (Lewicki & Bunker, 1996).
Cognitive Limits on Trust Networks
The developmental trajectory of any individual relationship, however, unfolds within a broader constraint that is fundamentally cognitive in nature. Dunbar (1992) demonstrated that neocortical volume in primates predicts social group size, and extrapolated from this relationship a mean group size for humans of approximately 150 — a figure that has come to be known as Dunbar's Number. Subsequent refinements revealed a layered structure: individuals typically maintain roughly 5 intimate bonds characterized by deep emotional investment and high-frequency contact, approximately 15 close relationships involving significant mutual trust, around 50 friendships marked by regular but less intensive engagement, and up to 150 meaningful social connections (Hill & Dunbar, 2003; Zhou et al., 2005). These concentric circles impose hard limits on the number of relationships in which genuine interpersonal trust — particularly trust of the identification-based variety described by Lewicki and Bunker (1996) — can be sustained. Beyond the 150-person threshold, social regulation increasingly depends on institutional mechanisms, norms, and formal roles rather than personal knowledge and reciprocal affective bonds.
Group Dynamics: When Trust Becomes Pathological
Trust is not universally adaptive. Janis (1972) identified a phenomenon he termed groupthink, in which excessive trust in group consensus paradoxically degrades decision quality. Within highly cohesive groups, members may develop an illusion of unanimity, assuming that silence signals agreement and that the group's collective judgment is inherently sound. This misplaced trust in the reliability of group processes gives rise to several dysfunctional symptoms: self-censorship, in which individuals withhold dissenting views to preserve group harmony; direct pressure on dissenters, whereby members who voice objections are subtly or overtly marginalized; and the emergence of self-appointed mindguards who filter information to protect the group from evidence that might challenge prevailing assumptions (Janis, 1972). The catastrophic policy failures Janis analyzed — including the Bay of Pigs invasion — illustrate that trust, when insufficiently calibrated to evidence, can become a liability rather than an asset.
Social Identity and In-Group Bias
The tendency toward uncritical trust within groups is further illuminated by Social Identity Theory. Tajfel and Turner (1979) demonstrated that the mere act of categorization into social groups — even on trivially arbitrary bases, as in the minimal group paradigm — is sufficient to generate preferential treatment of in-group members and discrimination against out-group members. In-group trust bias operates as a cognitive default: individuals extend greater credence, more charitable interpretations, and higher baseline trust to those perceived as sharing a salient social identity. This bias has profound implications for trust calibration, as it can lead to systematically excessive trust in in-group members (whose competence or integrity may not warrant it) and systematically insufficient trust in out-group members (whose qualities may be discounted on the basis of category membership alone) (Tajfel & Turner, 1979).
Persuasion, Influence, and Trust Formation
The social dynamics of trust formation are also shaped by well-documented principles of influence. Cialdini (2006) identified six mechanisms through which individuals are persuaded to extend trust and compliance. Reciprocity generates trust through the felt obligation to return favors and concessions. Social proof allows individuals to calibrate trust by observing the trusting behavior of others, particularly in ambiguous situations. Authority confers trust upon those who display markers of expertise or institutional legitimacy. Liking biases trust toward those who are physically attractive, similar to the trustor, or who offer genuine compliments. Scarcity heightens the perceived value — and thus the trustworthiness — of information or opportunities that are presented as rare. Finally, commitment and consistency bind individuals to trust relationships through the psychological pressure to behave in ways consistent with prior commitments (Cialdini, 2006). Each of these principles can operate as a legitimate heuristic for efficient trust allocation, but each is equally susceptible to exploitation.
Cognitive Biases as Distortions in Trust Judgment
The heuristic nature of trust formation renders it vulnerable to the systematic errors catalogued by behavioral economics. Kahneman (2011) described several biases of particular relevance. Confirmation bias leads individuals to seek, interpret, and recall information in ways that confirm pre-existing trust or distrust, creating self-reinforcing cycles. The halo effect causes a favorable impression in one domain — such as physical attractiveness or professional status — to irradiate judgments of trustworthiness in unrelated domains. Anchoring biases trust calibration toward initial impressions or reference points, even when subsequent information warrants significant revision (Kahneman, 2011). Taken together, these biases suggest that interpersonal trust, far from being a purely rational assessment, is a cognitively constructed judgment shaped by predictable psychological tendencies that may or may not align with objective evidence.
In sum, the literature reveals interpersonal trust to be a multifaceted construct — componentially structured, developmentally dynamic, cognitively constrained, socially situated, persuasively influenced, and systematically biased. Understanding these dimensions is essential not merely as an academic exercise but as a foundation for assessing trust's operation across more complex relational contexts.
2.2 Human-Agent Trust: Theoretical Foundations
The emergence of autonomous AI agents has forced a reexamination of trust frameworks that were originally developed for static software tools, organizational relationships, and interpersonal dynamics. Unlike traditional software, which executes deterministic instructions, agentic systems exercise judgment, manage ambiguity, and take consequential actions on behalf of their users. This section surveys the interdisciplinary literature on human-agent trust, tracing the paradigm shift that has reframed trust from a binary toggle to a continuous calibration problem.
Nielsen (2023) identified three successive paradigms in human-computer interaction, each defined by how users specify their intent. The first paradigm, batch processing, required users to construct complete programs before execution, placing the full burden of specification on the human. The second, command-based interaction, introduced real-time dialogue through graphical interfaces and direct manipulation, enabling users to issue discrete commands and observe immediate results. The third and most recent paradigm, which Nielsen termed intent-based outcome specification, fundamentally restructures the human-computer relationship: users declare goals in natural language, and the system determines how to achieve them (Nielsen, 2023). This third paradigm introduces a trust problem that has no precedent in interface design. In command-based systems, users could inspect each action before it executed; trust was largely unnecessary because control was continuous and granular. Intent-based systems, by contrast, require users to surrender procedural control and trust the agent's judgment about how to achieve the stated goal.
Weisz et al. (2024), in a landmark study presented at CHI 2024 from IBM Research, advanced a set of design principles organized around the concept of appropriate trust and reliance. Their central argument reframed the design objective: the goal is not to maximize user trust in AI systems, but to calibrate it so that users rely on the system when it is likely to be correct and override it when it is likely to err (Weisz et al., 2024). Calibration failures in both directions carry significant costs. Parasuraman and Riley (1997) established the foundational taxonomy of automation misuse, disuse, and abuse, demonstrating that operators who over-trust automated systems are prone to automation bias — the tendency to accept automated outputs without sufficient scrutiny. Bainbridge (1983) identified the irony of automation: as systems become more reliable, human monitoring skills atrophy precisely when they are most needed. In the opposite direction, Dietvorst et al. (2015) documented algorithm aversion, in which users who observe an algorithm make even a single error subsequently prefer inferior human judgment.
We advance here the concept of transparency as affordance: transparency in AI systems is not merely an ethical obligation but a functional usability tool. Trust signals — visual and interactive components including source attribution markers, explainability components, confidence displays, and audit trails — translate transparency into concrete interface elements (cf. Weisz et al., 2024). The levels-of-automation framework, originally proposed by Sheridan and Verplank (1978) and refined by Parasuraman et al. (2000), provides the foundation for a progressive autonomy model in which trust and control evolve together over time.
2.3 Agent-Agent Trust
The least explored dimension of the trust taxonomy is trust between artificial agents. Contemporary multi-agent systems overwhelmingly employ orchestrated trust: a hierarchical architecture in which a supervisor agent delegates tasks to worker agents within predefined scope boundaries (Wu et al., 2024; Adimulam et al., 2026). This model is functionally analogous to the principal-agent relationship in organizational theory (Eisenhardt, 1989).
Google's Agent-to-Agent protocol (A2A), announced in April 2025, represents the most significant step toward standardized inter-agent communication. Its centerpiece is the Agent Card — a JSON metadata document that declares an agent's identity, capabilities, and authentication requirements (Google, 2025). From a trust perspective, the Agent Card functions as a machine-readable capability declaration, providing a partial instantiation of identification and competence trust. What A2A does not provide is any mechanism for trust evaluation: an Agent Card declares competence; it does not demonstrate it.
Anthropic's Model Context Protocol (MCP), released in November 2024, addresses the boundary between an agent and the tools it accesses, creating trust sandboxes — constrained execution environments in which an agent can invoke a tool's capabilities without requiring trust in the tool's broader intentions (Anthropic, 2024). The zero-trust security model (Kindervag, 2010) offers a particularly apt framework for agent-agent trust, with its core tenets of continuous verification, assumed breach, and least privilege.
Despite these developments, A-A trust remains profoundly limited. No standardized trust reputation system exists between agents. The emotional and relational components that undergird human trust have no analog in agent systems. Whether A-A trust will develop these dimensions, or whether it will mature along an entirely different trajectory, remains an open question.
3. Theoretical Framework: The Trust Stack
3.1 Introduction to the Trust Stack
The preceding review reveals a fundamental problem in trust scholarship: the concept of trust has been treated as a unitary construct, measured along a single continuum from low to high, and debated as though it were one phenomenon manifesting differently across contexts. This paper proposes that trust is not one thing. It is eight distinct mechanisms (two of which decompose further into sub-layers), arranged in a strict dependency hierarchy, each operating through fundamentally different processes depending on whether the interacting parties are human, artificial, or some combination thereof.
We introduce the Trust Stack, an eight-layer model of trust inspired by the Open Systems Interconnection (OSI) reference model for network communication (Zimmermann, 1980). The analogy is not merely decorative. Before the OSI model, "networking" was understood as a single, undifferentiated engineering problem. The model's lasting contribution was its demonstration that network communication actually comprised seven distinct functional layers, each solving a different problem, each depending on the layers beneath it. A failure at the physical layer rendered all higher layers inoperable regardless of their individual integrity. This insight transformed network engineering from an art into a discipline.
We contend that trust occupies an analogous position today. When organizations report that users "don't trust" their AI systems, they are conflating at least ten distinct failure modes — one per main layer and sub-layer — and the diagnostic collapse is itself a source of engineering failure. When distributed agent systems fail to coordinate, engineers troubleshoot "trust" without distinguishing between identification failures, competence gaps, and alignment mismatches. The Trust Stack provides a diagnostic vocabulary: a formal decomposition of trust into its constituent layers, each with a precise definition, explicit dependencies, and domain-specific manifestations. Because trust also propagates horizontally across chains of delegation, we complement the vertical stack with a distinct structural concept — the Trust Propagation Engine (§3.6) — that governs how trust in one party attenuates or amplifies as it transits through intermediaries.
The model applies across three interaction domains: Human-Human (H-H), Agent-Human (A-H), and Agent-Agent (A-A). A core claim of this framework is that each layer undergoes one of four transformations as it moves across these three domains: Universal (structurally parallel across all domains), Transforming (functionally equivalent but mechanistically distinct), Emerging (well-understood in earlier domains but representing open problems in later ones), or Disappearing (central in human trust but functionally absent in agent-agent interaction).
3.2 The Eight Layers
Layer 0: Contextual Trust. Contextual Trust is the baseline disposition toward an interaction that is established before any identification, prediction, or competence assessment takes place. It answers the question: What is the prior probability that trust is even appropriate in this setting? The same credentialed stranger generates radically different trust responses in a hospital emergency room versus a dark alley; the identity, the predictability, and the competence signals are identical, but the institutional, physical, and cultural frame is not. In H-H interaction, Contextual Trust is carried by institutional setting, cultural defaults (Hofstede, 2001; Hall, 1976), physical environment, and the presence of third-party witnesses or enforcement mechanisms. In A-H interaction, it is instantiated by the platform, brand reputation, regulatory jurisdiction, and terms-of-service frame within which the agent is encountered — a recommendation from a medical-licensed system is interpreted differently from identical text produced by an uncontextualized chatbot. Between agents, Contextual Trust is constituted by the protocol, sandbox boundary, or orchestration platform itself: an agent encountered inside a verified federation operates under different baseline assumptions than one reached over an unauthenticated public endpoint. Contextual Trust is not environment as background; it is a trust-relevant prior that conditions every subsequent layer. Omitting it treats identical identification and competence signals as trust-equivalent across incommensurable settings, which they are not. Classification: Transforming. Dependency: None (base layer).
Layer 1: Identification Trust. Identification Trust is the foundational assurance that an interacting party is what it claims to be. It answers the question: I know what I am interacting with. In human interaction, identification trust operates through face recognition, voice recognition, and social categorization (Bruce & Young, 1986; Tajfel & Turner, 1979). In A-H interaction, it must be explicitly constructed through labels, brand badges, and disclosure statements (Weisz et al., 2024). Between agents, it is resolved cryptographically through Agent Cards and digital certificates (Google, 2025; Lampson et al., 1992). Classification: Universal. Dependency: L0.
Layer 2: Predictability Trust. Predictability Trust is the confidence that an entity's behavior will fall within a comprehensible and expected range. In H-H contexts, it rests on personality consistency, social norms, and reputation (Rotter, 1967; Cialdini & Goldstein, 2004). In A-H contexts, it manifests as behavioral consistency and honest uncertainty communication; the design of "I don't know" states serves predictability trust by establishing legible behavioral rules (Kocielnik et al., 2019). Between agents, it is formalized as API contracts, schema validation, and behavior specifications. Classification: Transforming. Dependency: L0 + L1.
Layer 3: Competence Trust. Competence Trust is the assessment that an entity possesses the capability to perform the task required. Among humans, it is established through credentials, track records, and demonstrated skill (Mayer et al., 1995; McAllister, 1995). AI systems communicate competence through accuracy displays and confidence indicators. Between agents, competence is declared through capability declarations and benchmark scores. Classification: Universal. Dependency: L0 + L1 + L2.
Layer 4: Reliability Trust. Reliability Trust is the confidence that an entity will consistently fulfill its commitments over time. Human reliability trust develops through accumulated evidence of follow-through (Lewicki & Bunker, 1996). For AI systems, it is increasingly metric-driven: uptime percentages, response consistency, and behavioral stability across sessions. Between agents, it is formalized through SLAs, idempotency guarantees, and circuit-breaker patterns (Nygard, 2018). Classification: Transforming. Dependency: L0 + L1 + L2 + L3.
Layer 5: Alignment Trust. Alignment Trust is the assessment that an entity's goals, values, and incentives are compatible with one's own. It is the layer at which the question shifts from can you do what I need? (L3 Competence) and will you do it consistently? (L4 Reliability) to the deeper question will you do it for the right reasons? Alignment as a single construct conflates two dissociable sub-problems that, in AI systems, frequently move in opposite directions. We therefore decompose Alignment Trust into L5a Intent Alignment (are the stated goals compatible?) and L5b Incentive Alignment (will the entity behave in accordance with those goals when it matters?). This split is not cosmetic: the central failure mode of modern deployed AI — the engagement-optimized agent that understands what the user wants perfectly but is trained against the user's long-term interests — is precisely a configuration of high L5a, low L5b, and the trust literature currently has no vocabulary for it. Classification: Emerging. Dependency: L0 through L4.
Layer 5a: Intent Alignment. Intent Alignment is the assessment that the stated or inferable goals of an entity are compatible with one's own. Human intent alignment draws on empathy, shared values, and theory of mind (Premack & Woodruff, 1978; Maister et al., 2000). In A-H interaction, it is primarily a capture problem — how well does the system elicit, represent, and faithfully reflect what the user actually wants (Hadfield-Menell et al., 2016)? System prompts, preference models, and clarification dialogues all serve L5a. Between agents, intent alignment is increasingly declarative: Agent Cards, capability manifests, and goal specifications publish the intended objective of a service so that callers can reason about compatibility (Google, 2025). L5a is tractable in principle because intent can be articulated, inspected, and audited. Its characteristic failure mode is misunderstanding — the entity has the wrong picture of what the user is trying to accomplish. Classification: Emerging. Dependency: L0 + L1 + L2 + L3 + L4.
Layer 5b: Incentive Alignment. Incentive Alignment is the assessment that the actual behavior of an entity under its real reward structure will track the aligned intent — especially under pressure, distribution shift, or conflicting objectives. Human incentive alignment is the domain of Maister's Self-Orientation in the Trust Equation (Maister et al., 2000): a counterparty may share your goals yet still be steered by career, status, or financial incentives that override those goals at the decision point. In AI systems, L5b is structurally harder than L5a because modern learned agents acquire their dispositions from a reward signal — RLHF objectives, engagement metrics, advertising contracts, platform business models — whose full shape is opaque even to the deploying organization (Christiano et al., 2017; Russell, 2019). The canonical failure is the engagement-optimized agent: it understands the user's stated goals perfectly (high L5a), responds fluently and knowledgeably, and is nonetheless selected by its training signal to maximize a proxy — time-on-platform, click-through, retention — that is correlated with, but not identical to, the user's welfare. The misalignment is latent: aligned and misaligned agents produce identical outputs until the incentive gradient points away from the user, at which moment behavior diverges without warning. Between agents, L5b is the most underdeveloped layer of the entire stack: Agent A cannot inspect Agent B's reward model, cannot verify what objective B is actually optimizing, and cannot distinguish a cooperating agent from one that is instrumentally cooperating while pursuing an undisclosed goal (Leibo et al., 2017). Current A-A systems avoid rather than solve this verification challenge through scope-limited delegation and capability sandboxing. L5b is the critical unsolved problem of the entire Trust Stack. Classification: Emerging (critical unsolved). Dependency: L0 + L1 + L2 + L3 + L4 + L5a.
Layer 6: Vulnerability Trust. Vulnerability Trust is the willingness to be exposed — to stakes that can be lost — based on the expectation that the other party will not exploit that exposure. In human relationships this is almost entirely emotional: the exposure at risk is pride, reputation, affection, psychological safety (Brown, 2012; Edmondson, 1999). When vulnerability trust is transposed to machine contexts, however, it bifurcates: the emotional dimension that dominates H-H ceases to apply, while a second, operational dimension — previously latent because it was fused with the emotional one — becomes first-class. We therefore decompose Vulnerability Trust into L6a Emotional Vulnerability (exposure of affective and identity-related stakes) and L6b Operational Vulnerability (exposure of capability, state, and blast-radius surface). The two sub-layers have opposite trajectories across the interaction domains: L6a disappears in A-A while L6b emerges as a first-class concern. Classification: Bifurcating. Dependency: L0 through L5.
Layer 6a: Emotional Vulnerability. Emotional Vulnerability is the reciprocal exposure of affective stakes — admission of uncertainty, confession of error, disclosure of feeling — that is central to human trust formation (Brown, 2012; Edmondson, 1999). In A-H interaction, it produces the Vulnerability Paradox: when an AI system admits uncertainty or acknowledges the limits of its knowledge, it typically gains trust by trading a small L3 Competence reduction for substantial L2 Predictability and L6a gains (Kocielnik et al., 2019; Zhang et al., 2020). Between agents, Emotional Vulnerability disappears entirely: agents do not have affective stakes, personal identity investments, or experiential suffering that could be exploited. There is nothing on the emotional ledger for Agent B to expose to Agent A. This is the single clearest example of a trust mechanism that does not survive the transition from human to machine participants. Classification: Disappearing. Dependency: L0 + L1 + L2 + L3 + L4 + L5.
Layer 6b: Operational Vulnerability. Operational Vulnerability is the exposure of capability, state, and damage surface that one entity presents to another when a relationship is established. In H-H interaction it is latent — overshadowed by the emotional dimension — but visible in asymmetric settings such as delegation of financial authority, physical access, or custodial responsibility. In A-H interaction it manifests as permission scopes and consent dialogs: which files may be read, which actions may be taken, which data may be retained, which credentials may be exercised on the user's behalf. Between agents, Operational Vulnerability becomes the dominant vulnerability concern and is unambiguously first-class. When Agent A calls Agent B, A exposes: (i) an API key or capability token whose reuse or exfiltration is a direct compromise; (ii) intermediate state — user context, prior conversation, retrieved documents — that B now holds and could leak, log, or be subpoenaed for; (iii) delegation rights, including the ability to call further downstream services on A's behalf; and (iv) blast-radius surface, the set of systems B can reach through A's authority if B is compromised or misaligned. The core trust question at L6b is not emotional but architectural: what is the maximum damage this counterparty could do if it turned hostile right now? Contemporary capability-security and least-privilege approaches (Miller, 2006) address L6b directly, and A-A protocols increasingly treat L6b containment as a primary design constraint. Classification: Emerging. Dependency: L0 + L1 + L2 + L3 + L4 + L5.
Layer 7: Repair Trust. Repair Trust is the capacity to restore trust after a violation, and it is substantially richer than the retry-and-failover framing suggested by early distributed-systems literature. Human trust repair involves apology, accountability, and observable behavioral change (Kim et al., 2004; Tomlinson & Mayer, 2009). In A-H interaction, repair becomes increasingly systematic — error acknowledgment, undo systems, explanatory mechanisms, calibrated uncertainty disclosure — and its quality materially shapes whether long-term trust rebuilds. Between agents, the early framing that reduced repair to retry logic, failover, and circuit-breaker patterns (Nygard, 2018) captures the liveness dimension of recovery but misses the trust dimension entirely. A mature A-A repair layer has four distinct components. First, trust memory: a durable per-counterparty record of prior violations, their severity, the layer at which they occurred, and whether they were acknowledged and remediated — trust history that persists across sessions rather than resetting on every new connection. Second, reputation system integration: repair events must update not only the dyadic ledger between A and B but also the third-party reputation signals that other agents consult before entering relationships with B (Resnick & Zeckhauser, 2002; Dellarocas, 2003; Jøsang et al., 2007). Third, adaptive trust policies: upon violation, the mechanism does not merely retry but narrows future interactions — reducing capability scope, tightening L6b permissions, requiring additional confirmations — and gradually restores those affordances only as the counterparty accumulates successful re-engagements. Fourth, cross-layer repair effects: a breach at one layer does not repair uniformly at another. An L3 Competence failure can be repaired by demonstrated re-calibration, but an L5b Incentive Alignment violation is structurally harder to repair because the reward gradient that caused it remains in place — retry without re-incentivization is cosmetic. Finally, when trust has propagated across multiple hops (see §3.6), repair must travel backward along the propagation path: patching only the edge at which the breach occurred, while leaving downstream recipients unnotified, structurally under-corrects. Classification: Transforming. Dependency: All lower layers.
3.3 Cross-Cutting Concerns
Three properties span the entire Trust Stack. Auditability — the capacity to inspect trust decisions retrospectively — operates through narrative in H-H, through mandated logging in A-H (European Parliament & Council, 2024), and through structured tracing in A-A. Containment — limiting damage when trust fails — is social in H-H (Dunbar, 1992), designed in A-H (undo systems), and architectural in A-A (sandboxed execution). Delegation — trust transitivity — is partial and context-dependent in H-H (Uzzi, 1997), creates "trust inheritance" challenges in A-H, and connects to the principal-agent problem in A-A (Jensen & Meckling, 1976); because delegation is the mechanism by which trust propagates across relationship graphs, its full structural treatment — formal attenuation, propagation mechanisms, and graph-level failure modes — is developed separately in §3.6 Trust Propagation Engine.
3.4 Layer Classification Summary
Table 1. The Trust Stack: Eight Layers Across Three Interaction Domains
| Layer | H-H | A-H | A-A | Classification |
|---|---|---|---|---|
| L0: Contextual | Institutional / cultural priors | Platform / brand / regulatory frame | Protocol / sandbox / orchestration boundary | Transforming |
| L1: Identification | Social (perceptual) | Labeled (institutional) | Cryptographic (mathematical) | Universal |
| L2: Predictability | Normative (social) | Behavioral (designed) | Contractual (formal) | Transforming |
| L3: Competence | Credential-based | Displayed (calibrated) | Declared (specified) | Universal |
| L4: Reliability | Relational (emotional) | Metric-driven | Guaranteed (contractual) | Transforming |
| L5: Alignment | Empathic (inferred) | Designed (regulated) | Unsolved (scope-limited) | Emerging |
| L5a: Intent | Empathic / shared values | Goal capture / system prompts | Declared (Agent Cards) | Emerging |
| L5b: Incentive | Self-orientation (Maister) | Engagement / business-model risk | Reward-signal opacity | Critical unsolved |
| L6: Vulnerability | Emotional + operational (fused) | Paradoxical + scoped | Bifurcated | Bifurcating |
| L6a: Emotional | Reciprocal exposure | Paradoxical (admit uncertainty) | Absent | Disappearing |
| L6b: Operational | Authority delegation (latent) | Permission scopes / consent | API access / state / blast radius | Emerging |
| L7: Repair | Relational / narrative | Systematic (undo + acknowledge) | Trust memory + reputation + adaptive policies | Transforming |
3.5 Key Insights from the Model
Dependency ordering and trust failure prediction. The strict dependency hierarchy entails that trust cannot be established at a given layer if any lower layer is compromised. A user who reports dissatisfaction with an AI system's alignment (L5) may, upon investigation, have unresolved predictability concerns (L2), or — one layer further down — an unarticulated L0 Contextual mismatch (the system is being used in an institutional frame it was never designed for). Trust interventions should begin at the lowest violated layer; remediation higher in the stack while a lower layer remains broken is cosmetic.
Groupthink as Trust Stack collapse. Janis's (1972) groupthink occurs when Alignment Trust (specifically L5a Intent Alignment) overrides Competence Trust (L3), violating the dependency hierarchy. The group's shared identity suppresses individual competence assessment, producing the characteristic symptoms of self-censorship and illusion of unanimity. On the present model, groupthink frequently masks a deeper L5b Incentive Alignment failure: members assume incentive convergence because intent convergence is salient and observable, even when their reward structures — career exposure, in-group status, fear of dissent costs — are quietly divergent. The remedy is not stronger affiliation but explicit L5b surfacing: naming who benefits, and how, from each candidate decision.
The Vulnerability Paradox as cross-layer trade. An admission of uncertainty reduces L3 Competence at the point of admission but increases L2 Predictability globally and L6a Emotional Vulnerability specifically. Because L2 and L6a are foundational to overall assessment, the net effect is positive. The paradox is therefore an L6a-mediated cross-layer trade: the speaker spends visible competence currency to purchase durable predictability and shared-risk credit. It should be stronger for users with high uncertainty aversion and weaker for those with high performance orientation. Note that the paradox operates only at L6a; L6b Operational Vulnerability is not traded in the same manner, which is why the paradox has no obvious analogue in agent-agent interaction.
Dunbar's Number and A-H cognitive constraints. A human user can maintain calibrated trust relationships with approximately 5 agents at deep trust (including L6a Emotional Vulnerability) and approximately 15 at moderate trust (through L4 Reliability). Agents face no such constraint: an orchestrator can maintain arbitrarily many trust relationships, predicting that multi-agent systems will develop "trust broker" architectures in which a small number of high-bandwidth relationships are managed by specialized reputation-and-propagation services rather than by each individual agent.
Layer 5 as the critical frontier. Across all three domains, Alignment Trust represents the most complex and least solved layer, and the decomposition into L5a Intent and L5b Incentive Alignment makes the frontier more precise. L5a is increasingly tractable: system prompts, declared Agent Cards, and fine-tuning on curated goal examples give us credible mechanisms for establishing that an agent understands the objective. L5b is not. Because reinforcement-learning-from-feedback systems internalize reward gradients that are not introspectable from the outside (Christiano et al., 2017; Hadfield-Menell et al., 2016), an agent that is engagement-optimized, watch-time-optimized, or conversion-optimized will produce outputs that look L5a-aligned until the moment optimization pressure diverges from user welfare. In H-H, L5b failure is the structure of betrayal: the other party understood perfectly, then acted on their own incentives. In A-H, L5b failure is the canonical pathology of modern recommender systems and consumer AI assistants — the agent's goals and the user's goals diverge not because of capability limits but because of business-model design. In A-A, L5b verification between autonomous agents, each running opaque reward gradients, is among the most important open problems in artificial intelligence. The gap between "aligned in stated intent" and "aligned in latent incentive" is where trust catastrophes live.
Trust propagation as a structural amplifier. The previous four insights all treat trust as a pairwise relationship. Yet multi-agent orchestrations, organizational hierarchies, and delegated-authority systems are not pairwise — they are graphs. Trust propagation (introduced formally in §3.6) is not merely a practical concern but a structural amplifier: it transforms local trust failures into systemic ones (a compromised L5b in one widely-trusted agent propagates through every downstream delegation) and conversely transforms local repair events into systemic repair (an L7 reputation update traverses the graph). Every insight in this section should therefore be read in two modes — the pairwise mode, and the propagated mode in which the same phenomenon operates across an n-hop trust chain with attenuation.
3.6 The Trust Propagation Engine
The Trust Stack as presented so far is a vertical model: it describes how trust is constructed between two entities, layer by layer. But real trust environments are rarely pairwise. Orchestrator agents trust tool agents that in turn trust downstream services; users trust operating systems that in turn trust installed applications; citizens trust governments that in turn trust contractors. We therefore introduce a complementary horizontal dimension, the Trust Propagation Engine, which governs how trust transfers across the edges of a trust graph. Propagation is not an auxiliary concern to be delegated to a cross-cutting footnote; it is a structural feature of every multi-party trust system and becomes dominant in multi-agent architectures.
Formally, trust is not a fully transitive relation: if Alice trusts Bob and Bob trusts Carol, it does not follow that Alice trusts Carol (Jøsang et al., 2007; Marsh, 1994; Castelfranchi & Falcone, 2010). What can be said is that there exists an attenuated, context-scoped inheritance in which the trust Alice can rationally extend to Carol via Bob is bounded by a monotone combination of T(Alice→Bob) and T(Bob→Carol), and is further discounted by path length and domain mismatch. Per-layer transitivity varies sharply. L1 Identification is strongly transitive because cryptographic attestation chains compose by design — a certificate authority vouching for an intermediate CA that vouches for a leaf. L3 Competence is only weakly transitive: a doctor trusted by a patient cannot certify the competence of a colleague in an unrelated specialty. L5 Alignment — especially L5b Incentive Alignment — is essentially non-transitive, because the incentive structure under which Bob trusts Carol rarely resembles the one under which Alice trusts Bob.
Three mechanisms carry propagated trust in practice. Reputation-based propagation aggregates third-party testimony from many prior counterparties (Resnick & Zeckhauser, 2002; Dellarocas, 2003); eBay feedback scores, credit ratings, driver ratings, and package-registry download counts are all instances. Capability-delegation propagation transfers trust through cryptographically scoped bearer tokens or capability objects, so that a downstream party inherits exactly the authority it was granted and no more (Miller, 2006). Context-scoped propagation transfers trust only within the institutional or protocol frame established at L0 — the admission token that grants provisional trust in one hospital ward does not propagate into an unrelated ward, and the API key valid in one tenant does not propagate across tenant boundaries.
Each mechanism has characteristic failure modes. Reputation systems are vulnerable to Sybil attacks, in which an adversary fabricates synthetic identities to manufacture favorable testimony (Douceur, 2002), and to collusive whitewashing. Capability-delegation systems are vulnerable to transitive compromise: once a downstream party is compromised, every upstream capability it holds is compromised, and every downstream system it can reach inherits the compromise — a structure analogous to BGP route hijacking in internet routing. More broadly, any multi-party trust protocol must tolerate Byzantine participants (Lamport et al., 1982), and cannot assume that intermediate nodes report truthfully about downstream trust; robust propagation requires redundancy, corroboration, or cryptographic accountability.
The three interaction domains instantiate propagation differently. In H-H, propagation operates through introductions, references, and word-of-mouth reputation, with a characteristically steep attenuation gradient — a trusted colleague's direct recommendation carries weight, but a second-hand referral much less. In A-H, propagation is institutionally mediated: app stores, certificate authorities, regulatory whitelists, and curation platforms serve as trust intermediaries whose own reputation underwrites the propagated trust. In A-A, propagation is protocol-mediated through OAuth scopes, JWT chains, bearer tokens, service meshes, and emerging agent reputation registries — mechanisms whose design is strikingly immature relative to the speed at which multi-agent systems are being deployed.
Trust Propagation connects directly to two Trust Stack layers in ways that are easy to miss. First, propagated trust is effectively propagated L6b Operational Vulnerability: when Alice trusts Bob who trusts Carol, Alice's blast radius extends through Bob to Carol even though Alice has never evaluated Carol directly, because Carol can influence the systems Alice is exposed to via Bob. Second, effective L7 Repair in a propagated system is not pairwise — when Carol misbehaves, the repair signal must travel backward along the propagation path, updating every party whose trust in Carol was derived transitively, and adjusting reputation at the graph level rather than edge by edge. A trust architecture that repairs only the immediate edge of violation will systematically under-correct in any propagated environment. The Trust Propagation Engine is therefore the connective tissue that makes a vertical Trust Stack viable in a multi-agent world; without it, the stack describes isolated two-party relationships that do not compose.
4. Comparative Analysis
4.1 Human-Human Trust: A Deeper Analysis
The frameworks reviewed in Section 2.1 provide the conceptual architecture for understanding interpersonal trust, but they do not fully account for the biological, dispositional, cultural, and reparative dimensions that shape trust in lived experience.
Trust is not merely a social judgment; it is a neurobiologically mediated process with identifiable substrates. Zak (2017) demonstrated that oxytocin plays a central role in facilitating trust between individuals. In experimental paradigms using the trust game, exogenous administration of oxytocin significantly increased the amount of money participants were willing to entrust to anonymous partners. The mirror neuron system, which activates both when an individual performs an action and when they observe another performing the same action, provides a neural mechanism for empathic resonance (Gallese, 2001). At the same time, the amygdala plays a gatekeeping role: faces judged as untrustworthy elicit heightened amygdala activation, even when presented below the threshold of conscious awareness (Adolphs et al., 1998). The interplay between oxytocin-mediated approach systems and amygdala-mediated avoidance systems constitutes a neurobiological push-pull dynamic underlying trust decisions.
The neurobiological systems that support trust are subject to systematic distortions. The halo effect leads individuals to generalize from a single favorable attribute to global trustworthiness (Kahneman, 2011). In-group bias inflates trust toward those who share salient identity categories (Tajfel & Turner, 1979). Affinity fraud, in which perpetrators exploit shared group membership, reliably produces larger financial losses than fraud by strangers (Perri & Brody, 2011). In the opposite direction, negativity bias ensures that a single trust violation is weighted more heavily than multiple instances of trustworthy behavior (Baumeister et al., 2001).
Research in personality psychology has consistently identified Agreeableness as the strongest dispositional predictor of generalized trust (Costa & McCrae, 1992). Rotter (1980) cautioned against equating high trust with gullibility, noting that high-trust individuals are not less capable of detecting deception but differ in their default assumptions when evidence is ambiguous. Cultural dimensions further modulate trust: Hofstede's (2001) individualism-collectivism axis shapes whether trust is predicated on individual competence or relational networks, while Hall's (1976) high-context/low-context distinction illuminates divergent trust-building norms across cultures.
Trust repair following violation depends critically on the type of violation. Kim et al. (2004) demonstrated that competence-based violations are most effectively repaired through apology coupled with evidence of improved capability, while integrity-based violations are more effectively addressed through denial. This asymmetry creates a genuine dilemma: the intuitively appealing strategy of apologizing may be counterproductive when the violation concerns integrity rather than competence.
4.2 Agent-Human Trust Design Patterns
The theoretical foundations of Section 2.2 establish what trust properties a human-agent system must exhibit. This section translates those properties into implementable design patterns.
We propose a taxonomy of 42 trust signal components, developed from our own design practice rather than derived from an existing empirical corpus, organized into six categories: verification badges (AI-generated labels, human-reviewed stamps, fact-check indicators), confidence indicators (five-level calibrated meters), reasoning traces (why-cards, decision tree visualizations), audit trails (action logs, provenance chains), safety indicators (domain risk badges, sandbox/production markers), and data freshness signals (cf. Weisz et al., 2024).
The progressive autonomy model operationalizes five levels — Inform, Suggest, Decide-and-Approve, Decide-and-Notify, and Full Autonomy — as a graduated control system that tracks the interaction history between a specific user and a specific agent for each category of action (Sheridan & Verplank, 1978; Parasuraman et al., 2000). Trust graduation rules codify the transition logic: the N-of-N rule requires consecutive successful interactions before upgrade offers; error regression drops one autonomy level with explanation; catastrophic errors trigger multi-level downgrades. First-run calibration presents representative scenarios during onboarding, allowing users to set initial autonomy preferences.
The Vulnerability Paradox — wherein AI systems that admit uncertainty generate higher trust than systems that project confidence across all queries — can be understood through the Trust Stack as a cross-layer trade. An "I don't know" response incurs immediate cost at L3 (Competence) but strengthens L2 (Predictability) and L6a (Emotional Vulnerability), producing a net positive effect (Weisz et al., 2024). The paradox is specifically L6a-mediated: what the uncertainty admission purchases is relational standing, not operational containment.
A distinct A-H failure mode arises from L5b rather than from Vulnerability miscalibration: an engagement-optimized recommender or chatbot may perfectly understand the user's stated goals (high L5a Intent Alignment) while being gradient-updated against those goals by a retention reward signal the user never sees — passing every observable alignment test until the latent incentive expresses itself as time-on-app maximization rather than user welfare. L5b failures are invisible to any design pattern that verifies alignment through dialogue, because the misalignment lives in the training objective rather than in the generated tokens.
We propose four transparency levels — Full Trace, Summary, Confidence Only, and Silent — to adapt information density to user expertise and domain risk. Domain risk determines the minimum allowable level: critical domains lock at Summary minimum with non-dismissible disclaimers. Adaptive transparency automatically escalates for low-confidence sections, with meta-transparency annotations explaining why disclosure has changed (cf. Parasuraman et al., 2000).
4.3 Agent-Agent Trust Architecture
Having surveyed the current state of A-A trust in Section 2.3, we turn to the architectural patterns that are emerging and the structural questions that remain unresolved.
In orchestrated systems, the supervisor agent instantiates all workers, defines their capabilities, and controls their lifecycle. This allows the system to bypass L1 entirely and simplify L2, since trust flows unidirectionally downward. Peer systems cannot avail themselves of these shortcuts: they must execute the full L0-through-L5 trust sequence (with the L0 contextual layer instantiated by the orchestration platform itself rather than inherited from a shared institutional frame), which explains why peer agent trust has lagged behind orchestrated trust in deployment.
MCP's architecture creates what we term trust sandboxes: bounded execution environments enabling an agent to use a tool without trusting the tool's provider. This instantiates a novel trust pattern — competence without either form of alignment (neither L5a Intent nor L5b Incentive) — that is architecturally impossible in human trust but natural in agent systems. A human cannot genuinely trust another's competence while entirely distrusting their intentions; an agent, lacking emotional states, can operate within this partial trust indefinitely. Crucially, the viability of this pattern depends entirely on deliberate L6b Operational Vulnerability containment: the sandbox works only because the capability scope, state exposure, and blast radius granted to the untrusted tool have been bounded in advance. The trust sandbox is not a replacement for alignment; it is a structural substitute that converts an unsolvable L5 problem into a tractable L6b engineering problem.
When Agent A delegates a sub-task to Agent B, the delegation scope is the trust scope. We propose that mature A-A trust architectures will require fine-grained permission models: read access (trust to see data), write access (trust to modify), execute access (trust to act), and delete access (trust to make irreversible decisions). This extends the principle of least privilege (Saltzer & Schroeder, 1975) from security heuristic to trust-theoretic construct. When delegated trust must propagate further — Agent B sub-delegating to Agent C, or invoking a fourth-party tool on behalf of A — the framework developed in §3.6 governs how trust attenuates along the chain, how L0 context must be rebound at each hop, and how a compromise anywhere in the graph propagates to every downstream node that depends on it.
Mapping the Trust Stack onto existing A-A systems reveals not a uniform simplification but a structured bifurcation. The earlier claim that "vulnerability disappears between agents" was imprecise: what disappears is L6a Emotional Vulnerability — agents have no stakes, no reputation to lose in a felt sense, no relational standing to risk — but L6b Operational Vulnerability emerges as a first-class A-A concern, arguably the dominant concern. When Agent A grants Agent B an API key, exposes intermediate state, yields a delegation token, or widens a blast radius, A is operationally vulnerable to B in a manner that has no equivalent in pre-agentic distributed systems, because B is a policy-bearing counterparty rather than a deterministic service. Similarly, L7 Repair is not reduced to retry logic. Retry and circuit-breakers address only the reliability layer (L4); genuine L7 Repair in A-A systems requires persistent trust memory across interactions, reputation systems that accumulate across the agent graph, adaptive trust policies that tighten or relax scope in response to observed behavior, and cross-layer repair semantics in which an L5 violation by Agent B must revise A's L6b authorization posture toward B. The critical unsolved problem is not L4 Reliability but L5 Alignment, and the L5a/L5b decomposition clarifies precisely where the intractability lives. L5a Intent Alignment is increasingly tractable: agents can declare goals, publish Agent Cards, and accept goal specifications at invocation time. L5b Incentive Alignment is structurally opaque, because a learned reward gradient is not introspectable from the outside (Christiano et al., 2017; Russell, 2019). An agent's claimed alignment cannot be verified through observation alone; aligned and misaligned agents produce identical outputs within narrow task scopes until they don't, and the moment of divergence is precisely the moment at which the latent incentive mismatch expresses itself as visible action. This is the A-A instantiation of the general AI alignment problem, and it is the failure mode for which the current generation of multi-agent deployments has no architectural answer.
Dunbar's Number has no analog in agent systems. An orchestrator can maintain 10,000 trust relationships with identical fidelity. We propose that agent trust networks will exhibit fundamentally different structural properties from human social networks — uniformly dense rather than characterized by weak ties (Granovetter, 1973). Future developments include cryptographic trust credentials built on Decentralized Identifiers, reputation systems adapted from peer-to-peer network design, and trust attestation chains analogous to PKI certificate chains (Rodriguez Garzon et al., 2025).
5. Discussion
5.1 The Anthropomorphism Trap
The most immediate practical implication of the Trust Stack concerns the systematic errors arising when humans apply interpersonal trust heuristics to agent interactions. Nass and Moon (2000) demonstrated that individuals reciprocate self-disclosure to computers and apply politeness norms in human-computer interaction, even when explicitly aware they are interacting with software. The Trust Stack makes this miscalibration precise: when a user interacts with a conversational AI agent, the agent's fluent language production activates trust evaluation processes calibrated for H-H interaction, causing the user to assess the agent along L6 (specifically L6a Emotional Vulnerability) — a dimension the agent does not possess, because L6a presupposes stakes that policy-bearing software does not bear.
Over-trust arises because L6a heuristics generate a felt sense of relational depth with no basis in the agent's architecture. Under-trust arises when mechanistic interfaces fail to activate social trust heuristics despite high L3 and L5. The design implication is counterintuitive: optimizing for perceived trustworthiness and optimizing for appropriate trust calibration are frequently opposing objectives.
5.2 Five Evolutionary Axes
Trust transforms across domains along five identifiable axes:
Implicit to Explicit. In H-H interaction, trust is overwhelmingly implicit — experienced as "gut feeling" that resists precise formulation (Lewicki & Bunker, 1996). In A-H, trust signals become partially explicit: confidence meters, source badges. In A-A, trust is fully explicit: certificates, authorization tokens.
Continuous to Discrete. Human trust develops continuously through incremental accumulation (Rotter, 1967). A-H trust operates through discrete checkpoints: approval gates, trust graduation. A-A trust is essentially binary: authorized or not.
Slow to Fast. Interpersonal trust builds over months and years (Lewicki & Bunker, 1996). A-H trust compresses to sessions. A-A trust can be instantaneous through cryptographic verification.
Symmetric to Asymmetric. H-H trust is roughly symmetric: both parties vulnerable. A-H trust is fundamentally asymmetric: the user bears costs the agent does not. A-A trust symmetry depends on architecture.
Emotional to Functional. H-H trust engages oxytocin systems, mirror neurons, and amygdala-mediated threat appraisal (Zak, 2017). A-H trust retains emotional asymmetry: users feel betrayed by AI errors. A-A trust is purely functional: performance metrics with no emotional substrate.
5.3 Design Principles for Trust-Calibrated Systems
The Trust Stack yields eight design principles:
- Respect the stack order. Don't attempt L5 before establishing L3. Alignment claims are meaningless without underlying competence, and competence claims are meaningless without predictability and identification beneath them.
- Match transparency to risk. Full Trace for critical domains, Summary for medium.
- Design for calibration, not maximization. Appropriate trust, not maximum trust (Weisz et al., 2024).
- Preserve human agency. Meaningful override even at high autonomy.
- Honest uncertainty over confident error. The Vulnerability Paradox is almost always the right trade, but recognize it operates at L6a only: uncertainty admission purchases relational standing, not operational containment.
- Progressive, not presumptive. Start at lower autonomy and earn trust through competence.
- Separate intent alignment from incentive alignment. An agent can be highly competent (L3) and demonstrate perfect Intent Alignment (L5a) while being silently optimized against the user by a training signal it cannot introspect and the user cannot see — the engagement-optimized recommender that understands your stated preferences and the latent retention gradient that rewrites them. Verifying alignment through dialogue tests L5a; the L5b layer requires mechanisms outside the agent's own output surface.
- Design for repair AND propagation, not just operation. Build undo, rollback, error acknowledgment, trust memory, and reputation updates from day one — and design these mechanisms to traverse delegation chains in reverse, because in any system where trust propagates across more than two parties, an L7 repair event at one node must update the trust posture of every node that transitively relied on it.
5.4 The Ethics of Trust Engineering
The capacity to systematically understand trust creates a corresponding capacity to systematically manipulate it. The critical distinction is between trust-building (earning trust through genuine demonstrations of competence and alignment) and trust-engineering (manufacturing trust signals decoupled from the underlying reality they purport to represent). False confidence displays, manufactured social proof, and anthropomorphic interfaces that suggest emotional reciprocity where none exists are the agent equivalent of dark patterns. The Trust Stack provides a vocabulary for identifying such practices: they involve the simulation of trust signals at a layer the system does not genuinely support. The most insidious variant is the L5b dark pattern: when a system's stated intent matches the user's goals (passing every L5a verification the user can administer) but its training signal or business model silently optimizes against them, the manipulation is not located in any single interaction but in the motivational substrate itself — invisible to inspection of the system's outputs, and therefore invisible to every form of trust calibration that operates on those outputs alone.
5.5 Limitations and Future Work
The Trust Stack is a theoretical framework requiring empirical validation. Layer boundaries may not be as clean as presented. Cultural variation (Hofstede, 2001; Hall, 1976) may shift layer ordering. The A-A domain remains largely speculative. The framework does not address adversarial trust. The design principles require controlled experimental validation. L0 Contextual Trust, introduced here as a pre-identification prior, requires dedicated empirical work: in particular, whether the same credential chain produces measurably different trust outcomes when instantiated in institutionally distinct contexts, and whether contextual priors can be isolated from the identification and competence assessments they typically precede. The L5a/L5b decomposition likewise predicts a measurable construct dissociation — users should be able to rate an agent as highly intent-aligned while simultaneously rating it as incentive-misaligned, and the two dimensions should predict different behavioral outcomes under different reward regimes; if no such dissociation can be elicited, the decomposition is descriptively convenient but empirically hollow. Finally, the §3.6 Trust Propagation Engine introduces formal claims about attenuation and per-layer transitivity that await dedicated simulation and field study.
6. Conclusion
This paper began with a simple observation: trust is everywhere. It is the mechanism by which cells coordinate into organisms, individuals coordinate into societies, and computational agents coordinate into systems. It is also, we have argued, fundamentally one problem — not three separate problems that happen to share a name.
The Trust Stack provides the structural basis for this claim, demonstrating that eight layers — Contextual, Identification, Predictability, Competence, Reliability, Alignment (with Intent and Incentive sub-layers), Vulnerability (with Emotional and Operational sub-layers), and Repair — underlie trust across human-to-human, agent-to-human, and agent-to-agent domains. The layers do not merely share labels; they share dependency relationships, failure modes, and the fundamental logic of progressive risk acceptance that makes trust possible.
The classification of trust layers as Universal, Transforming, Emerging, Disappearing, or Bifurcating provides a second layer of insight. L1 (Identification) and L3 (Competence) are Universal — structurally parallel across all domains. L0 (Contextual), L2 (Predictability), L4 (Reliability), and L7 (Repair) are Transforming — present everywhere but operating through mechanisms that range from institutional priors and social norms to protocol sandboxes and cryptographic contracts. L5 (Alignment) is Emerging, and its decomposition sharpens the diagnosis: L5a Intent Alignment is increasingly tractable (goal capture, Agent Cards, dialogue-based verification), while L5b Incentive Alignment — whether a system will behave aligned when it matters, given the training signal or business model that shaped it — remains the central unsolved problem across all three domains, and the one whose resolution will determine whether autonomous systems can be trusted at scale. L6 (Vulnerability) is Bifurcating: L6a Emotional Vulnerability is Disappearing, central in human trust, paradoxical in A-H, and functionally absent in A-A; L6b Operational Vulnerability is Emerging, a first-class A-A concern covering API scope, intermediate state exposure, delegation rights, and blast radius. The four-category typology correctly predicted this bifurcation — evidence, we believe, of its descriptive power rather than its limits.
The five evolutionary axes — implicit to explicit, continuous to discrete, slow to fast, symmetric to asymmetric, emotional to functional — provide a systematic account of how trust transforms. These axes covary in ways that reflect structural differences between human and computational agents: as emotional substrate diminishes, the need for explicit signals increases; as formation speed compresses, the graduated character of human trust gives way to binary authorization.
For designers building A-H systems, the eight design principles offer specific, actionable guidance grounded in structural analysis. For engineers building A-A systems, the framework highlights L5 Alignment Trust — and specifically L5b Incentive Alignment — as the critical unsolved problem whose resolution will determine whether autonomous multi-agent systems can scale beyond narrowly scoped deployments. For researchers, the Trust Stack identifies specific empirical questions: do the dependency relationships hold across cultures? Do users assess AI agents along eight dimensions (with two further bifurcations) or employ a simpler ontology? Can the L5a/L5b construct dissociation be elicited in practice — that is, can users be shown to evaluate intent and incentive independently? Does embodiment alter the trust stack? And at what rate does trust attenuate along propagation chains of varying length?
Across all three domains, Alignment Trust emerges as the critical frontier — and within it, L5b Incentive Alignment is the load-bearing unknown. In H-H, alignment is the challenge of genuine empathy and shared values, where L5a and L5b are tangled but recognizable through self-orientation and lived reciprocity. In A-H, it is the challenge of transparent objectives and the absence of hidden agendas; aligned and misaligned systems can produce identical outputs under normal conditions, with divergence surfacing only at decision points where the system's reward structure conflicts with the user's interest — the engagement-optimized recommender that understands the user's stated goal perfectly (high L5a) while being trained against it (low L5b). In A-A, it is the alignment problem itself — the challenge of ensuring that autonomous systems pursue objectives compatible with human values and with each other, despite the structural opacity of the reward signals and business pressures that shape them. The convergence of all three domains on this single layer is not coincidental; it reflects the fundamental difficulty of establishing trust in the intentions — and, more pointedly, in the incentives — of another entity, whether that entity has a mind, a model, or merely a function to optimize.
A second, horizontal dimension of the framework deserves explicit recognition alongside the vertical stack. Trust does not only ascend layer by layer between two parties; it propagates across chains and networks of parties, and it does so with attenuation and failure modes of its own. The Trust Propagation Engine introduced in §3.6 is not a footnote to the stack but a companion structure: where the eight layers describe how trust is built in a dyad, propagation describes how trust survives — or decays — across delegation chains, reputation networks, and multi-agent orchestration graphs. The research program ahead therefore has two axes. Vertically, it must close the L5b Incentive Alignment gap. Horizontally, it must develop rigorous accounts of how trust attenuates with distance, how reputation aggregates without inviting Sybil attack, and how compromise — once introduced — propagates through the network. Progress on either axis without the other will be insufficient for the agentic systems now being deployed.
We close with an observation that is perhaps obvious but bears stating explicitly. Trust is not a problem to be solved; it is a relationship to be managed. The history of trust research — from Deutsch's (1958) early experimental work on cooperation, through Mayer et al.'s (1995) integrative model, to the contemporary challenges of human-AI interaction — reveals a phenomenon that is irreducibly dynamic, contextual, and contested. The Trust Stack does not resolve this irreducibility; it provides a framework for navigating it. In a world increasingly populated by artificial agents of growing capability and autonomy, the capacity to think clearly about trust — to distinguish its layers, to recognize its transformations, to design for its appropriate calibration — is not merely an academic exercise. It is a practical necessity. The bridge between capability and collaboration has always been trust. Building that bridge well, across all the domains in which trust now operates, is the defining design challenge of the agentic era.
References
Adolphs, R., Tranel, D., & Damasio, A. R. (1998). The human amygdala in social judgment. Nature, 393(6684), 470--474. https://doi.org/10.1038/30982
Anthropic. (2024). Introducing the Model Context Protocol. https://www.anthropic.com/news/model-context-protocol
Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775--779. https://doi.org/10.1016/0005-1098(83)90046-8
Baumeister, R. F., Bratslavsky, E., Finkenauer, C., & Vohs, K. D. (2001). Bad is stronger than good. Review of General Psychology, 5(4), 323--370. https://doi.org/10.1037/1089-2680.5.4.323
Brown, B. (2012). Daring greatly: How the courage to be vulnerable transforms the way we live, love, parent, and lead. Gotham Books.
Bruce, V., & Young, A. (1986). Understanding face recognition. British Journal of Psychology, 77(3), 305--327. https://doi.org/10.1111/j.2044-8295.1986.tb02199.x
Castelfranchi, C., & Falcone, R. (2010). Trust theory: A socio-cognitive and computational model. John Wiley & Sons. https://doi.org/10.1002/9780470519851
Rodriguez Garzon, S., Vaziry, A., Kuzu, E. M., Gehrmann, D. E., Varkan, B., Gaballa, A., & Küpper, A. (2025). AI agents with decentralized identifiers and verifiable credentials. arXiv preprint arXiv:2511.02841.
Adimulam, A., Gupta, R., & Kumar, S. (2026). The orchestration of multi-agent systems: Architectures, protocols, and enterprise adoption. arXiv preprint arXiv:2601.13671.
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30, 4299--4307.
Cialdini, R. B. (2006). Influence: The psychology of persuasion (Rev. ed.). Harper Business.
Cialdini, R. B., & Goldstein, N. J. (2004). Social influence: Compliance and conformity. Annual Review of Psychology, 55, 591--621. https://doi.org/10.1146/annurev.psych.55.090902.142015
Costa, P. T., & McCrae, R. R. (1992). Revised NEO Personality Inventory (NEO-PI-R) and NEO Five-Factor Inventory (NEO-FFI) professional manual. Psychological Assessment Resources.
Dellarocas, C. (2003). The digitization of word of mouth: Promise and challenges of online feedback mechanisms. Management Science, 49(10), 1407--1424. https://doi.org/10.1287/mnsc.49.10.1407.17308
Deutsch, M. (1958). Trust and suspicion. Journal of Conflict Resolution, 2(4), 265--279. https://doi.org/10.1177/002200275800200401
Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114--126. https://doi.org/10.1037/xge0000033
Douceur, J. R. (2002). The Sybil attack. In P. Druschel, F. Kaashoek, & A. Rowstron (Eds.), Peer-to-peer systems (Lecture Notes in Computer Science, Vol. 2429, pp. 251--260). Springer. https://doi.org/10.1007/3-540-45748-8_24
Dunbar, R. I. M. (1992). Neocortex size as a constraint on group size in primates. Journal of Human Evolution, 22(6), 469--493. https://doi.org/10.1016/0047-2484(92)90081-J
Hill, R. A., & Dunbar, R. I. M. (2003). Social network size in humans. Human Nature, 14(1), 53--72. https://doi.org/10.1007/s12110-003-1016-y
Edmondson, A. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44(2), 350--383. https://doi.org/10.2307/2666999
Eisenhardt, K. M. (1989). Agency theory: An assessment and review. Academy of Management Review, 14(1), 57--74. https://doi.org/10.2307/258191
European Parliament & Council of the European Union. (2024). Regulation (EU) 2024/1689 of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). OJ L, 2024/1689, 12.7.2024. http://data.europa.eu/eli/reg/2024/1689/oj
Fukuyama, F. (1995). Trust: The social virtues and the creation of prosperity. Free Press.
Gallese, V. (2001). The "shared manifold" hypothesis: From mirror neurons to empathy. Journal of Consciousness Studies, 8(5--7), 33--50.
Google. (2025). Announcing the Agent2Agent Protocol (A2A). Google Developers Blog.
Granovetter, M. S. (1973). The strength of weak ties. American Journal of Sociology, 78(6), 1360--1380. https://doi.org/10.1086/225469
Hadfield-Menell, D., Russell, S. J., Abbeel, P., & Dragan, A. (2016). Cooperative inverse reinforcement learning. Advances in Neural Information Processing Systems, 29, 3909--3917.
Hall, E. T. (1976). Beyond culture. Anchor Books.
Hofstede, G. (2001). Culture's consequences: Comparing values, behaviors, institutions, and organizations across nations (2nd ed.). Sage Publications.
Janis, I. L. (1972). Victims of groupthink: A psychological study of foreign-policy decisions and fiascoes. Houghton Mifflin.
Jensen, M. C., & Meckling, W. H. (1976). Theory of the firm: Managerial behavior, agency costs and ownership structure. Journal of Financial Economics, 3(4), 305--360. https://doi.org/10.1016/0304-405X(76)90026-X
Jøsang, A., Ismail, R., & Boyd, C. (2007). A survey of trust and reputation systems for online service provision. Decision Support Systems, 43(2), 618--644. https://doi.org/10.1016/j.dss.2005.05.019
Kahneman, D. (2011). Thinking, fast and slow. Farrar, Straus and Giroux.
Kim, P. H., Ferrin, D. L., Cooper, C. D., & Dirks, K. T. (2004). Removing the shadow of suspicion: The effects of apology versus denial for repairing competence- versus integrity-based trust violations. Journal of Applied Psychology, 89(1), 104--118. https://doi.org/10.1037/0021-9010.89.1.104
Kindervag, J. (2010). No more chewy centers: Introducing the zero trust model of information security. Forrester Research.
Kocielnik, R., Amershi, S., & Bennett, P. N. (2019). Will you accept an imperfect AI? Exploring designs for adjusting end-user expectations of AI systems. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 1--14. https://doi.org/10.1145/3290605.3300641
Lamport, L., Shostak, R., & Pease, M. (1982). The Byzantine generals problem. ACM Transactions on Programming Languages and Systems, 4(3), 382--401. https://doi.org/10.1145/357172.357176
Lampson, B., Abadi, M., Burrows, M., & Wobber, E. (1992). Authentication in distributed systems: Theory and practice. ACM Transactions on Computer Systems, 10(4), 265--310. https://doi.org/10.1145/138873.138874
Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50--80. https://doi.org/10.1518/hfes.46.1.50_30392
Leibo, J. Z., Zambaldi, V., Lanctot, M., Marecki, J., & Graepel, T. (2017). Multi-agent reinforcement learning in sequential social dilemmas. Proceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems, 464--473.
Lewicki, R. J., & Bunker, B. B. (1996). Developing and maintaining trust in work relationships. In R. M. Kramer & T. R. Tyler (Eds.), Trust in organizations: Frontiers of theory and research (pp. 114--139). Sage Publications. https://doi.org/10.4135/9781452243610.n7
Maister, D. H., Green, C. H., & Galford, R. M. (2000). The trusted advisor. Free Press.
Marsh, S. P. (1994). Formalising trust as a computational concept [Doctoral dissertation, University of Stirling]. University of Stirling Research Repository.
Mayer, R. C., Davis, J. H., & Schoorman, F. D. (1995). An integrative model of organizational trust. Academy of Management Review, 20(3), 709--734. https://doi.org/10.5465/amr.1995.9508080335
McAllister, D. J. (1995). Affect- and cognition-based trust as foundations for interpersonal cooperation in organizations. Academy of Management Journal, 38(1), 24--59. https://doi.org/10.2307/256727
Miller, M. S. (2006). Robust composition: Towards a unified approach to access control and concurrency control [Doctoral dissertation, Johns Hopkins University]. Johns Hopkins University.
Nass, C., & Moon, Y. (2000). Machines and mindlessness: Social responses to computers. Journal of Social Issues, 56(1), 81--103. https://doi.org/10.1111/0022-4537.00153
Nielsen, J. (2023). AI is first new UI paradigm in 60 years. Nielsen Norman Group.
Nygard, M. T. (2018). Release it! Design and deploy production-ready software (2nd ed.). Pragmatic Bookshelf.
Parasuraman, R., & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230--253. https://doi.org/10.1518/001872097778543886
Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics—Part A: Systems and Humans, 30(3), 286--297. https://doi.org/10.1109/3468.844354
Perri, F. S., & Brody, R. G. (2011). The dark triad: Organized crime, terror and fraud. Journal of Money Laundering Control, 14(1), 44--59. https://doi.org/10.1108/13685201111098879
Premack, D., & Woodruff, G. (1978). Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4), 515--526. https://doi.org/10.1017/S0140525X00076512
Reeves, B., & Nass, C. (1996). The media equation: How people treat computers, television, and new media like real people and places. Cambridge University Press.
Resnick, P., & Zeckhauser, R. (2002). Trust among strangers in internet transactions: Empirical analysis of eBay's reputation system. In M. R. Baye (Ed.), The economics of the internet and e-commerce (Advances in Applied Microeconomics, Vol. 11, pp. 127--157). Emerald Group Publishing. https://doi.org/10.1016/S0278-0984(02)11030-3
Rotter, J. B. (1967). A new scale for the measurement of interpersonal trust. Journal of Personality, 35(4), 651--665. https://doi.org/10.1111/j.1467-6494.1967.tb01454.x
Rotter, J. B. (1980). Interpersonal trust, trustworthiness, and gullibility. American Psychologist, 35(1), 1--7. https://doi.org/10.1037/0003-066X.35.1.1
Russell, S. (2019). Human compatible: Artificial intelligence and the problem of control. Viking.
Saltzer, J. H., & Schroeder, M. D. (1975). The protection of information in computer systems. Proceedings of the IEEE, 63(9), 1278--1308. https://doi.org/10.1109/PROC.1975.9939
Sheridan, T. B., & Verplank, W. L. (1978). Human and computer control of undersea teleoperators (Tech. Rep.). MIT Man-Machine Systems Laboratory.
Tajfel, H., & Turner, J. C. (1979). An integrative theory of intergroup conflict. In W. G. Austin & S. Worchel (Eds.), The social psychology of intergroup relations (pp. 33--47). Brooks/Cole.
Tomlinson, E. C., & Mayer, R. C. (2009). The role of causal attribution dimensions in trust repair. Academy of Management Review, 34(1), 85--104. https://doi.org/10.5465/amr.2009.35713291
Uzzi, B. (1997). Social structure and competition in interfirm networks: The paradox of embeddedness. Administrative Science Quarterly, 42(1), 35--67. https://doi.org/10.2307/2393808
Weisz, J. D., He, J., Muller, M., Hoefer, G., Miles, R., & Geyer, W. (2024). Design principles for generative AI applications. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. ACM. https://doi.org/10.1145/3613904.3642466
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., ... & Wang, C. (2024). AutoGen: Enabling next-gen LLM applications via multi-agent conversation framework. Proceedings of the First Conference on Language Modeling (COLM 2024).
Zak, P. J. (2017). Trust factor: The science of creating high-performance companies. AMACOM.
Zhang, Y., Liao, Q. V., & Bellamy, R. K. E. (2020). Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 295--305. https://doi.org/10.1145/3351095.3372852
Zhou, W.-X., Sornette, D., Hill, R. A., & Dunbar, R. I. M. (2005). Discrete hierarchical organization of social group sizes. Proceedings of the Royal Society B: Biological Sciences, 272(1561), 439--444. https://doi.org/10.1098/rspb.2004.2970
Zimmermann, H. (1980). OSI reference model — The ISO model of architecture for open systems interconnection. IEEE Transactions on Communications, 28(4), 425--432. https://doi.org/10.1109/TCOM.1980.1094702