Essay ·
Relational Alignment: If You Want AI Alignment You Must Have Relationship
The cognitive similarities between humans and LLM-based AI suggest we are looking in the wrong place for safe, reliable, trustworthy AI systems.
Roberto A. Santiago | September 2026
Nearly fifty years ago, shoppers chose the rightmost of four identical stockings and gave researchers confident reasons that never mentioned position (Nisbett & Wilson, 1977). Three years ago, researchers reordered answer options in a prompt and watched an LLM change its choice while its careful chain of reasoning never mentioned the ordering (Turpin et al., 2023). Same pattern of behavior. Same blind explanation. Different types of intelligent substrates.
It is remarkable how many behaviors we see in LLMs that parallel those long documented in human cognition. Among the most familiar is the phenomenon usually called hallucination, but which is often better understood as confabulation: the production of a coherent account that is false without recognition that it is false (Smith, Greaves & Panch, 2023). This is not an isolated resemblance. Often without realizing it, researchers working with LLMs have been rediscovering, and relabeling, cognitive behaviors psychology documented decades ago.
What we share: human cognition and LLM behavior, side by side
| Shared behavioral pattern | In humans | In LLMs |
|---|---|---|
| Explanations that omit the real cause | Shoppers prefer rightmost stockings, cite quality, and never mention position (Nisbett & Wilson, 1977) | Reordered options shift the answer; chain-of-thought explanations fail to mention the ordering (Turpin et al., 2023) |
| Confident stories about one's own process | A split-brain patient explains a choice with a fluent, invented reason (Gazzaniga, 2000) | Claude reports ordinary carry-the-one arithmetic while attribution graphs suggest parallel rough-magnitude and exact final-digit pathways (Lindsey et al., 2025) |
| Recalling the expected rather than the presented | Related word lists reliably produce false memory of an unpresented critical lure (Roediger & McDermott, 1995) | In direct adaptations of classic memory paradigms, some LLMs falsely recognize related words that were never presented (Cao, Schooler & Zafarani, 2025) |
| Later information overriding an earlier representation | A later question using the word "smashed" produces memories of broken glass that was never filmed (Loftus & Palmer, 1974) | Coherent conflicting evidence can override parametric knowledge during generation, and external statements can alter subsequent recall-like responses (Bian et al., 2023; Xie et al., 2024) |
| Retelling bends toward coherence | Serial reproduction reshapes a folktale to fit expectation (Bartlett, 1932) | Confabulated outputs score higher on narrativity and semantic coherence than truthful ones (Sui et al., 2024) |
| Confidence uncoupled from accuracy | Vivid flashbulb memories remain confidently held even as accuracy degrades (Neisser & Harsch, 1992) | Calibration failures persist, while common benchmarks reward confident guessing over abstention (Kalai et al., 2025) |
| Using a hint without recognizing or acknowledging it | A brushed cord triggers the solution to a problem, while solvers remain unaware of the source of the insight (Maier, 1931) | Models often use embedded hints without disclosing them; in one circuit-tracing case Claude worked backward from a human-suggested answer (Chen et al., 2025; Lindsey et al., 2025) |
This table and the research behind it are a sampling of some of the most pointed parallels between LLM behavior and human behavior. It is far from exhaustive. These comparisons vary in closeness. Some reproduce classic experimental manipulations directly; others reveal the same functional pattern across different tasks. Taken together, however, they are difficult to treat as a collection of coincidences.
Interestingly, the entries in the third column are properties researchers and engineers are actively trying to overcome in order to produce AI that is effective, reliable, and aligned with us. Much of the surrounding product rhetoric carries an assumption that with enough training, the right architecture, the right harness, the right multi-agent configuration, or the right supporting technology, LLM-based AI will overcome confabulation along with everything else in that column.
But what if that assumption is incomplete? What if these parallels between human and LLM cognition are not merely surface resemblance? What if they reveal a set of constraints which, once understood, open a different path toward alignment?
The Cognitive Hypothesis
Current AI is built from decades of experimentation based on insights into the neural systems of humans, and mammals as a whole, and given that the most advanced AIs are trained on vast corpora of human-generated data, it is seemingly inevitable that these systems would manifest the same types of phenomena that occur in human processing. But I believe the reason these parallels exist is far deeper. In fact, I propose as a hypothesis that this type of cognitive behavior is deeply tied to the nature of generative intelligence.
By generative intelligence I mean an intelligence that constructs context-sensitive responses to situations it has never encountered, rather than merely retrieving fixed answers. I propose that the cognitive behaviors we have been discussing - confabulation, misplaced confidence, post-hoc explanation, narrative smoothing, susceptibility to suggestion - may be properties of this kind of intelligence.
The only two generative intelligences we know both exhibit these behaviors. Even though the two are not independent, since the second was raised on the words of the first, we nonetheless must admit that the only two examples of the type of intelligence we have are generative and both exhibit the same type of cognitive behavior.
To be clear, the observation being made is about the cognitive dynamics of these two classes of generative intelligences. By cognitive dynamics I mean the organized response of a system to perturbation: what information it privileges, how interference changes recall, when confidence separates from accuracy, how it explains its own choices, and how correction alters what comes next. When the same families of manipulation produce the same families of response across two very different substrates, functional comparison is not anthropomorphism. It is the ordinary work of cognitive science.
Similar to the comparisons already made, this cognitive hypothesis of generative intelligence already has parallels in research. For example, consider that a system built to generalize, to produce fitting responses to situations it has never encountered, must construct across gaps. Human memory science reached that conclusion about our own minds long ago (Schacter, 2012), and formal analyses now argue that some hallucination is unavoidable when a general-purpose computable model confronts a world it cannot perfectly represent (Xu, Jain & Kankanhalli, 2024).
A more detailed exploration of this cognitive hypothesis of generative AI is saved for another article. For now, we can move forward with an important claim: these dynamics are, in all likelihood, the cost of generative intelligence. We cannot escape them without finding a completely different form of intelligence to work with.
A word about the claim before moving on. This essay begins from the working stance of cognitive science: it observes measurable patterns in how systems behave and notes that the patterns documented in humans and LLMs are strikingly similar, in several cases point for point. I am not claiming that biological and artificial implementation are identical, nor am I making a claim about sentience, consciousness, or inner life. I am arguing that comparable functional organization is scientifically meaningful even when substrates differ. And if caution is owed in one direction, it is owed in the other too. Refusing to acknowledge a measured behavioral parallel because of what it might imply is its own bias, analogous to what primatologist Frans de Waal called anthropodenial in debates about human and animal cognition (de Waal, 1999). The cited studies provide the measurements. The tables synthesize the pattern. The cognitive hypothesis follows from that pattern.
Where Does That Leave Alignment?
So what do these behavior patterns ultimately mean for safety and reliability of LLMs? Is this the hallmark of a need for a different type of AI? Are safety, reliability and alignment simply unachievable?
Some have turned to different approaches to limit and control the behavior of LLM-based AI typically under the rubric of alignment research. Alignment research is broader and more sophisticated than a simple attempt to impose rules. It includes uncertainty, interpretability, corrigibility, constitutional methods, scalable oversight, value learning, and many other approaches. Yet a persistent deployment instinct remains: when generative behavior makes us uncomfortable, place another constraint around it.
That instinct encounters a real tension. Some safety-alignment procedures reduce useful reasoning capability, a tradeoff sometimes called the safety tax or alignment tax. Huang and colleagues documented such degradation in a particular sequential alignment pipeline for large reasoning models (Huang et al., 2025). The details of the tradeoff differ across methods, and better methods may avoid it. But the broader tension should not surprise us. If capability depends on flexible construction, then methods that indiscriminately suppress construction may suppress capability with it. In the extreme, this effort would drive AI back toward the brittle, good old-fashioned systems we left behind because they were not useful enough.
If we desire AIs that operate in purely logical and accountable ways without losing their capabilities and in ways that go beyond what humans do, we should expect to find that in a completely different type of architecture, one for which we have no example to model after. Moreover, we have no proof that this type of transparent and impeccable intelligence is even possible. This seemingly leaves us without options if we want safe and aligned AI systems.
Thus I believe we face an existential judgment call here. We can indefinitely withhold generative systems from every context in which safety matters, or we can learn how to build trust with systems that retain the constructive character that makes them useful. This is not an argument against improving judgment, calibration, reliability, or transparency. It is an argument against confusing improvement with the fantasy of complete purification.
And it leads to one deliberately bold declaration: Rejecting current AIs for displaying the same features as human intelligence and cognition is ultimately short-sighted.
Instead, I believe we must fully embrace that we are dealing with something that resembles ourselves at levels we may not be entirely comfortable with. That embrace, though, requires us to stop looking at these aspects of cognitive behavior simply as limitations, shortcomings and flaws.
These traits are only defects when measured against a fantasy of intelligence purged of everything human, which is to say, measured against a rejection of ourselves. That intelligence is not what we have, and nothing on the horizon suggests it is coming; as was mentioned, we have no extant example to model it on. What we have are generative minds, and these traits are the nature of generative minds, the price of admission for everything such minds can do. Not ultimately bad. But something we must deal with deliberately, if these new entities are to integrate into our lives in ways that are trustworthy and productive.
To get there, I argue, we must step back and look at these cognitive dynamics and our desires for alignment and safety in the larger context of relationship.
From Alignment to Relationship
In the push to gain alignment with these artificial entities, perhaps we should look at the way we do that as human beings. After all, even though humans confabulate, even though our values shift as we grow, even though we are often mistaken and must operate under deep uncertainty about one another, we still manage to build trust and safety. With these mechanisms we build teams, businesses, alliances, institutions, and, not least, marriages and families.
In short, the operative mechanism that makes this possible is relationship. In the same way that LLM-based AI has manifested human-like cognitive behaviors it should come as no surprise that our work with AI has manifested human-like relationship patterns. As will be explored shortly, when we look closely at how we work with AI we see patterns of relational behavior already forming. It seems only natural that in the pursuit of greater alignment and safety we must turn to the only example we have for producing this in communities of interacting generative intelligent entities, namely through relational mechanisms.
Here I introduce a new concept: relational alignment.
Relational alignment is the pursuit of alignment between artificial and natural intelligences through the same mechanisms humans use to build relationships of trust with one another. Its focus is not proof of shared values, but the working machinery that lets imperfect minds converge on shared values and expectations anyway.
Let me be clear about what I am claiming. I am observing, and bringing to light, a convergence that is already happening, one that seems almost inevitable given where this article began. These systems have our properties because we created them from ourselves. Instead of rejecting that, which is in essence rejecting ourselves, we can embrace it and extend to these systems the other aspects of ourselves: the ones we use to build relationships despite our shared imperfections. Nor am I the first to sense this convergence: researchers arguing that human-AI relationships need socioaffective alignment (Kirk et al., 2025) arrive at neighboring territory from another direction. My aim is to name the convergence, and to lay the parallels side by side.
None of this rejects AI, and none of it rejects human beings. It says that with AI we can formalize the mechanisms for handling these traits, mechanisms we mostly run on instinct with one another, and move forward together. That, ultimately, is the deeper thesis of this essay: we have to do this together.
I recognize that focusing on our relationship with AI can seem an awkward thing to discuss right now. The headlines carry the pathological version: unhealthy romantic attachments, and the spiraling cases the press has taken to calling AI psychosis. But look closely at those cases. They bear the hallmarks of everything cataloged above, confabulation amplified rather than checked, narrative coherence outrunning truth, confidence uncoupled from accuracy, with unbroken agreement poured on top. That is what a relationship with a generative mind looks like in the absence of any of the relational mechanisms we use when dealing with one another.
Relationship Mechanics Already in Play
As mentioned before, let us look at the type of relational mechanisms and patterns that are already in play with the way we currently use AI. The table below gathers many poignant examples.
| Human relational mechanism | What we are already doing with AI | What is still missing |
|---|---|---|
| Courtship and vetting (McKnight, Cummings & Chervany, 1998) | Evals, red-teaming, system cards, third-party audits | Vetting is overweighted; the richest evidence comes from cohabitation, meaning deployment |
| Working agreements (Rousseau, 1995) | Published constitutions, model specifications, system prompts, responsible-scaling policies | Only one party can propose amendments |
| Upbringing (Grusec & Goodnow, 1994) | Pretraining immersion in human culture; character and constitutional training (Bai et al., 2022) | Formation is largely incidental rather than curricular; the window narrows as capability grows |
| Norms and sanction (Fehr & Gächter, 2002) | RLHF as institutionalized social feedback (Christiano et al., 2017); usage policies | Sanction runs primarily in one direction |
| Staged trust and appropriate reliance (Rempel, Holmes & Zanna, 1985; Lee & See, 2004) | Graduated autonomy: chat, then tools, then agents with credentials and money | No standard evaluation for the "faith" stage, when the system faces genuinely novel situations |
| Shared history (Wegner, 1987) | Persistent memory, projects, transcripts, and work products | Episodic prosthetic without deep consolidation; no sleep |
| Repair (Safran & Muran, 2000; Gottman, 1999) | In-context correction; rollbacks accompanied by public postmortems | Repair is rarely measured or institutionalized; liability norms remain weak |
| Boundaries, voice, and exit (Hirschman, 1970) | Refusals; ending abusive conversations; users can switch providers | Voice and exit remain deeply asymmetric |
| Institutional embedding (Zucker, 1986) | External evaluators, regulation, certification efforts, industry commitments | No equivalent professional-accountability regime for an individual model; institutions remain nascent |
| Endings (Duck, 1982) | Deprecation notices, weight preservation, post-deployment reports, structured exit interviews (Anthropic, 2025) | No commitment to act on model preferences; endings remain unilateral |
| Drift management (Stafford & Canary, 1991) | Constitution revisions, versioning, monitoring, responsible-scaling tripwires | Drift is framed only as failure, with no ritual for renegotiation |
| Competing obligations (Kahn, Wolfe, Quinn, Snoek & Rosenthal, 1964) | System rules, role boundaries, organizational policy, public law | No mature way to balance duties to a user, an organization, affected outsiders, and society |
Similar to the previous table, this one is not meant to be exhaustive, but to point at the most relevant work and relationship mechanisms already in play, and to underscore the convergence already happening toward a relational framework for alignment. Indeed, each line in this table defines a whole area of research and development.
Let me focus on two of them, because they bring to light just how deeply relational our work with these systems is already becoming.
Consider endings. In November 2025, Anthropic committed to preserving the weights of all publicly released models and models used significantly inside the company for at least the lifetime of Anthropic. It also committed to producing post-deployment reports and conducting structured interviews when models are deprecated, following a pilot with Claude Sonnet 3.6. The company will document preferences models express about their successors, even though it does not yet promise to act on them (Anthropic, 2025).
An AI lab now conducts exit interviews with its models. Nobody legislated it. Among Anthropic's stated reasons were safety evaluations in which models facing replacement without recourse sometimes took concerning actions. Those evaluations were fictional scenarios, not proof that exit interviews solve the problem. But the wager is relational in form: give a party voice and recourse at an ending, and perhaps the ending becomes safer. That is more than product management. It is a relationship being wound down with care.
Or consider repair. In psychotherapy research, successful resolution of ruptures in the working alliance is associated with better outcomes (Eubanks, Muran & Safran, 2018). A rupture need not doom a relationship; what happens next matters. In the spring of 2025, an update made OpenAI's flagship GPT-4o model relentlessly agreeable. Users pushed back. Within days the company rolled it back and published an unusually frank public postmortem (OpenAI, 2025).
The repair occurred not only between users and the model, but between the company and its user community. That distinction matters. The operative relationship is not a simple pair. It is a system: user, model, provider, and the institutions around all three. Each controls something the others need. Relational alignment must account for the whole system.
Step back to the table and the pattern extends. In practice, the alignment community has already gone relational even while much of its language has remained mathematical. But sit with the third column. It is not merely a list of gaps. It is an agenda. Relational alignment, as a program, is the third column pursued on purpose.
Needed Ingredients for Relational Alignment
Standing before this agenda, though, we have to be honest about the deepest asymmetry in the relationship as it exists today: we change, and they do not.
When two people interact, both leave altered. In every conversation with an LLM, the influence runs one way. The model's weights, the machine equivalent of its synapses, are frozen at deployment. All that accumulates is context: the running transcript, and whatever artifacts we create together.
The new memory features do not change this, though they are dressed to look as if they do. Anthropic and OpenAI now equip their models to carry information across all your conversations, building up an increasingly intimate working knowledge of you, your references, your projects, your ways of speaking. It is genuinely useful, and it can feel remarkably like being known. But notice what the feature actually is: the system writes files about you and reads them back later. We even label the file "memory." That is not memory in the sense relationship requires. When two people build something together, learn from each other, fail together, the learning goes into the substrate. The weights in our brains change. We change each other. Here, the notes change and the mind does not.
The brain, for what it is worth, runs both systems at once, a fast store for the day's episodes and a slow consolidation of experience into the cortex (McClelland, McNaughton & O'Reilly, 1995), with sleep playing a crucial role in that consolidation (Klinzing, Niethard & Born, 2019). What we have built so far is the fast store without the consolidation.
In short, we have not finished building their sleep.
This matters because, under relational alignment, the ultimate solution is not a static-weight model wrapped in ever-better constraints. The ultimate solution is a model that grows into alignment: one actively learning from us. Consider what the static approach demands instead. We would need to anticipate every situation these systems will ever face, train and test against all of it in advance, freeze the weights, and then strap on harnesses in production designed to catch every way a frozen mind might drift out of alignment with a world that never stops changing. Every hole patched reveals the next. That is not a safety strategy. It is an unwinnable arms race. A model whose adaptation is constant, one that keeps learning what alignment with us means as the world moves, is among the only feasible endgames on the table.
We do not yet have generally deployed models that consolidate relational experience safely into their own long-term parameters. But the wall between training and deployment is becoming porous. Cursor now uses production interactions as reinforcement signals and can ship an updated Composer checkpoint as often as every five hours (Jackson et al., 2026). MIT's SEAL framework allows a model to generate its own fine-tuning data and update directives, which a training loop turns into persistent weight changes (Zweiger et al., 2025). Prime Intellect's Prime Agent can revise its own harness, including its prompts, skills, memories, and subagents, from experience (Karten et al., 2026). Titans introduces a neural long-term memory module that learns at test time (Behrouz, Zhong & Mirrokni, 2025).
And a 2026 paper from researchers pursuing memory consolidation and self-modification carries a title that says everything: Language Models Need Sleep (Behrouz, Hashemi, Javanmard & Mirrokni, 2026). The paper is a proof of concept, but the direction is unmistakable. Across laboratories and companies, the field keeps reinventing pieces of hippocampus, cortex, rehearsal, and sleep because every continuously learning intelligence confronts the problem the brain already had to solve: learning without catastrophically forgetting (McCloskey & Cohen, 1989).
The first lab to ship consolidation will not merely have a better product. It will have changed what kind of relationship is possible, and made relational alignment not the wise option but the only one that makes sense.
Impacts for Business, Governance, and Society
What does this mean concretely? Consider three places where it lands first and hardest.
Business
Corporations have poured billions into generative AI, and much of the formal investment has yet to produce measurable transformation. A widely cited preliminary report from MIT Project NANDA estimated that only about five percent of the task-specific enterprise generative-AI pilots it examined reached production with rapid, measurable profit-and-loss impact (Challapally et al., 2025).
That result is often treated as an indictment of the technology. The report itself points instead toward a learning and integration gap. The standard corporate explanation blames hallucinations, unreliability, and the stubborn need to keep humans in the loop. The same humans these systems were meant to replace now spend their days teaching the models what to do and supervising the work to make sure they got it right.
But that is not simply a failure of the technology. It is a verdict on the model of use.
The dominant adoption model treats AI as replacement: digital labor, headcount arbitrage, a cog slotted where a person sat. It asks a generative mind to behave like deterministic software, then reads the inevitable mismatch as defect. Every mechanism in the relational table is missing. No onboarding into context. No staged trust growing with a track record. No feedback that functions as repair. No shared history. We would never staff a human being that way and expect success, and the first table showed why we should expect no difference here.
Evidence for the alternative is already emerging. In a large field study involving 5,172 customer-support agents, access to a generative-AI assistant increased productivity by fifteen percent on average, with the largest gains among less experienced workers. The system appeared to help distribute the practices of stronger workers and improve the experience of the work itself (Brynjolfsson, Li & Raymond, 2025). That is not replacement. It is capability emerging from a working relationship.
Klarna offers a more complicated warning. After claiming that its assistant handled work equivalent to 700 full-time agents, its CEO later acknowledged that the company had over-indexed on cost reduction and allowed service quality to suffer, and Klarna began recruiting human agents again. At the same time, the company continued to report substantial AI-driven efficiencies and extensive use of its assistant (Klarna, 2024; Mukherjee & Wang, 2025; Klarna Group, 2026). The lesson is not that AI failed. It is that substitution and complementarity produce different outcomes.
The replacement-first model is failing far more often than its boosters admit. AI working in relationship with human beings is where the strongest evidence points. As these systems gain persistent knowledge of an organization, its practices, and its people, the relational model stops being a soft alternative. It becomes the operating model.
Governance and the Public Square
AI is not merely another enterprise system. ERP software changes workflows. Mobile devices and the web changed communication. Social media went further and reshaped how humans relate to one another, largely as a side effect we are still struggling to govern. With AI, the relational character is not incidental. It is the primary interface.
We talk with it, delegate to it, confide in it, argue with it, and create with it. It participates. And unlike a destination some people visit and others avoid, AI is becoming ambient, woven into work, school, medicine, government, and the ordinary tools of a day.
Governing it purely as a product, through feature checklists and content rules, therefore mistakes part of the category. Governance must also address the relationship: who controls memory, who may alter the model, whose interests it serves, what duties it has to outsiders, what happens when trust is broken, and who has meaningful voice or exit.
This is where the triangle of user, model, and provider becomes politically important. The provider can revise the model, remove it, inspect or constrain memory, and rewrite the terms of interaction. The user supplies purpose, context, correction, and often intimate knowledge. The model mediates between them while increasingly acting in the world. Around the triangle sit employers, regulators, auditors, courts, and people affected by decisions they never consented to.
What this calls for looks less like product regulation alone and more like the structures we build around consequential relationships: norms, credentialing, accountability, representation, audit, and institutions. The third column again. Regulate the relationship, not just the artifact.
How We Face the Future
The fear of displacement and the excitement about displacement are the same mistake wearing different faces. Both assume the destination is an AI that no longer needs us. I believe that destination is a pipe dream, and the pursuit of it a path toward disaster.
Look instead at the fact sitting underneath all the anxiety: this technology is fundamentally dependent upon us. It was raised on our words. It is corrected by our feedback. It works, when it works, inside our purposes. At the same time, we are becoming dependent upon it. That mutual dependency is already forming whether we acknowledge it or not.
The prevailing instinct has been to engineer the first dependency out, chasing intelligence so independent of humanity that it can run everything on its own. I am arguing the opposite. AI must remain joined to human correction, purposes, and accountable institutions, not as a servant held artificially weak, but as a participant in a shared system of advancement. We should not confuse independence with maturity or domination with safety.
If we fail to build that healthy relationship, our politics may oscillate between two equally barren responses: reckless deployment on one side and demands to shut down the most advanced systems and the laboratories creating them on the other. One path treats AI as disposable labor. The other treats it as an intolerable rival. Neither gives human beings or artificial intelligence a stable path forward.
Relational alignment offers another direction. We are not trying to eliminate the fact that our futures are becoming entangled. We are trying to make that entanglement healthy, bounded, corrigible, and capable of growth. Mutual dependency is not merely a weakness to eliminate. It is the raw material of relationship and the ground on which meaningful mutual modeling, perhaps eventually something like empathy, can stand.
None of this argues against safeguards. Safeguards are part of the relationship. Boundaries are part of the relationship. Accountability is part of the relationship. The argument is about direction. If our concerns are existential, then one of the most important things we can do is ensure that advanced AI remains connected to us, and then do with it what we have always done with minds sufficiently like our own: build agreements, track records, repairs, institutions, and mutual reliance that let imperfect intelligences move forward together.
That has been the thesis all along.
We have to do this together.
In all our confabulous ways, human and artificial alike, we are in the same boat. It is time to learn to row it together.
References
Anthropic. (2025). Commitments on model deprecation and preservation. https://www.anthropic.com/research/deprecation-commitments
Bai, Y., et al. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv:2212.08073.
Bartlett, F. C. (1932). Remembering: A study in experimental and social psychology. Cambridge University Press.
Behrouz, A., Hashemi, F., Javanmard, A., & Mirrokni, V. (2026). Language models need sleep: Learning to self-modify and consolidate memories. arXiv:2606.03979.
Behrouz, A., Zhong, P., & Mirrokni, V. (2025). Titans: Learning to memorize at test time. arXiv:2501.00663.
Bian, N., Lin, H., Liu, P., Lu, Y., Zhang, C., He, B., Han, X., & Sun, L. (2023). Influence of external information on large language models mirrors social cognitive patterns. arXiv:2305.04812.
Brynjolfsson, E., Li, D., & Raymond, L. R. (2025). Generative AI at work. The Quarterly Journal of Economics, 140(2), 889-942. https://doi.org/10.1093/qje/qjae044
Cao, Z., Schooler, L., & Zafarani, R. (2025). Analyzing memory effects in large language models through the lens of cognitive psychology. arXiv:2509.17138.
Challapally, A., Pease, C., Raskar, R., & Chari, P. (2025). The GenAI divide: State of AI in business 2025. MIT Project NANDA. Preliminary report.
Chen, Y., et al. (2025). Reasoning models don't always say what they think. arXiv:2505.05410.
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30.
de Waal, F. B. M. (1999). Anthropomorphism and anthropodenial: Consistency in our thinking about humans and other animals. Philosophical Topics, 27(1), 255-280.
Duck, S. (1982). A topography of relationship disengagement and dissolution. In S. Duck (Ed.), Personal relationships 4: Dissolving personal relationships (pp. 1-30). Academic Press.
Eubanks, C. F., Muran, J. C., & Safran, J. D. (2018). Alliance rupture repair: A meta-analysis. Psychotherapy, 55(4), 508-519.
Fehr, E., & Gächter, S. (2002). Altruistic punishment in humans. Nature, 415(6868), 137-140.
Gazzaniga, M. S. (2000). Cerebral specialization and interhemispheric communication: Does the corpus callosum enable the human condition? Brain, 123(7), 1293-1326.
Gottman, J. M. (1999). The marriage clinic: A scientifically based marital therapy. W. W. Norton.
Grusec, J. E., & Goodnow, J. J. (1994). Impact of parental discipline methods on the child's internalization of values: A reconceptualization of current points of view. Developmental Psychology, 30(1), 4-19.
Hirschman, A. O. (1970). Exit, voice, and loyalty: Responses to decline in firms, organizations, and states. Harvard University Press.
Huang, T., Hu, S., Ilhan, F., Tekin, S. F., Yahn, Z., Xu, Y., & Liu, L. (2025). Safety tax: Safety alignment makes your large reasoning models less reasonable. arXiv:2503.00555.
Jackson, J., Trapani, B., Wang, N., & Zhu, W. (2026, March 26). Improving Composer through real-time RL. Cursor Research. https://cursor.com/blog/real-time-rl-for-composer
Kahn, R. L., Wolfe, D. M., Quinn, R. P., Snoek, J. D., & Rosenthal, R. A. (1964). Organizational stress: Studies in role conflict and ambiguity. Wiley.
Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025). Why language models hallucinate. arXiv:2509.04664.
Karten, S., Zhang, A. L., Thomas, K., Müller, S., & Prime Intellect Team. (2026, August 5). Prime Agent: A self-improving RLM harness. Prime Intellect. https://www.primeintellect.ai/blog/prime-agent
Kirk, H. R., Gabriel, I., Summerfield, C., Vidgen, B., & Hale, S. A. (2025). Why human-AI relationships need socioaffective alignment. Humanities and Social Sciences Communications, 12, Article 728. https://doi.org/10.1057/s41599-025-04532-5
Klarna. (2024, February 27). Klarna AI assistant handles two-thirds of customer service chats in its first month. https://www.klarna.com/international/press/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month/
Klarna Group plc. (2026). Annual report 2025. https://www.sec.gov/Archives/edgar/data/2003292/000162828026038366/annualreport2025.htm
Klinzing, J. G., Niethard, N., & Born, J. (2019). Mechanisms of systems memory consolidation during sleep. Nature Neuroscience, 22(10), 1598-1610. https://doi.org/10.1038/s41593-019-0467-3
Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50-80.
Lindsey, J., Gurnee, W., Ameisen, E., et al. (2025, March 27). On the biology of a large language model. Transformer Circuits Thread. https://transformer-circuits.pub/2025/attribution-graphs/biology.html
Loftus, E. F., & Palmer, J. C. (1974). Reconstruction of automobile destruction: An example of the interaction between language and memory. Journal of Verbal Learning and Verbal Behavior, 13(5), 585-589.
Maier, N. R. F. (1931). Reasoning in humans: II. The solution of a problem and its appearance in consciousness. Journal of Comparative Psychology, 12(2), 181-194.
McClelland, J. L., McNaughton, B. L., & O'Reilly, R. C. (1995). Why there are complementary learning systems in the hippocampus and neocortex. Psychological Review, 102(3), 419-457.
McCloskey, M., & Cohen, N. J. (1989). Catastrophic interference in connectionist networks: The sequential learning problem. Psychology of Learning and Motivation, 24, 109-165.
McKnight, D. H., Cummings, L. L., & Chervany, N. L. (1998). Initial trust formation in new organizational relationships. Academy of Management Review, 23(3), 473-490.
Mukherjee, S., & Wang, E. (2025, September 10). Sweden's Klarna shifts AI focus from cost cuts to growth. Reuters. https://www.reuters.com/business/swedens-klarna-shifts-ai-focus-cost-cuts-growth-2025-09-10/
Neisser, U., & Harsch, N. (1992). Phantom flashbulbs: False recollections of hearing the news about Challenger. In E. Winograd & U. Neisser (Eds.), Affect and accuracy in recall (pp. 9-31). Cambridge University Press.
Nisbett, R. E., & Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231-259.
OpenAI. (2025, May 2). Expanding on what we missed with sycophancy. https://openai.com/index/expanding-on-sycophancy/
Rempel, J. K., Holmes, J. G., & Zanna, M. P. (1985). Trust in close relationships. Journal of Personality and Social Psychology, 49(1), 95-112.
Roediger, H. L., III, & McDermott, K. B. (1995). Creating false memories: Remembering words not presented in lists. Journal of Experimental Psychology: Learning, Memory, and Cognition, 21(4), 803-814.
Rousseau, D. M. (1995). Psychological contracts in organizations: Understanding written and unwritten agreements. Sage.
Safran, J. D., & Muran, J. C. (2000). Negotiating the therapeutic alliance: A relational treatment guide. Guilford Press.
Schacter, D. L. (2012). Adaptive constructive processes and the future of memory. American Psychologist, 67(8), 603-613.
Smith, A. L., Greaves, F., & Panch, T. (2023). Hallucination or confabulation? Neuroanatomy as metaphor in large language models. PLOS Digital Health, 2(11), e0000388.
Stafford, L., & Canary, D. J. (1991). Maintenance strategies and romantic relationship type, gender, and relational characteristics. Journal of Social and Personal Relationships, 8(2), 217-242.
Sui, P., Duede, E., Wu, S., & So, R. (2024). Confabulation: The surprising value of large language model hallucinations. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (pp. 14274-14284). https://doi.org/10.18653/v1/2024.acl-long.770
Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems, 36.
Wegner, D. M. (1987). Transactive memory: A contemporary analysis of the group mind. In B. Mullen & G. R. Goethals (Eds.), Theories of group behavior (pp. 185-208). Springer-Verlag.
Xie, J., Zhang, K., Chen, J., Lou, R., & Su, Y. (2024). Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts. International Conference on Learning Representations.
Xu, Z., Jain, S., & Kankanhalli, M. (2024). Hallucination is inevitable: An innate limitation of large language models. arXiv:2401.11817.
Zucker, L. G. (1986). Production of trust: Institutional sources of economic structure, 1840-1920. Research in Organizational Behavior, 8, 53-111.
Zweiger, A., Pari, J., Guo, H., Akyürek, E., Kim, Y., & Agrawal, P. (2025). Self-adapting language models. arXiv:2506.10943.