Research / Note 05
Personal AI Needs a Theory of the Person
- Category
- Personal AI
- Published
- Reading time
- 12 min
- Author
- Stéphane Benayoun, Co-founder and Chief Technology Officer
Abstract
Personal AI is often framed as a memory problem: retain enough of a user's history, retrieve the right information, and future interactions can become more relevant. Research on personalized language models and long-term conversational memory shows that these mechanisms are useful. It also exposes a second problem. As evidence accumulates across situations and over time, a system must decide what that evidence implies about the person it is trying to assist.
We use theory of the person in a deliberately limited sense: not a psychological model or digital replica, but a set of representational commitments about what should persist, what may change, how observations relate to one another, and how uncertainty should be preserved.
Our view is that this problem will become increasingly important as personal AI moves from isolated responses toward longitudinal interaction and action.
Large models know an extraordinary amount about the world.
They know remarkably little about the particular person in front of them.
This statement needs one qualification. A pretrained language model and a deployed personal AI system are not the same thing. Contemporary assistants can retain conversation histories, retrieve previous interactions, maintain profiles or use external memory systems. The limitation is therefore no longer simply that AI systems cannot retain information about an individual. Increasingly, the question is what form that information should take once it has accumulated.
For isolated interactions, the distinction may matter little. A model can translate a sentence, explain a mathematical concept or generate code without constructing a durable representation of the person asking. Personal AI creates a different requirement. If a system is expected to become more useful through repeated interaction, it needs some way to relate what is happening now to evidence acquired before: previous preferences, corrections, intentions, decisions and changes of circumstance.
Memory is an important part of that problem. It is not the whole of it.
Established research
Personalization can be achieved in several ways
Recent work on personalized language models provides useful evidence against treating personalization as a single technical problem.
The LaMP benchmark introduced seven personalized language tasks and showed that retrieving relevant items from a user’s profile could improve both classification and generation. Its experiments covered lexical, semantic and time-aware retrieval approaches, demonstrating that user history can be useful without modifying the underlying model for every individual (Salemi et al., 2024).
Other work reaches personalization through different representations. Persona-Plug constructs a user-specific embedding from historical material and reports stronger results than the personalization baselines it evaluates on LaMP (Liu et al., 2025). Personalized-RLHF instead introduces a learned user model into preference alignment so that different users need not be treated as if their feedback came from a single preference distribution (Li et al., 2024).
These results support a relatively narrow conclusion: representing or retrieving user-specific information can improve performance on personalized tasks. They do not establish that one representation is universally preferable, nor that success on those tasks amounts to a general understanding of the individual. In fact, their diversity is informative. Relevant aspects of a person can be exposed to a model through retrieved examples, textual profiles, learned representations, adapted parameters or combinations of these mechanisms.
The representational question begins when we ask what information each approach preserves and what distinctions future reasoning may require.
Long-term memory introduces temporal reasoning
Research on persistent conversational systems makes this question more visible.
Generative Agents combined stored observations, retrieval and reflection to support agents whose subsequent behaviour depended on accumulated experience (Park et al., 2023). MemGPT approached the constraint from a systems perspective, using hierarchical memory management to maintain useful context across interactions that exceeded the language model’s immediate context window (Packer et al., 2023).
Later benchmarks have made the evaluation problem more precise. LoCoMo contains conversations extending across many sessions and evaluates question answering, event summarization and multimodal dialogue generation over those histories. Maharana et al. found that long-context models and retrieval augmentation improve performance, but that models continue to have difficulty with long-range temporal and causal relationships (Maharana et al., 2024).
LongMemEval goes further in decomposing long-term interactive memory into information extraction, multi-session reasoning, temporal reasoning, knowledge updates and abstention (Wu et al., 2025). That decomposition is important. A system may successfully recover something a user said months ago and still fail to determine whether the information remains current, how it relates to later evidence, or whether there is enough evidence to answer at all.
The engineering problem of access to history therefore coexists with a reasoning problem about history.
Preferences are not always stable attributes
A similar distinction is emerging in work on personalization itself.
Gao et al. propose separating stable preferences from situational preferences in S2Pref, a 2026 benchmark designed around cases where contextual evidence changes how a user’s preference should be interpreted. Their taxonomy should not be treated as a definitive ontology of human preference, but it captures a practical failure mode of static profiles: apparently conflicting behaviours may be coherent once their situations are taken into account (Gao et al., 2026).
PersonalAgent approaches longitudinal personalization by continually refining a user profile from dialogue history rather than treating preference information as fixed after an initial interaction (Zhang et al., 2026). Research is also beginning to evaluate what happens when personalization affects actions rather than only generated text. ETAPP studies personalized tool invocation, while APOLLO tests whether agents can recover explicit and implicit preferences from interaction histories and apply them when selecting tools (Hao et al., 2025; Chen et al., 2026).
The practical consequence is that errors in user modelling need not remain confined to the wording of an answer. When an agent acts on behalf of a user, assumptions about that user can influence what the system chooses to do.
User modelling is an older problem than personal AI
The underlying problem substantially predates language models.
User-modelling research has long studied representations of knowledge, goals, interests, preferences and other user-specific properties so that interactive systems can adapt their behaviour. By 2001, Kobsa was already reviewing two decades of work on generic user-modelling systems and architectures (Kobsa, 2001).
An important strand of this literature also treated the user model as something more than hidden application state. Work on scrutable user models examined how users might inspect, understand and correct representations maintained about them. Kay and Kummerfeld later proposed a long-term personal user model intended to integrate evidence across a learner’s activities while retaining user control over the resulting representation (Kay & Kummerfeld, 2019).
The technical setting is now very different, but the central question is familiar: what should a computational system represent about a particular individual in order to adapt appropriately to that individual?
Personal AI makes the scope of that question much broader.
Interpretation
Access to the past and interpretation of the past are different problems
Consider a hypothetical system with essentially perfect retrieval. It retains every interaction and can surface any previous statement when required.
Suppose a user once says that they generally avoid crowded places. Months later they attend a large music festival, and at another point they reject a restaurant because it is too busy. All three observations can coexist without telling the same story. The festival may have been an exception; avoiding crowds may be a weak preference that is overridden for particular experiences; the person’s attitude may have changed; or the original statement may have been too general.
Retrieving the observations is useful because the system cannot reason about evidence it cannot access. Yet the observations do not contain their own interpretation. Their meaning depends partly on their temporal relationship, the circumstances in which they arose and what other evidence exists.
The same issue appears in less explicit forms of behaviour. A purchase may reflect convenience rather than preference. A search may concern another person. Repeated selection of an option may reveal a durable tendency, or merely the persistence of a constraint. An expressed interest may represent aspiration rather than intention.
For a personal system, accumulating more observations therefore creates a second-order problem: it must determine which conclusions, if any, are justified by them.
Every personalization mechanism makes representational commitments
Calling this a theory of the person does not require a theory of personality or a complete cognitive model.
The claim is more modest. Any system that adapts persistently to an individual has to make decisions about what counts as relevant user state.
A textual profile assumes that some useful properties can be expressed as propositions. A retrieval system assumes that parts of the historical record can become relevant again when sufficiently related to the current interaction. A learned embedding assumes that regularities useful for downstream tasks can be encoded in a compact latent representation. A recency-weighted model implicitly assumes something about the rate at which old evidence loses predictive value.
These assumptions may be entirely appropriate for their intended tasks. The difficulty arises when distinctions needed later have been removed during representation.
A compressed profile that says a user “dislikes crowded places,” for example, is easy to retrieve and use. It may also have discarded the observations from which the conclusion was derived, the contexts in which it held, and the degree of uncertainty that originally surrounded the inference.
This is not an argument against compression. Persistent systems cannot treat every historical event as equally important forever. It is an argument for treating compression as a modelling decision rather than as a neutral storage operation.
Larger context windows do not eliminate that decision
Increasing context capacity can reduce the amount of information that has to be discarded before inference. Retrieval can reduce it further by selecting material likely to be relevant to a particular query.
Neither mechanism removes the requirement to interpret evidence.
A system presented with several years of interaction history must still determine what matters now. It must decide whether later evidence supersedes earlier evidence, whether conflicting observations indicate change or context dependence, and whether a historical preference is relevant to the present intention. If those decisions are deferred to inference time, they remain modelling decisions; they have simply moved into the reasoning process of the model.
The distinction is consequently not between systems that “have a user model” and systems that do not. Even a system that reconstructs its view of a user from raw history at every interaction is performing user modelling when it selects evidence and draws conclusions from it.
TasteGraph perspective
Evidence should remain distinguishable from what is inferred from it
Our view begins with a simple representational distinction: an observation about a person and an inference about that person should not automatically be treated as equivalent.
This matters because evidence can support more than one interpretation. A choice may express preference, but it may also reflect price, availability, habit, social context or a temporary objective. Multiple observations increase the basis for inference without necessarily eliminating ambiguity.
For taste in particular, this distinction is fundamental. Someone may consistently favour one aesthetic direction in their own clothing and a different one when buying a gift. They may seek experimentation for a holiday and familiarity for everyday use. They may admire an object without wanting to own it. None of these cases requires the underlying person to be inconsistent.
A representation that immediately converts every trace into an unconditional preference may therefore discard precisely the structure that later makes the evidence intelligible.
Contradiction can carry information
Persistent systems also need a policy for conflicting evidence.
One possibility is that new evidence should simply replace old information. This is appropriate when a fact has changed: someone has moved to a different city, for example, or explicitly corrected an earlier statement.
Preference evidence is often harder. A contradiction may indicate genuine change, but it may also reveal context dependence, an incorrect previous inference, competing preferences or an exceptional decision. Treating every contradiction as an update risks losing historical structure; treating none of them as updates leaves the model unable to evolve.
LongMemEval’s explicit treatment of knowledge updates and abstention illustrates why this distinction matters operationally. A capable longitudinal system needs to do more than accumulate assertions. It needs some basis for deciding whether information should coexist, be revised or remain unresolved.
Persistence should preserve the possibility of change
The phrase persistent representation can easily be misunderstood as a request for a stable underlying profile.
We mean almost the opposite.
A persistent representation is useful because it maintains continuity across observations, not because it assumes that the properties inferred from those observations never change. A model of an individual over time should be capable of preserving durable patterns while also representing transitions, temporary states and uncertainty about whether an apparent change will persist.
This is particularly important for taste, where continuity and evolution coexist. Some preferences remain stable for long periods; others are acquired, abandoned, revived or expressed only in particular contexts. Modelling the person as a fixed endpoint would therefore be as limiting as modelling each interaction independently.
Persistence is valuable when it gives change a history.
Uncertainty changes what an AI can do
A persistent personal model also needs to distinguish between what it has evidence for and what it merely finds plausible.
This has behavioural consequences. When a system has collapsed ambiguous evidence into a confident user attribute, its only option may be to act on that attribute. When uncertainty remains representable, clarification becomes possible.
Recent personalization benchmarks increasingly test related capabilities. S2Pref includes cases in which contextual signals must be prioritized and ambiguity resolved, while APOLLO distinguishes preferences that are explicitly stated from those that must be inferred through prior behaviour. These are still bounded experimental settings, but they point toward a general design question for personal AI: when is inference justified, and when is obtaining another piece of evidence the better action?
The same question appears in older work on scrutable user models from another direction. If systems maintain increasingly consequential representations of individuals, there is a strong argument for making at least some conclusions inspectable and correctable rather than treating the model’s interpretation as irrevocable fact.
This does not require a digital twin
None of these requirements implies that personal AI needs a complete simulation of a human being.
That would be both an unnecessarily strong technical claim and an unclear objective.
A representation only needs to preserve distinctions relevant to the reasoning the system is expected to perform. Different applications may therefore require different representations, and there is no evidence that one technical form (textual profiles, latent representations, explicit structures, external memory or any particular combination) must dominate.
The important questions are functional.
Can the system distinguish an observation from an inference drawn from it? Can it preserve enough temporal information to recognise change? Can the same preference behave differently across situations without becoming an inconsistency that must immediately be removed? Can old evidence remain available without automatically controlling current behaviour? Can uncertainty survive long enough for the system to ask rather than guess?
These questions concern representation before they concern architecture.

From memory to personal intelligence
The development of personal AI is making increasingly sophisticated forms of memory technically practical. Systems can retain histories, construct profiles, retrieve relevant episodes, learn user representations and carry information across sessions. Research indicates that all of these mechanisms can improve personalization under the right conditions.
The next problem is not to decide which of them counts as “real” memory.
It is to determine what the accumulated evidence should mean.
As a system interacts with the same person across time, it has to preserve some distinctions between durable preference and temporary intention, between evidence and inference, between change and contradiction, and between absence of evidence and evidence of absence. Those distinctions need not be represented explicitly in every implementation, but the system’s behaviour will ultimately depend on how they are resolved.
Taste is one domain in which the problem becomes particularly visible because preference is structured, contextual and evolving. The same representational questions, however, apply more broadly whenever an AI is expected to assist the same individual repeatedly and to act on what it has learned.
A personal AI does not need a computational replica of the person.
It does need a sufficiently persistent account of what it has learned about them, how strongly that account is supported, and when its interpretation may need to change.
Personal intelligence requires persistent representation.
Selected references
Chen, Z.-Y., Lu, S., Xie, Q., Wang, X. & Lin, Y. (2026). “Towards Preference Following in Tool Calling Language Agents.” Findings of the Association for Computational Linguistics: ACL 2026, 33565–33581. doi:10.18653/v1/2026.findings-acl.1676.
Gao, C., Huang, Y., Yao, J., Wu, X. & Feng, J. (2026). “Beyond Static Profiles: Capturing the Fluidity of User Preferences in Diverse Scenarios.” Findings of the Association for Computational Linguistics: ACL 2026, 20618–20634. doi:10.18653/v1/2026.findings-acl.1033.
Hao, Y., Cao, P., Jin, Z., Liao, H., Chen, Y., Liu, K. & Zhao, J. (2025). “Evaluating Personalized Tool-Augmented LLMs from the Perspectives of Personalization and Proactivity.” Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 21897–21935. doi:10.18653/v1/2025.acl-long.1064.
Kay, J. & Kummerfeld, B. (2019). “From data to personal user models for life-long, life-wide learners.” British Journal of Educational Technology, 50(6), 2871–2884. doi:10.1111/bjet.12878.
Kobsa, A. (2001). “Generic User Modeling Systems.” User Modeling and User-Adapted Interaction, 11, 49–63. doi:10.1023/A:1011187500863.
Li, X., Zhou, R., Lipton, Z. C. & Leqi, L. (2024). “Personalized Language Modeling from Personalized Human Feedback.” arXiv:2402.05133.
Liu, J., Zhu, Y., Wang, S., Wei, X., Min, E., Lu, Y., Wang, S., Yin, D. & Dou, Z. (2025). “LLMs + Persona-Plug = Personalized LLMs.” Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 9373–9385. doi:10.18653/v1/2025.acl-long.461.
Maharana, A., Lee, D.-H., Tulyakov, S., Bansal, M., Barbieri, F. & Fang, Y. (2024). “Evaluating Very Long-Term Conversational Memory of LLM Agents.” Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 13851–13870. doi:10.18653/v1/2024.acl-long.747.
Packer, C., Wooders, S., Lin, K., Fang, V., Patil, S. G., Stoica, I. & Gonzalez, J. E. (2023). “MemGPT: Towards LLMs as Operating Systems.” arXiv:2310.08560.
Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P. & Bernstein, M. S. (2023). “Generative Agents: Interactive Simulacra of Human Behavior.” Proceedings of UIST '23. doi:10.1145/3586183.3606763.
Salemi, A., Mysore, S., Bendersky, M. & Zamani, H. (2024). “LaMP: When Large Language Models Meet Personalization.” Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 7370–7392. doi:10.18653/v1/2024.acl-long.399.
Wu, D., Wang, H., Yu, W., Zhang, Y., Chang, K.-W. & Yu, D. (2025). “LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory.” International Conference on Learning Representations (ICLR 2025).
Zhang, X., Wang, Y., Chen, R., Wang, Z., Hou, R. & Liu, Z. (2026). “Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues.” Findings of the Association for Computational Linguistics: ACL 2026, 3221–3240. doi:10.18653/v1/2026.findings-acl.159.