# Generative Agents: Interactive Simulacra of Human Behavior

**Authors:** Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein

**Affiliations:** Stanford University, Google Research, Google DeepMind

**Publication:** UIST '23, October 29-November 1, 2023, San Francisco, CA, USA

**arXiv:** 2304.03442v2 [cs.HC] - August 6, 2023

**DOI:** https://doi.org/10.1145/3586183.3606763

---

## Abstract

Believable proxies of human behavior can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication to prototyping tools. In this paper, we introduce generative agents: computational software agents that simulate believable human behavior. Generative agents wake up, cook breakfast, and head to work; artists paint, while authors write; they form opinions, notice each other, and initiate conversations; they remember and reflect on days past as they plan the next day.

To enable generative agents, we describe an architecture that extends a large language model to store a complete record of the agent's experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior. We instantiate generative agents to populate an interactive sandbox environment inspired by The Sims, where end users can interact with a small town of twenty-five agents using natural language.

In an evaluation, these generative agents produce believable individual and emergent social behaviors. For example, starting with only a single user-specified notion that one agent wants to throw a Valentine's Day party, the agents autonomously spread invitations to the party over the next two days, make new acquaintances, ask each other out on dates to the party, and coordinate to show up for the party together at the right time.

---

## 1 INTRODUCTION

How might we craft an interactive artificial society that reflects believable human behavior? From sandbox games such as The Sims to applications such as cognitive models and virtual environments, for over four decades, researchers and practitioners have envisioned computational agents that can serve as believable proxies of human behavior. In these visions, computationally-powered agents act consistently with their past experiences and react believably to their environments.

Such simulations of human behavior could populate virtual spaces and communities with realistic social phenomena, train people on how to handle rare yet difficult interpersonal situations, test social science theories, craft model human processors for theory and usability testing, power ubiquitous computing applications and social robots, and underpin non-playable game characters that can navigate complex human relationships in an open world.

However, the space of human behavior is vast and complex. Despite striking progress in large language models that can simulate human behavior at a single time point, fully general agents that ensure long-term coherence would be better suited by architectures that manage constantly-growing memories as new interactions, conflicts, and events arise and fade over time while handling cascading social dynamics that unfold between multiple agents.

Success requires an approach that can retrieve relevant events and interactions over a long period, reflect on those memories to generalize and draw higher-level inferences, and apply that reasoning to create plans and reactions that make sense in the moment and in the longer-term arc of the agent's behavior.

In this paper, we introduce generative agents—agents that draw on generative models to simulate believable human behavior—and demonstrate that they produce believable simulacra of both individual and emergent group behavior. Generative agents draw a wide variety of inferences about themselves, other agents, and their environment; they create daily plans that reflect their characteristics and experiences, act out those plans, react, and re-plan when appropriate; they respond when the end user changes their environment or commands them in natural language.

---

## Key Contributions

1. **Generative agents** - believable simulacra of human behavior that are dynamically conditioned on agents' changing experiences and environment

2. **A novel architecture** that makes it possible for generative agents to remember, retrieve, reflect, interact with other agents, and plan through dynamically evolving circumstances

3. **Two evaluations** - a controlled evaluation and an end-to-end evaluation that establish causal effects of the importance of components of the architecture

4. **Discussion** of the opportunities and ethical and societal risks of generative agents in interactive systems

---

## 2 RELATED WORK

### 2.1 Human-AI Interaction

Interactive artificial intelligence systems aim to combine human insights and capabilities in computational artifacts that can augment their users. A long line of work has explored ways to enable users to interactively specify model behavior. Recent advancements have extended these explorations to deep learning and prompt-based authoring.

Meanwhile, a persistent thread of research has advanced the case for language- and agent-based interaction in human-computer interaction. Formative work such as SHRDLU and ELIZA demonstrated the opportunities and the risks associated with natural language interaction with computing systems.

### 2.2 Believable Proxies of Human Behavior

Prior literature has described believability, or believable agents, as a central design and engineering goal. Believable agents are designed to provide an illusion of life and present a facade of realism in the way they appear to make decisions and act on their own volition, similar to the characters in Disney movies.

A diverse set of approaches to creating believable agents emerged over the past four decades:
- **Rule-based approaches** (finite-state machines, behavior trees) - manual authoring
- **Learning-based approaches** (reinforcement learning) - achieved success in adversarial games but not open worlds
- **Cognitive architectures** (SOAR, ICARUS) - maintained memories but limited to manually crafted procedures

Today, creating believable agents remains an open problem. Our argument is that large language models offer an opportunity to re-examine these questions, provided that we can craft an effective architecture to synthesize memories into believable behavior.

### 2.3 Large Language Models and Human Behavior

Generative agents leverage a large language model to power their behavior. The key observation is that large language models encode a wide range of human behavior from their training data.

Existing literature largely relies on first-order templates that employ few-shot prompts or chain-of-thought prompts. These templates are effective in generating behavior conditioned solely on the agent's current environment. However, believable agents require conditioning not only on their current environment but also on a vast amount of past experience, which is a poor fit using first-order prompting.

---

## 3 GENERATIVE AGENT BEHAVIOR AND INTERACTION

To illustrate the affordances of generative agents, we instantiate them as characters in a simple sandbox world reminiscent of The Sims. This sprite-based sandbox game world, Smallville, evokes a small town environment.

### 3.1 Agent Avatar and Communication

A community of 25 unique agents inhabits Smallville. Each agent is represented by a simple sprite avatar. We authored one paragraph of natural language description to depict each agent's identity, including their occupation and relationship with other agents, as seed memories.

**Example - John Lin:**
> John Lin is a pharmacy shopkeeper at the Willow Market and Pharmacy who loves to help people. He is always looking for ways to make the process of getting medication easier for his customers; John Lin is living with his wife, Mei Lin, who is a college professor, and son, Eddy Lin, who is a student studying music theory; John Lin loves his family very much; John Lin has known the old couple next-door, Sam Moore and Jennifer Moore, for a few years...

Each semicolon-delimited phrase is entered into the agent's initial memory as memories at the start of the simulation.

#### 3.1.1 Inter-Agent Communication

The agents interact with the world by their actions, and with each other through natural language. At each time step of the sandbox engine, the agents output a natural language statement describing their current action.

Agents communicate with each other in full natural language. They are aware of other agents in their local area, and the generative agent architecture determines whether they walk by or engage in conversation.

#### 3.1.2 User Controls

The user communicates with the agent through natural language by specifying a persona that the agent should perceive them as. To directly command one of the agents, the user takes on the persona of the agent's "inner voice"—this makes the agent more likely to treat the statement as a directive.

### 3.2 Environmental Interaction

Smallville features the common affordances of a small village, including a cafe, bar, park, school, dorm, houses, and stores. It also defines subareas and objects that make those spaces functional.

Users and agents can influence the state of the objects in this world, much like in sandbox games such as The Sims. End users can also reshape an agent's environment in Smallville by rewriting the status of objects surrounding the agent in natural language.

### 3.3 Example "Day in the Life"

Starting from the single-paragraph description, generative agents begin planning their days. As time passes in the sandbox world, their behaviors evolve as these agents interact with each other and the world, building memories and relationships, and coordinating joint activities.

We demonstrate the behavior of generative agents by tracing the output of our system over the course of one day for the agent John Lin:
- 7:00 AM - Wakes up, brushes teeth, showers, gets dressed, eats breakfast, checks news
- 8:00 AM - Catches up with son Eddy about his music composition
- 9:00 AM - Opens pharmacy counter at Willow Market and Pharmacy

### 3.4 Emergent Social Behaviors

By interacting with each other, generative agents in Smallville exchange information, form new relationships, and coordinate joint activities. These social behaviors are emergent rather than pre-programmed.

#### 3.4.1 Information Diffusion

As agents notice each other, they may engage in dialogue—as they do so, information can spread from agent to agent. For instance, Sam tells Tom about his candidacy in the local election, and gradually this becomes the talk of the town.

#### 3.4.2 Relationship Memory

Agents in Smallville form new relationships over time and remember their interactions with other agents. Sam meets Latoya Williams, and in a later interaction, remembers her photography project.

#### 3.4.3 Coordination

Generative agents coordinate with each other. Isabella Rodriguez is initialized with an intent to plan a Valentine's Day party. From this seed, the agent proceeds to invite friends, decorate the cafe, and on Valentine's Day, five agents show up to enjoy the festivities—all emergent behavior from a single seed intention.

---

## 4 GENERATIVE AGENT ARCHITECTURE

Generative agents aim to provide a framework for behavior in an open world: one that can engage in interactions with other agents and react to changes in the environment.

At the center of our architecture is the **memory stream**, a database that maintains a comprehensive record of an agent's experience. From the memory stream, records are retrieved as relevant to plan the agent's actions and react appropriately to the environment. Records are recursively synthesized into higher- and higher-level reflections that guide behavior.

Our current implementation utilizes the gpt3.5-turbo version of ChatGPT.

### 4.1 Memory and Retrieval

**Challenge:** Creating generative agents that can simulate human behavior requires reasoning about a set of experiences that is far larger than what should be described in a prompt.

**Approach:** The memory stream maintains a comprehensive record of the agent's experience. It is a list of memory objects, where each object contains a natural language description, a creation timestamp, and a most recent access timestamp.

The retrieval function scores all memories as a weighted combination of three elements:
- **Recency** - Higher score for recently accessed memories (exponential decay, factor 0.995)
- **Importance** - Distinguishes mundane from core memories (1-10 scale, asked via LLM)
- **Relevance** - Semantic similarity to current situation (cosine similarity of embeddings)

**Score formula:** `score = α_recency · recency + α_importance · importance + α_relevance · relevance`

### 4.2 Reflection

**Challenge:** Generative agents, when equipped with only raw observational memory, struggle to generalize or make inferences.

**Approach:** We introduce a second type of memory, called a reflection. Reflections are higher-level, more abstract thoughts generated by the agent. They are generated periodically when the sum of importance scores for the latest events exceeds a threshold (150 in our implementation—roughly 2-3 times per day).

The reflection process:
1. Query LLM with 100 most recent records to identify salient questions
2. Use questions as queries for retrieval, gather relevant memories
3. Prompt LLM to extract insights and cite evidence
4. Store as reflections in memory stream with pointers to source memories

This creates trees of reflections: leaf nodes are base observations, non-leaf nodes become more abstract and higher-level.

### 4.3 Planning and Reacting

**Challenge:** While LLMs can generate plausible behavior in response to situational information, agents need to plan over a longer time horizon to ensure coherent and believable sequences of actions.

**Approach:** Plans describe a future sequence of actions for the agent, and help keep the agent's behavior consistent over time. A plan includes a location, starting time, and duration.

The planning process starts top-down and recursively generates more detail:
1. Create high-level daily agenda (5-8 chunks)
2. Recursively decompose into hour-long chunks
3. Further decompose into 5-15 minute action chunks

#### Reacting and Updating Plans

Generative agents operate in an action loop where they perceive the world and decide whether to continue with their existing plan or react to new observations.

#### Dialogue

Agents converse as they interact with each other. We generate agents' dialogue by conditioning their utterances on their memories about each other.

---

## 5 SANDBOX ENVIRONMENT IMPLEMENTATION

The Smallville sandbox game environment is built using the Phaser web game development framework. We supplement it with a server that makes sandbox information available to generative agents and enables them to move and influence the sandbox environment.

### 5.1 From Structured World Environments to Natural Language, and Back Again

The architecture operates using natural language. We represent the sandbox environment as a tree data structure, with edges indicating containment relationships. Agents build individual tree representations as they navigate—they are not omniscient, and their tree may get out of date as they leave an area.

---

## 6 CONTROLLED EVALUATION

We evaluate generative agents in two stages: a controlled evaluation to test individual agent responses, and an end-to-end evaluation of emergent community behavior.

### 6.1 Evaluation Procedure

We "interview" agents to probe their ability to:
- Maintain **self-knowledge**
- Retrieve **memory** of past events
- Generate **plans** for future actions
- **React** to unexpected events
- **Reflect** on experiences

25 questions across 5 categories were asked to agents sampled from the end of two game days.

### 6.2 Conditions

We compared four architectures:
1. **Full architecture** - all components enabled
2. **No reflection** - observations and plans only
3. **No reflection or planning** - observations only
4. **No observation, reflection, or planning** - baseline (prior art)

Plus a human crowdworker-authored condition.

### 6.5 Results

**The full architecture produced the most believable behavior** (μ = 29.89; σ = 0.72). Performance degraded with each removed component:
- No reflection: μ = 26.88
- No reflection or planning: μ = 25.64
- Crowdworker: μ = 22.95
- No memory/planning/reflection: μ = 21.21

**Effect size:** Comparing prior work to full architecture produces d = 8.16 standard deviations.

**Common errors:**
- Failed to retrieve relevant memories
- Fabricated embellishments to memory
- Inherited overly formal speech from the language model

---

## 7 END-TO-END EVALUATION

We deployed 25 agents interacting continuously over two full game days in Smallville.

### 7.1 Emergent Social Behaviors

#### Information Diffusion

- Sam's candidacy spread from 1 → 8 agents (32%)
- Party news spread from 1 → 13 agents (52%)
- No hallucinated information

#### Relationship Formation

- Network density increased from 0.167 → 0.74
- Only 1.3% of responses were hallucinated

#### Coordination

- 5 agents attended the Valentine's Day party
- 3 agents cited conflicts for not attending
- 4 agents expressed interest but didn't plan to come

### 7.2 Boundaries and Errors

Three common modes of erratic behavior:

1. **Location confusion** - As agents learned more places, they sometimes chose less typical locations
2. **Norm violations** - Entering occupied bathrooms or closed stores
3. **Instruction tuning effects** - Overly formal dialogue and excessive cooperation

---

## 8 DISCUSSION

### 8.1 Applications of Generative Agents

- Social simulacra for prototyping social computing systems
- Virtual reality metaverses and social robots
- Human-centered design process tools
- Cognitive models for usability testing (GOMS, KLM)

### 8.2 Future Work and Limitations

- Enhance retrieval module for more relevant context
- Improve cost-effectiveness (current study cost thousands of dollars)
- Explore parallelization for real-time interactivity
- Extended evaluation over longer periods
- Robustness testing (prompt hacking, memory hacking, hallucination)

### 8.3 Ethics and Societal Impact

1. **Parasocial relationships** - Users may form inappropriate attachments
2. **Error impact** - Wrong inferences could cause harm in real applications
3. **Deepfakes and misinformation** - Recommend audit logs
4. **Over-reliance** - Should complement, not replace human stakeholders

---

## 9 CONCLUSION

This paper introduces generative agents, interactive computational agents that simulate human behavior. We describe an architecture for generative agents that provides a mechanism for storing a comprehensive record of an agent's experiences, deepening its understanding of itself and the environment through reflection, and retrieving a compact subset of that information to inform the agent's actions.

Looking ahead, we suggest that generative agents can play roles in many interactive applications, ranging from design tools to social computing systems to immersive environments.

---

## Resources

- **Demo:** https://reverie.herokuapp.com/UIST_Demo/
- **Code:** https://github.com/joonspk-research/generative_agents

---

*File processed: February 16, 2026*
*Source: arXiv:2304.03442v2 (22 pages, 1367 lines extracted)*