Autonomous AI agents can turn authors' goals into interactive narratives by independently organizing and carrying out generation and revision. As agents generate and revise extensive content, authors struggle to grasp its overall structure, local details, and relationships, complicating continued guidance. We present NarrativeSteward, an authoring environment that organizes outlines, worldbuilding, and narrative graphs as linked artifacts for agent implementation and author guidance. Agent dialogue and project-wide structural review help authors understand the evolving work and guide local and cross-layer revisions, while change records and execution verification help authors assess the resulting work. Technical tests validated the system's change records, recovery mechanisms, and execution diagnostics. In a 12-participant within-subject study, NarrativeSteward supported easier formulation of revision requests and inspection of changes, and greater perceived understanding of changes and story structure, than general-purpose agents. Qualitative findings show how reviewing the work and feedback helps authors develop requirements and guide subsequent delegation. We open-source NarrativeSteward at https://github.com/Tencent/NarrativeSteward.
Figures & tables
Figure 1. Interactive narrative authoring with general-purpose AI agents and NarrativeSteward. (a) With general-purpose AI agents, authors can delegate generation and revision but face the challenge of understanding an evolving narrative across scattered files and representations. (b) NarrativeSteward coordinates delegation, guidance, and verification through linked narrative artifacts. The agent organizes generation and revision across these artifacts, while a project-wide structural view helps authors understand their content and relationships. Together with change records and execution diagnostics, this view helps authors assess the evolving work and guide subsequent revisions through dialogue. A two-panel comparison of interactive narrative authoring. The left panel shows an author and a general-purpose AI agent in dialogue, with documents, graphs, state data, and runtime outputs feeding an incomplete mental representation of the story. The right panel shows NarrativeSteward: an author and agent exchange requests and responses above a project-wide structural view and linked narrative artifacts. A downward arrow from the author to the structural view indicates inspection and assessment. File cards and field snippets depict the structured project data on the right; a leftward visualization arrow connects them to outline, character, graph, and state-rule representations in the structural view. The agent reads, creates, and revises artifacts; the structural view displays their content and relationships and provides optional direct editing. A deterministic layer maintains structural references, change records and recovery, and execution verification, with changes and diagnostics returning to the structural view.
Figure 2. The NarrativeSteward authoring interface. Structural review, local inspection, and agent dialogue are connected in a shared workspace. Artifact navigation (A) provides access to the narrative layers. The structural view (B) shows the branching relationships of the selected event, The Fork , while the selected-object editor (C) exposes its content and references. In the agent dialogue (D), the author requests a more dangerous atmosphere in both the location card and the event while retaining two playable routes. A change record summarizes modifications to worldbuilding and the event graph, connecting the request with the resulting changes. An annotated screenshot of NarrativeSteward with four labeled regions. Artifact navigation (A) appears above the structural view (B) and selected-object editor (C), with agent dialogue (D) on the right. The event graph shows The Fork branching to End of the Meadow and Rest in the Woods. The editor displays The Fork's summary, character, and location references. The dialogue requests a more dangerous atmosphere in the location card and event while preserving two playable routes. Below the response, a change record lists one worldbuilding change and one event-graph change, with View changes, Keep this turn, and Undo this turn controls.
Figure 3. Linked narrative artifacts and how choices shape later events. (a) The event graph organizes the story’s events, while each event contains a beat graph of local narrative content and choices. Events and beats reference character and location cards from the worldbuilding; intent and outline provide creative context. These linked artifacts provide a shared basis for agent implementation and author review. (b) In The Fork , the player chooses woods and enters a beat that sets the state variable trail to woods , recording that choice. When the event ends, conditions on outgoing event edges check this value: the woods destination becomes available, while meadow remains locked. The player then selects woods and reaches its ending. The top panel links intent and outline to an event graph, character and location world cards, and a magnified beat graph within The Fork. Its meadow and woods event edges have matching conditions on trail. The bottom panel shows choosing woods, entering a beat that sets trail to woods, selecting the enabled woods event while meadow is locked, and reaching the woods ending. Purple and green arrows connect state writes and reads to the shared variable box.
Figure 4. Reviewing saved changes and guiding further revision. (a) The author asks the agent to revise an event named Crossing . After saving the revision, the system compares the previous and updated content to produce a changeset , a record of the turn’s actual changes. The example shows a change to the event’s summary. Keep retains the turn’s changes; Undo this turn reverses them together. (b) Locate selects the affected event in the graph and opens its summary field. The author can inspect the event and its connections, then delegate further changes or edit directly. Two panels show reviewing an agent revision and continuing work on the affected event. Panel a compares the Crossing summary before and after a saved turn and connects both versions to a system-computed changeset. The selected difference has a Locate button, while Keep and Undo this turn apply to the whole turn. Panel b locates Crossing in an event graph and opens its Summary field. The author can choose Edit or Delegate as optional refinement, leading to an illustrative revised summary, At dusk, one shore must wait.
Figure 5. Execution verification across narrative choices. (a) Choices within the event and beat graphs determine whether the player reaches the Checkpoint event, marked q , with a visitor pass, a key, or neither. A history is one feasible execution from Arrival to q . (b) Four histories produce three states: two histories yield the same visitor-pass state. Tuples give ( pass_type , has_key ). (c) Each gate column represents an outgoing edge of the checkpoint event; its condition is tested against every reachable state. A checkmark indicates an available exit, and a dash an unavailable one. With neither a pass nor a key, no exit is available, creating a dead end. The Staff gate is never enabled because no history supplies a staff pass. Three panels show how choices in event and beat graphs produce reachable states and checkpoint diagnostics. Panel a shows Arrival branching through Front Desk or Side Route before reconverging at Checkpoint q. Purple dashed connectors expand the two events into partial beat graphs, with ellipses for omitted beats. At Front Desk, asking directly or reading a notice leads to receiving a visitor pass. On Side Route, taking the key sets has_key to true, while leaving it changes no state. Panel b maps Histories 1 and 2 to (visitor, false), History 3 to (none, true), and History 4 to (none, false). Panel c tests each state against Visitor, Side, and Staff gate conditions. The visitor-pass state enables only Visitor, the key-only state enables only Side, and the state with neither enables no gate and is marked as a dead end. The Staff gate is never enabled.
Figure 6. Delegating and refining an interactive narrative in NarrativeSteward. In The Last Delivery , a courier must choose between delivering medicine to a clinic and a letter to a harbor. (a) The generated event graph connects Departure and Crossing to the two endings. (b) After reviewing Crossing’s summary, the author requests revisions to the summary and opening narration to emphasize the personal cost while preserving the rules and choices. The screenshot shows the summary change within the changeset, including its before-and-after text. (c) Locate opens the same summary in the event editor, where the author chooses to refine its wording directly, emphasizing the shore left waiting. The event graph connects Departure to Crossing, then to Clinic Ending and Harbor Ending. A changeset excerpt compares the Crossing summary before and after the agent's revision. The event inspector shows the author's subsequent summary about helping one shore and leaving the other waiting; purple annotations connect Locate to Summary and mark the selected words as an author edit.
Figure 7. Diagnosing an unreachable branch in The Last Delivery . (a) The author runs execution verification, which identifies a transition in Crossing that requires cargo=none (empty-handed). Earlier choices have already set cargo to medicine or letter , so this transition can never be taken. (b) The author uses Locate to inspect the corresponding edge condition. To preserve the intended choice between helping the clinic and the harbor, the author subsequently removes the unreachable beat and its incoming transition. A failed-verification excerpt lists the required none value and the actually reachable medicine and letter values. Beside it, the selected player-choice edge inspector shows the condition Cargo equals none and stable ID cross-e04.
Figure 8. Verifying and playtesting The Last Delivery after repair. (a) Execution verification passes after the unreachable branch is removed. (b) During playtesting, the author chooses to carry medicine at Departure. At Crossing, the clinic destination is available, while the harbor remains locked because it requires the letter. Cite for Agent attaches the selected option and current playtest context to the dialogue, where the author drafts a question about the locked destination. (c) Continuing the same playtest reaches Clinic Ending. A passed-verification summary reports complete coverage and zero main issues. A playtest excerpt offers the clinic and disables the harbor because cargo is medicine rather than letter; beside it, the agent dialogue panel displays attached Crossing context and the cited harbor choice while a question is drafted. A final excerpt shows the Clinic Ending reached by that route.
Check
Test unit
Result
Change detection
Expected change details
11/11 detected
Change precision
Reported change details
11/11 correct
Change localization
Locatable change details
11/11 matched
Failure / stop recovery
Failure or stop cases
4/4 restored
Whole-turn undo
Cases without conflicts
3/3 restored
Undo conflict protection
Cases with conflicts
3/3 protected
Table 1. Controlled evaluation of NarrativeSteward’s change records, recovery, and execution verification. Fractions indicate checks that matched predefined expectations out of the tested items or runs. The first three rows assess 11 individual change details across 4 editing cases. Recovery and undo checks assess restoration of prior work and protection of subsequent edits. Fault tests use paired faulty and corrected projects. Report invalidation and retention assess whether previous verification reports remain applicable after changes to checked content or to descriptions and image paths, respectively. Completion counts runs that finished with either a pass or a fault diagnosis.
Figure 9. Self-reported authoring experience with NarrativeSteward and participants’ selected general-purpose agents, rated on seven-point scales. Panels cover (a) delegation and revision, (b) understanding and refinement, and (c) feedback and decisions. Higher scores indicate more favorable experiences. Each item includes the same eligible participants in both conditions; n gives the number of pairs. Boxes show the interquartile range and median, whiskers extend to observations within 1.5 interquartile ranges of the box, and dots show individual ratings with horizontal offsets to separate responses. “Underst.” abbreviates understanding; “Keep dec. ease” and “Check dec. ease” refer to deciding whether to retain a modification or check further. Three panels compare thirteen experience measures for NarrativeSteward and general-purpose agents using vertical boxplots and individual scores. RQ1 spans the top row, with RQ2 and RQ3 below. Numeric-pair denominators vary with reported experience.
NarrativeSteward
Agent
RQ
Measure
Paired n
M
SD
M
SD
rrb
Adj. p
RQ1
Formulating revisions
12
6.08
0.90
3.42
1.38
1.00
0.003
Expressing requests
12
5.75
1.06
4.00
1.04
1.00
0.008
Inspecting changes
12
6.33
0.78
3.00
1.28
1.00
0.002
Revision request fit
12
5.00
0.95
4.50
0.80
0.52
0.375
RQ2
Change understanding
12
6.08
1.16
4.00
1.95
0.91
0.029
Table 2. Paired comparisons of self-reported authoring experience on seven-point scales. Higher scores indicate more positive ratings on each measure. NS denotes NarrativeSteward; Agent denotes participants’ selected general-purpose agents. Paired n counts eligible responses from the same participants in both conditions; M and SD give their mean and sample standard deviation. The paired rank-biserial effect size rrb is positive when NS ratings are higher. Adjusted p -values are from two-sided Wilcoxon signed-rank tests with Holm–Bonferroni correction within RQ1 (four measures), RQ2 (five), and RQ3 (four).
Category
Response
NS
Agent
Revision requests
Requested a revision
12
12
Considered revisions but made no request
0
0
Unfinished revisions
Had unfinished revisions
2
5
Had no unfinished revisions
10
7
Reasons for stopping
Did not want to invest further
1
2
Limited time / other activities
1
0
Table 3. Revision requests, unfinished revisions, and reasons for stopping. Counts refer to participants using NarrativeSteward (NS) or their selected general-purpose agent (Agent). The first two groups summarize whether participants requested revisions during the task and whether desired revisions remained unfinished at its end (12 participants per condition). Stopping reasons apply only to participants with unfinished revisions (NS: 2; Agent: 5). Each could select up to two reasons, so reason counts can overlap.
Figure 10. Helpfulness ratings for seven NarrativeSteward features, provided by participants who reported using each feature. Each dot represents one valid rating; n gives the number of ratings for that feature, and vertical offsets separate overlapping points. Ratings range from 1 (not at all helpful), through 4 (somewhat helpful), to 7 (extremely helpful). Seven rows show individual helpfulness ratings for the structural view, locating changes, inspecting content, direct editing, whole-turn undo, execution verification reports, and playtesting. Scores run from 1 to 7; each row displays its own valid rating count.
Authoring activity
NS
Agent
Both
Either
Neither
Unsure
Ideation
6
1
2
3
0
0
First playable version
6
3
1
1
0
1
Polishing
8
2
2
0
0
0
Checking / delivery
11
1
0
0
0
0
Table 4. Participants’ preferred tools for four interactive-narrative authoring activities. Each row counts one choice per participant ( N=12 ). NS denotes NarrativeSteward; Agent denotes participants’ selected general-purpose agents. Both means using the tools in combination, Either means either tool would suffice, Neither means neither would suit the activity, and Unsure indicates insufficient experience to choose.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Measure
Question
Anchors (1 / 4 / 7)
RQ1: Delegation experience and effort
Formulating revisions
When considering further revisions to your work, how easy was it to identify what specifically you wanted to change?
Very difficult / Neither difficult nor easy / Very easy
Expressing requests
When asking the agent to revise the work, how easy was it to express your requirements clearly?
Very difficult / Neither difficult nor easy / Very easy
Inspecting changes
After the agent completed a revision, how easy was it to determine what it had actually changed?
Very difficult / Neither difficult nor easy / Very easy
Revision request fit
To what extent did the agent’s result meet the revision requirements you had specified?
Not at all / Partially / Fully
RQ2: Understanding and continued refinement
Appendix
Table 5. Questionnaire items for the thirteen authoring-experience measures compared in Table 2 . Each question was asked for both tools. Anchors define scores of 1, 4, and 7 on seven-point scales.
Measure
Question
Response options
Revision experience
After obtaining the first version of the work, which of the following best describes your experience?
Select one: did not consider further revisions; considered revisions but did not request them; requested revisions from the agent; could not recall; had not completed the original task; had not obtained a first version.
Unfinished revisions
At the end of the task, were there revisions you wanted to make but did not continue to complete?
Select one: yes; no; could not judge; could not recall.
Reasons for unfinished revisions
What factors led you not to complete these revisions?
Select up to two: the current result already met task needs; limited time or other activities; insufficient understanding of the existing content or structure to identify revisions; understanding the work but lacking a specific revision idea; anticipated effort to explain requirements or communicate repeatedly; anticipated effort to inspect revised results; concern about affecting satisfactory content; unwillingness to invest further at that time; another reason; inability to recall.
Appendix
Table 6. Questions about revision experience and desired but unfinished revisions, asked for each tool. The reasons question allowed up to two selections from participants reporting unfinished revisions.
Feature
Purpose evaluated
Structural view
Understanding the story structure through event and beat graphs.
Locate from changes
Finding the content modified by the agent.
Content inspection
Judging whether content matched the author’s intent.
Direct editing
Making local revisions.
Whole-turn undo
Recovering from unsatisfactory agent revisions.
Execution verification report
Judging whether further checking or repair was needed.
Appendix
Table 7. NarrativeSteward features and the purposes assessed by the helpfulness questions. Participants first reported feature use; those reporting use rated helpfulness from 1 (not at all helpful) to 7 (extremely helpful), with 4 indicating somewhat helpful.
Measure
Question
Response options
Overall preference
If you undertook a similar interactive-narrative task again, which tool would you prefer to use?
Select one: definitely NarrativeSteward; probably NarrativeSteward; no clear preference; probably the selected general-purpose agent; definitely that agent; insufficient experience from the two tasks to judge.
Preference by authoring activity
Which approach would you choose for each of the following activities: developing a story direction; obtaining a first playable story; revising and polishing existing content; checking logic and preparing the final delivery?
Select one per activity: mainly NarrativeSteward; mainly the selected general-purpose agent; a combination of both; either tool; neither tool; insufficient experience to judge.
Appendix
Table 8. Questions about overall tool preference and preferences for individual authoring activities. Participants selected one response for overall preference and one for each activity.
Condition
Main story stages
Meaningful decision points
Reachable endings
NS
4.3 [4, 5]
22.4 [8, 49]
3.6 [3, 5]
Agent
4.0 [3, 5]
11.0 [5, 36]
2.9 [2, 3]
Appendix
Table 9. Scale of the 24 completed interactive narratives, with 12 stories per condition. NS denotes NarrativeSteward; Agent denotes participants’ selected general-purpose agents. Cells report mean [minimum, maximum] counts per story. Stages are major narrative phases; decision points are distinct choice locations that affect content, later actions, or endings; endings are distinct reachable narrative outcomes. Counts cover the whole story across its branches.
RQ
Measure
Paired n
NS: Mdn [Q1, Q3]
Agent: Mdn [Q1, Q3]
Paired difference
RQ1
Formulating revisions
12
6 [6, 7]
3.5 [3, 4.25]
2.5
RQ1
Expressing requests
12
6 [5, 6.25]
4 [4, 4.25]
1.5
RQ1
Inspecting changes
12
6.5 [6, 7]
3 [2, 4]
3.5
RQ1
Revision request fit
12
5 [4, 5.25]
4.5 [4, 5]
0
RQ2
Change understanding
12
6 [6, 7]
4 [2.75, 5.25]
1.5
RQ2
Edit stability
8
6 [6, 7]
4.5 [3.5, 5]
2
Appendix
Table 10. Paired rating distributions for the thirteen authoring-experience measures in Table 2 . NS denotes NarrativeSteward; Agent denotes the selected general-purpose agent. Both condition summaries use the same eligible pairs (n). Mdn is the median, and Q1 and Q3 are the first and third quartiles. Paired difference is the median of the within-participant differences, calculated as NS minus Agent on the seven-point scale.
Case
Verification time (ms)
Peak memory (MiB)
Execution states
Integrated reference
3.5 [3.3, 3.6]
36.8
136
Large scalar domain
59.4 [57.0, 60.5]
37.5
45,051
Many relevant Boolean variables
20.7 [20.2, 22.1]
37.3
16,374
Wide branch
5.9 [5.6, 6.7]
37.0
199
Distant fault
10.4 [10.1, 11.3]
37.9
261
Appendix
Table 11. Execution verification resources for one integrated reference story and four generated structures. Each case received one warmup and five measured runs in separate processes. Time is reported as median [minimum, maximum]; peak memory and explored execution states are medians. The distant-fault case completed with the expected diagnosis of a never-enabled edge.