Ream: Unfolding Mutual Awareness in Human-Agent Workspaces
Authors: Peiling Jiang, Sangho Suh, Varsha Kishore, Jonathan Bragg, Haijun Xia, Pao Siangliulue, Daniel S. Weld, Amy X. Zhang, +1 more
Organizations: University of California San Diego La Jolla, California, USA · Allen Institute for AI (Ai2) · Allen Institute for AI Seattle, WA, USA · University of Washington Seattle, Washington, USA
As AI agents work alongside humans in shared workspaces, a mutual awareness challenge arises: agents act at speeds that outpace human monitoring, and users' evolving interests are not always expressed in chat. This challenge is especially pressing in literature review, where both parties retrieve, read, and synthesize a growing body of papers. We present Ream, a literature review workspace that supports mutual awareness through structured artifacts, bidirectional engagement tracking, and localized visualizations. Users can see each party's activity within these documents, and agents can retrieve the same history to guide their work. In studies with eighteen researchers, participants used these traces to inspect evidence, steer agents, communicate through annotations, and reflect on their research focus. Shared histories also helped agents build on earlier work. These findings inform how engagement traces within shared documents can support transparency, personalized assistance, and coordination in human-agent knowledge work.
Figures & tables
Figure 1. Ream is a human-agent collaborative literature review workspace, tracking and visualizing both the user’s and agents’ actions to facilitate mutual awareness. The file tree (a) shows different amounts of activity on each file, with blue indicating user activity and orange indicating the agents’. The tabbed file viewer (b) shows a .report file with paper citation chips. The right chat panel (c) allows the user to create and instruct multiple agents across workspace folders at the same time. The Ream workspace has three panels: a file tree with blue user activity bars and orange agent activity bars, a tabbed report viewer with interactive citations and a paper preview card, and a chat panel for coordinating agents.
Element
Awareness question
Ream mechanisms
Who
Which party engaged with this artifact?
User/agent attribution in bars and histories
What
What work has occurred here?
Search, read, write, chat, and organize events
Where
Which sources and passages received attention?
Traces on files, references, and PDF passages
When
When was this artifact visited or changed?
Timestamped per-file histories
How
How was a source used to address a question?
Exploration questions, answers, and extracted quotes
Why
What goals or preferences help explain an action?
Notes, stars/downvotes, and agent questions
Table 1. Workspace-awareness elements adapted to human-agent artifacts, with the questions each mechanism helps collaborators address.
Figure 2. The .paper file view includes: (1) controls allowing users to Star or Downvote a paper to indicate their interests; (2) paper metadata and figures; (3) user notes (here, the user took notes about additional papers found while reading, which are in turn used by the agent to generate proactive suggestions); and (4) the Agent Explorations section with questions, answers, and quotes extracted from the paper (5). A paper document with Star and Downvote controls, paper metadata and figures, user notes, and an Agent Explorations section containing questions, answers, and evidence quotes from the paper.
Figure 3. When the chat input is focused, Ream proactively generates task suggestions based on user activities. Four task suggestions above the chat input propose reviewing a synthesis for unsupported claims, comparing interaction patterns across papers, adding a research question, and checking whether another paper adds a distinct form of user control.
Figure 4. Left: at the bottom of each file, Ream lists activities from the user and agents. Here, the agent first fetched this .paper file by searching. The user then mentioned it in chat, first asking a detailed question, then asking the agent to compare it with its citations. Right: the PDF view highlights passages matched to quotes returned by the agent’s paper-reading tool, connecting its questions and answers to source evidence for inspection. Three panels show a file's chronological signal history, the user's follow-up questions in chat, and passages in the PDF highlighted to match evidence quotes returned by the agent's paper-reading tool.
Figure 5. Two literature review workflows in Ream . P7 first worked with agents to gather papers and develop syntheses, then focused on reading. P9 interleaved searching, synthesis, organization, and reading throughout the session. Two stacked activity plots compare sessions P7 and P9, with user activity in blue and agent activity in orange. P7 concentrates agent work earlier and user activity later, while P9 interleaves user and agent activity.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Actor
Action
Strength
User
Hover over paper card
0.02
Open file
0.05
Active file focus (30 s)
0.80
Highlight PDF passage
0.50
Add / edit highlight note
0.30
Mention in chat
0.50
Appendix
Table 2. Action signal weights.
Figure 6. Activity traces for users and agents, and for different action types in user study sessions. Twelve panels show activity traces for participants P1 through P12. Each panel compares user and agent activity over the session and separates recorded actions into chat, organize, read, search, and write categories.
Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation imposes two structural costs: a Storytelling Tax, where failed experiments, rejected hypotheses, and the branching exploration process are discarded to fit a linear narrative; and an Engineering Tax, where the gap between reviewer-sufficient prose and agent-sufficient specification leaves critical implementation details unwritten. Tolerable for human readers, these costs become critical when AI agents must understand, reproduce, and extend published work. We introduce the Agent-Native Research Artifact (ARA), a protocol that replaces the narrative paper with a machine-executable research package structured around four layers: scientific logic, executable code with full specifications, an exploration graph that preserves the failures compilation discards, and evidence grounding every claim in raw outputs. Three mechanisms support the ecosystem: a Live Research Manager that captures decisions and dead ends during ordinary development; an ARA Compiler that translates legacy PDFs and repos into ARAs; and an ARA-native review system that automates objective checks so human reviewers can focus on significance, novelty, and taste. On PaperBench and RE-Bench, ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. On RE-Bench's five open-ended extension tasks, preserved failure traces in ARA accelerate progress, but can also constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities. Our code is open-sourced at https://github.com/Orchestra-Research/Agent-Native-Research-Artifact.
Jiachen Liu, Jiaxin Pei, Jintao Huang +34
University of Michigan · Stanford University · Ohio State University +22
In open-ended problem solving, collaborators often rely on discussion to surface concerns, challenge perspectives, and refine shared work as it evolves. While AI agents are increasingly used as discussion partners, existing multi-agent systems place a heavy burden on users to initiate and carefully orchestrate the discussions. We present DocuTeam, a mixed-initiative multi-agent discussion system in which both users and agents can initiate and steer conversations. Agents monitor document changes to proactively start and redirect discussions as the work evolves, while users can flexibly shape the conversation or adopt agent ideas. In a within-subjects study (N=20), participants using DocuTeam produced outcomes rated significantly more novel, relevant, and specific than with a baseline without any increase in cognitive load. Rather than using agents for one-off idea sourcing, participants engaged in an iterative refinement loop in which document changes prompted agent reactions, which led users to revisit and further develop their work.
Heechan Lee, Juhyeon Choi, Tae Soo Kim +2
School of Computing, KAIST, Daejeon, Republic of Korea · College of Liberal Studies, Seoul National University, Seoul, Republic of Korea · SkillBench, Santa Barbara, CA, USA
As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-and-development process behind this model as a case study of human--AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback. These observations indicate a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. Progress toward more autonomous AI research must therefore advance both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development.