NextMe-800: Anticipating Personal Behavior from Months of Egocentric Video
Organizations: The University of Hong Kong · Tongji University
Abstract
We often plan ambitiously yet act habitually and wonder, in retrospect, whether we would have planned differently had we known what we would actually do. Hindsight offers a valuable perspective on past decisions, although we often wish we could have simulated hindsight at the moment of choosing. If a system could generate plausible trajectories from one's personal history, such previews might help people formulate more realistic plans and make better informed decisions. We introduce NextMe-800, an approximately 800-hour first-person dataset from one volunteer over 126 days with 1 Hz images, gaze, and audio, captioned at five hierarchical abstraction levels from atomic actions to major activities. We formulate personalized action anticipation as open-vocabulary K-step sequence prediction and construct NextAct, a 1,500-point benchmark combining NextMe-800 with the multi-person EgoLife dataset. Using an embedding-based soft edit distance as the metric, we evaluate how well different models can anticipate personal behavior across abstraction levels and prediction horizons. NextMe-800 and NextAct provide a months-long resource and evaluation framework for studying how far ahead personal behavior can be anticipated from egocentric observation.
Figures & tables
| Mean | Ratio | |||
|---|---|---|---|---|
| Level | Events | dur. | Adj. | Cum. |
| L1 | 529,431 | 5 s | — | — |
| L2 | 48,093 | 1 min | 11 | 11 |
| L3 | 6,367 | 7.5 min | 7.6 | 83 |
| L4 | 1,988 | 24 min | 3.2 | 266 |
| L5 | 1,236 | 39 min | 1.6 | 428 |
| Dataset | Avg. rec. | Avg. | Awake | People | Total | Notes |
|---|---|---|---|---|---|---|
| time (h) | span | cov. (%) | time (h) | |||
| Ego4D ( Grauman et al., 2022 ) | 3.94 | – | – | 931 | 3,670 | POV video |
| EgoMonth ( Chen et al., 2026 ) | 15.06 | 66.5 d | 1.42 | 20 | 301.2 | POV video |
| EgoLife ( Yang et al., 2025 ) | 44.33 | 7 d | 39.58 | 6 | 266 | POV video |
| KrishnaCam ( Singh et al., 2016 ) | 70.20 | 9 mo | 1.63 | 1 | 70.2 | Outdoor POV |
| LongNAP ( Shaikh et al., 2026 ) | 91.85 | 28 d | 20.50 | 20 | 1,837 | Computer screen rec. |
| Predictor | L1 | L2 | L3 | L4 | L5 | Embedding score | Reranker score | |
|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro (preview) | 0.3402 | 0.3202 | 0.3348 | 0.3634 | 0.3567 | 0.3431 | 0.2030 |
| 1 | Codex + DeepSeek † | 0.3225 | 0.3252 | 0.3641 | 0.3509 | 0.3472 | 0.3419 | 0.1596 |
| 1 | GPT-5.6-sol | 0.3042 | 0.3144 | 0.3571 | 0.3599 | 0.3556 | 0.3382 | 0.1826 |
| 1 | DeepSeek-v4-flash | 0.3154 | 0.3114 | 0.3383 | 0.3564 | 0.3431 | 0.3329 | 0.1695 |
| 1 | Claude Opus 4.6 | 0.2851 | 0.3061 | 0.3329 | 0.3480 | 0.3401 | 0.3224 | 0.1509 |
| 1 | Gemma4-12B-it | 0.2963 | 0.2930 | 0.3284 | 0.3262 | 0.3096 | 0.3107 | – |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Value | Note |
|---|---|---|
| Source segments with any flag | 96 (15.66%) | 613 source segments |
| Equal-segment mean flag rate | 8.54% | Average within item, then across items |
| Post-hoc source segments with any flag | 72 (11.75%) | After removing one rater |
| Category | Segments | Fraction |
|---|---|---|
| Action/verb | 37 | 6.04% |
| Entity or text | 33 | 5.38% |
| Omitted event | 29 | 4.73% |
| Temporal displacement | 11 | 1.79% |
| Sequence-order reversal | 0 | 0.00% |
| L5 (2 total) | L2 (178 total) | L1 (2,287 total) |
|---|---|---|
| 14:23–17:54 | … (earlier tasks) … | |
| 3 h 31 min | 16:26–16:29 | 16:27:42–16:27:57 (15 s) |
| I work on RL | I compute target | I edit target_q = … |
| homework | Q-values | … |
| 16:29:33–16:29:36 (3 s) I scratch my nose | ||
| 16:29:43–16:29:44 (1 s) I glance at my phone | ||
| L1 | L2 | L3 | L4 | L5 | |
|---|---|---|---|---|---|
| 0.3997 | 0.3974 | 0.4010 | 0.3995 | 0.4310 | |
| 0.4173 | 0.3964 | 0.3916 | 0.4079 | 0.4386 |
| Method | 95% CI | Endpoint % | Near-tie % | Min. gap | |
|---|---|---|---|---|---|
| Embedding-8B (norm.) | 0.9724 | [0.9644, 0.9768] | 3.6 | 10.5 | 0.0590 |
| Claude Haiku 4.5 | 0.9612 | [0.9493, 0.9691] | 30.1 | 10.8 | 0.0606 |
| Gemini-3.7-flash | 0.9407 | [0.9282, 0.9487] | 34.3 | 10.7 | 0.0548 |
| GPT-5.6-luna | 0.9113 | [0.8969, 0.9219] | 31.6 | 11.0 | 0.0559 |
| Qwen3.5-flash | 0.8852 | [0.8704, 0.8961] | 37.5 | 13.1 | 0.0492 |
| ROUGE-L (character F1) | 0.8037 | [0.7866, 0.8179] | 12.4 | 19.8 | 0.0317 |
| Level | Ground truth | Prediction | Score |
| High-scoring | |||
| L1 | I read the Feishu document notes. | I read the Feishu reflection notes. | 0.92 |
| L3 | I disembark the airplane. | I disembark the airplane and navigate through the arrival terminal. | 0.79 |
| L5 | I perform a gym workout session with my companions. | I engage in a strength training session at the gym with my companions. | 0.86 |
| Low-scoring | |||
| L1 | I walk away from the table. | I write mathematical formulas on the notepad. | 0.00 |
| # | Ground truth | Prediction |
|---|---|---|
| High-scoring: L3, score 0.49 | ||
| 1 | I prepare my belongings in my apartment before departing. | I pack my laptop and study materials into my backpack at my apartment. |
| 2 | I travel from my apartment building to the university main library. | I travel from my apartment to the university library. |
| 3 | I study my math homework involving bipartite graphs at a library table. | I set up at a library desk and study diffusion-model lecture notes on my laptop. |
| 4 | I perform coursework and research-agent activities at the library desk. | I use ChatGPT and VS Code to work through programming tasks for my course assignment. |
| 5 | I manage a Zoom meeting and software configuration tasks from a hallway bench. | I take a break to browse WeChat and respond to messages. |