MotiveMob: Motivation as Semantic Action for Closed-Loop Human Mobility Generation
Organizations: The University of Osaka · The University of Tokyo · The Hong Kong University of Science and Technology · Institute of Science Tokyo
Abstract
Human mobility generation, an important task in urban research, synthesizes trajectory data for urban planning and transportation management. Human mobility can be characterized as a "why-where-when" decision process: people form an intention to move and then determine where and when the corresponding activity will take place. Trajectory generation under user-level and temporal distribution shifts may benefit from explicitly modeling this decision structure. However, many existing human mobility generation methods either represent behavioral intent at a coarse granularity, such as a daily plan or a trajectory-level description, or directly predict future locations without explicitly reasoning about a possible motivation for each movement step. We introduce MotiveMob, a motivation-driven autoregressive framework for human mobility generation that first forms a hypothesis about why the next movement may occur and then jointly generates where and when it may occur. At each step, a motivation predictor conditions on the current mobility state, a long-term behavioral report, and the mobility history to infer a plausible motivation or determine whether the trajectory should terminate. Given the hypothesized motivation, a state predictor grounds it in a candidate next location and arrival time. The candidate then undergoes speed-feasibility and repetition checks before being fed back for the next decision. We evaluate MotiveMob under distribution shifts involving unseen users and unseen temporal periods, including seasonal changes and the substantial behavioral disruption caused by the COVID-19 pandemic. Experiments show that MotiveMob consistently achieves better distributional fidelity than competitive pretraining-based and prompting-based methods under user-level and temporal distribution shifts, demonstrating robust generalization to out-of-distribution mobility patterns.
Figures & tables
| State predictor | Seen period | Regular unseen period | COVID-19 unseen period | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SD | SI | DARD | STVD | SD | SI | DARD | STVD | SD | SI | DARD | STVD | |
| Seen users | ||||||||||||
| MotiveMob (IWM ) | 0.0241 | 0.0159 | 0.0857 | 0.3841 | 0.0307 | 0.0212 | 0.1054 | 0.4558 | 0.0880 | 0.0309 | 0.1632 | 0.5745 |
| w/ positive-only | 0.0328 | 0.0115 | 0.0707 | 0.3870 | 0.0376 | 0.0185 | 0.0943 | 0.4469 | 0.0914 | 0.0390 | 0.1586 | 0.5602 |
| w/ rule-based | 0.0374 | 0.0101 | 0.0992 | 0.4286 | 0.0405 | 0.0165 | 0.1111 | 0.4932 | 0.0687 | 0.0198 | 0.1576 | 0.6230 |
| Unseen users | ||||||||||||
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Motivation | Behavioral interpretation | Destination venue categories |
|---|---|---|
| Return | Returning to a residence or temporary accommodation | Residential venues, hotels, hostels, inns, resorts, and other lodging venues |
| Personal Service | Obtaining a personal, household, financial, repair, or travel-related service | Salons, banks, repair shops, rental services, post offices, travel agencies, and related service venues |
| Entertainment | Participating in leisure, cultural, or recreational activities | Arts and entertainment, landmarks and outdoors, sports and recreation, cruises, auditoriums, and related venues |
| Dining | Eating or drinking | Dining and drinking venues, food courts, and cafeterias |
| Transit | Traveling or transferring between places | Airports, stations, public transport facilities, roads, and other transportation venues |
| Religious | Participating in religious or spiritual activities | Shrines, temples, churches, mosques, monasteries, and other religious venues |
| Parameter | Setting |
|---|---|
| Construction data | 25 seen users from January–October 2019, containing 3,559 daily trajectories and 31,063 observed transition instances |
| Alternative motivations | Three distinct motivations per observed transition, sampled without replacement while excluding the observed motivation |
| Prepared alternative cases | 93,189 cases before candidate generation and judging; 31,063 judged alternative transitions are used for balanced state-predictor training |
| Candidate-generator architecture | Three-layer pre-layer-normalized Transformer encoder with hidden dimension 128, 4 attention heads, feed-forward dimension 256, and dropout 0.1 |
| Prediction heads | Separate classification heads for location subcategory, spatial grid, and arrival-time bin |
| Candidate-generator objective | Equal-weighted cross-entropy losses for the subcategory, spatial-grid, and time-bin prediction heads |
| Parameter | State predictor | Motivation predictor |
|---|---|---|
| Initialization | Llama-3.1-8B-Instruct | Independently initialized from Llama-3.1-8B-Instruct |
| Training data | 31,063 observed transitions and 31,063 judged alternative transitions | 34,622 observed instances, comprising 31,063 step-level motivations and 3,559 END labels |
| Observed-to-alternative ratio | Not applicable | |
| Training objective | Assistant-token-only causal language-modeling loss | Assistant-token-only causal language-modeling loss |
| Quantization | 4-bit NF4 with double quantization and bfloat16 computation | 4-bit NF4 with double quantization and bfloat16 computation |
| QLoRA configuration | Rank , scaling factor , dropout 0, and no bias | Rank , scaling factor , dropout 0, and no bias |
| Parameter | Setting |
|---|---|
| Initial state | A START state; the first motivation and check-in are generated without applying a travel-speed constraint because no previous spatial state is available |
| Mobility-history context | Up to 30 check-ins from the three active days preceding the target day |
| Motivation decoding | Greedy decoding with at most 8 newly generated tokens |
| State decoding | Sampling with temperature 0.2, top- 0.9, and at most 32 newly generated tokens |
| State-generation attempts | At most 8 proposals per step, consisting of one initial proposal and up to seven retries |
| Retry decoding | Temperature increases from 0.3 to 0.9 in increments of 0.1; top- remains 0.9 and the maximum output length remains 32 tokens |
| Protocol component | Data usage and restriction |
|---|---|
| Predictor training | Trajectories from 25 seen users between January and October 2019. For the seen-user/seen-period evaluation, 20% of user–date trajectories are held out and excluded from predictor training. |
| User-level shift | The 50 unseen users are disjoint from the 25 users used to train and . |
| Temporal shift | The regular unseen period (November 2019–February 2020) and the COVID-19 period (April–October 2020) do not overlap with the predictor-training period. |
| Recent history | Contains at most 30 check-ins from the three active days strictly preceding target date . Target-day and future observations are never included. |
| Behavioral report | In the temporal-shift settings, the report is constructed only from observations preceding the evaluation period. In the seen-period settings, it is a static aggregate profile constructed from the query-disjoint support subset and may include support dates later than the target date. |
| COVID-19 update | For targets in April–October 2020, the report is updated using observations available through March 2020, without using any check-ins from the target period. |
| Method | Seen period (1186 trajectories) | Regular unseen period (466 trajectories) | COVID-19 unseen period (689 trajectories) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SD | SI | DARD | STVD | SD | SI | DARD | STVD | SD | SI | DARD | STVD | |
| LSTM | 0.0914 | 0.1375 | 0.2673 | 0.4161 | 0.1109 | 0.1698 | 0.3230 | 0.5779 | 0.1259 | 0.2073 | 0.3556 | 0.5725 |
| MHSA | 0.0908 | 0.1357 | 0.2468 | 0.4286 | 0.1210 | 0.1817 | 0.2706 | 0.5803 | 0.1262 | 0.2152 | 0.3044 | 0.5552 |
| DeepMove | 0.1777 | 0.1024 | 0.2348 | 0.4933 | 0.1645 | 0.1267 | 0.3180 | 0.6065 | 0.2308 | 0.1595 | 0.3340 | 0.6029 |
| GetNext | 0.0395 | 0.2820 | 0.2149 | 0.4635 | 0.0433 | 0.2834 | 0.2057 | 0.5601 | 0.0455 | 0.2952 | 0.1953 | 0.5294 |
| TrajGAIL | 0.2251 | 0.1256 | 0.0799 | 0.5836 | 0.2458 | 0.1227 | 0.1250 | 0.6576 | 0.2494 | 0.1118 | 0.1137 | 0.6424 |
| Method | Seen period (2804 trajectories) | Regular unseen period (1655 trajectories) | COVID-19 unseen period (1301 trajectories) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SD | SI | DARD | STVD | SD | SI | DARD | STVD | SD | SI | DARD | STVD | |
| LSTM | 0.1489 | 0.2437 | 0.3872 | 0.6207 | 0.1740 | 0.2420 | 0.3810 | 0.6569 | 0.1493 | 0.2645 | 0.4358 | 0.6551 |
| MHSA | 0.1897 | 0.1965 | 0.3578 | 0.5961 | 0.1715 | 0.2352 | 0.3762 | 0.6466 | 0.1582 | 0.2500 | 0.4182 | 0.6589 |
| DeepMove | 0.2754 | 0.2267 | 0.3548 | 0.6510 | 0.2746 | 0.2257 | 0.3876 | 0.6714 | 0.2709 | 0.2338 | 0.3918 | 0.6704 |
| GetNext | 0.0832 | 0.3220 | 0.2568 | 0.5752 | 0.0752 | 0.3215 | 0.2376 | 0.6408 | 0.0994 | 0.3351 | 0.2511 | 0.6438 |
| TrajGAIL | 0.1737 | 0.0825 | 0.1287 | 0.5750 | 0.1954 | 0.0912 | 0.1578 | 0.6513 | 0.1915 | 0.0943 | 0.1427 | 0.6532 |
| Backbone | Seen period | Regular unseen period | COVID-19 unseen period | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SD | SI | DARD | STVD | SD | SI | DARD | STVD | SD | SI | DARD | STVD | |
| Seen users | ||||||||||||
| Qwen3.5-9B | 0.0080 | 0.0088 | 0.0671 | 0.3552 | 0.0096 | 0.0116 | 0.0902 | 0.4464 | 0.0452 | 0.0173 | 0.1396 | 0.5728 |
| Gemma-4-E4B | 0.0081 | 0.0191 | 0.0828 | 0.4041 | 0.0114 | 0.0192 | 0.0953 | 0.4731 | 0.0425 | 0.0200 | 0.1300 | 0.5431 |
| Gemma-2-9B | 0.0099 | 0.0064 | 0.0641 | 0.3760 | 0.0097 | 0.0118 | 0.0889 | 0.4296 | 0.0482 | 0.0138 | 0.1319 | 0.5696 |
| Unseen users | ||||||||||||
| State generation | END predictor | Mean length | Max-step rate | SI | DARD |
|---|---|---|---|---|---|
| Direct, w/o (IWM) | Direct predictor | 7.52 | 1.48% | 0.0282 | 0.1401 |
| (IWM) | 8.08 | 2.27% | 0.0205 | 0.1087 | |
| (positive-only) | 7.42 | 2.06% | 0.0176 | 0.1132 | |
| Rule-based state predictor | 9.06 | 1.76% | 0.0215 | 0.1163 | |
| Ground truth | – | 6.82 | – | – | – |
| Initialization | Epoch | SD | SI | DARD | STVD | Mean rank |
|---|---|---|---|---|---|---|
| IL | 4 | 0.0200 | 0.0372 | 0.1550 | 0.4771 | 5.42 |
| IL | 8 | 0.0204 | 0.0208 | 0.1200 | 0.4898 | 5.04 |
| IL | 12 | 0.0203 | 0.0198 | 0.1043 | 0.4857 | 3.81 |
| IWM, | 4 | 0.0209 | 0.0378 | 0.1616 | 0.4742 | 5.42 |
| IWM, | 8 | 0.0213 | 0.0192 | 0.1134 | 0.4949 | 4.92 |
| IWM, | 12 | 0.0204 | 0.0207 | 0.1109 | 0.4969 | 5.02 |
| User | Period | Days | States | Resampling | Resampling | Correction | Feasibility | No feasible | Repetition |
|---|---|---|---|---|---|---|---|---|---|
| status | required | resolved | invoked | corrected | substitute | terminated | |||
| Seen | Seen | 908 | 7,647 | 932 (12.2%) | 528 (6.9%) | 404 (5.3%) | 231 (3.0%) | 173 (2.3%) | 121 (13.3%) |
| Seen | Regular unseen | 690 | 5,446 | 686 (12.6%) | 415 (7.6%) | 271 (5.0%) | 134 (2.5%) | 137 (2.5%) | 88 (12.8%) |
| Seen | COVID-19 unseen | 392 | 3,218 | 830 (25.8%) | 459 (14.3%) | 371 (11.5%) | 237 (7.4%) | 134 (4.2%) | 77 (19.6%) |
| Unseen | Seen | 1,345 | 10,066 | 2,582 (25.7%) | 1,534 (15.2%) | 1,048 (10.4%) | 629 (6.2%) | 419 (4.2%) | 207 (15.4%) |
| Unseen | Regular unseen | 795 | 6,016 | 1,610 (26.8%) | 929 (15.4%) | 681 (11.3%) | 432 (7.2%) | 249 (4.1%) | 150 (18.9%) |