Workplace sensing studies combine long-running behaviour traces with self-reports, yet the tools that collect those data often sit apart from the interface that returns results. We present TrustmeWatcher, the application built for the TRUST-ME project to connect this work. TrustmeWatcher reuses ActivityWatch's OS-level watchers for computer-activity collection and adds its own application layer. It turns the collected traces into an interactive screen-time dashboard, synchronizes responses from short questionnaires completed on the StreamDeck, and presents questionnaires alongside video highlights. Activity records and self-reports are aligned into labelled records for model development. The scope of this paper is limited to ActivityWatch data as model input. Artificial intelligence (AI) uses these activity records to predict six normalized state scores and an overall well-being score. The trained model runs locally, and the dashboard presents its predictions in semantic bands. Explainable artificial intelligence (XAI) helps users understand how recorded activity contributed to a prediction. Privacy Control lets users pause or resume the camera and eye tracker used by the study. We describe the workflow, its user-device and sensing-setup boundaries, and its use with records from 17 participants. The result is a deployed application and study workflow that integrates activity review, study data collection, privacy control, local prediction, and a participant-facing interface for XAI evaluation.
Figures & tables
Figure 1. TrustmeWatcher system overview and data flow. The diagram separates the participant's computer, the sensing setup, and an external model development pipeline. The sensing setup sends check-ins and video highlights to the participant's computer, while questionnaire responses are stored on the study server. Privacy commands from the participant control the camera and eye tracker. Study data are used to train the model, which is then deployed to the participant's computer for local prediction.
Figure 2. Architecture view. Three local software layers are stacked vertically. Reused ActivityWatch OS watchers form the bottom layer. The TrustmeWatcher local backend and dashboard form the two layers above it.
Figure 3. Screen-time summary using synthetic data. The complete screen-time summary showing a yearly timeline, category distribution, top categories, and top applications.
Figure 4. Daily feedback and sensing control using synthetic data. The user dashboard Home view showing daily semantic feedback cards beside direct control of the study sensors.
Figure 5. Counterfactual example. A counterfactual dialog for work productivity. It shows a current score of 3.14, a selected target of 4, a slider, and one model scenario that reallocates five minutes from communication to personal distraction, producing a model estimate of 4.28. A short disclaimer states that model predictions may be inaccurate and do not guarantee an outcome.
Wearable sensors capture fine-grained motion patterns that support rich behavioral understanding, yet most existing methods reduce these signals to activity labels. Recent LM-based approaches generate natural-language explanations for sensor data, but their reasoning is weakly grounded in the underlying signal, leading to fluent yet unverifiable explanations. We introduce TRACE-TS (Traceable Reasoning with Attribution-Grounded Evidence), a framework for structured and signal-grounded reasoning over wearable time series. TRACE-TS uses attribution from an expert classifier to identify salient spatio-temporal sensor regions, uses them to construct DAG reasoning traces with explicit evidence provenance, and trains a compact language model to generate these traces through gated cross-attention over sensor memory tokens. At inference, the adapted model jointly outputs the activity prediction and its reasoning trace, without requiring attribution computation or teacher guidance. We introduce Semantic Node Match(SNM), an LLM-as-judge metric that diagnoses reasoning fidelity at the observation, inference, and synthesis levels, localizing hallucinated observations and broken evidence chains missed by standard NLG metrics. Across seven wearable benchmarks, TRACE-TS achieves the best average accuracy and F1 among all evaluated methods (84.43%/81.24%), and outperforms the best LLM-based baseline by 17.96% in F1. Our code is available at https://github.com/SparshRastogi/TRACE-TS.
Sparsh Rastogi, Tanmay Kumar, Baiyu Chen +3
University of New South Wales, Sydney, Australia · Thapar Institute of Engineering and Technology, Patiala, India
Audio carries rich cues about human activities, and microphones are already built into most wearable devices. However, microphones also capture speech, and this privacy risk limits their use in Human Activity Recognition (HAR). We present AnomaSense, a sensor activation approach for wrist wearables that keeps the microphone off by default and turns it on for at most one second when an unsupervised anomaly detector flags an IMU segment that is likely to produce sound. The captured audio is further masked before it reaches the recognition model. We study 20 activities from 15 participants, organized into five groups in which activities share similar wrist motion but differ in the object or material involved. With IMU data alone, our recognition model reaches 78.98% accuracy in leave-one-participant-out validation. With the short, masked audio windows added, accuracy reaches 96.89% with no masking and stays above 86% when 90% of each one-second audio window is removed. On the same data, the anomaly detector triggers the microphone with 86.46% precision and 74.28% recall relative to sound events. We also report a small preliminary check of automatic speech recognition on masked speech, which shows that contiguous masking degrades recognition far more than point-wise masking at the same masking ratio. Our evaluation is a controlled, offline feasibility study. We describe the threat model, what the approach does and does not protect, and the steps needed before deployment.
Xue Wang, Yang Zhang
University of California, Los Angeles Los Angeles, CA, USA
Mobile sensing enables longitudinal monitoring of behavioral and physiological patterns in everyday settings. However, accurate prediction remains challenging in small-cohort health-sensing studies, where task-specific outcome supervision is limited relative to heterogeneous sensing data. Interpretability is also important, as model outputs should reflect meaningful behavioral and physiological patterns rather than predictive scores alone. We develop a Concept-Integrated Transformer (CIT) with LLM-guided concept supervision for explainable prediction from mobile sensing data. CIT uses a pretrained large language model to generate baseline-aware concept abnormality targets with confidence weights without manual concept annotation. Across two longitudinal datasets, CIT achieves the highest F1 score on AFFECT (0.756) and ties for the highest on a PHQ-9 dataset (0.765). The learned concept scores also reveal interpretable behavioral and physiological patterns; in AFFECT, sleep quantity and quality show the clearest difference between high and low negative affect groups. These findings support LLM-guided concept integration for accurate and interpretable prediction in small-cohort mobile sensing studies.
Yuning Wang, Iman Azimi, Amir M. Rahmani +1
University of Turku, Department of Computing, Turku, Finland · University of California, Irvine, Department of Computer Science, California, USA