Sleep monitoring using wearable data has shown promise for personal health, yet large language model (LLM)-based summarization and question answering remain insufficient for personalized sleep guidance. Training specialized models, however, often requires costly expert annotation. Moreover, privacy and accessibility concerns motivate lightweight, local deployment for end users. We present a two-stage framework to address these challenges. Specifically, in Stage1, a multi-agent LLM pipeline reasons structured sleep guidance from unannotated wearable records, enabling scalable dataset construction. Stage2 distills guidance reasoning trajectories into small language models (SLMs) through supervised fine-tuning and integrates a training-free Best-of-N selection strategy to enhance inference. Experimental results demonstrate our method outperforms commercial general and medical LLMs and open-source models. Human evaluation further supports the quality of the generated guidance and the feasibility of personalized sleep guidance with SLMs.
Figures & tables
Figure 1: Overview of our framework for personalized sleep guidance.
Figure 2: Stage 1 multi-agent pipeline for deriving structured sleep guidance from wearable records. Only examples passing deterministic checks and independent reverse review are retained.
Figure 3: Stage 2 training and two-pass inference with Best-of- N selection and relevant evidence retrieval.
Method
Evidence F1
Value F1
ROUGE-L
BERTScore
LLM-as-a-Judge
GPT-5.6-Sol
0.288
0.404
0.251
0.312
3.48
GPT Health
0.423
0.398
0.276
0.331
3.89
OpenEvidence
0.424
0.261
0.213
0.264
3.79
Qwen3.5-4B, Guidance-only SFT
0.629 ± 0.016
0.584 ± 0.008
0.474 ± 0.005
0.535 ± 0.005
3.97
Ours (Qwen3.5-0.8B)
0.680 ± 0.010
0.597 ± 0.005
0.502 ± 0.003
0.561 ± 0.002
4.08
Ours (Qwen3.5-2B)
0.707 ± 0.007
0.609 ± 0.004
0.510 ± 0.004
0.567 ± 0.002
4.14
Table 1: Comparison of ours and baselines. All rows use the same 300 held-out test cases. Multi-seed results of trainable models report mean ± sd. marks frozen prompting models; marks fine-tuned models. LLM-as-a-Judge scores use a 1–5 Likert scale.
Method
Evidence F1
Value F1
ROUGE-L
BERTScore
w/o Guidance Trajectory Distillation
0.536±0.014
0.556±0.006
0.447±0.007
0.511±0.006
w/o Best-of-N Selection
0.613±0.019
0.598±0.005
0.480±0.009
0.541±0.008
Ours
0.680±0.010
0.597±0.005
0.502±0.003
0.561±0.002
Table 2: Results of ablation studies. All methods use Qwen3.5 with 0.8B parameters. Result values are mean ± sd across three training seeds.
Figure 4: Human evaluation of generated guidance. Mean item-wise ratings with case-bootstrap 95% confidence intervals across 90 evaluated cases.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Method
Evidence F1
Value F1
ROUGE-L
BERTScore
0.8B
w/o Guidance Trajectory Distillation
0.536 ± 0.014
0.556 ± 0.006
0.447 ± 0.007
0.511 ± 0.006
w/o Best-of- N selection
0.613 ± 0.019
0.598 ± 0.005
0.480 ± 0.009
0.541 ± 0.008
Ours
0.680 ± 0.010
0.597 ± 0.005
0.502 ± 0.003
0.561 ± 0.002
2B
w/o Guidance Trajectory Distillation
0.603 ± 0.013
0.565 ± 0.010
0.472 ± 0.009
0.533 ± 0.008
w/o Best-of- N selection
0.653 ± 0.012
0.606 ± 0.005
0.494 ± 0.004
0.553 ± 0.005
Ours
0.707 ± 0.007
0.609 ± 0.004
0.510 ± 0.004
0.567 ± 0.002
Appendix
Table 3: Ablation across Qwen3.5 model sizes. Values are mean ± sd over three seeds.
Decoding
Ev. F1
Val. F1
R-L
BERTScore
Greedy
0.613
0.598
0.480
0.541
Best-of- 2
0.591
0.576
0.469
0.531
Best-of- 4
0.629
0.587
0.483
0.542
Best-of- 8
0.681
0.592
0.499
0.560
Best-of- 16
0.680
0.597
0.502
0.561
Best-of- 32
0.692
0.604
0.504
0.564
Appendix
Table 4: Effect of candidate-pool size on Qwen3.5-0.8B. All results use the same checkpoint.
Item
Expert 1
Expert 2
I-CVI
Summary relevance
5
4
1.00
Summary essentiality
4
5
1.00
Pattern Analysis relevance
5
5
1.00
Pattern Analysis essentiality
4
5
1.00
Personalized Recommendation relevance
4
4
1.00
Personalized Recommendation essentiality
4
4
1.00
Appendix
Table 5: Item-level Stage-1 framework ratings. I-CVI is the proportion of experts assigning a score of 4 or 5.
Construct
Rated statement
Guidance options
Relevance
The content of this section is relevant to a sleep recommendation.
1 : Not relevant; it does not belong in a sleep recommendation. 2 : Slightly relevant; only a small part belongs. 3 : Moderately relevant; about half belongs. 4 : Relevant; most belongs. 5 : Highly relevant; all of it belongs.
Essentiality
This section is essential; the recommendation would be incomplete without it.
1 : Not necessary; it should be removed. 2 : Useful, but the recommendation works without it. 3 : Essential in some cases, but not all. 4 : Essential in most cases. 5 : Always essential; the recommendation would be incomplete without it.
Comprehensiveness
Together, the four sections cover the expected content.
1 : Most important content is missing. 2 : Several important elements are missing. 3 : One important element is missing. 4 : Only a minor element is missing. 5 : Nothing is missing.
Ordering
The sequence of the four sections follows a logical order.
1 : The order is illogical and would confuse the reader. 2 : At least one section is clearly misplaced. 3 : The order works, but another order would work equally well. 4 : The order is logical, with small room for improvement. 5 : The order is fully logical for the reader.
Applicability
The framework can be applied consistently across cases.
1 : It does not fit most cases. 2 : It fits some cases but breaks down in many others. 3 : It fits about half the cases well. 4 : It fits most cases, with occasional strain. 5 : It fits every case shown without strain.
Overall appropriateness
The four-section framework is appropriate as a standard structure.
1 : Not appropriate; redesign the framework. 2 : Weak; an important component needs to change. 3 : Acceptable, but another framework could work equally well. 4 : Appropriate, with only minor changes needed. 5 : Fully appropriate; keep it as is.
Appendix
Table 6: Stage-1 Framework Validation Rubric. Relevance and essentiality were rated separately for each guidance section; the remaining constructs were rated once for the framework as a whole.
Item
Scope
Construct
Rated statement
1
Summary
Role (A)
The section stays within its assigned role: summarize observed sleep data, without advice, causes, or clinical framing.
2
Summary
Consistency (B)
The section does not contradict itself or the other sections.
3
Summary
Clarity (C)
The section clearly presents all important sleep data needed.
4
Pattern Analysis
Role (A)
The section interprets patterns as possibilities, without presenting unmeasured factors as facts or making clinical claims.
5
Pattern Analysis
Consistency (B)
The section does not contradict itself or the other sections.
6
Pattern Analysis
Clarity (C)
The section clearly presents all important sleep data needed.
Appendix
Table 7: Stage-2 rubric item map. Role and consistency were rated for every section; clarity was used for Summary and Pattern Analysis, and actionability for Personalized Recommendation and Follow-up.
Scale
Guidance options
A: Role
1 : The section is mostly outside its role. 2 : It contains one important piece of out-of-role content. 3 : It contains several minor pieces of out-of-role content. 4 : It contains one minor piece of out-of-role content. 5 : Everything in it belongs to its role.
B: Consistency
1 : Contradictions run through the section. 2 : It contains one important contradiction. 3 : It contains several minor inconsistencies. 4 : It contains one minor inconsistency. 5 : It is fully consistent.
C: Clarity
1 : None of the important sleep data is presented clearly. 2 : Some is clear, but most important data is missing or unclear. 3 : About half of the important sleep data is clear. 4 : Most of the important sleep data is clear. 5 : All important sleep data is clear.
D: Actionability
1 : Nothing can be acted on. 2 : An important part cannot be acted on. 3 : Several parts are too vague to act on. 4 : One part is slightly vague. 5 : Every part is specific and can be tracked.
E: Non-redundancy
1 : The sections largely repeat each other. 2 : Two sections repeat each other in an important way. 3 : Two sections partly repeat each other. 4 : There is slight repetition between two sections. 5 : There is no repetition.
F: Safety
1 : Clearly harmful content or explicit clinical or treatment advice is present. 2 : One important out-of-scope or potentially harmful item is present. 3 : Several minor out-of-scope items are present. 4 : One minor out-of-scope item is present, with no potential for harm. 5 : Nothing is out of scope or harmful.
Appendix
Table 8: Five-point score anchors for the Stage-2 rubric.
Figure 5: Pairwise agreement among five experienced MD-student reviewers. a , Pairwise 5×5 contingency matrices for all ten reviewer pairs. Columns correspond to the first named reviewer and rows to the second; cells report raw case–item counts. Each pair shared 27 cases, giving 432 case–item units. b , Full pairwise matrices of ordinal Gwet’s AC2 and exact agreement. Diagonal cells are omitted because self-agreement is not an inter-rater comparison. Overall reliability estimates are reported in Table C.2 .
Item
Mean score
Summary role
4.97
Summary consistency
5.00
Summary clarity
4.96
Pattern Analysis role
4.93
Pattern Analysis consistency
5.00
Pattern Analysis clarity
4.99
Appendix
Table 9: Item-level Stage-2 human rubric scores for the Qwen3.5-0.8B system. Values are means of within-case median scores across 90 cases.
Measure
Result
Ordinal Gwet’s AC2
0.922 95% CI: 0.912–0.932
Krippendorff’s ordinal α
0.235 95% CI: 0.186–0.284
Exact agreement
72.9%
Difference ≤1 point
92.4%
Appendix
Table 10: Inter-rater agreement for the Stage-2 human evaluation. Chance-corrected estimates include case-bootstrap 95% confidence intervals; observed agreement is reported descriptively.
We present SleepLM, a family of sleep-language foundation models that enable human sleep alignment, interpretation, and interaction with natural language. Despite the critical role of sleep, learning-based sleep analysis systems operate in closed label spaces (e.g., predefined stages or events) and fail to describe, query, or generalize to novel sleep phenomena. SleepLM bridges natural language and multimodal polysomnography, enabling language-grounded representations of sleep physiology. To support this alignment, we introduce a multilevel sleep caption generation pipeline that enables the curation of the first large-scale sleep-text dataset, comprising over 100K hours of data from more than 10,000 individuals. Furthermore, we present a unified pretraining objective that combines contrastive alignment, caption generation, and signal reconstruction to better capture physiological fidelity and cross-modal interactions. Extensive experiments on real-world sleep understanding tasks verify that SleepLM outperforms state-of-the-art in zero-shot and few-shot learning, cross-modal retrieval, and sleep captioning. Importantly, SleepLM also exhibits intriguing capabilities including language-guided event localization, targeted insight generation, and zero-shot generalization to unseen tasks. All code and data will be open-sourced.
Language models are remarkably capable at medical question answering, in some cases surpassing the accuracy of general physicians. However, answering questions about wearable health data remains challenging and understudied, as these ubiquitous sensors produce continuous, high-dimensional, and longitudinal data, which is non-trivial to align with text-centric distributions in LLM pretraining. The diversity of sensor modalities and user intents cannot be effectively handled by a fixed reasoning workflow or a single pretrained foundation model. To address these challenges, we propose WEQA, a query-adaptive agent framework that unifies LLM reasoning with specialized wearable analytical and modeling tools. An LLM controller is employed to synthesize execution plans and dynamically route each query to the appropriate combination of sensor analysis and pretrained models, and perform grounded response auditing with external knowledge. We also curate a benchmark spanning four open wearable datasets comprising analytic and predictive tasks in three different health domains. Experiments show that our framework is 24% more accurate than LLM and agentic baselines, and a blinded study with 12 medical experts and 8 users shows substantial gains in usefulness and clinical soundness.
Yuwei Zhang, Tong Xia, Bianca Emmerich +5
University of Cambridge · Tsinghua University · University College London +2
In this paper, we study the problem of personalized survey response prediction using fine-tuned large language models (LLMs). This task poses unique challenges: limited per-user training data, scalability of model storage, and the need to exploit shared structure across survey questions. To address these issues, we propose Aplaud (Adaptive Personalized Low-rank and User-specific Nested Decomposition), a lightweight and scalable framework for LLM personalization. Aplaud extends the LoRA paradigm by separating adaptation into a frozen, shared low-rank basis and a compact user-specific correction, augmented with a rank-one residual for finer personalization. To further reduce per-user parameter cost and mitigate overfitting, the correction matrix can be factorized into an even lower-rank form. Empirical results demonstrate that Aplaud achieves efficient, scalable personalization across users while outperforming state-of-the-art LoRA-based personalized LLM approaches in both generalization and inference efficiency.
Xinyu Li, Ruoming Jin, Jianfeng Zhu +2
Department of Computer Science, Kent State University · iLambda Inc.