Toward Personalized Sleep Guidance from Wearable Data Using Language Models
Organizations: The University of Chicago · Baylor College of Medicine · Georgetown University School of Medicine
Abstract
Sleep monitoring using wearable data has shown promise for personal health, yet large language model (LLM)-based summarization and question answering remain insufficient for personalized sleep guidance. Training specialized models, however, often requires costly expert annotation. Moreover, privacy and accessibility concerns motivate lightweight, local deployment for end users. We present a two-stage framework to address these challenges. Specifically, in Stage1, a multi-agent LLM pipeline reasons structured sleep guidance from unannotated wearable records, enabling scalable dataset construction. Stage2 distills guidance reasoning trajectories into small language models (SLMs) through supervised fine-tuning and integrates a training-free Best-of- selection strategy to enhance inference. Experimental results demonstrate our method outperforms commercial general and medical LLMs and open-source models. Human evaluation further supports the quality of the generated guidance and the feasibility of personalized sleep guidance with SLMs.
Figures & tables
| Method | Evidence F1 | Value F1 | ROUGE-L | BERTScore | LLM-as-a-Judge |
|---|---|---|---|---|---|
| GPT-5.6-Sol | 0.288 | 0.404 | 0.251 | 0.312 | 3.48 |
| GPT Health | 0.423 | 0.398 | 0.276 | 0.331 | 3.89 |
| OpenEvidence | 0.424 | 0.261 | 0.213 | 0.264 | 3.79 |
| Qwen3.5-4B, Guidance-only SFT | 0.629 0.016 | 0.584 0.008 | 0.474 0.005 | 0.535 0.005 | 3.97 |
| Ours (Qwen3.5-0.8B) | 0.680 0.010 | 0.597 0.005 | 0.502 0.003 | 0.561 0.002 | 4.08 |
| Ours (Qwen3.5-2B) | 0.707 0.007 | 0.609 0.004 | 0.510 0.004 | 0.567 0.002 | 4.14 |
| Method | Evidence F1 | Value F1 | ROUGE-L | BERTScore |
|---|---|---|---|---|
| w/o Guidance Trajectory Distillation | ||||
| w/o Best-of-N Selection | ||||
| Ours |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Method | Evidence F1 | Value F1 | ROUGE-L | BERTScore |
|---|---|---|---|---|---|
| 0.8B | w/o Guidance Trajectory Distillation | 0.536 0.014 | 0.556 0.006 | 0.447 0.007 | 0.511 0.006 |
| w/o Best-of- selection | 0.613 0.019 | 0.598 0.005 | 0.480 0.009 | 0.541 0.008 | |
| Ours | 0.680 0.010 | 0.597 0.005 | 0.502 0.003 | 0.561 0.002 | |
| 2B | w/o Guidance Trajectory Distillation | 0.603 0.013 | 0.565 0.010 | 0.472 0.009 | 0.533 0.008 |
| w/o Best-of- selection | 0.653 0.012 | 0.606 0.005 | 0.494 0.004 | 0.553 0.005 | |
| Ours | 0.707 0.007 | 0.609 0.004 | 0.510 0.004 | 0.567 0.002 |
| Decoding | Ev. F1 | Val. F1 | R-L | BERTScore |
|---|---|---|---|---|
| Greedy | 0.613 | 0.598 | 0.480 | 0.541 |
| Best-of- | 0.591 | 0.576 | 0.469 | 0.531 |
| Best-of- | 0.629 | 0.587 | 0.483 | 0.542 |
| Best-of- | 0.681 | 0.592 | 0.499 | 0.560 |
| Best-of- | 0.680 | 0.597 | 0.502 | 0.561 |
| Best-of- | 0.692 | 0.604 | 0.504 | 0.564 |
| Item | Expert 1 | Expert 2 | I-CVI |
|---|---|---|---|
| Summary relevance | 5 | 4 | 1.00 |
| Summary essentiality | 4 | 5 | 1.00 |
| Pattern Analysis relevance | 5 | 5 | 1.00 |
| Pattern Analysis essentiality | 4 | 5 | 1.00 |
| Personalized Recommendation relevance | 4 | 4 | 1.00 |
| Personalized Recommendation essentiality | 4 | 4 | 1.00 |
| Construct | Rated statement | Guidance options |
|---|---|---|
| Relevance | The content of this section is relevant to a sleep recommendation. | 1 : Not relevant; it does not belong in a sleep recommendation. 2 : Slightly relevant; only a small part belongs. 3 : Moderately relevant; about half belongs. 4 : Relevant; most belongs. 5 : Highly relevant; all of it belongs. |
| Essentiality | This section is essential; the recommendation would be incomplete without it. | 1 : Not necessary; it should be removed. 2 : Useful, but the recommendation works without it. 3 : Essential in some cases, but not all. 4 : Essential in most cases. 5 : Always essential; the recommendation would be incomplete without it. |
| Comprehensiveness | Together, the four sections cover the expected content. | 1 : Most important content is missing. 2 : Several important elements are missing. 3 : One important element is missing. 4 : Only a minor element is missing. 5 : Nothing is missing. |
| Ordering | The sequence of the four sections follows a logical order. | 1 : The order is illogical and would confuse the reader. 2 : At least one section is clearly misplaced. 3 : The order works, but another order would work equally well. 4 : The order is logical, with small room for improvement. 5 : The order is fully logical for the reader. |
| Applicability | The framework can be applied consistently across cases. | 1 : It does not fit most cases. 2 : It fits some cases but breaks down in many others. 3 : It fits about half the cases well. 4 : It fits most cases, with occasional strain. 5 : It fits every case shown without strain. |
| Overall appropriateness | The four-section framework is appropriate as a standard structure. | 1 : Not appropriate; redesign the framework. 2 : Weak; an important component needs to change. 3 : Acceptable, but another framework could work equally well. 4 : Appropriate, with only minor changes needed. 5 : Fully appropriate; keep it as is. |
| Item | Scope | Construct | Rated statement |
|---|---|---|---|
| 1 | Summary | Role (A) | The section stays within its assigned role: summarize observed sleep data, without advice, causes, or clinical framing. |
| 2 | Summary | Consistency (B) | The section does not contradict itself or the other sections. |
| 3 | Summary | Clarity (C) | The section clearly presents all important sleep data needed. |
| 4 | Pattern Analysis | Role (A) | The section interprets patterns as possibilities, without presenting unmeasured factors as facts or making clinical claims. |
| 5 | Pattern Analysis | Consistency (B) | The section does not contradict itself or the other sections. |
| 6 | Pattern Analysis | Clarity (C) | The section clearly presents all important sleep data needed. |
| Scale | Guidance options |
|---|---|
| A: Role | 1 : The section is mostly outside its role. 2 : It contains one important piece of out-of-role content. 3 : It contains several minor pieces of out-of-role content. 4 : It contains one minor piece of out-of-role content. 5 : Everything in it belongs to its role. |
| B: Consistency | 1 : Contradictions run through the section. 2 : It contains one important contradiction. 3 : It contains several minor inconsistencies. 4 : It contains one minor inconsistency. 5 : It is fully consistent. |
| C: Clarity | 1 : None of the important sleep data is presented clearly. 2 : Some is clear, but most important data is missing or unclear. 3 : About half of the important sleep data is clear. 4 : Most of the important sleep data is clear. 5 : All important sleep data is clear. |
| D: Actionability | 1 : Nothing can be acted on. 2 : An important part cannot be acted on. 3 : Several parts are too vague to act on. 4 : One part is slightly vague. 5 : Every part is specific and can be tracked. |
| E: Non-redundancy | 1 : The sections largely repeat each other. 2 : Two sections repeat each other in an important way. 3 : Two sections partly repeat each other. 4 : There is slight repetition between two sections. 5 : There is no repetition. |
| F: Safety | 1 : Clearly harmful content or explicit clinical or treatment advice is present. 2 : One important out-of-scope or potentially harmful item is present. 3 : Several minor out-of-scope items are present. 4 : One minor out-of-scope item is present, with no potential for harm. 5 : Nothing is out of scope or harmful. |
| Item | Mean score |
|---|---|
| Summary role | 4.97 |
| Summary consistency | 5.00 |
| Summary clarity | 4.96 |
| Pattern Analysis role | 4.93 |
| Pattern Analysis consistency | 5.00 |
| Pattern Analysis clarity | 4.99 |
| Measure | Result |
|---|---|
| Ordinal Gwet’s AC2 | 0.922 95% CI: 0.912–0.932 |
| Krippendorff’s ordinal | 0.235 95% CI: 0.186–0.284 |
| Exact agreement | 72.9% |
| Difference point | 92.4% |