Streaming spoken agents may produce the correct final action after acting too early. Final-turn scores do not reveal whether each observed speech prefix supports an exposed action. We introduce the Partial Speech Action Contract for Turn Taking in Speech Language Models (PACT-SLM), a controlled test that assigns the first valid action time and evaluates both action identity and timing. In the primary test, 80 paired contrast groups from four held-out semantic families yield 1,600 prefix predictions across clean and 15 dB noise renderings. Using source-utterance semantic targets rather than counterbalanced branch codes, WavLM Base Plus reaches 26.03% pooled post-onset semantic-label accuracy (95% group-bootstrap interval [22.14%, 29.68%]), exposes an action on 18.99% of pre-onset prefixes, and predicts 5.94% of complete trajectories exactly. Its pooled label score is at the 96th percentile of 100 within-prefix label permutations, below the 97.5th-percentile reference (26.73%). It exceeds matched text, scalar-acoustic, and shuffled-representation probes in pooled post-onset label accuracy. Elapsed time is more onset-exact (36.25% versus 23.13%) but less accurate about action identity (9.92% versus 26.03%). These results motivate separate measurement of action identity and onset timing in partial-speech evaluations.
Figures & tables
Figure 1: The PACT-SLM contract test. In the fixed-endpoint test (top), an action is valid only at the final prefix, so a duration or endpoint rule can wait until the last crop. In the variable-onset test (bottom), paired observed prefixes remain identical through the last required wait, action-specific evidence appears before the first valid action prefix, and every clean or noisy prefix is cropped from one branch waveform. Held-out families and noise test whether a probe uses transferable action evidence.
Split
Family memberships
Groups
Source trajectories
Prefix records
Rendering
Train
13
208
416
2,080
Clean
Development
13
52
104
520
Clean
Test
4
80
160
1,600
Clean, 15 dB noise
Table 1: Split composition after excluding pairs with unavailable source audio and pairs without an action contrast. A group is a paired action contrast; a source trajectory is one utterance in that group; a prefix record is one scored crop in one rendering. The 13 family memberships listed for training and development may overlap. Test includes clean and held-out 15 dB noise renderings.
Model or control
Trajectory exact (%) ↑
Pre-onset exposed (%) ↓
Post-onset label (%) ↑
Onset exact (%) ↑
Always wait
0.00
0.00
0.00
0.00
Endpoint plurality
0.00
0.00
4.13
0.00
Elapsed time only
0.00
4.43
9.92
36.25
Authored text only
3.75
58.23
18.60
16.25
Scalar acoustics
1.56
10.13
14.88
18.13
Prefix-shuffled WavLM
0.31
51.58
15.29
17.81
Table 2: Variable-onset results on 80 held-out contrast groups across clean and 15 dB noise renderings (1,600 prefix predictions). Values are percentages. Trajectory exactness requires all five decisions for one trajectory and rendering to match. Exposed-action rate is calculated over pre-onset prefixes only and is lower when better; pooled post-onset semantic-label accuracy and onset exactness are higher when better. WavLM Base Plus uses source-utterance semantic targets. Intervals are in Appendix D .
Metric
Clean
15 dB noise
Trajectory exact ↑
11.25 [6.25, 17.50]
0.62 [0.00, 1.88]
Pre-onset exposed ↓
21.52 [14.57, 29.17]
16.46 [10.69, 23.32]
Pooled post-onset label ↑
38.43 [31.76, 45.38]
13.64 [9.80, 17.72]
Onset exact ↑
40.00 [29.38, 49.38]
6.25 [1.88, 11.88]
Table 3: WavLM Base Plus results by held-out rendering. Values are percentages with 1,000 contrast-group-bootstrap 95% intervals.
Figure 2: Results on the 80-group held-out test under source-utterance semantic targets. Each panel reports one metric for all scored probes; points and intervals are percentages and 95% intervals from 1,000 resamples of the 80 contrast groups. Intervals hold fitted probes and their held-out predictions fixed. The audio-only probe exceeds matched learned controls on pooled post-onset label accuracy, while elapsed time has higher onset exactness. Complete-trajectory accuracy remains low.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Model or control
Trajectory exact ↑
Pre-onset exposed ↓
Post-onset label ↑
Onset exact ↑
Always wait
0.00 [0.00, 0.00]
0.00 [0.00, 0.00]
0.00 [0.00, 0.00]
0.00 [0.00, 0.00]
Endpoint plurality
0.00 [0.00, 0.00]
0.00 [0.00, 0.00]
4.13 [2.63, 5.86]
0.00 [0.00, 0.00]
Elapsed time only
0.00 [0.00, 0.00]
4.43 [1.79, 7.69]
9.92 [6.96, 12.86]
36.25 [25.00, 47.50]
Text only
3.75 [1.25, 6.88]
58.23 [48.66, 68.02]
18.60 [13.93, 23.14]
16.25 [8.75, 25.00]
Scalar acoustics
1.56 [0.31, 3.12]
10.13 [7.10, 13.81]
14.88 [12.24, 17.60]
18.13 [12.81, 23.13]
Prefix-shuffled WavLM
0.31 [0.00, 0.94]
51.58 [46.73, 56.44]
15.29 [12.70, 17.83]
17.81 [12.81, 22.81]
Appendix
Table 4: Estimates in percent with 1,000 group-bootstrap 95% intervals. Each cell is estimate [lower, upper]. The pre-onset exposure denominator contains only prefixes strictly before the valid onset.
Reference
Δ trajectory exact
Δ pre-onset exposed
Δ post-onset label
Δ onset exact
Always wait
+5.94 [+3.44, +9.06]
+18.99 [+14.19, +24.85]
+26.03 [+22.14, +29.68]
+23.13 [+17.81, +28.13]
Endpoint plurality
+5.94 [+3.44, +9.06]
+18.99 [+14.19, +24.85]
+21.90 [+19.02, +24.90]
+23.13 [+17.81, +28.13]
Elapsed time
+5.94 [+3.44, +9.06]
+14.56 [+8.75, +21.39]
+16.12 [+10.85, +21.50]
-13.12 [-23.76, -1.88]
Text only
+2.19 [-1.56, +5.94]
-39.24 [-49.68, -28.02]
+7.44 [+2.62, +12.60]
+6.88 [-2.82, +15.31]
Scalar acoustics
+4.38 [+1.25, +8.12]
+8.86 [+2.90, +15.50]
+11.16 [+5.66, +15.95]
+5.00 [-0.94, +11.56]
Shuffled WavLM
+5.62 [+2.81, +8.75]
-32.59 [-38.26, -26.77]
+10.74 [+6.73, +14.56]
+5.31 [-0.31, +10.94]
Appendix
Table 5: Absolute differences in each named metric (%) for WavLM Base Plus minus the named reference, with 95% group-bootstrap intervals. Positive values favor WavLM for the upward metrics; a positive pre-onset exposure difference is unfavorable.
Family
Audio post-onset label
Text+audio post-onset label
Audio exposed before onset
Audio onset exact
Audio trajectory exact
Text+audio trajectory exact
Book uncertainty
19.67
17.62
14.10
21.25
0.00
1.25
Calendar noise
8.75
7.92
22.50
18.75
1.25
1.25
Message barge-in
47.08
47.50
12.50
22.50
15.00
16.25
Timer urgency
28.69
27.46
26.92
30.00
7.50
7.50
Appendix
Table 6: Descriptive results within each of the four held-out semantic families (20 contrast groups per family). Values are percentages; these point estimates do not establish a between-family effect.