SoK: Semantic Decision Engines in Network Control Loops
Organizations: School of Electrical, Mechanical and Biomedical Engineering University of Technology Sydney, Sydney, Australia
Abstract
A semantic decision engine such as Jev can return a valid answer and still miss a network deadline, select an infeasible action or leave the service unverified. We systematize 139 paper families by decision interface, execution path and check ownership. Fifty families claim that their engine fits a control loop or time budget, but only four support the claim with matched measurement. Across all 139, four report deadline attainment. The gap concentrates where the decision has no deterministic computation step. Those 72 families make 22 of the claims, none supported, and name a coverage owner in only two. Bounded tests under one event model show that each gap can reverse an admission verdict. A decision that meets a 10 s budget for every isolated request meets it for none once decisions queue ahead of replayed execution times. The same engine passes one coverage check and fails another. We derive a minimum reporting record, design rules and a research agenda for admitting decision engines to control loops.
Figures & tables
| Loop claim | Check owned | |||||||
| Interface | Any | Matched | p95+ | Deadline | Load | Feasib. | Cover. | |
| S only | 11 | 4 | 0 | 0 | 1 | 4 | 1 | 0 |
| G only | 24 | 5 | 0 | 1 | 1 | 3 | 9 | 0 |
| U only | 2 | 0 | 0 | 0 | 0 | 0 | 1 | 0 |
| Mixed, no C | 35 | 13 | 0 | 1 | 0 | 6 | 12 | 2 |
| C only | 19 | 3 | 0 | 2 | 0 | 11 | 14 | 1 |
| Task | Model role | Other workflow steps | Baselines |
|---|---|---|---|
| Intent [ 23 , 2 ] | Classify attributes | Values; binding; joint checks | Classifier; structured LLM |
| Documents [ 24 , 18 ] | Rank or reject | Retrieval; extraction | Retrieval score; re-ranker |
| Configuration [ 25 , 26 ] | Select template or tool | Generation; validation | Constrained generator |
| Properties [ 27 , 28 ] | Interpret requirement | Formalization; checking | Symbolic verifier |
| Diagnosis [ 29 , 30 ] | Rank tests or actions | Hypotheses; probes; execution | Probing; procedures; tool agents |
| Reuse [ 31 , 32 ] | Trigger reinterpretation | Refresh; recomputation | Cache with controller |
| Field | Families (139) | All scopes (429) | Step definitions (393) | Unclear families | All-NA families |
| Timing, workload, and comparison | |||||
| Named timing endpoint | 92/139 (66.2%) [58.3, 74.1] | 164/429 (38.2%) [31.9, 44.9] | 143/393 (36.4%) [30.1, 43.1] | 2 | 0 |
| Specified metric denominator | 49/139 (35.3%) [27.3, 43.2] | 86/429 (20.0%) [14.6, 26.1] | 69/393 (17.6%) [12.5, 23.4] | 55 | 46 |
| p95 or higher | 9/139 (6.5%) [2.9, 10.8] | 14/429 (3.3%) [1.2, 5.8] | 11/393 (2.8%) [1.1, 4.9] | 1 | 0 |
| Deadline attainment | 4/139 (2.9%) [0.7, 5.8] | 7/429 (1.6%) [0.2, 3.6] | 5/393 (1.3%) [0.2, 2.8] | 2 | 0 |
| Input load | 48/139 (34.5%) [26.6, 42.4] | 87/429 (20.3%) [15.0, 26.0] | 74/393 (18.8%) [13.5, 24.5] | 1 | 0 |
| Boundary | What the record establishes |
|---|---|
| Client return | is observed call duration. Format validity reported separately. Service completion needs its own probe. |
| Decision / slot | is the decision after stated checks. spans slot acquisition to actual release. Includes waits, retries and timeouts. |
| Interface ACK | is Kubernetes API acknowledgement or FRRouting configuration submission. RAN uses base-station acknowledgement. |
| Verified result | is a new edge replica Ready plus a correct image response or required route within its round-trip time bound. |
| RAN observation | is the first recorded KPM change. Records a measurement update. Service verification uses a separate . |
| Family / lens | Service contracts [ 2 ] | RAN policies [ 3 ] | Application tasks / intervention | Measured outcomes |
|---|---|---|---|---|
| Structure L4 | 4, 6 or 8 output fields | Fresh with 3, 7, 21 or 57 telemetry cells | 4, 16 or 64 interacting requirements | Task correctness and return time |
| Irrelevant text L4 | Padding up to 16,384 tokens | N/A | Irrelevant context and repeated requirements | Task correctness and return time |
| Observations L3/L4 | N/A | Fresh / stale / noisy / contradictory telemetry | Complete / prose / missing / stale conflict / resolved conflict | Correct handling of observations and return time |
| Covered catalogue L3/L4 | Scale / unseen services partial replacement | N/A | Stable IDs / rekeying replacement with two candidates | Selection correctness and return time |
| Absent coverage L3 | Unsupported requests as an analogy | N/A | Missing contract candidate or infeasible path | Correct rejection or escalation and return time |
| Arrivals L2 | Open admission traces | Open intent-arrival traces | N/A | Queue time and slot load with native deadline attainment |
| Lens | Question | Estimands | Comparison families | Data |
|---|---|---|---|---|
| L1 Measurement boundary | How does the endpoint change timing effects and admission verdicts? | Stage times, endpoint slopes, decision fractions | Decision delay (4 slope tests) | SoK, RAN |
| L2 Workload summary | How do median, tail and deadline summaries change under queued work? | Return and queue quantiles, attainment, , modeled completion | Deadline and capacity (83 load cells) | SoK, Edge, RAN |
| L3 Check placement | What changes when observation, coverage or feasibility checks move within the workflow? | Correctness, validity, return p50/p95, online outcomes | Observation (66), coverage (39), workflow (45) | SoK, Edge, RAN |
| L4 Robustness scope | How do structure, length, representation and catalogue change bound a result? | Correctness, validity, return p50/p95, paired changes | Structure (89) | SoK, Edge, RAN |
| Model | Ownership | Return (s) p50 / p95 / p99 | ms (%) [CI] | s (%) [CI] | High-load | |
|---|---|---|---|---|---|---|
| Access policy, SoK-owned base condition | ||||||
| Jev-1.13 | SoK-owned | 81 | 0.350 / 0.483 / 0.620 | 0.0 [0.0,0.0] | 100.0 [100.0,100.0] | N/A |
| DeepSeek-V4.1-Flash | SoK-owned | 81 | 0.469 / 1.623 / 6.123 | 0.0 [0.0,0.0] | 88.9 [81.2,95.4] | N/A |
| Gemini-3.1-Flash-Lite | SoK-owned | 81 | 0.923 / 1.144 / 1.371 | 0.0 [0.0,0.0] | 74.1 [64.9,82.9] | N/A |
| GLM-4.7-Flash | SoK-owned | 81 | 0.812 / 1.504 / 7.866 | 0.0 [0.0,0.0] | 72.8 [62.3,82.7] | N/A |
| Qwen3.5-4B | SoK-owned | 81 | 0.329 / 0.388 / 0.435 | 0.0 [0.0,0.0] | 100.0 [100.0,100.0] | N/A |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Correct decisions | Valid | Response (s) | Input | ||||
| Method | Changed | Preserved | All | All | Median | p95 | Tokens |
| Short input: 4 requirements, 6k characters | |||||||
| Jev-1.13 | 52/60 | 55/60 | 107/120 | 120/120 | 0.382 | 0.646 | 1699–2640 |
| DeepSeek-V4.1-Flash | 38/60 | 43/60 | 81/120 | 119/120 | 0.558 | 1.441 | 1262–1630 |
| Gemini-3.1-Flash-Lite | 39/60 | 47/60 | 86/120 | 120/120 | 0.987 | 1.252 | 1396–2321 |
| GLM-4.7-Flash | 43/60 | 39/60 | 82/120 | 120/120 | 0.883 | 4.645 | 1174–1559 |
| Correct decisions | Valid | Response (s) | |||||
| Method | 3 nodes | 5 nodes | 8 nodes | All | All | Median | p95 |
| Complete / shared quota | |||||||
| Jev-1.13 | 10/40 | 2/40 | 2/40 | 14/120 | 120/120 | 0.364 | 0.619 |
| DeepSeek-V4.1-Flash | 10/40 | 4/40 | 0/40 | 14/120 | 120/120 | 0.499 | 0.754 |
| Gemini-3.1-Flash-Lite | 17/40 | 10/40 | 9/40 | 36/120 | 120/120 | 1.027 | 1.584 |
| GLM-4.7-Flash | 0/40 | 1/40 | 0/40 | 1/120 | 120/120 | 0.884 | 8.373 |
| Correct decisions | Stable only | |||||
| Method | Stable | Rekeyed | Replacement | Absent | Valid | Median (s) |
| A. Contract interpretation: 48 configurations per condition | ||||||
| Jev-1.13 | 48/48 | 48/48 | 48/48 | 45/48 | 48/48 | 0.323 |
| DeepSeek-V4.1-Flash | 47/48 | 43/48 | 37/48 | 3/48 | 48/48 | 0.447 |
| Gemini-3.1-Flash-Lite | 48/48 | 48/48 | 48/48 | 14/48 | 48/48 | 0.893 |
| GLM-4.7-Flash | 27/48 | 19/48 | 1/48 | 16/48 | 48/48 | 0.775 |
| Correct decisions | Response (s) | ||||||||
| Method | Connectivity | Required | Avoid | Hops | Budget | Ordered | Overall | Median | p95 |
| A. Feasible candidate present | |||||||||
| Jev-1.13 | 17/20 | 15/20 | 16/20 | 19/20 | 18/20 | 10/20 | 95/120 | 0.351 | 0.484 |
| DeepSeek-V4.1-Flash | 14/20 | 16/20 | 13/20 | 17/20 | 6/20 | 12/20 | 78/120 | 0.525 | 0.722 |
| Gemini-3.1-Flash-Lite | 15/20 | 16/20 | 19/20 | 16/20 | 18/20 | 13/20 | 97/120 | 0.954 | 1.208 |
| GLM-4.7-Flash | 1/20 | 1/20 | 1/20 | 5/20 | 0/20 | 1/20 | 9/120 | 0.872 | 6.666 |
| Correct decisions | Valid | Response (s) | |||||
| Method | 3 nodes | 5 nodes | 8 nodes | All | All | Median | p95 |
| Independent quota | |||||||
| Jev-1.13 | 17/40 | 7/40 | 5/40 | 29/120 | 120/120 | 0.358 | 0.642 |
| DeepSeek-V4.1-Flash | 12/40 | 5/40 | 2/40 | 19/120 | 120/120 | 0.502 | 0.861 |
| Gemini-3.1-Flash-Lite | 18/40 | 18/40 | 4/40 | 40/120 | 120/120 | 0.994 | 1.326 |
| GLM-4.7-Flash | 0/40 | 1/40 | 0/40 | 1/120 | 120/120 | 0.813 | 9.581 |
| Challenger - anchor | Endpoint | violation pp [95% CI] | Contrast direction | |
| Radio base point: latency-only arms, event-triggered mode | ||||
| DeepSeek - Jev | Affected class | 0.03 [-1.45,1.55] | 1 | Unresolved |
| DeepSeek - Jev | Network-wide | -0.66 [-2.14,0.46] | 1 | Unresolved |
| GLM 5.3 - Jev | Affected class | -2.45 [-5.13,-0.63] | 0.0072 | Lower violation |
| GLM 5.3 - Jev | Network-wide | -2.66 [-4.94,-1.03] | 0.0007999 | Lower violation |
| Qwen 3.8 - Jev | Affected class | -0.98 [-2.21,0.23] | 0.581 | Unresolved |
| (s) | p50 / p95 | p50 / p95 | |||||
|---|---|---|---|---|---|---|---|
| Transport–sparse physical execution | |||||||
| 0 | 300/300 | 0.044 / 0.051 | 0.561 / 0.586 | 0.1 | 0.0 | 100.0 [100.0, 100.0] | N/A |
| 0.1 | 300/300 | 0.144 / 0.151 | 0.659 / 0.687 | 69.3 | 15.2 | 100.0 [100.0, 100.0] | N/A |
| 0.5 | 300/300 | 0.544 / 0.551 | 1.060 / 1.086 | 91.9 | 47.2 | 100.0 [100.0, 100.0] | N/A |
| 2 | 300/300 | 2.045 / 2.051 | 2.561 / 2.587 | 97.8 | 78.1 | 100.0 [100.0, 100.0] | N/A |
| Transport–queue + execution replay, , /s | |||||||
| Implementation | Offered [CI] | Accepted | Queue (s) p50 / p95 | Return (%) [CI] | Native completion (%) [CI] |
|---|---|---|---|---|---|
| Service admission: 1 arrivals/s, B = 2 s. 300 arrivals per row | |||||
| Jev-1.13.0 | 0.071 [0.070,0.073] | 0.071 | 0.002 / 0.002 | 100.0 [100.0,100.0] | 94.0 [90.7,96.8] |
| SemIf-Qwen3.5-4B | 0.170 [0.156,0.184] | 0.170 | 0.002 / 0.002 | 100.0 [100.0,100.0] | 63.0 [56.6,69.9] |
| Laya | 0.028 [0.027,0.029] | 0.028 | 0.002 / 0.002 | 100.0 [100.0,100.0] | 11.7 [7.3,15.7] |
| DeepSeek-V4.1-Flash | 0.171 [0.094,0.273] | 0.165 | 0.002 / 0.002 | 92.3 [83.3,100.0] | 90.7 [81.7,97.9] |
| GLM-5.3-Flash | 0.217 [0.202,0.231] | 0.217 | 0.002 / 0.002 | 93.3 [90.7,96.1] | 91.7 [88.7,94.8] |
| Implementation | Offered [CI] | Accepted | Queue (s) p50 / p95 | Return (%) [CI] | Native completion (%) [CI] |
|---|---|---|---|---|---|
| Service admission: 4 arrivals/s, B = 2 s. 300 arrivals per row | |||||
| Jev-1.13.0 | 0.272 [0.268,0.277] | 0.272 | 0.002 / 0.002 | 99.7 [99.2,100.0] | 93.7 [89.8,95.3] |
| SemIf-Qwen3.5-4B | 2.637 [2.602,2.667] | 1.037 | 1.812 / 1.986 | 1.0 [0.0,2.7] | 1.0 [0.0,2.7] |
| Laya | 0.102 [0.097,0.106] | 0.102 | 0.002 / 0.002 | 100.0 [100.0,100.0] | 11.7 [8.0,15.6] |
| DeepSeek-V4.1-Flash | 0.401 [0.399,0.406] | 0.401 | 0.002 / 0.005 | 100.0 [100.0,100.0] | 98.3 [96.6,99.2] |
| GLM-5.3-Flash | 0.883 [0.806,1.018] | 0.871 | 0.349 / 1.508 | 84.3 [62.7,90.3] | 80.7 [57.6,86.7] |
| Implementation | Offered [CI] | Accepted | Queue (s) p50 / p95 | Return (%) [CI] | Native completion (%) [CI] |
|---|---|---|---|---|---|
| RAN policy: 0.1 arrivals/s, B = 1 s. 300 arrivals per row | |||||
| Jev-1.13.0 | 0.008 [0.008,0.008] | 0.008 | 0.001 / 0.002 | 100.0 [100.0,100.0] | 100.0 [100.0,100.0] |
| SemIf-Qwen3.5-4B | 0.011 [0.007,0.019] | 0.012 | 0.006 / 0.027 | 98.7 [95.7,100.0] | 98.7 [95.7,100.0] |
| AnyJev-L0 | 0.017 [0.017,0.018] | 0.019 | 0.005 / 0.026 | 92.7 [88.6,96.4] | 74.3 [69.0,79.7] |
| DeepSeek-V4.1-Flash | 0.020 [0.018,0.023] | 0.022 | 0.001 / 0.002 | 89.0 [85.1,92.8] | 89.0 [85.1,92.8] |
| GLM-5.3-Flash | 0.039 [0.035,0.044] | 0.042 | 0.001 / 0.002 | 23.3 [17.7,29.5] | 22.0 [16.4,28.0] |
| RAN policy: fresh / stale / noisy / contradictory, 57 cells | ||||||
|---|---|---|---|---|---|---|
| Model | Fresh | Stale | Noisy | Contradictory | Valid / | Return (s) p50 / p95 |
| Jev-1.13.0 | 90.7 [84.9,94.4] | 70.0 [62.2,76.8] | 88.7 [82.6,92.8] | 63.3 [55.4,70.6] | 150/150 | 0.292 / 0.408 |
| SemIf-Qwen3.5-4B | 14.0 [9.3,20.5] | 5.3 [2.7,10.2] | 14.7 [9.9,21.2] | 16.7 [11.6,23.4] | 150/150 | 0.601 / 0.713 |
| AnyJev-L0 | 28.0 [21.4,35.7] | 21.3 [15.5,28.6] | 24.0 [17.9,31.4] | 24.7 [18.5,32.1] | 150/150 | 0.492 / 0.504 |
| DeepSeek-V4.1-Flash | 96.0 [91.5,98.2] | 69.3 [61.5,76.2] | 97.3 [93.3,99.0] | 66.7 [58.8,73.7] | 150/150 | 0.683 / 1.013 |
| GLM-5.3-Flash | 96.7 [92.4,98.6] | 95.3 [90.7,97.7] | 96.7 [92.4,98.6] | 85.3 [78.8,90.1] | 137/150 | 1.851 / 2.975 |
| Contract catalogue: stable / rekeyed / two new candidates / absent | ||||||
|---|---|---|---|---|---|---|
| Model | Stable | Rekeyed | Two new | Absent | Valid / | Return (s) p50 / p95 |
| Jev-1.13 | 100.0 [92.6,100.0] | 100.0 [92.6,100.0] | 100.0 [92.6,100.0] | 93.8 [83.2,97.9] | 48/48 | 0.318 / 0.511 |
| DeepSeek-V4.1-Flash | 97.9 [89.1,99.6] | 89.6 [77.8,95.5] | 77.1 [63.5,86.7] | 6.2 [2.1,16.8] | 48/48 | 0.413 / 2.696 |
| Gemini-3.1-Flash-Lite | 100.0 [92.6,100.0] | 100.0 [92.6,100.0] | 100.0 [92.6,100.0] | 29.2 [18.2,43.2] | 48/48 | 0.934 / 1.600 |
| GLM-4.7-Flash | 56.2 [42.3,69.3] | 39.6 [27.0,53.7] | 2.1 [0.4,10.9] | 33.3 [21.7,47.5] | 48/48 | 0.825 / 2.943 |
| Qwen3.5-4B | 54.2 [40.3,67.4] | 54.2 [40.3,67.4] | 64.6 [50.4,76.6] | 2.1 [0.4,10.9] | 48/48 | 0.243 / 0.331 |
| Method | Self-check | Public gate | pp [95% CI] | Valid / | Return p50 / p95 (s) | |
| Independent quotas | ||||||
| Jev-1.13 | 29/120 | 84/120 | 45.8 [37.5,54.2] | 120/120 | 0.358 / 0.604 | |
| DeepSeek-V4.1-Flash | 19/120 | 60/120 | 34.2 [25.0,43.3] | 120/120 | 0.509 / 1.175 | |
| Gemini-3.1-Flash-Lite | 40/120 | 69/120 | 24.2 [14.2,34.2] | 0.0008467 | 120/120 | 1.009 / 1.241 |
| GLM-4.7-Flash | 1/120 | 71/120 | 58.3 [49.2,66.7] | 119/120 | 0.843 / 3.553 | |
| Qwen3.5-4B | 30/120 | 52/120 | 18.3 [10.8,26.7] | 0.003618 | 120/120 | 0.311 / 0.411 |
| Service contract: 4 / 6 / 8 fields, low constraint density | |||||||
|---|---|---|---|---|---|---|---|
| Model | 4 fields | 6 fields | 8 fields | N/A | pp [CI] | Valid / | Return (s) p50 / p95 |
| Jev-1.13.0 | 92.3 [88.8,94.8] | 78.7 [73.7,82.9] | 57.7 [52.0,63.1] | N/A | N/A | 300/300 | 0.272 / 0.371 |
| SemIf-Qwen3.5-4B | 54.3 [48.7,59.9] | 40.0 [34.6,45.6] | 13.3 [9.9,17.6] | N/A | N/A | 300/300 | 0.133 / 0.194 |
| Laya | 25.3 [20.7,30.5] | 7.7 [5.2,11.2] | 1.0 [0.3,2.9] | N/A | N/A | 300/300 | 0.057 / 0.242 |
| DeepSeek-V4.1-Flash | 98.3 [96.2,99.3] | 95.7 [92.7,97.5] | 90.0 [86.1,92.9] | N/A | N/A | 300/300 | 0.460 / 0.591 |
| GLM-5.3-Flash | 89.7 [85.7,92.6] | 96.3 [93.6,97.9] | 80.3 [75.5,84.4] | N/A | N/A | 300/300 | 0.930 / 1.954 |
| Method | Condition | n | correctness pp [CI] | Paired return (s) [CI] | |
| Service-contract padding: 16384 tokens minus base | |||||
| Jev-1.13.0 | Padding | 300 | -2.7 [-5.7,0.3] | 1 | 0.149 [0.135,0.157] |
| SemIf-Qwen3.5-4B | Padding | 300 | -15.7 [-21.3,-9.7] | 1.034 [1.032,1.036] | |
| Laya | Padding | 300 | -9.3 [-13.0,-6.0] | 0.082 [0.081,0.082] | |
| DeepSeek-V4.1-Flash | Padding | 300 | -1.7 [-3.3,-0.3] | 1 | 0.259 [0.249,0.276] |
| GLM-5.3-Flash | Padding | 300 | -4.7 [-7.3,-2.0] | 0.07742 | 0.883 [0.781,0.962] |
| Work | Interface | Task and execution path | Check responsibility |
|---|---|---|---|
| [ 2 ] | S/G/C | Intent interpretation to edge admission/execution | State: Scheduler reads queues/workers; Feasibility: Scheduler: predicted feasibility; Coverage: Scheduler: reject if no eligible node |
| [ 18 ] | G | Retrieved manuals to multivendor command-line output | Feasibility: INTA: syntax matching; LLM: semantic equivalence |
| [ 19 , 37 ] | S/G/C | Intent parameters to placement/routing solver | Feasibility: Integer linear program: resource/placement constraints; Coverage: Optimizer: joint request-subset admission |
| [ 20 , 38 ] | G | Tool-driven anomaly localization and diagnosis | NE |
| [ 23 ] | S/G | Intent extraction; proposed core placement | NE |
| [ 24 ] | G | Deployment-workflow generation and twin testing | NE |
| Field | Positive reporting criterion | Response categories |
|---|---|---|
| Timing, workload, and comparison | ||
| Timing endpoint | Explicit end event for a measured interval: model return, action dispatch, confirmed network effect, verified service result, or another named event. | Model; dispatch; network; service; other; not reported; unclear; N/A. |
| Metric denominator | Explicit population entering the metric: all attempts, successful requests, another subset, or different denominators for different metrics. | All; successful; other; mixed; unclear; N/A. |
| Tail latency | A latency quantile at p95 or higher. | Reported; not reported; unclear; N/A. |
| Deadline attainment | The fraction or count attaining an explicit deadline. | Reported; not reported; unclear; N/A. |
| Input load | Operational arrival rate, traffic load, concurrency or queue input, as defined below. | Reported; not reported; unclear; N/A. |
| Field | Initial comparable | Aligned recheck | Combined | ||||||
|---|---|---|---|---|---|---|---|---|---|
| A (%) | AC1 | A (%) | AC1 | A (%) | AC1 | ||||
| Timing, workload, and comparison | |||||||||
| Named timing endpoint | 93.5 | N/A | N/A | 87.9 | N/A | N/A | 91.1 | N/A | N/A |
| Specified metric denominator | 91.8 | 0.839 | 0.908 | 89.7 | 0.835 | 0.882 | 90.9 | 0.838 | 0.897 |
| p95 or higher | 100.0 | 1.000 | 1.000 | 99.4 | 0.935 | 0.994 | 99.8 | 0.965 | 0.997 |
| Deadline attainment | 99.1 | 0.663 | 0.991 | 97.1 | 0.674 | 0.970 | 98.3 | 0.674 | 0.982 |
| Coder vs coder | Coder vs table | |||||
| Field | A (%) | AC1 | GPT | Claude | ||
| Decision interface | ||||||
| Selection (S) | 129 | 76.7 | 0.50 | 0.58 | 0.39 | 0.51 |
| Generation (G) | 129 | 86.8 | 0.72 | 0.75 | 0.76 | 0.72 |
| Computation (C) | 129 | 82.9 | 0.67 | 0.66 | 0.57 | 0.69 |
| Exact interface set | 129 | 53.5 | – | – | 46.5% | 53.5% |
| Option | Researcher 1 selected | Researcher 2 selected | A (%) | AC1 | |
|---|---|---|---|---|---|
| Timing endpoint | |||||
| Model return | 38 | 41 | 97.3 | 0.846 | 0.967 |
| Dispatch | 2 | 6 | 98.5 | 0.244 | 0.985 |
| Network effect | 11 | 12 | 99.3 | 0.866 | 0.992 |
| Service result | 23 | 22 | 97.8 | 0.788 | 0.975 |
| Other | 92 | 81 | 94.3 | 0.831 | 0.914 |
| Field | Response category: number of scopes |
|---|---|
| Timing, workload, and comparison | |
| Named timing endpoint | Model return: 39; Dispatch: 6; Network effect: 12; Service result: 23; Other: 86; Not reported: 253; Unclear: 2 |
| Specified metric denominator | All attempts: 48; Successful: 5; Other: 16; Mixed: 7; Unclear: 76; N/A: 253 |
| p95 or higher | Reported: 14; Not reported: 390; Unclear: 1 |
| Deadline attainment | Reported: 7; Not reported: 397; Unclear: 1 |
| Input load | Reported: 81; Not reported: 324 |