Avatar-streaming systems are commonly evaluated with image and video quality assessment (IQA/VQA) metrics, implicitly treating visual fidelity as a proxy for communicative success. We test this assumption through a controlled behavioral study of rendered 3D avatars across a pristine condition and fourteen geometric, photometric, temporal, and combined distortions. Fifty-nine participants contributed 2,688 judgments of perceived action, response confidence, and visual quality. We identify Misleading Quality in this dataset as distorted renderings that retain above-average perceived quality but yield below-average action-recognition accuracy. We also derive an Intent Quality Score (IQS) combining recognition correctness and confidence as the behavioral target for objective metrics. Among 126 distorted content--condition cells, 31 (24.6%) exhibited Misleading Quality; temporal and geometric distortions showed the highest rates, at 50.0% and 31.1%, respectively. The results reveal a quality--accuracy dissociation where distortion families affect appearance and communication differently. Across 24 direct-scoring IQA/VQA metrics and three supervised feature-regression baselines, alignment with IQS remained limited; at λ=0.5, the best leave-one-content-out baseline reached PLCC =0.4435. Under this controlled protocol, visual fidelity alone is insufficient for avatar communication, motivating intent-aware quality assessment and streaming objectives.
Figure 2. Misleading Quality and quality–accuracy dissociation. (a) MOS and recognition accuracy by distortion. Dashed lines show grand means (MOS =3.29 , accuracy =0.77 ). Lower-right conditions are visually acceptable but behaviorally unsuccessful; upper-left conditions show the reverse. (b) Pristine and misleading-quality examples. Prevalence uses content–condition cells with these thresholds. A scatter plot compares mean visual-quality scores with action-recognition accuracy across distortion conditions. Dashed lines divide the plot at the grand-mean MOS of 3.29 and accuracy of 0.77. Several conditions fall in the lower-right misleading-quality region, where visual quality remains relatively high while recognition accuracy is low. Example pristine and misleading avatar renderings are shown beside the plot.
Evaluation
PLCC
95% CI
SRCC
95% CI
KRCC
95% CI
Feature-regression baselines
CONTRIQUE
0.2923
[ −0.0193 , 0.5627]
0.2742
[ −0.0117 , 0.5543]
0.1871
[ −0.0076 , 0.4087]
Re-IQA
0.4435
[0.1260, 0.6206]
0.4501
[0.0403, 0.6800]
0.3135
[0.0237, 0.4971]
DreamSim
0.3795
[ −0.1298 , 0.6242]
0.3470
[ −0.1204 , 0.6361]
0.2364
[ −0.0673 , 0.4628]
Behavioral reliability
IQS split-half agreement
0.6381
[0.5562, 0.7147]
0.6444
[0.5658, 0.7182]
0.4662
[0.4026, 0.5312]
Table 1. Leave-one-content-out feature-regression performance against IQS and split-half IQS agreement at λ=0.5 . Confidence intervals are 95% content-block bootstrap intervals. The table reports PLCC, SRCC, and KRCC correlations with 95 percent confidence intervals for CONTRIQUE, Re-IQA, DreamSim, and split-half IQS agreement. Re-IQA has the highest correlations among the three feature-regression baselines, while split-half IQS agreement is higher than all evaluated baselines.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
ID
Action
Description
E1
Waving
Greeting gesture with raised hand and outward-facing palm
E2
Thumbs-up
Positive acknowledgment with the thumb extended upward
E3
Angry yell
Aggressive expression with open mouth and tense posture
E4
Angry point
Accusatory gesture with an extended arm and finger
E5
Crazy gesture ∗
Erratic arm movements
E6
Begging
Pleading posture with clasped hands and lowered stance
Appendix
Table 2. Avatar–action content units used in the study. Ten avatar–action content units and their action descriptions. E5 was excluded from analysis because its pristine-condition consensus ratio was below the required threshold.
Family
Metric
PLCC
95% CI
FR-IQA
PSNR
0.2752
[0.090, 0.452]
FSIM
0.2429
[0.106, 0.412]
SSIM
0.2392
[0.123, 0.384]
VIF
0.2231
[0.142, 0.306]
LPIPS
0.0863
[ −0.017 , 0.219]
NR-IQA
NIQE
0.2491
[ −0.085 , 0.434]
Appendix
Table 3. Direct-scoring metric correlations with IQS by metric family. Brackets denote 95% content-block bootstrap confidence intervals; rows are ordered by PLCC within each family. Direct-scoring metric correlations with IQS, grouped into FR-IQA, NR-IQA, FR-VQA, and NR-VQA families. The table reports PLCC values and 95 percent content-block bootstrap confidence intervals for each metric.
Family
Condition
Generation setting
Geometry
GQ_LIGHT, GQ_MED, GQ_HEAVY
Vertex-position quantization with δ∈{0.004,0.008,0.016}
Geometry
DS_LIGHT, DS_MED
Mesh downsampling to 50% and 25% of the original geometry
Photometric
RD_720P, RD_360P
Spatial-resolution reduction to 720p and 360p
Temporal
LFR_15, LFR_10
Frame-rate reduction to 15 and 10 fps
Combined
MLBC_1, MLBC_2, HLBC_1, HLBC_2, SLBC_1
Fixed multi-operation presets used in the stimulus set
Appendix
Table 4. Summary of the controlled distortion families and condition labels. Combined conditions were implemented as fixed presets of geometric, photometric, and/or temporal operations. Controlled distortion conditions grouped into geometric, photometric, temporal, and combined families, with the rendering operation used for each condition.