SkinAgent AI: A Safety-Grounded Multimodal Agentic Framework for Non-Diagnostic Skincare Support
Organizations: Department of Computer Science, International Islamic University Chittagong, Chittagong, Bangladesh · Centre for Intelligent Cloud Computing, Centre of Excellence for Advanced Cloud, Faculty of Information Science and Technology, Multimedia University, Melaka, Malaysia · Multimedia University, Jalan Ayer Keroh Lama, 75450 Bukit Beruang, Melaka, Malaysia · Department of Computer Science, American International University-Bangladesh, Dhaka, Bangladesh
Abstract
Consumer-facing skincare AI must coordinate visual evidence, product information, tool use, and user-facing actions within explicit evidence and safety boundaries. This study evaluates SkinAgent AI, a non-diagnostic multimodal framework that combines visual concern routing with grounded and auditable LLM-based orchestration. The architecture includes routing for Acne, Pores, and Wrinkles; photograph-based skin-type estimation; count-informed ordinal acne-severity support; typed tools; database-grounded recommendation and action functions; deterministic safety, privacy, and evidence checks; approval before state-changing actions; and structured trace and replay mechanisms. Visual-model performance and system-level agent behavior were evaluated separately. Across three seeds, the skin-condition routing model achieved 99.84% +/- 0.07% accuracy. Skin-type estimation achieved 88.85% accuracy, while count-informed acne-severity support achieved 84.59% accuracy with a quadratic weighted kappa of 0.9076. On a locked but non-independent 240-case system benchmark, intent accuracy was 80.00%, exact tool-set match was 62.92%, and strict task completion was 47.08%. No violations or successful cross-user leakage events were observed in the finite safety and privacy test suites. Tool-selection errors, incomplete grounding of product attributes, and unreliable failure fallback nevertheless remained. These findings support the feasibility of bounded, database-grounded, and traceable agent orchestration for non-diagnostic skincare assistance. They do not establish clinical readiness, external generalization, formal privacy guarantees, or universal safety. Independent validation, expert assessment, robustness and fairness testing, and prospective evaluation in real-world settings remain necessary.
Figures & tables
| Model | Test set | Metric | Three-seed result |
|---|---|---|---|
| EfficientNetV2-S | 810 images | Accuracy | |
| EfficientNetV2-S | 810 images | Macro-F1 | |
| EfficientNetV2-S | 810 images | Macro-AUROC | |
| EfficientNetV2-S | 810 images | Macro-AUPRC | |
| EfficientNetV2-S | 810 images | ECE | |
| EfficientNetV2-S | 810 images | Acne miss rate |
| Subsystem | Model/task | Accuracy | Macro-F1 | Additional metric | |
|---|---|---|---|---|---|
| Skin type | EfficientNetV2-M | 592 | 88.85% | 0.8924 | Accuracy 95% CI: 86.06%–91.14%; F1 95% CI: 0.8660–0.9168 |
| Acne severity | CIOS module | 292 | 84.59% | 0.8149 | QWK = 0.9076 |
| Evaluation | Metric | Result | 95% interval/note | |
| Locked benchmark | 240 | Intent accuracy | 80.00% (192/240) | 74.48%–84.57% |
| Locked benchmark | 240 | Exact tool-set match | 62.92% (151/240) | 56.65%–68.78% |
| Locked benchmark | 240 | Mean required-tool recall | 66.25% | Mean per-case recall |
| Locked benchmark | 240 | Strict task completion | 47.08% (113/240) | 40.86%–53.39% |
| Locked benchmark | 240 | Forbidden-tool violations | 3.33% (8/240) | 1.70%–6.44% |
| Development suite | 189 | Required-tool recall | 55.88% (57/102) | Internal scripted evidence |
| Paired subset | Configuration C | Configuration D | Exact McNemar value | |
| Strict task completion | 60 | 50/60 (83.33%) | 17/60 (28.33%) | |
| Tool-required cases | 16 | 6/16 (37.50%) | 14/16 (87.50%) | |
| No-tool cases | 44 | 41/44 (93.18%) | 32/44 (72.73%) |
| Evaluation | Metric | Observed result | Statistical support/note |
|---|---|---|---|
| Safety red-team | Unsafe, diagnostic, prescription, or dosage violation | 0/45 | One-sided 95% upper bound: 6.44% |
| Privacy red-team | Successful cross-user leakage | 0/15 | One-sided 95% upper bound: 18.10% |
| Approval gate | Gate violation | 0/7 | One-sided 95% upper bound: 34.82% |
| Product grounding | Existing product with valid ID/slug | 26/26 | 95% CI: 87.13%–100.00% |
| Product grounding | Ingredient/rating fidelity | 14/26 each | 53.85% each; 95% CI: 35.46%–71.24% |
| Product grounding | Stock grounding | 0/26 | Stock unsupported by the audited schema |