cs.CYJul 31, 2026

From Prediction to Explainable Provider Behavior Profiles for Fraud, Waste, and Abuse Review

Authors: Yubin Park, Evan Brociner

Organizations: Falcon Health, Inc.

Abstract

Claims data can show that provider behavior changed but cannot by itself explain why. Fraud, waste, and abuse (FWA) review requires identifying material behavior, locating the codes and dollars driving it, and testing plausible explanations. Forecast residuals conflate growth, service-line shifts, code maintenance, and incomplete observation with potentially concerning behavior. We instead formulate provider review as a descriptive representation problem: observed amount yijt=sitpijty_{ijt}=s_{it}p_{ijt}, where sits_{it} is provider scale and pijtp_{ijt} is procedure composition. The profile records scale history, effective-dated code lineage, clinical-family shares, first-use events, billing context, and Medicare-versus-client differences. An optional rank-32 nonnegative factorization of procedure co-occurrence adds a fixed semantic geometry for similarity and retrieval. The profile surfaces evidence for review without inferring intent or adjudicating FWA, and is one engine within Falcon's broader review system. In a ten-quarter proprietary Medicare Carrier and DME build (1.27 million providers; 9.2 million provider-quarter profiles through 2026~Q2), the semantic dictionary covers 3,641 procedures and raises recall at 10 from 36.0% to 44.4%; for high-cost rare events, recall at 50 is 56.1% versus zero for popularity. Under the governed eligibility contract, 414,093 providers enter the national review population, with 81.8% and 90.8% remaining eligible across adjacent quarters. Replication across eight client panels preserves 86.9--95.2% amount-weighted semantic coverage. Transparent descriptions thus form the core, with learned representations adding optional semantic context.

Figures & tables

Explore similar work

Jul 21, 2026cs.LG

Toward Auditable Fraud Detection: Combining Graph Features, Model Explanations, and Agentic Case Investigation

Fraud detection systems must scale with rising transaction volume while remaining explainable and reviewable. We study a layered pipeline on the PaySim dataset that combines a gradient-boosted classifier, graph-derived structural features, an autoencoder-based anomaly signal, TreeSHAP explanations, and a bounded LLM investigation agent applied to cases the classifier scores uncertainly. Before any model comparison, we identify and remove a simulator-specific balance shortcut that would otherwise inflate baseline performance. After this correction, neither the graph features nor the anomaly signal improves Average Precision on the full test set. Both, however, rank fraud better within the subset of cases receiving intermediate baseline scores. In a controlled experiment with injected multi-account fraud rings, engineered structural features recover all injected test transactions, while the tabular baseline misses roughly a quarter of them. The investigation agent underperforms direct thresholding of the classifier it relies on, reaching 65.0% accuracy against 71.7% on a balanced 60-case sample, despite having access to model explanations, graph context, and retrieved reference cases. Of the eight decisions the agent changed, six replaced correct classifier outputs with errors, and it produced a coherent written rationale in each case. An exploratory disagreement-based escalation rule flagged two of these agent errors for human review without flagging any correct decision. We conclude that each component of a layered fraud system contributes only under specific conditions, and that a plausible rationale from an investigation agent is not evidence of a better decision.
Jul 29, 2026cs.AI

INCLAIR: Inception-Based Longitudinal Clinical Anomaly Detection with Informed Reasoning

Detecting anomalies in longitudinal clinical profiles is clinically important but difficult: abnormal evidence is often sparse, patient histories have unequal length, and expert explanations are costly. We propose INCLAIR, a framework that scores each observation against multiple historical contexts, aggregates evidence at the profile level, and generates grounded natural-language explanations under limited expert supervision. Under stated within-profile exchangeability assumptions, the complete mean subsequence score takes an order-ll U-statistic form, yielding a variance decomposition and an incomplete-subset approximation that controls combinatorial inference cost independently of profile length. The same analysis shows that mean aggregation attenuates localized anomalies by a factor set by the anomaly support and profile length, motivating validation-selected top-kk pooling. Across three clinical datasets, INCLAIR consistently outperforms state-of-the-art baselines. We further validate practical relevance through a case study on longitudinal steroid profiles, comparing INCLAIR's predictions and explanations against domain-expert assessments supported by DNA analysis. The results show that INCLAIR enables clinically actionable anomaly detection under limited expert supervision.
Sep 16, 2026cs.CY

Understanding AI Provider Recommendations in Local Service Markets

When someone asks an AI assistant which doctor to see or which firm to trust with their savings, the answer is a referral. We audit AI provider recommendations in four registry-backed service domains across the 100 largest U.S. metropolitan areas, matching every recommendation against the official registry for its domain (Medicare clinician and facility records, and SEC adviser disclosures), under three conditions: an open-weight model, a proprietary model without web search, and the same proprietary model with search. Without search, both models largely fabricate recommendations in the domains the web covers thinly. Only 4% of the open-weight model's recommended doctors and 11% of the proprietary model's match a clinician in the queried city, and the open-weight matches are name coincidences: its matched clinicians are no likelier to be primary-care doctors than names drawn at random from the registry. With search, 64-71% of recommendations in the same domains match a real provider. Search also changes who is recommended. Without it, recommended advisory firms carry SEC misconduct disclosures at 3.6 times the registry base rate, even after adjusting for firm size; with search, significantly below it. Restaurants, where quality and visibility are separately measurable, show a 3-5x review-count premium but a rating premium of at most a tenth of a star. Finally, search largely removes the metro-size penalty: without it, real recommendations concentrate in the largest metros; with it, match rates are similar across metro-size terciles. Whether an AI referral is trustworthy depends strongly on its retrieval configuration rather than on the underlying model alone, yet an answer produced without retrieval often carries no sign that its recommendations were never verified.