cs.AIMar 29, 2026

PeopleSearchBench: Evaluating AI-Powered People Search Platforms with Criteria-Grounded Verification

Authors: Tianyu ShiWei WangZequn XieShuai ZhangBoyang XiaChenyu ZengQi ZhangLynn Ai+5 more

Abstract

AI-powered people search platforms are increasingly deployed for recruiting, sales prospecting, and professional networking, yet no standardized benchmark exists for their rigorous evaluation. We present PeopleSearchBench, an open-source benchmark comprising 119 multilingual queries across four scenarios: corporate recruiting, B2B sales prospecting, expert search, and influencer discovery. A central contribution is Criteria-Grounded Verification, an evaluation methodology that decomposes each query into explicit, independently checkable criteria and verifies each returned individual via live web search, producing factual relevance judgments rather than subjective LLM-as-judge scores (Cohen's kappa = 0.84 with human annotators). We evaluate four architecturally diverse platforms along three complementary dimensions---Relevance Precision, Effective Coverage, and Information Utility---and find that multi-source search agents significantly outperform single-domain systems, particularly in influencer discovery where the performance gap is largest. Platform rankings are robust across ablations on scoring thresholds, dimension weights, and judge models. All code, queries, and evaluation prompts are publicly available.

Explore similar work

CardsList
  1. InfluMatch: Frontier-Quality KOL Search at 4B-Model Cost

    Jul 7, 2026Krittanon Kaewtawee, Petmongkon Pornpichitsuwan, Natchaya Temyingyong +4Social InfluenceFunnel Efficiency