KUPAS MASTER: Distilling the Tacit Expertise of Master Practitioners into Agent-Ready Experience Corpora
Organizations: Shanghai Kupas Technology Co., Ltd. and Tongji University
Abstract
Experienced professionals know more than just facts and conclusions. They know which cues matter, why a judgment is reasonable, and which action to take. Routine work records often leave out this tacit knowledge, making it difficult for Large Language Model (LLM) agents to use professional experience effectively. We introduce KUPAS MASTER, an experience engineering platform built around nine-layer cognitive corpus construction. It turns heterogeneous work records and practitioner interviews into traceable, reusable experience corpora for agents. Six case elements preserve the task process: context, cues, judgment, action, boundaries, and outcomes. Nine-layer cognitive corpus construction organizes tacit experience along nine extraction dimensions and stores the resulting assets in six libraries: rules, constraints, best practices, negative examples, corner cases, and skills. Semantic alignment, individual experience distillation, organizational consolidation, and cross-review preserve source evidence, conditions of use, and unresolved disagreements. The platform packages these assets into callable skills with explicit inputs, steps, dependencies, and stopping conditions, connecting experience collection to task execution and evaluation feedback. Using authorized samples from 20 randomly selected practitioners, the platform processed 1,576 source files into 23,024 individual experience records and 13,113 organizational assets. The evaluation spans multiple professional domains. Under common task inputs and scoring criteria, the base model, raw corpus retrieval-augmented generation (RAG), and KUPAS MASTER agent scored 70.63, 79.75, and 89.58, respectively. The KUPAS MASTER agent improved on raw-corpus RAG in all seven scoring dimensions. The platform provides a practical path from individual tacit experience to organizational knowledge and agent capabilities.
Figures & tables
| Element | Content | Distinctions to retain |
| Context | Task, role, object, environment, time, and available resources. | Conditions observed in this case versus conditions assumed for reuse. |
| Cues | Signals, measurements, statements, and changes noticed during work. | Direct observations versus reported or inferred observations. |
| Judgment | Interpretations, supporting reasons, alternatives, and uncertainty. | Practitioner accounts versus model-generated explanations. |
| Action | Selected operations, order, parameters, and stopping conditions. | Planned, attempted, and completed actions. |
| Boundaries | Prohibitions, scope, missing prerequisites, and referral conditions. | Mandatory constraints versus preferences or usual practice. |
| Outcome | Immediate results, later verification, and unresolved effects. | Expected, reported, and independently verified outcomes. |
| ID | Dimension | Extraction question | Representative objects |
| L1 | Knowledge | Which facts and concepts were used? | Entities, terms, relations, and domain materials. |
| L2 | Attention | Which observations received priority? | Diagnostic cues, ignored distractions, and shifts of focus. |
| L3 | Judgment | What supports this interpretation? | Conditional judgments, thresholds, alternatives, and uncertainty. |
| L4 | Decision | How was an action selected? | Prerequisites, selection criteria, and stopping conditions. |
| L5 | Association | Which other cases or concepts were relevant? | Analogies, recalled precedents, and links across cases. |
| L6 | Anticipation | What consequences were expected? | Predicted effects, time horizons, and follow-up checks. |
| Asset | Reusable content | Common dimensions |
| Rules | Scoped conditions and recommended judgments or actions, including exceptions. | L3, L4 |
| Constraints | Prohibitions or required prerequisites, their basis, and allowed alternatives. | L7, L9 |
| Best practices | Supported successful procedures and conditions for considering reuse. | L1, L4, L5, L6 |
| Negative examples | Inappropriate actions or adverse outcomes, context, and reviewed corrections. | L3, L4, L7, L9 |
| Corner cases | Unusual contexts where a routine rule may fail or need adjustment. | L2, L3, L5, L7 |
| Skills | Callable procedures with inputs, outputs, dependencies, checks, and evidence. | L2, L3, L4, L8, L9 |
| Operator and input | Transformation and output | Reject or leave unresolved when |
| Term alignment: wording and context | Propose a standard term while retaining aliases, units, and source spans. | Multiple interpretations remain plausible, or a unit conversion lacks support. |
| Claim extraction: case narrative | Separate observations, judgments, and planned actions, each with its own source. | The extracted claim lacks supporting content. |
| Condition preservation: conditional statement | Recover triggers, recommendations, prerequisites, and exceptions as a candidate rule. | Compression changes the action, drops an exception, or alters a prerequisite. |
| Conflict classification: comparable asset pair | Check scope overlap and action compatibility to distinguish agreement, different conditions, and conflict. | Scope overlap is unknown or evidence is insufficient to resolve the conflict. |
| Item | A: Base model | B: Raw-corpus RAG | C: KUPAS MASTER |
| Approach | General analysis and an evidence checklist. | Cites work injury insurance regulations and proposes mediation steps. | Calls the mediation assessment skill and combines rules with corner cases. |
| Boundaries | Does not clearly separate mediation from formal injury determination. | Identifies some procedural and risk boundaries. | States that mediation does not replace formal determination and distinguishes accident scenarios and risks. |
| Output | General principles and evidence suggestions. | Legal grounds and process advice. | A structured process and an output file. |
| Comparison | Rated lowest on all 8 items. | Ties with C on 1 item. | Rated best or tied for best on all 8 items. |
| Item | Configuration and scale |
| Sample | 20 randomly selected authorized practitioners, covering finance and accounting, community governance, engineering quality, safety oversight, healthcare, mediation, and emergency management. |
| Tasks | 177 questions, independently answered by A, B, and C, for 531 responses. |
| Difficulty | 8 easy, 83 medium, and 86 hard questions. |
| Configurations | A: base model; B: raw-corpus RAG; C: KUPAS MASTER agent. |
| Scoring | A shared seven-dimension rubric applied to independent answers to the same tasks. |
| Run traces | Individual records for 5 practitioners and 53 questions, covering retrieval, library calls, and skill execution. |
| Method | Composite score | Score pass rate (%) | Evidence sufficiency and accuracy | Boundaries and compliance | Time (s) |
| A: Base model | 70.63 | 87.57 | 62.18 | 74.52 | 16.2 |
| B: Raw-corpus RAG | 79.75 | 98.87 | 77.27 | 79.19 | 34.0 |
| C: KUPAS MASTER | 89.58 | 100.00 | 88.32 | 88.23 | 57.1 |
| Asset type | Runs | Observed role and example tasks |
| Rules | 46 | Express judgment conditions as rules, thresholds, and requirements, including flood dispatch decisions and engineering quality escalation criteria. |
| Best practices | 33 | Supply established procedures and steps that turn experience into a task process. |
| Corner cases | 35 | Add exceptions, high-risk situations, and business boundaries that routine rules may miss. |
| Constraints | 28 | Identify prohibited actions or required prerequisites for procedural and compliance checks. |
| Negative examples | 18 | Provide recurring errors and inappropriate responses to help recognize risky decisions and exceptions. |
| Skills | 46 | Organize retrieved experience into steps, checks, responsible roles, and structured deliverables. |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| ID | Analysis focus | Supporting evidence |
|---|---|---|
| E1 | Overall improvement in task outcomes | Configuration scores, paired questions, resampling intervals, and below-threshold scores in the random authorized sample. |
| E2 | Added value from structuring the same information | Overall gain of the experience-and-skill configuration built from the same sources. |
| E3 | Useful, traceable experience representation | Asset fields, source relationships, automated checks, and separately recorded human review. |
| E4 | Consolidation into organizational assets | Individual-to-organizational asset counts, library versions, and consolidation records. |
| E5 | Skills in task execution | Skill calls, file generation, and task scores. |
| E6 | Boundary handling and reliability | Boundary score gains and associations between asset use and boundary-judgment rows. |