5W1H+Which: Context-Valid Semantic Indexing with Progressive Ontology Binding
Organizations: PwC China AI Center · Tsinghua University
Abstract
Transforming raw data into queryable knowledge requires both early extraction of reusable information and explicit types, relations, and applicability conditions for particular tasks. If indexing selects content too early around a single business schema, later tasks may be unable to use information that was omitted. If the index retains only open-ended text, however, rule-based reasoning lacks checkable premises. We propose 5W1H+Which, a semantic indexing design that separates content extraction from ontology binding. The 5W1H questions organize source-grounded content units; Which points to versioned ontology elements and records mapping relations, scope, and validation status. Time, location, system environment, and participant roles are not merely retrieval labels: together, they constrain the contexts in which facts, bindings, and rules apply. Unbound content remains searchable, while bound content enters a formal reasoning path only after premise checks. The method further distinguishes business valid time, system knowledge time, and operational traces, and uses dependency records to support binding revalidation and the maintenance of derived conclusions. A worked example of migration from an on-premises server to a cloud environment illustrates the different treatment of world-state changes, ontology-version changes, and changes in rule applicability. We formulate three groups of falsifiable hypotheses concerning cross-task evidence coverage, control of contextual misuse, and incremental update cost. The planned evaluation includes a strong typed fact-graph baseline with the same evidence, temporal information, and budget, to test whether benefits arise from 5W1H organization, deferred binding, or additional information and engineering effort. The contribution is a testable indexing mechanism, not a claim to a new universal ontology or a demonstrated performance advantage.
Figures & tables
| Prior work | Relevant mechanism | What our evaluation must isolate |
|---|---|---|
| SocraticKG; XPEventCore [R1, R2] | 5W1H extraction and event modeling | Organization versus extra facts or better extraction prompts |
| OaK; OM4OV [R3, R5] | Dynamic ontology and mapping/version changes | Independent content retention and binding revalidation under equal information |
| DimMem [R12] | Dimensional memory and constrained retrieval | Additional value beyond a dimension-aware fact graph |
| RuleMem; InMind [R13, R14] | Rule reuse and implicit applicability | Candidate discovery versus validated admission; useful association versus scope leakage |
| Recuris [R15] | Experience, working state, and verified updates | Keep skill learning outside the core indexing claim |
| Fortunate Recall [R16] | Typed and generic lifecycle policies | Category labels versus metadata, routing, and enforced policies |
| Direction | Content-layer representation | Meanings that must remain distinct |
|---|---|---|
| What | Objects, events, relations, or claims with exact payloads | A document topic is not a specific claim that can be assessed as true or false. |
| Who | References to participants with explicit roles | Author, operator, responsible party, approver, and current user are different roles. |
| When | Instants, intervals, precision, time zone, and the basis for resolving relative times | Event occurrence, planned validity, ingestion, and publication are different times. |
| Where | Typed physical, organizational, and computational scopes | A server room, department, cluster, and location within a source are not one kind of place. |
| Why | Source-stated purposes and reasons, or separately recorded causal hypotheses | A stated reason is not a demonstrated causal mechanism. |
| How | Methods, tools, procedures, parameters, and execution references | A planned method is not necessarily the method actually executed. |
| Input | Source and system-record time | Business meaning | Index treatment |
|---|---|---|---|
| A | Verification record, 2026-09-01 | S runs on on-premises device L on that date. | Record the evidenced state for that date; continued validity requires a separately declared persistence policy. |
| B | Planning record, 2026-09-10 | S is scheduled to migrate to the cloud on 2027-01-01. | Mark as a plan; do not change actual deployment facts in advance. |
| C | Completion record ingested only on 2027-01-05 | Migration is verified as effective on 2027-01-03; the former production instance is retired. | Preserve both actual effective time and late-arriving knowledge time. |
| D | Responsibility rule verified separately from C | After migration, team T manages the application layer; infrastructure responsibility follows the contract. | Do not infer all responsibilities automatically from the CloudDeployment type. |
| Test | Hold fixed | Vary | Primary outcome |
|---|---|---|---|
| Extraction | Sources, model, token and storage budgets | 5W1H prompt versus open atomic-fact prompt | Evidence coverage and qualifier retention |
| Binding | Same extracted content and mapper | Early schema filtering versus retained unbound content | New-task evidence coverage at matched precision |
| Categories | Same facts, lifecycle metadata, gates and routing | Named categories versus a flat typed representation | Incremental calibration and coverage; report equivalence margins if tested |
| Admission | Same facts, time/scope metadata and access controls | Stored metadata versus enforced semantic gates | Contextual misuse at matched answer coverage |
| Applicability | Same evidence and query templates | Positive indirect-use cases versus scope-mismatched negative pairs | Useful activation and false activation |
| Maintenance | Same event stream, semantics and measured cost accounting | Full recomputation versus incremental repair | Withdraw/add/retain agreement and total cost |