ARGUS: Role-Aware Event Knowledge Graphs for U.S. Employment-Discrimination Complaints
Organizations: University of Massachusetts Amherst · Bloomberg
Abstract
U.S. employment-discrimination complaints describe complex event sequences that are not explicitly captured by lexical or embedding-based representations alone. We present ARGUS, a source-grounded pipeline that combines a 5W1H-inspired schema, legal-domain models, and LLM-based structured generation to construct document-level Event Knowledge Graphs (EKGs) from CourtListener complaints. ARGUS extracts fact-bearing statements, builds chunk-level event graphs with participant, temporal, and causal structure, and merges them into document-level representations. We evaluate graph quality through human and multi-model assessment and test downstream utility on claim classification and legal QA. The graph-structured classifier outperforms raw and linearized baselines on the held-out set, and EKG-only retrieval improves document-scoped QA, while open-retrieval gains remain limited by low first-stage candidate recall. These results suggest that EKGs are most useful for organizing and reasoning over evidence once relevant material has been retrieved.
Figures & tables
| District | Docs | Chunks | Avg/Doc | Sentences | Fact Rate |
|---|---|---|---|---|---|
| EAPD | 20 | 87 | 4.35 | – | – |
| KYED | 31 | 138 | 4.45 | 1,155 | 81.7% |
| NYSD | 46 | 318 | 6.91 | 4,808 | 92.7% |
| PAWD | 28 | 165 | 5.89 | 1,607 | 93.0% |
| RID | 28 | 118 | 4.21 | 1,631 | 90.9% |
| Model | Acc. | Prec. | Rec. | F1 |
|---|---|---|---|---|
| LEGAL-BERT | 0.9400 | 0.8889 | 0.6154 | 0.7273 |
| LexLM | 0.9300 | 0.8000 | 0.6154 | 0.6957 |
| GPT-5 | 0.8500 | 0.4615 | 0.9231 | 0.6154 |
| Metric | GPT-OSS | Qwen |
|---|---|---|
| Quality Score | 0.8272 | 0.8852 |
| Judge Score | 0.6642 | 0.7966 |
| Structural Score | 0.9902 | 0.9739 |
| Temporal Consistency | 0.9118 | 0.7647 |
| Events/Chunk | 7.09 | 7.50 |
| Temporal Edges/Chunk | 2.15 | 3.12 |
| Metric | Claude | Qwen |
|---|---|---|
| Aligned | 78.30% | 85.20% |
| Partially Aligned | 17.40% | 14.80% |
| Misaligned | 4.30% | 0.00% |
| Percent Agreement | 69.60% | 91.70% |
| Fleiss’ | 0.143 | 0.669 |
| Krippendorff’s | 0.153 | 0.673 |
| Case | Pre-Ev | Post-Ev | Loss | Pre-En | Post-En |
|---|---|---|---|---|---|
| Miczulski v. Alix | 27 | 26 | 1 | 31 | 13 |
| Poole v. Ampler | 40 | 40 | 0 | 34 | 17 |
| Sermarini v. Step Up | 35 | 35 | 0 | 26 | 10 |
| Bleiler v. Chester | 27 | 27 | 0 | 24 | 13 |
| Bacon v. M.A.G. | 5 | 5 | 0 | 6 | 6 |
| Case | Events | Temporal | Causal |
|---|---|---|---|
| Miczulski v. Alix Inc | 21 | 19 | 6 |
| Poole v. Ampler Pizza | 0 | 0 | 0 |
| Sermarini v. A Step Up | 16 | 14 | 20 |
| Bleiler v. Chester Co. | 0 | 0 | 0 |
| Bacon v. M.A.G. Ent. | 17 | 16 | 13 |
| Dimension | Win | Agree. | AC1 | |
|---|---|---|---|---|
| Event Grounding | 93.3% | 86.7% | 0.848 | 0.031 |
| Evidence Grounding | 100% | 100% | 1.0 | 0.031 |
| Edge Correctness | 100% | 100% | 1.0 | 0.031 |
| Overall | 100% | 100% | 1.0 | 0.031 |
| Representation | Micro-F1@0.5 |
|---|---|
| Raw-text LEGAL-BERT | 0.5897 |
| Linearized-EKG LEGAL-BERT | 0.5833 |
| EKG-GNN | 0.6415 |
| Context condition | Mean Token-F1 |
|---|---|
| Raw document | 0.323 |
| Hybrid text + EKG | 0.363 |
| EKG only | 0.446 |
| Retrieval method | Recall@10 |
|---|---|
| BM25 | 0.0100 |
| E5-large-v2 | 0.0100 |
| Hybrid Graph-RAG | 0.0800 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Stability | Flip | AC1 |
|---|---|---|---|
| DeepSeek | 80% | 20% | 1.0 |
| Qwen | 100% | 0% | 1.0 |
| Mistral | 100% | 0% | 1.0 |
| Mode | LEGAL-BERT | LexLM |
|---|---|---|
| Classifier | 0.8718 | 0.8616 |
| Hybrid | 0.8711 | 0.8627 |
| LLM | 0.7115 | 0.7186 |
| Feature set | Silhouette |
|---|---|
| TF-IDF baseline | 0.091 |
| EKG full (2k dims) | 0.363 |
| EKG Legacy7D | 0.465 |
| EKG Core6D ( ) | 0.649 |
| Run | Representation | Sizes ( ) | Pearson | Spearman |
|---|---|---|---|---|
| A | SALI(23) + Core6D | 111/14/6/2 | – | |
| B | Core6D + Hybrid PCA(16) | balanced, 30s–50s | ||
| C | A + Hybrid PCA + NN(12) | 31/61/19/22 |
| Feature | A | B | C |
|---|---|---|---|
| SALI claim tags | yes | no | yes |
| Core6D structure | yes | yes | yes |
| Hybrid PCA context | no | yes | yes |
| Graph neighbors ( NN) | no | no | yes |
| Question type | Raw F1 | EKG F1 |
|---|---|---|
| Temporal order | 0.381 | 0.538 |
| Causal/enablement | 0.387 | 0.522 |
| Temporal overlap | 0.241 | 0.384 |
| Party resolution | 0.152 | 0.197 |
| Comparison | F1 | Bootstrap 95% CI | Wilcoxon |
|---|---|---|---|
| EKG Raw | +0.123 | [+0.109, +0.137] | |
| EKG Hybrid | +0.083 | [+0.069, +0.096] | |
| Hybrid Raw | +0.040 | [+0.030, +0.052] |
| Question type | F1 (EKG Raw) | Wilcoxon | |
|---|---|---|---|
| Temporal order | 119 | +0.157 | |
| Causal/enablement | 99 | +0.134 | |
| Temporal overlap | 12 | +0.143 | |
| Party resolution | 66 | +0.045 | |
| Temporal relation | 3 | +0.019 | 0.25 |
| Retrieval condition | MC accuracy |
|---|---|
| Raw document (top 3) | 0.8547 |
| Chunked text (top 3) | 0.8632 |
| EKG graph (top 3) | 0.8632 |
| Gold EKG subset ( ) | 0.8600 |
| Comparison | McNemar | Bootstrap 95% CI | Wilcoxon |
|---|---|---|---|
| Document vs. chunk | 1.000 | pp | 0.813 |
| Document vs. EKG | 1.000 | pp | 0.820 |
| Chunk vs. EKG | 0.688 | pp | 1.000 |
| Context condition | MCQ accuracy |
|---|---|
| No passage | 0.7600 |
| BM25-RAG (open retrieval, top 10) | 0.7300 |
| Hybrid Graph-RAG (open retrieval, top 10) | 0.7200 |
| Gold passage | 0.7900 |
| Original oracle-document Graph-RAG | 0.7500 |
| Improved oracle-document Graph-RAG | 0.8200 |