Constructing Structured Decision Sources for Consensus-Based Pseudo-Label Learning
Abstract
Consensus can make pseudo-label learning more reliable, but only when its predictors contribute genuinely different evidence. Multiple models that repeat the same boundary provide additional votes without additional information. We address this problem by con structing decision sources through controlled changes to within-class structure. Starting from a shared graph representation, we vary center granularity and neighborhood mixing, reproduce each resulting source to test its stability, and select a complementary subset using node pair coassignment. Unanimous predictions from the selected sources are then ranked for student training. On the public fixed splits of Cora, CiteSeer, and PubMed, evaluated with five random seeds, the constructed sources improve fixed-budget training pseudo-label precision by 1.19 to 4.39 percentage points over three conventionally initialized GCN sources. Under matched structural filters, three-source consensus is more precise than each constituent source in all 45 dataset slot eed comparisons. The gains are strongest in pseudo-label quality: downstream accuracy remains competitive but does not lead on every dataset. These results identify source construction rather than model count alone as an important design problem for consensus-based pseudo-label learning.
Figures & tables
| Dataset | Nodes | Edge entries | Features | Classes | Train/val/test | |
|---|---|---|---|---|---|---|
| Cora | 2708 | 10556 | 1433 | 7 | 140/500/1000 | 2568 |
| CiteSeer | 3327 | 9104 | 3703 | 6 | 120/500/1000 | 3207 |
| PubMed | 19717 | 88648 | 500 | 3 | 60/500/1000 | 19657 |
| Dataset | Voting rule | Before | Before prec. | After | After prec. | After cov. | Removed correct/incorrect |
|---|---|---|---|---|---|---|---|
| Cora | Structured sources 3/3 | 2473.6 | 84.27 | 1452.4 | 92.53 | 56.56 | 741.0/280.2 |
| DiPat 3/3 + structural gate | 2437.4 | 83.23 | 1416.4 | 93.47 | 55.16 | 705.0/316.0 | |
| Joint 4/6 | 2388.4 | 85.36 | 1431.8 | 93.38 | 55.76 | 702.2/254.4 | |
| Joint 5/6 | 2298.8 | 86.87 | 1415.4 | 93.86 | 55.12 | 668.8/214.6 | |
| Joint 6/6 | 2208.8 | 88.21 | 1396.8 | 94.19 | 54.39 | 632.8/179.2 | |
| CiteSeer | Structured sources 3/3 | 3081.2 | 73.40 | 1322.0 | 85.86 | 41.22 | 1126.8/632.4 |
| Dataset | Method | Training pseudo-labels | Test Accuracy | Test Macro-F1 |
|---|---|---|---|---|
| Cora | Supervised GCN | 0 | ||
| Graph-only DiPat adaptation | 1734 | |||
| Structured sources | 1734 | |||
| CiteSeer | Supervised GCN | 0 | ||
| Graph-only DiPat adaptation | 2165 | |||
| Structured sources | 2165 |
| Dataset | Decision sources | All-vote prec. | q90 prec./cov. | Training PL prec. | Test Acc. | Test Macro-F1 |
|---|---|---|---|---|---|---|
| Cora | Structured sources | 84.27 | 92.53 /56.56 | 93.04 | 84.04 | 83.03 |
| Three random-seed GCNs | 84.80 | 91.94/57.06 | 91.30 | 83.30 | 81.95 | |
| CiteSeer | Structured sources | 73.40 | 85.86 /41.22 | 83.71 | 73.74 | 68.29 |
| Three random-seed GCNs | 71.55 | 84.63/42.03 | 79.33 | 73.08 | 69.21 | |
| PubMed | Structured sources | 79.36 | 84.35 /71.75 | 88.18 | 79.94 | 79.53 |
| Three random-seed GCNs | 79.96 | 84.06/71.10 | 87.00 | 79.82 | 79.35 |
| Dataset | Selected slot 0 | Selected slot 1 | Selected slot 2 |
|---|---|---|---|
| Cora | +0.61 (5/5) | +0.80 (5/5) | +0.42 (5/5) |
| CiteSeer | +0.66 (5/5) | +0.40 (5/5) | +0.38 (5/5) |
| PubMed | +0.34 (5/5) | +0.25 (5/5) | +0.32 (5/5) |
| Dataset | Mean reproduction stability | Mean co-assignment distance | Distance range |
|---|---|---|---|
| Cora | 0.833 | 0.222 | 0.140–0.303 |
| CiteSeer | 0.814 | 0.263 | 0.128–0.412 |
| PubMed | 0.918 | 0.234 | 0.088–0.329 |