This paper proposes a novel crowd counting approach, the Gaussian Density Splatting Network (GDSNet). Unlike methods that rely on conventional, grid-based density maps and are sensitive to spatial resolution, GDSNet represents a crowd as a superposition of continuous 2D Gaussian primitives. Our approach is built upon two key contributions. First, we introduce a control-point-based fitting mechanism to structure the prediction of the Gaussian parameters. We design a method to allocate a set of control points that define local regions, from which features are pooled to regress each primitive's parameters. Second, we adapt a differentiable Gaussian Splatting framework to the counting task by parameterizing each primitive with geometric parameters and a scalar density mass. This formulation allows the network to be trained end-to-end via spatial matching of differentiably rendered density maps, naturally providing both local density supervision and global count optimization. Extensive evaluations on four standard benchmarks show GDSNet consistently outperforms the state of the art.
Figures & tables
Figure 1: Comparison of grid-based, point-based, and our continuous Gaussian representation.
Paradigm
Magnitude ( αi )
Basis Function ( Ψi )
Spatial Support ( ξi )
State
Representation Formula D^(x)
Grid-based
Scalar density ( d^u )
Piecewise box ( ΠΩu )
Fixed grid cell ( s )
Discrete
∑u∈Gsd^uΠΩu(x)
Point-based
Binary indicator ( I )
Dirac delta ( δ )
Zero-volume point ( μi )
Discrete
∑iI(pi>η)δ(x−μi)
GDSNet
Density mass ( ρi )
2D Gaussian ( N )
Explicit geometry ( μi,Σi )
Continuous
∑iρiN(x;μi,Σi)
Table 1: Comparison of crowd counting paradigms from a basis-function perspective.
Figure 2: Pipeline of GDSNet.
Dataset
Venue & Year
ShTech A
ShTech B
JHU-Crowd++
UCF-QNRF
NWPU
Method
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
BL [ 15 ]
ICCV 2019
62.8
101.8
7.7
12.7
75.0
299.9
88.7
154.8
105.4
454.2
DMCount [ 16 ]
NeurIPS 2020
59.7
95.7
7.4
11.8
61.6
256.1
85.6
148.3
88.4
388.6
UOT [ 17 ]
AAAI 2021
58.1
95.9
6.5
10.2
60.5
252.7
83.3
142.3
87.8
387.5
GL [ 42 ]
CVPR 2021
61.3
95.4
7.3
11.7
59.5
259.5
84.3
147.5
79.3
346.1
P2PNet [ 6 ]
ICCV 2021
52.7
85.1
6.3
9.9
–
–
85.3
154.5
72.6
331.6
Table 2: Counting performance on ShTech A and B , JHU-Crowd++, UCF-QNRF, and NWPU.
Figure 3: Visualization of the structural prior, simplex heatmap, and rasterized density map. The simplex heatmap is obtained by assigning each Delaunay simplex the mass predicted by its associated Gaussian primitive, revealing how responses are spatially distributed over the triangulated support.
Table 6
Configuration
MAE
MSE
Triangle mass regression
82.21
133.85
Fixed Gaussian
81.42
133.24
Isotropic Gaussian
78.79
132.86
Full anisotropic Gaussian
75.86
130.52
Table 6: Gaussian geometry ablation on UCF-QNRF.
Figure 4: Simplex heatmap (top) and rasterized density map (bottom) of different CPA priors.
Figure 5: Performance w.r.t. train and test sampling strides.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Params (M)
GPU memory (MB)
Runtime (ms)
Baseline: backbone + regression head
21.50
130.50
12.54
Backbone
20.02
48.64
11.02
+ Topology computation
+0
+0.65
+9.00
+ Feature aggregation
+1.66
+11.04
+0.29
+ Constrained Gaussian parameterization
+0.04
+1.70
+0.20
+ Gaussian field rasterization
+0
+0.21
+0.11
Appendix
Table 7: Component-wise parameter, GPU memory, and runtime measurements.
Stride
Gaussian
Runtime (ms)
GPU memory (MB)
Topology time (ms)
Topology memory (MB)
6
1814
16.93
48.95
14.37
1.06
8
1002
13.65
48.93
11.15
0.56
10
639
12.27
48.92
9.72
0.55
12
451
10.94
48.92
8.42
0.55
16
244
10.07
48.91
7.59
0.51
Appendix
Table 8: Runtime and GPU memory measurements with varying Gaussian primitive allocation.
Method
CPA Strategy
Sampling
Support Region
SCP
TFA
Module Ablations
Full GDSNet
Structural Prior
Dither
Triangle Simplex
✓
✓
w/o TFA
Structural Prior
Dither
Triangle Simplex
✓
–
w/o (TFA & SCP)
Structural Prior
Dither
Fixed grid patch
–
–
w/o (TFA & SCP & CPA )
Structural Prior
Uniform grid
Fixed grid patch
–
–
CPA Ablation
Uniform Grid
Uniform
Dither
Triangle Simplex
✓
✓
Sobel Filter Response Prior
Sobel Filter Response Prior
Dither
Triangle Simplex
✓
✓
Appendix
Table 9: Implementation comparison of the full model and ablation variants. CPA, SCP, and TFA are the control point allocation, simplex-constrained parameterization, and topology-aware feature aggregation modules, respectively.
Figure 6: Comparison of Sobel filter response prior (top) and our structural prior (bottom).
Model / source
Setting
MAE
Molmo [ 46 ]
Direct prompting
7.26×108
Gemini-2.5 [ 46 ]
Direct prompting
517.17
LLaVA-OneVision-7B [ 47 ]
Direct prompting
336.80
LLaVA-OneVision-7B / WS-COC [ 47 ]
Counting-specific tuning
128.90
GDSNet
Specialized density estimation
48.80
Appendix
Table 10: Reported MLLM and GDSNet results on ShanghaiTech Part A.
Figure 7: Representative failure case of GDSNet on a highly congested scene. The model produces a smooth density map which lacks sharply localized peaks for individual heads.
Crowd counting must recover reliable local density under severe variations in perspective, head scale, occlusion, and background clutter. Although modern counting objectives provide strong spatial supervision, many multi-level decoders still use spatially invariant feature fusion and apply one receptive-field pattern to every location. We propose DCA-MoE, a framework that makes both decisions content dependent while retaining a frozen DINOv3 encoder. Spatially Adaptive Layer Fusion (SALF) predicts position-wise weights over four aligned backbone features, and Density-Routed Multi-Receptive-Field Experts (DR-MoE) assigns each location a soft mixture of local, mid-range, and large-context residual experts. An EBC-style head reconstructs block density, while DMCount supervision and an auxiliary routing-balance term train the decoder without updating the backbone. On the NWPU-Crowd validation split, the strongest paired configuration, based on DINOv3 ViT-L/16, obtains 31.7 MAE and 72.2 RMSE; the matched ViT-B/16 full model obtains a paired 32.2/75.9. Cross-dataset results remain mixed, and several component baselines currently report independently selected minima from a single seed. The evidence therefore supports the feasibility of spatially adaptive fusion and routing, while broader paired and multi-seed evaluation remains necessary for causal attribution.
While 3D Gaussian Splatting (3DGS) has demonstrated impressive real-time rendering performance, its efficacy remains constrained by a reliance on heuristic density control. Despite numerous refinements to these handcrafted rules, such methods inherently lack the flexibility to adapt to diverse scenes with complex geometries. In this paper, we propose a paradigm shift for density control from rigid heuristics to fully learnable policies. Specifically, we introduce \textbf{LeGS}, a framework that reformulates density control as a parameterized policy network optimized via Reinforcement Learning (RL). Central to our approach is the tailored effective reward function grounded in sensitivity analysis, which precisely quantifies the marginal contribution of individual Gaussians to reconstruction quality. To maintain computational tractability, we derive a closed-form solution that reduces the complexity of reward calculation from O(N2) to O(N). Extensive experiments on the Mip-NeRF 360, Tanks & Temples, and Deep Blending datasets demonstrate that \textbf{LeGS} significantly outperforms state-of-the-art methods, striking a superior balance between reconstruction quality and efficiency. The code will be released at https://github.com/AaronNZH/LeGS
Zhenhua Ning, Xin Li, Jun Yu +3
Pengcheng Laboratory, Shenzhen · Harbin Institute of Technology, Shenzhen
RGB-Thermal (T) crowd counting aims to integrate visible-spectrum and thermal infrared information to improve the robustness of crowd density estimation in complex scenes. Although existing studies generally improve counting accuracy through cross-modal feature fusion, most current methods rely on implicit cross-modal fusion strategies and lack explicit modeling of local spatial discrepancies as well as fine-grained characterization of modality reliability at the positional level, thereby limiting the accuracy and interpretability of the fusion process. To address these issues, this paper proposes a two-stage fusion framework, RACANet, a Reliability-Aware Crowd Anchor Network for RGB-T crowd counting. First, we introduce a lightweight cross-modal alignment pretraining stage, which explicitly learns cross-modal semantic correspondences through crowd-prior supervision and local bidirectional soft matching. Then, based on the priors learned during pretraining, a Local Anchor Fusion Module (LAFM) is introduced in the formal training stage. This module generates local semantic anchors by aggregating features from highly reliable regions and further enables adaptive pixel-level feature redistribution with a local attention mechanism. In addition, we propose a discrepancy-aware consistency constraint to dynamically coordinate the reliability of regions where modal representations are consistent. Experiments conducted on two widely used benchmark datasets, RGBT-CC and Drone-RGBT, demonstrate that RACANet outperforms existing methods. The anonymous code is available at https://anonymous.4open.science/r/RACANet-9985.
Jinghao Shi, Mengqi Lei, Kunliang He +3
School of Computer Science, China University of Geosciences, Wuhan, Wuhan, China