Traditional energy forecasting solutions rely on task-specific supervision and energy asset representations, limiting transferability and the ability to capture general temporal dynamics across heterogeneous assets. We address this by proposing a distributed Joint Embedding Predictive Architecture (JEPA) for self-supervised learning from heterogeneous energy time-series. The framework predicts latent representations of masked temporal segments while integrating temporal observations and contextual information within a shared embedding space. To prevent representation collapse, training combines a latent-space predictive objective with covariance and temporal variance regularization. The evaluation was conducted on energy consumption and generation datasets under data-degradation scenarios and compared with a Transformer forecasting baseline. The learned representations remained stable (cosine similarity
≈0.98; effective rank 185-235). JEPA achieved performance comparable to a Transformer on building energy data, higher
R2 in 3/5 consumer clusters, and outperformed the baseline on 9/10 unseen PVs (
R2=0.73-0.88 vs. <0.45), while showing greater robustness to missing data.