From the Drosophila Visual Connectome to General-Purpose Computer Vision
Authors: Zongyu Li, Akito Yamauchi, Huaizhi Liu, Vishwanatha Rao, Jia Guo, for the Frontotemporal Lobar Degeneration Neuroimaging Initiative, for the Alzheimer's Disease Neuroimaging Initiative
Organizations: Department of Biomedical Engineering, Columbia University, New York, NY, USA · Department of Electrical Engineering, Columbia University, New York, NY, USA · Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA · Department of Bioengineering, School of Engineering and Applied Science, University of Pennsylvania, Philadelphia, PA, USA · Department of Psychiatry, Columbia University, New York, NY, USA
Biological connectomes encode structured solutions to visual computation that may provide reusable inductive biases for artificial vision. We develop ConnectomeX around FlyVision, a trainable architecture that preserves parallel ON/OFF processing, recurrent computation and population-level graph interaction while scaling model capacity across tasks. FlyVision reached 99.34% accuracy on MNIST with 80,608 parameters and 78.03% on CIFAR-10 with 81,408 parameters. On ImageNet-1K, FlyVision Base and Large reached 60.79% and 66.25% top-1 accuracy with 1.8 and 3.7 million parameters, while a Large local-k7 model with a learned low-frequency branch reached 66.53%, compared with 69.25% for ResNet18 with 11.7 million parameters. On a 22-class skin-disease benchmark, FlyVision Large achieved 63.78% accuracy and 95.28% macro-AUROC with 2.99 million parameters. In four-class chest radiography, ImageNet-pretrained FlyVision Base and Large reached 92.60% and 92.76% accuracy with 1.33 and 2.97 million parameters, compared with 91.56% for ImageNet-pretrained ResNet18 with 11.18 million. BrainAGE extends FlyVision to volumetric T1-weighted MRI by applying a shared ImageNet-pretrained FlyVision Large encoder to 24 sagittal, coronal and axial slices per scan and combining slice-level age estimates by confidence-modulated Gaussian voting. On 433 held-out scans, three-axis fusion achieved a mean absolute error of 5.98 years and R^2 = 0.868. Across the 224x224 classification tasks, the best FlyVision configuration remained within three percentage points of ResNet18 on ImageNet-1K and skin-disease classification and exceeded it on chest radiography with substantially fewer parameters. These results show that a conserved connectome-informed computation can scale from compact recognition to large-scale natural and biomedical vision.
Figures & tables
Figure 1: Connectome-informed visual computation and cross-domain scaling. (A) Population-level schematic of the Drosophila visual-connectome prior. Representative neuronal populations are arranged across the retina, lamina, medulla, lobula and lobula plate. Solid arrows indicate sparse feed-forward projections between successive neuropils, curved arrows denote local recurrent processing, and red dashed arrows summarize inter-region feedback. (B) ImageNet-trained FlyVision Large backbone. The input is the dog image already shown in panel C. ON and OFF stem maps and stages S1–S3 show mean absolute activations across all channels from a forward pass through the trained global-surround checkpoint. Each map uses its own first–99th percentile display range; color intensity is not comparable across maps. Each stage has a feed-forward 3×3 convolution and recurrent updates combining a 3×3 convolution with a signed four-population mixer, followed by GroupNorm, SiLU, lateral competition and a gated state update. S1/S2/S3 use one/two/two updates with shared weights. Global pooling, a 768-dimensional embedding and the ImageNet head produce 1,000 logits for this example. (C) Cross-domain scaling tests spanning MNIST, CIFAR-10, ImageNet-1K, skin-disease imaging, chest radiography and volumetric BrainAGE regression. BrainAGE is represented as multi-view analysis of a three-dimensional T1-weighted MRI volume: the model processes 24 slices across sagittal, coronal and axial planes and fuses them to a scan-level age estimate.
Figure 2: Scaling the FlyVision architecture across tasks. (A) Conserved FlyVision computational motif. Parallel ON and OFF pathways feed three hierarchical stages (S1–S3), each combining a spatial operator, gated recurrence, population-graph interaction and normalization/nonlinearity before global pooling, embedding and a task-specific prediction head. (B) FlyVision architecture family. Compact FlyVision uses 16/32/64-channel stages and a 64-dimensional embedding for MNIST and CIFAR-10; FlyVision Base uses 64/128/256-channel stages and a 512-dimensional embedding for 224 × 224 inputs; FlyVision Large uses 96/192/384-channel stages and a 768-dimensional embedding. Parameter counts vary with the task interface, from approximately 80,000 parameters in the compact model to 1.33–1.84 million in Base and 2.97–3.74 million in Large. (C) Task progression and transfer. The study advances from MNIST and CIFAR-10 through ImageNet-1K, skin-disease classification and chest radiography to BrainAGE regression on a three-dimensional T1-weighted MRI volume. A shared FlyVision Large encoder processes eight sagittal, eight coronal and eight axial T1w slices; the age head predicts slice age, the confidence head supplies learned voting weights, and three-axis aggregation produces the scan-level estimate. (D) Representative outcomes across the architecture family, including paired ImageNet-1K validation top-1/top-5 accuracy of 60.79%/82.33% for Base global surround and 66.25%/86.07% for Large global surround, together with compact recognition, biomedical transfer and multi-view brain-age regression.
Figure 3: MNIST held-out performance in the same overview format used for the larger classification tasks. Accuracy is plotted against macro-F1 for all benchmarked models on the common 10,000-image test set; horizontal bars show 95% Wilson confidence intervals for accuracy. Filled markers indicate scratch training and open markers transferred initialization. Marker radius is linearly mapped to trainable parameter count, with a small minimum for visibility; exact counts are reported in Table 1 .
Figure 4: Compact recognition on MNIST. (A) Held-out top-1 accuracy with 95% Wilson confidence intervals for FlyVision and conventional baselines. (B) Accuracy plotted against trainable parameter count on a logarithmic axis. Scratch FlyVision reaches 99.34% top-1 accuracy with 80,608 parameters, placing it near the highest observed accuracy while remaining markedly smaller than ResNet18.
Figure 5: CIFAR-10 performance across scratch and transfer conditions. Held-out accuracy is plotted against macro-F1 for all ten training conditions on the common 10,000-image test set; horizontal bars show 95% Wilson confidence intervals for accuracy. Filled markers denote scratch training and open markers transferred initialization. Marker radius is linearly mapped to trainable parameter count, with a small minimum for visibility; exact counts and initialization routes are reported in Table 2 .
Figure 6: Natural-image recognition and transfer on CIFAR-10. (A) Test top-1 accuracy for models trained from scratch. (B) Transfer-induced change in top-1 accuracy relative to each architecture’s scratch baseline. Direct ImageNet initialization improves ResNet18 by 2.40 percentage points, whereas MNIST initialization changes FlyVision by -0.41 percentage points. The complete accuracy–macro-F1–parameter landscape across scratch and transferred conditions is shown in Fig. 5 .
Figure 7: ImageNet-1K validation performance and model size. Validation top-1 accuracy is plotted against paired top-5 accuracy for the completed models. Bubble radius is proportional to trainable parameter count. FlyVision variants are shown in blue and ResNet18 in gray. Here LF denotes the learned low-frequency pathway described in Methods. The Large local-k7/low-frequency checkpoint was evaluated on all 50,000 validation images. Exact top-1, top-5, parameter and available operation-count values are reported in Table 3 .
Figure 8: Skin-disease held-out performance and model size. Test accuracy is plotted against macro-F1 for six models evaluated on the same 1,538 test images. Horizontal bars show 95% Wilson intervals for accuracy, and bubble radius is proportional to trainable parameter count. Global and local-k7 refer to full-image and 7×7 neighborhood surrounds, respectively; LF is the learned low-frequency pathway. Full per-image probabilities support the one-versus-rest macro-AUROC values in Table 4 and the directional confusion analysis reported in the Results.
Figure 9: Skin-disease held-out confusion matrix for FlyVision Large local-k7 with the learned low-frequency branch. The 22-class matrix uses the same 1,538 image-level test samples as Fig. 8 ; rows are true classes and columns are predicted classes. Color indicates the percentage within each true-class row, while numerals give counts on the diagonal and for off-diagonal cells with at least five images. The accompanying CSV contains all cell counts and the paired per-image predictions. Overall accuracy is 981/1,538 (63.78%). The evaluation therefore quantifies image-level generalization.
Figure 10: Chest-radiography held-out performance in the cross-task overview format. Accuracy is plotted against macro-F1 for six models evaluated on the same 3,175-image test set; horizontal bars show bootstrap 95% accuracy intervals. Marker radius is linearly mapped to trainable parameter count, with a small minimum for visibility. The five high-performing models cluster at the upper right of the full-range panel; the right panel labels and enlarges them using exact model coordinates and complete confidence intervals. The left panel retains the lower-performing MLP and the overall scale. Exact model size and performance values are reported in Table 5 .
Figure 11: Held-out chest-radiography benchmark. (A) Accuracy and macro-F1 for six models on the 3,175-image test set; horizontal bars show 95% confidence intervals from 300 bootstrap resamples. (B) Test macro-AUROC and macro-AUPRC. (C) Macro-F1 versus total parameter count on a logarithmic x-axis; vertical bars show the same bootstrap macro-F1 intervals. IN1K, ImageNet-1K pretraining.
Figure 12: Chest-radiography classwise behavior and error structure. (A–C) Test-set F1, sensitivity/recall and specificity for COVID-19, lung opacity, normal and viral pneumonia. (D) Representative confusion matrices for FlyVision Base, FlyVision Large, ResNet18 and MLP. Cells report observed counts and row-normalized percentages.
Figure 13: Chest-radiography performance-efficiency trade-offs. Held-out macro-F1 is plotted against (A) total parameter count, (B) approximate GFLOPs per image measured from the executed chest-radiography forward graph, (C) single-GPU inference latency and (D) single-GPU throughput. The FlyVision radiography operation counts reflect its downstream stride/recurrent schedule; Table 3 reports the ImageNet-forward counts. X-axes are logarithmic; vertical lines show 95% bootstrap confidence intervals for macro-F1.
Figure 14: Hierarchical chest-radiography analysis. (A) Disease-versus-normal sensitivity, specificity, F1 and AUROC for the collapsed flat classifier and hierarchical detectors. (B) Accuracy and macro-F1 with 95% bootstrap confidence intervals for the final flat, two-stage and multi-head four-class systems. (C) Corresponding test confusion matrices. (D) Validation sensitivity-specificity curves with selected thresholds and the target disease sensitivity of 0.95.
Figure 15: Disease-subtype performance and radiography routing errors. (A) Accuracy and macro-F1 with 95% bootstrap confidence intervals for the disease-only stage-2 classifier and the multi-head subtype output. (B) Subtype confusion matrices. (C) Percentage of true-normal test images routed to each disease class by the flat, two-stage and multi-head systems. Blue, orange and green bars denote COVID-19, lung opacity and viral pneumonia, respectively; lung opacity is the dominant false-positive destination for true Normal images in all three schemes.
Figure 16: BrainAGE multi-view aggregation overview. The complete BrainAGE model contains 3.27 million trainable parameters. Axis-specific and three-axis fused predictions are positioned by held-out MAE and RMSE; the fused result is MAE 5.98 years, RMSE 7.67 years, R2=0.868 and Pearson r=0.931 on 433 held-out scans.
Figure 17: Multi-view FlyVision Large BrainAGE regression. (A) Predicted versus chronological age for 433 held-out scans; the dashed diagonal denotes identity and the inset reports the primary regression metrics. (B) Bland–Altman representation of brain-age delta against the mean of predicted and chronological age. The solid horizontal line marks mean delta and dashed lines mark the 95% limits of agreement. (C) MAE for sagittal, coronal and axial axis-specific estimates and for final three-axis fusion. Each axis estimate aggregates eight native-scale slices using the Gaussian positional prior and learned slice confidence. (D) MAE across five-year chronological-age bins; sex-stratified MAE is shown descriptively in the inset.
Model
Initialization
Parameters
Accuracy (%)
Macro-F1 (%)
ResNet18
ImageNet transfer
11,181,642
99.38
99.38
FlyVision
Scratch
80,608
99.34
99.33
ResNet18
Scratch
11,181,642
99.22
99.22
LeNet
Scratch
44,426
98.60
98.59
MLP
Scratch
109,386
97.50
97.48
Table 1: MNIST test performance.
Model
Initialization
Parameters
Accuracy (%)
Macro-F1 (%)
ResNet18
ImageNet direct
11,181,642
85.84
85.81
ResNet18
ImageNet → MNIST
11,181,642
84.04
84.10
ResNet18
Scratch
11,181,642
83.44
83.41
ResNet18
MNIST transfer
11,181,642
82.47
82.52
FlyVision
Scratch
81,408
78.03
77.99
FlyVision
MNIST transfer
81,408
77.62
77.55
Table 2: CIFAR-10 test performance and transfer conditions.
Model
Parameters (M)
GFLOPs
Top-1 (%)
Top-5 (%)
FlyVision Base, global
1.836
6.632
60.79
82.33
FlyVision Base, local-k7
1.836
–
60.25
81.93
FlyVision Large, global
3.737
14.803
66.25
86.07
FlyVision Large, local-k7 + low-frequency
3.744
–
66.53
86.31
ResNet18, scratch
11.690
3.628
69.25
88.57
Table 3: ImageNet-1K validation performance at best-top-1 checkpoints. Top-5 is paired with the selected top-1 checkpoint. GFLOPs count forward Conv/Linear operations for 224 × 224 inputs with two FLOPs per multiply-accumulate for the ImageNet execution graph (stride-2 stem; S1/S2/S3 strides 1/2/2 and recurrent-update counts 1/2/2); the later local/low-frequency model was not profiled under this convention.
Model
Parameters (M)
Accuracy (%)
Macro-F1 (%)
Balanced accuracy (%)
Macro-AUROC (%)
FlyVision Base, global
1.334
60.01
55.47
55.24
94.43
FlyVision Base, local-k7
1.334
58.45
54.34
53.88
94.56
FlyVision Base, local-k7 + LF
1.339
62.42
58.67
58.06
95.23
FlyVision Large, global
2.985
63.39
59.94
59.80
95.59
FlyVision Large, local-k7 + LF
2.992
63.78
60.23
58.28
95.28
ResNet18
11.188
66.19
61.96
62.27
95.96
Table 4: Skin-disease test performance on the identical 1,538-image held-out set. Macro-AUROC is the unweighted mean of 22 one-versus-rest class AUROCs.
Model
Parameters (M)
GFLOPs
Accuracy (%)
Macro-F1 (%)
Macro-AUROC (%)
FlyVision Large (pretrained)
2.97
2.79
92.76
92.85
98.68
FlyVision Base (pretrained)
1.33
1.28
92.60
92.75
98.68
FlyVision Base (scratch)
1.33
1.28
92.19
92.09
98.67
ResNet18
11.18
3.63
91.56
91.98
98.01
LeNet
0.46
0.14
89.80
89.83
97.68
MLP
19.28
0.04
56.91
50.98
79.88
Table 5: Chest-radiography test performance and model efficiency. All models were evaluated on the same 3,175-image held-out test split. GFLOPs are measured from the executed radiography forward graphs for 224 × 224 inputs with two FLOPs per multiply-accumulate. The checkpoint-loaded FlyVision radiography graph uses stage strides 2/2/2 and one recurrent update per stage, whereas the ImageNet graph uses stage strides 1/2/2 and recurrent-update counts 1/2/2; macro-AUROC uses one-versus-rest aggregation.
Understanding the extent to which measured synaptic wiring determines computation remains a central challenge. Here, we couple the proofread adult Drosophila melanogaster connectome to an anatomically faithful model of its eye. Visual information is inputted in the eye model, then passed to the connectome, and finally read from a Kenyon-cell-centered linear decoder. This creates a connectome-only model in which the anatomical graph and eye geometry are fixed and only scalar synaptic gains and neuronal thresholds may be learned. The model supports multitask vision, including color discrimination, shape classification, and numerical discrimination that follows a ratio-dependent scaling characteristic of approximate number perception. To test whether precise connectivity is consequential under wiring economy, we compare the biological graph to randomized ensembles that increasingly preserve biological synaptic constraints. At matched wiring cost, the biological network consistently yields higher accuracy, whereas less constrained rewiring surpasses it at the cost of inflated wiring. These findings indicate that the measured connectivity and eye geometry jointly set efficient operating points for visual computation.
Eudald Correig-Fraga, Roger Guimerà, Marta Sales-Pardo
Innovamat Education, Sant Cugat del Vallès, Catalonia · Department of Chemical Engineering, Universitat Rovira i Virgili, Tarragona, Catalonia · Center for Computational Science and Applied Mathematics (ComSCIAM), Universitat Rovira i Virgili, Tarragona, Catalonia +1
While deep learning models achieve state-of-the-art performance in complex tasks, they remain brittle when faced with new environments or sensory deprivation. In contrast, biological systems exhibit remarkable tolerance to these challenges. We address this vulnerability by developing a recurrent neural network (RNN) whose architecture is directly derived from the synaptic-resolution brain connectome of the fruit fly Drosophila melanogaster. We demonstrate the feasibility of training the fly connectome neural network (FLYNN) to perform vision-based navigation in MuJoCo, achieving performance comparable to modern hand-crafted networks of similar parameter counts. Crucially, FLYNN exhibits superior resistance to out-of-distribution (OOD) data and tolerance to sensory loss without further training. It remained functional even under total vision loss while hand-crafted networks largely failed, even when specifically trained with camera dropout. Principal Component Analysis (PCA) of the internal state of FLYNN suggests that it exhibits a particularly high degree of representational modularity, which might be related to its robustness. Our work provides a new direction for designing resilient artificial agents following the topology of biological brains.
Benquan Wang, Jingdao Chen
Department of Computer Science Engineering, Mississippi State University, Mississippi State, MS 39762, USA
Proofreading--correcting segmentation errors in 3D brain reconstructions--is the rate-limiting step in synapse-resolution connectomics. We release ConnectomeBench2, a unified multi-species dataset of over 716,485 expert-labeled proofreading decisions with >4,500,000 associated images spanning four major open connectomes (mouse, human, zebrafish, fly), spanning both split and merge error correction. Trained on this dataset, a single Vision Transformer with shared encoders for mesh geometry and electron microscopy reaches human-level accuracy across species for split error correction and merge error identification, with performance scaling with data size and modality. Beyond accuracy, we show that the model is well-calibrated within distribution, that measures of distribution distance predict where calibration and accuracy will degrade on unseen data, and that connectomics-specific pretraining and active learning-based sample selection show potential to substantially reduce the labeling effort needed to extend to new species and brain regions. The benchmark provides the infrastructure to train and evaluate increasingly capable vision models for connectomic proofreading. Data and code availability. The ConnectomeBench2 dataset is released on Hugging Face at https://huggingface.co/datasets/jeffbbrown2/ConnectomeBench2. The accompanying codebase is available on GitHub at https://github.com/timfarkas/ConnectomeBench2.