Real-world perception systems must adapt to changing environments, but manual image annotation cannot scale to field data volumes. We present BirdsEye, which shifts expert annotation from images to the field: an operator records target locations in world coordinates using RTK positioning and calibrated projective geometry propagates each observation to all frames where the target is visible. To quantify how well physical annotations align with image observations, we derive a first-order mapping from camera-pose uncertainty to pixel uncertainty and validate it against Monte Carlo simulation. This mapping is linear in the six per-axis pose variances, so it inverts into a sensor design tool: we give a sufficient condition converting an annotation tolerance into a convex set of admissible pose-noise budgets, a closed-form largest admissible scaling of a deployed sensor suite, and a unique per-axis pose specification under an equal-budget-share allocation. We also analyze the planar-surface approximation underlying the projection, which holds up to 10 degrees of terrain slope. By direct measurement, we show that system projection accuracy is sub-decimeter (sub-30 pixel) at AGL altitudes of 10-20m under conditions excluding sustained yawing. During an in-field case study across three agricultural sites, two field workers produced 12,524 annotated frames carrying 55,600 labels in roughly 12 hours (25.5x per-worker rate increase over manual labeling). Detectors trained on imagery collected by this workflow recovered 56-89% of in-view surveyed targets at a geographically distinct farm, at pre-registered operating points; human review of the leading configuration estimates detection precision at 83-87%, spanning three tie-break conventions for clusters carrying contradictory human verdicts.
Figures & tables
Figure 1: The BirdsEye UAV-based image annotation system. Pictured (bottom left) are the specialist in the field with a handheld annotation device, the RTK GPS base station, and the UAV surveying a field. System software (bottom right) then either collates expert annotations and UAV imagery via camera projective models, creating an annotated dataset, or passes UAV imagery directly to a trained neural model for feature detection and subsequent georeferencing of detection event pixels (top left).
Figure 3: An illustration of the variance within the general class of “burrow”. (Left) An obvious burrow with a small flattened mound. (Center) An occluded burrow and an obvious mound; is “mound” as important as “hole”? (Right) A very old, flattened mound with no visible hole; how degraded a target is allowable?
Figure 4: Illustration of the geometric relationship between a 3D geotag in the world and its 2D projection in the image. Also pictured are the position vectors relevant to the cross-products in the geotag detection method of Section 3.6.2 .
Figure 5: Our soft labeling strategy lets us use empirical or analytical projection uncertainty estimates to tune the shape of Gaussian patches, mitigating the chance that projective errors lead to truly erroneous positive annotations.
Altitude
Maneuver
Error Norm
Standard Deviation
Meters
Pixels
Meters
Pixels
10m
Hovering
0.038
18.8
0.019
9.7
Yawing
0.191
82.2
0.052
22.1
Head-on
0.045
19.6
0.024
10.4
Strafing
0.062
27.4
0.044
19.7
20m
Hovering
0.043
9.8
0.037
8.4
Table 1: Norms and uncertainties of error vectors from georeferencing benchmark tests. †
Figure 6: Perturbation study results showing pixel uncertainty sensitivity to individual pose degrees of freedom. Shaded regions denote the range of Monte Carlo estimates across all visible test points at a given altitude.
Figure 7: Scaling behavior of pixel-space uncertainty under covariance inflation. (Top) Growth of the major axis of the uncertainty ellipse, which exhibits the expected square-root scaling law (slope 0.5 in log–log space) predicted by first-order propagation, up to the point where linearization begins to break down. (Middle) Evolution of ellipse anisotropy, demonstrating relative insensitivity to noise scale. (Bottom) Similarity of the analytical model and Monte Carlo ensemble of covariance, by the symmetric Kullback-Leibler divergence (sKLD). The stated heuristics for indistinguishable ( 10−2 ) and clearly distinguishable ( 10−1 ) distributions are well above these curves, but the 5m is beginning to diverge at scale factors above 10.
Sufficient condition
AGL
f(sbase) [px 2 ]
κ⋆
Penalty
Verdict
Eq. 4 (trace)
10 m
108.61
0.9207
1.450 ×
reject ( 1.09× )
20 m
46.42
2.154
1.761 ×
accept
supXλmax(Σu) (exact)
10 m
74.92
1.335
1.000 ×
accept
20 m
26.36
3.794
1.000 ×
accept
Table 2: Admissible uniform scaling κ⋆ of the deployed pose-noise budget (Equation 5 ) under two criteria, at λu,target=100 px 2 ( r=30 px at 3σ ). f(sbase) is the bound’s estimate of the worst-case pixel variance at the deployed Σ0 ; the true value is supXλmax(Σu)=74.92 px 2 at h=10 m and 26.36 px 2 at h=20 m. The trace condition is an upper bound on λu,max and the eigenvalue condition is exact. The “penalty” column shows κ⋆ (exact) /κ⋆ (row). Suprema are evaluated over 2,307,121 sampled points per altitude at 1 px stride.
Axis
Deployed 1σ
Balanced, 10 m
Balanced, 20 m
Budget Share at sbase , 10 m
x
10 mm
13.64 mm
27.28 mm
16.7%
y
10 mm
13.45 mm
26.89 mm
16.7%
z
60 mm
49.12 mm
98.25 mm
42.9%
roll
0.04 ∘
0.0753 ∘
0.0753 ∘
8.5%
pitch
0.04 ∘
0.0741 ∘
0.0741 ∘
9.0%
yaw
0.13 ∘
0.3013 ∘
0.3013 ∘
6.1%
Table 3: Admissible per-axis pose uncertainty ( 1σ ) required to hold the soft-label radius used in training ( r=30 px at 3σ , λu,target=100 px 2 ) at every point in frame, from Equation 4 with the balanced allocation of Section 3.7.2 . Deployed values are the datasheet 1σ figures of Section 4.2 . Bold entries are exceeded by the deployed platform. Admissible translation tolerances scale linearly with altitude; rotation tolerances are altitude-invariant at nadir. The final column is each axis’s share of the deployed platform’s total worst-case pixel variance, gi⋆σi,dep2/∑jgj⋆σj,dep2 ; it describes the deployed Σ0 and is not the equal 1/6 share that the balanced allocation assigns by construction. Altitude accounts for 44.7% of that total, which is why z is the axis over its allocation.
Figure 8: Error propagation under violations of the annotation projection pipeline flat-world assumption. The horizontal axis represents the terrain slope angle and the vertical axis represents the distance, in meters, of a test point from the optical center, normalized by the altitude, in meters, of the camera and contour labels are in units of pixels. We note that terrain-induced error exceeds the magnitude of pose-induced error at slopes of approximately 10 degrees.
Location
Dataset Composition
Dataset Role
# Geotags
# Frames
Annotations
Campus Farm
1,181
10,034
15,000
Train
Campus Farm
308
6,514
10,578
Train
Partner Site 1
161
2,120
6,354
Test
Campus Farm
256
3,371
8,117
Validation
Partner Site 2
181
2,555
6,613
Train
Table 4: (Upper) Summary of principal flight campaign. (Middle) Summary of subsequent field timing experiments performed at the Campus Farm. (Lower) Comparison of annotation strategy productivity. All flights performed at nominally 10m AGL. The “Dataset Role” columns show which flights were partitioned into the training, validation, and test sets of the standard ML training pipeline.
Figure 9: A georeferenced image frame from the test flight, showing both human-made field annotations (Green Rings) and CNN detections (Red Patches). Observe that the model detects both the concrete, positive training examples established by human experts and also borderline cases. This difference in behavior, where humans isolate good training examples and CNNs indicate every detection, leads to the failure of the simple precision statistic as a performance assessment in our case study. To the right of the image frame is a georeferenced detection map under construction, where magenta squares represent human field annotations and green dots represent the centroids of model detection patches.
Arm
Ckpt
t / ε / m / εpred
Recall
GMCF
F1
Max pooled F1
014
0.08 / 0.2 / 2 / 0.2
0.466
0.480
0.472
Max pooled recall
030
0.04 / 0.1 / 2 / 0.2
0.805
0.148
0.250
Lowest review burden
013
0.035 / 0.2 / 2 / 0.2
0.434
0.417
0.425
Table 5: The 3 pre-registered confirmatory arms. Each arm fixes a checkpoint and a complete operating point ( t / ε / m / εpred ) from the exact-inference pooled sweep. Validation figures are the micro-averaged screening values that the rule selected on; F1 is the selection-only quantity.
Ckpt
t / ε / m / εpred
RPRE
GMCF PRE
013
0.035 / 0.2 / 2 / 0.2
0.583
0.189
014
0.08 / 0.2 / 2 / 0.2
0.563
0.223
030
0.04 / 0.1 / 2 / 0.2
0.887
0.058
015
0.08 / 0.1 / 2 / 0.35
0.815
0.173
017
0.08 / 0.15 / 5 / 0.2
0.695
0.235
021
0.12 / 0.15 / 5 / 0.2
0.702
0.164
Table 6: Held-out PRE statistics on the unseen site, whole flight, 2,117 frames processed (3 of the 2120 had duplicated timestamps), 151 of 161 geotags in view under Equation 1 . Recall is over in-view geotags. The top 3 rows are the pre-registered checkpoint configurations, while the bottom 9 are the remainder of the field.
GMCF POST
R POST ( ≥ )
Ckpt
Clusters
GMCF PRE
R PRE
pos
maj
neg
pos
maj
neg
013
132
0.432
0.490
0.780
0.742
0.705
0.615
0.609
0.603
014
113
0.522
0.536
0.867
0.832
0.832
0.634
0.630
0.630
030 (t=0.04)
445
0.189
0.868
0.404
0.339
0.279
0.922
0.914
0.909
015
201
0.373
0.682
0.726
0.682
0.632
0.787
0.780
0.773
017
225
0.324
0.715
0.684
0.627
0.578
0.819
0.809
0.803
Table 7: Adjudicated GMCF and recall per pre-registered arm, at the parameters the review was conducted at ( ε=0.20 m, m=10 , εpred=0.35 m, 3 cm grid dedup). The PRE column here is therefore not the pre-registered geometry of Table 6 . POST columns are given under all three tie-break conventions in the order positive-wins / majority-wins / hard-negative-wins. Every POST recall is an upper bound and is printed with ≥ . The top 3 rows are the pre-registered checkpoint configurations, while the bottom 9 are the remainder of the field.
Figure 10: Effect of adjudication on every reviewed checkpoint layer, each at the review’s parameters ( ε=0.20 m, m=10 , εpred=0.35 m). Filled and open circles are the PRE (geometry-only) points; arrows run to the adjudicated result; the thick segment spans the three tie-break conventions and carries one white-filled marker per convention, because there are exactly three discrete outcomes there and not a continuum. The segment is a convention, not an error bar. The three layers carrying pre-registered arms are drawn dark and labeled in bold: 014 (max F1), 013 (lowest review burden) and 030 at t=0.040 (max recall); the remaining layers are drawn light. The whole reviewed field is shown because it is the evidence that the arms are ordinary members of it rather than the points where adjudication happened to help: every layer moves up and to the right, and the largest precision gain belongs to a layer no arm selected. POST recall values are upper bounds ( a=0 ).
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11: Pooled recall against pooled GMCF for all 30 screened checkpoints, micro-averaged over the 3 validation flights. Filled circles are the 11 checkpoints that passed both health guards and are tagged with their checkpoint identifier; open squares are sparse and open triangles collapsed.
Figure 12: The two health guards as a decision rule, for all 30 screened checkpoints. Both axes are logarithmic. Shaded strips are the rejected regions: below the distinct-position cut (0.05, dashed) a checkpoint is collapsed, and left of the detection-rate cut (1.7 det/frame on its worst validation flight, dotted) it is too sparse to score. The double-headed arrows are the margins – the ratio between the nearest checkpoint on either side of each cut – and only the cut-bounding checkpoints are labeled, because they are the ones whose values decide where a cut may sit. A cut placed on a 30-point sample is an argument about its neighbors, not a natural boundary.
Ckpt
det/frame (pooled)
Recall
GMCF
F1 (selection only)
Relative review load
028
249.67
0.782
0.130
0.224
66.9 ×
021
176.80
0.853
0.086
0.156
47.4 ×
017
43.21
0.739
0.139
0.234
11.6 ×
031
34.46
0.667
0.159
0.256
9.2 ×
030
30.03
0.695
0.202
0.313
8.0 ×
044
14.20
0.440
0.375
0.405
3.8 ×
Appendix
Table 8: The 11 healthy checkpoints, micro-averaged over the 3 validation flights, sorted by pooled detections per frame descending. Review load is relative to the lightest checkpoint. The field spans 67 × in review load for 3.0 × in recall, which is an operational trade and not a ranking: no row in this table is a recommendation. F1 is the pooled validation quantity used for selection only (Appendix B ). The remaining 19 of 30 checkpoints were excluded by the health guards (11 collapsed, 8 sparse).