This book presents a code-first introduction to computer vision, spanning classical 2D image processing, classical 3D vision, and deep learning. Organized as 44 short chapters across three parts, the book builds each topic from first principles: image arithmetic and morphology; convolution, pyramids, and frequency-domain filtering; feature detection, optical flow, and stereo; projective geometry, camera calibration, and structure from motion; and the full arc of modern deep learning, from a single neuron through convolutional networks, backpropagation, classic architectures, transfer learning, object detection, and semantic and instance segmentation, concluding with engineering considerations like mixed-precision and parallel training. Every technique is implemented directly in Python and NumPy or PyTorch and checked numerically against the corresponding OpenCV or PyTorch library function, so readers see not just the mathematics but its concrete behavior on real and synthetic data. The material was distilled with AI assistance from freely available online course notes, condensing extensive working code into concise mathematical exposition while preserving verified, reproducible results throughout. It is intended as a self-contained reference for students and practitioners who want to understand computer vision algorithms and their Python implementations.
Figures & tables
Figure 1.1: A 100×100 grayscale array with a white square.
Figure 1.2: Inversion is a single elementwise array operation.
Figure 1.3: Red + green renders as yellow.
Figure 1.4: The six psychological primaries as a 2×3 RGB image.
Figure 1.5: Slicing the first axis selects a row; the second axis, a column.
Figure 1.6: Left: BGR array shown directly (wrong colors). Center and right: the two equivalent fixes, in agreement. Photo: Picryl .
Figure 2.1: Look closely at the circle: it became darker , not brighter.
Figure 2.2: Naive NumPy addition wraps (darkens); cv2.add saturates (clamps at 255) — the same idea applies at the bottom end for cv2.subtract , clamped at 0 instead of wrapping to a large positive value.
Figure 2.3: Photo, foreground mask, and the selectively brightened result. Photo: Wikimedia Commons .
Figure 2.4: cv2.addWeighted at five values of α , cross-dissolving between two dog photos. Photos: Wikimedia Commons , Wikimedia Commons .
Figure 2.5: The difference image is zero everywhere nothing changed, bright wherever the circle used to be or now is, and dark in the region where it is present in both frames.
Figure 3.1: A synthetic bright circle on a dark background, with added noise and salt-and-pepper specks.
Figure 3.2: Simple thresholding at t=128 .
Figure 3.3: Histogram with the Otsu-selected cutoff, compared against the hand-picked threshold; the two results are visually near-identical here.
Figure 3.4: Thresholded, eroded, dilated, and opened (erode then dilate) versions of the same binary image.
Figure 4.1: A photo of fruit, thresholded with Otsu’s method (Chapter 3 ), then cleaned up with a morphological opening. Photo: Stan Birchfield.
Figure 4.2: Flood fill from a single seed (red × ): the hand-written stack version (left) and cv2.floodFill (right) fill an identical region.
Figure 4.3: Each connected blob labeled with a distinct color and its integer ID.
Figure 4.4: All blobs (left) versus only blobs with area above a chosen cutoff (right), which discards the two smallest apples.
Figure 5.1: The centroid recovered from moments, marked on the rotated ellipse.
Figure 5.2: The major axis recovered purely from second-order central moments matches the ellipse’s true drawn orientation.
Figure 5.3: Original, translated, rotated + scaled, and a different shape entirely. Images: Public Domain Pictures .
Figure 6.1: A unit circle stretched by A ; the red lines are the eigenvectors scaled by their eigenvalues.
Figure 6.2: Major (thick) and minor (thin) axes recovered from a known ellipse’s eigenvectors: a=70.6 , b=25.5 , matching the drawn semi-axes (70, 25).
Figure 6.3: Equivalent-ellipse axes for a pair of eyeglasses. The ellipse summarizes the mass distribution rather than tracing the outline, so it cuts across the bridge between lenses. Photo: Stan Birchfield.
Figure 7.1: Distance fields for the three metrics, contoured. Manhattan distance forms diamonds, chessboard distance forms squares, and Euclidean distance forms circles — the shape most people intuitively think of as “equal distance away.”
Figure 7.2: Chamfer (mask size 3) vs. a more accurate distance transform (mask size 5) and their difference; the two agree closely (max error under half a pixel).
Figure 8.1: The test image: an asymmetric letter F.
Figure 8.2: Horizontal, vertical, and both-axis flips.
Figure 8.3: A crop, obtained by plain NumPy slicing.
Figure 8.6: A unit square under each transform. Only the affine square stops looking like a (possibly resized) square — its corners are no longer 90∘ , because affine transforms allow shear. All three keep opposite sides parallel. (The more general projective transforms, e.g. camera perspective, are covered in a later chapter.)
Figure 8.7: Original, a simulated skewed photo with 3 identified landmarks, and the rectified result — mean absolute pixel error against the true original is just 3.14 (out of 255).
Figure 9.1: The test image used throughout this chapter.
Figure 9.2: Forward mapping: visible holes (black speckles) where no source pixel happened to land.
Figure 9.3: Inverse mapping with nearest-neighbor sampling: no holes, but blocky.
Figure 9.4: Nearest-neighbor (left) vs. bilinear (right). Bilinear has smooth edges around the triangle and rectangle instead of jagged steps — the classic trade-off is a softer, slightly blurrier image in exchange for removing aliasing artifacts.
Figure 9.5: Our bilinear implementation vs. cv2.warpAffine(..., flags=cv2.INTER_LINEAR) : max pixel difference is 1 (out of 255), mean difference 0.092 . The tiny remaining gap comes from OpenCV using fixed-point rounding internally for speed rather than full floating-point arithmetic.
Figure 10.1: A ramp-up-then-down signal and its convolution with a derivative kernel: positive while rising, zero at the peak, negative while falling.
Figure 10.2: An asymmetric test shape, its convolution (kernel flipped), and its correlation (kernel not flipped) with the same asymmetric kernel — visibly different results.
Figure 10.3: A noisy synthetic photo, then box blur, Gaussian blur, sharpen, and Sobel- x edge detection, all implemented with the same convolve2d function above.
Figure 11.1: Naive subsampling turns the regular brick pattern into fine, noise-like speckle — a classic aliasing artifact (the same effect that makes a car’s wheels look like they’re spinning backwards on camera). Blurring first keeps the wall recognizable. Photo: publicdomainpictures.net .
Figure 11.2: Larger sigma averages over a wider neighborhood, removing finer detail, as long as the kernel is large enough to represent it.
Figure 11.3: A 3×3 kernel is far too small to represent sigma=5 : mean absolute difference against the correctly auto-sized kernel is 9.31 gray levels.
Figure 11.4: A 5-level Gaussian pyramid: 256×256 down to 16×16 .
Figure 11.5: pyrDown then pyrUp loses detail permanently: mean absolute reconstruction error is 1.59 gray levels. The sharp edges of the circle and square come back soft and shifted-looking — pyrUp (interpolation) can only guess, not restore, the discarded detail.
Figure 12.1: Gradient magnitude from the simple difference kernel and the three smoothed operators, on a clean test image. On a low-noise image like this, the four results look nearly identical.
Figure 12.2: Naively thresholding ∥∇I∥ gives thick, blobby edges; cv2.Canny gives thin, connected ones. Image source: National Gallery of Art .
Figure 12.3: With low thresholds on a noisy image, speckle noise survives as spurious edge fragments. At the highest thresholds, faint detail disappears completely, while the strongest outlines survive — Canny’s threshold trading off noise rejection against sensitivity to faint-but-genuine edges.
Figure 12.4: Three points on the line x+y=300 each trace a sinusoid in (ρ,θ) space; all three cross at the line’s true parameters, ρ=212.1 , θ=45∘ .
Figure 12.5: Input line, the log-scaled (ρ,θ) accumulator (true parameters marked in cyan), and the recovered line overlaid in red. The basic Hough transform returns an infinite line, not a line segment.
Figure 12.6: Two lines (true angles −45∘ and 0∘ ) plus a distractor circle with no straight edges. Every detected segment lands almost exactly on one of the two true angles; the circle correctly produces no line detections. Each line typically yields two or three near-duplicate detections, since Canny finds an edge on each side of a stroke with finite width.
Figure 12.7: Detected circle (red) overlaid on the true one (white).
Figure 12.8: 2D slices of the 3D circle accumulator: radius vs. cx (max over cy ) and radius vs. cy (max over cx ). Each slice shows a bright X where every edge point’s votes cross at the true center and radius.
Figure 13.1: A smoothed step edge, its first derivative (a peak), and its second derivative (a zero crossing), all at the same location.
Figure 13.2: Every edge produces a positive lobe on one side and a negative lobe on the other — the edge itself sits at the zero crossing between them. The response is shown on a red-blue diverging color scale specifically because the Laplacian is signed — red marks positive values, blue marks negative, making the sign flip at each zero crossing easy to see at a glance.
Figure 13.3: A noisy disk, its LoG response, and the resulting zero-crossing edge map (Marr-Hildreth edges). The LoG response uses the same red/blue diverging scale as Figure 13.2 , for the same reason.
Figure 13.4: LoG response vs. the scaled DoG approximation — visually almost indistinguishable.
Figure 13.5: A 5-level Laplacian pyramid, each level shifted by +128 gray levels so negative values are visible. Each level is mostly flat gray except right at edges — exactly where detail is lost by blurring and downsampling. The final, smallest level stores actual image content rather than a difference, since there’s nothing coarser left to compare it to.
Figure 13.6: Original vs. reconstructed-from-Laplacian-pyramid — pixel-perfect.
Figure 14.1: Gaussian blur smears each corrupted pixel’s extreme value into its neighbors, leaving faint speckle everywhere. The median filter simply outvotes an isolated bad pixel with its many good neighbors, removing nearly all of the noise while keeping edges sharp.
Figure 14.2: Top: a noisy step edge, Gaussian-blurred, and bilateral-filtered. Bottom: intensity profiles along a column crossing the edge. The Gaussian-blurred profile ramps gradually across many rows — the edge has been visibly softened. The bilateral-filtered profile stays flat on each side (noise removed) but still jumps sharply at the true edge — it survived because pixels across the boundary were too different in intensity to be averaged together, no matter how spatially close they were.
Figure 15.1: A square wave approximated with 1 harmonic and with 9 harmonics — more harmonics track the true square wave more closely, especially near the sharp transitions.
Figure 15.2: A filled square and its log-magnitude spectrum. The bright cross through the center comes from the square’s sharp horizontal and vertical edges — a hard step edge contains energy at every frequency along the direction perpendicular to it, which is why a sharp-edged shape has such a spread-out spectrum.
Figure 15.3: A grating with 12 cycles across the image and its spectrum. The brightest spectrum point lands at offset −12 from center, matching the grating’s true frequency of 12.
Figure 15.4: Ideal low-pass and high-pass masks in the frequency domain.
Figure 15.5: Original image, ideal low-pass result (blurred), and ideal high-pass result (edges only).
Figure 15.6: Ideal low-pass (sharp cutoff, visible ringing near edges) vs. a Gaussian mask (smooth falloff, no ringing).
Figure 16.1: db2 ’s scaling (low-pass) and wavelet (high-pass) filters, four taps each.
Figure 16.2: A test image and its four db2 subbands.
Figure 16.3: Gabor kernels across four orientations ( θ ) and four wavelengths ( λ ).
Figure 16.4: Test bars at 0∘ , 45∘ , 90∘ , 135∘ .
Figure 17.1: Reconstructing an 8×8 block from only the top-left n×n DCT coefficients, for n=8,4,2,1 . Even keeping just the top-left 2×2 (4 of 64 coefficients — a 16× reduction) preserves most of the block’s overall structure; only the finest detail is lost.
Figure 17.2: File size vs. JPEG quality (left), and the rate-distortion curve: PSNR vs. file size (right).
Figure 17.3: Photograph vs. graphic, original and JPEG-compressed at quality 75. Photo by Peter Herrmann on Unsplash .
Figure 17.4: A 4× zoomed crop, original vs. JPEG quality 10 — visible blocking and ringing at the roofline.
Figure 18.1: A real photo and a scatter plot of its R and G channel values — tightly correlated (correlation coefficient 0.954 on this image), because shading scales all three channels together.
Figure 18.2: A shaded disk (dark on the left, bright on the right), an HSV hue threshold, and a fixed RGB range. Hue thresholding recovers the disk regardless of the lighting gradient; the RGB range misses the darkened side entirely.
Figure 18.3: Original, chroma subsampled 8× (barely visible loss), and luma subsampled 8× (clearly visible loss) — the same detail loss is simply less objectionable when it happens to color rather than brightness.
Figure 18.4: Histogram of Lab* distances for 2000 color pairs all held at the same RGB distance. Any algorithm using raw RGB distance as a proxy for perceptual similarity (nearest-neighbor color matching, k-means color clustering, background-subtraction thresholds) inherits this same distortion; switching to Lab distance is a simple, standard fix.
Figure 19.1: A synthetic noisy 4-region color image, the true regions, and the k-means segmentation on Lab* pixel colors — purity 1.000 against the true regions.
Figure 19.2: True concentric rings vs. the k-means result (centers marked with × ). The clusters aren’t ambiguous to a human eye at all; the algorithm’s assumption (round, center-based clusters) is simply the wrong tool for this shape.
Figure 19.3: True rings, the k-means result (fails), and the DBSCAN result (succeeds). DBSCAN never assumed a center or a shape, only that points within a ring are densely connected to their neighbors, while the gap between rings is not.
Figure 20.1: A test image containing corners, edges, and flat regions.
Figure 20.2: The test image and its Harris response R (red = corner, blue = edge).
Figure 20.3: cv2.goodFeaturesToTrack (Shi-Tomasi) corners marked in red.
Figure 20.4: SIFT keypoints overlaid on the original (circle size = scale, line = orientation). Photo by Dawson Lovell on Unsplash .
Figure 20.5: 191 matches across a 35∘ rotation +0.7× scale change. Nearly all the ratio-test survivors are genuinely correct correspondences, despite the rotation and scale change — something a detector fixed to one scale and orientation simply couldn’t recover.
Figure 21.1: Two different true motions (right + down vs. down-only). Inside a small window straddling the rectangle’s top (horizontal) edge, the two motions look pixel-for-pixel identical — verified directly: a local measurement genuinely cannot tell them apart.
Figure 21.2: Corners (green) and estimated motion (red arrows, exaggerated 15× since the true motion is sub-pixel).
Figure 21.3: Tracked corners with a larger true motion, using cv2.calcOpticalFlowPyrLK .
Figure 21.4: Flow magnitude spreading outward from a single textured square, after 1, 20, and 200 Gauss-Seidel iterations. After just 1 iteration, flow is only nonzero right at the square’s edges; by 200 iterations it has spread well beyond, though still far from every pixel — a real solver would run more iterations (or use a pyramid) to converge everywhere.
Figure 21.5: A synthetic frame and its dense Farnebäck flow field, visualized with color = direction, brightness = speed. Every moving circle produces the same color here, since they all share the same true motion.
Figure 21.6: A frame from the Army sequence and its dense flow field. Image source: Middlebury Optical Flow .
Figure 22.1: Left image, right image, and ground-truth disparity (three depth planes).
Figure 22.2: A 100×100 crop, its true disparity, and the recovered disparity from a from-scratch block matcher — the three depth planes are cleanly separated away from the border and depth-boundary regions.
Figure 22.3: Left-to-right disparity, right-to-left disparity, and the result after the consistency check (gray = rejected). On this crop, 1440 of 10,000 pixels are marked inconsistent, concentrated at occlusions and depth-boundary windows.
Figure 22.4: Ground truth vs. cv2.StereoBM output at full resolution (gray = invalid).
Figure 22.5: blockSize 5, 15, and 31 on the same stereo pair — small windows preserve sharp edges but are noisier; large windows are smoother but blur depth boundaries.
Figure 22.6: Left-to-right, right-to-left, and consistency-checked disparity on the Tsukuba stereo pair. The lamp, the nearest object in the scene, stands out with the largest disparity; the bookshelves recede smoothly behind it. Only 1121 of 110,592 pixels ( 1.0% ) get flagged — concentrated exactly along depth discontinuities: book spines, shelf edges, the head’s silhouette. Image source: Middlebury Stereo Datasets (University of Tsukuba).
Figure 22.7: The synthetic scene’s disparity map converted to metric depth.
Figure 22.8: The depth map lifted into 3D: three flat terraces at three different heights, with holes where StereoBM had no confident match cut cleanly out of each one.
Figure 23.1: OLS (vertical error) vs. TLS (perpendicular error) when both x and y are noisy.
Figure 23.2: Ordinary least squares, corrupted by 15 outliers among 40 inliers.
Figure 23.3: RANSAC’s found inliers (blue) vs. rejected outliers (red), and the resulting fit.
Figure 23.4: Same contaminated data, three different fits: OLS, Huber-reweighted, and RANSAC.
Figure 24.1: Two parallel lines, before and after a homography with a nonzero bottom row — they now converge toward a vanishing point.
Figure 24.2: A document, a simulated photograph at an angle (strong keystoning), and the rectified result from 4 corner correspondences. Mean absolute pixel difference after the distort-then-rectify round trip is 4.59 — small resampling loss only.
Figure 25.1: The two input views. Photo: Stan Birchfield.
Figure 25.2: Feature matches in the overlap region (first 60 shown).
Figure 25.3: The stitched panorama. The roof (top, fully visible only in view 2) and the bush (bottom, fully visible only in view 1) both appear in the mosaic, confirming that the composite genuinely draws from both source images.
Figure 25.4: Reprojection error of the RANSAC inliers, histogrammed.
Figure 25.5: cv2.Stitcher in SCANS mode: status OK .
Figure 26.1: A 3D cube, projected through P=K[R∣t] onto the image plane.
Figure 26.2: All in focus (pinhole idealization) vs. focused at 5 m — near and far objects blur, the mid-depth object stays sharp.
Figure 26.3: A true-color scene, a 16×16 crop of the raw Bayer mosaic (grayscale readout, no color visible), and the same crop colorized by filter, revealing the RGGB pattern.
Figure 26.5: Correct vs. wrong Bayer pattern assumption — a strongly color-shifted image.
Figure 26.6: Naive blend of encoded black/white (left) vs. the physically correct blend (middle) vs. encoded 0.5 for reference (right) — the naive blend is visibly, dramatically darker.
Figure 26.7: True colors, under a warm tungsten cast, and gray-world corrected. Mean absolute pixel error drops from 18.98 to 7.43 after correction — substantial, though not to zero, since this small scene (dominated by a colored circle) isn’t perfectly neutral on average.
Figure 27.1: The flat checkerboard target, and cv2.findChessboardCorners ’ detections on one synthetic photo.
Figure 27.2: An ideal pinhole grid (straight lines) vs. the same grid through a lens with k1=−0.3 , k2=0.1 — visibly bowed.
Figure 27.3: Distorted grid (as captured) vs. undistorted using the estimated K and distortion — the lines straighten out, though a faint bow remains, since the estimated parameters are themselves only approximately correct.
Figure 28.1: Three candidate depths for the same pixel x1 , all lying in one epipolar plane; their projections into camera 2 all fall on the same epipolar line.
Figure 28.2: 40 3D points projected into both cameras.
Figure 28.3: 6 selected points in image 1, and their predicted epipolar lines l2=Fx1 in image 2 — every true match lands exactly on its predicted line, confirming a point’s search in the second image really does collapse from 2D to 1D, even without rectifying the images first.
Figure 29.1: The two input views. Image source: BlendedMVS (CC BY 4.0).
Figure 29.2: Feature matches between the two views.
Figure 29.3: The recovered 3D points (colored, estimated pose) alongside points triangulated with the true poses (gray), plus both camera centers.
Figure 30.1: A 2D cloud with its principal direction, and the same points projected onto that direction (1D).
Figure 30.2: Left: the decision boundary in the original 2D space, perpendicular to w . Right: the same points projected down onto w , with classification reduced to thresholding a single number at zero. The 2D picture is more intuitive, but the 1D picture is what actually generalizes — in 100 dimensions there’s no picture to draw of the “boundary,” but “project to a single number, then threshold” still works exactly the same way.
Figure 30.3: The best possible linear boundary between a disk and a surrounding ring — no rotation or shift of a single straight line can separate them, so the best any linear projection can manage is a mediocre compromise.
Figure 31.1: The same two-class dataset from Chapter 30 .
Figure 31.2: The sigmoid activation.
Figure 31.3: Binary cross-entropy: the penalty for one prediction, as a function of the predicted probability.
Figure 31.4: Training loss over 500 epochs of gradient descent.
Figure 32.1: The ring-and-disk dataset — no single linear projection separates these (Chapters 30 – 31 ).
Figure 32.2: Training loss (2-layer MLP) over 3000 epochs.
Figure 32.3: Each hidden unit is one projection direction wi , cutting the space along wi⋅x+bi=0 .
Figure 32.4: Original input space vs. the hidden representation (top 2 PCA directions of a 4D space).
Figure 32.5: Accuracy vs. capacity, across 5 random seeds at each width.
Figure 33.1: Plain gradient descent: the classic zigzag, 30 steps on a narrow bowl.
Figure 33.2: 30 steps of plain SGD, momentum, and Adam on the same narrow bowl.
Figure 33.3: Three learning rate schedules: step decay, cosine annealing, and linear warmup.
Figure 33.4: Noisy gradients: constant learning rate vs. step decay, loss on a log scale.
Figure 34.1: Sobel’s edge kernel, applied as a conv layer, on a simple square.
Figure 34.2: Train examples (centered positions) vs. test examples (unseen, corner positions).
Figure 35.1: The entire training set is only 12 noisy images like these.
Figure 35.2: The classic overfitting curve: training loss falls monotonically while validation loss bottoms out and rises again.
Figure 35.3: 6 of 240 augmented copies, all still “the same 12 base images.”
Figure 36.1: One example from each of the four classes.
Figure 36.2: The confusion matrix, visualized: rows are true class, columns are predicted class.
Figure 37.1: Gradient magnitude vs. depth, plain sigmoid network vs. residual network, log scale.
Figure 39.1: An input image and its saliency map — brightest where the gradient of the class score with respect to that pixel is largest.
Figure 39.2: Input image, saliency map, and Grad-CAM heatmap, side by side.
Figure 39.3: t-SNE of each test image’s 16-d pooled feature vector, colored by true class: untrained CNN features (left) vs. trained CNN features (right).
Figure 40.1: Synthetic face examples vs. non-face clutter.
Figure 40.2: The scene (true faces outlined in green) and the window classifier’s score map, indexed by window center (gray border = no window fits there).
Figure 40.3: 2 detections after NMS (red, dashed) vs. 2 true faces (green).
Figure 40.4: cv2.CascadeClassifier on a real photo: 4 detections. Image source: NASA (public domain), via Wikimedia Commons.
Figure 41.1: Sample training scenes, each with a single object and its ground-truth box.
Figure 41.2: Predicted boxes (red, dashed) vs. true boxes (green) on four test scenes, with per-example IoU.
Figure 41.3: cv2.FaceDetectorYN : 3 detections, each with a confidence score. Image source: NASA (public domain), via Wikimedia Commons.