Performing chess game position recognition solely from a single image of a three-dimensional board requires predicting the position and orientation of the board relative to the camera, the occupancy of squares and the piece type, which includes its colour. We propose an R-CNN-based framework with independent components for piece recognition and board geometry estimation, whose predictions are combined to reconstruct the position. For piece recognition, we adapt Faster R-CNN using a class-weighted objective and a deeper classification head. The detector operates directly on the input image, retaining alternative piece hypotheses that are subsequently refined using constraints on piece counts and square occupancy. For board detection, we introduce an octagonal arrangement of eight labelled boundary keypoints, predicted using the keypoint head of Mask R-CNN. These provide redundant correspondences for homography estimation and encode board orientation. The estimated homography maps representative points from the piece boxes to an 8x8 grid. On a synthetic dataset, the modifications to piece detection increase mean average precision from 61.59% to 90.14%. Of the predicted board keypoints, 97.11% are within 1% of the image diagonal of their labelled targets. Using ground-truth piece boxes with the predicted homographies gives correct square assignments for every test position. The complete framework recovers 76.61% of test positions exactly and 96.49% with at most one incorrect square.
Figures & tables
Figure 1 : Unfiltered outputs of the final piece detector. Boxes are labelled with predicted classes and confidence scores. Uppercase letters denote white pieces and lowercase letters denote black pieces. Several hypotheses can be retained for the same physical piece before the position is assembled.
Piece
White
Black
King
8516
8516
Queen
5494
5472
Rook
12458
12474
Bishop
9336
9306
Knight
8390
8484
Pawn
47182
47232
Table 1 : Piece counts in the augmented training data. Each image is also reflected horizontally, so these counts are twice the unaugmented counts.
Figure 2 : Projection from an image of a chessboard to a regular grid. Labelled correspondences determine both the board geometry and its orientation in the grid. The transformation applies to the board plane.
Figure 3 : The two keypoint layouts investigated. Blue points mark labelled locations. In the cross layout, the centre is collinear with each pair of opposite corners. In the octagonal layout, no three selected points are collinear. The connecting lines illustrate the layouts and are not additional predicted features.
Figure 4 : Failure of the cross layout when the predicted board box excludes an outer corner. The keypoint head cannot place a prediction outside the box. The blue crosses mark true locations, and the connected predictions show the resulting anomalous keypoint.
Figure 5 : Selection of a representative point for square assignment. The red point is the bottom midpoint of the box and the yellow point is its top midpoint. After projection, the green point is displaced by 0.2 square units from the red point towards the yellow point.
Figure 6 : Per-class average precision for the three piece detectors. The mAP values are 61.59%, 78.64% and 90.14%, respectively.
Configuration
Matched
Unmatched
Correctly
Conditional
pieces
predictions
classified
accuracy
Baseline
7335
399
6144
83.76%
Weighted loss
7352
1285
6715
91.34%
Weighted loss + count limits
7212
540
7002
97.09%
Table 2: Piece detection diagnostics on 7,387 ground-truth pieces. Correct classification is measured only among matched pieces. The comparison shows the effects of class weighting and piece-count filtering on the matched pieces and unmatched predictions.
Figure 7 : Classification confusion matrices for matched pieces. Columns are true classes and rows are predicted classes, with each column normalized by its number of matched instances. Uppercase letters denote white pieces and lowercase letters denote black pieces. These matrices exclude missed pieces and unmatched detections.
Figure 8 : Outputs of the final board detector. The eight labelled predictions are connected to show their order around the board. The predicted enclosing box is also shown. Some board features are occluded by pieces.
Figure 9 : Examples of complete position reconstruction. The board diagrams show the system’s predictions, with file and rank labels indicating the display orientation.
Metric
Training
Validation
Test
Mean incorrect squares per board
0.30
0.36
0.28
Boards with no mistakes (%)
76.44
70.42
76.61
Boards with at most one mistake (%)
95.12
94.37
96.49
Per-square error rate (%)
0.47
0.56
0.43
Correct boards with true piece boxes (%)
99.74
100.00
100.00
Table 3: Complete-system results. The final row evaluates square assignment using true piece boxes and classes with predicted geometry.
Metric
Chesscog
Present system
Mean incorrect squares per board
0.15
0.28
Boards with no mistakes (%)
93.86
76.61
Boards with at most one mistake (%)
99.71
96.49
Per-square error rate (%)
0.23
0.43
Table 4: Complete position recognition on the synthetic chess dataset. Chesscog results are from Wölflein and Arandjelović [ 11 ] . Board orientation is provided to Chesscog and predicted by our framework.