Do Papers Tell the Whole Story? A Benchmark and Framework for Uncovering Hidden Implementation Gaps in Bioinformatics
Organizations: Xi’an Jiaotong University
Abstract
As bioinformatics software is increasingly applied across a broader range of scenarios and the rapid development of large language models (LLMs) further lowers the barriers to software use and development, the composition of the bioinformatics research community is undergoing substantial change. Consequently, a growing number of researchers require a deeper understanding of methodological details and software behavior. In this context, systematically analyzing the relationship between paper descriptions and code implementations is emerging as an important new challenge in the field. To address this challenge, we introduce paper-code consistency analysis as a new research perspective and construct BioCon, the first benchmark dataset for paper-code consistency analysis in bioinformatics. Furthermore, we develop a unified cross-modal analysis framework to systematically investigate this problem from three perspectives: sentence-level detection, cross-modal retrieval, and project-level assessment. Experimental results demonstrate that the proposed framework can effectively model the semantic relationships between scientific publications and software implementations. Further case studies reveal that paper-code inconsistency is not a single phenomenon but arises from multiple underlying causes, among which Author-Perceived Non-Essential Details represents the most prevalent category. These findings suggest that paper-code consistency analysis is not merely a technical problem but also raises broader discussions regarding knowledge dissemination, the boundaries of code disclosure, and community norms. We hope that this work will encourage the bioinformatics community to re-examine the relationship between scientific publications and software implementations while providing a foundation for future research in paper-code consistency analysis.
Figures & tables
| Metrics | Value | Metrics | Value |
|---|---|---|---|
| Lines of Code | 19134 | Cyclomatic Complexity | 2732 |
| Number of Files | 76 | Number of Stars | 355 |
| Number of Functions | 681 | Citation Count | 603 |
| \botrule |
| Split | # Projects | # Consistent | # Inconsistent | # Total |
|---|---|---|---|---|
| Train | 38 | 957 | 4785 | 5742 |
| Validation | 5 | 90 | 450 | 540 |
| Test | 5 | 83 | 415 | 498 |
| Total | 48 | 1130 | 5650 | 6780 |
| \botrule |
| Encoder | Acc ( ) | F1 ( ) | MCC ( ) |
|---|---|---|---|
| CodeBERT | 0.8668 ±0.0041 | 0.7350 ±0.0117 | 0.4852 ±0.0102 |
| CodeT5+ | 0.8541 ±0.0018 | 0.7380 ±0.0054 | 0.4763 ±0.0109 |
| CodeGen | 0.8534 ±0.0065 | 0.7260 ±0.0108 | 0.4531 ±0.0220 |
| Qwen3 | 0.8728 ±0.0048 | 0.7303 ±0.0017 | 0.4829 ±0.0108 |
| UniXcoder | 0.8949 ±0.0064 | 0.7736 ±0.0139 | 0.5754 ±0.0272 |
| \botrule |
| Encoder | MRR ( ) | Hit@1 ( ) | Hit@5 ( ) | Hit@10 ( ) |
|---|---|---|---|---|
| CodeBERT | 0.5205 ±0.0072 | 0.4217 ±0.0120 | 0.6546 ±0.0106 | 0.6908 ±0.0080 |
| CodeT5+ | 0.5059 ±0.0044 | 0.4096 ±0.0070 | 0.6346 ±0.0040 | 0.6988 ±0.0070 |
| CodeGen | 0.5362 ±0.0057 | 0.4458 ±0.0139 | 0.6305 ±0.0106 | 0.6707 ±0.0080 |
| Qwen3 | 0.5346 ±0.0098 | 0.4618 ±0.0145 | 0.6104 ±0.0080 | 0.6627 ±0.0120 |
| UniXcoder | 0.5452 ±0.0040 | 0.4297 ±0.0106 | 0.7108 ±0.0070 | 0.7389 ±0.0040 |
| \botrule |
| Encoder | MRR ( ) | Hit@1 ( ) | Hit@5 ( ) | Hit@10 ( ) |
|---|---|---|---|---|
| CodeBERT | 0.3558 ±0.0069 | 0.2851 ±0.0080 | 0.4257 ±0.0145 | 0.5341 ±0.0040 |
| CodeT5+ | 0.3180 ±0.0053 | 0.2570 ±0.0040 | 0.4016 ±0.0080 | 0.4699 ±0.0070 |
| CodeGen | 0.2968 ±0.0053 | 0.2450 ±0.0106 | 0.3695 ±0.0106 | 0.4659 ±0.0040 |
| Qwen3 | 0.3353 ±0.0194 | 0.2892 ±0.0184 | 0.3655 ±0.0080 | 0.4377 ±0.0145 |
| UniXcoder | 0.3509 ±0.0041 | 0.2771 ±0.0070 | 0.4217 ±0.0070 | 0.5261 ±0.0040 |
| \botrule |
| Encoder | Acc ( ) | F1 ( ) | MAE ( ) | RMSE ( ) |
|---|---|---|---|---|
| CodeBERT | 0.7833 ±0.0167 | 0.6191 ±0.0112 | 0.2825 ±0.0219 | 0.3771 ±0.0279 |
| CodeT5+ | 0.7667 ±0.0167 | 0.6055 ±0.0215 | 0.3023 ±0.0077 | 0.3942 ±0.0170 |
| CodeGen | 0.7833 ±0.0167 | 0.5943 ±0.0136 | 0.3130 ±0.0146 | 0.4027 ±0.0228 |
| Qwen3 | 0.8000 ±0.0289 | 0.6632 ±0.0397 | 0.2869 ±0.0139 | 0.3815 ±0.0172 |
| UniXcoder | 0.8167 ±0.0167 | 0.6520 ±0.0441 | 0.2792 ±0.0054 | 0.3759 ±0.0073 |
| \botrule |
| Project | Acc ( ) | F1 ( ) | MCC ( ) |
| Project 1 | 0.9286 | 0.9048 | 0.8257 |
| Project 2 | 0.7778 | 0.7750 | 0.5500 |
| Project 3 | 0.5000 | 0.5000 | 0.1636 |
| Project 4 | 0.7692 | 0.7068 | 0.5394 |
| Project 5 | 0.7500 | 0.7334 | 0.5774 |
| Project 6 | 0.9333 | 0.9283 | 0.8660 |
| Cause | Paper Sentence | Probability | Human Validation |
|---|---|---|---|
| Operational Issues During Development and Release | During the generation of pseudo labels, a Spies Capture Rate (SCR) assesses the model’s performance of each bin in identifying the known positives of ”spies” from unlabeled dataset. | 0.1345 | Inconsistent. No implementation related to SCR was identified in the code. |
| (26.67%) | The weighted positive loss term penalizes the errors of the predictions in truly positive instances, while the BCE term trains the model using entire samples as a binary classifier. | 0.1733 | Inconsistent. The implementation of weighted positive loss was not identified in the code. |
| The whole slide images were segmented into smaller patches of 256 × 256 pixels using a sliding window approach. | 0.1749 | Inconsistent. The image size used in the code is 112 × 112. | |
| Author-Perceived Non-Essential Details (55%) | Model parameters were determined by Adam stochastic optimization (Kingma 2014) for 1000 iterations with a batch size of 200 using the binary cross-entropy loss function. | 0.1401 | Inconsistent. No implementation of Adam optimizer and BCE loss was identified in the code. |
| The second step, homology reduction, a critical process in preparing input data for a neural network model (Li et al.2023) , aimed to further eliminate redundancy, thereby minimizing bias, reducing overfitting risk, and enhancing the ability of the model to generalize to unseen data. | 0.2033 | Inconsistent. No implementation of homology reduction was identified in the code. | |
| Thus, we interrogated the relative contributions of different features toward the model’s output using SHAP (SHapley Additive exPlanations) analysis (Sundararajan and Najmi 2020). | 0.4503 | Inconsistent. No usage or implementation related to SHAP was identified in the code. |
| Type | Paper Sentence | Probability | Human Validation |
|---|---|---|---|
| False Positives Caused by Differences in Code Implementation Forms | From these, we derived accuracy and calculated the area under the ROC curve (AUROC) to assess the model’s ability to distinguish between binding and nonbinding proteins. | 0.1440 | Consistent. The corresponding implementation exists but is not encapsulated as a function. |
| Confusion matrices were used to calculate true positives, true negatives, false positives, and false negatives. | 0.3195 | Consistent. The corresponding implementation exists but is not encapsulated as a function. | |
| False Negatives Caused by Implicit Methodological Descriptions | To help address possible overfitting, L2 regularization was used with a value of 0.0001. | 0.7555 | Inconsistent. No L2 regularization setting was identified in the code. |
| A constant learning rate was used with a step-size of 0.1. | 0.5317 | Inconsistent. No learning rate configuration was identified in the code. | |
| Insufficient Filtering of Non-Implementation Sentences | Promising advances have been recently achieved thanks to computational and Artificial Intelligence (AI) methods, possibly opening new avenues for a faster and cheaper option in identifying G4BPs. | 0.1137 | The statement is descriptive rather than implementation-related. |
| However, a limitation of this metric arises when the model incorrectly predicts a large number of unlabeled samples as positive. | 0.1618 | The statement is descriptive rather than implementation-related. |